Molecule modeling
By incorporating an interatomic position representation that characterizes the spatial relationships between atoms into molecular modeling, the method addresses the limitations of existing techniques, resulting in improved accuracy in predicting molecular properties.
Patent Information
- Application Number
- PCT/US2024/053485
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-30
AI Technical Summary
Existing molecular modeling techniques often rely on limited representations, such as pairwise distances between atoms, which fail to capture the complexity of molecular structures, resulting in inadequate accuracy in predicting molecular properties.
The proposed solution involves determining an interatomic position representation that characterizes the relative spatial positions between individual pairs of atoms in a molecule, combined with an atomic attribute representation, to generate a feature representation that improves the accuracy of predicting molecular properties.
By considering the relative spatial positions between atoms, the proposed method introduces more abundant information into molecular modeling, leading to enhanced accuracy in predicting molecular properties.
Smart Images

Figure US2024053485_30052025_PF_FP_ABST
Abstract
Description
MOLECULE MODELING BACKGROUND
[0001] With the development of machine learning technology, it has been widely used in various technical fields. Molecular modeling is an important task in fields such as materials science, energy applications, biotechnology and drug research. Machine learning has been widely applied in such fields. Using machine learning to model and characterize molecules can predict their properties, providing support for further research. SUMMARY
[0002] According to implementations of the present disclosure, a solution for molecular modeling is proposed. In the solution, an interatomic position representation of a molecule is determined based on respective positions of a plurality of atoms in the molecule. The interatomic position representation characterizes relative spatial positions between individual pairs of atoms in the plurality of atoms. A feature representation of the molecule is determined based on an atomic attribute representation of the molecule and the interatomic position representation. The atomic attribute representation characterizes respective attributes of the plurality of atoms. A prediction of a target property for the molecule is determined based on the feature representation. According to embodiments of the present disclosure of the present disclosure, relative spatial positions between atoms are considered in molecular modeling to introduce more abundant information in modeling. In this way, the accuracy of predicting molecular properties can be improved.
[0003] This Summary is provided to introduce the selection of objects in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the technical solution as defined, nor is it intended to be used to limit the scope thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] FIG. 1 illustrates a block diagram of an example environment in which a plurality of implementations of the present disclosure can be implemented;
[0005] FIG. 2A illustrates a schematic diagram of an example architecture of a self- attention layer according to some implementations of the present disclosure;
[0006] FIG.2B illustrates a schematic diagram of an example molecular structure according to some implementations of the present disclosure;
[0007] FIG. 3A illustrates a schematic diagram of an example architecture of a machine learning model for molecular modeling according to some implementations of the present disclosure;
[0008] FIG. 3B illustrates an example architecture of an embedding layer in a machine learning model according to some implementations of the present disclosure;
[0009] FIG. 3C illustrates an example architecture of a self-attention block in a machine learning model according to some implementations of the present disclosure;
[0010] FIG.3D illustrates an example architecture of a decoder in a machine learning model according to some implementations of the present disclosure;
[0011] FIG. 4 illustrates a flowchart of the process of molecular modeling according to some implementations of the present disclosure; and
[0012] FIG. 5 illustrates a schematic block diagram of an electronic device capable of implementing a plurality of implementations of the present disclosure. DETAILED DESCRIPTION
[0013] The present disclosure is now explained with reference to several example implementations. It should be understood that these implementations are explained only for enabling those of ordinary skill in the art to better understand and therefore implement the present disclosure, rather than implying any limitations on the scope of the present disclosure.
[0014] As used herein, the term “comprise” and its variants are to be read as open terms that mean “include, but is not limited to.” The term “based on” is to be read as “based at least in part on.” The term “one implementation” is to be read as “at least one implementation.” The term “another implementation” is to be read as “at least one other implementation.” The terms “first,” “second” and so on may refer to different or the same objects. Other definitions, explicit and implicit, might be further included below.
[0015] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various implementations are described throughout this specification, and any type of implementation can be included under any section / subsection. In addition, the implementation described in any section / subsection can be combined in any way with any other implementation described in the same section / subsection and / or different sections / subsections.
[0016] In this specification, unless explicitly stated, performing a step “in response to A” does not mean performing the step immediately after A, but may include one or more intermediate steps.
[0017] As used herein, a group of elements, an element group or similar expressions may include zero, one or more such elements. The group of elements can be ordered or unordered. For example, a group of dividing lines can include zero, one or more dividing lines. As used herein, a sequence of elements or similar expressions may include one or more such elements, and the elements in the sequence are ordered.
[0018] As used herein, the term “model” can learn the corresponding association between inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for a given input. A model can be generated based on machine learning technology. Deep learning (DL) is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning based-model. As used herein, the "model" can also be referred to as “machine learning model”, “learning model”, “machine learning network” or “learning network”, and these terms are used interchangeably herein.
[0019] Usually, machine learning can roughly include three stages, namely training stage, testing stage, and usage stage (also referred to as inference stage). In the training stage, a given model can be trained using a large number of training data, and continue to iterate until the model can obtain consistent inference that meets the expected goal from the training data. Through training, the model can be considered to be able to learn the association between input and output (also referred to as input to output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In the inference stage, the model can be used to process the actual input and determine the corresponding output based on the trained parameter values. Example Environment and Basic Principles
[0020] FIG. 1 illustrates a schematic diagram of an example environment 100 in which implementations of the present disclosure can be implemented. As shown in FIG. 1, the environment 100 includes an electronic device 110. It is expected to use such an electronic device 110 to achieve molecular modeling and property prediction. To this end, in some implementations, a machine learning model 150 may be deployed in the electronic device 110 for molecular modeling and property prediction.
[0021] As shown in FIG. 1, the electronic device 110 may take information related to the molecular structure of a molecule 102 as input and output a predicted property 120. In the implementations of the present disclosure, any suitable properties may be predicted, such asbut not limited to the highest occupied molecular orbital (HOMO) energy level, lowest unoccupied molecular orbital (LUMO) energy level, zero-point vibration energy (ZPVE), and so on.
[0022] The molecule 102 may typically consist of a plurality of atoms, whose arrangement in space forms the molecular structure. The molecular structure may also affect molecular properties. Therefore, in order to predict the property for the molecule 102, it is necessary to model the molecular structure.
[0023] In FIG.1, the electronic device 110 may be any system with computing power, such as various computing devices / systems, terminal devices, servers, etc. The terminal device may be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a desktop computer, a laptop, a notebook, a netbook, a tablet, a media computer, a multimedia tablet, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. The server includes but is not limited to a mainframe, an edge computing node, a computing device in a cloud environment, etc.
[0024] It should be understood that the components and arrangements in the environment shown in FIG.1 are only examples, and the computing system suitable for implementing the implementations described herein may include one or more different components, other components, and / or different arrangements.
[0025] It should be understood that the structure and function of each element in the environment 100 are described for illustrative purposes only and does not imply any limitations on the scope of the present disclosure.
[0026] As mentioned above, it is necessary to model the molecular structure. To this end, some solutions use a pairwise distance between atoms as a relative position encoding. However, molecular structures are often complex, and the relative distance between atoms may only provide rather limited information. Therefore, such a representation is often inadequate for capturing complex interactions within molecules.
[0027] Based on the implementation of the present disclosure, a solution for molecular modeling based on Interatomic Position Encoding is proposed. In this solution, an interatomic position representation of a molecule is determined based on respective positions of a plurality of atoms in the molecule. The interatomic position representation characterizes relative spatial positions between individual pairs of atoms in the plurality of atoms. A feature representation of the molecule is determined based on an atomic attribute representation of the molecule and the interatomic position representation, the atomic attribute representation characterizing respective attributes of the plurality of atoms. Aprediction of a target property for the molecule is determined based on the feature representation. According to embodiments of the present disclosure of the present disclosure, relative spatial positions between atoms are considered in molecular modeling to introduce more abundant information in modeling. In this way, the accuracy of predicting molecular properties may be improved.
[0028] The description is presented below to example implementations of the present disclosure in conjunction with the accompanying drawings.
[0029] In order to conduct molecular modeling, it is necessary to describe the environment in which the atoms in the molecule are located. To better understand the implementation of the present disclosure, description is first presented to a way for representing the environment in which atoms are located. Atomic Cluster Expansion (abbreviated as ACE) may serve as a complete descriptor of the local atomic chemical environments, represented by hierarchical multi-body expansion. The key components of ACE include: ACE defines a complete set of basis functions for the environment of the centered atom (radial basis functions and spherical harmonics in practice); and ACE significantly reduces the computational efforts for body order computing to linear time complexity scaling with the number of atoms within the molecule. These advantages serve ACE as an accurate, fast, and transferable theory framework for molecular modeling. Example operations of ACE will be further described through formulas below.
[0030] ACE includes a set of orthogonal basis functions to describe the spatial relations between two atoms within a cluster (such as an atom i and atom j). The expansion of atomic potential energy of atom i may be represented by Formula (1):indicates a polynomial degree of the basis function, and indicates a expansionFormula (1) has an unrestricted summation.
[0031] The sum of surrounding atoms within the cluster will scale to the sum of K neighbors and will become numerically large. N indicates the number ofmolecule. ACE leverages the density trick to reduce the computational overhead. Thus, atomic base and A-basis may be defined, as shown in Formulas (2) and (3), respectively:andrepresents a geometry of . ACE may effectively represent an atomic expansion scaling with the complexityby firstly summing the neighboring basis functions and then applying is not rotationally invariant, additional Clebsch-Gordon coefficients are required to construct fully permutation and isometric- invariant basis(B-basis), and the potential energy of the cluster in Formula (1) is described as their linear combination with the coefficient , as shown in Formulas (4) and (5): (5)is therefore suitable for simulating molecules. Interatomic Position Encoding for Molecular Modeling
[0032] In the implementation of the present disclosure, there is proposed to model molecules using an interatomic position representation. The interatomic position representation characterizes the relative spatial positions between individual pairs of atoms in a plurality of atoms within a molecule and may be considered as an Interatomic Position Encoding (IPE). IPE may efficiently describe the multi-body contributions for geometric molecule modeling. The distinction from ACE is that IPE further takes the interactions between atomic clusters into account. IPE may be used for any suitable model. Description is presented below to ACE by taking attention blocks as an example.
[0033] FIG. 2A shows a schematic diagram of an example architecture of a self-attention layer 200A according to some implementations of the present disclosure. The self-attention layer 200A includes a plurality of linear layers 210, a plurality of linear layers 215, and a plurality of linear layers 220. The value input representation V may represent theeigenvectors of each molecular position. The value input representation Q may represent the eigenvector of the current position. The value input representation K may represent the eigenvectors at other positions. The electronic device 110 inputs the value input representation V into the plurality of linear layers 210, inputs the query input representation Q into the plurality of linear layers 215, and inputs the key input representation K into the plurality of linear layers 220. Through linear transformation, the electronic device 110 may calculate self-attention weights.
[0034] Furthermore, the self-attention layer 200A further includes a matrix multiplication layer 230, a scaling layer 240, and a matrix multiplication layer 225. The matrix multiplication layer 230 receives the output of the linear layer 215, the output of the linear layer 220, and an interatomic position representation 201. Thus, the electronic device 110 may apply the linearly transformed value input representation Q and value input representation K to each position representation in the interatomic position representation 201. The matrix multiplication layer 225 receives the output of the linear layer 210 and the output of the scaling layer 240. Thus, the electronic device 110 may update the interatomic position representation 201 using the output result of matrix multiplication layer 225.
[0035] Additionally, or alternatively, in some implementations, the self-attention layer 200A may further include a mask layer 235. The output of the matrix multiplication layer 230 passes through the mask layer 235 and is then input to the scaling layer 240. The mask layer 235 may be used to mask information or ignore padding, thereby better processing sequence data. The mask layer 235 is used to filter out the relative spatial information between atoms that are too far away. An example of the mask layer 235 will be further described with reference to FIG. 3C below.
[0036] In this way, an extended self-attention mechanism with the interatomic position representation may be achieved. This helps in the construction of self-attention weights, while such self-attention weights are updated through atomic attribute representations, effectively achieving molecular modeling.
[0037] In some implementations, for any pair of atoms (represented by a first and second atom) within a molecule, IPE may characterize the relative spatial position between any two atoms based on the interaction between a first atomic cluster including the first atom and a second atomic cluster including the second atom. The first atomic cluster is centered around the first atom, i.e., the first atomic cluster may indicate the local chemical environment in which the first atom is located. The second atomic cluster is centered around the second atom, i.e., the second atomic cluster may indicate the local chemical environment in whichthe second atom is located. Description is presented below with reference to FIG. 2B. The interactions between different atomic clusters may include multi-body extensions, for example, they may be represented by multi-body extensions. As an example, interactions between different atomic clusters may include one or more interaction terms, each corresponding to a body expansion order.
[0038] An example implementation of IPE will be described with reference to the molecular structure shown in FIG. 2B. This figure illustrates a schematic diagram of an example 200B of molecular structures according to some implementations of the present disclosure. The example 200B includes a cluster ^^^and a cluster ^^^. The cluster ^^^and the cluster ^^^include a group of atoms centered on the atom i and a group of atoms centered on the atom j (represented by circles with different filling patterns). The electronic device 110 may obtain a merged cluster ^^^^by merging the clusters ^^^and ^^^with a set of variable basis functions. The merged cluster ^^^^may describe the potential energy ^^^^^of interactions between atoms.
[0039] In one example, the cluster ^^^includes a group of atoms, as atoms ^^^, ^^ଶ, etc., with the centered atom of atom ^^. The cluster ^^^includes another group of suchas atoms ^^^, ^^ଶ, etc., with the centered atom of atom ^^. The clusters ^^^and ^^^are moved to overlap the atoms ^^ and ^^. Later, other atoms are merged into acluster σ୧୨, and the newly formed atomic base is as shown in Formula (6): (6)of a singleatomic bases and are as , which may still follow thedensity trick. Due to the may be derived. Anew A-basis may be constructed for the merged cluster ^^^^when taking the productof , as shown in Formula (7):(7) ents an expansion ingle cluster, represents an A base of a single cluster, represents an A base of the mergedrepresents a Hadamard product, anda Kronecker Product or tensor
[0040] As an example,. may describe both the- body and -contributes to-body expansion, while contributes to the- body expansion due to cluster merging.
[0041] For example, when taking , and represent 2-body (im) and 3-body expansion (im1m2) in the cluster σi,the 4-body expansion (mijn) in merged cluster σij, respectively. The 3-body expansion (im1m2) represents interactions between the atom i, the atom m1and the atom m2. The 4-body expansion (mijn) represents interactions between the atom m, the atom i, the atom j and the atom n, in merged cluster σij.
[0042] In some implementations, exists when ^ ≥ 2, i.e., considering at least 4-bodyexpansion within molecules. matrix of B-basis may be constructed, and the potential ^^^^^of cluster σij may be represented, asin Formula (8):space) and. In may be interpreted as the sum of cosine values of the,which represents the 3-bodycontributions in . Similarly, the base may be treated as the sum of proper dihedral angles two clusters ^^^and ^^^, i.e., , which representsthe 4-body contributions in . The base isfor the torsion potential between these two which considers the interaction within asingle cluster. Incorporating effectively enhances the representation of interatomic relations by capturing the between clusters, thereby contributing to the designof a Transformer as to be described below.
[0044] In the implementations of the present disclosure, based on a molecule with N atoms, an interatomic position encoding matrix (i.e., interatomic position representation) may be derived, whichpotentials. represents the order of the body expansion in Formula (2).
[0045] In some implementations, IPE may be used in self-attention mechanisms. For example, an initial attention map may be generated using atomic attribute representations for query input representations and key input representations respectively. Then, the attention map for the atomic attribute representation is derived by multiplying the atomic position representation with the initial attention map (e.g., direct product). Referring to FIG.2A, if the interatomic position representation 201 is , is directly multiplied with thevalue input representation Q and the value input representation K before passing through the scaling layer 240 (for example, through the softmax function), i.e., jointly input into the matrix multiplication layer 230. This may be used as a position encoding in Transformer, as shown in Formula (9):map; represents an atomic attribute representation in molecules (as to be described, represents a learnable weight matrix, F is a hidden layer dimension, and Qa value input representation and value input representation, respectively.
[0046] In some implementations, the basis function in Formula (3) may represent the local chemical environment of cluster ^^^withThus, it is possible to constructan A-basis matrix for a single cluster with an expansion order , and , where is an integer, andto Q and the value input representation K, and the attention(e.g. through the softmax function) may be modified as shown in Formula (10): of theKronecker Product.
[0047] In some implementations, the Clebsch-Gordon coefficient may be added to ensure the rotational invariance and . The linear expansion is as shown in Formula (11):order of a learnable weight matrix representing a -coefficient (weight). In somesuch operations may be achieved by a tensor broadcast. Therefore, the interatomic position representation may be shown in Formula (12):to Formula (9). As an example, it is appropriate to use IPE for the calculation of attention maps. Example architecture of machine learning models
[0049] An example architecture of a machine learning model will be described taking an attention mechanism-based Transformer as an example with reference to FIGs. 3A to 3D. The machine learning model may predict the properties of molecules based on their structure.
[0050] As shown in FIG.3A, an example architecture 300 generally includes an embedding layer 310, an encoder 302, and a decoder 330. The encoder 302 may further include a plurality of attention blocks, such as attention blocks 320-1, ..., attention blocks 320-L-1, and attention blocks 320-L, which are also collectively or individually referred to as attention blocks 320.
[0051] Overall, an initial atomic attribute representation and an initial interatomic position representation may be obtained through the embedding layer 310. As described above, the interatomic position representation characterizes the relative spatial position between individual pairs of atoms in a plurality of atoms within a molecule. The atomic attribute representation characterizes the corresponding attributes of these atoms. The attribute of an atom may be, for example, the type of atom. Furthermore, the atomic attribute representation and interatomic position representation may be updated through the encoder 302 to determine a feature representation of the molecule. The feature representation obtained as such may be regarded as being merged with the atomic attribute information and molecular spatial structure information. The decoder 330 may predict the properties of molecules based on such feature representations.
[0052] In the example of FIG. 3A, the example architecture 300 may take the atomic type and atomic coordinates from a molecule as inputs and output the corresponding molecularmolecule contains N atoms. The embedding layer 310, encoder 302, and decoder 330 will be described below in conjunction with FIGs. 3B to 3D, respectively.
[0053] FIG. 3B illustrates a schematic diagram of an example of the embedding layer 310 according to some implementations of the present disclosure. The embedding layer 310 receives molecular information from the molecule 102 as inputs. The molecular information includes the atomic type Z of the molecular structure 102 and its corresponding atomic coordinates.
[0054] The processing of the molecular information in the embedding layer 310 includes two branches. In one branch, the atomic type Z is processed by the embedding layer andnormalization layer to output an atomic attribute representation . In this example, theembedding layer 310 maps the atomic type Z to the atomic attribute representation , as shown in Formula (13):where embed() represents an embedding operation for the atomic type, and LayerNorm() represents layer normalization.
[0055] In another branch, the atomic coordinates are utilized. For example, the relative spatial positions between individual pairs of atoms in a plurality of atoms may be determined based on the molecular structural information. Then, the relative spatial positions may be transformed into atomic position representations by using a feature transformation network. As an example, such a feature transformation network may be based on any suitable basisfunction. In the example in FIG. 3B, the atomic position representation may beinitialized with the radial basis functions (abbreviated as RBF). As shown in FIG. 3B, the atomic coordinates (which may derive the relative spatial positions between individual pairs of atoms) are processed through RBF-based feature transformation and linear transformationto output the atomic position representation , as shown in Formulas (14) and (15):where represents a relative spatial position between two atoms andnorm. and areparameters that may beto specify a center and width of . is a smooth cosine cutoff function.is composed of the values of K radial basis functions. ismapping basis functions to the hidden layer dimension.
[0056] In some implementations, the initial atomic position representation and atomic attribute representation described above may be updated. Specifically, the atomic attribute representation may be updated based on the interatomic position representation. The interatomic position representation may be updated based on an intermediate result of updating the atomic attribute representation. The feature representation is determined based on the updated atomic attribute representation and the updated atomic position representation. Thus, the cross updating of the atomic position representation and the atomic attribute representation has been achieved, enabling different types of molecular information to be merged.
[0057] In some implementations, atomic attribute representations may be updated based on attention mechanisms. For example, an attention map for the atomic attribute representation may be determined based on the inter atomic position representation, and the atomic attribute representation may be updated based on the attention map.
[0058] Such updates may be achieved using any suitable attention mechanism-based network. In some implementations, at least one attention block may be utilized. The atomic attribute representation and the interatomic position representation are iteratively updated in at least one attention block. An example will be described below with reference to the encoder 320 shown in FIG. 3C.
[0059] In this example, the encoder 302 may be an encoder for extracting geometric features of molecules, including L attention blocks. Each attention block has a hidden layer dimension F and takes the atomic attribute representation X and the interatomic position representation output by the embedding layer 310 as inputs. Therefore, the interatomic position representation may be updated by the atomic attribute representation in each attention block. shows a schematic diagram of an example of the encoder 302 according tosome implementations of the present disclosure. In conjunction with FIG. 3A, the attention block 320-1 in the encoder 302 receives the atomic attribute representation and theinteratomic position representation from the embedding layer 310, and after processing,obtains updated intermediate atomic attribute representations and interatomic position representations in this attention block. The output of this attention block is fed into the nextattention block, and so on until the L-th attention block. The output obtained from theL-th attention block 320-L serves as a molecular feature representation.
[0061] Referring to FIG. 3B, for the l-th (l∈(1,L)) attention block, is received as inputs and will be sequentially processed through a self-attention layera normalization layer 322, a feedforward neural network 323 and a normalization layer 324 to output .
[0062] Compared with traditional Transformers, the encoder 302 introduces a learnable IPR matrix in each attention block, i.e., the interatomic position representation. Return to FIG. an example of the self-attention layer 321. The architecture shown in FIG. 2A may be used to implement the self-attention layer 321.
[0063] After passing through the matrix multiplication layer 230, an adjusted attention map may be obtained for the interatomic position representation, as shown in Formula (9). Any suitable activation function may be used to act on the attention map. For example, sigmoid linear unit (SiLU) activation may be used to enhance accuracy. In some implementations, the mask layer 235 may process the attention map with activation functions applied to filter out attention weights between at least one pair of atoms whose relative distance exceeds a predetermined distance, as shown in Formula (16): .Regarding the operation of mask layer 235, in some implementations, the relative distance between individual pairs of atoms may be determined based on their relative spatialpositions, such as . Then, the attention map may be updated based on the distancebetween individual pairs of atoms to filter out attention weights related to at least one pair of atoms whose relative distance exceeds the predetermined distance from the attention map. The updated attention map is as shown in Formula (16). In this implementation, the mask layer 235 may correspond to the implementation of an alternative attention mask, which may effectively filter out atoms with excessive distance, concentrate attention and improve the performance of the attention mechanism.
[0065] In this way, an intermediate atomic attribute representation for the l-th attention block may be determined based on the updated attention map and the value input representation for the l-th attention block. For example, after the operation of the matrix multiplication layer 225, the output of the self-attention layer (also referred to as self- attention representation) is obtained, as shown in Formula (17): (17)The l-th attention block further receives as inputs and adds it to itself after being processed by the self-attention layer. Therefore, may be updated after the self- attention layer.
[0066] An example of updating will be described below. In the update of , the self-attention representation output by the self-attention layer may be utilized. The relativespatial positions between individual pairs of atoms are transformed using basis functions. The interatomic position representation is updated based on the self-attention representation and transformed relative spatial positions. Additionally, in some implementations, the interatomic position representation is further updated based on the ACE-related body expansion order, as to described below.
[0067] As an example, under the instruction of Formula (7), A-bases for all merged clusters may be constructed, as shown in Formulas (18) and (19):represents aspherical harmonic function with an order and represent two learnable matrices without bias to ensure(18), the relative spatial positions between atoms are transformed with a spherical harmonic functionof a basis function with rotational invariance).
[0068] In one implementation, a new form of residual IPE may also be constructed within each attention block 230, as shown in Formula (20):
[0069] Furthermore, the residual connection may be applied to compute IPE for the next attention block, as shown in Formula (21): (21)of a gatedfilter. From the above description, by setting a value , the body expansion order considered for the relative spatial position betweenrepresented by IPE may be controlled.
[0070] The update of atomic features may follow traditional procedures, as shown in Formulas (22) and (23): (22)and represent two learnable matrices in the feedforward neural network 323. The output attention block is fed into the normalization layer 324, such as layer normalization.
[0071] It should be understood that although the above formulas describe complex processes, in practice, this model may use simplified settings to reduce computational complexity. For example, may be employed for body expansion, and for spherical harmonic it is possible to streamline the tensorcontraction and Clebsch-Gordon product calculations, making the model more efficient and easier to implement.
[0072] An example implementation of the encoder 302 has been described above. The molecular feature representation may be obtained using the encoder 302. In this example, the updated atomic attribute representation for the L-th attention block may be used feature representation is input to the decoder 330to predict the molecular property.
[0073] In some implementations, in the decoder, operations such as linear transformations and nonlinear activations may be performed on feature representations to derive predicted values of target properties. FIG. 3D illustrates a schematic diagram of an example of the decoder 330 according to some implementations of the present disclosure. In this example, the decoder 330 may be a lightweight decoder. This ensures that the Transformer remains computationally efficient while maintaining its ability to accurately predict molecular properties based on the geometric features captured by the encoder.
[0074] In one example, the decoder 330 includes two linear layers, an activation function layer (such as SiLU activation function) and an aggregation module (which is used to aggregate the outputs of various nodes in the linear layer) to predict specific properties. For example, the decoder 330 may receive atomic features output by the L-th attention block of the encoder 302, and sequentially processthrough the linear layer, theactivation function layer, the linear layer, and the aggregation module to output the molecular property 120.
[0075] It should be understood that the architecture of the model described above is only illustrative rather than limiting. Example Method
[0076] FIG. 4 shows a flowchart of a method 400 for molecular modeling according to some implementations of the present disclosure. The method 400 can be implemented at the electronic device 110 in FIG.1, for example, implemented using the machine learning model 150 in the electronic device 110.
[0077] As shown in FIG. 4, at block 410, the electronic device 110 determines an interatomic position representation of a molecule based on respective positions of a plurality of atoms in the molecule. The interatomic position representation characterizes relative spatial positions between individual pairs of atoms in the plurality of atoms.
[0078] At block 420, the electronic device 110 determines a feature representation of the molecule based on an atomic attribute representation of the molecule and the interatomic position representation. The atomic attribute representation characterizes respective attributes of the plurality of atoms.
[0079] At block 430, the electronic device 110 determines a prediction of a target property for the molecule based on the feature representation.
[0080] In some implementations, determining the feature representation of the molecule comprises: updating the atomic attribute representation based on the interatomic position representation; updating the interatomic position representation based on an intermediate result of updating the atomic attribute representation; and determining the feature representation based on the updated atomic attribute representation and the updated interatomic position representation.
[0081] In some implementations, updating the atomic attribute representation comprises: determining an attention map for the atomic attribute representation based on the interatomic position representation; and updating the atomic attribute representation based on the attention map.
[0082] In some implementations, determining the attention map for the atomic attribute representation comprises: generating an initial attention map by using the atomic attribute representation for a query input representation and a key input representation, respectively; and deriving the attention map for the atomic attribute representation by multiplying the interatomic position representation with the initial attention map.
[0083] In some implementations, the feature representation of the molecule is determined using a machine learning model, the machine learning model comprises at least one attention block, and both the atomic attribute representation and the interatomic position representation are iteratively updated in the at least one attention block.
[0084] In some implementations, updating the atomic attribute representation comprises: for a given attention block in the at least one attention block, obtaining a query input representation and a key input representation for the given attention block, the query input representation and the key input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block; obtaining a position input representation for the given attention block, the position input representation comprising the interatomic position representation or comprising an intermediate interatomic position representation updated in an attention block prior to the given attention block; determining an attention map for the given attention block based on the query input representation, the key input representation and the intermediate interatomic position representation; and determining an intermediate atomic attribute representation for the given attention block based on the attention map and a value input representation for the given attention block, the value input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block.
[0085] In some implementations, updating the atomic attribute representation based on the attention map comprises: determining relative distances between individual pairs of atoms in the plurality of atoms based on the relative spatial positions between individual pairs of atoms in the plurality of atoms; updating the attention map based on the relative distances between individual pairs of atoms in the plurality of atoms to filter out, from the attention map, an attention weight related to at least one pair of atoms for which the relative distance exceeds a predetermined distance; and determining an intermediate atomic attribute representation for the given attention block based on the updated attention map and a value input representation for the given attention block.
[0086] In some implementations, the intermediate result comprises: a self-attention representation generated in an update of the atomic attribute representation, and updating the interatomic position representation comprises: transforming the relative spatial positions between individual pairs of atoms in the plurality of pairs of atoms by using a basis function;and updating the interatomic position representation based on the self-attention representation and the transformed relative spatial positions.
[0087] In some implementations, determining the interatomic position representation of the molecule comprises: determining the relative spatial positions between individual pairs of atoms in the plurality of atoms based on structural information of the molecule; and transforming the relative spatial positions into the interatomic position representation by using a feature transformation network.
[0088] In some implementations, the feature transformation networks are at least based on a Radial Basis Function, RBF.
[0089] In some implementations, corresponding attributes of the plurality of atoms comprise corresponding types of the plurality of atoms.
[0090] In some implementations, determining the prediction of the target property for the molecule comprises: deriving a predicted value of the target property by performing a linear transformation and a nonlinear activation on the feature representation.
[0091] In some implementations, for a first atom and a second atom in the plurality of atoms, the interatomic position representation characterizes a relative spatial position between the first atom and the second atom based on an interaction between a first atomic cluster and a second atomic cluster, and the interaction comprises multi-body expansion, wherein the first atomic cluster comprises a plurality of atoms centered at the first atom, and the second atomic cluster comprises a plurality of atoms centered at the second atom.
[0092] Some example implementations of the present disclosure are listed below. Example Device
[0093] FIG. 5 illustrates a schematic block diagram of an electronic device capable of implementing a plurality of implementations of the present disclosure. It should be understood that the electronic device 500 shown in FIG. 5 is only illustrative and should not constitute any limitation on the functionality and scope of the implementation described herein.
[0094] As shown in FIG. 5, the electronic device 500 includes an electronic device 500 in the form of a general-purpose computing device. Components of the electronic device 500 may include, but are not limited to, one or more processors or processing devices 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560.
[0095] In some implementations, the electronic device 500 can be implemented as a computing device, computing system, server, mainframe, and other computing capable devices.
[0096] The processing device 510 can be a physical or virtual processor and can execute various processes based on the programs stored in the memory 520. In a multiprocessor system, a plurality of processing units executes computer executable instructions in parallel to improve the parallel processing capability of the electronic device 500. The processing device 510 may include a central processing unit (CPU), graphics processing unit (GPU), microprocessor, controller, and / or microcontroller, etc.
[0097] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any accessible media to the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 may include a volatile memory (such as a register, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 may include removable or non-removable media, and may include computer-readable media such as memory, flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within the electronic device 500.
[0098] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, a disk driver for reading from or writing to a removable, non-volatile disk and an optical disk driver for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each driver can be connected to a bus (not shown) via one or more data medium interfaces.
[0099] The communication unit 540 implements communication with another computing device via a communication medium. Additionally, functions of components of the electronic device 500 may be realized by a single computing cluster or a plurality of computing machines, and these computing machines may communicate through communication connections. Therefore, the electronic device 500 may operate in a networked environment using a logic connection to one or more other servers, a Personal Computer (PC) or a further general network node.
[0100] The input device 550 may be one or more various input devices, such as a mouse, a keyboard, a trackball, a voice-input device, and the like. The output device 560 may be one or more output devices, e.g., a display, a loudspeaker, a printer, and so on. The electronicdevice 500 may also communicate through the communication unit 540 with one or more external devices (not shown) as required, where the external device, e.g., a storage device, a display device, and so on, communicates with one or more devices that enable users to interact with the electronic device 500, or with any device (such as a network card, a modem, and the like) that enable the electronic device 500 to communicate with one or more other computing devices. Such communication may be executed via an Input / Output (I / O) interface (not shown).
[0101] In some implementations, apart from being integrated on an individual device, some or all of the respective components of the electronic device 500 may also be set in the form of a cloud computing architecture. In the cloud computing architecture, these components may be remotely arranged and may cooperate to implement the functions described by the subject matter described herein. In some implementations, the cloud computing provides computation, software, data access and storage services without informing a terminal user of physical locations or configurations of systems or hardware providing such services. In various implementations, the cloud computing provides services via a Wide Area Network (such as Internet) using a suitable protocol. For example, the cloud computing provider provides, via the Wide Area Network, the applications, which can be accessed through a web browser or any other computing component. Software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote location. The computing resources in the cloud computing environment may be merged or spread at a remote datacenter. The cloud computing infrastructure may provide, via a shared datacenter, the services even though they are shown as a single access point for the user. Therefore, components and functions described herein can be provided using the cloud computing architecture from a service provider at a remote location. Alternatively, components and functions may also be provided from a conventional server, or they may be mounted on a client device directly or in other ways.
[0102] The electronic device 500 may be used for implementing molecular modeling in a plurality of implementations of the present disclosure. The memory 520 may include one or more modules with one or more program instructions, which can be accessed and run by the processing unit 510 to implement the various functions described herein. For example, the memory 520 may include a molecular modeling module 522 for performing molecular modeling in one or more of the above implementations. As shown in FIG. 5, the electronic device 500 can obtain inputs required for molecular modeling through the input device 550 and can provide outputs of molecular modeling through the output device 560, such as thepredicted properties. In some implementations, the electronic device 500 may further receive inputs from other devices (not shown) via the communication unit 540. Example Implementations
[0103] Some example implementations of the present disclosure are listed below.
[0104] In an aspect, the present disclosure provides a computer implemented method. The method comprises: determining an interatomic position representation of a molecule based on respective positions of a plurality of atoms in the molecule, the interatomic position representation characterizing relative spatial positions between individual pairs of atoms in the plurality of atoms; determining a feature representation of the molecule based on an atomic attribute representation of the molecule and the interatomic position representation, the atomic attribute representation characterizing respective attributes of the plurality of atoms; and determining a prediction of a target property for the molecule based on the feature representation.
[0105] In some implementations, determining the feature representation of the molecule comprises: updating the atomic attribute representation based on the interatomic position representation; updating the interatomic position representation based on an intermediate result of updating the atomic attribute representation; and determining the feature representation based on the updated atomic attribute representation and the updated interatomic position representation.
[0106] In some implementations, updating the atomic attribute representation comprises: determining an attention map for the atomic attribute representation based on the interatomic position representation; and updating the atomic attribute representation based on the attention map.
[0107] In some implementations, determining the attention map for the atomic attribute representation comprises: generating an initial attention map by using the atomic attribute representation for a query input representation and a key input representation, respectively; and deriving the attention map for the atomic attribute representation by multiplying the interatomic position representation with the initial attention map.
[0108] In some implementations, the feature representation of the molecule is determined using a machine learning model, the machine learning model comprises at least one attention block, and both the atomic attribute representation and the interatomic position representation are iteratively updated in the at least one attention block.
[0109] In some implementations, updating the atomic attribute representation comprises: for a given attention block in the at least one attention block, obtaining a query inputrepresentation and a key input representation for the given attention block, the query input representation and the key input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block; obtaining a position input representation for the given attention block, the position input representation comprising the interatomic position representation or comprising an intermediate interatomic position representation updated in an attention block prior to the given attention block; determining an attention map for the given attention block based on the query input representation, the key input representation and the intermediate interatomic position representation; and determining an intermediate atomic attribute representation for the given attention block based on the attention map and a value input representation for the given attention block, the value input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block.
[0110] In some implementations, updating the atomic attribute representation based on the attention map comprises: determining relative distances between individual pairs of atoms in the plurality of atoms based on the relative spatial positions between individual pairs of atoms in the plurality of atoms; updating the attention map based on the relative distances between individual pairs of atoms in the plurality of atoms to filter out, from the attention map, an attention weight related to at least one pair of atoms for which the relative distance exceeds a predetermined distance; and determining an intermediate atomic attribute representation for the given attention block based on the updated attention map and a value input representation for the given attention block.
[0111] In some implementations, the intermediate result comprises: a self-attention representation generated in an update of the atomic attribute representation, and updating the interatomic position representation comprises: transforming the relative spatial positions between individual pairs of atoms in the plurality of pairs of atoms by using a basis function; and updating the interatomic position representation based on the self-attention representation and the transformed relative spatial positions.
[0112] In some implementations, determining the interatomic position representation of the molecule comprises: determining the relative spatial positions between individual pairs of atoms in the plurality of atoms based on structural information of the molecule; and transforming the relative spatial positions into the interatomic position representation by using a feature transformation network.
[0113] In some implementations, the feature transformation networks are at least based on a Radial Basis Function, RBF.
[0114] In some implementations, corresponding attributes of the plurality of atoms comprise corresponding types of the plurality of atoms.
[0115] In some implementations, determining the prediction of the target property for the molecule comprises: deriving a predicted value of the target property by performing a linear transformation and a nonlinear activation on the feature representation.
[0116] In some implementations, for a first atom and a second atom in the plurality of atoms, the interatomic position representation characterizes a relative spatial position between the first atom and the second atom based on an interaction between a first atomic cluster and a second atomic cluster, and the interaction comprises multi-body expansion, wherein the first atomic cluster comprises a plurality of atoms centered at the first atom, and the second atomic cluster comprises a plurality of atoms centered at the second atom.
[0117] In another aspect, the present disclosure provides an electronic device. The electronic device comprises: a processing unit; and a memory coupled to the processing unit and comprising instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising: determining an interatomic position representation of a molecule based on respective positions of a plurality of atoms in the molecule, the interatomic position representation characterizing relative spatial positions between individual pairs of atoms in the plurality of atoms; determining a feature representation of the molecule based on an atomic attribute representation of the molecule and the interatomic position representation, the atomic attribute representation characterizing respective attributes of the plurality of atoms; and determining a prediction of a target property for the molecule based on the feature representation.
[0118] In some example implementations, determining the feature representation of the molecule comprises: updating the atomic attribute representation based on the interatomic position representation; updating the interatomic position representation based on an intermediate result of updating the atomic attribute representation; and determining the feature representation based on the updated atomic attribute representation and the updated interatomic position representation.
[0119] In some example implementations, updating the atomic attribute representation comprises: determining an attention map for the atomic attribute representation based on the interatomic position representation; and updating the atomic attribute representation based on the attention map.
[0120] In some example implementations, determining the attention map for the atomic attribute representation comprises: generating an initial attention map by using the atomic attribute representation for a query input representation and a key input representation, respectively; and deriving the attention map for the atomic attribute representation by multiplying the interatomic position representation with the initial attention map.
[0121] In some example implementations, the feature representation of the molecule is determined using a machine learning model, the machine learning model comprises at least one attention block, and both the atomic attribute representation and the interatomic position representation are iteratively updated in the at least one attention block.
[0122] In some example implementations, updating the atomic attribute representation comprises: for a given attention block in the at least one attention block, obtaining a query input representation and a key input representation for the given attention block, the query input representation and the key input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block; obtaining a position input representation for the given attention block, the position input representation comprising the interatomic position representation or comprising an intermediate interatomic position representation updated in an attention block prior to the given attention block; determining an attention map for the given attention block based on the query input representation, the key input representation and the intermediate interatomic position representation; and determining an intermediate atomic attribute representation for the given attention block based on the attention map and a value input representation for the given attention block, the value input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block.
[0123] In some example implementations, updating the atomic attribute representation based on the attention map comprises: determining relative distances between individual pairs of atoms in the plurality of atoms based on the relative spatial positions between individual pairs of atoms in the plurality of atoms; updating the attention map based on the relative distances between individual pairs of atoms in the plurality of atoms to filter out, from the attention map, an attention weight related to at least one pair of atoms for which the relative distance exceeds a predetermined distance; and determining an intermediate atomic attribute representation for the given attention block based on the updated attention map and a value input representation for the given attention block.
[0124] In some example implementations, the intermediate result comprises: a self- attention representation generated in an update of the atomic attribute representation, and updating the interatomic position representation comprises: transforming the relative spatial positions between individual pairs of atoms in the plurality of pairs of atoms by using a basis function; and updating the interatomic position representation based on the self-attention representation and the transformed relative spatial positions.
[0125] In some example implementations, determining the interatomic position representation of the molecule comprises: determining the relative spatial positions between individual pairs of atoms in the plurality of atoms based on structural information of the molecule; and transforming the relative spatial positions into the interatomic position representation by using a feature transformation network.
[0126] In some example implementations, the feature transformation networks are at least based on a Radial Basis Function, RBF.
[0127] In some example implementations, corresponding attributes of the plurality of atoms comprise corresponding types of the plurality of atoms.
[0128] In some example implementations, determining the prediction of the target property for the molecule comprises: deriving a predicted value of the target property by performing a linear transformation and a nonlinear activation on the feature representation.
[0129] In some example implementations, for a first atom and a second atom in the plurality of atoms, the interatomic position representation characterizes a relative spatial position between the first atom and the second atom based on an interaction between a first atomic cluster and a second atomic cluster, and the interaction comprises multi-body expansion, wherein the first atomic cluster comprises a plurality of atoms centered at the first atom, and the second atomic cluster comprises a plurality of atoms centered at the second atom.
[0130] In a further aspect, the present disclosure provides a computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions which, when executed by a device, cause the device to perform acts comprising: determining an interatomic position representation of a molecule based on respective positions of a plurality of atoms in the molecule, the interatomic position representation characterizing relative spatial positions between individual pairs of atoms in the plurality of atoms; determining a feature representation of the molecule based on an atomic attribute representation of the molecule and the interatomic position representation, the atomic attribute representation characterizing respective attributes of the plurality of atoms; anddetermining a prediction of a target property for the molecule based on the feature representation.
[0131] In some example implementations, determining the feature representation of the molecule comprises: updating the atomic attribute representation based on the interatomic position representation; updating the interatomic position representation based on an intermediate result of updating the atomic attribute representation; and determining the feature representation based on the updated atomic attribute representation and the updated interatomic position representation.
[0132] In some example implementations, updating the atomic attribute representation comprises: determining an attention map for the atomic attribute representation based on the interatomic position representation; and updating the atomic attribute representation based on the attention map.
[0133] In some example implementations, determining the attention map for the atomic attribute representation comprises: generating an initial attention map by using the atomic attribute representation for a query input representation and a key input representation, respectively; and deriving the attention map for the atomic attribute representation by multiplying the interatomic position representation with the initial attention map.
[0134] In some example implementations, the feature representation of the molecule is determined using a machine learning model, the machine learning model comprises at least one attention block, and both the atomic attribute representation and the interatomic position representation are iteratively updated in the at least one attention block.
[0135] In some example implementations, updating the atomic attribute representation comprises: for a given attention block in the at least one attention block, obtaining a query input representation and a key input representation for the given attention block, the query input representation and the key input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block; obtaining a position input representation for the given attention block, the position input representation comprising the interatomic position representation or comprising an intermediate interatomic position representation updated in an attention block prior to the given attention block; determining an attention map for the given attention block based on the query input representation, the key input representation and the intermediate interatomic position representation; and determining an intermediate atomic attribute representation for the given attention block based on the attention map and a value input representation for the given attention block, the value inputrepresentation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block.
[0136] In some example implementations, updating the atomic attribute representation based on the attention map comprises: determining relative distances between individual pairs of atoms in the plurality of atoms based on the relative spatial positions between individual pairs of atoms in the plurality of atoms; updating the attention map based on the relative distances between individual pairs of atoms in the plurality of atoms to filter out, from the attention map, an attention weight related to at least one pair of atoms for which the relative distance exceeds a predetermined distance; and determining an intermediate atomic attribute representation for the given attention block based on the updated attention map and a value input representation for the given attention block.
[0137] In some example implementations, the intermediate result comprises: a self- attention representation generated in an update of the atomic attribute representation, and updating the interatomic position representation comprises: transforming the relative spatial positions between individual pairs of atoms in the plurality of pairs of atoms by using a basis function; and updating the interatomic position representation based on the self-attention representation and the transformed relative spatial positions.
[0138] In some example implementations, determining the interatomic position representation of the molecule comprises: determining the relative spatial positions between individual pairs of atoms in the plurality of atoms based on structural information of the molecule; and transforming the relative spatial positions into the interatomic position representation by using a feature transformation network.
[0139] In some example implementations, the feature transformation networks are at least based on a Radial Basis Function, RBF.
[0140] In some example implementations, corresponding attributes of the plurality of atoms comprise corresponding types of the plurality of atoms.
[0141] In some example implementations, determining the prediction of the target property for the molecule comprises: deriving a predicted value of the target property by performing a linear transformation and a nonlinear activation on the feature representation.
[0142] In some example implementations, for a first atom and a second atom in the plurality of atoms, the interatomic position representation characterizes a relative spatial position between the first atom and the second atom based on an interaction between a first atomic cluster and a second atomic cluster, and the interaction comprises multi-body expansion,wherein the first atomic cluster comprises a plurality of atoms centered at the first atom, and the second atomic cluster comprises a plurality of atoms centered at the second atom.
[0143] In a further aspect, the present disclosure provides a computer readable medium storing computer executable instructions that, when executed by a device, cause the device to perform one or more example implementations of the method of the aspect described above.
[0144] The functionality described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-Programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), and the like.
[0145] Program code for carrying out methods of the subject matter described herein may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or a server.
[0146] In the context of the present disclosure, a machine-readable medium may be any tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine- readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0147] Further, although operations are depicted in a particular order, it should be understood that the operations are required to be executed in the particular order shown orin a sequential order, or all operations shown are required to be executed to achieve the expected results. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, while several specific implementation details are contained in the above discussions, these should not be construed as limitations on the scope of the subject matter described herein. Certain features that are described in the context of separate implementations may also be implemented in combination in a single implementation. Conversely, various features that are described in the context of a single implementation may also be implemented in multiple implementations separately or in any suitable sub- combination.
[0148] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter specified in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
CLAIMS 1. A computer implemented method comprising: determining an interatomic position representation of a molecule based on respective positions of a plurality of atoms in the molecule, the interatomic position representation characterizing relative spatial positions between individual pairs of atoms in the plurality of atoms; determining a feature representation of the molecule based on an atomic attribute representation of the molecule and the interatomic position representation, the atomic attribute representation characterizing respective attributes of the plurality of atoms; and determining a prediction of a target property for the molecule based on the feature representation.
2. The method of claim 1, wherein determining the feature representation of the molecule comprises: updating the atomic attribute representation based on the interatomic position representation; updating the interatomic position representation based on an intermediate result of updating the atomic attribute representation; and determining the feature representation based on the updated atomic attribute representation and the updated interatomic position representation.
3. The method of claim 2, wherein updating the atomic attribute representation comprises: determining an attention map for the atomic attribute representation based on the interatomic position representation; and updating the atomic attribute representation based on the attention map.
4. The method of claim 3, wherein determining the attention map for the atomic attribute representation comprises: generating an initial attention map by using the atomic attribute representation for a query input representation and a key input representation, respectively; and deriving the attention map for the atomic attribute representation by multiplying the interatomic position representation with the initial attention map.
5. The method of claim 2, wherein the feature representation of the molecule is determined using a machine learning model, the machine learning model comprising at least one attention block, and wherein both the atomic attribute representation and the interatomic positionrepresentation are iteratively updated in the at least one attention block.
6. The method of claim 5, wherein updating the atomic attribute representation comprises: for a given attention block in the at least one attention block, obtaining a query input representation and a key input representation for the given attention block, the query input representation and the key input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block; obtaining a position input representation for the given attention block, the position input representation comprising the interatomic position representation or comprising an intermediate interatomic position representation updated in an attention block prior to the given attention block; determining an attention map for the given attention block based on the query input representation, the key input representation and the intermediate interatomic position representation; and determining an intermediate atomic attribute representation for the given attention block based on the attention map and a value input representation for the given attention block, the value input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block.
7. The method of claim 3, wherein updating the atomic attribute representation based on the attention map comprises: determining relative distances between individual pairs of atoms in the plurality of atoms based on the relative spatial positions between individual pairs of atoms in the plurality of atoms; updating the attention map based on the relative distances between individual pairs of atoms in the plurality of atoms to filter out, from the attention map, an attention weight related to at least one pair of atoms for which the relative distance exceeds a predetermined distance; and determining an intermediate atomic attribute representation for the given attention block based on the updated attention map and a value input representation for the given attention block.
8. The method of claim 2, wherein the intermediate result comprises: a self-attention representation generated in an update of the atomic attribute representation, and updatingthe interatomic position representation comprises: transforming the relative spatial positions between individual pairs of atoms in the plurality of atoms by using a basis function; and updating the interatomic position representation based on the self-attention representation and the transformed relative spatial positions.
9. The method of claim 1, wherein determining the interatomic position representation of the molecule comprises: determining the relative spatial positions between individual pairs of atoms in the plurality of atoms based on structural information of the molecule; and transforming the relative spatial positions into the interatomic position representation by using a feature transformation network.
10. The method of claim 9, wherein the feature transformation networks are at least based on a Radial Basis Function, RBF.
11. The method of claim 1, wherein determining the prediction of the target property for the molecule comprises: deriving a predicted value of the target property by performing a linear transformation and a nonlinear activation on the feature representation.
12. The method of claim 1, wherein for a first atom and a second atom in the plurality of atoms, the interatomic position representation characterizes a relative spatial position between the first atom and the second atom based on an interaction between a first atomic cluster and a second atomic cluster, and the interaction comprises multi-body expansion, wherein the first atomic cluster comprises a plurality of atoms centered at the first atom, and the second atomic cluster comprises a plurality of atoms centered at the second atom.
13. An electronic device, comprising: a processing unit; and a memory coupled to the processing unit and comprising instructions stored thereon, the instructions, when executed by the processing unit, causing the device to perform acts comprising: determining an interatomic position representation of a molecule based on respective positions of a plurality of atoms in the molecule, the interatomic position representation characterizing relative spatial positions between individual pairs of atoms in the plurality of atoms; determining a feature representation of the molecule based on an atomic attribute representation of the molecule and the interatomic position representation, theatomic attribute representation characterizing respective attributes of the plurality of atoms; and determining a prediction of a target property for the molecule based on the feature representation.
14. The device of claim 13, wherein determining the feature representation of the molecule comprises: updating the atomic attribute representation based on the interatomic position representation; updating the interatomic position representation based on an intermediate result of updating the atomic attribute representation; and determining the feature representation based on the updated atomic attribute representation and the updated interatomic position representation.
15. The device of claim 14, wherein updating the atomic attribute representation comprises: determining an attention map for the atomic attribute representation based on the interatomic position representation; and updating the atomic attribute representation based on the attention map.
16. The device of claim 15, wherein determining the attention map for the atomic attribute representation comprises: generating an initial attention map by using the atomic attribute representation for a query input representation and a key input representation, respectively; and deriving the attention map for the atomic attribute representation by multiplying the interatomic position representation with the initial attention map.
17. The device of claim 14, wherein the feature representation of the molecule is determined using a machine learning model, the machine learning model comprising at least one attention block, and wherein both the atomic attribute representation and the interatomic position representation are iteratively updated in the at least one attention block.
18. The device of claim 17, wherein updating the atomic attribute representation comprises: for a given attention block in the at least one attention block, obtaining a query input representation and a key input representation for the given attention block, the query input representation and the key input representation comprising the atomic attribute representation or comprising an intermediate atomic attributerepresentation updated in an attention block prior to the given attention block; obtaining a position input representation for the given attention block, the position input representation comprising the interatomic position representation or comprising an intermediate interatomic position representation updated in an attention block prior to the given attention block; determining an attention map for the given attention block based on the query input representation, the key input representation and the intermediate interatomic position representation; and determining an intermediate atomic attribute representation for the given attention block based on the attention map and a value input representation for the given attention block, the value input representation comprising the atomic attribute representation or comprising an intermediate atomic attribute representation updated in an attention block prior to the given attention block.
19. The device of claim 15, wherein updating the atomic attribute representation based on the attention map comprises: determining relative distances between individual pairs of atoms in the plurality of atoms based on the relative spatial positions between individual pairs of atoms in the plurality of atoms; updating the attention map based on the relative distances between individual pairs of atoms in the plurality of atoms to filter out, from the attention map, an attention weight related to at least one pair of atoms for which the relative distance exceeds a predetermined distance from the attention map; and determining an intermediate atomic attribute representation for the given attention block based on the updated attention map and a value input representation for the given attention block.
20. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions which, when executed by a device, cause the device to perform acts comprising: determining an interatomic position representation of a molecule based on respective positions of a plurality of atoms in the molecule, the interatomic position representation characterizing relative spatial positions between individual pairs of atoms in the plurality of atoms; determining a feature representation of the molecule based on an atomic attribute representation of the molecule and the interatomic position representation, the atomicattribute representation characterizing respective attributes of the plurality of atoms; and determining a prediction of a target property for the molecule based on the feature representation.