Isovariant neural network model for molecular conformation optimization task

Through isovariant neural network model, the molecular conformation is optimized, and the Transformer architecture driven by Wasserstein gradient flow and multi-task loss function are used to solve the problems of low efficiency and poor interpretability of molecular ground state conformation prediction in the prior art, and efficient and accurate molecular ground state conformation acquisition is achieved.

CN120373357APending Publication Date: 2025-07-25RENMIN UNIVERSITY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510362423.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is inefficient and lacks interpretability in molecular ground state conformation prediction, the generative and predictive methods are inaccurate, and the optimization method model architecture design is unreasonable.

Method used

The isovariant neural network model is adopted, and the isovariant Transformer architecture driven by Wasserstein gradient flow is used to optimize the molecular conformation through encoder and decoder, and the energy function based on the atomic latent hybrid model is optimized, and the model is trained in combination with the multi-task loss function.

Benefits of technology

Improve the accuracy and efficiency of molecular ground state conformation prediction, and enhance the interpretability of the model and reduce time cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373357A_ABST
    Figure CN120373357A_ABST
Patent Text Reader

Abstract

The invention provides an isovariant neural network model for a molecular conformation optimization task. Comprising the following four steps of: 1, giving a molecular map, and acquiring a corresponding low-quality conformation by using an open-source bioinformatics tool RDKit; 2, encoding the low-quality conformation to a hidden space by using an encoder; a third step of using an equivariant Transform architecture driven by a Wasserstein gradient flow to minimize an energy function based on an atomic potential hybrid model so as to optimize the low-quality conformation; and 4, decoding a final ground state conformation, namely the most stable molecular conformation, from the hidden space by using a decoder. Compared with the prior art, the method provided by the invention has the advantages that the interpretability of the model is remarkably improved, and the accuracy and efficiency of obtaining the molecular ground state conformation can be greatly improved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more particularly, to an equivariant neural network model for molecular conformation optimization tasks. Background Art

[0002] The molecular ground state conformation, as the most stable three-dimensional molecular structure, represents the energy minimization state on the potential energy surface. Therefore, it determines many important molecular properties and plays a key role in many downstream applications such as molecular docking and molecular property prediction. Traditionally, the molecular ground state conformation can be accurately obtained through various energy-based simulation methods (such as molecular dynamics simulation and density functional theory calculation), but these methods are inefficient and costly, and can no longer meet the growing demand. In recent years, with the wide application of artificial intelligence technology in the scientific field, many deep learning-based methods have been successively proposed for the molecular ground state conformation prediction task, and these methods can be roughly divided into three categories: generative methods, predictive methods, and optimization methods.

[0003] Generative methods mainly use various generative models to generate a large number of potential low-energy conformations, and then screen out the conformation with the lowest energy from these conformations. However, due to the uncertainty of the two links of generative model sampling and conformation screening itself, such methods actually cannot guarantee to obtain the molecular ground state conformation.

[0004] Predictive methods use the molecular 2D graph as input to directly predict the molecular ground state conformation, thus avoiding the uncertainty of generative methods. However, due to relying solely on the molecular 2D graph as input to predict the ground state conformation from scratch, such methods not only have limited prediction accuracy but also low efficiency.

[0005] Optimization methods deal with the molecular ground state conformation prediction task from the perspective of molecular conformation optimization, and obtain the molecular ground state conformation by using easily accessible low-quality molecular conformations as input and correspondingly optimizing these conformations. Essentially, optimization methods actually use neural networks to predict the gradient of the molecular conformation potential energy surface and then imitate gradient descent to update the input conformation. However, the model architectures of existing optimization methods are often designed empirically, making their forward calculation process not actually correspond to minimizing a physically meaningful energy function. Therefore, not only does the model lack interpretability, but its performance is also poor. Summary of the Invention

[0006] The purpose of the embodiments of the present disclosure is to provide an equivariant neural network model for molecular conformation optimization tasks, whose molecular conformation optimization process corresponds to minimizing an energy function based on an atomic potential hybrid model, greatly improving the interpretability of the model while enhancing the accuracy and efficiency of obtaining the molecular ground state conformation.

[0007] It includes four steps:

[0008] In the first step, given a 2D graph corresponding to a molecule composed of N atoms, denoted as where represents the set of nodes composed of atoms, represents the set of edges composed of chemical bonds between atoms, and the corresponding low-quality conformation is obtained using the open-source bioinformatics tool RDKit

[0009] In the second step, the encoder module is used to encode the low-quality conformation into the latent space, and the initial atomic representation and the atomic relationship representation

[0010]

[0011] Here, the function f encodes the atomic type v i of the i-th atom into the corresponding initial atomic representation is a Gaussian kernel function composed of a mean μ ∈ R H and a standard deviation σ ∈ R H constitutes a Gaussian kernel function, represents the Euclidean distance between the i-th atom and the j-th atom in 3D space, and respectively represent the learnable weights and biases associated with each atom pair type (v i , v j ).

[0012] In the third step, in the latent space, the L-layer WGFormer architecture is used to obtain the final atomic representation X (0) and the atomic relationship representation R (0) from the initial atomic representation X {(L)} and the atomic relationship representation R (L) , thereby minimizing an energy function based on the atomic potential mixture model to optimize the input low-quality conformation:

[0013] X (L) , R (L) = WGFormer L (X (0) , R (0) );

[0014] In the fourth step, the decoder is used to decode the final ground state conformation, i.e., the most stable molecular conformation, from the latent space

[0015]

[0016] The MLP here represents a multi-layer perceptron, and represent the initial relationship representation and the final relationship representation obtained through L layers of WGFormer between the i-th atom and the j-th atom, respectively.

[0017] Specifically, the l-th layer of WGFormer takes the atomic representation and relationship representation of the output of the previous layer, namely (X (l-1) , R (l -1) ) as the input, and outputs (X (l) , R (l) ) as the input for the next layer;

[0018] It includes H attention heads (where H > 0, and in our implementation, H is set to 64). For the h-th attention head, its calculation is as follows:

[0019] Q (l,h) = K (l,h) = X (l-1 )W (l,h)

[0020]

[0021] V (l,h) = -X (l-1 )A (l,h)

[0022]

[0023] T (l,h) = κ M (R (l,h) )V (l,h)

[0024] Here, W (l) = [W (l,1) , …, W (l,H) ∈ R D×D is a learnable matrix used to replace W Q , W K , W V in the standard Transformer architecture, and represents the h-th component of W (l) , X (l-1) ∈ R N×D represents the atomic representation of the output of the previous layer of WGFormer, R (l-1,h) ∈ R N×N represents the relationship representation of the output of the previous layer of WGFormer, R (l-1) ∈ R N×N×H is the h-th component of R M(·) represents the Sinkhorn operation, that is, first perform the exponential operation on the input matrix, then alternately normalize its rows and columns, and repeat the operation for M steps;

[0025] Finally, based on the output of each attention head WGFormer concatenates them to obtain the final output (X (l) , R (l) ):

[0026]

[0027] Here, Concat(·) represents the concatenation operation, that is, concatenate H matrices T with dimension into a matrix with dimension N×D, and concatenate H matrices R with dimension N×N (l,h) into a matrix with dimension N×N×H. (l,h)

[0028] During the training process, a multi-task loss function is used to enable the model to simultaneously fit the molecular ground state conformation C and the corresponding interatomic distance matrix D, thereby enhancing the model performance;

[0029] Specifically, given the molecular ground state conformation C and the predicted molecular ground state conformation First, use RDKit to align C with in 3D space to obtain the aligned molecular ground state conformation C * , and then simultaneously constrain the predicted ground state conformation and the corresponding interatomic distance matrix to make them simultaneously fit the true values C * and D:

[0030]

[0031] Among them, represents calculating the error loss for the distance part using the L1 norm, represents calculating the error loss for the coordinate part using the Frobenius norm, and λ is used to control the weight.

[0032] The technical effect to be achieved by the present invention is as follows:

[0033] Compared with the existing methods, the present invention can achieve better molecular ground state conformation prediction results at a lower time cost, and the conformation optimization process can correspond to minimizing an energy function based on an atomic potential mixture model, greatly enhancing the model interpretability. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The above and other objects and features of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings.

[0035] Figure 1 is a schematic diagram showing the structure of an equivariant neural network model for a molecular conformation optimization task according to an embodiment of the present disclosure;

[0036] Figure 2 is a schematic diagram showing the structure of an attention head in the equivariant neural network model according to an embodiment of the present disclosure;

[0037] Figure 3 is a comparison chart showing the prediction effects of different methods on two public datasets, Molecule3D and QM9, according to an embodiment of the present disclosure;

[0038] Figure 4 is a comparison chart showing the prediction efficiencies of different methods on two public datasets, Molecule3D and QM9, according to an embodiment of the present disclosure;

[0039] Figure 5 is a comparison chart showing the visualization effects of different methods according to an embodiment of the present disclosure. Detailed Embodiments

[0040] The following detailed embodiments are provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent after understanding the disclosure of the present application. For example, the order of operations described herein is merely illustrative and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.

[0041] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. On the contrary, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein, which will be apparent after understanding the disclosure of the present application.

[0042] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more thereof.

[0043] Although terms such as "first", "second", and "third" may be used herein to describe various components, elements, regions, layers, or parts, these components, elements, regions, layers, or parts should not be limited by these terms. Instead, these terms are only used to distinguish one component, element, region, layer, or part from another. Thus, without departing from the teachings of the examples, the first component, first element, first region, first layer, or first part referred to in the examples described herein may also be referred to as the second component, second element, second region, second layer, or second part.

[0044] In the specification, when an element (such as a layer, region, or substrate) is described as "on" another element, "connected to" or "coupled to" another element, the element can be directly "on" another element, directly "connected to" or "coupled to" another element, or there can be one or more other elements therebetween. In contrast, when an element is described as "directly on" another element, "directly connected to" or "directly coupled to" another element, there can be no other elements therebetween.

[0045] The terms used herein are only for describing various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" specify the presence of the stated features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0046] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs after understanding this disclosure. Unless explicitly defined herein, terms (such as those defined in a general dictionary) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and should not be interpreted in an idealized or overly formalized manner.

[0047] In addition, in the description of the examples, when a detailed description of a related structure or function that is considered well-known would cause an unclear interpretation of the disclosure, such a detailed description will be omitted.

[0048] Figure 1 is a schematic diagram showing an equivariant neural network model for a molecular conformation optimization task according to an embodiment of the present disclosure.

[0049] It includes the following steps:

[0050] The first step, given a 2D graph corresponding to a molecule composed of N atoms, denoted as where Denote the set of nodes composed of atoms, Denote the set of edges composed of chemical bonds between atoms. First, we use the open-source bioinformatics tool RDKit to obtain its corresponding low-quality conformation

[0051] In the second step, use the encoder module to encode the low-quality conformation into the latent space and obtain the initial atomic representation and the atomic relationship representation

[0052]

[0053] The function f here encodes the atomic type v of the i-th atom i into the corresponding initial atomic representation is a Gaussian kernel function composed of the mean μ ∈ R H and the standard deviation σ ∈ R H constituting it, denotes the Euclidean distance between the i-th atom and the j-th atom in 3D space, and respectively denote the learnable weights and biases associated with each atomic pair type (v i , v j ).

[0054] In the third step, we use the L-layer WGFormer (an equivariant Transformer architecture driven by the Wasserstein gradient flow) in the latent space to obtain the final atomic representation X (0) and the atomic relationship representation R (0) from the initial atomic representation X {(L)} and the atomic relationship representation R (L) , thereby minimizing an energy function based on the atomic potential mixture model to optimize the input low-quality conformation:

[0055] X (L) , R (L) = WGFormer L (X (0) , R (0) )

[0056] Specifically, for the l-th layer of WGFormer, it takes the output of the previous layer, i.e., (X (L-1) , R (L-1) ) as the input and outputs (X (l) , R (l) ) as the input of the next layer. It includes H attention heads. For the h-th attention head, its calculation is as Figure 2 shown:

[0057] Q (l,h) = K (l,h) = X (l-1 )W (l,h)

[0058]

[0059] V (l,h) = -X (l-1) A (l,h)

[0060]

[0061] T (l,h) = κ M (R (l,h) )V (l,h)

[0062] Here, W (l) = [W (l,1) , …, W (l,H) ∈ ℝ D×D is a learnable matrix used to replace W in the standard Transformer architecture Q , W K , W V , and represents the h-th component of W (l) , X (l-1) ∈ ℝ N×D represents the atomic representation of the output of the previous layer of WGFormer, and R (l-1,h) ∈ ℝ N×N represents the relational representation R of the output of the previous layer of WGFormer (l-1) ∈ ℝ N×N×H is the h-th component of R M κ(·) represents the Sinkhorn operation, that is, first perform the exponential operation on the input matrix, then alternately normalize its rows and columns respectively, and repeat the execution for M steps;

[0063] Finally, based on the output of each attention head WGFormer concatenates them to obtain the final output (X (l) , R (l) ):

[0064]

[0065] Here, Concat(·) represents the concatenation operation, that is, concatenate H matrices T with dimension (l,h) into a matrix with dimension N × D, and concatenate H matrices R (l,h) with dimension N × N into a matrix with dimension N × N × H.

[0066] It should be noted that, compared with the standard Transformer architecture, the WGFormer used in the third step adjusts the relationship of "QKV" in the attention mechanism and replaces the Softmax operation with the Sinkhorn operation, making it a kind of equivariant Transformer architecture driven by the Wasserstein gradient flow. Therefore, the molecular conformation optimization process in this step also corresponds to minimizing an energy function based on the atomic potential mixture model, greatly increasing the model interpretability.

[0067] In the fourth step, we use the decoder to decode the final ground state conformation, that is, the most stable molecular conformation, from the latent space.

[0068]

[0069] During the training process, we use a multi-task loss function to make the model simultaneously fit the molecular ground state conformation C and the corresponding inter-atomic distance matrix D, thereby enhancing the model performance. Specifically, based on the true molecular ground state conformation C and the molecular ground state conformation predicted by us We first use RDKit to align C with in 3D space to obtain the aligned molecular ground state conformation C * , and then simultaneously constrain the predicted ground state conformation and the corresponding inter-atomic distance matrix to make them simultaneously fit to the true values C * and D:

[0070]

[0071] Among them, represents calculating the error loss for the distance part using the L1 norm, represents calculating the error loss for the coordinate part using the Frobenius norm, and λ is used to control the weight.

[0072] The WGFormer used to optimize the molecular conformation in the latent space is an equivariant Transformer architecture driven by the Wasserstein gradient flow. Compared with the standard Transformer architecture, the WGFormer adjusts the relationship of "QKV" in the attention mechanism and replaces the Softmax operation with the Sinkhorn operation, making its molecular conformation optimization process correspond to minimizing an energy function based on the atomic potential mixture model, greatly increasing the model interpretability.

[0073] Experimental verification

[0074] Figure 3 It is a comparison of the prediction effects of different methods on two public datasets, Molecule3D and QM9. It can be seen that our method has achieved the best performance on these two public datasets;

[0075] Figure 4 It is a comparison of the prediction efficiencies of different methods on two public datasets, Molecule3D and QM9. It can be seen that compared with other methods, our method can improve the efficiency by at least 50%, greatly reducing the time cost of predicting the ground-state conformation of molecules.

[0076] Although some embodiments of the present disclosure have been shown and described, those skilled in the art should understand that these embodiments can be modified without departing from the principles and spirit of the present disclosure as defined by the claims and their equivalents.

Claims

1. An equivariant neural network model for molecular conformation optimization tasks, characterized in that, It includes four steps: It includes four steps: First step, given a 2D graph corresponding to a molecule composed of N atoms, denoted as where represents the set of nodes composed of atoms, represents the set of edges composed of chemical bonds between atoms, and use the open-source bioinformatics tool RDKit to obtain its corresponding low-quality conformation In the second step, use the encoder module to encode the low-quality conformation into the latent space to obtain the initial atomic representation and the atomic interaction relationship representation The function f here encodes the atomic type v of the i-th atom i into the corresponding initial atomic representation is a Gaussian kernel function composed of a mean μ ∈ R H and a standard deviation σ ∈ R H ; represents the Euclidean distance between the i-th atom and the j-th atom in 3D space, and represent the learnable weight and bias associated with each atomic pair type (v i , v i ), respectively; In the third step, use the L-layer WGFormer architecture in the latent space to obtain the final atomic representation X (0) from the initial atomic representation X (0) and the inter-atomic relationship representation R {(L)} to obtain the final atomic representation X (L) and the inter-atomic relationship representation R , thereby minimizing an energy function based on an atomic latent mixture model to optimize the input low-quality conformation: X (L) ,R (L) = WGFormer L (X (0) ,R (0) ); In the fourth step, use the decoder to decode the final ground state conformation, i.e., the most stable molecular conformation, from the latent space. Here, the MLP represents a multi-layer perceptron, and respectively represent the initial relational representation and the final relational representation obtained through L layers of WGFormer between the i-th atom and the j-th atom.

2. The equivariant neural network model for molecular conformation optimization task according to claim 1, characterized in that, The implementation method of the WGFormer architecture is that for the l-th layer, the atomic representation and relational representation of the output of the previous layer, namely (X (l-1) , R (l-1) ), are used as the input, and the output (X (l) , R (l) ) is used as the input for the next layer; It includes H attention heads (where H > 0, and in our implementation, H is set to 64). For the h-th attention head, its calculation is as follows: Q (l,h) = K (l,h) = X (l-1) W (l,h) V (l,h) = -X (l-1) A (l,h) T (l,h) = κ M (R (l,h) )V (l,h) The W here (l) = [W (l,1) , …, W (l,H) ∈ R D×D is a learnable matrix used to replace W Q , W K , W V in the standard Transformer architecture, and represents the h-th component of W (l) , X (l-1) ∈ R N×D denotes the atomic representation output by the previous layer of WGFormer, R (l-1,h) ∈ R N×N denotes the relational representation R output by the previous layer of WGFormer (l-1) ∈ R N×N×H 's h-th component, κ M (·) represents the Sinkhorn operation, that is, first perform the exponential operation on the input matrix, then alternately normalize its rows and columns, and repeat the execution for M steps; Finally, based on the output of each attention head WGFormer concatenates them to obtain the final output (X (l) , R (l) ): Here, Concat(·) represents the concatenation operation, that is, concatenating H matrices T of dimension into a matrix of dimension N×D, and concatenating H matrices R of dimension N×N (l,h) into a matrix of dimension N×N×H. (l,h) ​ 3. The equivariant neural network model for molecular conformation optimization task according to claim 2, wherein During the training process, a multi-task loss function is used to enable the model to simultaneously fit the molecular ground-state conformation C and the corresponding inter-atomic distance matrix D, thereby enhancing the model's performance; Given the ground-state conformation C of the molecule and the predicted ground-state conformation of the molecule First, use RDKit to align C with in 3D space to obtain the aligned ground-state conformation C of the molecule * . Then, simultaneously constrain the predicted ground-state conformation and the corresponding interatomic distance matrix to fit them simultaneously to the true values C * and D: Among them, indicates calculating the error loss for the distance part using the L1 norm, indicates calculating the error loss for the coordinate part using the Frobenius norm, and λ is used to control the weight.