Conformation prediction method

By acquiring ligand and target protein features through conformation prediction models and performing position embedding and cross-attention operations, the problem of ignoring molecular and binding pocket shape features in existing technologies is solved, and more efficient binding conformation prediction is achieved.

CN115862729BActive Publication Date: 2026-02-06LIANTAI CLUSTER (BEIJING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211542147.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-02
Publication Date
2026-02-06
Estimated Expiration
2042-12-02

AI Technical Summary

Technical Problem

In drug design, existing technologies rely on deep learning methods to independently learn molecular features and binding pocket features, neglecting important features such as the shape between the molecule and the binding pocket, resulting in insufficient performance in predicting binding conformations.

Method used

By using a conformation prediction model, features of ligands and target proteins are obtained, and position embedding, cross attention, and pair connection operations are performed to predict the binding conformation of ligands and target proteins.

Benefits of technology

It improves the predictive performance of binding conformations and can find the optimal binding conformation more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115862729B_ABST
    Figure CN115862729B_ABST
Patent Text Reader

Abstract

The application discloses a conformation prediction method. The method comprises the following steps: acquiring a conformation prediction model, and acquiring structure data of a complex to be predicted; wherein the structure data of the complex to be predicted comprises structure data of a ligand and structure data of a target protein; inputting the structure data of the ligand and the structure data of the target protein into the conformation prediction model to acquire ligand features and target protein features; performing position embedding, cross attention and connection operation on the ligand features and the target protein features to obtain a connection result matched with the ligand features and the target protein features; and predicting a target binding conformation of the structure data of the ligand and the structure data of the target protein according to the connection result. The technical scheme of the embodiment of the application provides a new conformation prediction method, learns important features such as shapes between ligands and target proteins, improves the prediction performance of the binding conformation, and helps to find the optimal binding conformation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of drug design, and particularly relates to a conformation prediction method. BACKGROUND

[0002] In the field of drug design, it is difficult to screen candidate molecules that bind to target proteins from databases and adjust them to the most favorable binding conformation. Only a small fraction of molecules can bind to target proteins and produce effective therapeutic effects. Traditional methods to solve this problem are divided into two categories: experimental methods and computational methods. Among them, experimental methods such as X-ray diffraction, nuclear magnetic resonance and other methods. Recently, low-temperature electron microscopy has also become an important method for observing and understanding molecular interactions. Compared with experimental methods, computational methods are more efficient and more convenient. At the same time, computational methods can also reveal some potential biological properties.

[0003] The prior art uses a deep learning method combined with a differential evolution algorithm to solve the docking problem. The existing method has advantages in speed performance, but the performance needs to be improved, and the disadvantage is that the important features between the molecules and the binding pocket (for example, shape complementarity) are ignored when learning the molecular features and the binding pocket features independently. SUMMARY

[0004] The present application provides a conformation prediction method to provide a new conformation prediction method, learn important features such as shape between ligands and target proteins, improve the prediction performance of the binding conformation, and help to find the optimal binding conformation.

[0005] According to an aspect of the present application, a conformation prediction method is provided, comprising:

[0006] Obtaining a conformation prediction model and obtaining to-be-predicted complex structure data; wherein the to-be-predicted complex structure data comprises ligand structure data and target protein structure data;

[0007] Inputting the ligand structure data and the target protein structure data into the conformation prediction model to obtain ligand features and target protein features;

[0008] Performing position embedding, cross-attention and connection operation on the ligand features and the target protein features to obtain a connection result matched with the ligand features and the target protein features;

[0009] According to the connection result, predicting a target binding conformation of the ligand structure data and the target protein structure data.

[0010] The technical scheme of the embodiment of the present application inputs ligand structure data and target protein structure data in the to-be-predicted complex structure data into a pre-trained conformation prediction model, obtains ligand features and target protein features, performs position embedding, cross attention, and connection operation on the ligand features and the target protein features, obtains a connection result matched with the ligand features and the target protein features, and predicts the target binding conformation of the ligand structure data and the target protein structure data according to the connection result. The technical means solves the problem that the prior art independently learns molecular features and binding pocket features and ignores important features such as shapes between molecules and binding pockets, provides a new conformation prediction method, learns important features such as shapes between ligands and target proteins, improves the prediction performance of the binding conformation, and helps to find the optimal binding conformation.

[0011] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0013] Figure 1 A flowchart of a conformation prediction method provided by the embodiment of the present application is shown in FIG. 1.

[0014] Figure 2 A specific application schematic flowchart of a conformation prediction method provided by the embodiment of the present application is shown in FIG. 2. DETAILED DESCRIPTION

[0015] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0016] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification and in the claims and the above-described drawings are intended to cover a non-exclusive inclusion, for example, a process, method, system, product or apparatus that includes a list of steps or units not necessarily limited to those explicitly listed, but can include other steps or units not expressly listed or inherent to such processes, methods, products or apparatus.

[0017] Embodiments

[0018] Figure 1 A flowchart of a conformation prediction method provided by an embodiment of the present application, which can be applicable to the case of predicting the binding conformation of a ligand and a target protein, can be executed by a conformation prediction device, which can be realized in the form of hardware and / or software, and can be configured in a server. As shown in the figure, the method comprises: Figure 1

[0019] S110, obtaining a conformation prediction model and obtaining structure data of a complex to be predicted; wherein the structure data of the complex to be predicted comprises ligand structure data and target protein structure data.

[0020] In this embodiment, the conformation prediction model and the structure data of the complex to be predicted can be obtained first.

[0021] In an optional implementation, obtaining the conformation prediction model can comprise: obtaining training set data and test set data; inputting the training set data into a deep learning model for feature learning to obtain a learned deep learning model; inputting the test set data into the learned deep learning model for verification testing to obtain the conformation prediction model.

[0022] For example, the training set data and the test set data can be obtained using the v2020 version of the pdbbind dataset, and the complexes that do not meet the pretreatment requirements and the complexes contained in CASF-2016 can be deleted, and 19080 complexes can be obtained. The training set can include 17000 complexes, and the test set can include 2080 complexes. The complex data of the training set is input into the deep learning model for feature learning to obtain a learned deep learning model; the complex data of the test set is used to verify the learned deep learning model, thereby obtaining the conformation prediction model.

[0023] S120, inputting the ligand structure data and the target protein structure data into the conformation prediction model to obtain ligand features and target protein features.

[0024] In this embodiment, the ligand structure data and the target protein structure data can be input into the conformation prediction model to obtain the ligand features and the target protein features.​

[0025] In one optional implementation, inputting ligand structure data and target protein structure data into a conformation prediction model to obtain ligand features and target protein features may include: processing the ligand structure data into a molecular map and processing the target protein structure data into a binding grid; inputting the molecular map and binding grid into the conformation prediction model to obtain ligand features and target protein features.

[0026] In this embodiment, the ligand structure data and target protein structure data can be preprocessed. The ligand structure data can be processed into a molecular map, and the target protein structure data can be processed into a binding grid. The molecular map and the binding grid are then input into the conformation prediction model to obtain the ligand features and target protein features.

[0027] Specifically, ligand structure data and target protein structure data can be preprocessed into G1 = (v1, ε1) and G2 = (v...) respectively. t , ε t The graph is represented by ) . For the ligand, each node v 1i ∈v1 is represented as a one-hot vector, indicating the atomic type of the node; similarly, each edge e 1i,j ∈ε1 represents a one-hot vector, indicating the connection of atoms v. 1i and atom v 1j The bond type (single, double, triple, and aromatic). For the target protein, the original input is a PDB file. Each PDB file of the target protein is processed into a polygonal mesh of binding points using MaSIF. Each point in this mesh is v ti ∈v t Represented as a vector of four properties (electrostatics, hydrophilicity, hydrogen bond donor / acceptor, and shape index), with each edge e ti,j ∈ε t Represented as connection point v ti and point v tj The vector.

[0028] S130. Perform position embedding, cross attention, and pair connection operations on the ligand features and target protein features to obtain pair connection results that match the ligand features and target protein features.

[0029] In this embodiment, a series of operations, including position embedding, cross attention, and pairing, can be performed on the ligand features and target protein features to obtain the pairing result.

[0030] In one alternative implementation, the ligand features and target protein features can first be embedded separately to obtain the embedded ligand features and the embedded target protein features.

[0031] Specifically, the position embedding can be performed on the ligand features according to degrees of atom nodes in a molecule graph to obtain embedded ligand features; and the position embedding can be performed on the target protein features according to degrees of points in a binding grid to obtain embedded target protein features. The molecule graph is a directed graph composed of atom nodes and atom connection edges; and the binding grid is a polygon grid directed graph composed of points and point connection edges. The position embedding is performed on the atom nodes in the molecule graph based on in-degrees and out-degrees of the atom nodes to describe positions of the atom nodes, and the position embedding based on the degrees of the atom nodes is added to the ligand features to obtain the embedded ligand features; and the position embedding is performed on the points in the binding grid based on in-degrees and out-degrees of the points to describe positions of the points, and the position embedding based on the degrees of the points is added to the target protein features to obtain the embedded target protein features.

[0032] Secondly, the cross-attention operation can be performed on the embedded ligand features and the embedded target protein features to obtain processed ligand features and processed target protein features.

[0033] Specifically, the cross-attention operation can be performed on the embedded ligand features according to Q vectors in the embedded ligand features, K vectors in the embedded target protein features and V vectors in the embedded target protein features to obtain the processed ligand features; and the cross-attention operation can be performed on the embedded target protein features according to Q vectors in the embedded target protein features, K vectors in the embedded ligand features and V vectors in the embedded ligand features to obtain the processed target protein features.

[0034] The formula of the cross-attention operation is The Q vectors, the K vectors and the V vectors of the embedded ligand features and the Q vectors, the K vectors and the V vectors of the embedded target protein features can be obtained. The cross-attention operation is performed on the embedded ligand features according to the Q vectors of the embedded ligand features, the K vectors of the embedded target protein features and the V vectors of the embedded target protein features to obtain the processed ligand features. That is, the point features of the target protein are queried according to the features of each atom node of the ligand, and as a result, the points of the target protein that need to be paid attention to when the binding occurs are obtained.

[0035] Similarly, the cross-attention operation is performed on the embedded target protein features according to the Q vectors of the embedded target protein features, the K vectors of the embedded ligand features and the V vectors of the embedded ligand features to obtain the processed target protein features.

[0036] The above-mentioned performing the position embedding and the cross-attention operation has the advantage that important features between the ligand molecule and the target protein binding pocket can be paid attention to, and the prediction performance of the binding conformation is improved.

[0037] Finally, the connection operation can be performed on the processed ligand features and the processed target protein features to obtain a connection result.

[0038] S140, predicting the target binding conformation of the ligand structure data and the target protein structure data according to the connection result.

[0039] The target binding conformation can be the optimal binding conformation between the ligand and the target protein.

[0040] Optionally, the target binding conformation of the ligand structure data and the target protein structure data can be predicted by a potential function according to the connection result.

[0041] The technical scheme of the embodiment of the present application inputs the ligand structure data and the target protein structure data in the to-be-predicted complex structure data into a pre-trained conformation prediction model to obtain ligand features and target protein features; performs position embedding, cross attention, and connection operations on the ligand features and the target protein features to obtain a connection result matched with the ligand features and the target protein features; and predicts the target binding conformation of the ligand structure data and the target protein structure data according to the connection result. The technical means solves the problem that the prior art independently learns molecular features and binding pocket features, ignores important features such as shapes between molecules and binding pockets, provides a new conformation prediction method, learns important features such as shapes between ligands and target proteins, improves the prediction performance of the binding conformation, and helps to find the optimal binding conformation.

[0042] Exemplarily, Figure 2 A specific application schematic flowchart of a conformation prediction method provided by the embodiment of the present application is provided. After preprocessing the ligand structure data and the target protein structure data, a molecule graph and a binding grid are obtained respectively; further, feature extraction is performed by a graph convolutional neural network and a residual graph convolutional neural network to obtain ligand features and target protein features; position embedding and cross attention are performed on the ligand features and the target protein features, and then connection operations are performed; after a series of operations such as multilayer perception, Gaussian mixture, and potential function on the connection result, a conformation prediction result is obtained.

[0043] It should be understood that the various forms of flowcharts shown above can be reordered, added, or deleted steps. For example, each step described in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical scheme of the present application can be achieved, which is not limited herein.

[0044] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A conformation prediction method characterized by, The method comprises the following steps: obtaining a conformation prediction model and obtaining to-be-predicted complex structure data; wherein the to-be-predicted complex structure data comprises ligand structure data and target protein structure data; processing the ligand structure data into a molecular graph and processing the target protein structure data into a binding grid; wherein the molecular graph is a directed graph composed of atomic nodes and atomic connection bonds; the binding grid is a polygon grid directed graph composed of points and point connection edges; inputting the molecular graph and the binding grid into the conformation prediction model to obtain ligand features and target protein features; performing position embedding on the ligand features according to the degrees of the atomic nodes in the molecular graph to obtain embedded ligand features; performing position embedding on the target protein features according to the degrees of the points in the binding grid to obtain embedded target protein features; performing cross-attention operation on the embedded ligand features according to a query Q vector in the embedded ligand features, a key K vector and a value V vector in the embedded target protein features to obtain processed ligand features; performing cross-attention operation on the embedded target protein features according to a Q vector in the embedded target protein features, a K vector and a V vector in the embedded ligand features to obtain processed target protein features; performing pair connection operation on the processed ligand features and the processed target protein features to obtain a pair connection result; predicting a target binding conformation of the ligand structure data and the target protein structure data according to the pair connection result.

2. The method of claim 1, wherein, The method for obtaining a conformation prediction model comprises the following steps: obtaining training set data and test set data; inputting the training set data into a deep learning model for feature learning to obtain a learned deep learning model; inputting the test set data into the learned deep learning model for verification testing to obtain a conformation prediction model.

3. The method of claim 1, wherein, predicting a target binding conformation of the ligand structure data and the target protein structure data according to the pair connection result, which comprises: predicting a target binding conformation of the ligand structure data and the target protein structure data according to the pair connection result through a potential function.