A method and device for training a molecular docking model

CN117153292BActive Publication Date: 2026-09-25HANGZHOU CARBON SILICON SMART TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310994545.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-08
Publication Date
2026-09-25
Estimated Expiration
2043-08-08

AI Technical Summary

Technical Problem

模型训练数据较少,从而导致训练的分子对接模型的预测效果较差,且较少的数据量已严重限制了深度学习对接方法的发展

Benefits of technology

[0086]本申请实施例中,通过基于物理模拟数据对待训练分子对接模型进行预训练,物理模拟数据为蛋白口袋和小分子对接得到的对接构象。基于微调样本数据对预训练后的分子对接模型进行微调训练,得到模型输出,微调样本数据是基于复合物晶体数据和基础样本数据混合而成的,基准样本数据包括:伪晶体数据和物理模拟数据中的至少一种。基于模型输出,计算得到损失值。在损失值处于预设范围内的情况下,得到分子对接模型。本申请实施例通过构造伪晶体数据,从而可以丰富分子对接模型的训练数据的数量,进而能够提高训练的分子对接模型的效果,解决较少的数据量已严重限制了深度学习对接方法的发展的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117153292B_ABST
    Figure CN117153292B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a training method and device of a molecular docking model. The method comprises: pre-training a to-be-trained molecular docking model based on physical simulation data, the physical simulation data being a docking conformation obtained by docking a protein pocket and a small molecule; fine-tuning training the pre-trained molecular docking model based on fine-tuning sample data to obtain a model output, the fine-tuning sample data being mixed based on complex crystal data and basic sample data, the basic sample data comprising at least one of pseudo crystal data and the physical simulation data; calculating a loss value based on the model output; and obtaining the molecular docking model in a case where the loss value is within a preset range. The embodiments of the present application can improve the effect of the trained molecular docking model and solve the problem that a small amount of data has seriously limited the development of a deep learning docking method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method and apparatus for a molecular docking model. Background Technology

[0002] Molecular docking is a computational chemistry method used to predict the binding patterns between two molecules. In drug design and molecular biology research, molecular docking technology is widely used to predict the binding patterns between drug molecules and protein molecules, as well as to study protein-protein and protein-nucleic acid interactions.

[0003] Molecular docking typically involves two molecules: a receptor and a ligand. The receptor is the target molecule, usually a protein, while the ligand is the candidate molecule, usually a drug molecule. The goal of molecular docking is to determine the optimal binding mode between the ligand and the receptor, i.e., at which positions and orientations the ligand binds to the receptor most stably.

[0004] Molecular docking is typically performed using computer simulations or deep learning methods. Molecular docking technology plays an important role in drug design and molecular biology research, and can accelerate the process of new drug discovery and protein structure research.

[0005] During the receptor-ligand binding process, the conformations of the receptor and ligand will change to varying degrees due to the influence of intermolecular forces. These changes are often very helpful. To simplify the problem and reduce the difficulty of prediction, molecular docking tasks are divided into three different docking modes: rigid docking, semi-flexible docking, and flexible docking.

[0006] Commonly used molecular docking models typically use experimentally obtained complex crystal data as training data. Existing work has collected and organized this data, with the most commonly used dataset being PDBBIND2020, containing 19,443 records. The limited training data results in poor prediction performance of trained molecular docking models, and this small amount of data has severely limited the development of deep learning docking methods. Summary of the Invention

[0007] The technical problem to be solved by the embodiments of this application is to provide a method and apparatus for training a molecular docking model, so as to improve the effect of training the molecular docking model and solve the problem that the limited amount of data has seriously restricted the development of deep learning docking methods.

[0008] In a first aspect, embodiments of this application provide a method for training a molecular docking model, the method comprising:

[0009] Pre-training is performed on the molecular docking model to be trained based on physical simulation data, wherein the physical simulation data is the docking conformation obtained by docking protein pockets and small molecules.

[0010] The pre-trained molecular docking model is fine-tuned based on fine-tuning sample data to obtain the model output. The fine-tuning sample data is a mixture of complex crystal data and basic sample data. The basic sample data includes at least one of pseudo-crystal data and physical simulation data.

[0011] Based on the model output, the loss value is calculated;

[0012] The molecular docking model is obtained when the loss value is within a preset range.

[0013] Optionally, before pre-training the molecular docking model to be trained based on the physical simulation data, the method further includes:

[0014] Obtain protein crystals from the PDB database;

[0015] The protein crystals are pocketed to obtain protein pockets;

[0016] For each of the protein pockets, small molecules are randomly selected from a preset database;

[0017] The protein pocket and the small molecule are docked to generate a docking conformation, and the docking conformation is used as the physical simulation data.

[0018] Optionally, when the pseudo-crystal data is included in the basic sample data,

[0019] Before fine-tuning the pre-trained molecular docking model based on fine-tuning sample data to obtain the model output, the following steps are also included:

[0020] Traverse the homologous protein data and extract the complex crystal data of the homologous protein data;

[0021] The target protein sequence in the homologous protein data is pocket-aligned to obtain aligned proteins;

[0022] The alignment protein is combined with the ligand to generate the pseudo-crystal data.

[0023] Optionally, the step of performing pocket alignment processing on the target protein sequence in the homologous protein data to obtain aligned proteins includes:

[0024] Compare the protein sequences of homologous proteins and remove crystal data whose sequences differ from other homologous protein sequences by a set value;

[0025] The remaining crystal data was pocket-aligned using a point cloud matching algorithm to obtain the aligned protein.

[0026] Optionally, the model output includes multiple docking configurations, each determined by a different model input.

[0027] The loss value calculated based on the model output includes:

[0028] Based on the conformational combination of the docking conformations, multiple first loss values ​​are calculated;

[0029] Based on the aforementioned multiple docking conformations and reference crystal conformation, multiple second loss values ​​are calculated;

[0030] The weighted sum of the plurality of first loss values ​​and the plurality of second loss values ​​is used to obtain the loss value.

[0031] Optionally, after obtaining the molecular docking model, the method further includes:

[0032] Obtain protein pockets and random initial conformations;

[0033] The protein pocket and the random initial conformation are input into the molecular docking model;

[0034] The molecular docking model is invoked to process the protein pocket and the random initial conformation to obtain the predicted docking conformation;

[0035] Based on the conformation rationality constraint algorithm, the predicted docking conformation is subjected to conformation rationality constraint processing to obtain the final target docking conformation.

[0036] Optionally, the conformational rationality constraint algorithm is used to perform conformational rationality constraint processing on the predicted docking conformation to obtain the final target docking conformation, including:

[0037] Based on preset tools, small molecule conformations that satisfy statistical laws are generated;

[0038] Initialize the initial values ​​for the conformational structure change action;

[0039] Based on the initial amount, the small molecule conformation is subjected to conformational structure change processing to obtain the updated small molecule conformation;

[0040] Based on the updated small molecule conformation and the predicted docking conformation, the conformational difference loss is calculated;

[0041] Based on the conformational difference loss, the update amount of the conformational structure change action is calculated;

[0042] The initial quantity is updated based on the updated quantity, and the conformational structure of the small molecule is processed by a conformational structure change action.

[0043] The update process is executed iteratively until the conformational difference loss is lower than the loss threshold, at which point the target docking conformation is output.

[0044] Optionally, the conformational structure change action includes at least one of rotation, translation, and torsion.

[0045] Secondly, embodiments of this application provide a training device for a molecular docking model, the device comprising:

[0046] The pre-training module is used to pre-train the molecular docking model to be trained based on physical simulation data, wherein the physical simulation data is the docking conformation obtained by docking protein pockets and small molecules.

[0047] The fine-tuning module is used to fine-tune the pre-trained molecular docking model based on fine-tuning sample data to obtain the model output. The fine-tuning sample data is a mixture of complex crystal data and baseline sample data. The baseline sample data includes at least one of pseudo-crystal data and physical simulation data.

[0048] The loss value calculation module is used to calculate the loss value based on the model output;

[0049] The molecular docking model acquisition module is used to obtain the molecular docking model when the loss value is within a preset range.

[0050] Optionally, the device further includes:

[0051] The protein crystal acquisition module is used to acquire protein crystals from the PDB database;

[0052] The protein pocket acquisition module is used to perform pocket segmentation processing on the protein crystal to obtain protein pockets;

[0053] The small molecule screening module is used to randomly select small molecules from a preset database for each protein pocket;

[0054] The pre-training data acquisition module is used to dock the protein pocket and the small molecule to generate a docking conformation, and use the docking conformation as the physical simulation data.

[0055] Optionally, when the pseudo-crystal data is included in the basic sample data,

[0056] The device further includes:

[0057] The complex crystal extraction module is used to traverse homologous protein data and extract complex crystal data of the homologous protein data.

[0058] The alignment protein acquisition module is used to perform pocket alignment processing on the target protein sequence in the homologous protein data to obtain the alignment protein.

[0059] The pseudo-crystal data generation module is used to combine the alignment protein with the ligand to generate the pseudo-crystal data.

[0060] Optionally, the alignment protein acquisition module includes:

[0061] The crystal data removal unit is used to compare protein sequences in homologous proteins and remove crystal data whose differences from other homologous protein sequences are greater than a set value.

[0062] The alignment protein acquisition unit is used to perform pocket alignment on the remaining crystal data using a point cloud matching algorithm to obtain the alignment protein.

[0063] Optionally, the model output includes multiple docking configurations, each determined by a different model input.

[0064] The loss value calculation module includes:

[0065] The first loss value calculation unit is used to calculate multiple first loss values ​​based on the conformational combination of the docking conformations;

[0066] The second loss value calculation unit is used to calculate multiple second loss values ​​based on the multiple docking conformations and the reference crystal conformation;

[0067] The loss value acquisition unit is used to perform a weighted summation of the plurality of first loss values ​​and the plurality of second loss values ​​to obtain the loss value.

[0068] Optionally, the device further includes:

[0069] The protein pocket acquisition module is used to acquire protein pockets and random initial conformations;

[0070] The protein pocket input module is used to input the protein pocket and the random initial conformation into the molecular docking model;

[0071] The predicted conformation acquisition module is used to call the molecular docking model to process the protein pocket and the random initial conformation to obtain the predicted docking conformation;

[0072] The target conformation acquisition module is used to perform conformation rationality constraint processing on the predicted docking conformation based on the conformation rationality constraint algorithm to obtain the final target docking conformation.

[0073] Optionally, the target conformation acquisition module includes:

[0074] Molecular conformation generation unit, used to generate small molecule conformations that satisfy statistical laws based on preset tools;

[0075] Initialization unit, used to initialize the initial quantities of conformational structure change actions;

[0076] The conformation acquisition unit is used to perform conformational structure change processing on the small molecule conformation based on the initial amount to obtain the updated small molecule conformation.

[0077] The difference loss calculation unit is used to calculate the conformational difference loss based on the updated small molecule conformation and the predicted docking conformation.

[0078] The update amount calculation unit is used to calculate the update amount of the conformational structure change action based on the conformational difference loss.

[0079] The structure change unit is used to update the initial quantity based on the update quantity, and to perform conformational structure change processing on the small molecule conformation;

[0080] The target conformation output unit is used to iteratively execute the update process until the conformation difference loss is lower than the loss threshold, and then outputs the target docking conformation.

[0081] Optionally, the conformational structure change action includes at least one of rotation, translation, and torsion.

[0082] Thirdly, embodiments of this application provide an electronic device, including:

[0083] A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the training method for the molecular docking model described in any of the preceding claims.

[0084] Fourthly, embodiments of this application provide a computer-readable storage medium that, when the instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to perform the training method for the molecular docking model described in any of the preceding claims.

[0085] Compared with the prior art, the embodiments of this application have the following advantages:

[0086] In this embodiment, the molecular docking model to be trained is pre-trained based on physical simulation data, which consists of docking conformations obtained from protein pockets and small molecule docking. The pre-trained molecular docking model is then fine-tuned using fine-tuning sample data, which is a mixture of complex crystal data and baseline sample data. The baseline sample data includes at least one of pseudo-crystal data and physical simulation data. A loss value is calculated based on the model output. When the loss value is within a preset range, the molecular docking model is obtained. This embodiment, by constructing pseudo-crystal data, enriches the amount of training data for the molecular docking model, thereby improving the performance of the trained molecular docking model and addressing the problem that limited data has severely restricted the development of deep learning docking methods.

[0087] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0088] Figure 1 A flowchart illustrating the steps of a training method for a molecular docking model provided in this application embodiment;

[0089] Figure 2 A schematic diagram illustrating a physical simulation data acquisition process provided in an embodiment of this application;

[0090] Figure 3 A schematic diagram illustrating a fine-tuning sample data acquisition process provided in an embodiment of this application;

[0091] Figure 4 A schematic diagram of a model fine-tuning process provided in an embodiment of this application;

[0092] Figure 5 A schematic diagram of a conformational rationality constraint process provided for an embodiment of this application;

[0093] Figure 6 A schematic diagram of the structure of a training device for a molecular docking model provided in an embodiment of this application;

[0094] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0095] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0096] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0097] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes said element.

[0098] Example 1

[0099] Reference Figure 1 The diagram illustrates a flowchart of the steps involved in training a molecular docking model according to an embodiment of this application. Figure 1 As shown, the training method for this molecular docking model may include the following steps:

[0100] Step 101: Pre-train the molecular docking model to be trained based on physical simulation data, wherein the physical simulation data is the docking conformation obtained by docking protein pockets and small molecules.

[0101] The embodiments of this application can be applied to scenarios where pseudo-crystal data is used to train molecular docking models.

[0102] The physical simulation data can be the docking conformations obtained from protein pockets and small molecule docking; this physical simulation data serves as the pre-training data for the molecular docking model. The process for acquiring the physical simulation data will be described in detail below with reference to the specific implementation method.

[0103] In one specific implementation of this application, before step 101 above, the following may also be included:

[0104] Step A1: Obtain protein crystals from the PDB database.

[0105] In this embodiment of the application, when preparing physical simulation data, a PDB database (containing various protein crystals), a drug database (containing various drug molecules), and a commercial molecule library (containing various synthetic molecules) can be downloaded.

[0106] First, protein crystals can be obtained from the PDB database.

[0107] Step A2: Perform pocket segmentation on the protein crystal to obtain protein pockets.

[0108] After obtaining protein crystals from the PDB database, the protein crystals can be segmented into pockets to obtain protein pockets. In specific implementations, pocketing of the protein crystals can be performed based on ligands or using algorithms such as FPocket.

[0109] Understandably, each protein pocket is only a part of a protein crystal, not a molecule.

[0110] After the protein crystals are pocketed to obtain protein pockets, step A3 is performed.

[0111] Step A3: For each protein pocket, randomly select small molecules from a preset database.

[0112] After segmenting the protein crystals into pockets to obtain protein pockets, small molecules can be randomly selected from a pre-defined database for each protein pocket. The pre-defined database can be any molecular database or a combination of multiple molecular databases, such as drugs, drugBank, zinc, etc.

[0113] In practice, the obtained pockets can be looped a set number of times (this parameter can be adjusted as needed), and each time a random number is used to control the random selection of small molecules from a commercial molecular library or from a Drugs database.

[0114] After randomly selecting small molecules from a pre-defined database for each protein pocket, proceed to step A4.

[0115] Step A4: Dot the protein pocket and the small molecule to generate a docking conformation, and use the docking conformation as the physical simulation data.

[0116] Conformation refers to the spatial configuration or shape of a molecule. It can refer to the conformation of a single molecule or the conformation of a group of molecules. In molecular conformation, the atoms in a molecule can take different arrangements, and these arrangements determine the molecule's physical and chemical properties.

[0117] After randomly selecting small molecules from a preset database for each protein pocket, the protein pocket and the small molecule can be docked to generate docking conformations, which are then used as physical simulation data. Specifically, existing molecular docking algorithms (such as AutoDock Vina, Glide, etc.) can be used to dock the protein pocket and the small molecule, and each protein pocket can obtain the same number of docking conformations as the set number of times.

[0118] Large-scale pre-training is one of the most important methods that has achieved major breakthroughs in the field of artificial intelligence in recent years. The BERT / GPT series models have accumulated a lot of knowledge during the pre-training process by pre-training on large-scale text data, which has greatly improved the performance of tools such as search and dialogue assistants. However, in the field of molecular docking, there is an extreme lack of data, let alone pre-training data. This application embodiment constructs large-scale pre-training data by utilizing existing physical docking methods and PDB databases.

[0119] The process for acquiring physical simulation data can be combined with... Figure 2 The following is a detailed description.

[0120] Reference Figure 2 The diagram illustrates a physical simulation data acquisition process provided in an embodiment of this application. Figure 2 As shown, the physical simulation data acquisition process may include the following steps:

[0121] Step 1: Download the PDB database, Drugs database, and commercial molecular libraries.

[0122] Step 2: Traverse the PDB database and cut pockets for each crystal encountered to obtain several protein pockets.

[0123] Step 3: For the obtained pocket, perform 3000 cycles (this parameter can be adjusted as needed), each time using a random number to control the random selection of small molecules from a commercial molecular library or from the Drugs database.

[0124] Step 4: The obtained pocket-small molecule pairs are docked using existing docking algorithms (such as AutoDock Vina, Glide, etc.). Each pocket can obtain docking conformations of 3000 different molecules.

[0125] Step 5: Determine whether the traversal of the PDB database is complete. If so, end the process and use all the obtained docking configurations as physical simulation data.

[0126] Pre-training refers to training a model using labeled data with low accuracy or automatically generated labels, enabling it to learn the basic patterns of the data. The resulting model is called a pre-trained model.

[0127] After obtaining the physical simulation data, it can be used to pre-train the molecular docking model to be trained. In practice, the physical simulation data can be divided into several batches of subsample data, and then each batch of subsample data can be combined to perform a set number of training rounds to complete the pre-training process of the docking model.

[0128] After pre-training the molecular docking model to be trained based on physical simulation data, step 102 is performed.

[0129] Step 102: Fine-tune the pre-trained molecular docking model based on the fine-tuning sample data to obtain the model output. The fine-tuning sample data is a mixture of complex crystal data and basic sample data. The basic sample data includes at least one of pseudo-crystal data and physical simulation data.

[0130] Fine-tuning sample data refers to sample data used to fine-tune the pre-trained molecular docking model. In this example, the fine-tuning sample data may be a mixture of complex crystal data and baseline sample data, wherein the baseline sample data includes at least one of pseudo-crystal data and physical simulation data.

[0131] When the fine-tuning sample data includes pseudo-crystal data, the process for obtaining the pseudo-crystal data can be described in detail below in conjunction with the specific implementation method.

[0132] In one specific implementation of this application, before step 102 above, the following may also be included:

[0133] Step B1: Traverse the homologous protein data and extract the complex crystal data of the homologous protein data.

[0134] In this embodiment, the complex can be a complex composed of multiple types of molecules. This scheme generally refers to a complex composed of proteins and small molecules.

[0135] By utilizing homologous protein information in the PDB database, "pseudo-crystal data" of homologous proteins and crystal ligands can be constructed. The "pseudo-crystal data" can be mixed with complex crystal data and supplemented with a certain proportion of pre-training data to obtain fine-tuned data.

[0136] Complex crystals represent a one-to-one correspondence between proteins and active ligands. In most cases, a homologous protein often has multiple active ligands. Different active ligands with similar modes of action can cause different and slight conformational changes in the protein. These data can be used to construct "pseudo-crystal data" of homologous proteins and active ligands to expand the amount of crystal data. Furthermore, this embodiment considers that training only active ligands during the fine-tuning phase can bias the model towards predicting stronger forces, which is not conducive to the scoring function distinguishing whether the docked molecule possesses activity. Therefore, for each homologous protein, in each training round, a pre-training dataset of the corresponding protein can be randomly selected and added to the fine-tuning dataset.

[0137] First, we can traverse homologous protein data and obtain their complex crystals.

[0138] After traversing the homologous protein data and extracting the crystal data of the complexes from the homologous protein data, step B2 is executed.

[0139] Step B2: Perform pocket alignment on the target protein sequence in the homologous protein data to obtain aligned proteins.

[0140] After traversing the homologous protein data and extracting the complex crystal data of the homologous protein data, pocket alignment can be performed on the target protein sequence in the homologous protein data to obtain aligned proteins. The target protein sequence refers to the protein sequence in the homologous protein data, excluding crystal data that differs significantly from other homologous protein sequences. Specifically, the alignment process can be as follows: compare the protein sequences in the homologous protein data, remove crystal data whose differences from other homologous protein sequences exceed a set value, and use a point cloud matching algorithm (such as Kabsh, ICP, or other point cloud alignment algorithms) to perform pocket alignment on the remaining crystal data to obtain aligned proteins.

[0141] In practice, the protein sequences of homologous proteins can be compared, and crystal data that differs too much from other homologous protein sequences can be removed. The remaining data can be aligned using a point cloud matching algorithm (since the protein data comes from different sources during the acquisition process, the alignment algorithm here rotates and translates each protein to overlap the pockets as much as possible).

[0142] After pocket alignment of the target protein sequence in the homologous protein data to obtain the aligned protein, step B3 is performed.

[0143] Step B3: Combine the alignment protein with the ligand to generate the pseudo-crystal data.

[0144] After obtaining the alignment protein, it can be combined with ligands to generate crystal data.

[0145] The above implementation process can be described as follows: Figure 3 As shown.

[0146] Reference Figure 3 This diagram illustrates a fine-tuning sample data acquisition process provided in an embodiment of this application. Figure 3 As shown, the fine-tuning sample data acquisition process may include the following steps:

[0147] Step 1: Traverse the homologous protein data to obtain the complex crystal.

[0148] Step 2: Determine if the protein sequences are sufficiently similar. If so, proceed to Step 3. If not, continue iterating through the homologous protein data.

[0149] Step 3: Align protein pockets using a point cloud matching algorithm. Specifically, compare the protein sequences in homologous proteins, remove crystal data that differs too much from other homologous protein sequences, and use a point cloud matching algorithm to align the remaining pockets (since the protein data come from different sources during acquisition, the alignment algorithm here rotates and translates each protein to overlap the pockets as much as possible).

[0150] Step 4: Combine the protein and ligand to obtain pseudo-crystal data.

[0151] Step 5: Randomly select one pre-training data for the homologous protein, and use the complex crystal data, pseudo-crystal data, and pre-training data together as fine-tuning sample data.

[0152] Step 6: Determine if the homologous protein data traversal is complete. If yes, end the process. If not, continue traversing the homologous protein data.

[0153] Fine-tuning refers to training a pre-trained model using precise data, enabling the model to predict results that are more closely aligned with reality.

[0154] After obtaining the fine-tuning sample data, the pre-trained molecular docking model can be fine-tuned based on the fine-tuning sample data to obtain the model output.

[0155] Understandably, in this example, the model output may include multiple docking configurations, each determined by a different model input. Specifically, it may include the following four docking configurations:

[0156] 1. Docking conformation 1: This model uses the crystal acceptor conformation and the ligand conformation after random rotation and translation as input. The model only needs to predict the rotation and translation of the ligand. The task is relatively simple, making docking conformation 1 quite similar to the crystal conformation.

[0157] 2. Docking conformation 2: This model uses the crystal acceptor conformation and random ligand conformations as input. In addition to predicting the rotation and translation of the ligands, it needs to further predict the bond length, bond angle, and torsion angle of the ligands. The task difficulty is moderate.

[0158] 3. Docking conformation 3: This model uses a noisy receptor conformation and a random ligand conformation as input. It needs to predict not only the ligand's rotation, translation, bond length, bond angle, and torsion angle, but also denoise the receptor. This is the most challenging task and results in docking conformation 3 being the least accurate. (Noisy receptor conformations include, but are not limited to, the following: 1. Protein conformations predicted by models such as AlphaFold2 / ESMFold; 2. Random noise added to the original receptor crystal conformation; 3. Conformations of homologous proteins induced by other active ligands.)

[0159] 4. Docking Conformation 4: The model is trained using pre-training data. Given a crystal acceptor conformation and random inactive ligand conformations, the model learns the docking results from other software. Since the force strength of the docking results from other software is often weaker than that of the crystal conformation, after this training, the model will impose certain restrictions on the force strength of the inactive conformation, thereby assisting the scoring function in distinguishing between active and inactive ligands.

[0160] Understandably, among the four docking conformations mentioned above, docking conformation 4 is optional, meaning that only docking conformations 1, 2, and 3 are predicted. Furthermore, any two of docking conformations 1, 2, and 3 can be selected for prediction; this embodiment does not impose any restrictions on this.

[0161] After fine-tuning the pre-trained molecular docking model based on fine-tuning sample data to obtain the model output, step 103 is executed.

[0162] Step 103: Calculate the loss value based on the model output.

[0163] After fine-tuning the pre-trained molecular docking model using fine-tuning sample data to obtain the model output, the loss value can be calculated based on the model output. In this example, the loss value can be obtained by summing multiple losses. Specifically, multiple first loss values ​​can be calculated based on the conformational combinations of docking conformations. Multiple second loss values ​​are then calculated based on the multiple docking conformations and their corresponding baseline crystal conformations. The weighted sum of the multiple first loss values ​​and multiple second loss values ​​yields the final loss value. The baseline crystal conformation can be a pre-configured standard conformation of a protein pocket.

[0164] The process for calculating the loss value can be combined with... Figure 5 The following is a detailed description.

[0165] like Figure 5 As shown, the crystal acceptor conformation and the randomly rotated and translated crystal ligand conformation are used as inputs, and the output after processing by the molecular docking model is: Docking Conformation 1. The crystal acceptor conformation and the random ligand conformation are used as inputs, and the output after processing by the molecular docking model is: Docking Conformation 2. The noisy acceptor conformation and the random ligand conformation are used as inputs, and the output after processing by the molecular docking model is: Docking Conformation 3. The crystal acceptor conformation and the random inactive ligand conformation are used as inputs, and the output after processing by the molecular docking model is: Docking Conformation 4. Among them, Docking Conformation 1 assists in predicting Docking Conformation 2, and the corresponding loss is Loss 1. Docking Conformation 2 assists in predicting Docking Conformation 3, and the corresponding loss is Loss 2. Docking Conformation 1 assists in predicting Docking Conformation 2 and Docking Conformation 3, and the corresponding loss is Loss 3. Learning from the crystal conformation through Docking Conformation 1, Docking Conformation 2, and Docking Conformation 3, the corresponding losses are Loss 4, Loss 5, and Loss 6, respectively. Learning from other software docking conformations through Docking Conformation 4, the corresponding loss is Loss 7. The final loss value can be obtained by weighted summation of these 7 losses.

[0166] Understandably, losses 1, 2, and 3 can be optional. That is, when calculating the loss value, only losses 4, 5, 6, and 7 are calculated, and the final loss value is obtained by weighted summation of these four losses.

[0167] After calculating the loss value, proceed to step 104.

[0168] Step 104: If the loss value is within a preset range, the molecular docking model is obtained.

[0169] After calculating the loss value, it can be determined whether the loss value is within the preset range.

[0170] If the loss value is not within the preset range, it indicates that the molecular docking model has not converged. In this case, the molecular docking model can be trained again using more training sample data. Specifically, the pre-training and fine-tuning training process can be performed again until the molecular docking model converges.

[0171] If the loss value is within the preset range, it means that the molecular docking model has converged. At this time, the fine-tuned molecular docking model can be used as the final molecular docking model, which can then be applied to the subsequent molecular docking inference process.

[0172] This application's embodiments construct pseudo-crystal data, thereby enriching the amount of training data for the molecular docking model and improving its performance. This addresses the problem that limited data has severely restricted the development of deep learning docking methods. Furthermore, the accuracy of the molecular docking model can be improved through pre-training and model fine-tuning processes.

[0173] After training the molecular docking model, it can be applied to the subsequent molecular docking conformation prediction process. After obtaining the predicted conformation output by the model, reasonable constraints can be applied to the predicted conformation to obtain the final docking conformation. This implementation process can be described in detail below with reference to the specific implementation method.

[0174] In one specific implementation of this application, after step 104 above, the following may also be included:

[0175] Step C1: Obtain the protein pocket and random initial conformation.

[0176] In the embodiments of this application, after training the molecular docking model, when performing molecular docking conformation prediction, protein pockets and random initial conformations can be obtained.

[0177] After obtaining the protein pocket and random initial conformation, proceed to step C2.

[0178] Step C2: Input the protein pocket and the random initial conformation into the molecular docking model.

[0179] After obtaining the protein pocket and random initial conformation, the protein pocket and random initial conformation can be input into the molecular docking model to predict the molecular docking conformation.

[0180] After inputting the protein pocket and random initial conformation into the molecular docking model, step C3 is performed.

[0181] Step C3: Call the molecular docking model to process the protein pocket and the random initial conformation to obtain the predicted docking conformation.

[0182] After inputting the protein pocket and random initial conformation into the molecular docking model, the model can be called to process the protein pocket and random initial conformation to obtain the predicted docking conformation. In other words, by processing the protein pocket and random initial conformation through the molecular docking model, the model outputs the predicted docking conformation.

[0183] After obtaining the predicted docking conformation, proceed to step C4.

[0184] Step C4: Based on the conformation rationality constraint algorithm, perform conformation rationality constraint processing on the predicted docking conformation to obtain the final target docking conformation.

[0185] After obtaining the predicted docking conformation, a conformation rationality constraint algorithm can be used to apply constraints to the predicted docking conformation to obtain the final target docking conformation. The specific constraint process can be as follows: Figure 5 As shown.

[0186] Reference Figure 5 This diagram illustrates a conformational rationality constraint process provided in an embodiment of this application. Figure 5 As shown, the process may include the following steps:

[0187] 1. Generate small molecule conformations that satisfy statistical laws based on preset tools. That is, generate reasonable small molecule conformations. In this example, preset tools can be RDKIT, OpenBabel, etc. Statistical laws refer to the general distribution of bond lengths and bond angles of various types of bond conformations observed in reality, such as the six bond angles within a benzene ring being always equal. The criterion for judging whether a conformation is reasonable or unreasonable is whether it satisfies this statistical law.

[0188] 2. Initialize the initial values ​​for the configuration structure change actions; In this example, the configuration structure change actions can include at least one of the following: rotation, translation, and torsion. In the specific implementation, if there are docking configurations, the rotation matrix and translation amount between the overall reasonable configuration and the docking configuration are calculated using a point cloud alignment algorithm. At this time, the torsion angle update amount is initialized to 0, that is, the torsion angle is not changed during initialization; if there are no docking configurations, all values ​​are initialized to 0.

[0189] 3. Based on the initial values, perform conformational structure modification on the small molecule conformation to obtain an updated small molecule conformation. That is, based on the updated values ​​of the rotation matrix, translation amount, and torsion angle obtained above, perform rotation, translation, and torsion on the reasonable conformation to obtain the updated reasonable conformation.

[0190] 4. Based on the updated small molecule conformation and the predicted docking conformation, the conformational difference loss is calculated. Specifically, the conformational difference loss can be calculated by comparing the updated reasonable conformation with the predicted results (such as the distance matrix) of the docking conformation or other models that represent the conformation.

[0191] 5. Based on the conformational difference loss, calculate the update amount of the conformational structure change action.

[0192] 6. Update the initial quantity based on the updated quantity, and perform conformational structure change processing on the small molecule conformation.

[0193] 7. Iterate through the update process until the conformational difference loss is lower than the loss threshold, then output the target docking conformation. That is, after multiple iterations where the loss is less than a certain threshold, end the iteration and output the updated conformation.

[0194] The molecular docking model training method provided in this application pre-trains the molecular docking model based on physical simulation data, which consists of docking conformations obtained from protein pockets and small molecule docking. The pre-trained molecular docking model is then fine-tuned using fine-tuning sample data, which is a mixture of complex crystal data and baseline sample data. The baseline sample data includes at least one of pseudo-crystal data and physical simulation data. A loss value is calculated based on the model output. When the loss value is within a preset range, the molecular docking model is obtained. This application embodiment, by constructing pseudo-crystal data, enriches the amount of training data for the molecular docking model, thereby improving the performance of the trained molecular docking model and addressing the problem that limited data has severely restricted the development of deep learning docking methods.

[0195] Example 2

[0196] Reference Figure 6 The diagram shows a schematic representation of a training device for a molecular docking model provided in an embodiment of this application. Figure 6 As shown, the training device 600 for the molecular docking model may include the following modules:

[0197] The pre-training module 610 is used to pre-train the molecular docking model to be trained based on physical simulation data, wherein the physical simulation data is the docking conformation obtained by docking protein pockets and small molecules.

[0198] The fine-tuning module 620 is used to fine-tune the pre-trained molecular docking model based on fine-tuning sample data to obtain the model output. The fine-tuning sample data is a mixture of complex crystal data and basic sample data. The basic sample data includes at least one of pseudo-crystal data and physical simulation data.

[0199] The loss value calculation module 630 is used to calculate the loss value based on the model output;

[0200] The molecular docking model acquisition module 640 is used to obtain the molecular docking model when the loss value is within a preset range.

[0201] Optionally, the device further includes:

[0202] The protein crystal acquisition module is used to acquire protein crystals from the PDB database;

[0203] The protein pocket acquisition module is used to perform pocket segmentation processing on the protein crystal to obtain protein pockets;

[0204] The small molecule screening module is used to randomly select small molecules from a preset database for each protein pocket;

[0205] The pre-training data acquisition module is used to dock the protein pocket and the small molecule to generate a docking conformation, and use the docking conformation as the physical simulation data.

[0206] Optionally, when the pseudo-crystal data is included in the basic sample data,

[0207] The device further includes:

[0208] The complex crystal extraction module is used to traverse homologous protein data and extract complex crystal data of the homologous protein data.

[0209] The alignment protein acquisition module is used to perform pocket alignment processing on the target protein sequence in the homologous protein data to obtain the alignment protein.

[0210] The pseudo-crystal data generation module is used to combine the alignment protein with the ligand to generate the pseudo-crystal data.

[0211] Optionally, the alignment protein acquisition module includes:

[0212] The crystal data removal unit is used to compare protein sequences in homologous proteins and remove crystal data whose differences from other homologous protein sequences are greater than a set value.

[0213] The alignment protein acquisition unit is used to perform pocket alignment on the remaining crystal data using a point cloud matching algorithm to obtain the alignment protein.

[0214] Optionally, the model output includes multiple docking configurations, each determined by a different model input.

[0215] The loss value calculation module includes:

[0216] The first loss value calculation unit is used to calculate multiple first loss values ​​based on the conformational combination of the docking conformations;

[0217] The second loss value calculation unit is used to calculate multiple second loss values ​​based on the multiple docking conformations and the reference crystal conformation;

[0218] The loss value acquisition unit is used to add the plurality of first loss values ​​and the plurality of second loss values ​​to obtain the loss value.

[0219] Optionally, the device further includes:

[0220] The protein pocket acquisition module is used to acquire protein pockets and random initial conformations;

[0221] The protein pocket input module is used to input the protein pocket and the random initial conformation into the molecular docking model;

[0222] The predicted conformation acquisition module is used to call the molecular docking model to process the protein pocket and the random initial conformation to obtain the predicted docking conformation;

[0223] The target conformation acquisition module is used to perform conformation rationality constraint processing on the predicted docking conformation based on the conformation rationality constraint algorithm to obtain the final target docking conformation.

[0224] Optionally, the target conformation acquisition module includes:

[0225] Molecular conformation generation unit, used to generate small molecule conformations that satisfy statistical laws based on preset tools;

[0226] Initialization unit, used to initialize the initial quantities of conformational structure change actions;

[0227] The conformation acquisition unit is used to perform conformational structure change processing on the small molecule conformation based on the initial amount to obtain the updated small molecule conformation.

[0228] The difference loss calculation unit is used to calculate the conformational difference loss based on the updated small molecule conformation and the predicted docking conformation.

[0229] The update amount calculation unit is used to calculate the update amount of the conformational structure change action based on the conformational difference loss.

[0230] The structure change unit is used to update the initial quantity based on the update quantity, and to perform conformational structure change processing on the small molecule conformation;

[0231] The target conformation output unit is used to iteratively execute the update process until the conformation difference loss is lower than the loss threshold, and then outputs the target docking conformation.

[0232] Optionally, the conformational structure change action includes at least one of rotation, translation, and torsion.

[0233] The molecular docking model training device provided in this application pre-trains the molecular docking model based on physical simulation data, which consists of docking conformations obtained from protein pockets and small molecule docking. The pre-trained molecular docking model is then fine-tuned using fine-tuning sample data, which is a mixture of complex crystal data and baseline sample data. The baseline sample data includes at least one of pseudo-crystal data and physical simulation data. A loss value is calculated based on the model output. When the loss value is within a preset range, the molecular docking model is obtained. This application embodiment, by constructing pseudo-crystal data, enriches the amount of training data for the molecular docking model, thereby improving the performance of the trained molecular docking model and addressing the problem that limited data has severely restricted the development of deep learning docking methods.

[0234] Example 3

[0235] This application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the above-described method for training the molecular docking model.

[0236] Figure 7 A schematic diagram of the structure of an electronic device 700 according to an embodiment of the present invention is shown. Figure 7 As shown, the electronic device 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RAM) 703. The RAM 703 can also store various programs and data required for the operation of the electronic device 700. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0237] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, microphone, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0238] The various processes and handling described above can be executed by processing unit 701. For example, the methods of any of the above embodiments can be implemented as computer software programs, which are tangibly contained in a computer-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by CPU 701, one or more actions of the methods described above can be performed.

[0239] Example 4

[0240] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for training the molecular docking model.

[0241] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0242] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0243] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminals (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0244] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0245] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal, causing a series of operational steps to be executed on the computer or other programmable terminal to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0246] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0247] The above provides a detailed description of the training method, apparatus, electronic device, and computer-readable storage medium for a molecular docking model provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A training method for a molecular docking model, characterized in that, The method includes: Pre-training is performed on the molecular docking model to be trained based on physical simulation data, wherein the physical simulation data is the docking conformation obtained by docking protein pockets and small molecules. The pre-trained molecular docking model is fine-tuned based on fine-tuning sample data to obtain model output. The fine-tuning sample data is a mixture of complex crystal data and basic sample data. The basic sample data includes at least one of pseudo-crystal data and physical simulation data. The model output includes multiple docking conformations, each of which is determined by different model inputs. The model inputs include: crystal acceptor conformation and ligand conformation after random rotation and translation, crystal acceptor conformation and random ligand conformation, and noisy acceptor conformation and random ligand conformation. Based on the model output, the loss value is calculated, including: calculating multiple first loss values ​​based on the conformational combination of the docking conformations; calculating multiple second loss values ​​based on the multiple docking conformations and the reference crystal conformation; and performing a weighted summation of the multiple first loss values ​​and the multiple second loss values ​​to obtain the loss value. The molecular docking model is obtained when the loss value is within a preset range.

2. The method according to claim 1, characterized in that, Before pre-training the molecular docking model to be trained based on the physical simulation data, the following is also included: Obtain protein crystals from the PDB database; The protein crystals are pocketed to obtain protein pockets; For each of the protein pockets, small molecules are randomly selected from a preset database; The protein pocket and the small molecule are docked to generate a docking conformation, and the docking conformation is used as the physical simulation data.

3. The method according to claim 1, characterized in that, When the pseudo-crystal data is included in the basic sample data, Before fine-tuning the pre-trained molecular docking model based on fine-tuning sample data to obtain the model output, the following steps are also included: Traverse the homologous protein data and extract the complex crystal data of the homologous protein data; The target protein sequence in the homologous protein data is pocket-aligned to obtain aligned proteins; The alignment protein is combined with the ligand to generate the pseudo-crystal data.

4. The method according to claim 3, characterized in that, The process of pocket-aligning the target protein sequence in the homologous protein data to obtain aligned proteins includes: Compare the protein sequences of homologous proteins and remove crystal data whose sequences differ from other homologous protein sequences by a set value; The remaining crystal data was pocket-aligned using a point cloud matching algorithm to obtain the aligned protein.

5. The method according to claim 1, characterized in that, After obtaining the molecular docking model, the process further includes: Obtain protein pockets and random initial conformations; The protein pocket and the random initial conformation are input into the molecular docking model; The molecular docking model is invoked to process the protein pocket and the random initial conformation to obtain the predicted docking conformation; Based on the conformation rationality constraint algorithm, the predicted docking conformation is subjected to conformation rationality constraint processing to obtain the final target docking conformation.

6. The method according to claim 5, characterized in that, The conformational rationality constraint algorithm applies conformational rationality constraint processing to the predicted docking conformation to obtain the final target docking conformation, including: Based on preset tools, small molecule conformations that satisfy statistical laws are generated; Initialize the initial values ​​for the conformational structure change action; Based on the initial amount, the small molecule conformation is subjected to conformational structure change processing to obtain the updated small molecule conformation; Based on the updated small molecule conformation and the predicted docking conformation, the conformational difference loss is calculated; Based on the conformational difference loss, the update amount of the conformational structure change action is calculated; The initial quantity is updated based on the updated quantity, and the conformational structure of the small molecule is processed by a conformational structure change action. The update process is executed iteratively until the conformational difference loss is lower than the loss threshold, at which point the target docking conformation is output.

7. The method according to claim 6, characterized in that, The conformational structural change actions include at least one of the following: rotation, translation, and torsion.

8. A training device for a molecular docking model, characterized in that, The device includes: The pre-training module is used to pre-train the molecular docking model to be trained based on physical simulation data, wherein the physical simulation data is the docking conformation obtained by docking protein pockets and small molecules. The fine-tuning module is used to fine-tune the pre-trained molecular docking model based on fine-tuning sample data to obtain model output. The fine-tuning sample data is a mixture of complex crystal data and basic sample data. The basic sample data includes at least one of pseudo-crystal data and physical simulation data. The model output includes multiple docking conformations, each of which is determined by different model inputs. The model inputs include: crystal acceptor conformation and ligand conformation after random rotation and translation, crystal acceptor conformation and random ligand conformation, and noisy acceptor conformation and random ligand conformation. A loss value calculation module is used to calculate a loss value based on the model output; the loss value calculation module includes: a first loss value calculation unit, used to calculate multiple first loss values ​​based on the conformational combination of the docking conformations; a second loss value calculation unit, used to calculate multiple second loss values ​​based on the multiple docking conformations and the reference crystal conformation; and a loss value acquisition unit, used to add the multiple first loss values ​​and the multiple second loss values ​​to obtain the loss value; The molecular docking model acquisition module is used to obtain the molecular docking model when the loss value is within a preset range.

9. The apparatus according to claim 8, characterized in that, The device further includes: The protein crystal acquisition module is used to acquire protein crystals from the PDB database; The protein pocket acquisition module is used to perform pocket segmentation processing on the protein crystal to obtain protein pockets; The small molecule screening module is used to randomly select small molecules from a preset database for each protein pocket; The pre-training data acquisition module is used to dock the protein pocket and the small molecule to generate a docking conformation, and use the docking conformation as the physical simulation data.

Citation Information

Patent Citations

  • Training method and device for in-pocket molecule generation model

    CN116206677A

  • Method for predicting docking posture between protein and ligand based on graph neural network

    CN116343910A