Method, device and application for constructing target pocket structure of target protein
By constructing a target pocket structure method and deep learning model for the target protein, the problem of identifying the dynamic and diverse structure of the p53 protein was solved, and efficient and accurate prediction of new pocket structures and binding molecule screening were achieved to activate mutant p53 protein and restore its tumor suppressor function.
Patent Information
- Application Number
- CN202410493578.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-04-23
AI Technical Summary
Existing technologies make it difficult to accurately identify and describe the structural dynamics and diversity of the p53 protein, resulting in insufficient understanding of its stabilization pockets, which affects the search for new stabilization pockets and the design of small molecule drugs.
By obtaining multiple crystal files, generating candidate pocket structures, selecting benchmark pocket structures, screening and summarizing based on similarity, combining deep learning models to predict the binding of target proteins, generating target pocket structures of target proteins, screening out highly similar and stable pocket structures, and using the Pocket2Mol model to predict binding molecules.
The accuracy and efficiency of constructing new pocket structures of target proteins have been improved, and the binding molecules of target proteins can be predicted more accurately, activating mutant p53 proteins and restoring their tumor suppressor function.
Smart Images

Figure CN118430638B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence and drug discovery technology, and in particular to a method for constructing a target pocket structure of a target protein, a method for predicting target protein binding molecules, target protein binding molecules, related devices, computer program products, computing equipment, and computer-readable storage media. Background Art
[0002] p53 is a collective term for a family of homologous proteins known as tumor suppressor proteins (also called p53 proteins or p53 tumor proteins). Encoded by the TP53 (human) and Trp53 (mouse) genes, p53 is one of the earliest tumor suppressor genes discovered. p53 proteins regulate the cell cycle, promoting apoptosis and cellular senescence, thereby preventing cancer. p53 proteins maintain genomic stability, preventing or reducing mutations. Hence, they are called the guardians of the genome.
[0003] In current cancer research, the p53 protein plays a key role as a tumor suppressor gene. However, structural changes caused by mutations cause p53 to lose its normal physiological function, limiting its ability to inhibit cancer cell growth. Although p53 mutant isoforms, such as Y220C and the L1 / S3 pocket, have been extensively studied, existing research has primarily focused on the activation of single mutant isoforms, and broader and more in-depth exploration is still insufficient. In addition, the structural complexity of p53 poses a significant challenge to research. Its high degree of dynamics and diverse mutational isoforms make the accurate identification and description of pockets extremely difficult, affecting the accurate and comprehensive understanding of new stabilization pockets.
[0004] Specifically, the p53 protein is highly structurally dynamic, making it difficult to fully capture its changes in different structural states using traditional experimental techniques. Methods such as X-ray crystallography and nuclear magnetic resonance may not provide sufficient spatial resolution, limiting a comprehensive understanding of the dynamics of p53 structure, especially when searching for new stabilizing pockets.
[0005] On the other hand, the identification and computational modeling of pockets are also fraught with uncertainty. Traditional structure prediction methods face uncertainties in identifying and describing additional stabilizing pockets in the p53 protein. The shape and properties of pockets can be difficult to accurately predict, while computational models can be limited in their prediction of novel pockets by structural complexity and interaction diversity without further molecular dynamics simulation and experimental validation to confirm their accuracy.
[0006] Therefore, current methods for finding the p53 stabilization pocket still need to be improved. Summary of the Invention
[0007] The present application aims to solve at least one of the existing problems. To this end, the present application proposes a method for constructing a new stabilizing pocket for a target protein.
[0008] Specifically, this application provides the following technical solutions:
[0009] In the first aspect of the present application, the present application proposes a method for constructing a target pocket structure of a target protein. According to an embodiment of the present application, the target pocket structure is used to predict the binding molecules of the target protein, and the method includes: obtaining multiple crystal files derived from the target protein or a mutation thereof, at least one of the multiple crystal files is derived from a mutant of the target protein; based on the multiple crystal files, respectively generating candidate pocket structures corresponding to the crystal files; selecting one of the multiple crystal files as a standard crystal file, and determining the candidate pocket structure corresponding to the standard crystal file as the benchmark pocket structure; based on the similarity between the benchmark pocket structure and the other candidate pocket structures, screening the other candidate pocket structures to obtain screened candidate pockets; and summarizing the benchmark pocket structure with the screened candidate pockets to obtain the target pocket structure. The above method can be used to accurately construct a new pocket structure of the target protein, while improving the construction efficiency, which helps to more accurately predict the binding molecules of the target protein. In small molecule drug design, the target pocket structure of the target protein is obtained by this method, thereby screening small molecule drugs that bind to the pocket structure.
[0010] In the second aspect of the present application, the present application proposes a method for predicting target protein binding molecules. According to an embodiment of the present application, the method comprises: constructing a target pocket structure of a target protein based on the method described in the first aspect; obtaining a crystal file corresponding to the target pocket structure, and inputting the crystal file into a trained machine learning model to obtain a target protein binding molecule. This method can accurately and reliably predict target protein binding molecules, avoiding speculation about the binding mode between the target protein and the binding molecule. In some examples of the present application, this method can be used to predict small molecule drugs that can bind to the target protein.
[0011] In a third aspect of the present application, the present application provides a target protein binding molecule. According to embodiments of the present application, the target protein binding molecule is predicted and obtained by the method of the second aspect. In some examples of the present application, the aforementioned target protein binding molecule can be used as a small molecule drug for disease treatment or scientific research.
[0012] In the fourth aspect of the present application, the present application proposes a device for constructing a target pocket structure of a target protein. According to an embodiment of the present application, the device includes: a crystal file acquisition unit for acquiring multiple crystal files derived from the target protein or a mutation thereof, at least one of the multiple crystal files being derived from a mutant of the target protein; a candidate pocket structure generation unit for generating candidate pocket structures corresponding to the crystal files based on the multiple crystal files; a reference pocket structure determination unit for selecting one of the multiple crystal files as a standard crystal file and determining the candidate pocket structure corresponding to the standard crystal file as the reference pocket structure; a candidate pocket screening unit for screening the other candidate pocket structures based on the similarity between the reference pocket structure and the other candidate pocket structures to obtain a screened candidate pocket; and a target pocket acquisition unit for summarizing the reference pocket structure with the screened candidate pocket to obtain the target pocket structure. In some examples of the present application, the aforementioned device can efficiently construct the target pocket structure of the target protein with high accuracy and operability. Based on the obtained accurate target pocket structure, accurate prediction of the binding molecules of the target protein can be achieved.
[0013] In the fifth aspect of the present application, the present application proposes a target protein binding molecule prediction system. According to an embodiment of the present application, the system includes: a target pocket construction device for constructing a target pocket structure of a target protein based on the device described in the third aspect; a target protein binding molecule prediction device for obtaining a crystal file corresponding to the target pocket structure, and inputting the crystal file into a trained machine learning model to obtain a target protein binding molecule. In some examples of the present application, the aforementioned system can comprehensively predict the binding molecules of the target protein with high accuracy, efficiency and ease of operation.
[0014] In a sixth aspect of the present application, a computer program product is provided. According to an embodiment of the present application, the computer program product comprises computer instructions; when some or all of the computer instructions are executed on a computer, the method for constructing a target pocket structure of a target protein as described in the first aspect of the present application or the method for predicting target protein binding molecules as described in the second aspect of the present application is executed.
[0015] In a seventh aspect, the present application provides a computing device. According to an embodiment of the present application, the computing device includes: a processor and a memory; the memory is configured to store a computer program; and the processor is configured to execute the computer program to implement the method for constructing a target pocket structure of a target protein as described in the first aspect of the present application or the method for predicting target protein binding molecules as described in the second aspect.
[0016] In an eighth aspect of the present application, the present application provides a computer-readable storage medium. According to an embodiment of the present application, the computer-readable storage medium stores computer instructions or a program that, when executed on a computer, causes the method for constructing a target pocket structure of a target protein as described in the first aspect of the present application or the target protein binding molecule prediction method as described in the second aspect of the present application to be executed.
[0017] In some examples of this application, the aforementioned computer program products, computing devices, and computer-readable storage media achieve efficient automation and improve efficiency and accuracy by automatically executing computer instructions to construct a target pocket structure for a target protein or a method for predicting target protein binding molecules. Furthermore, the instruction-based nature of these methods allows for high consistency and reliability across diverse environments.
[0018] Additional aspects and advantages of the present application will be given in part in the following description and in part will become obvious from the following description or will be learned through practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 Schematic diagram of a device for constructing a target pocket structure of a target protein provided in a specific embodiment of the present application;
[0021] Figure 2 Schematic diagram of the target protein binding molecule prediction system provided for the specific embodiment of this application;
[0022] Figure 3 Schematic diagram of the screening process for the P53 protein stabilization pocket and binding molecules provided in the examples of this application;
[0023] Figure 4 is a schematic diagram of the volume distribution of protein pockets provided in an embodiment of the present application; wherein A is pocket 1; B is pocket 2; C is pocket 3; and D is pocket 4;
[0024] Figure 5 Schematic diagram of the molecular dynamics simulation results of small molecules (target protein binding molecules) obtained based on pocket4 in the examples provided in this application;
[0025] Figure 6Schematic diagram of the molecular dynamics simulation results of a small molecule (target protein binding molecule) obtained based on pocket2 in the examples provided in the present application; wherein A is the molecular dynamics simulation result of the small molecule and the mutant type p53-G245S; B is the molecular dynamics simulation result of the small molecule and p53-H168R-R249S. DETAILED DESCRIPTION
[0026] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in a sequence other than those illustrated or described herein. In this application, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. In the description of this application, unless otherwise stated, "plurality" refers to two or more than two.
[0028] In this article, unless otherwise specified, the term "target pocket structure" refers to a spatial structure constructed for a specific target protein, which contains potential sites or binding regions for the interaction of protein molecules with other molecules.
[0029] As used herein, unless otherwise specified, the term "target protein binding molecule" refers to a molecular entity that interacts with a specific target protein and binds to the target pocket. These binding molecules can be drug candidates, ligands, protein ligands, or other biomolecules that interact with specific binding sites on the target protein to produce biological effects, such as modulating signal transduction, gene expression, or protein function.
[0030] Unless otherwise specified, the term "crystal file" in this article refers to data files containing protein crystal structures obtained through techniques such as X-ray crystallography. These files record the three-dimensional structure of a protein under specific crystallization conditions, including detailed information such as atomic coordinates, lattice parameters, and crystallographic symmetry. Crystal files are typically saved in a standard file format, such as the PDB format.
[0031] Unless otherwise specified, the term "Ratcliff / Obershelp algorithm" in this document refers to an algorithm for calculating the similarity between two strings. Its main concept is to evaluate the similarity between the two strings by comparing the number and length of common subsequences. This algorithm is commonly used in tasks such as text matching and string alignment. In some examples of this application, this algorithm is used in the pocket identification process to compare the similarity of atoms surrounding a pocket, thereby identifying pockets with similar structures.
[0032] In this article, unless otherwise specified, the term "machine learning model" is a computational model or algorithm that can automatically perform tasks such as prediction, classification, identification or decision-making by learning and analyzing input data. The learning process of the model is based on statistical principles and data pattern recognition, and a training data set is used to adjust model parameters and optimize the model to improve its prediction or reasoning ability. Machine learning models can use various algorithms and techniques, such as deep learning, neural networks, support vector machines, decision trees, random forests, etc. These models can be trained and optimized through supervised learning, unsupervised learning or reinforcement learning. In some examples of this application, the inventors use a trained deep learning model to predict target protein binding molecules. In this application, the aforementioned trained deep learning model is selected from the Pocket2Mol molecule generation model.
[0033] In this document, unless otherwise specified, the term "device" or "unit" refers to a computer program or part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall device or unit that includes the functions of the device or unit.
[0034] Existing technologies have many shortcomings in solving problems related to the p53 protein. First, traditional experimental techniques are unable to fully capture the changes in the p53 protein under different structural states, which limits the understanding of its structural dynamics. Second, traditional structure prediction methods have uncertainties in the accurate identification and description of pockets, especially for complex proteins such as p53, which may lead to misidentification or inaccurate description of pocket structures. Moreover, existing research mainly focuses on the study of single p53 mutant subtypes, and lacks broader and deeper exploration, which limits the development of treatment strategies for different subtypes.
[0035] To this end, the present application proposes a method for constructing a target pocket structure of a target protein, a method for predicting a target protein binding molecule, a target protein binding molecule, a related device, a computer program product, a computing device, and a computer-readable storage medium. Each of these is described below:
[0036] Method for constructing target pocket structure of target protein
[0037] In one aspect of the present application, the present application proposes a method for constructing a target pocket structure of a target protein, the method comprising:
[0038] 1) obtaining a plurality of crystal files derived from the target protein or a mutant thereof, wherein at least one of the plurality of crystal files is derived from a mutant of the target protein;
[0039] In some examples of the present application, the aforementioned multiple crystal files of the target protein or its mutation are obtained from public databases. In other examples of the present application, the crystal files can also be obtained through self-testing.
[0040] In some examples of the present application, the target protein is P53 protein, and the multiple crystal files are respectively derived from mutants of the P53 protein.
[0041] In other examples of the present application, the target protein may also be selected from other types of proteins, such as p16 protein, p21 protein, p63 protein, and p73 protein.
[0042] In some examples of the present application, the crystal file is a PDB file obtained by single-chain decomposition of the crystal structure data of the P53 protein mutant. In some preferred examples of the present application, the aforementioned PDB file does not contain information on water molecules and impurity atoms.
[0043] 2) based on the multiple crystal files, generating candidate pocket structures corresponding to the crystal files respectively;
[0044] In some examples of the present application, the above-mentioned multiple crystal files are input into the fpocket program to generate candidate pocket structures corresponding to the crystal files.
[0045] 3) selecting one of the plurality of crystal files as a standard crystal file, and determining the candidate pocket structure corresponding to the standard crystal file as a reference pocket structure;
[0046] In some examples of the present application, the standard crystal file is determined based on the similarity of the candidate pocket structures. Specifically, the aforementioned determination method specifically includes: classifying the candidate pocket structures; and selecting the crystal file containing the most candidate pocket types as the standard crystal file.
[0047] In some examples of the present application, the aforementioned classification is performed according to mutation type.
[0048] 4) Based on the similarity between the reference pocket structure and the other candidate pocket structures, screening the other candidate pocket structures to obtain screened candidate pockets;
[0049] In some examples of the present application, the similarity is determined using the Ratcliff-Obershelp algorithm, which is used to quickly identify similar pockets so that subsequent prediction of P53 protein-binding molecules can be performed using only one of the pockets as a reference.
[0050] In some examples of the present application, the screening further includes at least one of the following: performing a first screening on the candidate pocket structures to retain pocket structures having a similarity with the reference pocket structure greater than a first threshold as first-screened candidate pocket structures; and performing a second screening on the screened candidate pocket structures to remove candidate pocket structures having a mutation type ratio less than a second threshold, thereby obtaining second-screened candidate pocket structures. The accuracy of the obtained candidate pocket structures is improved through screening.
[0051] In some examples of the present application, the first threshold is selected from 76% to 82%. In a preferred example of the present application, the first threshold is selected from 80%.
[0052] In some examples of the present application, the second threshold is selected from 66% to 72%. In a preferred example of the present application, the second threshold is selected from 70%.
[0053] In some examples of the present application, the screening further includes: performing a third screening on the candidate pocket structures that have undergone the second screening, retaining the candidate pocket structure with the smallest absolute value of the difference from the average volume of the candidate pocket, and obtaining the candidate pocket structure that has undergone the third screening. In the present application, the candidate pocket structure with the smallest absolute value of the difference from the average volume of the candidate pocket is the candidate pocket closest to the average volume of the candidate pocket. By obtaining the candidate pocket closest to the average volume of the candidate pocket for binding protein prediction, its representativeness, accuracy and reliability are higher, and more in line with actual biological significance.
[0054] 5) Summarizing the reference pocket structure and the screened candidate pockets to obtain the target pocket structure.
[0055] In some examples of the present application, the candidate pockets obtained by the above summary may be one or more, and may be from a reference pocket structure or from other candidate pocket structures.
[0056] In some examples of this application, new stabilizing pockets of the target protein can be screened based on the above method, so that after a single protein pocket is inactivated, small molecule compounds can still be designed based on other pockets to activate the protein.
[0057] Target protein binding molecule prediction method
[0058] In another aspect of the present application, the present application proposes a method for predicting target protein binding molecules, which method includes: constructing a target pocket structure of the target protein based on any of the methods in the above examples; obtaining a crystal file corresponding to the target pocket structure, and inputting the crystal file into a trained machine learning model to obtain the target protein binding molecule.
[0059] In some examples of the present application, the machine learning model is selected from a deep learning model. In a preferred example of the present application, the deep learning model is selected from a Pocket2Mol molecular generation model.
[0060] The Pocket2Mol molecular generation model is an E(3)-isovariant generation network consisting of two modules. It can not only capture the spatial and bonding relationships between atoms in the binding pocket, but also sample target protein binding molecules from a tractable distribution conditioned on the pocket representation without relying on Markov chain Monte Carlo (MC) methods. The target protein pocket binding molecule design method is as follows:
[0061] 1) A new deep geometric neural network is designed to accurately model the 3D structure of the pocket;
[0062] 2) A new sampling strategy was designed to achieve more efficient conditional 3D coordinate sampling;
[0063] 3) A sign of the model's ability to sample the chemical bonds between a pair of atoms.
[0064] The Pocket2Mol model utilizes vector-based neurons and geometric vector perceptrons to learn the chemical and geometric constraints imposed by protein pockets. It jointly predicts frontier atoms, atomic positions, atom types, and chemical bonds through shared atomic-level embeddings, and samples molecules in an autoregressive manner. Thanks to the vector-based neurons, the model can directly generate a tractable distribution of relative atomic coordinates relative to the focal atom, avoiding the use of traditional MC algorithms. Experimental results show that the molecules generated by Pocket2Mol not only have better affinity and chemical properties, but also have more realistic and accurate structures. Furthermore, Pocket2Mol is faster than previous MC-based autoregressive sampling algorithms.
[0065] In some examples of this application, the inventors used the trained Pocket2Mol molecule generation model to predict target protein binding molecules and efficiently screen potential target protein binding molecules based on the Pocket2Mol molecule generation model.
[0066] In other examples of the present application, the Pocket2Mol molecule generation model can be trained by itself to predict the target protein molecule or the trained Pocket2Mol molecule generation model can be fine-tuned to achieve the target protein molecule prediction.
[0067] The following example shows the training and prediction process of the Pocket2Mol molecule generation model:
[0068] First, a large amount of known target protein binding molecule data is prepared for model training, including the target protein's pocket structure and known binding molecule information. This data is then used to train the Pocket2Mol model, allowing it to learn the target protein's binding patterns and characteristics. The model is then tested for generalization and robustness.
[0069] In actual prediction, the trained Pocket2Mol model is used to generate potential binding molecule candidates based on the target protein pocket structure. Finally, the generated candidate molecules are further evaluated and screened, and suitable molecules are selected for wet experiment verification.
[0070] In some examples of the present application, the above prediction method further includes: screening the target protein binding molecules based on the drug similarity quantitative estimation value and the synthesizable ease value.
[0071] Among them, the aforementioned quantitative estimate of drug similarity (QED) represents a quantitative indicator used to measure the degree of similarity between two or more drugs. This indicator is calculated based on the structure, chemical properties or biological activity of drug molecules, and uses numerical values to represent the level of similarity between drugs. In some examples of this application, the inventors set the threshold of the quantitative estimate of drug similarity to 0.5, and molecules with a value above 0.5 have higher drug-like properties. In some examples of this application, the aforementioned threshold can be adaptively adjusted. If a molecule with higher drug-like properties is required, the threshold can be raised.
[0072] The aforementioned synthesizability value represents a quantitative indicator of the synthetic accessibility (SA) of a molecule. In some examples of this application, the inventors set the threshold of the synthesizability value to 5, and molecules with a value below 5 have higher synthesizability. In some examples of this application, this threshold can also be adaptively adjusted. For example, if a molecule with higher synthesizability is required, the threshold can be lowered.
[0073] In some examples of the present application, the aforementioned prediction method further includes: based on Software is used to further screen the target protein binding molecules. The software screens target protein binding molecules including MMGBSA dG Bind and dock score parameters.
[0074] In some examples of this application, the aforementioned prediction method further includes: further screening the target protein binding molecules based on molecular dynamics simulation methods. By numerically simulating the temporal evolution of the molecular system, a deeper understanding of the molecular structure and dynamic behavior is achieved, while simultaneously evaluating the thermodynamic stability of the molecular binding process to predict and optimize the strength and characteristics of the intermolecular interaction, thereby obtaining candidate binding molecules with good binding stability to the target protein.
[0075] In one example of the present application, the screening condition is that the binding free energy value is not greater than -20 kcal / mol. In other examples of the present application, the screening condition can be adaptively adjusted based on experimental requirements.
[0076] In some examples of the present application, the target protein binding molecules obtained by screening based on the above method can reactivate at least one type of mutated target protein, thereby restoring the function of the target protein.
[0077] Target protein binding molecules
[0078] In another aspect of the present application, a target protein binding molecule is provided, which is obtained by any of the above-mentioned prediction methods. The target protein binding molecule obtained based on the above-mentioned prediction method can effectively bind to and activate the target protein.
[0079] In some examples of the present application, the target protein binding molecule has a structure shown in formula (I),
[0080]
[0081] In some examples of the present application, the target protein binding molecule having the structure shown in formula (I) can activate at least one mutant type of P53 protein, thereby restoring the function of the p53 tumor suppressor gene.
[0082] Device for constructing the target pocket structure of target protein
[0083] In another aspect of the present application, the present application proposes a device for constructing a target pocket structure of a target protein, referring to Figure 1 The device includes: a crystal file acquisition unit S100, a candidate pocket structure generation unit S200, a reference pocket structure determination unit S300, a candidate pocket screening unit S400 and a target pocket acquisition unit S500.
[0084] Wherein, the crystal file acquisition unit S100 is connected to the candidate pocket structure generation unit S200;
[0085] Unit S100 is used to obtain a plurality of crystal files derived from the target protein or a mutation thereof, wherein at least one of the plurality of crystal files is derived from a mutant of the target protein;
[0086] Unit S200 is used to generate candidate pocket structures corresponding to the crystal files based on the multiple crystal files;
[0087] The candidate pocket structure generating unit S200 is connected to the reference pocket structure determining unit S300;
[0088] Unit S300 is configured to select one of the plurality of crystal files as a standard crystal file, and determine the candidate pocket structure corresponding to the standard crystal file as a reference pocket structure;
[0089] The reference pocket structure determination unit S300 is connected to the candidate pocket screening unit S400;
[0090] Unit S400 is used to screen the other candidate pocket structures based on the similarity between the reference pocket structure and the other candidate pocket structures to obtain screened candidate pockets;
[0091] The candidate pocket screening unit S400 is connected to the target pocket acquiring unit S500.
[0092] Unit S500 is used to aggregate the reference pocket structure and the screened candidate pockets to obtain the target pocket structure.
[0093] In some examples of this application, the aforementioned device can efficiently construct the target pocket structure of the target protein with high accuracy and operability. Based on the obtained accurate target pocket structure, accurate prediction of the binding molecules of the target protein can be achieved.
[0094] Target protein binding molecule prediction system
[0095] In another aspect of the present application, the present application proposes a target protein binding molecule prediction system, referring to Figure 2 The system includes: a target pocket construction device S01 and a target protein binding molecule prediction device S02. The target pocket structure construction device S01 and the target protein binding molecule prediction device S02 are connected.
[0096] The S01 device is used to construct the target pocket structure of the target protein based on the above-mentioned device for constructing the target pocket structure of the target protein; wherein, the S01 device includes: a crystal file acquisition unit S100, a candidate pocket structure generation unit S200, a reference pocket structure determination unit S300, a candidate pocket screening unit S400 and a target pocket acquisition unit S500.
[0097] S02 device is used to obtain the crystal file corresponding to the target pocket structure, and input the crystal file into the trained machine learning model to obtain the target protein binding molecule.
[0098] In some examples of the present application, the aforementioned system is capable of comprehensively predicting the binding molecules of the target protein with high accuracy, efficiency and ease of operation.
[0099] It should be noted that the features and technical effects described in this article for different aspects can be used as reference for each other and will not be repeated here.
[0100] Computer program product, computing device, and computer-readable storage medium
[0101] In another aspect of the present application, the present application proposes a computer program product, which includes computer instructions. When part or all of the computer instructions are run on a computer, the method for constructing a target pocket structure of a target protein or the target protein binding molecule prediction method as described above in the present application is executed.
[0102] In another aspect of the present application, the present application proposes a computing device, comprising: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the computer program to implement the method for constructing a target pocket structure of a target protein or the target protein binding molecule prediction method as described above in the present application.
[0103] On the other hand, the present application proposes a computer-readable storage medium, which includes computer instructions. When the instructions are executed by a computer, the computer implements the method for constructing a target pocket structure of a target protein or the target protein binding molecule prediction method as described above in the present application.
[0104] In other words, when implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0105] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0106] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0107] Modules described as separate components may or may not be physically separate, and components displayed as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected based on actual needs to achieve the purpose of the present embodiment. For example, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module.
[0108] The embodiments of the present invention will be described in more detail below, examples of which are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to be used to explain the present invention, but should not be understood as limiting the present invention.
[0109] Example 1: Screening of P53 protein stabilization pockets and binding molecules
[0110] This example uses P53 protein as an example to illustrate the advantages of the method for constructing the target pocket structure of the target protein and the method for predicting target protein binding molecules. Figure 3 , the specific steps are as follows:
[0111] 1. Crystal structure data preprocessing
[0112] Obtain protein crystal structure data related to p53 from the PDB public database, about 5,000 related crystal structures, including those obtained through Homo sapiens and Refinement Resolution After screening, 333 p53 protein crystal structures remained, of which 103 showed mutations, including several common "hotspot" mutations: G245S, R248W, R273H, R175H, Y220C, and R282W. These 103 crystal structures were divided into 31 categories based on mutation type. To facilitate pocket discovery, the 103 crystal structures were split into single chains, resulting in a total of 293 single-chain crystal files (pdb files). These files were named according to the pdb ID and the corresponding A, B, C, D, ... chains. Water molecules and impurity atoms were removed from each single-chain p53 crystal file.
[0113] 2. Protein pocket prediction and screening
[0114] 293 pre-processed monomeric p53 crystal files were used to identify potential binding pockets on the surface of the p53 protein crystal structure using fpocket. A total of 2871 potential binding pocket regions were identified, each with relevant information (*_vert.pqr and *_atm.pdb). The crystal file containing the most candidate pocket types was selected as the standard crystal file, resulting in the 1uol_A_out / pockets file. The amino acid sequence contained in it is as follows:
[0115] SVPSQKTYQGSYGFRLGFLHSGTAKSVTCTYSPALNKLFCQLAKTCPVQLWVDS
[0116] TPPPGTRVRAMAIYKQSQHMTEVVRRCPHHERCSDSDGLAPPQHLIRVEGNLRAEYL
[0117] DDRNTFRHSVVVPYEPPEVGSDCTTIHYNYMCYSSCMGGMNRRPILTIITLEDSSGNL
[0118] LGRDSFEVRVCACCPGRDRRTEEENLR (SEQ ID NO: 1).
[0119] Taking all pockets generated in the 1uol_A_out / pockets file as a reference, we screened out the overlap in the sixth column of all ATOM rows in the corresponding pocket file (i.e., the overlap in the names of atoms around the pocket). If the overlap in the atoms around the pocket is not less than 80%, the two pockets are considered to be the same. This part uses the Ratcliff / Obershelp similarity algorithm, which is derived as follows:
[0120] Suppose there are two strings s1 and s2, of length n and m respectively. Define M[i][j] as the length of the longest common subsequence of s1[0,1,...,i-1] and s2[0,1,...,j-1].
[0121] Initialize, set M[i][0] = M[0][j] = 0, for all i and j, such that:
[0122] Construct the matching matrix,
[0123] Extract the longest common subsequence, starting from the lower right corner of the matrix, and backtrack to get the longest common subsequence. Finally, M[n][m] contains the length of the longest common subsequence of the entire sequence.
[0124] Similarity ratio calculation:
[0125] Similarity ratio = (2×LCS) / (n+m); where LCS is the length of the longest common subsequence.
[0126] The Ratcliff / Obershelp similarity algorithm uses dynamic programming to extract the longest common subsequence to efficiently and comprehensively measure the similarity between two strings.
[0127] After the above similarity calculation, the overlap of pockets with similarity of not less than 80% with 293 monomeric p53 crystal structures, 103 crystal structures, and 31 mutation types was calculated, and the results shown in Table 1 were finally screened out;
[0128] Table 1
[0129]
[0130] Pockets with a screening pocket mutation ratio of no less than 70% were retained. As can be seen from Table 1, pocket1, pocket2, pocket3, and pocket4 in the 1uol_A_ou / pockets folder have a degree of overlap of no less than 80% with 218, 94, 118, and 117 of the 293 monomeric p53 crystal structures, respectively, accounting for 74.4%, 32.1%, 40.3%, and 39.9%, respectively. They also overlap with 98, 55, 74, and 75 of the 103 crystal structures, accounting for 95.1%, 53.4%, 71.8%, and 72.8%, respectively. They contain 29, 23, 23, and 27 of the 31 mutation types, accounting for 93.5%, 74.2%, 74.2%, and 87.1%, respectively.
[0131] 3. Generation of target protein binding molecules
[0132] The volume of the monomeric p53 crystal structure overlapping with pockets 1, 2, 3, and 4 was calculated, and the following can be seen from the volume distribution diagram (Figure 4):
[0133] ①The volumes of pocket 1 and pocket 3 are concentrated around 100, and the pocket volumes are relatively small. The small molecules (target protein binding molecules) generated based on the information of these two pockets are simple and easy to synthesize, so pocket 1 and pocket 3 are discarded.
[0134] ②The average volumes of pocket2 and pocket4 are 255.4392 and 263.7367, respectively. The pockets closest to the average volume of the pockets in the monomeric p53 crystal structure are selected: pocket3 in the 2wgx_B.pdb crystal structure and pocket3 in the 2x0u_A.pdb crystal structure, with corresponding pocket volumes of 255.3667 and 264.0614, respectively. The coordinates of the pocket centers are pocket2 (120.371, 94.546, -37.520) and pocket4 (96.442, 76.595, -13.407), respectively, to determine the center position of the pocket; the deep learning molecular generation model Pocket2Mol is used to generate molecules based on the pocket information of pocket2 and pocket4.
[0135] Using pocket information from pocket 2 and pocket 4 as reference, 695 and 895 small molecules were generated, respectively. The quantitative estimate of drug-likeness (QED) and synthetic accessibility (SA) values for each were calculated. SA expresses the ease of synthesis (synthetic accessibility) of drug-like molecules. This method can characterize the synthetic accessibility of a molecule as a score between 1 (easy to manufacture) and 10 (very difficult to manufacture). The estimation method is based on a combination of the contribution of molecular fragments and a molecular complexity penalty. The recommended range is: when QED>0.5 and SA<5, small molecules that meet both conditions have better drugability and synthetic accessibility.
[0136] 4. Virtual screening and molecular docking of target protein binding molecules
[0137] The Pocket2Mol molecule generation model will be used to generate small molecules in pocket2 and pocket4. Virtual screening and molecular docking were performed to obtain 15 and 17 small molecules, respectively. The MMGBSA dG Bind and dock score values of each small molecule were also obtained.
[0138] 5. Screening of target protein binding molecules based on structural stability
[0139] The molecules with good virtual screening and docking scores were screened based on the structural stability of the compounds, and 8 and 9 small molecules remained in pocket 2 and pocket 4, respectively, as candidate small molecule compounds.
[0140] 6. Screening of target protein binding molecules based on molecular dynamics simulation
[0141] After virtual screening of the small molecules generated by pocket 2, eight candidate small molecules were obtained. To conduct more precise screening, molecular dynamics simulations were performed using complexes of different mutants and small molecules as models, and a total of 22 p53 mutations were identified.
[0142] S1. For the R273H single mutation, molecular dynamics simulation screening revealed that four small molecules (25, 290, 663, and 692) bind stably to pocket 2 of the p53 protein with high affinity.
[0143] S2. For the R273C single mutation, through molecular dynamics simulation screening, we found that two small molecules (25, 663) bound more stably to pocket 2 of the p53 protein and had higher affinity.
[0144] S3. For the N239Y single mutation, three small molecules (334, 676, and 692) were retained as candidate drugs through molecular dynamics simulation screening, but their binding free energy was relatively weak and further molecular optimization was required.
[0145] S4. For the M133L-V203A-Y220C-N239Y-N268D multiple mutations, through molecular dynamics simulation screening, we found that two small molecules (677 and 692) bound more stably to pocket 2 of the p53 protein and had higher affinity, with binding free energies of -25.16 kcal / mol and -28.85 kcal / mol, respectively.
[0146] The molecular dynamics simulation results of the 22 mutants and the eight candidate small molecules were summarized to obtain Table 2;
[0147] Table 2
[0148]
[0149]
[0150] After virtual screening of the small molecules generated in pocket 4, nine candidate small molecules were obtained. To conduct a more precise screening, the R280K single mutation was selected as a model for molecular dynamics simulation. The simulation results were not good, with only one small molecule stably bound to the pocket, and the binding free energy was poor, at -15.05 kcal / mol. Figure 5 As shown in Figure 3 , the pocket is almost completely exposed to the solvent from a conformational perspective, making it unsuitable as a drug pocket. This exposed conformation may weaken the pocket's interactions with drug molecules, thereby reducing drug efficacy. Furthermore, the exposed pocket may increase the risk of interactions with harmful substances, leading to unpredictable side effects.
[0151] Molecular dynamics simulations of small molecules generated from pockets 2 and 4 with the screened p53 protein crystal structure revealed that pocket 4, from the perspective of candidate small molecules, is almost completely solvent-exposed, making it unsuitable as a druggable pocket, and was therefore discarded. However, in pocket 2, eight candidate small molecules screened by experts were docked and screened against 22 mutation types through molecular dynamics simulations. Small molecule 692 exhibited excellent binding among the large group of small molecules, showing broad binding to multiple mutation types. This finding opens new possibilities for p53 research, suggesting that small molecule 692 may be a potential drug for treating a variety of p53 mutation-related diseases. Furthermore, molecules 25, 279, 334, and 676 also showed good binding trends to multiple mutations. These small molecules also have great potential as potential drugs and are important targets for future research.
[0152] The following is an example of the small molecule structure formula (Formula I, molecule No. 334) obtained based on pocket2 and the molecular dynamics simulation results ( Figure 6 ).
[0153]
[0154] Example 2: Verification of small molecule activation of mutant protein activity
[0155] This example aims to verify that the small molecule obtained in Example 1 can be used to activate the activity of at least one mutant type protein (P53 protein).
[0156] Taking small molecule No. 334 (structural formula shown in formula (I)) as an example, the smiles format of small molecule No. 334 is:
[0157] Cc1c[nH]nc1Oc1cc(CN2CC(O)C(O)C2)cc2nc(O)ccc12, in order to verify whether small molecule No. 334 can reactivate the activity of mutant p53-G245S and mutant p53-H168R-R249S respectively, the specific experimental steps are as follows:
[0158] S1. Cell selection and culture: Cell lines containing mutant p53-G245S and p53-H168R-R249S were selected and cultured in DMEM medium containing 10% fetal bovine serum.
[0159] S2. Small molecule preparation: Based on the structure and solubility of the small molecule, a small molecule stock solution of a certain concentration is prepared using an organic solvent (such as DMSO). Note that the final DMSO concentration should be controlled below 0.1% to avoid toxic effects on cells.
[0160] S3. Experimental grouping: The cells were divided into a control group (no small molecule added), a DMSO group (addition of the same volume of DMSO) and an experimental group (addition of small molecule solutions of different concentrations), with at least three replicates for each group.
[0161] S4. Cell treatment: ① Seed the cells in a 96-well plate with 5000 cells per well. ② After the cells adhered, add the small molecule solution and continue culturing for 24, 48, or 72 hours.
[0162] S5. Detection of p53 activity: ① Use Western blot to detect the expression levels of p53 and its downstream target genes; ② Use quantitative PCR (qPCR) to detect the mRNA levels of p53 target genes.
[0163] S6. Cell Proliferation and Apoptosis Assays: ① Use a cell proliferation assay kit such as MTT, CCK-8, or Brdu to measure cell proliferation and calculate the cell proliferation inhibition rate. ② Use flow cytometry and fluorescence microscopy to observe apoptosis-related markers (such as Annexin V staining and caspase activity assay) to assess the ability of small molecules to induce cell apoptosis.
[0164] S7. Statistical Analysis: ① Analyze data using GraphPad Prism, SPSS, or other statistical software to compare differences between the experimental and control groups. ② Perform significance tests using t-tests or one-way analysis of variance (ANOVA), with P < 0.05 considered statistically significant.
[0165] The results showed that small molecule No. 334 can effectively reactivate the activity of mutant p53-G245S and p53-H168R-R249S, and can be used to treat diseases caused by these two mutant p53.
[0166] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0167] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention without departing from the principles and purpose of the present invention.
Claims
1. A method for constructing a target pocket structure of a target protein, characterized in that: The target pocket structure is used to predict the binding molecules of the target protein, and the method comprises: Obtaining a plurality of crystal files derived from mutations of the target protein, wherein at least one of the plurality of crystal files is derived from a mutant of the target protein; Based on the multiple crystal files, respectively generate candidate pocket structures corresponding to the crystal files; Selecting one of the multiple crystal files as a standard crystal file, and determining the candidate pocket structure corresponding to the standard crystal file as a reference pocket structure; wherein the standard crystal file is determined based on the similarity of the candidate pocket structures, including: classifying the candidate pocket structures; selecting the crystal file containing the most candidate pocket types as the standard crystal file; Based on the similarity between the reference pocket structure and the other candidate pocket structures, screening the other candidate pocket structures to obtain screened candidate pockets; and The reference pocket structure and the screened candidate pockets are combined to obtain the target pocket structure.
2. The method according to claim 1, characterized in that The target protein is P53 protein, and the multiple crystal files are respectively derived from mutants of the P53 protein.
3. The method according to claim 2, characterized in that The crystal file is a PDB file obtained by performing single-chain decomposition on the crystal structure data of the P53 protein mutant.
4. The method according to claim 1, wherein The screening further includes at least one of the following: Performing a first screening on the candidate pocket structures to retain pocket structures having a similarity with the reference pocket structure higher than a first threshold as candidate pocket structures that have passed the first screening; and The screened candidate pocket structures are subjected to a second screening to remove the candidate pocket structures whose mutation type ratio is lower than a second threshold, thereby obtaining candidate pocket structures that have passed the second screening.
5. The method according to claim 4, characterized in that The screening further includes: The candidate pocket structures that have passed the second screening are subjected to a third screening, and the candidate pocket structure with the smallest absolute value of difference from the average volume of the candidate pocket is retained to obtain the candidate pocket structure that has passed the third screening.
6. The method according to claim 1, characterized in that The similarity is determined by the Ratcliff-Obershelp algorithm.
7. A method for predicting target protein binding molecules, characterized in that: include: Constructing a target pocket structure of a target protein based on the method according to any one of claims 1 to 6; Obtain a crystal file corresponding to the target pocket structure, and input the crystal file into a trained machine learning model to obtain the target protein binding molecule.
8. The method according to claim 7, characterized in that The machine learning model is selected from a deep learning model.
9. The method according to claim 8, characterized in that The deep learning model is selected from the Pocket2Mol molecular generation model.
10. The method according to claim 7, characterized in that Further including: The target protein binding molecules are screened based on the quantitative estimation of drug-likeness and ease of synthesis.
11. The method according to claim 10, characterized in that Further including: Based on the molecular dynamics simulation method, the target protein binding molecules are further screened.
12. A target protein binding molecule, characterized in that The target protein binding molecule is predicted and obtained by the method according to any one of claims 7 to 11.
13. The target protein binding molecule according to claim 12, characterized in that The target protein binding molecule has a structure shown in formula (I), Formula (I).
14. A device for constructing a target pocket structure of a target protein, characterized in that: include: a crystal file acquisition unit, configured to acquire a plurality of crystal files derived from the target protein mutation, wherein at least one of the plurality of crystal files is derived from a mutant of the target protein; a candidate pocket structure generating unit, configured to generate candidate pocket structures corresponding to the crystal files based on the plurality of crystal files; a reference pocket structure determining unit, configured to select one of the plurality of crystal files as a standard crystal file, and determine the candidate pocket structure corresponding to the standard crystal file as the reference pocket structure; wherein the standard crystal file is determined based on the similarity of the candidate pocket structures, including: classifying the candidate pocket structures; and selecting the crystal file containing the most candidate pocket types as the standard crystal file; a candidate pocket screening unit, configured to screen the other candidate pocket structures based on the similarity between the reference pocket structure and the other candidate pocket structures to obtain screened candidate pockets; and The target pocket acquisition unit is used to aggregate the reference pocket structure and the screened candidate pockets to obtain the target pocket structure.
15. A target protein binding molecule prediction system, characterized in that: include: A target pocket construction device for constructing a target pocket structure of a target protein based on the device according to claim 14; The target protein binding molecule prediction device is used to obtain the crystal file corresponding to the target pocket structure and input the crystal file into a trained machine learning model to obtain the target protein binding molecule.
16. A computer program product, characterized in that include: Computer instructions; When part or all of the computer instructions are executed on a computer, the method for constructing a target pocket structure of a target protein according to any one of claims 1 to 6 or the method for predicting target protein binding molecules according to any one of claims 7 to 11 is executed.
17. A computing device, characterized in that include: processor and memory; The memory is used to store computer programs; The processor is used to execute the computer program to implement the method for constructing a target pocket structure of a target protein according to any one of claims 1 to 6 or the method for predicting target protein binding molecules according to any one of claims 7 to 11.
18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions or programs. When the computer instructions or programs are executed on a computer, the method for constructing a target pocket structure of a target protein according to any one of claims 1 to 6 or the target protein binding molecule prediction method according to any one of claims 7 to 11 is executed.
Citation Information
Patent Citations
Keratinase mutants with thermostability and catalytic activity improved
CN106636042A
Protein pocket alignment evaluation method and system
CN112116020A