Drug target rapid screening system and method based on drug effect characteristic fingerprints
Through the drug target rapid screening system based on drug efficacy characteristic fingerprints, the problem of low target screening efficiency in existing technologies has been solved, efficient and accurate target screening and reverse docking have been achieved, and the computing efficiency and screening accuracy have been significantly improved.
Patent Information
- Application Number
- CN202510799606.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-26
AI Technical Summary
Existing target screening technologies rely on manually predefined pharmacophore models, which have incomplete coverage and inefficient matching algorithms and cannot capture all potential interactions of the target, resulting in low screening efficiency and limited accuracy.
A drug target rapid screening system based on pharmacodynamic characteristic fingerprints is adopted. Through the target structure database, drug-target interaction search engine module, small molecule preprocessing module and reverse molecular docking module, an inverted index algorithm is used to construct a pharmacodynamic characteristic fingerprint database to quickly screen out irrelevant targets and improve screening efficiency and accuracy.
It significantly improves the computational efficiency and accuracy of target screening, reduces invalid calculations, improves the efficiency and accuracy of reverse docking, and can comprehensively explore the potential interaction patterns between targets and drugs.
Smart Images

Figure CN120708762A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of drug research and development and computational chemistry, and specifically relates to a system and method for rapid screening of drug targets based on drug efficacy characteristic fingerprints. Background Art
[0002] Targets, as key molecules in drug action, serve as an important bridge between drugs and disease treatment, playing a central role in drug development. Researching drug targets is a critical step in new drug development. Traditional target screening methods rely heavily on complex pharmacological and molecular biology experiments, a process that requires continuous iteration and verification of hypotheses, leading to lengthy and costly drug development cycles.
[0003] Reverse molecular docking is a commonly used virtual screening technique used to predict target proteins that drug molecules may interact with. However, traditional reverse docking methods typically require the individual calculation of a large number of targets, resulting in a large computational load and low screening efficiency. In practice, however, due to drug specificity, most targets have no obvious interaction with drug molecules. However, due to the lack of an effective and rapid screening mechanism, these targets are still included in the docking calculation, wasting a large amount of computing resources, being time-consuming, and inefficient.
[0004] To improve the efficiency of reverse docking screening, researchers have attempted to narrow the range of candidate targets by building target databases and combining them with pre-screening methods. However, existing pre-screening methods are mostly based on simple similarity matching or chemical feature descriptions, which fail to fully consider the local interaction patterns between drug molecules and targets, resulting in limited accuracy and reliability of screening results.
[0005] In the prior art, it is a common target screening strategy to use pharmacophores for reverse target screening and combine it with molecular docking as a subsequent fine screening step. For example, after using the PharmMapper prediction platform to perform preliminary screening based on pharmacophore matching, it is combined with molecular docking technology for precise calculations. Although this combined approach can narrow the target range to a certain extent, it still faces significant efficiency challenges when processing massive amounts of data. Specifically, the pharmacophore matching algorithm of the PharmMapper prediction platform runs slowly, especially when dealing with large-scale protein databases, and its computational efficiency is significantly limited. In addition, the PharmMapper prediction platform is highly dependent on a pre-manufactured pharmacophore model database, which limits its potential in discovering new targets. More importantly, since these pharmacophores are screened based on specific criteria, they may not be able to fully cover the complex interaction patterns between drugs and targets, resulting in the omission of some potential targets.
[0006] Therefore, there is an urgent need for an efficient and accurate target screening method that can quickly screen out drug-irrelevant targets before reverse docking, thereby significantly improving screening efficiency and predicting potential drug targets. Summary of the Invention
[0007] In order to solve the problems of existing target screening technology that relies on artificially predefined pharmacophore models, has incomplete coverage, low matching algorithm efficiency, and cannot capture all potential interactions of targets, the present invention uses a drug target rapid screening system and method based on pharmacophoric feature fingerprints to identify all potential local interactions of targets. When predicting drug targets, it can quickly screen out irrelevant targets and narrow the screening scope of reverse docking, thereby improving screening efficiency and prediction accuracy.
[0008] The technical solutions of the present invention are as follows:
[0009] A rapid screening system for drug targets based on efficacy fingerprints, wherein the target is a target protein of the human body, animal or plant, comprising:
[0010] Target structure database: pre-processes the structure of the target, stores target information, and defines the target binding site;
[0011] Drug-target interaction search engine module: Generates a pharmacophore model based on the target binding site, and then converts the pharmacophore model into a pharmacophoric fingerprint; uses an inverted index algorithm to construct the pharmacophoric fingerprint into a drug-target interaction search engine database;
[0012] Small molecule preprocessing module: obtaining the chemical structure of the chemical small molecule to be queried, preprocessing the chemical structure of the chemical small molecule to obtain the preprocessed chemical small molecule, performing conformational search on the chemical small molecule to generate a three-dimensional conformational space, generating a new pharmacophore and a pharmacophoric characteristic fingerprint set based on each conformation and the pharmacophoric characteristic fingerprint of the chemical small molecule, inputting the new pharmacophoric characteristic fingerprint set into a drug-target interaction search engine for retrieval and screening, and obtaining a candidate target list;
[0013] Reverse molecular docking module: calls the molecular docking program to calculate the binding ability between candidate targets and chemical small molecules, scores and ranks them, and provides a list of the top targets;
[0014] Online data analysis module: displays the binding posture and interaction pattern of small chemical molecules and targets through histograms, descending tables and three-dimensional structure visualization, and displays the calculation results of the reverse molecular docking module.
[0015] Furthermore, the target information includes three-dimensional structure data information, related disease information and protein family type classification information;
[0016] The pharmacophore includes the types and spatial positions of hydrogen bond receptors or donors, hydrophobic interactions, aromatic ring interactions, and positive and negative charge interactions.
[0017] Furthermore, the generation of a pharmacophore model based on the target binding site is specifically: by analyzing the chemical properties of the target binding site, combining molecular probes, exploring all potential interactions between the drug and the target within the target binding site, and expressing them in the form of a pharmacophore model.
[0018] Furthermore, the method for converting the pharmacophore into a pharmacophoric characteristic fingerprint is specifically as follows: enumerating and combining pharmacophore features in groups of two or three to form pharmacophoric characteristic pairs, calculating the distance between the features and converting it into a fingerprint code to generate a pharmacophoric characteristic fingerprint.
[0019] Furthermore, the inverted index algorithm is used to construct the drug efficacy feature fingerprint to the drug-target interaction search engine database, specifically: all drug efficacy feature fingerprints are merged into a drug efficacy feature fingerprint set, the inverted index algorithm is used, each drug efficacy feature fingerprint in the drug efficacy feature fingerprint set is used as an index item, a unique identifier is assigned to each target as an index target, a mapping relationship between the index item and the index target is established and stored in the drug-target interaction search engine database, so as to realize the rapid positioning of associated targets through the drug efficacy feature fingerprint.
[0020] Furthermore, the pharmacophore and pharmacophoric fingerprint set are generated based on the pharmacophoric fingerprints of each conformation of the chemical small molecule, specifically:
[0021] A pharmacophore is generated for each conformation in the conformational space of the chemical small molecule, and a pharmacophore characteristic fingerprint is generated by using a method for converting the pharmacophore into a pharmacodynamic characteristic fingerprint, and the pharmacodynamic characteristic fingerprint set is obtained by summarizing the fingerprints.
[0022] Furthermore, the new pharmacodynamic characteristic fingerprint set is used to obtain a candidate target list, specifically as follows: according to the obtained pharmacodynamic characteristic fingerprint set, a search is performed in the drug-target interaction search engine module to obtain the target ID corresponding to each pharmacodynamic characteristic fingerprint, and a target ID summary table is obtained; the frequency of occurrence of each target ID is counted, that is, the number of times the target matches the pharmacodynamic characteristic fingerprint of the chemical small molecule, the more matches, the more potential interactions, and the higher the score, the target lists with high scores are sorted according to the scores and screened to obtain a candidate target list.
[0023] Furthermore, the binding ability of the candidate target with the chemical small molecule is calculated as follows:
[0024]
[0025] Among them, n is the number of matched efficacy feature fingerprints, wi is the weight, and ci is the number of times the fingerprint appears in the current target.
[0026] A rapid screening method for drug targets based on efficacy characteristic fingerprints is implemented by the above-mentioned screening system. The screening method is as follows:
[0027] Step 1: The user uploads the chemical structure data of the chemical small molecule to be searched, or manually constructs the chemical structure, or enters the SMILES sequence of the molecule;
[0028] Step 2: After the user submits the data, the small molecule preprocessing module is called to obtain a set of pharmacodynamic feature fingerprints of each conformation of the chemical small molecule;
[0029] Step 3: searching and screening in the drug-target interaction search engine module based on the obtained drug efficacy characteristic fingerprint set to obtain a candidate target list;
[0030] Step 4: Call the reverse docking module, calculate the binding ability between the target and the chemical small molecule in the candidate target list through the molecular docking program, and score and sort them to give a list of targets with the highest scores;
[0031] Step 5: Use the online data analysis module to perform data analysis and visual inspection on the reverse docking structure, and use the three-dimensional structure to display the binding posture and interaction mode of the chemical small molecule and the target. Combined with the score of the reverse docking module, the final target of the chemical small molecule is determined, and further verification is carried out in combination with experiments.
[0032] Compared with the prior art, the present invention has the following beneficial effects:
[0033] Compared with traditional reverse molecular docking technology, the present invention has the following beneficial effects: the present invention proposes a method and system for rapid screening of drug targets based on pharmacodynamic fingerprints, which significantly improves computational efficiency. Traditional reverse virtual screening usually adopts an inefficient strategy of traversing all targets. However, in practice, the number of targets that can bind well to small chemical molecules is limited, resulting in a large amount of computing resources being wasted on meaningless targets. The present invention uses a drug-target interaction search engine algorithm to quickly locate potential targets, screen out a large number of irrelevant targets, thereby reducing invalid calculations and significantly improving the efficiency and accuracy of reverse docking.
[0034] Compared with the existing pharmacophore-based target prediction technology (such as PharmMapper), the present invention has the following significant advantages: 1) By introducing molecular probe technology, it can fully explore all potential interaction patterns within the active site, avoiding the problem that the pharmacophore generated by the traditional method cannot reflect the potential interactions of the entire target; 2) Compared with the traditional pharmacophore matching method, the present invention adopts the pharmacophoric feature fingerprint coding technology, which greatly simplifies the calculation process. The traditional method requires a large number of overlapping comparisons to calculate the matching degree between the pharmacophore and the molecule, resulting in a large amount of calculation and low efficiency. The present invention only needs to compare the fingerprint coding consistency of the target and the molecule, and can screen out related targets by counting the frequency of occurrence of the same target, which significantly reduces the calculation complexity and improves the prediction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 The system architecture and method flow of the present invention;
[0036] Figure 2 Schematic diagram of pharmacophore fingerprint matching: By comparing the consistency of the pharmacophore fingerprints, it can be determined whether two groups of ternary pharmacophores can be matched;
[0037] Figure 3 The process of building a drug-target interaction engine based on drug efficacy fingerprints;
[0038] Figure 4 Schematic diagram of retrieving potential interaction targets of small molecules based on the inverted index table of pharmacodynamic fingerprints;
[0039] Figure 5 Schematic diagram of online data analysis module;
[0040] Figure 6 The system and method provided by the present invention have potential advantages over traditional traversal methods. DETAILED DESCRIPTION
[0041] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0042] See also Figure 1 , a drug target rapid screening system based on efficacy characteristic fingerprint, wherein the target is a target protein of the human body, animal or plant, including:
[0043] Target structure database: pre-processes the structure of the target, stores target information, and defines the target binding site;
[0044] Drug-target interaction search engine module: Generates a pharmacophore model based on the target binding site, and then converts the pharmacophore into a pharmacophoric fingerprint; uses an inverted index algorithm to construct the pharmacophoric fingerprint into a drug-target interaction search engine database;
[0045] Small molecule preprocessing module: obtaining the chemical structure of the chemical small molecule to be queried, preprocessing the chemical structure of the chemical small molecule to obtain the preprocessed chemical small molecule, performing conformational search on the chemical small molecule to generate a three-dimensional conformational space, generating a new pharmacophore and a pharmacophoric characteristic fingerprint set based on each conformation and the pharmacophoric characteristic fingerprint of the chemical small molecule, inputting the new pharmacophoric characteristic fingerprint set into a drug-target interaction search engine for retrieval and screening, and obtaining a candidate target list;
[0046] Reverse molecular docking module: calls the molecular docking program to calculate the binding ability between candidate targets and chemical small molecules, scores and ranks them, and provides a list of the top targets;
[0047] Online data analysis module: displays the binding posture and interaction pattern of small chemical molecules and targets through histograms, descending tables and three-dimensional structure visualization, and displays the calculation results of the reverse molecular docking module.
[0048] Among them, the target structure database also supports importing target structures from existing target libraries, such as batch importing from the sc-PDB target library. The database currently has nearly 18,000 target structures (including different binding sites) covering 4.8k different proteins. In addition, it can also be imported from existing target structure databases such as the DrugBank database, BindingDB database, and PDTD database; it can also download new protein structures that can be used as drug targets from the PDB database; for important targets that do not have crystal structures, AlphaFold2 is used to predict the three-dimensional structure of the target, and the molecular dynamics algorithm is used to appropriately optimize the structure before adding it to the target structure library;
[0049] In addition, the binding site is defined. The binding site is the location where the drug binds to the target. The accuracy of the definition has a significant impact on subsequent calculations. If the user provides binding site information, the user-defined binding site information is used. If the user does not define it, the target binding site prediction program, such as P2rank prediction program or FPOCKET prediction program, is automatically called to predict the location of the binding site, score the binding site, and retain the binding sites that meet the requirements.
[0050] The definition of binding sites can be based on existing drug-target complex information, or manually defined based on literature and experience, or predicted by algorithms to predict possible binding sites on the target; or molecular dynamics simulation sampling can be used to generate binding sites in different states to better obtain binding site information;
[0051] After the above targets have been preprocessed and the binding sites have been defined, the processed structures are saved, and relevant information such as target name, ID number, gene, protein type, source, and binding site information such as the XYZ position of the center point, binding site size, binding site score, etc. are saved in the database to facilitate subsequent data query and call.
[0052] In one embodiment of the present invention, the target information includes three-dimensional structure data information, related disease information and protein family type classification information;
[0053] The pharmacophore includes the types and spatial positions of hydrogen bond receptors or donors, hydrophobic interactions, aromatic ring interactions, and positive and negative charge interactions.
[0054] In one embodiment of the present invention, the preprocessing of the target structure database and the small molecule preprocessing module is mainly performed by calling the Rdkit software, including removing irrelevant atoms in the target structure, performing structural inspection, repairing missing defects, and performing hydrogenation and appropriate structural optimization.
[0055] In another embodiment of the present invention, the generation of a pharmacophore model based on the target binding site is specifically as follows: by analyzing the chemical properties of the target binding site, combining molecular probes, exploring all potential interactions between the drug and the target within the target binding site, and expressing them in the form of a pharmacophore model;
[0056] In the process of exploring the interaction types of the target binding site by the molecular probe, a score (energy score) can be performed according to the size of the binding energy, and interactions with poor binding force (pharmacodynamic characteristics) can be screened out;
[0057] During the construction of the structure-based pharmacophore model, appropriate cropping and clustering can be performed to remove redundant features and reduce the total number of pharmacophoric features, thereby facilitating subsequent analysis and use.
[0058] In one embodiment of the present invention, the method for converting the pharmacophore into a pharmacophoric characteristic fingerprint is specifically as follows: enumerating and combining pharmacophore features in groups of two or three to form pharmacophoric characteristic pairs, calculating the distance between the features and converting it into a fingerprint code to generate a pharmacophoric characteristic fingerprint.
[0059] Based on the pharmacophoric fingerprint of small chemical molecules, after generating the pharmacophore features, each three pharmacophoric features in the pharmacophore are combined to form a series of ternary feature pairs, and the distance between the two features (the side length of the triangle) is calculated. The vertices of the triangle represent three different interactions, and the side lengths represent the relative positions between each interaction.
[0060] Fingerprint coding is the fingerprint coding of pharmacophoric features. For example, the more common pharmacophore features can be represented by similar letters: hydrogen bond acceptor (A), hydrogen bond donor (D), hydrophobic (H), positive charge (P), negative charge (N) and aromatic ring (R). When encoding, the order of the three pharmacophoric features must also be determined. The pharmacophoric features are first sorted in alphabetical order. When the alphabetical order is the same, the pharmacophoric feature with the smaller variable length is preferred to ensure that the three pharmacophoric features have a unique order and obtain a unique code. Here, the three pharmacophoric features (vertices) after sorting are set to ABC. Then, the three edges are binned (equally spaced bins) according to the edge length in the order of AB, AC, and BC, so as to encode the ternary pharmacophoric features. At this time, the ternary pharmacophore is represented by a string or a digital ID.
[0061] like Figure 2 As shown, after encoding, it is only necessary to compare whether the encoding is consistent to determine whether the two ternary pharmacophores can be matched, without having to perform any rotation or translation operations on the pharmacophores. This is because ternary pharmacophores are definitely coplanar, and the comparison of triangles on a plane is relatively simple. When the characteristics of two ternary pharmacophores are consistent and the side lengths are not much different, they can match and overlap in space, otherwise they cannot overlap and match. In contrast, traditional pharmacophore matching algorithms require a large amount of spatial overlap and matching degree calculations, which are computationally intensive and complex algorithms and are not suitable for rapid search. The present invention can quickly achieve matching of pharmacophore substructures (ternary pharmacophores) through drug effect feature fingerprints.
[0062] In one embodiment of the present invention, the inverted index algorithm is used to construct the drug efficacy characteristic fingerprint to the drug-target interaction search engine database, specifically: all drug efficacy characteristic fingerprints are merged into a drug efficacy characteristic fingerprint set, the inverted index algorithm is used to use each drug efficacy characteristic fingerprint in the drug efficacy characteristic fingerprint set as an index item, a unique identifier is assigned to each target as an index target, a mapping relationship between the index item and the index target is established and stored in the drug-target interaction search engine database, so as to realize rapid positioning of associated targets through drug efficacy characteristic fingerprints;
[0063] Among them, the inverted index table is constructed, and an inverted index algorithm similar to that of a search engine is used. The fingerprint code is used as a "word", and the ID of the corresponding target is used as the "document number". That is, the pharmacodynamic fingerprint (word) is used as the index item, and the active pocket of the target is used as the index object to construct an inverted index table of "drug-target interaction". After the index is constructed, when querying the "rich interaction" target of a small chemical molecule, it is only necessary to generate the pharmacodynamic fingerprint of the small chemical molecule. Through the inverted index table, the relevant target can be quickly located, such as Figure 3 shown.
[0064] In addition, in a preferred embodiment of the present invention, the conformational search specifically comprises: generating a low-energy conformational set of the chemical small molecule by a conformational generation algorithm to construct a three-dimensional conformational space, wherein the three-dimensional conformational space covers all potential binding conformations of the chemical small molecule when binding to the target;
[0065] Specifically, the conformation generation algorithm provided by the Rdkit software can be used to generate small molecule conformations to generate a set of potentially active conformations. This technology is a conventional technique in the art.
[0066] In one embodiment of the present invention, the generation of a new pharmacophore and a pharmacophoric characteristic fingerprint set based on the pharmacophoric characteristic fingerprints of each conformation and chemical small molecule is specifically as follows:
[0067] For each conformation in the conformational space, a pharmacophore is generated using the chemical structure of the small chemical molecule. When the small chemical molecule and the target have the same pharmacophoric characteristic fingerprint, it is presumed that the two match in the local pharmacophore space and have potential interaction. At this time, the characteristics of the corresponding pharmacophore are extracted, and a pharmacophoric characteristic fingerprint is generated using the method of converting the pharmacophore into a pharmacophoric characteristic fingerprint;
[0068] Specifically, based on the active conformations generated by the Rdkit software, the pharmacophore characteristics of each activity of the molecule are calculated using Rdkit, resulting in a series of vector sets of pharmacophoric characteristics. A pharmacophore is a set of molecular features that play a key role in the interaction between a molecule and a biological target.
[0069] In one embodiment of the present invention, the method of obtaining a candidate target list using a new pharmacodynamic characteristic fingerprint set is as follows: searching the obtained pharmacodynamic characteristic fingerprint set in the drug-target interaction search engine module to obtain the target ID corresponding to each pharmacodynamic characteristic fingerprint, and summarizing the target ID summary table; counting the frequency of occurrence of each target ID, that is, the number of times the target matches the pharmacodynamic characteristic fingerprint of the chemical small molecule, the more matches, the more potential interactions, and the higher the score, sorting according to the score and screening several target lists with high scores to obtain a candidate target list.
[0070] The scoring and sorting function is primarily used for the following: Similar to a search engine, when there are many targets, each query may return a large number of results, so it is necessary to sort the results to prioritize targets with high matching scores. The more matching fingerprints there are, the more potential interactions there are, and the higher the score. The score of a target is calculated using formula (1), where n is the number of matching fingerprints, w is the weight, and c is the number of times the fingerprint appears in the current target.
[0071]
[0072] Preferably, the weight value will take into account the following factors: 1) TD-IDF: Considering that the frequency of some fingerprints may be very high and they will appear in many targets, such interactions should be removed or have their valuations lowered due to their lack of specificity. For better ranking, refer to the text search engine and use the TD-IDF algorithm to calculate the weight of the fingerprint to reduce the impact of high-frequency fingerprints. 2) Proximity of matching fingerprints: In the search engine, if the entire keyword phrase is fully matched, the ranking will be relatively high. Similarly, if several matching fingerprints are adjacent to each other in the receptor and the chemical small molecule pharmacophore, it means that the correlation is good and the weight value should be increased.
[0073] The calculation method of the binding ability between the candidate target and the chemical small molecule is:
[0074]
[0075] Among them, n is the number of matched efficacy feature fingerprints, w i is the weight, c i The number of times the fingerprint appears in the current target.
[0076] In a preferred embodiment of the present invention, the molecular docking program in the reverse molecular docking module can utilize Smina as its molecular docking engine. The Smina molecular docking engine is an extended version of AutoDock Vina, offering greater efficiency in energy minimization, supporting a wider range of input file formats, and enabling the development of scoring functions. Smina can be programmed to automatically perform molecular docking calculations between small molecules and active sites in a list of candidate targets, assessing their binding abilities and ranking them to generate a list of top-scoring targets for further analysis and verification.
[0077] The scoring function is one of the main bottlenecks of the current molecular docking algorithm. The existing scoring functions are not ideal. There is a poor correlation between the calculated score and the experimental binding affinity, or there is a target preference. Therefore, it is necessary to optimize the scoring function.
[0078] Consistency evaluation scoring is adopted, that is, multiple algorithms are used for scoring, and protein targets that rank relatively high in various scoring functions are selected during the final sorting.
[0079] Personalized scoring function: By collecting crystal structures of target-small molecule complexes and their experimental affinity values (Kd or Ki values), the interaction between small molecules and targets is analyzed, and new intermolecular forces (such as halogen bonds, sulfur bonds, phosphine bonds, and cation-cation interactions) are considered. Simultaneously, the calculation methods for descriptors such as surface physicochemical feature matching and small molecule entropy effects are optimized to generate corresponding descriptors as training sets. Using machine learning algorithms (such as support vector regression (SVR)), different training sets are used for fitting optimization for different types of targets / active sites, and personalized scoring functions are customized to improve the applicability and accuracy of molecular docking.
[0080] In one embodiment of the present invention, the online data analysis module deploys the relevant software to the website for sharing in order to facilitate researchers to use the reverse docking program developed by the present invention. PHP is used for website development, and Jquery and Bootstrap are used as front-end frameworks. In view of the particularity of drug design work, the system will also introduce an online molecular 3D display engine based on ngl to display the three-dimensional structure of biological macromolecules and small molecules, as well as molecular docking results, so as to build a visual graphical user interface that meets the requirements of online drug molecule design and has good interactive capabilities. And introduce the Cytoscape.js plug-in, combined with the GO database and KEGG signaling pathway, to perform visual analysis of network pharmacology and signaling pathways, which is convenient for multi-target data analysis of traditional Chinese medicine. Users only need to upload the small molecule structure of chemical small molecules to perform target recognition analysis. After the reverse docking is completed, users will be able to view the binding mode of chemical small molecules and receptors online through the 3D display window, and can analyze interactions (such as hydrogen bonds, hydrophobicity, halogen bonds, π interactions, etc.), and display active site pockets, such as Figure 5 shown.
[0081] The following is a method for constructing the drug-target interaction search engine database of the present invention:
[0082] Step 1: Users follow the instructions to upload target protein structure data and target information, including target name, protein type, gene information, related disease information, active site information, etc.
[0083] When the target structure does not exist, a new target can be constructed through methods such as AlphaFold or homology modeling, and relevant target information can be entered, including target name, protein type, gene information, related disease information, active site information, model accuracy information, etc.
[0084] Step 2: The target structure database automatically performs structural preprocessing on the target protein;
[0085] The target protein is subjected to structural preprocessing through the target structure database, including removing irrelevant atoms such as water, small molecules, etc., performing structural inspection, repairing missing defects, and performing hydrogenation and appropriate structural optimization to ensure the accuracy of the structure.
[0086] Step 3: When the target contains an active compound, the binding site of the target is preferably defined by the location of the active compound. After processing, the processed target structure data is saved, and various related information is entered and saved in the target structure database;
[0087] The binding site of the target can be manually defined based on literature and experience, or predicted by algorithms to predict possible binding sites on the target; or sampled using molecular dynamics simulations to generate binding sites in different states to better obtain information about the binding site.
[0088] Step 4: If Figure 3 As shown in the figure, after the target structure is saved, the drug-target interaction search engine module is called to analyze the chemical properties of the target binding site and combine molecular probes to explore all potential interactions between the drug and the target in the target binding site and express them in the form of a pharmacophore model;
[0089] Step 5: Figure 3 As shown, the drug-target interaction search engine module generates a pharmacodynamic feature fingerprint set, and the drug-target interaction search engine module stores the pharmacodynamic feature fingerprint set in the drug-target interaction search engine database; that is, all pharmacodynamic feature fingerprints are merged into a pharmacodynamic feature fingerprint set, and an inverted index algorithm is used to use each pharmacodynamic feature fingerprint in the pharmacodynamic feature fingerprint set as an index item, and a unique identifier is assigned to each target as an index target. A mapping relationship between the index item and the index target is established and stored in the drug-target interaction search engine database, so as to realize rapid positioning of associated targets through pharmacodynamic feature fingerprints.
[0090] During the generation of efficacy fingerprints, continuous variables such as intermolecular distances can be binned into equal-width bins (e.g., 0.5 angstrom bins). This discretization strategy not only improves the efficiency of fingerprint generation but also enhances the robustness of fingerprint matching by constructing tolerance intervals.
[0091] The screening method of the present invention is further described below with reference to a specific embodiment:
[0092] A rapid screening method for drug targets based on efficacy characteristic fingerprints, the screening method is as follows:
[0093] Step 1: If Figure 3 As shown, users upload the chemical structure data of the chemical small molecules to be retrieved, or manually construct the chemical structure, or input the SMILES sequence of the molecule;
[0094] Step 2: After the user submits the data, the small molecule preprocessing module is called to obtain a set of pharmacodynamic feature fingerprints of each conformation of the chemical small molecule;
[0095] The small molecule processing module pre-processes the chemical structure of the small molecule to be retrieved, including performing conformational search to obtain the possible conformational space of the molecule; based on the conformational space, it generates the pharmacophore of each conformation and extracts characteristic data;
[0096] Step 3: If Figure 4 As shown, the small molecule preprocessing module searches and screens the drug-target interaction search engine module based on the obtained drug efficacy feature fingerprint set to obtain a candidate target list;
[0097] Retrieval and screening are as mentioned above: obtain the target ID corresponding to each fingerprint and summarize it to obtain a target ID summary table. Count the frequency of each target ID, that is, the number of times the target matches the pharmacodynamic characteristic fingerprint of the chemical small molecule. The more matches, the more potential interactions there are, and the higher the score. After appropriate screening and sorting based on the degree of matching, a list of candidate targets can be obtained;
[0098] Step 4: Since the drug-target interaction search engine module only performs a rough screening of targets, in order to make the prediction results more accurate, it is necessary to call the reverse docking module for detailed calculation and screening;
[0099] At this time, the reverse docking module is called to calculate the binding ability between the targets in the candidate target list and the chemical small molecules through the molecular docking program, and then score and sort them to give a list of targets with the highest scores;
[0100] Step 5: Use the online data analysis module to perform data analysis and visual inspection on the reverse docking structure, and use the three-dimensional structure to intuitively display the binding posture and interaction mode of the chemical small molecule and the target. Combined with the score of the reverse docking module, the final target of the chemical small molecule is determined, and further verification is carried out in combination with experiments, such as Figure 5 shown.
[0101] like Figure 6 As shown, traditional reverse virtual screening typically uses an inefficient strategy of traversing all targets. However, due to the relatively small number of targets that can effectively bind to small chemical molecules, this strategy often leads to the inefficient consumption of large amounts of computing resources. In contrast, the present invention uses a drug-target interaction search engine algorithm to rapidly locate potential targets and screen out a large number of irrelevant targets, thereby significantly reducing wasted computations and greatly improving the efficiency and accuracy of reverse docking.
[0102] In summary, the present invention characterizes the interaction between drugs and targets based on pharmacodynamic characteristic fingerprints and uses them to construct an indexing method, which has the following advantages: 1) pharmacodynamic characteristics can well represent the potential interactions in target binding sites, including their interaction types, spatial positions, etc.; 2) small molecules can also easily generate pharmacophores and their pharmacodynamic characteristic fingerprints, thereby facilitating the calculation of the matching degree between small molecules and targets; 3) by converting the three-dimensional pharmacophore into a one-dimensional fingerprint code, the matching degree of drug-target related features can be conveniently compared, reducing the complexity of the calculation while having a certain tolerance and robustness; 4) by using fingerprint coding, the pharmacodynamic characteristic fingerprints can be treated as words, which facilitates the construction of the drug-target interaction search engine database using existing inverted indexing technology.
[0103] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention's description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A drug target rapid screening system based on efficacy characteristic fingerprint, wherein the target is a target protein of the human body, animal or plant, characterized in that: include: Target structure database: pre-processes the structure of the target, stores target information, and defines the target binding site; Drug-target interaction search engine module: Generates a pharmacophore model based on the target binding site and then converts the pharmacophore model into a pharmacodynamic fingerprint; An inverted index algorithm is used to construct a drug efficacy feature fingerprint to drug-target interaction search engine database; Small molecule preprocessing module: obtaining the chemical structure of the chemical small molecule to be queried, preprocessing the chemical structure of the chemical small molecule to be queried to obtain the preprocessed chemical small molecule, performing conformational search on the chemical small molecule to generate a three-dimensional conformational space, generating a new pharmacophore and a pharmacophoric characteristic fingerprint set based on each conformation and the pharmacophoric characteristic fingerprint of the chemical small molecule, inputting the new pharmacophoric characteristic fingerprint set into a drug-target interaction search engine for retrieval and screening to obtain a candidate target list; Reverse molecular docking module: calls the molecular docking program to calculate the binding ability between candidate targets and chemical small molecules, scores and ranks them, and provides a list of the top targets; Online data analysis module: displays the binding posture and interaction pattern of small chemical molecules and targets through histograms, descending tables and three-dimensional structure visualization, and displays the calculation results of the reverse molecular docking module.
2. A drug target rapid screening system based on drug efficacy characteristic fingerprint according to claim 1, characterized in that: The target information includes three-dimensional structure data information, related disease information and protein family type classification information; The pharmacophore includes the types and spatial positions of hydrogen bond receptors or donors, hydrophobic interactions, aromatic ring interactions, and positive and negative charge interactions.
3. A drug target rapid screening system based on drug efficacy characteristic fingerprint according to claim 1, characterized in that: The target-based binding site-based pharmacophore model is generated by analyzing the chemical properties of the target binding site, combining molecular probes, exploring all potential interactions between the drug and the target within the target binding site, and expressing them in the form of a pharmacophore model.
4. A drug target rapid screening system based on drug efficacy characteristic fingerprint according to claim 1, characterized in that: The method for converting the pharmacophore into a pharmacophoric characteristic fingerprint is specifically as follows: enumerating and combining pharmacophore features in groups of two or three to form pharmacophoric characteristic pairs, calculating the distance between the features and converting the distance into a fingerprint code to generate a pharmacophoric characteristic fingerprint.
5. The drug target rapid screening system based on drug efficacy characteristic fingerprint according to claim 1, characterized in that: The method of using an inverted index algorithm to construct a drug efficacy feature fingerprint to a drug-target interaction search engine database specifically comprises the following steps: merging all drug efficacy feature fingerprints into a drug efficacy feature fingerprint set, using an inverted index algorithm, taking each drug efficacy feature fingerprint in the drug efficacy feature fingerprint set as an index item, assigning a unique identifier to each target as an index target, establishing a mapping relationship between the index item and the index target and storing the result in the drug-target interaction search engine database, thereby realizing rapid positioning of associated targets through drug efficacy feature fingerprints.
6. A drug target rapid screening system based on drug efficacy characteristic fingerprint according to claim 4, characterized in that: The new pharmacophore and pharmacophoric characteristic fingerprint set are generated based on the pharmacophoric characteristic fingerprints of each conformation of the chemical small molecule, specifically: A pharmacophore is generated for each conformation in the conformational space of the chemical small molecule, and a pharmacophore characteristic fingerprint is generated by using a method for converting the pharmacophore into a pharmacodynamic characteristic fingerprint, and the pharmacodynamic characteristic fingerprint set is obtained by summarizing the fingerprints.
7. The drug target rapid screening system based on drug efficacy characteristic fingerprint according to claim 1, characterized in that: The method of obtaining a candidate target list by using the new pharmacodynamic characteristic fingerprint set is as follows: searching the obtained pharmacodynamic characteristic fingerprint set in the drug-target interaction search engine module to obtain the target ID corresponding to each pharmacodynamic characteristic fingerprint, and summarizing the target ID summary table; counting the frequency of occurrence of each target ID, that is, the number of times the target matches the pharmacodynamic characteristic fingerprint of the chemical small molecule, the more matches, the more potential interactions, and the higher the score, sorting according to the score and screening several target lists with high scores to obtain a candidate target list.
8. The drug target rapid screening system based on drug efficacy characteristic fingerprint according to claim 1, characterized in that: The calculation method of the binding ability between the candidate target and the chemical small molecule is: Among them, n is the number of matched efficacy feature fingerprints, w i is the weight, c i The number of times the fingerprint appears in the current target.
9. A method for rapid screening of drug targets based on drug efficacy fingerprints, characterized in that: The screening system according to claim 1 is implemented as follows: Step 1: The user uploads the chemical structure data of the chemical small molecule to be searched, or manually constructs the chemical structure, or enters the SMILES sequence of the molecule; Step 2: After the user submits the data, the small molecule preprocessing module is called to obtain a set of pharmacodynamic feature fingerprints of each conformation of the chemical small molecule; Step 3: searching and screening in the drug-target interaction search engine module based on the obtained drug efficacy characteristic fingerprint set to obtain a candidate target list; Step 4: Call the reverse docking module, calculate the binding ability between the target and the chemical small molecule in the candidate target list through the molecular docking program, and score and sort them to give a list of targets with the highest scores; Step 5: Use the online data analysis module to perform data analysis and visual inspection on the reverse docking structure, and use the three-dimensional structure to display the binding posture and interaction mode of the chemical small molecule and the target. Combined with the score of the reverse docking module, the final target of the chemical small molecule is determined, and further verification is carried out in combination with experiments.