Method and apparatus for drug combination prediction integrating molecular docking and expression profile data
By integrating molecular docking and expression profiling data and utilizing the Autodock and L1000 databases, binding energies and imprint matching scores were calculated, solving the problem of low accuracy in molecular docking prediction and achieving more efficient drug prediction.
Patent Information
- Application Number
- CN202211297021.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-21
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-10-21
AI Technical Summary
Existing molecular docking methods have low prediction accuracy and are difficult to reflect potential biological characteristics and downstream effects. Furthermore, existing combined strategies of molecular docking and expression profiling analysis have failed to achieve the highest prediction efficiency.
By integrating molecular docking and expression profiling data, molecular docking screening was performed using tools such as Autodock, and expression profiling data from the L1000 database was combined to calculate binding energy fraction and imprinting matching score. Z-score transformation and summation ranking were then used to construct a drug co-prediction method.
It significantly improved the accuracy of drug prediction, especially the hit rate for known drugs, by 2.8 to 1.9 times, and enhanced the ability to reflect intrinsic activity and downstream biological effects.
Smart Images

Figure CN116564434B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of drug informatics and bioinformatics, and in particular to a drug joint prediction method and device integrating molecular docking and expression profile data. BACKGROUND
[0002] Molecular docking is a computational method for predicting the interaction between two molecules by generating a binding model, which was first proposed by Kuntz et al. One of its important application scenarios is drug repositioning. In the research of anti-poison drugs, molecular docking is often used in the research of countermeasures for poisons with clear damage targets (such as sarin and VX, etc.). In recent years, more and more molecular docking-based drug repositioning has been reported. However, molecular docking is a method with low prediction accuracy. As reported by Lyu et al., in 100 million molecular docking tests, the prediction accuracy of the highest ranked molecule is only 22% to 26%. A common factor that is currently recognized as causing the low prediction accuracy of molecular docking is that the docking tool cannot accurately simulate the actual binding process between the receptor and the ligand, so many previous studies often focus on solving the problems of the algorithm itself and the data quality in the docking process, such as the bias of the scoring function and the inaccurate structure information, and the means adopted include using multiple scoring functions for joint scoring to reduce the bias of a single algorithm, or correcting or optimizing the inaccurate structure information through molecular dynamics simulation. However, improving the prediction accuracy of molecular docking is still a great challenge.
[0003] Data integration helps the system to understand the molecular and physiological processes after drug interference, so as to form a more reasonable drug discovery route. In recent years, many studies have integrated multiple data sets to better and more clearly understand the system they study, but they have focused on integrating multi-omics data (such as transcriptome, proteome, metabolome, etc.). Molecular docking is a structure-based method, which is difficult to reflect the potential biological characteristics (such as intrinsic activity or downstream effects). Omics data is the most abundant data source reflecting the actual system information of molecules, and is widely used in various studies. As one of the most mature omics, expression profile has rich drug interference data information, and is often used for drug repositioning research in the form of imprint matching. Improving the prediction accuracy of molecular docking technology is an urgent technical problem to be solved in the industry, so we consider whether we can use an integration method to combine molecular docking and expression profile analysis, so as to provide a simple and efficient joint drug prediction framework. SUMMARY
[0004] Therefore, the main purpose of the present application is to provide a drug joint prediction method and device integrating molecular docking and expression profile data, so as to at least partially solve the above technical problems.
[0005] In order to achieve the above-mentioned purpose, as a first aspect of the present application, a drug combination prediction method integrating molecular docking and expression profile data is provided, comprising the following steps:
[0006] Obtaining structural information of disease targets and FDA-approved drugs, and performing molecular docking targeting screening on the two to obtain binding energy scores;
[0007] Taking a gene set closely related to the disease as a reference signature, using a differentially expressed gene set induced by the disturbance of FDA-approved drugs on cells or tissues as a drug signature, calculating the correlation between the input reference signature (S i ) and the drug signature (S j ), and obtaining a signature matching score;
[0008] Z-score conversion is performed on the binding energy score and the signature matching score of the drug i to obtain a corrected docking score (SDi) and a corrected signature matching score (SSi), and then SDi and SSi are summed and sorted to obtain an integrated rank sequence Ri, wherein i=1, 2…n, n is the number of tested drugs; the rank sequence is the result of the drug combination prediction method based on molecular docking and expression profile data analysis.
[0009] As another aspect of the present application, a drug combination prediction device based on molecular docking and expression profile data analysis is also provided, comprising:
[0010] A data acquisition unit for receiving input data or querying a local / network database to obtain the required data;
[0011] A calculation processing unit for calculating and processing the data obtained by the data acquisition unit according to the drug combination prediction method as described above to obtain the result of the drug combination prediction method based on molecular docking and expression profile data analysis.
[0012] Based on the above technical solution, the drug combination prediction method and device integrating molecular docking and expression profile data of the present application have at least one of the following beneficial effects compared with the prior art:
[0013] The present application combines molecular docking and expression profile data analysis to establish a drug function prediction tool, which integrates intrinsic activity and downstream biological effect factors into the molecular docking strategy, thereby helping to obtain candidate drugs with target effects;
[0014] Seven targets (annotated by DPDR database) are selected in the application, including four nuclear receptors (NR3C1, ESR1, AR and PPARG) and three enzymes (CA4, PDE4B and HMGCR), by drawing the top N (N≤100) ROC-like curves, calculating the top N TPR and the area under the curve (AUC), to test the prediction performance of the three methods of DIORS, molecular docking and signature matching on known drugs; the results show that the AUC of StandardScaler (DIORS) is the highest, which is 32.3, which is significantly higher than the performance of any single method; compared with single molecular docking, the hit rate of the top 10, 30 and 100 positive drugs of StandardScaler is increased by 2.8 times, 2.5 times and 1.9 times respectively (see Figure 2 ). BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 It is a flowchart of the drug function prediction method based on molecular docking and expression profile data analysis of the application.
[0016] Figure 2 It is a ROC-like curve diagram for performance comparison of the drug function prediction method based on molecular docking and expression profile data analysis, molecular docking and signature matching. DETAILED DESCRIPTION
[0017] Some terms in the application have the following meanings:
[0018] Molecular docking is a theoretical simulation method for drug design through the characteristics of the receptor and the interaction mode between the receptor and the drug molecule, which mainly studies the interaction between molecules (such as ligand and receptor) and predicts the binding mode and affinity. In recent years, molecular docking method has become an important technology in the field of computer-aided drug research, but the prediction accuracy of molecular docking technology has yet to be improved.
[0019] Data integration is a data integration method that collects, sorts, cleans, and converts data from different data sources and loads them into a new data source to provide a unified data view for data consumers. Data integration has the following advantages: (1) transparency of underlying data structure: provides a unified interface for data access (consumer application), and the consumer application does not need to know where the data is saved, which type of access (XQuery, SQL) is supported by the source database, the physical structure of the data, the network protocol, etc.; (2) performance and scalability: data integration separates data integration and data access into two processes, so the data is ready for access; (3) providing a real single data view: the advantage of data integration is that the data is more real, accurate and reliable after data verification and cleaning; (4) good reusability: since there is actual physical storage, data can provide reusable data views for various applications without worrying about the availability of the underlying actual data source; (5) enhanced data management capability: the advantage of data integration is that data rules can be implemented during data loading and conversion to ensure data management.
[0020] Gene expression profile refers to constructing a non-biased cDNA library of cells or tissues in a specific state, collecting cDNA sequence fragments, qualitatively and quantitatively analyzing the mRNA population composition, and thus depicting the gene expression type and abundance information of the specific cells or tissues in the specific state. The data table prepared in this way is called gene expression profile.
[0021] Expression profile blotting is a kind of expression profile data reflecting the real gene expression changes of cells after drug interference.
[0022] Reference blotting is a comparison gene set used as input expression profile data in integrated analysis, which is usually a gene set recognized or known to be related to physiology, disease, drug action, etc. For example, when using integrated analysis to mine TLR4 inhibitors, the reference blotting can select a gene set of TLR4 pathway, or a gene set of TLR4 agonists / inhibitors disturbing cells.
[0023] Drug blotting is a gene set induced to change expression when a drug disturbs cells, tissues or organisms, which is used for comparison with reference blotting in integrated analysis to obtain drugs that can induce or reverse changes close to reference blotting. For example, when using integrated analysis to mine TLR4 inhibitors, the blotting of different drugs can be compared with the reference blotting of known TLR4 inhibitors to find drugs that can induce similar TLR4 inhibitor disturbance.
[0024] After the inventor consults literatures and experimental researches, it is found that data integration helps the system to understand the molecular and physiological processes after drug interference, so as to form a more reasonable drug discovery route. In recent years, many studies have integrated multiple data sets to better and more clearly understand the system they study, but they focus on integrating multi-omics data (such as transcriptome, proteome, metabolome, etc.), and do not involve the integration of structural data. Molecular docking technology is a structure-based method, which is difficult to reflect the potential biological characteristics (such as intrinsic activity or downstream effect). Both have certain shortcomings, but the inventor finds that using the idea of data integration to combine molecular docking and pharmacological activity data may help to obtain more reliable prediction, thereby making up for the shortcomings and deficiencies of the two and realizing a more accurate drug prediction method.
[0025] The inventor understands and finds that omics data is the most abundant data source reflecting the actual system information of molecules, which is widely used in various studies. As one of the most mature omics, expression profile has rich drug interference data information, and is often used for drug repositioning research in the form of imprint matching. As of the present date, a small number of studies have combined molecular docking and expression profile analysis for drug repositioning and mechanism prediction, mainly using the following two combined strategies: (1) molecular docking of the top-ranked drugs obtained by expression profile analysis, and selecting drugs with high docking scores as candidate drugs; (2) expression profile analysis of the top-ranked drugs obtained by molecular docking, and selecting drugs with the highest / lowest expression profile comparison scores as candidate drugs. These two strategies have also been fully applied in the recent research of COVID-19 treatment drugs. For example, Duarte et al. first used expression profile analysis and then used molecular docking to predict candidate compounds, and found that atorvastatin had a potential anti-COVID-19 effect. Raul et al. used expression profile analysis after molecular docking to screen drugs, and predicted anti-COVID-19 candidate drugs for three viral proteins (3CL viral protease, NSP15 endoribonuclease, and NSP12 RNA-dependent RNA polymerase). However, the existing combined strategies simply combine molecular docking and expression profile analysis in sequence (or in the form of preliminary screening and secondary screening), and it is not clear how to combine them to achieve the highest prediction efficiency, and there is also a lack of a simple and efficient combined framework.
[0026] The LINCS project has the largest cell expression profile database at present, i.e. L1000 database, which catalogues the reactions of a large number of human cell lines to compounds, covers 33,608 molecules, and a total of more than 100 million expression imprints, and is the best data resource as supplementary data of molecular docking. The present application combines the L1000 data and structural data on the basis of the StandardScaler comprehensive evaluation method, and finally constructs a method with higher prediction performance than molecular docking, and the intrinsic activity and downstream biological effect factors are included in the molecular docking strategy, thereby helping to obtain candidate drugs with target effects.
[0027] Specifically, the present application provides a drug joint prediction method based on molecular docking and expression profile data analysis, which can include intrinsic activity data in the molecular docking strategy, help to obtain candidate drugs with target effects, and thereby improve the drug prediction accuracy. The method specifically includes the following steps:
[0028] Target structure information of a target disease and FDA approved drug molecular structure information are obtained, and Autodock and other tools are used for molecular docking screening, to obtain the binding energy score (a parameter reflecting molecular docking) of each drug on the target;
[0029] The relevant gene set (reference imprint S i ) of the target disease and the differential gene set (drug imprint S j ) induced by the perturbation of the FDA approved drug molecules in cells or tissues are obtained, and the correlation between the input S i and S j is calculated to obtain the imprint matching score;
[0030] The binding energy score and the imprint matching score of the drug i are respectively subjected to Z-score conversion to obtain the corrected docking score (SDi) and the corrected imprint matching score (SSi), then SDi and SSi are summed and sorted to obtain the integrated rank sequence Ri, wherein i=1, 2…n, and n is the number of tested drugs; according to the rank sequence, the result of the drug joint prediction method based on molecular docking and expression profile data analysis of the present application can be obtained.
[0031] The binding energy score is directly obtained from the Autodock molecular docking software used, and the calculation formula is: binding energy = intermolecular energy + internal energy + torsional energy - unbonded extended energy, with the unit of kCal / mol. The internal energy is the difference in internal energy before and after ligand binding, not the absolute value of the internal energy.
[0032] The results can be used to measure the prediction performance of the method by checking whether the ranking of known effective drugs is at the top and calculating the area under the curve (AUC) of the class-ROC curve.
[0033] Autodock Vina can be used to target screen drugs in the intersection of target structure information and approved drug target structure information, with the binding pocket parameter set to the default value in this process.
[0034] The target structure information can come from self-test, literature or database, preferably from the PDB (Protein Data Bank) database or DRAR annotated disease target database.
[0035] The approved drug target and structure information can come from self-test, literature or database, preferably from the Drugbank database.
[0036] The association between the input reference fingerprint (S i ) and the drug fingerprint (S j ) is calculated by the bidirectional Jaccard score (Signed Jaccard score, SJ score) represented by the following formula.
[0037]
[0038]
[0039] S i is the significant differential gene of the reference fingerprint, wherein S i up is the up-regulated gene, S i dn is the down-regulated gene; S j is the significant differential gene of the drug, wherein S j up is the up-regulated gene, S j dnFor down-regulation genes; SJ value range in [-1, 1].
[0040] Wherein, the reference signatures of the tested target are mainly from Enrichr and CREEDS dataset.
[0041] Wherein, the expression profile data of approved drugs are from level 4 data in LINCS L1000 database, and the expression profile signatures of the compounds are obtained by CD algorithm.
[0042] Wherein, the differentially expressed genes can be obtained by Z-test algorithm according to threshold P < 0.01, for example.
[0043] In a preferred embodiment, as shown in Figure 1 the method comprises the following steps:
[0044] ①Obtain drugs containing both molecular structure information and expression profile information from FDA approved drugs on the market as background data for integrated analysis;
[0045] ②Molecular docking simulation is performed on the molecular structure of FDA approved drugs on the market and the target molecular structure of the target disease to obtain the affinity score (binding energy score) of each drug to the target, and Z-score conversion is performed;
[0046] ③Signature matching analysis is performed on the differentially expressed genes (drug signatures) of the expression profile of FDA approved drugs on the market and the related genes (reference signatures) of the target disease, the matching score (two-way Jaccard coefficient) of each drug to the reference signature is calculated, and Z-score conversion is performed;
[0047] ④Add the Z-scores of ② and ③, and convert them into rank sequences to obtain the final ranking of each drug.
[0048] More specifically, the method comprises the following steps:
[0049] Step 1: Virtual molecular docking screening
[0050] Target structure information is from DRAR database, and approved drug targets and structure information are from PDB (Protein Data Bank). Autodock Vina is used to perform target screening on drugs in the intersection of LINCS dataset and PDB dataset, and the binding pocket parameter is set to the default value in this process.
[0051] Step 2: Similarity comparison based on signature matching
[0052] Expression signatures are the expression profile data reflecting the real gene expression changes of cells after drug interference. The reference signatures of the tested targets are mainly from Enrichr and CREEDS datasets. The inventors take the signatures of representative marketed drugs with specific targets as drug signatures, and calculate the correlation between the input reference signatures (S i ) and drug signatures (S j ) by SJ scores.
[0053]
[0054]
[0055] S i are the significant differential genes of the reference signatures, where S i up are the up-regulated genes, and S i dn are the down-regulated genes. S j are the significant differential genes of the drug, where S j up are the up-regulated genes, and S j dn are the down-regulated genes. The SJ value ranges between [-1, 1], where the closer the SJ value is to 1, the higher the similarity of the differential genes of the reference signatures and the differential genes of the drug, and the closer the SJ value is to -1, the higher the opposite nature of the differential genes of the reference signatures and the differential genes of the drug, and SJ=0 represents no correlation between the two.
[0056] where the expression profile data of the approved drugs are from the level 4 data in the LINCS L1000 database, and the expression profile signatures of the compounds are obtained by CD algorithm. Z-test is used to obtain the differential expression genes with a threshold P<0.01.
[0057] Step 3: Integration of docking and expression signature matching scores
[0058] Summing and ranking the Z-score standardized indicators (Standard scaler): after the calculation of molecular docking and signature matching for all drugs is completed, the two scores of drug i are respectively converted into Z-score to obtain the corrected docking score (SDi) and the corrected signature matching score (SSi), and then SDi and SSi are summed and ranked to obtain the integrated rank sequence Ri, as shown in the following formula. Where i=1, 2…n, n is the number of tested drugs.
[0059] Ri=Rank(SDi+SSi); formula (3)
[0060] Step 4: Evaluation of the predictive performance of the DIORS method
[0061] The overall process uses the Scikit-learn compute class-Area Under Curve (AUC) of the ROC curve to measure the predictive performance of the method.
[0062] The application also provides a drug combination prediction device based on molecular docking and expression profile data analysis, comprising:
[0063] A data acquisition unit is configured to receive input data or query a local / network database to obtain required data.
[0064] A computing processing unit is configured to perform computing processing on the data obtained by the data acquisition unit according to the drug combination prediction method as described above, so as to obtain the result of the drug combination prediction method based on molecular docking and expression profile data analysis.
[0065] The data acquisition unit may be, for example, a local auxiliary input device such as a keyboard, a handwriting input board, a scanner, etc., or a network searching and collecting module including a network communication device, such as a network device connected to the Internet through dial-up, broadband, mobile base station, etc., and capable of performing network data acquisition and transmission through a corresponding network interface.
[0066] The computing processing unit may be, for example, in the form of a general-purpose computing device. It may include one processor or multiple processors working cooperatively. The application does not exclude distributed processing, i.e., the processors may be dispersed in different physical devices. The electronic device of the application is not limited to a single physical entity, but may also be the sum of multiple physical devices.
[0067] It should be understood that the computing processing unit described above is only an example of the application, and the computing processing unit of the application may also include elements or components not shown in the above examples. For example, some computing processing units also include a local memory for locally storing computing results, and a display module for outputting computing results.
[0068] The program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Python, Java, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider. Also, while a Python package is invoked in the above examples, the functions and libraries associated with the package can also be implemented in other programming languages.
[0069] From the above description of the embodiments, it is easy for those skilled in the art to understand that the present application can be implemented by hardware capable of executing specific computer programs, such as the device of the present application, and servers, clients, mobile phones, control units, processors, etc. contained in the device, and the present application can also be implemented by a server network containing the above-mentioned devices or components. The present application can also be implemented by other programmable computing devices that execute the method of the present application, such as single-chip computers, single-board computers, programmable logic controllers (PLCs), field programmable gate arrays (FPGAs), etc. executing software. It should be noted that the computer software executing the method of the present application is not limited to being executed by one or a specific hardware entity, but can also be implemented in a distributed manner by non-specific hardware. For computer software, the software product can be stored in a computer-readable storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), or can be distributed and stored on a network, as long as it can enable electronic devices to execute the method according to the present application.
[0070] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0071] Among them, the embodiments of the present application take the most targeted 7 targets of approved drugs as an example for prediction, but the present application is not limited thereto, which is only exemplary as the best embodiment.
[0072] The specific experimental steps of the embodiments of the present application are consistent with the above experimental methods.
[0073] Step 1: Virtual molecular docking screening
[0074] Target structure information is from DRAR database, approved drug targets and structure information is from PDB (Protein Data Bank). This part of the study first selects 7 kinds of targets with the most approved drugs, which are PPARG (PDB ID: 3B1M), MAOB (PDB ID: 1S3E), NR3C1 (PDB ID: 1NHZ), ESR1 (PDB ID: 1XPC), AR (PDB ID: 3B66), CA4 (PDB ID: 3FW3), PDE4B (PDB ID: 3FRG), ACE (PDB ID: 1O86), HMGCR (PDB ID: 2R4F), and then the AutodockVina is used to target screen 911 kinds of approved drugs contained in the LINCS dataset, and the binding pocket parameter is set to the default value in this process.
[0075] Step 2: Similarity alignment based on imprint matching
[0076] The reference imprint of the tested target mainly comes from Enrichr and CREEDS dataset. The inventors take the imprint of the representative marketed drug with a specific target as the drug imprint, and calculate the correlation between the input reference imprint (S i ) and the drug imprint (S j ) through SJ score.
[0077] Step 3: Integration of docking and expression profile imprint matching scores
[0078] Z-score standardization index summation ranking (Standard scaler): After the molecular docking and imprint matching calculation of all drugs are completed, the scores of drug i are respectively subjected to Z-score conversion to obtain the corrected docking score (SDi) and the corrected imprint matching score (SSi), then SDi and SSi are summed and ranked to obtain the integrated rank sequence Ri, as shown in the following formula (3). Wherein i = 1, 2…n, n is the number of tested drugs.
[0079] Ri = Rank (SDi + SSi); formula (3)
[0080] Step 4: Prediction performance evaluation of DIORS and molecular docking
[0081] Drug target annotation was referenced from Drug repurposing hub database. A known target drug was considered as true positive prediction if it was predicted correctly by the test target, otherwise it was considered as false positive prediction. The whole process used the roc_curve and auc functions of the Scikit-learn package in Python to calculate the area under the curve (AUC) of the class-ROC curve to measure the prediction performance of the method.
[0082] Experimental results:
[0083] The performance of each method was evaluated by detecting whether DIORS and two separate methods (molecular docking and signature matching) could accurately identify drugs targeting specific targets. Based on the number of drugs corresponding to each drug target, the present inventors selected seven common drug targets (annotated from the DPDR database), including four nuclear receptors (NR3C1, ESR1, AR, and PPARG) and three enzymes (CA4, PDE4B, and HMGCR), and the seven common drug targets each contained 5-20 drugs known to specifically bind to the target, which were used as true positive predictions. Next, by drawing the ROC curve of the top N (N≤100), the TPR of the top N and the area under the curve (AUC) were calculated to test the prediction performance of the above three methods (DIORS (StandardScaler), molecular docking, and signature matching) for known drugs. The AUC of StandardScaler (DIORS) was the highest, at 32.3 (as shown in Figure 2 ), which was significantly higher than the performance of any single method. Compared with molecular docking alone, the hit rate of the top 10, 30, and 100 positive drugs of StandardScaler increased by 2.8 times, 2.5 times, and 1.9 times, respectively (as shown in Figure 2 ).
[0084] The above specific embodiments further illustrate the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A drug combination prediction method based on molecular docking and expression profile data analysis, characterized in that, The method comprises the following steps: Obtaining structural information of disease targets and FDA-approved drugs, and performing molecular docking target screening on both to obtain binding energy scores; Using the FDA approved marketed drugs perturbation induced differential gene sets on cells or tissues as drug signatures, calculate the input reference signatures S i and drug signatures S j between the reference signatures and drug signatures, get signature matching scores; Z-score transformation of the binding energy scores and footprint match scores of drugs i to obtain corrected docking scores SDi and corrected footprint match scores SSi followed by SDi and SSi summing and ranking to obtain an integrated rank sequence Ri where i = 1,2... n, n is the number of drugs tested; According to the order sequence, the result of the drug combination prediction method based on molecular docking and expression profile data analysis is obtained. wherein the inputted reference footprint S i and the drug footprint S j The association between the inputted reference footprint and the drug footprint is calculated by the SJ score expressed by the following equation: ; formula (1, 2) wherein, S i is an up-regulated gene, S i up is an up-regulated gene, S i dn is a down-regulated gene; S j is a drug gene, S j up is an up-regulated gene, S j dn is a down-regulated gene; SJ value ranges between [-1, 1].
2. The drug combination prediction method of claim 1, wherein, Autodock Vina is used to perform target screening on drugs in the intersection of target structural information and approved drug target structural information, and the binding pocket parameter is set to the default value in this process.
3. The drug combination prediction method according to claim 1 or 2, characterized in that, The target structural information comes from the DRAR database.
4. The drug combination prediction method according to claim 1 or 2, characterized in that, The approved drug targets and structural information come from the Protein Data Bank database.
5. The drug combination prediction method of claim 1, wherein, The reference footprint of the tested target comes from the Enrichr and CREEDS datasets.
6. The drug combination prediction method of claim 1, wherein, The expression profile data of the approved drug comes from the level 4 data in the LINCS L1000 database, and the CD algorithm is used to obtain the expression profile footprint of the compound.
7. The drug combination prediction method of claim 1, wherein, Differentially expressed genes are obtained by using the Z-test algorithm according to the threshold P<0.
01.
8. The drug combination prediction method of claim 1, wherein, The prediction method further comprises calculating the area under the class-ROC curve AUC to measure the prediction performance of the drug combination prediction method by checking whether the ranking of the known effective drug is in the front.
9. The drug combination prediction method of claim 8, wherein, The drug target annotation refers to the Drug repurposing hub database, and the known target drug is predicted to be correctly tested as a true positive prediction.
10. The drug combination prediction method of claim 8, wherein, The entire process uses the roc_curve and auc functions of the Python Scikit-learn software package to calculate the area under the class-ROC curve AUC.
11. A drug combination prediction device based on molecular docking and expression profile data analysis, characterized by, It comprises: A data acquisition unit for receiving input data or querying local / network databases to obtain required data; A computing and processing unit for computing and processing the data obtained by the data acquisition unit according to the drug combination prediction method of any one of claims 1-10 to obtain the result of the drug combination prediction method based on molecular docking and expression profile data analysis.
Citation Information
Patent Citations
Drug prediction method, drug predication device and computer equipment
CN110310703A
System and method for selecting a set of candidate drug compounds
US20210287763A1