Drug new function prediction system and drug safety evaluation system
By building a cross-omics analysis system, using causal inference and quasi-randomized analysis methods, we can solve the problems of high cost and high failure rates in drug research and development, and realize the prediction and safety evaluation of new drug functions, improving the accuracy of causal inference and the accuracy of target recognition.
Patent Information
- Application Number
- CN202510419130.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
There are problems of high costs, long cycles and high failure rates in existing drug research and development, including non-causal associations of targets, bottlenecks of cross-species transformation, lack of omics data integration, and out-of-control clinical trial costs.
A cross-omic analysis method and quasi-randomized analysis method based on causal inference were constructed. By obtaining the risk factors of the target disease, the association relationship between drugs and diseases, and the target genes, an analysis database was constructed to predict new drug functions and evaluate safety.
Improve the accuracy of causal inference, accurately locate potential causal targets, reduce resource waste, improve the effectiveness of clinical transformation and disease-specific target discovery rate, and reduce the failure rate of new drug development.
Smart Images

Figure CN120340672A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug research and development, and particularly to a drug new function prediction system and a drug safety evaluation system. Background Art
[0002] In recent years, the drug research and development market has been highly competitive, and the investment in drug research and development has been increasing. The problem of high input and low output of drugs has always been a difficult problem in the process of new drug research and development.
[0003] Predicting the drug targets of existing drugs is a good way to solve this problem. By locating the potential targets of drugs, the potential functions of drugs can be predicted.
[0004] Currently, there are the following technical defects in predicting the new functions of old drugs:
[0005] (1) Insufficient control of confounding factors by statistical methods: Traditional correlation analysis (such as GWAS) relies on observational data and lacks methods to effectively control confounding variables such as environmental factors and population heterogeneity, resulting in false positive associations. For example, the association between statins and reduced CRP levels may be mediated by the lipid metabolism pathway rather than a direct causal relationship.
[0006] (2) Lack of causal inference methodology: Existing technologies do not integrate a causal reasoning framework and cannot distinguish the temporal interference between drug action and disease progression. For example, when analyzing the cardiovascular protective effect of methotrexate, it is difficult for traditional methods to exclude the confounding effect of the remission of rheumatoid arthritis on the results.
[0007] (3) Flux limitation in experimental verification: In vitro experiments need to verify the functions of candidate targets one by one, and the cycle is as long as several months (such as the determination of single-target IC50 requires 2-3 weeks / compound), and large-scale multi-omics integration analysis cannot be achieved.
[0008] (4) Insufficient biological interpretability of computational models: Molecular docking relies on static crystal structures (such as PDB 7SKQ), ignoring the dynamic conformational changes of proteins, resulting in a significant difference between the prediction results and in vivo activity (such as the in vitro inhibition rate of remdesivir against SARS-CoV-2 RdRp does not match the clinical efficacy).
[0009] Currently, there are the following technical defects in the research and development of new drugs:
[0010] (1) Non-causal association of targets: Traditional methods cannot distinguish disease-related genes from true causal targets, resulting in a high clinical failure rate (such as the phase III failure rate of drugs based on the β-amyloid hypothesis in Alzheimer's disease research reaches 99.6%).
[0011] (2) Cross-species transformation bottleneck: There are essential differences between animal models and human pathological mechanisms (for example, the homology of mouse liver metabolic enzyme CYP450 is only 70%), resulting in ineffective human trials for preclinically effective drugs (for example, the neuroprotective agent NAXY-7 was effective in the ALS model but did not reach the endpoint in phase III).
[0012] (3) Lack of omics data integration: Existing platforms lack causal inference across the genome, transcriptome, and proteome (such as the isolated analysis of eQTL and pQTL data), and cannot identify the core pathways driving phenotypes (such as the priority determination of the IL-6 / JAK / STAT3 cascade).
[0013] (4) Out-of-control clinical trial costs: Traditional randomized controlled trials (RCTs) need to enroll thousands of patients (such as the FOURIER trial of the PCSK9 inhibitor with n = 27,564), and the cost per case exceeds 50,000, resulting in a total R & D cost of over 250 million. Summary of the Invention
[0014] To address the deficiencies of the existing technology, the present invention provides a new drug function prediction system and a drug safety evaluation system; by means of a cross-omics analysis method and a quasi-randomization analysis method based on causal inference, a high-throughput integrated data analysis system is constructed to solve the problems of high cost, long cycle, and high failure rate in the drug R & D process.
[0015] On the one hand, a new drug function prediction system is provided, including:
[0016] A first acquisition module, configured to: acquire risk factors of a target disease;
[0017] A second acquisition module, configured to: acquire the association relationship between a target drug and a target disease;
[0018] A third acquisition module, configured to: acquire the target gene of a target drug;
[0019] A database construction module, configured to: construct an analysis database according to the risk factors of the target disease, the association relationship between the target drug and the target disease, and the target gene of the target drug;
[0020] An analysis and prediction module, configured to: analyze the analysis database to obtain the new function prediction result of the target drug.
[0021] On the one hand, a drug safety evaluation system is provided, including:
[0022] A fourth acquisition module, configured to: acquire the target gene of a target disease;
[0023] A fifth acquisition module, configured to: acquire the intersection of significant target genes based on the target gene of the target disease;
[0024] A safety assessment module, which is configured to: based on the intersection of significant target genes, perform a safety assessment on drugs targeting the significant target genes.
[0025] The above technical solution has the following advantages or beneficial effects:
[0026] The prediction of new functions of old drugs in the present invention has the following several technical advantages:
[0027] 1. Improve the accuracy of causal inference: By selecting gene variants closely related to drug targets as instrumental variables, it is possible to effectively simulate quasi-randomized trials, strip the causal relationship from complex confounding factors, and thus significantly enhance the scientificity and accuracy of causal inference;
[0028] 2. Reduce the interference of reverse causal effects: Using gene variants as instrumental variables, since the influence of gene variants precedes the occurrence of diseases, this "temporal order" relationship helps to eliminate reverse causal effects and makes causal inference more reliable;
[0029] 3. Simulate the design of randomized controlled trials: By simulating the design of randomized controlled trials and using the random distribution of gene variants in the population, it simulates the randomization effect in traditional clinical trials, thereby reducing the potential selection bias in traditional non-experimental designs. This quasi-randomized design method makes the causal inference process closer to the actual effect of drug intervention and provides reliable data support for accurately identifying drug targets;
[0030] 4. Wide applicability: With the wide application of genome-wide association study data and the accumulation of large-scale genomic data, the platform has wide applicability. It can not only be applied in various disease fields, but also be combined with multi-omics data to verify causal relationships through biological information at different levels, thereby further enhancing the comprehensiveness and accuracy of causal inference.
[0031] The new drug development of the present invention has the following technical advantages
[0032] 1. Precisely locate potential causal targets: Using genetic variants as instrumental variables can precisely locate potential causal targets in the occurrence and progression of diseases. This method helps to distinguish truly therapeutic potential targets, thereby greatly increasing the success rate of drug development and avoiding wasting resources due to incorrect target selection;
[0033] 2. Reduce resource waste: The platform excludes non-causal associations through causal inference, which can avoid resource waste and development risks caused by selecting incorrect targets. Research shows that the targets identified by the platform have a higher success conversion rate in the drug development process. This method can significantly optimize resource allocation and improve the efficiency of drug research and development;
[0034] 3. Improve the effectiveness of clinical translation: The platform is not only used for early drug target identification, but also to verify the causality of existing targets, thereby supporting the rational design of clinical trials and the re-development of drugs. Research shows that the targets screened by the platform can more effectively reduce the failure rate of new drug development. Compared with traditional target discovery methods, the targets screened by the platform have greater potential for clinical translation and can provide strong causal evidence for actual treatment, thus improving the efficacy of new drugs in clinical applications;
[0035] 4. Significantly increase the discovery rate of disease-specific targets: The platform combines multi-omics data, such as genomic, transcriptomic, and proteomic data, and can analyze the causality of potential targets at multiple levels. The integration of different omics data makes the platform more prominent in discovering disease-specific targets. Through multi-omics cross-validation, the platform can identify more specific and therapeutically potential targets, providing support for the precision and personalization of drug research and development. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0037] Figure 1 It is a system function module diagram of Example 1. DETAILED DESCRIPTION OF THE INVENTION
[0038] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs.
[0039] Currently, the prediction of new functions for old drugs is mainly achieved from the following aspects:
[0040] (1) Target prediction based on correlation analysis: Identify gene loci related to disease phenotypes through Genome-Wide Association Study (GWAS), and combine drug-gene interaction databases to speculate potential targets (such as DrugBank, PharmGKB). For example, predict by statistically analyzing the similarity between the chemical structure of the drug and known targets.
[0041] (2) Screening and verification through in vitro experiments: Use high-throughput cell experiments or molecular biology techniques (such as gene knockout / overexpression) to verify the effect of the drug on candidate targets. Typical cases include the screening of compound libraries based on kinase activity detection.
[0042] (3) Retrospective analysis of clinical data: Mining drug-disease associations using Electronic Health Records (EHRs), such as statistically analyzing the differences in disease incidence rates among drug-using populations through ICD coding (e.g., methotrexate unexpectedly reduces the risk of cardiovascular events in patients with rheumatoid arthritis).
[0043] (4) Molecular docking simulation technology: Predicting the binding affinity between small molecule compounds and protein structures through computational chemistry methods (such as AutoDock Vina), for example, virtually docking existing drugs with novel target proteins for screening.
[0044] In addition, the current new drug R & D is mainly achieved from the following aspects:
[0045] (1) High-throughput compound screening: In vitro activity screening of a million-level small molecule library based on a robotic automation platform (such as Tecan Freedom EVO), typical applications include the determination of IC50 for kinase inhibitor libraries.
[0046] (2) Animal model validation: Using gene-edited animals (such as APOE4 transgenic mice) to evaluate the improvement effect of drugs on pathological phenotypes, for example, testing cognitive function through the Morris water maze.
[0047] (3) Multi-omics biomarker analysis: Integrating transcriptome (RNA-seq), proteome (LC-MS / MS), and metabolome (NMR) data to identify treatment response biomarkers, such as the association between PD-L1 expression levels and the efficacy of immune checkpoint inhibitors.
[0048] (4) Adaptive clinical trial design: Dynamically adjusting the dosing regimen using Bayesian statistical methods (such as the CRM model), typical cases include the combined design of tumor drug dose escalation trials (Phase I) and efficacy verification (Phase II).
[0049] Example 1 provides a drug new function prediction system;
[0050] As Figure 1 shown, the drug new function prediction system includes:
[0051] The first acquisition module, which is configured to: acquire risk factors of the target disease;
[0052] The second acquisition module, which is configured to: acquire the association between the target drug and the target disease;
[0053] The third acquisition module, which is configured to: acquire the target gene of the target drug;
[0054] A database construction module, which is configured to: construct an analysis database according to the risk factors of a target disease, the association relationship between a target drug and the target disease, and the target gene of the target drug;
[0055] An analysis and prediction module, which is configured to: analyze the analysis database to obtain a prediction result of the new function of the target drug.
[0056] Further, the obtaining of the risk factors of the target disease includes:
[0057] (1-1): Based on the target disease, screen exposure factors and genetic instrumental variables;
[0058] The exposure factors include: biomarkers and physiological indicators such as hematological indicators (immunocyte characteristics such as lymphocyte count, neutrophil count, platelet count, etc.), biochemical parameters (rheumatoid factor level, C-reactive protein, lipid profile, etc.), metabolomics data (specific metabolites in serum or urine, etc.), and disease classification and clinical phenotypes such as autoimmune diseases (psoriasis, rheumatoid arthritis, hypothyroidism, etc.), cardiovascular diseases (deep vein thrombosis, hypertension, coronary heart disease, etc.), metabolic diseases (diabetes, obesity-related phenotypes, etc.), and lifestyle and environmental factors, etc.
[0059] The genetic instrumental variables include: single nucleotide polymorphism (SNP).
[0060] The screening of exposure factors and genetic instrumental variables based on the target disease is carried out according to the following principles:
[0061] Select genetic instrumental variables with a significance threshold P < 5e -8 and a linkage disequilibrium clumping (LD Clumping) threshold r 2 < 0.001;
[0062] (1-2): For the target disease, exposure factors and genetic instrumental variables, conduct an analysis using Mendelian randomization to obtain a Mendelian randomization analysis result;
[0063] When only one genetic instrumental variable is screened out for the exposure factor, the Wald ratio algorithm is used to estimate the causal effect; when two or more genetic instrumental variables are screened out for the exposure factor, the IVW, MR-Egger regression, and weighted median algorithms are used to estimate the causal effect, and finally the results with the same effect direction for each algorithm are retained;
[0064] (1-3): Perform multiple testing correction on the results of Mendelian randomization analysis. If more than one significant exposure factor (e.g., n) is identified through the analysis, Bonferroni correction is adopted, i.e., the genome-wide significance threshold is adjusted to P < 0.05 / n.
[0065] (1-4): According to the corrected results, output the result charts to obtain the risk factors of the target disease. When the P-value corresponding to the causal effect estimate of an exposure is less than the genome-wide significance threshold after Bonferroni correction, it is considered that the exposure is a risk factor for the target disease.
[0066] Exemplarily, based on the determined target outcome, exposure factors and genetic instruments are further screened. The phenotype data downloaded by the system includes a large amount of GWAS data such as physiological and biochemical indicators, anthropometrics, psychological and behavioral characteristics, and lifestyle factors from databases such as UK Biobank, IEU openGWAS, and FinnGen. The system selects the eligible instrumental variables for each exposure factor as instrumental variables according to the selected database and major exposure factor categories, with a significance threshold of P < 5e -8 (or other set thresholds) and an LD Clumping threshold of r 2 < 0.001, 10Mb distance.
[0067] Exemplarily, based on the selected exposure and outcome, the system automatically conducts a phenome-wide association study with Mendelian randomization (PheWAS-MR) analysis. According to the selected outcome indicator, selected exposure factors, and instrumental variables, the R package Two-Sample MR is used for phenome-wide analysis. When only one instrumental variable is screened for the exposure phenotype, the Wald ratio method is used to estimate the causal effect; when there are two or more instrumental variables, methods such as IVW, MR-Egger regression, and weighted median are used to estimate the causal effect. The IVW result is the main result, and the results of other methods are for sensitivity analysis. Finally, the results with consistent effect directions for each method are retained.
[0068] Exemplarily, based on the obtained analysis results, multiple testing correction is performed. According to the number of selected exposure phenotypes, the system conducts multiple testing correction with a genome-wide significance P-value of 0.05 / (number of exposure phenotypes) after Bonferroni correction and retains the results that meet the corrected P-value.
[0069] Exemplarily, based on the calibration results, a result chart is plotted and output. According to the final results, the system automatically plots a forest plot and collates the result table, listing the significant risk factors and protective factors of diseases screened out by PheWAS-MR.
[0070] It should be understood that the system utilizes the downloaded massive genome-wide association study (GWAS) data. Through phenome-wide analysis, it screens out possible risk factors for the target disease outcome from a large number of phenotypes at the fastest speed, providing ideas for further drug target analysis.
[0071] Furthermore, the obtaining of the association relationship between the target drug and the target disease includes:
[0072] (2-1): Determine the core variable database, which includes: variables of the target drug and variables of the target disease;
[0073] The variables of the target drug include: drug dose, dosing frequency, and duration of medication;
[0074] The variables of the target disease include: disease diagnosis time, disease severity, and disease complications;
[0075] (2-2): Clean the data of the patients in the core variable database;
[0076] (2-3): Construct an observational cohort based on the cleaned core variable database; the observational cohort includes: the target drug cohort and the control cohort;
[0077] The target drug cohort includes patients who have used the target drug, and the starting time of the target drug cohort is the time point when the patient first used the target drug;
[0078] The control cohort includes patients who have not used the target drug, and the age and gender of the patients in the control cohort are the same as those of the patients in the target drug cohort;
[0079] (2-4): Conduct descriptive statistical analysis on the basic characteristics of the target drug cohort and the control cohort:
[0080] For continuous variables, if the data conforms to a normal distribution, statistical description is carried out using the mean and standard deviation;
[0081] If the data shows a skewed distribution, the median and quartiles are used for description;
[0082] For categorical variables, frequency and percentage are used for description;
[0083] (2-5): Baseline feature comparison:
[0084] Use statistical methods to evaluate the differences in confounding factors between the target drug cohort and the control cohort;
[0085] The statistical methods include: chi-square test, t-test or non-parametric test;
[0086] Specifically, if it is a continuous variable with a normal distribution, an independent samples t-test is used; if it is a continuous variable with a non-normal distribution, a rank sum test is used; if it is a categorical variable, a chi-square test is used;
[0087] Through statistical methods, the P-value of each baseline variable between the two groups is calculated. If P > 0.05, it indicates that there is no significant difference between the groups and the baseline characteristics are balanced.
[0088] Among them, the t-test is used to compare the mean differences of two groups of continuous variables with a normal distribution. The formula is as follows:
[0089]
[0090] and are the sample means of the two groups respectively; and are the sample variances of the two groups respectively; n1 and n2 are the sample sizes of the two groups respectively.
[0091] Among them, the rank sum test is used to compare the distribution differences of two groups of continuous variables with a non-normal distribution. The formula is as follows:
[0092]
[0093] n1 and n2 are the sample sizes of the two groups respectively; R1 is the rank sum of the first group, that is, the sum of the ranks of all data in the first group in the combined dataset.
[0094] Among them, the chi-square test is used to compare the distribution differences of two groups of categorical variables. The formula is as follows:
[0095]
[0096] O is the observed frequency, that is, the number of times a specific classification combination appears (such as 30 males in the intervention group); the expected frequency, that is, the number of times a specific classification combination should appear theoretically under the assumption of "no difference between the two groups".
[0097] If significant differences in certain confounding factors (basic demographic variables, lifestyle and behavior habits, and clinically relevant variables) are found between the two groups, then all confounding factors that may affect the study outcome will be included as covariates in the regression model for adjustment in subsequent analyses, so as to achieve control of confounding factors.
[0098] (2 - 6): Based on the results of the baseline characteristics comparison, estimate the effect of the target drug on the target disease;
[0099] When the outcome indicator is a binary variable, use a logistic regression model to analyze the effect of the target drug on the probability of the outcome occurring;
[0100] When the outcome indicator is survival data and the proportional hazards assumption is satisfied, use the Cox proportional hazards regression model to analyze the effect of the target drug on the risk of disease occurrence;
[0101] If the outcome indicator is a continuous variable, use a linear regression model to analyze the effect of the target drug on the outcome indicator;
[0102] (2 - 7): Based on the results of the effect estimation of the target drug on the target disease, conduct a sensitivity analysis to obtain the association between the target drug and the target disease.
[0103] Specifically, when the sample size is sufficient and there may be multicollinearity among variables, stepwise regression analysis is adopted. The stepwise regression analysis optimizes the model structure by gradually introducing or removing variables and is applicable to situations where the relative contributions of each variable need to be explored; while when there are large differences in baseline characteristics or many confounding factors between the two groups of patients, propensity score matching is adopted to balance the confounding variables by matching patients with similar baseline characteristics, so that the target drug cohort and the control cohort are highly comparable.
[0104] Re - estimate the drug effect using the above - mentioned method, compare the effect indicators under each model, and test the consistency of the results. If the results are consistent, it indicates that the research data and model settings are relatively robust, and the platform will report the consistent results as the main research findings; when the results are inconsistent, different adjustment strategies (such as stratified matching methods) need to be adopted to obtain more robust effect estimates.
[0105] Exemplarily, (2-1): Determine the variables of the target drug and the variables of the target disease, including: The system accurately locates the variables closely related to the research question in the UK Biobank database and the Cheeloo LEAD database. For the target drug, it covers the usage of the drug, such as the dosage, frequency, and duration of medication; for the disease outcome, it includes the diagnosis time of the disease, the severity of the condition, related complications, etc. At the same time, the system also collects possible confounding factors, such as the patient's age, gender, race, lifestyle (smoking, drinking, etc.), comorbidities, etc., for subsequent adjustment analysis. Through this step, the system clarifies the core variables required for the research, laying a data foundation for formulating inclusion and exclusion criteria.
[0106] Exemplarily, the (2-2): Clean the data of the patients in the core variable database, including:
[0107] Retain patients who are alive and currently have complete registration records in the UK Biobank or Cheeloo LEAD database.
[0108] Retain patients who have a usage record of the target drug and the medication time meets the time range set by the research.
[0109] Retain patients with complete disease outcome data to accurately evaluate the impact of the drug on the disease.
[0110] Retain the basic information (age, gender, etc.) of the patients and the data of important confounding factors complete for subsequent statistical analysis and adjustment.
[0111] Delete patients who have used the target drug before the start of the research to ensure that the causal effect of the drug can be accurately evaluated.
[0112] Delete patients with severe data loss due to other reasons during the research period to ensure the reliability of the research results.
[0113] Delete patients with special diseases or conditions, such as severe hepatic and renal insufficiency, malignant tumors, etc., to avoid these factors interfering with the correlation analysis between the drug and the disease outcome.
[0114] By formulating clear retention and deletion criteria, it provides a clear screening basis for constructing an observational cohort, ensuring the accuracy and representativeness of the research subjects.
[0115] Exemplarily, (2-3): Construct an observational cohort based on the cleaned core variable database; the observational cohort includes: the target drug cohort and the control cohort.
[0116] Target drug cohort: Based on the established inclusion and exclusion criteria, the system screened eligible patients who had used the target drug from the UK Biobank and Cheeloo LEAD databases. The time point of the first use of the target drug was taken as the starting time of the target drug cohort to construct the target drug cohort.
[0117] Control cohort: The system selected patients who had not used the target drug and met the inclusion criteria as controls, and matched them according to the basic characteristics of the patients (such as age, gender, etc.) to improve the comparability between the two groups, and finally completed the construction of the control cohort. The starting time of the control cohort can be set as a time point close to that of the target drug cohort, or determined according to the specific research design.
[0118] By constructing the target drug cohort and the control cohort, the system provides a structured data framework for subsequent analysis steps, ensuring the scientificity and rigor of the research design.
[0119] Exemplarily, (2-4): Conduct descriptive statistical analysis on the basic characteristics of the target drug cohort and the control cohort: Based on the constructed observational cohort, the system conducts descriptive statistical analysis on the basic characteristics of the target drug cohort and the control cohort, including the distribution of confounding factors such as age, gender, race, comorbidities, etc., and the use of the target drug (such as drug dosage, medication duration, etc.). Specifically, for continuous variables (such as age, drug dosage, etc.), if the data conforms to a normal distribution, the system will use the mean ± standard deviation for statistical description; if the data is skewed, the median and quartiles will be used for description. For categorical variables (such as gender, race, etc.), the system uses frequency and percentage for description. This analysis not only provides a preliminary data overview for the research, but also provides an important reference basis for baseline characteristic comparison and effect estimation.
[0120] Exemplarily, (2-5): Baseline characteristic comparison, including: Before conducting the effect analysis, the system compares the characteristics of the target drug cohort and the control cohort at baseline to ensure the comparability of the two groups at the beginning of the analysis. Methods such as chi-square test, t-test or non-parametric test are used to evaluate the differences in confounding factors between the two groups. If significant differences in certain confounding factors are found between the two groups, adjustments need to be made in subsequent analyses to achieve control of the confounding factors. Through this step, the system ensures the balance of the research data and provides reliable baseline data support for effect estimation.
[0121] Exemplarily, (2-6): Estimate the effect of the target drug on the target disease based on the results of the baseline characteristic comparison, including: Based on the results of the baseline characteristic comparison, the system selects an appropriate statistical method according to the research design and data characteristics to estimate the effect of the target drug on the disease outcome. The statistical methods include:
[0122] Logistic regression model: When the outcome indicator is a binary variable, such as whether a disease occurs or whether a complication appears, etc., the Logistic regression model is used to analyze the impact of the target drug on the probability of the outcome occurring, and to estimate the Odds Ratio (OR) and its 95% confidence interval.
[0123] Cox proportional hazards regression model: When the outcome indicator is survival data and the proportional hazards assumption is satisfied, the Cox proportional hazards regression model is used to analyze the impact of the target drug on the risk of disease occurrence, and to estimate the Hazard Ratio (HR) and its 95% confidence interval.
[0124] Linear regression model: If the outcome indicator is a continuous variable, such as the change in the level of a disease-related biomarker, a linear regression model is used to analyze the impact of the target drug on the outcome indicator, and to estimate the regression coefficient and its 95% confidence interval.
[0125] Through effect estimation, the system obtained the specific impact of the target drug on the disease outcome, provided preliminary results for subsequent sensitivity analysis, and laid a foundation for the formation of research conclusions.
[0126] Exemplarily, (2-7): Based on the effect estimation results of the target drug on the target disease, including: To evaluate the robustness of the research results, the system conducts a series of sensitivity analyses based on the effect estimation results. Try different confounding factor adjustment methods, such as stepwise regression, propensity score matching, etc., to ensure the consistency of results under the same method; exclude certain patient groups that may have a greater impact on the results, such as elderly patients, patients with multiple severe diseases, etc. If the results change significantly, the system will systematically evaluate the possible sources of bias and quantify the degree of their impact. This step not only verifies the reliability of the research results but also further enhances the robustness of the research conclusions, providing strong support for the final conclusion.
[0127] It should be understood that based on the analysis results, the system integrates real-world data, conducts an observational cohort study, explores the potential association between the drug and the disease outcome, and provides strong evidence for the determination of drug targets and subsequent research.
[0128] Furthermore, the obtaining of the target gene of the target drug includes:
[0129] (3-1): Obtain the potential drug target genes of the risk factors of the target disease;
[0130] (3-2): According to the potential target genes, retrieve the transcriptional data and protein data, obtain the corresponding transcriptional level and protein expression level of the potential target genes, and use the transcriptional level and protein expression level as exposure data;
[0131] (3-3): According to the set screening rules, SNPs that are strongly correlated with the target gene are screened from the exposure data, and the screened SNPs are used as instrumental variables;
[0132] The set screening rules include:
[0133] Rule 1: Set the significance threshold to P < 5e -8 ;
[0134] Rule 2: There is no obvious linkage disequilibrium (non-random association between alleles at different loci, usually expressed as the correlation coefficient r) at a given genetic distance (the degree of genetic difference between different populations or species, measured by a certain value, i.e., parameter kb) 2 Set the parameters to kb = 10000 and r 2 =0.001, which means removing the most significant SNP within 10000 kb. 2 SNPs greater than 0.001;
[0135] Rule 3: Minor Allele Frequency (MAF)>1%;
[0136] Rule 4: Use F statistics to eliminate weak instrumental variables, requiring Among them, R 2 It reflects the proportion of the variance of the exposure variable that can be explained by SNPs, N is the sample size, and k is the number of SNPs;
[0137] The relationship between rule 1, rule 2, rule 3 and rule 4 is "and"; all four rules need to be satisfied at the same time;
[0138] (3-4): Conduct statistical analysis on the obtained instrumental variables:
[0139] When there is only one instrumental variable, the Wald ratio method is used to estimate the causal effect; when there are two or more instrumental variables, the IVW method is used to estimate the causal effect, and multiple testing correction is performed, and the significant P value is 0.05; the exposure-outcome pair corresponding to P<0.05 is regarded as a positive result;
[0140] (3-5): Perform sensitivity analysis on the positive results obtained from statistical analysis, and use the genes that pass the sensitivity analysis as target genes for drugs;
[0141] Specifically, the maximum likelihood method, weighted mode, weighted median, and MR-Egger regression were used to test the robustness and consistency of the test results; Cochran's Q test was used to evaluate heterogeneity; the intercept term of MR-Egger regression was used to test for horizontal pleiotropy; MR-PRESSO was used to identify outliers; leave-one-out analysis was used to evaluate the impact of individual instrumental variables on the causal effect estimates; and MR Steiger test was used to evaluate the degree of reverse causation.
[0142] (3-6): Visualize the results of statistical analysis and sensitivity analysis;
[0143] (3-7): Perform colocalization analysis on the genes and corresponding outcome data in the visualization results to evaluate whether the genes and the target disease share the same causal instrumental variables. If so, the gene is used as the target drug target gene.
[0144] Colocalization analysis uses the method of Bayesian inference to calculate the posterior probability under five given assumptions. The five assumptions are as follows:
[0145] H0: Phenotype 1 and phenotype 2 are not significantly correlated with all SNP sites in a genomic region;
[0146] H1 / H2: Phenotype 1 or phenotype 2 is significantly correlated with SNP sites in a genomic region;
[0147] H3: Phenotype 1 and phenotype 2 are significantly correlated with SNP sites in a genomic region, but are driven by different causal variant sites;
[0148] H4: Phenotype 1 and phenotype 2 are significantly correlated with SNP sites in a genomic region and are driven by the same causal variant site.
[0149] The formula for calculating the posterior probability is as follows:
[0150]
[0151] Where, D represents the observed data (i.e., the data to be analyzed), S represents the setting, and S h represents the setting under different assumptions H h h = {1, 2, 3, 4}.
[0152] If the same instrumental variable is shared (i.e., PP.H4 > 0.7), it indicates that the SNP affects the target disease by influencing the gene.
[0153] Genes that pass all of (3-4) statistical analysis, (3-5) sensitivity analysis, and (3-7) colocalization analysis are taken as the obtained target drug target genes.
[0154] It should be understood that for the obtained target disease risk factors, potential drug target genes are found by means of drug data or external drug databases. Evidence at the cross-omics level is sought from the transcriptional and protein expression levels through Drug target Mendelian Randomization (DTMR), and verification is carried out through colocalization analysis and pathway enrichment analysis.
[0155] Exemplarily, (3-1): Obtain the known target genes of the target drug, including: retrieving existing literature or the DGIdb5.0 database [CANNON M, STEVENSON J, STAHL K, et al. Nucleic Acids Research. 2024 Jan 5. DOI: https: / / doi.org / 10.1093 / nar / gkad1040.]. Specifically, the DGIdb database collects all drug-gene interaction information such as DrugBank, PharmGKB, ChEMBL, Drug Target Commons, etc., so as to obtain the target genes of the target drug.
[0156] Exemplarily, (3-2): According to the known target genes, retrieve transcriptional data and protein data to obtain the corresponding transcriptional and protein expression levels of the known target genes. Taking the transcriptional and protein expression levels as exposure data means retrieving the obtained target for transcriptional data (eQTL data from eQTLGen and GTEx) or protein data (pQTL data from UKB-PPP, ARIC, and deCODE), or performing GWAS meta on the database to obtain the corresponding transcriptional and protein expression levels of the target as exposure data.
[0157] Exemplarily, (3-3): According to the set screening rules, screen out SNPs strongly related to the target genes from the exposure data, and take the screened SNPs as instrumental variables, including:
[0158] Select SNPs strongly related to the target genes from the exposure data as instrumental variables according to the following screening steps: Set the significance threshold to P < 5e -8 (For the eQTLGen database in transcriptomics, set it to 2.02e -5 , for the GTEx database, set it to 1e -5 ); There is no obvious linkage disequilibrium at a certain genetic distance (kb = 10000, r 2<0.001); Minor Allele Frequency (MAF) > 1%; weak instrumental variables are excluded using the F statistic, requiring where R 2 reflects the proportion of the variance of the exposure variable that can be explained by SNPs, N is the sample size, and k is the number of SNPs. The SNPs passing the above significance and independence tests can be used as instrumental variables.
[0159] Exemplarily, (3-4): perform statistical analysis on the obtained instrumental variables: input the obtained instrumental variable information into the "TwoSampleMR" R package and perform statistical analysis with the built-in outcome data in the system to estimate the causal effects of target genes on various phenotypes. When there is only one instrumental variable, the Wald ratio method is used to estimate the causal effect; when there are two or more instrumental variables, the IVW method is used to estimate the causal effect, and multiple testing correction is performed, with the significance P value taken as 0.05.
[0160] Exemplarily, (3-6): visualize the results of statistical analysis and sensitivity analysis: visualize the results that all pass the analysis. Since there are a large number of disease phenotypes analyzed, a Manhattan Plot is used to display the results of association analysis, that is, a scatter plot with the chromosome as the horizontal axis and the transformation of the P values obtained from each association analysis as the vertical axis.
[0161] Exemplarily, (3-7): perform colocalization analysis on the genes and corresponding outcome data in the visualization results to evaluate whether the gene and the target disease share the same causal instrumental variable. If they share the same instrumental variable, it indicates that the SNP affects the target disease by influencing the gene, including: feeding back the genes in the plotted results and the corresponding outcome GWAS data to the "coloc" R package (using the coloc.abf function in the coloc package called by RStudio) for colocalization analysis to evaluate whether the gene and the disease share the same causal instrumental variable. If there is a shared instrumental variable, it indicates that the SNP affects the disease by influencing the gene rather than other biological pathways.
[0162] Set the prior probabilities P1 = P2 = 1e -4 ; the prior probability P that the SNP is related to both gene expression and disease risk 12 = 1e -5 ; it is considered that PP.H4 > 0.7 indicates evidence of colocalization. It should be noted that the lack of evidence of colocalization does not prove that the MR result analysis is invalid.
[0163] Further, constructing an analysis database according to the risk factors of the target disease, the association relationship between the target drug and the target disease, and the target gene of the target drug, including:
[0164] Combining the risk factors of the target disease, the association relationship between the target drug and the target disease, and the target gene of the target drug to obtain an analysis database.
[0165] It should be understood that it involves a data analysis method for repurposing old drugs based on simulating a target trial. By constructing a database required for simulating the target trial and using domestic and foreign population data for quasi-randomized analysis, the therapeutic effect of the target drug on the target disease is evaluated, providing data support for the strategy of repurposing old drugs.
[0166] Further, analyzing the analysis database to obtain the new function prediction result of the target drug, including:
[0167] (5-1): Conduct baseline information analysis, compare the characteristics of the intervention group and the control group at baseline, and output the baseline characteristic comparison result;
[0168] First, calculate descriptive statistics (such as mean and standard deviation) for the baseline covariates (demographic indicators, disease history, etc.); for continuous variables with a normal distribution, an independent samples t-test is used; for continuous variables with a non-normal distribution, a rank sum test is used; for categorical variables, a chi-square test is used. The above statistical methods calculate the P-value of each baseline variable between the two groups. Generally, it is considered that P>0.05 indicates no significant difference between the groups and balanced baseline characteristics.
[0169] t-test: Used to compare the mean differences of two groups of continuous variables with a normal distribution. The formula is as follows:
[0170]
[0171] and are the sample means of the two groups respectively; and are the sample variances of the two groups respectively; n1 and n2 are the sample sizes of the two groups respectively.
[0172] Rank sum test: Used to compare the distribution differences of two groups of continuous variables with a non-normal distribution. The formula is as follows:
[0173]
[0174] n1 and n2 are the sample sizes of the two groups respectively; R1 is the rank sum of the first group, that is, the sum of the rankings of all data in the first group in the combined dataset.
[0175] Chi-square test: The chi-square test is used to compare the distribution differences of two groups of categorical variables. The formula is as follows:
[0176]
[0177] O is the observed frequency, that is, the number of times a particular classification combination appears (for example, there are 30 males in the intervention group); the expected frequency is the number of times a particular classification combination should theoretically appear under the assumption of "no difference between the two groups".
[0178] (5 - 2): Conduct the main analysis:
[0179] If the outcome indicator is survival data and the proportional hazards assumption is satisfied, the COX proportional hazards regression model will be used for analysis;
[0180] If the proportional hazards assumption is not satisfied, the time - dependent COX proportional hazards regression model will be used for adjustment;
[0181] Among them, survival data refers to the data type that includes the time dimension (the time from a certain starting point, such as the start of treatment, to the occurrence of the outcome event) and whether the outcome event occurs. The COX proportional hazards regression model is used to analyze the relationship between survival data and multiple covariates (also called independent variables). It assumes that the hazard ratio (HR) of different individuals is constant over time, that is, it satisfies the "proportional hazards assumption". The model formula is as follows:
[0182] h(t|X) = h0(t)exp(X T β)
[0183] h(t|X) is the hazard function at time t, given the covariates X, representing the instantaneous probability of the event occurring at time t. h0(t) is the baseline hazard function, representing the hazard when all covariates X are zero. exp(X T β) is the HR, representing the multiple of the individual's hazard relative to the baseline hazard given the covariates X.
[0184] Among them, the time - dependent COX proportional hazards regression model introduces time - dependent covariates and regression coefficients into the COX proportional hazards regression model. The model formula is as follows:
[0185] h(t|X(t)) = h0(t)exp(X(t) T β(t))
[0186] If the outcome indicator is continuous data, linear regression analysis will be conducted.
[0187] It should be understood that continuous data refers to variables that can take any value within a set range, and their values are continuous, such as weight, blood glucose level, etc.
[0188] Linear regression analysis is used to study the linear relationship between a continuous outcome indicator and independent variables. The basic form of the model is as follows:
[0189] Y = β0 + β1X1 + β2X2 + … + β n X n + ∈
[0190] Y is the dependent variable, that is, the continuous outcome indicator; X is the independent variable; β0 is the intercept term, representing the expected value of the dependent variable when all independent variables are zero; β1 etc. are regression coefficients, representing the linear influence of each independent variable on the dependent variable; is the error term, representing the random variation not explained by the model.
[0191] (5 - 3): Both COX proportional hazards regression and linear regression can output charts on HR and 95% CI of the exposure to survival risk. That is, by simulating the effects of certain interventions on health using observational cohort data, the obtained effect estimates can be used to predict the potential effects of the target drug on disease treatment.
[0192] Finally, based on the results of COX proportional hazards regression or linear regression analysis, combined with statistical significance and effect direction, the possible new functions of the target drug are screened out. If the target drug shows a significant protective effect or increased risk in a certain disease outcome, it can be speculated that the drug has new potential indications or mechanisms of action in this disease area, thus obtaining the prediction results of the new functions of the target drug.
[0193] Exemplarily, S104 - 1: Determine the research question and define the target variables. According to the drug targets and disease outcomes of the DTMR study, the system interaction interface provides the following:
[0194] Target drug: A drug targeting a specific drug target and recommended as the first choice or standard treatment by the guidelines.
[0195] Control drug: A drug recommended as the first choice or standard treatment by the guidelines or a conventional placebo.
[0196] Target outcome: A disease outcome indicator related to the study.
[0197] After the system receives the selection instruction, it automatically generates the definition of the research question and uses it as the basis for subsequent analysis. The variables provided by the system all come from the UK Biobank database (UKB) or the Cheeloo LEAD database. According to user needs, other variables can also be selected on the system, such as demographic indicators (such as age, gender, race, etc.); disease diagnosis information (such as diagnosis time, diagnosis code, etc.); comorbidities (such as hypertension, diabetes, etc.); medication history (such as drug name, medication time, dose, etc.). The system integrates the query results into a structured data set and stores it in a local or cloud database for use in subsequent steps.
[0198] It should be understood that the UKB database [BYCROFT C, FREEMAN C, PETKOVA D, et al. The UK Biobank resource with deep phenotyping and genomic data [J]. Nature, 2018, 562(7726): 203 - 9. DOI: 10.1038 / s41586 - 018 - 0579 - z.]
[0199] S104 - 2: Formulate preliminary inclusion and exclusion criteria. After extracting demographic indicators, disease outcome diagnoses, comorbidities, past medical history and medication history, disease diagnosis time, etc. from the two major databases in S104 - 1, the system can form a preliminary database. Then, according to the user's exclusion criteria, the system will delete a series of unavailable data: In the first step, exclude samples of patients who are not currently alive and patients who are not registered in the database; in the second step, exclude patients who do not have the simulated indications of the target drug and lack complete medical history data information; in the third step, exclude data of patients who have previously used the target drug; in the fourth step, exclude data of patients who have previously been diagnosed with the target disease outcome under study; in the fifth step, exclude samples lacking key variables (such as demographic indicators, disease diagnosis information, etc.). After the above five steps, the system forms a database that meets the inclusion and exclusion criteria.
[0200] S104 - 3: Design a simulated target trial. After completing the inclusion and exclusion operations in S104 - 2, the system's next page provides various simulated target trial designs. The system compares the target drug usage data with the control drug usage data. If there is a suitable active control drug available and there is a sequential treatment pattern, the active control new user design is used. Among them, the target drug is a drug targeting a specific drug target and is recommended by the guidelines as the first - choice or standard treatment for a certain disease; the control drug is a drug recommended by the guidelines as the first - choice or standard treatment for the target disease or a conventional placebo.
[0201] After completing the above treatment strategy, the system conducts treatment allocation. Mark the time point t0 when the treatment strategy is first used, classify the patients according to the strategy they receive at this time, and simulate randomization through a high - throughput L1 - regularized propensity score matching model.
[0202] In addition to matching [DEB S, AUSTIN P C, TU J V, et al. A Review of Propensity-Score Methods and Their Use in Cardiovascular Research [J]. Can J Cardiol, 2016, 32(2): 259-65. DOI: 10.1016 / j.cjca.2015.05.015], common methods also include inverse probability weighting [HOGAN J W, LANCASTER T. Instrumental variables and inverse probability weighting for causal inference from longitudinal observational studies [J]. Stat Methods Med Res, 2004, 13(1): 17-48. DOI: 10.1191 / 0962280204sm351ra], g-estimation [TOH S, HERNáNDEZ-DíAZ S, LOGAN R, et al. Estimating absolute risks in the presence of nonadherence: an application to a follow-up study with baseline randomization [J]. Epidemiology, 2010, 21(4): 528-39. DOI: 10.1097 / EDE.0b013e3181df1b69], and doubly robust methods [LI J, HANDORF E, BEKELMAN J, et al. Propensity score and doubly robust methods for estimating the effect of treatment on censored cost [J]. Stat Med, 2016, 35(12): 1985-99. DOI: 10.1002 / sim.6842], etc.
[0203] S104-4: After determining the design strategy, construct a cohort database. In the first step of the system, construct the target cohort: screen out patients who have not used the control drug from the database formed above; extract the first drug use record of the target drug among these patients as the target cohort; store the relevant information of the target cohort (such as drug use time, dose, patient characteristics, etc.) in the target cohort database.
[0204] The second step is to construct a control cohort: Select patients who have not used the target drug from the database formed above; Extract the first drug use records of the control drug among these patients as the control cohort; Store the relevant information of the control cohort (such as drug use time, dose, patient characteristics, etc.) in the control cohort database.
[0205] The third step is to construct an outcome cohort: Select samples with the target disease outcome from the target cohort and the control cohort; Use the selected samples as the outcome cohort; Store the relevant information of the outcome cohort (such as disease diagnosis time, outcome type, etc.) in the outcome cohort database.
[0206] The fourth step is that the system defines the follow-up start time (t0) as the time when the patient first uses the target drug or the control drug. Starting from t0, until the earliest diagnosis of the target disease, death, or the end of the 3-month observation after t0, the follow-up stops. The system then completes the definition of the follow-up information. After completing the above steps, the system forms the final analysis database.
[0207] Embodiment 2 provides a drug safety evaluation system;
[0208] The drug safety evaluation system includes:
[0209] A fourth acquisition module, which is configured to: Acquire the target gene of the target disease;
[0210] A fifth acquisition module, which is configured to: Based on the target gene of the target disease, acquire the intersection of significant target genes;
[0211] A safety evaluation module, which is configured to: Based on the intersection of significant target genes, evaluate the safety of drugs targeting the significant target genes.
[0212] Furthermore, the acquisition of the target gene of the target disease includes:
[0213] (6-1): Acquire the target disease input by the user; Acquire the exposure database;
[0214] (6-2): Screen out SNPs with strong genetic associations with the target gene from the exposure database as instrumental variables; The SNP refers to Single Nucleotide Polymorphism, and the strong genetic association means that the significance P value of the association between the SNP and the target target gene at the genome-wide level is less than the set threshold;
[0215] (6-3): Based on the target disease, exposure database, and instrumental variables, perform analysis using Mendelian randomization to obtain the results of Mendelian randomization analysis. When only one instrumental variable is screened for the exposure target gene, the Wald ratio method is used to estimate the causal effect. When there are two or more instrumental variables, the IVW method is used to estimate the causal effect. And perform multiple testing correction, with the significance P-value taken as 0.05, and then obtain the corrected significant target genes in each database.
[0216] (6-4): Perform sensitivity analysis on the corrected significant target genes obtained in each database.
[0217] (6-5): Perform colocalization analysis on the significant target genes obtained from the sensitivity analysis to evaluate whether the genes and the target disease share the same instrumental variables, and obtain the target genes with significant causal effects on the target disease in each database.
[0218] Exemplarily, the obtaining of the target disease input by the user is to select the disease or phenotype data of interest as the outcome data.
[0219] Exemplarily, the obtaining of the exposure database is to select 5 databases for transcriptomic eQTL {eQTLGen, GTEx} and proteomic pQTL {UKB-PPP, ARIC, deCODE} as the exposure datasets. Or first perform GWAS meta on the omics databases of the same type to obtain the corresponding transcriptional levels and protein expression levels of the target genes as the exposure data.
[0220] Exemplarily, the screening of SNPs strongly related to the target gene from the exposure database as instrumental variables includes: selecting SNPs strongly related to the target gene as instrumental variables:
[0221] Set the significance threshold as P < 5e -8 (For the eQTLGen database of transcriptomics, set it as 2.02e-5, and for the GTEx database, set it as 1e -5 );
[0222] There is no obvious linkage disequilibrium (r 2 <0.01) at a certain genetic distance;
[0223] The minor allele frequency (MAF) > 1%;
[0224] Use the F statistic to eliminate weak instrumental variables, requiring where, R 2Reflect the proportion of the variance of the exposure variable that can be explained by SNPs. N is the sample size and k is the number of SNPs. SNPs that pass the above significance and independence tests can be used as instrumental variables.
[0225] Exemplarily, the sensitivity analysis of the corrected significant target genes in each database includes:
[0226] Use other different MR methods such as maximum likelihood method, weighted mode, weighted median, MR-Egger regression, etc. to test the robustness and consistency of the results; use Cochran’s Q test to evaluate heterogeneity; use MR-Egger regression intercept term test to evaluate horizontal pleiotropy; use MR-PRESSO to identify outliers; use Leave-one-out Analysis to evaluate the impact of a single instrumental variable on the causal effect estimate; use MR Steiger test to evaluate the degree of reverse causation. Finally, retain the significant target genes for which the effect directions calculated by multiple MR methods including the IVW method are consistent, pass the heterogeneity and horizontal pleiotropy tests, and there is no reverse causation.
[0227] Exemplarily, the (6-5): perform colocalization analysis on the significant target genes obtained from the sensitivity analysis to evaluate whether the genes and the target disease share the same instrumental variable, and obtain the target genes in each database that have a significant causal effect on the target disease, including:
[0228] Feed the GWAS data of the significant target genes and the target outcome into the "coloc" R package for colocalization analysis to evaluate whether the gene and the target share the same causal instrumental variable. Set the prior probability P1 = P2 = 1e -4 The prior probability P that the SNP is related to both gene expression and disease risk 12 = 1e -5 It is considered that PP.H4 > 0.7 indicates evidence of colocalization, suggesting that the SNP affects the target outcome through its effect on the gene rather than other biological pathways. After a series of tests and analyses, finally, the target genes in each database that have a significant causal effect on the target outcome are obtained.
[0229] It should be understood that through massive gene screening of 5 databases including transcriptomics and proteomics, and using Mendelian randomization method for cross-omics analysis, the target genes in each database that have a significant causal effect on the target outcome are obtained.
[0230] Furthermore, the obtaining of the intersection of significant target genes based on the target disease's target genes includes:
[0231] Intersect the target gene results with significant causal effects in different databases and output the analysis results.
[0232] Exemplarily, based on the target genes with significant causal effects on the target outcome in each obtained database, further intersect the significant results in 5 databases to obtain the intersection of target genes that are consistently significant from the transcriptomics to the proteomics level. Specifically, it includes the following steps:
[0233] Result analysis and visualization: First, the system permutes and combines the transcriptomics and proteomics databases, and intersects the target gene results with significant causal effects in each finally obtained database, aiming to find the significantly consistent target genes from the transcriptomics to the proteomics level; then, the system automatically outputs the analysis result chart (forest plot) and provides the target gene information and druggability information of this part to the user to help them judge whether the gene has the potential to be a drug target.
[0234] Pathway enrichment analysis: If the user expects to further understand its biological mechanism, the system can perform pathway enrichment analysis based on the selected significant target genes with the help of the GO database and the KEGG database, and use the STRING database to draw the protein-protein interaction network (PPI) to interpret, verify and supplement the results.
[0235] Furthermore, based on the intersection of significant target genes, perform safety assessment on drugs targeting the significant target genes, including:
[0236] (8-1): Determine the exposure data: Use the intersection of target genes that are consistently significant from the transcriptomics to the proteomics level as the exposure data;
[0237] (8-2): Screen the instrumental variables based on the exposure data;
[0238] (8-3): Determine the outcomes: Use more than 9,000 phenotypes included in the IEU OPEN GWAS database and 784 binary phenotypes with the number of cases greater than or equal to 500 in the UKB SAIGEGWAS database as the full phenotype outcomes;
[0239] (8-4): Test the causal effect: Based on the exposure data, instrumental variables, and the full phenotype outcome database, use the PheWAS-MR system to evaluate the causal effect of candidate target genes on non-target disease phenotypes and obtain potential threats to drug safety.
[0240] Specifically, draw a Manhattan plot based on the chromosomal position of the gene (horizontal axis) and the negative logarithm transformation -log 10 (P) (vertical axis) of the P-value of the gene's association analysis with each phenotype to visualize the association analysis results.
[0241] Combined with Bonferroni multiple testing correction, the significance threshold is set to P′ = 0.05 / n, where n is the total number of non-target disease phenotypes, and the vertical warning line -log 10 (P′) is marked on the Manhattan plot.
[0242] If all association signals are below the warning line, it indicates that the target gene has initially passed the safety assessment and has no significant causal effect on other phenotypes;
[0243] If there are significant signals breaking through the warning line, phenotypic-specific Mendelian randomization is used for further verification to clarify its potential threat to drug safety.
[0244] Based on the intersection of significant target genes, the PheWAS-MR method is used to explore the causal relationship between target genes and various clinical outcomes, and then to evaluate the safety of drugs targeting these significant genes in clinical applications, so as to obtain potential drug target genes that can be used as treatment target outcomes in the future.
[0245] In addition, the system also provides multi-causal evidence fusion analysis, including traditional meta-analysis, network meta-analysis, GWAS meta-analysis, etc. When there are multiple available independent databases in the above two parts of the proposed R & D analysis process, researchers can choose to summarize and upload the results of different databases provided by the system to this section, and fuse multiple results through the corresponding meta-analysis. The system will provide a final causal effect estimate. For example:
[0246] The first acquisition module, when the selected exposure factor is included in multiple databases in the system, can first perform GWAS meta-analysis on the exposure factor data of different databases in (1-2) to find more significant instrumental variables and enter (1-3) for analysis.
[0247] The second acquisition module, after obtaining the effect estimates corresponding to the UKB and Cheeloo LEAD databases in (2-7), can perform meta-analysis on the effects of the two databases to summarize more accurate analysis results.
[0248] The third acquisition module, in (3-2), the system can perform GWAS meta on eQTL data from eQTLGen and GTEx or pQTL data from UKB-PPP, ARIC, and deCODE, and use the corresponding transcriptional level and protein expression level as exposure data to find more significant instrumental variables and enter the subsequent (3-3) for analysis.
[0249] Fourth acquisition module: When multiple databases in the system contain the selected exposure factors, GWAS meta-analysis can be first performed on the exposure factor data of different databases in (6-2) to find more significant instrumental variables and then enter the subsequent (6-3) for analysis.
[0250] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A drug new function prediction system, characterized in that, Comprising: A first acquisition module configured to: acquire risk factors of a target disease; A second acquisition module configured to: acquire the association relationship between a target drug and the target disease; A third acquisition module configured to: acquire the target genes of the target drug; A database construction module configured to: construct an analysis database based on the risk factors of the target disease, the association relationship between the target drug and the target disease, and the target genes of the target drug; An analysis and prediction module configured to: analyze the analysis database to obtain a prediction result of the new function of the target drug.
2. The drug new function prediction system according to claim 1, characterized in that, The acquisition of the risk factors of the target disease includes: (1-1): Screen exposure factors and genetic instrumental variables based on the target disease; the principle for screening exposure factors and genetic instrumental variables based on the target disease is: set the significance threshold P < 5e -8 and the linkage disequilibrium clustering threshold r 2 < 0.001 and select the genetic instrumental variables; (1-2): Perform Mendelian randomization analysis on the target disease, exposure factors, and genetic instrumental variables to obtain Mendelian randomization analysis results; (1-3): Perform multiple test corrections on the Mendelian randomization analysis results. If more than 1 significant exposure factor is screened out by the analysis, Bonferroni correction is used, that is, the genome-wide significance threshold is adjusted to P < 0.05 / n; (1-4): According to the correction results, output a result chart to obtain the risk factors of the target disease; when the P value corresponding to the causal effect estimate of a certain exposure is less than the genome-wide significance threshold after Bonferroni correction, it is considered that this exposure is a risk factor of the target disease.
3. The drug new function prediction system according to claim 1, characterized in that, The acquisition of the association relationship between the target drug and the target disease includes: (2-1): Determine a core variable database, which includes: variables of the target drug and variables of the target disease; (2-2): Clean the data of the patients in the core variable database; (2-3): Construct an observational cohort based on the cleaned core variable database; the observational cohort includes: a target drug cohort and a control cohort; (2-4): Perform descriptive statistical analysis on the basic characteristics of the target drug cohort and the control cohort; (2-5): Baseline characteristic comparison: Use statistical methods to evaluate the differences in confounding factors between the target drug cohort and the control cohort; The statistical methods include: chi-square test, t-test or non-parametric test; Specifically, if it is a normally distributed continuous variable, an independent sample t-test is used; if it is a non-normally distributed continuous variable, a rank sum test is used; if it is a categorical variable, a chi-square test is used; By statistical methods, calculate the P value of each baseline variable between the two groups. If P > 0.05, it means there is no significant difference between the groups and the baseline characteristics are balanced; (2-6): Based on the baseline characteristic comparison results, estimate the effect of the target drug on the target disease; When the outcome indicator is a binary variable, use a logistic regression model to analyze the impact of the target drug on the probability of the outcome occurrence; When the outcome indicator is survival data and the proportional hazards assumption is satisfied, use a Cox proportional hazards regression model to analyze the impact of the target drug on the risk of disease occurrence; If the outcome indicator is a continuous variable, use a linear regression model to analyze the impact of the target drug on the outcome indicator; (2-7): Based on the effect estimation results of the target drug on the target disease, perform a sensitivity analysis to obtain the association relationship between the target drug and the target disease.
4. The drug new function prediction system according to claim 1, characterized in that, The obtaining of the target gene of the target drug includes: (3-1): Obtain the potential drug target genes of the risk factors of the target disease; (3-2): According to the potential target genes, retrieve the transcriptional data and protein data to obtain the corresponding transcriptional level and protein expression level of the potential target genes, and use the transcriptional level and protein expression level as exposure data; (3-3): According to the set screening rules, screen out the SNPs strongly related to the target gene from the exposure data, and use the screened SNPs as instrumental variables; The set screening rules include: Rule 1: Set the significance threshold as P < 5e -8 ; Rule two: There is no obvious linkage disequilibrium at the set genetic distance; Rule three: Minor allele frequency > 1%; Rule 4: Use the F statistic to exclude weak instrumental variables, requiring that where R 2 reflects the proportion of the variance of the exposure variable that can be explained by the SNPs, N is the sample size, and k is the number of SNPs; Rule one, rule two, rule three and rule four are in an AND relationship; (3-4): Perform statistical analysis on the obtained instrumental variables: When there is only one instrumental variable, use the Wald ratio method to estimate the causal effect; when there are two or more instrumental variables, use the IVW method to estimate the causal effect and perform multiple test corrections, and take the significance P value as 0.05; Take the exposure outcome pair corresponding to P < 0.05 as the positive result; (3-5): Perform a sensitivity analysis on the positive results obtained from the statistical analysis, and use the genes passing the sensitivity analysis as the target drug target genes; (3-6): Visualize the statistical analysis results and sensitivity analysis results; (3-7): Perform a colocalization analysis on the genes and the corresponding outcome data in the visualization results to evaluate whether the genes and the target disease share the same causal instrumental variable. If so, use the genes as the target drug target genes.
5. The drug new function prediction system according to claim 1, characterized in that, The construction of the analysis database according to the risk factors of the target disease, the association relationship between the target drug and the target disease, and the target gene of the target drug includes: Merge the risk factors of the target disease, the association relationship between the target drug and the target disease, and the target gene of the target drug to obtain the analysis database.
6. The drug new function prediction system according to claim 1, characterized in that, The analysis of the analysis database to obtain the new function prediction results of the target drug includes: (5-1): Perform a baseline information analysis, compare the characteristics of the intervention group and the control group at baseline, and output the baseline characteristic comparison results; (5-2): Perform a main analysis: If the outcome indicator is survival data and the proportional hazards assumption is met, use the COX proportional hazards regression model for analysis; if the proportional hazards assumption is not met, use the time-dependent COX proportional hazards regression model for adjustment; (5-3): Both COX proportional hazards regression and linear regression can output charts of HR and 95% CI regarding the exposure to the survival risk; According to the results of COX proportional hazards regression or linear regression analysis, combined with statistical significance and effect direction, screen out the possible new functions of the target drug.
7. A drug safety evaluation system, characterized in that, Include: The fourth acquisition module, which is configured to: acquire the target gene of the target disease; The fifth acquisition module, which is configured to: based on the target gene of the target disease, acquire the intersection of significant target genes; A safety assessment module, which is configured to: perform a safety assessment on drugs targeting significant target genes based on the intersection of significant target genes.
8. The drug safety evaluation system according to claim 7, wherein The obtaining of the target genes of the target disease includes: (6-1): Obtain the target disease input by the user; obtain the exposure database; (6-2): Screen out SNPs with strong genetic associations with the target genes from the exposure database as instrumental variables; the SNP refers to single nucleotide polymorphism, and the strong genetic association means that the significance P value of the association between the SNP and the target target gene at the genome-wide level is less than the set threshold; (6-3): Based on the target disease, exposure database and instrumental variables, perform analysis using Mendelian randomization to obtain the Mendelian randomization analysis results; when only one instrumental variable is screened out for the exposure target gene, the Wald ratio method is used to estimate the causal effect; when there are two or more instrumental variables, the IVW method is used to estimate the causal effect; and perform multiple testing correction, with the significance P value taken as 0.05, and then obtain the corrected significant target genes in each database; (6-4): Perform a sensitivity analysis on the corrected significant target genes obtained in each database; (6-5): Perform a colocalization analysis on the significant target genes obtained from the sensitivity analysis to evaluate whether the genes and the target disease share the same instrumental variables, and obtain the target genes with significant causal effects on the target disease in each database.
9. The drug safety evaluation system according to claim 7, wherein The performing of a safety assessment on drugs targeting significant target genes based on the intersection of significant target genes includes: (8-1): Determine the exposure data: Use the intersection of target genes that are consistently significant from transcriptomics to proteomics as the exposure data; (8-2): Screen the instrumental variables based on the exposure data; (8-3): Determine the outcomes: Use more than 9,000 phenotypes included in the IEU OPEN GWAS database and 784 binary phenotypes with the number of cases greater than or equal to 500 in the UKB SAIGE GWAS database as the full set of phenotypic outcomes; (8-4): Test the causal effect: Based on the exposure data, instrumental variables, and the full set of phenotypic outcome database, use the PheWAS-MR system to evaluate the causal effect of candidate target genes on non-target disease phenotypes, and obtain potential threats to drug safety.
10. The drug safety evaluation system according to claim 9, characterized in that, Using the PheWAS-MR system to evaluate the causal effect of candidate target genes on non-target disease phenotypes and obtaining potential threats to drug safety includes: Negative logarithm transformation -log of the P-value based on the chromosomal location of the gene and the association analysis between the gene and each phenotype 10 (P) Draw a Manhattan plot to visualize the results of the association analysis; Combined with Bonferroni multiple test correction, the significance threshold is set to P′ = 0.05 / n, where n is the total number of non-target disease phenotypes, and the vertical coordinate warning line -log 10 (P′); If all association signals are below the warning line, it is indicated that the target gene has initially passed the safety assessment and has not produced a significant causal effect on other phenotypes; If there are significant signals breaking through the warning line, further verify through phenotype-specific Mendelian randomization to clarify its potential threat to drug safety.