Method for initial screening of complex disease drugs based on whole genome association signals
By using the DESE algorithm to screen cell types for complex diseases and combining drug-induced gene expression profiles and GWAS data, the specific perturbations of genes by drugs can be predicted, solving the problem of drug screening for complex diseases and achieving precise narrowing of the drug candidate range.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SUN YAT SEN UNIV
- Filing Date
- 2022-09-28
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies are insufficient for effectively screening drugs for complex diseases, especially due to the lack of appropriate primary tissue samples and drug-induced changes in gene expression, resulting in poor performance of traditional methods in drug research for complex diseases such as mental illnesses.
The DESE algorithm was used to screen for complex disease-related cell types. Gene expression profiles were induced by drug blank control, and GWAS data and Mann-Whitney U test were combined to predict drug-specific perturbations to genes. Iterative analysis was then used to narrow down the drug screening range.
It enables the prediction of specific perturbations in susceptibility genes for complex diseases, narrowing the scope of drug screening and improving the accuracy and efficiency of drug screening.
Smart Images

Figure CN115547429B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-throughput drug screening technology, and in particular to a method for initial screening of drugs for complex diseases based on genome-wide association signals. Background Technology
[0002] In recent years, researchers have increasingly focused on finding new indications for existing drugs to reduce drug development costs and accelerate the research process. While there is already considerable research on cancer-related drugs, research on drugs for other complex diseases (such as mental illnesses) seems relatively limited. This difference is due to several factors. First, compared to cancer research, it is more difficult to collect appropriate primary disease tissues (such as brain tissue) from patients or healthy individuals for research on other complex diseases (such as mental illnesses). Second, patients with complex diseases often take medications to alleviate symptoms, which may alter the expression status of disease-related genes. Consequently, obtaining disease-related tissues (such as brain tissue from patients with mental illnesses) may not fully reflect the biological basis of the disease. With the rapid development of genomic technologies, some potential solutions have emerged for drug research on complex diseases (such as mental illnesses). Existing research techniques include the use of genome-wide association studies (GWAS) for complex diseases. Identified phenotype-associated genetic variants are further located to the genes containing these variants, and then analyzed to determine if these genes are target genes of known drugs. However, this technique is based on single gene targets and is not applicable to complex diseases with multiple susceptibility genes. Another approach is to calculate differentially expressed genes in patients relative to healthy controls and the fold change in expression of these differentially expressed genes based on transcriptomic data from the primary tissue of the disease. However, the primary tissues of many complex diseases are still unclear. In addition, some primary tissues are difficult to collect (such as brain tissue from patients with mental illness), which limits the use of this technique. Summary of the Invention
[0003] To address the aforementioned technical problems, the present invention aims to provide a method for preliminary screening of drugs for complex diseases based on genome-wide association signals, which can narrow down the screening range of candidate drugs for complex diseases by predicting drugs that specifically perturb multiple disease susceptibility genes.
[0004] The first technical solution adopted in this invention is: a method for preliminary screening of drugs for complex diseases based on genome-wide association signals, comprising the following steps:
[0005] Based on genome-wide association signals, cells with complex diseases are screened using the DESE algorithm to obtain cell lines with complex diseases. The DESE algorithm is an algorithm that uses the Mann-Whitney U rank sum test to examine the significance of disease-associated genes in tissue cells to infer cell types associated with complex diseases.
[0006] By inducing and analyzing cell lines of complex diseases using drug blank controls, the specific perturbation spectrum of genes by drugs can be obtained.
[0007] Based on genome-wide association signals and conditional data, cyclic prediction analysis and calculation of drug-specific perturbation spectra of genes are performed to obtain drugs for susceptibility genes of complex diseases.
[0008] Furthermore, the step of inducing and analyzing cell lines for complex diseases using a drug blank control to obtain a specific perturbation spectrum specifically includes:
[0009] Gene expression profiles were obtained by inducing cell lines with complex diseases using a drug blank control and then performing fold-labeling.
[0010] Gene expression profiles were calculated using a robust z-score scoring method to obtain the positive specific perturbation profiles of genes.
[0011] After inverting the positive-specific perturbation spectrum of the gene, a robust z-score scoring method is applied to obtain the reverse-specific perturbation spectrum of the gene.
[0012] By integrating the positive and negative specific perturbation spectra of genes, the specific perturbation spectra of drugs on genes are obtained.
[0013] Furthermore, the formula for calculating the positive specific perturbation spectrum of the drug on the gene is as follows:
[0014]
[0015] In the above formula, This indicates that the p-value test is adjusted to follow a uniform distribution. Indicates the first The fold change in gene expression level after a drug induces a specific gene. This represents the average fold change in gene expression after induction by all drugs. This represents the standard deviation of the fold change in gene expression after induction by all drugs. This indicates the positive specific perturbation spectrum of the drug on the gene.
[0016] Furthermore, the step of performing cyclic prediction analysis and calculation on the drug-specific perturbation spectrum of genes based on genome-wide association signals and conditional data to obtain drugs for complex disease susceptibility genes specifically includes:
[0017] Collect GWAS summary data, reference genotype data, and reference gene model data. Perform conditional association analysis based on the ranking of gene significance to phenotype to predict potential susceptibility genes associated with the phenotype and obtain genes that are significant under the conditional association pattern.
[0018] The trend of conditionally significant genes perturbing drug specificity was examined, and the p-value of the drug was obtained by Mann-Whitney U test.
[0019] The specific perturbation score of a gene is obtained by calculating the characteristic perturbation based on the drug's P-value and the gene's ranking in drug-specific perturbation.
[0020] Genes are sorted based on their specific perturbation scores, and a new batch of disease-associated genes is obtained by re-performing conditional association tests. The trend of their drug-specific perturbations is then re-examined. The above conditional association test, Mann-Whitney U test, and characteristic perturbation calculation steps are repeated until the loop termination condition is met, and drugs for complex disease susceptibility genes are output.
[0021] Furthermore, the formula for calculating the characteristic perturbation based on the drug's P-value and the gene's ranking in the drug's specific perturbation is shown below:
[0022]
[0023] In the above formula, This represents the sum of the number of disease susceptibility genes and the number of non-susceptibility genes. Indicates the first One drug, Indicates the first One inducer gene, Indicates the first Drug-induced genes Specific perturbation effect in drugs The induced sequence of all genes, express This drug express The p-value of the drug after the Mann-Whitney U test.
[0024] Furthermore, the loop is an iterative calculation, which includes repeatedly executing the prediction step, the verification step, and the perturbation calculation step until the following termination condition is met:
[0025]
[0026] In the above formula, express The drug in the Mann-Whitney U test on the 1st Each P-value.
[0027] The beneficial effects of the method of this invention are as follows: This invention uses the DESE algorithm to screen cells with complex diseases, and then uses a large number of drugs to induce the specific expression of the cell types with complex diseases. Among them, the therapeutic drugs tend to specifically perturb (upregulate or downregulate) the expression of some disease susceptibility genes. Finally, by combining conditional gene association analysis and drug-specific perturbation analysis, drugs that can specifically perturb multiple disease susceptibility genes can be predicted, thereby narrowing the screening range of candidate drugs for complex diseases. Attached Figure Description
[0028] Figure 1 This is a flowchart of the steps of the method for preliminary screening of drugs for complex diseases based on genome-wide association signals according to the present invention;
[0029] Figure 2 This is a schematic diagram of the process of screening complex disease cells using the DESE algorithm in this invention;
[0030] Figure 3 This is a schematic diagram of the process of inducing and analyzing cell lines of complex diseases using a large number of drugs in this invention;
[0031] Figure 4 This is a schematic diagram of the process by which the present invention analyzes and calculates the drug-specific perturbation spectrum of genes by combining GWAS summary data. Detailed Implementation
[0032] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. The step numbers in the following embodiments are only for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adapted according to the understanding of those skilled in the art.
[0033] Reference Figure 1 This invention provides a method for preliminary screening of drugs for complex diseases based on genome-wide association signals, the method comprising the following steps:
[0034] S1. Based on genome-wide association signals, the DESE algorithm is used to screen cells of complex diseases to obtain cell lines of complex diseases. The DESE algorithm is an algorithm that uses the Mann-Whitney U rank sum test to examine the significance of disease-associated genes in tissue cells to infer cell types associated with complex diseases.
[0035] Specifically, refer to Figure 2Based on gene expression profile data of M cell lines, the prediction algorithm DESE is used to infer the cell types associated with complex diseases in these M cell lines. The DESE algorithm uses a conventional nonparametric test (Mann-Whitney U rank-sum test) to examine the significance of disease-associated genes in tissue cells. The basic principle is that disease-associated genes generally tend to be specifically expressed in the primary tissue of the disease. DESE calculates a P-value for each cell type (representing the significance of the association between the cell type and the complex disease), and then uses a corrected P-value of less than 0.05 to screen cell types that are significantly associated with complex diseases.
[0036] S2. By inducing and analyzing cell lines of complex diseases using drug blank controls, the specific perturbation spectrum of genes by drugs can be obtained.
[0037] Specifically, refer to Figure 3 This invention performs differential gene expression analysis based on gene expression profiles generated by a large number of drugs and controls (such as dimethyl sulfoxide DMSO) induction of complex disease-related cell lines, and generates drug-induced differential gene expression profiles.
[0038] S21. Inducing cell lines with complex diseases using a drug blank control and performing fold labeling treatment to obtain gene expression profiles;
[0039] Specifically, if the number of cell types associated with complex diseases in step S1 is greater than 1, the gene expression profiles generated by the blank control-induced phenotype-related cell lines are averaged to generate the combined blank control-induced gene expression profiles; the gene expression profiles generated by the drug-induced cell lines of the same drug are averaged to generate the drug-induced gene expression profiles of the same drug-related cell lines (regardless of cell type, duration of drug treatment, or drug concentration).
[0040] S22. Gene expression profiles are calculated and processed using a robust z-score scoring method to obtain the positive specific perturbation profiles of genes for certain drugs.
[0041] Specifically, the robust z-score scoring method is a method for measuring the specific expression of genes in tissues, assuming that the gene... exist Different tissues have expression levels Using Huber robust regression pairs, calculate the weighted value of each gene with respect to the regression line. The greater the deviation from the regression line, the smaller the weight. Calculate the weighted mean based on these weighted values. and weighted standard deviation Then a certain gene The robust expression for a certain organization-specific expression is shown below:
[0042]
[0043] Genes in complex disease-related cell lines In medicine The fold change in gene expression caused by interference is denoted as ,in Indicates that it is made from drugs Inducible genes The resulting gene expression levels, This indicates that the gene was induced by a blank control (such as DMSO). To avoid infinite values in the base-2 logarithmic calculation, the present invention adds 1 to both the numerator and denominator when calculating the fold change in gene expression level.
[0044] Furthermore, the formula for calculating the positive specific perturbation spectrum of the drug on the gene is as follows:
[0045]
[0046] In the above formula, This indicates that the p-value test is adjusted to follow a uniform distribution. Indicates the first The fold change in gene expression level after a drug induces a specific gene. This represents the average fold change in gene expression after induction by all drugs. This represents the standard deviation of the fold change in gene expression after induction by all drugs. This indicates the positive specific perturbation spectrum of the drug on the gene;
[0047] In the above formula, the Z-score quantitatively describes the deviation between the fold change in gene expression induced by each drug and the fold change in gene expression induced by all drugs. This is used to adjust the (hypothesis test) P-value to make it follow a uniform distribution;
[0048] S23. After inverting the positive specific perturbation spectrum of the gene, the robust z-score scoring method is applied to obtain the reverse specific perturbation spectrum of the gene.
[0049] Specifically, based on the differential expression profile of genes induced by drugs, the specific perturbation profile of genes by drugs is calculated. In actual analysis, this invention assumes that therapeutic drugs may specifically upregulate or downregulate the expression of disease susceptibility genes, or both.
[0050] S24. Integrate the original specific perturbation spectrum and the reverse specific perturbation spectrum to obtain the drug-specific perturbation spectrum of the gene.
[0051] Specifically, the generated drug-to-gene specific perturbation spectrum is referred to as the original drug perturbation spectrum. Then, the present invention inverts the data matrix of the original perturbation spectrum to obtain the inverted drug-to-gene specific perturbation spectrum, which is the reversed drug perturbation spectrum.
[0052] S3. Based on genome-wide association signals and combined with conditional data, cyclic prediction analysis and calculation of the drug-specific perturbation spectrum of genes are performed to obtain drugs for susceptibility genes of complex diseases.
[0053] Specifically, refer to Figure 4 Based on the data collected by S31 and an iterative cycle, conditional gene association analysis and drug-specific perturbation analysis of susceptibility genes were performed.
[0054] S31. Collect GWAS summary data, reference genotype data, and reference gene model data. Perform conditional association tests based on the ranking of gene-phenotype significance to obtain genes that are significant under the conditional association pattern.
[0055] Specifically, the GWAS summary data is used to obtain disease-associated genes, the reference genotype data is used to calculate linkage disequilibrium between loci and genes, and the reference gene model data is used to assign GWAS loci to specific genes based on their physical locations.
[0056] Collect large-scale GWAS summary data for complex diseases (available from public databases), reference genotype data (available from the 1000 Genomes Project website for the corresponding ethnic group), reference gene model data (RefSeqGene data can be obtained from relevant web pages), gene type information from the HGNC database, and raw drug perturbation spectrum (reverse drug perturbation spectrum) data generated in step S24.
[0057] The first step of the first iteration loop is to perform conditional gene association analysis using the Effective Chi-Square Statistic (ECS) algorithm, combined with large-scale GWAS summary data, reference genotype data, and reference gene model data from complex diseases, to predict potential susceptibility genes with phenotypic associations. The specific prediction process involves first sorting genes according to a certain score, and then using ECS to perform a gene-level association test on the first gene to obtain the statistic. The conditional association statistic for the second gene is used to infer whether the gene is associated with a disease. The conditional association statistic for the third gene is And so on;
[0058] In this analysis, genes enter the conditional association process in order of their p-values. Generally speaking, genes that enter the process first have a higher chance of retaining significant p-values, which are obtained by the prediction algorithm DESE.
[0059] S32. Examine the trend of conditionally significant genes on drug-specific perturbation, and obtain the drug's P-value by performing the Mann-Whitney U test.
[0060] Specifically, the second step of the first iteration loop is to analyze whether the conditionally associated significant genes generated in S32 have statistically significant enrichment in the expression profile of each drug through specific perturbations. The statistical test used is the Mann-Whitney U test (i.e., the Wilcoxon rank-sum test). The Mann-Whitney U test assumes that there are n genes in the genome, of which m genes are disease-associated genes. The Mann-Whitney U test is used to compare whether the m genes have higher specific expression than the nm genes, and then the statistical significance P-value of each drug can be obtained.
[0061] S33. Based on the drug's P-value and the gene's ranking in drug-specific perturbation, perform characteristic perturbation calculations to obtain the gene's specific perturbation score.
[0062] Specifically, the third step of the first iteration loop is to calculate the drug-specific perturbation score of each gene across all drugs, thereby reordering all genes. Assuming... The p-value of the drug after the Mann-Whitney U test was Then for the first The invention relates to a drug, which is sorted in descending order based on the robust z-scores generated by the drug perturbing all genes. Assuming the [number]th [gene]... Drug-induced genes Specific perturbation effect in drugs The order of all genes induced is So, genes Drug-specific perturbation score among all drugs The following formula can be used to calculate it, and the calculation formula is shown below:
[0063]
[0064] In the above formula, This represents the sum of the number of disease susceptibility genes and the number of non-susceptibility genes. Indicates the first One drug, Indicates the first One inducer gene, Indicates the first Drug-induced genes Specific perturbation effect in drugs The induced sequence of all genes, express This drug express The p-value of the drug after the Mann-Whitney U test.
[0065] S34. Sort the genes according to the gene-specific perturbation score, re-perform the conditional association test to obtain a new batch of disease-associated genes, re-examine the trend of their drug-specific perturbation, and repeat the above conditional association test step, Mann-Whitney U test step and characteristic perturbation calculation step until the loop termination condition is met, and output the drugs for complex disease susceptibility genes.
[0066] Specifically, after the first iteration is completed, the second iteration will begin with a new conditional gene association analysis. Genes with higher drug-specific perturbation scores will be prioritized for the new conditional association test. Subsequent iterations will follow the steps and sequence of the first iteration until the next iteration. The significant differences in the perturbation effects of each drug on susceptible and non-susceptible genes of complex diseases during the next iteration and the first iteration The differences between the two orders are very small at the significance level, and the following termination condition is met:
[0067]
[0068] In the above formula, express The drug in the Mann-Whitney U test on the 1st One P-value;
[0069] Once the iteration process stops, each drug will generate a significant P-value associated with the phenotype. Drugs with a corrected P-value less than 0.05 are considered to be associated with complex diseases. It is important to note that the original perturbation spectrum can be used to predict drugs that can specifically upregulate potential susceptibility genes for complex diseases, while the reverse perturbation spectrum can be used to predict drugs that can specifically downregulate potential susceptibility genes for complex diseases.
[0070] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A method for preliminary screening of drugs for complex diseases based on genome-wide association signals, characterized in that, Includes the following steps: Based on genome-wide association signals, cells with complex diseases are screened using the DESE algorithm to obtain cell lines with complex diseases. The DESE algorithm is an algorithm that uses the Mann-Whitney U rank sum test to examine the significance of disease-associated genes in tissue cells to infer cell types associated with complex diseases. By inducing and analyzing cell lines of complex diseases using drug blank controls, the specific perturbation spectrum of genes by drugs can be obtained. Based on genome-wide association signals and conditional data, cyclic prediction analysis and calculation of drug-specific perturbation spectra of genes are performed to obtain drugs for susceptibility genes of complex diseases.
2. The method for preliminary screening of drugs for complex diseases based on genome-wide association signals according to claim 1, characterized in that, The step of inducing and analyzing cell lines for complex diseases using a drug blank control to obtain a specific perturbation spectrum specifically includes: Gene expression profiles were obtained by inducing cell lines with complex diseases using a drug blank control and then performing fold-labeling. Gene expression profiles were calculated using a robust z-score scoring method to obtain the positive specific perturbation profiles of genes. After inverting the positive-specific perturbation spectrum of the gene, a robust z-score scoring method is applied to obtain the reverse-specific perturbation spectrum of the gene. By integrating the positive and negative specific perturbation spectra of genes, the specific perturbation spectra of drugs on genes are obtained.
3. The method for preliminary screening of drugs for complex diseases based on genome-wide association signals according to claim 2, characterized in that, The formula for calculating the positive specific perturbation spectrum of the drug on genes is shown below: In the above formula, This indicates that the p-value test is adjusted to follow a uniform distribution. Indicates the first The fold change in gene expression level after a drug induces a specific gene. This represents the average fold change in gene expression after induction by all drugs. This represents the standard deviation of the fold change in gene expression after induction by all drugs. This indicates the positive specific perturbation spectrum of the drug on the gene.
4. The method for preliminary screening of drugs for complex diseases based on genome-wide association signals according to claim 3, characterized in that, The step of performing cyclic prediction analysis and calculation on the drug-specific perturbation spectrum of genes based on genome-wide association signals and conditional data to obtain drugs for susceptibility genes of complex diseases specifically includes: Collect GWAS summary data, reference genotype data, and reference gene model data. Perform conditional association analysis based on the ranking of gene significance to phenotype to predict potential susceptibility genes associated with the phenotype and obtain genes that are significant under the conditional association pattern. The trend of conditionally significant genes perturbing drug specificity was examined, and the p-value of the drug was obtained by Mann-Whitney U test. The specific perturbation score of a gene is obtained by calculating the characteristic perturbation based on the drug's P-value and the gene's ranking in drug-specific perturbation. Genes are sorted based on their specific perturbation scores, and conditional association tests are performed again to obtain a new batch of disease-associated genes. Their trends in drug-specific perturbation are re-examined, and the above prediction, testing, and perturbation calculation steps are repeated until the loop termination condition is met, and drugs for complex disease susceptibility genes are output.
5. The method for preliminary screening of drugs for complex diseases based on genome-wide association signals according to claim 4, characterized in that, The formula for calculating characteristic perturbation based on the drug's P-value and the gene's ranking in drug-specific perturbation is shown below: In the above formula, This represents the sum of the number of disease susceptibility genes and the number of non-susceptibility genes. Indicates the first One drug, Indicates the first One inducer gene, Indicates the first Drug-induced genes Specific perturbation effect in drugs The induced sequence of all genes, express This drug express The p-value of the drug after the Mann-Whitney U test.
6. The method for preliminary screening of drugs for complex diseases based on genome-wide association signals according to claim 4, characterized in that, The loop is an iterative calculation, which includes repeatedly executing the prediction step, the verification step, and the perturbation calculation step until the following termination condition is met: In the above formula, express The drug in the Mann-Whitney U test on the 1st Each P-value.
Citation Information
Patent Citations
Screening method for multi-target drugs and / or pharmaceutical combinations
CN104965998A
Method for screening disease phenotype related mutation sites and application thereof
CN112735594A