Solid tumor treatment target prediction method and system based on multi-modal omics data
Through self-developed computational medical algorithms, the multimodal omics data are integrated, and the therapeutic targets of solid tumors are identified and the drug reuse combination is optimized, which solves the problem of difficult to effectively integrate multimodal omics data in the existing technology, and achieves efficient therapeutic target identification and drug combination optimization for complex diseases such as pancreatic cancer.
Patent Information
- Application Number
- CN202510122596.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively integrate multimodalomic data to identify therapeutic targets and drug reuse combinations for solid tumors, especially in complex diseases such as pancreatic cancer.
Using a method based on self-developed computational medical algorithms, multimodal omics data is integrated, including traditional omics, clinical omics, imaging omics and pathological omics data, functional enrichment analysis is performed through Fisher’s method and KEGG pathway collection to identify potential therapeutic targets, and potential drug targets are determined through drug reuse analysis and perturbation removal analysis.
It has achieved efficient identification of solid tumor therapeutic targets and optimization of drug reuse combinations, improved the efficiency of precise treatment, and provided new therapeutic hope for complex diseases such as pancreatic cancer.
Smart Images

Figure CN120072033A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technologies of bioinformatics tool development and drug target prediction, and particularly to the field of computational medicine research for integrating and analyzing multi-omics data based on self-developed computational medicine algorithms to identify treatment targets for solid tumors and drug reutilization combinations. Background Art
[0002] Cancer is one of the main factors leading to the global disease burden, and the global cancer burden will continue to increase in the next 30 years, posing a major threat to human life safety and quality of life. Data shows that cancer accounts for about 1 / 6 of global deaths and nearly 1 / 4 of deaths caused by non-communicable diseases, and is an important factor leading to premature non-communicable deaths globally. In 2022, there were 19.96 million new cases and 9.74 million deaths. It is estimated that by 2050, the global cancer burden will increase to an alarming 35.3 million new cancer cases, highlighting the urgency of discovering effective treatment strategies.
[0003] A solid tumor is an abnormal tissue mass, usually without cysts or fluid areas, and can be either benign or malignant. They can occur in various parts of the body, such as the breast, lung, prostate, or colon, and are composed of a group of abnormal cells that grow uncontrollably, forming a mass. The occurrence of solid tumors involves complex interactions of factors such as genetics, environment, and lifestyle. Some solid tumors may be caused by genetic mutations, such as BRCA1 and BRCA2 gene mutations being associated with breast and ovarian cancers; environmental factors such as exposure to carcinogens and viral infections also increase the risk of solid tumor occurrence; poor eating habits, obesity, alcohol, smoking, and other lifestyle factors are also risk factors.
[0004] Taking pancreatic cancer, one of the deadliest tumors globally, as an example, the global cancer statistics in 2022 showed that the incidence rate of pancreatic cancer ranked 12th among global tumor incidences, but the mortality rate was as high as the 6th among malignant tumor deaths. It has a rapid development, extremely poor prognosis, and a very low 5-year survival rate. Pancreatic cancer has an insidious onset, with atypical early symptoms, often manifested as symptoms such as indigestion, diarrhea, upper abdominal discomfort, or back pain, and is easily confused with other digestive system diseases. Data from the China Pancreatic Disease Big Data Center shows that pancreatic cancer has the characteristics of low early diagnosis rate, low surgical resection rate, and low drug effectiveness rate.
[0005] 10% to 20% of pancreatic cancers are associated with genetic factors. Therefore, in-depth understanding of the genetic basis of pancreatic cancer is of great significance for revealing the genetic susceptibility of the disease and developing new treatment strategies. Retrospective analysis has shown that drug targets supported by genetic evidence can multiply the success rate of drug development, fully demonstrating the broad prospects of pancreatic cancer drug development based on genetic analysis. Studies have shown that individuals carrying mutations in genes such as BRCA2 and ATM have a significantly increased risk of pancreatic cancer. In addition, due to the rapid development of multi-omics detection technologies, multiple genetic loci related to pancreatic cancer have been discovered in the past decade, such as the ABO gene (9q34.2), KLF5 (13q22.1), and TERT (5p15.33). In 2020, a large-scale genome-wide association study involving multiple international cohorts identified 5 new susceptibility gene loci, including genes such as NOC2L and HNF1B, further explaining the genetic risk.
[0006] Although these discoveries from multi-modal omics data provide important clues for understanding the pathogenic mechanisms of solid tumors, they have not been fully utilized in the field of drug target identification. Future drug development requires the development of a target prioritization system that combines multi-modal omics integration and functional verification, including traditional omics sequencing technologies, epigenetics analysis, radiomics, pathomics, etc., to accelerate the transformation from multi-modal omics data to therapeutic targets, improve the efficiency of precision medicine, and bring new treatment hopes to solid tumor patients.
[0007] The integration of multi-modal omics data is expected to empower precision oncology research. Integrating multi-modal omics data and associating it with effector genes and related pathways is of great significance for the discovery of therapeutic targets, but this process is full of challenges. First, the pathogenesis of solid tumors is complex and involves multi-modal omics data, including traditional omics data (such as genomics, transcriptomics, proteomics, phosphoproteomics), clinical omics data, radiomics data, pathomics data, etc., making it difficult to effectively identify effector genes. To address these difficulties, a comprehensive drug target prioritization method is needed. On the one hand, it can achieve the efficient integration and full utilization of multi-modal omics data; on the other hand, integrating protein interaction information can better identify potential drug targets associated with known drug targets, promoting the development of multi-modal technologies in the discovery of solid tumor drug targets.
[0008] Existing drug target prioritization methods such as Open Targets integrate data such as eQTL, protein quantitative trait loci (pQTL), and epigenomics to prioritize disease targets; however, this method does not integrate network evidence, that is, protein interaction relationships, and cannot identify target genes that have no genetic evidence but are functionally related. Summary of the Invention
[0009] To solve the above technical problems in the prior art, the first aspect of the present application provides a method for predicting treatment targets of solid tumors based on multi-omics data, which includes:
[0010] Step S1: Input multi-omics data of solid tumors, regard each omics data as a column of predictors, calculate the affinity score, and construct a gene-predictor matrix;
[0011] Step S2: Based on the gene-predictor matrix, use Fisher's method to integrate predictors, rank all input genes by priority, perform functional enrichment analysis on the genes with higher ranking using the KEGG pathway set, and use Z-score and FDR to measure the enrichment situation;
[0012] Step S3: Perform pathway crosstalk network analysis based on the genes ranked by priority, construct a pathway crosstalk network related to the progression of solid tumors, and use the restart random walk algorithm to identify sub-networks enriched with high-score nodes in the pathway crosstalk network;
[0013] Step S4: Based on the information in the drug R & D database, perform drug repurposing analysis on the key target genes in the sub-network enriched with high-score nodes that have been identified, and perform perturbation removal analysis by attacking the nodes of the sub-network enriched with high-score nodes that have been identified individually or jointly to identify potential targets.
[0014] Further, in step S1, the priority index algorithm is used to optimize and integrate the multi-omics data of solid tumors, quantitatively rank these omics data in the context of the protein interaction network, construct multiple types of predictors, and the predictors include (i) mutation predictors (MUT) generated from gene mutation data; (ii) transcriptome predictors (RNA) obtained from transcriptome data analysis; (iii) imaging predictors (RAD) obtained from radiomics data analysis; (iv) pathological predictors (PAT) generated from pathological data, to obtain a prediction matrix with multiple genes and multiple predictors. Then, the affinity scores of multiple predictors are calculated using the random walk restart algorithm, and the supraHex software package is used to visually construct a predictor-specific graph;
[0015] The random walk restart algorithm is:
[0016] (Equation 1)
[0017] represents the restart probability of jumping back to the seed node, is the probability of moving to an adjacent node; A represents the normalized Laplacian adjacency matrix related to the network; is the starting probability vector containing the seed gene scores, which is 0 for non-seed genes; represents the probability vector of the algorithm accessing network nodes at the t-th iteration; represents the probability vector of the algorithm accessing network nodes at the (t + 1)-th iteration;
[0018] The multi-omics data includes traditional omics data, clinical omics data, radiomics data, and pathomics data. The traditional omics data includes genomics data, transcriptomics data, proteomics data, and phosphoproteomics data.
[0019] Further, in step S2, in the gene-predictor matrix, the affinity score of a given predictor is converted to a p-value-like value by the empirical cumulative distribution function, and the empirical cumulative distribution function is:
[0020] (Equation 2)
[0021] where, represents the affinity score of the i-th gene on the j-th predictor, is the converted p-value, and the eCDF is estimated based on all genes;
[0022] Subsequently, for each row of genes, the converted p-values of different predictors are combined by Fisher's combination method with reference to Equations 3 - 5. Target prioritization based on Fisher's combination method:
[0023] (Equation 3)
[0024] (Equation 4)
[0025] (Equation 5)
[0026] In the above formula, represents the converted p-value corresponding to the i-th gene in the j-th predictor, J is the number of predictors, represents the distribution with 2J degrees of freedom, CP i represents the combined p-value of the i-th gene, that is, the value of the eCDF of the chi-square distribution at x i ;
[0027] Finally, the combined p-value is normalized to a priority level from 0 to 5:
[0028] (Equation 6)
[0029] where, CP irepresents the combined p-value of the i-th gene, i.e., the value of the CDF at x i when; PR i represents the priority rating of the i-th gene, i.e., the order among the K genes.
[0030] Further, in step S3, using the dnet software package, pathway crosstalk network analysis is performed using 52 genes ranked in the top 1% of the prioritized gene list to identify gene interaction subsets containing highly prioritized and interconnected genes. The 52 genes include: FN1 (fibronectin 1), THBS2 (thrombospondin 2), PLA2G1B (phospholipase A2 group IB), LSP1 (lymphocyte-specific protein 1), FYB1 (FYN-binding protein 1), RUNX2 (RUNX family transcription factor 2), TP53 (tumor protein p53), KRAS (KRAS proto-oncogene, GTPase), COL1A1 (type I collagen alpha 1 chain), COL3A1 (type III collagen alpha 1 chain), ITGA5 (integrin alpha 5 subunit), CD74 (CD74 molecule), HLA-DRA (major histocompatibility complex, class II, DR alpha), ACTA2 (alpha-2 actin, smooth muscle), THBS1 (thrombospondin 1), ACTB (beta-actin), PLCB2 (phospholipase C beta 2), MMP9 (matrix metalloproteinase 9), SLC2A1 (solute carrier family 2 member 1), HLA-DRB1 (major histocompatibility complex, class II, DR beta 1), ITGA2 (integrin alpha 2 subunit), CALD1 (caldesmon 1), COL1A2 (type I collagen alpha 2 chain), ICAM1 (intercellular adhesion molecule 1), ITGB2 (integrin beta 2 subunit), MYLK (myosin light chain kinase), RAC2 (Rac family small GTPase 2), EGR1 (early growth response 1), GRB2 (growth factor receptor-bound protein 2), PDGFRB (platelet-derived growth factor receptor beta), CREBBP (CREB-binding protein), SMAD4 (SMAD family member 4), CTSB (caspase B), EGF (epidermal growth factor), LEF1 (lymphoid enhancer-binding factor 1), ITGAV (integrin alpha V subunit), CTNNB1 (beta-catenin), ITGA1 (integrin alpha 1 subunit), LCN2 (lipocalin 2), MYL9 (myosin light chain 9), JUN (Jun proto-oncogene, AP-1 transcription factor subunit), ITPR1 (inositol 1,4,5-trisphosphate receptor type 1), TSC1 (TSC complex subunit 1), CSF1R (colony-stimulating factor 1 receptor), IRS1 (insulin receptor substrate 1), HLA-DQB1 (major histocompatibility complex, class II, DQ beta 1), MMP13 (matrix metalloproteinase 13), ITGAM (integrin alpha M subunit), MAPK3 (mitogen-activated protein kinase 3), LAT (linker for activation of T cells), SYK (spleen tyrosine kinase), CD247 (CD247 molecule);From the gene interaction relationships merged from the KEGG EIP pathway, a subset is selected for the identification of pathway intersections. The process of identifying pathway intersections is accomplished by heuristically solving the set prize Steiner tree problem, which mainly includes the following steps:;
[0031] (i) Graph transformation: Combine connected positive nodes into isolated meta-nodes, connect these meta-nodes through a single negative node, and assign positive scores to the meta-nodes and negative scores to the single nodes;
[0032] (ii) Assigning edge weights: In the transformed graph, assign weights to the edges. Define two types of edges, including "single node - single node" edges and "single node - meta node" edges, ensuring that all weights are non-negative. A "single node - single node" edge connects two single nodes, and the absolute sum of the scores of the two nodes normalized by their degrees is used to obtain the edge weight; a "single node - meta node" edge connects a single node and a meta node, and its weight value is the score of the single node normalized by its degree;
[0033] (iii) Minimum spanning tree MST: Use Prim's greedy algorithm to find the MST in the weighted transformed graph, that is, the subgraph that connects all nodes and has the minimum edge weights;
[0034] (iv) Shortest paths between meta-nodes: Find all shortest paths between any pair of meta-nodes in the MST;
[0035] (v) Generating a subgraph: Generate a subgraph that contains the nodes in the shortest paths and the edges between the nodes;
[0036] (vi) Identifying "connecting nodes": "Connecting nodes" are single nodes directly connected to meta-nodes, and their absolute scores are not greater than the sum of the scores of their neighboring meta-nodes;
[0037] (vii) Identifying the MST of "connecting nodes": Find the MST in the "connecting nodes" graph, where the "connecting nodes" graph only contains "connecting nodes" and the edges connecting the "connecting nodes";
[0038] (viii) Optimal path and subgraph nodes: From the "connecting nodes" MST and the connected meta-nodes, find the optimal path. Among all possible paths, find the path with the largest sum of the scores of the nodes and their connected meta-nodes. The nodes in the optimal path and their connected meta-nodes are the subgraph nodes;
[0039] (ix) Extracting the subgraph: Extract the subgraph from the input graph. This subgraph or pathway intersection is the maximum-scoring subgraph with as many positive nodes as possible and as few negative nodes as possible. Here, the input graph only contains the subgraph nodes and the edges connecting the subgraph nodes;
[0040] Perform a degree-preserving node permutation test for 100 iterations, allowing the specification of the number of nodes or genes in the resulting pathway intersections. The required output is obtained through a refined iterative search;
[0041] The identified pathway intersections are represented using a gene network or presented at the pathway level. Here, each pathway is a node, and the edges represent the inferred associations between pathways. Pathways significantly enriched with intersection genes are determined as nodes only through Fisher's exact test. The edges in the network are first inferred based on the member genes shared between pathways and then filtered using the igraph package to identify the minimum spanning tree. Only the edges present in the spanning tree are retained, and the thickness of the edges is adjusted proportionally according to the number of member genes shared between the two endpoint pathways.
[0042] Furthermore, in step S4, (1) the method for drug repurposing analysis is as follows: Based on the intervention target genes identified from the pathway intersection network, use the ChEMBL database to explore evidence of drug repurposing for these genes and discover target genes with clinical trials or approved drugs; the ChEMBL database summarizes the treatment data of approved therapies, and the treatment data includes drugs, development stages, target genes, mechanisms of drug action, and disease indications; use a one-sided Fisher's exact test to evaluate the statistical significance of approved drug targets enriched in pathway intersection genes; (2) the method for perturbation removal analysis is as follows: According to the intervention target genes of the pathway intersection network, perform pathway interference analysis on individual genes and combined genes through targeted interference analysis using network betweenness centrality to select the optimal combination of targeted genes, and then the drug combination corresponding to the optimal combination of targeted genes can be used as a potential drug target for solid tumor treatment.
[0043] The second aspect of this application provides a solid tumor treatment target prediction system based on multi-omics data, which includes:
[0044] A prediction matrix construction module, which is used to input multi-omics data of solid tumors, regard each omics data as a column of predictors, calculate the affinity score, and construct a gene-predictor matrix;
[0045] A target gene prioritization and functional enrichment analysis module, which is used to integrate predictors using Fisher's method based on the gene-predictor matrix, prioritize all input genes, and perform functional enrichment analysis on the genes with higher priorities using the KEGG pathway set, and use the Z-score and FDR to measure the enrichment situation;
[0046] Pathway intersection network construction module, which is used to perform pathway intersection network analysis based on prioritized genes, construct a pathway intersection network related to solid tumor progression, and identify subnetworks enriched with high-score nodes in the pathway intersection network using the weighted random walk algorithm;
[0047] Analysis module, which is used to perform drug repurposing analysis on key target genes in the identified subnetworks enriched with high-score nodes based on drug R & D database information, and perform perturbation removal analysis by means of single or combined attacks on the nodes of the identified subnetworks enriched with high-score nodes to identify potential targets.
[0048] Furthermore, the prediction matrix construction module is used to: optimize and integrate multi-modal omics data of solid tumors using the priority index algorithm, quantitatively rank these omics data in the context of the protein-protein interaction network, construct multiple types of predictors, including (i) mutation predictors (MUT) generated from gene mutation data; (ii) transcriptome predictors (RNA) obtained from transcriptome data analysis; (iii) radiomics predictors (RAD) obtained from radiomics data analysis; (iv) pathology predictors (PAT) generated from pathology data, to obtain a prediction matrix with multiple genes and multiple predictors, then calculate the affinity scores of multiple predictors using the random walk with restart algorithm, and visualize and construct predictor-specific graphs using the supraHex software package;
[0049] The random walk with restart algorithm is:
[0050] (Equation 1)
[0051] represents the restart probability of jumping back to the seed node, is the probability of moving to an adjacent node; A represents the normalized Laplacian adjacency matrix related to the network; is the starting probability vector containing the scores of seed genes, which is 0 for non-seed genes; represents the probability vector of the algorithm accessing network nodes at the t-th iteration; represents the probability vector of the algorithm accessing network nodes at the (t + 1)-th iteration;
[0052] The multi-modal omics data includes traditional omics data, clinical omics data, radiomics data, and pathology omics data, and the traditional omics data includes genomics data, transcriptomics data, proteomics data, and phosphoproteomics data.
[0053] Further, the target gene prioritization and functional enrichment analysis module is used for: in the gene-predictor matrix, the affinity score of a given predictor is converted into a pseudo p-value by the empirical cumulative density function, and the empirical cumulative density function is:
[0054] (Equation 2)
[0055] where, represents the affinity score of the i-th gene on the j-th predictor, is the converted p-value, and eCDF is estimated based on all genes;
[0056] Subsequently, for each row of genes, the converted p-values of different predictors are combined by Fisher's method of combination with reference to Equations 3-5, and the target prioritization based on Fisher's method of combination:
[0057] (Equation 3)
[0058] (Equation 4)
[0059] (Equation 5)
[0060] In the above formula, represents the converted p-value corresponding to the i-th gene in the j-th predictor, J is the number of predictors, represents the distribution with 2J degrees of freedom, CP i represents the combined p-value of the i-th gene, that is, the value of the eCDF of the chi-square distribution at x i ;
[0061] Finally, the combined p-value is normalized to a priority level from 0 to 5:
[0062] (Equation 6)
[0063] where, CP i represents the combined p-value of the i-th gene, that is, the value of the CDF at x i ; PR i represents the priority rating of the i-th gene, that is, the order among K genes.
[0064] Furthermore, the pathway intersection network construction module is used for: using the dnet software package, performing pathway intersection network analysis with the top 1% of 52 genes in the prioritized gene list to identify gene interaction subsets containing highly prioritized and interconnected genes. The 52 genes include: FN1 (fibronectin 1), THBS2 (thrombospondin 2), PLA2G1B (phospholipase A2 group IB), LSP1 (lymphocyte-specific protein 1), FYB1 (FYN-binding protein 1), RUNX2 (RUNX family transcription factor 2), TP53 (tumor protein p53), KRAS (KRAS proto-oncogene, GTPase), COL1A1 (collagen type I alpha 1 chain), COL3A1 (collagen type III alpha 1 chain), ITGA5 (integrin alpha 5 subunit), CD74 (CD74 molecule), HLA-DRA (major histocompatibility complex, class II, DR alpha), ACTA2 (alpha-2 actin, smooth muscle), THBS1 (thrombospondin 1), ACTB (beta-actin), PLCB2 (phospholipase C beta 2), MMP9 (matrix metalloproteinase 9), SLC2A1 (solute carrier family 2 member 1), HLA-DRB1 (major histocompatibility complex, class II, DR beta 1), ITGA2 (integrin alpha 2 subunit), CALD1 (caldesmon 1), COL1A2 (collagen type I alpha 2 chain), ICAM1 (intercellular adhesion molecule 1), ITGB2 (integrin beta 2 subunit), MYLK (myosin light chain kinase), RAC2 (Rac family small GTPase 2), EGR1 (early growth response 1), GRB2 (growth factor receptor-bound protein 2), PDGFRB (platelet-derived growth factor receptor beta), CREBBP (CREB-binding protein), SMAD4 (SMAD family member 4), CTSB (caspase B), EGF (epidermal growth factor), LEF1 (lymphoid enhancer-binding factor 1), ITGAV (integrin alpha V subunit), CTNNB1 (beta-catenin), ITGA1 (integrin alpha 1 subunit), LCN2 (lipocalin 2), MYL9 (myosin light chain 9), JUN (Jun proto-oncogene, AP-1 transcription factor subunit), ITPR1 (inositol 1,4,5-trisphosphate receptor type 1), TSC1 (TSC complex subunit 1), CSF1R (colony-stimulating factor 1 receptor), IRS1 (insulin receptor substrate 1), HLA-DQB1 (major histocompatibility complex, class II, DQ beta 1), MMP13 (matrix metalloproteinase 13), ITGAM (integrin alpha M subunit), MAPK3 (mitogen-activated protein kinase 3), LAT (linker for activation of T cells), SYK (spleen tyrosine kinase), CD247 (CD247 molecule);From the gene interaction relationships merged from the KEGG EIP pathway, a subset is selected for the identification of pathway intersections. The process of identifying pathway intersections is completed by heuristically solving the set prize Steiner tree problem, mainly including the following steps:;
[0065] (i) Graph transformation: Combine connected positive nodes into isolated meta-nodes, connect these meta-nodes through a single negative node, and assign positive scores to the meta-nodes and negative scores to the single nodes;
[0066] (ii) Assign edge weights: In the transformed graph, assign weights to the edges. Define two types of edges, including "single node - single node" edges and "single node - meta node" edges, ensuring that all weights are non-negative. A "single node - single node" edge connects two single nodes, and the absolute sum of the scores of the two nodes normalized by degree is used to obtain the edge weight; a "single node - meta node" edge connects a single node and a meta node, and its weight value is the score of the single node normalized by degree;
[0067] (iii) Minimum spanning tree MST: Use Prim's greedy algorithm to find the MST in the weighted transformed graph, that is, a subgraph that connects all nodes and has the minimum edge weight;
[0068] (iv) Shortest paths between meta-nodes: Find all shortest paths between any pair of meta-nodes in the MST;
[0069] (v) Generate a subgraph: Generate a subgraph containing the nodes in the shortest paths and the edges between the nodes;
[0070] (vi) Identify "connecting nodes": "Connecting nodes" are single nodes directly connected to meta-nodes, and their absolute scores are not greater than the sum of the scores of neighboring meta-nodes;
[0071] (vii) Identify the MST of "connecting nodes": Find the MST in the "connecting nodes" graph, where the "connecting nodes" graph only contains "connecting nodes" and the edges connecting "connecting nodes";
[0072] (viii) Optimal path and subgraph nodes: From the "connecting nodes" MST and the connected meta-nodes, find the optimal path. Among all possible paths, find the path with the largest sum of the scores of the nodes and their connected meta-nodes. The nodes in the optimal path and their connected meta-nodes are subgraph nodes;
[0073] (ix) Extract the subgraph: Extract the subgraph from the input graph. This subgraph or pathway intersection is the maximum scoring subgraph with as many positive nodes as possible and as few negative nodes as possible. Among them, the input graph only contains subgraph nodes and the edges connecting subgraph nodes;
[0074] Perform a 100-iteration degree-preserving node permutation test, allowing the specification of the number of nodes or genes in the resulting pathway intersections, and the required output is obtained through a refined iterative search;
[0075] The identified pathway intersections are represented using a gene network or presented at the pathway level. Here, each pathway is a node, and the edges represent the inferred associations between pathways. Pathways significantly enriched with intersection genes are determined as nodes only through Fisher's exact test. The edges in the network are first inferred based on the member genes shared between pathways and then filtered using the igraph package to identify the minimum spanning tree. Only the edges present in the spanning tree are retained, and the thickness of the edges is adjusted proportionally according to the number of member genes shared between the two endpoint pathways.
[0076] Furthermore, the analysis module is used for: (1) Drug repurposing analysis: Based on the intervention target genes identified from the pathway intersection network, use the ChEMBL database to explore evidence of drug repurposing for these genes and discover target genes with clinical trials or approved drugs; the ChEMBL database summarizes the treatment data of approved therapies, and the treatment data includes drugs, development stages, target genes, mechanisms of drug action, and disease indications; use one-sided Fisher's exact test to evaluate the statistical significance of approved drug targets enriched in pathway intersection genes; (2) Perturbation removal analysis: According to the intervention target genes of the pathway intersection network, perform pathway interference analysis on single genes and combined genes through targeted interference analysis using network betweenness centrality to select the optimal combination of targeted genes, and the drug combination corresponding to the optimal combination of targeted genes can be used as a potential drug target for solid tumor treatment.
[0077] After adopting the above technical solutions, the present application has the following technical effects:
[0078] The present invention not only integrates multi-modal omics datasets for the identification of effector genes but also incorporates protein-protein interaction networks for the identification of peripheral genes, fully leveraging the value of clinical evidence, omics evidence, and network evidence in solid tumor drug target mining. Based on the priority index algorithm, it realizes the integrated analysis of multi-modal omics summary data to prioritize treatment targets, thus facilitating the rapid development of computational medicine for solid tumors.
[0079] The present invention proposes a method for integrating multi-modal omics data analysis to identify solid tumor treatment targets and drug repurposing combinations based on self-developed computational medicine algorithms, aiming to determine treatment candidate drugs, including treatment targets and repurposed drugs, which includes a series of application functions capable of integrating multi-modal omics data, where multi-modal omics data includes traditional omics data (such as genomics, transcriptomics, proteomics, phosphoproteomics), clinical omics data, radiomics data, pathomics data, etc.
[0080] The present invention integrates multi - omics data, and further integrates regulatory genomics, protein - protein interaction networks, and drug - targetable information, realizing a general process for mining therapeutic targets from two perspectives of "clinical evidence" and "omics evidence", with strong systematicity. Users only need to provide a gene set of interest and its ranking information data, and then they can obtain quantitative recommendation results, functional enrichment analysis, therapeutic drug recommendations, and drug repurposing analysis results displayed in the pathway intersection network.
[0081] (1) The present invention integrates multi - omics data, including traditional omics data (such as genomics, transcriptomics, proteomics, phosphoproteomics), clinical omics data, radiomics data, pathomics data, etc. It uses Fisher's method to combine various predictors to generate a gene prediction matrix of affinity scores, associating the genetic loci of solid tumors with candidate target genes. The present invention innovates on the basis of self - developed computational medicine algorithms, fully utilizes the information of multi - omics data, and integrates network interaction relationships to perform unsupervised prediction of therapeutic targets. The prioritized genes identified may serve as potential drug targets.
[0082] (2) Another highlight of the present invention is the development of advanced analysis functions based on the pathway intersection network, including perturbation removal analysis and drug repurposing analysis, which can more effectively illustrate the potential of prioritized genes to become drug targets from multiple levels of information such as network connectivity, disease specificity, molecular pathways, and target gene clusters, and also explore a new method of knowledge interpretation for mining drug targets from multi - omics data.
[0083] (3) The present invention can provide a user - friendly interface and an open - source R package, accompanied by a detailed user manual providing step - by - step instructions, which enables users to easily use and browse, promotes integrated and efficient knowledge discovery, and all analyses are reproducible.
[0084] (4) The present invention has functions of openness, reproducibility, and scalability, and is expected to be flexibly applied to the prediction of multi - omics data targets in various disease fields.
[0085] In summary, as a target prediction tool dominated by multi - omics data and driven by network interaction relationships, the present invention is aimed at a series of diseases of solid tumors, which can help narrow the gap between the genetic loci of solid tumors and the identification of therapeutic targets, thus making decisions for candidate drugs in subsequent pre - clinical verification and clinical trials, and accelerating the development process of computational medicine for identifying therapeutic targets using human disease multi - omics data, functional genomics, and protein - protein interaction networks. Brief Description of the Drawings
[0086] Figure 1 It is a schematic diagram of the principle overview of the method for predicting therapeutic targets of solid tumors based on multi - omics data.
[0087] Figure 2 It is a flowchart of a method for predicting treatment targets for pancreatic cancer.
[0088] Figure 3 It is a graph showing the enrichment analysis results of prioritized target genes. Figure 3 In Figure A of [reference], it is a bubble chart display of enrichment analysis, where the node size is determined by the number of target genes and the color is determined by FDR; Figure 3 In Figure B of [reference], it is a Manhattan plot of prioritized genes in two significantly enriched pathways, showing the priority index (y-axis) of prioritized target genes arranged along the genomic position, and the pathway genes on each chromosome (x-axis) are annotated.
[0089] Figure 4 It is a graph of pathway intersection results for pancreatic cancer. Using pancreatic cancer multi-omics data to identify the pathway intersection network related to pancreatic cancer progression and visualize it, that is, a highly prioritized and interconnected gene network, where nodes are classified according to the type of predictor, and edges come from KEGG interaction data and are colored according to priority.
[0090] Figure 5 It is a kite-like graph of the KEGG pathways enriched by pathway intersection genes. The enrichment significance (FDR) is calculated using one-sided Fisher's exact test. The size of each kite is determined by the number of member genes and is represented by blue dots.
[0091] Figure 6 It is a graph of the analysis results of drug repurposing based on pathway intersection. Figure 6 In Figure A of [reference], it is an enrichment forest plot of pathway intersection genes targeted by approved or phase drugs. The enrichment analysis is performed using one-sided Fisher's exact test, and the significance level (adjp), odds ratio, and 95% confidence interval (represented by a line) are returned. Figure 6 In Figure B of [reference], it is a dot plot showing the pathway intersection genes (y-axis) targeted by currently approved disease drugs. The dots are integers.
[0092] Figure 7 It is a perturbation removal analysis based on pathway intersection. The x-axis represents the node removal situation (represented by the solid circles below), and the y-axis shows the proportion of nodes disconnected individually or jointly. For removing two or three nodes simultaneously, only the removal combination with the greatest impact effect is shown. Inserted on the left is a visualization of the network intersection after removing the corresponding nodes in the pathway, and its layout is the same as that of Figure 4 but only the genes with the best removal effect are labeled.
[0093] Figure 8For the summary of various analyses, the results of the aforementioned target gene function enrichment analysis, drug target enrichment analysis, drug repurposing analysis, etc. (different columns) are displayed through a dot plot, and the most potential treatment targets and repurposed drugs for pancreatic cancer can be located. Detailed implementation manners
[0094] The advantages of the present invention are further elaborated below in conjunction with the accompanying drawings and specific embodiments. Those skilled in the art should understand that the content specifically described below is illustrative rather than restrictive, and should not be used to limit the protection scope of the present invention.
[0095] This embodiment provides a system and method for predicting treatment targets of solid tumors based on multi-modal omics data. The system for predicting treatment targets of solid tumors based on multi-modal omics data includes: a prediction matrix construction module, a target gene prioritization and function enrichment analysis module, a pathway intersection network construction module, and an analysis module.
[0096] The prediction matrix construction module is used to input multi-modal omics data of solid tumors, regard each omics data as a column of predictors, calculate the affinity score, and construct a gene-predictor matrix.
[0097] The target gene prioritization and function enrichment analysis module is used to integrate predictors based on the gene-predictor matrix using Fisher's method, prioritize all input genes, and perform function enrichment analysis on the prioritized genes using the KEGG pathway set, and use the Z-score and FDR to measure the enrichment situation.
[0098] The pathway intersection network construction module is used to perform pathway intersection network analysis based on the prioritized genes, construct a pathway intersection network related to the progression of solid tumors, and use the restart random walk algorithm to identify subnetworks enriched with high-score nodes in the pathway intersection network.
[0099] The analysis module is used to perform drug repurposing analysis on the key target genes in the identified subnetworks enriched with high-score nodes based on drug R & D database information, and perform perturbation removal analysis by attacking the nodes of the identified subnetworks enriched with high-score nodes individually or jointly to identify potential targets.
[0100] As Figure 1 and Figure 2 shown, taking the multi-modal omics data of pancreatic cancer as an example below, the method for predicting treatment targets of solid tumors based on multi-modal omics data includes the following steps S1-S4.
[0101] Step S1: Input multi-modal omics data of solid tumors, regard each omics data as a column of predictors, calculate the affinity score, and construct a gene-predictor matrix.
[0102] The main purpose of this step is data preprocessing to generate a prediction matrix.
[0103] The input file of the present invention is the omics summary data composed of two columns of information. The first column of the summary data is the gene name, and the second column is the gene importance score. The higher the score, the more important the gene. The score can be obtained by statistical analysis, such as obtaining the gene score data of each omics through differential gene analysis. Users input the summary data of each omics according to the characteristics of their own projects. For example, taking the multi-modal omics data of pancreatic cancer as the input, the summary data of each omics can be regarded as a column of predictors, and the predictors of multiple omics are combined to form a prediction matrix.
[0104] The present invention optimizes the previously developed priority index algorithm (R package: Pi, http: / / bioconductor.org / packages / release / bioc / html / Pi.html). The original Pi algorithm takes GWAS data as input, determines the seed genes related to the input disease based on the proximity of the genome to disease-related single nucleotide polymorphisms, chromatin conformation evidence, and expression quantitative trait locus information, then scores the seed genes using ontology knowledge, and then uses network interaction information to identify non-seed genes, and finally constructs a gene-predictor matrix. The prediction matrix generates Pi priority scores (scores range from 0 to 5) and a ranking list. The present invention updates the Pi algorithm, broadens its data analysis structure, makes it face multi-modal omics data, and breaks the limitations of the original GWAS data, including traditional omics data (such as genomics, transcriptomics, proteomics, phosphoproteomics), clinical omics data, radiomics data, pathomics data, etc., increasing the universality of the algorithm.
[0105] The present invention converts the summary data of each omics into different predictors, calculates their gene affinity scores, and constructs a predicted gene-predictor matrix; then ranks all the input genes based on the prediction matrix, and performs functional enrichment analysis on the genes ranked in the top 1% (which can also be customized). Subsequently, the restart random walk algorithm is used to identify the sub-network enriched with high-score nodes in the pathway intersection network; finally, drug repositioning research is carried out on the key target genes in the identified pathway intersection network based on the drug R & D database built in the present invention (including information such as drugs, disease indications, targets, and mechanisms of action). The present invention adds a perturbation removal analysis module based on the pathway intersection network, and uses the method of single / joint attack on the nodes in the network to perform perturbation removal analysis to visualize important nodes.
[0106] Specifically, the present invention takes multi-modal omics data as input, quantifies and ranks these omics data in the context of a protein-protein interaction network to form multiple gene predictors, including: (i) based on gene mutation data, using the weighted values of the top 50 high-frequency mutated genes with the highest mutation frequencies to generate a mutation predictor (MUT); (ii) based on the transcriptome data between solid tumor patients and paired samples, calculating the differential gene expression levels as the transcriptome predictor (RNA), i.e., -log10(FDR) * abs(fold changes); (iii) based on the differential radiomics data between solid tumor patients and paired samples to predict the radiomics predictor (RAD), using a quantification method similar to that in (ii); (iv) based on the differential pathology data between solid tumor patients and paired samples to predict the pathology predictor (PAT), using a quantification method similar to that in (ii). Among them, the protein-protein interaction data is sourced from the STRING database. This network contains approximately 14,000 genes and approximately 201,000 edges. The network nodes are divided into seed nodes (weighted by corresponding scores) and non-seed nodes.
[0107] Restart the random walk algorithm to start a random walk from each seed gene, iteratively traverse neighbor nodes along the network edges. In each iteration, the algorithm faces a choice: either randomly move to an adjacent node or jump back to the seed node. This choice is controlled by the restart probability, enabling the algorithm to not only maintain the connection with the seed gene but also explore peripheral genes located far away (Equation 1). After a finite number of iterations, the algorithm obtains a probability vector in the steady state. This vector stores the affinity scores of each gene relative to the core gene, ranging from 0 to 1, which is used to measure the overall influence exerted by the seed nodes on the network.
[0108] , Equation 1
[0109] In the above formula, represents the restart probability of jumping back to the seed node (correspondingly, is the probability of moving to an adjacent node), A represents the normalized Laplacian adjacency matrix related to the network, is the starting probability vector containing the scores of the seed genes (i.e., the fractions in Equation 2 and Equation 4), which is 0 for non-seed genes, represents the probability vector of the algorithm accessing the network nodes at the t-th iteration. This probability vector represents the steady state. After reaching the steady state, the affinity scores of all nodes in the network are assigned to the seed nodes.
[0110] For different omics data, the vectors containing these affinity scores respectively constitute different predictors. Generally speaking, seed genes are more likely to obtain higher affinity scores than non-seed genes. However, if a peripheral gene has high connectivity with most core genes, it may also obtain a high affinity score, indicating that it is greatly affected by the network.
[0111] The present invention constructs a gene-predictor matrix using the above scoring process, with genes as rows and predictors as columns. In this matrix, the affinity scores can be combined to integrate omics evidence, clinical evidence, and network evidence. The present invention also uses the supraHex software package to visually construct predictor-specific graphs.
[0112] The present invention integrates the pancreatic cancer sequencing data (genomic, transcriptomic, radiomic, and pathological data) previously generated by the research group, quantitatively ranks these omics data in the context of the protein-protein interaction network, and constructs four types of predictors, including (i) mutation predictors (MUT) generated from gene mutation data; (ii) transcriptomic predictors (RNA) obtained from transcriptomic data analysis; (iii) radiomic predictors (RAD) parsed from radiomic data; (iv) pathological predictors (PAT) generated from pathological data, obtaining a prediction matrix of 16,000 genes and four predictors. Then, the affinity scores of the four predictors are calculated using the random walk with restart algorithm, and predictor-specific graphs are visually constructed using the supraHex software package.
[0113] Step S2: Based on the gene-predictor matrix, use Fisher's method to integrate the predictors, rank all input genes by priority, and perform functional enrichment analysis on the genes with higher ranks using the KEGG pathway set, and use the Z-score and FDR to measure the enrichment situation.
[0114] The main purpose of this step is to rank all input genes by priority based on the prediction matrix and perform functional enrichment analysis on the top 1% (which can also be customized) of the ranked genes.
[0115] Based on the gene-predictor matrix obtained above, the present invention uses Fisher's method to integrate the predictors, ranks all input genes by priority, and performs functional enrichment analysis on the top 1% (which can also be customized) of the ranked genes, and uses the Z-score and FDR (one-sided Fisher's exact test) to measure the enrichment situation. In the gene-predictor matrix, the affinity scores of a given predictor are converted into pseudo p-values by the empirical cumulative distribution function (eCDF) (Equation 2), and the eCDF is estimated from the affinity scores of all genes on this predictor.
[0116] , Equation 2
[0117] Here, represents the affinity score of the i-th gene on the j-th predictor, is the transformed p-value, and eCDF is estimated based on all genes.
[0118] Subsequently, for each row of genes, the transformed p-values of different predictors are combined by Fisher's method referring to Equations 3-5. Target prioritization based on Fisher's method:
[0119] , Equation 3
[0120] , Equation 4
[0121] , Equation 5
[0122] In the above formula, represents the transformed p-value corresponding to the i-th gene in the j-th predictor, J is the number of predictors, represents the distribution with 2J degrees of freedom, CP i represents the combined p-value of the i-th gene (i.e., the value of the eCDF of the chi-square distribution at x i ).
[0123] Finally, the combined p-value is normalized to a priority level from 0 to 5 (Equation 6).
[0124] , Equation 6
[0125] where CP i represents the combined p-value of the i-th gene (i.e., the value of the CDF at x i ), and PR i represents the priority rating of the i-th gene (the order among K genes).
[0126] Based on the gene-predictor matrix, the present invention uses Fisher's method to integrate predictors and predicts the prioritized genes of about 16,000 pancreatic cancer multi-omics. Enrichment analysis is performed on the top 1% of the prioritized genes using the KEGG pathway set, as Figure 3 shown, and it is found that the top-ranked genes are all enriched in pathways related to pancreatic cancer mechanisms and treatment targets, such as pancreatic secretion and focal adhesion. This result indicates that the optimized priority index algorithm can support work for disease mechanism and treatment target discovery.
[0127] Step S3: Based on the prioritized genes, perform pathway crosstalk network analysis to construct a pathway crosstalk network related to solid tumor progression, and use the restart random walk algorithm to identify subnetworks enriched with high-score nodes in the pathway crosstalk network.
[0128] The main purpose of this step is to construct a pathway crosstalk network based on the prioritized genes.
[0129] The present invention uses the algorithm originally proposed in the dnet software package to perform pathway crosstalk network analysis to identify subsets of gene interactions (merged from KEGG pathways) containing highly prioritized and interconnected genes. To focus on molecular interactions involved in signal transduction and independent of specific species, cells, or diseases, the present invention selects subsets from the gene interaction relationships merged from the KEGG EIP pathway for the identification of pathway crosstalk, which contain a total of 2158 genes and 15375 interaction relationships, and each interaction relationship or edge in the network exists in at least one pathway.
[0130] The process of identifying pathway crosstalk is completed by heuristically solving the set prize Steiner tree problem, mainly including the following steps:
[0131] (i) Graph transformation: Combine connected positive nodes into isolated meta-nodes, connect these meta-nodes through a single negative node, and assign positive scores to the meta-nodes and negative scores to the single nodes.
[0132] (ii) Assign edge weights: In the transformed graph, assign weights to the edges, define two types of edges, including single-single edges (connecting two single nodes, and the absolute sum of the scores of the two nodes normalized by degree is used to obtain the edge weight) and single-meta edges (connecting a single node and a meta-node, and its weight value is the score of the single node normalized by degree), ensuring that all weights are non-negative.
[0133] (iii) Minimum spanning tree (MST): Use Prim's greedy algorithm to find the MST in the weighted transformed graph, that is, a subgraph that connects all nodes and has the minimum edge weight.
[0134] (iv) Shortest paths between meta-nodes: Find all shortest paths between any pair of meta-nodes in the MST.
[0135] (v) Generate a subgraph: Generate a subgraph containing the nodes and the edges between the nodes in the shortest paths.
[0136] (vi) Identify linkers: Linkers are single nodes directly connected to meta-nodes, and their absolute scores are not greater than the sum of the scores of neighboring meta-nodes.
[0137] (vii) Identify the MST of the linker: Find the MST in the linker graph (only including the linker and the edges connecting the linkers).
[0138] (viii) Optimal path and subgraph nodes: From the linker MST and the connected meta-nodes, find the optimal path. Among all possible paths, find the path with the largest sum of scores of the nodes and their connected meta-nodes. The nodes within the optimal path and their connected meta-nodes are "subgraph nodes".
[0139] (ix) Extract the subgraph: Extract the subgraph from the input graph (only including the subgraph nodes and the edges connecting the subgraph nodes). This subgraph (or "pathway intersection") is the maximum-scoring subgraph, having as many positive nodes and as few negative nodes as possible.
[0140] To evaluate the statistical significance (p-value) of the identified pathway intersections, the present invention performs a degree-preserving node permutation test with 100 iterations. In addition, this analysis allows specifying the number of nodes or genes in the resulting pathway intersections, and the required output is obtained through a refined iterative search.
[0141] The identified pathway intersections can be represented not only by gene networks but also presented at the pathway level, where each node is a pathway and the edges represent the inferred associations between pathways. Pathways significantly enriched with intersection genes are determined as nodes only through Fisher's exact test. The edges in the network are first inferred based on the member genes shared between pathways and then filtered using the igraph package to identify the minimum spanning tree. Only the edges present in the spanning tree are retained, and the thickness of the edges is adjusted proportionally according to the number of member genes shared between the two endpoint pathways.
[0142] The present invention uses the top 1% of genes in the prioritized gene list to perform pathway intersection network analysis to identify the pathway intersection network related to pancreatic cancer disease progression, so as to determine potential therapeutic intervention targets, such as Figure 4As shown, the results show a pathway convergence network composed of 52 highly confident and interconnected genes, which is also the final result list according to the prioritization strategy. The 52 genes include: FN1 (fibronectin 1), THBS2 (thrombospondin 2), PLA2G1B (phospholipase A2 group IB), LSP1 (lymphocyte-specific protein 1), FYB1 (FYN-binding protein 1), RUNX2 (RUNX family transcription factor 2), TP53 (tumor protein p53), KRAS (KRAS proto-oncogene, GTPase), COL1A1 (type I collagen alpha 1 chain), COL3A1 (type III collagen alpha 1 chain), ITGA5 (integrin alpha 5 subunit), CD74 (CD74 molecule), HLA-DRA (major histocompatibility complex, class II, DR alpha), ACTA2 (alpha 2 actin, smooth muscle), THBS1 (thrombospondin 1), ACTB (beta actin), PLCB2 (phospholipase C beta 2), MMP9 (matrix metalloproteinase 9), SLC2A1 (solute carrier family 2 member 1), HLA-DRB1 (major histocompatibility complex, class II, DR beta 1), ITGA2 (integrin alpha 2 subunit), CALD1 (calponin 1), COL1A2 (type I collagen alpha 2 chain), ICAM1 (intercellular adhesion molecule 1), ITGB2 (integrin beta 2 subunit), MYLK (myosin light chain kinase), RAC2 (Rac family small GTPase 2), EGR1 (early growth response 1), GRB2 (growth factor receptor-bound protein 2), PDGFRB (platelet-derived growth factor receptor beta), CREBBP (CREB-binding protein), SMAD4 (SMAD family member 4), CTSB (caspase B), EGF (epidermal growth factor), LEF1 (lymphoid enhancer-binding factor 1), ITGAV (integrin alpha V subunit), CTNNB1 (beta-catenin), ITGA1 (integrin alpha 1 subunit), LCN2 (lipocalin 2), MYL9 (myosin light chain 9), JUN (Jun proto-oncogene, AP-1 transcription factor subunit), ITPR1 (inositol 1,4,5-trisphosphate receptor type 1), TSC1 (TSC complex subunit 1), CSF1R (colony-stimulating factor 1 receptor), IRS1 (insulin receptor substrate 1), HLA-DQB1 (major histocompatibility complex, class II, DQ beta 1), MMP13 (matrix metalloproteinase 13), ITGAM (integrin alpha M subunit), MAPK3 (mitogen-activated protein kinase 3), LAT (linker for activation of T cells), SYK (spleen tyrosine kinase), CD247 (CD247 molecule).
[0143] Next, as Figure 5As shown, the present invention explores the pathways significantly enriched with the intersection genes, so as to conduct functional characterization and display of the identified intersections at the pathway level. Most of the intersection genes in this pathway intersection network are functionally related to the PI3K-AKT signaling pathway (COL1A1 / 2, CSF1R, EGF, FN1, GRB2, IRS1, ITGA1 / 2 / 5, ITGAV, KRAS, MAPK3, PDGFRB, SYK, THBS1 / 2, TP53 and TSC1) and the Rap1 signaling pathway (ACTB, CSF1R, CTNNB1, EGF, FYB1, ITGAM, ITGB2, KRAS, LAT, MAPK3, PDGFRB, PLCB2, RAC2 and THBS1), among which 6 genes (CSF1R, EGF, KRAS, MAPK3, PDGFRB and THBS1) interact between these two signaling pathways. The research results based on pathway intersection highlight the importance of the PI3K-AKT signaling pathway and the Rap1 signaling pathway in the treatment of pancreatic cancer.
[0144] Step S4: Based on the information in the drug R & D database, conduct drug repurposing analysis on the key target genes in the identified sub-network enriched with high-score nodes, and adopt the method of single or combined attack on the nodes of the identified sub-network enriched with high-score nodes for perturbation removal analysis to identify potential targets.
[0145] The main purposes of this step are drug quantitative recommendation, drug repurposing analysis and perturbation removal analysis.
[0146] The present invention conducts drug repurposing research on the key target genes in the sub-network based on the information in the drug R & D database built into the software (including information such as drugs, disease indications, targets and mechanisms of action), and adopts the method of single / combined attack on the nodes in the network for perturbation removal analysis of drugs to identify potential targets.
[0147] (i) Drug repurposing analysis based on pathway intersection.
[0148] Drug repurposing analysis relies on the information extracted from ChEMBL, which summarizes the treatment data of approved therapies (including drugs, development stages, target genes, mechanisms of drug action and disease indications). For a specific disease indication, only the target genes with clear mechanisms of action and capable of explaining the efficacy of drug treatment for the disease will be selected as drug targets (i.e., the genes targeted by any approved drug). We define the genes targeted by drugs in development stages greater than or equal to Phase II as "clinically theoretically supported targets", indicating that there is sufficient evidence to prove that these genes can be effectively used for the treatment of solid tumors. The one-sided Fisher's exact test is used to evaluate the statistical significance of the enrichment of approved drug targets to pathway intersection genes.
[0149] As Figure 6 and Figure 8 shown, the present invention further focuses on whether the pathway intersection genes are targeted by drugs approved for treating diseases other than pancreatic cancer, in order to search for the possibility of drug repurposing. Based on the intervention target genes identified by the pathway intersection network, the ChEMBL database is used to explore the evidence of drug repurposing for these genes, and it is found that 8 genes are target genes with clinical trials or approved drugs. Notably, among these 8 approved drug targets, 6 genes (FN1, KRAS, PDGFRB, ITGAV, CSF1R, SYK) are important members of the PI3K-AKT signaling pathway, highlighting the potential of drug repurposing targeting this pathway.
[0150] (ii) Perturbation removal analysis based on pathway intersection.
[0151] By using the targeted interference analysis of network betweenness centrality, the intervenability of the genes in the pathway intersection network is evaluated to select the optimal targeting combination to maximize the impact of removing nodes (alone or in combination) on pathway interactions. The scope of the perturbation removal analysis aims to evaluate the impact of nodes on pathway intersection. This analysis includes two dimensions: (i) removing nodes alone (i.e., single-node removal), and (ii) removing combined nodes simultaneously (i.e., combined removal). If the specified removed nodes are crucial to the network, their removal will cause most nodes to disconnect from the main components in the network. On the other hand, we explore the optimal combination of target removal such that removing this node combination (e.g., removing two nodes simultaneously) maximizes the effect on the network, i.e., having the highest proportion of disconnected nodes. We visually display the effects of single-node or multi-node removal through the upset plot of the ggupset package.
[0152] The perturbation removal analysis based on pathway intersection can be used to explore other potential therapeutic targets, aiming to identify the key genes that have the greatest impact on pathway intersection. The perturbation removal analysis can measure the vulnerability of pathway intersection to the removal of nodes alone or in combination. Removing key network nodes will cause most nodes in the pathway intersection to disconnect. According to the intervention target genes of the pathway intersection network, pathway interference analysis is performed on single genes and combined genes. As Figure 7 and Figure 8 shown, the results indicate that removing a single gene does not have strong robustness to the stability of the entire intersection network, while interfering with the combined nodes, the drug-binding energy of SYK and ITGAV can produce the greatest interference effect (about 38.5% of the nodes disconnect), and this result indicates that the drug combination corresponding to these two genes can be used as potential drug targets for the treatment of pancreatic cancer.
[0153] It should be noted that the embodiments of the present invention have better implementability and do not impose any form of limitation on the present invention. Any person skilled in the art may use the technical content disclosed above to modify or transform it into equivalent effective embodiments. However, as long as it does not depart from the technical solution of the present invention, any modification, equivalent change or modification made to the above embodiments based on the technical essence of the present invention still falls within the scope of the technical solution of the present invention.
Claims
1. A method for predicting therapeutic targets of solid tumors based on multimodal omics data, characterized in that: include: Step S1: Input the multimodal omics data of solid tumors, treat each omics data as a column of predictors, calculate the affinity score, and construct a gene-predictor matrix; Step S2: Based on the gene-predictor matrix, Fisher's method was used to integrate predictors, prioritize all input genes, and the KEGG pathway set was used to perform functional enrichment analysis on the prioritized genes, using Z score and FDR to measure enrichment; Step S3: Pathway intersection network analysis is performed based on the prioritized genes to construct a pathway intersection network related to solid tumor progression, and a restarted random walk algorithm is used to identify subnetworks enriched with high-score nodes in the pathway intersection network; Step S4: Based on the information of the drug development database, drug reuse analysis is performed on the key target genes in the identified sub-networks enriched with high-score nodes, and disturbance removal analysis is performed by single or joint attack on the identified sub-network nodes enriched with high-score nodes to identify potential targets.
2. The method for predicting solid tumor therapeutic targets based on multimodal omics data according to claim 1, characterized in that: In step S1, the optimized priority index algorithm is used to integrate the multimodal omics data of solid tumors, and these omics data are quantitatively ranked in the context of the protein interaction network to construct various types of predictors, including (i) mutation predictors (MUT) generated by gene mutation data; (ii) transcriptome predictors (RNA) obtained by transcriptome data analysis; (iii) imaging predictors (RAD) obtained by imaging data analysis; (iv) pathology predictors (PAT) generated by pathology data. A prediction matrix consisting of multiple genes and four predictors is obtained, and then the random walk restart algorithm is used to calculate the affinity scores of multiple predictors, and the supraHex software package is used to visualize and construct predictor-specific graphs; The random walk restart algorithm is: (Equation 1) represents the restart probability of jumping back to the seed node, is the probability of moving to a neighboring node; A represents the standardized Laplacian adjacency matrix associated with the network; is the starting probability vector containing the scores of seed genes and is 0 for non-seed genes; Represents the probability vector of the algorithm visiting the network node at the tth iteration; Represents the probability vector of the algorithm visiting the network node at iteration t+1; The multimodal omics data include traditional omics data, clinical omics data, imaging omics data and pathological omics data, and the traditional omics data include genomics data, transcriptomics data, proteomics data and phosphorylation omics data.
3. The method for predicting solid tumor therapeutic targets based on multimodal omics data according to claim 1, characterized in that: In step S2, the affinity score of a given predictor in the gene-predictor matrix is converted to a class p-value using the empirical cumulative density function: (Equation 2) in, represents the affinity score of the ith gene on the jth predictor, is the transformed p-value, and the eCDF is estimated based on all genes; Then, for each row of genes, the transformed p-values of different predictors were combined by Fisher's combination method according to Equations 3 to 5, and the targets were prioritized based on Fisher's combination method: (Equation 3) (Equation 4) (Equation 5) In the above formula, represents the transformed p-value corresponding to the i-th gene in the j-th predictor, where J is the number of predictors, Represents a degree of freedom of 2J Distribution, CP i represents the combined p-value of the ith gene, i.e., the eCDF of the chi-square distribution at x i The value at time; Finally, the combined p-values were normalized to a priority scale from 0 to 5: (Equation 6) Among them, CP i represents the combined p-value of the ith gene, i.e., CDF at x i The value of PR i represents the priority rating of the ith gene, that is, its order among the K genes.
4. The method for predicting solid tumor therapeutic targets based on multimodal omics data according to claim 1, characterized in that: In step S3, pathway intersection network analysis was performed using the dnet software package using the top 1% of the prioritized gene list to identify a subset of gene interactions containing highly prioritized and interconnected genes: FN1, THBS2, PLA2G1B, LSP1, FYB1, RUNX2, TP53, KRAS, COL1A1, COL3A1, ITGA5, CD74, HLA-DRA, ACTA2, THBS1, ACTB, PLCB2, MMP9, SLC2A 1. HLA-DRB1, ITGA2, CALD1, COL1A2, ICAM1, ITGB2, MYLK, RAC2, EGR1, GRB2, PDGFRB, CREBBP, SMAD4, CTSB, EGF, LEF1, ITGAV, CTNNB1, ITGA1, LCN2, MYL9, JUN, ITPR1, TSC1, CSF1R, IRS1, HLA-DQB1, MMP13, ITGAM, MAPK3, LAT, SYK, CD247; From the gene interaction relationship merged by KEGG EIP pathway, a subset was selected for the identification of pathway intersections. The process of identifying pathway intersections was completed by heuristically solving the set-prize Steiner tree problem, which mainly included the following steps: (i) Graph transformation: combine connected positive nodes into isolated meta-nodes, connect these meta-nodes through a single negative node, and assign positive scores to meta-nodes and negative scores to single nodes; (ii) Assigning edge weights: In the converted graph, weights are assigned to edges. Two types of edges are defined, including "single node-single node" edges and "single node-meta node" edges. All weights are ensured to be non-negative. "Single node-single node" edges connect two single nodes. The absolute sum of the scores of the two nodes normalized by degree is used to obtain the edge weight. "Single node-meta node" edges connect a single node and a meta node. The weight value is the score of the single node normalized by degree. (iii) Minimum spanning tree (MST): Use Prim's greedy algorithm to find the MST in the weighted transition graph, i.e., the subgraph connecting all nodes with the minimum edge weights; (iv) Shortest paths between meta-nodes: Find all shortest paths between any pair of meta-nodes in the MST; (v) Generate subgraph: Generate a subgraph containing the nodes in the shortest path and the edges between the nodes; (vi) Identify "connected nodes": A "connected node" is a single node that is directly connected to a meta-node and whose absolute score is no greater than the sum of the scores of its neighboring meta-nodes; (vii) Identify the MST of a "connected node": Find the MST in a "connected node" graph, which only contains "connected nodes" and the edges connecting "connected nodes"; (viii) Optimal path and subgraph nodes: Find the optimal path from the "connected nodes" MST and the connected meta-nodes. Among all possible paths, find the path with the largest sum of scores of the node and its connected meta-nodes. The nodes in the optimal path and its connected meta-nodes are subgraph nodes. (ix) Extract subgraph: Extract a subgraph from the input graph. This subgraph or path intersection is the maximal score subgraph with as many positive nodes as possible and as few negative nodes as possible, where the input graph only contains subgraph nodes and edges connecting subgraph nodes. A degree-preserving node permutation test was performed for 100 iterations, allowing the specification of the number of nodes or genes in the resulting pathway intersections. The required output was obtained through a perfect iterative search; the identified pathway intersections were represented using a gene network representation or displayed based on the pathway level, where pathways were each node and edges represented the inferred associations between pathways. Pathways that were significantly enriched in intersection genes were identified as nodes only if they passed Fisher's exact test. The edges in the network were first inferred based on the member genes shared between pathways, and then filtered using the igraph package to identify the minimum spanning tree. This process only retained the edges present in the spanning tree, and the thickness of the edges was adjusted proportionally according to the number of shared member genes between the two endpoint pathways.
5. The method for predicting solid tumor therapeutic targets based on multimodal omics data according to claim 1, characterized in that: In step S4, (1) the method of drug reuse analysis is: based on the intervention target genes identified by the pathway intersection network, the ChEMBL database is used to explore the evidence of drug reuse of these genes and find target genes with clinical trials or approved drugs; the ChEMBL database summarizes the treatment data of approved therapies, and the treatment data includes drugs, development stages, target genes, drug mechanisms of action and disease indications; a one-sided Fisher's exact test is used to evaluate the statistical significance of the enrichment of approved drug targets to pathway intersection genes; (2) the method of perturbation removal analysis is: based on the intervention target genes of the pathway intersection network, a pathway interference analysis is performed on single genes and combined genes by using the targeted interference analysis of the network betweenness centrality to select the optimal targeted gene combination, and the drug combination corresponding to the optimal targeted gene combination can be used as a potential drug target for the treatment of solid tumors.
6. A solid tumor therapeutic target prediction system based on multimodal omics data, characterized in that: include: A prediction matrix construction module, wherein the prediction matrix construction module is used to input multimodal omics data of solid tumors, regard each omics data as a column of predictors, calculate affinity scores, and construct a gene-predictor matrix; A target gene prioritization and functional enrichment analysis module, which is used for a gene-predictor matrix, uses Fisher's method to integrate predictors, prioritizes all input genes, and performs functional enrichment analysis on prioritized genes based on the KEGG pathway dataset, using Z-scores and FDR to measure enrichment; A pathway intersection network construction module, wherein the pathway intersection network construction module is used to perform pathway intersection network analysis based on the prioritized genes, construct a pathway intersection network related to solid tumor progression, and use a restarted random walk algorithm to identify a subnetwork enriched with high-score nodes in the pathway intersection network; An analysis module is used to perform drug reuse analysis on key target genes in an identified sub-network enriched with high-score nodes based on drug development database information, and to perform disturbance removal analysis by attacking the identified sub-network nodes enriched with high-score nodes individually or in combination to identify potential targets.
7. The solid tumor therapeutic target prediction system based on multimodal omics data according to claim 6, characterized in that: The prediction matrix construction module is used to: integrate multimodal omics data of solid tumors using an optimized priority index algorithm, quantitatively rank these omics data in the context of a protein interaction network, and construct multiple types of predictors, including (i) mutation predictors (MUT) generated by gene mutation data; (ii) transcriptome predictors (RNA) obtained by transcriptome data analysis; (iii) imaging predictors (RAD) obtained by imaging data analysis; (iv) pathology predictors (PAT) generated by pathology data, and obtain a prediction matrix with multiple genes and multiple predictors, then use a random walk restart algorithm to calculate the affinity scores of multiple predictors, and use the supraHex software package to visualize and construct predictor-specific graphs; The random walk restart algorithm is: (Equation 1) represents the restart probability of jumping back to the seed node, is the probability of moving to a neighboring node; A represents the standardized Laplacian adjacency matrix associated with the network; is the starting probability vector containing the scores of seed genes and is 0 for non-seed genes; Represents the probability vector of the algorithm visiting the network node at the tth iteration; Represents the probability vector of the algorithm visiting the network node at iteration t+1; The multimodal omics data include traditional omics data, clinical omics data, imaging omics data and pathological omics data, and the traditional omics data include genomics data, transcriptomics data, proteomics data and phosphorylation omics data.
8. The solid tumor therapeutic target prediction system based on multimodal omics data according to claim 6, characterized in that: The target gene priority sorting and functional enrichment analysis module is used to convert the affinity score of a given predictor into a p-value in the gene-predictor matrix by the empirical cumulative density function, where the empirical cumulative density function is: (Equation 2) in, represents the affinity score of the ith gene on the jth predictor, is the transformed p-value, and the eCDF is estimated based on all genes; Then, for each row of genes, the transformed p-values of different predictors were combined by Fisher's combination method according to Equations 3 to 5, and the targets were prioritized based on Fisher's combination method: (Equation 3) (Equation 4) (Equation 5) In the above formula, represents the transformed p-value corresponding to the i-th gene in the j-th predictor, where J is the number of predictors, Represents a degree of freedom of 2J Distribution, CP i represents the combined p-value of the ith gene, i.e., the eCDF of the chi-square distribution at x i The value at time; Finally, the combined p-values are normalized and scaled to a priority level from 0 to 5: (Equation 6) Among them, CP i represents the combined p-value of the ith gene, i.e., CDF at x i The value of PR i represents the priority rating of the ith gene, that is, its order among the K genes.
9. The solid tumor therapeutic target prediction system based on multimodal omics data according to claim 6, characterized in that: The pathway intersection network building module was used to perform pathway intersection network analysis using the top 1% of the prioritized gene list using the dnet software package to identify a subset of gene interactions containing highly prioritized and interconnected genes, including: FN1, THBS2, PLA2G1B, LSP1, FYB1, RUNX2, TP53, KRAS, COL1A1, COL3A1, ITGA5, CD74, HLA-DRA, ACTA2, THBS1, ACTB, PLCB2, MMP9, S LC2A1, HLA-DRB1, ITGA2, CALD1, COL1A2, ICAM1, ITGB2, MYLK, RAC2, EGR1, GRB2, PDGFRB, CREBBP, SMAD4, CTSB, EGF, LEF1, ITGAV, CTNNB1, ITGA1, LCN2, MYL9, JUN, ITPR1, TSC1, CSF1R, IRS1, HLA-DQB1, MMP13, ITGAM, MAPK3, LAT, SYK, CD247; From the gene interaction relationship merged by KEGG EIP pathway, a subset was selected for the identification of pathway intersections. The process of identifying pathway intersections was completed by heuristically solving the set-prize Steiner tree problem, which mainly included the following steps: (i) Graph transformation: combine connected positive nodes into isolated meta-nodes, connect these meta-nodes through a single negative node, and assign positive scores to meta-nodes and negative scores to single nodes; (ii) Assigning edge weights: In the converted graph, weights are assigned to edges. Two types of edges are defined, including "single node-single node" edges and "single node-meta node" edges. All weights are ensured to be non-negative. "Single node-single node" edges connect two single nodes. The absolute sum of the scores of the two nodes normalized by degree is used to obtain the edge weight. "Single node-meta node" edges connect a single node and a meta node. The weight value is the score of the single node normalized by degree. (iii) Minimum spanning tree (MST): Use Prim's greedy algorithm to find the MST in the weighted transition graph, i.e., the subgraph connecting all nodes with the minimum edge weights; (iv) Shortest paths between meta-nodes: Find all shortest paths between any pair of meta-nodes in the MST; (v) Generate subgraph: Generate a subgraph containing the nodes in the shortest path and the edges between the nodes; (vi) Identify "connected nodes": A "connected node" is a single node that is directly connected to a meta-node and whose absolute score is no greater than the sum of the scores of its neighboring meta-nodes; (vii) Identify the MST of a "connected node": Find the MST in a "connected node" graph, which only contains "connected nodes" and the edges connecting "connected nodes"; (viii) Optimal path and subgraph nodes: Find the optimal path from the "connected nodes" MST and the connected meta-nodes. Among all possible paths, find the path with the largest sum of scores of the node and its connected meta-nodes. The nodes in the optimal path and its connected meta-nodes are subgraph nodes. (ix) Extract subgraph: extract a subgraph from the input graph, which is a subgraph or a path intersection with the highest score, as many positive nodes as possible and as few negative nodes as possible, where the input graph only contains subgraph nodes and edges connecting subgraph nodes; A degree-preserving node permutation test was performed for 100 iterations, allowing specification of the number of nodes or genes in the resulting pathway intersection, and the desired output was obtained through a refined iterative search; The identified pathway intersections are represented using gene networks or displayed based on the pathway level, where each pathway is a node and the edges represent the inferred associations between pathways. Pathways that are significantly enriched in intersection genes are identified as nodes only if they pass Fisher's exact test. The edges in the network are first inferred based on the member genes shared between pathways, and then filtered using the igraph package to identify the minimum spanning tree. This process only retains the edges present in the spanning tree, and the thickness of the edges is adjusted proportionally according to the number of shared member genes between the two endpoint pathways.
10. The solid tumor therapeutic target prediction system based on multimodal omics data according to claim 6, characterized in that: The analysis module is used for: (1) drug reuse analysis: based on the intervention target genes identified by the pathway intersection network, the ChEMBL database is used to explore the evidence of drug reuse of these genes and discover target genes with clinical trials or approved drugs; the ChEMBL database summarizes the treatment data of approved therapies, which include drugs, development stages, target genes, drug action mechanisms and disease indications; a one-sided Fisher's exact test is used to evaluate the statistical significance of the enrichment of approved drug targets to pathway intersection genes; (2) perturbation removal analysis: based on the intervention target genes of the pathway intersection network, pathway interference analysis is performed on single genes and combined genes by using targeted interference analysis of network betweenness centrality to select the optimal targeted gene combination. The drug combination corresponding to the optimal targeted gene combination can be used as a potential drug target for the treatment of solid tumors.
Citation Information
Patent Citations
Cross-disease treatment target prediction method, electronic equipment, information platform, storage medium and program product
CN118098341A
Network biology approach for identifying targets for combination therapies
US20110119259A1
Method of discovering novel anticancer drug using co-essentiality network, and an apparatus thereof
US20240212786A1
Cited By
Targeted senile degenerative bone disease key lesion regulation factor mRNA therapy recommendation evaluation method, electronic equipment and program product
CN121483362A