Systems and methods for target or biomarker identification
Patent Information
- Application Number
- JP2026512719
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-29
- Filing Date
- 2024-08-27
- Publication Date
- 2026-09-08
Smart Images

Figure 2026530467000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure is directed toward a computer implementation method for identifying disease targets or biomarkers. In particular, this disclosure is directed toward identifying disease targets or biomarkers using causal analysis, but is not limited thereto. In particular, this disclosure is directed toward identifying disease targets or biomarkers using causal analysis and a knowledge graph, but is not limited thereto. [Background technology]
[0002] Identifying cell type-specific contributions to specific diseases is difficult, especially since patient populations can be highly heterogeneous due to differences in, for example, genetic contributions to the disease, the way symptoms manifest, and the progression of the disease.
[0003] Many groups have attempted to elucidate the cell type-specific genetic contribution to disease (or other phenotypes), but have been unable to assess causal relationships between cell type and disease characteristics, and / or have relied on cell type prediction.
[0004] One approach to elucidating cell type-specific genetic contributions to disease is to use quantitative trait loci (QTLs) of a single cell type as exposure variables for analysis, using methods such as Mendelian randomization. For example, Dang et al., Mov. Disord., 2022;37(12):2451-2456 performed a two-sample Mendelian randomization (2SMR) analysis to identify causative genes in Parkinson's disease using dopaminergic neuron-specific expression quantitative trait loci (eQTLs).
[0005] Several studies have placed 2SMR results (generated using bulk omics data as exposure variables) within secondary cell type-specific datasets. For example, McGowan et al., Human Molecular Genetics, 28, 2019, 3293-3300, performed 2SMR on circulating cytokines and immune-mediated diseases, then examined gene expression in three immune cell type datasets (i.e., monocytes, neutrophils, and T cells) to determine whether the identified causal associations were attributable to a single cell type.
[0006] Other studies have used bulk omics data to deconvolve data and predict cell type-specific contributions. For example, Westra et al., PLoS Genet 11(5):e1005223 (2015) performed a gene-environment interaction meta-analysis to infer cell type-specific contributions derived from whole blood eQTLs.
[0007] Numerous studies, including Richardson et al., Nat. Commun. 11, 185 (2020), have evaluated multiple tissue-level exposure variables in relation to disease / trait outcomes using 2SMR analysis, which performed 2SMR on 48 tissue-specific eQTLs and 395 complex traits. Taylor et al., Genome Med. 11, 6 (2019) performed 2SMR on 44 tissue-specific eQTLs and 14 cardiovascular traits. Porcu et al., Nat. Commun. 10, 3300 (2019) developed a transcriptome-wide summary statistics-based MR approach using multiple SNPs as instrumental variables and multiple gene expression traits as exposures simultaneously, applying this method to 48 tissues and 43 traits.
[0008] Cellular specificity and function in healthy and diseased states are determined by interactions between different genes or proteins at the levels of transcription, translation, and post-translational modification. Due to the high degree of interdependence between genes, diseases rarely arise from single-gene dysfunction; they usually reflect larger abnormalities in normally interconnected genes. This logic, combined with the vast accumulation of biological knowledge, makes the use of knowledge graphs and biological networks in discovering new disease genes, modules, and ultimately targets very appealing. For example, Jensen et al., Circulation: Cardiovascular Genetics. 2011;4:549-556 identified novel susceptibility genes for coronary artery disease by analyzing neighboring genes of genes that significantly hit in genome-wide association studies (GWAS) via protein-protein interaction networks. Vlietstra et al., PLoS ONE 17(7):e0271395(2022) used a protein knowledge graph to identify genes targeted by disease-associated non-coding SNPs. Himmelstein et al., PLoS Comput. Biol., 11(7):e1004259. (2015) and Binder et al., Commun. Biol., 5, 125 (2022) used knowledge graphs for prioritizing and predicting novel disease-related genes, respectively.
[0009] The development of single-cell (scRNAseq) and single-nucleus (snRNAseq) RNA sequencing technologies has enabled the identification of cell-type-specific disease genes. However, the development of computational methods involving these types of datasets has been limited to group comparisons of individual datasets based on state-of-the-art statistical methods, without considering extensions to include data integration or biological networks.
[0010] Numerous experimental and / or computational methods exist to analyze cell type-specific contributions to disease phenotypes, including single-cell gene expression profiling (e.g., Human Cell Atlas), multiplex imaging / spatial transcriptomics (e.g., Allen Brain Atlas-Aging, Dementia, and TBI Study), or cell deconvolution (e.g., cell fractionation of AML samples). [Overview of the initiative]
[0011] A first aspect of this disclosure provides a method for identifying key targets or biomarkers for a disease. The method includes: acquiring data on genetic risk variants associated with a first feature of the disease using one or more processors; acquiring cell type-specific molecular weight trait locus (molQTL) data for a first cell type, each of which is associated with a molecular feature, using one or more processors; and determining one or more genetic risk variants causally related to the first feature of the disease and one or more associated molecular features, using one or more processors. The method further includes: identifying multiple candidate targets or biomarkers for the disease based on one or more associated molecular features using one or more processors; and generating a first cell type-specific knowledge graph, each of which is represented by one of the multiple candidate targets or biomarkers, a molecular pathway associated with a candidate target or biomarker, or one or more common interactors, using one or more processors. The method further includes calculating the level of interconnectivity for each node in the first cell type-specific knowledge graph using one or more processors, and identifying candidate targets or biomarkers, or common interactors, represented by nodes with high levels of interconnectivity, as primary targets or biomarkers.
[0012] Identifying a primary target enables therapies to be offered to treat or prevent a disease, and these therapies target the primary target, or, in some embodiments, a biological pathway associated with the primary target. A biological pathway associated with the primary target is a biological pathway that is directly affected by modifications of the primary target and propagates the effects that modify the primary target. Therefore, targeting another part of a biological pathway can also be used to treat or prevent a disease.
[0013] Identifying key biomarkers enables the detection of biomarkers and facilitates disease diagnosis. The identification of key biomarkers can be used to stratify patient populations, allowing for more personalized therapeutic interventions.
[0014] One aspect of the methods of this disclosure aims to improve target identification in drug discovery programs by elucidating the genetic contribution to disease at a cell type-specific level, as known genetic mechanisms of drug action are associated with an increased likelihood of drug approval (see Nelson et al., Nat. Genet., 2015, 47(8):856-60). The methods of this disclosure provide cell type resolution of targets and the relationship between the target and specific disease features(s), directly incorporating specific patient groups or subgroups that can be identified into the target discovery process. The methods of this disclosure enable the generation of a results matrix that defines the causal relationships between cell types and disease features, thereby enabling prioritization of which cell types should be further evaluated using a cell type-specific knowledge graph. The methods can also accelerate downstream drug optimization processes, such as biological prioritization, assay development, patient selection / stratification, and biomarker identification.
[0015] The method described herein uses a cell type-specific knowledge graph to position candidate targets or manufacturers, thereby deriving additional biological insights and leading to the identification of broader biological variations in pathways and / or functions. This offers a significant advantage over methods that analyze single genes or proteins as primary targets / biomarkers.
[0016] Genes or proteins associated with genetic risk mutations may or may not be therapeutic targets, and may be undesirable as therapeutic targets (for example, due to known side effects or the difficulty in producing drugs specific to those targets). If genes or proteins associated with genetic risk mutations are expressed in multiple cell types, the methods of this disclosure provide insights into which cell types should be prioritized, for example, based on the incidence and / or effect size of the causal relationship between the cell type and the disease feature. If genes or proteins associated with genetic risk mutations cannot be therapeutic targets, the methods of this disclosure can be used to identify and rank drug-potential genes or proteins related to genetically supported disease biology, while maintaining cell type specificity within a curated knowledge graph.
[0017] According to a second aspect, the disclosure provides a method for treating a subject suffering from a disease, wherein a primary target identified by the method according to the first aspect of the disclosure is targeted by the therapy to treat or prevent the disease. In certain circumstances, multiple lead targets may be identified, and one or more of these lead targets may be targeted by the therapy to treat or prevent the disease. The therapy can be any suitable therapy (i.e., any suitable therapeutic modality) depending on the nature of the target. If the target is a gene sequence, gene therapy can be used to knock out the gene or reduce its expression. If the target is a protein, molecules that can interact with the protein to inhibit or reduce its function, such as antibody molecules or small molecules, can be used. In some embodiments, the primary target is not a known target of a disease; that is, the target has not been previously disclosed as a target of a disease. The primary target may be known as a target of another disease, or it may not be known as a target of any disease. The therapy can be provided in vivo or ex vivo.
[0018] According to a third aspect, the Disclosure provides a method for determining whether a subject has a disease, comprising determining the presence of a major biomarker in a sample obtained from the subject, wherein the major biomarker is identified by the method described in the first aspect of the Disclosure. In situations where multiple major biomarkers have been identified, the Method may include determining the presence of multiple major biomarkers in a sample to determine whether the subject has a disease. Depending on the major biomarker and the specific cell type for which a causal relationship between the biomarker and the disease has been determined, different methods and samples may be used to detect the presence of the biomarker.
[0019] In a fourth aspect, the Disclosure discloses a system comprising one or more processors and memory for storing instructions, wherein when an instruction is executed by one or more processors, the system causes one or more processors to carry out a method according to the first aspect of the Disclosure.
[0020] Additional aspects of this disclosure provide a method for identifying key targets or biomarkers for a disease. The method includes: acquiring data on genetic risk mutations associated with multiple features of the disease using one or more processors; acquiring cell-type-specific molecular weight trait locus (molQTL) data for multiple cell types, where each molQTL is associated with a molecular feature, using one or more processors; and determining a matrix of gene mutations causally related to one or more of the multiple features of the disease and one or more relevant molecular features for multiple cell types, using one or more processors. The method further includes: identifying multiple candidate targets or biomarkers for the disease based on one or more relevant molecular features using one or more processors; and generating a first cell-type-specific knowledge graph, each of the first multiple nodes, where each of the first multiple nodes represents one of the multiple candidate targets or biomarkers, a molecular pathway associated with a candidate target or biomarker, or one or more common interactors. The method further includes calculating the level of interconnectivity for each node in the first cell type-specific knowledge graph using one or more processors, and identifying candidate targets or biomarkers, or common interacting factors, represented by nodes with high levels of interconnectivity in the first cell type-specific knowledge graph, as primary targets or biomarkers.
[0021] Additional aspects and embodiments of the System are disclosed, and these aspects and embodiments should not be construed as limiting this disclosure.
[0022] In accordance with the foregoing and the disclosures herein, this disclosure includes improvements to computer capabilities, or at least improvements to other technologies, because the disclosures herein disclose a process for creating and storing cell type-specific knowledge graphs used to identify primary targets or biomarkers. When deployed on an underlying system, such a process for creating and storing knowledge graphs allows the systems and methods of this disclosure to be performed with fewer iterations and use fewer computational resources than systems and methods relating to the prior art. In other words, this disclosure describes the capabilities of the computer itself as an improvement to "any other art or technical field," because the cell type-specific knowledge graph of this disclosure allows the underlying computer system to utilize fewer processing and memory resources compared to systems and methods relating to the prior art. This is at least because the cell type-specific knowledge graph provides a compact and useful representation of the cell network relating to candidate targets or markers without requiring further experiments, analyses, or computer simulations involving multiple computational cycles and extensive testing with data. Therefore, using the cell type-specific knowledge graph of this disclosure reduces the number of computational cycles or iterations and mitigates the impact on the underlying computing power compared to conventional prior art systems and methods (which do not use a compact and useful representation of the cell network). In other words, the systems and methods of this disclosure are an improvement over the prior art, at least because the prior art systems and methods require an empirical or trial-and-error approach, which may involve real-world trials with respect to the identification of key targets or biomarkers, resulting in and potentially requiring large database and memory utilization as well as processor usage to arrive at similar real-world or simulated results with the same or similar outcomes.On the other hand, the disclosed systems and methods describe the use of a compact and informative representation of a cellular network including a cell-type-specific knowledge graph, which is performed for candidate target or marker identification within the graph itself without requiring additional experiments, and this requires less memory usage and / or processing usage compared to conventional approaches where or in which a large set of unknown, potentially irrelevant data is used or required. Furthermore, the disclosure herein enables the identification and utilization of high-reliability datasets as part of the cell-type-specific knowledge graph, thereby reducing the need for additional computational cycles of an underlying computing platform.
[0023] Furthermore, the present disclosure involves converting or reducing a particular article to a different state or object, for example, converting or reducing a cell-type-specific knowledge graph to its graphical representation. The graphical representation of the cell-type-specific knowledge graph provides visual feedback on a graphical user interface (GUI) or other display (e.g., a display screen) regarding the relationships of the target(s) or biomarker(s) (including the primary target(s) or biomarker(s)) within the knowledge graph, thereby helping to improve the identification of primary targets or biomarkers, or common interacting factors, and such relationships may not be identified or inferred from non-graphical representations. Furthermore, the generated graphical representation efficiently uses screen space to reduce display requirements (e.g., size requirements of a display screen), while advantageously reducing the visual impact of nodes that are unlikely to be identified as primary targets or biomarkers.
[0024] Further, the present disclosure includes specific features that are well understood in the art, are routine, and differ from conventional activities, and / or adds unconventional steps that limit the present disclosure to particular useful applications. For example, the present disclosure relates to systems and methods for identifying key targets or biomarkers of a disease, systems and methods for generating cell type-specific knowledge graphs, and / or systems and methods for identifying key targets or biomarkers, which can be used, for example, for carrying out drug discovery and drug development.
[0025] Advantages will become more apparent to those skilled in the art from the following description of the preferred embodiments which are illustrated and described by way of example. As will be understood, the present embodiments are capable of other different embodiments, and various details are capable of modification in various aspects. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not restrictive.
[0026] Hereinafter, embodiments of the present disclosure will be described, by way of example only, with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] [Figure 1] 1 shows a method for identifying a key target or biomarker for a disease according to an aspect of the present disclosure. [Figure 2] 2 shows steps that can be performed in the method of FIG. 1 according to an embodiment of the present disclosure. [Figure 3] 3 shows two-sample Mendelian randomization. [Figure 4] 4 shows a causal relationship matrix according to an embodiment of the present disclosure. [Figure 5A] 5A shows a knowledge graph according to an embodiment of the present disclosure. [Figure 5B] 5B shows a cell type-specific knowledge graph generated from the knowledge graph shown in FIG. 6A according to an embodiment of the present disclosure. [Figure 6A] 6 shows a visualization of a matrix of results of two-sample Mendelian randomization according to an example of the present disclosure. [Figure 6B] A visualization of the matrix resulting from two-sample Mendelian randomization according to one embodiment of this disclosure is shown. [Figure 6C] A visualization of the matrix resulting from two-sample Mendelian randomization according to one embodiment of this disclosure is shown. [Figure 6D] A visualization of the matrix resulting from two-sample Mendelian randomization according to one embodiment of this disclosure is shown. [Figure 7-1] Figures 6A to 6C show a knowledge graph related to the matrix shown in Figures 6A to 6C, according to one embodiment of this disclosure. [Figure 7-2] Figures 6A to 6C show a knowledge graph related to the matrix shown in Figures 6A to 6C, according to one embodiment of this disclosure. [Figure 7-3] Figures 6A to 6C show a knowledge graph related to the matrix shown in Figures 6A to 6C, according to one embodiment of this disclosure. [Figure 8] Figures 6A to 6D show a knowledge graph related to the matrices shown in one embodiment of this disclosure. [Figure 9] An exemplary computing system for performing the method of this disclosure is shown. [Modes for carrying out the invention]
[0028] The ability to provide a process for identifying cell-type-specific and directional candidate targets / biomarkers causally associated with disease features (e.g., clinical features of Parkinson's disease) enables the identification of key targets / biomarkers with high confidence. To provide such a process, this disclosure is directed towards integrating molecular features within a broader biological context using causal analysis and a knowledge graph. Cell-type-specific candidate targets / biomarkers with genetically defined causal relationships to one or more disease features are identified by performing causal analysis (e.g., Mendelian randomization). The identified candidate targets / biomarkers are then ranked by analyzing their interconnections with common interactors / regulators in a customized cell-type-specific knowledge graph. Advantageously, this disclosure enables the identification of broader biological changes in pathways and / or molecular functions (rather than focusing solely on candidate targets of a single gene / protein) by situating causally identified candidate targets / biomarkers within a broader biological context. Primary targets and / or biomarkers may be used, for example, to carry out drug discovery or drug development.
[0029] Figure 1 shows a method 100 for identifying a primary target or biomarker for a disease according to one aspect of the present disclosure.
[0030] Method 100 includes the steps of obtaining genetic risk variant data (step 102), obtaining cell type-specific molecular weight trait locus (molQTL) data (step 104), determining genetic risk variants causally related to one or more molecular features (step 106), identifying candidate targets or biomarkers (step 108), generating a cell type-specific knowledge graph (step 110), and calculating the level of interconnectivity (step 112). Method 100 also includes the optional steps of identifying a major target or biomarker or common interactor (step 114), and outputting the major target or biomarker or common interactor (step 116). Method (100) may also include identifying one or more biological pathways associated with a major target or multiple major targets. Biological pathways associated with a major target(s) are biological pathways directly affected by modifications of the major targets.
[0031] Generally, Method 100 identifies and prioritizes cell type-specific targets / biomarkers relevant to the treatment or diagnosis of a disease. More specifically, it determines the genetically defined causal relationships between candidate targets / biomarkers, specific cell types, and disease features, and determines whether the candidate targets / biomarkers and corresponding disease features are unique or conserved across cell types and / or disease features. The cell type-derived candidate targets / biomarkers for the disease indication are then positioned as nodes within a customized cell type-specific knowledge graph (KG). In this way, the initial candidate targets / biomarkers are expanded to include other targets / biomarkers (e.g., genes or proteins) identified as common interacting factors / regulators connected to them. This helps to expand the pool of cell type-specific candidate targets / biomarkers (e.g., 5-10 times). Furthermore, expanded biological pathways can be identified within these expanded networks. Given the sparse nature of cell-type-specific candidate targets / biomarkers with significant causal associations, this step improves the ability to interpret the function and / or significance of targets / biomarkers within the context of broader cell networks, and enables more reliable and appropriate guidance for target / biomarker identification. As a result, prioritization of targets / biomarkers and pathways becomes possible. Importantly, by applying graph-based analysis to cell-specific KG data, candidate targets / biomarkers can be analyzed based on their interconnectivity / centrality, thereby enabling the identification of key (i.e., strong) targets / biomarkers with greater confidence. Overall, this method provides an iterative and modular framework for target / biomarker identification.
[0032] In step 102, data on genetic risk variants is obtained. The data on genetic risk variants is associated with a first characteristic of the disease. In some embodiments, data on genetic risk variants associated with multiple characteristics of the disease is obtained.
[0033] A “genetic risk variant” refers to any genetic variation in a nucleic acid sequence associated with one or more features of a disease. In certain embodiments, a genetic risk variant may be a single nucleotide polymorphism / variant (SNP / SNV), an insertion and / or deletion variant, a copy number variant, a translocation and / or inversion, or a combination thereof. In certain embodiments, a genetic risk variant is an SNP. In certain embodiments, a genetic risk variant is identified from a genome-wide association study (GWAS) that associates the genetic risk variant with one or more features of a disease. In certain embodiments, a “genetic risk variant” is any genetic variation in a nucleic acid sequence associated with one or more features of a disease for a specific patient subgroup, e.g., a patient subgroup identified by biomarkers (e.g., genetic markers) and / or comorbidities.
[0034] "Disease features" refer to the phenotypic characteristics of a disease, and in some embodiments, include the onset of the disease, i.e., the presence of the disease. A disease is characterized by a set of features, such as symptoms, morphological changes, functional changes, etc. Many diseases have common features, and a diagnosis of a disease is made based on a combination of these features. In this disclosure, a genetic risk variant may be associated with the presence of a disease and / or with one of several features used to diagnose a disease, i.e., a disease feature. In certain embodiments, disease features refer to phenotypic characteristics of a disease that, in themselves, cannot definitively diagnose the disease. For example, in the case of Parkinson's disease, "disease features" may refer to the severity of motor scores, age of onset, cognitive impairment, depression, daytime sleepiness, dyskinesia, olfactory impairment, insomnia, RBD, MMSE score, and Hoehn and Yahr scale. In certain embodiments, the term "disease features" refers only to the phenotypic characteristics of a disease and does not include the onset of the disease, i.e., the presence of the disease.
[0035] In the context of this disclosure, the term “disease” means any disease, including cardiac diseases such as ischemic heart disease and coronary artery disease; cancer; bacterial and viral infections; cerebrovascular diseases such as stroke; respiratory diseases including chronic obstructive pulmonary disease; diabetes mellitus; autoimmune diseases such as ulcerative colitis, Crohn’s disease, inflammatory bowel disease, rheumatoid arthritis, Guillain-Barré syndrome, Sjögren’s syndrome, scleroderma, and Graves’ disease; and neurological diseases including Alzheimer’s disease, multiple sclerosis, Parkinson’s disease, amyotrophic lateral sclerosis (ALS) and frontotemporal dementia. In some embodiments, the disease is Parkinson’s disease.
[0036] Data on genetic risk variants can be obtained from any suitable source, including GWAS summary statistics. Examples of GWAS summary statistics databases are listed in Table I. [Table I] In one embodiment, GWAS data are preprocessed such that data for genetic risk variants are associated with one or more features of a disease, and that the association meets a statistically significant threshold, i.e., the relationship or association between two or more features is statistically significant. Any preferred method, as is well known to those skilled in the art, can be used to determine statistical significance. The threshold can be set to a standard significance level, for example, p<0.05, or to a higher level of statistical significance, for example, p<0.01, p<0.001, or p<0.0001. In one particular embodiment, the statistically significant threshold p is p<0.0001. Those skilled in the art will understand that such preprocessing can be performed using the ieugwasr package in the R programming language, or any preferred software such as vcf2gwas, PyGWAS, or Hail in the Python programming language.
[0037] In step 104, cell type-specific molecular weight phenotypic locus (molQTL) data are obtained for a first cell type. Each molQTL is associated with a molecular feature. In a particular embodiment, cell type-specific molecular weight phenotypic locus (molQTL) data are obtained for multiple cell types.
[0038] A molecular weight-based phenotypic locus (molQTL) refers to any gene mutation associated with molecular features such as expression level (eQTL), DNA methylation level (meQTL), histone modification (hQTL), chromatin accessibility (caQTL), alternative gene splicing (sQTL), protein-level pQTL, microRNA expression (mirQTL), and ribosome occupancy (rQTL). See Vandiedonck, Clin. Genet., 93, 520-532 (2018). As used herein, molQTL data are cell-type specific molQTL data. In other words, molQTL data are related to a particular cell type. In certain embodiments, gene mutations associated with molecular features may be single single nucleotide polymorphisms / mutations (SNPs / SNVs), insertions and / or deletions, copy number variations, translocations and / or inversions, or combinations thereof. In certain embodiments, the gene mutation is an SNP.
[0039] In certain embodiments, molQTL data are cell-specific eQTL data. Gene mutations associated with expression levels may be single single nucleotide polymorphisms / mutations (SNPs / SNVs), insertions and / or deletions, copy number variations, translocations and / or inversions, or combinations thereof. In certain embodiments, molQTL data are cell-specific eQTL data, and gene mutations are SNPs associated with expression levels.
[0040] Cell type-specific molQTL data can be obtained from one or more sources, including Uffelmann et al., Biological Psychiatry, 2021, 89:41-53, the DICE database, BIOSQTL, BloodeQTL, EQTLcatalogue, EQTLGen, GTEx, scRNA_eQTLs, BRAINEAC, CMC, PsychENCODE, and xQTLServer.
[0041] Similar to the gene mutation data mentioned above, molQTL data can also be preprocessed to obtain data for specific cell types, thereby ensuring that the correlation between molQTL data and molecular features meets the statistically significant threshold (as defined above).
[0042] As used throughout this disclosure, “specific cell types” refers to defined cell types that are phenotypic and / or functionally distinct. This term may refer to cell types, cell subtypes, or cell phenotypes. For example, this term may refer to cell lineages such as T cells, B cells, natural killer cells, and monocytes; cell subtypes such as T cell subtypes (e.g., CD4 naive, stimulated CD4, CD8 naive, stimulated CD8, TFH, TH1, TH2, TH1 / 17, TH17, Treg naive, and Treg memory cells); or cell phenotypes such as M1 macrophages and M2 macrophages.
[0043] In a particular embodiment, a particular cell type is a specific cell type and / or a specific cell subtype, rather than a general class of cells such as immune cells, blood cells, or muscle cells.
[0044] Using cell type-specific molecular weight phenotypic locus (molQTL) data allows for the precise identification of causal effects at a cell type-specific level. For example, when whole blood-derived molQTL data is used (instead of cell type-specific data), the mean effect size is observed for targets causally associated with multiple cell types, and this is dominated by the strongest effect, potentially obscuring differences in the magnitude and / or direction of the association. For instance, in the 2SMR for PD onset, KANSL1 showed positive and negative causal associations with PD risk depending on the cell type.
[0045] In step 106, one or more genetic risk mutations are identified that are causally related to the primary characteristic of the disease and one or more associated molecular characteristics.
[0046] This step enables the identification of genetic risk variants that have a causal relationship with disease characteristics and one or more associated molecular characteristics. In certain embodiments, the identification of causality enables the identification of directionality between genetic risk variants and disease characteristics.
[0047] By integrating data on one or more genetic risk variants with cell type-specific molQTL data using a causal analysis process to determine causal relationships, one or more genetic risk variants that are causally related to disease characteristics and one or more associated molecular characteristics are identified. In one embodiment, causality is determined using a Mendelian randomization approach such as two-sample Mendelian randomization (2SMR), which will be described in more detail later in relation to Figure 3.
[0048] Causal analysis (e.g., 2SMR) can identify causal relationships between one or more genetic risk variants and one or more disease features at a cell-specific level. The causal relationship satisfies a statistically significant threshold, e.g., p < 0.0001. In embodiments using Mendelian randomization, the statistical significance of the causal relationship is determined as part of the implementation of the Mendelian randomization analysis. Preferred methods include those described in Richardson et al., 2020 (cited above), Taylor et al., 2019 (cited above), Porcu et al., 2019 (cited above), and the R module of TwoSampleMR. In certain embodiments, the R module of TwoSampleMR is used to perform the Mendelian randomization analysis. In certain embodiments, the use of Mendelian randomization analysis enables the identification of directionality between genetic risk variants and disease features.
[0049] Identifying causal relationships between one or more genetic risk variants and one or more disease features at the cell-specific level is based on identifying genetic risk variants that are associated with disease features and molecular weight traits.
[0050] In one embodiment, data for one or more genetic risk variants and cell type-specific molQTL data are harmonized and / or filtered before causal analysis is performed. Harmonizing the data ensures that the genetic risk variant data and cell type-specific molQTL data are comparable and allows for the identification of common genetic features, such as identical SNPs, in the two datasets. Data harmonization can be achieved by mapping the data to a single reference genome. Suitable genome coordinate transformation and remapping tools for performing such harmonization include CrossMap and the Bioconductor rtracklayer package for the R programming language.
[0051] The step of filtering data before performing causal analysis includes identifying genetic features shared between the two datasets, such as identical SNPs. Non-shared genetic features can be filtered out, or removed, from the datasets before causal analysis integrates matching data for one or more genetic risk variants with matched cell type-specific molQTL data. Such filtering steps can be easily performed using a suitable software package, such as TwoSampleMR.
[0052] In embodiments where genetic risk variant data for multiple different disease features are obtained in step 102 and / or molQTL data for multiple different cell types are obtained in step 104, step 106 includes generating a matrix of gene variants. The gene variants are causally related to one or more features of the disease or one or more relevant molecular features for the cell type.
[0053] As will be explained in more detail in relation to Figure 4 below, the genetic risk variation matrix includes a table of values that quantify the causal relationship between cell type-molecular feature combinations and disease feature outcomes. Thus, the genetic risk variation matrix can be generated by estimating causal relationships for each combination of gene variations, disease features, and cell type-specific molecular features in the acquired data (for example, by calculating causal effects using Mendelian randomization). For example, a matrix may be generated for a specific disease feature to quantify the causal relationship between gene variations and several specific cell types (see Figure 4). Alternatively, a matrix may be generated for a specific cell type to quantify the causal relationship between gene variations and several different disease features. A 3D matrix may be generated to quantify the causal relationships between several specific disease features, gene variations, and several specific cell types.
[0054] In step 108, multiple candidate targets or biomarkers for the disease are identified based on one or more relevant molecular characteristics. Furthermore, in embodiments in which a matrix of genetic risk variations is generated, step 108 may include identifying candidate targets or biomarkers for the disease from the matrix of genetic risk variations (described in more detail in relation to Figure 4 below).
[0055] Candidate disease targets or biomarkers are nucleic acid sequences or proteins associated with molecular features identified as having a causal relationship between genetic risk mutations and disease characteristics. For example, if the molecular feature is the expression level of a specific gene in a particular cell, the candidate target or biomarker could be a specific gene or encoding protein, a cis-acting nucleic acid sequence, or a trans-acting nucleic acid sequence or protein.
[0056] The step of identifying candidate disease targets or biomarkers based on one or more genetic risk variants identified as causally related to one or more disease features is achieved by determining the location and / or association of the genetic risk variants. If the genetic risk variants reside within a specific gene, the candidate target / biomarker may be determined to be that gene. For genetic risk variants located outside the coding region, molQTL data or other data can be used to map the genetic variants to the genes affected by the genetic risk variants.
[0057] In certain embodiments, hierarchical clustering is used to evaluate candidate targets or biomarkers. Distance methods for performing such hierarchical clustering include Euclidean distance, Pearson correlation, Spearman correlation, and Manhattan distance.
[0058] In step 110, a first cell type-specific knowledge graph is generated. The first cell type-specific knowledge graph contains a first set of nodes, each representing either one of several candidate targets or biomarkers, a biological pathway associated with a candidate target or biomarker, or one or more common interacting factors.
[0059] Here, a knowledge graph is a network or graph of molecular pathways known to be associated with each of the candidate targets or biomarkers. Molecular pathways that form part of a knowledge graph include gene expression pathways and protein-protein interaction pathways. A cell type-specific knowledge graph is specific to the cell type from which molQTL data was obtained, or to a group of cell types closely related to and encompassing a specific cell type(s) from which molQTL data was obtained, i.e., to the cell subtypes that have the ability to transform into the specific cell type from which molQTL data was obtained. In a particular embodiment, the corresponding cell type-specific knowledge graph is specific to the cell type from which molQTL data was obtained.
[0060] A cell type-specific knowledge graph includes nodes for a particular cell type and edges connecting those nodes. A node may represent one of the following: a candidate target or maker (e.g., a gene identified as causally related to a disease feature), a molecular pathway associated with a candidate target or biomarker, or a common interacting factor (e.g., a gene associated with a molecular pathway related to a candidate target or biomarker and / or a candidate target or maker).
[0061] Cell type-specific knowledge graphs can be generated from open-source databases such as Ingenuity Pathways Analysis (IPA), Hetionet, CKG (Molecular quantitative trait loci. Nat Rev Methods Primers3,5(2023)), and PrimeKG (Chandak,P.,Huang,K.&Zitnik,M. Building a knowledge graph to enable precision medicine. Sci Data10,67(2023)). Cell type-specific knowledge graphs can be created by adjusting the data to remove any genes or pathways that are not expressed in the cell type from which molQTL data was obtained, i.e., genes that do not exist in the cell type-specific transcriptome profile. As will be described in more detail in relation to Figures 5A and 5B below, a cell type-specific knowledge graph related to a particular cell type can be generated from a “parent” knowledge graph, i.e., a knowledge graph that includes both genes and pathways expressed and not expressed in the cell type from which molQTL data was obtained. The parent knowledge graph is pruned according to locality constraints (i.e., nodes and edges are removed), and the pruned parent knowledge graph becomes the cell type-specific knowledge graph. The parent knowledge graph is pruned according to locality constraints with respect to nodes in the parent knowledge graph that represent candidate targets or biomarkers. In this way, only nodes in the parent knowledge graph that satisfy locality constraints (in relation to nodes representing candidate targets or biomarkers) are retained in the cell type-specific knowledge graph.
[0062] In one embodiment, method 100 further includes outputting a cell type-specific knowledge graph (not shown).
[0063] Outputting a cell type-specific knowledge graph involves storing the cell type-specific knowledge graph in the computing device's memory. This memory may be the computing device's memory, such as random access memory (RAM). Alternatively, it may be a persistent memory location, such as a hard disk drive or cloud storage, which are non-transient computer-readable media. Storing the cell type-specific knowledge graph in the computing device's memory improves the predictive and identification capabilities of the computing device through the stored cell type-specific knowledge graph. Furthermore, pruning the cell type-specific knowledge graph from the "parent" knowledge graph (as described above) can create and store compact representations of pathways related to candidate targets or biomarkers, or one or more common interactors, thereby enabling more efficient use of computing resources, such as memory and storage requirements.
[0064] Additionally, or alternatively, outputting a cell type-specific knowledge graph includes displaying a graphical representation of the first cell type-specific knowledge graph within a graphical user interface for user review. An example of a graphical representation is shown in Figure 7, which is described in more detail below. The graphical representation provides the user with important visual feedback regarding relationships identified within the cell type-specific knowledge graph that may not be identified or inferred from the non-graphical representation. Therefore, outputting such a graphical representation may improve the identification of key targets or biomarkers, thereby contributing to improved patient outcomes. The graphical representation is generated by applying a graph plotting algorithm to the cell type-specific knowledge graph, such as a force direction plotting algorithm or a spring / rebound model. The size of nodes in the graphical representation may vary depending on their interconnectivity (as described in relation to step 112 below), with nodes with a high level of interconnectivity appearing larger in the graphical representation than nodes with low interconnectivity. Advantageously, this improves the identification of potential primary targets or biomarkers by reducing the amount of display area allocated to nodes that are less likely to be identified as primary targets or biomarkers, and thus reducing the visual impact of those nodes.
[0065] In step 112, the level of interconnectivity is calculated for each node in the first cell type-specific knowledge graph. Candidate targets or biomarkers, or common interacting factors, represented by nodes with high levels of interconnectivity are identified as primary targets or biomarkers.
[0066] The level of interconnectivity associated with a node is a measure of the node's relationship within the overall context of a cell type-specific knowledge graph. A node with a high level of interconnectivity indicates that the node (e.g., a gene, pathway, etc.) has a significant (e.g., greater) biological effect in the overall biology of the cell type, particularly with respect to relevant biological effects, i.e., biological effects predicted to influence disease features. Therefore, the level of interconnectivity can be used as a measure to rank the relationship between nodes and their disease features with respect to arbitrarily relevant biological effects. The level of interconnectivity is determined for each node in the first cell type-specific knowledge graph. In one embodiment, the level of interconnectivity is a measure of node centrality, such as the node's degree or its intrinsic center. Those skilled in the art will understand that such measures can be calculated using a suitable software package such as NetworkX.
[0067] Candidate targets or biomarkers with a high level of interconnectivity, or common interactors with a high level of interconnectivity, are identified as primary targets or biomarkers. As used throughout this disclosure, common interactors are genes or proteins present in the knowledge graph that have a high level of interconnectivity with candidate targets or biomarkers but have not been identified as candidate targets or biomarkers. In certain embodiments, common interactors are identified as primary targets or biomarkers.
[0068] The term “primary target” refers to any target that can be targeted by a therapy to treat or prevent a disease. In certain embodiments, the primary target is a gene sequence or a protein. If the target is a gene sequence, gene therapies such as miRNA therapy can be used to knock out that gene sequence or reduce its expression. If the target is a protein, molecules that can interact with the protein to inhibit or reduce its function, such as antibody molecules or small molecules, can be used. In some embodiments, the primary target is not a known target of the disease. The primary target may be known as a target of another disease, or it may not be known as a target of any disease. In certain embodiments, the therapy can be delivered in vivo or ex vivo.
[0069] The term “primary biomarker” refers to any biomarker that can be used to diagnose a disease and / or to monitor the progression and / or severity of the disease. Primary biomarkers can be used alone to diagnose a disease or in combination with other biomarkers. Biomarkers may be gene sequences, proteins, or other molecular traits. For example, a primary biomarker may be the presence or level of a nucleic acid sequence or protein, or the level of nucleic acid methylation or phosphorylation. In certain embodiments, the primary biomarker is the presence or level of a nucleic acid sequence or protein.
[0070] In a particular embodiment, the primary target or biomarker is the most highly interconnected, targetable, or detectable target, biomarker, or common interactor. Other factors may also be considered in identifying the primary target or biomarker, including whether the target or biomarker is specific to or conserved across cell types, whether it is causally associated with multiple disease features, the accessibility of the target or biomarker, and / or whether it is known or predicted to be associated with other diseases. For example, if a highly interconnected candidate target cannot be used as a therapeutic target, the methods of this disclosure can be used to identify and rank drug-worthy targets, i.e., candidate targets or common interactors that are highly interconnected while maintaining desired cell type specificity.
[0071] In one embodiment, the determination of whether a node has a high level of interconnectivity is based on the percentile level of interconnectivity of all nodes in the cell type-specific knowledge graph. Nodes with an interconnectivity level within the top 75 percentile of the interconnectivity level of all nodes in the cell type-specific knowledge graph are identified as nodes representing a major candidate target or biomarker, or a common interacting factor. In a particular embodiment, nodes with an interconnectivity level within the top 90 percentile or top 95 percentile of the interconnectivity level of all nodes in the cell type-specific knowledge graph are identified as nodes representing a major candidate target or biomarker, or a common interacting factor. Additionally or alternatively, the determination of whether a node has a high level of interconnectivity is based on a predetermined threshold. For example, if a node has an interconnectivity level above a predetermined threshold (e.g., 100, 150, 200 or more connections), that node is identified as a node representing a major candidate target or biomarker, or a common interacting factor.
[0072] In an optional step 114, a first candidate target or biomarker, or a first common interactor, is identified based on the level of interconnectivity of the respective nodes representing the first candidate target or biomarker, or the first common interactor.
[0073] In an optional step 116, a first candidate target or biomarker, or a first common interacting factor, is output.
[0074] In one embodiment, outputting a first candidate target or biomarker, or a first common interaction factor, includes storing or saving the first candidate target or biomarker, or the first common interaction factor, in persistent storage such as non-volatile memory or a non-transient medium. Additionally or alternatively, outputting a first candidate target or biomarker, or a first common interaction factor, includes transmitting the first candidate target or biomarker, or the first common interaction factor, over a network (e.g., a local area network, a wide area network, etc.) and displaying the first candidate target or biomarker, or the first common interaction factor, or a representation thereof, for user review.
[0075] Method 100 relates to an effective method for identifying a primary target or primary biomarker of a disease. The primary target can then be used as a therapeutic target to treat the disease. The primary biomarker can be used as a biomarker for disease diagnosis. Method 100 may further include validating the primary target or primary biomarker by performing suitable in silico, in vitro, or in vivo experiments. For example, in vitro or animal models of the disease can be treated by using the primary target as a therapeutic target. Patient-derived samples or data derived from such samples can be tested to confirm that the primary biomarker is effective in the diagnosis or monitoring of the disease.
[0076] Figure 2 shows steps that may be carried out as part of Method 100 shown in Figure 1 according to one embodiment of the present disclosure.
[0077] Figure 2 shows step 302 for acquiring a second cell type-specific knowledge graph, step 304 for generating a third cell type-specific knowledge graph, and step 306 for calculating the level of interconnectivity. Figure 2 further shows optional step 310 for identifying major targets or manufacturers or common interactors, and optional step 312 for outputting major targets or biomarkers or common interactors. In one embodiment, the steps shown in Figure 2 are performed as part of method 100 shown in Figure 1. That is, steps 302 for acquisition and 304 for generation can be performed after or together with step 110 for generation in method 100.
[0078] In step 302, a second knowledge graph consisting of a second set of nodes is obtained. At least one node of the first set of nodes of the first cell type-specific knowledge graph (generated in generation step 110 of method 100) and at least one node of the second set of nodes are identical or have a known biological relationship, for example, being functionally related within a known biological pathway.
[0079] The second knowledge graph and the first cell type-specific knowledge graph share at least one common node; that is, they both contain at least one node representing a common molecular feature and / or pathway. The second knowledge graph is associated with cell type-specific molecular weight trait locus (molQTL) data for a second cell type; that is, the second knowledge graph is a cell type-specific knowledge graph associated with a cell type different from the cell type associated with the first cell type-specific knowledge graph. Alternatively, the second knowledge graph is associated with data on genetic risk variants associated with a second feature of the disease; that is, the second knowledge graph is a knowledge graph associated with a disease feature different from the first feature of the disease associated with the first cell type-specific knowledge graph.
[0080] In one embodiment, the second knowledge graph is obtained by generating the second knowledge graph using the process described above in relation to method 100. Alternatively, the second knowledge graph is obtained by loading, receiving, or otherwise acquiring the second knowledge graph from a computer-readable medium.
[0081] In step 304, a third knowledge graph is generated. The third knowledge graph includes a first cell type-specific knowledge graph and a second knowledge graph. The third knowledge graph includes at least one edge connecting at least one node of the first plurality of nodes and at least one node of the second plurality of nodes.
[0082] The third knowledge graph is generated by "bridging" the first cell type-specific knowledge graph and the second knowledge graph to form a combination or union of the two graphs. To bridge the two graphs, a connection can be formed by common nodes present in the first and second knowledge graphs. Alternatively, the two graphs can be bridged by forming edges between nodes of the first cell type-specific knowledge graph and nodes of the second knowledge graph, where these nodes have known biological relationships, for example, functionally related within known biological pathways. These two graphs can also be bridged by using common interaction factors between nodes of the first cell type-specific knowledge graph and nodes of the second knowledge graph.
[0083] If a second knowledge graph is associated with molQTL data for a second cell type, a third knowledge graph can be used to identify key targets or biomarkers cell-specifically or across multiple cell types. This provides additional information that allows for better identification of desired key targets or biomarkers. For example, if a key target or biomarker (e.g., a gene or protein) is identified in multiple cell types, this disclosure provides insights into which cell types should be prioritized, for example, based on the incidence and / or effect size of the causal relationship between the key target or biomarker and disease features. Thus, key targets or biomarkers can be identified cell-specifically or across multiple cell types.
[0084] If a second knowledge graph relates to another feature of the disease, a third knowledge graph may identify, at a cell-specific level, key targets or biomarkers that are causally associated with one or more disease features of the disease. This provides additional information that allows for better identification of desired key targets or biomarkers. For example, if a key target or biomarker (e.g., a gene or protein) is identified as being causally associated with one or more disease features, this disclosure provides insight, for example, into which disease features(s) to target and which targets or biomarkers should be prioritized based on the incidence and / or effect size of the causal relationship between the key target or biomarker and each disease feature. Thus, key targets or biomarkers may be identified with respect to one or more disease features and / or in a cell-specific manner or across multiple cell types. If a second knowledge graph relates to a second or another feature of the disease, it is self-evident that at least one of the disease features is not the onset of the disease, i.e., the presence of the disease. In a particular embodiment, the first and second features of the disease are not the onset of the disease, i.e., the presence of the disease.
[0085] Step 306, which calculates the level of interconnectivity, corresponds to step 112 in Method 100, except that interconnectivity is calculated for each node in the third knowledge graph. Therefore, for further details of the operations performed in step 306, please refer to the description of step 112. Similarly, the optional steps of identifying step 310 and outputting step 312 correspond to identifying step 114 and outputting step 116 as described in relation to Method 100 (but using the third knowledge graph).
[0086] Those skilled in the art will understand that by repeating the steps described above in Figure 2 for two or more cell types and / or two or more disease features, key targets or biomarkers can be identified from knowledge graphs that include combinations of several cell type-specific knowledge graphs and / or knowledge graphs related to several features of a disease. In one embodiment, a cell type-specific disease feature-specific knowledge graph is generated for each combination of cell type and disease feature. For example, if molQTL data is obtained for two different cell types (e.g., cell types A and B) and genetic risk variant data is obtained for two different disease features (e.g., disease features X and Y), then cell type-specific disease feature-specific knowledge graphs are generated for cell type A and disease feature X, cell type A and disease feature Y, cell type B and disease feature X, and cell type B and disease feature Y (using the approach described above in relation to method 100 shown in Figure 1). These knowledge graphs may be analyzed individually to identify key targets or biomarkers for each cell type and disease feature combination, or they may be combined using the approach described in relation to Figure 2. Combined or bridged knowledge graphs can be used to identify key candidate targets or biomarkers, or common interacting factors.
[0087] Figure 3 shows two-sample Mendelian randomization (2SMR).
[0088] Figure 3 shows the instrumental variable 402, Z, exposure 404, X, outcome 406, Y, and one or more confounding factors 408, U. The instrumental variable 402, Z corresponds to a genetic risk variant (e.g., a single nucleotide polymorphism / mutation, SNP / SNV) associated with disease features. Exposure 404, X corresponds to one or more molecular features, and outcome 406, Y corresponds to disease features. The association between genetic risk variant Z and molecular feature X is as follows:
number
number
number
[0089] connection
number
number
number
[0090] Therefore, 2SMR is used to test the causal hypothesis that molecular feature X is causally related to disease feature Y, and at the same time, it also provides an estimate of the causal effect (i.e., the amount or extent of the positive or negative causal effect that molecular feature has on disease feature).
[0091] As described above, 2SMR may be used as part of the identification step 106 of Method 100 or the identification step 206 of Method 200. Any suitable software package or library, such as the TwoSampleMR module of the R programming language, can be used to implement 2SMR.
[0092] Figure 4 shows a causal matrix 502 according to the embodiments of this disclosure.
[0093] Matrix 502 (also called the causal matrix, effect matrix, or causal / effect matrix) contains multiple causal effect values determined in relation to multiple cell types (horizontal axis of Matrix 502) and multiple molecular features (vertical axis of Matrix 502) and disease features or outcomes (e.g., motor score severity for Parkinson's disease). Alternatively, although not shown in Figure 4, the matrix may contain multiple causal effect values determined in relation to multiple disease features or outcomes (e.g., horizontal axis of the matrix), multiple molecular features (e.g., vertical axis of the matrix), and specific cell types (e.g., specific T cell subtypes). The strength of causality is represented by the intensity of various causal effect values. No intensity indicates no causality, while dark intensity indicates a strong causal relationship.
[0094] Figure 4 shows the first value 504, the second value 506, the third value 508, and the fourth value 510 of matrix 502.
[0095] In Figure 4, each value in Matrix 502 can be understood as a quantification of the causal relationship between cell type-molecular feature combinations and disease outcomes. The first value 504 corresponds to a measure of the causal relationship between cell type CT1, molecular feature MF1 and disease features or outcomes. For example, the first value 504 may correspond to the magnitude of the causal relationship between the gene CCDC92 (MF1) in Th1 / 17 cells (CT1) and the severity of the Parkinson's disease motor score (UPDRS III). The second value 506 corresponds to cell type CT1, molecular feature MF i This corresponds to a measure of the causal relationship between disease characteristics or outcomes. The third value, 508, corresponds to cell type CT. i Molecular characteristics MF i This corresponds to a measure of the causal relationship between disease characteristics or outcomes. The fourth value, 510, corresponds to cell type CT. n Molecular characteristics MF i This corresponds to a measure of the causal relationship between disease characteristics or outcomes.
[0096] The first value 504 and the fourth value 510 are approximately equal to 0, meaning that MF1 of CT1 and CT n of MF i have no causal association with a disease characteristic or outcome. The second value 506 is a negative value, meaning that the MF of CT1 i has a causal association with the risk of onset / presence of a disease characteristic or outcome, and has a negative association. The magnitude of the second value 506 represents the magnitude of the causal association between MF i , CT1 and the disease characteristic or outcome. The third value 508 is a positive value, meaning that the MF of CT i of MF i has a causal association with the risk of onset / presence of a disease characteristic or outcome, and has a positive association. The magnitude of the third value 508 represents the magnitude of the causal association between MF i , CT i and the disease characteristic or outcome.
[0097] The matrix 502 may be generated by iteratively applying a causal analysis approach (e.g., Mendelian randomization) to all cell type-molecular feature pairs associated with a given disease characteristic or outcome. As described above with reference to Figure 1, a matrix such as matrix 502 can be used to identify candidate targets or biomarkers. For example, a molecular feature MF i may be identified as a candidate target or biomarker in association with cell type CT i and a disease characteristic or outcome, because the third value 508 is higher compared to other values in the column of values associated with CT i and the disease characteristic or outcome. Here, a high causal effect value for a cell type is determined by identifying the maximum causal effect value across all molecular features for that cell type. Alternatively, a high causal effect value is determined by identifying any causal effect value above a predetermined threshold across all module features for that cell type.
[0098] Figure 5A shows a knowledge graph 600 according to an embodiment of the present disclosure.
[0099] Knowledge Graph 600 includes multiple nodes, including nodes 602-614 representing molecular features (e.g., genes), nodes 616-622 representing molecular pathways, node 624 representing a first candidate target or biomarker, and node 626 representing a second candidate target or biomarker. Both nodes 624 and 626 represent molecular features identified as candidate targets or biomarkers.
[0100] Graph 600 corresponds to a “parent” knowledge graph from which a cell type-specific knowledge graph is generated. Therefore, knowledge graph 600 includes genes and pathways expressed and not expressed in the cell types to which the cell type-specific knowledge graph relates. To generate a cell type-specific knowledge graph, knowledge graph 600 is pruned to remove nodes that are not biologically associated with / annotated to cell types and candidate targets or biomarkers. The pruning step is used to remove genes, proteins, and pathways not annotated to a particular cell type from the parent knowledge graph. The pruning step may also be used to remove genes, proteins, and pathways that are active in a particular cell type but are also active in different, unrelated cell types and / or are not of interest for their involvement in cellular homeostasis. Knowledge graph 600 is also pruned according to one or more locality constraints. For example, nodes in the knowledge graph 600 that are identified as not representing a candidate target or biomarker, or that are not within a certain distance of a candidate target or biomarker, are removed by pruning. A cell type-specific knowledge graph is generated from the knowledge graph 600 by pruning or removing nodes that are not identified and / or do not satisfy one or more locality constraints. Advantageously, locality constraints help ensure that relationships between nodes are not significantly affected by a large number of intermediate or bystander nodes.
[0101] The first locality constraint is based on a shortest path constraint such that the shortest path distance between pairs of nodes representing candidate targets or biomarkers in the cell type-specific knowledge graph is less than a given distance (e.g., less than 4 or less than 3). In the example shown in Figure 5A, nodes in knowledge graph 600 representing candidate targets or biomarkers are identified (i.e., nodes 624 and 626), and the shortest paths between each pair of these nodes are calculated. Between nodes 624 and 626, there are two shortest paths: one containing node 616 and another containing node 620. Since the shortest path between nodes 624 and 626 is less than 4 (i.e., there are less than 4 nodes on the path connecting nodes 624 and 626), the nodes on these paths are retained in the cell type-specific knowledge graph.
[0102] The second locality constraint is based on the direct proximity of the node representing the candidate target or biomarker. That is, any node in knowledge graph 600 that is directly adjacent to (i.e., directly connected to) a node representing a candidate target or biomarker is retained in the cell type-specific knowledge graph. In the example shown in Figure 5A, applying the second locality constraint (in addition to the first locality constraint) results in nodes 602, 604, 614, and 618 all being retained in the cell type-specific knowledge graph because they are direct neighbors of node 624 or node 626.
[0103] The third locality constraint involves evaluating the specificity and significance of nodes in the knowledge graph using a one-sample substitution analysis. More specifically, the one-sample substitution analysis is performed based on the node order and the corresponding false detection rate (FDR), which is calculated by comparing the node order in the cell type-specific knowledge graph with the random graph. Any node with a p-value < 0.05 is retained in the cell type-specific knowledge graph (i.e., any node with a p-value ≥ 0.05 is pruned from the cell type-specific knowledge graph).
[0104] Nodes in the knowledge graph 600 that do not satisfy one or more of the above locality constraints are pruned or deleted, resulting in the generation of a cell type-specific knowledge graph.
[0105] Figure 5B shows a cell type-specific knowledge graph 600-1 generated from the knowledge graph 600 shown in Figure 5A, according to one embodiment of the present disclosure.
[0106] Cell type-specific knowledge graph 600-1 corresponds to knowledge graph 600 shown in Figure 5A, pruned according to one or more locality constraints, with one or more nodes in knowledge graph 600 that do not satisfy the locality constraints removed. That is, nodes 606-614 and node 622 in knowledge graph 600 are pruned and do not appear in cell type-specific knowledge graph 600-1. In this way, cell type-specific knowledge graph 600-1 includes nodes that are closely related to, biologically relevant, and statistically significant to candidate targets or biomarkers.
[0107] Subsequently, the level of interconnectivity is calculated for each node in the cell type-specific knowledge graph 600-1, and potential primary targets or biomarkers can be identified. As mentioned above, the node order can be used as a measure of the level of interconnectivity of the nodes. In the cell type-specific knowledge graph 600-1, nodes 602, 614, and 618 have an order of 1, nodes 604, 616, 620, and 626 have an order of 3, and node 624 has an order of 5. Node 624 has a high level of interconnectivity (compared to the interconnectivity levels of other nodes in the cell type-specific knowledge graph 600-1) and can therefore be identified as a primary target or biomarker.
[0108] In certain embodiments, genetic risk variant data can be additionally acquired for one or more comorbidities. The genetic risk variant data for one or more comorbidities is determined to be causally related to one or more comorbidities and one or more associated molecular features in a manner similar to that described above for one or more disease features. This step allows for the identification of multiple candidate targets or biomarkers for one or more comorbidities. Knowledge graphs of molecular pathways known to be associated with each of the candidate targets or manufacturers are generated in the manner described above. The knowledge graphs for comorbidities can be combined with one or more knowledge graphs acquired for disease features in the manner described above. Thus, based on one or more cell types, one or more disease features, and comorbidities, major targets or manufacturers can be identified. Such methods provide further information regarding major targets or biomarkers for diseases with comorbidities.
[0109] Comorbidities are additional medical conditions that occur alongside the primary disease. Examples include depression accompanied by diabetes, or Parkinson's disease accompanied by diabetes. [Examples]
[0110] Data source All data used for the 2SMR analysis were obtained from the following published papers: Schmiedel et al., Cell, 175, 1701-1715, 2018; Nalls et al., Lancet Neurol., 18, 1091-1102, 2019; Blauwendraat et al., Mov Disord., 34, 866-875, 2019; Iwaki et al., Mov Disord., 34, 1839-1850, 2019; and Liu et al., Nature Genetics, 47, 979-986, 2015. Furthermore, the data used to construct the KG were obtained from publicly available databases (i.e., Entrez Gene, Gene Ontology, Wikipathways, Reactome, and / or Disease Ontology), and relationships were extracted from Hetionet and / or PrimeKG.
[0111] Immune cell type-specific eQTL summary statistics Filtered eQTL data for 15 purified human immune cell types (including B, CD4 naive, stimulated CD4, CD8 naive, stimulated CD8, classical monocytes, non-classical monocytes, NK, TFH, TH1, TH2, TH1 / 17, TH17, naive Treg, and Treg memory cells) were available in variant call format (VCF) through the DICE database. All data were mapped to the GRCh37 human reference genome. Genetic polymorphisms listed as having a significant association with the expression of neighboring genes (may include multiple genes) were pre-filtered based on significance as follows: corrected p-value <0.05, raw p-value <0.0001, and transcripts per million (TPM) >1.0. For information on how these data were generated, see Schmeidel et al., 2018 (op.).
[0112] To prepare these immune cell type-specific eQTLs for use as exposure variables in 2SMR analysis, the following data fields were extracted from all VCFs using Python (version 3.8.8): rsID, ensembleID, geneSymbol, pvalue (p<0.0001), beta, stat, and FDR. Furthermore, given the beta value and statistic (t-statistic), the standard error (SE) was calculated as follows: me$all$eqtls$beta_se=me$all$eqtls$beta / me$all$eqtls$statistic
[0113] The extracted eQTL summary statistics were saved in CSV format.
[0114] GWAS Summary Statistics All GWAS summary statistics were processed using R (versions 3.6.1-4.2.2) through the IDE RStudio (v2021.09.0-v2022.07.2) unless otherwise specified. While two of the three datasets required manual curation (described below), all final GWAS summary statistics outputs were mapped to the GRCh37 human reference genome and formatted according to the IEU OpenGWAS project guidelines. This included rsID, chromosome, location, beta, SE, sample size, p-value, effect allele frequency (EAF), effect allele, other allele, outcome, outcome ID, original outcome name, deprecated outcome, outcome retained in MR, and 16 data fields from the data source. All final GWAS summary statistics were saved in CSV format. The following describes the processing used to generate the final GWAS summary statistics.
[0115] GWAS summary statistics for Parkinson's disease (PD) incidence (Nalls et al., 2019 (ibid.)) and inflammatory bowel disease prevalence (Liu et al., 2015 (ibid.)) were processed directly through the IEU OpenGWAS project using the R modules TwoSampleMR (version 0.5.6) and ieugwasr (version 0.1.5). Specifically, statistically significant SNPs (p<0.0001) and their associated data fields were extracted from the PD dataset "ieu-b-7" using an ieugwasr access token.
[0116] The PD age-onset GWAS summary statistics (Blauwendraat et al., 2019 (cited above)) were available for local download as a txt file through the International Parkinson's Disease Genomics Consortium. Of the 15 original data fields provided, the following seven were selected for further processing: MarkerName (concatenation of chromosomal and positional information for each SNP), Allele1 (author-confirmed effect allele), Allele2, Freq1 (EAF), Effect (β), StdErr (SE), and P-value. Only entries with a P-value less than 0.0001 were retained. Allele1 and Allele2 data entries were reformatted to uppercase. Chromosomal and positional fields were extracted from "MarkerName". Sample size, outcome, outcome ID, original outcome name, deprecated outcome, outcome retained in MR, and data source fields were manually added to reflect the information described in the original paper. Because the original GWAS summary statistics data file did not contain information on rsIDs or reference alleles (i.e., the other allele), chromosome and position coordinates were lifted over from GRCh37 to the GRCh38 human reference genome using the CrossMap Python function (version 3.8.8) to ensure consistency between the identification of rsIDs and reference alleles using the following four methods. In the first method, the dbSNP GRCh38 reference file was used as a dictionary, and Python functions were used to identify rsIDs, reference alleles, and their associated GRCh38 coordinates for most SNPs. In the second method, if the rsID was not found in the dbSNP GRCh38 reference file, the rsID was identified using the R module SNPlocs.Hsapiens.dbSNP151.GRCh38 (version 0.99.23). In the third method, if the reference allele was not found in the dbSNP GRCh38 reference file, the reference allele was identified using the R module RMySQL (version 0.10.25).Finally, for the three SNPs for which rsIDs could not be found using either method, the information was manually entered by cross-referencing the dbSNP and Ensembl VEP websites.
[0117] The PD UPDRSIII GWAS summary statistics (Iwaki et al., 2019 (cited above)) were available for local download as a txt file via the PD Progression Meta-data GWAS Browser. Of the seven original data fields provided, the following five were selected for further processing: SNP (chromosome-position linkage), BETA, SE, P, and N (sample size). Only entries with P < 0.0001 were retained. The chromosome and position fields were extracted from "SNP". The sample size, outcome, outcome ID, original outcome name, deprecated outcome, outcome retained in MR, and data source fields were manually added to reflect the information described in the original paper. Similar to Blauwendraat et al., 2019 (cited above), Iwaki et al., 2019 (cited above) did not provide the original GWAS summary statistics, rsID, and reference allele (as well as effect allele). Therefore, these data were identified after lifting over genomic coordinates from GRCh37 to the GRCh38 human reference genome using one or more of the following methods: (1) cross-referencing of dbSNP GRCh38 reference files, (2) use of the R module SNPlocs.Hsapiens.dbSNP151.GRCh38, or (3) manual curation using the dbSNP and Ensembl VEP websites.
[0118] Two-sample Mendelian randomization (2SMR) Using the R module TwoSampleMR, 2SMR analyses were iteratively performed for each of the 15 immune cell types across three disease features. Briefly, shared SNPs were identified for each immune cell type and disease feature combination and stored in their respective lists. Next, CSV files of eQTLs extracted from immune cells were filtered based on the shared SNP lists and saved as new CSV files (e.g., PD_incidence_x_B_cell_naive_extracted_se.csv). Finally, the newly generated clinical_outcome_x_immune_cell_extracted_se.csv file was loaded as exposure data, the data was reconciled with disease feature GWAS summary statistics, and 2SMR analyses were performed. Each results summary was labeled with an immune cell type-specific label and contained the following 11 data fields: exposure(gene symbol), id.exposure(gene symbol), id.outcome, outcome, method, nsnp, b(β), se, pval, cell type, and SNP(rsID). Since only a single SNP was identified as a valid instrumental variable for each gene of interest in all runs, the Wald ratio method was used. Finally, the immune cell type-specific results for the disease features of interest were integrated and saved as a CSV file.
[0119] 2SMR visualization To visualize 2SMR results for each clinical outcome as a heatmap, the integrated immune cell type-specific results (including the variables exposure, cell type, and b) were reformatted into a wide-format matrix. Any "NA" values were replaced with "0" as placeholders to allow for hierarchical clustering. The heatmap was generated using the R module ComplexHeatmap (version 2.14.0).
[0120] Figures 6A-6D show the causal effect matrix, i.e., visualization of 2SMR, obtained for disease characteristics / onset outcomes of PD (Figure 6A), age of onset (Figure 6B), exercise score severity (Figure 6C), and the onset of inflammatory bowel disease (Figure 6D).
[0121] As mentioned above, using cell type-specific molecular weight phenotypic locus (molQTL) data allows for the accurate identification of causal effects at a cell type-specific level. For example, when whole blood-derived molecular weight phenotypic locus (molQTL) data is used (instead of cell type-specific data), the average effect size is observed for targets causally associated with multiple cell types, and this is dominated by the strongest effect, potentially obscuring differences in the magnitude and / or direction of the association. For instance, in the 2SMR for PD onset, KANSL1 showed positive and negative causal associations with PD risk depending on the cell type. Table II shows the relationship data for six different cell types and the overall mean for the gene KANLS1, extracted from Figure 6A. [Table II]
[0122] T cell lineage-specific knowledge graph The most numerous and robust causal relationships were identified between T cells and the three disease features of Parkinson's disease (PD). Therefore, T cells (including CD4 naive, stimulated CD4, CD8 naive, stimulated CD8, TFH, TH1, TH2, TH1_17, TH17, Treg naive, and Treg memory cells) were prioritized for further evaluation using knowledge graph (KG) analysis. While it is possible to generate and analyze KGs for individual T cell types, it was determined that evaluating all T cells would provide better insights because many T cell effector cells possess the ability / plasticity to change between subtypes depending on the microenvironment (compared to other immune cell lineages). To create T-cell lineage-specific knowledge graphs (KGs), gene features from open-source databases (e.g., Entrez Gene, Gene Ontology, Reactome, etc.) were removed if the gene was not present in any T-cell type-specific transcriptome profile (see Schmeidel et al., 2018 (cited above)), by using Hetionet or PrimeKG as the KG framework, focusing only on the metanodes and corresponding edges of "Biological Process," "Cellular Component," "Disease," "Gene," "Molecular Function," and / or "Pathway." For example, using this T-cell lineage-specific knowledge graph, a disease feature KG for PD motor score severity (UPDRSIII scoring) was then created, and 21 causative genes identified by 2SMR analysis were input. However, only 12 of these genes were identified within this KG (9 genes were not present in the KG used in this example).
[0123] Cell type-specific knowledge graph analysis and visualization Graph relationship calculations were generated using NetworkX (version 2.7) in Python (version 3.10). To generate a disease feature network of cell-specific connectivity for one or more disease features, the shortest paths between pairs of input genes (i.e., candidate targets / biomarkers) from the 2SMR results were calculated, and the analysis was restricted to include only the first neighbor or nodes connecting two causative genes with a path length of less than 3. The resulting network shows the connectivity between genes (nodes). To evaluate the specificity and significance of the genes (nodes) in this graph, a one-sample substitution analysis was performed based on the node degree and the corresponding false detection rate (FDR), which is calculated by comparing the node degree in the final graph with a random graph. Finally, the network was further refined to include only nodes with a p-value less than 0.05, and the node degree was calculated to determine the node with the highest interconnectivity in the graph. The graph was visualized using Gephi (version 00.9.7 202208031831).
[0124] Figure 7 shows the visualization of a cell type-specific knowledge graph obtained by applying the above method to a single disease feature. Figure 8 shows the visualization of a cell type-specific knowledge graph obtained by applying the above method to multiple disease features.
[0125] result This method provides a cell type-specific and direction-aware process for identifying candidate targets / biomarkers causally associated with disease features (e.g., clinical characteristics of PD patients). A total of 126 unique candidate targets / biomarkers were identified across three PD disease features (incidence, age of onset, and exercise score severity), and the directionality (positive or negative effect size) of these targets / biomarkers and their association with one or more specific immune cell populations was elucidated (Figures 6A-6C). Furthermore, for inflammatory bowel disease (IBD), 264 unique immune cell-specific and direction-aware candidate targets / biomarkers were identified, seven of which were shared with 2SMR results for one or more PD disease features (Figure 6D).
[0126] This method involved positioning causative genes (i.e., candidate targets / biomarkers) within cell type-specific key gardens (KGs), identifying common interacting factors (i.e., additional candidate targets / biomarkers), identifying augmented pathways, and ranking targets / biomarkers based on their degree of interconnectivity. For example, in a T-cell type-specific KG for motor score severity in PD, 12 causative genes were input, forming a network containing 80 KG genes and 49 biological pathways (see Figure 7). Genes ranked highly had 30 or more connections with other nodes in the KG and were identified as targets or biomarkers for PD. As another example, an assessment of shared biology between IBD and PD disease features identified 13 genes (one from 2SMR execution and 12 discovered through KGs) and 5 biological pathways (see Figure 8).
[0127] This method derives additional biological insights by positioning 2SMR results within cell type-specific key groups (KGs), thereby identifying broader biological alterations in pathways and / or molecular function (rather than focusing solely on a single gene / protein candidate target as the primary output). This allows for the identification of highly interconnected targets / biomarkers and confidently demonstrating their involvement in the disease. Such targets / biomarkers can therefore be considered potent, i.e., primary, targets / biomarkers of the disease. Furthermore, the identification of broader biological alterations in pathways and / or molecular function allows for the identification of common interactors with high interconnectivity to 2SMR results, which can be confidently considered targets / biomarkers of the disease.
[0128] These relationships are graphically illustrated by Figure 8, which shows an example of a graphical output for display via a graphical user interface (GUI) or display screen. The output (as shown in Figure 8) further represents a compact and useful representation of the cell network related to candidate targets and markers, which can be efficiently stored in computer memory (e.g., memory 906) as a cell type-specific knowledge graph. In some embodiments, the cell type-specific knowledge graph may include a link list, one or more database tables linked by key values, or other storage links designed to reduce memory usage of computer memory for efficient storage. As shown in the example in Figure 8, the inputs to the cell type-specific knowledge graph were PD incidence, PD age onset, PD motor score severity, and IBD. As shown in Figure 8, each of these values is linked via one or more genes or pathways, forming a compact and useful representation of the cell network.
[0129] Figure 9 shows an exemplary computing system for carrying out the method of the present disclosure. Specifically, Figure 9 shows a block diagram of one embodiment of the computing system according to an exemplary embodiment of the present disclosure. The computing system shown in Figure 9 may correspond to some or all of any of the functional units described above.
[0130] The arithmetic system 900 can be configured to perform any of the operations disclosed herein, such as any of the operations described with reference to the methods shown in Figures 1 and 2. The arithmetic system includes one or more computing devices 902. One or more arithmetic units 902 of the arithmetic system 900 comprises one or more processors 904 and a memory 906. One or more processors 904 can be any general-purpose processors configured to execute a set of instructions. For example, one or more processors 904 can be one or more general-purpose processors, one or more field-programmable gate arrays (FPGAs), and / or one or more application-specific integrated circuits (ASICs). In one embodiment, one or more processors 904 includes one processor. Alternatively, one or more processors 904 includes multiple operably connected processors. One or more processors 904 are communicably coupled to the memory 906 via an address bus 908, a control bus 910, and a data bus 912. The memory 906 may be random access memory (RAM), read-only memory (ROM), persistent storage devices such as hard drives, and / or erasable programmable read-only memory (EPROM). One or more arithmetic units 902 further include an I / O interface 914 that is communicatively coupled to an address bus 908, a control bus 910, and a data bus 912.
[0131] Memory 906 can store information that can be accessed by one or more processors 904. For example, memory 906 (e.g., one or more non-transient computer-readable storage media, memory devices) can include computer-readable instructions (not shown) that can be executed by one or more processors 904. These computer-readable instructions may be software written in any preferred programming language or implemented in hardware. Additionally or alternatively, computer-readable instructions may be executed logically and / or virtually in separate threads on one or more processors 904. For example, memory 906 can store instructions (not shown) that, when executed by one or more processors 904, cause one or more processors 904 to perform operations such as any of the operations and functions configured to be performed by the arithmetic system 900 as described herein. Additionally or alternatively, memory 906 can store data (not shown) that can be acquired, received, accessed, written, manipulated, created, and / or stored. In some embodiments, one or more arithmetic units 902 can retrieve data from one or more memory devices located away from the arithmetic system 900, and / or store data in one or more memory devices.
[0132] The arithmetic system 900 further includes a storage unit 916, a network interface 918, an input controller 920, and an output controller 922. The storage unit 916, the network interface 918, the input controller 920, and the output controller 922 are coupled to a central control unit (i.e., memory 906, address bus 908, control bus 910, and data bus 912) via an I / O interface 914.
[0133] The storage unit 916 is a computer-readable medium, preferably a non-transient computer-readable medium containing one or more programs, which, when executed by one or more processors 904, cause the arithmetic system 900 to perform the method steps of the present disclosure. Alternatively, the storage unit 916 is a transient computer-readable medium. The storage unit 916 can be a persistent storage device such as a hard drive, a cloud storage device, or any other suitable storage device.
[0134] The network interface 918 can be a Wi-Fi module, a network interface card, a Bluetooth module, and / or any other suitable wired or wireless communication device. In one embodiment, the network interface 918 is configured to connect to a network such as a local area network (LAN) or wide area network (WAN), the Internet, or an intranet.
[0135] The various embodiments and examples described above provide an overview for understanding the embodiments and examples of the disclosed method. The figures provided herein illustrate exemplary embodiments of the system and method and are not intended to limit the scope of this disclosure.
[0136] Unless otherwise specified, all technical terms used herein have the same meaning as those commonly understood by those skilled in the art. The singular forms “a,” “an,” and “the” include plural references unless the context of this disclosure expressly indicates otherwise. The term “or” includes “and / or” unless expressly indicated otherwise.
[0137] All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent that each individual publication, patent, or patent application is specifically and individually indicated to be incorporated by reference.
[0138] The nature of this disclosure The following aspects of this disclosure are illustrative and are not intended to limit the scope of this disclosure.
[0139] Embodiment 1. A method for identifying a primary target or biomarker for a disease, comprising: acquiring data on a first characteristic of the disease and a genetic risk variant associated with that characteristic by one or more processors; acquiring cell type-specific molecular weight trait locus (molQTL) data for a first cell type, wherein each molQTL is associated with a molecular characteristic, by one or more processors; determining one or more genetic risk variants that are causally related to the first characteristic of the disease and one or more associated molecular characteristics, by one or more processors; and identifying a plurality of candidate targets or biomarkers for the disease based on the one or more associated molecular characteristics, by one or more processors. A method comprising: identifying; generating a first cell type-specific knowledge graph containing a first plurality of nodes using one or more processors, wherein each of the first plurality of nodes represents one of the plurality of candidate targets or biomarkers, a molecular pathway associated with the candidate target or biomarker, or one or more common interactors; and calculating the level of interconnectivity for each node in the first cell type-specific knowledge graph using one or more processors, wherein a candidate target or biomarker, or a common interactor, represented by a node having a high level of interconnectivity in the first cell type-specific knowledge graph, is identified as a primary target or biomarker.
[0140] Embodiment 2. A method according to Embodiment 1, further comprising outputting the first cell type-specific knowledge graph by the one or more processors.
[0141] Embodiment 3. A method according to Embodiment 2, wherein the step of outputting the first cell type-specific knowledge graph includes storing the first cell type-specific knowledge graph in a memory location of a computing device.
[0142] Embodiment 4. A method according to Embodiment 2, wherein the step of outputting the first cell type-specific knowledge graph includes displaying a graphical representation of the first cell type-specific knowledge graph within a graphical user interface.
[0143] Embodiment 5. A method according to any one of the preceding embodiments, further comprising identifying a first candidate target or biomarker, or a first common interactor, by the one or more processors based on the level of interconnectivity of each node representing the first candidate target or biomarker, or the first common interactor.
[0144] Embodiment 6. A method according to Embodiment 5, further comprising outputting the first candidate target or biomarker, or the first common interaction factor, by the one or more processors.
[0145] Embodiment 7. A method according to any one of the preceding embodiments, wherein the causal relationship is determined using Mendelian randomization from data on the genetic risk variant and the molQTL data.
[0146] Embodiment 8. A method according to any one of the preceding embodiments, further comprising acquiring a second knowledge graph, comprising a second plurality of nodes, by one or more processors, wherein at least one node of the first plurality of nodes and at least one node of the second plurality of nodes are related.
[0147] Apparatus 9. The method according to Apparatus 8, wherein the second knowledge graph is associated with cell type-specific molecular weight trait locus (molQTL) data for a second cell type.
[0148] Embodiment 10. The method according to Embodiment 8, wherein the second knowledge graph relates to data on genetic risk variants associated with the second characteristic of the disease.
[0149] Aspect 11. The method according to Aspect 8, further comprising generating a third knowledge graph comprising a first cell type-specific knowledge graph and the second knowledge graph by one or more processors, wherein the third knowledge graph includes at least one edge connecting the at least one node of the first plurality of nodes and the at least one node of the second plurality of nodes, and a candidate target or biomarker, or common interacting factor, represented by nodes having a high level of interconnectivity in the third knowledge graph, is identified as a primary target or biomarker.
[0150] Embodiment 12. A method according to Embodiment 11, wherein each at least one edge connects nodes representing a candidate target or biomarker, a molecular pathway, or a common interacting factor involved.
[0151] Embodiment 13. A method according to any one of the preceding embodiments, wherein the first cell type-specific knowledge graph is generated from a parent knowledge graph.
[0152] Embodiment 14. The method according to Embodiment 13, wherein the first cell type-specific knowledge graph is generated by pruning the parent knowledge graph in accordance with locality constraints relating to nodes in the parent knowledge graph that represent the candidate target or biomarker.
[0153] Embodiment 15. The method according to Embodiment 14, wherein the locality constraint is a shortest path constraint such that the shortest path distance between pairs of nodes representing candidate targets or biomarkers in the first cell type-specific knowledge graph is less than a predetermined distance.
[0154] Embodiment 16. The method according to Embodiment 15, wherein the predetermined distance is less than 4.
[0155] Embodiment 17. A method according to any one of the preceding embodiments, wherein the level of interconnectivity of nodes is either the degree of the nodes or the intrinsic centrality of the nodes.
[0156] Embodiment 18. A method according to any one of the preceding embodiments, wherein a node is determined to have a high level of interconnectivity if the level of interconnectivity of the node falls within the upper quartile of the interconnectivity levels of all nodes in the first cell type-specific knowledge graph.
[0157] Embodiment 19. A method according to any one of the preceding embodiments, wherein the nodes in the first cell type-specific knowledge graph relating to a molecular pathway include nodes representing gene expression pathways and / or nodes representing protein-protein interaction pathways.
[0158] Embodiment 20. A method according to any one of the preceding embodiments, wherein the primary target is a target that can be targeted by therapy to treat or prevent the disease.
[0159] Embodiment 21. The method according to Embodiment 20, wherein the primary target is a nucleic acid sequence or a protein.
[0160] Embodiment 22. A method according to any one of the preceding embodiments, wherein the primary biomarker for the disease indicates the presence or severity of the disease.
[0161] Apparatus 23. The method according to Apparatus 22, wherein the primary biomarker is the presence or level of a nucleic acid sequence or protein, or the level of nucleic acid methylation or phosphorylation.
[0162] Embodiment 24. A method according to any one of the preceding embodiments, wherein the genetic risk mutation associated with the first characteristic of the disease is a single nucleotide polymorphism / mutation (SNP / SNV), insertion and / or deletion mutation, copy number mutation, translocation and / or inversion, or a combination thereof.
[0163] Embodiment 25. The method of Embodiment 24, wherein the genetic risk mutation is an SNP.
[0164] Embodiment 26. A method according to any one of the preceding embodiments, wherein the disease is a heart disease such as ischemic heart disease and coronary artery disease; cancer; bacterial or viral infection; cerebrovascular disease such as stroke; respiratory disease such as chronic obstructive pulmonary disease; diabetes; autoimmune diseases such as ulcerative colitis, Crohn's disease, inflammatory bowel disease, rheumatoid arthritis, Guillain-Barré syndrome, Sjögren's syndrome, scleroderma, and Graves' disease; or a neurological disease such as Alzheimer's disease, multiple sclerosis, Parkinson's disease, amyotrophic lateral sclerosis (ALS), and frontotemporal dementia.
[0165] Embodiment 27. A method according to any one of the preceding embodiments, wherein the molQTL data relates gene mutations to molecular features, and the gene mutations are selected from the group consisting of single nucleotide polymorphisms / mutations (SNPs / SNVs), insertions and / or deletions, copy number variations, translocations and / or inversions, or combinations thereof.
[0166] Embodiment 28. The method according to Embodiment 27, wherein the genetic risk mutation is an SNP.
[0167] Embodiment 29. A method according to any one of the preceding embodiments, wherein the molQTL data is cell type-specific expression level trait locus (eQTL) data, cell type-specific DNA methylation level trait locus (meQTL) data, cell type-specific histone modification level trait locus (hQTL) data, cell type-specific chromatin accessibility level trait locus (caQTL) data, cell type-specific alternative gene splicing level trait locus (sQTL) data, cell type-specific protein level level trait locus (pQTL) data, cell type-specific microRNA expression level trait locus (mirQTL) data, or cell type-specific ribosome occupation level trait locus (rQTL) data.
[0168] Embodiment 30. A method according to any one of the preceding embodiments, wherein, prior to the step of determining the causal relationship, the data of the one or more genetic risk variants are harmonized with the molQTL data.
[0169] Embodiment 31. A method according to any one of the preceding embodiments, wherein the candidate target or biomarker for the disease is a nucleic acid sequence or protein relating to one or more associated molecular features.
[0170] Embodiment 32. A method according to Embodiment 31, wherein one or more relevant molecular features are nucleic acid sequences and / or proteins, and the candidate target or biomarker of the disease is the nucleic acid sequence and / or protein.
[0171] Embodiment 33. A method according to any one of the preceding embodiments, wherein the knowledge graph is limited to a first cell type from which the molQTL data was obtained, or to a group of cell types including the first cell type from which the molQTL data was obtained and cell types having the ability to transform into the first cell type.
[0172] Embodiment 34. A method for treating a subject suffering from a disease, wherein a primary target identified by the method of any one of the preceding embodiments is targeted by the therapy for the treatment or prevention of the disease.
[0173] Embodiment 35. A method for determining whether a subject has a disease, wherein the method comprises determining the presence of a major biomarker in a sample obtained from the subject, the major biomarker being identified by the method of any one of the preceding embodiments.
[0174] Embodiment 36. A system comprising one or more processors and a memory for storing instructions, wherein when an instruction is executed by one or more processors, the system causes one or more processors to acquire data on genetic risk mutations associated with a first characteristic of the disease, to acquire molQTL data for a first cell type, wherein each of the molQTLs is associated with a molecular characteristic, to determine one or more genetic risk mutations that have a causal relationship with the first characteristic of the disease and one or more associated molecular characteristics, and to determine a plurality of candidate targets for the disease based on the one or more associated molecular characteristics. A system in which a biomarker is identified by one or more processors, a first cell type-specific knowledge graph including a first plurality of nodes is generated by one or more processors, each of the first plurality of nodes represents one of the plurality of candidate targets or biomarkers, a molecular pathway associated with the candidate target or biomarker, or one or more common interactors, and the level of interconnectivity for each node in the first cell type-specific knowledge graph is calculated by one or more processors, and a candidate target or biomarker, or a common interactor, represented by a node with a high level of interconnectivity in the first cell type-specific knowledge graph is identified as the primary target or biomarker.
[0175] Embodiment 37. A non-transient computer-readable medium for storing non-transient instructions, wherein when an instruction is executed by a processing unit including one or more processors, the one or more processors are caused to acquire data on genetic risk mutations associated with a first characteristic of the disease; to acquire molQTL data for a first cell type, wherein each of the molQTLs is associated with a molecular characteristic; to determine one or more genetic risk mutations that have a causal relationship with the first characteristic of the disease and one or more associated molecular characteristics; and based on the one or more associated molecular characteristics, a plurality of candidate targets or batches of the disease. A non-transient, computer-readable medium that identifies an iomarker, generates a first cell type-specific knowledge graph containing a first plurality of nodes, each of the first plurality of nodes representing one of the plurality of candidate targets or biomarkers, a molecular pathway associated with the candidate target or biomarker, or one or more common interactors, calculates the level of interconnectivity for each node in the first cell type-specific knowledge graph, and identifies candidate targets or biomarkers, or common interactors, represented by nodes with a high level of interconnectivity in the first cell type-specific knowledge graph, as the primary target or biomarker.
[0176] Embodiment 38. A method for identifying a primary target or biomarker for a disease, comprising: acquiring data on genetic risk mutations associated with multiple characteristics of the disease using one or more processors; acquiring molQTL data for multiple cell types, wherein each molQTL is associated with a molecular characteristic, using one or more processors; determining a matrix of gene mutations causally related to one or more of the multiple characteristics of the disease and one or more associated molecular characteristics for the multiple cell types, using one or more processors; and identifying multiple candidate targets or biomarkers for the disease based on the one or more associated molecular characteristics, using one or more processors. A method comprising: identifying; generating a first cell type-specific knowledge graph, comprising one or more processors, comprising: each of the first multiple nodes representing one of the multiple candidate targets or biomarkers, a molecular pathway associated with the candidate target or biomarker, or one or more common interactors; and calculating the level of interconnectivity for each node in the first cell type-specific knowledge graph, wherein a candidate target or biomarker, or a common interactor, represented by a node having a high level of interconnectivity in the first cell type-specific knowledge graph, is identified as a primary target or biomarker.
Claims
1. A method for identifying key targets or biomarkers for a disease, The data on the first characteristic of the aforementioned disease and the associated genetic risk mutations are obtained by one or more processors, The method involves acquiring molQTL data for a first cell type, wherein each molQTL is associated with a molecular characteristic, using one or more processors. The first characteristic of the disease and one or more related molecular characteristics, and one or more genetic risk mutations having a causal relationship with them, are determined by the one or more processors. Based on the aforementioned one or more related molecular characteristics, multiple candidate targets or biomarkers for the disease are identified by the aforementioned one or more processors. The method involves generating a first cell type-specific knowledge graph, comprising a plurality of first nodes, using one or more processors, wherein each of the plurality of first nodes represents one of the plurality of candidate targets or biomarkers, a molecular pathway associated with the candidate target or biomarker, or one or more common interacting factors. This includes calculating the level of interconnectivity for each node in the first cell type-specific knowledge graph using one or more processors, A method for identifying candidate targets or biomarkers, or common interacting factors, represented by nodes with a high level of interconnectivity within the first cell type-specific knowledge graph, as primary targets or biomarkers.
2. The method according to claim 1, A method further comprising outputting the first cell type-specific knowledge graph by the one or more processors.
3. A method according to claim 2, wherein the step of outputting the first cell type-specific knowledge graph includes storing the first cell type-specific knowledge graph in a memory location of a computing device.
4. A method according to claim 2, wherein the step of outputting the first cell type-specific knowledge graph includes displaying a graphical representation of the first cell type-specific knowledge graph within a graphical user interface.
5. The method according to claim 1, A method further comprising identifying a first candidate target or biomarker, or a first common interactor, by one or more processors based on the level of interconnectivity of each node representing the first candidate target or biomarker, or the first common interactor.
6. The method according to claim 5, A method further comprising outputting the first candidate target or biomarker, or the first common interaction factor, by one or more processors.
7. A method according to claim 1, wherein the causal relationship is determined using Mendelian randomization from data on the genetic risk variant and the molQTL data.
8. The method according to claim 1, A method for acquiring a second knowledge graph, which includes a second plurality of nodes, by one or more processors, wherein at least one node of the first plurality of nodes and at least one node of the second plurality of nodes are related.
9. A method according to claim 8, wherein the second knowledge graph is associated with cell type-specific molecular weight trait locus (molQTL) data for a second cell type.
10. A method according to claim 8, wherein the second knowledge graph relates to data on genetic risk variants associated with the second characteristic of the disease.
11. The method according to claim 8, The process further includes generating a third knowledge graph, comprising a first cell type-specific knowledge graph and the second knowledge graph, by one or more processors, wherein the third knowledge graph includes at least one edge connecting the at least one node of the first plurality of nodes and the at least one node of the second plurality of nodes. A method for identifying candidate targets or biomarkers, or common interacting factors, represented by nodes with a high level of interconnectivity in the third knowledge graph, as primary targets or biomarkers.
12. A method according to claim 11, wherein each at least one edge connects nodes representing a related candidate target or biomarker, a related molecular pathway, or a related common interacting factor.
13. A method according to claim 1, wherein the first cell type-specific knowledge graph is generated from a parent knowledge graph.
14. A method according to claim 13, wherein the first cell type-specific knowledge graph is generated by pruning the parent knowledge graph in accordance with locality constraints relating to nodes in the parent knowledge graph that represent the candidate target or biomarker.
15. A method according to claim 14, wherein the locality constraint is a shortest path constraint such that the shortest path distance between pairs of nodes representing candidate targets or biomarkers in the first cell type-specific knowledge graph is less than a predetermined distance.
16. The method according to claim 15, wherein the predetermined distance is less than 4.
17. A method according to claim 1, wherein the level of interconnectivity of nodes is either the order of the nodes or the intrinsic center of the nodes.
18. A method according to claim 1, wherein a node is determined to have a high level of interconnectivity if the level of interconnectivity of the node falls within the upper quartile of the interconnectivity levels of all nodes in the first cell type-specific knowledge graph.
19. A method according to claim 1, wherein the nodes in the first cell type-specific knowledge graph relating to a molecular pathway include nodes representing gene expression pathways and / or nodes representing protein-protein interaction pathways.
20. A method according to claim 1, wherein the primary target is a target that can be targeted by therapy to treat or prevent the disease.
21. The method according to claim 20, wherein the primary target is a nucleic acid sequence or a protein.
22. A method according to claim 1, wherein the primary biomarker for the disease indicates the presence or severity of the disease.
23. The method according to claim 22, wherein the primary biomarker is the presence or level of a nucleic acid sequence or protein, or the level of nucleic acid methylation or phosphorylation.
24. A method according to claim 1, wherein the genetic risk mutation associated with the first characteristic of the disease is a single nucleotide polymorphism / mutation (SNP / SNV), insertion and / or deletion mutation, copy number mutation, translocation and / or inversion, or a combination thereof.
25. The method according to claim 24, wherein the genetic risk mutation is an SNP.
26. The method according to claim 1, wherein the disease is a heart disease such as ischemic heart disease and coronary artery disease; cancer; bacterial or viral infection; cerebrovascular disease such as stroke; respiratory disease such as chronic obstructive pulmonary disease; diabetes; autoimmune diseases such as ulcerative colitis, Crohn's disease, inflammatory bowel disease, rheumatoid arthritis, Guillain-Barré syndrome, Sjögren's syndrome, scleroderma, and Graves' disease; or a neurological disease such as Alzheimer's disease, multiple sclerosis, Parkinson's disease, amyotrophic lateral sclerosis (ALS), and frontotemporal dementia.
27. A method according to claim 1, wherein the molQTL data associates gene mutations with molecular features, and the gene mutations are selected from the group consisting of single nucleotide polymorphisms / mutations (SNP / SNV), insertions and / or deletions, copy number variations, translocations and / or inversions, or combinations thereof.
28. The method according to claim 27, wherein the genetic risk mutation is an SNP.
29. The method according to claim 1, wherein the molQTL data is cell type-specific expression level trait locus (eQTL) data, cell type-specific DNA methylation level trait locus (meQTL) data, cell type-specific histone modification level trait locus (hQTL) data, cell type-specific chromatin accessibility level trait locus (caQTL) data, cell type-specific alternative gene splicing level trait locus (sQTL) data, cell type-specific protein level level trait locus (pQTL) data, cell type-specific microRNA expression level trait locus (mirQTL) data, or cell type-specific ribosome occupation level trait locus (rQTL) data.
30. A method according to claim 1, wherein, prior to the step of determining the causal relationship, the data of the one or more genetic risk variants are harmonized with the molQTL data.
31. A method according to claim 1, wherein the candidate target or biomarker of the disease is a nucleic acid sequence or protein relating to one or more associated molecular features.
32. A method according to claim 31, wherein one or more relevant molecular features are nucleic acid sequences and / or proteins, and the candidate target or biomarker of the disease is the nucleic acid sequence and / or protein.
33. A method according to claim 1, wherein the knowledge graph is limited to a first cell type from which the molQTL data was obtained, or a group of cell types including the first cell type from which the molQTL data was obtained and cell types having the ability to transform into the first cell type.
34. A method according to claim 1, wherein the primary target identified in the first cell type-specific knowledge graph is targeted by therapy to treat or prevent the disease in a subject suffering from the disease.
35. A method according to claim 1, further comprising determining the presence of a major biomarker in a sample obtained from a subject, wherein the major biomarker of the subject is identified in the first cell type-specific knowledge graph.
36. It is a system, One or more processors, A memory that stores instructions, and the memory, when executed by one or more processors, Data on genetic risk mutations associated with the first characteristic of the disease is obtained by one or more processors. Cell type-specific molecular weight trait locus (molQTL) data for a first cell type, wherein each molQTL is associated with a molecular feature, is acquired by one or more processors. The first characteristic of the disease and one or more related molecular characteristics, and one or more genetic risk mutations having a causal relationship with them, are determined by the one or more processors. Based on the one or more relevant molecular characteristics, the one or more processors identify multiple candidate targets or biomarkers for the disease. A first cell type-specific knowledge graph, comprising a first plurality of nodes, is generated by one or more processors, each of which represents one of the plurality of candidate targets or biomarkers, a molecular pathway associated with the candidate target or biomarker, or one or more common interacting factors, and The level of interconnectivity for each node in the first cell type-specific knowledge graph is calculated by one or more processors. A system in which candidate targets or biomarkers, or common interacting factors, represented by nodes with a high level of interconnectivity within the first cell type-specific knowledge graph, are identified as primary targets or biomarkers.
37. A non-transient computer-readable medium for storing instructions, wherein when an instruction is executed by a processing unit including one or more processors, the one or more processors... To obtain data on the first characteristic of the aforementioned disease and the associated genetic risk mutations, Cell type-specific molecular weight trait locus (molQTL) data for a first cell type, wherein each of the molQTLs is associated with a molecular characteristic, thereby obtaining molQTL data. Determine one or more genetic risk mutations that have a causal relationship with the first characteristic of the disease and one or more related molecular characteristics. Based on the aforementioned one or more related molecular characteristics, multiple candidate targets or biomarkers for the disease are identified. A first cell type-specific knowledge graph is generated, comprising a first set of nodes, each of which represents one of the multiple candidate targets or biomarkers, a molecular pathway associated with the candidate target or biomarker, or one or more common interacting factors, and The level of interconnectivity is calculated for each node in the first cell type-specific knowledge graph. A non-transient, computer-readable medium in which candidate targets or biomarkers, or common interacting factors, represented by nodes with a high level of interconnectivity within the first cell type-specific knowledge graph, are identified as primary targets or biomarkers.
38. A method for identifying key targets or biomarkers for a disease, The data on genetic risk mutations associated with multiple characteristics of the aforementioned disease is acquired by one or more processors. The process involves acquiring cell type-specific molecular weight trait locus (molQTL) data for multiple cell types, where each molQTL is associated with a molecular characteristic, using one or more processors. Determining, by one or more processors, a matrix of gene mutations having a causal relationship with one or more of the aforementioned multiple characteristics of the disease and one or more related molecular characteristics of the aforementioned multiple cell types, Based on the aforementioned one or more related molecular characteristics, multiple candidate targets or biomarkers for the disease are identified by the aforementioned one or more processors. The method involves generating a first cell type-specific knowledge graph, comprising a plurality of first nodes, using one or more processors, wherein each of the plurality of first nodes represents one of the plurality of candidate targets or biomarkers, a molecular pathway associated with the candidate target or biomarker, or one or more common interacting factors. This includes calculating the level of interconnectivity for each node in the first cell type-specific knowledge graph using one or more processors, A method for identifying candidate targets or biomarkers, or common interacting factors, represented by nodes with a high level of interconnectivity within the first cell type-specific knowledge graph, as primary targets or biomarkers.