Systems and methods to evaluate condensate knowledge graphs
By integrating genomic, transcriptomic, and proteomic data through gene-network knowledge graphs, the methods identify condensate-related target genes, addressing the limitations of existing studies and enabling novel drug targets and biomarkers for disease treatment.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- DEWPOINT THERAPEUTICS INC
- Filing Date
- 2026-01-22
- Publication Date
- 2026-07-30
AI Technical Summary
Genomic and transcriptomic studies fail to provide insights into how phase separation of genes into condensates relates to disease mechanisms, missing disease-relevant genes that can be identified through a condensate lens, hindering understanding of disease mechanisms and treatment development.
Computer-implemented methods using gene-network knowledge graphs integrate genomic, transcriptomic, and proteomic data to identify condensate-related target genes by analyzing phase separation in disease contexts, calculating integration values, and incorporating condensate information to prioritize and validate target genes.
Enables rapid and high-throughput identification of novel condensate-related target genes, uncovering new condensates associated with diseases and treatment mechanisms, facilitating drug screening and biomarker development for personalized medicine.
Smart Images

Figure US2026012200_30072026_PF_FP_ABST
Abstract
Description
Docket No.: 185992002640SYSTEMS AND METHODS TO EVALUATE CONDENSATE KNOWLEDGE GRAPHSCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority benefit of United Stated Provisional Patent Application Serial No. 63 / 748,738, filed January 23, 2025, the contents of which are incorporated herein by reference in their entirety.FIELD
[0002] The present application relates to, in certain aspects, computer implemented methods and systems for identifying one or more condensate related target gene for a disease using an integrated data structure. In certain aspects, the application also relates to assessing (such as measuring) phase separation in disease contexts and the integration of such information to guide selection techniques for one or more condensate related target genes, and methods for validating condensate related target genes.BACKGROUND
[0003] Condensates are membrane-less molecular assemblies formed through liquid-liquid phase separation and they enable key biochemical processes. Condensates present a new avenue by which disease can be explained and treatments with mechanisms of actions through modification in condensate phenotypes can be discovered.
[0004] Genomic and transcriptomic studies have identified genes associated with diseases and have started to uncover high-level mechanisms driving disease phenotypes. For example, genome wide associate studies (GWAS) identify genes associated with disease by assessing correlation between natural genetic variation and disease phenotypes. Associations between the same genetic variants and variation in gene expression in a diseased tissue can help explain the connection between the gene and the disease. However, these studies do not provide the necessary insight to understand how phase separation of the genes, or expression products thereof, into condensates plays a role. Further, these studies may miss disease relevant genes that can only be understood from a condensate lens. Thus, methods for identification of potential condensate forming genes that are relevant in a disease context is a necessary step for better understanding disease mechanisms and finding novel treatments for disease through, e.g., modification of condensate phenotypes.1MF-366315044Docket No.: 185992002640BRIEF SUMMARY
[0005] Provided herein are methods for identifying and validating genes as condensate related target genes. Also provided herein are high throughput methods for obtaining condensate information about a gene in a disease model by assessing phase separation in disease contexts and the use of the condensate information in identifying condensate related target genes. A condensate related target gene may be defined according to a change in phenotype between healthy and disease states, wherein a condensate forms in at least one of the phenotypes. It is appreciated that a condensate may form in both states and the difference in the condensate phenotype may be the size or location of the condensate. As a non-limiting example, in a cell in a healthy state, the protein product of the gene may not phase separate into a condensate but in a disease state the protein product does phase separate into a condensate. A condensate related target gene can be described as a gene wherein modification of the gene (at the DNA, RNA, or protein level) contributes to a change in a condensate phenotype for a condensate comprising the target gene or another gene in disease and non-disease state. The DNA, RNA, and / or protein product of a condensate related target gene may be in the condensate in one or both states or the effect may be secondary wherein a modification to the condensate related target gene leads to the DNA, RNA and / or protein of another gene to move in and out of a condensate.
[0006] Together, the identification of condensate related target genes can help to identify condensates that play a role in disease formation or treatment mechanisms. By anchoring the analysis on condensate related target genes rather than known condensates, the methods described herein can be used to identify new condensates or elucidate previously unappreciated roles that a known condensate plays in a disease. For example, the methods described herein can also be used to identify new gene targets for use in high throughput drug screening designed to identify drug compounds that are capable of changing disease condensate phenotypes to healthy condensate phenotypes.
[0007] The methods described herein can be used to identify and validate condensate related target genes, in a high throughput, multidimensional, and user-friendly way. The methods use graphical database solutions to store and analyze genetic and transcriptomic information about a disease and methods for identifying genes that phase separate in disease states. Once a knowledge graph has been constructed for a disease and condensate information obtained from analyzing phase separation in the disease state are integrated, graphical user interfaces are2MF-366315044Docket No.: 185992002640created to aid a user in traversing the knowledge graph to select condensate related target genes for validation.
[0008] Provided herein are computer implemented methods for identifying one or more condensate related target genes for a disease, comprising: generating a gene-network knowledge graph comprising at least a plurality of nodes, wherein the plurality of nodes comprise at least one gene node, at least one disease node, and edges representing a relationship between each node; calculating an integration value for one or more gene nodes in the genenetwork knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes of the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph; obtaining condensate information for one or more gene nodes in the gene-network knowledge graph; identifying one or more condensate related target genes from the gene-network knowledge graph using the condensate information and the integration value for the corresponding gene node.
[0009] In some aspects, the plurality of nodes further comprises at least one gene pathway node. In some aspects, the plurality of nodes further comprises at least one condensate node. In some aspects, the plurality of nodes further comprises at least one drug node.
[0010] In some aspects, the relationship between each node comprises a relationship selected from a group consisting of gene-disease association, a protein-protein association, an RNA-protein association, a protein-pathway interaction, a tissue specific expression relationship, a drug-gene interaction, a gene-condensate interaction, and an RNA-condensate interaction. In some aspects, the relationship between each node comprises a relationship between two nodes obtained from results of a statistical genomics analysis a statistical transcriptomics analysis, or a proteomic analysis. In some aspects, the statistical genomics analysis comprises GWAS, eQTL, pQTL, PheWAS, Polygenic gene prioritization (POP), or rare variant aggregate testing. In some aspects, the statistical transcriptomics analysis comprises differential gene expression, pathway level interference, gene dependence analysis, gene perturbation analysis, foundational model gene interaction mapping, or gene regulatory network mapping.
[0011] In some aspects, calculating an integration value comprises weighting the number of edges connecting the corresponding gene node to a node in one or more selected nodes the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph by characteristics of3MF-366315044Docket No.: 185992002640the edges connecting each node. In some aspects, calculating an integration value comprises performing a personalized page rank analysis.
[0012] In some aspects, the one or more selected nodes have been selected based on a known relationship to the disease. In some aspects, the one or more selected nodes comprise one or more nodes connected to the disease node associated with the disease by less than a predetermined number of edges. In some aspects, the predetermined number of edges is 1, 5, 10, 15, or 20. In some aspects, the integration value represents connectivity and selectivity of the gene node to the disease node representing the disease.
[0013] In some aspects, the condensate information comprises predicted phase separation for each of one or more polypeptides corresponding to one or more gene nodes. In some aspects, the predicted phase separation is generated from a protein sequence corresponding to the gene node.
[0014] In some aspects, the condensate information comprises a disease state phase transition characteristic for each of one or more polypeptides corresponding to one or more gene nodes. In some aspects, the condensate information has been obtained using a method comprising: determining the disease state phase transition characteristic for each of the one or more polypeptides.
[0015] In some aspects, the determining the disease state phase transition characteristic for a first polypeptide of the one or more polypeptides comprises performing one or more of the following comparisons: the quantity of the first polypeptide in an insoluble fraction of a disease model cell lysate as compared to the quantity of the first polypeptide in an insoluble fraction of a non-disease model cell lysate; or the quantity of the first polypeptide in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the first polypeptide in a soluble fraction of a non-disease model cell lysate versus a nondisease model cell lysate. In some aspects, the methods further comprise quantifying the first polypeptide to perform the one or more comparisons. In some aspects, the methods further comprise quantifying the first polypeptide in the insoluble fraction of the disease model cell lysate and the first polypeptide in the insoluble fraction of the non-disease model cell lysate. In some aspects, the methods further comprise quantifying the first polypeptide in the soluble fraction of the disease model cell lysate and the first polypeptide in the soluble fraction of the non-disease model cell lysate.4MF-366315044Docket No.: 185992002640
[0016] In some aspects, the methods further comprise quantifying the first polypeptide in the disease model cell lysate and the first polypeptide in the non-disease model cell lysate. In some aspects, the quantifying comprises performing a quantitative mass spectrometry technique. In some aspects, the quantitative mass spectrometry technique comprises use of isobaric labeling. In some aspects, the quantitative mass spectrometry technique comprises use of tandem mass tag (TMT) or isobaric tags for relative and absolute quantification (iTRAQ). In some aspects, the quantitative mass spectrometry technique is multiplexed.
[0017] In some aspects, the methods further comprise fractionating a disease model cell lysate and / or a non-disease model cell lysate. In some aspects, the methods further comprise obtaining a cell lysate from the disease model and / or the non-disease model. In some aspects, the methods further comprise lysing, separately, a sample from the disease model to obtain the disease model cell lysate and a sample from the non-disease model to obtain the non-disease model cell lysate.
[0018] In some aspects, the disease model comprises a cell model for a disease in a disease state and / or the non-disease model comprises a cell model for a disease in a non-disease state.
[0019] In some aspects, one or more condensate related target genes comprises identifying one or more gene nodes with a significant integration value. In some aspects, the one or more condensate related target genes comprises identifying one or more gene nodes that cluster together based on integration values and condensate information.
[0020] In some aspects, identifying one or more condensate related target genes comprises identifying one or more gene nodes using condensate phenotype information. In some aspects, identifying one or more condensate related target genes comprises identifying one or more gene nodes using genetic perturbation information. In some aspects, identifying one or more condensate related target genes comprises identifying one or more gene nodes using compound modulating condensates information.
[0021] Also provided herein are methods of identifying one or more condensate related target genes for a disease, the method comprising: generating a gene-network knowledge graph comprising at least a plurality of nodes, wherein the plurality of nodes comprise at least one gene node, at least one disease node, and edges representing a relationship between each node; calculating an integration value for one or more gene nodes in the gene-network knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes the gene-network knowledge graph and the number of edges connecting5MF-366315044Docket No.: 185992002640the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph; obtaining a disease state phase transition characteristic for one or more gene nodes in the gene-network knowledge graph by a method comprising: obtaining a disease model cell lysate and a non-disease model cell lysate; fractionating, separately, the disease model cell lysate and the non-disease model cell lysate to obtain respective insoluble and soluble fractions thereof; performing a quantitative mass spectrometry technique on one or more of: the soluble fraction of the disease model cell lysate, the disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate; or the insoluble fraction of the disease model cell lysate, the soluble fraction of the disease model cell lysate, the disease model cell lysate, the insoluble fraction of the non-disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate; determining a disease state phase transition characteristic for each of one or more polypeptides corresponding to the one or more gene nodes based on one or more of the following comparisons: the quantity of a polypeptide in an insoluble fraction of a disease model cell lysate as compared to the quantity of the polypeptide in an insoluble fraction of a non-disease model cell lysate; or the quantity of a polypeptide in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the polypeptide in a soluble fraction of a non-disease model cell lysate versus a non-disease model cell lysate; and identifying one or more condensate related target genes from the gene-network knowledge graph using the disease state phase transition characteristic and the integration value for the corresponding gene node.
[0022] Also provided herein are methods identifying a one or more condensate related target genes for a disease, the method comprising: obtaining a disease model cell lysate and a non-disease model cell lysate; fractionating, separately, the disease model cell lysate and the non-disease model cell lysate to obtain respective insoluble and soluble fractions thereof; performing a quantitative mass spectrometry technique on one or more of: the insoluble fraction of the disease model cell lysate and the insoluble fraction of the non-disease model cell lysate; the soluble fraction of the disease model cell lysate, the disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate; or the insoluble fraction of the disease model cell lysate, the soluble fraction of the disease model cell lysate, the disease model cell lysate, the insoluble fraction of the non-disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate; determining a disease environment phase transition characteristic for each of one6MF-366315044Docket No.: 185992002640or more polypeptides based on one or more of the following comparisons: the quantity of a polypeptide in an insoluble fraction of a disease model cell lysate as compared to the quantity of the polypeptide in an insoluble fraction of a non-disease model cell lysate; or the quantity of a polypeptide in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the polypeptide in a soluble fraction of a non-disease model cell lysate versus a non-disease model cell lysate; and filtering a gene-network knowledge graph using the disease environment phase transition characteristics to identify the one or more condensate related target genes from the disease.
[0023] Also provided herein are systems and computer-readable non-transitory storage mediums with instructions for performing the computer implemented methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] FIG. 1 describes an example flowchart describing a method for identifying one or more condensate related target gene for a disease according to various embodiments.
[0025] FIG. 2 describes an example flowchart describing a method for determining a cell fraction profile for a disease.
[0026] FIG. 3 illustrates an example system in accordance with various embodiments.
[0027] FIG. 4 illustrates an example computer system used to implement some or all of the techniques described herein.
[0028] FIG. 5 illustrates an exemplary section of gene network knowledge graph according to various embodiments. The disease node is represented as a black circle, genes nodes are represented by striped circles, and the checkered circles represent condensate nodes. The arrows represent the edges.
[0029] FIG. 6 illustrates a heatmap of gene nodes clustered according to aspects of the T2D knowledge graph.
[0030] FIG. 7 illustrates perturbation scores for a subset of condensate related target genes generated for healthy and / or disease cells from the T2D disease model.
[0031] FIGs. 8A-8B illustrate genes on the CRC gene network knowledge graph and condensate prediction scores. FIG. 8A illustrates genes plotted by number of directly connected CRC genes and condensate prediction scores. FIG. 8B illustrates genes plotted by a7MF-366315044Docket No.: 185992002640personalized page rank calculated from the CRC gene network knowledge graph and condensate prediction scores.
[0032] FIGs.9A-9B illustrates the relationship between compounds and Bcat in the CRC gene network knowledge graph. FIG. 9A illustrates a section of the CRC gene network knowledge graph containing Bcat and three compounds. FIG. 9B illustrates the edges between the drug nodes and the gene nodes in a table.
[0033] FIGs. 10A-10B illustrates the results on Bcat condensates following treatment of HCT116 cells. FIG. 10A shows HCT116 cells marked with a Bcat condensate marker following treatment with DMSO, Capecitabine, NCB-0846, and SSTC3 at 30µM. FIG. 10B illustrates a dose response curve for Bcat condensates in HCT116 cells treated with NCB-0846 and SSTC3 for 4 hours and 24 hours.DETAILED DESCRIPTION
[0034] Provided herein are computer implemented methods that leverage multi-omics data, phase separation information in disease relevant contexts and complex data structures to identify condensate related genes for a disease. As described herein, the methods incorporate condensate information in multiple forms and at a variety of steps allowing for unique discoveries about disease condensatopathy.
[0035] Condensate information may be a known condensate or how likely a macromolecule associated with a gene is to phase separate and / or form a condensate. The condensate information may be related to a disease such that the condensate information is how likely a macromolecule associated with a gene is to phase separate in a disease state and / or how likely characteristics associated with the phase separation are to be different in a disease and healthy state. Condensate information may come from literature searches, models that predict phase separation from amino acid sequences or from proteomics experiments, such as those described herein. The variability in condensate information and the formats that it may come in requires complex data structures. The methods described herein use knowledge graphs to incorporate condensate information in multiple ways. The condensate information may be stored in nodes and edges of the gene-network knowledge graphs described herein. The condensate information may also be used in analyzing the information stored in the gene-network knowledge graphs in order to identify the gene nodes most likely to be condensate related target genes.8MF-366315044Docket No.: 185992002640
[0036] Such methodology is based on, at least in part, the inventors surprising findings of over 60 novel condensate target genes for insulin resistance in type-2 diabetes and unique perspectives for leveraging massive quantities of information contained in different formats to provide confident identification of condensate related genes for a disease that can then be subsequently validated. The findings were made possible in part by using the methods described herein for identification and understanding of biomolecular interactions and disease links that are not identifiable without a condensate lens.
[0037] The identification of novel condensate related target genes has allowed for the identification of new condensates that are associated with disease as well as new disease associations for previously identified condensates. Disease associated condensates can be further studies to understand the role the condensate plays in a disease or treatment mechanism. For example, a new condensate may form the basis for a drug screening assay or the molecular community of the condensate can be studied to understand and find new drug targets. The condensate approach may unlock polygenic drug targets that differ from the conventional one target one drug approach often used in the pharmaceutical drug development. Condensate related target genes or condensates may also provide novel biomarkers for the disease. Biomarkers can be used to identify subpopulations of individuals with a disease who may respond more favorably to a treatment. Accordingly, the methods described herein and the resulting discoveries open the door for new diagnostics, prognostics, therapeutics and companion therapeutics.
[0038] Described are systems and methods that utilized computer generated knowledge graphs that connect disease relevant information with condensate information. These computer generate knowledge graphs allow for rapid and high throughput identification of condensate related target genes. The knowledge graph data structures are used to represent information extracted from different data sources and in different data structures because the nodes and edges are flexibly designed. Once genes, diseases, drugs, condensates, and gene networks are organized as a knowledge graph, the relationships between the various data can be understood using the spatial relationship between nodes and edges.
[0039] Prioritization and validation of condensate related target genes using the system and methods provided herein, provide novel targets for treating the disease. The methods in part rely on novel methods for using proteomics to analyze condensates in a disease state. Integrating solubility information about a gene in a disease context provides the condensate lens previously missing from other methods used to identify druggable targets for disease.9MF-366315044Docket No.: 185992002640I. Definitions
[0040] Unless otherwise defined, all of the technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field to which this disclosure belongs.
[0041] As used herein, “condensate” means a non-membrane-encapsulated compartment formed by phase separation of one or more proteins and / or other macromolecules such as nucleic acids (including all stages of phase separation).
[0042] The terms “polypeptide” and “protein,” as used herein, may be used interchangeably to refer to a polymer comprising amino acid residues, and are not limited to a minimum length. Such polymers may contain natural or non-natural amino acid residues, or combinations thereof, and include, but are not limited to, peptides, polypeptides, oligopeptides, dimers, trimers, and multimers of amino acid residues. Full-length polypeptides or proteins, and fragments thereof, are encompassed by this definition. The terms also include modified species thereof, e.g., post-translational modifications of one or more residues, including but not limited to, methylation, phosphorylation glycosylation, sialylation, or acetylation.
[0043] The term “treating” or “treatment,” as used herein, is an approach for obtaining beneficial or desired results including clinical results. For purposes of this application, beneficial or desired clinical results include, but are not limited to, one or more of the following: alleviating one or more symptoms resulting from the disease, diminishing the extent of the disease, stabilizing the disease (e.g., preventing or delaying the worsening of the disease), preventing or delaying the spread (e.g., metastasis) of the disease, preventing or delaying the recurrence of the disease, delay or slowing the progression of the disease, ameliorating the disease state, providing a remission (e.g., partial or total) of the disease, decreasing the dose of one or more other medications required to treat the disease, delaying the progression of the disease, increasing the quality of life, and / or prolonging survival.
[0044] The term “individual” refers to a mammal and includes, but is not limited to, human, bovine, horse, feline, canine, mouse, rodent, or primate.
[0045] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.
[0046] As used herein, the terms “comprising” (and any form or variant of comprising, such as “comprise” and “comprises”), “having” (and any form or variant of having, such as “have”10MF-366315044Docket No.: 185992002640and “has”), “including” (and any form or variant of including, such as “includes” and “include”), or “containing” (and any form or variant of containing, such as “contains” and “contain”), are inclusive or open-ended and do not exclude additional, un-recited additives, components, integers, elements, or method steps.
[0047] Throughout this disclosure, various aspects of the claimed subject matter are presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the claimed subject matter. Accordingly, the description of a range should be considered to have specifically disclosed all the possible sub-ranges as well as individual numerical values within that range. For instance, where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit, unless the context clearly dictate otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure. In some embodiments, two opposing and open-ended ranges are provided for a feature, and in such description it is envisioned that combinations of those two ranges are provided herein. For example, in some embodiments, it is described that a feature is greater than about 10 units, and it is described (such as in another sentence) that the feature is less than about 20 units, and thus, the range of about 10 units to about 20 units is described herein.
[0048] The term “about” as used herein refers to the usual error range for the respective value readily known in this technical field. Reference to “about” a value or parameter herein includes (and describes) variations that are directed to that value or parameter per se. For example, description referring to “about X” includes description of “X.” Exemplary degrees of error are within 20 percent (%), such as within 15%, within 10%, or within 5% of a given value or range of values.
[0049] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.II. Computer implemented methods for identifying condensate related target genes
[0050] Provided herein are computer implemented methods for identifying one or more condensate related target gene for a disease using a gene-network knowledge graph and condensate information. Condensate information may be integrated into the methods for 11MF-366315044Docket No.: 185992002640building the gene-network knowledge graph and for identifying one or more condensate related target genes from the gene network knowledge graph. Characteristics of the condensate phenotype in a healthy and / or disease cell model for the condensate related target gene can also be used for downstream validation. Condensate information may come from published literature, be predicted from the sequence of a protein, or be generated using proteomics methods. The condensate information may represent how likely a gene is to phase separate in a healthy and / or diseased cell. The condensate information may be used to add a condensate node and / or condensate edge to the gene-network knowledge graph. A condensate edge may be added to connect a gene node to a disease node if it is discovered that the protein product of the gene exhibits a different condensate phenotype in healthy and diseased cell states. The methods may be implemented iteratively, wherein condensate nodes and / or condensate edges can be added to the gene-network knowledge graph automatically building evidence for a gene node to represent a condensate related target gene.
[0051] The gene-network knowledge graphs can be any of the gene-network knowledge graphs described herein and can be used to integrate information such as genomic, transcriptomic, molecular, and pharmaceutical related to one or more diseases of interest. A condensate related target gene may be defined according to a change in phenotype between healthy and disease states, wherein a condensate forms in at least one of the phenotypes. It is appreciated that a condensate may form in both states and the difference in the condensate phenotype may be the size or location of the condensate.
[0052] The change in condensate phenotype may be due to a modification in the DNA, RNA, or protein of the gene and thus the methods described herein can be used for understanding the mechanisms connecting genetic variants to disease and identifying potential treatments the disease. The methods described herein comprise generating a data structure capable of storing information and the relationship between the information. The data structure allows for flexibility in the data that can be included and combined for relationship-based analysis. In certain aspects, as described herein, the data structure is a knowledge graph. A knowledge graph is constructed by providing a graph database (e.g. Neo4J) program, as described herein, with at least information about genes and a disease to form nodes of each category. The flexibility of the program allows for the information about a gene or a disease in a node to comprise a variety of information such as a name, a common identifier, and potential synonyms for each. Nodes, such as gene nodes, can be supplemented with information about a sequence (DNA, RNA, protein). As nodes (e.g., from node categories as described herein) are added to12MF-366315044Docket No.: 185992002640the gene network knowledge graph, the relationship between the node and other nodes in the graph is provided to the graph database program (e.g. Neo4J) as information stored in edges between the nodes. The edges and relationship the edges may represent are described herein.
[0053] In some embodiments, the gene-network knowledge graph comprises nodes from a plurality of node categories. Categorizing the nodes aids in analysis that requires the computer to understand how different nodes represent different aspects of disease biology that may be important when querying the gene network knowledge graph and identifying condensate related target genes. For example, a condensate node may be provided more weight in an analysis used to identify a condensate related target gene. The node categories may be gene nodes, disease nodes, pathway nodes, condensate nodes, and drug nodes, as described herein. The name of the node and the category of the node may be used for constructing the gene network knowledge graph.
[0054] In some embodiments, the gene nodes are encoded with condensate information, such as the condensate information described herein. Such information may be encoded in the node as a modification to nodes physical appearance (e.g., size) in a graphical user interface output for the gene network knowledge graph. The condensate information that may be encoded in a gene node differs from a condensate node in that the condensate information encoded in a gene node represents a condensate characteristic for the gene in a disease state. If a knowledge graph comprises multiple disease nodes, a gene node may be encoded with condensate information for one or more of the diseases. Exemplary condensate information may comprise relative changes in the amount of a macromolecule associated with the gene node in a condensate in a disease state, or the ratio of the amount of the macromolecule in a condensate versus the total amount of the macromolecule in a cell in a disease state. In some aspects, condensate information may comprise the presence of, or lack of a macromolecule associated with a gene node in a condensate.
[0055] The methods may comprise, generating a gene-network knowledge graph comprising at least a plurality of nodes, wherein the plurality of nodes comprise at least one gene node, at least one disease node, and edges representing a relationship between each node, calculating an integration value for one or more gene nodes in the gene-network knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph; obtaining condensate information for one or more gene nodes in the gene-network knowledge 13MF-366315044Docket No.: 185992002640graph; identifying one or more condensate related target genes from the gene-network knowledge graph using the condensate information and the integration value for the corresponding gene node. In some embodiments, the condensate information is stored in the gene-network knowledge graph data structure along with the gene node that that the condensate information relates to.
[0056] The methods may comprise, generating a gene-network knowledge graph comprising at least a plurality of nodes, wherein the plurality of nodes comprise at least one gene node, at least one disease node, and edges representing a relationship between each node wherein the gene-network graph comprises at least one condensate node or condensate edge; calculating an integration value for one or more gene nodes in the gene-network knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes of the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph; and identifying one or more condensate related target genes from the gene-network knowledge graph using integration value for the corresponding gene node. In some embodiments, the integration value relates to condensate information because the addition of condensate nodes and / or condensate edges in the gene-network knowledge graph increases the density of nodes and edges between the disease node and the gene node.
[0057] The methods may comprise, generating a gene-network knowledge graph comprising at least a plurality of nodes, wherein the plurality of nodes comprise at least one gene node, at least one disease node, and edges representing a relationship between each node wherein the gene-network graph comprises at least one condensate node or condensate edge; calculating an integration value for one or more gene nodes in the gene-network knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes of the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph; obtaining condensate information for one or more gene nodes in the gene-network knowledge graph; and identifying one or more condensate related target genes from the gene-network knowledge graph using the condensate information and the integration value for the corresponding gene node. In some embodiments, the methods adding a condensate node and / or condensate edge into the gene-network knowledge graph using the condensate information.
[0058] FIG. 1 illustrates an exemplary method according to the methods described herein for identifying one or more condensate related target gene for a disease. The methods are 14MF-366315044Docket No.: 185992002640implemented on a computer system such as the computer system of FIG. 4 and can be implemented within the exemplary system of FIG. 3.
[0059] At block 102, a gene-network knowledge graph is generated. The gene-network knowledge graph comprises at least a plurality of nodes, wherein the plurality of nodes comprises at least one gene node, at least one disease node, and edges representing a relationship between each node. In some embodiments, the gene-network knowledge graph comprises at least one condensate node or condensate edge. In some embodiments, the plurality of nodes comprises pathway nodes, condensate nodes, or drug nodes. The contents of the nodes and the associated edges may be collected from public sources such as databases and scientific publications or by performing experiments in the lab as described herein.
[0060] In some embodiments, the plurality of nodes comprises gene nodes. A gene node may represent the DNA, RNA, or protein associated with a gene. A gene node may be represent by and be encoded into the graph according the scientific name of the gene, a common name for a gene, a sequence (e.g. DNA RNA, protein). In some embodiments, the gene node may be a gene associated with a disease. Gene nodes may be connected by edges to one or more gene nodes and / or any other known type as described herein. As additional nodes are added to the gene network knowledge graph, gene nodes can be added for genes that are not originally on the list of genes associated with a disease. In some embodiments, a pathway node as described herein may be added to the gene network knowledge graph and all the genes known to be part of the pathway can also be added to the gene network knowledge graph as gene nodes. Gene nodes may also connect to other gene nodes with edges representing known interactions between genes at the DNA, RNA, or protein level. For example, an edge “interacts with” may be used to connect two gene nodes when there is a known protein interaction between the gene products.
[0061] In some embodiments, the plurality of nodes comprises disease nodes. A disease node may represent and be encoded into the graph according to a clinical indication, a common name for the disease, or a phenotype associated with a disease a sequence. In some embodiments, the disease node may be the disease that the method is being used to look for condensate related target genes for. In some embodiments, the plurality of disease nodes may comprise additional diseases or phenotypes associated with a disease. For example, a disease node may be type 2 diabetes or insulin resistance. Additional disease nodes may be incorporated into the gene network knowledge graph as other nodes are added to the graph. For example, if a gene node is added because the gene is known to interact with a gene associated with a first disease, and 15MF-366315044Docket No.: 185992002640the added gene node represents a gene known to be associated with a second disease, a disease node can be added for the second disease. Disease nodes may be connected by one or more edges to any other node type as described herein. Edges that may connect a gene node to a disease node may include, but are not limited to, “mutation causes”, “is GWAS target”, “is disregulated.” Edges that may connect a drug node to a disease node may include by are not limited to, “used to treat.”
[0062] The methods described herein can be used for identifying one or more condensate related target genes for a disease. The disease nodes comprise the disease that the methods is being used to identify a target for as well as any other diseases that may be included in the gene network knowledge graph on account of the nodes and edges included.
[0063] In some embodiments, the disease is a cancer, neurological disease, cardiac disease, or metabolic disease. In some embodiments, the disease is a cancer. In some embodiments, the cancer is a solid tumor cancer, such as but not limited to colorectal cancer or ovarian cancer. In some embodiments, the disease is a neurological disease. In some embodiments, the neurological disease, such as amyotrophic lateral sclerosis (ALS), multiple sclerosis, frontotemporal disorder, Parkinson’s disease, and Alzheimer’s disease. In some embodiments, the disease is a cardiac disease, such as familial or non-familial dilated cardiomyopathy (DCM), e.g., Desmoplakin (DSP), Desmoglein-2 (DSG2), and alpha-protein kinase 3 (ALPK3). In some embodiments, the disease is a metabolic disease. In some embodiments, the metabolic disease is type two diabetes (T2D).
[0064] In some embodiments, the plurality of nodes comprises at least one pathway node. Pathway nodes may be connected by one or more edges to any other node type as described herein. In some embodiments, a pathway node is connected to gene nodes by edges, wherein the edges represent the inclusion of a gene in the pathway. For example, an edge “in_pathway” may connect a gene involved in the cellular pathway represented in the pathway node. The pathway data may be sourced from a database such as the Reactome database.
[0065] In some embodiments, the plurality of nodes comprises at least one condensate node. A condensate node may represent the condensate information described herein. Condensate information incorporated as nodes in the gene network knowledge graph creates connections in the graphical structure which may increase the integration value for a gene. A condensate node may represent a condensate that has been characterized in the literature. The condensate may be any of the condensate types described in Table 1. Condensate nodes may be connected16MF-366315044Docket No.: 185992002640by one or more edges to any other node type as described herein. Edges that may connect a condensate node to a gene node may include, but are not limited to, “in condensate,” which may represent a relationship wherein a macromolecule associated with the gene node is a known member of the condensate represented in the condensate node. In some embodiments, a condensate node may be connected by an edge to a disease node because the phenotype associated with the condensate is known to be different in individuals with the disease compared to in non-disease individuals.Table 1: Example condensate typesCondensate Type DefinitionNuclear condensate involved in promoting cellular dormancy and A body control of local protein synthesis under stress conditions.Perinuclear condensate that serves as storage of organelles and RNA Balbiani body in oocytes.Cajal body Nuclear condensate involved in mRNA processing.The main center for organization of microtubules in animal cells. It is also involved in regulation of mitosis, cellular motility, polarity, and Centrosome adhesion.Cleavage body Nuclear condensate involved in protein degradation.Core fermentation Core fermentation factories for glycolytic proteins and glycolytic granules mRNAs. mRNAs are actively translated in these condensates.Nuclear condensate, also known as Gemini of Cajal bodies, involved Gem in snRNP maturation.Perinuclear ribonucleoprotein-granules found in C. elegans germ line Germ granule cells. They are also known as Nuage bodies and germ granules.Cytoplasmic granules, the assembly site of glycolityc enzymes. Their formation in yeast correlates with increased glucose consumption and Glycolytic bodies cell survival.Nuclear condensate that controls gene silencing and chromatin Heterochromatin compaction.Histone locus Nuclear condensate with function in transcriptional regulation and body processing of histone mRNA.Nuclear condensates involved in regulation of transcription, apoptotic signaling and antiviral response. They have also been referred to as Kremer body PML body and nuclear domain 10.Cytoplasmic ribonucleoprotein granules involved in spatial Localization body regulation of RNA in developing Xenopus oocytes.Perinuclear ribonucleoprotein-granules found in C. elegans germ line Nuage body cells. They are also known as germ granules and P-granules.Located on the nuclear envelope, they regulate nucleo-cytoplasmic Nuclear pore traffic by providing a selectivity barrier between the nuclear and complex cytoplasmic components.Nuclear condensates, also referred to as interchromatin granuleNuclear speckles clusters, regulate gene splicing and genome organization.17MF-366315044Docket No.: 185992002640Nuclear stress Nuclear condensates that transiently sequester proteins to protect body them from aggregation during stress.Nuclear ribonucleoprotein condensate that functions as the site of Nucleolus ribosome biogenesis, in stress sensing, and protein quality control.Cytoplasmic condensates enriched in p62 / SQSTM1 protein, involved p62 body in autophagy and SUMOylation.Paraspeckles Nuclear condensates with roles in transcriptional regulation.P-body Cytoplasmic condensates involved in RNA processing.Perinuclear ribonucleoprotein-granules found in C. elegans germ line P-granule cells. They are also known as Nuage bodies and germ granules.Nuclear condensates involved in regulation of transcription, apoptotic signaling and antiviral response. They have also been referred to as PML body Kremer body and nuclear domain 10.Proteasome Cytoplasmic condensates with roles in protecting proteasomes from storage granule degradation.Located on the plasma membrane, locally concentrates signaling Signaling cluster molecules to control signal transduction.Cytoplasmic condensates consisting of RNA and proteins that form upon stress. They serve as protective depots for RNA-binding Stress Granule proteins and translationally arrested mRNA.Synaptic density Located on the plasma membrane, regulate neuronal communication.Associated with the endoplasmic reticulum, function in regulation of TIS granule translation.Chromatin-associated nuclear condensates that recruit the Transcriptional transcriptional machinery to control active transcription. They have condensate also been referred to as transcriptional foci and super-enhancers.Cytoplasmic condensates associated with P-bodies, involved in U-body assembly and storage of U snRNPs.Bacterial ribonucleoprotein bacterial bodies (BR-bodies) are biomolecular condensates that form through phase separation of BR-body RNA degradosomes and are involved in bacterial mRNA decay.Prokaryotic enzymatic biomolecular condensate involved in Carboxysome carboxylation.Nuclear condensates that form in response to DNA double strand DNA damage foci breaks to concentrate repair proteins at the damaged DNA sites.Biomolecular condensate located in the mitochondrial matrix enriched in dsRNA, nascent mRNA and proteins involved in RNA Mitochondrial processing and maturation. They play roles in RNA regulation in RNA granule mitochondria.Biomolecular condensate found in neurons, involved in packaging of Neuronal granule mRNA and its transport to dendritic synapses.Perinucleolar Nuclear biomolecular condensate enriched in RNA-binding proteins compartment and RNA polymerase III, located at the perifery of the nucleolus.Nuclear biomolecular condensate, enriched in polycomb group Polycomb body proteins, involved in gene silencing.Eukaryotic enzymatic biomolecular condensate involved in Pyrenoid carboxylation.Sam68 nuclear Prinucleolar biomolecular condensate enriched in Sam68 protein andbody is scaffolded by architectural RNAs (arcRNAs).18MF-366315044Docket No.: 185992002640
[0066] In some embodiments, the plurality of nodes comprises at least one drug node. In some embodiments, a drug node may be a therapeutic compound or candidate. A drug node may be a chemical compound, such as a chemical compound that is used in treating a disease. In some embodiments, the drug node may be a biologic. It is appreciated that the drug category represents treatment modalities including but not limited to antibodies, siRNAs, gene editing methods, and cell therapy methods. In some embodiments, information from the ChEMBL databased may be encoded in one or more drug nodes. Drug nodes may be connected by one or more edges to any other node type as described herein. Drug nodes may be connected by one or more edges to any other node type as described herein. Edges that may connect a drug node to a gene node may include, but are not limited to, “has target.”
[0067] At block 102, the nodes of the gene network knowledge graph are connected by edges. Edges link nodes of the gene network knowledge graph. Edges may be directional to provide relevant information about the relationship between the nodes that they connect. In some embodiments, edges are labeled according to the relationship between the nodes as described herein. In some embodiments, the edges are not labeled. In some embodiments, a relationship between each node represented as an edge can be used to expand the gene network knowledge graph. In some embodiments, a node is added based on a potential edge between the node and a node already in the knowledge graph as described herein.
[0068] In some embodiments, the nodes of the gene-network knowledge graph are connected by one or more condensate edges. Incorporation of a condensate edge may be an additional way to incorporate condensate information into the gene-network knowledge graph. A condensate edge may be added to the gene-network knowledge graph to represent a relationship between two nodes, such as a gene node and a disease node related to a condensate. For example, it may be known or discovered using the methods for determining a cell fraction profile, as described herein that a macromolecule associated with the gene node phase separates in a disease state associated with the disease node and a condensate edge may be added connecting the gene node and disease node. Condensate information may be encoded in the system as a condensate edge rather than a condensate node wherein the condensate information does not relate to a canonical or known condensate. This feature of the method allows for discovery of new condensate phenotypes even when a condensate comprising a macromolecule associated with a gene node has not previously been described.
[0069] In some embodiments, the relationship between each node is stored in the system as an edge. In some embodiments, the relationship between each node comprises a relationship 19MF-366315044Docket No.: 185992002640selected from a group consisting of a gene-disease association, a protein-protein association, an RNA-protein association, a protein-pathway interaction, a tissue-specific expression relationship, a drug-gene interaction, a gene-condensate interaction, and an RNA-condensate interaction. It is appreciated that the relationship between each node may be any relationship between the two nodes obtained from scientific publications, databases, and experiments performed by one of skill in the art to understand the relationship between two nodes.
[0070] In some embodiments, the relationship between each node between two nodes obtained from results of a statistical genomics analysis, a statistical transcriptomics analysis, or a proteomics analysis. Statistical genomics and statistical transcriptomic methods provide information about genes, pathways, and diseases at high throughput. Statistical genomics analyses can be used to identify associations between genetic variants and disease, or gene regulation as described herein. Statistical transcriptomic analyses can be used to identify gene regulatory differences in different contexts such as between non-disease (e.g. healthy) and disease cell models. The gene regulatory differences may be at the mRNA expression level, protein level, gene regulatory network, or gene pathway level. Proteomics analysis can be used to identify proteins that change phase between disease and non-disease states, as described herein.
[0071] In some embodiments, the statistical genomics analysis comprises GWAS, eQTL studies, pQTL studies, PheWAS, Polygenic gene prioritization (POP), or rare variant aggregate testing.
[0072] In some embodiments, the statical genomics analysis comprises GWAS. Genome wide association studies (GWAS) can be used to identify genes associated with a disease by associating significant genetic variants identified in a GWAS for the disease to a gene. The variant may be in the coding region of the gene or may be in a non-coding region of the genome that is involved in regulation of the gene.
[0073] In some embodiments, statistical genomics analyses such as eQTL studies, pQTL studies, PheWAS, POP, and rare variant aggregate testing is used to explain the connection between a significant GWAS variant hit and the disease. For example, a genetic variant may correlate with differences in expression of a gene. Thus, the gene with variable expression may be associated with the disease and an edge can be used to connect the gene to the disease.20MF-366315044Docket No.: 185992002640
[0074] In some embodiments, the statistical genomics analysis comprises eQTL studies. An eQTL is a genetic variant associated with expression differences for a gene. In some embodiments, the genetic variation is associated with variation in mRNA expression levels.
[0075] In some embodiments, the statistical genomic analysis comprises a pQTL study. A pQTL is a genetic variant associated with expression differences for a gene. In some embodiments, the genetic variation is associated with variation in protein expression levels.
[0076] In some embodiments, the statistical genomics analysis comprises a PheWAS. A phenome-wide association study (PheWAS) is an analysis of the association between genetic variation and different phenotypes. In some embodiments, the different phenotypes are phenotypes that are known contributors for the disease, such as but not limited to a molecular trait, a biochemical trait, a cellular trait, or a related clinical diagnosis. Like GWAS, the genetic variants identified can be connected to a gene using any of the methods described herein or known in the art.
[0077] In some embodiments, the statistical genomic analysis comprises Polygenic gene prioritization (POP) analysis. Weeks et. al. Leveraging polygenic enrichments of gene features to predict gene underlying complex traits and disease. 55 Nature Genetics 1267-1276 (2023). A POP analysis can be used for prioritizing genes for the genetic variants identified in GWAS as associated with a disease. In some embodiments, POP may be used to identify a gene associated with a significant GWAS SNP. POP may be used to establish a relationship between a gene node and a disease node.
[0078] In some embodiments, the statistical genomics analysis comprises rare variant aggregate testing. In some embodiments, rare variant aggregate testing comprises rare variant burden testing. Rare variant aggregate testing can be used to identify genes associated with a disease when the individual genetic variants are not at a high enough allele frequency to be tested individually in a GWAS. Rare variants tested in these methods may be coding or noncoding variants. In some embodiments, rare variants are identified in a gene and an edge can be defined between the gene node and the disease node.
[0079] In some embodiments, the statistical transcriptomic analysis comprises a differential gene expression (DEG) analysis. DEG analyses can be used to identify genes that are upregulated or down regulated between two or more contexts. DEG analysis can be used to identify genes that are differentially regulated in disease or between tissues. Differences in gene expression can be measured at the mRNA or protein level. If a gene is up or down21MF-366315044Docket No.: 185992002640regulated in a cell model for disease or tissue from an individual with a disease, an edge may be used to connect the gene node to the disease node.
[0080] In some embodiments, the statistical transcriptomic analysis comprises pathway level interference. In some embodiments, pathway level interference may comprise gene set enrichment analysis (GSEA) or over-representation analysis (ORA). Using the genes that are differentially regulated in a context, such as in a disease state, these methods can be used to identify gene sets, such as pathways, associated with the disease. The genes that are differentially regulated may be identified using methods described herein and known in the art such as DEG analysis. The methods identify pathways where more genes than expected by chance are differentially regulated in the pathway. Once a pathway is identified, the pathway and the genes in the pathway can be connected to each other and to the disease with edges in the gene network knowledge graph.
[0081] In some embodiments, the statistical transcriptomic analysis comprises gene dependence analysis. A gene dependence analysis may be used to understand the effect of intentional dysregulation of a gene. If intentionally dysregulation of gene through up regulation or down regulation of the gene effects a phenotype associated with a disease, an edge may be used to connect the gene node with the disease node in the gene network knowledge graph. In some embodiments, CRISPR knockouts can be used to assay the effect of knocking out a genes on a phenotype associated with a disease in a disease cell model. In some embodiments, the gene dependence analysis can be used to define an edge between a gene node and a disease node.
[0082] In some embodiments, the statistical transcriptomic analysis comprises gene perturbation analysis. Gene perturbation analysis may comprise knocking down or knocking out a gene in a disease state and measuring the resulting transcriptomic effects. In some embodiments, the perturbation may contribute to a transcriptomic profile changing from more similar to a disease state to more similar to a non-disease state. In some embodiments, gene perturbation analysis can be done in silico. In some embodiments, in silico gene perturbation analysis comprises the methods described in Theodoris et al. Transfer Learning enables predictions in network biology. 618 Nature 616-624 (2023) In some embodiments, the gene perturbation analysis may use foundational models. In some embodiments, a gene perturbation analysis can ve used to define an edge between a gene node and a disease node.22MF-366315044Docket No.: 185992002640
[0083] In some embodiments, the statistical transcriptomic analysis comprises foundational model gene interaction mapping. In some embodiments, the foundational model gene interaction mapping may be used to predict an interaction between genes. In some embodiments, foundational models are fine-tuned with disease relevant data, e.g. RNAseq data collected from a disease cell model. In some embodiments, the fine-tuned foundational models can be used to infer gene networks linked to a disease state but not a non-disease state. In some embodiments, the output of the fine-tuned foundational models can be used to define an edge between two or more gene nodes and / or disease nodes.
[0084] In some embodiments, the statistical transcriptomic analysis comprises gene regulatory network mapping. Gene regulatory network mapping uses computational methods for annotating the relationship between the gene function. In some embodiments, ARACNE can be used for gene regulatory mapping. Margolin et al. ARACNE: An algorithm for the reconstruction of gene regulatory networks in a mammalian cellular context. 7 BMC Bioinformatics (2006). ARACNE uses correlations between expression of genes to learn relationships between the genes. In some embodiments, the output relationships from ARACNE can be used to define an edge between two or more gene nodes.
[0085] Other statistical genomics or transcriptomics analyses known in the art can be used to identify and explain the relationship between genes and a disease and thus to form edges between the gene and the disease in the gene network knowledge graph. In some embodiments, the results of statistical genomics or transcriptomic analyses are obtained from published scientific publications or databases.
[0086] At block 104, an integration value is calculated for one or more gene nodes in the genenetwork knowledge graph generated at block 102. The methods comprise calculating an integration value for one or more gene nodes in the gene-network knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph. In some embodiments, the gene-network knowledge graph is any of the knowledge graphs described herein.
[0087] The integration value captures the connectivity of the gene node to other nodes of interest in the gene network knowledge graph, such as the disease node. In principle, the integration value calculated at block 104 takes account of (1) how strongly and (2) how23MF-366315044Docket No.: 185992002640selectively each gene node is linked to the disease gene network. The gene nodes with the highest integration values have several unique connections to the disease node, and that have few connections to other genes not connected to the disease node. In some embodiments, genes with the highest integration values may have connections to a condensate node or may be connected to the disease node through a condensate edge.
[0088] In some embodiments, calculating the integration value comprises weighting the number of edges connecting the corresponding gene node to a node in one or more selected nodes the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph by characteristics of the edges connecting each node. In some embodiments, the characteristic of the edges comprises the directionality of the edges. In some embodiments, edges directed to the gene node are weighted higher than edges directed away from the gene node. In some embodiments, the characteristic of the edges comprises the number of edges to the disease of interest. In some embodiments, an edge is weighted higher if the edge is closer to the disease. In some embodiments, the characteristic of the edges comprises the characterization of one of then nodes connected to the edge. In some embodiments, an edge may be weighted higher if the edge is connected to a condensate node. In some embodiments, an edge may be weighted higher if the edge is a condensate edge. In some embodiments, calculating the integration value comprises performing a personalized page rank analysis.
[0089] In some embodiments, the one or more selected nodes have been selected based on a known relationship to the disease. Selecting nodes based on a known relationship to the disease may comprise selecting nodes that are part of a pathway known to be important for a disease. In some embodiments, the one or more selected nodes may have been selected based on condensate information. In some embodiments, the one or more selected nodes may comprise condensate nodes or nodes connected to the disease by a condensate edge.
[0090] In some embodiments, the one or more selected nodes comprise one or more nodes connected to the disease node associated with the disease by less than a predetermined number of edges. In some embodiments, nodes connected to the disease by less than the predetermined number of edges may be direct connections to the disease.
[0091] In some embodiments, the predetermined number of edges is 1, 5, 10, 15, or 20 edges. In some embodiments, the predetermined number of edges is less than 5, less than 10, less than 15, or less than 20 edges. In some embodiments, the predetermined number of edges is between24MF-366315044Docket No.: 1859920026401 and 20, 1 and 15, 1 and 10, or 1 and 5 edges. In some embodiments, the predetermined number of edges is between 5 and 20, 10 and 20, or 15 and 20 edges. It is appreciated that the predetermined number of edges may be based on the complexity of the gene network knowledge graph, such as the number of total edges in the graph. The predetermined number of edges may be larger for a gene network with more total edges than a graph with fewer.
[0092] At block 106, condensate information is obtained for one or more gene nodes in the gene-network knowledge graph generated at block 102. In some embodiments, the condensate information is generated using any of the methods for generating a cell fraction profile as described herein. In some embodiments, the condensate information may be generated by searching for evidence the DNA, RNA or protein associated with the gene node has been identified in a known condensate. In some embodiments, the methods may comprise integration of a condensate database, wherein condensate information can automatically be retrieved from the condensate database. The condensate database may be a database curated from published literature related to genes, condensates, and diseases. In some embodiments, the condensate information is generated by predicting if a protein will phase separate using the protein sequence associated with the gene node. In some embodiments, the prediction may be related to an estimate of disorder for an amino acid sequence. Models that can be used to predict phase separation from a protein sequence include but are not limited to PScore and PICNIC. In some embodiments, the gene-network knowledge graph is any of the gene-network knowledge graphs as described herein. It is appreciated that block 106 can be before or after block 104.
[0093] In some embodiments, the condensate information comprises a probability a polypeptide associated with the gene node is to phase separate in a cell model. In some embodiments, the condensate information is obtained from a proteomic experiment as described herein. In some embodiments, the condensate information is a binary indication of whether the protein product for the gene node phase separates in the disease context. In some embodiments, the condensate information represents the proportion of the protein product that phase separatees in the disease context. In some embodiments, the condensate information represents the difference between the amount of the protein product that phase separatees in the healthy and disease context. In some embodiments, the condensate information comprises a disease state phase transition characteristic for a polypeptide associated with a gene node. The disease state phase transition characteristic can be generated using the methods described25MF-366315044Docket No.: 185992002640herein. For example, a disease state phase transition characteristic for a polypeptide may be determined using a cell fraction profile for a disease, as described herein.
[0094] In some embodiments, the condensate information has been obtained using a method comprising: determining a disease state phase transition characteristic for each of one or more macromolecules, such as the polypeptides associated with a gene node in the gene network knowledge graph. In some embodiments, the disease state phase transition characteristics for a macromolecule comprises a comparison of the quantity of the macromolecules in fractionated cell lysates collected from disease model or a non-disease model, as described herein.
[0095] In some embodiments, the condensate information is stored in the gene-network knowledge graph. In some embodiments, the condensate information for a gene node is stored in the gene-network knowledge graph by scaling the size of the gene node.
[0096] In some embodiments, the condensate information relates to the probability the macromolecule associated with the gene node can be found in a condensate identified in a disease state or the relationship between the macromolecule and condensate change in a disease state. For example, the condensate information mat related to the probability the macromolecule associated with the gene node can no longer be found in a condensate in a disease state compared to in a non-disease state.
[0097] At block 108, one or more condensate related target genes are selected from the genenetwork knowledge graph using the condensate information obtained at block 106 and the integration values calculated at block 104.
[0098] In some embodiments, identifying one or more condensate related target genes comprises scaling the integration value by the condensate information. Scaling the integration value by the condensate information may comprise scaling the integration value by the probability of the protein gene product to phase separate in the disease context.
[0099] In some embodiments, identifying one or more condensate related target gene comprises choosing a cutoff for the integration value and the condensate information. The cutoffs may be selected based on the integration values and condensate information for positive control genes or negative control genes. A positive control gene may be a gene known to be associated with a disease or known to be a condensate related target gene for the disease. A negative control gene may be a gene known to not be associated with a disease or a gene associated with the disease but is not a condensate related target gene for the disease. As the methods described herein can be iterative, the positive and negative controls may be selected 26MF-366315044Docket No.: 185992002640based on the methods described herein. If a gene is identified and validated using the methods described herein as a condensate related target gene, other gene nodes with similar or stronger integration values can be selected as well. Positive and negative control genes may be identified based on scientific publications or databases.
[0100] In some embodiments, identifying one or more condensate related target genes comprises identifying one or more gene nodes with condensate information suggesting that a protein gene product is more likely to phase separate in the disease context than in the healthy context.
[0101] In some embodiments, identifying one or more condensate related target genes comprises identifying one or more gene nodes with a significant integration value. The significance of an integration value may be assessed using statistical methods known in the art. Such methods may include, but are not limited to generating a multiple comparison-adjusted p-value calculation by bootstrapped randomization of all the edges in the gene network knowledge graph. The randomization can be performed while keeping the number of edges per gene node constant.
[0102] In some embodiments, identifying one or more condensate related target genes comprises identifying one or more gene nodes with an integration value that is higher or more significant than an integration values for a gene node defined as a positive or negative control.
[0103] In some embodiments, identifying one or more condensate related target gene comprises identifying genes nodes with similar integration values and condensate information. In some embodiments, the methods may comprise clustering the genes according to the condensate information (e.g. probability the protein gene product phase separates in the disease context) and the integration value. In some embodiments, clustering comprises generating a heatmap. In some embodiments, genes can be selected based on analyzing clusters of genes that are likely to be related to the disease. These clusters may have an over representation of genes involved in gene sets or pathways of interest for the disease. In some embodiments, the clustering analysis may be used to de-select genes that would have been selected as a condensate related target gene using other prioritization methods. For example, clusters comprising genes with known potential safety issues can be deprioritized and genes with similar integration scores and condensate information can be de-selected. The positive and negative control genes as described herein can be annotated in the clusters and genes can be selected and de-selected as condensate related target genes based on their relationship to the27MF-366315044Docket No.: 185992002640positive and negative control genes. Genes that cluster with a positive control gene may be selected and genes that cluster with negative control genes may be deselected.
[0104] In some embodiments, identifying one or more condensate related target gene may comprise identifying one or more condensate related gene using the results of one or more validation experiments. In some embodiments, the validation experiments generate condensate phenotype information, genetic perturbation information, and compound modulating condensate information as described herein. In some embodiments, the one or more validation experiments are performed as part of the systems described herein.
[0105] In some embodiments, a validation experiment comprises imaging (e.g., high content). In some embodiments, the validation experiment comprises tagging the condensate related target gene with a fluorescence marker and performing high content fluorescent imaging of cells of from a cell model of disease. The imaging technique may be an in situ hybridization (ISH) technique, e.g., fluorescent ISH (FISH) technique, such as using a nucleic acid probe that specifically binds to a marker, e.g., a biological marker. In some embodiments, the biological marker, such as a polypeptide, a DNA, an RNA (coding or non-coding), or any modifications thereof, such as a post-translational modification of a polypeptide (e.g., phosphorylation, glycosylation, O-GlcNAcylation, UBL-protein conjugation (e.g., sumoylation), methylation, sialylation, acetylation, ADP-ribosylation, farnesylation, prenylation, deamidation, proteolysis, geranylgeranylation, hydroxylation, ubiquitylation, nitrosylation, lipidation), an epigenetic modification (e.g., histone acetylation or methylation, DNA methylation, etc.), or a modification to a nucleic acid (e.g., RNA capping). In some embodiments, the biological marker binds to the condensate related target gene. Thus, in some embodiments, the biological marker comprises a label. In some embodiments, the biological marker is labeled (such as via an affinity reagent, e.g., an antibody). In some embodiments, the label is selected from the group consisting of a radioactive label, a colorimetric label, a luminescent label, a chemically-reactive label (such as a component moiety used in click chemistry), and a fluorescent label. In some embodiments, the label is a small molecule, such as a compound having a molecular weight of 1000 Da or less. In some embodiments, the label is a small molecule comprising a fluorophore. In some embodiments, the label is associated with, such as covalently or non-covalently, a marker. In some embodiments, the label can be, but not limited to, Halo, dendra2, GFP, RFP, or mCherry.
[0106] The methods described herein may comprise comparing tagging the condensate related target gene in non-diseased and diseased cells with a biological marker as described herein for 28MF-366315044Docket No.: 185992002640imaging. The non-disease and disease cells may be cells from the cell model of the disease in the healthy and disease state. The non-disease cells may be cells from the non-disease cell model. The disease cells may be cells from the disease cell model. Using the images of the cells, the strength of the difference in the condensate phenotype in healthy cells and disease cells can be measured. A condensate related target gene is validated if there is a difference in a condensate phenotype between disease cells and healthy cells.
[0107] In some embodiments, a condensate phenotype comprises one or more observable or measurable characteristics or phenotypic identifiers associated with a condensate in a cell model. For example, observable or measurable characteristics or phenotypic identifiers associated with a condensate may be determined by imaging a composition comprising cells of a cell model. Observable or measurable characteristics of a condensate phenotype include, but are not limited to, presence (including absence and level / amount), location, distribution, kinetics (such as kinetics of formation or dissolution), morphological (e.g., size, shape, sphericity), material (e.g., fluidity or rigidity), and compositional properties of a condensate.
[0108] In some embodiments, the condensate phenotype is characterized by one or more phenotypic identifiers, such as an identifier selected from the group consisting of a condensate presence, absence, level, morphological feature, location, behavior, composition, and material property. In some embodiments, the condensate phenotype comprises the presence of a condensate of interest. In some embodiments, the condensate phenotype comprises the absence (including disappearance or dissolution) of a condensate of interest. In some embodiments, the condensate phenotype comprises the amount of a condensate of interest, including amount based on number of individual condensates and / or a size feature. In some embodiments, the condensate phenotype comprises the amount of a condensate of interest comprising and / or not comprising a component (e.g., marker such as biological marker, or one or more other biomolecules that become components of the condensate under certain conditions). In some embodiments, the condensate phenotype comprises the level (e.g., amount and / or strength) of association of a marker, such as a biomolecule (e.g., polypeptide, DNA, RNA), with a condensate of interest. In some embodiments, the condensate phenotype comprises the level (e.g., amount and / or strength) of association of a first biomolecule (e.g., polypeptide, DNA, RNA) with a second biomolecule in a cell model, wherein one or both of the biomolecules are associated with a condensate of interest, or the two biomolecules associate with different condensates. In some embodiments, the condensate phenotype comprises the abundance (or level of association) of a component of the condensate of interest within the condensate of29MF-366315044Docket No.: 185992002640interest. In some embodiments, the condensate phenotype comprises the location of a condensate of interest or component thereof, such as the subcellular location. For example, a condensate or a component thereof moves to a location where the condensate or component thereof would not normally locate during healthy condition (e.g., translocate to cytoplasm under disease condition). In some embodiments, the condensate phenotype comprises the distribution of a condensate of interest or component thereof (e.g., relative to other cellular organelles, other condensates, or other biomolecules). For example, condensates or components thereof distribute more densely at a subcellular location (e.g., densely distributed around the Golgi apparatus) compared to how they distribute during healthy condition. In some embodiments, the condensate phenotype comprises a morphological feature of a condensate of interest in a cell model, such as size, shape, volume, surface area, and / or sphericity. In some embodiments, the condensate phenotype comprises the number of condensates per cell. In some embodiments, the condensate phenotype comprises the composition of a condensate of interest. In some embodiments, the condensate phenotype comprises the behavior or material property of a condensate of interest, such as dynamic property, liquidity, solidity, or fiber formation. In some embodiments, the condensate phenotype comprises information regarding the kinetics of condensate formation. In some embodiments, the condensate phenotype comprises information regarding the kinetics of condensate dissolution. In some embodiments, the condensate phenotype comprises changes in a phenotypic identifier, such as a formation or dissolution characteristic, in response to an external stimulus.
[0109] In some embodiments, the condensate phenotype demonstrates that a condensate of interest is present in, or derived from, a cell model. In some embodiments, the condensate phenotype demonstrates that a condensate of interest is absent in, or not derived from, a cell model.
[0110] In some embodiments, the validation experiment comprises a genetic perturbation experiment. The genetic perturbation experiment may comprise modifying the expression of the condensate related target gene and comparing the global transcriptomic effect between healthy and disease cells. The healthy and disease cells may be cells from the cell model of the disease in the healthy and disease state. Modifying the expression of the condensate related target gene may comprise performing a CRISPR knockdown or an siRNA knockdown. Comparing the global transcriptomic effect of the genetic perturbation comprise performing RNA-sequencing, featuring the expression profiles from the RNA-sequencing data in the nondisease (e.g. healthy) state, disease state, the native and perturbed state, and performing30MF-366315044Docket No.: 185992002640differential gene expression analyzes. In some embodiments, methods such as DGE, PCA, and Boruta, Kursa and Rudnicki, Feature Selection with the Boruta Package, 36 J. of Statistical Software (2010), may be used to reduce transcriptomic profiles defined by RNA-sequencing genes that are relevant for the change between profiles. Comparisons between the non-disease (e.g. healthy) state, disease state, the native and perturbed state may be analyzed by measuring the cosine similarity of a vector of gene expression in different states. In some embodiments, each value in the vector is gene expression measured by RNA sequencing of a gene. In some embodiments, the vector may comprise measurements for the relevat genes identified with PCA, or Boruta.
[0111] In some embodiments, the validation experiment comprises a proteomic experiment. In some embodiments, a proteomics experiment can be used to measure the difference in abundance of a protein for a condensate related gene in an insoluble fraction, as described herein between a non-disease (e.g. health) and a disease state.
[0112] In some embodiments, identifying one or more condensate related target genes comprises identifying one or more gene nodes using compound modulating condensate information. Compound modulating condensate information may relate to whether a compound is known to modify the condensate phenotype of cells when the compound is applied. If a compound is known to modify a condensate phenotype of cells the compound is applied and is known to bind to or modify the condensate related target gene, it is more likely the condensate related target gene is a druggable target for the disease.III. Method for determining a cell fraction profile for a disease
[0113] The methods described herein can be used for determining a cell fraction profile for a disease. The cell fraction profile comprises quantification of one or more macromolecules (e.g., polypeptides) in one or more cell lysate fractions collected from cells from a disease model and / or a non-disease model. The cell fraction profile may comprise relative quantification of the one or more polypeptides in the one or more cell lysate fractions. Disease states can impact environments inside cells, or portions thereof, and change, amongst other things, cellular genomes, transcriptomes, proteomes, and metabolomes. The multivariable nature contributing to a disease state complicates predictions of how a macromolecule, such as a polypeptide, behaves or exists in a specific disease state including characteristics regarding phrase separation properties of said macromolecule or the propensity for said macromolecule to exhibit a modulation in the amount that is found in a condensate. For example, in certain disease31MF-366315044Docket No.: 185992002640states a polypeptide may be found in increased amounts or percentages in a condensate as compared to a non-disease state. The methods taught herein provide a direct measure of a macromolecule’s association (or lack thereof) with a condensate that can then be used in the methods taught herein for identifying genes as condensate related target genes.
[0114] In some aspects, for the methods provided herein involving condensate information, the condensate information associated with a node comprises a disease state phase transition characteristic for the macromolecules represented by the node, such as a gene or a polypeptide. In some embodiments, the method comprises determining the disease state phase transition characteristic for a macromolecule, such as a polypeptide. In some embodiments, the determining the disease state phase transition characteristic for a macromolecule, such as a polypeptide, comprises performing one or more of the following comparisons: the quantity of the macromolecule in an insoluble fraction of a disease model cell lysate as compared to the quantity of the macromolecule in an insoluble fraction of a non-disease model cell lysate; or the quantity of the macromolecule in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the macromolecule in a soluble fraction of a non-disease model cell lysate versus a non-disease model cell lysate. In some embodiments, the method further comprises quantifying the macromolecule, such as a polypeptide, to perform one or more comparisons. In some embodiments, the method further comprises quantifying the macromolecule, such as polypeptide, in the insoluble fraction of the disease model cell lysate and the macromolecule, such as polypeptide, in the insoluble fraction of the non-disease model cell lysate. In some embodiments, the method further comprises quantifying the macromolecule, such as polypeptide, in the soluble fraction of the disease model cell lysate and the macromolecule, such as polypeptide, in the soluble fraction of the non-disease model cell lysate. In some embodiments, the method further comprises quantifying the macromolecule, such as polypeptide, in the disease model cell lysate and the macromolecule, such as polypeptide, in the non-disease model cell lysate.
[0115] Also provided herein are methods of determining a cell fraction profile for a disease, comprising: fractionating cell lysate from non-disease cells and disease cells to generate one or more cellular fractions for the non-disease cells and the disease cells; obtaining relative quantification mass spectrometry data for the one or more cellular fractions from the non-disease cells and the disease cells; and identifying one or more polypeptides enriched in one or more of the cellular fractions from the non-disease cells and the disease cells to generate a cell fraction profile for the disease.32MF-366315044Docket No.: 185992002640
[0116] FIG. 2 illustrates an exemplary method according to the methods described herein for determining a cell fraction profile for a disease.
[0117] At block 202, a cell lysate from a non-disease cells and disease cells are fractionated to generate one or more cellular fractions for the non-disease cells and the disease cells. In some embodiments, the cell fractions comprise an insoluble fraction and / or a soluble fraction. In some embodiments, the cell fractions comprise a total fraction. In some embodiments, a total fraction is a combination of an insoluble fraction and a soluble fraction. In some embodiments, a total fraction is a non-fractionated cell lysate.
[0118] In some embodiments, the insoluble fraction comprises one or more phase separated polypeptides. In some embodiments, the insoluble fraction comprises one or more condensates. In some embodiments, the non-disease cells are cells from a non-disease cell model. In some embodiments, the disease cells are cells from a disease cells model.
[0119] In some embodiments, the cell model is a cell model for a disease (e.g., disease cell model) wherein the cell model comprises one or more disease-associated factors attributable to the disease. In some embodiments, the disease is a multifactorial disease having a plurality of disease-associated factors, wherein a cell model of the disease comprises one or more disease-associated factors attributable to the disease. In some embodiments, the cell model is a cell model for non-disease state, wherein the non-disease cell model does not comprise one or more disease-associated factors attributable to a disease. In some embodiments, the cell model for the disease comprises disease cells and the non-disease cell model comprises non-disease cells.
[0120] In some embodiments, the method further comprises fractionating a disease model cell lysate and / or a non-disease model cell lysate such as to obtain one or more fractions of a soluble and / or insoluble phase such as by centrifugation. In some embodiments, a detergent, e.g. mild (NP-40) and / or strong (SDS) detergent can be used to solubilize a fraction after centrifugation.
[0121] In some embodiments, the method further comprises obtaining a cell lysate from the disease model and / or the non-disease model. In some embodiments, the method further comprises lysing, separately, a sample from the disease model to obtain the disease model cell lysate and a sample from the non-disease model to obtain the non-disease model cell lysate. In some embodiments, the disease model comprises a cell model for a disease in a disease state and / or the non-disease model comprises a cell model for a disease in a non-disease state.33MF-366315044Docket No.: 185992002640Techniques for obtaining disease model and non-disease model cell lysates, fractioning said samples into soluble and insoluble phases are known in the field, e.g., Hu et al., Protein Cell, 8, 2017, and Hallegger et al., Cell, 184, 2021, which are hereby incorporated herein by reference in their entirety.
[0122] In some embodiments, the disease model comprises or is a cell model for a disease or a disease state (e.g., a disease cell model), or an aspect thereof. The disease models encompassed herein, may, in certain embodiments, comprises one or more disease-associated factors attributable to the disease. In some embodiments, the disease is a multifactorial disease having a plurality of disease-associated factors, wherein a cell model of the disease comprises one or more of the plurality of disease-associated factors attributable to the disease. In some embodiments, the disease model is a tissue (such as a xenograft) or an organism model (such as a mammal). In some embodiments, the disease model is characterized as having a disease, such as a cancer or a neurological disorder. In some embodiments, the non-disease model is a cell model serving as a control or representative of a healthy state. In some embodiments, wherein a disease model comprises one or more disease-associated factors attributable to a disease, the non-disease model does not comprise one or more disease-associated factors attributable to a disease. In some embodiments, the non-disease model is not characterized as having the disease of an associated disease model. In some embodiments, the disease model is disease tissue (or cells) and the non-disease model is non-disease tissue (or cells), such as healthy tissue. In some embodiments, the disease model and the non-disease model are of the same cell type. In some embodiments, the disease model and the non-disease model are of the same tissue type.
[0123] At block 204, the cellular fractions from block 202 are used to obtain relative quantification mass spectrometry data for the one or more cellular fractions. The relative quantification mass spectrometry data for the one or more cellular fractions comprises relative quantifications for one or more polypeptides in the cellular fractions from the non-disease and the disease cells. In some embodiments, the mass spectrometry data comprises relative quantifications for more than 10, more than 100, or more than 1000 polypeptides.
[0124] In some embodiments, the mass spectrometry data comprises relative quantifications for more than 10, more than 20, or more than 30, more than 40, more than 50, more than 60, more than 70, more than 80, or more than 90 polypeptides. In some embodiments, the mass spectrometry data comprises relative quantifications for more than 100, more than 200, or more than 300, more than 400, more than 500, more than 600, more than 700, more than 800, or 34MF-366315044Docket No.: 185992002640more than 900 polypeptides. In some embodiments, the mass spectrometry data comprises relative quantifications for more than 1000, more than 2000, or more than 3000, more than 4000, more than 5000, more than 6000, more than 7000, more than 8000, or more than 9000 polypeptides.
[0125] In some embodiments, the mass spectrometry data comprises relative quantifications for between 10 and 100, 100 and 1000, or 1000 and 10000 polypeptides. In some embodiments, the mass spectrometry data comprises relative quantifications for between 10 and 100, 10 and 90, 10 and 80, 10 and 70, 10 and 60, 10 and 50, 10 and 40, 10 and 30, or 10 and 20 polypeptides. In some embodiments, the mass spectrometry data comprises relative quantifications for between 10 and 100, 20 and 100, 30 and 100, 40 and 100, 50 and 100, 60 and 100, 70 and 100, 80 and 100, or 90 and 100 polypeptides.
[0126] In some embodiments, the mass spectrometry data comprises relative quantifications for between 100 and 1000, 100 and 900, 100 and 800, 100 and 700, 100 and 600, 100 and 500, 100 and 400, 100 and 300, or 100 and 200 polypeptides. In some embodiments, the mass spectrometry data comprises relative quantifications for between 100 and 1000, 200 and 1000, 300 and 1000, 400 and 1000, 500 and 1000, 600 and 1000, 700 and 1000, 800 and 1000, or 900 and 1000 polypeptides.
[0127] In some embodiments, the mass spectrometry data comprises relative quantifications for between 1000 and 10000, 1000 and 9000, 1000 and 8000, 1000 and 7000, 1000 and 6000, 1000 and 5000, 1000 and 4000, 1000 and 3000, or 1000 and 2000 polypeptides. In some embodiments, the mass spectrometry data comprises relative quantifications for between 1000 and 10000, 2000 and 10000, 3000 and 10000, 4000 and 10000, 5000 and 10000, 6000 and 10000, 7000 and 10000, 8000 and 10000, or 9000 and 10000 polypeptides.
[0128] In some embodiments the quantifying comprises performing a quantitative mass spectrometry technique. Quantitative mass spectrometry techniques are known in the field and include absolute quantification techniques (e.g., SRM and MRM), relative quantification techniques, label-free quantification techniques (e.g., spectral counting), and label -based quantification techniques (e.g., isobaric labeling), and combinations thereof. The quantitative mass spectrometry techniques described herein include those performed via a liquid chromatography mass spectrometry (LC-MS) technique. In some embodiments, the LC-MS technique is a bottom-up approaching comprises separating the peptide products obtained from a sample (such as via proteolytic digestion of an insoluble and / or soluble fraction from a35MF-366315044Docket No.: 185992002640disease model and / or non-disease model) via a liquid chromatography technique. Liquid chromatography techniques contemplated by the present application include methods for separating polypeptides and liquid chromatography techniques compatible with mass spectrometry techniques. In some embodiments, the liquid chromatography technique comprises a high-performance liquid chromatography technique. Thus, in some embodiments, the liquid chromatography technique comprises an ultra-high performance liquid chromatography technique. In some embodiments, the liquid chromatography technique comprises a high-flow liquid chromatography technique. In some embodiments, the liquid chromatography technique comprises a low-flow liquid chromatography technique, such as a micro-flow liquid chromatography technique or a nanoflow liquid chromatography technique. In some embodiments, the liquid chromatography technique comprises an online liquid chromatography technique coupled to a mass spectrometer. In some embodiments, the online liquid chromatography technique is a high-performance liquid chromatography technique. In some embodiments, the online liquid chromatography technique is an ultra-high performance liquid chromatography technique.
[0129] In some embodiments, capillary electrophoresis (CE) techniques, or electrospray or MALDI techniques may be used to introduce the sample to the mass spectrometer.
[0130] In some embodiment, the mass spectrometry technique comprises an ionization technique. Ionization techniques contemplated by the present application include techniques capable of charging, e.g., polypeptides or nucleic acids. Thus, in some embodiments, the ionization technique is electrospray ionization. In some embodiments, the ionization technique is nano-electrospray ionization. In some embodiments, the ionization technique is atmospheric pressure chemical ionization. In some embodiments, the ionization technique is atmospheric pressure photoionization. In some embodiments, the ionization technique is matrix-assisted laser desorption ionization (MALDI). In some embodiment, the mass spectrometry technique comprises electrospray ionization, nanoelectrospray ionization, or a matrix-assisted laser desorption ionization (MALDI) technique.
[0131] In some embodiments, the LC-MS technique comprises analyzing macromolecules, or products therefrom (e.g, peptide digests) via a mass spectrometry technique. Mass spectrometers contemplated by the present invention, to which an online liquid chromatography technique is coupled, include high-resolution mass spectrometers and low- resolution mass spectrometers. Thus, in some embodiments, the mass spectrometer is a time-of-flight (TOF) mass spectrometer. In some embodiments, the mass spectrometer is a 36MF-366315044Docket No.: 185992002640quadrupole time-of-flight (Q-TOF) mass spectrometer. In some embodiments, the mass spectrometer is a quadrupole ion trap time-of-flight (QIT-TOF) mass spectrometer. In some embodiments, the mass spectrometer is an ion trap. In some embodiments, the mass spectrometer is a single quadrupole. In some embodiments, the mass spectrometer is a triple quadrupole (QQQ) In some embodiments, the mass spectrometer is an orbitrap. In some embodiments, the mass spectrometer is a quadrupole orbitrap. In some embodiments, the mass spectrometer is a fourier transform ion cyclotron resonance (FT) mass spectrometer. In some embodiments, the mass spectrometer is a quadrupole fourier transform ion cyclotron resonance (Q-FT) mass spectrometer. In some embodiments, the mass spectrometry technique comprises positive ion mode. In some embodiments, the mass spectrometry technique comprises negative ion mode. In some embodiments, the mass spectrometry technique comprises a time-of-flight (TOF) mass spectrometry technique. In some embodiments, the mass spectrometry technique comprises a quadrupole time-of-flight (Q-TOF) mass spectrometry technique. In some embodiments, the mass spectrometry technique comprises an ion mobility mass spectrometry technique. In some embodiments a low-resolution mass spectrometry technique, such as an ion trap, or single or triple-quadrupole approach is appropriate.
[0132] In some embodiments, the LC-MS technique comprises processing the obtained MS signals. In some embodiments, the LC-MS technique comprises peak detection. In some embodiments, the LC-MS technique comprises determining ionization intensity. In some embodiments, the LC-MS technique comprises determining peak height. In some embodiments, the LC-MS technique comprises determining peak area. In some embodiments, the LC-MS technique comprises determining peak volume. In some embodiments, the LC-MS technique comprises identifying a polypeptide by amino acid sequence. In some embodiments, the LC-MS technique comprises identifying a protein by amino acid sequence. In some embodiments, the LC-MS technique comprises manually validation. In some embodiments, the LC-MS technique comprises identifying an associated identifier, such as a protein identifier.
[0133] In some embodiments, the quantitative mass spectrometry technique comprises use of isobaric labeling. Isobaric labeling techniques are known in the art and include sets of tags used to label sets of samples, wherein each tag is of the same or similar molecular weight and each tag in the set of tags can be distinguished based on a fragmentation pattern of products following fragmentation occurring in a mass spectrometer, such as collision induced dissection (CID). In some embodiments, the quantitative mass spectrometry technique comprises use of tandem mass tag (TMT) or isobaric tags for relative and absolute quantification (iTRAQ). In37MF-366315044Docket No.: 185992002640some embodiments, the quantitative mass spectrometry technique is multiplexed, e.g., 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 11x, 12x, 13x, 14x, 15x, 16x, 17x, or 18x.
[0134] In certain aspects provided herein, the mass spectrometry techniques disclosed herein comprise analysis of soluble aspects of a sample, such as via soluble proteome profiling. In some embodiments, such methods comprise solubilizing a sample comprising polypeptides (e.g., a cell lysate) using a detergent, such as NP-40 or SDS. In some embodiments, such methods comprise solubilizing a sample comprising polypeptides following centrifugation of a sample comprising polypeptides (e.g. a cell lysate). In some embodiments, said solubilizing produces separate polypeptide populations, such as by using different detergents, e.g., a mild detergent such as NP-40 and a strong detergent such as SDS. In some embodiments, the separate polypeptide populations can then be prepared for mass spectrometry analysis.
[0135] At block 206, the relative quantification mass spectrometry data from block 204 is used to identify one or more polypeptides enriched in one or more of the cellular fractions from the non-disease cells and the disease cells to generate a cell fraction profile for the disease.
[0136] In some embodiments, identifying one or more polypeptides in one or more of the cellular fractions comprises identifying one or more polypeptides enriched in the insoluble fraction of a disease model cell lysate compared to the insoluble fraction of a non-disease cell lysate. Comparing the quantity of one or more polypeptides in the insoluble fraction of the disease fraction to the quantity of one or more polypeptides in the insoluble fraction the non-disease fraction helps to identify one or more polypeptides that may aggregate to form condensates more readily in either the non-disease or disease state. The one or more polypeptides may aggregate to form condensates in cells from the cell model when the cell model is in a disease state, such as due to stress. The one or more polypeptides may contribute to disease dysregulation in humans.
[0137] In some embodiments, the methods comprise identifying one or more polypeptides enriched in the soluble fraction of a disease model cell lysate compared to the soluble fraction of a non-disease cell model. In some embodiments, the quantification of the one or more polypeptides in the soluble fraction of the disease model is divided by the quantification of the one or more polypeptides in cells from the disease model before fractionation. In some embodiments, the quantification of the one or more polypeptides in the soluble fraction of the disease model is divided by the quantification of the one or more polypeptides in cells from the disease model before fractionation (e.g., total fraction) In some embodiments, the38MF-366315044Docket No.: 185992002640quantification of the one or more polypeptides in the soluble fraction of the non-disease model is divided by the quantification of the one or more polypeptides in cells from the non-disease model before fractionation (e.g., total fraction). In some embodiments, quantifications for the total fraction (e.g., in disease or non-disease cell lysate), can be generated by adding the quantifications of the one or more polypeptides in the soluble and insoluble fractions. The comparison of the ratio of the one or more polypeptides in the soluble to total fraction in the disease model to the ratio of the one or more polypeptides in the soluble to total fraction in the non-disease model may reveal single proteins that shift from soluble to insoluble or vice versal. The comparison can be used to understand the mechanisms of how disease states impact polypeptide solubility.
[0138] The methods disclosed herein can be used to identify genes likely to phase separate in a disease state or those of which have a different profile in a disease state. In some embodiments, a cell fraction profile as described herein comprises determining a disease state phase transition characteristic for each of one or more polypeptides. In some embodiments, the cell fraction profile comprises condensate information for one or more polypeptide related to one or more gene node of a gene network knowledge graph as described herein. In some embodiments, the cell fraction profile comprises condensate information for use in identifying condensate related target genes using the methods described herein.IV. Systems
[0139] FIG. 3 illustrates an example system for identifying one or more condensate related target genes for a disease, in accordance with various embodiments. System 300 may include a computing system 302, user devices 330-1 to 330-N (also referred to collectively as “user devices 330” and individually as “user device 330”), databases 340 (e.g., genomics database 342, condensate data database 344,), or other components. In some embodiments, components of system 300 may communicate with one another using network 350, such as the Internet.
[0140] User devices 330 may communicate with one or more components of system 300 via network 350 and / or via a direct connection. User devices 330 may be a computing device configured to interface with various components of system 300 to control one or more tasks, cause one or more actions to be performed, or effectuate other operations. For example, user device 330 may be configured to receive and display a gene-network knowledge graph as described herein. Example computing devices that user devices 330 may correspond to include, but are not limited to, which is not to imply that other listings are limiting, desktop computers,39MF-366315044Docket No.: 185992002640servers, mobile computers, smart devices, wearable devices, cloud computing platforms, or other client devices. In some embodiments, each user device 330 may include one or more processors, memory, communications components, display components, audio capture / output devices, image capture components, or other components, or combinations thereof. Each user device 330 may include any type of wearable device, mobile terminal, fixed terminal, or other device.
[0141] It should be noted that while one or more operations are described herein as being performed by particular components of computing system 302, those operations may, in some embodiments, be performed by other components of computing system 302 or other components of system 300. As an example, while one or more operations are described herein as being performed by components of computing system 302, those operations may, in some embodiments, be performed by aspects of user devices 330.
[0142] Computing system 302 may be configured to communicate with one another, one or more other devices, systems, and / or servers, using network 350 (e.g., the Internet, an Intranet). System 300 may also include one or more databases 340 used to store data used by one or more components of system 300. In some embodiments, the computing system comprises a graph database program. In some embodiments, the graph database program is Neo4J. This disclosure anticipates the use of one or more of each type of system and component thereof without necessarily deviating from the teachings of this disclosure.
[0143] Although not illustrated, other intermediary devices (e.g., data stores of a server connected to computing system 302) can also be used. The components of system 300 of FIG.3 can be used in a variety of contexts where scanning and evaluating digital pathology images, such as whole slide images, are essential components of the work. As an example, system 300 can be associated with a clinical environment where a user is evaluating cells for drug discovery and evaluation. The user can review the gene network knowledge graph using user device 330 and can provide additional information to computing system 302 that can be used to guide or direct the analysis of the gene network knowledge graph. Persons of ordinary skill in the art will recognize that, in some examples, no user review of the gene network knowledge graph may be needed.
[0144] In some embodiments, the user can interact with the gene network knowledge graph using user device 330 by providing an input query into the system. In some embodiments, computing system, use the input query to automatically generate a response. The input query40MF-366315044Docket No.: 185992002640may be a disease represented in the gene network knowledge graph as a disease node. The input query may be a natural language input or a selection according to the disease nodes that are stored in the gene network knowledge graph. Using the input query, the computing system may automatically rank genes represented as gene nodes by the integration score and condensate information as described herein. The computing system may traverse between the disease node associated with the input query and other nodes in the gene network knowledge graph using the edges connecting the nodes and may calculate the integration score as described herein. The user may use the response for identifying condensate related target genes.
[0145] The input query may be a gene represented in the gene network knowledge graph as a gene node. The input query may be a natural language input or a selection according to the gene nodes that are stored in the gene network knowledge graph. Using the input query, the computing system may automatically output condensate information for the gene that has been stored in the graphical structure and an integration score for one or more diseases represented as disease nodes. The computing system may traverse between the gene node associated with the input query and disease nodes in the gene network knowledge graph using the edges connecting the nodes and may calculate the integration score as described herein. The user may use the response for identifying condensate related target genes.
[0146] In some embodiments, the system can cause a user interface to display all or part of the gene network knowledge graph or the responses to a user query as described herein. In some embodiments, the interfaces are graphical user interfaces that allow users to interact with and control the gene network knowledge graph and responses to input queries as well as control system devices. In some embodiments, the system may also generate user interactable gene network knowledge graphs generated according to the methods described herein. In some embodiments, a user can interact with the graphical user interface to select a condensate related target gene according to an interaction score and condensate information for a gene.
[0147] In some embodiments, the graphical user interface may be configured to display the gene network knowledge graph, for example using the representation in FIG. 5. In some embodiments, the condensate information for a gene may be represented in the graphical representation of the gene-network knowledge graph as a relative node size. A gene that is more likely to form a condensate may appear larger than a gene less likely to form a condensate.
[0148] Other intermediary devices may include devices for performing validation assays as described herein automatically. In some embodiments, the components comprise automated41MF-366315044Docket No.: 185992002640cell culture or automated liquid handler devices. Persons of ordinary skill in the art will recognize that, in some examples, no user intervention may be needed to perform the functional assay as described herein.
[0149] FIG.4 illustrates an example computer system 400. In some embodiments, one or more computer systems 400 perform one or more steps of one or more methods described or illustrated herein. In some embodiments, one or more computer systems 400 provide functionality described or illustrated herein. In some embodiments, software running on one or more computer systems 400 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 400. Herein, reference to a computer system may encompass a computing device, and vice versa, where appropriate. Moreover, reference to a computer system may encompass one or more computer systems, where appropriate.
[0150] This disclosure contemplates any suitable number of computer systems 400. This disclosure contemplates computer system 400 taking any suitable physical form. As example and not by way of limitation, computer system 400 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (such as, for example, a computer-on-module (COM) or system -on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these. Where appropriate, computer system 400 may include one or more computer systems 400; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 400 may perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. As an example, and not by way of limitation, one or more computer systems 300 may perform in real time or in batch mode one or more steps of one or more methods described or illustrated herein. One or more computer systems 300 may perform at various times or at different locations one or more steps of one or more methods described or illustrated herein, where appropriate.
[0151] In some embodiments, computer system 400 includes a processor 402, memory 404, storage 406, an input / output (I / O) interface 408, a communication interface 410, and a bus 412. Although this disclosure describes and illustrates a particular computer system having a 42MF-366315044Docket No.: 185992002640particular number of components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.
[0152] In some embodiments, processor 402 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions, processor 402 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 404, or storage 406; decode and execute them; and then write one or more results to an internal register, an internal cache, memory 404, or storage 406. In some embodiments, processor 402 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 402 including any suitable number of any suitable internal caches, where appropriate. As an example, and not by way of limitation, processor 402 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches may be copies of instructions in memory 404 or storage 406, and the instruction caches may speed up retrieval of those instructions by processor 402. Data in the data caches may be copies of data in memory 404 or storage 406 for instructions executing at processor 402 to operate on; the results of previous instructions executed at processor 402 for access by subsequent instructions executing at processor 402 or for writing to memory 404 or storage 406; or other suitable data. The data caches may speed up read or write operations by processor 402. The TLBs may speed up virtual-address translation for processor 402. In some embodiments, processor 402 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 402 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 402 may include one or more arithmetic logic units (ALUs); be a multi-core processor; or include one or more processors 402. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.
[0153] In some embodiments, memory 404 includes main memory for storing instructions for processor 402 to execute or data for processor 402 to operate on. As an example, and not by way of limitation, computer system 400 may load instructions from storage 406 or another source (such as, for example, another computer system 400) to memory 404. Processor 402 may then load the instructions from memory 404 to an internal register or internal cache. To execute the instructions, processor 402 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 40243MF-366315044Docket No.: 185992002640may write one or more results (which may be intermediate or final) to the internal register or internal cache. Processor 402 may then write one or more of those results to memory 404. In some embodiments, processor 402 executes only instructions in one or more internal registers or internal caches or in memory 404 (as opposed to storage 406 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory 404 (as opposed to storage 406 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple processor 402 to memory 404. Bus 412 may include one or more memory buses, as described below. In some embodiments, one or more memory management units (MMUs) reside between processor 402 and memory 404 and facilitate access to memory 404 requested by processor 402. In some embodiments, memory 404 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Moreover, where appropriate, this RAM may be single-ported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 404 may include one or more memories 404, where appropriate. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.
[0154] In some embodiments, storage 406 includes mass storage for data or instructions. As an example, and not by way of limitation, storage 406 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magneto-optical disc, magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. Storage 406 may include removable or non-removable (or fixed) media, where appropriate. Storage 406 may be internal or external to computer system 400, where appropriate. In some embodiments, storage 406 is non-volatile, solid-state memory. In some embodiments, storage 406 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination of two or more of these. This disclosure contemplates mass storage 406 taking any suitable physical form. Storage 406 may include one or more storage control units facilitating communication between processor 402 and storage 406, where appropriate. Where appropriate, storage 406 may include one or more storages 406. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.
[0155] In some embodiments, I / O interface 408 includes hardware, software, or both, providing one or more interfaces for communication between computer system 400 and one or44MF-366315044Docket No.: 185992002640more I / O devices. Computer system 400 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and computer system 400. As an example, and not by way of limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device, or a combination of two or more of these. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interfaces 408 for them. Where appropriate, I / O interface 408 may include one or more device or software drivers enabling processor 402 to drive one or more of these I / O devices. I / O interface 408 may include one or more I / O interfaces 408, where appropriate. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.
[0156] In some embodiments, communication interface 410 includes hardware, software, or both providing one or more interfaces for communication (such as, for example, packet-based communication) between computer system 400 and one or more other computer systems 400 or one or more networks. As an example, and not by way of limitation, communication interface 410 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 410 for it. As an example, and not by way of limitation, computer system 400 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 400 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), or other suitable wireless network or a combination of two or more of these. Computer system 400 may include any suitable communication interface 410 for any of these networks, where appropriate. Communication interface 410 may include one or more communication interfaces 410, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.45MF-366315044Docket No.: 185992002640
[0157] In some embodiments, bus 412 includes hardware, software, or both coupling components of computer system 400 to each other. As an example and not by way of limitation, bus 412 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Bus 412 may include one or more buses 412, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.
[0158] Herein, a computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, non-volatile, or a combination of volatile and non-volatile, where appropriate.
[0159] The scope of this disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the example embodiments described or illustrated herein that a person having ordinary skill in the art would comprehend. The scope of this disclosure is not limited to the example embodiments described or illustrated herein. Moreover, although this disclosure describes and illustrates respective embodiments herein as including particular components, elements, feature, functions, operations, or steps, any of these embodiments may include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that a person having ordinary skill in the art would comprehend. Furthermore, reference in the appended claims to an apparatus or system or a component of an apparatus or system being adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function encompasses that apparatus, system, component, whether or not it or that particular function is activated, turned on, or unlocked, as long as that apparatus, system, or component46MF-366315044Docket No.: 185992002640is so adapted, arranged, capable, configured, enabled, operable, or operative. Additionally, although this disclosure describes or illustrates particular embodiments as providing particular advantages, particular embodiments may provide none, some, or all of these advantages.EXEMPLARY EMBODIMENTS
[0160] Embodiments disclosed herein may include:
[0161] Embodiment 1. A computer implemented method for identifying one or more condensate related target genes for a disease, comprising:generating a gene-network knowledge graph comprising at least a plurality of nodes, wherein the plurality of nodes comprise at least one gene node, at least one disease node, and edges representing a relationship between each node;calculating an integration value for one or more gene nodes in the gene-network knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes of the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the genenetwork knowledge graph;obtaining condensate information for one or more gene nodes in the gene-network knowledge graph;identifying one or more condensate related target genes from the gene-network knowledge graph using the condensate information and the integration value for the corresponding gene node.
[0162] Embodiment 2. The computer implemented method of embodiment 1, wherein the plurality of nodes further comprises at least one gene pathway node.
[0163] Embodiment s. The computer implemented method of embodiment 1 or embodiment 2, wherein the plurality of nodes further comprises at least one condensate node.
[0164] Embodiment 4. The computer implemented method of any of embodiments 1-3, wherein the plurality of nodes further comprises at least one drug node.
[0165] Embodiment 5. The computer implemented method of any of embodiments 1-4, wherein the relationship between each node comprises a relationship selected from a group consisting of gene-disease association, a protein-protein association, an RNA-protein47MF-366315044Docket No.: 185992002640association, a protein-pathway interaction, a tissue specific expression relationship, a druggene interaction, a gene-condensate interaction, and an RNA-condensate interaction.
[0166] Embodiment 6. The computer implemented method of any of embodiments 1-4, wherein the relationship between each node comprises a relationship between two nodes obtained from results of a statistical genomics analysis a statistical transcriptomics analysis, or a proteomic analysis.
[0167] Embodiment 7. The computer implemented method of embodiment 6, wherein the statistical genomics analysis comprises GWAS, eQTL, pQTL, PheWAS, Polygenic gene prioritization (POP), or rare variant aggregate testing.
[0168] Embodiment 8. The computer implemented method of embodiment 6 or embodiment 7, wherein the statistical transcriptomics analysis comprises differential gene expression, pathway level interference, gene dependence analysis, gene perturbation analysis, foundational model gene interaction mapping, or gene regulatory network mapping.
[0169] Embodiment 9. The computer implemented method of any of embodiments 1-8, wherein calculating an integration value comprises weighting the number of edges connecting the corresponding gene node to a node in one or more selected nodes the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph by characteristics of the edges connecting each node.
[0170] Embodiment 10. The computer implemented method of any of embodiments 1-8, wherein calculating an integration value comprises performing a personalized page rank analysis.
[0171] Embodiment 11. The computer implemented method of any of embodiments 1- 10, wherein the one or more selected nodes have been selected based on a known relationship to the disease.
[0172] Embodiment 12. The computer implemented method of any of embodiments 1- 11, wherein the one or more selected nodes comprise one or more nodes connected to the disease node associated with the disease by less than a predetermined number of edges.
[0173] Embodiment 13. The computer implemented method of embodiment 12, wherein the predetermined number of edges is 1, 5, 10, 15, or 20.48MF-366315044Docket No.: 185992002640
[0174] Embodiment 14. The computer implemented method of any of embodiments 1- 13, wherein the integration value represents connectivity and selectivity of the gene node to the disease node representing the disease.
[0175] Embodiment 15. The computer implemented method of any of embodiments 1- 14, wherein the condensate information comprises predicted phase separation for each of one or more polypeptides corresponding to one or more gene nodes.
[0176] Embodiment 16. The computer implemented method of embodiment 15, wherein the predicted phase separation is generated from a protein sequence corresponding to the gene node.
[0177] Embodiment 17. The computer implemented method of any of embodiments 1-16, wherein the condensate information comprises a disease state phase transition characteristic for each of one or more polypeptides corresponding to one or more gene nodes.
[0178] Embodiment 18. The computer implemented method of embodiment 17, wherein the condensate information has been obtained using a method comprising: determining the disease state phase transition characteristic for each of the one or more polypeptides.
[0179] Embodiment 19. The computer implemented method of embodiment 18, wherein the determining the disease state phase transition characteristic for a first polypeptide of the one or more polypeptides comprises performing one or more of the following comparisons:the quantity of the first polypeptide in an insoluble fraction of a disease model cell lysate as compared to the quantity of the first polypeptide in an insoluble fraction of a nondisease model cell lysate; orthe quantity of the first polypeptide in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the first polypeptide in a soluble fraction of a non-disease model cell lysate versus a non-disease model cell lysate.
[0180] Embodiment 20. The computer implemented method of embodiment 19, further comprising quantifying the first polypeptide to perform the one or more comparisons.
[0181] Embodiment 21. The computer implemented method of embodiment 19, further comprising quantifying the first polypeptide in the insoluble fraction of the disease model cell lysate and the first polypeptide in the insoluble fraction of the non-disease model cell lysate.49MF-366315044Docket No.: 185992002640
[0182] Embodiment 22. The computer implemented method of embodiment 19, further comprising quantifying the first polypeptide in the soluble fraction of the disease model cell lysate and the first polypeptide in the soluble fraction of the non-disease model cell lysate.
[0183] Embodiment 23. The computer implemented method of embodiment 19, further comprising quantifying the first polypeptide in the disease model cell lysate and the first polypeptide in the non-disease model cell lysate.
[0184] Embodiment 24. The computer implemented method of any of embodiments 20-23, wherein the quantifying comprises performing a quantitative mass spectrometry technique.
[0185] Embodiment 25. The computer implemented method of embodiment 24, wherein the quantitative mass spectrometry technique comprises use of isobaric labeling.
[0186] Embodiment 26. The computer implemented method of embodiment 24, wherein the quantitative mass spectrometry technique comprises use of tandem mass tag (TMT) or isobaric tags for relative and absolute quantification (iTRAQ).
[0187] Embodiment 27. The computer implemented method of any of embodiment 24- 26, wherein the quantitative mass spectrometry technique is multiplexed.
[0188] Embodiment 28. The computer implemented method of any of embodiment 19- 27, further comprising fractionating a disease model cell lysate and / or a non-disease model cell lysate.
[0189] Embodiment 29. The computer implemented method of any of embodiment 19- 28, further comprising obtaining a cell lysate from the disease model and / or the non-disease model.
[0190] Embodiment 30. The computer implemented method of embodiments 19-29, further comprising lysing, separately, a sample from the disease model to obtain the disease model cell lysate and a sample from the non-disease model to obtain the non-disease model cell lysate.
[0191] Embodiment 31. The computer implemented method of any of embodiments 19- 30, wherein the disease model comprises a cell model for a disease in a disease state and / or the non-disease model comprises a cell model for a disease in a non-disease state.
[0192] Embodiment 32. The computer implemented method of any of embodiments 1- 31, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes with a significant integration value.50MF-366315044Docket No.: 185992002640
[0193] Embodiment 33. The computer implemented method of any of embodiments 1- 32, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes that cluster together based on integration values and condensate information.
[0194] Embodiment 34. The computer implemented method of any of embodiments 1- 33, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes using condensate phenotype information.
[0195] Embodiment 35. The computer implemented method of any of embodiments 1- 34, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes using genetic perturbation information.
[0196] Embodiment 36. The computer implemented method of any of embodiments 1- 35, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes using compound modulating condensates information.
[0197] Embodiment 37. A method of identifying one or more condensate related target genes for a disease, the method comprising:generating a gene-network knowledge graph comprising at least a plurality of nodes, wherein the plurality of nodes comprise at least one gene node, at least one disease node, and edges representing a relationship between each node;calculating an integration value for one or more gene nodes in the gene-network knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the genenetwork knowledge graph;obtaining a disease state phase transition characteristic for one or more gene nodes in the gene-network knowledge graph by a method comprising:obtaining a disease model cell lysate and a non-disease model cell lysate; fractionating, separately, the disease model cell lysate and the non-disease model cell lysate to obtain respective insoluble and soluble fractions thereof;performing a quantitative mass spectrometry technique on one or more of:51MF-366315044Docket No.: 185992002640the soluble fraction of the disease model cell lysate, the disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate; orthe insoluble fraction of the disease model cell lysate, the soluble fraction of the disease model cell lysate, the disease model cell lysate, the insoluble fraction of the non-disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate;determining a disease state phase transition characteristic for each of one or more polypeptides corresponding to the one or more gene nodes based on one or more of the following comparisons:the quantity of a polypeptide in an insoluble fraction of a disease model cell lysate as compared to the quantity of the polypeptide in an insoluble fraction of a non- disease model cell lysate; orthe quantity of a polypeptide in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the polypeptide in a soluble fraction of a non-disease model cell lysate versus a non-disease model cell lysate; andidentifying one or more condensate related target genes from the gene-network knowledge graph using the disease state phase transition characteristic and the integration value for the corresponding gene node.
[0198] Embodiment 38. A method of identifying one or more condensate related target genes for a disease, the method comprising:obtaining a disease model cell lysate and a non-disease model cell lysate; fractionating, separately, the disease model cell lysate and the non-disease model cell lysate to obtain respective insoluble and soluble fractions thereof;performing a quantitative mass spectrometry technique on one or more of:the insoluble fraction of the disease model cell lysate and the insoluble fraction of the non-disease model cell lysate;the soluble fraction of the disease model cell lysate, the disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non- disease model cell lysate; or52MF-366315044Docket No.: 185992002640the insoluble fraction of the disease model cell lysate, the soluble fraction of the disease model cell lysate, the disease model cell lysate, the insoluble fraction of the non-disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate;determining a disease environment phase transition characteristic for each of one or more polypeptides based on one or more of the following comparisons:the quantity of a polypeptide in an insoluble fraction of a disease model cell lysate as compared to the quantity of the polypeptide in an insoluble fraction of a non-disease model cell lysate; orthe quantity of a polypeptide in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the polypeptide in a soluble fraction of a non-disease model cell lysate versus a non-disease model cell lysate; andfiltering a gene-network knowledge graph using the disease environment phase transition characteristics to identify the one or more condensate related target genes from the disease.
[0199] Embodiment 39. A system for identifying one or more condensate related target genes for a disease, the system comprising:one or more processors,a user input device, anda memory communicatively coupled to the one or more processors configured to store instructions that, when executed by the one or more processors, cause the systems to:perform the computer implemented method of any one of embodiments 1-36.
[0200] Embodiment 40: A computer-readable non-transitory storage medium storing one or more programs, the one or more programs comprising instructions that when executed by one or more processors of a system cause the system to:perform the computer implemented method of any one of embodiments 1-36.53MF-366315044Docket No.: 185992002640EXAMPLESExample 1-Identification of condensate targets for insulin resistance in Type 2 Diabetes (T2D)
[0201] This example demonstrates a method of identifying and validating condensate targets for insulin resistance in type-2 diabetes (T2D).
[0202] Omics-based methods were used to identify genes associated with type-2 diabetes. The methods relied on human genetics, transcriptomics, and proteomics. First, several statistical genomics approaches were used to identify T2D disease genes. A Genome Wide Associate Study (GWAS) was performed using human genetics and T2D disease status from UK Biobank. Using the GWAS summary statistics, statistical genomics approaches were used to identify disease genes. Significant single nucleotide polymorphisms (SNPs) were mapped to gene using and genes using polygenic gene prioritization (PoPS). The genes with significant SNPs were identified as genes significantly associated with T2D. In addition, genes with significant SNPs in a Phenome wide association study (PheWAS), genes with significant expression quantitative trait loci (eQTL) or protein quantitative trait loci (pQTL), and genes identified to harbor rare variants associated with T2D using aggregate variant burden testing and rare variant testing were all identified as T2D disease genes.
[0203] Second, statistical transcriptomics approaches were used to further identify T2D disease genes. The results of differential gene expression (DEG) analyses were obtained and genes significantly up regulated or down regulated in individuals with T2D were identified as T2D disease genes. Gene regulatory networks and pathways associated with T2D were identified using Pathway-level inferences (GSEA, ORA) and ARACNE methods on the DEG results. Genes in the significant pathways and networks were identified as T2D disease genes.
[0204] The results of the statistical genomic and statistical transcriptomic approaches were combined and overall, 826 disease genes were identified as likely associated with T2D. A T2D knowledge graph was created using the 826 T2D disease genes, and T2D as nodes. The edges represented the connection between the genes and each other according to data generated using the statistical genomic and statistical transcriptomic approaches as well as the connection between the genes and T2D. For example, connections between genes inferred with ARACANE were encoded as edges.54MF-366315044Docket No.: 185992002640
[0205] Publicly available data was used to expand the number and characteristics of nodes in the T2D knowledge graph and insert the edges between the nodes based on the relationship between the nodes. Additional gene nodes and disease nodes as well as the edges between them were added using the EMBL-EBI GWAS catalog. The String database was used to encode edges between gene nodes for established protein-protein interactions. Condensate nodes and the edges between condensate nodes, gene nodes, disease nodes, and drug nodes were encoded with information from CD-code and condensate.com. Gene nodes and drug nodes were expanded and connected with edged using information from ChEMBL. FIG. 5 illustrates an example section of a gene network knowledge graph for T2D (filled circle). The gene nodes are depicted as striped circles. The arrows between the nodes represent the edges. The publicly available data included gene-disease associations, protein-protein interactions, protein pathways interactions, tissue-specific expression, and drug-gene interactions. Additional node types included known condensates, gene pathways, druggable compounds, pathways, and other diseases. For example, the checkered circles in the exemplary section of the graph in FIG. 5 may represent known condensate.
[0206] An integration value was calculated for the gene nodes in the T2D knowledge graph. The integration score took into account how strongly the gene node is linked to the T2D node and how selective the gene node is linked to the network formed by the 826 gene nodes associated with the T2D node. Genes with high integrator values had many connections to the T2D network and fewer connections to genes outside of the network. The integration value calculation included an algorithmic scoring including personalized page rank scoring, calculating the number of direct connections to the T2D network, and the proportion of direct connection compared to all connection for the gene node. A connection between the disease node and the gene node was classified as a direct connection if the gene node connected to the T2D network with a single edge. For example, if Gene X is known to interact with Gene Y, and Gene Y is in the disease gene network, then Gene X has (at least one) direct connection.
[0207] The significance of each integration value was calculated by performing a bootstrapped randomization of all edges in the graph and normalizing the resulting p-values using multiple comparison adjustments. The number of edges per gene were kept consistent for the bootstrapping randomization. The null hypothesis was that there is no significant difference between the page rank score for the gene node and the mean value for page rank scores generated for the gene node in 500 randomized graphs. Genes represented by gene nodes were considered significant if the corrected p-value was equal to or smaller than a corrected55MF-366315044Docket No.: 185992002640significance level (0.01 divided by number of observations). Exemplary results for 26 significant gene nodes are shown in Table 2. The integration values and p-values were benchmarked compared to a defined T2D positive control gene node.Table 2: Integration value significance for exemplary T2D genesGen PPR s pval adj Number_T2D_ Number_Total_ T2D_to_Total_ e core Neighbors Neighbors Neighbors Gen 0.8325 0 38 640 0.059375 e T 3Gen 0.8096 0 22 494 0.044534 e M 46Gen 0.7075 0 16 113 0.141593 e R 59Gen 0.6375 0 29 399 0.072682 e Z 59Gen 0.6354 0 27 158 0.170886 e E 56Gen 0.6277 0 32 365 0.087671 e S 04Gen 0.5847 0 21 118 0.177966 e H 52Gen 0.5563 0 14 122 0.114754 e L 04Gen 0.5517 0 29 566 0.051237 e I 84Gen 0.5286 1.15156265540 30 255 0.117647 e U 88 999e-111Gen 0.5285 0 16 116 0.137931 e Y 43Gen 0.4874 0 14 111 0.126126 e A 69Gen 0.4823 0 40 495 0.080808 e V 94Gen 0.4777 4.32869464854 23 218 0.105505 e B 75 489e-189Gen 0.4245 0 19 464 0.040948 e O 93Gen 0.4201 0 14 82 0.170732 e G 93Gen 0.4106 0 36 403 0.08933 e J 84Gen 0.3915 0 20 467 0.042827 e P 21Gen 0.3878 0 25 433 0.057737 e D 07Gen 0.3865 0 18 233 0.077253e K 356MF-366315044Docket No.: 185992002640Gen 0.3751 0 25 303 0.082508 e W 1Gen 0.3725 0 17 403 0.042184 e Q 59Gen 0.3698 0 16 351 0.045584 e X 92Gen 0.3674 0 14 276 0.050725 e C 62Gen 0.3639 0 20 337 0.059347 e F 7Gen 0.3595 0 18 350 0.051429eN 91* PPR score: personalized pagerank score with nodes corresponding to 826 T2D genes * pval adj: Bonferroni corrected p-value* Number_T2D_Neighbors: Number of direct neighbors in the knowledge graph to nodes that are part of the 826 T2D genes* Number Total Neighbors: Total number of direct neighbors to gene node in knowledge graph* T2D_toTotal_Neighbors: Number of T2D neighbors divided by Number Total Neighbors
[0208] Condensate information was obtained for the gene nodes. This information included the likelihood of phase separation for the gene as a proxy for condensate participation. The information was a prediction of phase separation for each gene node from PICNIC, Hardarovich et al, PICNIC accurately predicts condensate-forming proteins regardless of their structural disorder across organisms. 15 Nature Communication (2024). The genes nodes associated with gene with over a 0.5 predicted phase separation score from PICNIC were included for potential selection as a potential target genes for T2D.
[0209] Multidimensional prioritization of the potential target genes for validation using wet-lab experiments was performed. The strength of the genomics and transcriptomics based evidence for association between the gene and T2D was considered in combination with the values above. The nature and character of the gene’s connectivity in the knowledge graph was also analyzed and considered for prioritization and to expand the number of potential target genes that could be tested. This analysis was performed by clustering the potential target genes based on the strength and connectivity of the gene in the knowledge graph. FIG. 6 is a representative heatmap showing the clusters of potential target genes. The genes with high ranking integrator scores were clustered based on connections in the T2D knowledge graph. A gene set enrichment analysis was performed to discover if any clusters represented known gene sets, networks, or pathways. Sampling from diverse clusters for further validation helped derisk and increased the complimentary of programs where multiple targets were pursued. Clusters with known potential safety issued were selected to reduce the risk of on-pathway toxicity in 57MF-366315044Docket No.: 185992002640drug development. Clusters with novel connectivity’s were selected to increase the potential for identification of novel therapeutic options. The methods resulted in 200 potential target genes for further validation.Example 2- Image and omics based validation of condensate targets for insulin resistance T2D
[0210] Various validation tests were performed to confirm the effect predicted by the methods in example 1. The validation assays described herein are an exemplary subset of the assays that can be used.
[0211] The HepG2 cell line was modified to make it insulin resistant in order to model T2D. The non-disease (healthy) and disease state were modeled by growing the modified Hep2G cells with insulin in corresponding media for at least 24 hours. The non-disease (healthy) state received 0.1nm insulin and the disease condition received lOnM insulin, glucose, and palmitate. In some experiments an additional cell IGF 1 -receptor knock out cell line (IGF1R-KO HepG2) was also used to model the non-disease (healthy) and disease states. The 200 target genes identified in example 1 were tagged with fluorescent tagged antibodies in cells representing the non-disease (healthy) state and cells representing the disease state and high content fluorescent imaging was conducted. The intensity of the fluorescence in the cell images was used as a proxy for condensate strength in each image. The relative phenotypic change of condensate in the disease model was calculated for each of the target genes as the difference between the fluorescence in non-disease (healthy) and disease cells. The significance strength of phenotype driving feature was calculated using a Z prime value. 66 T2D target genes had Z prime values greater than 0.3 and were considered validated by the imaging experiment. These genes were expected to influence T2D through condensate activity of their expressed proteins.
[0212] siRNA methods were used to knockdown and overexpress the T2D target genes individually in both non-disease and diseased cells. RNA sequencing (RNAseq) was collected for cells from each condition. Featurization of each state and perturbation was performed using dimensionality reduction methods, such as PCA, Boruta, and DEG. Cosine similarity of the knockdown expression vs expression in the non-disease and disease state was calculated. The “on target” score was the location of the T2D target gene on the vector between the disease state and the non-disease state and the “Off target” score was how far the T2D target gene was in any other direction in the feature space. T2D genes were validated in the siRNA study if58MF-366315044Docket No.: 185992002640they result in a shift of the expression feature vector from the disease state toward the nondisease (healthy) state. (FIG. 7).Example 3- Identification of Beta catenin (Beat) as a condensate target
[0213] This example demonstrates the method of identifying and validating condensate target genes as described herein is effective at identifying a known condensate related target gene.
[0214] Constitutive expression of Beat is a driver of malignancy in colorectal cancer (CRC). Beat is a known condensate related target gene for CRC because Beat condensates can lead to selective inducement of cancer cell death and reversal of the oncogenic effects of Beat gene expression.
[0215] CRC related genes were identified as genes with recurrent somatic mutation in CRC, The Cancer Genome Atlas Network. Nature 487, 330-337 (2012), and genes with proteomic differences between paired tumor and normal adjacent tissues, Vasaikar et al., Cell 177(4) (2019). The identified CRC genes were, ACVR1B, AMER1, APC, ATP6V0D2, BRAF, CASP8, CDC27, CTNNB1 (Beat), DMD, EDNRB, FBXW7, FZD3, CALNT17, GPC6, GRIK3, KRAS, MAP3K21, MAP7, MIER3, MY01B, NRAS, PIK3CA, PTPN12, SLC9A9, SMAD2, SMAD4, SOX9, TCERG1, TCF7L2, TP53, TTN, ACVR2A, CASP5, RFX5, RNF43, LTN1, SNRNP40, TGFBR2.
[0216] A gene network knowledge graph was created using the CRC genes as gene nodes. Additional nodes were added to the gene network knowledge graph by adding genes with direct protein interactions to the CRC genes, canonical condensates containing any of the genes represented in the gene nodes, and compounds that have been shown to bind to or modify DNA, RNA, or protein for the genes in the gene nodes. The CRC gene network knowledge graph comprised 1322 genes and 38 condensates.
[0217] An integration value was calculated for the gene nodes using the number of CRC neighbors (FIG. 8A) and a personalized page rank (FIG.8B) and the value was plotted against a condensate prediction score outputted from PICNIC for each gene. A gene was considered a CRC neighbor if there was a direct connection between the gene and a CRC -associated gene in the graph. In FIG. 8A and FIG. 8B, the CRC genes are represented by name and in red and the genes added by expanding the knowledge graph are represented in yellow. CTNNB1 (Beat) was identified as a gene with a high integration value using both methods and a gene likely to form a condensate (condensate prediction score above 0.5).59MF-366315044Docket No.: 185992002640
[0218] The section of the CRC gene network knowledge graph is illustrated in FIG. 9A.Three previously known compounds were found to have connections to the CRC gene network knowledge graph. The three compounds target distinct gene nodes connected to Beat (FIG.9B). This suggested different mechanism of actions for the compounds.
[0219] HCT116 cells were treated with the Capecitabin, NCB-0846, SSTC3, and DMSO as a negative control at a dose of 30pM. The cells were incubated in a humidity-controlled incubator at 37°C for 24 hours prior to compound treatment and 24 hours after compound treatment. Following incubation, cells were fixed using a 3% formaldehyde solution. The fixed cells were stained with DAPI and CellMask Blue dyes to visualize DNA and cytoplasm respectively. Antibody staining was used to target Phospho-beta catenin condensates (Thermo Fischer 23H16L13). The condensates were visualized using a secondary antibody.
[0220] High-content confocal microscopy was used for high-throughput microscopy. The nuclei are marked with DAPI (blue) and Beat condensates are marked with Phospho-beta catenin (red). As shown in FIG. 10A, the treatment with NCB-0846 and SSTC3 lead to Beat condensates in the nucleus, but treatment with DMSO or Capectabine did not.
[0221] A dose response study was performed for NCB-0846 and SSTC3. HCT116 cells were treated with the two compounds at different doses for either 4 or 24 hours. After incubation, fixing, and tagging as described above, the number of Beat condensates in the nucleus (spot No) normalized by spot counts in DMSO negative controls were calculated. As shown in FIG. 10B, NCB-0846 and SSTC3 induced a Beat condensates in a dose dependent manner as would be expected if the drug caused the phenotypic change. Additional assays were performed to show treatment with NCB-0846 and SSTC3 killed CRC cells and led to decreased expression of Beat target genes (data not shown).
[0222] The results showed the CRC knowledge graph approach could be used to identify Beat as a condensate related target gene and to identify drugs that had known links to other CRC genes using the Beat condensate phenotype. The drugs were experimentally validated to effect Beat, down regulate Beat target genes and have cell killing effect on colorectal cancer.Example 4- Cell fraction profiles for T2D
[0223] This example demonstrates mass spectrometry protein quantification according to methods described herein. Mass spectrometry was used to determine cell fraction profiles for type 2 diabetes (T2D). The non-disease (healthy) and disease cells as described in example 260MF-366315044Docket No.: 185992002640were subjected to soluble proteome profiling to identify condensate-bound protein subpopulations.
[0224] Cells were lysed and the lysate was subjected to centrifugation. The supernatants containing soluble proteins were obtained and the insoluble proteins in the pellet were subjected to two detergents, one mild (NP-40) and one strong (SDS) and solubilized for processing (see solubility profiling methods in Sridiharan et al., Systemic discovery of biomolecular condensate-specific protein phosphorylation, 18 Nature Chemical Biology, 1104-1114 (2022)). The subpopulations (previously soluble and insoluble) were separately analyzed by mass spectrometry to determined protein levels. The insoluble proteins were assumed to be in condensates.
[0225] Using the cell fraction profiles, gene nodes with significant changes in protein levels of the insoluble fraction (when comparing measurements in lysate from the non-disease (healthy) and disease conditions) that had at least one disease relevant neighbor in the T2D knowledge graph with a significant change in protein levels in the insoluble fractions themselves were prioritized for further analysis. It was found that of the about 40 T2D genes with significant integration scores from the T2D knowledge graph, about 22 of the genes were validated using the cell fraction profile (data not shown).61MF-366315044
Claims
Docket No.: 185992002640CLAIMS1. A computer implemented method for identifying one or more condensate related target genes for a disease, comprising:generating a gene-network knowledge graph comprising at least a plurality of nodes, wherein the plurality of nodes comprise at least one gene node, at least one disease node, and edges representing a relationship between each node;calculating an integration value for one or more gene nodes in the gene-network knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes of the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the genenetwork knowledge graph;obtaining condensate information for one or more gene nodes in the gene-network knowledge graph;identifying one or more condensate related target genes from the gene-network knowledge graph using the condensate information and the integration value for the corresponding gene node.
2. The computer implemented method of claim 1, wherein the plurality of nodes further comprises at least one gene pathway node.
3. The computer implemented method of claim 1 or claim 2, wherein the plurality of nodes further comprises at least one condensate node.
4. The computer implemented method of any of claims 1-3, wherein the plurality of nodes further comprises at least one drug node.
5. The computer implemented method of any of claims 1-4, wherein the relationship between each node comprises a relationship selected from a group consisting of gene-disease association, a protein-protein association, an RNA-protein association, a protein-pathway interaction, a tissue specific expression relationship, a drug-gene interaction, a genecondensate interaction, and an RNA-condensate interaction.62MF-366315044Docket No.: 1859920026406. The computer implemented method of any of claims 1-4, wherein the relationship between each node comprises a relationship between two nodes obtained from results of a statistical genomics analysis a statistical transcriptomics analysis, or a proteomic analysis.
7. The computer implemented method of claim 6, wherein the statistical genomics analysis comprises GWAS, eQTL, pQTL, PheWAS, Polygenic gene prioritization (POP), or rare variant aggregate testing.
8. The computer implemented method of claim 6 or claim 7, wherein the statistical transcriptomics analysis comprises differential gene expression, pathway level interference, gene dependence analysis, gene perturbation analysis, foundational model gene interaction mapping, or gene regulatory network mapping.
9. The computer implemented method of any of claims 1-8, wherein calculating an integration value comprises weighting the number of edges connecting the corresponding gene node to a node in one or more selected nodes the gene-network knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph by characteristics of the edges connecting each node.
10. The computer implemented method of any of claims 1-8, wherein calculating an integration value comprises performing a personalized page rank analysis.
11. The computer implemented method of any of claims 1-10, wherein the one or more selected nodes have been selected based on a known relationship to the disease.
12. The computer implemented method of any of claims 1-11, wherein the one or more selected nodes comprise one or more nodes connected to the disease node associated with the disease by less than a predetermined number of edges.
13. The computer implemented method of claim 12, wherein the predetermined number of edges is 1, 5, 10, 15, or 20.63MF-366315044Docket No.: 18599200264014. The computer implemented method of any of claims 1-13, wherein the integration value represents connectivity and selectivity of the gene node to the disease node representing the disease.
15. The computer implemented method of any of claims 1-14, wherein the condensate information comprises predicted phase separation for each of one or more polypeptides corresponding to one or more gene nodes.
16. The computer implemented method of claim 15, wherein the predicted phase separation is generated from a protein sequence corresponding to the gene node.
17. The computer implemented method of any of claims 1-16, wherein the condensate information comprises a disease state phase transition characteristic for each of one or more polypeptides corresponding to one or more gene nodes.
18. The computer implemented method of claim 17, wherein the condensate information has been obtained using a method comprising: determining the disease state phase transition characteristic for each of the one or more polypeptides.
19. The computer implemented method of claim 18, wherein the determining the disease state phase transition characteristic for a first polypeptide of the one or more polypeptides comprises performing one or more of the following comparisons:the quantity of the first polypeptide in an insoluble fraction of a disease model cell lysate as compared to the quantity of the first polypeptide in an insoluble fraction of a nondisease model cell lysate; orthe quantity of the first polypeptide in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the first polypeptide in a soluble fraction of a non-disease model cell lysate versus a non-disease model cell lysate.
20. The computer implemented method of claim 19, further comprising quantifying the first polypeptide to perform the one or more comparisons.64MF-366315044Docket No.: 18599200264021. The computer implemented method of claim 19, further comprising quantifying the first polypeptide in the insoluble fraction of the disease model cell lysate and the first polypeptide in the insoluble fraction of the non-disease model cell lysate.
22. The computer implemented method of claim 19, further comprising quantifying the first polypeptide in the soluble fraction of the disease model cell lysate and the first polypeptide in the soluble fraction of the non-disease model cell lysate.
23. The computer implemented method of claim 19, further comprising quantifying the first polypeptide in the disease model cell lysate and the first polypeptide in the non-disease model cell lysate.
24. The computer implemented method of any of claims 20-23, wherein the quantifying comprises performing a quantitative mass spectrometry technique.
25. The computer implemented method of claim 24, wherein the quantitative mass spectrometry technique comprises use of isobaric labeling.
26. The computer implemented method of claim 24, wherein the quantitative mass spectrometry technique comprises use of tandem mass tag (TMT) or isobaric tags for relative and absolute quantification (iTRAQ).
27. The computer implemented method of any of claims 24-26, wherein the quantitative mass spectrometry technique is multiplexed.
28. The computer implemented method of any of claims 19-27, further comprising fractionating a disease model cell lysate and / or a non-disease model cell lysate.
29. The computer implemented method of any of claims 19-28, further comprising obtaining a cell lysate from the disease model and / or the non-disease model.
30. The computer implemented method of claims 19-29, further comprising lysing, separately, a sample from the disease model to obtain the disease model cell lysate and a sample from the non-disease model to obtain the non-disease model cell lysate.65MF-366315044Docket No.: 18599200264031. The computer implemented method of any of claims 19-30, wherein the disease model comprises a cell model for a disease in a disease state and / or the non-disease model comprises a cell model for a disease in a non-disease state.
32. The computer implemented method of any of claims 1-31, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes with a significant integration value.
33. The computer implemented method of any of claims 1-32, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes that cluster together based on integration values and condensate information.
34. The computer implemented method of any of claims 1-33, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes using condensate phenotype information.
35. The computer implemented method of any of claims 1-34, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes using genetic perturbation information.
36. The computer implemented method of any of claims 1-35, wherein identifying one or more condensate related target genes comprises identifying one or more gene nodes using compound modulating condensates information.
37. A method of identifying one or more condensate related target genes for a disease, the method comprising:generating a gene-network knowledge graph comprising at least a plurality of nodes, wherein the plurality of nodes comprise at least one gene node, at least one disease node, and edges representing a relationship between each node;calculating an integration value for one or more gene nodes in the genenetwork knowledge graph based on the number of edges connecting the corresponding gene node to a node in one or more selected nodes the gene-network66MF-366315044Docket No.: 185992002640knowledge graph and the number of edges connecting the gene to a node outside of the one or more selected nodes in the gene-network knowledge graph;obtaining a disease state phase transition characteristic for one or more gene nodes in the gene-network knowledge graph by a method comprising:obtaining a disease model cell lysate and a non-disease model cell lysate;fractionating, separately, the disease model cell lysate and the nondisease model cell lysate to obtain respective insoluble and soluble fractions thereof;performing a quantitative mass spectrometry technique on one or more of:the soluble fraction of the disease model cell lysate, the disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate; orthe insoluble fraction of the disease model cell lysate, the soluble fraction of the disease model cell lysate, the disease model cell lysate, the insoluble fraction of the non-disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non- disease model cell lysate;determining a disease state phase transition characteristic for each of one or more polypeptides corresponding to the one or more gene nodes based on one or more of the following comparisons:the quantity of a polypeptide in an insoluble fraction of a disease model cell lysate as compared to the quantity of the polypeptide in an insoluble fraction of a non-disease model cell lysate; or the quantity of a polypeptide in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the polypeptide in a soluble fraction of a non-disease model cell lysate versus a non-disease model cell lysate; andidentifying one or more condensate related target genes from the gene-network knowledge graph using the disease state phase transition characteristic and the integration value for the corresponding gene node.67MF-366315044Docket No.: 18599200264038. A method of identifying a one or more condensate related target genes for a disease, the method comprising:obtaining a disease model cell lysate and a non-disease model cell lysate; fractionating, separately, the disease model cell lysate and the non-disease model cell lysate to obtain respective insoluble and soluble fractions thereof; performing a quantitative mass spectrometry technique on one or more of:the insoluble fraction of the disease model cell lysate and the insoluble fraction of the non-disease model cell lysate;the soluble fraction of the disease model cell lysate, the disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate; orthe insoluble fraction of the disease model cell lysate, the soluble fraction of the disease model cell lysate, the disease model cell lysate, the insoluble fraction of the non-disease model cell lysate, the soluble fraction of the non-disease model cell lysate, and the non-disease model cell lysate;determining a disease environment phase transition characteristic for each of one or more polypeptides based on one or more of the following comparisons: the quantity of a polypeptide in an insoluble fraction of a disease model cell lysate as compared to the quantity of the polypeptide in an insoluble fraction of a non-disease model cell lysate; orthe quantity of a polypeptide in a soluble fraction of the disease model cell lysate versus the disease model cell lysate as compared to the quantity of the polypeptide in a soluble fraction of a non-disease model cell lysate versus a non-disease model cell lysate; and filtering a gene-network knowledge graph using the disease environment phase transition characteristics to identify the one or more condensate related target genes from the disease.
39. A system for identifying one or more condensate related target genes for a disease, the system comprising:one or more processors,a user input device, and68MF-366315044Docket No.: 185992002640a memory communicatively coupled to the one or more processors configured to store instructions that, when executed by the one or more processors, cause the systems to:perform the computer implemented method of any one of claims 1-36.
40. A computer-readable non-transitory storage medium storing one or more programs, the one or more programs comprising instructions that when executed by one or more processors of a system cause the system to:perform the computer implemented method of any one of claims 1-36.69MF-366315044