Methods and systems for identifying compounds that form, stabilize, or disrupt molecular complexes

By combining size exclusion chromatography and mass spectrometry with metabolomic analysis, the problem of difficult to identify protein-protein interaction regulators in the prior art is solved, and efficient identification of protein complex components and regulatory compounds is achieved.

CN120303045APending Publication Date: 2025-07-11ENVEDA THERAPEUTICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380081283.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-11-26
Filing Date
2023-11-22
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, when identifying regulators of protein-protein interactions, there is a lack of high-throughput and unbiased method, and it is difficult to query regulators of multiple protein complexes at the same time.

Method used

Size exclusion chromatography is used to fractionate biological samples and compound libraries or natural extracts, analyze fraction shifts, combine mass spectrometry and metabolomics to identify components and regulatory compounds of protein complexes, and identify complex formation or dissociation caused by compound by fraction shifts.

Benefits of technology

The identification of high-throughput, unbiased protein-protein complex regulators is achieved, and the impact of multiple compounds on the complex is simultaneously recognized, improving the identification efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303045A_ABST
    Figure CN120303045A_ABST
Patent Text Reader

Abstract

Methods for identifying components of protein-protein complexes and methods for identifying one or more compounds that cause the formation, stabilization, or dissociation of the protein complex are described herein. Systems for performing such methods are also described. The method may include fractionating a first sample containing a first portion of a biological sample and a second sample containing a second portion of the biological sample in combination with a drug, a library of compounds, or a natural extract. The elution fraction may be analyzed using a proteomic or metabonomic method to identify one or more binding proteins that form a complex with the target protein, or one or more compounds that cause complex formation, stabilization, or dissociation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the priority benefit of Indian Application No. 202211068104, filed on November 26, 2022, the content of which is incorporated herein by reference for all purposes. Technical field

[0003] Systems and methods for identifying components of protein - protein complexes are described herein. Systems and methods for identifying compounds that cause complex formation, stabilization, or dissociation are also described herein. Background art

[0004] Protein - protein interactions (PPIs) are central to a variety of biological processes, and their dysfunction is associated with the pathogenesis of a range of human diseases and disorders. The contact interface between two proteins is the structural basis of their interaction. Similar or overlapping protein interfaces can be promiscuous and are adopted multiple times in hub proteins. PPIs can be transient or permanent, homophilic or heterophilic, and specific or non - specific, and can be regulated through signal transduction (biochemical) cascades. Thus, the ability to modulate disease - related protein - protein interactions (PPIs) using small - molecule inhibitors is an important diagnostic and therapeutic strategy.

[0005] The prior art for querying PPI modulators relies on low - throughput methods, such as the "bait and prey" assay. Thus, there is increasing interest in developing high - throughput, multiplexed methods that query multiple modulators of protein complexes in an unbiased manner. Summary of the invention

[0006] Systems and methods for identifying components of protein - protein complexes are described herein. Systems and methods for identifying compounds that cause complex formation, stabilization, or dissociation are also described herein.

[0007] A method for identifying components of a protein-protein complex may include: fractionating a first sample using size exclusion chromatography to produce a first plurality of fractions, the first sample comprising a first protein-containing portion of a biological sample; fractionating a second sample using size exclusion chromatography to produce a second plurality of fractions, the second sample comprising (i) a second portion of the biological sample and (ii) a compound library, a drug, or a natural extract; analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions; for a target protein, identifying a fraction shift between the first plurality of fractions and the second plurality of fractions, wherein the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds or a drug in the compound library or natural extract; and identifying one or more binding proteins that form a complex with the target protein. Co-elution of one or more binding proteins with the target protein in the second plurality of fractions but not in the first plurality of fractions indicates that a compound or a drug in the compound library or natural extract causes formation or stabilization of a complex comprising the target protein and at least one of the one or more binding proteins. Co-elution of one or more binding proteins with the target protein in the first plurality of fractions but not in the second plurality of fractions indicates that a compound or a drug in the compound library or natural extract causes dissociation of a complex comprising the target protein and at least one of the one or more binding proteins. Co-elution may be determined, for example, based on the peak elution fractions of the target protein and the one or more binding proteins. Co-elution of one or more binding proteins with the target protein may be based on, for example, the peak elution fractions of the one or more binding proteins and the peak elution fraction of the target protein. The method may further include selecting, as members of the complex, one or more of the one or more binding proteins based on the molecular weights of the one or more putative binding proteins and the target protein and the fraction numbers of the fractions comprising the target protein and the one or more putative binding proteins.

[0008] A method for identifying one or more compounds that cause protein complex formation, stabilization, or dissociation may include: fractionating a first sample using size-exclusion chromatography to produce a first plurality of fractions, the first sample comprising a protein-containing first portion of a biological sample; fractionating a second sample using size-exclusion chromatography to produce a second plurality of fractions, the second sample comprising (i) a second portion of the biological sample and (ii) a compound library or natural extract; analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions; for a target protein, identifying a fraction shift between the first plurality of fractions and the second plurality of fractions, wherein the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds in the compound library or natural extract; and identifying one or more compounds that cause the fraction shift, including analyzing: (1) for a fraction shift indicating complex formation or stabilization caused by one or more compounds, the fraction in the second plurality of fractions that contains the target protein, to identify one or more compounds that co-elute with the target protein in the fraction, or (2) for a fraction shift indicating complex dissociation caused by one or more compounds, (i) the fraction in the second plurality of fractions that contains the target protein, to identify one or more compounds that co-elute with the target protein in the fraction, or (ii) the fraction in the second plurality of fractions that contains a binding protein that forms a complex with the target protein in the absence of the one or more compounds, to identify one or more compounds that co-elute with the binding protein in the fraction.

[0009] Identifying one or more compounds that cause the fraction shift may include, for example, obtaining a metabolomic profile of the fractions in the second plurality of fractions; obtaining a metabolomic profile of the compound library or natural extract; and identifying one or more compounds present in the fractions and the compound library or natural extract. In some embodiments, identifying one or more compounds that cause the fraction shift further includes obtaining a metabolomic profile of the first sample; and filtering the metabolomic profile of the fractions in the second plurality of fractions to exclude compounds present in the first sample.

[0010] The metabolomic profile may be obtained using mass spectrometry. For example, in some embodiments, the metabolomic profile is obtained using liquid chromatography and tandem mass spectrometry (LC-MS / MS). In some embodiments, the metabolomic profile is obtained using computer nuclear magnetic resonance (NMR).

[0011] Identifying one or more compounds that cause fraction shift includes confirming the co-elution of one or more compounds and a target protein or one or more binding proteins, which can be based on the peak elution fractions of the one or more compounds and the peak elution fractions of the target protein or binding protein. In some specific embodiments, the method can include identifying a binding protein that forms a complex with the target protein, wherein the co-elution of the binding protein and the target protein in a first plurality of fractions rather than a second plurality of fractions indicates that one or more compounds in a compound library or a natural extract cause the dissociation of a complex comprising the target protein and the binding protein.

[0012] In some specific embodiments of the above method, analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions includes proteomic analysis. In some specific embodiments of the above method, analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions includes using mass spectrometry. For example, analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions can include using liquid chromatography tandem mass spectrometry (LC-MS / MS).

[0013] In some specific embodiments of the above method, the compound library or natural extract is substantially free of proteins.

[0014] In some specific embodiments of the above method, the biological sample includes a cell-free biological sample, a tissue extract, a cell extract, or a subcellular extract.

[0015] In some specific embodiments of the above method, the second sample contains a compound library.

[0016] In some specific embodiments of the above method, the second sample contains a natural extract.

[0017] In some specific embodiments of the above method, the natural extract is a plant extract.

[0018] In some specific embodiments of the above method, the biological sample containing proteins is obtained from a cell lysate.

[0019] In some specific embodiments of the above method, the biological sample containing proteins is obtained from animal tissue.

[0020] In some specific embodiments of the above method, the biological sample containing proteins is obtained from mammalian tissue.

[0021] In some specific embodiments of the above method, the biological sample containing proteins is obtained from brain, liver, lung, or kidney tissue.

[0022] In some specific implementations, the system includes one or more processors; and a non-transitory computer-readable storage medium storing one or more programs, the one or more programs, when executed by the one or more processors, cause the system to: receive first proteomics spectral data of a first plurality of fractions obtained by fractionating a first sample using size-exclusion chromatography, the first sample comprising a protein-containing portion of a biological sample; receive second proteomics spectral data of a second plurality of fractions obtained by fractionating a second sample using size-exclusion chromatography, the second sample comprising (i) a second portion of the biological sample and (ii) a compound library, a drug, or a natural extract; identify proteins in the first plurality of fractions and the second plurality of fractions based on the first proteomics spectral data and the second proteomics data; for a target protein, identify a fraction shift between the first plurality of fractions and the second plurality of fractions, wherein the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds or a drug in the compound library or natural extract; and identify one or more binding proteins that form a complex with the target protein. Co-elution of one or more binding proteins with the target protein in the second plurality of fractions but not in the first plurality of fractions indicates the formation or stabilization of a complex comprising the target protein and at least one of the one or more binding proteins caused by a compound or a drug in the compound library or natural extract. Co-elution of one or more binding proteins with the target protein in the first plurality of fractions but not in the second plurality of fractions indicates the dissociation of a complex comprising the target protein and at least one of the one or more binding proteins caused by a compound or a drug in the compound library or natural extract.

[0023] In some embodiments, the system includes one or more processors; and a non-transitory computer-readable storage medium storing one or more programs, which when executed by the one or more processors cause the system to: receive first proteomic spectral data of a first plurality of fractions obtained by fractionating a first sample using size-exclusion chromatography, the first sample comprising a protein-containing portion of a biological sample; receive second proteomic spectral data of a second plurality of fractions obtained by fractionating a second sample using size-exclusion chromatography, the second sample comprising (i) a second portion of the biological sample and (ii) a compound library or a natural extract; and identify proteins in the first plurality of fractions and the second plurality of fractions based on the first proteomic spectral data and the second proteomic data; for a target protein, identify a fraction shift between the first plurality of fractions and the second plurality of fractions, wherein the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds in the compound library or the natural extract; receive metabolomic data of the compound library or the natural extract; receive metabolomic data of the fraction in the second plurality of fractions that contains the target protein or a binding protein that forms a complex with the target protein in the absence of one or more compounds; and identify one or more compounds that cause the fraction shift by analyzing the metabolomic data of the compound library or the natural extract and the metabolomic data of the fraction in the second plurality of fractions that contains the target protein or a binding protein that forms a complex with the target protein in the absence of one or more compounds. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1A FIG. shows an exemplary method for identifying protein-protein interactions according to some embodiments.

[0025] Figure 1B FIG. shows an exemplary method for determining complex formation or complex dissociation according to some embodiments, which may include in Figure 1A the exemplary method shown.

[0026] Figure 2A FIG. shows an exemplary method for identifying compounds that modulate protein complexes according to some embodiments.

[0027] Figure 2B FIG. shows an exemplary method for determining complex formation or complex dissociation according to some embodiments.

[0028] Figure 3 FIG. shows an exemplary schematic diagram visualizing compound-protein interactions according to some embodiments. Interactions are prioritized based on algorithm scores of metabolite-protein interaction confidence.

[0029] Figure 4Shows an exemplary method for identifying components of a protein-protein complex and compounds that cause protein complex formation from a pool of known or novel compounds, according to some embodiments.

[0030] Figure 5 Shows a visualization of compound-protein interaction maps of RNF114 in the presence or absence of natural extract P45, according to an exemplary experiment. Interactions are prioritized based on an algorithm score of compound-protein interaction confidence.

[0031] Figure 6 Shows a visualization of compound-protein interaction maps of RNF114 and binding proteins (gray circles) in the presence or absence of natural extract P45, according to an exemplary experiment. Interactions are prioritized based on an algorithm score of metabolite-protein interaction confidence.

[0032] Figure 7 Shows a peak analysis graph of RNF114 and binding proteins in the presence or absence of natural extract P45, according to some embodiments.

[0033] Figure 8 Shows an exemplary system that can be used in conjunction with the methods described herein, according to some embodiments.

[0034] Figure 9 Shows an exemplary system for use in conjunction with the methods described herein, according to some embodiments. Detailed Description

[0035] Methods and systems for identifying components of protein-protein interactions are described herein. Proteins in a sample of a biological sample containing proteins (e.g., a lysate from a mammalian source, such as a tissue, cell, or subcellular lysate, or a cell-free biological sample, such as an extract from saliva, cerebrospinal fluid, plasma, etc.) are fractionated based on apparent molecular weight or size. Another portion of the same biological sample is combined with a compound library, natural extract, or drug from a different source (e.g., a plant) and similarly fractionated based on apparent molecular weight or size. The compound library, natural extract, or drug combined with the biological sample preferably does not contain proteins or is substantially free of proteins. Proteins that are complexed with each other migrate together at the complexed apparent size. The fractions are analyzed to determine the identity of the proteins contained in each fraction.

[0036] Each fraction is associated with an elution volume (i.e., the volume of buffer eluted from the column prior to that fraction), which may be related to the apparent weight or size of the protein (usually assumed based on the hydrodynamic diameter). Fraction shift is the change in the elution volume (or fraction count, which is associated with the elution volume) of a protein under different conditions. As described herein, the fraction shift of a target protein is the difference in the elution volume or fraction count of the target protein in the presence and absence of a compound library, natural extract, or drug. Thus, fraction shift indicates protein-protein association (e.g., formation and / or stabilization) or dissociation events caused by a compound library, natural extract, or drug.

[0037] Binding proteins can be more confidently identified based on co-migration with the target protein, and the identities of these binding proteins can be determined by proteomics. Methods and systems for identifying compounds that cause protein complex formation or stabilization or dissociation are also described herein. In some specific embodiments of the method, proteins in a sample of a biological sample (e.g., a lysate from a mammalian source, such as a tissue, cell, or subcellular extract, or a cell-free biological sample, such as an extract from saliva, cerebrospinal fluid, plasma, etc.) are fractionated based on apparent molecular weight or size. Another portion of the same biological sample is combined with a compound library, natural extract, or drug from a different source (e.g., a plant extract) and similarly fractionated based on apparent molecular weight or size. Proteins that complex with each other migrate together at the complexed apparent size. The fractions are separated for further analysis. One portion of the fractions is analyzed to determine the identity of the proteins contained in each fraction. Another portion of the fractions is analyzed to determine the identity of the compounds contained in the fractions.

[0038] In some specific embodiments of a method for identifying components of a protein-protein complex, the method includes fractionating a first sample using size-exclusion chromatography to produce a first plurality of fractions, the first sample comprising a protein-containing first portion of a biological sample (e.g., a cell-free biological sample, a tissue extract, a cell extract, or a subcellular extract). A second sample is also fractionated using size-exclusion chromatography to produce a second plurality of fractions, the second sample comprising (i) a second portion of the biological sample and (ii) a compound library, a drug, or a natural extract. The first plurality of fractions and the second plurality of fractions are analyzed to identify proteins in the first plurality of fractions and the second plurality of fractions. For a target protein, a fraction shift between the first plurality of fractions and the second plurality of fractions can then be identified. The fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds or a drug in the compound library or natural extract. One or more binding proteins that form a complex with the target protein can then be identified. For example, co-elution of one or more binding proteins with the target protein in the second plurality of fractions but not in the first plurality of fractions indicates that a compound or a drug in the compound library or natural extract causes formation or stabilization of a complex comprising the target protein and at least one of the one or more binding proteins. Co-elution of one or more binding proteins with the target protein in the first plurality of fractions but not in the second plurality of fractions indicates that a compound or a drug in the compound library or natural extract causes dissociation of a complex comprising the target protein and at least one of the one or more binding proteins. Co-elution can be determined, for example, based on the peak elution fractions of the target protein and the one or more binding proteins. In some cases, the binding proteins are identified as binding proteins. Thus, in some specific embodiments, the method further includes selecting one or more of the one or more binding proteins as members of the complex based on the molecular weights of the one or more binding proteins and the target protein and the fraction numbers of the fractions comprising the target protein and the one or more binding proteins.

[0039] A method for identifying one or more compounds that cause the formation or dissociation of a protein complex can include fractionating a first sample containing a proteinaceous first portion of a biological sample (e.g., a cell-free biological sample, a tissue extract, a cell extract, or a subcellular extract) using size exclusion chromatography to produce a first plurality of fractions. A second sample containing (i) a second portion of the biological sample and (ii) a compound library or a natural extract can also be fractionated using size exclusion chromatography to produce a second plurality of fractions. The first plurality of fractions and the second plurality of fractions can be analyzed to identify the proteins in the first plurality of fractions and the second plurality of fractions. For a target protein, a fraction shift between the first plurality of fractions and the second plurality of fractions can be identified. The fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds in the compound library or the natural extract. Thereby, one or more compounds that cause the fraction shift can be identified. For example, for a fraction shift indicating complex formation or stabilization caused by one or more compounds, the fractions in the second plurality of fractions containing the target protein can be analyzed to identify one or more compounds that co-elute with the target protein in the fraction. For a fraction shift indicating complex dissociation caused by one or more compounds, (i) the fractions in the second plurality of fractions containing the target protein can be analyzed to identify one or more compounds that co-elute with the target protein in the fraction, or (ii) the fractions in the second plurality of fractions containing a binding protein that forms a complex with the target protein in the absence of one or more compounds can be analyzed to identify one or more compounds that co-elute with the binding protein in the fraction. Identifying one or more compounds that cause the fraction shift can include obtaining a metabolomics profile of the fractions in the second plurality of fractions; obtaining a metabolomics profile of the compound library or the natural extract; and identifying one or more compounds present in both the fractions and the compound library or the natural extract. Optionally, identifying one or more compounds that cause the fraction shift further includes obtaining a metabolomics profile of the first sample; and filtering the metabolomics profile of the fractions in the second plurality of fractions to exclude compounds present in the first sample. Identifying one or more compounds that cause the fraction shift can optionally include confirming the co-elution of one or more compounds and the target protein or one or more binding proteins based on the peak elution fractions of the one or more compounds and the peak elution fractions of the target protein or the binding protein. The method can also optionally include identifying a binding protein that forms a complex with the target protein, wherein the co-elution of the binding protein and the target protein in the first plurality of fractions but not in the second plurality of fractions indicates that one or more compounds in the compound library or the natural extract cause the dissociation of the complex containing the target protein and the binding protein.

[0040] This document further describes systems that can be used to implement the methods described herein. Such systems can include one or more processors and a non-transitory computer-readable storage medium storing one or more programs, which when executed by the one or more processors cause the system to perform the method steps described herein. The system can also include a liquid chromatography system (which can include a size exclusion chromatography column and / or a reverse phase chromatography column) and / or a mass spectrometry (or tandem mass spectrometry) system.

[0041] Definitions

[0042] As used in this specification and the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include plural referents. Thus, for example, a reference to "a cell" optionally includes a combination of two or more such cells, and the like.

[0043] As used herein, the terms "about" and "approximately" refer to the usual error range of the corresponding values that are readily known to those skilled in the art. Exemplary degrees of error are within 20% of a given value or range of values, typically within 10%, and more typically within 5%. A reference herein to a "about" or "approximately" value or parameter includes (and describes) embodiments that relate to that value or parameter itself.

[0044] It should be understood that the aspects and embodiments of the invention described herein include "comprising" aspects and embodiments, "consisting of" aspects and embodiments, and "consisting essentially of" aspects and embodiments.

[0045] A "compound library" is any collection of a plurality of compounds. A compound library can contain any number of compounds of two or more, such as five or more, fifty or more, one hundred or more, five hundred or more, or one thousand or more small molecule compounds. A compound library can be a small molecule library.

[0046] An "extract" is a biological material that has been processed to remove or substantially remove one or more components of the material. For example, an extract can be processed to remove one or more of fat, carbohydrates, or proteins. An extract can contain protein or can be substantially protein-free. An extract is considered to be "substantially protein-free" if it contains 5% (by mass) or less of its original protein content (i.e., the natural state of the biological material).

[0047] As used herein, the term "sample" refers to a composition obtained or derived from a subject and / or individual of interest, which contains cells and / or other molecular entities that will be characterized and / or identified, for example, based on physical, biochemical, chemical, and / or physiological characteristics. A sample can be tissue, cells, subcellular structures (e.g., organelles), or cell-free biological samples (e.g., saliva, plasma, cerebrospinal fluid, etc.) or can be an extract from them.

[0048] As used herein, the terms "individual", "patient", or "subject" are used interchangeably and refer to any single animal in need of treatment, such as a mammal (including non-human animals, such as dogs, cats, horses, rabbits, zoo animals, cows, pigs, sheep, and non-human primates). In a specific embodiment, the patient herein is a human.

[0049] A "small molecule" is any molecule having a molecular weight of 1000 daltons or less.

[0050] Where a range of values is provided, it is understood that each intermediate value between the upper and lower limits of that range, as well as any other stated value or intermediate value within that range, is encompassed within the scope of the present disclosure. Where the range includes an upper or lower limit, ranges excluding either of those included limits are also included in the present disclosure.

[0051] It should be understood that one, some, or all of the features of the various embodiments described herein can be combined to form other embodiments of the present invention. The subsection headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0052] The features and preferences described above regarding "embodiments" are different preferences and are not limited to that particular embodiment; where technically feasible, they can be freely combined with features from other embodiments and can form preferred combinations of features. This description is presented to enable a person of ordinary skill in the art to make and use the present invention, and this description is provided in the context of a patent application and its claims. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles herein can be applied to other embodiments. Accordingly, the present invention is not intended to be limited to the embodiments shown, but rather to the broadest scope consistent with the principles and features described herein.

[0053] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference in their entirety, to the extent that each individual publication, patent, or patent application is specifically and individually indicated to be incorporated by reference in its entirety. In the event of a conflict between the terms herein and those of the incorporated references, the terms herein shall control.

[0054] Methods for Identifying Protein-Protein Complexes

[0055] Methods for identifying components of protein-protein complexes are described herein. Protein-protein complexes can be formed, stabilized, or dissociated in the presence of a compound (e.g., a small molecule or a drug), which can be part of a compound library or a natural extract, or a single drug tested alone. Samples such as biological samples (e.g., cell-free biological samples, tissue extracts, cell extracts, or subcellular extracts) can be mixed with a compound library, a drug (e.g., a drug under investigation), or a natural extract (such as a plant extract). Protein complexes can be identified by fractionating the sample (e.g., by size-exclusion chromatography) to obtain a plurality of fractions (i.e., a first plurality of fractions of a first sample and a second plurality of fractions of a second sample); analyzing the first plurality of fractions and the second plurality of fractions to identify the proteins in the fractions (e.g., using proteomic analysis); for a target protein, identifying a fraction shift between the first plurality of fractions and the second plurality of fractions; and identifying one or more binding proteins that form a complex with the target protein.

[0056] For a target protein, comparing protein migration (e.g., elution profiles) in the presence or absence of a compound library, a drug, or a natural extract can determine whether the complex is formed, stabilized, or dissociated by one or more compounds or a drug in the compound library or the natural extract. That is, a fraction shift caused by a compound library, a drug, or a natural extract can be identified by comparing the elution of the target protein in the presence and absence of the compound library, the drug, or the natural extract. The fraction shift can be used to determine whether a compound library, a drug, or a natural extract causes the formation of a stable molecular complex or disrupts the formation of a molecular complex. Binding proteins that form a complex with the target protein can also be identified by analyzing the fractions of other proteins that migrate similarly (e.g., co-elute or co-migrate) with the target protein in the presence or absence of a compound library or a natural extract. Co-elution of one or more binding proteins with the target protein in the second plurality of fractions rather than the first plurality of fractions indicates that a compound or a drug in the compound library or the natural extract causes the formation or stabilization of a complex comprising the target protein and at least one of the one or more binding proteins. Co-elution of one or more binding proteins with the target protein in the first plurality of fractions rather than the second plurality of fractions indicates that a compound or a drug in the compound library or the natural extract causes the dissociation of a complex comprising the target protein and at least one of the one or more binding proteins.

[0057] Exemplary methods for identifying protein-protein complexes are shown in Figure 1AFirst, a portion of a biological sample (such as a cell-free biological sample, tissue extract, cell extract, or subcellular extract) is analyzed to identify the baseline elution characteristics of proteins and protein complexes within the sample. A drug, compound library, or natural extract is added to another portion of the biological sample. The sample (i.e., containing the drug, compound library, or natural extract) is analyzed separately to identify the elution characteristics of proteins and protein complexes within the sample in the presence of the drug, compound library, or natural extract. Proteins eluting in different fractions between the two samples are determined to have a fraction shift. Depending on the direction of the fraction shift, the shift can indicate that the protein associates or dissociates with its binding partner in the presence of the drug, compound library, or natural extract. The method described uses the fraction shift to identify binding proteins (e.g., binding partners) of target proteins that associate or dissociate in the presence of the drug, compound library, or natural extract.

[0058] The samples (e.g., a first sample and a second sample) are fractionated separately, such as in Figure 1A 102 and 104. The first sample can be a portion of a biological sample (e.g., a cell-free biological sample, tissue extract, cell extract, or subcellular extract). The second sample is another portion of the biological sample combined with a drug, compound library, or natural extract (such as a plant extract). The combination preferably includes mixing or incubating the biological sample with the drug, compound library, or natural extract prior to fractionation. The incubation of the biological sample with the drug, compound library, or natural extract can be carried out at various temperatures and / or durations, but under conditions that preferably maintain the integrity of the sample. The incubation can also be carried out under physiological conditions. The incubation can be carried out while the sample is non-static (e.g., rotating or stirring) or static. Optionally, one or more protease inhibitors can be added to the sample during incubation. The sample integrity can be determined by assaying proteolysis, protein unfolding, and / or protein aggregation.

[0059] Biological samples can be obtained by lysing cells from tissues or subcellular structures isolated from cells. The cells, tissues, and / or subcellular structures can be lysed, for example, by sonication or detergents. The lysed material can be processed, for example, by centrifugation and / or filtration (e.g., to remove cell solids), by enzymes or chemicals (e.g., to remove nucleic acids), dialysis or buffer exchange, or other processing techniques. For subcellular lysates, standard separation techniques such as density gradient centrifugation can be used to separate the target subcellular structures (e.g., mitochondria, nuclei, and other organelles) from other cellular components. Before sample fractionation, the volume of the cell extract sample can be further reduced. Preferably, such extract processing does not denature the proteins in the lysate or induce their proteolysis or aggregation. As further discussed herein, biological samples are not limited to cell extracts, as in some specific embodiments, the biological sample can be a cell-free biological sample (e.g., saliva, cerebrospinal fluid, plasma, etc.) or an extract of a cell-free biological sample.

[0060] In some embodiments, biological samples containing proteins are obtained from animal tissues. In some embodiments, biological samples containing proteins are obtained from mammalian tissues. In some embodiments, the biological samples are obtained from brain, liver, lung, or kidney tissues.

[0061] A biological sample can be divided such that a first portion and a second portion can be used for a first sample and a second sample, respectively. The first sample contains a protein-containing portion of the biological sample. In some embodiments, the first sample contains the biological sample, which contains proteins obtained from tissues, cells, or subcellular organelles. In some embodiments, biological samples containing proteins (e.g., tissue, cell, or subcellular extracts) are obtained from animal tissues. In some embodiments, biological samples containing proteins are obtained from mammalian tissues. In some embodiments, biological samples containing proteins are obtained from brain, liver, lung, or kidney tissues. In some embodiments, methods for identifying components of protein-protein complexes include fractionating the first sample to produce a first plurality of fractions. In some embodiments, methods for identifying proteins include fractionating the first sample by size exclusion chromatography to produce a first plurality of fractions.

[0062] The second sample comprises a portion of the biological sample in combination with a compound library (e.g., a small molecule library), a drug, or a natural extract (e.g., a plant, animal, bacterial, or fungal extract). The compound library can comprise fully synthetic compounds, fully natural compounds, or a mixture of synthetic and natural compounds. The natural extract can be from an organism of a different taxonomic classification than the organism from which the biological sample is derived. For example, a tissue, cell, or subcellular extract can be from a different kingdom (e.g., the animal kingdom, the plant kingdom, the fungal kingdom, the protist kingdom, the archaebacterial kingdom, or the eubacterial kingdom), phylum, class, order, family, genus, or species than the source of the natural extract. The combined components of the second sample can be mixed and / or incubated to allow the components of the compound library, drug, or natural extract to interact or bind with the components of the biological sample. In some embodiments, the drug, compound library, or natural extract is substantially free of proteins. Preferably, the difference between the first sample and the second sample is the presence of the compound library, drug, or natural extract in the second sample.

[0063] Methods for identifying components of a protein-protein complex can include fractionating a first sample and a second sample to produce a first plurality of fractions (i.e., for the first sample) and a second plurality of fractions (i.e., for the second sample), as Figure 1A shown in 102 and 104 of. A fractionation method such as size exclusion chromatography or high performance liquid chromatography (HPLC) can be used to fractionate the samples to obtain a plurality of fractions. Fractionation is a separation process in which a sample is separated into a plurality of smaller amounts or fractions based on one or more physical characteristics of the sample components, such as hydrodynamic diameter (e.g., when separating components based on size exclusion chromatography) or molecular weight. Hydrodynamic diameter is a suitable approximation of molecular weight when used in accordance with the methods described herein. Fractions are collected based on one or more differences in the specific properties of the individual components. To ensure that protein complexes are not disrupted during fractionation, fractionation of the samples should be performed under non-denaturing conditions. For example, fractionation can be performed under physiological conditions, such as at a pH between about 6 and about 8. Fractionation can also be performed over a temperature range that may or may not be physiological. In some specific implementations of the methods described herein, fractionation is performed in phosphate buffered saline. In some embodiments, the method for identifying components of a protein-protein complex includes fractionating the sample using size exclusion chromatography.

[0064] The samples can be fractionated at a constant flow rate and / or a set volume. The final sample is diluted into several fractions obtained after fractionation. The fractionated samples do not contain the same proteins and / or compounds in all fractions because they will be separated based on one or more criteria selected by the fractionation technique. For example, size exclusion chromatography separates a mixture based on physical properties such as the size and shape (hydrodynamic size) of proteins or protein complexes. The fractions can also contain one or more compounds.

[0065] Collect the fractions and analyze them to identify the proteins in each fraction. As shown in 106 of Figure 1A , analyze the fractions from the first sample to identify the proteins in the multiple fractions of the first sample, and as shown in 108, analyze the fractions from the second sample to identify the proteins in the multiple fractions of the second sample. This identification can be performed, for example, using proteomic analysis. Exemplary proteomic analysis techniques can include using mass spectrometry (e.g., liquid chromatography tandem mass spectrometry (LC-MS / MS)). Before identification, the samples can be further processed. For example, the samples can also be separated by SDS-PAGE, and proteins of a specific size can be excised from the gel for further analysis. One or more proteases can also be used to intentionally digest the liquid- or gel-based samples to fragment the proteins into shorter polypeptides before protein identification. In some embodiments, the protease is an amino acid-specific protease. In some embodiments, the protease cleaves only at the N-terminus of the amino acid. In some embodiments, the protease cleaves only at the C-terminus of the amino acid. The fractions can also be further processed to prepare the samples, such as buffer exchange to remove salts and other buffer components that may interfere with the analysis. For example, to prepare samples for LC-MS / MS, the proteins in the fractions are precipitated from the solution using a kit (containing components such as buffers) or chemicals and then resuspended in a compatible buffer. The samples can be analyzed using protein identification techniques including Western blotting or LC-MS / MS. The collected fractions can be further separated by using liquid chromatography to separate the peptides, and then the eluate is directed to a mass spectrometer with an ionization source. Methods that can be used to identify the proteins in the fractions include proteomic methods such as immunoassays, mass spectrometry, and / or combinations thereof. In some embodiments, identifying the proteins in the first multiple fractions and the second multiple fractions includes using mass spectrometry. In some embodiments, the mass spectrometry is liquid chromatography mass spectrometry (LC-MS / MS). In some embodiments, the method used to identify the proteins in the first fraction and the second fraction is the same method.

[0066] Proteins of interest can be identified within a biological sample. The protein of interest can be, for example, a drug target or a potential drug target. The protein can be designated as the target protein. The target protein eluting in different fractions (e.g., with a fraction shift) in the presence of a drug, compound library, or natural extract indicates that it has undergone a change in the complex state (e.g., through complex formation or stabilization or dissociation), as evidenced by a change in hydrodynamic size. For example, a binding protein co-elutes with the target protein because the complex remains associated during fractionation. In Figure 1AAt 110, a fraction shift of the target protein between the first sample and the second sample can be identified. Identifying a fraction shift of the target protein indicates that the target protein interacts with other proteins (e.g., binding proteins) differently in the presence or absence of a drug, compound library, or natural extract, e.g., by complex formation or stabilization or complex dissociation.

[0067] Identifying the fraction shift can include converting fractionation information and protein identity data into a two-dimensional matrix, e.g., Figure 3 as shown. On one axis, fraction numbers corresponding to one sample (e.g., the first sample) are presented. On the other axis, fraction numbers corresponding to the other sample (e.g., the second sample) are presented. Data points representing the identified proteins are assigned coordinates on the graph based on the fraction in which they elute in either sample. While protein elution is a distribution and can cover a range of fractions, the peak elution fraction is assigned based on the maximum protein abundance (e.g., maximum intensity) observed between all fractions.

[0068] Evaluating the fraction shift of a protein such as the target protein can indicate whether a compound in the compound library or a natural extract or a drug causes complex formation or stabilization or dissociation. If a protein elutes in a different fraction in one sample than in the other, the data point representing that protein will be off the diagonal ( Figure 3 ). If the target protein elutes in a later fraction in the first sample (a biological sample without a drug, compound library, or natural extract) than in the second sample (a biological sample with a drug, compound library, or natural extract), it can be concluded that the drug, compound library, or natural extract causes the formation or stabilization of a complex containing the target protein. This is because the drug, compound library, or natural extract causes the target protein to associate with a substance of higher molecular weight. In contrast, if the target protein elutes in an earlier fraction in the first sample (a biological sample without a drug, compound library, or natural extract) than in the second sample (a biological sample with a drug, compound library, or natural extract), it can be concluded that the drug, compound library, or natural extract causes the dissociation of a complex containing the target protein. If a protein does not change its peak elution fraction between the two samples, the data point representing that protein will be located along the diagonal (e.g., assigned to the same fraction in both samples).

[0069] At Figure 1A 108 of the method shown, one or more binding proteins that form a complex with the target protein (in the first sample or the second sample) are identified. Exemplary methods for identifying one or more binding proteins are shown in further detail in Figure 1BIn 202, the assay fraction shift is evaluated to determine whether a drug, compound library, or natural extract causes complex formation or stabilization or complex dissociation. Co-elution of one or more additional proteins with the target protein in the first sample or the second sample (but not both) indicates that one or more compounds or drugs in the compound library or natural extract cause the formation or stabilization of a complex comprising the target protein and one or more additional proteins or complex dissociation. Thus, one or more of the proteins in the complex (other than the target protein) can be referred to as "binding proteins" because they are involved in the complex containing the target protein in the first sample or the second sample. Proteins that co-elute in the same fraction as the target protein in the presence of a drug, compound library, or natural extract (but do not co-elute in the absence of a drug, compound library, or natural extract) indicate that the additional one or more proteins are binding proteins that form a complex with the target protein in the presence of a drug, compound library, or natural extract. Proteins that co-elute in the same fraction as the target protein in the absence of a drug, compound library, or natural extract (but do not co-elute in the presence of a drug, compound library, or natural extract) indicate that the additional one or more proteins are binding proteins that form a complex with the target protein in the absence of a drug, compound library, or natural extract. If a drug, compound library, or natural extract causes complex formation or stabilization, proteins that co-elute with the target protein in the second plurality of fractions (i.e., the plurality of fractions of the second sample comprising the biological sample and the drug, compound library, or natural extract) can be identified as binding proteins of the target protein, as shown in 204. Optionally, equivalent fractions (e.g., the same fraction number) in the first plurality of fractions (i.e., the plurality of fractions of the first sample comprising the biological sample and not comprising a drug, compound library, or natural extract) can be analyzed, and proteins present in said fractions are excluded as binding proteins, as shown in 206. If a drug, compound library, or natural extract causes complex dissociation, proteins that co-elute with the target protein in the first plurality of fractions (i.e., the plurality of fractions of the first sample comprising the biological sample and not comprising a drug, compound library, or natural extract) can be identified as binding proteins of the target protein, as shown in 208.

[0070] Method for identifying compounds that modulate protein complexes

[0071] The present invention describes methods for identifying components of compounds that modulate protein complexes. Protein-protein complexes can be formed or stabilized or dissociated in the presence of one or more compounds, such as one or more compounds from a compound library or a natural extract. A sample, such as a biological sample (e.g., a cell-free biological sample, a tissue extract, a cell extract, or a subcellular extract), can be mixed with a compound library or a natural extract, such as a plant extract. The protein complex can be identified by fractionating the sample to obtain a plurality of fractions. Fractionating the sample can be accomplished by a variety of methods, such as size exclusion chromatography. Proteomics analysis can be applied to the fractions to identify proteins. Metabolomics analysis, such as LC-MS / MS or computational NMR, can also be used to identify one or more compounds in the fractions. A combination of proteomics and metabolomics analysis can be used to identify one or more compounds that cause complex formation or stabilization or dissociation.

[0072] For example, a compound can cause a target protein to associate with one or more binding partners. In some embodiments, when the sample is fractionated, the compound will remain bound to the newly formed protein complex. Thus, identification of one or more compounds that co-elute with the target protein and its binding partner indicates that the compound modulates protein complex formation. Identification of one or more compounds that cause a fraction shift can be accomplished by obtaining a metabolomics profile, such as by tandem mass spectrometry (LC-MS / MS) of the fraction and identifying compounds present in both the fraction and the compound library or natural extract. In some embodiments, the metabolomics profile of a fraction of a first sample (e.g., the first fraction) is compared to the metabolomics profile of a fraction from a second sample. For example, an increase in the amount of a compound in a second fraction (e.g., the second fraction) relative to the first sample indicates a new co-eluting compound. Co-elution can be determined based on the fraction in which the peak elution (e.g., maximum abundance, maximum intensity) of the protein and / or compound is detected.

[0073] In some specific embodiments, the compound can cause dissociation of the target protein from one or more of its binding partners. For example, when a sample is fractionated, the compound can remain bound to the target protein or another member of the complex (e.g., one or more binding partners). Thus, identification of one or more compounds that co-elute with the target protein or its binding partner indicates that the compound modulates protein complex dissociation. Identification of one or more compounds that cause a fraction shift can be accomplished by obtaining a metabolomics profile, e.g., by tandem mass spectrometry of the fractions (LC-MS / MS) and identifying the compounds present in both the fraction and a compound library or natural extract. Another method for determining whether a compound co-elutes in a new fraction is to compare the metabolomics profile of the fractions of a first sample (e.g., the first fraction) with the metabolomics profile of the fractions from a second sample. For example, enrichment of the compound in a second fraction (e.g., the second fraction) relative to the first sample indicates a new co-eluting compound. Co-elution can be defined by the fraction in which a peak elution (e.g., maximum abundance, maximum intensity) of the protein and / or compound is detected.

[0074] For a target protein, comparing protein migration (e.g., elution characteristics) in the presence or absence of a compound library or natural extract can determine whether a complex is formed or stabilized or dissociated by one or more compounds in the compound library or natural extract. That is, a fraction shift caused by a compound library or natural extract can be identified by comparing the target protein elution in the presence and absence of the compound library or natural extract. Fraction shifts can be used to determine whether a compound library or natural extract causes the formation of a stable molecular complex or disrupts the formation of a molecular complex. Binding proteins that form a complex with the target protein can also be identified by analyzing the fractions of other proteins that migrate similarly (e.g., co-elute or co-migrate) with the target protein in the presence or absence of a compound library or natural extract. Co-elution of one or more binding proteins with the target protein in a second plurality of fractions rather than a first plurality of fractions indicates that a compound in the compound library or natural extract or a drug causes the formation or stabilization of a complex comprising the target protein and at least one of the one or more binding proteins. Co-elution of one or more binding proteins with the target protein in a first plurality of fractions rather than a second plurality of fractions indicates that a compound in the compound library or natural extract or a drug causes the dissociation of a complex comprising the target protein and at least one of the one or more binding proteins.

[0075] Compounds responsible for complex formation or stabilization or complex dissociation are expected to co-elute with the target protein and / or one or more binding proteins. For example, if a compound is responsible for complex formation or stabilization, it is expected that the compound will bind to the target protein and / or binding protein and thus co-elute with the target protein and one or more binding proteins in multiple fractions of a sample containing a biological sample and a compound library or natural extract. If a compound is responsible for complex dissociation, it is expected that the compound will bind to the target protein or binding protein and thus co-elute with the target protein or binding protein in multiple fractions of a sample containing a biological sample and a compound library or natural extract.

[0076] Exemplary methods for identifying protein-protein complexes are shown in Figure 2A . First, a portion of a biological sample (such as a cell-free biological sample, tissue extract, cell extract, or subcellular extract) is analyzed to identify the baseline elution characteristics of proteins and protein complexes within the sample. A compound library or natural extract is added to another portion of the biological sample. This sample (i.e., containing the drug, compound library, or natural extract) is analyzed separately to identify the elution characteristics of proteins and protein complexes within the sample in the presence of the drug, compound library, or natural extract. Proteins eluting in different fractions between the two samples are determined to have a fraction shift. Depending on the direction of the fraction shift, the shift can indicate that the protein associates or dissociates with its binding partner in the presence of the compound library or natural extract.

[0077] The samples (e.g., the first sample and the second sample) are fractionated separately, such as in Figure 2A 302 and 304. The first sample can be a portion of a biological sample (e.g., a cell-free biological sample, tissue extract, cell extract, or subcellular extract). The second sample is another portion of the biological sample combined with a compound library or natural extract (such as a plant extract). The combination preferably includes mixing or incubating the biological sample with the compound library or natural extract prior to fractionation. The incubation of the biological sample and the compound library or natural extract can be carried out at various temperatures and / or durations, but under conditions that preferably maintain the integrity of the sample. The incubation can also be carried out under physiological conditions. The incubation can be carried out while the sample is non-static (e.g., rotating or stirring) or static. Optionally, one or more protease inhibitors can be added to the sample during incubation. The integrity of the sample can be determined by assaying proteolysis, protein unfolding, and / or protein aggregation.

[0078] Biological samples can be obtained by lysing cells from tissue or subcellular structures isolated from cells. Cells, tissues, and / or subcellular structures can be lysed, for example, by sonication or detergents. The lysed material can be processed, for example, by centrifugation and / or filtration (e.g., to remove cell solids), by enzymes or chemicals (e.g., to remove nucleic acids), dialysis or buffer exchange, or other processing techniques. For subcellular lysates, standard separation techniques such as density gradient centrifugation can be used to separate the target subcellular structures (e.g., mitochondria, nuclei, and other organelles) from other cellular components. The volume of the cell extract sample can be further reduced prior to sample fractionation. Preferably, such extract processing does not denature the proteins in the lysate or induce their proteolysis or aggregation. As further discussed herein, biological samples are not limited to cell extracts, as in some embodiments, the biological sample can be a cell-free biological sample (e.g., saliva, cerebrospinal fluid, plasma, etc.) or an extract of a cell-free biological sample.

[0079] In some embodiments, a biological sample containing proteins is obtained from animal tissue. In some embodiments, a biological sample containing proteins is obtained from mammalian tissue. In some embodiments, the biological sample is obtained from brain, liver, lung, or kidney tissue.

[0080] A biological sample can be divided such that a first portion and a second portion can be used for a first sample and a second sample, respectively. The first sample contains the protein-containing portion of the biological sample. In some embodiments, the first sample contains the biological sample that contains proteins obtained from tissue, cells, or subcellular organelles. In some embodiments, a biological sample containing proteins (e.g., a tissue, cell, or subcellular extract) is obtained from animal tissue. In some embodiments, a biological sample containing proteins is obtained from mammalian tissue. In some embodiments, a biological sample containing proteins is obtained from brain, liver, lung, or kidney tissue. In some embodiments, a method for identifying components of a protein-protein complex includes fractionating the first sample to produce a first plurality of fractions. In some embodiments, a method for identifying a protein includes fractionating the first sample by size exclusion chromatography to produce a first plurality of fractions.

[0081] The second sample comprises a portion of the biological sample in combination with a compound library (e.g., a small molecule library) or a natural extract (e.g., a plant, animal, bacterial, or fungal extract). The compound library can comprise fully synthetic compounds, fully natural compounds, or a mixture of synthetic and natural compounds. The natural extract can be from an organism of a different taxonomic classification than the organism from which the biological sample is produced. For example, a tissue, cell, or subcellular extract can be from a different kingdom (e.g., the animal kingdom, the plant kingdom, the fungal kingdom, the protist kingdom, the archaebacterial kingdom, or the eubacterial kingdom), phylum, class, order, family, genus, or species than the source of the natural extract. The combined components of the second sample can be mixed and / or incubated to allow the components of the compound library, drug, or natural extract to interact or bind with the components of the biological sample. In some embodiments, the compound library, drug, or natural extract is substantially free of protein. Preferably, the difference between the first sample and the second sample is the presence of the compound library or natural extract in the second sample.

[0082] Methods for identifying compounds that modulate complex formation, stabilization, or dissociation can include fractionating a first sample and a second sample to produce a first plurality of fractions (i.e., for the first sample) and a second plurality of fractions (i.e., for the second sample), as Figure 2A shown in 302 and 304. The samples can be fractionated using fractionation methods such as size exclusion chromatography or high performance liquid chromatography (HPLC) to obtain a plurality of fractions. Fractionation is a separation process in which a sample is divided into a plurality of smaller amounts or fractions based on one or more physical characteristics of the sample components, such as hydrodynamic diameter (e.g., when separating components based on size exclusion chromatography) or molecular weight. When used in accordance with the methods described herein, the hydrodynamic diameter is a suitable approximation of molecular weight. Fractions are collected based on one or more differences in the specific properties of the individual components. To ensure that protein complexes are not disrupted during fractionation, the fractionation of the samples should be performed under non-denaturing conditions. For example, fractionation can be performed under physiological conditions, such as at a pH between about 6 and about 8. Fractionation can also be performed over a temperature range that may or may not be physiological. In some specific implementations of the methods described herein, fractionation is performed in phosphate buffered saline. In some embodiments, methods for identifying components of a protein-protein complex include fractionating a sample using size exclusion chromatography.

[0083] The samples can be fractionated at a constant flow rate and / or a set volume. The final sample is diluted into a number of fractions obtained after fractionation. The fractionated samples do not contain the same proteins and / or compounds in all fractions because they will be separated based on one or more criteria selected by the fractionation technique. For example, size exclusion chromatography separates a mixture based on physical properties such as the size and shape (hydrodynamic size) of proteins or protein complexes. The fractions can also contain one or more compounds.

[0084] Collect the fractions and analyze them to identify the proteins in each fraction. As Figure 1A shown at 306, analyze the fractions from the first sample to identify the proteins in the multiple fractions of the first sample, and as shown at 308, analyze the fractions from the second sample to identify the proteins in the multiple fractions of the second sample. This identification can be performed, for example, using proteomic analysis. Exemplary proteomic analysis techniques can include using mass spectrometry (e.g., liquid chromatography tandem mass spectrometry (LC-MS / MS)). Before identification, the samples can be further processed. For example, the samples can also be separated by SDS-PAGE, and proteins of a specific size can be excised from the gel for further analysis. One or more proteases can also be used to deliberately digest the liquid- or gel-based samples to fragment the proteins into shorter polypeptides before protein identification. In some embodiments, the protease is an amino acid-specific protease. In some embodiments, the protease cleaves only at the N-terminus of the amino acid. In some embodiments, the protease cleaves only at the C-terminus of the amino acid. The fractions can also be further processed to prepare the samples, such as buffer exchange to remove salts and other buffer components that may interfere with the analysis. For example, to prepare a sample for LC-MS / MS, the proteins in the fraction are precipitated from the solution using a kit (containing components such as buffers) or chemicals and then resuspended in a compatible buffer. The samples can be analyzed using protein identification techniques, including Western blotting or LC-MS / MS. The collected fractions can be further separated by using liquid chromatography to separate the peptides, and then the eluate is directed to a mass spectrometer with an ionization source. Methods that can be used to identify the proteins in the fractions include proteomic methods, such as immunoassays, mass spectrometry, and / or combinations thereof. In some embodiments, identifying the proteins in the first multiple fractions and the second multiple fractions includes using mass spectrometry. In some embodiments, the mass spectrometry is liquid chromatography mass spectrometry (LC-MS / MS). In some embodiments, the method used to identify the proteins in the first fraction and the second fraction is the same method.

[0085] A protein of interest can be identified within a biological sample. The protein of interest can be, for example, a drug target or a potential drug target. This protein can be designated as the target protein. The target protein eluting in different fractions (e.g., with a fraction shift) in the presence of a compound library or natural extract indicates that it has undergone a change in the complex state (e.g., through complex formation or stabilization or dissociation), as evidenced by a change in hydrodynamic size. For example, a binding protein co-elutes with the target protein because the complex remains associated during fractionation. In Figure 2AAt 310, a fraction shift of the target protein between the first sample and the second sample can be identified. Identifying a fraction shift of the target protein indicates that the target protein interacts differently with other proteins (e.g., binding proteins), such as through complex formation or stabilization or complex dissociation, in the presence or absence of a drug, compound library, or natural extract.

[0086] Identifying the fraction shift can include converting fractionation information and protein identity data into a two-dimensional matrix, such as Figure 3 shown. On one axis, fraction numbers corresponding to one sample (e.g., the first sample) are presented. On the other axis, fraction numbers corresponding to the other sample (e.g., the second sample) are presented. Data points representing identified proteins are assigned coordinates on the graph based on the fraction in which they elute in either sample. While protein elution is a distribution and can cover a range of fractions, the peak elution fraction is assigned based on the maximum protein abundance (e.g., maximum intensity) observed between all fractions.

[0087] Evaluating the fraction shift of a protein, such as the target protein, can indicate whether a compound in a compound library or a natural extract or a drug causes complex formation or stabilization or dissociation. If a protein elutes in different fractions in one sample rather than the other, the data point representing that protein will be off the diagonal ( Figure 3 ). If the target protein elutes in a later fraction in the first sample (a biological sample without a compound library or natural extract) than in the second sample (a biological sample with a compound library or natural extract), it can be concluded that the compound library or natural extract causes the formation or stabilization of a complex containing the target protein. This is because the compound library or natural extract causes the target protein to associate with substances of higher molecular weight. In contrast, if the target protein elutes in an earlier fraction in the first sample (a biological sample without a compound library or natural extract) than in the second sample (a biological sample with a compound library or natural extract), it can be concluded that the compound library or natural extract causes the dissociation of a complex containing the target protein. If a protein does not change its peak elution fraction between the two samples, the data point representing that protein will be located along the diagonal (e.g., assigned to the same fraction in both samples).

[0088] At Figure 2A 308, one or more compounds that cause the fraction shift (i.e., cause complex formation or stabilization or complex dissociation) are identified. Identification of the one or more compounds can include metabolomic analysis of fractions containing the target protein and / or binding protein in a second plurality of fractions. The process for identifying the one or more compounds that cause the fraction shift can vary depending on whether the compound library or natural extract causes complex formation or stabilization or complex dissociation.

[0089] Exemplary processes for identifying one or more compounds are shown in Figure 2B . At Figure 2B 402, fraction shifts are evaluated to determine whether a compound library or natural extract causes complex formation or stabilization or complex dissociation. If the fraction shift indicates complex formation or stabilization, one or more compounds co-eluting with the target protein are identified at 404, for example by obtaining a metabolomic profile of the elution fraction containing the target protein. In some embodiments, co-elution is based on the fraction in which the peaks containing the target protein and the compound elute. For example, metabolomic and proteomic analyses of the target protein elution fraction and one or more adjacent fractions can be performed to determine the elution fractions of the compound and the target protein. Optionally, a metabolomic profile of the compound library or natural extract can be obtained (e.g., by performing metabolomic assays on the compound library or natural extract), as shown at 406. The metabolomic profile of the compound library or natural extract can be used to identify compounds co-eluting with the target protein. Optionally, a metabolomic profile of a biological sample can be obtained (e.g., by performing metabolomic assays on the biological sample). The metabolomic profile of the biological sample can be used to exclude compounds that cause complex formation or stabilization. One or more compounds causing the fraction shift can be confirmed by identifying one or more compounds present in both the fraction containing the target protein and the compound library or natural extract, as shown at 408.

[0090] If fraction shift indicates complex dissociation, one or more compounds that cause complex dissociation may bind to the target protein or one or more binding proteins (and co-elute with them). Thus, if fraction shift indicates complex dissociation, one or more binding proteins can be identified (i.e., one or more proteins that form a complex with the target protein in the absence of one or more compounds from a compound library or natural extract), as shown at 410. As described above, proteins that co-elute in the same fraction as the target protein in the absence of a compound library or natural extract (but do not co-elute in the presence of a compound library or natural extract) indicate that the additional one or more proteins are binding proteins that form a complex with the target protein in the presence of a compound library or natural extract. Thus, one or more binding proteins can be identified by identifying one or more proteins that co-elute with the target protein in fractions of the first plurality of fractions, for example, by performing proteomic analysis on one or more fractions of the first plurality of fractions (i.e., fractions associated with a sample containing a biological sample without a natural extract or compound library). Once one or more binding proteins are identified, the elution fraction (or fractions) of one or more binding proteins from the second plurality of fractions (i.e., fractions associated with a sample containing a biological sample and a compound library or natural extract) can be identified, as shown at 412, for example, by proteomic analysis of the fractions in the second plurality of fractions. Then one or more compounds that co-elute with one or more binding proteins or with the target protein can be identified at 414, for example, using metabolomic analysis. In some embodiments, co-elution is based on the elution fraction in which a peak containing the target protein or binding protein and the compound elutes. For example, metabolomic and proteomic analysis of the target protein elution fraction and / or the binding protein elution fraction and one or more adjacent fractions can be performed to determine the elution fraction of the compound and the target protein or binding protein. Optionally, a metabolomic profile of the compound library or natural extract can be obtained (e.g., by performing metabolomic assays on the compound library or natural extract), as shown at 416. The metabolomic profile of the compound library or natural extract can be used to identify compounds that co-elute with the target protein. Optionally, a metabolomic profile of the biological sample can be obtained (e.g., by performing metabolomic assays on the biological sample). The metabolomic profile of the biological sample can be used to rule out compounds that cause complex formation or stabilization. One or more compounds that cause fraction shift can be confirmed by identifying one or more compounds present in both the fraction containing the target protein (or binding protein) and the compound library or natural extract, as shown at 418.

[0091] sample

[0092] A sample as described herein (e.g., a first sample) can be a biological sample (e.g., tissue, cell or subcellular extract or cell-free biological sample). The biological sample can be divided into a first sample and a second sample. The first sample can contain the biological sample, but should not contain a drug, compound library or natural extract being studied according to the methods described herein. The second sample contains the same biological sample (i.e., the second part of the biological sample), and also contains at least one drug, compound library or natural extract. The second sample is combined with the drug, compound library or natural extract (e.g., by mixing). Preferably, the first sample and the second sample differ only in the presence or absence of the drug, compound library or natural extract. The sample can be proteins and polypeptides suspended in a buffer. The sample can have proteins, polypeptides and compounds.

[0093] In some embodiments, the sample contains proteins. In some embodiments, the sample contains compounds. In some embodiments, the sample contains both proteins and compounds. In some embodiments, the sample can contain extracts from different sources. In some embodiments, the natural extract is a plant extract.

[0094] In some embodiments, the sample can be fractionated by any of the methods disclosed herein. In some embodiments, the sample is fractionated by size exclusion chromatography. In some embodiments, the sample is fractionated to obtain multiple fractions. Fractionation is a separation process in which a mixture is divided into multiple smaller amounts or fractions, where the composition is different. The fractions are collected based on one or more differences in the specific properties of the individual components. In some embodiments, the sample can be fractionated to obtain multiple fractions. In some embodiments, the multiple fractions have the same, similar or equal volume. In some embodiments, the multiple fractions can be further analyzed by proteomics and / or metabolomics methods. In some embodiments, the individual fractions are collected and then analyzed. In some embodiments, the individual fractions are collected, combined and then analyzed. In some embodiments, the individual fractions are collected, further separated and then analyzed. In some embodiments, the individual fractions are collected, further separated and then analyzed using two different methods. Exemplary types of samples that can be used in the present invention are described below.

[0095] Biological sample

[0096] Biological samples can be derived from or obtained from biological materials isolated from any organism such as a human or a rodent. Biological samples can be from tissue extracts, cell extracts, or subcellular extracts (e.g., organelle extracts) or cell-free biological samples. Exemplary cell-free biological samples include but are not limited to plasma samples, cerebrospinal samples, saliva samples, milk samples, sputum samples, and fecal samples. Tissues can be isolated from an organism and mechanically digested, enzymatically digested, or both to release cells. The cells can then be applied to further lysis protocols to generate cell extracts. Prior to cell lysis, specific cell types can be isolated or sorted to obtain specific cell type extracts. Subcellular extracts can also be obtained by first gently lysing the cells, then applying a series of centrifugation or purification (such as tag purification) methods to isolate intact organelles of interest, and then completely lysing to obtain subcellular extracts.

[0097] Several methods are commonly used to extract proteins, including mechanical disruption, liquid homogenization, high-frequency sound waves (sonication), freeze / thaw cycles, and manual grinding. The choice of cell lysis method depends on the starting material, volume, and sensitivity of the protein being extracted.

[0098] Physical disruption is an effective method for lysing a variety of cells and has high lysis efficiency. Methods of physical extraction include but are not limited to any one of the following or combinations of the following: Dounce homogenizer, sonicator, blender, mortar and pestle, freezing with reagents such as dry ice with ethanol or liquid nitrogen, and French press. In physical disruption methods, the material is physically decomposed by shear forces or external forces to release cell components.

[0099] Detergents solubilize proteins and disrupt lipid-lipid, protein-protein, and protein-lipid interactions. It can be used to extract total proteins or subcellular fractions or organelles from various sample types. Detergent-based lysis is easily applicable to small volumes or larger samples and is a gentler alternative to physical disruption of cell membranes, although when preparing protein samples from tissues, it is often used in combination with homogenization and mechanical grinding to achieve complete cell disruption.

[0100] In some embodiments, the methods disclosed herein fractionate a first sample. In some embodiments, the first sample comprises a protein-containing portion of a cell extract. In some embodiments, the first sample is a cell extract comprising proteins obtained from a tissue, cell, or subcellular extract. In some embodiments, the cell extract containing proteins is obtained from animal tissue. In some embodiments, the cell extract containing proteins is obtained from mammalian tissue. In some embodiments, the extract containing proteins is obtained from brain, liver, lung, or kidney tissue.

[0101] Natural extract

[0102] An extract is a mixture of secondary metabolites. Different classes of compounds are present in plants and their extracts; however, most bioactive compounds come from four main classes: alkaloids, glycosides, polyphenols, and terpenes. Various traditional and modern methods are used to prepare plant extracts from different parts of plants, such as Soxhlet extraction, reflux extraction, sonication, decoction, maceration, pressurized liquid extraction, solid-phase extraction, microwave-assisted extraction, hydrodistillation, and enzyme-assisted extraction. Sample preparation first decomposes the matrix and then separates the target analyte.

[0103] In some embodiments, the methods disclosed herein fractionate a second sample. In some embodiments, the second sample comprises a natural extract. In some embodiments, the second sample comprises (i) a second portion of a tissue, cell, or subcellular extract and (ii) a natural extract. In some embodiments, the natural extract is a plant extract. In some embodiments, the natural extract is substantially free of proteins.

[0104] Compound library

[0105] A chemical library or compound library is a collection of stored chemicals that are typically ultimately used for high-throughput screening or industrial manufacturing. A chemical library can simply consist of a series of stored chemicals. Each chemical has associated information stored in a database, including information such as the chemical structure, purity, quantity, and biochemical characteristics of the compound. For example, in the drug discovery process, a variety of organic chemicals are needed for testing against disease models in high-throughput screening. A chemical library can contain fully synthetic compounds, fully natural compounds, or a mixture of synthetic and natural compounds. A compound library can contain one or more drugs and / or one or more natural extracts.

[0106] In some embodiments, the methods disclosed herein include fractionating a second sample that comprises a natural extract. In some embodiments, the second sample comprises (i) a second portion of a cell extract and (ii) a compound library. In some embodiments, the compound library is substantially free of proteins.

[0107] Drug

[0108] A drug can be a chemical substance or compound of completely natural, completely synthetic, or semi-synthetic origin. As described above, it can be isolated from natural sources (e.g., plant extracts). A drug can also be synthetic or partially synthetic.

[0109] In some embodiments, the methods disclosed herein include fractionating a second sample that comprises a drug. In some embodiments, the second sample comprises (i) a second portion of a cell extract and (ii) a drug. In some embodiments, the drug is substantially free of proteins.

[0110] Analysis method

[0111] The method of the present invention utilizes one or more analytical methods to separate and / or analyze complex mixtures of proteins, peptides, and compounds.

[0112] Liquid chromatography

[0113] One method of separating a sample into simplified or individual parts is liquid chromatography. Liquid chromatography includes a mobile phase and a stationary phase. A sample having proteins and / or separating the sample into its individual parts. This separation is based on the components of the sample in the presence of a mobile phase and a stationary phase. The mobile phase (liquid phase) can be a buffer, preferably a physiologically relevant buffer. Liquid chromatography uses a pump to cause a pressurized liquid and a sample mixture to flow through a column filled with an adsorbent (stationary phase), which separates the sample components based on the interaction of the sample components with the stationary phase. In some embodiments, the buffer can be pH 6-8. In some embodiments, the buffer contains phosphate buffered saline. The stationary phase contains a resin that forms a matrix. Molecules will enter the column and interact with the stationary phase in different ways based on their various properties. For example, in ion exchange chromatography, the stationary phase can be charged to separate proteins or polypeptides based on the charge or isoelectric point of the protein or polypeptide.

[0114] The stationary phase can also be made of different resins to form a matrix that retains proteins based on hydrodynamic size or molecular weight. Size exclusion chromatography (SEC) is a chromatographic method in which molecules in solution (such as proteins and polypeptides) are separated according to their size and, in some cases, according to their molecular weight. A size exclusion column can be selected based on the resolution of the apparent molecular weight within a given molecular weight range. In some embodiments, the sample is fractionated by size exclusion chromatography. When the protein exits the column, one or more detectors can be used to identify the presence of the protein relative to a separate buffer.

[0115] High performance liquid chromatography (HPLC), high pressure liquid chromatography is similar to SEC in that it uses a stationary phase and a mobile phase to separate, identify, and quantify the components in a mixture. This technique is applied to analytical chemistry and also relies on a pump to cause the sample to pass through the column. The components within the sample will interact differently with the stationary phase, resulting in their elution at different times. The mobile phase can be a mixture of solvents (such as water, acetonitrile, and / or methanol).

[0116] HPLC is different from size exclusion chromatography in that it operates at high pressure, with a shorter column length, and smaller resins to produce a high resolution separation of compounds. When the compound exits the column, one or more detectors can be used to identify the presence of the compound relative to a separate buffer. In some embodiments, a sample containing the compound is fractionated by HPLC.

[0117] Proteomics

[0118] Proteomics is the large-scale study of proteins. As applied in the present application, proteomics analysis includes identifying all proteins within a sample or each fraction, such as multiple fractions obtained from a fractionated sample.

[0119] Mass spectrometry (MS)-based proteomics (e.g., measuring and analyzing the mass and quantity of proteins in a sample or fraction) can be used to analyze known proteins or peptides to calculate possible fragmentation of molecules or compounds. Proteins and / or peptides in the samples of the present invention can be identified by comparing the retention time / index (IR), the mass-to-charge ratio (m / z) of the ions, and the MS fragmentation pattern with known proteins and / or peptides.

[0120] In some embodiments, a first plurality of fractions and a second plurality of fractions are analyzed to identify proteins in the first plurality of fractions and the second plurality of fractions. In some embodiments, the analysis includes proteomics analysis. In some embodiments, the analysis includes using mass spectrometry. In some embodiments, the analysis includes using liquid chromatography tandem mass spectrometry (LC-MS / MS).

[0121] Metabolomics

[0122] Metabolomics methods can be used according to the methods described, e.g., to identify drugs, compounds, or small molecules in a sample.

[0123] MS-based metabolomics

[0124] MS-based metabolomics (e.g., measuring and analyzing the mass, identity, and quantity of compounds in a sample or fraction) can be used to analyze the known chemical structure of a compound to calculate possible fragmentation of the compound. Compounds in the samples of the present invention can be identified by comparing the retention time / index (IR), the mass-to-charge ratio (m / z) of the ions, and the MS fragmentation pattern with known compounds (such as those present in a compound library or natural extract).

[0125] In some embodiments, identifying one or more compounds causing fraction shift includes obtaining a metabolomics profile of the fraction. In some embodiments, identifying one or more compounds causing fraction shift includes obtaining a metabolomics profile from a second plurality of fractions. In some embodiments, identifying one or more compounds causing fraction shift includes obtaining a metabolomics profile of a compound library or a natural extract. In some embodiments, identifying one or more compounds causing fraction shift includes identifying one or more compounds present in both the fraction and the compound library or natural extract. In some embodiments, identifying one or more compounds causing fraction shift includes: (i) obtaining a metabolomics profile of the fraction in the second plurality of fractions; (ii) obtaining a metabolomics profile of the compound library or natural extract; and (iii) identifying one or more compounds present in both the fraction and the compound library or natural extract. In some embodiments, identifying one or more compounds causing fraction shift further includes obtaining a metabolomics profile of a first sample. In some embodiments, identifying one or more compounds causing fraction shift further includes filtering the metabolomics profile of the fraction in the second plurality of fractions to exclude compounds present in the first sample. In some embodiments, identifying one or more compounds causing fraction shift further includes (i) obtaining a metabolomics profile of the first sample; and (ii) filtering the metabolomics profile of the fraction in the second plurality of fractions to exclude compounds present in the first sample. In some embodiments, the metabolomics profile is obtained using mass spectrometry. In some embodiments, the metabolomics profile is obtained using liquid chromatography and tandem mass spectrometry (LC-MS / MS). In some embodiments, identifying one or more compounds causing fraction shift includes confirming co-elution of one or more compounds and a target protein or one or more binding proteins based on the peak elution fraction of the one or more compounds and the peak elution fraction of the target protein or binding protein(s).

[0126] In some embodiments, identifying one or more compounds causing fraction shift includes determining a fragmentation profile of a natural extract. In some embodiments, identifying one or more compounds causing fraction shift includes determining a fragmentation profile of a compound library. In some embodiments, identifying one or more compounds causing fraction shift includes determining a fragmentation profile of a cell extract and a natural extract. In some embodiments, identifying one or more compounds causing fraction shift includes determining a fragmentation profile of a cell extract and a compound library. In some embodiments, identifying one or more compounds causing fraction shift includes determining a fragmentation profile of a natural extract and comparing it with the fragmentation profile of the cell extract and the natural extract. In some embodiments, identifying one or more compounds causing fraction shift includes determining a fragmentation profile of a compound library and comparing it with the fragmentation profile of the cell extract and the compound library.

[0127] Computational NMR

[0128] Computational (computer) nuclear magnetic resonance (NMR) can be used to analyze the known chemical structure of a compound or drug to calculate the likelihood of the compound binding to a protein and / or a peptide. This method predicts stable compound - protein complexes in a liquid solution by calculating NMR chemical shifts.

[0129] In some embodiments, identifying one or more compounds that cause a fraction shift includes obtaining a metabolomics profile of the fraction. In some embodiments, identifying one or more compounds that cause a fraction shift includes obtaining a metabolomics profile from a second plurality of fractions. In some embodiments, identifying one or more compounds that cause a fraction shift includes obtaining a metabolomics profile of a compound library or a natural extract. In some embodiments, identifying one or more compounds that cause a fraction shift includes identifying one or more compounds present in both the fraction and the compound library or the natural extract. In some embodiments, identifying one or more compounds that cause a fraction shift includes: (i) obtaining a metabolomics profile of the fraction in the second plurality of fractions; (ii) obtaining a metabolomics profile of the compound library or the natural extract; and (iii) identifying one or more compounds present in both the fraction and the compound library or the natural extract. In some embodiments, identifying one or more compounds that cause a fraction shift further includes obtaining a metabolomics profile of a first sample. In some embodiments, identifying one or more compounds that cause a fraction shift further includes filtering the metabolomics profile of the fraction in the second plurality of fractions to exclude compounds present in the first sample. In some embodiments, identifying one or more compounds that cause a fraction shift further includes (i) obtaining a metabolomics profile of the first sample; and (ii) filtering the metabolomics profile of the fraction in the second plurality of fractions to exclude compounds present in the first sample. In some embodiments, a metabolomics profile is obtained using computational NMR. In some embodiments, identifying one or more compounds that cause a fraction shift includes confirming the co - elution of one or more compounds and a target protein or one or more binding proteins based on the peak elution fractions of the one or more compounds and the peak elution fractions of the target protein or the one or more binding proteins. In some embodiments, confirming co - elution includes using computational NMR.

[0130] System

[0131] The present disclosure also describes a system for performing the methods described herein. The system can include one or more processors and a non-transitory computer-readable storage medium storing one or more programs, which when executed by the one or more processors cause the system to perform the methods. In some specific embodiments, the system can be configured to identify components of a protein-protein complex. In some specific embodiments, the system can be configured to identify one or more compounds that cause the formation, stabilization, or dissociation of a protein complex. The system can also include one or more analytical components for obtaining data used in the methods performed, such as a chromatography system configured to fractionate one or more samples (which can include, for example, a size-exclusion chromatography column), one or more mass spectrometers (which can be configured to obtain proteomics data and / or metabolomics data). A system including one or more mass spectrometers can also include a liquid chromatography system (which can include, for example, a reversed-phase liquid chromatography column). For example, the system can include a tandem mass spectrometer, which can be further equipped with a liquid chromatography system, for example to perform LC-MS / MS.

[0132] Exemplary systems configured to identify components of a protein-protein complex can include one or more processors; and a non-transitory computer-readable storage medium storing one or more programs that, when executed by the one or more processors, cause the system to: receive first proteomic profile data of a first plurality of fractions obtained by fractionating a first sample (e.g., using size-exclusion chromatography), the first sample comprising a protein-containing portion of a biological sample; receive second proteomic profile data of a second plurality of fractions obtained by fractionating a second sample (e.g., using size-exclusion chromatography), the second sample comprising (i) a second portion of the biological sample and (ii) a compound library, a drug, or a natural extract; identify proteins in the first plurality of fractions and the second plurality of fractions based on the first proteomic profile data and the second proteomic data; for a target protein, identify a fraction shift between the first plurality of fractions and the second plurality of fractions, wherein the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds or a drug in the compound library or natural extract; and identify one or more binding proteins that form a complex with the target protein. Co-elution of one or more binding proteins with the target protein in the second plurality of fractions but not in the first plurality of fractions indicates that a compound or a drug in the compound library or natural extract causes formation or stabilization of a complex comprising the target protein and at least one of the one or more binding proteins. Co-elution of one or more binding proteins with the target protein in the first plurality of fractions but not in the second plurality of fractions indicates that a compound or a drug in the compound library or natural extract causes dissociation of a complex comprising the target protein and at least one of the one or more binding proteins. Co-elution of one or more binding proteins and the target protein can be determined based on, for example, the peak elution fractions of the target protein and the one or more binding proteins. For example, co-elution of one or more binding proteins with the target protein can be based on the peak elution fractions of the one or more binding proteins and the peak elution fraction of the target protein.

[0133] When executed by the one or more processors, the one or more programs can further cause the system to select, as members of the complex, one or more of the one or more binding proteins based on the molecular weights of the one or more putative binding proteins and the target protein and the fraction numbers of the fractions comprising the target protein and the one or more putative binding proteins.

[0134] The system can further include an analytical system for obtaining the first proteomic profile data and / or the second proteomic profile data. For example, the system can include one or more mass spectrometers. In some embodiments, the system can include a liquid chromatography system and a tandem mass spectrometer, which can be configured to generate proteomic profile data using LC-MS / MS.

[0135] Exemplary systems configured to identify one or more compounds that cause protein complex formation, stabilization, or dissociation can include one or more processors; and a non-transitory computer-readable storage medium storing one or more programs that, when executed by the one or more processors, cause the system to: receive first proteomic profile data for a first plurality of fractions obtained by fractionating a first sample (e.g., using size-exclusion chromatography), the first sample comprising a protein-containing portion of a biological sample; receive second proteomic profile data for a second plurality of fractions obtained by fractionating a second sample (e.g., using size-exclusion chromatography), the second sample comprising (i) a second portion of the biological sample and (ii) a compound library or natural extract; identify proteins in the first plurality of fractions and the second plurality of fractions based on the first proteomic profile data and the second proteomic data; for a target protein, identify a fraction shift between the first plurality of fractions and the second plurality of fractions, where the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds in the compound library or natural extract; receive metabolomic data for the compound library or natural extract; receive metabolomic data for a fraction of the second plurality of fractions that contains the target protein or a binding protein that forms a complex with the target protein in the absence of one or more compounds; and identify one or more compounds that cause the fraction shift by analyzing the metabolomic data for the compound library or natural extract and the metabolomic data for the fraction of the second plurality of fractions that contains the target protein or a binding protein that forms a complex with the target protein in the absence of one or more compounds.

[0136] In some embodiments, the system is configured to identify one or more compounds that cause the fraction shift by receiving metabolomic profiles for fractions of the second plurality of fractions; receiving metabolomic profiles for the compound library or natural extract; and identifying one or more compounds present in both the fractions and the compound library or natural extract. The system can be further configured to receive metabolomic profiles for the first sample; and filter the metabolomic profiles for fractions of the second plurality of fractions to exclude compounds present in the first sample.

[0137] The system can also include an analytical system for obtaining the first proteomic profile data and / or the second proteomic data profiles. For example, the system can include one or more mass spectrometers. In some embodiments, the system can include a liquid chromatography system and a tandem mass spectrometer, which can be configured to generate proteomic profile data using LC-MS / MS.

[0138] The system may also include an analysis system for obtaining metabolomics profile data and / or second proteomics data profiles. For example, the system may include one or more mass spectrometers. In some specific embodiments, the system may include a liquid chromatography system and a tandem mass spectrometer, which may be configured to generate proteomics profile data using LC-MS / MS. In some embodiments, the system may include a nuclear magnetic resonance (NMR) system. For example, the system may be configured to obtain metabolomics profile data using computer nuclear magnetic resonance (NMR).

[0139] In some specific embodiments, identifying one or more compounds that cause fraction shifts may include, for example, confirming the co-elution of one or more compounds and a target protein or one or more binding proteins based on the peak elution fractions of the one or more compounds and the peak elution fractions of the target protein or binding protein(s).

[0140] In some specific embodiments, the system is configured to identify binding proteins that form complexes with a target protein. For example, the co-elution of a binding protein and a target protein in a first plurality of fractions but not in a second plurality of fractions indicates that one or more compounds in a compound library or natural extract cause the dissociation of a complex comprising the target protein and the binding protein.

[0141] Figure 8 An example of a computing device or system according to one embodiment is shown. Device 800 may be a host computer connected to a network. Device 800 may be a client computer or a server. As Figure 8 shown, device 800 may be any suitable type of microprocessor-based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device), such as a phone or a tablet. The device may include, for example, one or more processors 810, an input device 820, an output device 830, a memory or storage device 840, a communication device 860, and one or more analysis systems 870 (e.g., one or more liquid chromatography systems and / or one or more mass spectrometers). Software 850 residing in memory or storage device 840 may include, for example, an operating system and software for performing the methods described herein. Input device 820 and output device 830 may generally correspond to those described herein and may be connected to or integrated with the computer.

[0142] Input device 820 may be any suitable device that provides input, such as a touch screen, a keyboard or keypad, a mouse, or a voice recognition device. Output device 830 may be any suitable device that provides output, such as a touch screen, a tactile device, or a speaker.

[0143] The storage device 840 can be any suitable device that provides storage (e.g., electrical, magnetic, or optical memory, including RAM (volatile and non-volatile), cache, hard disk drive, or removable storage disk). The communication device 860 can include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a wired medium (e.g., physical system bus 880, Ethernet connection, or any other wired transmission technology) or wirelessly (e.g., or any other wireless technology).

[0144] The software module 850, which can be stored as executable instructions in the storage device 840 and executed by the processor 810, can include, for example, an operating system and / or processes embodying the functionality of the methods of the present disclosure (e.g., embodied in the devices as described herein).

[0145] The software module 850 can also be stored and / or transmitted within any non-transitory computer-readable storage medium for use by or in connection with an instruction execution system, apparatus, or device (such as those described herein), from which instructions associated with the software can be obtained and executed. In the context of the present disclosure, a computer-readable storage medium can be any medium, such as the storage device 840, that can contain or store a process for use by or in connection with an instruction execution system, apparatus, or device. Examples of computer-readable storage media can include memory units, such as hard disk drives, flash drives, and distributed modules operating as a single functional unit. Additionally, the various processes described herein can be embodied as modules configured to operate in accordance with the above-described embodiments and techniques. Further, although the processes can be shown and / or described separately, those skilled in the art will understand that the above processes can be routines or modules within other processes.

[0146] The software module 850 can also be propagated within any transmission medium for use by or in connection with an instruction execution system, apparatus, or device (such as those described above), from which instructions associated with the software can be obtained and executed. In the context of the present disclosure, a transmission medium can be any medium that can communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. Transmission-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.

[0147] The device 800 can be connected to a network (e.g., network 904, such as Figure 9As shown and / or described below, the network can be any suitable type of interconnected communication system. The network can implement any suitable communication protocol and can be secured by any suitable security protocol. The network can include network links arranged in any suitable manner for transmitting and receiving network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.

[0148] Device 800 can be implemented using any operating system, such as an operating system suitable for operating on a network. Software module 850 can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, the application software embodying the functionality of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or as a web-based application or web service via a web browser. In some embodiments, the operating system is executed by one or more processors (e.g., processor 810).

[0149] Device 800 may also include one or more analysis systems 870 (e.g., one or more liquid chromatography systems and / or one or more mass spectrometers).

[0150] Figure 9 An example of a computing system according to one embodiment is shown. In system 900, device 800 (e.g., as described above and shown in Figure 8 is connected to network 904, which is also connected to device 906. In some embodiments, device 906 is an analysis system (which may include, for example, one or more liquid chromatography systems and / or one or more mass spectrometers).

[0151] Devices 800 and 906 may communicate, for example, via a suitable communication interface over a network 904 such as a local area network (LAN), a virtual private network (VPN), or the Internet. In some embodiments, network 904 may be, for example, the Internet, an intranet, a virtual private network, a cloud network, a wired network, or a wireless network. Devices 800 and 906 may communicate partially or entirely via wireless or hardwired communication such as Ethernet, IEEE 802.11b wireless, etc. Additionally, devices 800 and 906 may communicate, for example, via a suitable communication interface over a second network such as a mobile / cellular network. Communication between devices 800 and 906 may also include or communicate with various servers such as mail servers, mobile servers, media servers, telephone servers, etc. In some embodiments, devices 800 and 906 may communicate directly (as an alternative to or in addition to communicating over network 904), for example, via wireless or hardwired communication such as Ethernet, IEEE 802.11b wireless, etc. In some embodiments, devices 800 and 906 communicate via communication 908, which may be a direct connection or may be via a network (e.g., network 904).

[0152] One or both of devices 800 and 906 generally include logic (e.g., http web server logic) or are programmed to format data accessed from local or remote databases or other data and content sources to provide and / or receive information over network 904 according to the various examples described herein.

[0153] Example

[0154] Example 1: High-throughput mapping of metabolite-host protein interactions for scalable drug discovery.

[0155] Cell lysates were mixed with natural extracts and analyzed. Several proteins were identified as co-eluting with RNF114 in the presence of the compound. This example demonstrates an exemplary high-throughput method for identifying compound-protein interactions.

[0156] Lysate and extract preparation: Frozen mouse tissues (pooled from brain, liver, lung, and kidney) were homogenized (30 s x 5 cycles) in a 2 mL screw-cap tube with 1X PBS and zirconia beads in a bead mill (speed 3450 rpm). After homogenization, the lysate was centrifuged at 21,000 x g for 15 minutes to obtain a clarified supernatant. Protein concentration was determined using a BCA protein assay kit. Plant extracts were separately dissolved in DMSO and then combined with mouse tissue lysates (total 60 mg protein).

[0157] Sample separation: The binding reaction / cotemperature incubation was incubated at 25 °C for 1 hour, and then size exclusion chromatography was performed. Size exclusion chromatography was performed on an AKTA AVANT 150 (GE), Cytiva (UNICORN TM software version 7.6) using a SUPERDEX 200pg 16 / 600 column (GE). The running buffer for separation was 50 mM ammonium bicarbonate + 150 mM NaCl in milliQ water. Before sample analysis, the column was equilibrated with 2 column volumes of the running buffer at a flow rate of 1 mL / min. After loading the sample onto the column, 80 fractions were collected at a flow rate of 0.8 mL / min with the running buffer. The amount of protein in each fraction was quantified using the BCA assay.

[0158] Sample preparation for metabolomics: For each fraction, a volume corresponding to 20 μg of total protein was sampled from the fraction for metabolomics (800 μL in total). The sample was dried using a speedvac and then methanol was added. The sample was sonicated for 10 minutes, then vortexed for 5 minutes, and centrifuged at room temperature for 10 minutes. The aliquot was dried using a concentrator.

[0159] Sample preparation for proteomics: Fractions equivalent to 20 μg of protein were taken for proteomics and mixed with 20% SDS. The protein was reduced with 5 mM tris(2-carboxyethyl)phosphine hydrochloride at 37 °C for 1 hour and alkylated with methyl methanethiosulfonate (MMTS) at room temperature in the dark for 30 minutes. At 37 °C, the protein was further digested overnight with trypsin / lysC (1:100) on a thermomixer. After digestion, the sample was removed from the thermomixer and 0.5 pmol of PREMIS TM (Promega) peptide mixture was added to the sample. Then the peptides were loaded onto an S-TRAP column and centrifuged at 10,000 xg for 30 s. The peptides were eluted with triethylammonium bicarbonate buffer (TEABC), 0.1% formic acid, and 50% acetonitrile (ACN).

[0160] The peptides were resuspended in 0.1% formic acid and analyzed by liquid chromatography tandem mass spectrometry (LC-MS / MS). Peptides were resolved on an Ultimate 3000 RSLCnano system coupled to an Orbitrap Eclipse. 1 μg was loaded onto a C18 column 50 cm, 3.0 μm Easy-Spray column (Thermo Fisher Scientific). Peptides were eluted at a flow rate of 300 nl / min with a 0-40% gradient of buffer B (80% acetonitrile, 0.1% formic acid) and injected for MS analysis. The LC gradient was run for 100 minutes. MS1 spectra were acquired in the Orbitrap (R = 240k; AGC target = 400000; max IT = 50 ms; RF lens = 30%; mass range = 400 - 2000; centroid data). Dynamic exclusion was applied for 10 seconds to exclude all charge states of a given precursor. MS2 spectra were collected in the linear ion trap (rate = turbo; AGC target = 20,000; max IT = 50 ms; NCE HCD = 35%).

[0161] Data analysis: The mass spectrometry.raw files were used for proteomics database searching. Proteomics database searches were performed using FragPipe software [version 17.1; Releases Nesvilab / FragPipe (github.com)] against human and mouse proteome databases (UniProtKB Release 2021_03) downloaded from UniProt, including contaminants and decoy proteins, for raw data searching. Searches were performed with precursor and fragment tolerances of 10 ppm and 0.05 Da, respectively. A false discovery rate of 0.1% was maintained at both the peptide and protein levels. The combined peptide output file was used for dose-response curve (DRC) peptide trend analysis.

[0162] Proteomics and metabolomics LC-MS / MS data were further filtered according to the described cleaning strategy. Contaminants and human proteins such as keratin were removed from the proteomics data, and then proteins that appeared in the 75% fraction were removed. We retained proteins with at least ≥2 unique peptides and a total of ≥2 identified peptide spectrum matches (PSMs) for downstream correlation analysis. The metabolomics dataset was preprocessed by removing metabolites (features) that appeared in all 75% fractions, including metabolites (features) with abundance values >10000, and selecting metabolites / features with relevant MS / MS spectra. The false discovery rate (FDR) was calculated to be 5% based on the Pearson R2 value.

[0163] Results: Compound-protein interactions were plotted to visualize peak shifts in the presence of a compound library ( Figure 5) Proteins that do not change their peak elution fractions in the presence of metabolites are visualized on the diagonal. Proteins that move to the right of the diagonal indicate that the protein elutes in a lower-numbered fraction and increases in apparent molecular weight in the presence of the native extract. Proteins that move to the left of the diagonal indicate that the protein elutes in a higher-numbered fraction and loses apparent molecular weight in the presence of the native extract.

[0164] The exemplary protein RNF114 elutes in fraction 37 of the mouse lysate in the absence of the metabolite library ( Figure 5 ). When the lysate is incubated with the P45 plant extract, RNF114 is observed to elute in fraction 16 ( Figure 5 ). This indicates that RNF114 can both form complexes of larger apparent molecular weight and dissociate into complexes of smaller apparent molecular weight in the presence of P45 ( Figure 5 ).

[0165] Other proteins were observed to elute in higher-numbered fractions in the control lysate but co-elute with RNF114 in lower-numbered fractions in the presence of the metabolite ( Figure 6 ). The shift in the peak elution volume indicates that both RNF114 and the binding protein increase in apparent molecular weight in the presence of the compound. This indicates that the compound can mediate the formation of protein complexes between RNF114 and the binding protein.

[0166] Results: To verify compound-dependent complex formation, mouse protein lysates were incubated with or without P45. In the absence of the compound, RNF114 and the binding protein (a K-Ras signaling regulator) migrated in separate fractions ( Figure 7 ). In the presence of the compound, the peaks corresponding to the unbound proteins became undetectable in fractions 31 and 37, respectively, but appeared in fraction 16. This indicates that the P45 native extract mediates protein-protein association between RNF114 and the binding protein.

Claims

1. A method for identifying components of a protein-protein complex, the method comprising: fractionating a first sample using size exclusion chromatography to generate a first plurality of fractions, the first sample comprising a first protein-containing portion of a biological sample; fractionating a second sample using the size exclusion chromatography to generate a second plurality of fractions, the second sample comprising (i) a second portion of the biological sample and (ii) a compound library, a drug, or a natural extract; analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions; for a target protein, identifying a fraction shift between the first plurality of fractions and the second plurality of fractions, wherein the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds in the compound library, the drug, or the natural extract; and identifying one or more binding proteins that form a complex with the target protein, wherein: co-elution of the one or more binding proteins with the target protein in the second plurality of fractions but not in the first plurality of fractions indicates formation or stabilization of a complex comprising the target protein and at least one of the one or more binding proteins caused by a compound in the compound library, the drug, or the natural extract; and co-elution of the one or more binding proteins with the target protein in the first plurality of fractions but not in the second plurality of fractions indicates dissociation of a complex comprising the target protein and at least one of the one or more binding proteins caused by a compound in the compound library, the drug, or the natural extract.

2. The method of claim 1, the method further comprising selecting, as members of the complex, one or more of the one or more binding proteins based on the molecular weights of the one or more putative binding proteins and the target protein and the fraction numbers of the fractions comprising the target protein and the one or more putative binding proteins.

3. The method of claim 1 or 2, wherein co-elution is determined based on the peak elution fractions of the target protein and the one or more binding proteins.

4. The method of any one of claims 1 to 3, wherein co-elution of the one or more binding proteins with the target protein is based on the peak elution fractions of the one or more binding proteins and the peak elution fraction of the target protein.

5. A method for identifying one or more compounds that cause protein complex formation, stabilization, or dissociation, the method comprising: fractionating a first sample using size exclusion chromatography to generate a first plurality of fractions, the first sample comprising a first protein-containing portion of a biological sample; fractionating a second sample using the size exclusion chromatography to generate a second plurality of fractions, the second sample comprising (i) a second portion of the biological sample and (ii) a compound library or a natural extract; analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions; For a target protein, identify a fraction shift between the first plurality of fractions and the second plurality of fractions, wherein the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds in the compound library or the natural extract; and Identify the one or more compounds that cause the fraction shift, including analyzing the following substances: (1) for a fraction shift indicating complex formation or stabilization caused by the one or more compounds, analyze the fraction in the second plurality of fractions that contains the target protein to identify one or more compounds that co-elute with the target protein in the fraction, or (2) for a fraction shift indicating complex dissociation caused by the one or more compounds, analyze (i) the fraction in the second plurality of fractions that contains the target protein to identify one or more compounds that co-elute with the target protein in the fraction, or (ii) the fraction in the second plurality of fractions that contains a binding protein that forms a complex with the target protein in the absence of the one or more compounds to identify one or more compounds that co-elute with the binding protein in the fraction.

6. The method according to claim 5, wherein identifying the one or more compounds that cause the fraction shift includes: Obtaining a metabolomics profile of the fraction in the second plurality of fractions; Obtaining a metabolomics profile of the compound library or the natural extract; and Identifying one or more compounds present in both the fraction and the compound library or the natural extract.

7. The method according to claim 6, wherein identifying the one or more compounds that cause the fraction shift further includes: Obtaining a metabolomics profile of the first sample; and Filtering the metabolomics profile of the fraction in the second plurality of fractions to exclude compounds present in the first sample.

8. The method according to claim 6 or 7, wherein the metabolomics profile is obtained using mass spectrometry.

9. The method according to claim 8, wherein the metabolomics profile is obtained using liquid chromatography and tandem mass spectrometry (LC-MS / MS).

10. The method according to claim 6 or 7, wherein the metabolomics profile is obtained using computer nuclear magnetic resonance (NMR).

11. The method according to any one of claims 5 to 10, wherein identifying the one or more compounds that cause the fraction shift includes confirming the co-elution of the one or more compounds and the target protein or the one or more binding proteins based on the peak elution fractions of the one or more compounds and the peak elution fractions of the target protein or the binding protein.

12. The method according to any one of claims 5 to 11, the method includes identifying the binding protein that forms the complex with the target protein, wherein the co-elution of the binding protein and the target protein in the first plurality of fractions rather than in the second plurality of fractions indicates that one or more compounds in the compound library or the natural extract cause the dissociation of the complex containing the target protein and the binding protein.

13. The method according to any one of claims 1 to 12, wherein analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions comprises proteomic analysis.

14. The method according to claim 13, wherein analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions comprises using mass spectrometry.

15. The method according to claim 13, wherein analyzing the first plurality of fractions and the second plurality of fractions to identify proteins in the first plurality of fractions and the second plurality of fractions comprises using liquid chromatography tandem mass spectrometry (LC-MS / MS).

16. The method according to any one of claims 1 to 15, wherein the compound library, the drug or the natural extract is substantially free of proteins.

17. The method according to any one of claims 1 to 16, wherein the biological sample comprises a cell-free biological sample, a tissue extract, a cell extract or a subcellular extract.

18. The method according to any one of claims 1 to 17, wherein the second sample comprises the compound library.

19. The method according to any one of claims 1 to 17, wherein the second sample comprises the natural extract.

20. The method according to claim 19, wherein the natural extract is a plant extract.

21. The method according to any one of claims 1 to 20, wherein the biological sample containing proteins is obtained from a cell lysate.

22. The method according to any one of claims 1 to 21, wherein the biological sample containing proteins is obtained from an animal tissue.

23. The method according to any one of claims 1 to 22, wherein the biological sample containing proteins is obtained from a mammalian tissue.

24. The method according to any one of claims 1 to 23, wherein the biological sample containing proteins is obtained from brain, liver, lung or kidney tissue.

25. A system, the system comprising: one or more processors; and a non-transitory computer-readable storage medium storing one or more programs, the one or more programs, when executed by the one or more processors, cause the system to: receive first proteomic spectral data of a first plurality of fractions obtained by fractionating a first sample using size exclusion chromatography, the first sample comprising a protein-containing portion of a biological sample; receive second proteomic spectral data of a second plurality of fractions obtained by fractionating a second sample using the size exclusion chromatography, the second sample comprising (i) a second portion of the biological sample and (ii) a compound library, a drug or a natural extract; identify proteins in the first plurality of fractions and the second plurality of fractions based on the first proteomic spectral data and the second proteomic data; For a target protein, identify a fraction shift between the first plurality of fractions and the second plurality of fractions, wherein the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds in the compound library, the drug, or the natural extract; and Identify one or more binding proteins that form a complex with the target protein, wherein: Co-elution of the one or more binding proteins with the target protein in the second plurality of fractions but not in the first plurality of fractions indicates formation or stabilization of a complex comprising the target protein and at least one of the one or more binding proteins caused by a compound in the compound library, the drug, or the natural extract; and Co-elution of the one or more binding proteins with the target protein in the first plurality of fractions but not in the second plurality of fractions indicates dissociation of a complex comprising the target protein and at least one of the one or more binding proteins caused by a compound in the compound library, the drug, or the natural extract.

26. A system, the system comprising: One or more processors; and A non-transitory computer-readable storage medium storing one or more programs, the one or more programs, when executed by the one or more processors, cause the system to: Receive first proteomic spectral data of a first plurality of fractions obtained by fractionating a first sample using size-exclusion chromatography, the first sample comprising a protein-containing portion of a biological sample; Receive second proteomic spectral data of a second plurality of fractions obtained by fractionating a second sample using the size-exclusion chromatography, the second sample comprising (i) a second portion of the biological sample and (ii) a compound library or a natural extract; and Identify proteins in the first plurality of fractions and the second plurality of fractions based on the first proteomic spectral data and the second proteomic data; For a target protein, identify a fraction shift between the first plurality of fractions and the second plurality of fractions, wherein the fraction shift indicates complex formation or stabilization or complex dissociation caused by one or more compounds in the compound library or the natural extract; Receive metabolomic data of the compound library or the natural extract; Receive metabolomic data of a fraction in the second plurality of fractions that contains the target protein or a binding protein that forms a complex with the target protein in the absence of the one or more compounds; and Identify the one or more compounds causing the fraction shift by analyzing the metabolomic data of the compound library or the natural extract and the metabolomic data of the fraction in the second plurality of fractions that contains the target protein or the binding protein that forms a complex with the target protein in the absence of the one or more compounds.

Citation Information

Patent Citations

  • Full-nutrient concocted rice suitable for plateau self-heating rice

    CN116076664A