Method for assaying protein
Patent Information
- Application Number
- JP2024150202
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-05-02
- Filing Date
- 2024-08-30
- Publication Date
- 2026-08-26
AI Technical Summary
Current protein identification techniques relying on mass spectrometry are inefficient and require extensive sample preparation, limiting their speed and specificity, especially for identifying a wide range of proteins.
A method involving affinity reagents that bind to unique spatial addresses on a substrate, allowing for rapid identification and quantification of proteins by observing binding patterns and using deconvolution methods to determine protein identity, even in complex mixtures.
Enables rapid and accurate identification of multiple proteins with high accuracy, potentially identifying over 90% of proteins in a sample with confidence, reducing the need for extensive sample preparation and improving throughput.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] cross reference This application is a continuation of U.S. Provisional Patent Application No. 62 / 429,063, filed December 1, 2016, and pending until May 2, 2017. This application claims the benefit of U.S. Provisional Patent Application No. 62 / 500,455, filed on Oct. 13, 2003, the entire disclosure of which is hereby incorporated by reference. No. 6,399,433, filed on Oct. 23, 2003, and is here [Background technology]
[0002] 2. Background of the Invention Current techniques for protein identification typically involve highly specific and sensitive antibodies. Binding of peptides and subsequent readout of information, or peptide readout data from a mass spectrometer The nucleotide sequence depends on one of the following: Summary of the Invention
[0003] The present disclosure provides methods and systems for assaying proteins. In some embodiments, the disclosure may be very incomplete and / or may not be specific to a particular protein. From a series of measurements that are not specific for a protein, the identity of the protein in the mixture, i.e. The methods and methods described herein provide an approach to predicting the sequences of these genes. The method and system also include the steps of: characterizing and / or identifying biopolymers, including proteins; In addition, the methods and systems described herein may be used to This method is more rapid than techniques for protein identification that rely on data from quantitative spectrometers. In some instances, the methods described herein can be used to identify and the system identifies at least 400 different proteins with at least 50% accuracy. At least 10% higher than techniques for protein identification that rely on data from mass spectrometry In some instances, the methods described herein may be used to more quickly identify The methods and systems described herein include detecting at least 1000 different proteins, % accuracy compared to techniques for protein identification that rely on data from mass spectrometry. It can be used to identify at least 10% more quickly.
[0004] One aspect of the invention provides a method for determining a characteristic of a protein, the method comprising: obtaining a unique, resolvable spatial address for each individual protein moiety; A portion of one or more proteins is conjugated to a substrate so as to have a In some cases, each individual protein moiety is optically resolvable. The method may include first determining whether or not a first affinity reagent is capable of binding to a target molecule. In some embodiments, the method further comprises applying a fluid comprising set n to the substrate. The affinity reagent may include or be linked to an identifiable tag. The method includes, after each application of one or more affinity reagents to the substrate of the first to nth sets, performing the steps of: observing the affinity reagent or identifiable tag. one or more unique spatial addresses of the substrate having one or more observation signals; and identifying one or more species having the identified unique spatial address. Each portion of the protein is associated with one or more observed signals. In some instances, determining that the antibody contains one or more epitopes. Each of the conjugated moieties of several proteins has a unique spatial address on the substrate. In some examples, the first to third affinity reagents are associated with the Each affinity reagent of the nth set is specific for an individual protein. In some cases, affinity assays are not specific for a particular protein family. The binding epitopes of drugs are not known and are not specific to individual proteins. Nor are they specific for any particular protein family.
[0005] In some cases, the methods of the present disclosure also include methods for detecting multiple proteins bound to a single site. may be used with a substrate that has been modified to detect at least about 50% of the protein at a single location. , 60%, 70%, 80%, 90%, or more than 90% contain a common amino acid sequence. In the disclosed method, a substrate having multiple proteins bound to a single site is also used. and at least about 50%, 60%, 70%, 80%, 90%, %, or greater than 90% includes at least 95% amino acid sequence identity.
[0006] In some embodiments, the one or more proteins comprise a single protein molecule. In some embodiments, the one or more proteins may comprise bulk proteins. In some embodiments, the one or more proteins may comprise a homogenate on the substrate. It may contain multiple identical proteins conjugated to the same unique spatial address.
[0007] In some embodiments, each affinity of the first to nth sets of one or more affinity reagents The specificity reagent identifies a family of one or more epitopes present in two or more proteins. In some embodiments, the method comprises identifying a portion of one or more proteins. The identity of the portion is determined based on one or more epitopes in the portion. and determining according to an accuracy threshold. In some examples, one or The first to nth sets of the multiple affinity reagents include more than 100 affinity reagents. In one embodiment, the method comprises detecting only one protein or only one protein isoform. It further includes the use of an affinity reagent that binds to
[0008] In some embodiments, the method includes identifying a portion of one or more proteins. determining the identity of the affinity reagent based on the pattern of binding of the affinity reagent according to an accuracy threshold; In some examples, the substrate is a flow cell. A portion of the one or more proteins is conjugated to the substrate using a photoactivatable linker. In some instances, a portion of one or more proteins is photocleaved. It is conjugated to the substrate using a degradable linker.
[0009] In some examples, at least a portion of at least one set of affinity reagents comprises: In some instances, the antibody is modified to be conjugated to an identifiable tag. In some examples, the identifiable tag is a fluorescent tag. In some examples, the identifiable tag is a magnetic tag. In some examples, the identifiable tag is a nucleic acid barcode. In some cases, the identifiable tag is an affinity tag (e.g., biotin, Flag, myc). In the example, the number of spatial addresses occupied by the identified portion of the protein is The proteins are then counted to quantify the levels of the protein in the sample. The identity of one or more protein moieties is determined by deconvolution. In some instances, the one or more proteins are determined using a The identity of a portion of a protein is determined by an epitope associated with a unique spatial address. In some examples, the combination of The method includes conjugating a portion of one or more proteins to a substrate prior to conjugating the one or more proteins to a substrate. In some instances, the method further comprises denaturing the substrate or the plurality of proteins. The portion of one or more proteins that are specific to the material is a complex mixture of multiple proteins. In some instances, the method is used to identify multiple proteins. can be.
[0010] An additional aspect of the present invention provides a method for identifying a protein comprising the steps of: Each is specific to only one protein or to only one protein family. Obtaining a panel of affinity reagents, which may or may not be specific; determining the binding characteristics of the antibodies in the panel; determining the identity of the protein; repeatedly exposing the protein to a panel of antibodies; determining a set of antibodies that bind; and matching the set of antibodies with the sequence of the protein. To achieve this, one or more deconvolution methods based on the known binding properties of the antibody are used. thereby determining the identity of the protein. In some instances, the protein to be identified is a sample that includes multiple different proteins. In some instances, the method includes identifying multiple proteins in a single sample. It is possible to identify them simultaneously.
[0011] Another aspect of the invention provides a method for identifying a protein, any of which is specific to only one protein or to only one protein family. Obtaining a panel of antibodies, which may or may not be present; determining the binding characteristics of the antibodies in the panel. Repeatedly exposing a protein to a panel of antibodies, and selecting antibodies that do not bind to the protein. and matching the set of antibodies to a protein sequence. using one or more deconvolution methods based on the known binding properties of the antibody. thereby determining the identity of the protein.
[0012] Another aspect of the invention is to use m affinity reagents to identify n proteins in a mixture of proteins. The present invention provides a method for uniquely identifying and quantifying a protein, where n is greater than m and n and m are less than m. and m are positive integers greater than 1, and proteins are separated by intrinsic properties. In some examples, n is about 5, 10, 20, 50, 100, or more times greater than m. 500, 1,000, 5,000, or 10,000 times larger.
[0013] Another aspect of the invention is to use m binding reagents to bind n proteins in a mixture of proteins. The present invention provides a method for uniquely identifying and quantifying a protein, wherein n is greater than m and In some instances, proteins are separated according to a size-based separation method. They have not been separated by any method based on charge or charge.
[0014] Another aspect of the invention is to use m affinity reagents to identify n target proteins in a mixture of protein molecules. The present invention provides a method for uniquely identifying and quantifying single protein molecules, the method comprising: and the protein single molecule is conjugated to a substrate and the individual protein molecules are The protein molecules are arranged such that each has a unique, optically resolvable spatial address. It further includes that the protein monomolecules are spatially separated.
[0015] Another aspect of the invention is to identify unknown proteins 1 to 10 from a pool of n possible proteins. The present invention provides a method for identifying molecules with greater than a threshold degree of certainty. using a panel, where the number of affinity reagents in the panel is m, and m is 10 times smaller than n. is less than 1.
[0016] Another aspect of the invention is to provide a method for identifying an unknown protein selected from a pool of n possible proteins. The present invention provides a method for selecting a panel of m affinity reagents capable of identifying a protein, where m is less than n-1.
[0017] Another aspect of the invention is to provide a method for identifying an unknown protein selected from a pool of n possible proteins. The present invention provides a method for selecting a panel of m affinity reagents capable of identifying a protein, Here, m is less than one tenth of n.
[0018] Another aspect of the invention is that a panel of less than 4000 affinity reagents can be used to identify 20,000 different proteins. said panel of less than 4000 affinity reagents such that each of said affinity reagents can be uniquely identified. A method for selecting
[0019] Another aspect of the invention is to use m binding reagents to bind n proteins in a mixture of proteins. The present invention provides a method for uniquely identifying and quantifying a quality, where m is less than n-1, and each tandem Proteins are identified by their unique profiles of binding by a subset of m binding reagents. do.
[0020] In some examples, the method comprises identifying proteins in a human proteome from a human protein sample. More than 20% of the proteins could be identified, where the proteins were not substantially destroyed during processing. In some instances, the method includes having a protein sequence database available. The proteome of any organism (e.g., yeast, E. coli, C. elegans) In some cases, the protein sequence can be identified. The sequence database contains genome, exome, and / or transcriptome sequences. In some instances, the method may generate more than 4000 affinity sequences. No reagents are required. In some instances, the method can be used to measure more than 100 mg of protein sample. Not needed.
[0021] Another aspect of the present invention provides a method for uniquely identifying a protein molecule. obtaining a panel of affinity reagents; and isolating said protein molecule from one of the affinity reagents in the panel. Each affinity reagent is exposed to one molecule of the protein, and each affinity reagent either binds to one molecule of the protein or does not bind to one molecule of the protein. and collecting the protein to determine its identity. In addition, in some embodiments, the protein is The identity of the affinity molecule is determined by the binding of any individual affinity reagent in a panel of affinity reagents. In some cases, overlapping binding features cannot be determined by the combined data. Affinity reagents having the formula: may be used to enhance affinity for any particular target.
[0022] Another aspect of the invention provides a method for determining a characteristic of a protein, the method comprising: or a portion of a plurality of proteins to a substrate, Alternatively, each of the conjugated portions of the multiple proteins may be attached to a unique space on the substrate. In some examples, the unique spatial address is associated with a unique The method may also include: applying a first to nth set of one or more affinity reagents to the substrate, Each affinity reagent of the first to nth sets of one or more affinity reagents is a nucleic acid having a length of 1 to 10 residues. Each of the affinity reagents of the first to nth sets recognizes an epitope and is The reagent is linked to an identifiable tag. In addition, the method includes one or more affinity After each application of the first to nth sets of reagents to the substrate, the following steps are performed: observing the identifiable tag; detecting one or more of the substrates having one or more observed signals; identifying a plurality of unique spatial addresses; and Each of the one or more protein portions has one or more observed signals. Determining that the polypeptide contains one or more epitopes associated with the polypeptide.
[0023] Another aspect of the invention provides a method for determining a characteristic of a protein, the method comprising: or a portion of a plurality of proteins to a substrate, wherein one Alternatively, each of the conjugated portions of the multiple proteins may be attached to a unique space on the substrate. The method also includes: determining whether the first to second affinity reagents are associated with a first address of the one or more affinity reagents; n sets of one or more affinity reagents, Each affinity reagent in the set is a nucleotide sequence that is linked to one or more of the nucleotides present in one or more proteins. A set of one to n affinity reagents that recognize a family of epitopes and Each affinity reagent in the sample is linked to an identifiable tag. After each application of the plurality of affinity reagents to the substrate of the first to nth sets, the following steps are performed: observing the identifiable tag; detecting one or more of the substrates having an observed signal; identifying a plurality of unique spatial addresses; and Each of the portions of one or more proteins having the formula A determining process.
[0024] A further aspect of the invention provides a method for identifying a protein, the method comprising the steps of: The method includes the steps of: obtaining a panel of affinity reagents having a known degree of nonspecificity; determining the binding characteristics of the affinity reagents, by repeatedly exposing the protein to the panel of affinity reagents; determining a set of affinity reagents that bind to the protein; and Matching a set of drugs to a protein sequence based on the known binding properties of affinity reagents using one or more deconvolution methods to identify the protein The process of determining the identity of a protein.
[0025] In addition, another aspect of the present invention provides a method for identifying a protein, the method comprising the steps of: The method includes the steps of: obtaining a panel of affinity reagents having known degrees of nonspecificity; determining the binding characteristics of the affinity reagents by repeatedly exposing the protein to the panel of affinity reagents; determining a set of affinity reagents that do not bind to the protein; and To match a set of affinity reagents to a protein sequence, we use the known binding properties of the affinity reagents. using one or more deconvolution methods based on Determining the identity of the protein.
[0026] INCORPORATION BY REFERENCE All publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety. Each such patent or patent application is specifically and individually indicated as being incorporated by reference. No. 6,393,311, which are incorporated herein by reference to the same extent as if set forth herein. [Brief description of the drawings]
[0027] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the invention and its advantages will be readily apparent from the following detailed description which sets forth illustrative embodiments in which the principles of the invention are utilized. Further details can be had by reference to the detailed description and accompanying drawings, in which:
[0028] [Figure 1] 1 shows a first illustration of protein quantification by anti-peptide antibody decoding, according to some embodiments. [Diagram 2] 1 shows a second illustration of protein quantification by anti-peptide antibody decoding, according to some embodiments. [Diagram 3] 1 shows flow cell conjugation according to some embodiments. [Figure 4] 1 shows a grid of unique spatial addresses on a flow cell, according to some embodiments. [Diagram 5] 1 shows the deconstruction of a protein as a set of peptides that can be matched with d-coded antibodies, according to some embodiments. [Figure 6] 1 shows a schematic of protein identification / quantification by anti-peptide antibody decoding, according to some embodiments. [Figure 7] 1 shows observations of a first set of anti-peptide antibodies, according to some embodiments. [Figure 8] 1 shows observations of a second set of anti-peptide antibodies, according to some embodiments. [Figure 9] 1 shows observations of a third set of anti-peptide antibodies, according to some embodiments. [Figure 10] 1 shows computer decoding of antibody measurement data according to some embodiments. [Figure 11] 1 shows quantification of the proteome, according to some embodiments. [Figure 12] 1 illustrates an example of an exception list, according to some aspects. [Figure 13] FIG. 1 shows the coverage of 3-mer d-coded antibody sampling that may be required for quantification according to an embodiment. [Figure 14] 1 illustrates a computer control system programmed or otherwise configured to carry out the methods provided herein. [Figure 15] 1 shows an example of the impact of the number of 3-mer d-coded probes on identifiability versus proteome coverage, according to embodiments herein. [Figure 16A] 1 shows an image depicting a single protein molecule conjugated to a substrate according to embodiments herein. [Figure 16B] 16B shows an image depicting a magnified portion of the indicated area of FIG. 16A, with conjugated proteins indicated by arrows, according to embodiments herein. [Figure 17] 1 shows the identification of proteins according to embodiments of the present disclosure. [Figure 18] 1 shows an illustration of protein identification according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0029] Detailed Description of the Invention In some examples, the approach may include three aspects: 1) protein and 2) an addressable substrate to which the peptide and / or protein fragment can be conjugated; a set of affinity reagents, e.g., each affinity reagent capable of binding to a peptide with a variety of specificities; and 3) to predict the identity of proteins at precise spatial addresses in the matrix. Prior knowledge of the binding characteristics of the affinity reagent, the affinity at each address in the substrate, Specific patterns of binding of reagents and / or in a mixture (e.g., the human proteome) a database of possible protein sequences; In some examples, the precise spatial address is a unique spatial address. It may be.
[0030] sample The sample may be any biological sample that contains proteins. The sample may be a tissue or cell sample. The antibody may be taken from a cell or from the environment of a tissue or cell. Samples include tissue biopsies, blood, plasma, extracellular fluids, cultured cells, culture media, discarded tissues, plants Materials, synthetic proteins, archaeal, bacterial and / or viral samples, mycelial tissue In some instances, the protein may be from a sample preparation. During production, they are isolated from their primary source (cells, tissues, body fluids such as blood, environmental samples, etc.). The protein may or may not be purified from its primary source. In some cases, the primary source is homogenized prior to further processing. In some cases, the cells are lysed using a buffer such as RIPA buffer. A buffer may also be used at this stage. The sample is diluted to remove lipids and particulate matter. The sample may also be purified to remove nucleic acids. The sample may be treated with RNase and DNase. It may comprise a protein, a protein fragment, or a partially degraded protein.
[0031] The sample may be taken from a subject having a disease or disorder. The disease or disorder may be an infectious disease or disorder. Disease, immune disorders or disorders, cancer, genetic disorders, degenerative diseases, lifestyle-related diseases, injuries, rare Infectious diseases may be caused by bacteria, viruses, fungi, and Non-limiting examples of cancers include bladder cancer, lung cancer, and / or cancer of the lungs. Cancer, brain cancer, melanoma, breast cancer, non-Hodgkin's lymphoma, cervical cancer, ovarian cancer, colorectal Cancer, pancreatic cancer, esophageal cancer, prostate cancer, kidney cancer, skin cancer, leukemia, thyroid cancer, liver Some examples of genetic diseases or disorders include, but are not limited to, cancer of the liver, pancreas, kidney, and uterus. However, cystic fibrosis, Charcot-Marie-Tooth disease, Huntington's disease, and Peutzfeldt-Jakob disease are not - Includes Jeghers syndrome, Down syndrome, rheumatoid arthritis, and Tay-Sachs disease. Non-limiting examples of lifestyle-related diseases include obesity, diabetes, arteriosclerosis, heart disease, stroke, high blood pressure, and liver cirrhosis. These include kidney disease, kidney inflammation, cancer, chronic obstructive pulmonary disease (COPD), hearing problems, and chronic back pain. Some examples include, but are not limited to, abrasions, brain injuries, contusions, burns, concussions, Congestive heart failure, injuries at construction sites, dislocations, flail chest, fractures, hemothorax, herniated discs, hip Propointer, hypothermia, lacerations, pinched nerve, pneumothorax, ribs Fractures, sciatica, spinal cord injuries, tendon, ligament and fascial injuries, traumatic brain injury and whiplash injuries The sample may be taken before and / or after treatment of a subject with a disease or disorder. The sample may be taken before and / or after treatment. The sample may be taken during or after treatment. Multiple samples may be taken during the regimen to monitor the effects of treatment over time. The sample may be taken from a subject known to have an infectious disease for which a diagnostic antibody is not available. The antibody may be taken from a subject known or suspected to have a pulmonary embolism.
[0032] The sample may be taken from a subject suspected of having a disease or disorder. You may experience unexplained symptoms such as fatigue, nausea, weight loss, aches and pains, weakness, or amnesia. The sample may be taken from a subject who is undergoing a study. The sample may be taken from a subject who has symptoms with a known cause. Samples may be taken based on family medical history, age, environmental exposures, lifestyle risk factors, or The risk of developing a disease or disorder due to factors such as the presence of other known risk factors It may be taken from a subject.
[0033] The sample may be taken from an embryo, a fetus, or a pregnant woman. In some examples, the sample is In some instances, the protein may be isolated from maternal plasma. Protein isolated from circulating fetal cells in fluid.
[0034] The protein may be treated to remove modifications that may interfere with epitope binding. For example, proteins may be post-translationally treated with glycosidases to remove glycosylation. Proteins are treated with a reducing agent to reduce disulfide bonds in the protein. The protein may be treated with a phosphatase to remove the phosphate groups. Other non-limiting examples of post-translational modifications that can be removed include acetate, amide groups, methyl groups, lipids, and the like. , ubiquitination, myristoylation, palmitoylation, isoprenylation or prenylation ( For example, farnesol and geranylgeraniol), farnesyl, geranylgeraniol glypiation, lipoylation, flavin moiety attachment, phosphoprotein The samples also included post-translational protein cleavage, endothelinylation, and retinylidene Schiff base formation. In some instances, the protein may be treated with a phosphatase inhibitor. In some instances, an acid may be added to the sample to protect disulfide bonds. A shaping agent may be added.
[0035] The protein may then be fully or partially denatured. Proteins can be completely denatured. Proteins can be denatured by detergents, strong acids or bases, concentrated condensed inorganic salts, organic solvents (e.g. alcohol or chloroform), irradiation, Alternatively, the protein may be denatured by application of an external stress such as heat. Proteins may also be denatured by precipitation, lyophilization, and denaturation in a denaturing buffer. The protein may be denatured by heating. The protein may be chemically modified. A modification method that is less likely to cause artifacts may be chosen.
[0036] The sample protein is subjected to conjugation to produce shorter polypeptides. The remaining protein may be processed either before or after the synthesis to generate fragments. It may be partially digested with an enzyme such as ribozyme K or left intact. In a further example, the protein may be exposed to a protease, such as trypsin. Additional examples of proteases include serine proteases, cysteine proteases, threoproteases, and ribozymes. Proteinases, Aspartic acid proteases, Glutamic acid proteases, Metallopharyngoproteinases protease, and asparagine peptide lyase.
[0037] In some cases, extremely large and small proteins (e.g. titin) It may be useful to remove such proteins by filtration or other suitable methods. In some instances, extremely large proteins can be removed by , proteins greater than 500 kD, 600 kD, 650 kD, 700 kD, 750 kD, 800 kD, or 850 kD In some examples, extremely large proteins may include about 8,000 amino acids, Approximately 8,500 amino acids, approximately 9,000 amino acids, approximately 9,500 amino acids, approximately 10,000 amino acids, approximately 10,500 amino acids The polypeptide may comprise a protein of more than about 11,000 amino acids, or more than about 15,000 amino acids. In some instances, a small protein is less than about 10 kD, less than 9 kD, less than 8 kD, less than 7 kD, < 6 kD, < 5 kD, < 4 kD, < 3 kD, < 2 kD, or < 1 kD In some instances, a small protein may include less than about 50 amino acids, 45 amino acids, or A protein having less than 40 amino acids, less than 35 amino acids, or less than about 30 amino acids. Extremely large or small proteins may be isolated by size exclusion chromatography. Extremely large proteins can be removed by size exclusion chromatography. isolated and treated with a protease to produce intermediate sized polypeptides, It may be recombined with intermediate sized proteins.
[0038] In some cases, the proteins may be arranged in size order. In this method, proteins are arrayed by sorting the proteins into microwells. In some cases, the proteins may be sorted into nanowells. In some cases, proteins may be aligned by SD The alignment may be performed by passing the gel through a gel such as an S-PAGE gel. In some cases, proteins may be aligned by other methods of size fractionation. In some cases, proteins may be separated based on charge. Proteins may be separated on the basis of hydrophobicity. In some cases, proteins may be separated from other In some cases, proteins may be separated based on their physical characteristics. In some cases, proteins may be separated under non-denaturing conditions. In some cases, different fractions of the fractionated protein may be obtained by eluting different fractions of the substrate. In some cases, different portions of the separated proteins may be located in different regions. The portions may be disposed on different regions of the substrate. In some cases, the protein sample is , separated on an SDS-PAGE gel, and the proteins are sorted by size in a continuum. As classified, the material may be transferred from an SDS-PAGE gel to a substrate. The protein sample may be divided into three fractions based on size, and the three fractions may be They may be applied to first, second, and third regions on the substrate, respectively. The proteins used in the systems and methods described herein are In some cases, the systems described herein may be classified as And the proteins used in the methods need not be classified.
[0039] Proteins may be tagged, for example with identifiable tags, to allow for sample multiplexing. , may be tagged. Some non-limiting examples of identifiable tags include: Fluorophores or nucleic acid barcoded base linkers. The fluorophores used are: GFP, YFP, RFP, eGFP, mCherry, tdtomato, FITC, Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 680, Alexa Fluor 750, Pacific Blue, Coumarin, BODIPY FL, Pacific Green, Oregon Green, Cy3, Cy5, Pacific Orange e, TRITC, Texas Red, R-phycoerythrin, allophycocyanin, etc. The fluorophores may include any fluorescent protein or other fluorophore known in the art. .
[0040] Any number of protein samples can be multiplexed. For example, multiplexed reactions can be 2, 3, 4 , 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, about 20, about 25, about 30, about 35 , about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100 The identifiable tags may comprise proteins from 100 or more initial samples. This may provide a means to examine proteins with respect to the sample from which they were derived, or to compare proteins from different samples. Proteins may be induced to segregate to different areas on a solid support.
[0041] Base material In some embodiments, the protein is then chemically attached to the substrate. In some cases, the protein is applied to a functionalized substrate to bind to the In some cases, the protein may be attached to the substrate via the attachment of a thiol. The protein may be attached to the substrate via the attachment of a nucleic acid. A mediator may be applied, after which the mediator adheres to the substrate. The protein is then captured onto a surface (e.g., a thiolated surface). In some cases, the protein may be conjugated to a protein (e.g., a gold bead). One protein may be conjugated to each bead. may be conjugated to beads (e.g., one protein per bead), and The beads may be captured on a surface (e.g., in microwells and / or nanowells). in).
[0042] The substrate may be any substrate capable of forming a solid support. As used herein, a substrate, or solid substrate, is a material to which a protein is covalently or non-covalently attached. Solid substrates may refer to any solid surface to which the solid can be attached. Non-limiting examples of solid substrates include particles, bilayers, and the like. Tablets, slides, surfaces of device components, membranes, flow cells, wells, chambers, macros It contains a fluid chamber and may be flat or curved or have other shapes. and may be smooth or may have irregularities. In some cases, the surface of the substrate In some cases, the surface of the substrate may include nanowells. In some cases, the surface of the substrate may be incorporated with one or more nanowells. In some embodiments, the substrate may comprise one or more microwells. The materials are glass, sugars such as dextran, and plastics such as polystyrene or polypropylene. Stick, polyacrylamide, latex, silicone, gold or other metals, or cellulose The nucleic acid may be composed of a nucleic acid sequence and may allow for covalent or non-covalent attachment of oligonucleotides. For example, the surface of the substrate may be modified with a maleimide or a maleimide. The carboxylates are functionalized by modification with specific functional groups such as phosphate moieties or succinate moieties. or may have chemically reactive groups such as amino, thiol, or acrylate groups. The silane may be derivatized by modification with, for example, silanization. The drugs are aminopropyltrimethoxysilane, aminopropyltriethoxysilane, and The substrate contains 4-aminobutyltriethoxysilane. The substrate is N-hydroxysuccinimide (NHS The glass surface may also be functionalized with, for example, epoxy silanes, acrylates, etc. Tosilane or acrylamide silane is used to form acrylate or epoxy. The substrate and treatment for oligonucleotide attachment are preferably or is stable to repeated binding, washing, imaging, and elution steps. In some instances, the substrate may be a slide or a flow cell.
[0043] Ordered arrays of functional groups can be fabricated using techniques such as photolithography, dip-pen nanolithography, and lithography, nanoimprint lithography, nanosphere lithography, nanoball lithography lithography, nanopillar arrays, nanowire lithography, scanning probe lithography, Thermochemical lithography, Thermal scanning probe lithography, Localized oxidation nanolithography, Molecular May be fabricated by self-assembly, stencil lithography, or electron beam lithography The functional groups in the ordered array are such that each functional group is 200 nanometers away from any other functional group. (nm) or less than about 200 nm, about 225 nm, about 250 nm, about 275 nm, about 30 0 nm, approx. 325 nm, approx. 350 nm, approx. 375 nm, approx. 400 nm, approx. 425 nm, approx. 450 nm, approx. 475 nm, approx. 50 0 nm, approx. 525 nm, approx. 550 nm, approx. 575 nm, approx. 600 nm, approx. 625 nm, approx. 650 nm, approx. 675 nm, approx. 70 0 nm, approx. 725 nm, approx. 750 nm, approx. 775 nm, approx. 800 nm, approx. 825 nm, approx. 850 nm, approx. 875 nm, approx. 90 0 nm, approx. 925 nm, approx. 950 nm, approx. 975 nm, approx. 1000 nm, approx. 1025 nm, approx. 1050 nm, approx. 1075 nm, approx. 1100 nm, approx. 1125 nm, approx. 1150 nm, approx. 1175 nm, approx. 1200 nm, approx. 1225 nm, approx. 1250 nm, approx. 1 275 nm, approx. 1300 nm, approx. 1325 nm, approx. 1350 nm, approx. 1375 nm, approx. 1400 nm, approx. 1425 nm, approx. 1450 nm, approx. 1475 nm, approx. 1500 nm, approx. 1525 nm, approx. 1550 nm, approx. 1575 nm, approx. 1600 nm, approx. 1625 nm , about 1650 nm, about 1675 nm, about 1700 nm, about 1725 nm, about 1750 nm, about 1775 nm, about 1800 nm, approx. 1825 nm, approx. 1850 nm, approx. 1875 nm, approx. 1900 nm, approx. 1925 nm, approx. 1950 nm, approx. 1975 nm, approx. 2 The randomly spaced functional groups may be spaced apart from each other such that the distance between the functional groups is greater than 10,000 nm, or greater than 2000 nm. The functional groups are spaced apart from any other functional groups by at least about 50 nm, about 100 nm, about 150 nm, about 2 00 nm, approx. 250 nm, approx. 300 nm, approx. 350 nm, approx. 400 nm, approx. 450 nm, approx. 500 nm, approx. 550 nm, approx. 6 00 nm, approx. 650 nm, approx. 700 nm, approx. 750 nm, approx. 800 nm, approx. 850 nm, approx. 900 nm, approx. 950 nm, approx. 1 000 nm, or in a dense state such that it is greater than 100 nm.
[0044] The substrate may be indirectly functionalized. For example, the substrate may be PEGylated and have a functional group may be applied to all of the PEG molecules, or to a subset of the PEG molecules. Additionally, in some cases, beads (e.g., gold beads) may be conjugated. and the beads may then be captured onto a surface (e.g., a thiolated surface). In some cases, one protein may be conjugated to each bead. In this case, the protein may be conjugated to a bead (e.g., one bead). The beads may be captured on a surface (e.g., a microsphere). in wells and / or nanowells).
[0045] The substrate may include microscale or nanoscale structures (e.g., microwells, nanoparticles, etc.). Nanowells, micropillars, single molecule arrays, nanoballs, nanopillars, or nanowa The fibers may be functionalized using techniques suitable for the application of the fluoroscopy (e.g., ordered structures such as ears). In some cases, the substrate may have microwells of different sizes. In the microwell, the size may be 1 micrometer (μm), about 2 μm, about 3 μm, μm, approx. 4 μm, approx. 5 μm, approx. 6 μm, approx. 7 μm, approx. 8 μm, approx. 9 μm, approx. 10 μm, approx. 15 μm, approx. 20 μm, approx. 25 μm, approx. 30 μm, approx. 35 μm, approx. 40 μm, approx. 45 μm, approx. 50 μm, approx. 55 μm, Approximately 60 μm, approximately 65 μm, approximately 70 μm, approximately 75 μm, approximately 80 μm, approximately 85 μm, approximately 90 μm, approximately 95 μm, Approx. 100 μm, Approx. 105 μm, Approx. 110 μm, Approx. 115 μm, Approx. 120 μm, Approx. 125 μm, Approx. 130 μm, Approx. 1 35 μm, approx. 140 μm, approx. 145 μm, approx. 150 μm, approx. 155 μm, approx. 160 μm, approx. 165 μm, approx. 170 μm, approx. 175 μm, approx. 180 μm, approx. 185 μm, approx. 190 μm, approx. 195 μm, approx. 200 μm, approx. 205 μm , approximately 210 μm, approximately 215 μm, approximately 220 μm, approximately 225 μm, approximately 230 μm, approximately 235 μm, approximately 240 μm, Approx. 245 μm, Approx. 250 μm, Approx. 255 μm, Approx. 260 μm, Approx. 265 μm, Approx. 270 μm, Approx. 275 μm, Approx. 2 80 μm, approx. 285 μm, approx. 290 μm, approx. 295 μm, approx. 300 μm, approx. 305 μm, approx. 310 μm, approx. 315 μm, approx. 320 μm, approx. 325 μm, approx. 330 μm, approx. 335 μm, approx. 340 μm, approx. 345 μm, approx. 350 μm , approximately 355 μm, approximately 360 μm, approximately 365 μm, approximately 370 μm, approximately 375 μm, approximately 380 μm, approximately 385 μm, Approx. 390 μm, approx. 395 μm, approx. 400 μm, approx. 405 μm, approx. 410 μm, approx. 415 μm, approx. 420 μm, approx. 4 25 μm, approx. 430 μm, approx. 435 μm, approx. 440 μm, approx. 445 μm, approx. 450 μm, approx. 455 μm, approx. 460 μm, approx. 465 μm, approx. 470 μm, approx. 475 μm, approx. 480 μm, approx. 485 μm, approx. 490 μm, approx. 495 μm , about 500 μm, or greater than 500 μm. In some cases, the substrate may be The microwells may have diameters ranging from 5 μm to 500 μm. In some cases, the substrate may have microwells ranging in size from about 5 μm to about 500 μm. In some cases, the substrate contains microwells ranging in size from 10 μm to 100 μm. In some cases, the substrate may have a size ranging from about 10 μm to about 100 μm. In some cases, the substrate may have microwells of different sizes. A range of different sizes are provided so that proteins can be sorted into microwells of different sizes. In some cases, the microwells in the substrate may have a size of 10 μm. may be distributed by size (e.g., the larger microwells are in the first region (In some cases, the microwells are distributed in a first region and the smaller microwells are distributed in a second region.) In some cases, the substrate may have about 10 different sized microwells. In this case, the substrates may be about 20 different sizes, about 25 different sizes, about 30 different sizes, about 35 different sizes,about 40 different sizes,about 45 different sizes,about 50 different sizes, About 55 different sizes,About 60 different sizes,About 65 different sizes,About 70 different sizes, About 75 different sizes, about 80 different sizes, about 85 different sizes, about 90 different sizes, About 95 different sizes, about 100 different sizes, or more than 100 different sizes of micro The substrate may have a well.
[0046] In some cases, the substrate may have nanowells of different sizes. In this case, the nanowell may be about 100 nanometers (nm), about 150 nm, about 200 nm, about 250 nm, or about 300 nm. nm, approx. 300 nm, approx. 350 nm, approx. 400 nm, approx. 450 nm, approx. 500 nm, approx. 550 nm, approx. 600 nm, approx. 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, or In some cases, the substrate may have a size between 50 nm and 1 micrometer. In some cases, the nanowells may have diameters in the range of 100 nm to 1 micrometer. In the method, the substrate may have nanowells with sizes ranging from 100 nm to 500 nm. In some cases, the substrate allows different sized proteins to be placed in different sized nanowells. In some cases, the nanowells may have a range of different sizes so that they can be categorized. In this case, the nanowells in the substrate may be distributed by size (e.g., larger The larger nanowells are distributed in a first region and the smaller nanowells are distributed in a second region. In some cases, the substrate may have about 10 nanowells of different sizes. In some cases, the substrate may include about 20 different sizes, or more than 30 different sizes. The nanowell may have
[0047] In some cases, the substrate may be a matrix of different sized proteins or different sized nanoparticles. A range of different sizes of nanowells may be used so that they can be grouped into wells and / or microwells. In some cases, the substrate may have wells and / or microwells. The nanowells and / or microwells may be distributed by size (e.g., For example, larger microwells are distributed in a first region and smaller nanowells are distributed in a second region. In some cases, the substrate contains about 10 nanoparticles of different sizes. In some cases, the substrate may have wells and / or microwells. 20 different sizes,about 25 different sizes,about 30 different sizes,about 35 different sizes,about 40 different sizes,about 45 different sizes,about 50 different sizes,about 55 different sizes,about 60 different sizes, about 65 different sizes, about 70 different sizes, about 75 different sizes, about 80 different sizes, about 85 different sizes, about 90 different sizes, about 95 different sizes, about Nanowells and / or microwells of 100 or more different sizes The substrate may have a well.
[0048] The substrate may include metal, glass, plastic, ceramic, or a combination thereof. In some preferred embodiments, the solid substrate is a flow cell. The flow cell may be composed of a single layer or multiple layers. The cell consists of a base layer (e.g., made of borosilicate glass), a channel layer (e.g., For example, made of etched silicon), and a cover or top layer. When the layers are assembled together, an enclosed channel can be formed, which , with inlets / outlets at both ends through the cover. The thickness of each layer is variable, but is preferred. The layer may be made of, but is not limited to, photosensitive glass, boron, The present invention relates to a method for manufacturing a semiconductor device, comprising the steps of: The different layers may be constructed from any suitable material known in the art. or may be constructed from different materials.
[0049] In some embodiments, the flow cell has an opening for the channel at the bottom of the flow cell. The flow cell can contain millions of attached probes at locations that can be visualized separately. In some embodiments, the method of the present invention may include a target conjugation site. The various flow cells used have different numbers of channels (e.g., 1 channel, 2 or more). 1 channel, 3 or more channels, 4 or more channels, 6 or more channels, 8 or more channels channels, 10 or more channels, 12 or more channels, 16 or more channels, or more than 16 Various flow cells may contain channels of different depths or widths. These may vary between channels in one flow cell or may be The flow cell channels may also have different depths. and / or width may vary. For example, a channel may have one or more In several places, the depth is less than about 50 μιη, the depth is less than about 50 μιη, the depth is less than about 100 μιη, Depth of about 100 μιη, depth of about 100 μιη to about 500 μιη, depth of about 500 μιη, The channels may be greater than about 500 μm deep. The channels may be, but are not limited to, circular, Can have any cross-sectional shape, including semicircular, rectangular, trapezoidal, triangular, or oval cross-sections. possible.
[0050] The protein may be spotted, dropped or pipetted onto the substrate; It may be poured, washed, or otherwise applied. In the case of moiety-functionalized substrates, protein modification is not required. A substrate functionalized with a moiety (e.g., a sulfhydryl, an amine, or a linker nucleic acid). In the case of , a cross-linking reagent (e.g., disuccinimidyl suberate, NHS, sulfonamide In the case of a substrate functionalized with a linker nucleic acid, the sample protein may be , may be modified with a complementary nucleic acid tag.
[0051] In some cases, the protein may be conjugated to a nucleic acid. The nucleic acid nanoballs can then be formed, thereby linking the nucleic acid nanoballs to the proteins. When nucleic acid nanoballs are attached to a substrate, the proteins attached to the nucleic acid nanoballs The DNA nanoballs are attached to the substrate (e.g., by adsorption or conjugation). The substrate may be functionalized with amines to which the nucleic acid nanoballs can be attached. The substrate may have a functionalized surface.
[0052] In some cases, the nucleic acid nanoballs may have functionally active ends (e.g., maleimines). The protein may then be conjugated to the nanoball. The nucleic acid nanoball may be conjugated to a protein, thereby linking the nucleic acid nanoball to the protein. When the nucleic acid nanoball is attached to the substrate, the protein attached to the nucleic acid is transferred to the substrate via the nucleic acid nanoball. The DNA nanoballs are attached to the substrate (e.g., by adsorption or conjugation). The substrate can be functionalized with amines to which the nucleic acid nanoballs can be attached. It may have a surface.
[0053] Photoactivatable crosslinkers may be used to direct crosslinking of samples to specific areas on a substrate. Photoactivatable crosslinkers detect proteins by attaching each sample to a known area of a substrate. Photoactivatable crosslinkers may be used to allow for sample multiplexing. By detecting the fluorescent tag before cross-linking the protein, successfully tagged proteins can be identified. Examples of photoactivatable crosslinkers include, but are not limited to, N- 5-Azido-2-nitrobenzoyloxysuccinimide, Sulfosuccinimidyl 6-(4'-azido Dihydro-2'-nitrophenylamino)hexanoate, Succinimidyl 4,4'-azidopentanoate Succinimidyl 6-(4,4'-azidotanoate, sulfosuccinimidyl 4,4'-azidotanoate Dipentanamide)hexanoate, Sulfosuccinimidyl 6-(4,4'-azipentanamide hexanoate, succinimidyl 2-((4,4'-azipentanamido)ethyl)-1,3'-dithio Opropionate, and sulfosuccinimidyl 2-((4,4'-azipentanamido)ethyl )-1,3'-dithiopropionate.
[0054] Samples can also be multiplexed by restricting the binding of each sample to a separate area on the substrate. For example, the substrate may be organized into lanes. Another method for multiplexing is to Repeatedly apply the sample across the entire substrate, removing non-specific protein binding after each sample application. This is followed by a protein detection step utilizing a reagent or dye. Examples are SYPRO® Ruby, SYPRO® Orange, SYPRO® Red, SYPRO® Fluorescent proteins, such as O® Tangerine, and Coomassie™ Fluor Orange The gel may include a gel stain.
[0055] By tracking the location of all proteins after each addition of sample, each location on the substrate It is possible to determine the stage which first contained the protein, and therefore It is possible to determine the sample from which the sample originates. This method also allows the determination of the amount of the sample from the substrate after each application of the sample. Saturation conditions can be determined and allow for maximization of protein binding on the substrate. Assume that only 30% of activated positions are occupied by protein after the first application of sample. Then, either a second application of the same sample or a different sample may be applied. stomach.
[0056] The polypeptide may be attached to the substrate by another moiety. In some examples, The polypeptide may be linked via the N-terminus, C-terminus, both termini, or via internal residues. It may be attached.
[0057] In addition to permanent crosslinkers, the use of photocleavable linkers for some applications; and thereby allowing selective extraction of proteins from the substrate after analysis. In some cases, it may be appropriate to make the photocleavable crosslinker In some cases, the photocleavable crosslinker may be used for multiplexing different samples. One or more samples may be used in the multiplexed reaction. In some cases The multiplexed reactions included a control sample that was crosslinked to the substrate via a permanent crosslinker, and a photocleavable sample. The experimental sample may include a crosslinked substrate via a crosslinking agent.
[0058] Each conjugated protein is optically Spatially separated from each other conjugated protein so that it can be resolved Proteins may therefore be individually labeled with a unique spatial address. In some embodiments, this may be achieved by integrating each protein molecule with other protein molecules. The low concentration of protein and the low density of attachment on the substrate are spatially separated from each other. This can be achieved by conjugation using attachment sites such as photoactivatable crosslinkers. When a photoactivation method is used, photoactivation is performed so that the protein is added to a predetermined position. A pattern may be used.
[0059] In some methods, the purified bulk protein is purified by To determine the identity of the marker, conjugated to a substrate and using the methods described herein, Bulk proteins include purified proteins collected together. In some instances, the bulk protein may be Conjugated proteins or bulk proteins are optically resolvable. at a position spatially separated from each of the other bulk proteins The protein or bulk protein may thus be conjugated to In some embodiments, the nucleosomes may be individually labeled with a spatial address of A low concentration such that each protein molecule is spatially separated from each of the other protein molecules by conjugation using the low density of attachment sites on the protein and substrate. By way of example, when a photoactivatable crosslinker is used, one or more proteins may be crosslinked. Light patterns may be used so that texture is added at predetermined locations.
[0060] In some embodiments, each protein may be associated with a unique spatial address. For example, if proteins are attached to a substrate at spatially separated locations, each protein A protein can be assigned an indexed address, such as by coordinates. In this example, a grid of pre-assigned unique spatial addresses is In some embodiments, the location of each protein may be determined by a fixed mark on the substrate. The substrate may include easily identifiable fixed marks so that the relative position of the substrate may be determined. In some examples, the substrate may have grid lines and / or In some instances, the surface of the substrate may be crosslinked. To provide a reference for locating the protein, The conjugated polypeptide may be marked with a pattern such as the outer edge of the conjugated polypeptide. Its shape is also used as a reference to determine the unique position of each spot. good.
[0061] The substrate may also include conjugated protein standards and controls. The gated protein standards and controls were prepared by conjugating known positions of the proteins. In some instances, the conjugate may be a peptide or protein of known sequence. The incorporated protein standards and controls can serve as internal controls in the assay. The protein may be applied to the substrate from a purified protein stock, or Nucleic Acid-Programmable Protein Array (N The polymer may be synthesized on the substrate by a process such as APPA.
[0062] In some instances, the substrate may include fluorescent standards. These fluorescent standards may be These fluorescent standards may also be used to calibrate the intensity of the fluorescent signal between samples. In addition, to correlate the intensity of the fluorescent signal with the number of fluorophores present in a given area, Fluorescent standards may be used to identify the different types of fluorophores used in the assay. It may include some or all of the fore.
[0063] Affinity Reagents Once proteins from the sample are conjugated to the substrate, multiple affinity reagent measurements can be performed. The measurement process described herein can be used to measure a variety of affinities. A reactive reagent may be used.
[0064] Affinity reagents are any suitable reagent that binds with reproducible specificity to proteins or peptides. For example, the affinity reagent may be an antibody, an antibody fragment, an aptamer, or a peptide. In some instances, a monoclonal antibody may be selected. In some cases, an antibody fragment such as a Fab fragment may be selected. The reagent may be a commercially available affinity reagent, such as a commercially available antibody. A preferred affinity reagent would be one that screens commercially available affinity reagents to identify those with useful characteristics. In some cases, the affinity reagent may be selected by screening. Alternatively, the antibodies may be screened for their ability to bind to a single protein. In some cases, affinity reagents are those that bind to an epitope or amino acid sequence. In some cases, the affinity reagents may be screened for their ability to The group is able to distinguish similar proteins (e.g., those with highly similar sequences) by differential binding. Several In some cases, affinity reagents increase binding specificity for a particular protein. Therefore, affinity reagents may be screened for overlapping binding characteristics. Screening may be performed in a variety of different formats. One example is NAPPA or epitope typing. The aim of the present invention may be to screen affinity reagents against a multi-layer array. In this study, protein-specific affinity reagents designed to bind to protein targets are used. In some cases, multiple targets may be used (e.g., commercially available antibodies or aptamers). The protein-specific affinity reagents may be mixed prior to the binding assay. For each step, a new mixture of protein-specific affinity reagents is It may be selected to include a randomly selected subset of the available affinity reagents. For example, each subsequent mixture may contain more than one affinity reagent present in the mixture. In some cases, tamper-evident messages may be generated in the same random fashion, with the hope that Protein identification is accomplished more rapidly using a mixture of protein-specific affinity reagents. In some cases, such mixtures of protein-specific affinity reagents may be used. At each step of the experiment, the percentage of the unknown protein that is bound by the affinity reagent is increased. The affinity reagent mixture may be added at 1%, 5%, 10%, 20%, or 40% of all available affinity reagents. It may consist of 0%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or more.
[0065] Affinity reagents may have high, medium, or low specificity. In some instances, the affinity reagent may recognize several different epitopes. The affinity reagent may recognize an epitope present on two or more different proteins. In some instances, the affinity reagent recognizes an epitope that is present on many different proteins. In some cases, the affinity reagents used in the methods of the present disclosure may be In some cases, the present invention may be highly specific for only one epitope. The affinity reagent used in the method shown is directed to only one epitope that contains a post-translational modification. The nucleic acid sequence may be highly specific.
[0066] In some embodiments, the affinity reagent directed to identifying a target amino acid sequence is As used in the methods described herein, they may also be distinguished from one another. It may actually include a group of different components that may not even be distinguishable from one another. In particular, Different components that can be used to identify the same target amino acid sequence may be used to identify the same target amino acid. The same detection moiety may be used to identify a nucleic acid sequence. Affinity reagents that bind to the trimeric amino acid sequence (AAA) can bind to any of the adjacent sequences. It may contain only one probe that binds to the trimeric AAA sequence without any effect, or it may contain only one probe that binds to the trimeric AAA sequence without any effect, Aβ (wherein α and β can be any amino acid) is a five amino acid epitope with different structure. The second case may include a group of 400 probes, each of which binds to a different top. In some cases, the 400 probes were each present in equal amounts. In some cases of the second case, the 400 probes may be The characteristic binding sites for each probe are determined so that they have an equal probability of binding a given 5 amino acid epitope. They may be combined such that the amount of each probe can be weighted according to binding affinity.
[0067] Novel affinity reagents may be generated by any method known in the art. Methods for developing affinity reagents include SELEX, phage display, and seeding. In some instances, affinity reagents may be designed using structure-based drug design methods. Structure-based drug design (or direct drug design) involves identifying an epitope of interest and a parent molecule. It utilizes knowledge of the three-dimensional structure of the binding site of the compatibility agent.
[0068] In some cases, the affinity reagent may be labeled with a nucleic acid barcode. In examples, the nucleic acid barcodes may be used to purify the affinity reagents after use. In some instances, the nucleic acid barcode labels the affinity reagents for repeated use. In some cases, the affinity reagent may be used to The molecule may be labeled with a fluorophore that can be used to classify the molecule.
[0069] In some cases, multiple affinity reagents labeled with nucleic acid barcodes are multiplexed. The affinity reagents may be multiplexed and then detected using a complementary nucleic acid probe. The other group demonstrated a single cycle assay using multiple complementary nucleic acids with distinct detection moieties. In some cases, multiplexed groups of affinity reagents may be used to detect The method comprises multiple cycles using a single complementary nucleic acid conjugated to a detection moiety. In some cases, multiplexed groups of affinity reagents may be used for detection. a plurality of complementary nucleic acids each conjugated to a distinct detection moiety; In some cases, multiplexing of affinity reagents may be performed. The group comprises multiple phases each conjugated to a separate group of detection moieties. Detection may occur in multiple cycles using complementary nucleic acids.
[0070] In some cases, one or more affinity reagents labeled with a nucleic acid barcode. The affinity reagent may be crosslinked to the bound protein. Once crosslinked to the substrate, the affinity reagent is then subjected to barco analysis to determine the identity of the crosslinked affinity reagent. In some cases, multiple bound proteins may be sequenced. The antibody may be exposed to one or more types of affinity reagent. In some cases, multiple bound antibodies may be used. When the bound protein is crosslinked with one or more affinity reagents, the bound affinity reagents A drug-associated barcode is associated with each of the multiple bound proteins. The resulting cross-linked affinity reagent may be sequenced to determine its identity.
[0071] A family of affinity reagents may include one or more types of affinity reagents. For example, the disclosed methods include the use of antibodies, antibody fragments, Fab fragments, aptamers, peptides, and proteins. A family of affinity reagents comprising one or more of the proteins may be used.
[0072] The affinity reagent may be modified, including but not limited to, the attachment of a detection moiety. The detection moiety may be attached directly or indirectly. For example, the detection moiety The affinity reagent may be covalently attached directly to the affinity reagent or may be attached via a linker. or a complementary nucleic acid tag or a parent tag such as a biotin-streptavidin pair. The attachment may be via a compatibility reaction. Any suitable attachment method can be chosen.
[0073] Detection moieties include, but are not limited to, fluorophores, bioluminescent proteins, a nucleic acid segment comprising a mutation region and a barcode region, or a nanoparticle, such as a magnetic particle. The detection moiety comprises a chemical tether for linking the detection moiety to a different pattern of excitation or emission. The fluorescent substance may comprise several different fluorophores, having the following structure:
[0074] The detection moiety may be cleavable from the affinity reagent, which is no longer of interest. Reducing signal contamination by a process in which the detection moiety is removed from the affinity reagent that is not This may enable the following:
[0075] In some cases, the affinity reagent is unmodified. For example, if the affinity reagent is an antibody If the affinity reagent is a marker, the presence of the antibody may be detected by atomic force microscopy. A modification, for example an antibody specific for one or more of the affinity reagents. For example, if the affinity reagent is a mouse antibody, If present, mouse antibodies may be detected using an anti-mouse secondary antibody. The reagent may be an aptamer that is detected by an antibody specific for the aptamer. The antibody may be modified with a detection moiety as described above. In some cases, a secondary antibody The presence of may be detected by atomic force microscopy.
[0076] In some instances, the affinity reagent may have the same modification, e.g., a conjugated green It may contain a fluorescent protein or may contain two or more different modifications. For example: Each affinity reagent contains several different fluorophores, each with a different excitation or emission wavelength. Several different affinity reagents may be combined. This may allow for multiplexing of affinity reagents, since multiple affinity reagents can be identified and / or differentiated. In one example, the first affinity reagent may be conjugated to green fluorescent protein; The second affinity reagent may be conjugated to a yellow fluorescent protein, and the third affinity reagent Drugs may be conjugated to red fluorescent proteins, thus allowing the identification of these three affinity assays. The drugs can be multiplexed and identified by their fluorescence. The second, fourth, and seventh affinity reagents may be conjugated to green fluorescent protein, and the third, fourth, and seventh affinity reagents may be conjugated to green fluorescent protein. the fifth, fifth, and eighth affinity reagents may be conjugated to a yellow fluorescent protein; and The third, sixth, and ninth affinity reagents may be conjugated to a red fluorescent protein; In this case, the first, second, and third affinity reagents may be multiplexed together, while the second, fourth, and the seventh, as well as the third, sixth, and ninth affinity reagents, are then subjected to two further multiplex reactions. The number of affinity reagents that can be multiplexed together is used to distinguish between them. For example, the detection moiety may vary depending on the fluorophore-labeled affinity reagent. Multiplexing may be limited by the number of unique fluorophores available. For this purpose, multiplexing of affinity reagents labeled with nucleic acid tags is determined by the length of the nucleic acid barcode. It is okay to do so.
[0077] The specificity of each affinity reagent can be determined prior to use in an assay. The binding specificity can be determined in control experiments using known proteins. Experimental methods may be used to determine the specificity of the affinity reagent. Known protein standards are loaded at known locations to assess the specificity of multiple affinity reagents. In another example, the specificity of each affinity reagent may be determined by binding to a control and a standard. The substrate is a sample of the experimental specimen, so that the experimental sample ID can be calculated from the The panel may include both samples and a panel of controls and standards. Affinity reagents of known specificity may be included along with affinity reagents of known specificity, Data from affinity reagents of known specificity may be used to identify proteins, and the pattern of binding of affinity reagents of unknown specificity to the proteins to be identified. The binding specificity of each of the affinity reagents may be determined by determining which proteins bind to each of the affinity reagents. Any individual binding data can be used to assess whether the binding was consistent with the known binding data of other affinity reagents. It is also possible to reconfirm the specificity of the affinity reagent. By multiple uses of the same panel, the specificity of the affinity reagent can be increasingly refined with each iteration. Although affinity reagents uniquely specific for a particular protein may be used, The methods described herein may not require them. In addition, the methods may require a range of specific In some instances, the methods described herein may be effective in treating is that the affinity reagent is not specific for any particular protein, but instead Particularly useful when specific for an amino acid motif (e.g., the tripeptide AAA) It could be.
[0078] In some examples, one or more affinity reagents are of a given length, e.g., 2, 3, 4 , 5, 6, 7, 8, 9, 10 amino acids, or more than 10 amino acids. In some examples, the one or more affinity reagents may be selected to be bispecific. The nucleotide sequences were selected to bind to amino acid motifs ranging in length from 1 amino acid to 40 amino acids. It's okay to be found out.
[0079] In some instances, the affinity reagent has high, intermediate, or low binding affinity. In some cases, a parent having low or intermediate binding affinity may be selected. In some cases, the affinity reagent may be selected to be about 10 -3 M, 10 -4 M, 10 -5 M, 10 -6 M, 10 -7 M, 10 -8 M, 10 -9 M, 10 -10 Dissociation constant of M or less In some cases, the affinity reagent may have a molecular weight of about 10 -10 M, 10 -9 M, 10 -8 M, 10 - 7 M, 10 -6 M, 10 -5 M, 10 -4 M, 10 -3 M, 10 -2 Dissociation constants greater than or equal to M The number may be any number.
[0080] Some of the affinity reagents contain amino acid sequences that are phosphorylated or ubiquitinated. In some examples, the 1 The species or species of affinity reagent may be selected from the group consisting of epitopes that may be carried by one or more proteins. The nucleotide sequence may be chosen to be broadly specific for a family of topes. Thus, one or more affinity reagents may bind to two or more different proteins. In some instances, one or more affinity reagents are selectively selective for their target(s). For example, the affinity reagent may bind less than 10%, less than 10%, less than 15%, less than 20%, less than 25%. Less than, less than 30%, less than 35%, or less than 35% may bind to their target(s). In some instances, the one or more affinity reagents are capable of binding to one or more of their targets. For example, the affinity reagent may bind more than 35%, more than 40%, more than 45%, more than 6 More than 0%, More than 65%, More than 70%, More than 75%, More than 80%, More than 85%, More than 90%, More than 91%, More than 92%, More than 93%, More than 94%, 95% greater than 96%, greater than 97%, greater than 98%, or greater than 99% of the IgG antibodies can bind to one or more of their targets. .
[0081] To compensate for weak binding, an excess of affinity reagent may be applied to the substrate. Approximately 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1 or 10:1 to the sample protein The affinity reagent may be applied in excess of about 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1 or 10:1, which is an excess over the expected occurrence of the epitope in the sample protein. may be applied excessively.
[0082] The affinity reagent may also include a magnetic component. To manipulate some or all binding affinity reagents in the same image plane or z-stack It may be useful to operate some or all of the affinity reagents in the same image plane. , which may improve the quality of the imaging data and reduce noise in the system.
[0083] binding measurements Given a set of modified affinity reagents and a conjugated substrate, affinity The reactive reagent may be repeatedly applied to the substrate. Each measurement cycle consists of several steps. In a first step, an affinity reagent is applied to a substrate, where the affinity reagent binds to a conjugate. The carboxylate may be adsorbed to the jugate protein.
[0084] The substrate may then be washed briefly to remove non-specific binding. This process may be carried out under conditions that do not elute the affinity reagent bound to the immobilized protein. Some examples of buffers that can be used in the process are phosphate buffered saline, Tris buffered saline Water, phosphate-buffered saline with Tween 20, and Tris-buffered saline with Tween 20 Contains saline.
[0085] Following adsorption, the binding address of each modified affinity reagent is conjugated directly to the affinity reagent. Measurement of fluorophores that are conjugated to affinity reagents or For example, by measuring a fluorophore conjugated to a complementary nucleic acid to a nucleic acid strand that is The detection method is determined by the choice of detection moiety. and bioluminescent moieties can be detected optically, and in some cases, a secondary detection reagent is used. The unique address of each protein immobilized on the substrate must be determined prior to binding measurements. A list of addresses containing immobilized proteins may be determined and used to measure binding. It may be generated accordingly.
[0086] The affinity reagent can then be desorbed by a more stringent wash. The process may remove some or all of the affinity reagent from the immobilization substrate. In some cases, affinity reagents are designed to have low to intermediate binding affinity to facilitate removal. The affinity reagents used may be recaptured for reuse or may be selected to be By way of example, affinity reagents having cleavable detection moieties may be used. If so, the detection moiety may be cleaved and removed at this stage. Following thorough washing, in some instances, any remaining fluorescence can be quenched, further Stringent washes may be applied to remove residual affinity reagents. / Contamination can be detected by re-imaging the substrate before applying the next affinity reagent Contamination can also be prevented by continuously monitoring the images for recurring signals. This completes one cycle of the analysis.
[0087] In some embodiments, the fluorescently tagged affinity reagent is responsive to intense light at an activation wavelength. Upon prolonged exposure, the fluorescent tag may be quenched. Quenching of the fluorescent tag is a step to remove the affinity reagent. In some embodiments, the signal is determined based on the n-1 cycles of the previous signal. Repeating n fluorophores to identify which were derived from the clone may be desirable.
[0088] The cycle continues for each affinity reagent, or multiplex thereof. The results list the binding coordinates of each affinity reagent, or the affinity reagent bound at each coordinate position. This is a very large table, see for example Figure 10.
[0089] analysis The final step in protein identification is to identify the substrate from the information about the affinity reagents bound to the coordinates. To determine the most likely identity of each protein at each coordinate in The software may include a software tool for determining the binding characteristics of each affinity reagent. For example, if a given affinity reagent is Assume that the affinity reagent selectively binds to proteins containing A. The binding characteristics of each affinity reagent and the target protein in the sample are A database of proteins, a list of their binding coordinates, and information on their binding patterns Given a,software tool can provide a probable,verification of the identity as well as the reliability of,the identity. Assign an identity to each coordinate. Precise 1-1 match between affinity reagent and protein. In the extreme case of mapping, this can be accomplished with a simple lookup table. However, when the combination is more complex, this can be achieved by solving an appropriate satisfaction problem. In cases where the binding characteristics are highly complex, an expectation maximization approach may be employed. That's fine.
[0090] The software also identifies the number of positions at which each affinity reagent did not bind, either in some or all positions. The list can be used to determine the proteins present, for the absence of epitopes. The software can also use this information to determine whether an affinity reagent has bound to each address. , and information about which did not bind. Both the presence and absence of epitopes The software may include a database. The database may include: It may contain sequences of some or all of the known proteins in the species from which the sample was obtained. For example, if the sample is known to be of human origin, some For example, a database with the sequences of all human proteins may be used. If known, some or all of the protein sequence databases are used. The database also contains some or all of the known protein variants and mutations. The sequences of the somatic proteins, as well as some or all of the possible mutations that may result from DNA frameshift mutations The database may also include sequences of all possible proteins. It may include sequences of possible truncated proteins that may result from the donor or from degradation.
[0091] The software covers machine learning, deep learning, statistical learning, supervised learning, unsupervised learning, and cross-platform computing. Rastering, Expectation Maximization, Maximum Likelihood Estimation, Bayesian Inference, Linear Regression, Logistic Regression, One or more algorithms, such as binary classification, multiclass classification, or other pattern recognition algorithms For example, the software may include an algorithm for (i) analyzing information about the binding characteristics of each affinity reagent; (ii) information on a database of proteins in the sample; (iii) information on a list of binding coordinates; and and / or (iv) analyzing information about the pattern of binding of the affinity reagent to the protein (e.g., (a) one or more algorithms that analyze the (b) the probable identity and / or authenticity of the coordinates (e.g. identity Confidence levels and / or confidence intervals for one or more algorithms Machine learning algorithms may be implemented for the purpose of generating or assigning Examples of algorithms are support vector machines (SVM), neural networks, and convolutional neural networks. Neural Network (CNN), Deep Neural Network, Cascade Neural Network networks, k-nearest neighbor (k-NN) classification, random forests (RF), and other types of classification trees and regression trees (CART).
[0092] The software predetermines the identity of the protein at each address. The method may be performed on a substrate that has been trained using the disclosed method. The software is a nucleic acid-programmable protein array (Nucleic Acid-Programmable Protein Array). n Array) or epitope tiling array as a training dataset. It is okay to do so.
[0093] Characterization of samples Once the decoding is complete, the probability of the protein conjugated to each address is Their abundance in the mixture is then determined by the counting Thus, a list of each protein present in the mixture, and and the number of observations of that protein can be collected.
[0094] Additionally, photocleavable linkers, or other types of specifically cleavable linkers, can be attached to the protein. If a protein is used to attach to a substrate, the particular protein of interest can then be attached to the substrate. The protein may be removed from the substrate and collected for further investigation. The disclosed method also provides for the isolation of desired compounds from a mixture. In some cases, the method may serve as a means to purify and / or isolate novel proteins. In the method, a specific isotype or post-translationally modified protein is purified and It may be possible to isolate and / or isolate potential proteins and associated In samples where a complete list of sequences is not available, the method It may be possible to distinguish different proteins within a protein group, which can then be further For example, gut microbiome samples can be extracted for further investigation. For highly complex samples containing many unknown proteins, the methods described herein can be used The methods described above may be used to fractionate samples prior to mass spectrometry. Thus, once their identity has been determined, proteins are eluted from the substrate. Once proteins have been identified, their identity can be confirmed by removing them from the substrate. Subsequent rounds of affinity reagent binding for proteins whose identity has not yet been determined This allows the player to continue playing the game while being aware of background noise and the remaining rounds. In some cases, specificity for a particular protein may be reduced. One or more affinity reagents having the formula Used as a first round to identify abundant proteins such as purine These abundant proteins may be removed early in the process. In this case, a subset of proteins on the substrate are bound after each round of affinity reagent binding. or every 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20 rounds of affinity reagent binding, or may be removed after every 20 or more rounds. The signal to noise ratio is It may increase after each round.
[0095] In some cases, unidentified proteins were classified into groups based on their binding patterns. For example, in some cases, An existing protein may not be represented in the sequence database. The proteins are grouped into groups, each of which contains a set of unknown proteins with the same sequence in the sample. To achieve this goal, the researchers classified the analytes into groups based on their binding patterns to affinity probes. The amount of protein may be estimated for each group and the limiting Although not intended to be used for differential quantification between healthy and disease states, longitudinal analysis, or may be included in quantitative analyses, including biomarker discovery. In some cases, unidentified The groups may be selectively removed from the substrate for identification by mass spectrometry. In this study, the unidentified groups were specifically designed to generate reliable identifications. This may be identified by performing further binding affinity measurement experiments.
[0096] In some cases, after a protein or set of proteins is removed, the substrate It may be possible to add additional sample to the material. For example, serum albumin may be added to the sample. It is a protein that is abundant in serum and accounts for approximately half of the total protein in the serum. Removing serum albumin after the first round of binding allows for the addition of additional blood sample to the substrate. In some embodiments, prior to immobilizing the sample on the substrate, Removal of abundant proteins, for example via immunoprecipitation or affinity column purification It may be chosen.
[0097] Protein modifications can be identified using the methods of the present disclosure. For example, post-translational modifications can be identified by enzyme Detection using detection reagents specific for the modification, incorporating treatment (e.g., phosphatase treatment) By repeated cycles of extraction, affinity reagents specific for different modifications may be identified. To determine the presence or absence of such modifications in the modified protein, The method may also include determining the number of instances in which each protein has and does not have a given modification. Allows for quantification.
[0098] Mutations in proteins are important for determining the binding pattern of the sample protein and the predicted protein identity. This may be detected by checking for discrepancies between the identity of Matches the affinity reagent binding profile of a known protein except for affinity reagent binding of The protein or polypeptide immobilized on the substrate may have one amino acid substitution. Since affinity reagents may have overlapping epitopes, the immobilized proteins may be 1-amino Even with only acid substitutions, there are some inconsistencies with the predicted affinity binding pattern. DNA mutations that cause a frameshift of a premature stop codon can also be detected.
[0099] The number of affinity reagents required may be less than the total number of epitopes present in the sample. For example, suppose affinity reagents are designed such that each affinity reagent recognizes one unique three-peptide epitope. When a drug is selected, it must have an affinity to recognize all possible epitopes in the sample. The total set of reagents is 20 x 20 x 20 = 8000 types. However, the method of the present disclosure of affinity reagent, approximately 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000 , 2500, 3000, 3500, 4000, 4500, 5000, 5500, or 6000 species may be required. In some cases, the method comprises: 13 shows the results of each affinity reagent. Given a set of x affinity reagents specific for a unique amino acid 3-mer as a function of binding ability, A simulation demonstrating the percentage of known human proteins that can be identified if As shown in FIG. 13, 98% of human proteins are composed of 8000 3-mer parent proteins. With compatibility reagents and a binding probability of 10%, they can be uniquely identified.
[0100] The disclosed methods can be highly accurate. The disclosed methods have been shown to be accurate to approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 11 %, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.9%, or greater than 99.9% accuracy It may be possible to identify the protein.
[0101] The methods of the present disclosure provide a method for preventing or reducing the incidence of pulmonary circulation disorders, including ... Predict the identity of each protein with 5%, 99.9%, or >99.9% confidence. The confidence level may be different for different proteins in the sample. For example, proteins with highly unique sequences are those that are highly similar to other proteins. Proteins can be identified with greater confidence than previously identified proteins. A protein can be identified with a high degree of confidence as a member of a protein family. However, the exact identity of the protein is less certain. In some cases, extremely large or extremely small proteins may be predicted with less confidence than more moderately sized proteins.
[0102] In some cases, proteins are highly expressed as members of a protein family. The exact identity of the protein can be identified with little confidence, however. may be predicted with less confidence. For example, Proteins are difficult to distinguish reliably from standard forms of the protein. In this case, neither the standard sequence nor the version containing the single amino acid mutation may have high reliability. However, the unknown protein is a member of a protein group that contains both sequences. A high degree of confidence can be assessed that similar situations exist where the protein is similar. In instances where a sequence may have multiple related isoforms, this may occur.
[0103] The disclosed methods can identify some or all of the proteins in a given sample. The disclosed method may be capable of detecting approximately 90%, 91%, 92%, 93%, 94%, or 100% of the proteins in a sample. %, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.9%, or greater than 99.9% can be identified It may be possible.
[0104] The methods of the present disclosure may be capable of rapidly identifying proteins in a sample. The methods shown can produce greater than about 100, greater than about 1000, greater than about 5000, or greater than about 1000 per flow cell per day. More than 0,000 pieces, more than about 20,000 pieces, more than about 30,000 pieces, more than about 40,000 pieces, more than about 50,000 pieces, more than about 100,000 pieces , over 1,000,000 pieces, about 10,000,000 pieces, about 100,000,000 pieces, about 1,000,000,000 pieces, about 10, More than 000,000,000 proteins, more than about 100,000,000,000 proteins, and more than about 1,000,000,000,000 proteins. The disclosed method can achieve approximately 1000 cells / day per flow cell. 10 Over 10 pieces 11 Over 10 pieces 12 Over 10 pieces 13 Over 10 pieces 14 Over 10 pieces 15 Over 10 pieces 16 Over 10 pieces 17 More than one, is about 10 17 It may be possible to identify more than one protein per day. Approximately 10 per flow cell 10 ~10 12 pieces, 10 11 ~10 14 pieces, 10 12 ~10 16 Pieces or 10 13 ~1 0 17 The disclosed method may be capable of identifying 100 proteins per day. Approximately 10 pg, 20 pg, 30 pg, 40 pg, 50 pg, 60 pg, 70 pg, About 80 pg, about 90 pg, about 100 pg, about 300 pg, about 300 pg, about 400 pg, about 500 pg, about 600 pg, about 700 pg, about 800 pg, about 900 pg, about 1 ng, about 2 ng, about 3 ng, about 4 ng, about 5 ng, about 6 ng, about 7 ng, about 8 ng, about 8 ng, about 10 ng, about 10 ng, about 20 ng, about 30 ng, about 40 ng, about 50 ng, about 60 n g, about 70 ng, about 80 ng, about 90 ng, about 100 ng, about 300 ng, about 300 ng, about 400 ng, about 500 ng, Approximately 600 ng, approximately 700 ng, approximately 800 ng, approximately 900 ng, approximately 1 μg, approximately 2 μg, approximately 3 μg, approximately 4 μg, approximately 5 μg, approx. 6 μg, approx. 7 μg, approx. 8 μg, approx. 8 μg, approx. 10 μg, approx. 10 μg, approx. 20 μg, approx. 30 μg , about 40 μg, about 50 μg, about 60 μg, about 70 μg, about 80 μg, about 90 μg, about 100 μg, about 300 μg, approx. 300 μg, approx. 400 μg, approx. 500 μg, approx. 600 μg, approx. 700 μg, approx. 800 μg, approx. 900 μg It is possible to identify more than 95% of proteins in approximately 1 mg of protein or more than 1 mg of protein. obtain.
[0105] The methods of the present disclosure can be used to evaluate the proteome following an experimental treatment. The methods described can be used to assess the efficacy of a therapeutic intervention.
[0106] The methods of the present disclosure can be used for biomarker discovery. By monitoring proteome expression in disease- and non-disease-affected subjects, biomarkers can be identified. The present invention can identify cancer cells in subjects prior to onset of disease or at risk of developing disease. By monitoring proteome expression in subjects with Markers can be identified. By assessing the proteomic expression of a subject, the health status of the subject can be determined. The disclosed methods may indicate a patient's risk of developing a disease or disorder. used to assess the risk of developing a disease or to distinguish responders from non-responders to a drug / therapy The methods of the present disclosure may be of particular use for personalized medicine.
[0107] The methods of the present disclosure can be used to diagnose diseases. Different diseases or disease stages can be diagnosed by: The different panels of protein expression may be associated with , may be associated with different treatment outcomes for each given treatment. The data may be used to diagnose a subject and / or to select the most appropriate treatment. It can be used.
[0108] The methods of the present disclosure can be used to identify the individual or species from which a sample originates. For example, the disclosed methods can determine whether a sample is in fact from the species or source claimed. The methods described herein can be used to determine whether a protein is abundant or not. In samples poor in nucleic acids, this may have advantages over PCR-based methods. Identification of the origin of a honey sample. As a further example, the methods of the present disclosure can be used to verify food safety and food It can be used to evaluate quality control.
[0109] The disclosed method uses affinity reagents that are fewer than the number of possible proteins to detect proteins. It can be used to identify any single protein molecule from a pool of protein molecules. For example, the method may involve using a panel of affinity reagents to generate a pool of n potential proteins. may be identified with a certainty above a threshold, The number of affinity reagents in the panel is m, and m is less than n. may be known proteins corresponding to known protein and gene sequences, or It may be an unknown protein with no known protein or gene sequence. In the case of proteins, the method can identify unknown protein signatures and thus The presently disclosed method can identify the presence and amount of an unknown protein, but not its amino acid sequence. , identifying an unidentified protein selected from a pool of n possible proteins. The method may be used to select a panel of m affinity reagents that can be used to identify the m affinity reagents. The disclosed method also involves using m binding reagents to bind n proteins in a mixture of proteins. Proteins can be uniquely identified and quantified, where each protein is a member of m species. The binding agents are identified by their unique profiles of binding by a subset of the binding agents. is less than approximately one-half, one-third, one-quarter, one-fifth, one-sixth, or one-seventh of n , less than a tenth, less than a twentieth, less than a fiftieth, or less than a hundredth. As an example, the present disclosure provides a panel of affinity reagents that is at least about 100, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 12,000, 14,000, 16,000, 18,000 , 20,000, 25,000, 30,000, 35,000, 40,000, 50,000, 100,000, 200,000, 300,000, 400 ,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 2,000,000, 3,000,0 Uniquely identify each of 0, 4,000,000, or 5,000,000 different proteins In order to make it possible to do so, the following categories are available: less than approximately 100 species, less than 200 species, less than 300 species, less than 400 species, less than 500 species, less than 600 species Full, Less than 700 species, Less than 800 species, Less than 900 species, Less than 1000 species, Less than 1500 species, Less than 2000 species, 2500 species To select a panel of affinity reagents with less than 3000, less than 3500, or less than 4000 species may be used for.
[0110] The methods of the present disclosure may be capable of identifying a large proportion of proteins in a proteome. The disclosed methods may be used to detect and treat mammalian, avian, fish, amphibian, reptile, vertebrate, and insect species. of proteins in the vertebrate, plant, fungal, bacterial, or archaeal proteomes The disclosed methods may be capable of identifying the majority of proteins in a proteome. of approximately more than 5%, more than 10%, more than 15%, more than 20%, more than 25%, more than 30%, more than 35%, more than 40%, more than 45%, more than 50%, more than 55% , over 60%, over 65%, over 70%, over 75%, over 80%, over 85%, over 90%, over 95%, over 96%, over 97%, over 98%, 9 It may be possible to identify more than 9%. EXAMPLES
[0111] Example 1: Identification of proteins using antibodies that bind to unique 3-mer peptides the percentage of coverage of the set of all epitopes in the proteome, and Determine the relationship between the percentage of the proteome that can be identified using the methods of the present disclosure To investigate this, we carried out a computational experiment. A set of amino acid epitopes was selected. Protein modifications were not considered. 20 naturally occurring Since there are 20 amino acids, the total set of all 3-mer epitopes is 20 x 20 x 20 = 8000 For the simulation, x is the number of possible epitopes of a species screened in one experiment. The number of epitopes to be grouped is set as the number of epitopes to be grouped, and for each value of x from 1 to 8000, A set of topes was randomly selected and the percentage of the proteome that could be identified was calculated. Figure 13 shows the results of this simulation.
[0112] Example 2: Identification of proteins using antibodies that bind to unique 3-mer peptides To determine the impact of the number of affinity reagents on identifiability and coverage, Further experiments were carried out using a computer. For this, affinity is plotted to indicate the percentage of the proteome that can be identified (y-axis). Data series were calculated for a range of reagent pool sizes and the results are shown in Table 1. For example, a 100 amino acid protein has 98 3-mer amino acid epitope "landing sites". If 20% of these 3-mer amino acid epitopes are bound, it is considered that the protein This may or may not be sufficient to identify the quality. Then, a pool of 250 3-mer specific affinity reagents was used to tentatively determine the target site of each protein. If 20% of the points are bound, only about 7% of the proteome can be identified. For an affinity reagent pool, binding of 20% of the landing sites results in approximately 98% of the proteome being bound. can be identified.
[0113] Table 1. Identifiability vs. proteome coverage of 3-mer d-coded probes The power of numbers TIFF2024170480000002.tif98163
[0114] Example 3: Photoprotein molecules conjugated onto a substrate Phycoerythrin, a fluorescent protein sample, was incubated in an incubation chamber for 4 h. Direct conjugation to NHS-ester coated coverslips at 37 °C for 4 h. The fluorescent protein samples were then imaged on a Leica DMi8 equipped with a Hamamatsu orca flash 4.0 camera. The resulting capture images were taken at 100 nm and 100 ms exposure time. 16A and 16B show the capture images (colors inverted for clarity). Each dark spot represents an area of fluorescent signal indicating the presence of a protein. An enlarged view of FIG. 16A. The arrows in FIG. 16B indicate a clearly distinguishable marker from the background noise. A, signals representing proteins are shown.
[0115] The second protein sample, green fluorescent protein, was denatured and incubated for 1 h. Conjugates were then directly attached to NHS-ester coated coverslips in a 50-well plate at 4°C for 4 hours. The first image showed no baseline residual fluorescence, which was due to the presence of green fluorescent protein. The protein was then denatured with an anti-Alexa-Fluor 647 antibody. The anti-peptide antibody was then rinsed with 0.1% Tween-20. This was then imaged using TIRF on a Nikon Eclipse Ti equipped with an Andor NEO sCMOS camera. FIG. 17 shows the resulting captured image (colors inverted for clarity). did).
[0116] Example 4: Protein Identification Four potential proteins: green fluorescent protein, RNASE1, LTF, and GSTM1. The proteome of the quality is represented in Figure 18. In this example, unknowns from this proteome A single protein molecule is conjugated to a certain location on the substrate. , which are then probed with a panel of nine different affinity reagents. Each recognizes different amino acid trimers [AAA, AAC, AAD, AEV, GDG, QSA, LAD, TRK, DGD]. The unknown protein is labeled with the affinity reagent DGD It is determined that the four types of proteome bind to AEV, LAD, GDG, and QSA. Analysis of the protein sequence revealed that only GFP contains all five of these three amino acid motifs. These motifs are underlined in the sequence of Figure 18. Therefore, the unknown protein molecule is determined to be the GFP protein.
[0117] Computer Control System The present disclosure also relates to a computer controlled system programmed to carry out the methods of the present disclosure. FIG. 14 provides a program for characterizing and identifying biopolymers such as proteins. 1 shows a computer system 1401 that has been programmed or otherwise configured. The computer system 1401 may, for example, observe signals at unique spatial addresses of the substrate. Step 2. Identifying a portion of the biopolymer at a unique spatial address based on the observed signal. determining the presence of an identifiable tag linked to the molecule; determining a characteristic of the portion of the biopolymer To identify the identifiable tags, the determined tags are compared to a database of sequences of biopolymers. Various aspects of the present disclosure may be governed by the steps of evaluating and analyzing a sample, such as the step of evaluating. The computer system 1401 may be located on the user's electronic device or remotely to the electronic device. The electronic device may be a computer system installed in a computer system. The electronic device may be a portable electronic device.
[0118] The computer system 1401 includes a central processing unit (CPU, also referred to as a “processor” in this specification). "Computer Processor" 1405, which may be a single core or It may be a multi-core processor, or multiple processors for parallel processing. The computer system 1401 also includes a memory or memory location 1410 (e.g., a random access memory, read-only memory, flash memory), electronic storage unit unit 1415 (e.g., a hard disk) for communication with one or more other systems. A communication interface 1420 (e.g., a network adapter), as well as a cache , other memory, data storage and / or electronic display adapters, etc. The peripheral device 1425 includes a memory 1410, a storage unit 1415, an interface 1420, The peripheral device 1425 communicates with the CPU 1405 via a communication bus (solid line) on the motherboard or the like. The storage unit 1415 is a data storage unit for storing data. (or a data repository). The computer system 1401 may include a communication interface. With the assistance of the interface 1420, a computer network ("Network") 1430 The network 1430 may be the Internet, an internet and / or an extranet, or It may be an intranet and / or an extranet that communicates with the Internet The network 1430 may, in some cases, be a telecommunications and / or data network. The network 1430 is a distributed computing network such as cloud computing. The network may include one or more computer servers that may enable The computer 1430 may, in some cases, be assisted by the computer system 1401. The devices connected to the computer system 1401 may function as clients or servers. A peer-to-peer network may be implemented that may enable
[0119] The CPU 1405 may be embodied in a program or software form, and may be a machine-readable The instructions may be stored in a memory location, such as memory 1410. The instructions may be directed to the CPU 1405, which then executes the program. The CPU 1405 may be programmed or otherwise configured to perform the illustrated methods. Examples of operations performed by the CPU 1405 include fetch, decode, execute, and writeback. may include.
[0120] The CPU 1405 may be part of a circuit such as an integrated circuit. Other components may be included in the circuit. In some cases, the circuit is application specific. It is an ASIC (Application Specific Integrated Circuit).
[0121] The storage unit 1415 stores drivers, libraries, and stored programs. The storage unit 1415 can store files such as user preferences. It is possible to store user data such as references and user programs. The computer system 1401 may, in some cases, be connected to an intranet or the Internet. to a remote server that communicates with computer system 1401 via the Internet. One or more additional devices external to the computer system 1401, such as those located The device may include a data storage unit.
[0122] The computer system 1401 communicates with one or more remote For example, computer system 1401 can be , can communicate with a user's remote computer system. Examples of systems are personal computers (e.g., handheld PCs), slates or tablets. Tablet PCs (e.g. Apple iPad, Samsung Galaxy Tab), mobile phones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®, or personal digital assistant. , the computer system 1401 is accessible via a network 1430 .
[0123] The methods as described herein may be implemented, for example, by storing, in memory 1410 or electronic storage. on an electronic storage location of the computer system 1401, such as on the storage unit 1415 by machine (e.g., computer processor) executable code stored in Machine executable or machine readable code may be called software. During use, the code may be executed by the processor 1405. In some cases, the code may be retrieved from the storage unit 1415 and processed. The data may be stored in memory 1410 for quick access by the processor 1405. In some cases, the electronic storage unit 1415 may be omitted and the machine-executable instructions may be stored in memory. -1410 is stored.
[0124] The code may be precompiled and a processor adapted to execute the code may be It can be configured for use on a machine with The code can be precompiled or compiled at run time (as-compiled ) in a programming language that can be selected to enable execution of the code in can be provided.
[0125] The systems and methods provided herein, such as computer system 501 Aspects of technology can be embodied in the form of programming. A program that is carried or embodied in some type of readable medium, typically a machine (or processor) "Articles" in the form of executable code and / or associated data on a processor Machine executable code may be considered as a program or an "article of manufacture." (e.g. read-only memory, random access memory, flash memory) or hardware The software may be stored on an electronic storage unit, such as a hard disk. The media for this software may be any tangible memory or Various semiconductor memories, tape drives, disk drives and the like and any or all of its associated modules, which may be implemented as software. You can provide non-transient storage at any time during programming. All of the software or any part thereof, the Internet or various other telecommunications networks Such communication may occur, for example, between one computer and from one computer or processor to another, for example, a management server or host From your computer to the application server's computing platform, It may therefore be possible to carry software elements. Another type of media that can be used is one that passes through a physical interface between local devices. those used over wired and optical landline networks, and Includes light waves, radio waves, and electromagnetic waves, such as those used over various air links. A means for carrying such waves, such as a wired or wireless link, optical link, or the like. A physical element may also be considered as a medium for carrying the software. As used herein, unless limited to non-transitory, tangible "storage" media, The terms "computer-readable medium" and "machine-readable medium" refer to any medium that is accessible to a processor for execution. Refers to any medium involved in providing instructions.
[0126] Therefore, machine-readable media, such as computer executable code, Including, but not limited to, tangible storage media, carrier wave media, or physical transmission media Non-volatile storage media may take many forms, including, for example, any computer Any storage device in a computer or the like, or the like, as shown in the drawings including optical or magnetic disks, such as those that may be used to implement databases, etc. Volatile storage media include the main memory of a computer platform. The tangible transmission medium is the bus in the computer system. Coaxial cable; copper wire; and optical fiber. , in the form of electrical or electromagnetic signals, or at radio frequency (RF) and infrared (IR) It may take the form of sound or light waves, such as those generated during data communications. Common forms of computer-readable media thus include, for example: floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic Media, CD-ROM, DVD or DVD-ROM, any other optical media, punch cards, paper tape, perforated Any other physical storage medium that has a pattern of RAM, ROM, PROM and EPROM, F LASH-EPROM, any other memory chip or cartridge, carrying data or instructions A carrier wave carrying such a carrier wave, or a cable or link carrying such a carrier wave, or a computer-readable Any other medium from which programming code and / or data may be embedded. Many of these forms of computer readable media are written to a processor for execution. The instruction sequence may be responsible for conveying one or more sequences of one or more instructions.
[0127] The computer system 1401 includes an electronic display that includes a user interface (UI) 1440. Examples of UIs include, but are not limited to, a display 1435. , Graphical User Interface (GUI) and Web-based User Interface Includes the interface.
[0128] The methods and systems of the present disclosure may be performed by one or more algorithms. The algorithm is performed by the software when executed by the central processing unit 1405. The algorithm obtains a characteristic of a portion of a biopolymer, e.g., a portion of a protein. and / or identity. For example, the algorithm may The most likely identity of a portion of a candidate biopolymer, such as a portion of a complementary protein, This may be used to determine the identity of the
[0129] In some embodiments, the compound recognizes a short epitope that is present in many different proteins. Aptamers or peptamers that achieve this are called digital aptamers or digital peptamers. One aspect of the present invention is the detection of digital aptamers or digital peptamers. and a set of at least about 15 digital aptamers or digital peptamers, each of the 15 digital aptamers or digital peptamers , specifically bind to different epitopes consisting of 3, 4, or 5 consecutive amino acids. Each digital aptamer or digital peptamer is characterized by Multiple distinct peptides containing the same epitope that is bound by the tal aptamer or digital peptamer In some embodiments, the digital aptamer or digital The set of digital peptamers consists of 100 peptides that bind to epitopes consisting of three consecutive amino acids. In some embodiments, the digital aptamer or digital peptamer is A set of tal aptamers or digital peptamers consists of four consecutive amino acids. In some embodiments, the method further comprises the step of: A set of digital aptamers or digital peptamers consists of five consecutive amino acids. We have also developed 100 digital aptamers or digital peptamers that bind to the epitopes of interest. In some cases, the digital affinity reagent includes an antibody, an aptamer, a peptamer, The antibody may be a polypeptide, a peptide, or a Fab fragment.
[0130] In some embodiments, the set of digital aptamers comprises at least about 20, 30, 40, 50 , 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 types of data In some embodiments, the set of digital aptamers comprises a set of linked At least 1000 digital aptamers that bind to epitopes consisting of four consecutive amino acids In some embodiments, the set of digital aptamers comprises five consecutive amino acids. The method further includes at least 100 digital aptamers that bind to epitopes consisting of acids. The set of digital aptamers binds to epitopes consisting of three consecutive amino acids. In some embodiments, the digital aptamer further comprises at least 100 digital aptamers. In some embodiments, the surface is an array. It is.
[0131] In another aspect, the present invention provides a method for determining the protein binding profiles of a sample containing a plurality of different proteins. The present invention provides a method for generating a profile, the method comprising the steps of: contacting the sample with a set of digital aptamers under conditions that The set of digital aptamers includes at least about 15 digital aptamers, Each of the aptamers binds a different epitope consisting of 3, 4, or 5 consecutive amino acids. Each digital aptamer is characterized as specifically binding to the digital aptamer. Recognizes multiple distinct proteins that contain the same epitope that the tal aptamer binds optionally removing unbound protein; and detecting binding of the protein binding protein of the sample to the aptamer, The process where the profile is generated.
[0132] In some embodiments, the method further comprises: subjecting the sample to a digital aptamer under conditions that permit binding. treating the sample with a protein cleaving agent prior to step (a) of contacting the sample with the set of mers; Further includes:
[0133] In another aspect, the invention provides a method for detecting a protein comprising the steps of: detecting a protein comprising: The method includes a library of protein binding profiles for the detection of a target protein, the method comprising the steps of: contacting the sample with the set of digital aptamers under conditions that permit binding; wherein the set of digital aptamers comprises at least about 15 digital aptamers; Each of the 15 digital aptamers consists of 3, 4, or 5 consecutive amino acids. Each digital aptamer is characterized as specifically binding to a different epitope. The digital aptamer is a set of multiple distinct and different aptamers that contain the same epitope to which the digital aptamer binds. Optionally, removing unbound protein; The protein of the sample being tested is detected by detecting binding to the digital aptamer. generating a binding profile, whereby the protein binding profile is and repeating the above steps with at least two samples.
[0134] In some embodiments, the method further comprises: subjecting the sample to a digital aptamer under conditions that permit binding. and further comprising treating the sample with a protein cleaving agent prior to contacting the sample with the set of mers. include.
[0135] In another aspect, the invention includes a method for characterizing a test sample, the method comprising: The method includes the steps of: coupling a test sample to a set of digital aptamers under conditions that allow binding. The step of contacting, wherein the set of digital aptamers is at least about 15 digital aptamers. Each of the 15 digital aptamers has 3, 4, or 5 consecutive characterized as specifically binding to distinct epitopes consisting of amino acids; and Each digital aptamer comprises a plurality of digital aptamers each having the same epitope to which the digital aptamer binds. and optionally removing unbound proteins. The process detects the binding between the protein and the digital aptamer, thereby determining whether the protein in the test sample is a protein or not. generating a protein binding profile; and characterizing the test sample by analyzing the test sample. The generated protein binding profile of the sample was compared with the protein binding profile of the reference sample. A process of comparing.
[0136] In another aspect, the present invention provides a method for determining the presence or absence of bacteria, viruses, or cells in a test sample. The present invention includes a method for binding a test reagent comprising the steps of: contacting the sample with a set of digital aptamers; The set includes at least about 15 digital aptamers, and the 15 digital aptamers Each of them specifically binds to a different epitope consisting of 3, 4, or 5 consecutive amino acids. Each digital aptamer is characterized in that it binds to recognizing multiple distinct proteins that contain the same epitope that overlaps with each other; a step of removing unbound proteins; and detecting binding between the proteins and the digital aptamer. generating a protein binding profile for the test sample by thereby generating a protein binding profile; and comparing the protein binding profile to a protein binding profile of a reference sample, The presence or absence of bacteria, viruses, or cells in the test sample is thereby determined by comparison. Process.
[0137] In another aspect, the invention includes a method for identifying a test protein in a sample, comprising: The method comprises the steps of: The suspected sample is subjected to a digital aptamer detection assay comprising at least about 15 digital aptamers. wherein each of the 15 digital aptamers is a contiguous set of three or characterized as specifically binding to distinct epitopes consisting of four or five amino acids. and each digital aptamer binds to the same epitope. and detecting a plurality of distinct proteins comprising the test protein and a nucleotide sequence; The identity of the test protein is determined by detecting binding to a set of target aptamers. determining the identity of at least about six digital aptamers to the test protein; and the presence of binding indicates the presence of at least about six epitopes in the test protein. The identity of at least about six epitopes identifies the test protein. A process used to.
[0138] Although preferred embodiments of the present invention are shown and described herein, such embodiments It will be apparent to those skilled in the art that the above-mentioned methods are provided as examples only. It is not intended to be limiting to the specific examples provided herein. The invention has been described with reference to the above specification, but the description and drawings of the embodiments herein are not intended to limit the scope of the invention. The above aspects should not be construed in a limiting sense. Numerous modifications, alterations, and substitutions are possible. It will now occur to those skilled in the art without departing from the invention. All aspects of the present invention are subject to a variety of conditions and variables, and may differ materially from the specific illustrations described herein. It should be understood that the present invention is not limited to the photographs, configurations, or relative proportions shown. Various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It should be understood that the present invention therefore also encompasses any such alternatives. It is intended to cover all such modifications, variations, and equivalents. The scope defines the scope of the invention and the methods and structures within the scope of these claims. , and their equivalents, are intended to be covered thereby.
[0139] Notwithstanding the claims that follow, the disclosures described herein also include As defined by the clause: 1. A set of digital aptamers comprising at least about 15 digital aptamers. Each of the 15 digital aptamers consists of 3, 4, or 5 consecutive amino acids. Each digital antigen is characterized as specifically binding to a different epitope. The digital aptamer is a set of multiple distinct, different aptamers that contain the same epitope to which the digital aptamer binds. A set of digital aptamers that recognize proteins. 2. Contains 100 digital aptamers that bind to epitopes consisting of three consecutive amino acids. , a set of digital aptamers according to clause 1. 3. 100 digital aptamers that bind to epitopes consisting of four consecutive amino acids were generated. A set of digital aptamers according to clause 1, comprising: 4. 100 digital aptamers that bind to epitopes consisting of five consecutive amino acids were generated. A set of digital aptamers according to clause 3, comprising: 5. At least about 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 70 A digital aptamer according to clause 1, containing 0, 800, 900 or 1000 digital aptamers. Tamar set. 6. At least 1000 digital applications that bind to epitopes consisting of four consecutive amino acids. A set of digital aptamers according to clause 1, including a set of aptamers. 7. At least 100 digital aptamers that bind to epitopes consisting of 5 consecutive amino acids The set of digital aptamers according to clause 6, further comprising: 8. At least 100 digital aptamers that bind to epitopes consisting of three consecutive amino acids The set of digital aptamers according to clause 7, further comprising: 9. A digital aptamer according to any one of clauses 1 to 8, in which the digital aptamer is immobilized on a surface. A set of luaptamers. 10. The set of digital aptamers according to clause 9, wherein the surface is an array. 11. A method for determining the protein binding profile of a sample containing a plurality of different proteins, comprising the steps of: Method to generate the file: a) contacting the sample with a set of digital aptamers under conditions that allow binding. wherein the set of digital aptamers comprises at least about 15 digital aptamers. Each of the 15 digital aptamers consists of 3, 4, or 5 consecutive amino acids. Each digital approach is characterized as specifically binding to a different epitope. The digital aptamer is a set of multiple distinct and different aptamers that contain the same epitope to which the digital aptamer binds. recognizing the protein; b) optionally removing unbound protein; and c) detecting the binding of a protein to the digital aptamer, generating a protein binding profile for the sample. 12. (a) contacting the sample with the set of digital aptamers under conditions that allow binding 12. The method of claim 11, further comprising the step of treating said sample with a protein cleaving agent prior to said step. 13. A method for detecting a plurality of proteins comprising the steps of: A method for generating a library of protein binding profiles comprising: a) contacting the sample with a set of digital aptamers under conditions that allow binding. wherein the set of digital aptamers comprises at least about 15 digital aptamers. Each of the 15 digital aptamers consists of 3, 4, or 5 consecutive amino acids. Each digital approach is characterized as specifically binding to a different epitope. The digital aptamer is a set of multiple distinct and different aptamers that contain the same epitope to which the digital aptamer binds. recognizing the protein; b) optionally removing unbound protein; c) The protein is tested by detecting the binding of the digital aptamer to the protein. generating a protein binding profile for the sample, whereby A binding profile is generated; and d) Repeating steps (a)-(c) with at least two samples. 14. (a) contacting the sample with the set of digital aptamers under conditions that allow binding 14. The method of clause 13, further comprising the step of treating said sample with a protein cleaving agent prior to said step. 15. A library of protein binding profiles prepared using the method of clause 13. 16. A method for characterizing a test sample, comprising the steps of: a) contacting a test sample with a set of digital aptamers under conditions that permit binding; The set of digital aptamers comprises at least about 15 digital aptamers. Each of the 15 digital aptamers is a sequence of 3, 4, or 5 consecutive amino acids. Each digital peptide is characterized as specifically binding to a different epitope consisting of: The aptamer is a set of multiple distinct, distinct antigens that contain the same epitope to which the digital aptamer binds. recognizing a protein that is b) optionally removing unbound protein; c) detecting the binding of a protein to the digital aptamer, thereby determining whether the test sample is a aptamer. generating a protein binding profile; and d) determining the protein binding profile of the test sample to characterize the test sample; comparing the protein binding profile to a reference sample. 17. A method for determining the presence or absence of bacteria, viruses, or cells in a test sample, comprising the steps of: How to do this: a) contacting a test sample with a set of digital aptamers under conditions that permit binding; The set of digital aptamers comprises at least about 15 digital aptamers. Each of the 15 digital aptamers is a sequence of 3, 4, or 5 consecutive amino acids. Each digital peptide is characterized as specifically binding to a different epitope consisting of: The aptamer is a set of multiple distinct, distinct antigens that contain the same epitope to which the digital aptamer binds. recognizing a protein that is b) optionally removing unbound protein; c) detecting the binding of a protein to the digital aptamer, thereby determining whether the test sample is a aptamer. generating a protein binding profile, whereby a file is generated; and d) the protein binding profile of the test sample and the protein binding profile of the reference sample thereby determining the presence of bacteria, viruses, or cells in the test sample. A process in which nothing is determined by comparison. 18. A method for identifying a test protein in a sample, comprising the steps of: a) A sample containing or suspected of containing the test protein is subjected to a small contacting the set of digital aptamers, the set including at least about 15 digital aptamers; Each of the 15 digital aptamers has 3 or 4 or 5 consecutive amino acids. Each nucleotide sequence has been characterized as specifically binding to a different epitope consisting of a nucleotide sequence. The digital aptamer is a multiple distinct aptamer that contains the same epitope that the digital aptamer binds. recognizing different proteins; and b) by detecting binding between a test protein and the set of digital aptamers; Determining the identity of the test protein, comprising: the presence of binding indicates that the aptamer binds to the test protein; and Both of these showed the presence of about six epitopes, and the identity of at least about six epitopes. is used to identify the test protein. 19. A method for determining a characteristic of a protein, comprising the steps of: Obtaining a substrate, each of the individual protein moieties (at a molecular level) being optically The method comprises: locating one or more proteins so that they have unique spatial addresses that can be resolved quantitatively. a moiety is conjugated to a substrate; A fluid comprising a first to an nth (ordered) set of one or more affinity reagents, applying to the substrate, each of the one or more affinity reagents comprising one or more One epitope (a continuous or non-contiguous amino acid sequence) of a portion of several proteins each affinity reagent of the first to nth sets of one or more affinity reagents specific to is linked to an identifiable tag; Each application of one or more affinity reagents to the substrate of the first and subsequent sets up to n. carrying out the steps of: observing the identifiable tag; One or more unique spatial addresses of the substrate having one or more observation signals. identifying the A portion of one or more proteins having an identified unique spatial address One or more epitopes, each associated with one or more observed signals determining that the sample contains Each protein segment is characterized based on one or more epitopes The process of doing this. 20. A method for determining a characteristic of a protein, comprising the steps of: Obtaining a substrate, each location containing one protein or at least six of the proteins. 0% of the proteins share the same amino acid sequence. A portion of one or more proteins is conjugated to the substrate so as to Circumstances; Based on a fluid containing a first to an nth (ordered) set of one or more affinity reagents, applying to the material one or more affinity reagents, each of which comprises one or more An epitope (a continuous or discontinuous amino acid sequence) of a part of a protein of a species each affinity reagent of the first to nth sets of one or more affinity reagents is specific linked to an identifiable tag; Each application of one or more affinity reagents to the substrate of the first and subsequent sets up to n. carrying out the steps of: observing the identifiable tag; One or more unique spatial addresses of the substrate having one or more observation signals. identifying the A portion of one or more proteins having an identified unique spatial address One or more epitopes, each associated with one or more observed signals determining that the sample contains Each protein segment is characterized based on one or more epitopes The process of doing this. 21. At least 400 different proteins will be identified using protein sequencing that relies on data from mass spectrometry. can be used to identify proteins at least 10% more quickly than techniques for protein identification; Methods under clauses 19 or 20. 22. Identify at least 400 distinct proteins with at least 50% accuracy; How to. 23. Whether the method identifies a particular protein by itself within a confidence threshold of greater than 10% Regardless of whether the particular protein is a member of a particular protein family, The method of clause 22, which identifies the 24. A portion of one or more proteins is located on the substrate based on the size of the protein. Separated, clause 19 or 20 methods. 25. A portion of one or more proteins is separated on the substrate based on the charge of the protein. Separated, by the methods of clauses 19 or 20. 26. The method of clause 19 or 20, wherein the substrate comprises microwells. 27. The method of clause 19 or 20, wherein the substrate comprises microwells of different sizes. 28. The method of clause 19 or 20, wherein the protein is attached to the substrate via a biotin attachment. 29. The method of clause 19 or 20, wherein the protein is attached to the substrate via a nucleic acid. 30. The method of clause 29, wherein the protein is attached to the substrate via nucleic acid nanoballs. 31. The method of clause 19 or 20, wherein the protein is attached to the substrate via nanobeads. 32. The process of obtaining a substrate to which one or more protein moieties are attached comprises the steps of: Obtaining a substrate with the correct array and each functional group corresponds to one protein molecule from the sample. Clause 19 or 20, comprising applying a protein sample to the How to. 33. The process for obtaining a substrate having an ordered array of functional groups includes photolithography, Drop-pen nanolithography, nanoimprint lithography, nanosphere lithography Thermal scanning probe lithography, local oxidation nanolithography, molecular self-assembly, stainless steel and electron beam lithography. The methods of clause 32, including the process. 34. The method of claim 3, wherein each functional group is positioned at least about 300 nm away from each other functional group. Method 2. 35. The method of claim 19 or 2, wherein the substrate comprises an ordered array of microwells of different sizes. 0 ways. 36. The step of obtaining a substrate comprises the steps of conjugating a first sample of protein to the substrate; using a protein dye to detect each protein-bound site from the sample. conjugating the second sample, and collecting each of the proteins bound from the second sample. 21. The method of clause 19 or 20, comprising using a protein dye to detect location. 37. The step of obtaining a substrate comprises the steps of conjugating a first sample of protein to the substrate; using a protein dye to detect each protein-bound site from the sample. From the number of bound proteins, the percentage of functional groups on the substrate to which proteins are not bound is determined. 21. The method of clause 19 or 20, comprising the step of determining 38. At least one component is a binding threshold for any instance of a core sequence, regardless of the adjacent sequence. The affinity reagents may be modified to have binding affinities greater than or equal to 100% by weight based on the same core sequence but with different flanking sequences. The method of clause 19 or 20, which may comprise a pool of components that bind to the sequence.
Claims
1. A method for determining polypeptide characteristics, including the following steps: A step of providing a plurality of polypeptides bound to the surface of a substrate, wherein each polypeptide molecule of the plurality of polypeptides is bound to an individually addressable position on the substrate, and the plurality of polypeptides bound to the surface of the substrate comprises at least 1,000,000 individual polypeptide molecules; A step of contacting a plurality of polypeptides on an array with at least 100 different affinity reagents by sequentially applying a solution containing at least one affinity reagent selected from at least 100 different affinity reagents to the plurality of polypeptides, wherein each of the at least 100 different affinity reagents has a binding affinity to two or more different epitopes present in the plurality of polypeptides; A step of individually detecting the binding of each of the at least 100 different affinity reagents to individual polypeptide molecules on the substrate; and A step of determining the characteristics of each of the plurality of polypeptides based on different affinity reagents that bind to each of the individual polypeptide molecules.
2. The method according to claim 1, wherein the plurality of polypeptides bonded to the surface of the substrate comprises at least 100,000,000 individual polypeptide molecules.
3. The method according to claim 1, wherein the plurality of polypeptides bound to the surface of the array comprises at least 100 different polypeptides.
4. The method according to claim 1, wherein the plurality of polypeptides bound to the surface of the array comprises at least 500 different polypeptides.
5. The method according to claim 1, wherein the plurality of polypeptides bound to the surface of the array comprises at least 1,000 different polypeptides.
6. The method according to claim 1, wherein the plurality of polypeptides bound to the surface of the array comprises at least 10,000 different polypeptides.
7. The method according to claim 1, wherein each polypeptide molecule of the plurality of polypeptides bound to the surface of the array comprises at least 20,000 distinct polypeptide molecules.
8. The method according to claim 1, wherein the individual polypeptide molecules are bonded to the surface of the substrate via an intermediate.
9. The method according to claim 8, wherein the intermediate comprises nanoparticles.
10. The method according to claim 8, wherein the intermediate comprises a nucleic acid molecule.
11. The method according to claim 1, wherein each of the polypeptide molecules is bonded to the surface of the substrate in a plurality of nanowells on the surface of the substrate.
12. The method according to claim 1, wherein the affinity reagent is selected from the group consisting of antibodies, antibody fragments, and aptamers.
13. The method according to claim 1, wherein the step of individually detecting each of the at least 100 different affinity reagents optically detects the binding of each of them to individual polypeptide molecules on the substrate.
14. The method according to claim 1, wherein a plurality of the at least 100 different affinity reagents have binding specificity to different epitopes consisting of three, four, or five consecutive amino acid residues.
15. The method according to claim 1, wherein the step of determining the characteristics of each of the plurality of polypeptides on the surface of the substrate includes identifying each of the at least 100 different affinity reagents bound to each of the individual polypeptide molecules in the plurality of polypeptides on the surface of the array in order to determine the characteristics of each of the plurality of polypeptides on the surface of the substrate.
16. The method according to claim 1, wherein the step of individually detecting the binding of each of the at least 100 different affinity reagents bound to each of the individual polypeptide molecules on the substrate includes optically detecting the position on the surface of the substrate to which each individual polypeptide molecule is bound and to which each of the affinity reagents is bound.
17. The method according to claim 1, wherein the step of determining the characteristics of each of the plurality of polypeptides includes identifying each of the plurality of polypeptides on the surface of the substrate.
18. The method according to claim 1, wherein the step of determining the characteristics of the plurality of polypeptides includes quantifying each of the different individual polypeptide molecules in the plurality of polypeptides on the surface of the substrate.
19. The method according to claim 1, wherein determining the characteristics of each of the plurality of polypeptides includes determining the probable identity of each of the individual polypeptide molecules of the plurality of polypeptides.
20. The method according to claim 19, further comprising determining the characteristics of each of the plurality of polypeptides to determine the reliability of the probable identity.