Decoding approach for protein identification

By combining computer methods with various empirical measurements and database comparisons, the accuracy and quantification challenges of identifying unknown proteins in existing technologies have been solved, achieving efficient and accurate protein identification and quantification.

JP2025163121APending Publication Date: 2025-10-28NAUTILUS SUBSIDIARY INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025127776
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-10-20
Filing Date
2025-07-30
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies have high specificity and sensitivity requirements for protein identification, making it difficult to accurately identify unknown proteins, and the quantification process is prone to errors.

Method used

A computer-implemented method is used to perform various empirical measurements on unknown protein samples, combined with information such as affinity reagent binding measurements, protein length, hydrophobicity, and isoelectric point. An algorithm is then used to calculate the probability of the presence of candidate proteins, and a database is used for comparison to select the most likely protein.

Benefits of technology

It improves the accuracy and efficiency of unknown protein identification, reduces errors, and achieves efficient quantification and identification of unknown proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025163121000001_ABST
    Figure 2025163121000001_ABST
Patent Text Reader

Abstract

To provide a decoding approach method for accurate and efficient identification of proteins.SOLUTION: A decoding approach method for identification of proteins includes the steps of: receiving information of a plurality of empirical measurements on unknown proteins in a sample; comparing the information of empirical measurements to a database comprising a plurality of protein sequences, each protein sequence corresponding to a candidate protein among a plurality of candidate proteins; calculating, for each candidate protein, a probability that the candidate protein generates the observed set of measurement outcomes, based on the comparison of the information of empirical measurements to the database; and selecting a candidate protein with the highest probability and calculating a probability that the candidate protein is an exact identity.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] cross reference This application is a continuation of U.S. Provisional Patent Application No. 62 / 611,979, filed December 29, 2017, and U.S. Provisional Patent Application No. 62 / 611,979, filed October 20, 2018. This application claims the benefit of International Application No. PCT / US2018 / 056807, each of which is hereby incorporated by reference in its entirety. The bodies are incorporated herein by reference. [Background technology]

[0002] background Current techniques for protein identification typically involve highly specific and sensitive parental sequences. Binding of compatible reagents (such as antibodies) and subsequent readout, or from a mass spectrometer It relies on either peptide read data (typically 12-30 amino acids long). Such techniques allow for the transfer of highly specific and sensitive affinity reagents to proteins of interest. determining the presence, absence, or amount of the candidate protein based on analysis of the binding assay to the It can be applied to unknown proteins in a sample. Summary of the Invention

[0003] overview Provided herein is a method for improved identification and characterization of proteins in samples of unknown proteins. The need for quantification is recognized. The methods and systems provided herein provide for the determination of This significantly reduces or eliminates errors in identifying proteins, and Such methods and systems improve the quantification of unknown proteins. Accurate and efficient identification of candidate proteins in protein samples can be achieved. The identification of parent molecules that are configured to selectively bind to one or more candidate proteins is carried out. Measurement of binding of affinity reagent probes, protein length, protein hydrophobicity, and isoelectric point In some embodiments, the calculation may be based on the information obtained from the unknown protein. The charge can be applied to individual affinity reagent probes, pooled affinity reagent probes, or individual The affinity reagent probes are exposed to a combination of pooled affinity reagent probes. Identification is the determination of the confidence level that each of one or more candidate proteins is present in the sample. It may involve speculation.

[0004] The methods and systems provided herein are directed to the production of fully intact proteins, or to identify proteins based on a series of experiments performed on protein fragments. Each experiment may include an algorithm for performing empirical measurements on a protein. and may provide information that may be useful for identifying the protein. Examples include binding of affinity reagents (e.g., antibodies or aptamers), protein length, protein Information about the experimental outcome includes measuring the protein's hydrophobicity and isoelectric point. To calculate the probability or likelihood of a quality candidate and / or observed experimental outcome We select proteins from the list of protein candidates that maximize the likelihood of The methods provided herein can be used to infer the identity of proteins. The method and system also provides a population of protein candidates and the relationships between these protein candidates. and algorithms for calculating the probability of experimental outcomes from each.

[0005] In one aspect, the present disclosure provides a method for identifying proteins in a sample of unknown proteins. The present invention provides a computer-implemented method for detecting a variance in a variance of ... ) collecting information on a plurality of empirical measurements performed on the unknown protein in the sample; (b) receiving by said computer said plurality of said empirical measurements; At least a portion of the information is compared to a database containing a plurality of protein sequences. a step of computer-based comparison, wherein each protein sequence is compared with a plurality of candidate proteins; (c) selecting a candidate protein from the plurality of protein sequences; the plurality of empirical measurements of the information for the database including a column. one or more candidate proteins among said plurality of candidate proteins based on at least a portion of said comparison For each of the coproteins, (i) combining the information of the plurality of empirical measurements with the candidate coproteins. (ii) the probability that the candidate protein will be produced when the candidate protein is present in the sample. (iii) the probability that the candidate protein is and the probability that the compound is present in the sample. The process.

[0006] In some embodiments, two or more of the plurality of empirical measurements comprise: (i) a pre-existing condition in the sample; For each of the one or more affinity reagent probes for the unknown protein, Binding assays, wherein each affinity reagent probe binds to one of said plurality of candidate proteins. a binding assay configured to selectively bind to one or more candidate proteins; (i) the length of one or more of the unknown proteins in the sample; (iii) the length of one or more of the unknown proteins in the sample; (iv) the hydrophobicity of one or more of said unknown proteins; and and the isoelectric point of the unknown protein.

[0007] In some embodiments, generating the plurality of probabilities comprises generating a plurality of additional affinity reagents. receiving additional information of the binding measurement for each of the probes, Each of the additional affinity reagent probes is directed to one or more of the plurality of candidate proteins. In some embodiments, the method is configured to selectively bind to the candidate protein. The method comprises, for each of said one or more candidate proteins, determining whether said candidate protein is generating a confidence level matching one of said unknown proteins in said sample. nothing.

[0008] In some embodiments, the plurality of affinity reagent probes comprises 50 or fewer affinity reagent probes. In some embodiments, the plurality of affinity reagent probes comprises 100 or fewer probes. In some embodiments, the plurality of affinity reagent probes comprises In some embodiments, the plurality of affinity reagent probes comprises 200 or fewer affinity reagent probes. In some embodiments, the affinity reagent probes comprise 300 or fewer affinity reagent probes. The plurality of affinity reagent probes includes 500 or fewer affinity reagent probes. In embodiments, the plurality of affinity reagent probes comprises more than 500 affinity reagent probes. In some embodiments, the method comprises generating a paper or Further included is the step of generating an electronic report.

[0009] In some embodiments, the sample comprises a biological sample. In some embodiments, the method further comprises: and further comprising identifying a disease state in the subject based at least on the probability of .

[0010] In some embodiments, (c) comprises selecting one or more candidate proteins from said plurality of candidate proteins. For each of the proteins, (i) combining the information of the plurality of empirical measurements into the candidate protein. The method further comprises generating, by the computer, the probability that a protein will be produced. In some embodiments, (c) comprises selecting one or more candidate proteins from said plurality of candidate proteins. (ii) for each of the candidate proteins, when the candidate protein is present in the sample, generating the probabilities that the plurality of empirical measurements are not observed by the computer; In some embodiments, (c) comprises selecting one or more of said plurality of candidate proteins. for each of a plurality of candidate proteins, (iii) determining whether the candidate protein is present in the sample; generating the probability of the occurrence of a disease by the computer. In some embodiments, the measured outcome includes binding of the affinity reagent probe. Thus, the measured outcome includes non-specific binding of the affinity reagent probe. In some embodiments, the measured outcome comprises binding of the affinity reagent probe. The measured outcome includes non-specific binding of the affinity reagent probe. In some embodiments, the empirical measurement comprises binding of an affinity reagent probe. The empirical measurement includes non-specific binding of the affinity reagent probe.

[0011] In some embodiments, the method produces a certain sensitivity of protein identification with a predetermined threshold. In some embodiments, the predetermined threshold is inaccurate. In some embodiments, the protein in the sample is In some embodiments, the Proteins do not start at the end of the protein.

[0012] In some embodiments, the empirical measurement comprises determining the presence or absence of one or more of the unidentified nucleotides in the sample. In some embodiments, the empirical measurement includes the length of the sample. In some embodiments, the hydrophobicity of one or more of the unknown proteins is The empirical determination includes the isoelectric point of one or more of the unknown proteins in the sample. In some embodiments, the empirical measurements are measurements performed on a mixture of antibodies. In some embodiments, the empirical measurements are performed on samples obtained from multiple species. In some embodiments, the empirical measurements include measurements performed in non-synonymous In the presence of single amino acid variations (SAVs) caused by single nucleotide polymorphisms (SNPs), This includes measurements performed using

[0013] Additional aspects and advantages of the present disclosure are set forth in the following detailed description of exemplary embodiments of the present disclosure. It will be readily apparent to those skilled in the art from the following detailed description. The disclosure is capable of other and different embodiments, and certain details thereof may be modified in various ways. Obviously, modifications can be made without departing from the disclosure. Therefore, the drawings and description are to be regarded as illustrative in nature and as restrictive. It should not be regarded as something.

[0014] INCORPORATION BY REFERENCE All publications, patents, and patent applications mentioned herein are incorporated by reference in their entirety. Each patent or patent application is specifically and individually identified as being incorporated by reference. All such disclosures are hereby incorporated by reference to the same extent as if set forth in the appended claims. Any publications and patents or patent applications incorporated by reference that conflict with this disclosure are included herein. To the extent applicable, the present specification shall take precedence over any such conflicting matters, and / or Or it is intended to be superior. [Brief explanation of the drawings]

[0015] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the invention and its advantages will be apparent from the following detailed description, which sets forth illustrative embodiments in which the principles of the invention are utilized. The detailed description and accompanying drawings (also referred to herein as "Figures") illustrate exemplary embodiments of the present invention. " and "FIG."

[0016] [Figure 1] 1 shows an exemplary flow chart for protein identification of an unknown protein in a biological sample according to disclosed aspects. [Figure 2]Figure 1 shows the sensitivity of affinity reagent probes (e.g., percent of substrates identified with a false discovery rate (FDR) of less than 1%) according to disclosed embodiments, plotted against the number of probe recognition sites (e.g., trimer-binding epitopes) in the affinity reagent probes (ranging up to 100 probe recognition sites or trimer-binding epitopes) for three different experimental cases (using 50, 100, and 200 probes, represented by gray circles, black circles, and white circles, respectively). [Figure 3] Figure 1 shows the sensitivity of affinity reagent probes (e.g., percent of substrates identified with a false discovery rate (FDR) of less than 1%) according to disclosed embodiments, plotted against the number of probe recognition sites (e.g., trimer-binding epitopes) in the affinity reagent probes (ranging up to 700 probe recognition sites or trimer-binding epitopes) for three different experimental cases (using 50, 100, and 200 probes, represented by gray circles, black circles, and white circles, respectively). [Figure 4] 1 shows plots demonstrating the sensitivity of protein identification from experiments using 100 probes (left), 200 probes (center), or 300 probes (right) according to disclosed embodiments. [Figure 5-1] 5 shows plots illustrating experimental protein identification sensitivity using various protein fragmentation approaches. In the top and bottom rows, respectively, protein identification performance according to disclosed embodiments is shown with 50, 100, 200, and 300 affinity reagent measurements (from left to right in the four panels) at maximum fragment length values ​​of 50, 100, 200, 300, 400, and 500 (represented by a hexagon, downward-pointing triangle, upward-pointing triangle, diamond, square, and circle, respectively). [Figure 5-2]5 shows plots illustrating experimental protein identification sensitivity using various protein fragmentation approaches. In the top and bottom rows, respectively, protein identification performance according to disclosed embodiments is shown for 50, 100, 200, 300, 400, and 500 maximum fragment length values ​​(represented by a hexagon, downward-pointing triangle, upward-pointing triangle, diamond, square, and circle, respectively) with 50, 100, 200, and 300 affinity reagent measurements (from left to right in the four panels). [Figure 6] 1 shows a plot illustrating the sensitivity (percent of substrates identified with an FDR of less than 1%) of human protein identification from experiments using various combinations of measurement types according to disclosed embodiments. [Figure 7] 1 shows plots illustrating the sensitivity of experimental protein identification using 50, 100, 200, or 300 affinity reagent probe paths for unknown proteins from either E. coli, yeast, or human (represented by circles, triangles, and squares, respectively), according to disclosed embodiments. [Figure 8] 1 shows a plot showing binding probability (y-axis, left) and sensitivity of protein identification (y-axis, right) versus replicates (x-axis), according to disclosed embodiments. [Figure 9] Comparison of the estimated false positive rate with the true false positive rate for a simulated 200 probe experiment according to disclosed embodiments demonstrates accurate estimation of the false positive rate. [Figure 10] 1 depicts a computer control system programmed or otherwise configured to perform the methods provided herein. [Figure 11] The performance of censored versus uncensored protein identification approaches is shown. [Figure 12] The tolerance of censored and non-censored protein identification approaches to random "false negative" binding outcomes is shown. [Figure 13] 1 shows the tolerance of censored and non-censored protein identification approaches to random "false positive" binding outcomes. [Figure 14] 1 shows the performance of censored and non-censored protein identification approaches using over- and under-estimated affinity reagent binding probabilities. [Figure 15] 1 shows the performance of truncated and non-censored protein identification approaches using affinity reagents with unknown binding epitopes. [Figure 16] 1 shows the performance of truncated and non-censored protein identification approaches using affinity reagents lacking binding epitopes. [Figure 17] We show the performance of censored and non-censored protein identification approaches using affinity reagents targeting the top 300 most abundant trimers in the proteome, 300 randomly selected trimers in the proteome, or the 300 least abundant trimers in the proteome. [Figure 18] 1 shows the performance of censored and non-censored protein identification approaches using affinity reagents with random or biosimilar off-target sites. [Figure 19] The performance of censored and non-censored protein identification approaches using an optimal set of affinity reagents (probes) is shown. [Figure 20] 1 shows the performance of censored and uncensored protein identification approaches using unmixed candidate affinity reagents and mixtures of candidate affinity reagents. [Figure 21] 1 illustrates two hybridization steps in enhancing binding between an affinity reagent and a protein, according to some embodiments. [Figure 22]1 shows the performance of protein identification using a collection of reagents for selective modification and detection of four amino acids (K, D, C, and W) according to some embodiments. [Figure 23]

[0023] Figure 1 shows the performance of protein identification using a collection of reagents for selective modification and detection of 20 amino acids (R, H, K, D, E, S, T, N, Q, C, G, P, A, V, I, L, M, F, Y, and W) according to some embodiments. [Figure 24] 1 shows the performance of protein identification using an ordered measurement of amino acids, according to some embodiments, where every amino acid is measured with a probability of detection (equal to the efficacy of the reaction) shown on the x-axis, and the y-axis shows the percent of proteins in a sample that are identified with a false discovery rate below 1%. DETAILED DESCRIPTION OF THE INVENTION

[0017] Detailed Description While various aspects of the present invention have been shown and described herein, such Those skilled in the art will appreciate that the embodiments are provided by way of example only. Variations and substitutions may occur to those skilled in the art without departing from the invention. It will be understood that various alternatives to the embodiments of the invention described herein may be employed. It should be.

[0018] The term "sample," as used herein, generally refers to a biological sample (e.g., A sample may be derived from tissue or cells, or from tissue or cells. The sample may be taken from the environment of the cell. In some examples, the sample may be a tissue biopsy, blood, Plasma, extracellular fluid, dried blood spots, cultured cells, culture media, discarded tissues, plant material , synthetic proteins, bacterial and / or viral samples, fungal tissues, archaea, if The sample may contain or be derived from a protozoan or a protozoan. The sample may be isolated from a sample of forensic evidence. fingerprints, saliva, urine, blood, stool, semen, or other bodily fluids isolated from a primary source prior to In some instances, proteins are isolated from their primary source (cells) during sample preparation. The samples may be from extinct species. Proteins may be extracted from their primary source, including, but not limited to, samples derived from fossils. It may or may not be purified from, or otherwise part of, It may or may not be enriched from a secondary source. Therefore, the primary source is homogenized prior to further processing. The cells are lysed using a buffer such as RIPA buffer. Denaturing buffers can also be used at this stage. The sample may be filtered or centrifuged to remove lipids and particulate matter. The sample may also be purified to remove nucleic acids or to remove RNases and The sample may be treated with untreated proteins, denatured proteins, The protein may comprise a protein fragment, or a partially degraded protein.

[0019] The sample may be taken from a subject having a disease or disorder. The disease or disorder may be an infectious disease or disorder. Diseases, immune disorders or immune diseases, cancer, genetic diseases, degenerative diseases, lifestyle-related diseases, wounds, rare diseases Infectious diseases can be bacterial, viral, fungal, and Non-limiting examples of cancers include bladder cancer, lung cancer, and / or cancer of the lungs. Cancer, brain cancer, melanoma, breast cancer, non-Hodgkin's lymphoma, cervical cancer, ovarian cancer, colon, Rectal cancer, pancreatic cancer, esophageal cancer, prostate cancer, kidney cancer, skin cancer, leukemia, and thyroid cancer Some examples of genetic diseases or disorders include: Diseases that may be affected include, but are not limited to, multiple sclerosis (MS), cystic fibrosis, Charcot-Marie disease Tooth disease, Huntington's disease, Peutz-Jeghers syndrome, Down syndrome, rheumatoid arthritis Non-limiting examples of lifestyle-related diseases include obesity, diabetes, arteriosclerosis, and Tay-Sachs disease. diabetes, heart disease, stroke, high blood pressure, cirrhosis of the liver, nephritis, cancer, chronic obstructive pulmonary disease (COPD), hearing loss Some examples of wounds include, but are not limited to, abrasions, Accidental injury, brain injury, contusion, burns, concussion, congestive heart failure, construction site injuries, dislocation, upset Chest, fractures, hemothorax, herniated disc, hip pointer, hypothermia, lacerations, compressed nerves pinched nerve, pneumothorax, rib fractures, sciatica, spinal cord injury, tendon, ligament, and fascial The sample may be collected from a subject with a disease or disorder, including a brain injury, traumatic brain injury, and whiplash injury. The sample may be taken before and / or after the treatment. Samples may be taken during treatment or during a treatment regimen. Multiple samples may be taken during treatment. Samples may be taken from the subject to monitor the effectiveness of the treatment over time. Subjects known or suspected to have a communicable disease for which no diagnostic test is available It may be taken from

[0020] The sample may be taken from a subject suspected of having a disease or disorder. Description of symptoms such as fatigue, nausea, weight loss, aches and pains, weakness, or amnesia The test may be taken from a subject experiencing unexplained symptoms. The sample may be taken from a subject with an explained symptom. , family medical history, age, environmental exposures, lifestyle risk factors, or other known risk factors collected from subjects at risk of developing a disease or disorder due to factors such as the presence of a child It is okay to do so.

[0021] The sample may be taken from an embryo, a fetus, or a pregnant woman. In some examples, the sample is In some instances, the protein may be isolated from maternal plasma. Protein isolated from circulating fetal cells in fluid.

[0022] The sample may be taken from a healthy individual. In some cases, the sample may be taken from the same individual. In some cases, samples obtained longitudinally may be collected from individuals. With the goal of monitoring physical health and early detection of health problems, In some embodiments, samples may be collected at home or in a clinical setting. , and may subsequently be transported by mail, courier, or other transport method prior to analysis. For example, a home user may collect a blood spot sample by finger prick. The blood spot samples may be dried and subsequently delivered by mail prior to analysis. In some cases, samples obtained longitudinally may be transported in a healthy To monitor performance or responses to stimuli that are expected to affect cognitive performance. Non-limiting examples include assessing response to medication, diet, or exercise therapy. include.

[0023] The sample proteins are treated to remove modifications that may interfere with binding to the epitope. For example, the protein may be treated with an enzyme. The protein may be treated with glycosidases to remove post-translational glycosylation. The protein may be treated with a reducing agent to reduce disulfide bonds in the protein. Proteins may be treated with phosphatases to remove phosphate groups. Other non-limiting examples of post-translational modifications that can be removed include acetate, amide groups, methyl groups, Lipid, ubiquitin, myristoylation, palmitoylation, isoprenylation or prenylation (e.g., farnesol and geranylgeraniol), farnesylation, geranylgeraniol Ranylation, glypiation, lipoylation, flavin moiety binding, phospholipid This includes hopantetheinylation, as well as retinylidene Schiff base formation.

[0024] The sample protein may have one or more residues that are more amenable to binding or detection by the affinity reagent. The residues may be modified to make them more amenable to the In this case, the sample protein may facilitate or enhance binding to the epitope. In some instances, the protein may be processed to preserve post-translational protein modifications that result in In some instances, a phosphatase inhibitor may be added to the sample. To protect the bond, an oxidizing agent may be added.

[0025] The proteins of the sample may be fully or partially denatured. Proteins can be completely denatured. Proteins can be denatured by detergents, strong acids or bases, concentrated condensed inorganic salts, organic solvents (e.g., alcohol or chloroform), irradiation, Alternatively, the protein may be denatured by application of an external stress such as heat. Proteins may also be denatured by precipitation, lyophilization, and denaturation in a denaturing buffer. The protein may be denatured by heating. The protein may be chemically modified. A modification method that is less likely to cause artifacts may be chosen.

[0026] The sample protein is subjected to conjugation to produce shorter polypeptides. The remaining protein may be processed either before or after the enzyme is digested to generate fragments. It may be partially digested with an enzyme such as proteinase K or left intact. In a further example, the protein may be exposed to a protease, such as trypsin. Additional examples of proteases include serine proteases, cysteine ​​proteases, threoproteases, and threoproteases. Nerve proteases, aspartic acid proteases, glutamic acid proteases, metalloproteinases protease, and asparagine peptide lyase.

[0027] In some cases, extremely large and small proteins (e.g., titin) It may be useful to remove such proteins, for example by filtration or other suitable means. In some instances, extremely large proteins can be removed by a variety of methods. At most about 400 kilodaltons (kD), 450 kD, 500 kD, 600 kD, 650 kD, 700 kD, 750 kD, In some instances, the protein may be extremely large. A large protein is at least about 8,000 amino acids, about 8,500 amino acids, about 9,000 amino acids, about 9,500 amino acids, about 10,000 amino acids, about 10,500 amino acids, about 11,000 amino acids, or about 1 In some instances, small proteins may include proteins that are 5,000 amino acids long. The quality is approximately less than 10 kD, less than 9 kD, less than 8 kD, less than 7 kD, less than 6 kD, less than 5 kD, less than 4 kD, less than 3 In some examples, the protein may comprise less than 1 kD, less than 2 kD, or less than 1 kD. , small proteins are less than about 50 amino acids, less than 45 amino acids, less than 40 amino acids, less than 35 amino acids It may contain proteins of less than about 30 amino acids or less than about 30 amino acids. Proteins can be removed by size exclusion chromatography. Proteins were isolated by size exclusion chromatography to isolate intermediate-sized polypeptides. It is treated with proteases to produce intermediate-sized proteins, which then recombine with the intermediate-sized proteins in the sample. It is okay to do so.

[0028] The proteins of the sample may have identifiable tags, which allow for, for example, sample multiplexing. Some non-limiting examples of identifiable tags include: based on fluorophores, fluorescent nanoparticles, quantum dots, magnetic nanoparticles, or DNA barcodes. The fluorophores used are GFP, YFP, RFP, eGFP, mCherry, tdtomato, FITC, Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, A lexa Fluor 680, Alexa Fluor 750, Pacific Blue, Coumarin, BODIPY FL, Pacific Gree n, Oregon Green, Cy3, Cy5, Pacific Orange, TRITC, Texas Red, phycoerythrin, and fluorescent proteins such as allophycocyanin.

[0029] Any number of protein samples can be multiplexed. For example, multiplexed reactions can be performed in 2, 3, 4 , 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, about 20, about 25, about 30, about 35 , about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100 The identifiable tag may comprise proteins from about 100 or more initial samples. This may provide a means to examine proteins with respect to the sample from which they originate, or to compare proteins from different samples. Proteins may be directed to segregate to different areas or solid supports. In this embodiment, the protein is then subjected to a chemical reaction to chemically bind the protein to the substrate. applied to a functionalized substrate.

[0030] Any number of protein samples without tagging or multiplexing For example, multiplexed reactions may be performed in groups of 2, 3, 4, 5, or 6 pieces, 7 pieces, 8 pieces, 9 pieces, 10 pieces, 11 pieces, 12 pieces, 13 pieces, 14 pieces, 15 pieces, 16 pieces, 17 pieces, 18 pieces, 19 pieces, About 20 pieces, about 25 pieces, about 30 pieces, about 35 pieces, about 40 pieces, about 45 pieces, about 50 pieces, about 55 pieces, about 60 pieces, about 65 pieces, Approximately 70, approximately 75, approximately 80, approximately 85, approximately 90, approximately 95, approximately 100, or more than approximately 100 initial samples For example, diagnostics for rare conditions may involve pooled samples. Analysis of the individual samples may then be performed on the samples tested, and the diagnostic results may be used to identify the individual ... The samples may be run only from the samples in the pool that were combinatorially pooled. may be multiplexed without being tagged using a tagging design, in which individual The signals from the samples are separated from the pool to be analyzed using computer-assisted multiplexing. The samples are mixed into pools in a manner that allows them to be distinguished from one another.

[0031] The term "substrate," as used herein, generally refers to a material that forms a solid support. A substrate, or solid substrate, refers to a substrate to which a protein can be covalently bound or A non-limiting example of a solid substrate is a particle. Substrates, beads, slides, surfaces of device components, membranes, flow cells, wells, chambers, Macrofluidic chambers, microfluidic chambers, channels, microfluidic channels, or or any other surface. The surface of the substrate can be flat or curved, and The surface of the substrate may have other shapes and may be smooth or rough. In some embodiments, the substrate may comprise a microwell. Any carbohydrate, plastic such as polystyrene or polypropylene, polyacrylic may be composed of amide, latex, silicone, metal such as gold, or cellulose; and Further modifications may be made to allow or enhance covalent or non-covalent binding of proteins. For example, the surface of the substrate may be decorated with maleic acid moieties or succinic acid moieties. or may be functionalized by modification with specific functional groups such as amino groups, thiol groups, or by modification with chemically reactive groups such as acrylate groups, e.g., by silanization. Suitable silane reagents include aminopropyltrimethoxysilane, The groups include 4-aminopropyltriethoxysilane, 4-aminobutyltriethoxysilane, and 4-aminobutyltriethoxysilane. The material may be functionalized with N-hydroxysuccinimide (NHS) functional groups. For example, epoxy silane, acrylate silane, or acrylamide silane may be used. and can be derivatized with other reactive groups such as acrylate or epoxy. The substrate and treatment for this purpose preferably undergoes repeated bonding, washing, imaging, and dissolving. In some instances, the substrate is suitable for use as a slide, flow cell, or microscale or nanoscale structures (e.g., microwells, Ordered nanopillars, single molecule arrays, nanoballs, nanopillars, or nanowires (correct structure).

[0032] The spacing of the functional groups on the substrate can be regular or random. Regular arrays can be fabricated using techniques such as photolithography, dip-pen nanolithography, Nanoimprint lithography, nanosphere lithography, nanoball lithography, Nanopillar array, nanowire lithography, scanning probe lithography, thermochemical lithography lithography, thermal scanning probe lithography, local oxidation nanolithography, molecular self-assembly, It may be fabricated by stencil lithography or electron beam lithography. The functional groups in the array are spaced 200 nanometers (nm) from any other functional group. ) or less than about 200 nm, about 225 nm, about 250 nm, about 275 nm, about 300 nm, about 32 5 nm, approx. 350 nm, approx. 375 nm, approx. 400 nm, approx. 425 nm, approx. 450 nm, approx. 475 nm, approx. 500 nm, approx. 52 5 nm, approx. 550 nm, approx. 575 nm, approx. 600 nm, approx. 625 nm, approx. 650 nm, approx. 675 nm, approx. 700 nm, approx. 72 5 nm, approx. 750 nm, approx. 775 nm, approx. 800 nm, approx. 825 nm, approx. 850 nm, approx. 875 nm, approx. 900 nm, approx. 92 5 nm, approx. 950 nm, approx. 975 nm, approx. 1000 nm, approx. 1025 nm, approx. 1050 nm, approx. 1075 nm, approx. 1100 nm , about 1125 nm, about 1150 nm, about 1175 nm, about 1200 nm, about 1225 nm, about 1250 nm, about 1275 nm, approx. 1300 nm, approx. 1325 nm, approx. 1350 nm, approx. 1375 nm, approx. 1400 nm, approx. 1425 nm, approx. 1450 nm, approx. 1 475 nm, approx. 1500 nm, approx. 1525 nm, approx. 1550 nm, approx. 1575 nm, approx. 1600 nm, approx. 1625 nm, approx. 1650 nm, approx. 1675 nm, approx. 1700 nm, approx. 1725 nm, approx. 1750 nm, approx. 1775 nm, approx. 1800 nm, approx. 1825 nm , about 1850 nm, about 1875 nm, about 1900 nm, about 1925 nm, about 1950 nm, about 1975 nm, about 2000 nm, The randomly spaced functional groups may be arranged such that the functional groups are spaced apart from each other by a distance of 2000 nm or greater. , on average at least about 50 nm, about 100 nm, about 150 nm, about 200 nm from any other functional group; approx. 250 nm, approx. 300 nm, approx. 350 nm, approx. 400 nm, approx. 450 nm, approx. 500 nm, approx. 550 nm, approx. 600 nm, approx. 650 nm, approx. 700 nm, approx. 750 nm, approx. 800 nm, approx. 850 nm, approx. 900 nm, approx. 950 nm, approx. 1000 nm , or in a dense state such that it is greater than 100 nm.

[0033] The substrate may be indirectly functionalized. For example, the substrate may be PEGylated and have a functional group The substrate may be applied to all of the PEG molecules or to a subset of the PEG molecules. Micro- or nano-scale structures (e.g., microwells, micropillars, single molecules) ordered structures such as nanoparticle arrays, nanoballs, nanopillars, or nanowires) The functionalized material may be functionalized using techniques suitable for the

[0034] The substrate comprises metal, glass, plastic, ceramic, or a combination thereof. In some preferred embodiments, the solid substrate is a flow cell. The flow cell may be composed of a single layer or multiple layers. The cell consists of a base layer (e.g., made of borosilicate glass), a channel layer (e.g., may include a cover layer or top layer, e.g., made of etched silicon When the layers are assembled together, an enclosed channel can be formed, which The thickness of each layer is variable, but is preferably The thickness is less than about 1700 μm. The layer is made of photosensitive glass, borosilicate glass, fused silicate glass, The substrate may be constructed from a suitable material such as fused silicate, PDMS, or silicone. The layers may be constructed from the same material or from different materials.

[0035] In some embodiments, the flow cell has an opening for the channel at the bottom of the flow cell. The flow cell can contain millions of attached microtubules in locations that can be individually visualized. In some embodiments, the target conjugation moiety used in the embodiments of the present invention may include a target conjugation moiety. The various flow cells used have different numbers of channels (e.g., 1 channel, 2 or more). 1 channel, 3 or more channels, 4 or more channels, 6 or more channels, 8 or more channels Channels, 10+ channels, 12+ channels, 16+ channels, or more than 16 Various flow cells may contain channels of different depths or widths. These may vary between channels in one flow cell or may be different. The flow cell channels may also have different depths. and / or width may vary. For example, one channel may have one or more Multiple locations: less than about 50 μm deep, about 50 μm deep, less than about 100 μm deep, and about 100 μm deep depth, about 100 μm to about 500 μm depth, about 500 μm depth, or more than about 500 μm depth The channels can be, but are not limited to, circular, semicircular, rectangular, pedestal, The cross-sectional shape may be any cross-sectional shape, including rectangular, triangular, or oval cross-sections.

[0036] The protein may be spotted, dropped, or pipetted onto the substrate; It may be poured on, washed or otherwise applied. In the case of moiety-functionalized substrates, protein modification is not required. Substrates functionalized with moieties (e.g., sulfhydryls, amines, or linker nucleic acids) In the case of , a cross-linking reagent (e.g., disuccinimidyl suberate, NHS, sulfonamide In the case of a substrate functionalized with a linker nucleic acid, the sample protein may be , may be modified with a complementary nucleic acid tag.

[0037] Photoactivatable crosslinkers may be used to direct crosslinking of samples to specific areas on a substrate. Photoactivatable crosslinkers allow proteins to be detected by attaching each sample to a known area of ​​the substrate. Photoactivatable crosslinkers may be used to allow for sample multiplexing. By detecting the fluorescent tag before cross-linking the protein, successfully tagged proteins can be identified. Examples of photoactivatable crosslinkers include, but are not limited to, N- 5-Azido-2-nitrobenzoyloxysuccinimide, Sulfosuccinimidyl 6-(4'-azidobenzoyloxysuccinimide) Dihydro-2'-nitrophenylamino)hexanoate, succinimidyl 4,4'-azipentanoate ate, sulfosuccinimidyl 4,4'-azipentanoate, succinimidyl 6-(4,4'-azipentanoate Dipentanamido)hexanoate, sulfosuccinimidyl 6-(4,4'-azipentanamido) hexanoate, succinimidyl 2-((4,4'-azipentanamido)ethyl)-1,3'-dithio propionate, and sulfosuccinimidyl 2-((4,4'-azipentanamido)ethyl )-1,3'-dithiopropionate.

[0038] The polypeptide may be attached to the substrate by one or more residues. In the present invention, the polypeptide may be linked via the N-terminus, C-terminus, both termini, or via internal residues. and may be attached.

[0039] In addition to permanent crosslinkers, the use of photocleavable linkers for some applications; and thereby allowing for selective extraction of proteins from the substrate after analysis. In some cases, it may be appropriate to make the photocleavable crosslinker In some cases, the photocleavable crosslinker may be used for multiplexing different samples. It may be used from one or more samples in a multiplexed reaction. The multiplexed reactions included a control sample crosslinked to the substrate via a permanent crosslinker, and a photocleavable sample. The sample may include an experimental sample that is crosslinked to a substrate via a crosslinking agent.

[0040] Each conjugated protein is optically Spatially separated from each other conjugated protein so that it can be resolved Proteins may therefore be individually labeled with unique spatial addresses. In some embodiments, this may be achieved by integrating each protein molecule with other protein molecules. The low concentration of proteins and the low density of attachment on the substrate are spatially separated from each other. This can be achieved by conjugation using attachment sites such as photoactivatable crosslinkers. When a protein is added to a predetermined location, photolysis is performed. A pattern may be used.

[0041] In some embodiments, each protein may be associated with a unique spatial address. For example, if proteins are attached to a substrate at spatially separated locations, each protein Proteins can be assigned indexed addresses, such as by coordinates. In this example, a grid of pre-assigned unique spatial addresses is In some embodiments, the location of each protein may be determined by a fixed mark on the substrate. The substrate may include easily identifiable fixed marks so that the accuracy of the measurement can be determined. In some instances, the substrate may have grid lines and / or In some instances, the surface of the substrate may have a "point of origin" or other reference point. Permanently or semi-permanently to provide a reference for locating the protein The conjugated polypeptide may be marked with a pattern such as the outer edge of the conjugated polypeptide. Its shape is also used as a reference to determine the unique position of each spot. good.

[0042] The substrate may also include conjugated protein standards and controls. Conjugated protein standards and controls are pre-conjugated proteins conjugated to known positions. In some instances, the conjugate may be a peptide or protein of known sequence. Protein standards and controls can serve as internal controls in the assay. The protein may be applied to the substrate from a purified protein stock, or Nucleic Acid-Programmable Protein Array (N The polymer may be synthesized on the substrate by a process such as APPA.

[0043] In some instances, the substrate may include fluorescent standards. These fluorescent standards may also be used to calibrate the intensity of the fluorescent signal between samples. Also, to correlate the intensity of the fluorescent signal with the number of fluorophores present in a given area, Fluorescent standards may be used to identify the different types of fluorophores used in the assay. It may include some or all of the fore.

[0044] Once the substrate is conjugated with proteins from the sample, multiple affinity reagent measurements are performed. The measurement process described herein can be carried out using a variety of affinity assays. In some embodiments, multiple affinity reagents may be mixed together. and the measurement is of the binding of the affinity reagent mixture to the protein-substrate conjugate. In some cases, binding of an affinity reagent mixture may be performed. The measurements obtained are based on different solvent conditions and / or different protein folding conditions. therefore, repeated measurements may vary depending on the same affinity reagent or affinity The same set of reactive reagents may be used under such altered solvent conditions and / or protein Several sets of binding measurements may be performed under different folding conditions to obtain different sets of binding measurements. In some cases, a different set of binding measurements may be performed to determine whether the protein in question is bound to the enzyme (e.g., processed (e.g., by glycosidases, phosphorylases, or phosphatases) Repeated measurements were performed on samples that had been treated with enzymes or samples that had not been treated with enzymes. It can be obtained by

[0045] The term "affinity reagent," as used herein, generally refers to a protein or This refers to a reagent that binds to a target molecule or peptide with reproducible specificity. For example, an affinity test The drug could be an antibody, antibody fragment, aptamer, mini-protein binder, or peptide. In some embodiments, the mini-protein binder may be 30 to 210 amino acids in length. In some embodiments, the miniaturized cellulose may contain a protein binder that can be used between the cellulose and the acid. Protein binders may be designed. For example, protein binders may be designed to bind large peptides. Cyclic molecules (e.g., those described in Hosseinzadeh et al., which are incorporated herein by reference in their entirety) al., “Comprehensive computational design of ordered peptide macrocycles,” Sci ence, 2017 Dec. 15; 358(6369): 1461-1466). In some embodiments, monoclonal antibodies, such as Fab fragments, may be selected. In some embodiments, the affinity reagent is a commercially available antibody, such as a commercially available antibody. In some embodiments, the desired affinity reagent is a useful By screening commercially available affinity reagents to identify those with the characteristic , may be selected.

[0046] Affinity reagents may have high, medium, or low specificity. In some instances, the affinity reagent may recognize several different epitopes. The affinity reagent may recognize epitopes present on two or more different proteins. In some instances, the affinity reagent recognizes an epitope present on many different proteins. In some cases, the affinity reagents used in the methods of the present disclosure may be epitaxial. In some cases, the present disclosure may be highly specific for only one of the topes. The affinity reagent used in the method is directed against only one epitope containing a post-translational modification. In some cases, affinity reagents may be highly specific. In some cases, the specificity of the epitope is highly similar to that of the target epitope. Affinity reagents with specificity for highly similar protein candidate sequences (e.g., 1A, 2B, 3C, 4C, 5C, 6C, 7C, 8C, 9C, 10C, 11C, 12C, 13C, 14C, 15C, 16C, 17C, 18C, 19C, 20C, 21C, 22C, 23C, specifically designed to distinguish between candidates with amino acid mutations, or isoforms In some cases, affinity reagents are used to maximize coverage of the protein sequence. In some embodiments, the antibody may have specificity for a highly diverse epitope. However, results may vary due to the stochastic nature of probe binding to protein-substrates. We anticipate that this may be useful and therefore may provide additional information regarding protein identification. Thus, experiments may be performed repeatedly using the same affinity probe.

[0047] In some cases, the specific, single epitope recognized by the affinity reagent Alternatively, multiple epitopes may not be completely known. For example, an affinity reagent may be used to identify one or more epitopes. or specifically for multiple full-length proteins, protein complexes, or protein fragments. Can be designed or selected for binding without knowledge of the specific binding epitope The qualitative process may have refined the binding profile of this reagent. Even if the specific binding epitope is unknown, binding assays using the affinity reagents can be performed. The determination can be used to determine the identity of a protein. Commercially available antibodies or aptamers designed to bind to protein targets are used as affinity reagents. Assay conditions (e.g., fully folded, partially denatured, After characterization under conditions of denaturation (or completely denatured), this affinity to the unknown protein The binding of the reactive reagent can provide information about the identity of the unknown protein. In some cases, a population of protein-specific affinity reagents (e.g., commercially available antibodies or aptamers), along with knowledge of the specific epitopes they target, can be used to generate protein identifications either with or without the knowledge of the In this case, the population of protein-specific affinity reagents is about 50, 100, 200, 300 , 400 pieces, 500 pieces, 600 pieces, 700 pieces, 800 pieces, 900 pieces, 1000 pieces, 2000 pieces, 3000 pieces, 4000 pieces, 5000 pieces The affinity reagents may include 10,000, 10,000, 20,000, or more than 20,000 affinity reagents. In the present invention, a population of affinity reagents has been shown to be reactive to a target in a particular organism, Any commercially available affinity reagent may be included, for example, a collection of protein-specific affinity reagents. may be assayed sequentially, with binding measurements made individually for each affinity reagent. In some cases, a subset of protein-specific affinity reagents may be used in binding assays. For example, a new mixture of affinity reagents may be prepared for each binding measurement pass. The mixture is selected to contain a subset of affinity reagents randomly selected from the complete set. For example, each subsequent mixture may be selected such that many of the affinity reagents are present in multiple mixtures. may be generated in the same random fashion, expecting them to be present in In this study, protein identification was more rapid using a mixture of protein-specific affinity reagents. In some cases, such mixtures of protein-specific affinity reagents may be produced. The compound is the percentage of unknown protein that the affinity reagent binds in any individual pass. The affinity reagent mixture may be approximately 100% of all available affinity reagents. It may comprise 1%, 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more than 90%. The mixture of affinity reagents evaluated in a single experiment may be identical even if the individual affinity reagents are the same. may or may not be in common. In some cases, they bind to the same protein. There may be multiple different affinity reagents in the population. Each affinity reagent may bind to a different protein. If multiple affinity reagents bind to a single unknown protein, the common affinity of the affinity reagents Confidence in the identity of the target unknown protein can be increased. In some cases, the use of multiple protein affinity reagents targeting the same protein may The affinity reagents bind to different epitopes on the same protein, and Only a subset of affinity reagents targeting the α- and β-terminal domains bind, and post-translational modifications or modifications of the binding epitope may be involved. or other steric hinderance. In some cases, the binding of an affinity reagent whose binding epitope is unknown may be Protein identification in combination with binding measurements of affinity reagents whose binding epitopes are known can be used to generate

[0048] In some examples, the one or more affinity reagents are 2, 3, 4, 5, 6, 7, or Binding to an amino acid motif of a given length, such as 1, 8, 9, 10, or more than 10 amino acids. In some examples, the one or more affinity reagents may be selected to combine two It appears to bind to a range of amino acid motifs of different lengths, from 10 to 40 amino acids. You may be selected as one of the following.

[0049] In some cases, the affinity reagents may be labeled with nucleic acid barcodes. In the example of In some instances, nucleic acid barcodes sort affinity reagents for repeated use. In some cases, the affinity reagent may be used to The reagents may be labeled with fluorophores that can be used to sort the reagents.

[0050] A family of affinity reagents may include one or more types of affinity reagents. For example, the methods of the disclosure can be used to detect antibodies, antibody fragments, Fab fragments, aptamers, peptides, and proteins. A family of affinity reagents containing one or more of the proteins may be used.

[0051] The affinity reagent may be modified. Examples of modifications include, but are not limited to, the addition of a detection moiety The detection moiety may be directly or indirectly bound. For example, The moiety may be directly covalently attached to the affinity reagent or may be attached via a linker. or a parent molecule such as a complementary nucleic acid tag or a biotin-streptavidin pair. The binding may be via affinity reactions. Any suitable bonding method can be chosen.

[0052] Affinity reagents can be used, for example, to identify or quantify binding events (e.g., fluorescently identify binding events). The identifiable tag may be tagged with an identifiable tag that allows for optical detection. Some non-limiting examples include: fluorophores, magnetic nanoparticles, or nucleic acid bars. The fluorophores used are GFP, YFP, RFP, eGFP, mCher ry, tdtomato, FITC, Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 488, Alexa Flu or 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alex a Fluor 647, Alexa Fluor 680, Alexa Fluor 750, Pacific Blue, Coumarin, BODIPY FL , Pacific Green, Oregon Green, Cy3, Cy5, Pacific Orange, TRITC, Texas Red, It may also contain fluorescent proteins such as coerythrin and allophycocyanin. The affinity reagent may be used to detect binding events, e.g., surface plasmon resonance (SPR) binding events. ) detection, the tag may be removed, such as when detected directly.

[0053] Examples of detection moieties include, but are not limited to, fluorophores, bioluminescent proteins, , a nucleic acid segment comprising an invariant region and a barcode region, or a nanoparticle such as a magnetic particle. For example, the affinity reagent may be tagged with a DNA barcode. These may then be unambiguously sequenced at those locations. As another example, a set of different fluorophores can be used for fluorescence resonance energy transfer (FRET) detection. The detection moiety may be a compound having different excitation or emission properties. The fluorescent dye may contain several different fluorophores, with a pattern that varies.

[0054] The detection moiety may be cleavable from the affinity reagent, which allows the detection moiety to be cleaved from the affinity reagent, where the detection moiety is no longer of interest. Signal contamination can be reduced by a step in which the detection moiety is removed from the affinity reagent that is not present. This may enable the following.

[0055] In some cases, the affinity reagent is unmodified. For example, if the affinity reagent is an antibody If the affinity reagent is an antibody, the presence of the antibody may be detected by atomic force microscopy. The antibody may be a modification and may be specific for one or more of the affinity reagents. For example, if the affinity reagent is a mouse antibody, If present, mouse antibodies may be detected using an anti-mouse secondary antibody. The reagent may be an aptamer that is detected by an antibody specific for the aptamer. The antibody may be modified with a detection moiety as described above. In some cases, a secondary antibody The presence of may be detected by atomic force microscopy.

[0056] In some instances, the affinity reagent may be modified with the same modification, e.g., a conjugated green It may contain a fluorescent protein or may contain two or more different modifications, for example: Each affinity reagent may contain several different fluorophores, each with a different excitation or emission wavelength. Several different affinity reagents may be combined. This may allow for multiplexing of affinity reagents, as multiple affinity reagents may be identified and / or distinguished. In one example, the first affinity reagent may be conjugated to green fluorescent protein; The second affinity reagent may be conjugated to a yellow fluorescent protein, and the third affinity reagent Drugs may be conjugated to red fluorescent proteins, thus enabling these three affinity assays. Drugs can be multiplexed and identified by their fluorescence. The second, fourth, and seventh affinity reagents may be conjugated to green fluorescent protein, and the third, fourth, and seventh affinity reagents may be conjugated to green fluorescent protein. the fifth, fifth, and eighth affinity reagents may be conjugated to yellow fluorescent protein; and The third, sixth, and ninth affinity reagents may be conjugated to a red fluorescent protein; In this case, the first, second, and third affinity reagents may be multiplexed together, while the second, fourth, and the seventh affinity reagent, as well as the third, sixth, and ninth affinity reagents, are two further Form a multiplexed reaction. The number of affinity reagents that can be multiplexed together depends on the number of For example, the detection moiety used for the fluorophore-labeled parent molecule may vary. Multiplexing of compatibility reagents can be limited by the number of unique fluorophores available. As an example, multiplexing of affinity reagents labeled with nucleic acid tags can be performed using a method that scales with the length of the nucleic acid barcode. The nucleic acid may be deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). obtain.

[0057] The specificity of each affinity reagent can be determined prior to use in an assay. The binding specificity can be determined in control experiments using known proteins. Experimental methods may be used to determine the specificity of the affinity reagent. Load known protein standards at known locations to assess the specificity of multiple affinity reagents In another example, the specificity of each affinity reagent may be determined by binding to a control and a standard. The substrate is then used to identify the experimental sample so that the In some cases, the panel may include both a sample and a panel of controls and standards. Affinity reagents of known specificity may be included along with affinity reagents of known specificity, Data from affinity reagents of known specificity may be used to identify proteins, The binding patterns of affinity reagents of unknown specificity to the proteins to be identified are The binding specificity of each protein may be determined by the affinity reagent. Use known binding data of other affinity reagents to assess whether any individual It is also possible to reconfirm the specificity of the affinity reagent. The frequency of binding of the affinity reagent to each known protein conjugated to the substrate is can be used to derive the probability of binding to any of the above proteins. In some cases, a known protein containing the epitope (e.g., amino acid sequence or post-translational modification) is used. The frequency of binding to a protein determines the probability of binding of an affinity reagent to a particular epitope. Therefore, the use of multiple affinity reagent panels can improve the characteristics of affinity reagents. The isomerism can be refined with each iteration. Uniquely specific for a particular protein. Although various affinity reagents may be used, the methods described herein do not require them. In addition, the method may be effective over a range of specificities. In examples, the methods described herein may be used to identify any particular protein. They are not specific for proteins, but instead contain amino acid motifs (e.g., tripeptides). It can be particularly effective when it is specific for AAA.

[0058] In some instances, the affinity reagent has high, intermediate, or low binding affinity. In some cases, a parent molecule with low or intermediate binding affinity may be selected. In some cases, the affinity reagent may be selected to be about 10 -3 M, 10 -4 M, 10 -5 M, 10 -6 M, 10 -7 M, 10 -8 M, 10 -9 M, 10 -10 M's, or about 10 -10 Dissociation constant smaller than M In some cases, the affinity reagent may have a number of about 10 -10 M, 10 -9 M, 10 -8 M , 10 -7 M, 10 -6 M, 10 -5 M, 10 -4 M, 10 -3 M, 10 -2 More than M or 10 -2 Larger than M It may have a dissociation constant, in some cases a low or intermediate k off Fast or Medium Intermediate or high k on Rapid affinity reagents may be preferred.

[0059] Some affinity reagents contain amino acid sequences such as phosphorylated or ubiquitinated amino acid sequences. In some instances, the modified amino acid sequence may be selected to bind to the modified amino acid sequence. The species or species of affinity reagent may comprise an epitope that may be carried by one or more proteins. The nucleotide sequence may be chosen to be broadly specific for a family of topes. Thus, one or more affinity reagents may bind to two or more different proteins. In some instances, one or more affinity reagents are selectively bound to their target(s). For example, the affinity reagent may bind less than 10%, less than 10%, less than 15%, less than 20%, less than 25% Less than 30%, or less than 35% may bind to their target(s). In this example, one or more affinity reagents have a moderate affinity for their target or targets. For example, the affinity reagent may bind more than 35%, more than 40%, more than 45%, more than 60%, more than 65% More than %, More than 70%, More than 75%, More than 80%, More than 85%, More than 90%, More than 91%, More than 92%, More than 93%, More than 94%, More than 95%, More than 96% , greater than 97%, greater than 98%, or greater than 99% may bind to their target(s).

[0060] To compensate for weak binding, an excess of affinity reagent may be applied to the substrate. Approximately 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1 or 10:1 to the sample protein The affinity reagent may be applied in excess of about 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1 ​​or 10:1, which is in excess of the expected occurrence of the epitope in the sample protein. may be applied excessively.

[0061] To compensate for the fast dissociation rate of the affinity reagents, a linker moiety is attached to each affinity reagent. and a linker moiety can be provided to connect the bound affinity reagent to the substrate to which it binds. For example, a DNA tag may be used to reversibly link a target protein to a target gene or unknown protein. A different DNA tag can be attached to the end of each affinity reagent, and a different DNA tag can be attached to the substrate or each unknown tag. After the affinity reagent hybridizes to the unknown protein, the Anchor DNA, which is complementary at one end to the DNA tag that is attached to the affinity reagent and is complementary at the other end to a tag attached to the substrate, can be added to the chip to bind affinity reagents to the substrate, which The reagents are prevented from dissociating prior to measurement. After binding, the linked affinity reagents are attached to the DNA linker. The proteins can be released by washing in the presence of heat or high salt concentrations to disrupt the bond.

[0062] FIG. 21 illustrates a method for enhancing binding between an affinity reagent and a protein, according to some embodiments. In particular, step 1 in Figure 21 shows the hybridization of the affinity reagent. As shown in step 1, the affinity reagent 2110 binds to the protein. The protein 2130 hybridizes to the slide 2105. The protein 2130 is bound to the slide 2105. As shown in Figure 1, the affinity reagent 2110 has a DNA tag 2120 attached. In this embodiment, the affinity reagent may have multiple DNA tags attached. In some embodiments, the affinity reagent may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more than 20 The DNA tag 2120 may have a DNA tag attached thereto. The DNA tag 2120 may be a single-stranded DNA having a recognition sequence 2125. In addition, the protein 2130 contains two DNA tags 2140. In this embodiment, the DNA tag is attached using a chemical reaction that reacts with cysteines in the protein. In some embodiments, the protein may contain multiple bound DNA tags. In some embodiments, the protein may have 1, 2, 3, 4, 5, 6 pieces, 7 pieces, 8 pieces, 9 pieces, 10 pieces, 11 pieces, 12 pieces, 13 pieces, 14 pieces, 15 pieces, 16 pieces, 17 pieces, 18 pieces, 19 pieces, 2 0 pieces, 25 pieces, 30 pieces, 35 pieces, 40 pieces, 45 pieces, 50 pieces, 55 pieces, 60 pieces, 65 pieces, 70 pieces, 75 pieces, 80 pieces, 85 Each DNA may have 90, 95, 100, or more than 100 DNA tags attached. The tag 2140 comprises a ssDNA tag having a recognition sequence 2145 .

[0063] As shown in step 2, the DNA linker 2150 binds to the affinity reagent 2110 and the protein 2130. The DNA linker 21 hybridizes to the attached DNA tags 2120 and 2140, respectively. 50 contains ssDNA having sequences complementary to recognition sequences 2125 and 2145, respectively. As shown in step 2, the recognition sequences 2125 and 2145 are linked to the DNA tag 2120 by the DNA linker 2150. and 2140 simultaneously to the DNA linker 2150 In particular, the first region 2152 of the DNA linker 2150 selectively hybridizes with the recognition sequence 2125. and the second region 2154 of the DNA linker 2150 selectively hybridizes with the recognition sequence 2145. In some embodiments, the first region 2152 and the second region 2154 are In particular, in some embodiments, the DNA linker The first region of the DNA linker and the second region of the DNA linker are located between the first region and the second region. The sequences may be spaced apart by non-hybridizing spacer sequences. In some embodiments, the sequence of the recognition sequence is not perfectly complementary to the DNA linker. However, it may still be possible to bind to a DNA linker sequence. In the present study, the length of the recognition sequence is less than 5 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, or Nucleotide, 8th nucleotide, 9th nucleotide, 10th nucleotide, 11th nucleotide, 12th nucleotide nucleotide, 13 nucleotide, 14 nucleotide, 15 nucleotide, 16 nucleotide, 17 nucleotide nucleotide, 18 nucleotide, 19 nucleotide, 20 nucleotide, 21 nucleotide, 22 nucleotide nucleotide, 23 nucleotide, 24 nucleotide, 25 nucleotide, 26 nucleotide, 27 nucleotide nucleotide, 28 nucleotide, 29 nucleotide, or 30 nucleotide, or In some embodiments, the recognition sequence is directed against a complementary DNA tag sequence. In some embodiments, the recognition sequence may have one or more mismatches. Approximately one of the nucleotides may be mismatched with respect to the complementary DNA tag sequence. and still be capable of hybridizing to a complementary DNA tag sequence. The recognition sequence has a mismatch with the complementary DNA tag sequence by less than 1 in 10 nucleotides. However, they can still hybridize to a complementary DNA tag sequence. In this manner, the recognition sequence is such that approximately 2 out of 10 nucleotides are complementary to the DNA tag sequence. may be mismatched but still be able to hybridize to a complementary DNA tag sequence. In some embodiments, the recognition sequence has more than two of the ten nucleotides that are complementary to the DNA tag. The sequence may be mismatched but still hybridize to the complementary DNA tag sequence. It can be easily identified.

[0064] The affinity reagent may also include a magnetic component. Scan some or all of the bound affinity reagents in the same image plane or z-stack It may be useful to operate some or all affinity reagents in the same image plane. This can improve the quality of the imaging data and reduce noise in the system.

[0065] The term "detector," as used herein, generally refers to a device that detects a signal. The present invention refers to an apparatus capable of detecting a signal of a binding event of an affinity reagent to a protein. The signal includes a signal indicating the presence or absence of a target molecule. The signal may be a surface plasmon resonance (SPR) signal. The signal may be a direct signal indicating the presence or absence of a binding event, such as a null. A signal is an indirect signal that indicates the presence or absence of a binding event, such as a fluorescent signal. In some cases, the detector may be capable of detecting a signal. The term "detector" refers to a detection method that may include optical and / or electronic components. Non-limiting examples of detection methods include optical detection, spectroscopic detection, electrostatic detection, and the like. pneumatic detection, electrochemical detection, magnetic detection, fluorescence detection, surface plasmon resonance (SPR), etc. Examples of optical detection methods include, but are not limited to, fluorometry and ultraviolet-visible light absorbance. Examples of spectroscopic detection methods include, but are not limited to, mass spectrometry, nuclear magnetic resonance spectroscopy, and the like. Examples of electrostatic detection methods include, but are not limited to, nuclear magnetic resonance (NMR) spectroscopy, and infrared spectroscopy. Examples of electrochemical detection methods include, but are not limited to, gel-based techniques such as gel electrophoresis. Examples of the amplification method include, but are not limited to, amplification of the amplified product after high performance liquid chromatography separation. Electrochemical detection of the product is included.

[0066] Protein identification in samples Proteins are important components of the cells and tissues of living organisms. An organism produces a large set of different proteins, typically referred to as a proteome. The proteome can change over time and reflect the various processes that a cell or organism experiences. It may also vary as a function of stage (e.g., cell cycle stage or disease state). Large-scale studies or measurements of the proteome (e.g., experimental analysis) are called proteomics. In proteomics, multiple methods for identifying proteins are available. The method is based on immunoassays (e.g., enzyme-linked immunosorbent assays (ELISAs)). and Western blot), mass spectrometry-based methods (e.g., matrix-assisted laser -desorption ionization (MALDI) and electrospray ionization (ESI), hybrid methods (e.g., mass spectrometry immunoassays (MSIA)), as well as protein microarrays. For example, single-molecule proteomics approaches have evolved from direct functionalization of amino acids to affinity The identity of protein molecules in a sample can be determined by a variety of approaches, ranging from the use of reagents to Information gathered from such an approach may be used to estimate the Alternatively, the measurement may be performed using a suitable assay to identify proteins present in the sample. It is analyzed by algorithm.

[0067] Accurate quantification of proteins is also difficult due to lack of sensitivity, lack of specificity, and detector noise. In particular, accurate quantification of proteins in a sample can be challenging. and random and random variations in detector signal levels that can cause errors in quantification. Challenges can be faced due to unpredictable systemic fluctuations. The equipment and detection schemes should provide equipment diagnostics and monitor common-mode behavior. However, protein binding (e.g. Affinity reagent probes inherently offer less than ideal sensitivity and specificity of binding. It is a stochastic process that can have

[0068] The present disclosure provides methods and systems for accurate and efficient identification of proteins. The methods and systems provided herein identify proteins in a sample. Such a method and system can significantly reduce or eliminate errors in determining the Thus, accurate and efficient identification of candidate proteins in samples of unknown proteins can be achieved. Protein identification is the calculation of unknown proteins in a sample using information from empirical measurements. For example, empirical measurements can be based on the selection of one or more candidate proteins. the binding information of the affinity reagent probes that are designed to specifically bind to the protein; the length of the protein; Protein identification may include the hydrophobicity and / or isoelectric point of the protein. Protein identification can be performed using one or more may include an estimate of the confidence level that each of the plurality of candidate proteins is present in the sample.

[0069] In one aspect, the present invention provides a method for identifying proteins in a sample of unknown proteins. A computer-implemented method 100 (e.g., shown in FIG. 1) for determining Disclosed is a method for analyzing a sample to generate a population of proteins identified in the sample. The amount of protein can be independently applied to each unknown protein in the candidate protein. The number of proteins can be calculated by counting the number of identifications for each protein. The method for identification involves combining information from multiple empirical measurements of unknown proteins in a sample. The empirical measurement may include receiving by a computer (e.g., step 105). i) one or more affinity reagents directed against one or more unknown proteins in the sample; (ii) binding measurements for each of the probes; and (iii) binding measurements for one or more unknown proteins. (iii) the hydrophobicity of one or more unknown proteins; and / or (iv) one Alternatively, it may include the isoelectric points of multiple unknown proteins. The affinity reagent probes may comprise a pool of multiple individual affinity reagent probes. For example, the affinity reagent probe pool may be two, three, four, five, six, or more. , 7, 8, 9, 10, or more than 10 affinity reagent probes. In some embodiments, the pool of affinity reagent probes may comprise two types of affinity reagents. The combination may comprise an affinity reagent probe in a pool of affinity reagent probes. In some embodiments, the pool of affinity reagent probes comprises the majority of the composition of the probes. The combination may include three types of affinity reagent probes, In some embodiments, the affinity reagent comprises the majority of the composition of the affinity reagent probes in the The drug probe pool may include four affinity reagent probes, the combination of which may be The affinity reagent probes make up the majority of the composition of the affinity reagent probes in the pool. In one embodiment, the pool of affinity reagent probes may include five different affinity reagent probes. The combination comprises a majority of the composition of the affinity reagent probes in the pool of affinity reagent probes. In some embodiments, the pool of affinity reagent probes comprises more than five affinity reagents. The combination may comprise an affinity reagent in a pool of affinity reagent probes. Each of the affinity reagent probes is designed to bind to multiple candidate proteins. The affinity of the target protein can be determined by the affinity of the target protein. The affinity reagent probe can be a k-mer affinity reagent probe. Each affinity reagent probe of the mer is directed to one or more of the candidate proteins. The information from the empirical measurements is used to identify unknown proteins. The assay may involve measuring the binding of a set of probes believed to be bound to a protein.

[0070] Next, at least a portion of the information from the empirical measurement of the unknown protein is used to identify multiple proteins. The protein sequences can be compared by computer to a database containing the protein sequences (e.g., Step 110). Each of the protein sequences is a candidate protein among a plurality of candidate proteins. The plurality of candidate proteins may correspond to at least 10, at least 20, or at least At least 30, at least 40, at least 50, at least 60, at least 70, At least 80, at least 90, at least 100, at least 150, at least 200, At least 250, at least 300, at least 350, at least 400, at least 450 pieces, at least 500 pieces, at least 600 pieces, at least 700 pieces, at least 800 pieces, at least The candidate protein set may include as few as 900, at least 1000, or more than 1000 different candidate proteins.

[0071] Next, for each of one or more candidate proteins among the plurality of candidate proteins, , the probability that an empirical measurement on a candidate protein will produce the observed measurement outcome is , may be calculated or generated by a computer (e.g., in step 115). The term "measured outcome," as used herein, refers to the outcome observed when a measurement is made. For example, the measured outcome of an affinity reagent binding experiment may be the binding or This can be a positive or negative outcome, such as either non-binding or non-binding. The measured outcome of the experiment measuring the length of the protein could be 417 amino acids. Alternatively, each of one or more candidate proteins among the plurality of candidate proteins For example, empirical measurements on a candidate protein must produce an observable measurement outcome. The probability may be calculated or generated by a computer. The probability that an empirical measurement on a candidate protein will produce an unobserved measurement outcome is Additionally or alternatively, the candidate tanks may be computed or generated by a computer. The probability that a set of empirical measures of quality will produce a set of outcomes is calculated by a computer. It may be calculated or generated by the

[0072] "Outcome set," as used herein, refers to a set of outcomes associated with a protein. Refers to multiple independent measured outcomes that correlate with one another. For example, a series of empirical affinity reagent binding Measurements can be performed on the unknown protein. Binding measurements for each individual affinity reagent contains the measured outcomes, and the set of all measured outcomes is called the outcome set. In some cases, the outcome set is a set of all observed outcomes. In some cases, the outcome set may be a subset of the may consist of unobserved measurement outcomes. Additionally or alternatively, multiple candidate For each of one or more candidate proteins in the protein The probability that is a candidate protein can be calculated or generated by a computer. The calculation or generation of 115 and / or 120 may be performed iteratively or non-iteratively. The probability in step 115 is calculated by dividing all the empirical measurement outcomes of the unknown protein by Based on comparison of candidate proteins against a database containing multiple protein sequences Thus, the input to the algorithm is the data of the candidate protein sequences. base, as well as empirical measurements of unknown proteins (e.g., The probes likely bound to the protein, the length of the unknown protein, and the hydrophobicity of the unknown protein. In some cases, the set of In this case, the input to the algorithm is any of the affinity reagents and any of the candidate proteins. The probability of generating any binding measurement with respect to the number of affinity reagents (e.g., the number of trimeric The output of the algorithm may include parameters related to estimating the probability of a given combination of the two. (i) Given the identity of a hypothesized candidate protein, the measured (ii) the probability that the measured outcome or outcome set is observed; Given a set of cams, we can find a set of candidate proteins for an unknown protein. The most likely identity selected from the list, and the probability that the identification is correct (e.g., (e.g., in step 120), and / or (iii) identifying high-probability candidate proteins. A group of proteins, and the unknown protein is one of the proteins in the group. The probability that the candidate protein is the measured protein may be included. and the probability of observing the measured outcome can be expressed as: P(measured outcome | protein)

[0073] In some embodiments, P(measured outcome | protein) is calculated entirely in silico. In some embodiments, P(measured outcome | protein) is the protein It is calculated based on or derived from the characteristics of the amino acid sequence. In this study, P(measured outcome | protein) is independent of knowledge of the amino acid sequence of the protein. For example, P(measured outcome | protein) is calculated independently for each protein candidate. Measurements were obtained in replicate experiments in the same sample, and P(measured outcome | tamper By calculating the quality of the outcome from the frequency: (number of measurements with the outcome / total number of measurements) , can be empirically determined. In some embodiments, P(measured outcome | protein) is , derived from a database of past measurements for the protein. , P(measurement outcome | protein) is the unknown protein with a censored measurement result. From this population, a set of confident protein identifications is generated, and then candidate proteins are identified. The measured output of a set of unknown proteins that were reliably identified as proteins In some embodiments, the unknown tongue is calculated by calculating the frequency of the tongue. A protein population can be identified using a seed value of P(measured outcome | protein), and The seed value is the number of measured values ​​among unknown proteins that reliably match the candidate protein. The process can be refined based on the frequency of the outcomes. In some embodiments, the process Repeat with new identifications generated based on updated probabilities of measured outcomes. The probability of a new measured outcome is then increased to increase the confidence of the identification. Generated from a dated set.

[0074] If the candidate protein is the protein being measured, the measured outcome is observed. The probability of not being hit can be expressed as: P(non-measured outcome | protein) = 1 - P(measured outcome | protein)

[0075] If the candidate protein is the protein being measured, then N individual measurement outputs The probability of observing a set of measured outcomes consisting of the outcome measures is can be expressed as a product of the probabilities for P(outcome set | protein) = P(measured outcome 1 | protein) * P(measured outcome Tocam2 | Protein) * … * P(Measured OutcomeN | Protein)

[0076] The unknown protein is a candidate protein (protein i ) is the probability that a possible candidate A calculation can be made based on the probability of a set of outcomes for each protein.

[0077] In some embodiments, the set of measured outcomes includes binding of affinity reagent probes In some embodiments, the set of measured outcomes includes non-specific binding of affinity reagent probes. This includes cases where

[0078] In some embodiments, the proteins in the sample are cleaved or degraded. In some embodiments, the protein in the sample contains the C-terminus of the original protein. In some embodiments, the proteins in the sample do not contain the N-terminus of the original protein. In some embodiments, the proteins in the sample do not contain the N-terminus of the original protein. First, it does not contain the C-terminus of the original protein.

[0079] In some embodiments, empirical measurements include measurements performed on a mixture of antibodies. In some embodiments, empirical measurements are performed on samples containing proteins from multiple species. In some embodiments, empirical measurements include measurements performed in human-derived samples. In some embodiments, empirical measurements include measurements performed on samples containing human In some embodiments, the present invention includes measurements performed on samples from species other than the The quantitative measurement of single amino acid variation (SAV) caused by nonsynonymous single nucleotide polymorphisms (SNPs) is In some embodiments, the empirical measurement includes a measurement performed on a sample in the presence of The term refers to insertions, deletions, translocations, inversions, and segmental deletions that affect the sequence of proteins in a sample. in the sample in the presence of genomic structural variants such as duplications or copy number variations (CNVs). Includes measurements.

[0080] In some embodiments, the method comprises: In some embodiments, the method further comprises applying the method to one or more For each of the candidate proteins, determine whether the candidate protein is a protein of interest in the unknown protein to be measured in the sample. The method further includes generating a confidence level that matches the protein. The confidence level is a probability value. Alternatively, the confidence level may include a probability value of having an error. Alternatively, the confidence level may be a measure of a certain degree of confidence (e.g., about 90%, about 95%, about 96%, about 97%, about 98%, approximately 99%, approximately 99.9%, approximately 99.99%, approximately 99.999%, approximately 99.9999%, approximately 99.99999% , approximately 99.999999%, approximately 99.9999999%, approximately 99.99999999%, approximately 99.99999999%, approximately 99.99999 99999%, approximately 99.99999999999%, approximately 99.999999999999%, approximately 99.999999999999% confidence The probability may include a range of probability values, optionally with a confidence level of 99.9999999999999% or greater. do.

[0081] In some embodiments, the method generates a probability that a candidate protein is present in a sample. The method further includes the steps of:

[0082] In some embodiments, the method comprises determining protein identifications, and associated probabilities, based on the number of proteins in a sample. For each unknown protein, a step of independently generating the unknown protein and identifying the unknown protein in the sample. In some embodiments, the method further comprises generating a list of all unique proteins found. The method comprises: detecting unique candidate proteins to determine the amount of each candidate protein in the sample; The method further includes counting the number of identifications generated for each protein. In this case, the collection of protein identifications and associated probabilities is given a high score, a high confidence, and The data may be filtered to include only identifications with low false discovery rates and / or low false discovery rates.

[0083] In some embodiments, the binding probability is determined for an affinity reagent to a full-length candidate protein. In some embodiments, binding probabilities can be generated by dividing a protein fragment (e.g., a complete Affinity reagents can be generated for the target sequence (subsequences of the entire protein sequence). For example, Only the first 100 amino acids of each unknown protein are conjugated. In this manner, when an unknown protein is treated and conjugated to a substrate, the first 10 All binding probabilities for epitope binding other than 0 amino acid are zero, or The binding probability is set to a very low probability representing the error rate, and the binding probability is calculated for each protein candidate. A similar approach can be used to generate the first 10, 20, or 30 amino acids of each protein. Amino acids, 50 amino acids, 100 amino acids, 150 amino acids, 200 amino acids, 300 amino acids, 400 amino acids amino acids, or more than 400 amino acids, can be used when conjugated to a substrate. A similar approach was used to sequence the last 10, 20, 50, 100, and 150 amino acids. amino acids, 200 amino acids, 300 amino acids, 400 amino acids, or more than 400 amino acids are contained in the substrate. It can be used when adjuvanted.

[0084] In some embodiments, one protein candidate as one half of a match pair is an unknown protein. If the protein cannot be assigned to a specific protein, the potential protein candidate is selected as the matched pair. A complementary group can be assigned to the unknown protein. The confidence level is determined by the probability of the protein being in the group. The unknown protein can be assigned to one of the protein candidates. The confidence level may include a probability value. Alternatively, the confidence level may include a probability value of having an error. Alternatively, the confidence level may be a measure of a certain degree of confidence (e.g., about 90%, about 95%, about 96%, about 97%, about 98%, approximately 99%, approximately 99.9%, approximately 99.99%, approximately 99.999%, approximately 99.9999%, approximately 99.99999% , approximately 99.999999%, approximately 99.9999999%, approximately 99.99999999%, approximately 99.999 9999999% confidence, approximately 99.99999999999% confidence, approximately 99.999999999999% confidence, approximately 99.9999999999999% confidence A range of probability values, optionally with a confidence level of greater than 99.9999999999999% For example, an unknown protein may be strongly matched to two protein candidates. The protein candidates may have high sequence similarity to each other (e.g., two protein candidates may have high sequence similarity to each other). isoforms, e.g., proteins with a single amino acid mutation compared to the canonical sequence In these cases, individual protein candidates can be assigned with high confidence. However, high confidence is given to two protein candidates that are strongly matched. A single, but unknown, member of a "protein group" containing the This may be due to a known protein.

[0085] In some embodiments, to detect a state in which unknown proteins are optically indistinguishable, For example, in the rare event that two or more proteins are present on the substrate, The possibility of binding to the same "well" or location of the same molecule is also a challenge. Regardless, in some cases, the conjugated protein is a non-specific The two may be treated with different dyes and the signal from the dyes may be measured. In situations where these proteins are optically indistinguishable, the signal generated by the dye The binding can be stronger than a site containing one protein, and multiple proteins bind. This can be used to indicate the location of the

[0086] In some embodiments, the plurality of candidate proteins is selected from a sample of unknown proteins. Human or biological DNA or RNA obtained or derived from They are generated or modified by analysis or analysis.

[0087] In some embodiments, the method derives information about post-translational modifications of unknown proteins. The information about post-translational modifications may be obtained by extracting information about the nature of the particular modification. The database can be considered as an exponential product of PTMs. For example, once a candidate protein sequence has been assigned to an unknown protein, it can be assayed. The pattern of affinity reagent binding for the selected protein was compared with that for the same candidate from previous experiments. The data can be compared to a database containing binding assays of affinity reagents. For example, The database is a nucleic acid programmer containing unmodified proteins of known sequence at known locations. Binding to Nucleic Acid Programmable Protein Array (NAPPA) It may be derived from.

[0088] Additionally or alternatively, the database of binding measurements may be used to identify proteins for which candidate protein sequences are unknown. The assayed protein may be derived from previous experiments that have been reliably assigned to the protein. The discrepancy in binding measurements between the proteins tested and the database of existing measurements is due to post-translational For example, affinity agents may be added to a database to provide information about the likelihood of modification. In this study, we found that the nucleotide sequence of the nucleotides that bind to the candidate protein frequently but do not bind to the assayed protein. If not, there is a higher likelihood that a post-translational modification is present somewhere on the protein. For affinity reagents for which there is a binding discrepancy, if the binding epitope is known, In this case, the location of the post-translational modification is at or near the binding epitope of the affinity reagent. In some embodiments, information about a particular post-translational modification may be Before treating the protein-substrate conjugate with an enzyme that specifically removes certain post-translational modifications and later by performing repeated affinity reagent measurements. For example, binding measurements can be performed for a series of affinity reagents prior to treatment of the substrate with phosphatase. This may be obtained and then repeated after treatment with phosphatase. Before phosphatase treatment, it binds to an unknown protein, but after phosphatase treatment, it binds to an unknown protein. Affinity reagents that do not bind (differentially) can provide evidence of phosphorylation. If the epitope recognized by the affinity reagent is known, phosphorylation can be performed by It may be located at or near the site of a binding epitope for the drug.

[0089] In some cases, the number of specific post-translational modifications may be related to the affinity for a specific post-translational modification. This can be determined using binding assays using reagents that recognize phosphorylation events, for example. Antibodies may be used as affinity reagents. The binding of these reagents is In some cases, the presence of at least one phosphorylation of an unknown protein may be observed. The number of distinct post-translational modifications of a particular type in a protein is specific for that particular post-translational modification. The binding affinity can be determined by counting the number of binding events measured for a specific affinity reagent. For example, a phosphorylation-specific antibody may be conjugated to a fluorescent reporter. In this case, the intensity of the fluorescent signal is determined by the amount of phospho-specific affinity reagent bound to the unknown protein. can be used to determine the number of phosphorylation-specific affinity bound to an unknown protein. The number of reactive reagents is then used to determine the number of phosphorylation sites in the unknown protein. In some embodiments, more precise number, identity, or location of post-translational modifications can be used. To derive this, evidence from affinity reagent binding experiments has been used to support the possibility of post-translational modifications. Existing knowledge of specific amino acid sequence motifs or specific protein locations (e.g. For example, from dbPTM, PhosphoSitePlus, or UniProt. For example, if the location of a post-translational modification cannot be precisely determined from affinity measurements alone, Positions containing amino acid sequence motifs often associated with post-translational modifications of the polypeptide may be advantageous. do.

[0090] In some embodiments, the probabilities are generated iteratively until a predetermined condition is met. In some embodiments, the predetermined condition is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, At least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least At least 99.9%, at least 99.99%, at least 99.999%, at least 99.9999%, at least 99.9 9999%, at least 99.999999%, at least 99.9999999%, at least 99.99999999%, at least at least 99.999999999%, at least 99.9999999999%, at least 99.99999999999%, at least with at least 99.9999999999999% confidence, or 99.9999999999999% confidence. generating each of the plurality of probabilities with greater than 9% confidence.

[0091] In some embodiments, the method comprises identifying one or more unknown proteins in a sample. The method further includes generating a paper or electronic report. For each of the candidate proteins, the rate is a probability that the candidate protein is present in the sample. A confidence level may further be indicated. The confidence level may comprise a probability value. Alternatively, the confidence level may be Alternatively, the confidence level may include a certain degree of confidence (e.g., about 90%). of, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.9%, about 99.99%, about 99.999 %, approximately 99.9999%, approximately 99.99999%, approximately 99.999999%, approximately 99.9999999%, approximately 99.9999999 %, approximately 99.999999999%, approximately 99.9999999999%, approximately 99.99999999999%, approximately 99.9999999999 99%, approximately 99.9999999999999% confidence, or greater than 99.9999999999999% confidence) A paper or electronic report may include a range of probability values ​​that indicate the expected false positives. Detection rate threshold (e.g., <10%, <9%, <8%, <7%, <6%, <5%, <4%) , less than 3%, less than 2%, less than 1%, less than 0.5%, less than 0.4%, less than 0.3%, less than 0.2%, or less than 0.1% A list of protein candidates that are identified below a false discovery rate (false discovery rate) may further be presented. The rate can be estimated by first sorting the protein identifications in descending order of confidence The estimated false discovery rate at any point in the sorted list is then 1 - av g_c_prob, where avg_c_prob is the current point in the list or a higher point. The average candidate probability for all proteins before (e.g., higher confidence) The list of protein identifications that fall below a desired false discovery rate threshold is then sorted All points in the given list before the earliest point for which the false discovery rate is higher than the threshold Alternatively, a desired false discovery rate threshold can be set by The list of protein identifications below the desired false discovery rate in the sorted list All previous tangents, including the most recent point, that are below or equal to the threshold It can be produced by reconstituting the protein.

[0092] In some embodiments, the sample comprises a biological sample. A biological sample is obtained from a subject. In some embodiments, the method comprises determining the subject based at least on a plurality of probabilities. In some embodiments, the method further comprises identifying a disease state or disorder in the subject. The method identifies protein candidates by counting the number of identifications generated for each protein candidate. For example, the absolute amount of protein present in a sample (e.g., The number of protein molecules (e.g., the number of protein molecules) is the number of confident identifications generated from protein candidates. In some embodiments, the amount can be calculated by counting the number of It can be calculated as a percentage of the total number of unknown proteins. Therefore, raw identification numbers are used to eliminate systematic errors from the instrument and detection system. In some embodiments, the amount can be calibrated by the variation in detectability of the protein candidate. The detectability of a protein can be calculated by the following equation: It can be estimated from experimental measurements or computer simulations.

[0093] The disease or disorder may be an infectious disease, an immune disorder or disease, cancer, a genetic disease, a degenerative disorder, The disease may be a disease related to lifestyle, wound, rare disease, or ageing. It can be caused by bacteria, viruses, fungi, and / or parasites. Non-limiting examples include bladder cancer, lung cancer, brain cancer, melanoma, breast cancer, non-Hodgkin's lymphoma, and pediatric cancer. Cervical cancer, ovarian cancer, colon and rectal cancer, pancreatic cancer, esophageal cancer, prostate cancer, kidney cancer, Includes skin cancer, leukemia, thyroid cancer, liver cancer, and uterine cancer. Genetic or inherited diseases Some examples of genetic disorders include, but are not limited to, multiple sclerosis (MS), cystic fibrosis (CF) and leukemia (LE). fibrosis, Charcot-Marie-Tooth disease, Huntington's disease, Peutz-Jeghers syndrome, Lifestyle-related diseases include, but are not limited to, Down syndrome, rheumatoid arthritis, and Tay-Sachs disease. Examples include obesity, diabetes, arteriosclerosis, heart disease, stroke, high blood pressure, cirrhosis of the liver, nephritis, cancer, and chronic Some examples of injuries include obstructive pulmonary disease (COPD), hearing problems, and chronic back pain. Although not specified, abrasions, brain injuries, contusions, burns, concussions, congestive heart failure, Injuries at construction sites, dislocations, flail chest, fractures, hemothorax, herniated discs, hip pointers, low body fever, lacerations, compressed nerves, pneumothorax, rib fractures, sciatica, spinal cord injuries, tendons, ligaments , fascial injuries, traumatic brain injury, and whiplash.

[0094] In some embodiments, the method comprises administering to a subject a small molecule rather than or in addition to a protein. The method also includes identifying and quantifying the molecules (e.g., metabolites) or glycans. For example, lectins or antibodies that bind to sugars or combinations of sugars with diverse properties. Affinity reagents such as those listed above may be used to identify glycans. The nature of the affinity reagent binding to the combination was determined by analyzing its binding to a commercially available glycan array. For example, unknown glycans can be characterized by the presence of hydroxyl-reactive groups. Using chemical reactions, the glycans can be conjugated to functionalized substrates and binding measurements can be performed using glycans. The affinity reagents can be used to bind to unknown glycans on the substrate. Drug binding measurements are based on the number of glycans with a particular sugar or a particular combination of sugars. Alternatively, the structure of each unknown glycan can be determined by To identify, one or more binding assays are performed using the methods described herein. The binding measurements can then be compared with those predicted from a database of candidate glycan structures. In some embodiments, the protein is bound to a substrate and the glycan affinity reagent is used. Binding measurements are generated to identify glycans attached to the protein. The measurement allows identification of the protein backbone sequence and conjugated glycans in a single assay. Both glycan affinity reagents and protein affinity reagents were used to generate As another example, sulfhydryl, carbonyl, amine, or active hydrogen may be used. which function using chemical reactions that target coupling groups commonly found in metabolites. Metabolites can be conjugated to modified substrates. Binding measurements can be performed by measuring specific functional groups, structural moieties, and This may be done using affinity reagents with different properties for the chemotherapeutic or metabolite. The resulting binding measurements are compared with predicted binding measurements for a database of candidate small molecules. and the method described herein can be compared to the measurement at each location on the substrate. It can be used to identify metabolites. [Example]

[0095] Example 1: Protein Identification by Affinity Reagent Binding The methods described herein are directed to analyzing and / or identifying proteins in a sample. In order to measure the binding of affinity binding reagents (e.g., aptamers or antibodies), In this case, the probability of the measured outcome to be calculated is the probability of the protein candidate A binding event or event of an affinity binding reagent (e.g., affinity reagent or affinity probe) to is the probability of a non-binding event. The binding probability is the probability of a protein being recognized by an affinity binding reagent. may be modeled, adjusted for the presence of epitopes present in the protein sequence. For example, an epitope may be a "trimer" (a sequence of three amino acids). Drugs may be designed to target specific epitopes (e.g., GAV). Off-target binding of a drug (e.g., binding of an affinity reagent to an epitope other than its target epitope) The probability of binding to an additional epitope is calculated by including a non-zero probability of binding to the additional epitope. can be modeled as follows.

[0096] For example, an affinity reagent can be designed to bind to trimeric GAV, but with three additional recognition This affinity reagent may have off-target binding to the recognition sites: CLD, TYL, and IAD. Therefore, the joint probability can be modeled as follows: P(affinity probe binding | protein) = {GAV, CLD, TYL, or IAD is the protein sequence 0.25 if present; 0 otherwise}

[0097] There may be a small probability that the affinity reagent will bind non-specifically to the protein, which It can be expressed as: P(affinity probe binding | protein) = {GAV, CLD, TYL, or IAD is the protein sequence 0.25 if present; 0.00001 otherwise} Here, probability measures the outcome of detecting antibody binding.

[0098] As an example, consider the case where proteins from a sample of human origin are analyzed. The proteins are from the human "reference" proteome (e.g., canonical proteins as found in the Uniprot database of sequence and functional information That is, the list of protein candidates is considered to be represented in the UniProt database. This is a set of approximately 21,000 proteins and related sequences in the The population is derived from a sample, and each unknown protein is measured and recorded. In a series of affinity reagent binding experiments with a defined outcome (binding or non-binding), For example, such experiments can be carried out by sequentially adding different affinity reagents and and observing binding of the affinity reagent to the unknown protein. Or "probes" are proteins in a list of candidate proteins (approximately 800 possible triplicates). The target trimer is selected to target the most frequently observed trimer (among the target trimers). In addition to the target, each probe also targets several additional randomly selected trimers. The probability that a probe will bind to a protein sequence is expressed as obtain: P(affinity probe binding | protein) = 1 - [P(no nonspecific binding) * P(specific binding)] none)]

[0099] It is considered as follows: n = length of the protein candidate sequence; q = length of recognition site (e.g., 3); s = probability of nonspecific binding to trimers (e.g., 10 -5 ); p = probability of specific binding (e.g., 0.25); The terms P (no nonspecific binding) and P (no specific binding) are expressed as follows: obtain: P(no nonspecific binding) = (1 - s) n - q + 1 = (1 - 10 -5 ) n - 3 + 1 and P(no specific binding) = Π 各認識部位について (1-p) タンパク質において出現する部位の数

[0100] Finally, the probability that the probe does not bind to the protein can be expressed as: P(unbound affinity probe | protein) = 1 - P(bound affinity probe | protein) )

[0101] Figure 2 shows the sensitivity of the affinity reagent probes (e.g., identified with a false discovery rate (FDR) of less than 1%). The sensitivity is shown for three different experimental cases (grey circles, black circles, 50, 100, and 200 probes are used, represented by circles and open circles. ) in the affinity reagent probe, the recognition site of the probe (e.g., the trimer binding enzyme The number of recognition sites or trimer-binding epitopes of the probe can range up to 100. As shown in Figure 2, the number of probes used is has a significant impact on the ability to accurately identify proteins. Plotted on the y-axis The sensitivity is the accuracy at a threshold (e.g., upper limit) where less than 1% of identifications are incorrect. For example, if each probe is 5 recognition site or trimeric binding epitope (one target site, and four off-target sites) When 50 probes are used, the sensitivity of protein identification is less than 10%. It is about 60% when 100 probes are used, and about 60% when 200 probes are used. In fact, when 300 probes are used, the sensitivity is greater than 95% (results (The results are not shown on the plot.) This protein identification approach has been -Probes with target binding sites are available. 60 recognition sites or trimer binding epitopes (1 target site and 59 off-target sites) Even so, the sensitivity of identification was about 55% in experiments with 100 probes and 200 probes. In this experiment, it was about 90%.

[0102] However, as shown in Figure 3, the ability to identify proteins is reduced with 100 probes. The effect of ATP on the parental antibody is rapidly exacerbated when the parental antibody has multiple binding sites or a trimer-binding epitope. The sensitivity of the compatibility reagent probe (e.g., the percentage of substrates identified with a false discovery rate (FDR) of less than 1%) The sensitivity is shown in three different experimental cases (grey circles, black circles, and For the cases where 50, 100, and 200 probes are used, represented by white circles , the recognition site of the probe (e.g., a trimeric binding epitope) in the affinity reagent probe (probe recognition sites or trimer-binding epitopes range up to 700) For example, if each probe has 100 recognition sites or trimer-binding epitopes, When a target site is included in the target sequence, the probability of protein identification is high. The frequency is approximately 1% when 50 probes are used and 4% when 100 probes are used. Approximately 30% when 200 probes are used, and approximately 70% when 200 probes are used. The lobe contains 200 recognition sites or trimeric binding epitopes (1 target site, 199 off-targets). When 50 probes are used, the sensitivity of protein identification is Less than 1% when 100 probes are used, and less than 20% when 200 probes are used is used less than 40% of the time.

[0103] Example 2: Protein affinity assay for cleaved or degraded proteins Drug binding The methods described herein allow for the analysis and characterization of proteins in a sample that have been cleaved. In such experiments, affinity probes can be used to identify and / or identify The probability that a protein will bind to a protein is calculated based on the truncated portion of the protein, not the full-length protein sequence. For example, Figure 4 is modified to consider only binding to protein sequences that contain 100 Experiments using 100 probes (left), 200 probes (center), or 300 probes (right) Figure 1 shows plots demonstrating the sensitivity of experimental protein identification. Sensitivity of the probe (e.g., percent of substrates identified with a false discovery rate (FDR) of less than 1%) is determined for an experiment in which four lengths of substrate are measured: (1) intact (2) a 50-kDa (full-length) protein, (3) an N-terminal or C-terminal fragment of the protein, (4) N-terminal or C-terminal fragments of proteins of length 100; and (5) N-terminal or C-terminal fragments of proteins of length 200. N-terminal and C-terminal fragments of proteins. N-terminal and C-terminal fragments are represented by a solid bar and a solid Each probe is represented by a striped bar. Each probe contains one target trimer and four other random oligonucleotides. As shown in Figure 4, for example, a protein consisting of only 100 amino acids binds to a target trimer. If the protein is cleaved into fragments containing Even so, a significant proportion of the proteins (approximately 40%) can be identified.

[0104] When 300 probes are used, the protein is cleaved into fragments containing only 100 amino acids. In this case, approximately 70-75% of the proteins can be identified. Truncated proteins containing the C-terminal fragment are easier to identify than fragments containing the C-terminal fragment. These results demonstrate that the method is more efficient and easier to use (e.g., has higher sensitivity for protein identification).

[0105] Example 3: A protein that does not contain the C-terminus or N-terminus of the intact protein from which it is derived piece The methods described herein involve detecting protein fragments in a sample, the fragments being derived from the protein fragments. Protein fragments that do not contain either of the original two ends of the intact protein are separated. In such experiments, affinity profiles can be used to analyze and / or identify The probability that a probe will bind to a protein is calculated based on the truncated protein sequence, not the full-length protein sequence. Figure 5 shows the binding of various protein fragments to the target protein. 1 shows plots demonstrating the sensitivity of experimental protein identification using the fragmentation approach. In each of the columns below, the performance of protein identification is shown for 50, 100, 200, and and 300 affinity reagent measurements (from left to right in the four panels) of 50, 100, 200, 30 Maximum fragment length values ​​of 0, 400, and 500 (respectively for hexagons, downward-pointing triangles, and upward-pointing triangles) (represented by triangles, diamonds, squares, and circles).

[0106] Looking at the top row of Figure 5, each point in each subplot indicates the starting position of the fragment and the fragment length. The sensitivity (tandem repeat length) of a particular fragment generation approach is defined by the fragment length. The fragments represent the distance in amino acids from the N-terminus (e.g., 1000 to 10000 amino acids away from the N-terminus). The specific starting position (x) in each protein is represented by the number of amino acids (AA) present. The protein fragments are generated from the 50-amino acid sequence (plotted on the axis). acid, 100 amino acids, 200 amino acids, 300 amino acids, 400 amino acids, or 500 amino acids in length (maximum fragment length, or maximum fragment length value), They are respectively a hexagon, a downward-facing triangle, an upward-facing triangle, a rhombus, a square, and a circle. The protein is too short to generate fragments of a given specified length. If this is not possible, fragments shorter than the indicated length, including the C-terminus, will be retained. When experiments are performed with 0 affinity reagents, a small percentage of proteins However, only 200 affinities may be identified (plotted on the y-axis). Experiments were performed using reactive reagent probes and fragments with a maximum length of 200 amino acids. When the fragments are separated, approximately 50% to 85% of the protein fragments are isolated, depending on the starting site of the fragment (plotted on the x-axis). Proteins can be identified (plotted on the y-axis). There is an overall trend for the sensitivity of protein identification to decrease as the number of proteins increases. The tendency is that the further away from the N-terminus the fragment starts, the more likely it is to contain the C-terminus and fragment. This can be explained by the fact that many more fragments are generated that are less than the maximum length of the fragment. .

[0107] Looking at the bottom row of Figure 5, the four subplots here show fragments that do not match the maximum fragment length. Any fragments (e.g., fragments not containing the C-terminus) were analyzed prior to calculation of sensitivity and false discovery rate. Results are shown that are similar to those in the row above except that the protein is excluded from the The sensitivity of is calculated only among those proteins that could generate valid fragments. As shown in the bottom row of Figure 5, without fragment length adjustment, the fragment length is There is no statistically significant change in the sensitivity of protein identification with respect to the position of the fragment start site. The length of the fragment, rather than its position in the protein sequence, is the primary determinant of protein identification rate. This is a factor.

[0108] Example 4: Protein Identification by Measurement of Length, Hydrophobicity, and / or Isoelectric Point The methods described herein allow for the determination of length, hydrophobicity, and / or isoelectric point (pI). The information from the protein measurements is used to analyze and For a given protein query candidate, a specific The probability of measuring the length of can be expressed by: TIFF2025163121000002.tif8128 here σ = | CV * expected outcome value | u = (measured outcome value - expected outcome value) / σ

[0109] In this case, the measured outcome is the measured length of the unknown protein, and the expected The outcome value measured is the length of the protein query candidate. The coefficient of variation (CV) value is used to represent the expected precision of the approach. The probability of measuring a particular hydrophobicity was calculated using the same formula and the predicted outcome value was calculated. is the grand average of hydropathy calculated from the sequences of the protein candidates. The gravy score is set to the hydropathy (gravy) score. For example, the Kyte-Doolittle calculation method (e.g., [Kyte et al., “A simple method for displaying the hydropathic character of a protein,” J. Mol. Biol., 1982 May 5; 157(1):105-32). can be calculated using Biopython tools for computational molecular biology. The points (pI) are modeled using predicted pI values, which are calculated using the method of Bjellqvist (e.g. See, for example, Audain et al., "Accurate est imation of isoelectric point of protein and peptide based on amino acid sequence s,” Bioinformatics, 2015 November 14; 32(6):821-27]) in its entirety. Tabb, David L., "An algorithm for isoelectri c point estimation,” <http: / / fields.scripps.edu / DTASelect / 20010710-pI-Algorithm .pdf>, 2003 June 28]. , calculated from the candidate protein sequence. In all cases, the accuracy of the experimental measurements is , was set to a CV value of 0.1.

[0110] Figure 6 shows the results of experiments using various combinations of measurement types for the identification of human proteins. A plot showing the sensitivity (percent of substrates identified with an FDR of less than 1%) is shown. In fact, if only protein length measurements, only hydrophobicity measurements, or only pI measurements are used, Furthermore, it is not possible to identify proteins (e.g., sensitivity of less than 1%). Combining all of these (length + hydrophobicity + pI) still resulted in virtually no identification. However, measurements of protein length, hydrophobicity, or pI are not possible with affinity reagent probes. For example, proteins may be used to enhance measurements from binding experiments. The fractions may be divided based on any of these characteristics, and each fraction may be a different spatial location on the substrate. Following this fractionation and conjugation, the parent Binding measurements of affinity reagents may be made, as well as measurements of hydrophobicity, protein length, or pI. The location of the protein can be determined by its spatial address. Separation by molecular weight based on gel filtration (SDS-PAGE) or size exclusion chromatography The length of a protein can be determined by dividing the molecular weight by the average molecular weight of an amino acid (111 Da). Proteins can be estimated from their molecular weight by hydrophobic interaction chromatography. Proteins may be fractionated by hydrophobicity using ion exchange chromatography. For example, the tandem separation may be performed by fractionating the tandem separation by pI using a CV of 0.1. Additional measurements of protein length were performed using 100 probes (one probe per probe). The sensitivity of the identification using the experiment (target trimer, and four additional off-target sites) was approximately 5 Improved from 5% (without measuring protein length) to approximately 65% ​​(with measuring protein length) Similarly, an additional measurement of the protein length was performed with a CV of 0.1. 200 probes (one target trimer and four additional off-target sites per probe) The sensitivity of identification using the experiment ranged from about 90% (without measuring protein length) to about 95% (without measuring protein length). This has been improved to include protein length measurement.

[0111] Example 5: Protein identification by assay using a mixture of antibodies The method described herein involves measuring a mixture of affinity reagents in each binding experiment. To analyze and / or identify proteins in a sample using information from the experiment Consistent with the disclosed embodiments, the identification of 1,000 unknown human proteins can be Binding assays were performed using a pool of antibodies commercially available from Santa Cruz Biotechnology, Inc. A benchmark test was conducted by obtaining 1,000 proteins. , randomly selected from the Uniprot protein database, which contains approximately 21,005 proteins. The Santa Cruz Biotechnology catalogue shows that the antibody is reactive to human proteins. A list of available monoclonal antibodies can be downloaded from the online antibody registry. The list contained 22,301 antibodies and was based on the Uniprot human protein database. The results were filtered to a list of 14,566 antibodies that matched proteins in the database. The complete population of antibodies modeled in this study includes these 14,566 antibodies. Experimental evaluation of the binding of the antibody mixture to 1,000 unknown protein candidates was It was carried out as follows:

[0112] First, 50 mixtures of antibodies were modeled. To create any one mixture, 5,000 antibodies from the total population of antibodies were randomly selected.

[0113] Then, for each mixture, the binding probability is calculated for the mixture for any of the unknown proteins. Proteins were identified based on the goal of predicting their identity. Although each protein is "unknown" in the sense that it is a protein with a specific structure, the algorithm Note that the identity is known. If the mixture contained an antibody that binds to the unknown protein, a binding probability of 0.99 was assigned. When the antibody against the target was not included, a binding probability of 0.0488 was assigned. The probability of the joint outcome for the mixture was modeled as follows: P(binding outcome | protein) = {0.99 if the mixture contains an antibody against the protein} ; otherwise 0.0488} The value of 0.0488 indicates non-specific (off-target) activity against the protein for this mixture. The probability of a binding event occurring is expressed as the probability of a nonspecific binding event occurring for a mixture. The expected probability that an individual antibody will bind to a protein other than its target, and the probability that it will bind to a protein other than its target in a mixture The nonspecific binding index for a mixture of antibodies was modeled based on the number of proteins in the mixture. The probability of a bent is the probability that any one antibody in the mixture will bind nonspecifically. The ratio depends on the number of antibodies in the mixture (n) and the probability of nonspecific binding for any one antibody. It is calculated based on the rate (p) and can be expressed by the following equation: Probability of nonspecific binding in the mixture = 1 - (1 - p) n

[0114] In this case, the individual antibodies bind to something other than their target protein. Specific binding events were assigned a value of 0.00001 (10 -5 ) probability. Therefore, any The probability of nonspecific binding (p) for one antibody is 10 -5 which gives: Probability of nonspecific binding of the mixture = 1 - (1 - 10 -5 ) 5000 = 0.0488

[0115] In addition, the probability of a non-binding outcome to the protein was calculated as follows: P(unbound outcome | protein) = 1 - P(bound outcome | protein)

[0116] For each unknown protein, binding is calculated as the probability of binding of the antibody mixture to the unknown protein. Each antibody mixture measured was evaluated based on a minimum of 0 and a maximum of 1. A uniform distribution with values ​​is randomly sampled, and the resulting number is , the probability of binding of the antibody mixture to the unknown protein is less than Otherwise, the experiment yielded a non-binding event for the mixture. With all binding events evaluated, protein prediction was performed as follows: It will be carried out.

[0117] For each unknown protein, a set of evaluated binding events (50 in total, one per mixture) was (one per protein) for each of the 21,005 protein candidates in the Uniprot database. More specifically, the probability of observing a sequence of binding events was calculated for each candidate. The probability was determined for each individual mixture across all 50 mixtures measured. The probability of binding was calculated by multiplying the probability of a binding event / a non-binding event. It is calculated in the same manner as above, and the probability of non-binding is 1 minus the probability of binding. The protein query candidate with the highest binding probability is selected for the unknown protein. The identity is predicted based on the individual proteins. The rate is calculated as the probability of the top individual candidate divided by the combined probability of all candidates. was done.

[0118] with the predicted identities for each of the 1,000 unknown proteins. The unknown proteins were sorted in descending order of their probability of identification. A top-off is when the percentage of incorrect identifications among all preceding identifications in the list is 1%. Overall, 551 of the 1,000 unknown proteins were selected to be 1% of the total. Protein identification was therefore performed with a sensitivity of 55.1%. Ta.

[0119] Example 6: Protein identification in multiple species The methods described herein allow for the detection of proteins in samples obtained from many different species. For example, a series of affinity reagents can be applied to analyze and / or identify The results from the experiments are shown for E. coli, Saccharomyces cerevisiae, and Bacillus subtilis, which are represented by circles, triangles, and squares, respectively. Saccharomyces cerevisiae (yeast), or Homo sapiens The method can be used to identify proteins in Homo sapiens (humans). To tailor the analysis method to each species, lists of protein candidates were compiled from Uniprot species-specific sequences, such as a reference proteome for that species downloaded from It must be generated from the database.

[0120] Figure 7 shows the results of the analysis of E. coli, yeast, or human (represented by circles, triangles, and squares, respectively). 50, 100, 200, or 300 unknown proteins from either 1 shows a plot demonstrating the sensitivity of experimental protein identification using affinity reagent probe paths. Each probe binds to one target trimer and four additional off-target sites with a probability of 0.25. Sensitivity for experiments with 200 probes (identified with a false positive rate of less than 1%) The percentage of unknown proteins detected (i.e., the percentage of unknown proteins detected) was calculated for each of the three species tested. It was about 90%.

[0121] Example 7: Protein identification in the presence of SNPs The methods described herein involve detecting mutations caused by non-synonymous single nucleotide polymorphisms (SNPs). Analyze and / or identify proteins in a sample for the presence of a single amino acid mutation (SAV) It can be applied to proteins that have the same sequence except for a small number of single amino acid variations (SAVs). For example, in experiments using a series of affinity reagent measurements, In this study, the canonical form of a protein is highly selective for polymorphic regions of the protein. Unless affinity reagents are included in the experiment, it can be nearly impossible to distinguish from its variants. In cases where the polymorphic region is not distinguished by either of the affinity reagent measurements, Measurement of some protein types is performed for both canonical and variant protein query candidates. It is thought that similar probabilities (likelihoods) are returned for each of these (e.g., L (canonical protein | evidence) = 0.8 and L (mutant protein | evidence) = 0.8).

[0122] In such cases, no individual protein candidate returns a probability higher than 0.5. This may not be the case, for example, for a canonical protein, as represented below: (where cprot = canonical protein and vprot = variant protein) (see below): TIFF2025163121000003.tif8128Here L other is used for all proteins except for canonical and variant proteins. the aggregated likelihood of the quality query candidate and is a number greater than or equal to zero is.

[0123] In this case, a group of potential protein identifications may be returned for the unknown protein. For example, the probability of finding the top two most likely protein query candidates is The rate can be expressed as: TIFF2025163121000004.tif8129 Although this approach does not distinguish between canonical and variant proteins, Thus, a reliable identification can be derived from an unknown protein. other is close to zero If there is no significant difference, a reliable identification is likely to result.

[0124] Example 8: Iterative refinement of a probabilistic model with empirical results A probabilistic model used in one or more of the methods described herein The data is used to calculate protein identification using an expectation-maximization or related approach. It can be iteratively improved using empirical measurements during calculation. One such approach is , as described herein with respect to affinity reagent binding experiments.

[0125] First, the binding probability for each affinity reagent probe is initialized to a guess value. For example, a population of 200 probes could each target one trimer and have an estimated β of 0.5. Proteins can be identified using the approaches disclosed elsewhere herein. (See, e.g., Example 1). Then, for each probe: The joint probability of It will be improved.

[0126] (1) To update the joint probability, we use the results identified with an estimated false discovery rate of less than 0.01. A population of unknown proteins is used.

[0127] For each probe, the number of probes in the population containing the binding site (trimer) recognized by the probe is The proportion of proteins containing the protein is used to calculate the updated binding probability: TIFF2025163121000005.tif17128

[0128] Increase the probability of a probe having "more than 20 proteins with binding sites in the population" Update.

[0129] Updated probability is 10 -5 If it is less than 10 -5 (the probability of 0 is (to avoid being assigned a

[0130] (2) Perform another protein identification using the updated binding probabilities.

[0131] Repeat multiple iterations of steps 1 and 2 (e.g., 1, 2, 3, 4, 5 times total) , 6, 7, 8, 9, 10, or more than 10 repetitions).

[0132] This iterative approach yielded 200 nucleotides, each recognizing one trimer with a binding probability of 0.25. The binding measurements of 200 probes were performed with a set of 0.5. Using the initial guesses of the probe binding probabilities, we performed a set of 2000 unknown proteins. After five iterations of this iterative algorithm, the update The binding probability of the probes tested is more accurate (closer to 0.25) and less susceptible to tampering. The sensitivity of protein identification was increased.

[0133] Figure 8 shows the binding probability (y-axis, left) and sensitivity of protein identification (y-axis, The plots shown in Figure 8 are for the individual probes. The dark line in the light line indicates the binding probability of the probe. The median and thick lines indicate the sensitivity of protein identification in each replicate.

[0134] Example 9: Estimating false discovery rates of identification from protein candidate match probabilities of a protein used in one or more of the methods described herein. Probabilistic models for prediction or identification have as a direct product the ability to predict the identity of unknown proteins. For each, a list of protein sequence matches and whether the sequence matches are correct Often, only a subset of protein identifications are correct. Therefore, it is possible to estimate and control the false identification rate for a set of proteins. Methods useful for this purpose are described below.

[0135] First, the complete set of protein identifications is sorted in descending order by protein identification probability. rearranged, which is shown below (where prot = protein): Probability of prot1 (p1): 0.99 Probability of prot2 (p2): 0.97 Probability of prot3 (p3): 0.92 Probability of prot4 (p4): 0.9 Probability of prot5 (p5): 0.8 Probability of prot6 (p6): 0.75 Probability of prot7 (p7): 0.6 Probability of prot8 (p8): 0.5

[0136] Then, at each point in the list, the expected false discovery rate is Calculated as TIFF2025163121000006.tif4128, where TIFF2025163121000007.tif4128 is the average of all probabilities at a given point in the list and those before it ( shown below): TIFF2025163121000008.tif64128

[0137] As shown in Figure 9, the estimated Comparison of the estimated false positive rate with the true false positive rate demonstrates accurate estimation of the false positive rate. When looking at lots, the sensitivity of identification is compared to the true and estimated false identification rates. Looking at the bottom plot of Figure 9, the estimated false positive rate is 1.5 times lower than the true false positive rate. The dashed line represents the ideal, perfectly accurate false identification. Rate estimates are given.

[0138] The estimated false identification (ID) rate is a function of the risk of protein identification, depending on the tolerance for false identifications. This can be used as the threshold for

[0139] Example 10: Derivation of a false discovery rate estimation approach Each protein identification identifies the most likely protein for an unknown protein. A list of protein identifications containing matches and associated statements that the matches are accurate. Consider the probability (P(protein | evidence)). For example: TIFF2025163121000009.tif27128.

[0140] The expected number of false discoveries in this list is 1 - all proteins in the list The average matching probability for . In this case: TIFF2025163121000010.tif10128.

[0141] The rationale behind this approach is as follows: A list of N protein identifications and a random variable, each protein identification, prot i Consider the following, where the identification is prot if accurate i = 1 and the identification is inaccurate, then prot i = 0 In this case, the number of correct identifications (correctids) in any list is determined by these random The sum of the variables is: TIFF2025163121000011.tif14128

[0142] The expected value for each individual protein identification is equivalent to the probability of correct identification: TIFF2025163121000012.tif5128

[0143] Due to linearity of expectation, we have: TIFF2025163121000013.tif14128

[0144] The expected true discovery rate (number of correct identifications / number of identifications) is, on average, the probability of a candidate being: TIFF2025163121000014.tif14128

[0145] The false discovery rate is 1 - the true discovery rate, i.e.: The file is TIFF2025163121000015.tif4128.

[0146] Example 11: Protein Identification Using Binding Measurement Outcomes The methods described herein involve binding of affinity reagents to unidentified proteins and and / or may be applied to a different subset of data related to non-binding. In the method described herein, the identification of the measured combined outcomes It can be applied to experiments in which a subset of the variables is not examined (e.g., uncoupled measurement outcomes). These methods, in which a subset of the defined combined outcomes is not considered, are not included in this specification. In the literature, a "censored" estimation approach (such as the approach described in Example 1) In the results shown in Figure 10, the results of the truncated estimation approach The resulting protein identifications are based on binding events associated with specific unidentified proteins. Therefore, the censoring estimation approach is based on assessing the occurrence of unknown Non-binding outcomes are not considered when determining the identity of the protein.

[0147] This type of censoring estimation approach considers all available combined outcomes. (e.g., binding and non-binding outcomes associated with specific unidentified proteins) This contrasts with the "uncensored" approach, which involves multiple measures (both for a single outcome and for a combined outcome). In embodiments, a particular binding measure or binding measurement outcome may be more susceptible to error. or the expected binding outcome for the protein (e.g. the probability of deviation from the binding outcome (e.g., the probability that the protein will produce the binding outcome) A truncation approach may be applicable when it is expected that In affinity reagent binding experiments, binding and non-binding outcomes were measured. The probability can be calculated based on binding to a denatured protein with a mostly linear conformation. Under these conditions, the epitope may be readily accessible to affinity reagents. However, in some embodiments, binding measurements in the assayed protein sample can be collected under non-denaturing or partially denaturing conditions, under which the protein They exist in a "folded" state with significant three-dimensional structure, which is often The affinity reagent binds to epitopes on the protein that are accessible in the linear form, and the folded In the folded state, steric hindrance can make it inaccessible. For proteins, the epitopes recognized by the affinity reagents are structurally related to the folded protein. Experimental binding measurements obtained in unknown samples when they are in structurally accessible areas is expected to be consistent with the calculated probability of binding derived from the linearized protein. However, for example, if the epitope recognized by the affinity reagent is structurally If inaccessible, a predicted binding probability is calculated from the linearized protein. It can be expected that there are more non-binding outcomes than expected. Based on the specific conditions surrounding the protein, the three-dimensional structure can take several different possible forms. configurations, and each of the different possible configurations can be used to create a desired affinity reagent. Based on the degree of accessibility of the target molecule, a unique prediction can be made regarding binding to a particular affinity reagent. do.

[0148] Therefore, non-binding outcomes deviate from the calculated binding probability for each protein. censoring estimates that can be expected to be A "censored" estimation approach such as that provided in Figure 10 may be appropriate. In this study, only the joint outcomes measured were considered (in other words, non-joint outcomes). either not measured or non-bound outcomes that were measured were not considered), and Thus, the M measured joint outcomes that resulted in the joint measure All N measured joint outcomes, including both outcomes and non-joint measured outcomes This is a subset of the outcomes, but only this is considered in the probability of the combined outcome set. This can be described by the following expression: P(outcome set | protein) = P(binding event 1 | protein) * P(binding event event2 | protein) * … * P(binding eventM | protein)

[0149] When applying the truncation approach, a scale factor P(combined It may be appropriate to apply a longer Proteins generally have a higher probability of producing a potential binding outcome (e.g. (e.g., because they contain more potential binding sites). The scaled likelihood SL is used to calculate P(binding outcome set | protein) as a function of M outcomes. The number of unique combinations of binding sites, based on the number of potential binding sites on the protein , which can be generated from the protein, and by dividing by this, each candidate protein For a protein of length L that has a trimer recognition site, L -There may be two potential binding sites (e.g., the complete protein sequence, for each subsequence of length L), so: TIFF2025163121000016.tif12155

[0150] The probability of any candidate protein being selected from a population of Q possible candidate proteins is ,Given a set of outcomes, it can be expressed as: TIFF2025163121000017.tif12128

[0151] Censored vs. uncensored protein estimation approaches The performance of the embodiment is plotted in Figure 10. The data plotted in Figure 10 is provided in Table 1. do.

[0152] (Table 1) TIFF2025163121000018.tif83128

[0153] In the comparison shown in Figure 10, the sensitivity of protein identification (e.g., the number of unique proteins identified) The percent of protein (%) was used as the truncation estimate for the linear protein substrate. Plotted against the number of affinity reagent groups measured, for both the and uncensored estimates. The affinity reagents used target the highest and most abundant trimers in the proteome. and each affinity reagent has off-target affinity to four additional random trimers. If 100 affinity reagent sets are used, the non-censored approach It outperforms the truncated approach by more than a factor of two. The degree to which it outperforms the original estimate decreases when more groups are used.

[0154] Example 12: Protein Identification for Random False Negative and False Positive Affinity Reagent Binding Acceptability of In some cases, there are many false-negative binding assay outcomes for affinity reagent binding. "False negative" combined outcomes occur less frequently than expected. Such "false negative" outcomes manifest as affinity reagent binding measurements occurring in For example, binding detection methods, binding conditions (e.g., temperature, buffer composition, etc.), protein sample This may occur due to degradation or problems with the affinity reagent stock. The impact of false-negative measurements on identification and non-censored protein identification approaches To determine the impact, a subset of affinity reagent assays was performed, such as 1 in 10, 1 in 100, or 1 in 1,000, 1 in 10,000, or 1 in 100,000 By in silico exchanging artificial, observed binding events into non-binding events The affinity reagents were intentionally degraded by the following methods: 0, 1, 50, 100, and 20 out of a total of 300 affinity reagents. Either 0 or 300 particles were degraded in this manner. The results are plotted in Figure 11. As shown, the censored protein identification approach and the non-censored protein identification approach Both quality identification approaches tolerate this type of random false negative binding. The plotted data are provided in Table 2.

[0155] (Table 2) TIFF2025163121000019.tif186154TIFF2025163121000020.tif200154

[0156] Similarly, "false positive" binding outcomes occur more frequently than expected in affinity assays. The tolerance for "false positive" binding outcomes is a key factor in determining the binding outcome. Subsets were evaluated by swapping non-combined outcomes for combined outcomes. The results of this evaluation are provided in Table 3.

[0157] (Table 3) TIFF2025163121000021.tif234154TIFF2025163121000022.tif152154

[0158] These results, plotted in Figure 12, show that increasing the occurrence of random false positive measurements The performance of the truncated protein identification approach was significantly higher than that of the uncensored protein identification approach. However, both approaches show that the affinity of each group of affinity reagents deteriorates rapidly. False positive rate of 1 in 1000, or 1 in 100 for a subset of affinity reagents The ratio of is allowed.

[0159] Example 13: Protein Identification Using Over- or Under-Estimated Affinity Reagent Binding Probabilities Quality estimation performance The sensitivity of protein identification depends on the accurately predicted binding probability of the affinity reagent to the trimer, and and protein identification using over- or under-estimated affinity reagent binding probabilities - Patents.com The true probability of association was 0.25. The underestimated probability of association was: The overestimated joint probabilities were 0.30, 0.50, 0.75, and 0.2. A total of 300 affinity reagent measurements were obtained. None (0), all 300, or a subset (1, 50, 100, 200) are over-represented. All other proteins were assigned over- or under-estimated binding probabilities. In the identification, an exact joint probability (0.25) was used. The results of the analysis are provided in Table 4. .

[0160] (Table 4) TIFF2025163121000023.tif187170TIFF2025163121000024.tif234170TIFF2025163121000025.tif187170

[0161] These results, plotted in Figure 13, suggest that joint probabilities may not be accurately estimated. In some cases, censored protein identification may be the preferred approach. Shows.

[0162] Example 14: Protein prediction approach using affinity reagents with unknown binding epitopes Performance In some cases, the affinity reagent may have several unknown binding sites (e.g., epitope). A truncated protein identification approach using affinity reagent binding measurements. The sensitivity of the threonine and non-threonine protein identification approach was evaluated using five trimerization sites (e.g., , one target trimer, and four random off-target sites) for protein identification algorithms. The results were compared using affinity reagents that each bind with a probability of 0.25 entered into the algorithm. Affinity reagent subsets (0 of 300, 1 of 300, 50 of 300, 300 100 of 300, 200 of 300, or 300 of 300) are 1, 4, or 40 Each of the sites has an additional extra binding site, either The results of the analysis are shown in Table 5. .

[0163] (Table 5) TIFF2025163121000026.tif152170TIFF2025163121000027.tif241170TIFF2025163121000028.tif241170TIFF2025163121000029.tif83170

[0164] These results, plotted in Figure 14, show that the uncensored estimates indicate additional hidden binding sites. The higher tolerance to inclusion of and the performance of both estimation approaches are 0 affinity reagents are significantly impaired when they contain 40 additional binding sites This indicates that...

[0165] Example 15: Performance of protein prediction approaches using affinity reagents lacking binding epitopes In some cases, some of the annotated binding epitopes are not present. Inadequately characterized with a specific target (e.g., extra predicted binding sites) Affinity reagents may be present, i.e., generate predicted binding probabilities for the affinity reagents. The model used for affinity reagent binding includes an extra predicted site that does not exist. Censored and non-censored protein identification approaches using measurements The sensitivity of the approach is demonstrated by the use of random trimer sites (e.g., one target trimer and four random (unintentional off-target sites) with a probability of 0.25 when input into the protein identification algorithm. A subset of affinity reagents (300) were used for comparison. 0 of them, 1 of 300, 50 of 300, 100 of 300, 200 of 300, or 300 out of 300) are either 1, 4, or 40 extra expected results. The binding sites have binding probabilities of 0.05, 0.1, and 0.4 for a random trimer, respectively. or 0.25 for the affinity reagents used by the protein prediction algorithm The results of the analysis are shown in Table 6.

[0166] (Table 6) TIFF2025163121000030.tif234170TIFF2025163121000031.tif241170TIFF2025163121000032.tif220170

[0167] These results, plotted in Figure 15, demonstrate that uncensored estimates of affinity reagent binding are consistent with the model higher tolerance to the inclusion of extra predicted binding sites in The performance of both protein identification approaches was such that the majority of affinity reagents were 40 extra predicted This indicates that the binding site is impaired to some extent when the target protein is present.

[0168] Example 16: Estimating Truncation for Affinity Reagent Binding Analysis Using Alternative Scaling Strategies The methods described herein can be combined with various scaling strategies for probabilities. and estimating the identity of proteins (e.g., unknown proteins) using affinity reagent binding measurements. The censoring estimation approach described in Example 11 can be applied to the identification of proteins. The number of potential binding sites in the protein (protein length - 2) and the observed Based on the number of binding outcomes (M) observed, the probability of a protein being Scaling the probabilities: TIFF2025163121000033.tif12128

[0169] The method described herein provides an alternative approach for calculating the scaled likelihood. An example of this is the affinity assay used to measure proteins. Model the probability of generating N binding events for a protein of length k from a set of drugs. Another approach to normalization is to use the probability First, for each probe, the probe is applied to the unknown identities in the sample. The probability of binding to a trimer of TIFF2025163121000034.tif15128Here, P(trimer j ) compared to the total number of all 8,000 trimers in the proteome For any protein of length k, the frequency at which probe i is present is The probability of binding to a protein can be expressed as: P(protein binding | probe i , k) = 1 - (1 - P(trimer binding | probe i )) k-2

[0170] The number of successful binding events observed for a protein of length k is given by may also follow a Poisson-binomial distribution, where n is the number of probes made on the protein. is the number of combined measurements and the parameter p of the distribution プローブ, k is the probability of success for each trial. Showing rates: p プローブ, k = [P(binding | probe1, k), P(binding | probe2, k), P(binding | probe B3, k) … P(Binding | Probe n , k)]

[0171] A specific set of probes is used to generate N binding events from a protein of length k. The probability of a Poisson binomial distribution parameterized by p and evaluated at N is Rate Mass Function (PMF PoiBin ) can be represented as: P(N binding events | probe, k) = PMF PoiBin (N, p プローブ, k )

[0172] The scaled likelihood of a particular set of outcomes is calculated based on this probability. : TIFF2025163121000035.tif11128

[0173] Example 17: Use of randomly selected affinity reagents The methods described herein can be applied to any set of affinity reagents. For example, protein identification approaches target the most abundant trimers in the proteome. This can be applied to a set of affinity reagents that target a single target or random trimers. Affinity reagents targeting the top 300 most abundant trimers in the proteome Affinity reagents targeting 300 randomly selected trimers in the Results from a putative human protein analysis using affinity reagents targeting 300 of the most abundant trimers The results are shown in Tables 7A to 7C, respectively.

[0174] Table 7A~Table 7C Table 7A. 300 affinity reagents targeting the most scarce trimers in the proteome. TIFF2025163121000036.tif49128

[0175] Table 7B. 300 affinity reagents targeting random trimers in the proteome. TIFF2025163121000037.tif62128TIFF2025163121000038.tif241122TIFF2025163121000039.tif241122TIFF2025163121000040.tif241122 TIFF2025163121000041.tif241122TIFF2025163121000042.tif241122TIFF2025163121000043.tif241122TIFF2025163121000044.tif241122 TIFF2025163121000045.tif241122TIFF2025163121000046.tif241122TIFF2025163121000047.tif241122TIFF2025163121000048.tif241122 TIFF2025163121000049.tif241122TIFF2025163121000050.tif241122TIFF2025163121000051.tif241122TIFF2025163121000052.tif227122

[0176] Table 7C. 300 affinity reagents targeting the most abundant trimers in the proteome. TIFF2025163121000053.tif49128

[0177] These results are plotted in Figure 16. In all cases, each affinity reagent A randomly selected additional trimer has a binding probability of 0.25 to the target trimer. The performance of each affinity reagent set was evaluated by sensitivity (for example, The percentage of identified proteins is measured based on the affinity of each affinity reagent set. is evaluated in five iterations, where the performance of each iteration is plotted as a point, and Vertical lines connect replicate measurements from the same set of affinity reagents. The results from the affinity reagent set of 300 most abundant affinity reagents are blue, The bottom 300 are green. A total of 300 affinity reagents targeting random trimers are shown. 100 different sets were generated and evaluated. Each of the sets was a gray It is represented by a set of five grey dots (one for each iteration) connected by vertical lines. Based on the uncensored estimates used in this analysis, the more abundant trimer is targeted. Targeting the target trimer improves the performance of the identification compared to targeting a random trimer.

[0178] Example 18: Affinity Reagents with Biosimilar Off-Target Sites The methods described herein involve targeting different types of off-target binding sites (epitopes). This example can be applied to affinity reagent binding experiments using affinity reagents with a nucleotide sequence (e.g., a nucleotide sequence). In this study, the performance of two classes of affinity reagents is compared: random affinity reagents, and "Biosimilar" affinity reagents. Results from these evaluations are shown in Tables 8A-8D.

[0179] Table 8A~Table 8D Table 8A: Biosimilars with off-target sites and the highest number of nucleotides in the proteome Performance of censoring estimates using affinity reagents targeting the most abundant 300 trimers TIFF2025163121000054.tif35128

[0180] Table 8B: Biosimilars with off-target sites and the highest number of nucleotides in the proteome Performance of uncensored estimates using affinity reagents targeting the most abundant 300 trimers TIFF2025163121000055.tif35128

[0181] (Table 8C) Proteins with random off-target sites and most abundant in the proteome Performance of censoring estimates using affinity reagents targeting 300 different trimers TIFF2025163121000056.tif35128

[0182] (Table 8D) Proteins with random off-target sites and most abundant in the proteome Performance of uncensored estimates using affinity reagents targeting 300 identical trimers TIFF2025163121000057.tif35128

[0183] Unlike random affinity reagents, biosimilar affinity reagents have a unique affinity for the target epitope and biotin. have biologically similar, off-target binding sites. Both mirror affinity reagents target their target epitopes (e.g., trimers) with a binding probability of Each affinity reagent in the random class is randomly selected with a binding probability of 0.25. It has four selected off-target trimer binding sites. In contrast, "biosimilars" The four off-target binding sites of the affinity reagent are most similar to the target trimer of the affinity reagent. These four trimers are combined with a probability of 0.25. For reagents, the similarity between the trimer sequences is BL for the amino acid pair at each sequence position. The OSUM62 coefficients are calculated by summing the OSUM62 coefficients. Both sets of iosimilar affinity reagents target the top most abundant triads in the human proteome. The target was 300 trimers, where the abundance level included one or more instances of trimers. Figure 17 shows the number of unique proteins that are targeted to random off-target sites. Affinity reagents with (blue) or biosimilar off-target sites Percentage of proteins identified in human samples when (orange) is used Regarding the truncated protein estimation approach (dashed line) and the uncensored protein The performance of the estimation approach (solid line) is shown.

[0184] In this comparison, the uncensored estimates outperformed the censored estimates, Censoring estimates performed better with biosimilar affinity reagents, and The estimate of censoring performs better with random affinity reagents.

[0185] Alternatively, affinity reagents could be used that target the most abundant trimers in the proteome. Instead, an optimal set of trimeric targets can be determined by measuring candidate proteins (e.g., human proteins). the type of protein estimation performed (censored or uncensored), and the Based on the type of affinity reagent used (random or biosimilar), specific applications A "greedy" algorithm can be used to find the best affinity locus, as described below. To select a set of reagents, one can use: 1) Initialize an empty list of affinity reagents (ARs) to be selected. 2) A set of candidate ARs (e.g., each of which has a random off-target site) A population of 8,000 ARs (each targeting a unique trimer) is initialized. 3) (e.g., all human proteins in the Uniprot reference proteome) A set of protein sequences is selected for optimization. 4) Repeat the following until the desired number of ARs has been selected: a. For each candidate AR: i. Simulate the binding of the candidate AR to the set of proteins. ii. Simulated binding measurements from the candidate AR and all previously selected A Protein estimation was performed for each protein using simulated binding measurements from R. To carry out. iii. The exact protein sequence for each protein as determined by protein prediction. A score is calculated for the candidate AR by summing the probabilities of protein identification. b. Add the AR with the highest score to the set of selected ARs and call it the candidate AR Remove from list.

[0186] The greedy approach targets the top 4,000 most abundant trimers in the human proteome either a random population of affinity reagents or a population of biosimilar affinity reagents. This was used to select the 300 best affinity reagents. The optimal values ​​were calculated for both the censored and uncensored protein estimates. The results from the analysis are provided in Tables 9A-9D.

[0187] Table 9A~Table 9D Table 9A: Biosimilars with off-target sites and the highest concentration of nucleotides in the proteome Estimating performance of censoring using affinity reagents targeting 300 suitable trimers TIFF2025163121000058.tif35128

[0188] Table 9B: Biosimilar off-target sites and the most abundant proteins in the proteome Uncensored estimated performance using affinity reagents targeting 300 suitable trimers TIFF2025163121000059.tif35128

[0189] (Table 9C) Random off-target sites and optimal triplet in the proteome Performance of censoring estimates using affinity reagents targeting 300 mers TIFF2025163121000060.tif35128

[0190] Table 9D. Random off-target sites and optimal triplet in the proteome. Performance of uncensored estimates using affinity reagents targeting 300 mers TIFF2025163121000061.tif35128

[0191] Optimal for both censored and uncensored protein estimates The performance of the synthesized probe sets is plotted in FIG.

[0192] The use of the set of affinity reagents selected by the greedy optimization algorithm is Both censored and uncensored protein prediction approaches were used. The performance of both random and biosimilar affinity reagent sets was evaluated. In addition, the random set of affinity reagents improves the greedy approach to find affinity reagents. When used to select for affinity reagents, they perform nearly identically to biosimilar affinity reagent sets. do.

[0193] Example 19: Protein Prediction Using Binding Mixtures of Affinity Reagents The methods described herein involve the measurement of protein activity using a mixture of affinity reagents. The affinity reagents can be used to analyze and / or identify the quality of the product. The probability that a particular protein will produce a binding outcome when assayed is: can be calculated as: 1) The average probability of nonspecific epitope binding for each affinity reagent in the mixture Calculate TIFF2025163121000062.tif4128. 2) The number of binding sites in a protein is determined by the length of the protein (L) and the number of affinity reagents. Calculated based on epitope length (K): Number of binding sites = L - K + 1. Nonspecific binding The probability that the event does not occur is TIFF2025163121000063.tif5128. 3) For each affinity reagent in the mixture, the probability that an epitope-specific binding event will not occur. is calculated as follows: TIFF2025163121000064.tif191284) For proteins, the probability that a mixture will produce a non-binding outcome is: TIFF2025163121000065.tif111285) The probability that a mixture will produce the combined outcome is: P(bound | protein) = 1 - P(unbound | protein)

[0194] To calculate the probability of a binding or nonbinding outcome from a protein mixture This approach is useful for analyzing the performance of affinity reagent mixtures for protein identification. This was used in combination with the methods described herein to Each affinity reagent binds to its target trimeric epitope with a probability of 0.25 and The four trimers that are most similar to the target TOP have a probability of binding of 0.25. Regarding the similarity of the trimers, the amino acid at each sequence position in the compared trimers is It is calculated by summing the coefficients from the BLOSUM62 substitution matrix. Affinity reagents are designed to detect off-target and target sites, as calculated using the BLOSUM62 substitution matrix. Twenty additional pairs were selected with scaled binding probabilities according to sequence similarity between the target trimers. Binding to off-target sites. The probability for these additional off-target sites is As follows: TIFF2025163121000066.tif4128Here S OT is the BLOSUM62 similarity between the off-target site and the target site, and S self is the BLOSUM62 similarity between the target sequence and itself. 2.45 x 10 8 The joint probability below Any off-target site has a binding probability of 2.45 x 10 8 is adjusted to have The probability of nonspecific epitope binding is 2.45 x 10 in this example. 8 is.

[0195] The optimal set of 300 affinity reagent mixtures was determined to be both censored and uncensored. For protein estimation, a greedy approach was used to generate: 1) Initialize an empty list of affinity reagent (AR) mixtures to be selected. 2) Candidate affinity reagents (in this example, using the greedy approach detailed in Example 18) Initialize a list of the 300 best ones (calculated using the . 3) (e.g., all human proteins in the Uniprot reference proteome) A set of protein sequences is selected for optimization. 4) Repeat the following until the desired number of AR mixtures is generated: a. Initialize an empty mixture. b. For each candidate AR: i. Simulate the combined outcome using the current mixture with the candidate AR added to it. To simulate. ii. Simulated binding measurements from i. and simulated binding measurements from previously generated mixtures Protein prediction is performed for each protein using simulated binding measurements. iii. The exact protein sequence for each protein as determined by protein prediction. A score is calculated for this mixture with this candidate AR by summing the probabilities of the protein identification. Determine. c. The highest scoring candidate AR is added to the mixture. d. For each candidate AR not previously in the mixture, add the following to the mixture to which the AR is added: iii) and the highest scoring candidate is selected as the mixture. If a candidate has a higher score than the previous candidate added to the mix, add it to the mix, and Repeat this process. The mixture is then split into two groups, with the highest-scoring candidate AR added previously. If all candidate ARs are present in the mixture, the score of the mixture decreases compared to the candidate that was present in the mixture. If it is added, it is completed.

[0196] Figure 19 shows the results of unmixed candidate affinity reagents and mixtures of candidate affinity reagents with truncated protein estimates and demonstrates the sensitivity of protein identification when used with uncensored protein estimates The data plotted in Figure 19 is shown in Tables 10A-10B.

[0197] Table 10A~Table 10B (Table 10A) Binding of individual probes (unmixed) or mixtures of probes (mixed) Performance of censoring estimates using measurements made in TIFF2025163121000067.tif56141

[0198] (Table 10B) Binding of individual probes (unmixed) or mixtures of probes (mixed) Performance of uncensored estimates using measurements made in TIFF2025163121000068.tif56142

[0199] The use of mixtures improves performance when uncensored estimates are used, but not when censored estimates are used. If estimation is used, it may adversely affect performance.

[0200] Example 20 - Glycan Identification Using a Database of Seven Candidate Glycans Consider the situation where the database contains seven candidate glycans: TIFF2025163121000069.tif61162

[0201] In addition, the experiment identified four affinity nucleotides, each of which has a 25% likelihood of binding to a given disaccharide. The other disaccharides to which these reagents bind are listed in the database. It is not found in any glycans.

[0202] A hit table is constructed for the affinity reagents for each sequence in the database. (Rows = affinity reagents #1-#4, columns = SEQ ID) TIFF2025163121000070.tif38170

[0203] Notably, this information arrives gradually and may therefore be calculated iteratively. From the hit table, P(glycan_i | AR_j) is calculated as the probability matrix For a given entry, the hit table is evaluated to generate a If P_Landing_AR_n = true landing rate = 0.25 is used; otherwise, hit table = 0 Note that if P(detector error) = 0.00001 is used. TIFF2025163121000071.tif67166

[0204] Note that many cells contain a probability of 0.00001. This small probability is due to the detector First, the unnormalized glycan probabilities are calculated for each candidate. For the coglycan, it is calculated as the product of probabilities: TIFF2025163121000072.tif16153

[0205] A size normalization is then calculated, which indicates that some number of affinity reagents are required for a given glycoprotein. The number of ways that a glycan can land is expressed as a function of the number of potential binding sites on the glycan. The normalization of the size is expressed as Choose(site_i, n). For example, candidate ID 52 is disaccharide moieties, and a size normalization of [6 choose 4], which is 15. If there are more binding events than the number of disaccharide moieties, the size normalization factor is set to 1. The unnormalized probability of each glycan is calculated as follows to take this size correction into account: It is normalized by dividing by the size normalization, which is shown below: TIFF2025163121000073.tif23165

[0206] Then, the probabilities are calculated so that the entire set of probabilities across the entire database sums to 1. is normalized. This sums up the size-normalized probability to 0.00390641, and This normalization normalizes the size to achieve a final balanced probability. This is achieved by dividing each of the given probabilities: TIFF2025163121000074.tif23170

[0207] Example 21: Censoring Protein Identification in Samples Containing Protein Isoforms performance The protein identification approach described herein is directed to protein isoforms. The isoforms of canonical proteins can be applied to samples containing canonical proteins. formed by alternative splicing of the same gene as the protein, or canonical proteins, which are formed by other genes in the same gene family as the protein A protein isoform can refer to a variant of a protein. It may be structurally similar, typically sharing most of its sequence with the canonical protein. We share.

[0208] Protein samples and affinity reagents Affinity assays were performed to determine the effect of the presence of isoform sequences on protein identification. Drug binding analysis was performed on 20,374 unique canonical human proteins and their canonical tandems. The study was carried out on a protein population consisting of 21,987 unique protein isoforms. Canonical and isoform proteins are listed in the Uniprot database. These are listed in the reference human proteome available as part of the manual It is used to mean that the protein is annotated and reviewed in Only proteins labeled "Swiss-Prot" were included in the analysis. The number of isoforms contained in each non-specific protein is The canonical proteins in this set ranged from 0 to 36. The average number of isoforms is 1.08. Samples were analyzed using a panel of 384 affinity reagents. where each group represents the binding output of a unique affinity reagent to each of the proteins in the sample. Each affinity reagent binds to the target trimer with a probability of 0.25 and The four most similar trimers bind with a probability of 0.25. The other off-target trimers bind with a probability of 2.4 5 x 10 -8 The larger number and 0.25 * 1.5 -x In the latter case, x is the trimeric ratio of the off-target trimer subtracted from the similarity of the target trimer to itself. Similarity between trimer sequences is, for example, similarity between three sequence positions The BLOSUM62 coefficients are calculated by summing the amino acid pairs in each of the Trimeric targets of affinity reagents can be optimized for the human proteome. Selection was performed using a greedy approach as described in Example 18.

[0209] Performance of protein identification using unknown isoform sequences The data set contains only the sequences for 20,374 canonical proteins in the protein sample. Using a database, censoring protein estimates were performed on binding outcomes from samples. The database used for protein prediction consisted of 21,987 proteins in the sample. Because the sequences of protein isoforms are missing, the results of this analysis are limited to the potential protein sequences in the sample. This demonstrates the performance when the sequences of protein isoforms are not known. Using protein estimation, the correct protein family was identified for 83.9% of the proteins in the sample. The term "protein family" is used herein to mean a family of proteins. As used generally, canonical protein sequences and the canonical protein Refers to the set of sequences that includes all isoforms of a sequence. The identities of the predicted proteins in the protein families have been analyzed. The protein is identified if it is within the same protein family as the protein being identified.

[0210] Performance of protein identification using known isoform sequences Protein predictions are based on all protein sequences in the sample (canonical protein sequences and This was performed using a sequence database consisting of both nucleotide and isoform protein sequences. In this case, accurate protein sequences are obtained for 60.9% of the proteins in the sample with a false discovery rate of 1%. When the correct sequence for a protein is identified, the protein is identified. The exact protein sequence is identified. Furthermore, the exact protein family is identified based on the sample. The protein family identification rate and correct identification rate were The discrepancy between the identification rate of a protein sequence and the identification rate of a protein with multiple isoform candidates with similar sequences is due to the This may arise due to the difficulty in distinguishing the identity of proteins between complements.

[0211] Performance of protein identification using a priori defined protein families Protein family of canonical and isoform protein sequences for protein families, where the grouping into Identification rates can be improved by directly calculating protein family probabilities. For each individual protein being measured, the protein is a member of a protein family. The probability of a family being a bar is calculated by adding up the probabilities of each of the individual protein sequences that make up the family. The protein with the highest probability for the protein being analyzed can be calculated by The protein family that corresponds to the protein family is assigned as the protein family identification. When protein family probabilities are calculated in this manner, the exact protein family is , 97.2% of the proteins in the sample are identified with a false discovery rate of 1%. If the protein family probability is not calculated directly, the exact protein family is 89.8% of the proteins in the sample are identified with a false discovery rate of 1%.

[0212] Example 22: Censoring Times in Samples Containing Proteins with Single Amino Acid Variants (SAV) Protein identification performance The protein identification approach described herein identifies proteins with single amino acid mutations. It can be applied to samples containing proteins. What is a single amino acid variant (SAV) of a canonical protein? As used herein, a canonical protein that generally differs by one amino acid. A single amino acid mutant protein is typically a mutation in the It can arise from missense single nucleotide polymorphisms (SNPs) in genes.

[0213] Protein samples and affinity reagents To determine the effect of the presence of SAV proteins on protein identification, affinity reagent binding was performed. The combined analysis identified 20,374 unique canonical human proteins and their This was performed on a protein population consisting of 12,827 unique SAVs of high quality. Proteins are based on the reference human proteolytic protein available as part of the Uniprot database. For each canonical protein, If one or more of all SAVs are present in the SAV database, a randomly selected SAV The SAV database used was Uniprot human polymorphisms. and disease mutations index. Proteins labeled "Swiss-Prot" are used to indicate that the protein is Only proteins were included in the analysis. Samples were analyzed using a panel of 384 affinity reagents. Here, each group represents a unique affinity reagent binding outcome for each of the proteins in the sample. Each affinity reagent binds to the target trimer with a probability of 0.25 and has a maximum affinity to the target trimer. The other off-target trimers bind with a probability of 2.45 x 10 -8 The larger number and 0.25 * 1.5 -xIn the latter case, x is The trimer of the off-target trimer subtracted from the similarity of the target trimer to itself Similarity between trimer sequences is, for example, the similarity of the three sequence positions to the target. The BLOSUM62 coefficients can be calculated by summing the BLOSUM62 coefficients for each amino acid pair in each sequence. The trimer target of the affinity reagent was optimized for the human proteome as described in Example 18. The selection was performed using a greedy approach as described in.

[0214] Performance of protein identification using known SAV sequences Data containing only sequences for 20,374 canonical proteins in the protein sample Using a database, censoring protein estimates were performed on binding outcomes from samples. The database used for protein prediction was 12,827 SAVs in the sample. Due to the lack of protein sequences, the results of this analysis do not provide a complete picture of all potential SAV sequences in the sample. Protein prediction performed in this manner can be used to accurately predict the performance of the protein when it is not known. The SAV protein family was identified for 96.0% of the proteins in the sample with a false discovery rate of 1%. The term "SAV protein family" as used herein is defined as Generally, a canonical protein sequence and all SAs of the canonical protein sequence The exact SAV protein family for a protein is , the deduced protein identity is the same as the protein being analyzed If it is within a protein family, it is identified.

[0215] Performance of protein identification using known SAV sequences Protein predictions are based on all protein sequences in the sample (canonical protein sequences and When performed using a sequence database consisting of both HIV-1 and HIV-2 protein sequences, Potential protein sequences were identified for 27.1% of the proteins in the sample, with a false discovery rate of 1%. When the correct sequence for a protein is identified, the correct sequence for that protein is identified. The protein sequence is identified. Furthermore, the exact SAV protein family is identified based on the type of protein in the sample. The identification rate of SAV protein families and the correct typing rate were The discrepancy between the identity of the canonical protein sequence and the identification rate of the protein sequence is Due to the difficulty in distinguishing between the identities of highly similar SAV sequences, You can.

[0216] Performance of protein identification using a priori defined SAV protein families The identification rate for SAV protein families directly correlates with the probability of SAV protein families. This can be improved by calculating the individual proteins being measured. The probability that a protein is a member of the SAV protein family is The probability of each individual protein sequence can be calculated by summing the probabilities. The SAV protein family with the highest probability for the protein being identified is the SAV protein family. The probability of the SAV protein family being assigned as an identification of the protein family is When calculated in this manner, the exact SAV protein family is 96. 5% of the proteins are identified with a false discovery rate of 1%. If not calculated directly, the exact SAV protein family is estimated to be 96. 1% will be identified with a false discovery rate of 1%.

[0217] Example 23: Protein estimation of censoring in samples containing proteins from a mixture of species Performance In some cases, the protein sample contains proteins from each of multiple species. The protein sample may include proteins derived from external sources such as fossils. In some embodiments, the protein sample is a recombinant protein or an in vitro protein. Proteins synthesized, modified, or derived from other organisms, such as proteins synthesized by transcription and translation may comprise genetically engineered proteins. In some embodiments, synthetic Modified or engineered proteins contain non-naturally occurring sequences (e.g., CRISPR-C These may include those resulting from modifications with as9 or other artificial genetic constructs. Each may be, for example, a mammal (e.g., a human, mouse, rat, primate, or monkey). ), livestock (beef cattle, dairy cattle, poultry, horses, pigs, etc.), sport animals, companion animals ( pet or support animal); plant, protist, bacterium, virus, or archaea That's fine.

[0218] In this example, samples from tumor xenograft mouse models were obtained from both murine and human origin. The protein estimates may contain substantial amounts of proteins of both species origin. To determine the performance of protein prediction in samples with proteins from the mixture, affinity Reagent binding assays were performed on 2,000 unique mouse proteins and 2,000 unique human proteins. This was performed on a protein population consisting of human and mouse proteins. Both proteins were compared with the canonical sequence in the Uniprot reference proteome of each species. The samples were randomly selected from a population of ss-Prot sequence entries. where each group has a unique affinity for each of the proteins in the sample. The binding outcome of the reagents is measured. Each affinity reagent binds to the target trimer with a probability of 0.25. and binds to the four trimers that are most similar to the target trimer with a probability of 0.25. The ribonucleotide trimer is 2.45 x 10 -8 The larger number and 0.25 * 1.5 -x The probability of bonding is where x is the off-target trimer subtracted from the similarity of the target trimer to itself. The similarity between the trimer sequences is the similarity of the target trimer to the trimer target. For example, summing the BLOSUM62 coefficients for the amino acid pairs at each of the three sequence positions The trimeric target of the affinity reagent is optimal for the human proteome. To achieve this, a greedy approach was used to select the desired sequences, as described in Example 18.

[0219] Protein predictions were compared with the human proteome (Uniprot human reference proteome, Contains only sequences for candidate proteins from the canonical Swiss-Prot sequence entry When performed on mixture samples using the database, results showed a false discovery rate of 1%. There were no protein identifications in the sample below a threshold (e.g., 0% identification rate). In comparison, protein predictions were performed for the human proteome and We used a database containing sequences for candidate proteins from both the mouse proteome and the mouse genomic DNA. When performed with 85.3% of the proteins in the sample were below the 1% false discovery rate threshold. This difference in performance was identified in samples containing proteins from multiple species (e.g., For mixture samples, the performance of protein identification is evaluated based on the analysis of protein predictions. A database containing sequences for candidate proteins from all species represented in the data was used. This shows that there is a significant improvement when implemented in

[0220] Example 24: Design of affinity reagent sets against a panel of protein targets Affinity reagents optimized for the identification of specific subsets of proteins in a sample. For example, an optimal set of affinity reagents can be designed for proteome-wide identification. Target binding with fewer affinity reagent binding groups compared to using a set optimized for In this example, the parent A set of compatible reagents is a potential biomarker for clinical response to cancer immunotherapy treatment. For optimal identification of 25 human proteins, a target panel was generated. Proteins are listed in Table 11.

[0221] Table 11. Proteins included in the target panel for response to cancer immunotherapy TIFF2025163121000075.tif180155

[0222] To generate a set of affinity reagents optimized for complete proteome identification, A greedy selection approach was applied as described in Example 18. This selection of affinity reagents The set can be referred to as a "proteome-optimized" affinity reagent set. To generate a set of affinity reagents optimized for protein identification, the methods described in Example 18 were used. A modified version of step 4) i) is carried out, in which the protein is determined by protein estimation. , by summing each of the probabilities of correct protein identification for each protein. Instead of calculating a score for the candidate affinity reagent, The score is a measure of the probability of correct protein identification for only the proteins in the target panel. This affinity reagent set is called a "panel-optimized" set. A set of affinity reagents optimized for the proteome can be referred to as a "proteome-optimized" affinity reagent set. The performance of the panel-optimized affinity reagent set was verified using Swiss-Pro from Uniprot. All unique canonical proteins (20,374 proteins) in the human reference proteome The samples were tested in a human proteome sample containing a target panel of Both affinity reagent sets were used to analyze protein samples. and the truncation estimate is used to determine the censoring for all proteins in the sample. was used to generate protein identification.

[0223] Proteins in a panel of targets identified by a proteome-optimized affinity reagent set The number of proteins and the number of target proteins identified by the panel-optimized affinity reagent set The number of proteins in the panel is shown in Table 12. The number of target proteins that should be counted as successful identifications is For proteins in the panel, the proteins were identified in the samples with a false discovery rate below 1%. The identification must be in the list of all proteins that are identified. For example, 150 affinity reagents were used to predict protein activity. from either the roteome-optimized or panel-optimized sets. The data set included analyses using the top 150 affinity reagents. where each affinity reagent is analyzed in an individual group.

[0224] Table 12. Performance of protein identification for a 25 target panel of target proteins TIFF2025163121000076.tif80158

[0225] The results shown in Table 12 demonstrate that application of panel-optimized affinity reagents significantly improved the affinity of target panels. The results show that the panel-optimized affinity reagent set successfully increased the protein identification rate. <1% for both the ELISA kit and the proteome-optimized affinity reagent set The percentage of all proteins identified with false discovery rates is shown in Table 13.

[0226] Table 13. Performance of protein identification for all proteins in the sample TIFF2025163121000077.tif83158

[0227] The results shown in Table 13 demonstrate the ability to identify sets of proteins in specific target panels. To improve this, we show that panel-optimized affinity reagent sets can be generated. However, trade-offs can occur, and the overall protein content of the panel-optimized reagents in Table 13 The resulting panel-optimized The set of affinity reagents is optimal for identifying proteins outside the target panel. It is possible that this is not the case.

[0228] Example 25: Performance of protein prediction using detection of the presence, number, or order of individual amino acids The protein prediction approach described herein is directed to proteins and peptides. For example, the method can be applied to the determination of specific amino acids in proteins or peptides. Presence or absence (binary) of an amino acid, In proteins, the number (number) of amino acids indicates the order (order) of certain amino acids in a protein. In this example, the proteins are each associated with a specific antigen. The amino acids are selectively modified by a series of reactions. Each reaction in the series is The reaction has a probability of success between 0 and 1, which means that the reaction is successful at any one amino acid in the protein. The probability of successfully modifying an acid substrate is shown. After administration, the presence or absence of the selectively modified amino acid can be detected, The number of modified amino acids can be detected and / or the number of selectively modified amino acids in the protein can be determined. The order of a particular set of selected amino acids can be detected.

[0229] Detection from the determination of the presence and absence of amino acids To generate protein identifications from a series of binary measurements indicating the presence or absence of amino acids. Therefore, the probability Pr(detection of the presence of an amino acid | protein) is 1 - (1 - R aa ) Caa Shown as can be expressed as aa is the availability of the reaction for that amino acid, and Caa is the The probability Pr(non-detection of the presence of an amino acid) is the number of times that an amino acid occurs in a protein. Protein) can be expressed as 1 - Pr(detection of the presence of amino acids | protein). When a series of multiple measurements of amino acid detection are made, the candidate tag is Given a protein, the probability is calculated as , can be multiplied by: Pr(outcome set | protein) = Pr(measured outcome for amino acid 1 | protein) Protein) * Pr(Amino Acid 2 Measured Outcome | Protein) * … Pr(Amino Acid N Measured Outcomes for | Protein)

[0230] A particular candidate protein is the correct identity for the protein being measured. The probability is TIFF2025163121000078.tif9128, where TIFF2025163121000079.tif5128 is a possible protein sequence in a protein sequence database consisting of P proteins. It is the sum of the probabilities of the outcome set for each quality.

[0231] Detection from measuring the number of amino acids To generate a protein identification from a set of amino acid number measurements, we use the probability Pr(amino acid Measuring the number of acids | proteins) TIFF2025163121000080.tif6128, where R aa is the availability of the reaction for that amino acid, and Caa is is the number of times the amino acid occurs in the protein, and M is the number of times the amino acid occurs in the protein. is the number measured for the acid. If M > Caa, a probability of 0 is returned. When multiple measurements of a series of amino acids are made, the candidate protein is Given a quality, the probability is multiplied to determine the probability of the complete set of N measurements. Possible: Pr(outcome set | protein) = Pr(measured outcome for amino acid 1 | protein) Protein) * Pr(Measured outcome for amino acid 2 | Protein) * … Pr(Measured outcome for amino acid N Measured Outcomes | Protein

[0232] A particular candidate protein is the correct identity for the protein being measured. The probability is It can be represented as TIFF2025163121000081.tif9128, where TIFF2025163121000082.tif5128 is a possible protein sequence in a protein sequence database consisting of P proteins. It is the sum of the probabilities of the outcome set for each quality.

[0233] Detection from measurement of amino acid order In some embodiments, the order of selectively modified amino acids in the protein is determined by measuring For example, a protein with the sequence TINYPRTEIN can be modified at amino acids I and N. When detected and measured, it can produce a measured outcome, ININ. The quality is determined by the measurement outcome if a subset of amino acid modifications and / or measurements are not successful. The probability Pr(measured outcome | protein) can be calculated as Pr(aa_co unts | protein) * NUMORDER. TIFF2025163121000083.tif6128, where R aai is the availability of reaction for amino acid i, and M i is the amino acid i measured (For example, for the measured outcome INN, N is the number of times measured twice.) ), C aaiis the number of times that amino acid i appears in the sequence of the candidate protein, and ~L are all of the unique amino acids measured in the protein (e.g., the measured amino acids For ATP, I and N are measured. If the number of occurrences is greater than the number of times the amino acid appears in the candidate protein sequence, the probability Pr(aa_counts | protein) is set to zero. NUMORDER is the number of aa counts from the protein sequence. The number of variants of a particular outcome that can be generated. For example, the measured outcome IN can be produced from the protein TINYPRTEIN in the following manner: {T IN YPRTEIN, T I NYPRTEI N , TINYPRTE IN} Thus, NUMORDER is 3 for this particular outcome and protein sequence. NUMORDER is when it is not possible to generate a specific outcome from a protein Note that in (These proteins cannot be generated from the protein TINYPRTEIN.) The probability of a correct identification for a given protein is It can be expressed as TIFF2025163121000084.tif8128, where TIFF2025163121000085.tif5128 is a possible protein sequence in a protein sequence database consisting of P proteins. It is the sum of the probabilities of the measured outcomes for each quality. In the case where TIFF2025163121000086.tif5128 is equal to zero, the probability of the candidate protein is set to zero.

[0234] A group of reagents for selective modification and detection of the amino acids K, D, C, and W. The performance of protein identification is shown in Figure 22 and Table 14. The response is expressed as: Conducted with varying effectiveness. The type of detection (either "binary", "numerical", or "order"). These can be used to detect the presence or absence of an amino acid, the number of amino acids, or The detection of the amino acid sequence (or the order of amino acids) is indicated by the shading of each bar. The height of each bar represents 1%. The percentage of proteins in the samples that are identified with a false discovery rate below 0.01 is shown. was a human protein sample containing 1,000 proteins. With the effectiveness of the method, a significant number of proteins can be identified using the determination of the amino acid sequence. When a measure of the number of amino acids is used, a significant number of The presence or absence of an amino acid can be determined by the conditions tested. And there wasn't enough of it to generate protein detection.

[0235] Table 14. Selective modification and detection of four amino acids (K, D, C, and W) Protein identification performance TIFF2025163121000087.tif105128

[0236] As shown in Figure 23, a collection of reagents for selective modification and detection of amino acids is available. The types of amino acids are R, H, K, D, E, S, T, N, Q, C, G, P, A, V, I, L, M, F, Y, and W The type of detection is indicated by the shading of the line, and the effectiveness of the response is is shown on the x-axis. The y-axis is the percentage of proteins in the sample identified with a false discovery rate below 1%. Show percentage.

[0237] The results shown in Figure 23 and Table 15 indicate that such a population of reagents has a reaction efficiency of about 0.6 or better. It is very useful in protein identification when a measure of the number of amino acids is used. However, the presence or absence of an amino acid is not determined by the amino acid. When used instead of measuring acid counts, only a small percentage of proteins are identified It will not be done.

[0238] (Table 15) 20 amino acids (R, H, K, D, E, S, T, N, Q, C, G, P, A, V, I, L, M) Performance of Protein Identification Using Selective Modification and Detection of α, F, Y, and W TIFF2025163121000088.tif150128

[0239] FIG. 24 shows the performance of protein identification using the determination of the order of amino acids, where the amino acids are , measured by the probability of detection (equal to the effectiveness of the response) shown on the x-axis. The y-axis shows the probability of detection below 1%. The false discovery rate (false discovery rate) indicates the percentage of proteins in the sample that are identified. The amino acid sequence was measured in the N-terminal 25, 50, 100, or 200 amino acids of the The database of candidate protein sequences is based on the Uniprot Reference Human Protein The first 25, 50, and 100 of each canonical protein sequence in the quality database, respectively or consisted of 200 amino acids.

[0240] The performance shown in Figure 24 and Table 16 shows that at least the highest number of each protein was detected with a probability of detection of approximately 0.3. Sequencing the first 100 amino acids is shown to be optimal. At this rate, sequencing the first 25 or more amino acids is considered sufficient.

[0241] Table 16. Performance of protein identification using amino acid order determination TIFF2025163121000089.tif230154TIFF2025163121000090.tif141154

[0242] Figure 25 shows the results of a tryptic digest of a sample of 1,000 unique human proteins. The performance of the various approaches is shown. Samples are generated from these proteins, which fail All fully tryptic peptides greater than 12 in length that were cleaved without any tryptic activity were The dark lines represent all the results measured at varying detection probabilities (equivalent to the efficacy of the response). Performance is shown when protein identification is performed using measurements of all amino acid sequences. The lighter lines indicate that only the order of the amino acids K, D, W, and C affects the detection probability (effectiveness of the reaction). The performance is shown as measured at a sparse (equivalent) sequence database used for the estimation. The sequences were compared in the human reference proteome database downloaded from Uniprot. The full canonical protein sequence in Completely tryptic peptides with a length greater than 12 that were cleaved without The solid line indicates the percentage of peptides in the sample identified with a false discovery rate below 1%. The dashed line indicates the percentage of proteins in a sample that are identified with a false discovery rate below 1%. Percentages are shown. Proteins contain 1% or less of peptides with sequences unique to that protein. These results indicate that the amino acids K, D, and W are identified when the false discovery rate is below 0. Measurement of C alone is not sufficient for protein detection from tryptic digest samples. Furthermore, the probability of detection (equivalent to the effectiveness of the reaction) is approximately 0.5 or higher. The determination of all amino acids in order identifies the majority of proteins in the tryptic digest. is sufficient.

[0243] Computer Control System The present disclosure also provides a computer-controlled system programmed to carry out the methods of the present disclosure. Figure 10 shows a computer system 1001, which: receiving empirical measurement information for the protein; The observed protein sequences are compared against a database containing multiple protein sequences corresponding to the protein. to generate a probability that the candidate protein will produce a set of measured outcomes, and and / or to generate a probability that a candidate protein will be correctly identified in a sample. It is programmed or otherwise configured to do so.

[0244] The computer system 1001 is capable of implementing various aspects of the disclosed methods and systems, e.g. For example, receiving information on empirical measurements of unknown proteins in a sample; The information is compared with a database containing multiple protein sequences corresponding to the candidate protein. and comparing the probability that the candidate protein produces the observed set of measured outcomes. generating a probability that a candidate protein will be correctly identified in a sample; It is possible to control the process of forming the material.

[0245] The computer system 1001 may be connected to a user's electronic device or remotely to the electronic device. The electronic device may be a computer system located at a computer system location. The electronic device may be a portable electronic device. Computer system 1001 includes a central processing unit (CPU, also referred to herein as a "processor"). and "computer processors") 1005, which may be single-core or It can be a multi-core processor, or multiple processors for parallel processing. The computer system 1001 also includes a memory or memory location 1010 (e.g., a random access memory, read-only memory, flash memory), electronic storage unit unit 1015 (e.g., hard disk) for communication with one or more other systems a communication interface 1020 (e.g., a network adapter), and a cache , other memory, data storage and / or electronic display adapters The peripheral device 1025 includes a memory 1010, a storage unit 1015, an interface 1020, and The peripheral device 1025 communicates with the CPU 1 via a communication bus (physical wiring) such as a motherboard. 005. The storage unit 1015 is a data storage unit for storing data. The computer system 1001 may be a communication unit (or data repository). With the aid of an interface 1020, a computer network ("network") ) 1030. The network 1030 may be the Internet , an internet and / or extranet, or Intranet and / or extranet communicating with the Internet Network 1030 may, in some cases, be a telecommunications and / or data The network 1030 is a distributed computing system such as a cloud computing system. may include one or more computer servers that may enable computing For example, one or more computer servers may be configured to perform the analysis, calculations, and A cloud computing system on a network 1030 ("cloud") is used to perform various aspects of the production. This aspect may enable cloud computing, for example, to identify unknown proteins in a sample. receiving empirical measurement information, the empirical measurement information corresponding to the candidate protein; a step of comparing the observed measurement values ​​against a database containing a plurality of protein sequences corresponding to the observed measurement values; generating a probability that the candidate protein will produce an outcome set, and / or Such a process generates a probability that a protein will be correctly identified in a sample. Cloud computing is a technology that is used by, for example, Amazon Web Services (AWS), Microsoft Cloud services such as Microsoft Azure, Google Cloud Platform, and IBM Cloud The network 1030 may be provided by a network computing platform. In some cases, with the aid of the computer system 1001, Allows devices coupled to the system 1001 to function as either clients or servers. A peer-to-peer network may be implemented.

[0246] The CPU 1005 is a machine-readable The instructions may be stored in a memory location, such as memory 1010. The instructions may be directed to the CPU 1005, which then executes the program. CPU 1005 may be programmed or otherwise configured to perform the methods shown. Examples of operations performed by the CPU 1005 are fetch, decode, execute, and writeback. may include:

[0247] The CPU 1005 may be part of a circuit such as an integrated circuit. Other components may be included in the circuit. In some cases, the circuit is application specific. It is an application specific integrated circuit (ASIC).

[0248] The storage unit 1015 contains the drivers, libraries, and stored programs. The storage unit 1015 can store files such as user preferences. It can store user data such as references and user programs. The computer system 1001 may, in some cases, be connected to an intranet or the Internet. to a remote server that communicates with computer system 1001 via the Internet. One or more additional devices external to the computer system 1001, such as those located The device may include a data storage unit.

[0249] The computer system 1001 communicates with one or more remote computers via a network 1030. For example, computer system 1001 is , can communicate with a user's remote computer system. Examples of systems are personal computers (e.g., handheld PCs), slates, or tablets. Tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab), mobile phones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®, or personal digital assistant. , accessible to the computer system 1001 via a network 1030 .

[0250] The methods as described herein may be implemented, for example, on memory 1010 or on electronic storage. on an electronic storage location of computer system 1001, such as on storage unit 1015 by machine-executable code (e.g., a computer processor) stored in Machine executable or machine readable code may be called software. During use, the code may be executed by the processor 1005. In some cases, the code may be retrieved from the storage unit 1015 and The data may be stored in memory 1010 for quick access by the sensor 1005. In some cases, the electronic storage unit 1015 may be omitted and the machine-executable instructions may be stored in memory. It is stored in -1010.

[0251] The code may be precompiled and a processor adapted to execute the code may be It can be configured for use on a machine with The code can be pre-compiled or compiled at run time (as-compiled) ) style. can be provided.

[0252] The systems and methods provided herein, such as computer system 1001 Aspects of technology can be embodied in the form of programming. carried or embodied in some type of removable medium, typically a machine (or processor) "Article" in the form of executable code and / or associated data on a processor Machine executable code may be stored in memory (e.g., Read-only memory, random access memory, flash memory) or hardware It may be stored on an electronic storage unit, such as a hard disk. The media for the data may be the tangible memory or various semiconductor memories, tape drives, disk drives and the like may include any or all of the associated modules of the software It can provide non-transient storage at any time during programming. All of the software or portions thereof, the Internet or various other telecommunications networks Such communication may occur, for example, between one computer and another. from one computer or processor to another, for example, a management server or host From your computer to the application server's computing platform, It may be possible to load software onto the device, thus carrying software elements. Another type of media that can be used is one that passes through a physical interface between local devices. those used via wired and optical landline networks, and Includes light waves, radio waves, and electromagnetic waves, such as those used over various air links . Carrying such waves, such as wired or wireless links, optical links, or the like. A physical element may also be considered a medium for carrying the software. As used, unless limited to non-transitory, tangible "storage" media, The terms "computer-readable medium" and "machine-readable medium" refer to any medium that is accessible to a processor for execution. Refers to any medium involved in providing instructions.

[0253] Therefore, machine-readable media, such as computer executable code, Including, but not limited to, tangible storage media, carrier wave media, or physical transmission media Non-volatile storage media can take many forms, including, for example, any computer Any storage device in a computer or the like, or the like, or the data shown in the drawings including optical or magnetic disks, such as may be used to implement databases, etc. Volatile storage media include the main memory of a computer platform. The tangible transmission medium is the bus in the computer system. Coaxial cable; copper wire; and optical fiber. , in the form of electrical or electromagnetic signals, or at radio frequency (RF) and infrared (IR) It may take the form of sound or light waves, such as those generated during data communications. Common forms of computer-readable media thus include, for example: floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic Media, CD-ROM, DVD or DVD-ROM, any other optical media, punched card paper tape, perforated Any other physical storage medium that has a pattern of RAM, ROM, PROM and EPROM, F LASH-EPROM, any other memory chip or cartridge, carrying data or instructions a carrier wave carrying such a carrier wave, a cable or link carrying such a carrier wave, or a computer-readable medium Any other medium from which programming code and / or data may be embedded. Many of these forms of computer-readable media are transmitted to a processor for execution. The instruction sequence may be responsible for conveying one or more sequences of one or more instructions.

[0254] The computer system 1001 may include, for example, an algorithm, combined measurement data, candidate tanks, and the like. User interface for providing user selection of databases and data The device may include or be in communication with an electronic display 1035 including a user interface (UI) 1040. Examples of UI include, but are not limited to, a graphical user interface. It includes GUI and web-based user interfaces.

[0255] The methods and systems of the present disclosure may be performed by one or more algorithms The algorithm is performed by software when executed by the central processing unit 1005. The algorithm receives, for example, information on empirical measurements of unknown proteins in a sample. The information from the empirical measurements can be used to identify multiple proteins corresponding to the candidate protein. A set of observed measurement outcomes that can be compared against a database containing sequences and / or the probability that the candidate protein will be generated. The probability that is correctly identified in the sample can be generated.

[0256] While preferred embodiments of the present invention are shown and described herein, such embodiments It will be apparent to those skilled in the art that the above examples are provided by way of example only. It is not intended to be limited by the specific examples provided herein. The invention has been described with reference to the foregoing specification, but the description and drawings of the embodiments herein are not intended to be limiting. The terms should not be construed in a limiting sense. Numerous variations, modifications, and substitutions are possible. It will now occur to those skilled in the art without departing from the invention. All aspects of the present invention are subject to a variety of conditions and variables. It should be understood that the present invention is not limited to the specific depiction, composition, or relative proportions. Various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It should be understood that the present invention can also be used in any such It is intended to encompass any suitable alternatives, modifications, variations, or equivalents. The following claims define the scope of the invention and all inventions within the scope of these claims are to be construed as limiting the scope of the invention. Define that methods and compositions, and their equivalents, are encompassed thereby. , intend.

Claims

1. A computer-assisted method for identifying proteins in a sample of unknown proteins.

1. A method for carrying out a method comprising the steps of: (a) information about multiple empirical measurements performed on the unknown protein in the sample; receiving, by said computer, information; (b) analyzing at least a portion of said information from said plurality of said empirical measurements for a plurality of proteins. comparing said sequence by said computer to a database containing said sequence; , each protein sequence corresponds to one candidate protein among multiple candidate proteins; process; and (c) comparing said plurality of said instances against said database containing said plurality of protein sequences; Based on the comparison of the at least a portion of the information of the evidence measurement, for each of one or more candidate proteins in the protein, (i) the probability that the candidate protein generates the information of the plurality of empirical measurements; (ii) measuring the plurality of empirical measurements when the candidate protein is present in the sample; is the probability that is not observed, and (iii) the probability that the candidate protein is present in the sample generating one or more of the following by said computer:

2. Two or more of the plurality of empirical measurements: (i) one or more affinity reagent probes for said unknown protein in said sample; a binding measurement for each of said plurality of candidates, The co-protein is configured to selectively bind to one or more candidate proteins. binding measurements; (ii) the length of one or more of said unknown proteins in said sample; (iii) the hydrophobicity of one or more of the unknown proteins in the sample; and (iv) the isoelectric point of one or more of the unknown proteins in the sample.

2. The method of claim 1, selected from the group consisting of:

3. The step of generating a plurality of probabilities comprises generating a plurality of probabilities for each of a plurality of additional affinity reagent probes. and receiving additional information from the additional affinity reagent probes and the binding measurements of the additional affinity reagent probes. Each of the plurality of candidate proteins selectively binds to one or more candidate proteins. The method of claim 1 , wherein the plurality of sensors are configured to be coupled together.

4. for each of said one or more candidate proteins, the confidence level that the candidate protein matches one of the unknown proteins in the sample The method of claim 1 further comprising the step of generating:

5. 10. The method of claim 1, wherein the plurality of affinity reagent probes comprises 50 or fewer affinity reagent probes. The method described.

6. 10. The method of claim 1, wherein the plurality of affinity reagent probes comprises 100 or fewer affinity reagent probes. The method described.

7. 10. The method of claim 1, wherein the plurality of affinity reagent probes comprises 200 or fewer affinity reagent probes. The method described.

8. 10. The method of claim 1, wherein the plurality of affinity reagent probes comprises 300 or fewer affinity reagent probes. The method described.

9. 10. The method of claim 1, wherein the plurality of affinity reagent probes comprises 500 or fewer affinity reagent probes. The method described.

10. 10. The method of claim 1, wherein the plurality of affinity reagent probes comprises more than 500 affinity reagent probes. How to post.

11. generating a paper or electronic report identifying said protein in said sample.

10. The method of claim 1, further comprising:

12. The method of claim 1 , wherein the sample comprises a biological sample.

13. The method of claim 12, wherein the biological sample is obtained from a subject.

14. identifying a disease state in the subject based at least on the plurality of probabilities.

14. The method of claim 13, further comprising:

15. (c) is for each of one or more candidate proteins in said plurality of candidate proteins, (i) the probability that the candidate protein produces the information of the plurality of empirical measurements. by the computer.

2. The method of claim 1, comprising:

16. (c) is for each of one or more candidate proteins in said plurality of candidate proteins, (ii) measuring the plurality of empirical measurements when the candidate protein is present in the sample; The probability that by the computer.

2. The method of claim 1, comprising:

17. (c) is for each of one or more candidate proteins in said plurality of candidate proteins, (iii) the probability that the candidate protein is present in the sample by the computer.

2. The method of claim 1, comprising:

18. 16. The method of claim 15, wherein the measured outcome comprises binding of an affinity reagent probe.

19. 16. The method of claim 15, wherein the measured outcome comprises non-specific binding of the affinity reagent probe. 。

20. 17. The method of claim 16, wherein the measured outcome comprises binding of an affinity reagent probe.

21. 17. The method of claim 16, wherein the measured outcome comprises non-specific binding of the affinity reagent probe. 。

22. 18. The method of claim 17, wherein the empirical measurement comprises binding of an affinity reagent probe.

23. 18. The method of claim 17, wherein the empirical measurement comprises non-specific binding of affinity reagent probes. Law.

24. 10. The method of claim 1, further comprising generating a sensitivity of protein identification with a predetermined threshold. The method described below.

25. 25. The method of claim 24, wherein the predetermined threshold is less than 1% inaccuracy. 。

26. 10. The method of claim 1, wherein the protein in the sample is cleaved or degraded. The method described below.

27. The protein in the sample does not start from the end of the protein. The method described in 1.

28. The empirical measurement includes the length of one or more of the unknown proteins in the sample. The method according to any one of claims 15 to 17.

29. The empirical measurement determines the hydrophobicity of one or more of the unknown proteins in the sample. The method according to any one of claims 15 to 17, comprising:

30. The empirical determination determines the isoelectric point of one or more of the unknown proteins in the sample. The method according to any one of claims 15 to 17, comprising:

31. 10. The method of claim 1, wherein the empirical measurements include measurements performed on a mixture of antibodies. method.

32. wherein the empirical measurements include measurements performed on samples obtained from multiple species. The method described in item 1.

33. The empirical measurement is performed to determine whether a single amino acid variation caused by a non-synonymous single nucleotide polymorphism (SNP) 2. The method of claim 1, comprising measurements performed on a sample in the presence of a specific antigen (SAV).

Citation Information

Patent Citations

  • Methods and systems for identifying proteins

    US20030054408A1