Highly multiplexable analysis of proteins and proteomes

JP2024539610A5Pending Publication Date: 2025-10-16NAUTILUS SUBSIDIARY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024521263
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-11
Filing Date
2022-10-07
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Current proteomics technologies are limited in sensitivity and throughput, failing to comprehensively analyze the human proteome, with approximately 10% of proteins remaining unobserved and existing methods lacking both single molecule sensitivity and high throughput.

Method used

A method involving a binding profile of proteins to multiple affinity reagents, using a database and a decoding algorithm to identify proteins based on binding probabilities, enabling identification of proteins through empirical binding profiles and accounting for non-specific binding and stochastic fluctuations.

Benefits of technology

The method achieves high-throughput and single molecule sensitivity, effectively identifying a large fraction of the human proteome, overcoming limitations of existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

1. A method comprising: (a) providing inputs including: (i) a binding profile comprising a plurality of binding results for binding of an existing protein to a plurality of different affinity reagents, each binding result of the plurality of binding results comprising a measure of binding between the existing protein and a different affinity reagent of the plurality of different affinity reagents; (ii) a database comprising information characterizing or identifying a plurality of candidate proteins; and (iii) a binding model; (b) determining, according to the binding model, a probability for each of the affinity reagents to bind to each of the candidate proteins in the database; and (c) identifying the existing protein as the selected candidate protein having a probability of binding each of the affinity reagents that best matches the binding profile for the existing protein.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 254,420, filed October 11, 2021, which is incorporated by reference in its entirety.

[0002] Some embodiments relate to methods of performing protein binding assays. More particularly, some embodiments relate to methods of performing protein binding assays to identify existing proteins by using a binding profile that includes a plurality of binding results for the binding of an existing protein to a plurality of different affinity reagents. [Background technology]

[0003] The proteome is one of the most dynamic and valuable sources of biological insight. Current proteomic technologies are limited in their sensitivity and throughput, covering at best 35% of the human proteome in a single experiment (see Blume et al., Nat Commun 11, 3662 (2020) and Clark et al., Cell 180, 207 (2020), each of which is incorporated herein by reference). Despite the wealth of insights gained from now routine genomics and transcriptomics studies in biomedical research, a large gap remains between genome / transcriptome and phenotype. Since proteins constitute the main structural and functional components of the cell, proteomics is crucial to bridge this gap. However, in part due to the complex nature of proteins and the proteome and the high dynamic range (approximately 10% of the total) of the amounts of different proteins present in any given cell at any given time, the amount of protein present in any given cell is limited. 9Protein sequencing technology has lagged behind DNA sequencing technology due to the large number of proteins that are predicted to make up the human proteome (see Aebersold et al., Nat Chem Biol 14, 206-214 (2018), which are incorporated herein by reference). Furthermore, approximately 10% of the proteins predicted to make up the human proteome have never been clearly observed (see Omenn et al., J Proteome Res 19, 4735-4746 (2020) and Adhikari et al., Nat Commun 11, 5301 (2020), each of which is incorporated herein by reference).

[0004] Recently, single molecule identification has been envisioned as a method for analyzing small samples (including single cells) and rare proteins (see Alfaro et al., Nat Methods 18, 604-617 (2021) and Restrepo-Perez et al., Nat Nanotechnol 13, 786-796 (2018), each of which is incorporated herein by reference). Traditional bulk identification techniques such as mass spectrometry and immunoassays have been modified for the detection of single proteins (see Keifer & Jarrold, Mass Spectrom Rev 36, 715-733 (2017) and Risin et al., Nat Biotechnol 28, 595-599 (2010), each of which is incorporated herein by reference). Several concepts have been proposed to achieve single molecule protein sequencing. All of these use sequential processes, such as Edman-type degradation, to determine the positional information of amino acids within proteins (Swaminathan, et al. Nat Biotechnol (2018) and Swaminathan, et al., PLoS Comput Biol 11, e1004080 (2015), each of which is incorporated herein by reference) or directed protein translocation through nanopore channels (Kolmogorov et al., PLoS Comput Biol 13, e1005356 (2017), each of which is incorporated herein by reference). However, current methods do not achieve both single molecule sensitivity and high throughput at a level commensurate with the complexity of the human proteome. Thus, there is a need for comprehensive proteome analysis. The present disclosure fulfills this need and provides other advantages. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Blume et al., Nat Commun 11,3662(2020) [Non-Patent Document 2] Clark et al., Cell 180, 207 (2020) [Non-Patent Document 3] Aebersold et al., Nat Chem Biol 14, 206 - 214 (2018) [Non-Patent Document 4] Omenn et al., J Proteome Res 19, 4735 - 4746 (2020) [Non-Patent Document 5] Adhikari et al., Nat Commun 11, 5301 (2020) [Non-Patent Document 6] Alfaro et al., Nat Methods 18, 604 - 617 (2021) [Non-Patent Document 7] Restrepo-Perez et al., Nat Nanotechnol 13, 786 - 796 (2018) [Non-Patent Document 8] Keifer&Jarrold, Mass Spectrom Rev 36, 715 - 733 (2017) [Non-Patent Document 9] Risin et al., Nat Biotechnol 28, 595 - 599 (2010) [Non-Patent Document 10] Swaminathan, et al. Nat Biotechnol (2018) [Non-Patent Document 11] Swaminathan, et al., PLoS Comput Biol 11, e1004080 (2015) [Non-Patent Document 12] Kolmogorov et al., PLoS Comput Biol 13, e1005356 (2017) [Summary of the Invention]

[0005] The present disclosure provides a method for identifying an existing protein, the method comprising: (a) providing input to a computer processor, the input comprising: (i) a binding profile, the binding profile comprising a plurality of binding outcomes for binding of the existing protein to a plurality of different affinity reagents, each binding outcome of the plurality of binding outcomes comprising a measure of binding between the existing protein and a different affinity reagent of the plurality of different affinity reagents, the binding profile comprising positive binding outcomes and negative binding outcomes, (ii) a database comprising information characterizing or identifying a plurality of candidate proteins, and (iii) a binding model for each of the different affinity reagents, (b) determining a probability for each of the affinity reagents to bind each of the candidate proteins in the database according to the binding model, and (c) identifying the existing protein as a selected candidate protein, the selected candidate protein being the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the binding profile for the existing protein. The input may further include (iv) a non-specific binding rate, which includes the probability that a non-specific binding event occurs for one or more of the different affinity reagents.

[0006] 1. A method of identifying an existing protein, comprising the steps of: (a) contacting a plurality of different affinity reagents with a plurality of existing proteins in a sample; (b) obtaining binding data from step (a), said binding data comprising a plurality of binding profiles, each of said binding profiles comprising a plurality of binding results for binding of the existing protein of step (a) to a plurality of different affinity reagents, each binding result of said plurality of binding results comprising a measure of binding between the existing protein of step (a) and a different affinity reagent of said plurality of different affinity reagents, each of said binding profiles comprising a positive binding result and a negative binding result; and (c) obtaining a plurality of binding profiles comprising a plurality of binding results for binding of the existing protein of step (a) to a plurality of different affinity reagents. Also provided is a method comprising the steps of: (d) providing a database containing information characterizing or identifying a number of candidate proteins; (d) providing a binding model for each of the different affinity reagents; (e) determining, for each of the affinity reagents, a probability of binding to each of the candidate proteins in the database according to the binding model; and (f) identifying the existing protein as a selected candidate protein, wherein the selected candidate protein is the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the plurality of binding results for the existing protein.

[0007] The present disclosure provides a detection system comprising: (a) a detector configured to acquire signals from a plurality of binding reactions occurring between a plurality of different affinity reagents and a plurality of present proteins in a sample, (b) a database containing information characterizing or identifying a plurality of candidate proteins, and (c) a computer processor (i) in communication with the database and (ii) processing the signals to generate a plurality of binding profiles, each of the binding profiles comprising a plurality of binding results for the binding of the present protein of (a) to the plurality of different affinity reagents, each of the binding results representing a binding reaction between the present protein of (a) and a different one of the plurality of different affinity reagents. and (iii) process the binding profiles to determine, for each of the affinity reagents, a probability of binding to each of the candidate proteins in the database according to a binding model for each of the affinity reagents; and (iv) output an identification of a selected candidate protein, the selected candidate protein being the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the plurality of binding results for the existing protein.

[0008] A method for identifying existing proteins may be carried out in a detection system, the method comprising: (a) acquiring signals from a plurality of binding reactions carried out in the detection system, the binding reactions comprising contacting a plurality of different affinity reagents with a plurality of existing proteins in a sample; (b) processing the signals in the detection system to generate a plurality of binding profiles, each of the binding profiles comprising a plurality of binding results for binding of the existing protein of step (a) to the plurality of different affinity reagents, each binding result of the plurality of binding results comprising a measure of binding between the existing protein of step (a) and a different affinity reagent of the plurality of different affinity reagents, each of the binding profiles comprising a positive binding result and a negative binding result; and (c) contacting the existing protein of step (a) with a plurality of different affinity reagents. The method may include providing as an input a database containing information characterizing or identifying a plurality of candidate proteins; (d) providing as input to the detection system a binding model for each of the different affinity reagents; (e) processing the plurality of binding profiles in the detection system to determine, according to the binding model, a probability for each of the affinity reagents to bind to each of the candidate proteins in the database; and (f) outputting from the detection system an identification of a selected candidate protein, the selected candidate protein being the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the plurality of binding results for the existing protein. INCORPORATION BY REFERENCE

[0009] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that a publication, patent, or patent application incorporated by reference conflicts with a disclosure contained herein, the specification is intended to supersede and / or take precedence over such conflicting material. [Brief description of the drawings]

[0010] [Figure 1A] FIG. 1A shows the workflow, from sample preparation to data analysis, of the protein identification method.

[0011] [Figure 1B] FIG. 1B shows a representation of the protein decoding that identifies the protein at position A1 as EGFR.

[0012] [Figure 1C] FIG. 1C shows repeated sequential affinity reagent measurements against EGFR showing five unique binding patterns and one off-target binding event.

[0013] [Figure 1D] FIG. 1D shows the number of affinity reagents sufficient for 90% human proteome coverage, varying epitope length (dimer, trimer, tetramer) and the number of epitopes bound by each multiaffinity reagent (asterisks indicate values ​​above 2,000).

[0014] [Figure 1E] FIG. 1E shows the proteome coverage achieved when affinity reagent cycles are measured using either affinity reagents targeting trimer epitopes optimized for the human proteome or one of 20 random sets of trimer targets.

[0015] [Figure 1F] FIG. 1F shows proteome coverage for the human, mouse, yeast and E. coli proteomes measured using an affinity reagent set optimized for human proteome coverage.

[0016] [Figure 2A] FIG. 2A shows coverage of the human proteome for affinity reagents of various binding affinities.

[0017] [Figure 2B] Figure 2B shows the coverage of the human proteome along with non-specific binding to the array surface for affinity reagents of various binding affinities. The area of ​​the circle is proportional to the proteome coverage (also indicated on the circle).

[0018] [Figure 2C] Figure 2C shows the impact of mischaracterization of affinity reagent binding on proteome coverage for different proportions of unknown high affinity epitope targets. All error bars are standard deviations across five replicates.

[0019] [Figure 2D] Figure 2D shows the impact of mischaracterization of affinity reagent binding on proteome coverage for different percentages of identified false high affinity epitope targets. All error bars are standard deviations across five replicates.

[0020] [Figure 2E] Figure 2E shows the effect of mischaracterization of affinity reagent binding on proteome coverage, relative to systematic measurement error in binding probability. All error bars are standard deviations over five replicates.

[0021] [Figure 2F] Figure 2F shows the effect of mischaracterization of affinity reagent binding on proteome coverage, for random measurement errors in binding probability. All error bars are standard deviations across five replicates.

[0022] [Figure 3A]FIG. 3A shows the dynamic range of plasma protein quantification at different protein array sizes. Data are plotted in order of decreasing protein abundance from top to bottom. The dynamic range is the protein abundance divided by the most abundant protein in the sample. The outer width of the contour indicates the percentage of proteins at that abundance (one or multiple copies) deposited on the protein array. The inner width of the contour indicates the percentage of proteins at that abundance detected by the decoding method. The percentages are calculated over a rolling window of 51 proteins. The horizontal grey bar indicates 100%.

[0023] [Figure 3B] Figure 3B shows the dynamic range of protein quantification in HeLa cells at different protein array sizes. Data are presented as described above for Figure 3A.

[0024] [Figure 3C] FIG. 3C shows the reproducibility of quantification (coefficient of variation calculated across five replicates) compared to protein abundance in plasma as a contour plot (density isoproportional contour) with histograms outside the frame.

[0025] [Figure 3D] Figure 3D shows the reproducibility of quantification (coefficient of variation calculated across five replicates) relative to protein abundance in HeLa cells as a contour plot (density isoproportional contour) with histograms outside the frame.

[0026] [Figure 3E] FIG. 3E shows the agreement of protein abundance (identified copy number) measured by the decoding method with the true count of the protein on the array for a single experimental replicate of plasma.

[0027] [Figure 3F]FIG. 3F shows the agreement of protein abundance (identified copy number) measured by the decoding method with the true count of the protein on the array for a single experimental replicate in HeLa cells.

[0028] [Figure 4A] Figure 4A shows the effect of mischaracterization of affinity reagent binding on proteome coverage for different proportions of unknown high affinity (primary) and low to medium affinity (secondary) epitope targets. All coverage measurements are averages over five replicates.

[0029] [Figure 4B] Figure 4B shows the different percentages of identified false high affinity (primary) and low to medium affinity (secondary) epitope targets. All coverage measurements are averages over five replicates.

[0030] [Figure 4C] Figure 4C shows the systematic measurement error of binding probability for a total of 300 affinity reagents affected by different percentages of error occurrences. All coverage measurements are averages over five replicates.

[0031] [Figure 4D] Figure 4D shows random measurement errors of binding probability for a total of 300 affinity reagents affected by different percentages of error occurrences. All coverage measurements are averages over five replicates.

[0032] [Figure 5A]FIG. 5A shows the distribution of protein abundance of proteins in samples quantified by the decoding method in plasma deposited on a protein array and measured on an array with addresses occupied by 1010 proteins. The histogram counts of each group are averaged over five simulated replicate experiments. The Non-Specific Quant Rate shown is the maximum percentage of proteins observed in replicates with poor quantification (more than 10% of signals due to incorrect identification). The percentage of proteins in the sample quantified is shown as a gray line. The average proteome coverage is the percentage of the proteome present in the sample detected by the decoding method (averaged over five replicates). Error bars indicate standard deviation.

[0033] [Figure 5B] Figure 5B shows the distribution of protein abundances of proteins in samples quantified by the decoding method in depleted plasma deposited on a protein array and measured on an array with addresses occupied by 10 proteins. Data were processed and displayed as in Figure 5A.

[0034] [Figure 5C] Figure 5C shows the distribution of protein abundances of proteins in samples deposited on a protein array and quantified by the decoding method in a HeLa cell line measured on an array with addresses occupied by 10 proteins. Data were processed and displayed as in Figure 5A.

[0035] [Figure 5D] Figure 5D shows the distribution of protein abundances of proteins in samples quantified by the decoding method in plasma deposited on a protein array and measured on an array with addresses occupied by 108 proteins. Data were processed and displayed as in Figure 5A.

[0036] [Figure 5E] Figure 5E shows the distribution of protein abundances of proteins in samples quantified by the decoding method in depleted plasma deposited on a protein array and measured on an array with addresses occupied by 108 proteins. Data were processed and displayed as in Figure 5A.

[0037] [Figure 5F] Figure 5F shows the distribution of protein abundances of proteins in samples deposited on a protein array and quantified by the decoding method in a HeLa cell line measured on an array with addresses occupied by 108 proteins. Data were processed and displayed as in Figure 5A.

[0038] [Figure 6A] Figure 6A shows the sensitivity and specificity of the decoding method for non-depleted plasma. The probability threshold for protein identification was varied; log(threshold)=0, -1e-20, -1e-16, -1e-14, -1e-12, -1e-11, -1e-10, -1e-9, -1e-8, -1e-7, -1e-6, -1e-5, -1e-4, -1e-3, -1e-2, -0.1, -0.2 and -0.3. Lower thresholds resulted in higher sensitivity (quantified proteins) but also a higher proportion of non-specific quantification (signals with more than 10% of identifications being false). Points are plotted showing these metrics for each threshold evaluated for each of the five replicate samples (shown as different shapes). Simulations were performed with a data set containing addresses occupied by 1010 proteins and addresses occupied by 108 proteins.

[0039] [Figure 6B] Figure 6B shows the sensitivity and specificity of the decoding method for depleted plasma. Data were processed and displayed as in Figure 6A.

[0040] [Figure 6C] Figure 6C shows the sensitivity and specificity of the decoding method for the HeLa cell line. Data were processed and displayed as in Figure 6A.

[0041] [Figure 7A] Figure 7A shows the dynamic range of protein abundances deposited on arrays of different sizes for non-depleted plasma. Data are plotted in order of decreasing protein abundance from top to bottom. The dynamic range is the ratio of protein abundance to the abundance of the most abundant protein in the sample. The outer width of the contour indicates the percentage of proteins at that abundance (one or more copies) deposited on the array, with the top bar of each contour corresponding to 100%. The percentages are calculated over a rolling window of 51 proteins.

[0042] [Figure 7B] Figure 7B shows the dynamic range of protein abundances deposited on arrays of different sizes for depleted plasma. Data were processed and displayed as in Figure 7A.

[0043] [Figure 7C] Figure 7C shows the dynamic range of protein abundances deposited on arrays of different sizes for HeLa cells. Data was processed and displayed as in Figure 7A.

[0044] [Figure 8A] FIG. 8A shows the dynamic range of protein quantification for depleted blood samples assessed using the decoding method. Protein abundance data is plotted in order of decreasing abundance from top to bottom. The dynamic range is the ratio of protein abundance to the abundance of the most abundant protein in the sample. The outer width of the contour indicates the percentage of proteins at that abundance (one or multiple copies) deposited on the array. The inner width of the contour indicates the percentage of proteins at that abundance detected by the decoding method. The percentages are calculated over a rolling window of 51 proteins. The horizontal bar indicates 100%.

[0045] [Figure 8B] Figure 8B shows the reproducibility of quantification (CV% across five replicates) compared to protein abundance using contour plots with out-of-frame histograms (density isoproportional contours) for depleted blood samples assessed using the decoding method.

[0046] [Figure 8C] FIG. 8C shows the concordance of protein abundance (detected copy number) with the true count of the protein on the array for a single replicate of depleted blood sample assessed using the decoding method.

[0047] [Figure 8D] Figure 8D shows the distribution of fold change error, which is the count of protein copies detected by the decoding method divided by the copies of depleted plasma protein deposited on the array. Detected copies and deposited copies are averaged over the five replicates measured.

[0048] [Figure 9A] Figure 9A shows the quantification reproducibility and precision demonstrated for non-depleted plasma samples assayed in five replicates on an array with addresses occupied by 108 proteins. Quantification reproducibility (CV% across five replicates) is compared to protein abundance using a contour plot (density isoproportional contour) with an outer histogram. Detected and deposited copies are averaged over the five replicates measured.

[0049] [Figure 9B] Figure 9B shows the agreement of the protein amounts (identified copy numbers) measured by the decoding method with the true counts of the proteins on the array, shown for a single replicate of non-depleted plasma. Detected and deposited copies are averaged over the five replicates measured.

[0050] [Figure 9C] Figure 9C shows the distribution of fold change error, which is the count of protein copies identified by the decoding method divided by the copies of the protein deposited on the array, for non-depleted plasma. Copies detected and copies deposited are averaged over the five replicates measured.

[0051] [Figure 9D] Figure 9D shows the quantification reproducibility and precision demonstrated for depleted plasma assayed in five replicates on an array with addresses occupied by 108 proteins. Quantification reproducibility (CV% across five replicates) is compared to protein abundance using a contour plot (density isoproportional contour) with an outer histogram. Detected and deposited copies are averaged over the five replicates measured.

[0052] [Figure 9E] Figure 9E shows the agreement of the protein amounts (identified copy numbers) measured by the decoding method with the true counts of the proteins on the array, shown for a single replicate of depleted plasma. Detected and deposited copies are averaged over the five replicates measured.

[0053] [Figure 9F] Figure 9F shows the distribution of fold change error, which is the count of protein copies identified by the decoding method divided by the copies of the protein deposited on the array, for depleted plasma. Copies detected and copies deposited are averaged over the five replicates measured.

[0054] [Figure 9G]Figure 9G shows the quantification reproducibility and precision demonstrated for HeLa cells assayed in five replicates on an array with addresses occupied by 108 proteins. Quantification reproducibility (CV% across five replicates) is compared to protein abundance using a contour plot (density isoproportional contour) with an outer histogram. Detected and deposited copies are averaged over the five replicates measured.

[0055] [Figure 9H] Figure 9H shows the agreement of protein amounts (identified copy numbers) measured by the decoding method with the true counts of proteins on the array, shown for a single replicate of HeLa cells. Detected copies and deposited copies are averaged over the five replicates measured.

[0056] [Figure 9I] Figure 9I shows the distribution of fold change error, which is the count of protein copies identified by the decoding method divided by the copies of the protein deposited on the array, for HeLa cells. Detected copies and deposited copies are averaged over the five replicates measured.

[0057] [Figure 10A] FIG. 10A shows the reproducibility of protein deposition and protein quantification across five replicates for non-depleted plasma measured on an array with addresses occupied by 1010 proteins. The amount of protein deposited is the total number of proteins successfully deposited on the array. The amount of protein measured is the number of times a protein was identified by the decoding method. To demonstrate the agreement of the variation in the number of proteins deposited with the variation in the number of proteins measured, the CV (%) of each of these amounts across the five replicates is calculated for each unique protein detected in the sample and plotted using a contour plot.

[0058] [Figure 10B] Figure 10B shows the reproducibility of protein deposition and protein quantification across five replicates for HeLa cells measured on an array with addresses occupied by 10 proteins. Data were processed and displayed as described for Figure 10A.

[0059] [Figure 11] Figure 11 shows the fold change measurement error distribution of proteins detected in plasma samples measured on addresses occupied by 10 proteins. The fold change error is the count of protein copies detected by the decoding method divided by the copies of the protein deposited on the array. The detected copies and deposited copies are averaged over the five replicates measured.

[0060] [Figure 12] FIG. 12 illustrates a computer system that is programmed or otherwise configured to carry out the methods described herein.

[0061] [Figure 13] FIG. 13 shows the predicted non-join probability by sequence length for different semi-truncated decoding approaches.

[0062] [Figure 14] FIG. 14 shows the disjoint probability prediction for sequences of any length using different semi-truncated decoding approaches. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0063] Proteins can be detected using one or more affinity reagents that have a known or measurable binding affinity for the protein. For example, the affinity reagent can bind the protein to form a complex, and the signal generated by the complex can be detected. The protein detected by binding to a known affinity reagent can be identified based on the known or predicted binding properties of the affinity reagent. For example, to identify a candidate protein in a sample by simply observing a binding event, an affinity reagent that is known to selectively bind the candidate protein suspected to be present in the sample without substantially binding to other proteins in the sample can be used. This one-to-one correlation of affinity reagent to candidate protein can be used for the identification of one or more proteins. However, as the complexity of the proteins in a sample (i.e., the number and diversity of different proteins) increases, the time and resources to create a corresponding variety of affinity reagents with one-to-one specificity for these proteins approaches the limit of practicality.

[0064] The present disclosure provides methods, systems and compositions that can be advantageously used to overcome these limitations. In certain configurations, the number of different proteins identified can exceed the number of affinity reagents used. For example, the number of proteins identified can be at least 5 times, 10 times, 25 times, 50 times, 100 times, or more than the number of affinity reagents used. As described in further detail herein, one or more extant proteins can be identified by (1) performing a binding reaction using promiscuous affinity reagents that bind to multiple different candidate proteins suspected to be present in a given sample, (2) subjecting one or more extant proteins to a set of promiscuous affinity reagents that, taken as a whole, generate an empirical binding profile for each extant protein, and (3) performing a decoding method that evaluates the empirical binding profile according to a binding model for the binding of the promiscuous affinity reagents to multiple candidate proteins, thereby identifying individual extant proteins based on their compatibility with each candidate protein.

[0065] Promiscuity of an affinity reagent is a characteristic that can be understood in relation to a given population of proteins. Promiscuity can occur due to an affinity reagent recognizing an epitope present in multiple different proteins known or suspected to be present in a sample, such as a human proteome sample. For example, a promiscuous affinity reagent may recognize an epitope having a relatively short amino acid length, such as a dimer, trimer, tetramer, pentamer or hexamer, which is expected to be present in a significant number of different proteins in the proteome of a human or other species. Alternatively or additionally, a promiscuous affinity reagent can recognize different epitopes (i.e., epitopes having a variety of different structures), which are present in multiple different proteins in a proteome sample. For example, a promiscuous affinity reagent can have a high probability of binding to a primary epitope target and a lower probability of binding to one or more secondary epitope targets, which have a different sequence of amino acids compared to the primary epitope target. The secondary epitope target can be biosimilar to the primary epitope target, for example, according to the BLOSUM62 scoring matrix.

[0066] A single binding reaction between a promiscuous affinity reagent and a complex protein sample, such as a human proteome sample, may result in ambiguous results regarding the identity of the different proteins to which the promiscuous affinity reagent binds, but the ambiguity can be resolved when the results are evaluated with the decoding method described herein. The multiple binding results obtained from measuring the binding of multiple affinity reagents to one or more existing proteins may be input into the decoding method of the present disclosure to identify the most likely identity of that protein among a set of candidate proteins. The multiple binding results may be input into the decoding method along with information characterizing or identifying the multiple candidate proteins (e.g., amino acid sequences of the candidate proteins) and a binding model. The probability of each affinity reagent binding to all possible candidate proteins may be evaluated using the binding model, and the decoding method may output the identity of each existing protein. For example, the decoding algorithm may output the most likely identity for each existing protein as the candidate protein that best matches the observed binding results for the existing protein according to the binding model.

[0067] The binding model of the present disclosure may be constructed based on the assumption that the properties of affinity reagents that bind to existing proteins in a sample, even if unknown, can be treated as quantifiable random variables, and the uncertainty regarding the binding properties can be described by a probability distribution. The parameters of the affinity reagents may be determined, for example, based on prior knowledge of the affinity reagents (e.g., expected binding affinity for a particular epitope) and / or based on preliminary reactions performed using the affinity reagents (e.g., measurements of binding between the affinity reagents and one or more epitopes). The parameters of the affinity reagents may be treated as "priors" that are input to the decoding algorithm of the present disclosure. The parameters of the affinity reagents, when combined with empirically determined binding results and evaluated using the decoding method of the present disclosure, may output "posterior values," the calculation of which includes the calculation of a distribution of likelihoods for the identity of each existing protein used for the empirical determination. The posterior values ​​output by the decoding method may be used to update the prior values ​​used as input to subsequent evaluations using the decoding method. Thus, the influence of unknowns and artifacts in the initial evaluation of affinity reagents can be reduced as further empirical measurements are made and the results evaluated by the decoding method. This update cycle can provide the advantage of facilitating iterative improvements to the decoding method, thereby improving the accuracy with which existing proteins are identified or characterized.

[0068] An advantage of the decoding methods described herein is that they take into account features of the binding reaction that may otherwise adversely affect the accuracy with which a protein can be identified. For example, binding reactions performed at a single molecule scale (e.g., detection of binding of affinity reagents to individually resolved proteins on a protein array) will result in probabilistic results. In addition, non-specific binding of affinity reagents to the surface of the array to which the protein being observed is attached, for example, may also result in erroneous results. Another example is bias or distortion that may arise due to different lengths of the proteins analyzed in the decoding methods described herein. The decoding methods may be configured to take into account chance, non-specific binding, differences in protein length, or other factors to improve accuracy when identifying or characterizing proteins. For example, chance may be taken into account by estimating protein likelihood using the decoding method. Similarly, differences in protein length may be taken into account by calculating a normalization factor that depends on both the candidate protein length and the number of positive binding results observed.

[0069] For ease of explanation, the compositions, systems and methods of the present disclosure are often illustrated herein in the context of characterizing proteins using binding measurements. The examples described herein can be readily extended to characterizing other analytes (e.g., as an alternative or addition to proteins) or to performing other reactions (e.g., as an alternative or addition to binding reactions).

[0070] The present disclosure provides compositions, systems and methods that may be useful in various configurations for characterizing analytes, such as proteins, nucleic acids, cells or portions thereof, by obtaining multiple distinct and non-identical measurements of the analyte. In certain configurations, an individual measurement may not be accurate or specific enough to characterize alone, but a collection of multiple non-identical measurements may allow characterization with a high degree of accuracy, specificity and reliability. In some cases, a collection of multiple measurements using the same affinity reagent (e.g., repeating a binding reaction in triplicate) may allow characterization with a high degree of accuracy, specificity and reliability. In some cases, multiple promiscuous reagents may be reacted with a given analyte, and the observed reaction results for each of the promiscuous reagents may be detected. Promiscuous reagents may exhibit both low specificity for a variety of different recognized analytes and high reactivity for some or all of those analytes. Taking a binding reaction as an example, a promiscuous affinity reagent may exhibit both low specificity for a variety of different recognized analytes and high affinity for some or all of those analytes. For any of a variety of reactions, including but not limited to binding reactions, a first reaction performed using a first promiscuous reagent may recognize a first subset of analytes in a sample without distinguishing one analyte in the subset from another analyte in the sample. A second reaction performed using a second promiscuous reagent may recognize a second subset of analytes in a sample, again without distinguishing one analyte in the second subset from another analyte. However, the combination of measurements from the first and second reactions may distinguish between (i) an analyte that is uniquely present in the first subset but not in the second subset; (ii) an analyte that is uniquely present in the second subset but not in the first subset; (iii) an analyte that is uniquely present in both the first and second subsets; or (iv) an analyte that is uniquely absent in the first and second subsets.The number of promiscuous reagents used, the number of separate measurements taken, and the degree of promiscuousness of the reagents (e.g., the diversity of components recognized by the reagents) can be adjusted to suit the known or suspected diversity of different analytes in a given sample.

[0071] The compositions, systems, or methods described herein may be used to characterize an analyte or portion thereof with respect to any of a variety of properties or characteristics, including, for example, presence, absence, quantity (e.g., amount or concentration), chemical reactivity, molecular structure, structural integrity (e.g., full length or fragmentation), maturation state (e.g., presence or absence of a pre- or pro-sequence in a protein), location (e.g., in an analytical system such as an array, a subcellular compartment, a cell, or the natural environment), association with another analyte or moiety, binding affinity for another analyte or moiety, biological activity, chemical activity, etc. Analytes may be characterized by a common structural feature (e.g., for proteins, amino acid sequence length, overall charge, or overall pK a Analytes may be characterized with respect to relatively general features, such as the presence or absence of a common portion (e.g., for proteins, a short primary sequence motif or post-translational modification). Analytes may be characterized with respect to relatively specific features, such as a unique amino acid sequence (e.g., for the full length of a protein or a motif), an RNA or DNA sequence encoding the protein (e.g., for the full length of a protein or a motif), or an enzymatic or other activity that specifies the protein. Characterization may be sufficiently specific to identify the analyte, for example, at a level deemed appropriate or unambiguous by one of skill in the art. Analytes may be identified with a probability or score above a desired threshold for reliable identification.

[0072] The disclosed methods, compositions and systems can be advantageously deployed in situations where proteins produce different empirical binding profiles despite having the same primary structure and being subjected to the same set of affinity reagents. For example, the methods, compositions and systems are well suited for single molecule detection and other formats prone to stochastic variation. Certain configurations of the compositions, systems and methods herein can overcome ambiguity and errors in observed binding results to provide accurate identification and characterization of proteins. The methods can be advantageously deployed for complex samples including proteomes or subfractions thereof.

[0073] Terms used herein are understood to have their ordinary meaning in the relevant art unless otherwise specified. Some terms used herein and their meanings are listed below.

[0074] As used herein, the term "address" refers to a location within an array where a particular analyte (e.g., a protein, peptide, or unique identification label) is present. An address can contain a single analyte or can contain a population of several analytes of the same species (i.e., a collection of analytes). Alternatively, an address can contain a population of different analytes. Addresses are typically separated. Separate addresses can be adjacent or separated by interstitial spaces. Arrays useful herein can have addresses separated by, for example, 100 microns, 10 microns, 1 micron, 100 nm, less than 10 nm, or less. Alternatively or additionally, arrays can have addresses separated by at least 10 nm, 100 nm, 1 micron, 10 microns, or 100 microns. Addresses can each have an area of ​​less than 1 square millimeter, 500 square microns, 100 square microns, 10 square microns, 1 square micron, 100 square nm, or less. Arrays can have an area of ​​at least about 1x10 4 , 1x10 5 , 1x10 6 , 1x10 7 , 1x10 8, 1x10 9 , 1x10 10 , 1x10 11 , 1x10 12 Or it may contain more addresses.

[0075] As used herein, the term "affinity reagent" or "binding reagent" refers to a molecule or other substance that can specifically or reproducibly bind to an analyte (e.g., a protein). An affinity reagent can be larger than, smaller than, or the same size as the analyte. An affinity reagent can form a reversible or irreversible bond with the analyte. An affinity reagent can bind to the analyte in a covalent or non-covalent manner. An affinity reagent can include reactive affinity reagents, catalytic affinity reagents (e.g., kinases, proteases, etc.), or non-reactive affinity reagents (e.g., antibodies or fragments thereof). An affinity reagent can be non-reactive and non-catalytic, thereby not permanently altering the chemical structure of the analyte to which it binds. Affinity reagents that may be particularly useful for binding to proteins include, but are not limited to, antibodies or functional fragments thereof (e.g., Fab' fragments, F(ab')2 fragments, single chain variable fragments (scFv), di-scFv, tri-scFv, or microantibodies), affibodies, affilins, affimers, affitins, alphabodies, anticalins, avimers, DARPins, monobodies, nanoCLAMPs, nucleic acid aptamers, protein aptamers, lectins or functional fragments thereof.

[0076] As used herein, the term "array" refers to a collection of analytes (e.g., proteins) associated with unique identifiers such that the analytes can be distinguished from one another. The unique identifier can be, for example, a solid support (e.g., a particle or bead), a spatial address on a solid support, a tag, a label (e.g., a luminophore), or a barcode (e.g., a nucleic acid barcode) associated with the analyte and distinct from other identifiers in the array. The analytes can be associated with the unique identifiers by attachment, for example, via covalent or non-covalent bonds (e.g., ionic bonds, hydrogen bonds, van der Waals forces, electrostatics, etc.). An array can include different analytes each attached to a different unique identifier. An array can include different unique identifiers attached to the same or similar analytes. An array can include separate solid supports or separate addresses, each with a different analyte, and the different analytes can be identified according to their location on the solid support or address.

[0077] As used herein, the term "binding profile" refers to multiple binding results for a protein or other analyte. The binding results can be obtained from independent binding observations, e.g., independent binding results can each be obtained using a different affinity reagent. Alternatively, the results can be statistical measures, such as probability, likelihood, a measure of uncertainty, or a measure of variance. In some cases, the binding results can be generated in silico, e.g., derived from the modification of empirically obtained binding results. A binding profile can include empirical measurements, candidate measurements, putative measurements, calculated measurements, theoretical measurements, or combinations thereof. A binding profile can exclude one or more of the empirical measurements, candidate measurements, calculated measurements, or theoretical measurements, or putative measurements. A binding profile can include a vector of binding results. The elements of the vector can be digital values ​​(e.g., binary values ​​representing positive and negative binding results, respectively) or analog values ​​(e.g., probability values ​​ranging from 0 to 1).

[0078] As used herein, the term "comprising" is intended to be open-ended, including not only the recited elements, but further encompassing any additional elements.

[0079] As used herein, the term "each," when used in reference to a collection of items, is intended to identify each individual item in the set, but does not necessarily refer to every item in the set. Exceptions may occur where express disclosure or context clearly dictates otherwise.

[0080] As used herein, the term "epitope" refers to an affinity target within a protein, polypeptide, or other analyte. An epitope may include a sequence of amino acids that are contiguous in the primary structure of a protein. An epitope may include amino acids that are not contiguous in the primary sequence of a protein, but are structurally contiguous in the secondary, tertiary, or quaternary structure of the protein. An epitope may be or include a portion of a protein that arises by post-translational modification, such as phosphate, phosphotyrosine, phosphoserine, phosphothreonine, or phosphohistidine. An epitope may be capable of being recognized by an antibody or bound by an antibody. However, an epitope does not necessarily have to be recognized by any antibody, for example, it may instead be recognized by an aptamer, miniprotein, or other affinity reagent. An epitope may be capable of binding an antibody to elicit an immune response. However, an epitope does not necessarily have to be involved in eliciting an immune response, nor is it necessary that an epitope be capable of eliciting an immune response.

[0081] As used herein, the term "measurement" refers to information obtained from observation, simulation, or testing of a process. For example, a measurement on contacting an affinity reagent with an analyte can be referred to as a "binding outcome." A measurement can be positive or negative. For example, an observation of binding is a positive binding outcome, and an observation of non-binding is a negative binding outcome. If no positive or negative outcome is evident from a given measurement, the measurement can be a null outcome. An "empirical" measurement includes information based on the observation of a signal from an analytical technique. A "presumptive" measurement includes information based on theoretical or a priori evaluation of an analytical technique or analyte. A "candidate" measurement can include an empirical or presumptive measurement for a candidate analyte (e.g., for a candidate protein) known or suspected to be present in a sample or assay. Measurements can be expressed in binary terms, such as zero (0) for a negative binding outcome and one (1) for a positive binding outcome. In some cases, a ternary representation may be used, for example, where zero (0) represents a negative binding result, one (1) represents a positive binding result, and two (2) represents no result. It is also possible to use continuous or analog values, rather than integer or discrete values, to represent the different measurement results.

[0082] As used herein, the term "promiscuous" when used with respect to a reagent means that the reagent is known or suspected to react with a variety of different analytes in a given sample. For example, an affinity reagent that is known or suspected to recognize a variety of different analytes (e.g., a variety of proteins with different primary sequences) is promiscuous. A promiscuous reagent may be known or suspected to have high reactivity with one or more of the different analytes to which it reacts. For example, a promiscuous affinity reagent may have high affinity for one or more of the different analytes that it recognizes. A promiscuous reagent may be composed of a single type of reagent, such as a single affinity reagent, or a promiscuous reagent may be composed of two or more different types of reagents. For example, a promiscuous affinity reagent may be composed of a single type of antibody that recognizes a variety of different proteins in a sample, or a promiscuous affinity reagent may be composed of a pool containing several different antibody species that collectively recognize a variety of different proteins in a sample.

[0083] As used herein, the term "protein" refers to a molecule that comprises two or more amino acids linked by peptide bonds. A protein may also be referred to as a polypeptide, oligopeptide, or peptide. A protein may be a naturally occurring or synthetic molecule. A protein may comprise one or more non-natural amino acids, modified amino acids, or non-amino acid linkers. A protein may contain D-amino acid enantiomers, L-amino acid enantiomers, or both. The amino acids of a protein may be modified naturally or synthetically, for example, by post-translational modification. In some circumstances, different proteins may be distinguished from each other based on different genes from which the different proteins are expressed in an organism, different primary sequence lengths, or different primary sequence compositions. Nevertheless, proteins expressed from the same gene may be different proteoforms, distinguished, for example, based on non-identical length, non-identical amino acid sequence, or non-identical post-translational modification. Different proteins may be distinguished based on one or both of the gene of origin and the proteoform state.

[0084] As used herein, the term "single" when used in reference to an object such as an analyte means that the object is individually manipulated or distinguished from other objects. A single analyte can be a single molecule (e.g., a single protein), a single complex of two or more molecules (e.g., a multimeric protein having two or more separable subunits, a single protein attached to a structured nucleic acid particle or a single protein attached to an affinity reagent), a single particle, etc. Reference herein to a "single analyte" in the context of a composition, system, or method herein does not necessarily exclude application of the composition, system, or method to multiple single analytes that are individually manipulated or distinguished, unless contextually or explicitly indicated to the contrary.

[0085] As used herein, the term "single-analyte resolution" refers to the detection or ability to detect an analyte individually, for example as distinguished from its nearest neighbors in an array.

[0086] As used herein, the term "solid support" refers to a substrate that is insoluble in aqueous liquids. The substrate may be rigid. The substrate may be non-porous or porous. The substrate may be capable of taking up liquid (e.g., by porosity), but is typically, but not necessarily, sufficiently rigid so that the substrate does not substantially swell when taking up liquid and does not substantially shrink when the liquid is removed by drying. Non-porous solid supports are generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene with other materials, polypropylene, polyethylene, polybutylene, polyurethane, Teflon, cyclic olefins, polyimides, etc.), nylon, ceramics, resins, Zeonor™, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, fiber optic bundles, gels, and polymers. In certain configurations, the flow cell contains a solid support such that fluid introduced into the flow cell can interact with the surface of the solid support to which one or more components of a binding event (or other reaction) are attached.

[0087] The aspects described and claimed below can be understood in light of the above definitions.

[0088] The present disclosure provides a method for identifying an existing protein, the method comprising: (a) providing input to a computer processor, the input comprising: (i) a binding profile, the binding profile comprising a plurality of binding outcomes for binding of the existing protein to a plurality of different affinity reagents, each binding outcome of the plurality of binding outcomes comprising a measure of binding between the existing protein and a different affinity reagent of the plurality of different affinity reagents, the binding profile comprising positive binding outcomes and negative binding outcomes, (ii) a database comprising information characterizing or identifying a plurality of candidate proteins, and (iii) a binding model for each of the different affinity reagents, (b) determining a probability for each of the affinity reagents to bind each of the candidate proteins in the database according to the binding model, and (c) identifying the existing protein as a selected candidate protein, the selected candidate protein being the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the binding profile for the existing protein. The input may further include (iv) a non-specific binding rate, which includes the probability that a non-specific binding event occurs for one or more of the different affinity reagents.

[0089] 1. A method of identifying an existing protein, comprising the steps of: (a) contacting a plurality of different affinity reagents with a plurality of existing proteins in a sample; (b) obtaining binding data from step (a), said binding data comprising a plurality of binding profiles, each of said binding profiles comprising a plurality of binding results for binding of the existing protein of step (a) to a plurality of different affinity reagents, each binding result of said plurality of binding results comprising a measure of binding between the existing protein of step (a) and a different affinity reagent of said plurality of different affinity reagents, each of said binding profiles comprising a positive binding result and a negative binding result; and (c) obtaining a plurality of binding profiles comprising a plurality of binding results for binding of the existing protein of step (a) to a plurality of different affinity reagents. Also provided is a method comprising the steps of: (d) providing a database containing information characterizing or identifying a number of candidate proteins; (d) providing a binding model for each of the different affinity reagents; (e) determining, for each of the affinity reagents, a probability of binding to each of the candidate proteins in the database according to the binding model; and (f) identifying the existing protein as a selected candidate protein, wherein the selected candidate protein is the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the plurality of binding results for the existing protein.

[0090] The methods, compositions, and systems of the present disclosure are particularly well suited for use with proteins. Although proteins are exemplified throughout this disclosure, it will be understood that other analytes can be used as well. Exemplary analytes include, but are not limited to, biomolecules, polysaccharides, nucleic acids, lipids, metabolites, hormones, vitamins, enzyme cofactors, therapeutic agents, candidate therapeutic agents, or combinations thereof. Analytes can be non-biological atoms or molecules, such as synthetic polymers, metals, metal oxides, ceramics, semiconductors, minerals, or combinations thereof.

[0091] The one or more proteins used herein may be derived from natural or synthetic sources. Exemplary sources include, but are not limited to, biological tissues, fluids, cells or subcellular compartments (e.g., organelles). For example, a sample may be derived from a tissue biopsy, a biological fluid (e.g., blood, sweat, tears, plasma, extracellular fluid, urine, mucus, saliva, semen, vaginal fluid, synovial fluid, lymphatic fluid, cerebrospinal fluid, peritoneal fluid, pleural fluid, amniotic fluid, intracellular fluid, extracellular fluid, etc.), a fecal sample, a hair sample, cultured cells, culture medium, a fixed tissue sample (e.g., fresh frozen or formalin-fixed paraffin-embedded) or the product of a protein synthesis reaction. A protein source may include any sample in which a protein is a natural or expected constituent. For example, a primary source of a cancer biomarker protein may be a tumor biopsy sample or a bodily fluid. Other sources include environmental or forensic samples.

[0092] Exemplary organisms from which proteins or other analytes may be derived include, for example, mammals such as rodents, mice, rats, rabbits, guinea pigs, ungulates, horses, sheep, pigs, goats, cows, cats, dogs, primates, non-human primates, or humans; plants such as Arabidopsis thaliana, tobacco, corn, sorghum, oats, wheat, rice, rapeseed, and soybean; algae such as Chlamydomonas reinhardtii; nematodes such as Caenorhabditis elegans; insects such as Drosophila melanogaster, mosquitoes, fruit flies, honeybees, or spiders; fish such as zebrafish; reptiles; frogs and Xenopus Examples of such proteins include amphibians such as Dictyostelium discoideum, Pneumocystis carinii, Takifugu rubripes, yeasts, fungi such as Saccharomyces cerevisiae or Schizosaccharomyces pombe, or Plasmodium falciparum. Proteins may also be derived from bacteria, prokaryotes such as Escherichia coli, staphylococci or Mycoplasma pneumoniae, archae, viruses such as hepatitis C virus, influenza virus, coronavirus or human immunodeficiency virus, or viroids. The proteins may be derived from a homogenous culture or population of the above organisms, or from a collection of several different organisms, for example within a community or ecosystem.

[0093] In some cases, the protein or other biomolecule may be derived from an organism collected from a host organism. For example, the protein may be derived from a parasite, pathogen, symbiont, or cryptic organism collected from a host organism. The protein may be derived from an organism, tissue, cell, or biological fluid known or suspected to be associated with a disease state or disorder (e.g., cancer). Alternatively, the protein may be derived from an organism, tissue, cell, or biological fluid known or suspected not to be associated with a particular disease state or disorder. For example, proteins isolated from such sources can be used as a control to compare results obtained from a source known or suspected to be associated with a particular disease state or disorder. The sample may include a microbiota or a substantial portion of a microbiota. In some cases, one or more proteins used in the methods, compositions, or devices described herein may be obtained from a single source and at most from a single source. A single source can be, for example, a single organism (e.g., an individual human), a single tissue, a single cell, a single organelle (e.g., the endoplasmic reticulum, Golgi apparatus, or nucleus), or a single protein-containing particle (e.g., a virus particle or vesicle).

[0094] The disclosed methods, compositions, or devices may use or include a plurality of proteins having any of a variety of compositions, such as a plurality of proteins or a portion thereof that constitutes a proteome. For example, the plurality of proteins may include solution phase proteins, such as proteins or a portion thereof in a biological sample, or the plurality of proteins may include immobilized proteins, such as proteins attached to a particle or solid support. As a further example, the plurality of proteins may include proteins that are detected, analyzed, or identified in connection with the disclosed methods, compositions, or devices. The content of a plurality of proteins may be understood according to any of a variety of characteristics, such as those described below or elsewhere herein.

[0095] The plurality of proteins may be characterized in terms of total protein mass. The total mass of protein in 1 liter of plasma is estimated to be 70 g, and the total mass of protein in a human cell is estimated to be 100 pg-500 pg depending on the cell type. See Wisniewski et al., Molecular & Cellular Proteomics 13:10.1074 / mcp.M113.037309, 3497-3506 (2014), incorporated herein by reference. The plurality of proteins used or included in the methods, compositions or systems described herein may comprise at least 1 pg, 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 1 mg, 10 mg, 100 mg, 1 mg, 10 mg, 100 mg, or more of protein by mass. Alternatively or additionally, the plurality of proteins may contain at most 100 mg, 10 mg, 1 mg, 100 mg, 10 mg, 1 mg, 100 ng, 10 ng, 1 ng, 100 pg, 10 pg, 1 pg or less of protein by mass.

[0096] The plurality of proteins can be characterized in terms of mass percent relative to a given source, such as a biological source (e.g., cells, tissues, or biological fluids such as blood). For example, the plurality of proteins can comprise at least 60%, 75%, 90%, 95%, 99%, 99.9% or more of the total protein mass present in the source from which the plurality of proteins is derived. Alternatively or additionally, the plurality of proteins can comprise at most 99.9%, 99%, 95%, 90%, 75%, 60% or less of the total protein mass present in the source from which the plurality of proteins is derived.

[0097] A plurality of proteins can be characterized in terms of a total number of protein molecules. The total number of protein molecules in a Saccharomyces cerevisiae cell has been estimated to be about 42 million protein molecules. See Ho et al., Cell Systems (2018), DOI: 10.1016 / j.cels.2017.12.004, incorporated herein by reference. A plurality of proteins used or included in a method, composition or system described herein may be at least 1 protein molecule, 10 protein molecules, 100 protein molecules, 1×10 4 protein molecules, 1 x 10 6 protein molecules, 1 x 10 8 protein molecules, 1 x 10 10 of protein molecules, 1 mole (6.02214076 × 10 23 Alternatively or additionally, the plurality of proteins may comprise at most 100 moles of protein molecules, 10 moles of protein molecules, 1 mole of protein molecules, 1 x 10 10 protein molecules, 1 x 10 8 protein molecules, 1 x 10 6 protein molecules, 1 x 10 4 The nucleic acid may contain 1 protein molecule, 100 protein molecules, 10 protein molecules, 1 protein molecule or less.

[0098] The plurality of proteins can be characterized in terms of the diversity of full-length primary protein structures in the plurality. For example, the diversity of full-length primary protein structures in the plurality of proteins can be equated to the number of genes encoding different proteins in the source for the plurality of proteins. Whether the proteins are derived from a known genome or any genome, the diversity of full-length primary protein structures can be counted independently of the presence or absence of post-translational modifications in the proteins. The human proteome is estimated to have about 20,000 different protein-encoding genes, and a plurality of proteins from humans can include up to about 20,000 different primary protein structures. See Aebersold et al., Nat. Chem. Biol. 14:206-214 (2018), incorporated herein by reference. Other genomes and proteomes in nature are known to be larger or smaller. The plurality of proteins used or included in the methods, compositions or systems described herein may be at least 2, 5, 10, 100, 1×10 3 , 1×10 4 , 2×10 4 , 3×10 4 Alternatively or additionally, the plurality of proteins can have a complexity of at most 3×10 4 , 2×10 4 , 1×10 4 , 1×10 3 , 100, 10, 5, 2 or fewer different full-length primary protein structures.

[0099] In relative terms, the plurality of proteins used or included in a method, composition or system described herein may contain at least one representative for at least 60%, 75%, 90%, 95%, 99%, 99.9% or more of the proteins encoded by the genome of the source from which the sample is derived. Alternatively or additionally, the plurality of proteins may contain a representative for at most 99.9%, 99%, 95%, 90%, 75%, 60% or less of the proteins encoded by the genome of the source from which the sample is derived.

[0100] The plurality of proteins can be characterized in terms of the diversity of primary protein structures among the plurality, including transcribed splice variants. The human proteome is estimated to contain approximately 70,000 different primary protein structures, including splice variants. See Aebersold et al., Nat. Chem. Biol. 14:206-214 (2018), incorporated herein by reference. Additionally, due to fragmentation occurring in the sample, the number of partial length primary protein structures may increase. The plurality of proteins used or included in the methods, compositions or systems described herein may be at least 2, 5, 10, 100, 1×10 3 , 1×10 4 , 1×10 5 , 1×10 6 , 1×10 8 , 1×10 10 Alternatively or additionally, the plurality of proteins may have a complexity of at most 1×10 10 , 1×10 8 , 1×10 6 , 1×10 5 , 5×10 4 , 1×10 4 , 1×10 3 , 100, 10, 5, 2 or fewer different primary protein structures.

[0101] The plurality of proteins can be characterized in terms of the diversity of protein structure among the plurality, including different primary structures and different proteoforms among the primary structures. Different molecular forms of a protein expressed from a given gene are considered to be different proteoforms. Proteoforms can differ, for example, due to differences in primary structure (e.g., shorter or longer amino acid sequence), different arrangements of domains (e.g., transcriptional splice variants), or different post-translational modifications (e.g., the presence or absence of phosphoryl, glycosyl, acetyl, or ubiquitin moieties). The human proteome is estimated to contain hundreds of thousands of proteins, counting different primary structures and proteoforms. See Aebersold et al., Nat. Chem. Biol. 14:206-214 (2018), incorporated herein by reference. The plurality of proteins used or included in the methods, compositions, or systems described herein may be at least 2, 5, 10, 100, 1×10 3 , 1×10 4 , 1×10 5 , 1×10 6 , 5×10 6 , 1×10 7 Alternatively or additionally, the plurality of proteins may have a complexity of at most 1×10 7 , 5×10 6 , 1×10 6 , 1×10 5 , 1×10 4 , 1×10 3 , 100, 10, 5, 2 or fewer different protein structures.

[0102] The plurality of proteins can be characterized in terms of the dynamic range for different protein structures in a sample. The dynamic range can be a measure of the range of abundance for all different protein structures in the plurality of proteins, the range of abundance for all different primary protein structures in the plurality of proteins, the range of abundance for all different full-length primary protein structures in the plurality of proteins, the range of abundance for all different full-length gene products in the plurality of proteins, the range of abundance for all different proteoforms expressed from a given gene, or the range of abundance for any other set of different proteins described herein. The dynamic range for all proteins in human plasma is estimated to span more than 10 orders of magnitude, from the most abundant protein, albumin, to the rarest clinically measured protein. See Anderson and Anderson Mol Cell Proteomics 1:845-67 (2002), incorporated herein by reference. The dynamic range of the plurality of proteins described herein can be at least 10, 100, 1×10 3 , 1×10 4 , 1×10 6 , 1×10 8 , 1×10 10 Alternatively or additionally, the dynamic range of the proteins described herein may be at most 1×10 10 , 1×10 8 , 1×10 6 , 1×10 4 , 1×10 3 , 100, 10 or a multiple less than that.

[0103] The present disclosure provides assays useful for detecting one or more analytes. An exemplary assay format is shown diagrammatically in FIG. 1A. Proteins can be extracted from a sample and attached to an array. The unique identifier of the array can be an address. The array can be configured to have multiple addresses, with each address being attached to a respective individual protein from the sample. The proteins attached to the array can be in a denatured or native state. A structured nucleic acid particle (SNAP) may be capable of mediating the attachment of each protein to its respective address. Other linkers or attachment chemistries that can be used in addition to or instead of SNAP include, but are not limited to, those described in U.S. Patent Application Publication No. 2021 / 0101930, WO 2021 / 087402, or U.S. Patent Application Publication No. 63 / 159,500, each of which is incorporated herein by reference.

[0104] Typically, the identity of the protein at any given address is not known (thus the protein may be referred to as an "unknown" protein). The methods described herein can be used to identify proteins at one or more addresses in an array. Thus, the methods can be used to locate existing proteins within an array. Continuing with the example diagrammed in FIG. 1A, multiple affinity reagents (e.g., antibodies, aptamers, or small proteins) tagged with fluorophores can be contacted with the array and fluorescence can be detected from individual addresses to determine binding outcomes. Affinity reagents can be delivered to the array and detected sequentially, as shown, such that each cycle detects a binding outcome for an individual affinity reagent. In some configurations of the methods described herein, multiple different affinity reagents can be delivered in one cycle. The different affinity reagents delivered in a given cycle can be configured as a pool of indistinguishably labeled reagents (or the different affinity reagents can lack labels) such that the different reagents are not distinguished in the detection step. Alternatively, two or more different affinity reagents delivered in a given cycle can be distinguishably labeled. Thus, the affinity reagents can be distinguishably detected when bound to proteins on the array. The use of fluorescent labels and fluorescent detection is exemplary. Other labels and other detectors can be used, such as those described herein or known in the art.

[0105] Further examples of reagents and techniques that can be used to detect proteins in the methods, systems, or compositions of the present disclosure are described, for example, in U.S. Patent No. 10,473,654 or U.S. Patent Application Publication No. 2020 / 0318101 or U.S. Patent Application Publication No. 2020 / 0286584; or Egertson et al., BioRxiv (2021), DOI: 10.1101 / 2021.10.11.463967, each of which is incorporated herein by reference. Exemplary methods, systems, and compositions are described in further detail below.

[0106] Some configurations of the compositions, systems, or methods described herein can distinguish between different proteoforms, such as proteins that have the same primary structure (i.e., the same amino acid sequence) but differ with respect to the number, type, or location of post-translational modifications. The disclosed methods can be configured to identify the number, type, or location of one or more post-translational modifications in one or more proteins of a sample. Exemplary post-translational modifications include, but are not limited to, phosphoryl, glycosyl (e.g., N-acetylglucosamine or polysialic acid), ubiquitin, acyl (e.g., myristoyl or palmitoyl), isoprenyl, prenyl, farnesyl, geranylgeranyl, lipoyl, acetyl, alkyl (e.g., methyl or ethyl), flavin, heme, phosphopantetheinyl, C-terminal amidation, hydroxyl, nucleotidyl, adenylyl, uridylyl, proprionyl, S-glutathionyl, sulfate, succinyl, carbamyl, carbonyl, SUMOyl, or nitrosyl moieties.

[0107] Any of a variety of affinity reagents may be used in the compositions, systems, or methods described herein. An affinity reagent may be characterized, for example, for its binding properties prior to use in the methods described herein. Exemplary binding properties that may be characterized include specificity, strength of binding; equilibrium binding constants (e.g., K A Or K D ); binding rate constant, e.g., association rate constant (k on ) or dissociation rate constant (k off ); binding probability, etc. Binding properties can be determined for an epitope, a set of epitopes (e.g., a set of proteins with structural similarity), a protein, a set of proteins (e.g., a set of proteins with structural similarity), or a proteome.

[0108] The affinity reagent may include a label. Exemplary labels include, but are not limited to, fluorophores, luminophores, chromophores, nanoparticles (e.g., gold, silver, carbon nanotubes), heavy atoms, radioisotopes, mass labels, charge labels, spin labels, receptors, ligands, nucleic acid barcodes, polypeptide barcodes, polysaccharide barcodes, and the like. Labels can generate any of a variety of detectable signals, including, for example, optical signals such as absorbance of radiation, luminescence (e.g., fluorescence or phosphorescence) emission, luminescence lifetime, luminescence polarization, Rayleigh scattering and / or Mie scattering, magnetic properties, electrical properties, charge, mass, radioactivity, and the like. Label components may generate signals with characteristic frequencies, intensities, polarities, durations, wavelengths, sequences, or fingerprints. Labels need not generate signals directly. For example, labels can bind to receptors or ligands that have moieties that generate characteristic signals. Such labels may include, for example, nucleic acids encoded by specific nucleotide sequences, avidin, biotin, non-peptide ligands of known receptors, and the like.

[0109] The methods described herein can be carried out in fluid phase or on solid phase. In the case of fluid phase configuration, a fluid containing one or more proteins can be mixed with another fluid containing one or more affinity reagents. In the case of solid phase configuration, one or more proteins or affinity reagents can be attached to a solid support. One or more components involved in the binding event can be contained in a fluid, and the fluid can be delivered to a solid support, which is attached to one or more other components involved in the binding event.

[0110] The disclosed method may be performed with single analyte separation. A single analyte (e.g., a single protein) may be separated from other analytes, for example, based on spatial or temporal separation from other analytes. An alternative to single analyte separation is ensemble separation or bulk separation. Bulk separation configurations obtain a composite signal from multiple different analytes or affinity reagents in a container or on a surface. For example, a composite signal may be obtained from a population of different protein-affinity reagent complexes in a well or cuvette or on a solid support surface such that the individual complexes are not separated from each other. Ensemble separation configurations obtain a composite signal from a first collection of proteins or affinity reagents in a sample such that the composite signal is distinguishable from the signal generated by a second collection of proteins or affinity reagents in the sample. For example, the ensembles may be located at different addresses in an array. Thus, the composite signal obtained from each address is an average of the signals from the ensemble, but the signals from different addresses can be distinguished from each other.

[0111] The compositions, systems, or methods described herein may be configured to contact one or more proteins (e.g., an array of different proteins) with a plurality of different affinity reagents. For example, the plurality of affinity reagents (whether configured separately or as a pool) may include at least 2, 5, 10, 25, 50, 100, 250, 500, 1000, or more types of affinity reagents, each type of affinity reagent being different from the other types in terms of the epitope(s) it recognizes. Alternatively or additionally, the plurality of affinity reagents may include at most 1000, 500, 250, 100, 50, 25, 10, 5, or 2 types of affinity reagents, each type of affinity reagent being different from the other types in terms of the epitope(s) it recognizes. The different types of affinity reagents in the pool may be uniquely labeled so that the different types can be distinguished from each other. In some configurations, at least two, and up to all, of the different types of affinity reagents in the pool may be indistinguishably labeled. Alternatively or in addition to the use of unique labels, when assessing one or more proteins (e.g. in an array), different types of affinity reagents can be delivered and detected sequentially.

[0112] The disclosed methods can be performed on a single analyte (e.g., a single protein gene product) or in a multiplexed format. In a multiplexed format where the analyte is a protein, different proteins to be detected can be attached to different unique identifiers (e.g., addresses in an array), allowing for manipulation and detection of the proteins in parallel. For example, a fluid containing one or more different affinity reagents can be delivered to the array such that the proteins of the array are contacted with the affinity reagent(s) simultaneously. Furthermore, multiple addresses can be observed in parallel, allowing for rapid detection of binding events. The multiple different proteins can be at least 5, 10, 100, 1×10 3 , 1×10 4 , 2×10 4 , 3×10 4Alternatively or additionally, the proteome or proteome subfraction analyzed in the methods described herein may have a complexity of at most 3×10 4 , 2×10 4 , 1×10 4 , 1×10 3 The number of proteins detected, characterized, or identified in a sample may be of a complexity of 100, 10, 5, or less different native full length protein primary sequences. The plurality of proteins may constitute a proteome or a subfraction of a proteome. The total number of proteins detected, characterized, or identified in a sample may differ from the number of different primary sequences in the sample, for example, due to the presence of multiple copies of at least some protein species. Furthermore, the total number of proteins detected, characterized, or identified in a sample may differ from the number of candidate proteins presumed to be present in the sample, for example, due to the presence of multiple copies of at least some protein species, the absence of some proteins in the source of the sample, the presence of unexpected proteins in the source of the sample, or the loss of some proteins prior to analysis.

[0113] A particularly useful multiplex format uses an array of proteins and / or affinity reagents. Proteins can be attached to unique identifiers (e.g., addresses in an array) using any of a variety of means. Attachment can be covalent or non-covalent. Exemplary covalent attachments include chemical linkers, such as those achieved using click chemistry, or other bonds known in the art or described in U.S. Patent Application Publication No. 2021 / 0101930, which is incorporated herein by reference. Non-covalent attachment can be mediated, for example, by receptor-ligand interactions (e.g., (strept)avidin-biotin, antibody-antigen, or complementary nucleic acid strands) in which a receptor is attached to a unique identifier and a ligand is attached to a protein, or vice versa. In certain configurations, proteins are attached to a solid support (e.g., to addresses in an array) via structured nucleic acid particles (SNAPs). Proteins can be attached to SNAP, which can interact with a solid support, for example, by non-covalent interactions of DNA with the support and / or through covalent attachment of SNAP to the support. Nucleic acid origamis or nucleic acid nanoballs are particularly useful. The use of SNAP and other moieties to attach proteins to unique identifiers, such as tags or addresses in arrays, includes, but is not limited to, those described in U.S. Patent Application Publication No. 2021 / 0101930, WO 2021 / 087402, or U.S. Patent Application Publication No. 63 / 159,500, each of which is incorporated herein by reference.

[0114] The disclosed method can include assaying binding between a protein and an affinity reagent to determine a measurement result. For example, a measurement result on contacting an affinity reagent with an analyte can be observed as a binding result. The binding result can be positive or negative. For example, an observation of binding is a positive binding result and an observation of no binding is a negative binding result. For example, if a positive binding result cannot be distinguished from a negative binding result, the binding result can be a no binding result.

[0115] Binding can be detected using any of a variety of techniques appropriate for the reaction components used. For example, binding can be detected by obtaining a signal from a label attached to the affinity reagent if the affinity reagent is bound to the observed protein, by obtaining a signal from a label attached to the protein if the protein is bound to the observed affinity reagent, or by obtaining a signal(s) from a label attached to the affinity reagent and the protein if they are bound to each other. In some configurations, the protein-affinity reagent complex need not be detected directly, for example in a format in which a nucleic acid tag or other moiety is created or modified as a result of binding between the protein and the affinity reagent. Optical detection techniques such as luminescence intensity detection, luminescence lifetime detection, luminescence polarization detection, or surface plasmon resonance detection can be useful. Other detection techniques include, but are not limited to, electronic detection, such as techniques utilizing field effect transistors (FETs), ion-sensitive FETs, or chemically sensitive FETs. Exemplary methods are described in U.S. Pat. No. 10,473,654 or U.S. Patent Application No. 63 / 112,607 or 63 / 132,170, each of which is incorporated herein by reference.

[0116] The present disclosure provides a decoding method, for example in the form of a decoding algorithm, that can be used to evaluate the results of a binding reaction. The results can be used to identify or otherwise characterize proteins. In some configurations, clear and reproducible binding profiles can be observed for some or even the vast majority of the proteins to be identified in a sample. However, in many cases, one or more binding events can result in inconclusive or even anomalous results, which can then result in ambiguous binding profiles. For example, the observation of binding results in single molecule separations can be particularly prone to ambiguity due to chance in the behavior of single molecules when observed individually. The present disclosure provides a decoding method that provides accurate protein identification despite ambiguity and imperfections that can occur in single molecule formats or other situations.

[0117] In some configurations, a method for identifying or characterizing one or more extant proteins in a sample utilizes a decoding method that analyzes empirical binding profiles obtained for multiple binding reactions performed between each extant protein in the sample and multiple affinity reagents, and the empirical binding profiles are then evaluated for the binding behavior of the affinity reagents to multiple candidate proteins. The multiple candidate proteins can include proteins known or suspected to be present in the sample. Thus, the multiple candidate proteins can include multiple naturally occurring amino acid sequences. The decoding algorithm can output the identity of the extant protein as the candidate protein having binding characteristics that best match the empirical binding profile. This match can be determined based on a binding model that represents the affinity of each of the candidate proteins for each of the affinity reagents used to generate the empirical binding profile. A strong candidate protein can be identified as one whose modeled binding results are more consistent with the empirical binding profile compared to other candidate proteins that were evaluated.

[0118] The decoding method of the present disclosure may be configured to evaluate positive binding results. In a censored decoding configuration, the decoding method may evaluate positive binding results without evaluating negative binding results. In a non-censored decoding configuration, a strong candidate protein may be identified as one whose combination of positive and negative binding results is more consistent with the empirical binding profile compared to other candidate proteins evaluated. A candidate protein may be identified as weak or even inaccurate based on having many instances where positive and / or negative binding results are not consistent with the empirical binding profile being evaluated. The strongest candidate protein may be considered the most likely identity for the existing protein, and the confidence in this identification may be calculated as a relative measure of the suitability of the most likely protein compared to all of the other candidate proteins.

[0119] The computer processor may be configured to perform a decoding method that outputs an identity for one or more existing proteins based on various inputs. A particularly useful input is empirical binding data for the binding of an existing protein to a plurality of different affinity reagents. The binding data may be in the form of an empirical binding profile that includes a plurality of binding results. The empirical binding profile may include positive or negative binding results. The same may be true for the candidate result profile. In some configurations, the binding profile will include both positive and negative binding results. For example, the decoding may be performed in a "non-censored" configuration, where both positive and negative binding results are taken into account. Alternatively, the decoding may be performed in a "censored" configuration, where some or certain types of binding results are not taken into account. For example, the censored configuration may take into account positive binding results and omit negative binding results. The censored approach may be useful, for example, in situations where a particular binding measurement or binding result is expected to be prone to unacceptable or undesirable levels of error or artifacts.

[0120] Non-truncated decoding may be configured to equally utilize both positive and negative binding results when calculating the likelihood that a given existing protein has the identity of one or more candidate proteins. For example, the likelihood that each probe binds to each candidate protein may be known from empirical results and / or predicted from prior determination. The likelihood that each probe does not bind to each candidate protein may simply be determined as 1 minus the binding probability. The present disclosure provides a "semi-truncated" decoding configuration in which positive and negative binding results are evaluated independently of each other. Semi-truncated decoding may be configured to treat negative binding results as less informative than positive binding results. Instead of treating negative binding results as informative about the amino acid sequence of the existing protein, negative binding results may be treated as informative about the length of the existing protein that was not bound. In some configurations of the methods described herein, semi-truncated decoding is premised on the assumption that shorter proteins will have fewer positive binding results for a given set of affinity reagents compared to the number of positive binding results for longer proteins.

[0121] In the case of a semi-censored configuration, the negative binding probability can be calculated independently of the calculation of the positive binding probability. The semi-censored configuration offers the advantage of using a separate method for updating the protein likelihood from the negative binding results compared to the method used for the positive binding results. In the semi-censored configuration, the positive binding results can be weighted more heavily than the negative binding results. Alternatively, in the semi-censored configuration, the negative binding results can be weighted more heavily than the positive binding results. Different weights can be applied to offset expected or suspected biases in the binding reactions being evaluated, such as a high rate of off-target binding by one or more affinity reagents.

[0122] The empirical binding profile can be input to the decoding method described herein. For example, the empirical binding profile can be input to a computer processor that executes the decoding method. The series of empirical binding results that make up the empirical binding profile can be obtained using binding reactions, such as those described herein or known in the art. Alternatively, the binding profile can be obtained from a simulation and used similarly to the empirical binding profile. Each empirical binding result in the binding profile can result from one of a number of binding reactions performed between an existing protein and a number of affinity reagents. The empirical binding profile can be decoded after all binding results have been obtained for a given existing protein. Alternatively, for example, if the binding results are obtained sequentially, the decoding can be performed in real time, such that evaluation of empirical binding results from earlier binding reactions in the series begins and is possibly completed prior to or during the acquisition of empirical binding results for subsequent binding reactions in the series. The multiple empirical binding results do not necessarily need to be obtained sequentially, for example, instead, some or all of the binding results in the empirical binding profile are obtained from binding reactions occurring in parallel.

[0123] Another useful input to the decoding method is information about a plurality of candidate proteins. For example, information about a plurality of candidate proteins (e.g., a database of candidate protein information) can be input to a computer processor that executes the decoding method. The plurality of candidate proteins can be at least 10, 25, 50, 75, 100, 500, 1×10 3 , 1×10 4 , 1×10 6 , 1×10 8The database may include at least 10%, 25%, 50%, 75%, 90%, 95%, 99% or more of the proteins known or suspected to be present in the proteomes described herein or known in the art. The database may include candidate proteins from more than one organism. For example, the database may include organisms from a given ecosystem, such as a microbiome or environmental sample, organisms from a species of a particular family, class or genus, or all known proteins from all known species.

[0124] Information that can be included in the database of candidate proteins includes, but is not limited to, primary structure (i.e., amino acid sequence), secondary structure, tertiary structure, quaternary structure, name, or other information about the candidate protein. A text-based format for representing amino acid sequences may be used as the database in the methods or systems described herein. Information provided in FASTA format is particularly useful as the database. Information other than amino acid sequences may be included in the database. Particularly useful information that can be included in the database includes, for example, binding properties for the binding of one or more affinity reagents to the protein. However, such information need not be included in the database, but may instead be provided by a binding model. For example, the information may include the probability that each of a plurality of affinity reagents binds to each of a plurality of candidate proteins. In some configurations, such binding probabilities or other binding properties are derived empirically, for example, from binding experiments performed between one or more known candidate proteins and one or more known affinity reagents. In some embodiments, the binding probabilities or other binding properties are derived based on prior information, such as the presence of a predicted epitope sequence in the primary structure (e.g., amino acid sequence) of the candidate protein. Any of a variety of publicly available databases may be used, such as those described in Example I herein.

[0125] The database may include the probability or likelihood that the candidate protein will produce a positive binding result. Such information may be useful for several decoding configurations, including, for example, censored, uncensored, or semi-censored configurations. The database may further include the probability or likelihood that the candidate protein or pseudo protein will produce a negative binding result. Such information may be useful for uncensored or semi-censored decoding configurations.

[0126] The binding model can be input into the decoding method described herein. For example, the binding model can be input into a computer processor that executes the decoding method. The binding model may include a function for determining the probability of a specific binding event occurring between the protein and each of a plurality of affinity reagents. In some configurations, the binding model can include a function for determining the probability of a specific binding event occurring between a protein epitope and each of a plurality of affinity reagents. The epitopes evaluated by the model can have any of a variety of properties of interest. For example, the epitopes can have a defined length (e.g., the epitope length is no more than 2, 3, 4, 5, or 6 amino acids in the protein primary sequence) or chemical composition (e.g., the sequence of amino acids in the protein primary sequence). In some cases, the chemical composition can be relatively common in terms of chemical properties of amino acid side chains (or other moieties), such as charge, polarity, hydropathic index, steric size, steric shape, etc. For example, the chemical composition of an epitope can be expressed in terms of biosimilarity with another epitope.

[0127] The decoding method described herein can include a function for calculating the probability that each affinity reagent binds to some or all possible candidate proteins among a plurality of candidate proteins in a given database. The function can take into account positive binding outcomes. For example, if the function is used in an uncensored or semi-censored configuration, the function may be able to further take into account negative binding outcomes. The binding probabilities may be configured as a matrix. As demonstrated in Example I, the positive binding outcomes can be included in an M×N binding probability matrix B. In an uncensored configuration, the probability that a probe does not bind to a protein can be expressed as P(affinity probe does not bind|protein)=1-P(affinity probe binds|protein). When using a binding probability matrix, the non-binding probability matrix U can be calculated as U=1-B. However, the uncensored approach can be adversely affected by one or more non-binding events that have a very large impact on the decoding. For example, an affinity reagent may not bind to a particular site for many difficult to predict reasons (e.g., protein structure, the presence of unexpected post-translational modifications that prevent binding, etc.).

[0128] In some cases, the decoding may be overly biased towards short or long proteins. A normalization factor can be used to avoid overly biasing the decoding results towards short or long proteins, thereby shifting the possible identifications to overcome sequence length bias. In some cases, the binding probability can be normalized to the protein length by dividing it by a normalization constant. Another approach is to use a blinded non-censored approach, where the non-censored decoding is adapted to be more resilient to missing binding events. This can be done by adjusting the probability of negative binding results. For example, the probability of not binding a trimer of unknown identity can be calculated for each affinity reagent.

number

[0129] The above approach can be used to normalize proteins by length without considering the specific trimer composition of each protein. The above approach can be easily adjusted for epitopes with other lengths. In another configuration, we use multiple different proteins as training points to normalize θ for θ. Nj = (1-P(Probe binds protein j )) (NB "probe" means "affinity reagent" in this context), blind, uncensored decoding can be calculated as above for trimers. For example, 20,000 proteins can be used as training points, where j = 1...20,000. The above analysis can be modified for use with epitopes of sizes other than trimers, including, for example, dimers, tetramers, pentamers, etc.

[0130] For length normalization, a binomial approximation can be used. This approximation counts the total number of possible specific binding events S and the total number of possible non-specific binding events NS; the average binding probability among the possible specific binding events:

number

number

number

[0131] The length normalization can be done using the Poisson binomial (e.g., the exact or estimated Poisson binomial). The normalization can be done as follows: The joint probability p={p1, p1, p1...p 300 For a protein with {\displaystyle \mathbf{p}}, the pmf of the Poisson binomial distribution with p as a parameter is used to calculate the probability of observing N binding events, and for each candidate protein, the likelihood of an observed binding event is multiplied by PoiBin(p).pmf(N). The Poisson binomial pmf can be calculated using "exact" computational methods or an exact normal approximation (normal distribution + skew) (see Hong et al., Computational Statistics & Data Analysis 59:41-51 (2013), incorporated herein by reference).

[0132] Length normalization can also be performed via a semi-censored approach as described herein. The semi-censored configuration can allow for the total number of non-binding events to be considered more than the specific identity of the observed non-binding events. Example I demonstrates a semi-censored configuration in which the non-binding probability is adjusted to consider salient features of the candidate protein, such as the length of the candidate protein and the relative frequency of all possible unique epitopes of a particular amino acid length (e.g., dimers, trimers, tetramers, etc.). A vector of average non-binding probabilities for affinity reagents can be calculated. For example, the probability that a given affinity reagent does not bind to a trimer epitope, averaged over all 8000 trimers and weighted by the relative frequency of each trimer in the candidate protein database, can be calculated.

[0133] Another approach that can be used to avoid overly biased decoding results towards short or long proteins is to configure the semi-truncated decoding method to predict the probability of a negative binding outcome based on the length of the proteins suspected to be present in the sample, but independent of the amino acid sequence of these proteins. The prediction may also be made independent of knowledge of the epitope for the affinity reagent used to assay the sample. For example, the probability of a negative binding outcome can be predicted independently of the sequence length of the epitope. Thus, the decoding can be based on an algorithm that is equally applicable to the use of dimeric, trimeric, tetrameric or other length epitopes. As described in more detail below, a set of pseudo proteins can be generated and this set can be used to predict the negative binding probability.

[0134] The semi-truncated decoding method may be configured to use a plurality of candidate proteins that include amino acid sequences known or suspected to be present in a given sample. For example, a decoding method configured to evaluate proteins of human origin may utilize a plurality of candidate proteins that include amino acid sequences unique to humans. The semi-truncated decoding method may further be configured to use a set of pseudo proteins that may differ from the set of candidate proteins. A plurality of candidate proteins with unique sequences may be useful for determining the probability for a positive binding outcome between an affinity reagent and a candidate protein. A plurality of pseudo proteins may be useful for determining the probability for a negative binding outcome between an affinity reagent and a candidate protein.

[0135] In some configurations, the set of pseudo proteins can include full-length amino acid sequences that are known or suspected not to be present in a given sample. For example, none of the full-length amino acid sequences in the set of pseudo proteins need be present in the set of candidate proteins, or vice versa. Alternatively, a single full-length amino acid sequence or a subset of amino acid sequences can be present in both the set of pseudo proteins and the set of candidate proteins. In some configurations, a partial amino acid sequence can be present in both the set of pseudo proteins and the set of candidate proteins. A partial sequence present in both sets can contain at most 50, 40, 30, 20, 10, 9, 8, 7, 6, 5, 4, or 3 consecutive amino acids. Alternatively or additionally, a partial sequence present in both sets can contain at least 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, or 50 consecutive amino acids. In yet other configurations, the same amino acid sequence, whether full-length or partial, can be present in both the set of pseudo proteins and the set of candidate proteins.

[0136] Turning to the example of a decoding method configured to evaluate proteins from a particular organism, a set of pseudo proteins can be utilized that includes amino acid sequences that do not naturally occur in the organism. For example, the set of pseudo proteins can include amino acid sequences that are uniquely present in one or more organisms other than the organism being evaluated. The plurality of candidate proteins can lack full-length amino acid sequences that are not naturally present in a given sample (e.g., not naturally present in a particular organism), and the plurality of pseudo proteins can lack amino acid sequences that are uniquely present in a given sample (e.g., uniquely present in a particular organism).

[0137] When performing a semi-truncated decoding method, the number of pseudo proteins can be substantially the same as the number of candidate proteins. For example, the multiple candidate proteins can include native sequences for proteins known or suspected to be present in a given sample, and the multiple pseudo proteins can include amino acid sequences related to each of the native sequences in the multiple candidate proteins. The pseudo amino acid sequences can be associated with each unique sequence by each pseudo amino acid sequence having the same total length as the total length for the unique amino acid sequence in the candidate protein. However, each pseudo sequence may differ from its associated unique sequence in terms of the amino acid content of the sequence.

[0138] In an alternative configuration, the number of pseudo-proteins utilized in the semi-truncated decoding method can be greater than the number of candidate proteins utilized. For example, the plurality of candidate proteins can include unique sequences for proteins known or suspected to be present in a given sample, and the plurality of pseudo-proteins can include a plurality of pseudo-sequences associated with each of the unique sequences. Each individual unique sequence in the plurality of candidate proteins can be associated with at least 2, 3, 4, 5, 10, 25 or more pseudo-sequences in the plurality of pseudo-proteins. Similarly, the pseudo-sequences can be associated with each unique sequence in terms of the length of the two sequences. However, each pseudo-sequence can differ from its associated unique sequence in terms of amino acid content.

[0139] The set of pseudo proteins can be generated using any of a variety of methods. For example, the pseudo amino acid sequences can be selected randomly. As a more specific example, a pseudo sequence can be generated for each unique sequence by shuffling the order of amino acids in the unique sequence. Another option is to generate a pseudo sequence for each unique sequence by randomly assigning one of the 20 unique amino acids to each position along the length of the unique sequence.

[0140] A set of pseudo sequences may be generated to bias or weight the pseudo amino acid sequences to reflect the characteristics of multiple unique amino acid sequences present in a proteome or other sample to be evaluated using the decoding methods described herein. For example, a binning approach can be used in which all candidate proteins for a given sample (e.g., all proteins in a proteome) are aggregated into bins according to their amino acid sequence length. Within each bin, an uncensored unbound likelihood can be predicted for each protein, and the median value can be used as the semi-censored unbound likelihood for the entire bin. Thus, the proteins in a bin are representative of the sequence bias in the sample.

[0141] Another approach that can be used is to create a set of pseudosequences that are representative of sequence bias in the proteome (or other sample) of interest, and predict the non-binding probability for the pseudosequences. For example, a Markov model can be used. A Markov model is a statistical technique that can be used to model a sequence such that the probability of a sequence element is based on the limited context that precedes the element. A Markov model can be used to factorize the probability of observing an amino acid sequence with respect to the context-dependent probability of the amino acids in the sequence. A set of pseudosequences can be generated by Markov chain Monte Carlo sampling of amino acid sequences in a plurality of unique sequences, as described in Example II below.

[0142] The Markov chain can be adjusted to suit a particular assay condition or sample. For example, the transition probabilities can be modified to account for the over- or under-representation of one or more proteins in a sample. This approach can be useful, for example, when a sample has been experimentally enriched for one or more protein sequences. Thus, a protein sample can be fractionated, for example, via immunoprecipitation, chromatography, or other known separation techniques, and the assay results for the fractionated sample can be decoded using a set of pseudo-proteins derived from the use of appropriately modified transition probabilities in the Markov chain algorithm. Similarly, modified transition probabilities can be used to account for changes in the proteome due to over- or under-expression of one or more proteins, such as may result from a particular disease (e.g., cancer) or from genetic engineering.

[0143] Another algorithm that can be used is a generative adversarial network (GAN). For example, a GAN can generate a set of pseudo proteins from a set of candidate proteins such that the set of pseudo proteins has similar amino acid sequence characteristics as the set of candidate proteins. In some cases, a GAN can generate a set of pseudo proteins from a set of proteins other than the set of candidate proteins used for the decoding method. For example, a GAN can generate a set of pseudo proteins based on a subset of amino acid sequences in the set of candidate proteins used for decoding, based on a larger set of amino acid sequences that includes some or all of the sequences in the set of candidate proteins used for decoding, or based on a set of amino acid sequences from an organism other than the organism for the candidate proteins used for decoding. An expectation-maximization algorithm can also be used to generate a set of pseudo proteins.

[0144] The plurality of pseudo-proteins can have a total amino acid composition substantially equivalent to that of the plurality of candidate proteins. In another example, the plurality of pseudo-proteins can have a total composition of amino acid k-mers (e.g., dimers, trimers, tetramers, pentamers, etc.) substantially equivalent to the total composition of amino acid k-mers in the plurality of candidate proteins. The plurality of pseudo-proteins can have a sequence bias substantially equivalent to that in the plurality of candidate proteins. For example, the dependency of a particular k-mer on its sequence context can be the same in the plurality of pseudo-proteins as in the plurality of candidate proteins. In this example, the sequence context can refer to a type of single amino acid that is upstream or downstream of the k-mer. In some cases, the sequence context can refer to a subsequence of two or more amino acids that are upstream or downstream of the k-mer.

[0145] Thus, a method of identifying an existing protein includes the steps of: (a) providing input to a computer processor, the input including: (i) a binding profile, the binding profile including a plurality of binding results for binding of the existing protein to a plurality of different affinity reagents, each binding result of the plurality of binding results including a measure of binding between the existing protein and a different affinity reagent of the plurality of different affinity reagents, the binding profile including positive binding results and negative binding results; (ii) a database including information characterizing or identifying a plurality of candidate proteins; and (iii) a binding model for each of the different affinity reagents. (b) determining a probability for each of the affinity reagents binding to a candidate protein in the database according to the binding model, said determining comprising calculating probabilities for the positive binding outcomes and for the negative binding outcomes, the positive binding outcomes being weighted more heavily than the negative binding outcomes, and (c) identifying the existing protein as a selected candidate protein, the selected candidate protein being the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the binding profile for the existing protein. Step (b) may comprise (i) calculating the probability of a positive binding outcome occurring between each of the candidate proteins and each of the affinity reagents, and (ii) calculating the probability of a negative binding outcome occurring between each pseudo protein in a plurality of pseudo proteins and each of the affinity reagents.

[0146] In any configuration of the above method, the amino acid sequences in the plurality of pseudo-proteins have a total length that is identical to the total length of the amino acid sequences in the plurality of candidate proteins. As a further option, the plurality of pseudo-proteins can lack some or all of the full-length amino acid sequences present in the plurality of candidate proteins. Furthermore, the amino acid sequences in the plurality of pseudo-proteins may be generated by sampling the amino acid sequences in the plurality of candidate proteins using Markov chains, generative adversarial networks, or length-based binning.

[0147] The plurality of candidate proteins used in the methods described herein can include amino acid sequences that are naturally present in the sample from which the existing protein of interest is derived, whereas the plurality of pseudo proteins can include amino acid sequences that are not naturally present in the sample. Each individual pseudo protein of the plurality of pseudo proteins may be capable of having the same overall length as the overall length of one candidate protein in the plurality of candidate proteins.

[0148] The decoding method described herein may include a function for determining the probability of a non-specific binding event occurring between a protein and a plurality of affinity reagents. The model may take into account the context of one or more epitopes in a given candidate protein. For example, the function for determining the probability may be normalized to the length of a given candidate protein. Alternatively or additionally, the binding model used in the method or system described herein may include a function for determining the probability of a specific binding event occurring between a candidate protein and each of the affinity reagents. Also, the model may take into account the context of one or more epitopes in a given candidate protein. For example, the function may be normalized to the length of a given candidate protein.

[0149] In some configurations, the decoding method can include a function for determining the probability that a binding event occurs between each of the affinity reagents and the specific epitope for each affinity reagent and the epitope that is biosimilar. In a biosimilar model, an affinity reagent can be considered to target a particular epitope to which the affinity reagent binds with a certain probability. For example, the probability can be at least 0.01, 0.05, 0.1, 0.25 0.5, 0.75, 0.9, 0.99 or more. Alternatively or additionally, the probability can be at most 0.99, 0.9, 0.75, 0.5, 0.25, 0.1, 0.05, 0.01 or less. An affinity reagent can also be considered to bind one or more additional primary off targets with a probability within the above ranges. The number of additional primary targets can be at least 1, 3, 5, 7, 9, 15, 20 or more epitopes that are biosimilar to the targeted epitope. Alternatively or in addition, the number of additional primary targets can be at most 20, 15, 9, 7, 5, 3 or 1 epitopes that are biosimilar to the targeted epitope. Biosimilar epitope targets can be selected by calculating the pairwise similarity scores of the target epitope to all other possible epitopes of the same length, and then selecting one or more of the other epitopes with high similarity scores. The similarity scores can be calculated by summing the similarities between pairs of residues at each sequence position, for example, using BLOSUM62 or other functions for determining biological similarity.

[0150] In the decoding method of the present disclosure, a parameterized binding model can be used. For example, an affinity reagent can be modeled by assigning a binding probability to each unique target epitope recognized by the affinity reagent. A non-specific binding rate may be assigned to each affinity reagent. The non-specific binding rate may represent, for example, the probability that a given affinity reagent binds non-specifically to any epitope in a protein. The probability of an affinity reagent binding to a given candidate protein can be calculated by first calculating the probability that a specific binding event occurs. The model may take into account the number of each epitope in a given protein sequence. The binding model parameters may include a vector of the probability that a given affinity reagent binds to each recognized epitope. In addition, the model may include a function for calculating the probability that a non-specific protein binding event occurs. The model may take into account the length of each candidate protein sequence, the length of the epitope recognized by the affinity reagent, or both. The probability that an affinity reagent binds to a protein and generates a detectable signal may be expressed as the probability that one or more specific or non-specific binding events occur. An exemplary binding model is provided in Example I herein.

[0151] In some configurations of the systems or methods described herein, a non-specific binding rate can be provided as an input. The input can be in the form of one fixed non-specific binding rate for all affinity reagents, or a unique non-specific binding rate for each affinity reagent. Also, the non-specific binding rate can be iteratively and / or adaptively learned, just like other parameters in the affinity reagent binding model. The non-specific binding event can be the binding of an affinity reagent to a substance other than a protein. The substance can be a solid support attached to an existing protein. For example, the non-specific binding event can occur in a region of the array where the protein of interest is not present, such as an address or a location near where the protein of interest is present. In some cases, the non-specific binding event can occur at an empty address where no protein is present, or in a gap region on the array that separates one address from another address. Optionally, as illustrated in Example I herein, the input can be a surface non-specific binding rate that describes the probability of a surface non-specific binding event occurring in any given cycle during a series of binding reactions.

[0152] Execution of the decoding algorithm can include calculating a probability matrix including the probability of a positive binding outcome for each individual affinity reagent that binds to each candidate protein used in the binding reaction. Optionally, the method can further include calculating a probability matrix including the probability of a negative binding outcome for each individual affinity reagent that binds to each candidate protein used in the binding reaction. For example, the adjusted non-binding probability can be calculated as described in Example I or Example II herein. In an alternative configuration of the systems and methods described herein, the probability of a negative binding outcome can be calculated by subtracting the probability of a positive binding outcome from 1, where the probability is represented by a value between 0 and 1. Positive and negative binding outcomes can be weighted equally. Alternatively, positive binding outcomes can be weighted more heavily relative to negative binding outcomes. In other cases, negative binding outcomes can be weighted more heavily relative to positive binding outcomes. The latter weighting can be particularly desirable to account for the numerous and difficult to predict mechanisms by which affinity reagents may non-specifically bind to proteins.

[0153] Decoding can be performed by calculating a vector of likelihoods for multiple candidate proteins. The most likely candidate protein can be selected. For example, the selected candidate protein can be the one with the highest probability of binding an affinity reagent that matches most of the binding results obtained for a given existing protein. In another example, the candidate protein can be selected by multiplying the probability of the observed binding results. Optionally, if there are matches for the top proteins, one of the top proteins can be selected randomly or by another desired criterion. The probability of the correct identification can be based on the likelihood that the top protein is correct divided by the sum of the likelihoods that all other candidate proteins are correct. The identity of the protein can be output from the decoding system or method. Optionally, the probability of the correct identification can be output. The probability can be calculated as the quotient of the likelihood of the selected candidate protein divided by the sum of the likelihoods determined for all other candidate proteins evaluated by the decoding algorithm.

[0154] Exemplary algorithms and methods for characterizing proteins that may be used in combination with the methods or systems described herein include, for example, those described in U.S. Patent Application Publication No. 2020 / 0286584 or Egertson et al., BioRxiv (2021), DOI: 10.1101 / 2021.10.11.463967, each of which is incorporated by reference herein.

[0155] The decoding method may output information regarding the identity of one or more existing proteins. The information output for a given protein may be in the form of a determined identity for the protein or in the form of a probability or likelihood for one or more identities of the protein. For example, the most likely identity for an existing protein, the likelihood or probability of an existing protein having a particular identity, or both may be output by the decoding method. The decoding method may output a non-digital or non-binary score for a given existing protein's identity or for the likelihood of an existing protein having a particular identity. For example, the probability or likelihood score may be output in the form of an analog value between 0 and 1, or a percentage value between 0% and 100%. In some configurations, a digital or binary score indicating one of two discrete states may be output to indicate the identity of the protein or at least a subset of proteins to which the protein belongs (e.g., a family of proteins sharing a common structural motif).

[0156] One or more steps of the methods described herein may be performed in a detection system. Thus, the detection system may be configured to perform one or more steps of the methods described herein. For example, the detection system may be configured to perform one or more steps of the decoding methods described herein. The decoding methods described herein may be configured to improve the accuracy of the detection system. For example, the detection system may provide an initial identity or signature for one or more existing proteins, and may use the decoding methods described herein to output a subsequent identity or signature that is more accurate or otherwise improved compared to the initial identity or signature.

[0157] The disclosure provides a method for detecting a protein in a sample, the method comprising: (a) a detector configured to acquire signals from a plurality of binding reactions occurring between a plurality of different affinity reagents and a plurality of existing proteins in a sample; (b) a database containing information characterizing or identifying a plurality of candidate proteins; and (c) a computer processor (i) in communication with the database and (ii) processing the signals to generate a plurality of binding profiles, each of the binding profiles comprising a plurality of binding results for binding of the existing protein of (a) to the plurality of different affinity reagents, each of the binding results being a binding result between the existing protein of (a) and a different affinity reagent of the plurality of different affinity reagents. and a computer processor configured to: (i) generate a binding profile comprising a measure of binding between the affinity reagent and each of the existing proteins, the binding profiles including a measure of binding between the affinity reagent and each of the existing proteins, the binding profiles including a positive binding result and a negative binding result; (ii) process the binding profiles to determine, for each of the affinity reagents, a probability of binding to each of the candidate proteins in the database according to a binding model for each of the affinity reagents; and (iii) output an identification of a selected candidate protein, the selected candidate protein being the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the plurality of binding results for the existing protein.

[0158] A method for identifying existing proteins may be carried out in a detection system, the method comprising: (a) acquiring signals from a plurality of binding reactions carried out in the detection system, the binding reactions comprising contacting a plurality of different affinity reagents with a plurality of existing proteins in a sample; (b) processing the signals in the detection system to generate a plurality of binding profiles, each of the binding profiles comprising a plurality of binding results for binding of the existing protein of step (a) to the plurality of different affinity reagents, each binding result of the plurality of binding results comprising a measure of binding between the existing protein of step (a) and a different affinity reagent of the plurality of different affinity reagents, each of the binding profiles comprising a positive binding result and a negative binding result; and (c) contacting the existing protein of step (a) with a plurality of different affinity reagents. The method may include providing as an input a database containing information characterizing or identifying a plurality of candidate proteins; (d) providing as input to the detection system a binding model for each of the different affinity reagents; (e) processing the plurality of binding profiles in the detection system to determine, according to the binding model, a probability for each of the affinity reagents to bind to each of the candidate proteins in the database; and (f) outputting from the detection system an identification of a selected candidate protein, the selected candidate protein being the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the plurality of binding results for the existing protein.

[0159] The detection system can include a detector, such as a detector known in the art for detecting a label or analyte described herein. The detector can be configured to collect a signal (e.g., an optical signal) from an array or other vessel containing present proteins or other analytes. A camera, such as a complementary metal oxide semiconductor (CMOS) or charge-coupled device (CCD) camera, can be particularly useful for detecting optical labels, such as luminophores, for example. The detection system can further include an excitation source configured to excite present proteins, affinity reagents, or other analytes, for example, in the array or other vessel. The detection system can include a scanning mechanism configured to effect relative movement between the detector and the array or other vessel containing present proteins. In some cases, the scanning mechanism can be configured for time delay integration. For example, a detector capable of separating proteins on the array surface, including single molecule separation, can be particularly useful. Detectors used in DNA sequencing systems can be modified for use in the detection systems or other devices described herein. Exemplary detectors are described, for example, in U.S. Pat. Nos. 7,057,026; 7,329,492; 7,211,414; 7,315,019 or 7,405,281, or U.S. Patent Application Publication No. 2008 / 0108082, each of which is incorporated herein by reference.

[0160] The detection system may further include a fluidics device configured to contact reaction components for the reaction or other steps of the methods described herein. In certain embodiments, the reaction occurs on an array. Any of a variety of arrays, such as those described herein, may be present in the system. The proteins to be detected, e.g., proteins attached to an array, may be contained in any of a variety of reaction vessels. A particularly useful reaction vessel is a flow cell. The flow cell or other vessel may be present in the system in a permanent or removable manner, e.g., removable by hand or without the use of auxiliary tools. The flow cell or other vessel may have a detection window through which a detector observes one or more proteins (e.g., an array of proteins) or other analytes on the array. For example, an optically transparent window may be used in combination with an optical detector, such as a fluorometer or luminescence detector.

[0161] The fluidic device may include one or more reservoirs fluidly connected to the inlet of a flow cell or other container. The reservoir may include reagents for use in the methods described herein. The system may further include a pump, pressure supply, or other fluid movement device for delivering reagents from the reservoir to the container. The system may include a waste reservoir fluidly connected to the outlet of the container for removing used reagents. In the example embodiment where the container is a flow cell, the reagents may be delivered to the flow cell through the flow cell inlet, and then the reagents may flow through the flow cell and out of the flow cell outlet to the waste reservoir. Thus, the flow cell may be in fluid communication with one or more reservoirs of the system. The fluidic system may include at least one manifold and / or at least one valve for directing the reagents from the reservoir to the container where detection takes place. Exemplary fluidic devices that may be used in the systems of the present disclosure include those configured for cyclic delivery of reagents, such as those deployed in nucleic acid sequencing reactions. Exemplary fluidic devices are described in U.S. Patent Application Publication Nos. 2009 / 0026082; 2009 / 0127589; 2010 / 0111768; 2010 / 0137143; or 2010 / 0282617; or U.S. Patent Nos. 7,329,860; 8,951,781; or 9,193,996, each of which is incorporated herein by reference.

[0162] The present disclosure provides a computer system (e.g., a computer control system) programmed to perform the methods, algorithms, or functions described herein. In some cases, the computer system described herein can be a component of a detection system. The computer system can be programmed or otherwise configured to (a) receive inputs described herein, such as a binding profile, a database including information characterizing or identifying a plurality of candidate proteins, a binding model, and / or a non-specific binding rate of an affinity reagent, (b) determine a probability that the affinity reagent binds to the candidate protein, e.g., based on the binding model, and (c) identify an existing protein as the selected candidate protein.

[0163] FIG. 12 illustrates an exemplary computer system 1001. The computer system 1001 can be an electronic device of a detection system, the electronic device being integral with the detection system or located remotely with respect to the detection system. For example, the electronic device can be a portable electronic device. The computer system 1001 includes a computer processing unit (CPU, herein "processor" and "computer processor") 1005, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 1001 also includes memory or memory locations 1010 (e.g., random access memory, read-only memory, flash memory), electronic storage 1015 (e.g., hard disk), communication interface 1020 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 1025, such as cache, other memory, data storage, and / or electronic display adapters. The memory 1010, storage 1015, interface 1020, and peripheral devices 1025 communicate with the CPU 1005 via a communication bus (solid lines), such as a motherboard. The storage device 1015 may be a data storage device (or data repository) for storing data. The computer system 1001 may be operatively coupled to a computer network ("network") 1030 with the aid of a communication interface 1020. The network 1030 may be the Internet, an Internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. The network 1030 is, in some cases, a telecommunications and / or data network. The network 1030 may include one or more computer servers, which may enable distributed computing, such as cloud computing.For example, one or more computer servers may enable cloud computing ("cloud") on the network 1030 to perform various aspects of the analysis, calculation, and generation of the present disclosure, such as, for example, receiving information of empirical measurements of existing proteins in a sample; processing the information of the empirical measurements against a database including a plurality of protein sequences corresponding to candidate proteins, for example, using a binding model or function described herein; generating a probability of the candidate protein generating the empirical measurement, and / or generating a probability that the existing protein is correctly identified in the sample. Such cloud computing may be provided, for example, by a cloud computing platform such as Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM Cloud. The network 1030 may, in some cases, implement a peer-to-peer network, which may enable devices coupled to the computer system 1001 to act as clients or servers with the aid of the computer system 1001.

[0164] The CPU 1005 can execute a series of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 1010. The instructions may be directed to the CPU 1005, which may then be programmed or otherwise configured to implement the methods of the present disclosure. Examples of operations performed by the CPU 1005 may include fetch, decode, execute, and writeback.

[0165] The CPU 1005 may be part of a circuit, such as an integrated circuit. One or more other components of the system 1001 may be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).

[0166] The storage device 1015 can store files such as drivers, libraries, and stored programs. The storage device 1015 can store user data, such as user preferences and user programs. The computer system 1001 can include one or more additional data storage devices that are external to the computer system 1001, such as located on a remote server that communicates with the computer system 1001 via an intranet or the Internet in some cases.

[0167] The computer system 1001 can communicate with one or more remote computer systems via the network 1030. For example, the computer system 1001 can communicate with a remote computer system of a user. Examples of remote computer systems include a personal computer (e.g., a portable PC), a slate or tablet PC (e.g., an Apple® iPad, a Samsung® Galaxy Tab), a phone, a smartphone (e.g., an Apple® iPhone, an Android-enabled device, a Blackberry®), or a personal digital assistant (Apple, iPad, Samsung, Galaxy, iPhone, Android, and Blackberry are registered trademarks). A user can access the computer system 1001 via the network 1030.

[0168] The methods described herein may be implemented by machine (e.g., a computer processor) executable code stored in an electronic storage location of the computer system 1001, such as the memory 1010 or the electronic storage 1015. The machine executable or machine readable code may be provided in the form of software. In use, the code may be executed by the processor 1005. In some cases, the code may be read from the storage 1015 and stored in the memory 1010 for easy access by the processor 1005. In some situations, the electronic storage 1015 may be omitted and the machine executable instructions are stored in the memory 1010.

[0169] The code may be pre-compiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled at run-time. The code may be provided in a programming language that can be selected to allow the code to be executed in a pre-compiled or as-compiled fashion.

[0170] Aspects of the systems and methods provided herein, such as the computer system 1001, may be embodied in programming. Various aspects of the technology may be considered as a "product" or "article of manufacture," typically in the form of machine (or processor) executable code and / or associated data carried on or embodied in some type of machine-readable medium. The machine-executable code may be stored in electronic storage, such as memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium may include any tangible memory of a computer, processor, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives, etc., that may provide non-transitory storage for software programming at any time. All or part of the software may at times be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, for example, from a management server or host computer to a computer platform of an application server. Thus, other types of media that may carry software elements include light waves, radio waves, and electromagnetic waves over wired and optical terrestrial communications networks and through various air links, such as those used across physical interfaces between local devices. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media that carry software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0171] Thus, a machine-readable medium such as a computer executable code may take many forms, including but not limited to tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include, for example, optical or magnetic disks, any storage device in any computer, such as may be used to implement the databases shown in the figures, etc. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire and optical fibers, including the wires that make up a bus in a computer system. Carrier wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Thus, common forms of computer readable media include, for example, a floppy disk, a flexible disk, a hard disk, a magnetic tape, any other magnetic medium, a CD-ROM, a DVD or a DVD-ROM, any other optical medium, punch cards, punched tape, any other physical storage medium having a pattern of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave carrying data or instructions, a cable or link carrying such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.

[0172] The computer system 1001 can include or communicate with an electronic display 1035 that includes a user interface (UI) 1040 for providing, for example, user selection of algorithms, binding measurement data, candidate proteins, and databases. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0173] The methods and systems of the present disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software when executed by the central processing unit 1005. The algorithms may, for example, receive empirical measurements of proteins present in a sample, compare the empirical measurements against a database containing a plurality of protein sequences corresponding to candidate proteins, and generate a probability that the candidate protein produces the observed set of measurements and / or generate a probability that the candidate protein is correctly identified in the sample.

[0174] The present disclosure provides a non-transitory information storage medium having encoded thereon instructions for performing one or more steps of the methods described herein, for example, when these instructions are non-abstractly executed by an electronic computer. The present disclosure further provides a computer processor (i.e., not a human mind) configured to non-abstractly perform one or more of the methods described herein. It is understood that all methods, compositions, devices, and systems described herein can be embodied in physical, tangible, and non-abstract forms. The claims are intended to encompass physical, tangible, and non-abstract subject matter. Any explicit limitation of any claim to physical, tangible, and non-abstract subject matter, when taken as a whole, is understood to limit the claim to encompass only non-abstract subject matter. References to "non-abstract" subject matter exclude and are distinct from "abstract" subject matter as construed by binding precedents of the Supreme Court and the Federal Circuit at the priority date of this application. [Example 1] Single molecule protein identification using multi-affinity protein affinity reagents

[0175] This example describes the basis of high-throughput single molecule protein identification. The approach uses multi-affinity reagents that bind short linear epitopes with low specificity and a decoding algorithm that embraces the expected serendipity of single molecule binding. In simulations, the approach achieved high proteome coverage in a wide range of organisms and was robust to potential experimental confounders. In a simulated human plasma proteome experiment, the approach supported a dynamic range of detection spanning at least eight orders of magnitude. Results showed that, when implemented experimentally, the approach can quantitatively decode more than 90% of the human proteome in a single experiment, potentially revolutionizing proteomics research. Results and Discussion

[0176] As a preliminary matter, this example describes methods that can be used to identify and distinguish proteins based on their primary structure (i.e., amino acid sequence). In this context, reference to different proteins, whether implicit or explicit, relates to differences in their primary structure. Notwithstanding the above, the methods exemplified herein may in some cases be useful for identifying proteins based on differences such as the presence, number, type or location of post-translational modifications, with modifications that will be apparent to one of skill in the art.

[0177] Figure 1A shows the experimental setup for detecting multiple proteins with single molecule separation. Proteins were extracted from a sample and each protein was conjugated to structured nucleic acid particles (SNAP) under denaturing conditions, followed by 10 10Protein-conjugated SNAP is deposited on a solid support with addresses. No more than one protein-conjugated SNAP is bound per address, creating an ultra-high density single molecule array with each address having a protein optically resolvable from adjacent addresses. A series of affinity reagents tagged with fluorophores (e.g., antibodies, aptamers or small proteins) are contacted with the array. One affinity reagent is used per series of cycles, the presence or absence of binding is detected at each address, and the affinity reagent is washed off the array before the next affinity reagent is added via the next cycle. Integration of fluidics and image processing on the instrument allows high-resolution multi-cycle imaging of the addresses in the presence of affinity reagents. Thus, binding of affinity reagents to proteins results in a series of bound / unbound results for each protein, which can be used to infer the identity of the protein. Since there is only one protein per address, direct counting of addresses can be used to quantify each protein identified in the sample.

[0178] To identify many different proteins in the human proteome or other complex proteomes, a prohibitively large number of highly specific affinity reagents would be required. The present method overcomes this by using affinity reagents that bind short linear epitopes (e.g., trimers) with moderate specificity, such that each affinity reagent binds many different proteins. Binding of a single affinity reagent is insufficient to identify a specific protein with these promiscuous affinity reagents, but a series of affinity reagents can decode many different proteins. By detecting each new affinity reagent bound at each address over an increasing number of cycles, the list of possible protein identities at each address is gradually narrowed (Figure 1B).

[0179] In a typical single molecule binding reaction format, the binding is stochastic, since the affinity reagent is not always observed to bind the protein containing its epitope (see Chang et al., J Immunol Methods 378, 102-115 (2012), incorporated herein by reference). In addition, each affinity reagent may be observed to bind off-target epitopes. Thus, when the same series of single molecule binding reactions is repeated multiple times, multiple distinct binding patterns are typically observed (Figure 1C).

[0180] Taking this contingency into account, a binding model was devised in which each affinity reagent binds to a protein containing one copy of its target epitope with a first probability and binds to a protein containing one copy of an off-target epitope with an equal or lower probability. Since there are many factors that may prevent the affinity reagent from binding to its epitope, such as residual or transient protein structure due to partial denaturation, the presence of post-translational modifications, binding contingency, etc., a fairly low probability of 0.5 for on-target binding to its primary epitope and a probability of 0.5 for binding to off-target epitopes were initially selected. To determine affinity reagent selectivity that provides high coverage of the human proteome with a manageable number of different affinity reagents, affinity reagents with various target epitope lengths (dimers, trimers, or tetramers) and various numbers of off-target epitopes were evaluated. As shown in FIG. 1D, the analysis showed that 100 affinity reagents would facilitate the unique identification of 90% of the human proteome, with each affinity reagent binding a single trimer and 9 additional primary off-target trimers. In this scenario, each affinity reagent would bind approximately 23.7% of the proteins in the human proteome (note that the percentage is based on the number of unique protein sequences, independent of the variability of the expression level of each protein), and on average, approximately 24 binding events would be sufficient to identify a given protein (Table 1). Targeting tetramer epitopes would lower the number of binding events, but increase the number of affinity reagents sufficient to achieve similar coverage. Targeting dimer epitopes would allow a similar number of affinity reagents, but it may be difficult to generate affinity reagents that recognize dimers independently of the variability of the sequences surrounding the dimer. Therefore, a "trimer with 10 epitopes" affinity reagent selectivity model was used for this analysis. [Table 1]

[0181] It is also possible to use affinity reagents that are more specific, for example, that bind to a single epitope or a single protein. In some cases, multiple different affinity reagents can be combined to create a pool of affinity reagents that bind with apparent indiscrimination. For example, a pool of three different affinity reagents that are detected indiscriminately in the binding step will appear to indiscriminately bind the proteins targeted by the pool. As a more specific example, a pool of three different affinity reagents can apparently bind at least three different proteins, a pool of five different affinity reagents can apparently bind at least five different proteins, a pool of ten different affinity reagents can apparently bind at least ten different proteins, etc.

[0182] In addition to having a primary binding epitope, affinity reagents are likely to bind other off-target epitopes, albeit with lower probability. A "biosimilar" affinity reagent model (see Methods section below) was used, where each affinity reagent had a "tail" of up to 20 additional secondary off-target epitopes, with binding probability proportional to the similarity of the off-target epitopes to the target epitope. Using this model with target epitopes randomly selected from targets present in the human proteome, the decoding algorithm was able to uniquely identify approximately 98% of the proteins in the human proteome in 300 cycles (modeling a sample with one copy of each protein) (Figure 1E). Performance with fewer than 200 affinity reagents improved when a greedy selection algorithm (see Methods section below) was used to determine an optimal set of 300 trimer epitopes to achieve high human proteome coverage with as few affinity reagent cycles as possible (Figure 1E). This optimal set of epitopes was used for subsequent analyses.

[0183] To test whether the decoding strategy can be applied to proteomes from species other than human, the analysis of proteomes from mouse, S. cerevisiae, and E. coli was simulated using the same parameters with the same set of optimized affinity reagents (Figure 1F). Surprisingly, there were few differences between species, indicating that although smaller proteomes are slightly easier to decode, it is the protein sequence diversity that is the main driver of improved decoding performance. Thus, despite the stochastic nature of single molecule binding, the decoding strategy has the potential to decode over 90% of the proteome for a wide range of organisms.

[0184] Potential experimental confounders were evaluated. A first scenario was considered in which the probability of binding of the affinity reagent to the epitope was even lower than 0.5, for example due to poor binding affinity or kinetics. Even with a probability of 0.1, the decoding method achieved over 85% proteome coverage using 300 cycles (i.e., 300 different affinity reagents), whereas with a binding probability of 0.05, the proteome coverage dropped to about 55% (Figure 2A). Options for increasing coverage include, for example, using more affinity reagents, multiplexing several affinity reagents in a single run (e.g., using different fluorescent labels for each probe in a multiplexed set); running the affinity reagents in repeated cycles to improve the chances of observing binding; increasing the concentration of the affinity reagents; increasing the duration of the binding reaction; or attaching multiple copies of the affinity reagent to a scaffold such as a fluorescent particle or a structured nucleic acid particle. Thus, the decoding method may be feasible using affinity reagents over a range of binding probabilities, some of which are relatively low.

[0185] The effect of non-specific binding of affinity reagents to the surface of the array at locations close enough to the protein address to generate false binding signals was evaluated. Assuming a binding probability of 0.5, a non-specific binding rate of 0.05 or less gave a detection sensitivity of approximately 90%, as demonstrated by Figure 2B. For subsequent analyses, a non-specific binding rate of 0.001 was assumed. If experimentally the rate was found to be higher, the binding conditions (e.g., ionic strength, temperature, polarity, pH, osmolality, concentration of affinity reagent or surface tension) can be adjusted to reduce non-specific binding. For each affinity reagent, the same or different conditions can be used.

[0186] The impact of affinity reagent characterization (e.g., identification of target and off-target epitopes and their respective binding probabilities) was also evaluated. Such characterization can be performed in a straightforward manner using conventional epitope mapping approaches (Beyer et al., Science 318, 1888 (2007), incorporated herein by reference). For example, if each affinity reagent binds an additional number of epitopes that are unknown to the inference algorithm, trimer epitopes may be "missed" during affinity reagent characterization (Figure 2C, Figure 4A). However, the impact was small unless high probability (0.5) binding epitopes were consistently missed. When up to 20% of these epitopes were missed, proteome coverage remained above 92%. Trimer epitopes may also be erroneously identified as targets during affinity reagent characterization (Figure 2D, Figure 4B). The decoding method appeared to be robust against this type of error, since it achieved nearly 70% coverage even though half of all primary epitopes were incorrect. Given that the decoding method appeared to be more robust against having false positive epitopes in the affinity reagent model than against "missing" epitopes, the techniques used to characterize affinity reagents could be tuned toward sensitivity rather than specificity to achieve improved results. Evaluation of the impact of consistent over- or underestimation of the binding probability of affinity reagents to epitopes showed that the impact of such errors was small, except for large (>-0.2) underestimation of binding probability (Figure 2E, Figure 4C). The decoding method appeared to be extremely robust against noisy affinity reagent characterization, indicating that affinity reagent characterization does not need to be perfect, and that the method tolerates variability in affinity reagent binding properties that may arise from other potential experimental confounding factors such as temperature (Figure 2F, Figure 4D). In summary, the decoding method appeared to be robust to errors in the characterization of the affinity reagents.

[0187] Plasma is a good example of one of the major challenges for proteomics, since plasma protein concentrations can vary over 12 orders of magnitude, and typical mass spectrometry-based approaches typically identify only 8% of the proteome (see Anderson & Anderson, Mol Cell Proteomics 1, 845-867 (2002), incorporated herein by reference). To evaluate the theoretical performance of protein decoding strategies, 10 6 , 10 8 and 10 10 Simulations were performed to assay a non-depleted plasma sample with 300 affinity reagents on an array with 10 addresses. The simulations modeled running the same sample across five technical replicates. Some random noise in the binding probability of the affinity reagents to the trimers simulated the variability of affinity reagent binding across replicates. On average, 10 10 Simulations running the decoding algorithm with the address array demonstrated a dynamic range of detection spanning over 11.5 orders of magnitude, from the most abundant to the least abundant proteins detected (Figure 3A, Figures 5A-5F). The decoding method was able to quantify 59.4% of the 20,235 proteins in the modeled plasma samples. Almost all proteins were quantified with high specificity (Figures 6A-6C). Over 99.6% of the proteins measured had a quantitative specificity of over 90% (i.e., over 90% of protein identifications were true positives). Proteins within the top 9 orders of magnitude dynamic range were detected with 90% consistency. No bias in identifiability correlating with protein concentration was observed. Overall, 90% of proteins deposited on the array were detected, indicating that the ability to deposit low abundance proteins on the array, rather than the ability to decode the proteins, is the primary limiting factor in the dynamic range. Modeling demonstrated that the number of addresses could be increased to 10 11 or 10 12 These results suggest that increasing the concentration of IgG to IgG increases the identification of proteins deposited on the array from 66% to 79% and 92%, respectively (Figures 7A-7C).

[0188] Experimentally, the dynamic range can be compressed by depleting the most abundant proteins in plasma samples, for example, using affinity columns. Plasma samples modeled with 99% depletion of the top 20 proteins had, on average, 65.7% proteome coverage (Figures 8A-8D). When HeLa cell line samples were modeled, coverage was significantly higher (92.6%) but had a lower dynamic range (detection spanned 9.5 orders of magnitude) (Figure 3B).

[0189] Because detectability is not only a factor of abundance but also of sequence similarity, some proteins with relatively high abundance were not detected in all samples. If the sequence of a protein is very similar to another protein in the database, it may be difficult for the decoding algorithm to generate a reliable identification for these proteins. More selective affinity reagents can be used to detect these more difficult targets.

[0190] A strategy to increase throughput is to add 10 8 A more cost-effective approach would be to use arrays of protein addresses (e.g., multiplexing multiple proteomic samples onto one array, or running multiple smaller arrays in parallel). In this situation, low abundance proteins would become undetectable, resulting in a compressed dynamic range spanning 7.5 orders of magnitude in plasma (for consistently detected proteins), but with high coverage within that range (Figures 9A-9I).

[0191] Measurement reproducibility was evaluated across five technical replicates of modeled plasma and HeLa samples (Figures 3C and 3D). For mid- to high-abundance proteins, the coefficient of variation (CV) was less than 10%. Proteins within the top 5 orders of magnitude in terms of abundance in plasma samples generally had CVs less than 1%. As modeled, factors contributing to irreproducibility were stochastic variation in affinity reagent binding and protein deposition as well as variation in affinity reagent binding properties. Although these estimates do not account for many sources of experimental and biological variability such as sample preparation, they demonstrate the potential of this analytical platform and decoding algorithm to provide minimal variation relative to more common sources of variation. Indeed, the observed CVs in measured counts were not significantly different from the CVs of actual counts, indicating that measurement reproducibility can be improved by increasing throughput (Figures 10A and 10B).

[0192] The detected protein counts correlated with the number of proteins modeled on the array (Figures 3E and 3F). 76% of the plasma proteins had a fold change error of detected counts versus counts on the array within + / - 10% (Figure 11). In some cases, proteins with only a single copy on the chip were detected. Some proteins were significantly undercounted due to sequence similarity to other proteins in the sequence database. The linearity of detected counts versus counts on the array was consistent with the linearity of detected counts versus counts on the array when the array was scaled to 10. 11 We have shown that the dynamic range can be further expanded by extending to individual addresses or evaluating samples across multiple arrays.

[0193] In conclusion, the results presented in this example provide a theoretical basis for a single-molecule protein identification method that is invariant across proteomes and can be used to analyze the entire human proteome in a single experiment. It has important advantages over other proteome analysis methods. It is unique among emerging single-molecule peptide sequencing methods in that it employs a non-destructive affinity reagent approach rather than chemically intensive or cleavage-based sequencing approaches. It is robust against false negatives (i.e., the affinity reagent does not bind its epitope) and is optimized for non-specific affinity reagents. Thus, this decoding method turns a common weakness of affinity-based proteomics approaches into a strength. This decoding method is scalable to full proteome quantification and, unlike mass spectrometry, can be quantified over a wide dynamic range. By using intact proteins, this decoding method avoids the loss of information (such as proteoforms) that limits approaches based on detecting peptide fragments of proteins and partially mitigates the dynamic range challenge as the sample complexity is reduced by about two orders of magnitude. If successfully implemented experimentally, this decoding method will provide an easy-to-use, rapid, highly sensitive and reproducible method for analyzing and quantifying the proteome, even from a single cell, and is expected to open countless new opportunities in scientific discovery, not only in basic research but also in clinical research, such as molecular diagnostics and biomarker discovery.

[0194] The simulations described in this example demonstrate the potential of implementing a sensitive and rapid imaging platform. A particularly useful detection system would have rapid imaging and cycle speeds, since the dynamic range of the illustrated decoding method is directly related to the number of intact protein molecules measured. Preliminary estimates suggest that it is possible to profile 10 billion protein molecules within about one day using 300 affinity reagents and a cycle time of about 10 minutes. Successful experimental implementation of this decoding method would provide an easy-to-use, rapid, highly sensitive, and reproducible method for analyzing and quantifying the proteome, even from a single cell. Successful experimental implementation of this decoding method would open the way to countless new opportunities in scientific discovery, not only in basic research but also in clinical research, such as molecular diagnostics and biomarker discovery. method Protein Sequence Databases

[0195] Protein sequence databases were downloaded from Uniprot (www.uniprot.org). For each species, a "reference" proteome was selected by including "reference: yes" in the search query string for the proteome. The reference proteome was then filtered to include only curated (Swiss-prot) sequences (query string "curated: yes"). Sequence data were then downloaded in uncompressed fasta format (canonical sequences only). The specific proteomes and filter strings used were as follows: E. coli (K12 strain): Scrutinized: Yes & Biology: E. coli (K12 strain)

[83333] & Proteome: up000000625 (downloaded June 30, 2021) Saccharomyces cerevisiae (s288c): vetted: yes and organism: "Saccharomyces cerevisiae (ATCC 204508 / S288c strain) (baker's yeast) [559292]" and proteome: up000002311 (downloaded June 30, 2021) M. musculus (C57BL): Reviewed: Yes & Biology: "Mus musculus (Mouse)

[10090] " & Proteome: up000000589 (downloaded June 30, 2021) H. sapiens: scrutinized: yes and biology: "Homo sapiens (Human)

[9606] " and proteome: up000005640 (downloaded on July 6, 2021)

[0196] The proteome was further processed to remove any duplicated sequences and any sequences not entirely composed of the 20 standard amino acids. Additionally, sequences less than 30 in length were removed from each FASTA. Modeling the binding of affinity reagents to proteins

[0197] We modeled affinity reagents targeting epitopes of length k (e.g., for trimers, k=3) by assigning a binding probability θ to each unique target epitope j of length k recognized by the reagent. In addition, the protein nonspecific binding rate is denoted by p nsbエピトープ Given the primary sequence of a protein of length M, the probability that an affinity reagent will bind to the protein was calculated as follows:

[0198] First, the probability of a specific binding event occurring was calculated:

number

number

[0199] The probability that an affinity reagent will bind to a protein and generate a detectable signal is the probability that one or more specific or non-specific binding events will occur. p タンパク質結合 =1-(1-p 特異的 )*(1-p 非特異的 )

[0200] Where noted, the probability of binding to each protein was adjusted to account for additional random surface nonspecific binding (NSB); i.e., binding of affinity reagent to an array close enough to a protein address will generate a false positive binding event. The incidence of surface NSB is the probability of such a surface NSB event occurring during the acquisition of a single affinity reagent measurement at a single protein location on the array, 0 ≤ p 表面nsb Defined as <1. The adjusted probability of a protein binding event taking into account surface NSB was: p 調整された結合 =1-(1-p タンパク質結合 )*(1-p 表面nsb ) Biosimilar affinity reagent models

[0201] Unless otherwise stated, affinity reagents were modeled using the "biosimilar" model. In this model, affinity reagents target a particular epitope that the affinity reagent binds with probability 0.5. The affinity reagent also binds 9 additional primary off-target epitopes that are biosimilar to the targeted epitope with probability 0.5. Biosimilar targets were selected by calculating the pairwise similarity score of the target epitope to all other possible epitopes of the same length. Similarity scores were calculated by summing the BLOSUM62 similarities between pairs of residues at each sequence position. For example, when calculating the similarity between trimer SLL and trimer YLH, the score would be BLOSUM62(S,Y)+BLOSUM62(L,L)+BLOSUM62(L,H). Once all pairwise similarity scores were calculated, the top 9 epitopes most similar to the target were selected as the primary off-target epitopes. In the case of ties where multiple potential off-target epitopes had the same score, a random epitope was selected. In addition to the target epitope and the four off-target epitopes, up to 20 additional biosimilar secondary off-target epitopes with lower binding probabilities were added to the affinity reagent. The 20 secondary off-target epitopes bind to the next 20 most biosimilar epitopes beyond those already included in the affinity reagent model. These 20 additional epitopes have probabilities calculated as follows: b*(1.5 ot-ss ) Where: b = probability of binding of the affinity reagent to its target, ot = BLOSUM62 similarity score between the affinity reagent target and this off-target epitope, · ss = BLOSUM62 similarity score between the affinity reagent target and itself. If any of these additional off-target epitopes had a binding probability less than the affinity reagent epitope nonspecific binding rate, then that epitope was not included. -8 was set to. Simulation of stochastic affinity reagent binding

[0202] To simulate the binding of a set of affinity reagents to a single protein, we first calculate the binding probability θ of each affinity reagent i to the protein, using the method described in the Modeling the binding of affinity reagents to proteins section above. i To simulate the binding results for each affinity reagent, θ i We performed a single random draw from a Bernoulli distribution with parameter . A result of 1 is merging, a result of 0 is no merging. Protein Decoding overview

[0203] The protein decoding algorithm analyzed the set of affinity reagent binding measurements obtained for an extant protein and determined the most likely identity of that protein among the candidate set. The most likely protein identity was the one that best matched the observed binding measurements. This match was determined based on the binding model of each affinity reagent in the experiment that was used to estimate how likely each affinity reagent was to bind each potential protein. A strong candidate protein was one where most of the observed binding events matched affinity reagents that were likely to bind the protein. A weak candidate protein would have many instances where binding was observed to affinity reagents that were not expected to bind the candidate. The strongest candidate protein was considered the most likely identity for the extant protein, and the confidence in this identification could be calculated as a relative measure of the match of the most likely protein compared to all of the other candidates. input

[0204] The inputs to the decoding algorithm were: Combined data: D=[d1,d2,d3,..d N ], where d∈{0 (non-binding), 1 (binding)}. A series of binding measurements, one for each affinity reagent for the extant protein. A sequence database of length M that contains the primary sequence and name of each potential protein that may be present in the sample (e.g., the Human Protein Sequence Database described in the Protein Sequence Databases section above). ·Parameterized binding models for each of the N affinity reagents used in the experiment (see section Modeling the binding of affinity reagents to proteins above). Any surface non-specific binding rate (r), which describes the probability that a surface non-specific binding event occurs at any one address in any given cycle. Joint Probability Calculation

[0205] Calculate the M × N binding probability matrix B that describes the probability that each affinity reagent binds to all possible candidate proteins, and define the matrix b i,j The components of x are the probabilities of affinity reagent j binding to candidate protein i. These probabilities were calculated using the methods described in the Modeling the Binding of Affinity Reagents to Proteins section above.

[0206] Next, an M × N matrix U with the adjusted non-binding probabilities for each affinity reagent for each protein was calculated as follows: S=[s1,s2,s3,...s M ], where s i = Protein i Length - 2. F=[f1,f2,f3,...f 8000 ] calculate the relative frequency of all possible unique trimers among the set of all candidate protein sequences, where:

number

number

number

[0207] To avoid a single non-binding event having a very large impact on the protein, the adjusted non-binding probability was calculated this way (rather than U=1-B). The rationale was that there are great difficulties in predicting why an affinity reagent may not bind to a specific epitope (e.g., protein structure, post-translational modifications), and therefore the total number of non-binding events should be expected to be greater than the specific identity of the observed non-binding events. Decryption

[0208] A vector of likelihoods for each protein in the candidate database was calculated by multiplying the likelihood of each observed binding event: L=[L1,L2,L3...L M ]In the formula,

number

number

[0209] To calculate proteome coverage, a set of affinity reagents was defined as in the Modeling Affinity Reagent Binding to Proteins section above. For each protein in the human proteome defined in the Protein Sequence Database section above, affinity reagent binding was simulated (see the Simulation of Stochastic Affinity Reagent Binding section above). The binding data was then passed to a decoding algorithm along with the affinity reagent definitions and the FASTA sequence database. The output of the decoding algorithm was a single protein identification for each simulated protein and an estimated probability that the identification was correct. To calculate the percentage coverage, the number of proteins identified above a 1% true / false discovery rate threshold (see False Discovery Rate Calculation and Thresholding section below) was divided by the total number of proteins simulated. The percentage coverage was calculated by multiplying the percentage coverage by 100. This method was applied to all analyses except for modeling of cell, plasma and depleted plasma samples, which used the method described in the Quantitative Statistics section below. False discovery rate calculation and threshold setting

[0210] Given a list of decoded protein identities (protein identities and associated probabilities), we first calculated the false discovery rate by annotating each protein identification as correct or incorrect based on its match to the protein's true identity in the simulation. For each unique identification probability in the list, we calculated the false discovery rate (FDR) as the proportion of proteins that were incorrectly identified at or below that probability. To set the false discovery rate as a threshold, we determined the lowest probability score threshold that had an FDR smaller than the desired FDR. Identifications at or above this probability score met the FDR criterion and were considered "identified" at the desired FDR threshold. Demonstration of probabilistic connections

[0211] Stochastic binding of a series of 10 affinity reagents to the protein EGFR was simulated six times (Figure 1C). Affinity reagents with a binding sequence present in EGFR have a probability of binding of 0.5, and affinity reagents without a binding sequence in EGFR have a probability of binding of 0. Binding was simulated as described in the Simulation of Stochastic Affinity Reagent Binding section above. Assessment of affinity reagent requirements for efficient decoding

[0212] Affinity reagents with various target epitope lengths (2, 3 or 4, i.e. dimer, trimer, tetramer, respectively) with various numbers of primary off-target epitopes were modeled. In each case, the target binding probability was 0.5. "Number of epitopes per affinity reagent" = 1 represents an affinity reagent targeting a single epitope with no primary off-target epitopes. Other scenarios were modeled using affinity reagents with several biologically similar (see Biosimilar Affinity Reagent Model section above) primary off-target epitopes. For example, an affinity reagent described as targeting "five" epitopes has binding affinity for its target and four primary off-target sites. The affinity reagent did not have any secondary off-target epitopes (see Biosimilar Affinity Reagent Model section above). Targets for affinity reagents were randomly selected from targets present in the proteome. Off-target binding epitopes were not required to be present in the proteome.

[0213] To determine the number of affinity reagents required to achieve 90% coverage of the proteome, binding of excess affinity reagents (i.e., more than required for 90% coverage) to each protein in the proteome was simulated. For any number N of affinity reagents, the first N affinity reagents in the set were used to calculate proteome coverage. The number of affinity reagents required to achieve 90% proteome coverage was the lowest N that had 90% or greater coverage. Values ​​of N tested were in increments of 10.

[0214] The number of affinity reagents (N) required for 90% coverage was calculated, and the number of binding events observed for each simulated protein was recorded, and the average of these values ​​was reported as the "average number of binding events per protein." Additionally, the percent of proteins giving rise to binding events for each affinity reagent was recorded, and the average of these values ​​was reported as the "percent of protein bound per affinity reagent." Selection and evaluation of optimal affinity reagent trimer targets

[0215] In this analysis with trimer-targeting affinity reagents, the standard biosimilar affinity reagent model (see Biosimilar Affinity Reagent Model section above) was used. To achieve high proteome coverage with as few affinity reagents as possible, one set of "optimal" affinity reagent targets was calculated by estimating the optimal set of 300 targets using a greedy selection algorithm. In addition, 20 sets of 300 targets were randomly selected among the trimers present in the proteome (excluding any trimers containing cysteine). Proteome coverage for each of the 21 affinity reagent sets was evaluated as described in the Calculation of Proteome Coverage section above. To evaluate the scaling of proteome coverage with the number of affinity reagents used, proteome coverage was also evaluated for multiple 1st to Nth reagent subsets of each affinity reagent set.

[0216] An optimal set of trimer targets was selected as follows: 1. Initialize an empty list of selected affinity reagents (AR). 2. Initialize a set of candidate ARs (e.g., a collection of 6,859 ARs, each targeting a unique trimer that does not contain a cysteine ​​in it). 3. Select a set of protein sequences against which to optimize (e.g., all human proteins in the UniProt reference proteome). 4. Repeat the following until the desired number of ARs are selected: a. For each candidate AR, i. Simulate binding of the candidate AR to the set of proteins. ii. Perform decoding for each protein using the simulated binding measurements from the candidate AR and all previously selected ARs. iii. Calculate a score for the candidate AR by summing the probability of correct protein identification for each protein as determined by protein inference. b. Add the AR with the highest score to the set of selected ARs and remove it from the candidate AR list. Assessing proteome coverage in multiple organisms

[0217] Proteome coverage was evaluated for four different organisms using 300 affinity reagents targeting the optimal trimer set designed against the human proteome (see section Selecting and Evaluating Optimal Affinity Reagent Trimer Targets above). The sequence database for each organism is described in the Protein Sequence Database section above. For each organism, binding was simulated using an affinity reagent epitope binding affinity of 0.5 for each affinity reagent to each protein in the sequence database for that organism. The binding data was then decoded using the appropriate sequence database for that organism, and various 1-N subsets of the 300 affinity reagent set were used to evaluate proteome coverage as described in section Calculating Proteome Coverage. For example, to calculate coverage with 100 affinity reagents for a given organism, only data from the first 100 of the total 300 affinity reagents were considered during decoding. Application of noise to affinity reagent binding probabilities

[0218] We devised a method to model random perturbations in affinity reagent binding properties. The method applied random "noise" to the trimer (or other short linear epitope) binding probability while keeping the probability of binding between 0 and 1. Given a binding probability p, we determined the perturbed probability by drawing a sample from the distribution: Φ(Φ -1 (p)+N(0,σ 2 )) During the ceremony, N is normally distributed, σ 2 is a parameter used to adjust the intensity of the disturbance, ·Φ is the cumulative distribution function of the standard normal distribution.

[0219] Parameter σ 2was set so that the mean absolute deviation (MAD) of the distribution divided by the trimer probability p was equal to the desired target. This adjustment parameter is called the "fractional MAD." The fractional MAD was used to adjust for noise because of its conceptual similarity to the coefficient of variation (standard deviation divided by the mean), which is often used to describe measurement noise or repeatability for normally distributed measurements.

[0220] σ for the probability p of yielding the desired fractional MAD 2 A numerical approximation method was used to find the value of . First, given p and the desired fractional MAD, the target MAD was calculated as fractional MAD*p. p target MAD and the proposed σ 2 Given the values ​​of p and σ 2 A function optim is defined that generates 10,000 random samples from a noise distribution with parameters σ and returns the absolute value of the difference between the MAD of the 10,000 random samples and the target MAD. The minimize_scalar function from the scipy Python package finds the function σ that minimizes this function. 2 This process is repeated 50 times, and the best median σ^2 value among the 50 trials is deemed to be the appropriate value for generating a noise distribution with the desired MAD. Experimental confounding modelling Poor binding affinity

[0221] Proteome coverage (see Calculating Proteome Coverage section above) was assessed using 300 affinity reagents targeting the optimal trimer set (see Selecting and Evaluating Optimal Affinity Reagent Trimer Targets section above) that binds to each unique protein in the human proteome (Figure 2A). However, to simulate various affinity reagent binding affinities, affinity reagents were modeled using various target epitope binding rates ranging from 0.01 to 0.99. To model the relationship between the number of affinity reagents used and proteome coverage, various 1st to Nth subsets of the 300 affinity reagent set were used to assess proteome coverage as described in Calculating Proteome Coverage section. Binding simulation and decoding were repeated five times to generate replicate analyses. Non-specific binding to the array surface

[0222] Proteome coverage was evaluated using various combinations of affinity reagent binding affinities and nonspecific binding rates. In all cases, 300 affinity reagents targeting the optimal trimer set (see section Selecting and Evaluating Optimal Affinity Reagent Trimer Targets above) were used. However, affinity reagents were modeled using various target epitope binding rates ranging from 0.05 to 0.95 to simulate various affinity reagent binding affinities and various surface nonspecific binding ranging from 0 to 0.3. After modeling binding using surface NSB, proteome coverage was calculated as described in section Calculating Proteome Coverage above. Trimers overlooked during affinity reagent characterization

[0223] Binding measurements (see Simulation of Stochastic Affinity Reagent Binding section above) were generated for each of the optimal set of affinity reagents for each of the proteins in the human FASTA database (see Protein Sequence Database section above) with a surface NSB rate of 0.1% (see Nonspecific Binding to Array Surface section above). Prior to decoding the binding measurements to generate protein IDs, errors were introduced into the affinity reagent model by removing some of the primary epitopes. Such errors can occur in experimental situations, for example, if the method used to determine the epitopes to which affinity reagents bind misses some number of epitopes. The error-introduced affinity reagent model was used when decoding the binding measurements to generate protein IDs, which was expected to reduce the decoding performance. The severity of the errors was adjusted by adjusting the percentage of missed primary epitopes. To build a model in which 20% of the primary epitopes were missed, a random 20% of the primary epitopes (among all affinity reagents combined) were selected for removal. Since an optimal affinity reagent has 10 primary epitopes, this means that on average, two primary epitopes were missed in each affinity reagent, although due to random chance some may have more than one removed and others may not have any removed. In some analyses, a percentage of secondary epitopes were removed as well. Misidentification of trimeric epitopes during affinity reagent characterization

[0224] Similar to the section on trimers overlooked during affinity reagent characterization above, affinity reagent binding to proteins in the proteome was simulated with a surface NSB of 0.1% to generate errors in the affinity reagent model before decoding. For this analysis, false positive epitopes were added to the affinity reagents before decoding. This simulates a scenario in which the method used to characterize the epitopes bound by each affinity reagent incorrectly identifies some number of trimer epitopes that the affinity reagent does not bind. The severity of the error generation was adjusted by adding false primary epitopes such that the complete set contains a certain percentage of false epitopes. For example, 20% false epitopes means that false primary epitopes were added until 20% of the primary epitopes in the affinity reagent set were false. The additional epitopes were randomly distributed among the affinity reagents. The trimer identities of the additional epitopes were randomly selected with replacement. In some analyses, secondary epitopes were also affected by the error generation. Any added secondary epitopes must not coincide with existing or added primary epitopes. For example, an affinity reagent targeting primary epitopes HNW, HDW and HHW and secondary epitopes HRW and HGW can have LWW added as either an erroneous primary or secondary epitope, but HGW can only be added as an erroneous primary epitope, in which case its binding probability is updated to that of the primary epitope. Consistent over- or underestimation of affinity reagent trimer binding

[0225] Similar to the section on trimers overlooked during affinity reagent characterization above, affinity reagent binding to proteins in the proteome was simulated with a surface NSB of 0.1% to generate errors in the affinity reagent model prior to decoding. In this analysis, epitope binding probabilities were adjusted systematically higher or lower than the true value. This models a situation where the affinity reagent characterization method determines the correct trimer epitope targeted by the affinity reagent, but systematically overestimates or underestimates the strength of binding (as modeled by the binding probability). The manipulation involved applying some fold-change shift to the epitope binding probability such that the affinity reagent's primary epitope was shifted by the desired amount. For example, to model a shift of +0.25 for an affinity reagent with a true primary epitope binding probability of 0.25, the binding probabilities of all epitopes in the affinity reagent were multiplied by 2. In this case, a primary epitope with a true binding probability of 0.25 is assumed to bind with a probability of 0.5 when performing decoding. Similarly, this same multiplicative shift can be applied to secondary binding epitopes. For example, a secondary epitope with a binding probability of 0.2 has a binding probability of 0.4. Similarly, adjustments can be made to adjust the binding probability to a smaller value. In some analyses, the severity of the error was adjusted by only introducing errors into a portion of the affinity reagents. For example, 50% of the affinity reagents can be affected, meaning that half of the affinity reagents have a systematic error in their binding probability and the rest are unaffected. Characterization of noisy affinity reagents

[0226] Similar to the section on overlooked trimers during affinity reagent characterization above, binding of affinity reagents to proteins in the proteome was simulated with a surface NSB of 0.1% to introduce errors into the affinity reagent model before decoding. In this analysis, random noise was applied to the characterized epitope binding probabilities. Random noise was applied to a random percentage of affinity reagents in the set. For any affinity reagent affected by noise, all primary and secondary epitopes were subjected to some degree of noise and nonspecific binding rate of the affinity reagent. The binding probabilities were perturbed according to the method described in the section on applying noise to affinity reagent binding probabilities above, with amounts of noise ranging from fractional MAD 0 to 0.75. Simulation of cell line and plasma experiments Protein abundance database processing

[0227] Protein abundances downloaded from PaxDb v4.1 (Wang et al., Molecular Cellular Proteomics, 8:492-500 (2012) doi:10.1074 / mcp.O111.014704, incorporated herein by reference) were used to model the protein composition of each sample. Specifically, plasma protein abundances were from the "H. sapiens-Plasma (integrated)" dataset (https: / / pax-db.org / downloads / 4.1 / datasets / 9606 / 9606-PLASMA-integrated.txt downloaded September 2021). Cell line abundances were from the dataset "H. sapiens-cell lines, Hela, SC (Nagaraj, MSB, 2011)" (pax-db.org / downloads / 4.1 / datasets / 9606 / 9606-hela_Nagaraj_2011.txt) constructed from high-resolution mass spectrometry of HeLa cells (Nagaraj Molecular Systems, incorporated herein by reference). Biology, 7:548(2011).doi:10.1038 / msb.2011.81). Protein identities in the PaxDb data were mapped to protein identities in the Uniprot human protein sequence database (see Protein Sequence Databases section above) using the PaxDb to Uniprot mapping available from the PaxDb maintainers at https: / / pax-db.org / downloads / 4.1 / mapping_files / uniprot_mappings / full_uniprot_2_paxdb.04.2015.tsv.zip (downloaded September 2021). Any proteins present in the PaxDb database that could not be mapped to the UniProt sequence database were removed from the samples. 4,342 of 4,492 entries in the plasma database (97%) were successfully mapped, and no unmapped proteins accounted for more than 1% of the samples.Of the 8,817 entries in the cell database, 8,554 (97%) were successfully mapped, with no unmapped proteins accounting for more than 1% of the samples. In some cases, more than one entry in the PaxDb database was mapped to a single UniProt identifier in the sequence database. In these cases, only the first entry was retained. In the plasma database, 99 database entries were removed as a result of this operation (4,243 entries remaining). In the cell line database, 145 entries were removed (8,409 entries remaining). None of these operations removed any entries that accounted for more than 1% of the corresponding samples. 25 and 97 proteins with an abundance of 0 were removed from the plasma and cell line databases, respectively. After filtering, the abundance databases were normalized to sum to 1. Protein Abundance (Plasma) Data Complement

[0228] Proteins in the human protein sequence database that were not represented in the modeled plasma sample (see Protein Abundance Database Processing section above) were imputed with abundance. This process resulted in a "complete" plasma sample containing 20,235 proteins with a dynamic range of abundance of 12 orders of magnitude. The distribution of abundance in the complete plasma sample was modeled as a semi-Gaussian distribution (Eriksson, Nature Biotechnology, 25:651-655 (2007) doi:10.1038 / nbt1315, incorporated herein by reference): Let f(a|μ,σ) be the normal distribution probability density function with mean μ and standard deviation σ evaluated at x.

number

[0229] Next, the probability density function for the abundance of the protein requiring complementation was estimated. Based on the inference that any protein with log 10 (abundance) > t present in a "complete" plasma sample is accurately represented in PaxDb (i.e., not affected by detection bias), the threshold was set to t = A max - 4 for "high - abundance" proteins. A histogram (50 bins) was calculated for the log - 10 transformed abundances of PaxDb proteins, and the probability density of PaxDb proteins was estimated by normalizing the values in each bin so that the total area of the histogram is 1.

[0230] To adjust the high - abundance tail of the complete sample abundance distribution g(x) to match the probability density of protein abundances > t in PaxDb, the scaling factor α was calculated:

Equation

[0231] A kernel density estimator, K, was fitted to the log10-transformed plasma abundance values ​​using a Gaussian kernel with σ = 0.2 and subtracted from the scaled semi-Gaussian distribution to estimate a function proportional to the density of the probability distribution over the imputed protein abundance: h(x) = αg(x) - K(x). 10 (A max )-12 and log10 abundance log 10 (A max The function h(x) was evaluated at 500 abundance values ​​equally spread in base 10 logarithmic space between x and y. All points where h(x) evaluated to be less than 0 were set to 0. A continuous probability distribution was fitted to this grid of sample points using linear interpolation, and then normalized so that the total probability of the distribution was 1. The abundances of 16,017 proteins in the UniProt database not represented in the processed PaxDb dataset were set to a random sample from the above distribution. The resulting abundances were converted to mole fraction estimates by dividing each abundance by the sum of all abundances. Protein Abundance (Cell Lines) Data Complement

[0232] Abundances were imputed for proteins in the human protein sequence database (see Protein Abundance Database Processing section above) that were not represented in the modeled cell line sample. This process resulted in a "complete" cell line sample containing 20,235 proteins, with a dynamic range of abundances of 10 orders of magnitude. The "complete" cell line sample was modeled as an adjusted skew-normal distribution for log10-transformed abundances: ·g(x)=2.45*skewnorm.pdf(x|a=-2.12,μ=4.5,σ=2.55) where skewnorm.pdf is the probability density function of the skew-normal distribution.

[0233] A kernel density estimator K (Gaussian kernel, σ = 0.2) was fitted to the log10 transformed abundances of all entries in the processed PaxDb database for cell line samples. Log10 abundance log 10 (A max )-10 and log10 abundance log 10 (A max The function h(x) was evaluated at 500 abundance values ​​equally spread in base 10 logarithmic space between x and y. All points where h(x) evaluated to be less than 0 were set to 0. A continuous probability distribution was fitted to this grid of sample points using linear interpolation, and then normalized so that the total probability of the distribution was 1. The abundances of 11,923 proteins in the Uniport database not represented in the processed PaxDb dataset were set to a random sample from the above distribution. The resulting abundances were converted to mole fraction estimates by dividing each abundance by the sum of all abundances. Depleted plasma samples

[0234] To model plasma samples in which the most abundant proteins had been depleted from the sample (e.g., using a commercially available affinity column), the abundances of the top 20 most abundant proteins in the imputed plasma sample (see Data Imputation for Protein Abundance (Plasma) section above) were reduced by 99% and the abundances were re-normalized to a sum of 1 to serve as estimates of molar fractions. Simulation of protein deposition

[0235] Abundance {a1,a2,a3,...a n The deposition of a sample containing n proteins, {{, n, n}}, onto an array was modeled as a multinomial distribution. Protein abundances were assigned a probability summing to 1,

number

[0236] For each sample type (cells, plasma, depleted plasma), binding was simulated for five technical replicate protein arrays. The 300 affinity reagents used for binding were targeted to the first 300 optimal targets (see section on selecting and evaluating optimal affinity reagent trimer targets above), using the binding model described in section on biosimilar affinity reagent models above, with a surface non-specific binding rate of 0.001. To simulate random variation in binding between replicates, the affinity reagent binding probability was perturbed for each replicate with an absolute deviation of 0.1 of the mean of the fraction, using the method described in section on applying noise to affinity reagent binding probabilities above. Binding to each flow cell was then simulated as described in section on simulating stochastic affinity reagent binding above. Decrypting Joint Data

[0237] Protein decoding was performed separately for each replicate as described in the Protein Decoding section above. The human FASTA sequence database (see Protein Sequence Database section above) was used to define the protein candidate sequences. The affinity reagent model used for decoding of all replicates was the original affinity reagent set referenced in the Simulation of Binding Data section above before the application of random noise. The decoding method assumed a surface nonspecific binding rate of 0.001. Determining probability thresholds for protein quantification

[0238] Given an identification probability threshold p t In the sample, the protein is tAt t, the probability threshold can be quantified by counting the number of identifications for that protein in the decoding output. However, setting the probability threshold too low can result in many false positive identifications, resulting in low quantitative specificity. Setting the probability threshold too high can result in false negative identifications, resulting in low quantitative sensitivity. For each replicate flow cell analyzed, the decoding results were processed with the following probability thresholds: log(p)=0, -1x10^(-20), -1x10^(-16), -1x10^-14, -1x10^-12, -1x10^-11, -1x10^-10, -1x10^-9, -1x10^-8, -1x10^-7, -1x10^-6, -1x10^-5, -1x10^-4, -1x10^-3, -1x10^-2, -0.1, -0.2 and -0.3. For each threshold evaluated: · For all unique proteins identified at least once in the dataset; - Calculate the number of identifications reported for that protein that were true positives (i.e., correct identifications) and false positives (i.e., spots incorrectly identified as that protein) -Calculate the specificity of the quantification for this protein;

number

[0239] The lowest threshold resulting in a nonspecific identification rate of less than 0.1% for all replicates analyzed was used for downstream quantitative analysis. Quantitative Statistics After setting thresholds by probability of identification, the following statistics were calculated for each analysis; The specificity of protein identification was calculated as described in the section Determination of probability thresholds for protein quantification above. · Proteins with at least one identification in a given replicate were considered “identified” in that replicate.

[0240] Proteome coverage for a replicate was the percentage of all proteins present in the sample that were identified at least once in that replicate. The number of counts for that protein in each replicate was used to calculate the quantification reproducibility (CV%) for a protein across replicates:

number

[0241] This example describes a Markov model that is useful for predicting non-binding probability for use in semi-censored decoding. Advantageously, the Markov model facilitates prediction of non-binding probability in a manner that considers the length of proteins in a given proteome but is independent of the variability of the amino acid sequences of those proteins. The Markov model is used to generate a set of pseudosequences for each unique protein length L in the proteome of interest. For each pseudosequence, the non-binding probability of the affinity reagent can be predicted, and the average or median non-binding prediction of the set of pseudosequences of length L can be used as the predicted semi-censored non-binding probability for candidate proteins with any amino acid sequence of the same length.

[0242] A Markov model can be characterized as a finite set of states with transition probabilities between these states. These transition probabilities depend only on the current state. The example model used is described by the following transition matrix, where a given row represents a potential current trimmer state, and the elements of that row represent the transition probability from the row's current state to the state represented by the column's label. Next state AAA, AAC, AAD, …, CYY, …, YYY AAA 0.1,0.4,0.0,…,0.0,…,0.0 AAC 0.0,0.0,0.0,…,0.0,…,0.0 Current state AAD 0.0,0.0,0.0,…,0.0,…,0.0 … … , … , … ,…, … ,…, … CYY 0.0,0.0,0.0,…,0.0,…,0.2 … … , … , … ,…, … ,…, … YYY 0.0,0.0,0.0,…,0.0,…,0.2

[0243] In a trimer parameterization of the Markov model, the first two amino acids of any valid next state must remain the last two amino acids of the current state, and therefore many state transitions are not possible and have a transition probability of zero. As an example, given a current state "AAA" as represented by row 1, a transition to state "CYY" is not possible because the last two amino acids "AA" of the current state are not maintained as the first two amino acids of the next state. Potentially valid transitions may also have a transition probability of zero if the training data does not contain such transitions. Purely as an example, a valid transition from "AAA" to "AAD" is shown as having a transition probability of zero. Samples may be generated from the Markov model by first probabilistically selecting an initial state and history. Further states are then determined by probabilistically selecting a next state based on the transition probability of the current state. This random walk may be terminated after a predetermined number of transitions.

[0244] For each state, a transition probability is learned based on observed transitions in the proteome. Sequences generated from this model mimic the sequence characteristics (e.g., amino acid composition) of the real proteome. The proteome can be decoded with reference to a first set of candidate proteins that contain natural amino acid sequences predicted to be present in the proteome. The pseudosequences are amino acid sequences that do not naturally occur in the proteome. Each pseudosequence has an amino acid sequence length identical to the unique sequence it represents in the set of candidate proteins. When pseudoproteins are used for uncensored decoding, the average of the predicted non-binding probabilities of the pseudoproteins (the uncensored non-binding probability is simply 1-the predicted binding probability) approximates the predicted non-binding probability of the "average" sequence - i.e., one that represents the amino acid composition of the proteome of interest.

[0245] As is evident from the above discussion, the non-binding probability can be determined in a strictly length-dependent manner, such that variability in amino acid sequence does not affect the calculation. Two proteins of the same length will always have the same non-binding likelihood for a given affinity reagent, using these methods.

[0246] Similar models can be constructed based on sequence regions other than trimers. For example, trimers can be replaced with monomers, dimers, tetramers or pentamers in the above model. As the length of the sequence region increases, the effectiveness of the model can improve if sufficient training data is available. For proteomes that are similar in size to or smaller than the human proteome, shorter lengths such as monomers, dimers and trimers may be preferred.

[0247] The Markov model was compared to a binning approach, which was performed as follows; essentially all proteins in the human proteome were aggregated into bins of proteins of similar length. Within each bin, the uncensored unbound likelihood was predicted for each protein (i.e., (1-P(bound|protein))). The median was used as the semi-censored unbound likelihood for the entire bin.

[0248] Figure 13 shows the predicted non-binding probability by sequence length for different semi-censored decoding approaches. The results show that when compared to using trimer-based probability adjustments, the fit of the Markov model-based approach outperforms the binning approach by lowering the R-squared value. The probability adjustments were determined from:

number

[0249] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided herein. Although the present invention has been described with reference to the above specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes and substitutions will occur to those skilled in the art without departing from the present invention. It is understood that various alternatives to the embodiments of the present invention described herein may be used in practicing the present invention. It is therefore contemplated that the present invention shall encompass any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the present invention, and that methods and structures within the scope of these claims and their equivalents are encompassed by the scope of the present invention.

Claims

1. 1. A method for identifying an existing protein, comprising: (a) providing an input to a computer processor, said input comprising: (i) a binding profile, the binding profile comprising a plurality of binding results for binding of the existing protein to a plurality of different affinity reagents, each binding result of the plurality of binding results comprising a measure of binding between the existing protein and a different affinity reagent of the plurality of different affinity reagents, the binding profile comprising positive binding results and negative binding results; (ii) a database containing information characterizing or identifying a plurality of candidate proteins; (iii) a binding model for each of the different affinity reagents; and providing a method for producing a medicament for a patient, comprising: (b) determining a probability for each of the affinity reagents that binds to the candidate proteins in the database according to the binding model, wherein said determining includes calculating a probability for the positive binding outcome and for the negative binding outcome, wherein the positive binding outcome is weighted more heavily than the negative binding outcome; (c) identifying the existing protein as a selected candidate protein, wherein the selected candidate protein is the candidate protein in the database that has a probability of binding each of the affinity reagents that best matches the binding profile for the existing protein; A method comprising:

2. 2. The method of claim 1, wherein the input further comprises: (iv) a non-specific binding rate comprising the probability of a non-specific binding event occurring for one or more of the different affinity reagents.

3. The method of claim 2 , wherein the non-specific binding event comprises binding of the one or more of the different affinity reagents to a substance other than a protein.

4. 4. The method of claim 3, wherein the substance is a solid support attached to the existing protein.

5. 3. The method of claim 2, wherein the non-specific binding event comprises binding of one or more of the different affinity reagents to an unexpected moiety in a protein.

6. The method of claim 5 , wherein the unexpected portion comprises a post-translational modification of a protein.

7. 2. The method of claim 1, wherein the calculating the probability of the positive binding result comprises determining the probability of a positive binding event occurring between each candidate protein in the plurality of candidate proteins and each of the affinity reagents.

8. The method of claim 7, wherein the probability of a positive binding event is normalized with respect to the length of the candidate protein.

9. 9. The method of claim 8, wherein the probability of a positive binding event is normalized using a binomial approximation, an exact Poisson binomial, or an estimated Poisson binomial.

10. 8. The method of claim 7, wherein the calculating the probability of a negative binding outcome comprises determining the probability of a negative binding event occurring between each candidate protein in the plurality of candidate proteins and each of the affinity reagents.

11. The method of claim 10, wherein the probability of a negative binding event is normalized with respect to the length of the candidate protein.

12. 12. The method of claim 11, wherein the probability of the negative binding event is normalized using a binomial approximation, an exact Poisson binomial, or an estimated Poisson binomial.

13. 8. The method of claim 7, wherein the calculating the probability of a negative binding outcome comprises determining the probability of a negative binding event occurring between each pseudo-protein in a plurality of pseudo-proteins and each of the affinity reagents.

14. 14. The method of claim 13, wherein the amino acid sequences in said plurality of pseudo-proteins have total lengths that are identical to the total lengths of the amino acid sequences in said plurality of candidate proteins.

15. The method of claim 14, wherein the plurality of pseudo-proteins lacks any full-length amino acid sequence present in the plurality of candidate proteins.

16. 15. The method of claim 14, wherein the plurality of pseudo-proteins lacks a portion of the full-length amino acid sequence present in the plurality of candidate proteins.

17. 14. The method of claim 13, wherein the amino acid sequences in the plurality of pseudo-proteins are generated by sampling amino acid sequences in the plurality of candidate proteins using Markov chains, generative adversarial networks, or length-based binning.

18. The method of claim 10 , wherein the binding model further comprises a function for determining the probability of a positive binding event occurring between an epitope in a candidate protein and each of the affinity reagents.

19. 19. The method of claim 18, wherein the function for determining the probability of a negative binding event occurring between an epitope in a candidate protein and each of the affinity reagents is independent of the function for determining the probability of a positive binding event occurring between an epitope in a candidate protein and each of the affinity reagents.

20. 19. The method of claim 18, wherein the probability of a negative binding event occurring between an epitope in the candidate protein and each of the affinity reagents is determined independently of the probability of a positive binding event occurring between an epitope in the candidate protein and each of the affinity reagents.

21. 10. The method of claim 1, further comprising determining the probability that the existing protein identified in step (c) is the selected candidate protein.

22. 22. The method of claim 21 , wherein the probability is the quotient of the probability of the selected candidate protein determined in step (b) divided by the sum of the probabilities determined in step (b) for all other candidate proteins in the database.

23. The method of claim 1 , wherein the selected candidate protein has the greatest probability of binding the affinity reagent that is consistent with the majority of the binding results in the binding profile.

24. The method of any one of claims 1 to 23, wherein the positive and negative binding results are represented by non-binary values ​​in the binding profile.

25. The method of claim 1 , wherein the information in step (a)(ii) comprises the primary sequence of the candidate protein.

26. The method of claim 1 , wherein the binding model comprises a function for determining the probability of a specific binding event occurring between a protein epitope and each of the affinity reagents.

27. 27. The method of claim 26, wherein the epitope consists essentially of an amino acid trimer.

28. The method of claim 1 , wherein the binding model comprises a function for determining the probability of a non-specific binding event occurring between a protein epitope and each of the affinity reagents.

29. 29. The method of claim 28, wherein the epitope consists essentially of an amino acid trimer.

30. 2. The method of claim 1, wherein the binding model comprises a function for determining the probability of a binding event occurring between each of the affinity reagents and a specific epitope for the respective affinity reagent and a biosimilar epitope.

31. 2. The method of claim 1, wherein step (b) comprises calculating a probability matrix containing the probability of a positive binding outcome for each of the affinity reagents binding to each of the candidate proteins in the database.

32. 32. The method of claim 31 , wherein step (b) further comprises calculating a probability matrix containing the probability of a negative binding outcome for each of the affinity reagents binding to each of the candidate proteins in the database.

33. 1. A method for identifying an existing protein, comprising: (a) contacting a plurality of different affinity reagents with a plurality of proteins present in a sample; (b) obtaining binding data from step (a), the binding data comprising a plurality of binding profiles, each of the binding profiles comprising a plurality of binding results for binding of the existing protein of step (a) to the plurality of different affinity reagents, each binding result of the plurality of binding results comprising a measure of binding between the existing protein of step (a) and a different affinity reagent of the plurality of different affinity reagents, each of the binding profiles comprising a positive binding result and a negative binding result; (c) providing a database containing information characterizing or identifying a plurality of candidate proteins; (d) providing a binding model for each of said different affinity reagents; (e) determining, for each of the affinity reagents, a probability of binding to each of the candidate proteins in the database according to the binding model, wherein said determining includes calculating a probability for the positive binding outcome and for the negative binding outcome, wherein the positive binding outcome is weighted more heavily than the negative binding outcome; (f) identifying the existing protein as a selected candidate protein, wherein the selected candidate protein is the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the plurality of binding results for the existing protein; A method comprising:

34. 34. The method of claim 33, further comprising providing a non-specific binding rate comprising the probability of a non-specific binding event occurring for one or more of the different affinity reagents.

35. 35. The method of claim 34, wherein the non-specific binding event comprises binding of one or more of the different affinity reagents to a solid support attached to the existing protein.

36. 36. The method of any one of claims 33 to 35, wherein the calculating the probability of the positive binding outcome comprises determining the probability of a positive binding event occurring between each candidate protein in the plurality of candidate proteins and each of the affinity reagents.

37. 37. The method of claim 36, wherein the probability of a positive binding event is normalized with respect to the length of the candidate protein.

38. 38. The method of claim 37, wherein the probability of a positive binding event is normalized using a binomial approximation, an exact Poisson binomial, or an estimated Poisson binomial.

39. 37. The method of claim 36, wherein the calculating the probability of a negative binding outcome comprises determining the probability of a negative binding event occurring between each candidate protein in the plurality of candidate proteins and each of the affinity reagents.

40. 40. The method of claim 39, wherein the probability of a negative binding event is normalized with respect to the length of the candidate protein.

41. 41. The method of claim 40, wherein the probability of the negative binding event is normalized using a binomial approximation, an exact Poisson binomial, or an estimated Poisson binomial.

42. 37. The method of claim 36, wherein the calculating the probability of a negative binding outcome comprises determining the probability of a negative binding event occurring between each pseudo-protein in a plurality of pseudo-proteins and each of the affinity reagents.

43. 43. The method of claim 42, wherein the amino acid sequences in said plurality of pseudo-proteins have total lengths that are identical to the total lengths of the amino acid sequences in said plurality of candidate proteins.

44. 44. The method of claim 43, wherein said plurality of pseudo-proteins lacks any full-length amino acid sequence present in said plurality of candidate proteins.

45. 44. The method of claim 43, wherein said plurality of pseudo-proteins lacks a portion of said full-length amino acid sequence present in said plurality of candidate proteins.

46. 43. The method of claim 42, wherein the amino acid sequences in the plurality of pseudo-proteins are generated by sampling the amino acid sequences in the plurality of candidate proteins using Markov chains, generative adversarial networks, or length-based binning.

47. 47. The method of claim 46, further comprising determining the probability that the existing protein identified in step (f) is the selected candidate protein.

48. 48. The method of claim 47, wherein the positive and negative binding results are represented by non-binary values ​​in the binding profile.

49. 49. The method of claim 48, wherein step (e) comprises calculating a probability matrix containing the probability of a positive binding outcome for each of the affinity reagents binding to each of the candidate proteins in the database.

50. 50. The method of claim 49, wherein step (e) further comprises calculating a probability matrix comprising the probability of a negative binding outcome for each of the affinity reagents binding to each of the candidate proteins in the database.

51. 1. A method for identifying an existing protein using a detection system, comprising: (a) acquiring signals from a plurality of binding reactions carried out in a detection system, the binding reactions comprising contacting a plurality of different affinity reagents with a plurality of proteins present in the sample; (b) processing the signal in the detection system to generate a plurality of binding profiles, each of the binding profiles comprising a plurality of binding results for binding of the existing protein of step (a) to the plurality of different affinity reagents, each binding result of the plurality of binding results comprising a measure of binding between the existing protein of step (a) and a different affinity reagent of the plurality of different affinity reagents, each of the binding profiles comprising a positive binding result and a negative binding result; (c) providing as input to the detection system a database containing information characterizing or identifying a plurality of candidate proteins; (d) providing a binding model for each of the different affinity reagents as an input to the detection system; (e) processing the plurality of binding profiles in the detection system to determine, for each of the affinity reagents, a probability of binding to each of the candidate proteins in the database according to the binding model; (f) outputting from the detection system the identification of a selected candidate protein, the selected candidate protein being the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the plurality of binding results for the existing protein; A method comprising:

52. 1. A detection system comprising: (a) a detector configured to acquire signals from a plurality of binding reactions occurring between a plurality of different affinity reagents and a plurality of proteins present in the sample; (b) a database containing information characterizing or identifying a plurality of candidate proteins; (c) a computer processor, (i) communicating with said database; (ii) processing the signals to generate a plurality of binding profiles, each of the binding profiles comprising a plurality of binding outcomes for binding of the present protein of (a) to the plurality of different affinity reagents, each binding outcome of the plurality of binding outcomes comprising a measure of binding between the present protein of (a) and a different affinity reagent of the plurality of different affinity reagents, each of the binding profiles comprising a positive binding outcome and a negative binding outcome; (iii) processing the binding profile to determine, for each of the affinity reagents, a probability of binding to each of the candidate proteins in the database according to a binding model for each of the affinity reagents; and (iv) outputting an identification of a selected candidate protein, the selected candidate protein being the candidate protein in the database having a probability of binding each of the affinity reagents that best matches the plurality of binding results for the existing protein; a computer processor configured to A detection system comprising: