Compositions and methods for polypeptide analysis
Patent Information
- Application Number
- JP2024537910
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-04
- Filing Date
- 2022-12-22
- Publication Date
- 2026-01-07
Abstract
Description
[Technical field]
[0001] The present invention relates to compositions and methods for the analysis of polypeptides. [Background technology]
[0002] Measurement of the proteome provides deep and valuable insights into important biological processes. In adjacent fields such as genomics, advances in DNA sequencing technology have proven extremely beneficial in improving understanding of the progression of complex human diseases. Applying similar approaches to proteomics has been difficult for several reasons, including the large number of different proteins and even larger number of proteoforms, the wide dynamic range of protein abundance in cells and biological fluids, and the inability to copy or amplify proteins. Therefore, improved approaches are needed. Summary of the Invention [Means for solving the problem]
[0003] Methods and systems for determining the chemical properties of polypeptides are generally described. In some embodiments, the present application provides a method for determining a chemical property of a polypeptide. In some embodiments, the method includes contacting the polypeptide with one or more amino acid recognizers. In certain embodiments, the one or more amino acid recognition factors constitute a first set of one or more amino acid recognition factors that bind to the polypeptide. In some embodiments, the method includes detecting a first series of signal pulses indicative of a first series of binding events between the first set of one or more amino acid recognition factors and the polypeptide. In some embodiments, the method includes determining at least one chemical property of a first set of at least two amino acids of the polypeptide based on at least one property of the first series of signal pulses.
[0004] In some embodiments, the present application provides a device comprising at least one processor and at least one non-transitory computer-readable storage medium having encoded instructions that, when executed by the at least one processor, cause the at least one processor to perform a method for determining a chemical property of a polypeptide. In some embodiments, the method comprises detecting a first series of signal pulses indicative of a first series of binding events between a first set of one or more amino acid recognition factors and the polypeptide. In some embodiments, the method comprises determining at least one chemical property of at least two amino acids of the polypeptide based on at least one property of the first series of signal pulses.
[0005] In some embodiments, the present application provides at least one non-transitory computer-readable storage medium having encoded instructions that, when executed by at least one processor, cause the at least one processor to perform a method for determining a chemical property of a polypeptide. In some embodiments, the method includes detecting a first series of signal pulses indicating a first series of binding events between a first set of one or more amino acid recognition factors and the polypeptide. In some embodiments, the method includes determining at least one chemical property of at least two amino acids of the polypeptide based on at least one property of the first series of signal pulses.
[0006] In some aspects, the present application provides a method comprising obtaining data during a degradation process of a polypeptide. In some embodiments, the method comprises analyzing the data to determine portions of the data, each portion corresponding to at least one amino acid of the polypeptide. In certain embodiments, at least a first portion of the data corresponds to a first amino acid and comprises a first plurality of signal pulses indicative of a series of binding events between a first type of amino acid recognition factor and the first amino acid. In certain embodiments, a second portion of the data corresponds to a second amino acid and does not comprise a signal pulse indicative of a binding event between any type of amino acid recognition factor and the second amino acid. In some embodiments, the method comprises determining at least one chemical property of the first amino acid and / or the second amino acid based on at least one property of the first portion of the data and at least one property of the second portion of the data. In some embodiments, at least one non-transitory computer-readable medium is provided having encoded instructions that, when executed by at least one process, cause at least one processor to perform the method. In some embodiments, a device is provided that includes at least one processor and at least one non-transitory computer-readable medium.
[0007] In some embodiments, the present application provides a method for determining a chemical characteristic of a polypeptide. In some embodiments, the method includes detecting a first series of signal pulses indicative of a first series of binding events between a first set of one or more amino acid recognition factors and the polypeptide. In some embodiments, the method includes determining at least one characteristic of the first series of signal pulses. In some embodiments, the method includes comparing at least one characteristic of the first series of signal pulses with known characteristics of a plurality of amino acid segments comprising at least two amino acids. In some embodiments, the method includes determining at least one chemical characteristic of at least two amino acids of the polypeptide based on the comparison. In some embodiments, at least one non-transitory computer-readable medium is provided having encoded instructions that, when executed by at least one process, cause at least one processor to perform the method. In some embodiments, a device is provided that includes at least one processor and at least one non-transitory computer-readable medium.
[0008] In some aspects, the present application provides a method, comprising obtaining data during the degradation process of a polypeptide. In some embodiments, the method comprises analyzing the data to determine at least three portions of data, each portion corresponding to an amino acid of the polypeptide and comprising a plurality of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and the amino acid. In some embodiments, the method comprises determining one or more properties of each of the at least three portions of data. In some embodiments, the method comprises identifying the polypeptide based on the order of the at least three portions of data and the one or more properties of each of the at least three portions of data. In some embodiments, at least one non-transitory computer-readable medium is provided having encoded instructions that, when executed by at least one process, cause at least one processor to perform the method. In some embodiments, a device is provided, comprising at least one processor and at least one non-transitory computer-readable medium.
[0009] In some embodiments, the present application provides a method for determining at least one chemical property of an amino acid of a polypeptide. In some embodiments, the method includes detecting a first series of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and a first amino acid of the polypeptide. In some embodiments, the method includes determining at least one chemical property of a second amino acid of the polypeptide based on at least one property of the first series of signal pulses. In some embodiments, at least one non-transitory computer-readable medium is provided having encoded instructions that, when executed by at least one process, cause at least one processor to perform the method. In some embodiments, a device is provided that includes at least one processor and at least one non-transitory computer-readable medium.
[0010] In some embodiments, the present application provides a method for determining at least one chemical property of an amino acid of a polypeptide. In some embodiments, the method includes detecting a first series of signal pulses indicative of a series of binding events between a first set of one or more amino acid recognition factors and a first amino acid of the polypeptide. In some embodiments, the method includes detecting a second series of signal pulses indicative of a series of binding events between a second set of one or more amino acid recognition factors and a second amino acid of the polypeptide. In some embodiments, the method includes determining at least one chemical property of a second amino acid of the polypeptide based on at least one characteristic of the first series of signal pulses and at least one characteristic of the second series of signal pulses. In some embodiments, at least one non-transitory computer-readable medium is provided having encoded instructions that, when executed by at least one process, cause at least one processor to perform the method. In some embodiments, a device is provided that includes at least one processor and at least one non-transitory computer-readable medium.
[0011] In some embodiments, the application provides a method for identifying a disease or disorder in a subject. In some embodiments, the method includes digesting proteins in a sample from a subject to produce a plurality of polypeptides. In some embodiments, the method includes contacting a polypeptide of the plurality of polypeptides with one or more amino acid recognition factors and a cleavage agent. In some embodiments, the method includes detecting one or more series of signal pulses indicative of a binding event between the one or more amino acid recognition factors and the polypeptide as amino acids are progressively cleaved from the end of the polypeptide by the cleavage agent. In some embodiments, the method includes determining at least one chemical property of the polypeptide based on at least one property of the one or more series of signal pulses. In certain embodiments, the at least one chemical property is indicative of a modification of the protein. In certain embodiments, the modification of the protein is indicative of a disease or disorder in the subject. In some embodiments, at least one non-transitory computer-readable medium is provided having encoded instructions that, when executed by at least one process, cause at least one processor to perform the method. In some embodiments, a device is provided that includes at least one processor and at least one non-transitory computer-readable medium.
[0012] The details of certain embodiments of the present disclosure are set forth in the detailed description. Other features, objects, and advantages of the invention will be apparent from the examples, drawings, and claims. [Brief description of the drawings]
[0013] [Figure 1]An exemplary overview of real-time dynamic protein sequencing is shown. Protein samples are digested into polypeptides, immobilized in a nanoscale reaction chamber, and incubated with a mixture of freely diffusing N-terminal amino acid (NAA) recognition factors and cleavage agents (e.g., aminopeptidases) that perform the sequencing process. Labeled recognition factors turn on and off binding to the polypeptide when one of their cognate NAA is exposed at the N-terminus, thereby generating a characteristic pulse pattern. The NAA is cleaved by the cleavage agent, exposing the next amino acid for recognition. The temporal order and binding kinetics of NAA recognition allow for polypeptide identification and are sensitive to features that modulate the binding kinetics, such as post-translational modifications (PTMs). [Figure 2A] Examples of NAA recognition and dynamic sequencing are shown. Figures 2A-C show example traces demonstrating single molecule N-terminal recognition by PS610 (Figure 2A), PS961 (Figure 2B), and PS691 (Figure 2C); scatter plots of the number of pulses per recognition segment (RS) versus RS mean pulse duration (PD) are shown in Figures 2A-C for each peptide, with the median PD shown. Figure 2D shows example traces from dynamic sequencing of the synthetic peptide FAAWAAYAAAADDD (SEQ ID NO: 813). The median PD is shown above each RS. Figures 2E-G show dynamic sequencing of the synthetic peptide LAQFASIAAYASDDD (SEQ ID NO: 793) with PS610 and PS961. Figure 2E shows example traces. Figure 2F shows scatter plots of RS mean PD versus bin ratio, showing discrimination of recognition factors by bin ratio and of NAA by pulse duration. Figure 2G shows a scatter plot of the number of pulses per RS versus the RS mean PD, grouped by the amino acid label assigned to the RS. [Figure 2B]Examples of NAA recognition and dynamic sequencing are shown. Figures 2A-C show example traces demonstrating single molecule N-terminal recognition by PS610 (Figure 2A), PS961 (Figure 2B), and PS691 (Figure 2C); scatter plots of the number of pulses per recognition segment (RS) versus RS mean pulse duration (PD) are shown in Figures 2A-C for each peptide, with the median PD shown. Figure 2D shows example traces from dynamic sequencing of the synthetic peptide FAAWAAYAAAADDD (SEQ ID NO: 813). The median PD is shown above each RS. Figures 2E-G show dynamic sequencing of the synthetic peptide LAQFASIAAYASDDD (SEQ ID NO: 793) with PS610 and PS961. Figure 2E shows example traces. Figure 2F shows scatter plots of RS mean PD versus bin ratio, showing discrimination of recognition factors by bin ratio and of NAA by pulse duration. Figure 2G shows a scatter plot of the number of pulses per RS versus the RS mean PD, grouped by the amino acid label assigned to the RS. [Figure 2C] Examples of NAA recognition and dynamic sequencing are shown. Figures 2A-C show example traces demonstrating single molecule N-terminal recognition by PS610 (Figure 2A), PS961 (Figure 2B), and PS691 (Figure 2C); scatter plots of the number of pulses per recognition segment (RS) versus RS mean pulse duration (PD) are shown in Figures 2A-C for each peptide, with the median PD shown. Figure 2D shows example traces from dynamic sequencing of the synthetic peptide FAAWAAYAAAADDD (SEQ ID NO: 813). The median PD is shown above each RS. Figures 2E-G show dynamic sequencing of the synthetic peptide LAQFASIAAYASDDD (SEQ ID NO: 793) with PS610 and PS961. Figure 2E shows example traces. Figure 2F shows scatter plots of RS mean PD versus bin ratio, showing discrimination of recognition factors by bin ratio and of NAA by pulse duration. Figure 2G shows a scatter plot of the number of pulses per RS versus the RS mean PD, grouped by the amino acid label assigned to the RS. [Figure 2D]Examples of NAA recognition and dynamic sequencing are shown. Figures 2A-C show example traces demonstrating single molecule N-terminal recognition by PS610 (Figure 2A), PS961 (Figure 2B), and PS691 (Figure 2C); scatter plots of the number of pulses per recognition segment (RS) versus RS mean pulse duration (PD) are shown in Figures 2A-C for each peptide, with the median PD shown. Figure 2D shows example traces from dynamic sequencing of the synthetic peptide FAAWAAYAAAADDD (SEQ ID NO: 813). The median PD is shown above each RS. Figures 2E-G show dynamic sequencing of the synthetic peptide LAQFASIAAYASDDD (SEQ ID NO: 793) with PS610 and PS961. Figure 2E shows example traces. Figure 2F shows scatter plots of RS mean PD versus bin ratio, showing discrimination of recognition factors by bin ratio and of NAA by pulse duration. Figure 2G shows a scatter plot of the number of pulses per RS versus the RS mean PD, grouped by the amino acid label assigned to the RS. [Figure 2E] Examples of NAA recognition and dynamic sequencing are shown. Figures 2A-C show example traces demonstrating single molecule N-terminal recognition by PS610 (Figure 2A), PS961 (Figure 2B), and PS691 (Figure 2C); scatter plots of the number of pulses per recognition segment (RS) versus RS mean pulse duration (PD) are shown in Figures 2A-C for each peptide, with the median PD shown. Figure 2D shows example traces from dynamic sequencing of the synthetic peptide FAAWAAYAAAADDD (SEQ ID NO: 813). The median PD is shown above each RS. Figures 2E-G show dynamic sequencing of the synthetic peptide LAQFASIAAYASDDD (SEQ ID NO: 793) with PS610 and PS961. Figure 2E shows example traces. Figure 2F shows scatter plots of RS mean PD versus bin ratio, showing discrimination of recognition factors by bin ratio and of NAA by pulse duration. Figure 2G shows a scatter plot of the number of pulses per RS versus the RS mean PD, grouped by the amino acid label assigned to the RS. [Figure 2F]Examples of NAA recognition and dynamic sequencing are shown. Figures 2A-C show example traces demonstrating single molecule N-terminal recognition by PS610 (Figure 2A), PS961 (Figure 2B), and PS691 (Figure 2C); scatter plots of the number of pulses per recognition segment (RS) versus RS mean pulse duration (PD) are shown in Figures 2A-C for each peptide, with the median PD shown. Figure 2D shows example traces from dynamic sequencing of the synthetic peptide FAAWAAYAAAADDD (SEQ ID NO: 813). The median PD is shown above each RS. Figures 2E-G show dynamic sequencing of the synthetic peptide LAQFASIAAYASDDD (SEQ ID NO: 793) with PS610 and PS961. Figure 2E shows example traces. Figure 2F shows scatter plots of RS mean PD versus bin ratio, showing discrimination of recognition factors by bin ratio and of NAA by pulse duration. Figure 2G shows a scatter plot of the number of pulses per RS versus the RS mean PD, grouped by the amino acid label assigned to the RS. [Figure 2G] Examples of NAA recognition and dynamic sequencing are shown. Figures 2A-C show example traces demonstrating single molecule N-terminal recognition by PS610 (Figure 2A), PS961 (Figure 2B), and PS691 (Figure 2C); scatter plots of the number of pulses per recognition segment (RS) versus RS mean pulse duration (PD) are shown in Figures 2A-C for each peptide, with the median PD shown. Figure 2D shows example traces from dynamic sequencing of the synthetic peptide FAAWAAYAAAADDD (SEQ ID NO: 813). The median PD is shown above each RS. Figures 2E-G show dynamic sequencing of the synthetic peptide LAQFASIAAYASDDD (SEQ ID NO: 793) with PS610 and PS961. Figure 2E shows example traces. Figure 2F shows scatter plots of RS mean PD versus bin ratio, showing discrimination of recognition factors by bin ratio and of NAA by pulse duration. Figure 2G shows a scatter plot of the number of pulses per RS versus the RS mean PD, grouped by the amino acid label assigned to the RS. [Figure 3A]Examples of dynamic sequencing of various peptides with high accuracy kinetic output are shown. Figures 3A-3E show dynamic sequencing of peptide DQQRLIFAG (SEQ ID NO: 794). Figure 3A shows an example trace of DQQRLIFAG (SEQ ID NO: 794). Figure 3B shows a scatter plot of RS average PD vs. bin ratio. Figure 3C shows further example traces of dynamic sequencing of DQQRLIFAG (SEQ ID NO: 794). Figure 3D shows the distribution of durations of each RS and non-recognized segment (NRS) obtained during sequencing, with the average duration shown. Figure 3E shows a kinetic signature plot summarizing the characteristic sequencing behavior of DQQRLIFAG (SEQ ID NO: 794) peptide. Figures 3F-G show dynamic sequencing of synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 795) (top), RLAFSALGAADDD (SEQ ID NO: 796) (middle), and EFIAWLV (SEQ ID NO: 797) (bottom). Figure 3F shows example traces for each peptide. Figure 3G shows the corresponding kinetic signature plots of DQQIASSRLAASFAAQQY (SEQ ID NO: 856), RLAFSAL (SEQ ID NO: 857), and EFIAWLV (SEQ ID NO: 797). [Figure 3B]Examples of dynamic sequencing of various peptides with high accuracy kinetic output are shown. Figures 3A-3E show dynamic sequencing of peptide DQQRLIFAG (SEQ ID NO: 794). Figure 3A shows an example trace of DQQRLIFAG (SEQ ID NO: 794). Figure 3B shows a scatter plot of RS average PD vs. bin ratio. Figure 3C shows further example traces of dynamic sequencing of DQQRLIFAG (SEQ ID NO: 794). Figure 3D shows the distribution of durations of each RS and non-recognized segment (NRS) obtained during sequencing, with the average duration shown. Figure 3E shows a kinetic signature plot summarizing the characteristic sequencing behavior of DQQRLIFAG (SEQ ID NO: 794) peptide. Figures 3F-G show dynamic sequencing of synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 795) (top), RLAFSALGAADDD (SEQ ID NO: 796) (middle), and EFIAWLV (SEQ ID NO: 797) (bottom). Figure 3F shows example traces for each peptide. Figure 3G shows the corresponding kinetic signature plots of DQQIASSRLAASFAAQQY (SEQ ID NO: 856), RLAFSAL (SEQ ID NO: 857), and EFIAWLV (SEQ ID NO: 797). [Figure 3C]Examples of dynamic sequencing of various peptides with high accuracy kinetic output are shown. Figures 3A-3E show dynamic sequencing of peptide DQQRLIFAG (SEQ ID NO: 794). Figure 3A shows an example trace of DQQRLIFAG (SEQ ID NO: 794). Figure 3B shows a scatter plot of RS average PD vs. bin ratio. Figure 3C shows further example traces of dynamic sequencing of DQQRLIFAG (SEQ ID NO: 794). Figure 3D shows the distribution of durations of each RS and non-recognized segment (NRS) obtained during sequencing, with the average duration shown. Figure 3E shows a kinetic signature plot summarizing the characteristic sequencing behavior of DQQRLIFAG (SEQ ID NO: 794) peptide. Figures 3F-G show dynamic sequencing of synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 795) (top), RLAFSALGAADDD (SEQ ID NO: 796) (middle), and EFIAWLV (SEQ ID NO: 797) (bottom). Figure 3F shows example traces for each peptide. Figure 3G shows the corresponding kinetic signature plots of DQQIASSRLAASFAAQQY (SEQ ID NO: 856), RLAFSAL (SEQ ID NO: 857), and EFIAWLV (SEQ ID NO: 797). [Figure 3D]Examples of dynamic sequencing of various peptides with high accuracy kinetic output are shown. Figures 3A-3E show dynamic sequencing of peptide DQQRLIFAG (SEQ ID NO: 794). Figure 3A shows an example trace of DQQRLIFAG (SEQ ID NO: 794). Figure 3B shows a scatter plot of RS average PD vs. bin ratio. Figure 3C shows further example traces of dynamic sequencing of DQQRLIFAG (SEQ ID NO: 794). Figure 3D shows the distribution of durations of each RS and non-recognized segment (NRS) obtained during sequencing, with the average duration shown. Figure 3E shows a kinetic signature plot summarizing the characteristic sequencing behavior of DQQRLIFAG (SEQ ID NO: 794) peptide. Figures 3F-G show dynamic sequencing of synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 795) (top), RLAFSALGAADDD (SEQ ID NO: 796) (middle), and EFIAWLV (SEQ ID NO: 797) (bottom). Figure 3F shows example traces for each peptide. Figure 3G shows the corresponding kinetic signature plots of DQQIASSRLAASFAAQQY (SEQ ID NO: 856), RLAFSAL (SEQ ID NO: 857), and EFIAWLV (SEQ ID NO: 797). [Figure 3E]Examples of dynamic sequencing of various peptides with high accuracy kinetic output are shown. Figures 3A-3E show dynamic sequencing of peptide DQQRLIFAG (SEQ ID NO: 794). Figure 3A shows an example trace of DQQRLIFAG (SEQ ID NO: 794). Figure 3B shows a scatter plot of RS average PD vs. bin ratio. Figure 3C shows further example traces of dynamic sequencing of DQQRLIFAG (SEQ ID NO: 794). Figure 3D shows the distribution of durations of each RS and non-recognized segment (NRS) obtained during sequencing, with the average duration shown. Figure 3E shows a kinetic signature plot summarizing the characteristic sequencing behavior of DQQRLIFAG (SEQ ID NO: 794) peptide. Figures 3F-G show dynamic sequencing of synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 795) (top), RLAFSALGAADDD (SEQ ID NO: 796) (middle), and EFIAWLV (SEQ ID NO: 797) (bottom). Figure 3F shows example traces for each peptide. Figure 3G shows the corresponding kinetic signature plots of DQQIASSRLAASFAAQQY (SEQ ID NO: 856), RLAFSAL (SEQ ID NO: 857), and EFIAWLV (SEQ ID NO: 797). [Figure 3F]Examples of dynamic sequencing of various peptides with high accuracy kinetic output are shown. Figures 3A-3E show dynamic sequencing of peptide DQQRLIFAG (SEQ ID NO: 794). Figure 3A shows an example trace of DQQRLIFAG (SEQ ID NO: 794). Figure 3B shows a scatter plot of RS average PD vs. bin ratio. Figure 3C shows further example traces of dynamic sequencing of DQQRLIFAG (SEQ ID NO: 794). Figure 3D shows the distribution of durations of each RS and non-recognized segment (NRS) obtained during sequencing, with the average duration shown. Figure 3E shows a kinetic signature plot summarizing the characteristic sequencing behavior of DQQRLIFAG (SEQ ID NO: 794) peptide. Figures 3F-G show dynamic sequencing of synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 795) (top), RLAFSALGAADDD (SEQ ID NO: 796) (middle), and EFIAWLV (SEQ ID NO: 797) (bottom). Figure 3F shows example traces for each peptide. Figure 3G shows the corresponding kinetic signature plots of DQQIASSRLAASFAAQQY (SEQ ID NO: 856), RLAFSAL (SEQ ID NO: 857), and EFIAWLV (SEQ ID NO: 797). [Figure 3G]Examples of dynamic sequencing of various peptides with high accuracy kinetic output are shown. Figures 3A-3E show dynamic sequencing of peptide DQQRLIFAG (SEQ ID NO: 794). Figure 3A shows an example trace of DQQRLIFAG (SEQ ID NO: 794). Figure 3B shows a scatter plot of RS average PD vs. bin ratio. Figure 3C shows further example traces of dynamic sequencing of DQQRLIFAG (SEQ ID NO: 794). Figure 3D shows the distribution of durations of each RS and non-recognized segment (NRS) obtained during sequencing, with the average duration shown. Figure 3E shows a kinetic signature plot summarizing the characteristic sequencing behavior of DQQRLIFAG (SEQ ID NO: 794) peptide. Figures 3F-G show dynamic sequencing of synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 795) (top), RLAFSALGAADDD (SEQ ID NO: 796) (middle), and EFIAWLV (SEQ ID NO: 797) (bottom). Figure 3F shows example traces for each peptide. Figure 3G shows the corresponding kinetic signature plots of DQQIASSRLAASFAAQQY (SEQ ID NO: 856), RLAFSAL (SEQ ID NO: 857), and EFIAWLV (SEQ ID NO: 797). [Figure 4A]Examples of detection of single amino acid changes and PTMs are shown. Figures 4A-B show dynamic sequencing of synthetic peptides differing in a single amino acid: RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figure 4A shows an example trace. Figure 4B shows a scatter plot of RS mean PD vs. bin ratio. Figures 4C-D show detection of oxidized methionine with peptide RLMFAYPDDD (SEQ ID NO: 801). Figure 4C shows the distribution of mean PD for leucine. Labels indicate populations of leucine followed by methionine (LM) or methionine sulfoxide (LMo). Figure 4D shows examples of traces where methionine is recognized by PS961 and leucine is long PD (top, RLMFAYPDDD (SEQ ID NO: 801)) or where methionine is not recognized by oxidation and leucine is short PD (bottom, RLMoFAYPDDD (SEQ ID NO: 858)). Figure 4E shows scatter plots of RS mean PD vs. bin ratio for runs where oxidation was uncontrolled (top) or where methionine was fully oxidized (bottom). [Figure 4B] Examples of detection of single amino acid changes and PTMs are shown. Figures 4A-B show dynamic sequencing of synthetic peptides differing in a single amino acid: RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figure 4A shows an example trace. Figure 4B shows a scatter plot of RS mean PD vs. bin ratio. Figures 4C-D show detection of oxidized methionine with peptide RLMFAYPDDD (SEQ ID NO: 801). Figure 4C shows the distribution of mean PD for leucine. Labels indicate populations of leucine followed by methionine (LM) or methionine sulfoxide (LMo). Figure 4D shows examples of traces where methionine is recognized by PS961 and leucine is long PD (top, RLMFAYPDDD (SEQ ID NO: 801)) or where methionine is not recognized by oxidation and leucine is short PD (bottom, RLMoFAYPDDD (SEQ ID NO: 858)). Figure 4E shows scatter plots of RS mean PD vs. bin ratio for runs where oxidation was uncontrolled (top) or where methionine was fully oxidized (bottom). [Figure 4C] Examples of detection of single amino acid changes and PTMs are shown. Figures 4A-B show dynamic sequencing of synthetic peptides differing in a single amino acid: RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figure 4A shows an example trace. Figure 4B shows a scatter plot of RS mean PD vs. bin ratio. Figures 4C-D show detection of oxidized methionine with peptide RLMFAYPDDD (SEQ ID NO: 801). Figure 4C shows the distribution of mean PD for leucine. Labels indicate populations of leucine followed by methionine (LM) or methionine sulfoxide (LMo). Figure 4D shows examples of traces where methionine is recognized by PS961 and leucine is long PD (top, RLMFAYPDDD (SEQ ID NO: 801)) or where methionine is not recognized by oxidation and leucine is short PD (bottom, RLMoFAYPDDD (SEQ ID NO: 858)). Figure 4E shows scatter plots of RS mean PD vs. bin ratio for runs where oxidation was uncontrolled (top) or where methionine was fully oxidized (bottom). [Figure 4D]Examples of detection of single amino acid changes and PTMs are shown. Figures 4A-B show dynamic sequencing of synthetic peptides differing in a single amino acid: RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figure 4A shows an example trace. Figure 4B shows a scatter plot of RS mean PD vs. bin ratio. Figures 4C-D show detection of oxidized methionine with peptide RLMFAYPDDD (SEQ ID NO: 801). Figure 4C shows the distribution of mean PD for leucine. Labels indicate populations of leucine followed by methionine (LM) or methionine sulfoxide (LMo). Figure 4D shows examples of traces where methionine is recognized by PS961 and leucine is long PD (top, RLMFAYPDDD (SEQ ID NO: 801)) or where methionine is not recognized by oxidation and leucine is short PD (bottom, RLMoFAYPDDD (SEQ ID NO: 858)). Figure 4E shows scatter plots of RS mean PD vs. bin ratio for runs where oxidation was uncontrolled (top) or where methionine was fully oxidized (bottom). [Figure 4E] Examples of detection of single amino acid changes and PTMs are shown. Figures 4A-B show dynamic sequencing of synthetic peptides differing in a single amino acid: RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figure 4A shows an example trace. Figure 4B shows a scatter plot of RS mean PD vs. bin ratio. Figures 4C-D show detection of oxidized methionine with peptide RLMFAYPDDD (SEQ ID NO: 801). Figure 4C shows the distribution of mean PD for leucine. Labels indicate populations of leucine followed by methionine (LM) or methionine sulfoxide (LMo). Figure 4D shows examples of traces where methionine is recognized by PS961 and leucine is long PD (top, RLMFAYPDDD (SEQ ID NO: 801)) or where methionine is not recognized by oxidation and leucine is short PD (bottom, RLMoFAYPDDD (SEQ ID NO: 858)). Figure 4E shows scatter plots of RS mean PD vs. bin ratio for runs where oxidation was uncontrolled (top) or where methionine was fully oxidized (bottom). [Figure 5A] Examples of peptide discrimination in a mixture and mapping of peptides to the human proteome are shown. Figure 5A shows example traces from sequencing a mixture of peptides DQQRLIFAG (SEQ ID NO: 794) and RLAFSALGAADDD (SEQ ID NO: 796) on the same chip. The chip window shows the position of the reaction chambers that generate the sequencing readout for each peptide. Figure 5B shows example traces from dynamic sequencing of two peptides, DQQRLIFAGK (SEQ ID NO: 802) (top) and EFIAWLVK (SEQ ID NO: 803) (bottom), isolated from recombinant human proteins ubiquitin and GLP-1, respectively. Figure 5C shows a diagram showing the identification of the protein ubiquitin as a match to the kinetic signature from the DQQRLIFAGK (SEQ ID NO: 802) peptide in an in silico digestion of the human proteome based on kinetic information. Sequence numbers 804 (IVNFSRLIFHHLK), 805 (DIRLIFSNAK), 806 (GQSRLIFTYGLTNSGK), 807 (DQQRLLIFAGK), and 808 (DEHCLRLIFLK). [Figure 5B]Examples of peptide discrimination in a mixture and mapping of peptides to the human proteome are shown. Figure 5A shows example traces from sequencing a mixture of peptides DQQRLIFAG (SEQ ID NO: 794) and RLAFSALGAADDD (SEQ ID NO: 796) on the same chip. The chip window shows the position of the reaction chambers that generate the sequencing readout for each peptide. Figure 5B shows example traces from dynamic sequencing of two peptides, DQQRLIFAGK (SEQ ID NO: 802) (top) and EFIAWLVK (SEQ ID NO: 803) (bottom), isolated from recombinant human proteins ubiquitin and GLP-1, respectively. Figure 5C shows a diagram showing the identification of the protein ubiquitin as a match for the kinetic signature from the DQQRLIFAGK (SEQ ID NO: 802) peptide in an in silico digestion of the human proteome based on kinetic information. Sequence numbers 804 (IVNFSRLIFHHLK), 805 (DIRLIFSNAK), 806 (GQSRLIFTYGLTNSGK), 807 (DQQRLLIFAGK), and 808 (DEHCLRLIFLK). [Figure 5C] Examples of peptide discrimination in a mixture and mapping of peptides to the human proteome are shown. Figure 5A shows example traces from sequencing a mixture of peptides DQQRLIFAG (SEQ ID NO: 794) and RLAFSALGAADDD (SEQ ID NO: 796) on the same chip. The chip window shows the position of the reaction chambers that generate the sequencing readout for each peptide. Figure 5B shows example traces from dynamic sequencing of two peptides, DQQRLIFAGK (SEQ ID NO: 802) (top) and EFIAWLVK (SEQ ID NO: 803) (bottom), isolated from recombinant human proteins ubiquitin and GLP-1, respectively. Figure 5C shows a diagram showing the identification of the protein ubiquitin as a match to the kinetic signature from the DQQRLIFAGK (SEQ ID NO: 802) peptide in an in silico digestion of the human proteome based on kinetic information. Sequence numbers 804 (IVNFSRLIFHHLK), 805 (DIRLIFSNAK), 806 (GQSRLIFTYGLTNSGK), 807 (DQQRLLIFAGK), and 808 (DEHCLRLIFLK). [Figure 6A] Examples of chip operation are shown. Figure 6A shows an exploded view of a custom semiconductor chip and a small benchtop instrument designed to support protein sequencing assays. Figure 6B shows the chip achieves electronic rejection by discarding photoelectrons from a pulsed laser, then shifts to collect fluorescent photoelectrons from bound NAA recognition factors. The timing of the rejection and collection windows cycles between two modes (bin 1 and bin 0, example waveforms shown) in alternating frames to provide bin ratio estimates of the dye's fluorescence lifetime. Figure 6C shows the chip achieves a >10,000-fold attenuation of the incident laser light within 1 ns of the start of the rejection mode. Figure 6D shows example pulses for dyes with short and long fluorescence lifetimes, showing the difference in signal collection in bin 0 and bin 1. Figure 6E shows the distribution of average RS bin ratios collected for three dyes with different fluorescence lifetimes. Figure 6F shows that dye channel identification accuracy improves with the number of pulses captured per RS. [Figure 6B] Examples of chip operation are shown. Figure 6A shows an exploded view of a custom semiconductor chip and a small benchtop instrument designed to support protein sequencing assays. Figure 6B shows the chip achieves electronic rejection by discarding photoelectrons from a pulsed laser, then shifts to collect fluorescent photoelectrons from bound NAA recognition factors. The timing of the rejection and collection windows cycles between two modes (bin 1 and bin 0, example waveforms shown) in alternating frames to provide bin ratio estimates of the dye's fluorescence lifetime. Figure 6C shows the chip achieves a >10,000-fold attenuation of the incident laser light within 1 ns of the start of the rejection mode. Figure 6D shows example pulses for dyes with short and long fluorescence lifetimes, showing the difference in signal collection in bin 0 and bin 1. Figure 6E shows the distribution of average RS bin ratios collected for three dyes with different fluorescence lifetimes. Figure 6F shows that dye channel identification accuracy improves with the number of pulses captured per RS. [Figure 6C]Examples of chip operation are shown. Figure 6A shows an exploded view of a custom semiconductor chip and a small benchtop instrument designed to support protein sequencing assays. Figure 6B shows the chip achieves electronic rejection by discarding photoelectrons from a pulsed laser, then shifts to collect fluorescent photoelectrons from bound NAA recognition factors. The timing of the rejection and collection windows cycles between two modes (bin 1 and bin 0, example waveforms shown) in alternating frames to provide bin ratio estimates of the dye's fluorescence lifetime. Figure 6C shows the chip achieves a >10,000-fold attenuation of the incident laser light within 1 ns of the start of the rejection mode. Figure 6D shows example pulses for dyes with short and long fluorescence lifetimes, showing the difference in signal collection in bin 0 and bin 1. Figure 6E shows the distribution of average RS bin ratios collected for three dyes with different fluorescence lifetimes. Figure 6F shows that dye channel identification accuracy improves with the number of pulses captured per RS. [Figure 6D] Examples of chip operation are shown. Figure 6A shows an exploded view of a custom semiconductor chip and a small benchtop instrument designed to support protein sequencing assays. Figure 6B shows the chip achieves electronic rejection by discarding photoelectrons from a pulsed laser, then shifts to collect fluorescent photoelectrons from bound NAA recognition factors. The timing of the rejection and collection windows cycles between two modes (bin 1 and bin 0, example waveforms shown) in alternating frames to provide bin ratio estimates of the dye's fluorescence lifetime. Figure 6C shows the chip achieves a >10,000-fold attenuation of the incident laser light within 1 ns of the start of the rejection mode. Figure 6D shows example pulses for dyes with short and long fluorescence lifetimes, showing the difference in signal collection in bin 0 and bin 1. Figure 6E shows the distribution of average RS bin ratios collected for three dyes with different fluorescence lifetimes. Figure 6F shows that dye channel identification accuracy improves with the number of pulses captured per RS. [Figure 6E]Examples of chip operation are shown. Figure 6A shows an exploded view of a custom semiconductor chip and a small benchtop instrument designed to support protein sequencing assays. Figure 6B shows the chip achieves electronic rejection by discarding photoelectrons from a pulsed laser, then shifts to collect fluorescent photoelectrons from bound NAA recognition factors. The timing of the rejection and collection windows cycles between two modes (bin 1 and bin 0, example waveforms shown) in alternating frames to provide bin ratio estimates of the dye's fluorescence lifetime. Figure 6C shows the chip achieves a >10,000-fold attenuation of the incident laser light within 1 ns of the start of the rejection mode. Figure 6D shows example pulses for dyes with short and long fluorescence lifetimes, showing the difference in signal collection in bin 0 and bin 1. Figure 6E shows the distribution of average RS bin ratios collected for three dyes with different fluorescence lifetimes. Figure 6F shows that dye channel identification accuracy improves with the number of pulses captured per RS. [Figure 6F] Examples of chip operation are shown. Figure 6A shows an exploded view of a custom semiconductor chip and a small benchtop instrument designed to support protein sequencing assays. Figure 6B shows the chip achieves electronic rejection by discarding photoelectrons from a pulsed laser, then shifts to collect fluorescent photoelectrons from bound NAA recognition factors. The timing of the rejection and collection windows cycles between two modes (bin 1 and bin 0, example waveforms shown) in alternating frames to provide bin ratio estimates of the dye's fluorescence lifetime. Figure 6C shows the chip achieves a >10,000-fold attenuation of the incident laser light within 1 ns of the start of the rejection mode. Figure 6D shows example pulses for dyes with short and long fluorescence lifetimes, showing the difference in signal collection in bin 0 and bin 1. Figure 6E shows the distribution of average RS bin ratios collected for three dyes with different fluorescence lifetimes. Figure 6F shows that dye channel identification accuracy improves with the number of pulses captured per RS. [Figure 7A]Examples of recognition factor properties are shown. Figures 7A-7E show recognition factor kinetic characterization using a polarization assay (Example 1, Methods). Figures 7A-7B show the affinity (KD) (Figure 7A) and off-rates (koff) (Figure 7B) of PS610 for peptides with N-terminal phenylalanine, tyrosine, and tryptophan. SEQ ID NOs: 809 (FAKLK(FITC)DEESILKQ), 810 (YAKLK(FITC)DEESILKQ), and 811 (WAKLK(FITC)DEESILKQ) are shown in Figure 7B. Figure 7C shows the affinity of PS961 for peptides with N-terminal leucine, isoleucine, and valine. Figures 7D-7E show the affinity of PS691 for peptides with N-terminal arginine (Figure 7D) and single point polarization data measured for peptides with N-terminal arginine, lysine, and histidine using 2000 nM PS691 (Figure 7E). Figure 7F shows the binding energies calculated using a computer model (Example 1, Methods) for peptides with initial sequences LAX and LXA (where X = all 20 amino acids). Box plots show the fraction of the total binding energy contributed by amino acids at positions 1 (P1), 2 (P2), and 3 (P3), which tend to decrease exponentially from P1 to P3 (R2 > 0.97). Figure 7G shows the RS average PD determined in single molecule assays for LXA and LAX peptides using PS961 and for FXA and FAX peptides using PS610. Figure 7H shows that the non-polar solvation energy terms from the computational binding model with PS961 show high correlation with the actual RS average PD values observed in single molecule assays with peptides containing an N-terminal leucine and various amino acids at the P2 position. Peptides LVFA (SEQ ID NO: 859), LIFA (SEQ ID NO: 860), LVAR (SEQ ID NO: 861), LAFA (SEQ ID NO: 862), LQAR (SEQ ID NO: 863), LDAA (SEQ ID NO: 864), LCAR (SEQ ID NO: 865), LGAA (SEQ ID NO: 866), LMFA (SEQ ID NO: 867), LSAR (SEQ ID NO: 868), and LEFA (SEQ ID NO: 869) are shown. [Figure 7B]Examples of recognition factor properties are shown. Figures 7A-7E show recognition factor kinetic characterization using a polarization assay (Example 1, Methods). Figures 7A-7B show the affinity (KD) (Figure 7A) and off-rates (koff) (Figure 7B) of PS610 for peptides with N-terminal phenylalanine, tyrosine, and tryptophan. SEQ ID NOs: 809 (FAKLK(FITC)DEESILKQ), 810 (YAKLK(FITC)DEESILKQ), and 811 (WAKLK(FITC)DEESILKQ) are shown in Figure 7B. Figure 7C shows the affinity of PS961 for peptides with N-terminal leucine, isoleucine, and valine. Figures 7D-7E show the affinity of PS691 for peptides with N-terminal arginine (Figure 7D) and single point polarization data measured for peptides with N-terminal arginine, lysine, and histidine using 2000 nM PS691 (Figure 7E). Figure 7F shows the binding energies calculated using a computer model (Example 1, Methods) for peptides with initial sequences LAX and LXA (where X = all 20 amino acids). Box plots show the fraction of the total binding energy contributed by amino acids at positions 1 (P1), 2 (P2), and 3 (P3), which tend to decrease exponentially from P1 to P3 (R2 > 0.97). Figure 7G shows the RS average PD determined in single molecule assays for LXA and LAX peptides using PS961 and for FXA and FAX peptides using PS610. Figure 7H shows that the non-polar solvation energy terms from the computational binding model with PS961 show high correlation with the actual RS average PD values observed in single molecule assays with peptides containing an N-terminal leucine and various amino acids at the P2 position. Peptides LVFA (SEQ ID NO: 859), LIFA (SEQ ID NO: 860), LVAR (SEQ ID NO: 861), LAFA (SEQ ID NO: 862), LQAR (SEQ ID NO: 863), LDAA (SEQ ID NO: 864), LCAR (SEQ ID NO: 865), LGAA (SEQ ID NO: 866), LMFA (SEQ ID NO: 867), LSAR (SEQ ID NO: 868), and LEFA (SEQ ID NO: 869) are shown. [Figure 7C]Examples of recognition factor properties are shown. Figures 7A-7E show recognition factor kinetic characterization using a polarization assay (Example 1, Methods). Figures 7A-7B show the affinity (KD) (Figure 7A) and off-rates (koff) (Figure 7B) of PS610 for peptides with N-terminal phenylalanine, tyrosine, and tryptophan. SEQ ID NOs: 809 (FAKLK(FITC)DEESILKQ), 810 (YAKLK(FITC)DEESILKQ), and 811 (WAKLK(FITC)DEESILKQ) are shown in Figure 7B. Figure 7C shows the affinity of PS961 for peptides with N-terminal leucine, isoleucine, and valine. Figures 7D-7E show the affinity of PS691 for peptides with N-terminal arginine (Figure 7D) and single point polarization data measured for peptides with N-terminal arginine, lysine, and histidine using 2000 nM PS691 (Figure 7E). Figure 7F shows the binding energies calculated using a computer model (Example 1, Methods) for peptides with initial sequences LAX and LXA (where X = all 20 amino acids). Box plots show the fraction of the total binding energy contributed by amino acids at positions 1 (P1), 2 (P2), and 3 (P3), which tend to decrease exponentially from P1 to P3 (R2 > 0.97). Figure 7G shows the RS average PD determined in single molecule assays for LXA and LAX peptides using PS961 and for FXA and FAX peptides using PS610. Figure 7H shows that the non-polar solvation energy terms from the computational binding model with PS961 show high correlation with the actual RS average PD values observed in single molecule assays with peptides containing an N-terminal leucine and various amino acids at the P2 position. Peptides LVFA (SEQ ID NO: 859), LIFA (SEQ ID NO: 860), LVAR (SEQ ID NO: 861), LAFA (SEQ ID NO: 862), LQAR (SEQ ID NO: 863), LDAA (SEQ ID NO: 864), LCAR (SEQ ID NO: 865), LGAA (SEQ ID NO: 866), LMFA (SEQ ID NO: 867), LSAR (SEQ ID NO: 868), and LEFA (SEQ ID NO: 869) are shown. [Figure 7D]Examples of recognition factor properties are shown. Figures 7A-7E show recognition factor kinetic characterization using a polarization assay (Example 1, Methods). Figures 7A-7B show the affinity (KD) (Figure 7A) and off-rates (koff) (Figure 7B) of PS610 for peptides with N-terminal phenylalanine, tyrosine, and tryptophan. SEQ ID NOs: 809 (FAKLK(FITC)DEESILKQ), 810 (YAKLK(FITC)DEESILKQ), and 811 (WAKLK(FITC)DEESILKQ) are shown in Figure 7B. Figure 7C shows the affinity of PS961 for peptides with N-terminal leucine, isoleucine, and valine. Figures 7D-7E show the affinity of PS691 for peptides with N-terminal arginine (Figure 7D) and single point polarization data measured for peptides with N-terminal arginine, lysine, and histidine using 2000 nM PS691 (Figure 7E). Figure 7F shows the binding energies calculated using a computer model (Example 1, Methods) for peptides with initial sequences LAX and LXA (where X = all 20 amino acids). Box plots show the fraction of the total binding energy contributed by amino acids at positions 1 (P1), 2 (P2), and 3 (P3), which tend to decrease exponentially from P1 to P3 (R2 > 0.97). Figure 7G shows the RS average PD determined in single molecule assays for LXA and LAX peptides using PS961 and for FXA and FAX peptides using PS610. Figure 7H shows that the non-polar solvation energy terms from the computational binding model with PS961 show high correlation with the actual RS average PD values observed in single molecule assays with peptides containing an N-terminal leucine and various amino acids at the P2 position. Peptides LVFA (SEQ ID NO: 859), LIFA (SEQ ID NO: 860), LVAR (SEQ ID NO: 861), LAFA (SEQ ID NO: 862), LQAR (SEQ ID NO: 863), LDAA (SEQ ID NO: 864), LCAR (SEQ ID NO: 865), LGAA (SEQ ID NO: 866), LMFA (SEQ ID NO: 867), LSAR (SEQ ID NO: 868), and LEFA (SEQ ID NO: 869) are shown. [Figure 7E]Examples of recognition factor properties are shown. Figures 7A-7E show recognition factor kinetic characterization using a polarization assay (Example 1, Methods). Figures 7A-7B show the affinity (KD) (Figure 7A) and off-rates (koff) (Figure 7B) of PS610 for peptides with N-terminal phenylalanine, tyrosine, and tryptophan. SEQ ID NOs: 809 (FAKLK(FITC)DEESILKQ), 810 (YAKLK(FITC)DEESILKQ), and 811 (WAKLK(FITC)DEESILKQ) are shown in Figure 7B. Figure 7C shows the affinity of PS961 for peptides with N-terminal leucine, isoleucine, and valine. Figures 7D-7E show the affinity of PS691 for peptides with N-terminal arginine (Figure 7D) and single point polarization data measured for peptides with N-terminal arginine, lysine, and histidine using 2000 nM PS691 (Figure 7E). Figure 7F shows the binding energies calculated using a computer model (Example 1, Methods) for peptides with initial sequences LAX and LXA (where X = all 20 amino acids). Box plots show the fraction of the total binding energy contributed by amino acids at positions 1 (P1), 2 (P2), and 3 (P3), which tend to decrease exponentially from P1 to P3 (R2 > 0.97). Figure 7G shows the RS average PD determined in single molecule assays for LXA and LAX peptides using PS961 and for FXA and FAX peptides using PS610. Figure 7H shows that the non-polar solvation energy terms from the computational binding model with PS961 show high correlation with the actual RS average PD values observed in single molecule assays with peptides containing an N-terminal leucine and various amino acids at the P2 position. Peptides LVFA (SEQ ID NO: 859), LIFA (SEQ ID NO: 860), LVAR (SEQ ID NO: 861), LAFA (SEQ ID NO: 862), LQAR (SEQ ID NO: 863), LDAA (SEQ ID NO: 864), LCAR (SEQ ID NO: 865), LGAA (SEQ ID NO: 866), LMFA (SEQ ID NO: 867), LSAR (SEQ ID NO: 868), and LEFA (SEQ ID NO: 869) are shown. [Figure 7F]Examples of recognition factor properties are shown. Figures 7A-7E show recognition factor kinetic characterization using a polarization assay (Example 1, Methods). Figures 7A-7B show the affinity (KD) (Figure 7A) and off-rates (koff) (Figure 7B) of PS610 for peptides with N-terminal phenylalanine, tyrosine, and tryptophan. SEQ ID NOs: 809 (FAKLK(FITC)DEESILKQ), 810 (YAKLK(FITC)DEESILKQ), and 811 (WAKLK(FITC)DEESILKQ) are shown in Figure 7B. Figure 7C shows the affinity of PS961 for peptides with N-terminal leucine, isoleucine, and valine. Figures 7D-7E show the affinity of PS691 for peptides with N-terminal arginine (Figure 7D) and single point polarization data measured for peptides with N-terminal arginine, lysine, and histidine using 2000 nM PS691 (Figure 7E). Figure 7F shows the binding energies calculated using a computer model (Example 1, Methods) for peptides with initial sequences LAX and LXA (where X = all 20 amino acids). Box plots show the fraction of the total binding energy contributed by amino acids at positions 1 (P1), 2 (P2), and 3 (P3), which tend to decrease exponentially from P1 to P3 (R2 > 0.97). Figure 7G shows the RS average PD determined in single molecule assays for LXA and LAX peptides using PS961 and for FXA and FAX peptides using PS610. Figure 7H shows that the non-polar solvation energy terms from the computational binding model with PS961 show high correlation with the actual RS average PD values observed in single molecule assays with peptides containing an N-terminal leucine and various amino acids at the P2 position. Peptides LVFA (SEQ ID NO: 859), LIFA (SEQ ID NO: 860), LVAR (SEQ ID NO: 861), LAFA (SEQ ID NO: 862), LQAR (SEQ ID NO: 863), LDAA (SEQ ID NO: 864), LCAR (SEQ ID NO: 865), LGAA (SEQ ID NO: 866), LMFA (SEQ ID NO: 867), LSAR (SEQ ID NO: 868), and LEFA (SEQ ID NO: 869) are shown. [Figure 7G]Examples of recognition factor properties are shown. Figures 7A-7E show recognition factor kinetic characterization using a polarization assay (Example 1, Methods). Figures 7A-7B show the affinity (KD) (Figure 7A) and off-rates (koff) (Figure 7B) of PS610 for peptides with N-terminal phenylalanine, tyrosine, and tryptophan. SEQ ID NOs: 809 (FAKLK(FITC)DEESILKQ), 810 (YAKLK(FITC)DEESILKQ), and 811 (WAKLK(FITC)DEESILKQ) are shown in Figure 7B. Figure 7C shows the affinity of PS961 for peptides with N-terminal leucine, isoleucine, and valine. Figures 7D-7E show the affinity of PS691 for peptides with N-terminal arginine (Figure 7D) and single point polarization data measured for peptides with N-terminal arginine, lysine, and histidine using 2000 nM PS691 (Figure 7E). Figure 7F shows the binding energies calculated using a computer model (Example 1, Methods) for peptides with initial sequences LAX and LXA (where X = all 20 amino acids). Box plots show the fraction of the total binding energy contributed by amino acids at positions 1 (P1), 2 (P2), and 3 (P3), which tend to decrease exponentially from P1 to P3 (R2 > 0.97). Figure 7G shows the RS average PD determined in single molecule assays for LXA and LAX peptides using PS961 and for FXA and FAX peptides using PS610. Figure 7H shows that the non-polar solvation energy terms from the computational binding model with PS961 show high correlation with the actual RS average PD values observed in single molecule assays with peptides containing an N-terminal leucine and various amino acids at the P2 position. Peptides LVFA (SEQ ID NO: 859), LIFA (SEQ ID NO: 860), LVAR (SEQ ID NO: 861), LAFA (SEQ ID NO: 862), LQAR (SEQ ID NO: 863), LDAA (SEQ ID NO: 864), LCAR (SEQ ID NO: 865), LGAA (SEQ ID NO: 866), LMFA (SEQ ID NO: 867), LSAR (SEQ ID NO: 868), and LEFA (SEQ ID NO: 869) are shown. [Figure 7H]Examples of recognition factor properties are shown. Figures 7A-7E show recognition factor kinetic characterization using a polarization assay (Example 1, Methods). Figures 7A-7B show the affinity (KD) (Figure 7A) and off-rates (koff) (Figure 7B) of PS610 for peptides with N-terminal phenylalanine, tyrosine, and tryptophan. SEQ ID NOs: 809 (FAKLK(FITC)DEESILKQ), 810 (YAKLK(FITC)DEESILKQ), and 811 (WAKLK(FITC)DEESILKQ) are shown in Figure 7B. Figure 7C shows the affinity of PS961 for peptides with N-terminal leucine, isoleucine, and valine. Figures 7D-7E show the affinity of PS691 for peptides with N-terminal arginine (Figure 7D) and single point polarization data measured for peptides with N-terminal arginine, lysine, and histidine using 2000 nM PS691 (Figure 7E). Figure 7F shows the binding energies calculated using a computer model (Example 1, Methods) for peptides with initial sequences LAX and LXA (where X = all 20 amino acids). Box plots show the fraction of the total binding energy contributed by amino acids at positions 1 (P1), 2 (P2), and 3 (P3), which tend to decrease exponentially from P1 to P3 (R2 > 0.97). Figure 7G shows the RS average PD determined in single molecule assays for LXA and LAX peptides using PS961 and for FXA and FAX peptides using PS610. Figure 7H shows that the non-polar solvation energy terms from the computational binding model with PS961 show high correlation with the actual RS average PD values observed in single molecule assays with peptides containing an N-terminal leucine and various amino acids at the P2 position. Peptides LVFA (SEQ ID NO: 859), LIFA (SEQ ID NO: 860), LVAR (SEQ ID NO: 861), LAFA (SEQ ID NO: 862), LQAR (SEQ ID NO: 863), LDAA (SEQ ID NO: 864), LCAR (SEQ ID NO: 865), LGAA (SEQ ID NO: 866), LMFA (SEQ ID NO: 867), LSAR (SEQ ID NO: 868), and LEFA (SEQ ID NO: 869) are shown. [Figure 8A]Examples of binding and cleavage rates are shown. Figures 8A-8B show that inter-pulse duration (IPD) decreases with increasing recognition factor concentration. Scatter plots of RS mean PD vs. RS mean IPD for PS961 binding to LIF (Figure 8A) and IFA (Figure 8B) in dynamic sequencing assays at concentrations of 125 nM (orange) or 250 nM (blue). Median IPD values are shown. Recognition factor concentration did not affect RS mean PD. Figure 8C shows single exponential decay curves fitted to RS duration distributions for arginine, leucine, isoleucine, and phenylalanine obtained from dynamic sequencing of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794). Figures 8D-8E show that increasing aminopeptidase concentration in dynamic sequencing runs of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794) resulted in a decrease in NRS (Figure 8D) and RS (Figure 8E) duration. Median RS duration values are shown. [Figure 8B] Examples of binding and cleavage rates are shown. Figures 8A-8B show that inter-pulse duration (IPD) decreases with increasing recognition factor concentration. Scatter plots of RS mean PD vs. RS mean IPD for PS961 binding to LIF (Figure 8A) and IFA (Figure 8B) in dynamic sequencing assays at concentrations of 125 nM (orange) or 250 nM (blue). Median IPD values are shown. Recognition factor concentration did not affect RS mean PD. Figure 8C shows single exponential decay curves fitted to RS duration distributions for arginine, leucine, isoleucine, and phenylalanine obtained from dynamic sequencing of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794). Figures 8D-8E show that increasing aminopeptidase concentration in dynamic sequencing runs of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794) resulted in a decrease in NRS (Figure 8D) and RS (Figure 8E) duration. Median RS duration values are shown. [Figure 8C]Examples of binding and cleavage rates are shown. Figures 8A-8B show that inter-pulse duration (IPD) decreases with increasing recognition factor concentration. Scatter plots of RS mean PD vs. RS mean IPD for PS961 binding to LIF (Figure 8A) and IFA (Figure 8B) in dynamic sequencing assays at concentrations of 125 nM (orange) or 250 nM (blue). Median IPD values are shown. Recognition factor concentration did not affect RS mean PD. Figure 8C shows single exponential decay curves fitted to RS duration distributions for arginine, leucine, isoleucine, and phenylalanine obtained from dynamic sequencing of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794). Figures 8D-8E show that increasing aminopeptidase concentration in dynamic sequencing runs of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794) resulted in a decrease in NRS (Figure 8D) and RS (Figure 8E) duration. Median RS duration values are shown. [Figure 8D] Examples of binding and cleavage rates are shown. Figures 8A-8B show that inter-pulse duration (IPD) decreases with increasing recognition factor concentration. Scatter plots of RS mean PD vs. RS mean IPD for PS961 binding to LIF (Figure 8A) and IFA (Figure 8B) in dynamic sequencing assays at concentrations of 125 nM (orange) or 250 nM (blue). Median IPD values are shown. Recognition factor concentration did not affect RS mean PD. Figure 8C shows single exponential decay curves fitted to RS duration distributions for arginine, leucine, isoleucine, and phenylalanine obtained from dynamic sequencing of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794). Figures 8D-8E show that increasing aminopeptidase concentration in dynamic sequencing runs of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794) resulted in a decrease in NRS (Figure 8D) and RS (Figure 8E) duration. Median RS duration values are shown. [Figure 8E]Examples of binding and cleavage rates are shown. Figures 8A-8B show that inter-pulse duration (IPD) decreases with increasing recognition factor concentration. Scatter plots of RS mean PD vs. RS mean IPD for PS961 binding to LIF (Figure 8A) and IFA (Figure 8B) in dynamic sequencing assays at concentrations of 125 nM (orange) or 250 nM (blue). Median IPD values are shown. Recognition factor concentration did not affect RS mean PD. Figure 8C shows single exponential decay curves fitted to RS duration distributions for arginine, leucine, isoleucine, and phenylalanine obtained from dynamic sequencing of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794). Figures 8D-8E show that increasing aminopeptidase concentration in dynamic sequencing runs of the synthetic peptide DQQRLIFAG (SEQ ID NO: 794) resulted in a decrease in NRS (Figure 8D) and RS (Figure 8E) duration. Median RS duration values are shown. [Figure 9A]Examples of kinetic signatures from single amino acid changes and PTMs are shown. Figure 9A shows kinetic signature plots for three peptides, RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figures 9B-9C show incomplete RS information observed in dynamic sequencing of RLIFAYPDDD (SEQ ID NO: 799) peptide. Figure 9B shows the percentage of reads of each type of observed deletion of one or more RSs in traces starting with arginine recognition and ending with tyrosine recognition, and example traces. RLIFY (SEQ ID NO: 812). Figure 9C shows the percentage of reads of each type of observed truncation of one or more RSs in traces starting with arginine, and example traces. RLIFY (SEQ ID NO: 812). Figure 9D shows the affinity of PS961 for peptides with N-terminal methionine measured in a polarization assay (Example 1, Methods). Figure 9E shows binding energy predictions for peptides with N-terminal methionine and methionine sulfoxide (Mo) from computer modeling using PS961 (Example 1, Methods). Figure 9F shows kinetic signature plots for DQQRLIFAG (SEQ ID NO:794, residues 1-7 shown) and RLAFSALGAADDD (SEQ ID NO:796, residues 1-7 shown) peptides mixed and run on the same chip. Figure 9G shows kinetic signature plots for DQQRLIFAGK (SEQ ID NO:802, residues 1-7 shown) and EFIAWLVK (SEQ ID NO:803, residues 1-6 shown) peptides obtained from digestion of recombinant human ubiquitin and GLP-1. [Figure 9B]Examples of kinetic signatures from single amino acid changes and PTMs are shown. Figure 9A shows kinetic signature plots for three peptides, RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figures 9B-9C show incomplete RS information observed in dynamic sequencing of RLIFAYPDDD (SEQ ID NO: 799) peptide. Figure 9B shows the percentage of reads of each type of observed deletion of one or more RSs in traces starting with arginine recognition and ending with tyrosine recognition, and example traces. RLIFY (SEQ ID NO: 812). Figure 9C shows the percentage of reads of each type of observed truncation of one or more RSs in traces starting with arginine, and example traces. RLIFY (SEQ ID NO: 812). Figure 9D shows the affinity of PS961 for peptides with N-terminal methionine measured in a polarization assay (Example 1, Methods). Figure 9E shows binding energy predictions for peptides with N-terminal methionine and methionine sulfoxide (Mo) from computer modeling using PS961 (Example 1, Methods). Figure 9F shows kinetic signature plots for DQQRLIFAG (SEQ ID NO:794, residues 1-7 shown) and RLAFSALGAADDD (SEQ ID NO:796, residues 1-7 shown) peptides mixed and run on the same chip. Figure 9G shows kinetic signature plots for DQQRLIFAGK (SEQ ID NO:802, residues 1-7 shown) and EFIAWLVK (SEQ ID NO:803, residues 1-6 shown) peptides obtained from digestion of recombinant human ubiquitin and GLP-1. [Figure 9C]Examples of kinetic signatures from single amino acid changes and PTMs are shown. Figure 9A shows kinetic signature plots for three peptides, RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figures 9B-9C show incomplete RS information observed in dynamic sequencing of RLIFAYPDDD (SEQ ID NO: 799) peptide. Figure 9B shows the percentage of reads of each type of observed deletion of one or more RSs in traces starting with arginine recognition and ending with tyrosine recognition, and example traces. RLIFY (SEQ ID NO: 812). Figure 9C shows the percentage of reads of each type of observed truncation of one or more RSs in traces starting with arginine, and example traces. RLIFY (SEQ ID NO: 812). Figure 9D shows the affinity of PS961 for peptides with N-terminal methionine measured in a polarization assay (Example 1, Methods). Figure 9E shows binding energy predictions for peptides with N-terminal methionine and methionine sulfoxide (Mo) from computer modeling using PS961 (Example 1, Methods). Figure 9F shows kinetic signature plots for DQQRLIFAG (SEQ ID NO:794, residues 1-7 shown) and RLAFSALGAADDD (SEQ ID NO:796, residues 1-7 shown) peptides mixed and run on the same chip. Figure 9G shows kinetic signature plots for DQQRLIFAGK (SEQ ID NO:802, residues 1-7 shown) and EFIAWLVK (SEQ ID NO:803, residues 1-6 shown) peptides obtained from digestion of recombinant human ubiquitin and GLP-1. [Figure 9D]Examples of kinetic signatures from single amino acid changes and PTMs are shown. Figure 9A shows kinetic signature plots for three peptides, RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figures 9B-9C show incomplete RS information observed in dynamic sequencing of RLIFAYPDDD (SEQ ID NO: 799) peptide. Figure 9B shows the percentage of reads of each type of observed deletion of one or more RSs in traces starting with arginine recognition and ending with tyrosine recognition, and example traces. RLIFY (SEQ ID NO: 812). Figure 9C shows the percentage of reads of each type of observed truncation of one or more RSs in traces starting with arginine, and example traces. RLIFY (SEQ ID NO: 812). Figure 9D shows the affinity of PS961 for peptides with N-terminal methionine measured in a polarization assay (Example 1, Methods). Figure 9E shows binding energy predictions for peptides with N-terminal methionine and methionine sulfoxide (Mo) from computer modeling using PS961 (Example 1, Methods). Figure 9F shows kinetic signature plots for DQQRLIFAG (SEQ ID NO:794, residues 1-7 shown) and RLAFSALGAADDD (SEQ ID NO:796, residues 1-7 shown) peptides mixed and run on the same chip. Figure 9G shows kinetic signature plots for DQQRLIFAGK (SEQ ID NO:802, residues 1-7 shown) and EFIAWLVK (SEQ ID NO:803, residues 1-6 shown) peptides obtained from digestion of recombinant human ubiquitin and GLP-1. [Figure 9E]Examples of kinetic signatures from single amino acid changes and PTMs are shown. Figure 9A shows kinetic signature plots for three peptides, RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figures 9B-9C show incomplete RS information observed in dynamic sequencing of RLIFAYPDDD (SEQ ID NO: 799) peptide. Figure 9B shows the percentage of reads of each type of observed deletion of one or more RSs in traces starting with arginine recognition and ending with tyrosine recognition, and example traces. RLIFY (SEQ ID NO: 812). Figure 9C shows the percentage of reads of each type of observed truncation of one or more RSs in traces starting with arginine, and example traces. RLIFY (SEQ ID NO: 812). Figure 9D shows the affinity of PS961 for peptides with N-terminal methionine measured in a polarization assay (Example 1, Methods). Figure 9E shows binding energy predictions for peptides with N-terminal methionine and methionine sulfoxide (Mo) from computer modeling using PS961 (Example 1, Methods). Figure 9F shows kinetic signature plots for DQQRLIFAG (SEQ ID NO:794, residues 1-7 shown) and RLAFSALGAADDD (SEQ ID NO:796, residues 1-7 shown) peptides mixed and run on the same chip. Figure 9G shows kinetic signature plots for DQQRLIFAGK (SEQ ID NO:802, residues 1-7 shown) and EFIAWLVK (SEQ ID NO:803, residues 1-6 shown) peptides obtained from digestion of recombinant human ubiquitin and GLP-1. [Figure 9F]Examples of kinetic signatures from single amino acid changes and PTMs are shown. Figure 9A shows kinetic signature plots for three peptides, RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figures 9B-9C show incomplete RS information observed in dynamic sequencing of RLIFAYPDDD (SEQ ID NO: 799) peptide. Figure 9B shows the percentage of reads of each type of observed deletion of one or more RSs in traces starting with arginine recognition and ending with tyrosine recognition, and example traces. RLIFY (SEQ ID NO: 812). Figure 9C shows the percentage of reads of each type of observed truncation of one or more RSs in traces starting with arginine, and example traces. RLIFY (SEQ ID NO: 812). Figure 9D shows the affinity of PS961 for peptides with N-terminal methionine measured in a polarization assay (Example 1, Methods). Figure 9E shows binding energy predictions for peptides with N-terminal methionine and methionine sulfoxide (Mo) from computer modeling using PS961 (Example 1, Methods). Figure 9F shows kinetic signature plots for DQQRLIFAG (SEQ ID NO:794, residues 1-7 shown) and RLAFSALGAADDD (SEQ ID NO:796, residues 1-7 shown) peptides mixed and run on the same chip. Figure 9G shows kinetic signature plots for DQQRLIFAGK (SEQ ID NO:802, residues 1-7 shown) and EFIAWLVK (SEQ ID NO:803, residues 1-6 shown) peptides obtained from digestion of recombinant human ubiquitin and GLP-1. [Figure 9G]Examples of kinetic signatures from single amino acid changes and PTMs are shown. Figure 9A shows kinetic signature plots for three peptides, RLAFAYPDDD (SEQ ID NO: 798) (top), RLIFAYPDDD (SEQ ID NO: 799) (middle), and RLVFAYPDDD (SEQ ID NO: 800) (bottom). Figures 9B-9C show incomplete RS information observed in dynamic sequencing of RLIFAYPDDD (SEQ ID NO: 799) peptide. Figure 9B shows the percentage of reads of each type of observed deletion of one or more RSs in traces starting with arginine recognition and ending with tyrosine recognition, and example traces. RLIFY (SEQ ID NO: 812). Figure 9C shows the percentage of reads of each type of observed truncation of one or more RSs in traces starting with arginine, and example traces. RLIFY (SEQ ID NO: 812). Figure 9D shows the affinity of PS961 for peptides with N-terminal methionine measured in a polarization assay (Example 1, Methods). Figure 9E shows binding energy predictions for peptides with N-terminal methionine and methionine sulfoxide (Mo) from computer modeling using PS961 (Example 1, Methods). Figure 9F shows kinetic signature plots for DQQRLIFAG (SEQ ID NO:794, residues 1-7 shown) and RLAFSALGAADDD (SEQ ID NO:796, residues 1-7 shown) peptides mixed and run on the same chip. Figure 9G shows kinetic signature plots for DQQRLIFAGK (SEQ ID NO:802, residues 1-7 shown) and EFIAWLVK (SEQ ID NO:803, residues 1-6 shown) peptides obtained from digestion of recombinant human ubiquitin and GLP-1. [Figure 10A]Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10B] Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10C]Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10D] Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10E]Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10F] Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10G]Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10H] Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10I]Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10J] Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10K]Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10L] Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 10M]Examples of peptide identification using modeled proteome-wide kinetic signatures are shown. Figures 10A-C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. Figure 10H shows plots demonstrating high correlation between predicted and actual pulse durations from on-chip experiments for PS961 (left plot) and PS610 (right plot). Figures 10I-K show results from analysis of the human proteome. 10L-10M show results from an analysis of the E. coli proteome. [Figure 11A]Exemplary results showing direct identification of arginine PTMs are shown. Figure 11A shows different arginine PTMs, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine. Figure 11B shows an exemplary workflow for collecting samples, preparing a library of digested peptides, loading onto the chip, and performing on-chip sequencing and data analysis. Figure 11C shows sequencing data demonstrating that kinetic signatures distinguish peptides containing arginine, ADMA, and SDMA. Figure 11C-A shows example protein sequencing traces for three synthetic P38MAPKα-derived peptides containing arginine, ADMA, or SDMA at position 2. The full-length peptide sequence is shown for each example trace: YRELRLLK (SEQ ID NO: 834) (top), YRADMAELRKKL (SEQ ID NO: 894) (middle), YRSDMAELRLLK (SEQ ID NO: 895) (bottom). FIG. 11C-B shows the distribution of recognition segment (RS) mean pulse duration (PD) for RS corresponding to the first four residue sequence of each peptide: YREL (SEQ ID NO: 814) (left), YRADMAEL (SEQ ID NO: 815) (center), and YRSDMAEL (SEQ ID NO: 816) (right). Median values are shown for each distribution. FIG. 11C-C shows inter-pulse duration (IPD) for arginine vs. ADMA detection by PS621. FIG. 11D shows sequencing data demonstrating that the kinetic signature distinguishes peptides containing arginine and citrulline. FIG. 11D-A shows example protein sequencing traces for two synthetic peptides containing arginine or citrulline at position 2: peptide sequence LRLAFAYPDDDK (SEQ ID NO: 817) (QP707) and citrullinated peptide sequence LRCitLAFAYPDDDK (SEQ ID NO: 818) (QP789). Full length peptide sequences are shown for each example trace. Figures 11D-B show the distribution of RS average PD for RS corresponding to the first 5 residue sequences of each peptide: LRLAF (SEQ ID NO: 819) (left) and LCitLAF (SEQ ID NO: 820) (right). The median is shown for each distribution. [Figure 11B]Exemplary results showing direct identification of arginine PTMs are shown. Figure 11A shows different arginine PTMs, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine. Figure 11B shows an exemplary workflow for collecting samples, preparing a library of digested peptides, loading onto the chip, and performing on-chip sequencing and data analysis. Figure 11C shows sequencing data demonstrating that kinetic signatures distinguish peptides containing arginine, ADMA, and SDMA. Figure 11C-A shows example protein sequencing traces for three synthetic P38MAPKα-derived peptides containing arginine, ADMA, or SDMA at position 2. The full-length peptide sequence is shown for each example trace: YRELRLLK (SEQ ID NO: 834) (top), YRADMAELRKKL (SEQ ID NO: 894) (middle), YRSDMAELRLLK (SEQ ID NO: 895) (bottom). FIG. 11C-B shows the distribution of recognition segment (RS) mean pulse duration (PD) for RS corresponding to the first four residue sequence of each peptide: YREL (SEQ ID NO: 814) (left), YRADMAEL (SEQ ID NO: 815) (center), and YRSDMAEL (SEQ ID NO: 816) (right). Median values are shown for each distribution. FIG. 11C-C shows inter-pulse duration (IPD) for arginine vs. ADMA detection by PS621. FIG. 11D shows sequencing data demonstrating that the kinetic signature distinguishes peptides containing arginine and citrulline. FIG. 11D-A shows example protein sequencing traces for two synthetic peptides containing arginine or citrulline at position 2: peptide sequence LRLAFAYPDDDK (SEQ ID NO: 817) (QP707) and citrullinated peptide sequence LRCitLAFAYPDDDK (SEQ ID NO: 818) (QP789). Full length peptide sequences are shown for each example trace. Figures 11D-B show the distribution of RS average PD for RS corresponding to the first 5 residue sequences of each peptide: LRLAF (SEQ ID NO: 819) (left) and LCitLAF (SEQ ID NO: 820) (right). The median is shown for each distribution. [Figure 11C-1]Exemplary results showing direct identification of arginine PTMs are shown. Figure 11A shows different arginine PTMs, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine. Figure 11B shows an exemplary workflow for collecting samples, preparing a library of digested peptides, loading onto the chip, and performing on-chip sequencing and data analysis. Figure 11C shows sequencing data demonstrating that kinetic signatures distinguish peptides containing arginine, ADMA, and SDMA. Figure 11C-A shows example protein sequencing traces for three synthetic P38MAPKα-derived peptides containing arginine, ADMA, or SDMA at position 2. The full-length peptide sequence is shown for each example trace: YRELRLLK (SEQ ID NO: 834) (top), YRADMAELRKKL (SEQ ID NO: 894) (middle), YRSDMAELRLLK (SEQ ID NO: 895) (bottom). FIG. 11C-B shows the distribution of recognition segment (RS) mean pulse duration (PD) for RS corresponding to the first four residue sequence of each peptide: YREL (SEQ ID NO: 814) (left), YRADMAEL (SEQ ID NO: 815) (center), and YRSDMAEL (SEQ ID NO: 816) (right). Median values are shown for each distribution. FIG. 11C-C shows inter-pulse duration (IPD) for arginine vs. ADMA detection by PS621. FIG. 11D shows sequencing data demonstrating that the kinetic signature distinguishes peptides containing arginine and citrulline. FIG. 11D-A shows example protein sequencing traces for two synthetic peptides containing arginine or citrulline at position 2: peptide sequence LRLAFAYPDDDK (SEQ ID NO: 817) (QP707) and citrullinated peptide sequence LRCitLAFAYPDDDK (SEQ ID NO: 818) (QP789). Full length peptide sequences are shown for each example trace. Figures 11D-B show the distribution of RS average PD for RS corresponding to the first 5 residue sequences of each peptide: LRLAF (SEQ ID NO: 819) (left) and LCitLAF (SEQ ID NO: 820) (right). The median is shown for each distribution. [Figure 11C-2]Exemplary results showing direct identification of arginine PTMs are shown. Figure 11A shows different arginine PTMs, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine. Figure 11B shows an exemplary workflow for collecting samples, preparing a library of digested peptides, loading onto the chip, and performing on-chip sequencing and data analysis. Figure 11C shows sequencing data demonstrating that kinetic signatures distinguish peptides containing arginine, ADMA, and SDMA. Figure 11C-A shows example protein sequencing traces for three synthetic P38MAPKα-derived peptides containing arginine, ADMA, or SDMA at position 2. The full-length peptide sequence is shown for each example trace: YRELRLLK (SEQ ID NO: 834) (top), YRADMAELRKKL (SEQ ID NO: 894) (middle), YRSDMAELRLLK (SEQ ID NO: 895) (bottom). FIG. 11C-B shows the distribution of recognition segment (RS) mean pulse duration (PD) for RS corresponding to the first four residue sequence of each peptide: YREL (SEQ ID NO: 814) (left), YRADMAEL (SEQ ID NO: 815) (center), and YRSDMAEL (SEQ ID NO: 816) (right). Median values are shown for each distribution. FIG. 11C-C shows inter-pulse duration (IPD) for arginine vs. ADMA detection by PS621. FIG. 11D shows sequencing data demonstrating that the kinetic signature distinguishes peptides containing arginine and citrulline. FIG. 11D-A shows example protein sequencing traces for two synthetic peptides containing arginine or citrulline at position 2: peptide sequence LRLAFAYPDDDK (SEQ ID NO: 817) (QP707) and citrullinated peptide sequence LRCitLAFAYPDDDK (SEQ ID NO: 818) (QP789). Full length peptide sequences are shown for each example trace. Figures 11D-B show the distribution of RS average PD for RS corresponding to the first 5 residue sequences of each peptide: LRLAF (SEQ ID NO: 819) (left) and LCitLAF (SEQ ID NO: 820) (right). The median is shown for each distribution. [Figure 11C-3]Exemplary results showing direct identification of arginine PTMs are shown. Figure 11A shows different arginine PTMs, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine. Figure 11B shows an exemplary workflow for collecting samples, preparing a library of digested peptides, loading onto the chip, and performing on-chip sequencing and data analysis. Figure 11C shows sequencing data demonstrating that kinetic signatures distinguish peptides containing arginine, ADMA, and SDMA. Figure 11C-A shows example protein sequencing traces for three synthetic P38MAPKα-derived peptides containing arginine, ADMA, or SDMA at position 2. The full-length peptide sequence is shown for each example trace: YRELRLLK (SEQ ID NO: 834) (top), YRADMAELRKKL (SEQ ID NO: 894) (middle), YRSDMAELRLLK (SEQ ID NO: 895) (bottom). FIG. 11C-B shows the distribution of recognition segment (RS) mean pulse duration (PD) for RS corresponding to the first four residue sequence of each peptide: YREL (SEQ ID NO: 814) (left), YRADMAEL (SEQ ID NO: 815) (center), and YRSDMAEL (SEQ ID NO: 816) (right). Median values are shown for each distribution. FIG. 11C-C shows inter-pulse duration (IPD) for arginine vs. ADMA detection by PS621. FIG. 11D shows sequencing data demonstrating that the kinetic signature distinguishes peptides containing arginine and citrulline. FIG. 11D-A shows example protein sequencing traces for two synthetic peptides containing arginine or citrulline at position 2: peptide sequence LRLAFAYPDDDK (SEQ ID NO: 817) (QP707) and citrullinated peptide sequence LRCitLAFAYPDDDK (SEQ ID NO: 818) (QP789). Full length peptide sequences are shown for each example trace. Figures 11D-B show the distribution of RS average PD for RS corresponding to the first 5 residue sequences of each peptide: LRLAF (SEQ ID NO: 819) (left) and LCitLAF (SEQ ID NO: 820) (right). The median is shown for each distribution. [Figure 11D-1]Exemplary results showing direct identification of arginine PTMs are shown. Figure 11A shows different arginine PTMs, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine. Figure 11B shows an exemplary workflow for collecting samples, preparing a library of digested peptides, loading onto the chip, and performing on-chip sequencing and data analysis. Figure 11C shows sequencing data demonstrating that kinetic signatures distinguish peptides containing arginine, ADMA, and SDMA. Figure 11C-A shows example protein sequencing traces for three synthetic P38MAPKα-derived peptides containing arginine, ADMA, or SDMA at position 2. The full-length peptide sequence is shown for each example trace: YRELRLLK (SEQ ID NO: 834) (top), YRADMAELRKKL (SEQ ID NO: 894) (middle), YRSDMAELRLLK (SEQ ID NO: 895) (bottom). FIG. 11C-B shows the distribution of recognition segment (RS) mean pulse duration (PD) for RS corresponding to the first four residue sequence of each peptide: YREL (SEQ ID NO: 814) (left), YRADMAEL (SEQ ID NO: 815) (center), and YRSDMAEL (SEQ ID NO: 816) (right). Median values are shown for each distribution. FIG. 11C-C shows inter-pulse duration (IPD) for arginine vs. ADMA detection by PS621. FIG. 11D shows sequencing data demonstrating that the kinetic signature distinguishes peptides containing arginine and citrulline. FIG. 11D-A shows example protein sequencing traces for two synthetic peptides containing arginine or citrulline at position 2: peptide sequence LRLAFAYPDDDK (SEQ ID NO: 817) (QP707) and citrullinated peptide sequence LRCitLAFAYPDDDK (SEQ ID NO: 818) (QP789). Full length peptide sequences are shown for each example trace. Figures 11D-B show the distribution of RS average PD for RS corresponding to the first 5 residue sequences of each peptide: LRLAF (SEQ ID NO: 819) (left) and LCitLAF (SEQ ID NO: 820) (right). The median is shown for each distribution. [Figure 11D-2]Exemplary results showing direct identification of arginine PTMs are shown. Figure 11A shows different arginine PTMs, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrullinated arginine. Figure 11B shows an exemplary workflow for collecting samples, preparing a library of digested peptides, loading onto the chip, and performing on-chip sequencing and data analysis. Figure 11C shows sequencing data demonstrating that kinetic signatures distinguish peptides containing arginine, ADMA, and SDMA. Figure 11C-A shows example protein sequencing traces for three synthetic P38MAPKα-derived peptides containing arginine, ADMA, or SDMA at position 2. The full-length peptide sequence is shown for each example trace: YRELRLLK (SEQ ID NO: 834) (top), YRADMAELRKKL (SEQ ID NO: 894) (middle), YRSDMAELRLLK (SEQ ID NO: 895) (bottom). FIG. 11C-B shows the distribution of recognition segment (RS) mean pulse duration (PD) for RS corresponding to the first four residue sequence of each peptide: YREL (SEQ ID NO: 814) (left), YRADMAEL (SEQ ID NO: 815) (center), and YRSDMAEL (SEQ ID NO: 816) (right). Median values are shown for each distribution. FIG. 11C-C shows inter-pulse duration (IPD) for arginine vs. ADMA detection by PS621. FIG. 11D shows sequencing data demonstrating that the kinetic signature distinguishes peptides containing arginine and citrulline. FIG. 11D-A shows example protein sequencing traces for two synthetic peptides containing arginine or citrulline at position 2: peptide sequence LRLAFAYPDDDK (SEQ ID NO: 817) (QP707) and citrullinated peptide sequence LRCitLAFAYPDDDK (SEQ ID NO: 818) (QP789). Full length peptide sequences are shown for each example trace. Figures 11D-B show the distribution of RS average PD for RS corresponding to the first 5 residue sequences of each peptide: LRLAF (SEQ ID NO: 819) (left) and LCitLAF (SEQ ID NO: 820) (right). The median is shown for each distribution. [Figure 12A]12A shows an exemplary method for using kinetic signature information. FIG. 12A shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12B shows an exemplary method for determining a chemical property of a polypeptide, where one or more amino acids of the polypeptide are unrecognizable. FIG. 12C shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12D shows an exemplary method for identifying a protein from which a polypeptide is derived based on a pulse pattern comprising at least three recognition segments. FIG. 12E shows an exemplary method for characterizing an amino acid based on a pulse pattern emitted by one or more recognition factors bound to a first amino acid. FIG. 12F shows an exemplary method for determining at least one chemical property of an amino acid of a polypeptide. FIG. 12G shows an exemplary method for identifying a disease or disorder in a subject based on at least one chemical property of a polypeptide. [Figure 12B] 12A shows an exemplary method for using kinetic signature information. FIG. 12A shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12B shows an exemplary method for determining a chemical property of a polypeptide, where one or more amino acids of the polypeptide are unrecognizable. FIG. 12C shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12D shows an exemplary method for identifying a protein from which a polypeptide is derived based on a pulse pattern comprising at least three recognition segments. FIG. 12E shows an exemplary method for characterizing an amino acid based on a pulse pattern emitted by one or more recognition factors bound to a first amino acid. FIG. 12F shows an exemplary method for determining at least one chemical property of an amino acid of a polypeptide. FIG. 12G shows an exemplary method for identifying a disease or disorder in a subject based on at least one chemical property of a polypeptide. [Figure 12C]12A shows an exemplary method for using kinetic signature information. FIG. 12A shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12B shows an exemplary method for determining a chemical property of a polypeptide, where one or more amino acids of the polypeptide are unrecognizable. FIG. 12C shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12D shows an exemplary method for identifying a protein from which a polypeptide is derived based on a pulse pattern comprising at least three recognition segments. FIG. 12E shows an exemplary method for characterizing an amino acid based on a pulse pattern emitted by one or more recognition factors bound to a first amino acid. FIG. 12F shows an exemplary method for determining at least one chemical property of an amino acid of a polypeptide. FIG. 12G shows an exemplary method for identifying a disease or disorder in a subject based on at least one chemical property of a polypeptide. [Figure 12D] 12A shows an exemplary method for using kinetic signature information. FIG. 12A shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12B shows an exemplary method for determining a chemical property of a polypeptide, where one or more amino acids of the polypeptide are unrecognizable. FIG. 12C shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12D shows an exemplary method for identifying a protein from which a polypeptide is derived based on a pulse pattern comprising at least three recognition segments. FIG. 12E shows an exemplary method for characterizing an amino acid based on a pulse pattern emitted by one or more recognition factors bound to a first amino acid. FIG. 12F shows an exemplary method for determining at least one chemical property of an amino acid of a polypeptide. FIG. 12G shows an exemplary method for identifying a disease or disorder in a subject based on at least one chemical property of a polypeptide. [Figure 12E]12A shows an exemplary method for using kinetic signature information. FIG. 12A shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12B shows an exemplary method for determining a chemical property of a polypeptide, where one or more amino acids of the polypeptide are unrecognizable. FIG. 12C shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12D shows an exemplary method for identifying a protein from which a polypeptide is derived based on a pulse pattern comprising at least three recognition segments. FIG. 12E shows an exemplary method for characterizing an amino acid based on a pulse pattern emitted by one or more recognition factors bound to a first amino acid. FIG. 12F shows an exemplary method for determining at least one chemical property of an amino acid of a polypeptide. FIG. 12G shows an exemplary method for identifying a disease or disorder in a subject based on at least one chemical property of a polypeptide. [Figure 12F] 12A shows an exemplary method for using kinetic signature information. FIG. 12A shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12B shows an exemplary method for determining a chemical property of a polypeptide, where one or more amino acids of the polypeptide are unrecognizable. FIG. 12C shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12D shows an exemplary method for identifying a protein from which a polypeptide is derived based on a pulse pattern comprising at least three recognition segments. FIG. 12E shows an exemplary method for characterizing an amino acid based on a pulse pattern emitted by one or more recognition factors bound to a first amino acid. FIG. 12F shows an exemplary method for determining at least one chemical property of an amino acid of a polypeptide. FIG. 12G shows an exemplary method for identifying a disease or disorder in a subject based on at least one chemical property of a polypeptide. [Figure 12G]12A shows an exemplary method for using kinetic signature information. FIG. 12A shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12B shows an exemplary method for determining a chemical property of a polypeptide, where one or more amino acids of the polypeptide are unrecognizable. FIG. 12C shows an exemplary method for determining a chemical property of a polypeptide. FIG. 12D shows an exemplary method for identifying a protein from which a polypeptide is derived based on a pulse pattern comprising at least three recognition segments. FIG. 12E shows an exemplary method for characterizing an amino acid based on a pulse pattern emitted by one or more recognition factors bound to a first amino acid. FIG. 12F shows an exemplary method for determining at least one chemical property of an amino acid of a polypeptide. FIG. 12G shows an exemplary method for identifying a disease or disorder in a subject based on at least one chemical property of a polypeptide. [Figure 13] 1 shows an exemplary schematic diagram of a pixel of an integrated device. [Figure 14A] Examples of results showing the identification of threonine PTMs are shown. Figures 14A-B show results from sequencing reactions using recognition factors PS691, PS610, and PS961 for peptides RLTFIAYPDDD (SEQ ID NO: 821) (Figure 14A); and RLpTFIAYPDDD (SEQ ID NO: 822), where pT is phosphothreonine (Figure 14B). Figure 14C shows the recognition segment (RS) duration for leucine recognition in the sequencing reactions of Figures 14A (left panel) and 14B (right panel). [Figure 14B] Examples of results showing the identification of threonine PTMs are shown. Figures 14A-B show results from sequencing reactions using recognition factors PS691, PS610, and PS961 for peptides RLTFIAYPDDD (SEQ ID NO: 821) (Figure 14A); and RLpTFIAYPDDD (SEQ ID NO: 822), where pT is phosphothreonine (Figure 14B). Figure 14C shows the recognition segment (RS) duration for leucine recognition in the sequencing reactions of Figures 14A (left panel) and 14B (right panel). [Figure 14C]Examples of results showing the identification of threonine PTMs are shown. Figures 14A-B show results from sequencing reactions using recognition factors PS691, PS610, and PS961 for peptides RLTFIAYPDDD (SEQ ID NO: 821) (Figure 14A); and RLpTFIAYPDDD (SEQ ID NO: 822), where pT is phosphothreonine (Figure 14B). Figure 14C shows the recognition segment (RS) duration for leucine recognition in the sequencing reactions of Figures 14A (left panel) and 14B (right panel). [Figure 15A] Example results are shown showing the identification of tyrosine PTMs in sequencing reactions with recognition factors PS691, PS610, and PS961 for the peptides RLYFIAYPDDD (SEQ ID NO: 823) (Figure 15A); and RLpYFIAYPDDD (SEQ ID NO: 824), where pY is phosphotyrosine (Figure 15B). [Figure 15B] Example results are shown showing the identification of tyrosine PTMs in sequencing reactions with recognition factors PS691, PS610, and PS961 for the peptides RLYFIAYPDDD (SEQ ID NO: 823) (Figure 15A); and RLpYFIAYPDDD (SEQ ID NO: 824), where pY is phosphotyrosine (Figure 15B). [Figure 16A] Example results are shown showing the identification of lysine PTMs in sequencing reactions with recognition factors PS691, PS610, PS961 and PS1165 for the peptides RLYFKAYPDDD (SEQ ID NO: 825) (Figure 16A); and RLK{acetyl}FIAYPDDD (SEQ ID NO: 826), where K{acetyl} is acetylated lysine (Figure 16B). [Figure 16B] Example results are shown showing the identification of lysine PTMs in sequencing reactions with recognition factors PS691, PS610, PS961 and PS1165 for the peptides RLYFKAYPDDD (SEQ ID NO: 825) (Figure 16A); and RLK{acetyl}FIAYPDDD (SEQ ID NO: 826), where K{acetyl} is acetylated lysine (Figure 16B). [Figure 17A]An embodiment of an exemplary application of the present technology to the identification of β-amyloid variants is shown. FIG. 17A shows an example of a β-amyloid variant. FIG. 17B shows an example of a workflow for β-amyloid variant detection. FIG. 17C-G show examples of pulse patterns of β-amyloid wild type LVFFAE (SEQ ID NO: 827) versus variants (LVFFAK (SEQ ID NO: 828), LVFFGK (SEQ ID NO: 829), LVFFAG (SEQ ID NO: 830), LVPFAE (SEQ ID NO: 831)). [Figure 17B] An embodiment of an exemplary application of the present technology to the identification of β-amyloid variants is shown. FIG. 17A shows an example of a β-amyloid variant. FIG. 17B shows an example of a workflow for β-amyloid variant detection. FIG. 17C-G show examples of pulse patterns of β-amyloid wild type LVFFAE (SEQ ID NO: 827) versus variants (LVFFAK (SEQ ID NO: 828), LVFFGK (SEQ ID NO: 829), LVFFAG (SEQ ID NO: 830), LVPFAE (SEQ ID NO: 831)). [Figure 17C] An embodiment of an exemplary application of the present technology to the identification of β-amyloid variants is shown. FIG. 17A shows an example of a β-amyloid variant. FIG. 17B shows an example of a workflow for β-amyloid variant detection. FIG. 17C-G show examples of pulse patterns of β-amyloid wild type LVFFAE (SEQ ID NO: 827) versus variants (LVFFAK (SEQ ID NO: 828), LVFFGK (SEQ ID NO: 829), LVFFAG (SEQ ID NO: 830), LVPFAE (SEQ ID NO: 831)). [Figure 17D] An embodiment of an exemplary application of the present technology to the identification of β-amyloid variants is shown. FIG. 17A shows an example of a β-amyloid variant. FIG. 17B shows an example of a workflow for β-amyloid variant detection. FIG. 17C-G show examples of pulse patterns of β-amyloid wild type LVFFAE (SEQ ID NO: 827) versus variants (LVFFAK (SEQ ID NO: 828), LVFFGK (SEQ ID NO: 829), LVFFAG (SEQ ID NO: 830), LVPFAE (SEQ ID NO: 831)). [Figure 17E]An embodiment of an exemplary application of the present technology to the identification of β-amyloid variants is shown. FIG. 17A shows an example of a β-amyloid variant. FIG. 17B shows an example of a workflow for β-amyloid variant detection. FIG. 17C-G show examples of pulse patterns of β-amyloid wild type LVFFAE (SEQ ID NO: 827) versus variants (LVFFAK (SEQ ID NO: 828), LVFFGK (SEQ ID NO: 829), LVFFAG (SEQ ID NO: 830), LVPFAE (SEQ ID NO: 831)). [Figure 17F] An embodiment of an exemplary application of the present technology to the identification of β-amyloid variants is shown. FIG. 17A shows an example of a β-amyloid variant. FIG. 17B shows an example of a workflow for β-amyloid variant detection. FIG. 17C-G show examples of pulse patterns of β-amyloid wild type LVFFAE (SEQ ID NO: 827) versus variants (LVFFAK (SEQ ID NO: 828), LVFFGK (SEQ ID NO: 829), LVFFAG (SEQ ID NO: 830), LVPFAE (SEQ ID NO: 831)). [Figure 17G] An embodiment of an exemplary application of the present technology to the identification of β-amyloid variants is shown. FIG. 17A shows an example of a β-amyloid variant. FIG. 17B shows an example of a workflow for β-amyloid variant detection. FIG. 17C-G show examples of pulse patterns of β-amyloid wild type LVFFAE (SEQ ID NO: 827) versus variants (LVFFAK (SEQ ID NO: 828), LVFFGK (SEQ ID NO: 829), LVFFAG (SEQ ID NO: 830), LVPFAE (SEQ ID NO: 831)). [Figure 18A] Examples of results from sequencing reactions using recognition agents PS610, PS1220, and PS1223 for peptide fragments containing unmodified arginine or citrulline are shown. Figure 18A shows a plot of bin ratio versus pulse duration for peptide fragment VRFLEQQNK (SEQ ID NO: 841). Figure 18B shows a plot of bin ratio versus pulse duration for peptide fragment VCitFLEQQNK (SEQ ID NO: 842), where Cit is citrulline. [Figure 18B]Examples of results from sequencing reactions using recognition agents PS610, PS1220, and PS1223 for peptide fragments containing unmodified arginine or citrulline are shown. Figure 18A shows a plot of bin ratio versus pulse duration for peptide fragment VRFLEQQNK (SEQ ID NO: 841). Figure 18B shows a plot of bin ratio versus pulse duration for peptide fragment VCitFLEQQNK (SEQ ID NO: 842), where Cit is citrulline. [Figure 19A] 19A and 19B show example results from sequencing reactions using recognition factors PS610, PS1220, and PS1223 for the peptide fragment VRFLEQQNK (SEQ ID NO: 841). Figures 19A and 19B show example traces, and Figures 19C and 19D show example plots of intensity versus bin ratio. [Figure 19B] 19A and 19B show example results from sequencing reactions using recognition factors PS610, PS1220, and PS1223 for the peptide fragment VRFLEQQNK (SEQ ID NO: 841). Figures 19A and 19B show example traces, and Figures 19C and 19D show example plots of intensity versus bin ratio. [Figure 19C] 19A and 19B show example results from sequencing reactions using recognition factors PS610, PS1220, and PS1223 for the peptide fragment VRFLEQQNK (SEQ ID NO: 841). Figures 19A and 19B show example traces, and Figures 19C and 19D show example plots of intensity versus bin ratio. [Figure 19D] 19A and 19B show example results from sequencing reactions using recognition factors PS610, PS1220, and PS1223 for the peptide fragment VRFLEQQNK (SEQ ID NO: 841). Figures 19A and 19B show example traces, and Figures 19C and 19D show example plots of intensity versus bin ratio. [Figure 20A] Figures 20A and 20B show example results from sequencing reactions using recognition factors PS610, PS1220, and PS1223 for the peptide fragment VCitFLEQQNK (SEQ ID NO: 842), where Cit is citrulline. Figures 20A and 20B show example traces, and Figures 20C and 20D show example plots of intensity versus bin ratio. [Figure 20B]Figures 20A and 20B show example results from sequencing reactions using recognition factors PS610, PS1220, and PS1223 for the peptide fragment VCitFLEQQNK (SEQ ID NO: 842), where Cit is citrulline. Figures 20A and 20B show example traces, and Figures 20C and 20D show example plots of intensity versus bin ratio. [Figure 20C] Figures 20A and 20B show example results from sequencing reactions using recognition factors PS610, PS1220, and PS1223 for the peptide fragment VCitFLEQQNK (SEQ ID NO: 842), where Cit is citrulline. Figures 20A and 20B show example traces, and Figures 20C and 20D show example plots of intensity versus bin ratio. [Figure 20D] Figures 20A and 20B show example results from sequencing reactions using recognition factors PS610, PS1220, and PS1223 for the peptide fragment VCitFLEQQNK (SEQ ID NO: 842), where Cit is citrulline. Figures 20A and 20B show example traces, and Figures 20C and 20D show example plots of intensity versus bin ratio. [Figure 21A-1] Figure 21 shows an example of the results of mapping kinetic signatures to the human proteome. In Figure 21A, the kinetic signatures from each of the five clusters of traces from brain dopamine neurotrophic factor (CDNF) sequencing are mapped to a database of predicted kinetic signatures for more than 300000 peptides derived from in silico Lys-C digestion of about 20000 human proteins, and candidate matching peptides are shown for each cluster. Figure 21B shows the matching peptides for each kinetic signature aligned against the full-length sequence of human CDNF protein. [Figure 21A-2]Figure 21 shows an example of the results of mapping kinetic signatures to the human proteome. In Figure 21A, the kinetic signatures from each of the five clusters of traces from brain dopamine neurotrophic factor (CDNF) sequencing are mapped to a database of predicted kinetic signatures for more than 300000 peptides derived from in silico Lys-C digestion of about 20000 human proteins, and candidate matching peptides are shown for each cluster. Figure 21B shows the matching peptides for each kinetic signature aligned against the full-length sequence of human CDNF protein. [Figure 21A-3] Figure 21 shows an example of the results of mapping kinetic signatures to the human proteome. In Figure 21A, the kinetic signatures from each of the five clusters of traces from brain dopamine neurotrophic factor (CDNF) sequencing are mapped to a database of predicted kinetic signatures for more than 300000 peptides derived from in silico Lys-C digestion of about 20000 human proteins, and candidate matching peptides are shown for each cluster. Figure 21B shows the matching peptides for each kinetic signature aligned against the full-length sequence of human CDNF protein. [Figure 21A-4] Figure 21 shows an example of the results of mapping kinetic signatures to the human proteome. In Figure 21A, the kinetic signatures from each of the five clusters of traces from brain dopamine neurotrophic factor (CDNF) sequencing are mapped to a database of predicted kinetic signatures for more than 300000 peptides derived from in silico Lys-C digestion of about 20000 human proteins, and candidate matching peptides are shown for each cluster. Figure 21B shows the matching peptides for each kinetic signature aligned against the full-length sequence of human CDNF protein. [Figure 21B-1]Figure 21 shows an example of the results of mapping kinetic signatures to the human proteome. In Figure 21A, the kinetic signatures from each of the five clusters of traces from brain dopamine neurotrophic factor (CDNF) sequencing are mapped to a database of predicted kinetic signatures for more than 300000 peptides derived from in silico Lys-C digestion of about 20000 human proteins, and candidate matching peptides are shown for each cluster. Figure 21B shows the matching peptides for each kinetic signature aligned against the full-length sequence of human CDNF protein. [Figure 21B-2] Figure 21 shows an example of the results of mapping kinetic signatures to the human proteome. In Figure 21A, the kinetic signatures from each of the five clusters of traces from brain dopamine neurotrophic factor (CDNF) sequencing are mapped to a database of predicted kinetic signatures for more than 300000 peptides derived from in silico Lys-C digestion of about 20000 human proteins, and candidate matching peptides are shown for each cluster. Figure 21B shows the matching peptides for each kinetic signature aligned against the full-length sequence of human CDNF protein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] The accompanying drawings, which constitute a part of this specification, illustrate several embodiments of the present disclosure and, together with the accompanying description, serve to explain the principles of the disclosure. Aspects of the present application relate to methods and systems for obtaining information about a plurality of amino acids in a polypeptide based on the binding interactions between the polypeptide and one or more amino acid recognition factors. For example, kinetic signature information can be obtained from a series of signal pulses that indicate a series of binding events between one or more amino acid recognition factors and amino acids (e.g., terminal amino acids, internal amino acids) of the polypeptide. The kinetic signature information (e.g., pulse duration, inter-pulse duration, recognition segment (RS) duration, inter-segment duration) can be used to determine one or more chemical properties (e.g., identity, modification) of a plurality of amino acids of the polypeptide.
[0015] Protein characterization has several important applications, including determining the presence or absence of a protein (e.g., a disease-related protein) in a biological sample, identifying unknown proteins in a biological sample, and identifying proteins involved in biological activity in an isolated protein fraction. However, traditional methods of protein characterization, such as mass spectrometry and affinity-based methods, often face significant challenges, including the inability to identify unknown proteins and / or to distinguish unmodified proteins from proteins with post-translational modifications (PTMs). In contrast, the methods and systems described herein can provide highly accurate characterization of a wide range of proteins. In some embodiments, the methods and systems described herein use single molecule protein sequencing to identify and / or otherwise characterize proteins based on the kinetic signature of binding between a recognition factor and a polypeptide fragment of the protein. This approach provides the necessary resolution to distinguish between polypeptides with similar sequences or physicochemical properties.
[0016] Kinetic signature information can be useful for mapping peptides to their derived proteins, at least because kinetic signature information related to one amino acid can provide information about the chemical properties of multiple amino acids. In some cases, when amino acid recognition factors bind to a polypeptide, they contact not only one amino acid, but one or more upstream and / or downstream amino acids. This contact with one or more upstream and / or downstream amino acids can affect the kinetic signature information (e.g., pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration). This sensitivity of the kinetic signature information to upstream and / or downstream amino acids can provide rich information about peptide sequence composition and can facilitate mapping of peptides to proteomes. In some embodiments, the recognition factor can bind to a terminal amino acid of a polypeptide and one or more downstream amino acids. In some embodiments, the recognition factor can bind to an internal amino acid of a polypeptide and one or more upstream and / or downstream amino acids. In this way, the recognition factor can directly or indirectly sense all 20 amino acids found in the human body (i.e., the building blocks of the human proteome), and this information can be encoded in the average pulse duration, inter-pulse duration, recognition segment (RS) duration, and / or inter-segment duration. Furthermore, adjacent visible residues in a polypeptide can generally be represented by directly adjacent RSs (i.e., a consensus gap between two RSs can only occur if there is at least one invisible amino acid between them).
[0017] As described herein, signal pulses from a dye-labeled first type of amino acid recognition factor that binds to an amino acid, such as a terminal amino acid or an internal amino acid, of a polypeptide can be used to determine one or more chemical properties of a plurality of amino acids of a polypeptide. The inventors have recognized that such a technique is advantageous. For example, such a technique may allow for the determination of chemical properties of unrecognized amino acids. Such amino acids may, in some cases, not be recognized by any amino acid recognition factor present in the reaction chamber. Such a technique may also be time-saving, require fewer amino acid recognition factors, and / or require fewer signal collections. Thus, it is advantageous to obtain information about a plurality of amino acids based on a fewer series of signal pulses and / or use fewer recognition factors.
[0018] In some embodiments, kinetic signature information obtained from a first series of binding events between one or more amino acid recognition factors and a first amino acid (e.g., terminal amino acid, internal amino acid) of a polypeptide can be used to determine one or more chemical properties of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, or more amino acids of the polypeptide. In certain embodiments, kinetic signature information obtained from the first series of binding events can be used to determine one or more chemical properties of 1 amino acid, 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 10 amino acids, 15 amino acids, 20 amino acids, 50 amino acids, or 100 amino acids. In certain embodiments, the kinetic signature information obtained from the first set of binding events can be used to determine one or more chemical properties of 1-2 amino acids, 1-3 amino acids, 1-4 amino acids, 1-5 amino acids, 1-10 amino acids, 1-15 amino acids, 1-20 amino acids, 1-50 amino acids, 1-100 amino acids, 2-3 amino acids, 2-4 amino acids, 2-5 amino acids, 2-10 amino acids, 2-15 amino acids, 2-20 amino acids, 2-50 amino acids, 2-100 amino acids, 3-5 amino acids, 3-10 amino acids, 3-15 amino acids, 3-20 amino acids, 3-50 amino acids, 3-100 amino acids, 5-10 amino acids, 5-15 amino acids, 5-20 amino acids, 5-50 amino acids, 5-100 amino acids, 10-20 amino acids, 10-50 amino acids, 10-100 amino acids, 20-50 amino acids, 20-100 amino acids, or 50-100 amino acids.
[0019] In some embodiments, kinetic signature information obtained from a first series of binding events between one or more amino acid recognition factors and a first amino acid (e.g., terminal amino acid, internal amino acid) of a polypeptide can be used to determine one or more chemical properties of at least a second amino acid of the polypeptide. In certain embodiments, the first amino acid is a terminal amino acid and the second amino acid is downstream of the first amino acid. In certain embodiments, the first amino acid is an internal amino acid and the second amino acid is upstream or downstream of the first amino acid. In some examples, the second amino acid is adjacent to the first amino acid. In some cases, for example, the second amino acid is 10 amino acids or less, 5 amino acids or less, 4 amino acids or less, 3 amino acids or less, 2 amino acids or less, 1 amino acid or less, or 0 amino acids away from the first amino acid (i.e., the first and second amino acids are directly adjacent). In some cases, the second amino acid is at least 1 amino acid, at least 2 amino acids, at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, or at least 10 amino acids away from the first amino acid. In some cases, the second amino acid is separated from the first amino acid by 1-2 amino acids, 1-3 amino acids, 1-4 amino acids, 1-5 amino acids, 1-10 amino acids, 2-3 amino acids, 2-4 amino acids, 2-5 amino acids, 2-10 amino acids, 3-5 amino acids, 3-10 amino acids, or 5-10 amino acids.
[0020] In some embodiments, kinetic signature information obtained from a first series of binding events between one or more amino acid recognition factors and a first amino acid (e.g., terminal amino acid, internal amino acid) of a polypeptide can be used to determine one or more chemical properties of at least a second amino acid and a third amino acid of the polypeptide. In certain embodiments, the first amino acid is a terminal amino acid, and the second and third amino acids are downstream of the first amino acid. In certain embodiments, the first amino acid is an internal amino acid, and the second and third amino acids are independently upstream or downstream of the first amino acid. In some examples, the second amino acid and / or the third amino acid are adjacent to the first amino acid. In some cases, for example, the second amino acid and / or the third amino acid are 10 amino acids or less, 5 amino acids or less, 4 amino acids or less, 3 amino acids or less, 2 amino acids or less, 1 amino acid or less, or 0 amino acids away from the first amino acid (i.e., the first amino acid is directly adjacent to the second amino acid and / or the third amino acid). In some cases, the second amino acid and / or the third amino acid are separated from the first amino acid by at least 1 amino acid, at least 2 amino acids, at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, or at least 10 amino acids. In some cases, the second amino acid and / or the third amino acid are separated from the first amino acid by 1-2 amino acids, 1-3 amino acids, 1-4 amino acids, 1-5 amino acids, 1-10 amino acids, 2-3 amino acids, 2-4 amino acids, 2-5 amino acids, 2-10 amino acids, 3-5 amino acids, 3-10 amino acids, or 5-10 amino acids. In certain cases, the second amino acid is adjacent to the third amino acid. The second amino acid may or may not be adjacent to the third amino acid. In some embodiments, the second amino acid is separated from the third amino acid by 10 amino acids or less, 5 amino acids or less, 4 amino acids or less, 3 amino acids or less, 2 amino acids or less, 1 amino acid or less, or 0 amino acids (i.e., the second amino acid is directly adjacent to the third amino acid).In some cases, the second amino acid is separated from the third amino acid by at least 1 amino acid, at least 2 amino acids, at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, or at least 10 amino acids. In some cases, the second amino acid is separated from the third amino acid by 1-2 amino acids, 1-3 amino acids, 1-4 amino acids, 1-5 amino acids, 1-10 amino acids, 2-3 amino acids, 2-4 amino acids, 2-5 amino acids, 2-10 amino acids, 3-5 amino acids, 3-10 amino acids, or 5-10 amino acids.
[0021] In some embodiments, the determined one or more chemical properties include the identity of a first amino acid, a second amino acid, and / or a third amino acid of the polypeptide. In some embodiments, the determined one or more chemical properties include a modification (e.g., post-translational modification, mutation, conjugation to a binding entity) of the first amino acid, the second amino acid, and / or the third amino acid of the polypeptide. In certain embodiments, the determined one or more chemical properties can be used to identify the first amino acid, the second amino acid, and / or the third amino acid. In certain embodiments, the identified amino acids can be used to identify the protein from which the polypeptide is derived.
[0022] In some embodiments, the compositions, methods, and systems of the present disclosure may be utilized in dynamic peptide sequencing reactions. In this technique, structural information of a polypeptide may be determined by evaluating single molecule binding interactions between an amino acid recognition factor and a polypeptide while amino acids are progressively cleaved from the end of the polypeptide. FIG. 1 shows an example of a dynamic peptide sequencing reaction in which each on-off binding event produces a signal pulse in the signal output. As shown on the left, a protein sample may be fragmented into polypeptides that are immobilized in the reaction chamber of an array, and the immobilized polypeptides are exposed to one or more amino acid recognition factors and one or more cleavage agents (e.g., aminopeptidases). As shown on the right, the amino acid recognition factors reversibly bind to the end of the polypeptide, and a detectable signal is generated while the recognition factor binds to the polypeptide. Because the on-off binding of the recognition factor generally occurs at a faster rate than amino acid cleavage, the binding event preceding the amino acid cleavage produces a series of pulses in the signal output that may be used to determine structural information about the amino acids of the polypeptide.
[0023] Compositions, systems and methods for performing dynamic polypeptide sequencing and analyzing data obtained therefrom are more fully described in International Publication No. WO 2020102741(A1), filed November 15, 2019, and International Publication No. WO 2021236983(A2), filed May 20, 2021, each of which is incorporated by reference in its entirety.
[0024] As used herein, in some embodiments, the term "bond" or "bonds" refers to any non-covalent (e.g., hydrogen bonds, van der Waals interactions, aromatic interactions, electrostatic interactions) or covalent interactions between a particular binding component or any plurality thereof, and the terms "bind," "binding," "bound," and the like refer to the formation and / or existence of any such bonds. As an illustrative example, a binding event between an amino acid recognition factor and an amino acid can include the formation of one or more non-covalent or covalent interactions between the amino acid recognition factor and the amino acid.
[0025] In some embodiments, the term includes identifying one or more amino acids of a polypeptide. As used herein, in some embodiments, the terms "identifying," "determining identity," and the like, with respect to an amino acid, include determining the definite identity of an amino acid, as well as determining the probability of the definite identity of an amino acid. For example, in some embodiments, an amino acid is identified by determining the probability (e.g., 0% to 100%) that the amino acid is of a particular type, or by determining the probability for each of a plurality of particular types. Thus, in some embodiments, the terms "amino acid sequence," "polypeptide sequence," and "protein sequence" as used herein may refer to the polypeptide or protein material itself, and are not limited to the specific sequence information (e.g., a sequence of letters representing the order of amino acids from one end to another) that biochemically characterizes a particular polypeptide or protein.
[0026] Exemplary Techniques for Obtaining Information About Amino Acids As described herein, the inventors have developed a technique for obtaining information about multiple amino acids in a polypeptide based on a series of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and amino acids of the polypeptide. Figures 12A-12G show an exemplary method for determining and using kinetic signature information to characterize a polypeptide.
[0027] The methods described herein may be implemented by a system. For example, in some embodiments, the system comprises at least one non-transitory computer-readable medium having encoded thereon instructions that, when executed, cause a processor to perform one or more of the methods described herein. In some embodiments, the system further comprises a processor. The system may comprise any of the components of the integrated device described herein.
[0028] The method may facilitate obtaining information about multiple amino acids. For example, a polypeptide comprising a chain of amino acids may be used in the techniques described herein. The chain of amino acids may include at least one amino acid to which a dye-labeled recognition factor is attached. In some embodiments, the chain of amino acids includes a terminal amino acid and one or more downstream amino acids (e.g., amino acids at positions 1, 2, 3, 4, and / or 5 relative to the polypeptide terminus). In some embodiments, one or more amino acid recognition factors may be attached to the terminal amino acid. In some embodiments, one or more amino acid recognition factors may be attached to one or more amino acids downstream of the terminal amino acid in addition to the terminal amino acid of the peptide. In some embodiments, one or more amino acid recognition factors may be attached to an internal amino acid and one or more amino acids upstream or downstream of the internal amino acid.
[0029] A polypeptide can contain any number of amino acids. In some embodiments, a polypeptide contains at least 5 amino acids, at least 10 amino acids, at least 15 amino acids, at least 20 amino acids, at least 50 amino acids, or at least 100 amino acids. In some embodiments, a polypeptide contains 5-10, 5-15, 5-20, 5-50, 5-100, 10-15, 10-20, 10-50, 10-100, 15-20, 15-50, 15-100, 20-50, 20-100, or 50-100 amino acids.
[0030] In some embodiments, to obtain information about a chain of amino acids, a sample containing at least a portion of a polypeptide (e.g., all or a fragment thereof) may be loaded into an integrated device, such as an integrated device described herein. In particular, the polypeptide may be loaded into a reaction chamber of the integrated device. In some cases, the polypeptide may be bound to a surface of the chamber via a covalent or non-covalent bond (e.g., streptavidin-biotin bond, click chemistry bond) that immobilizes the polypeptide within the chamber.
[0031] In some embodiments, multiple polypeptides may be loaded onto an integrated device, and multiple chambers of an integrated device may receive one or more polypeptides. The techniques described herein for obtaining information about polypeptides may be performed in a parallel manner (e.g., concurrently, simultaneously).
[0032] 12A shows an exemplary method 1200 for determining a chemical property of a polypeptide. In some embodiments, method 1200 may begin with operation 1202. In some embodiments, one or more additional or alternative operations may be performed before operation 1202, such as any of the loading steps described herein.
[0033] In operation 1202, the polypeptide may be contacted with one or more amino acid recognition factors. As described herein, the polypeptide may include multiple amino acids. In some embodiments, the polypeptide includes a first amino acid (e.g., a terminal amino acid, an internal amino acid) to which one or more amino acid recognition factors may bind, and at least one other (e.g., upstream, downstream) amino acid (e.g., a second amino acid).
[0034] The one or more amino acid recognition factors may constitute a first set of one or more amino acid recognition factors that bind to the polypeptide. In some embodiments, the first set of one or more amino acid recognition factors may bind to an amino acid of the polypeptide, and in some embodiments, may bind to one or more additional amino acids. In certain embodiments, the first set of one or more amino acid recognition factors may bind to a terminal amino acid of the polypeptide, and in some embodiments, may bind to one or more downstream amino acids. In certain embodiments, the first set of one or more amino acid recognition factors may bind to an internal amino acid of the polypeptide, and in some embodiments, may bind to one or more upstream and / or downstream amino acids. At least one (and in some embodiments each) of the one or more amino acid recognition factors may be labeled with a fluorescent dye that emits an emission light when excited with an excitation light, as described herein. In some cases, the one or more amino acid recognition factors include multiple types of amino acid recognition factors. In certain cases, each type of amino acid recognition factor can only bind to a certain amino acid. As an illustrative example, the first type of amino acid recognition factor may preferentially bind to leucine, isoleucine, and valine. As another illustrative example, the second type of amino acid recognition factor may preferentially bind to phenylalanine, tyrosine, and tryptophan. As another illustrative example, the third type of amino acid recognition factor may preferentially bind to arginine. In some embodiments, the first set of one or more amino acid recognition factors includes one type of amino acid recognition factor. In some embodiments, the first set of one or more amino acid recognition factors includes two or more types of amino acid recognition factors. In some embodiments, each type of recognition factor is labeled with a unique dye and / or several unique dyes. Thus, the emitted light from the dye-labeled amino acid recognition factors can be used to obtain information (e.g., identify) about the amino acid to which the dye-labeled amino acid recognition factor is bound, and in some embodiments, about one or more additional amino acids.
[0035] In some embodiments, contacting the polypeptide with one or more amino acid recognition factors includes introducing one or more amino acid recognition factors onto the device (e.g., by loading a solution containing one or more amino acid recognition factors onto an integrated device containing the polypeptide). The one or more amino acid recognition factors may periodically bind to the polypeptide (e.g., to at least one amino acid of the polypeptide). The rate at which the one or more amino acid recognition factors bind to the polypeptide is referred to herein as the binding rate. In some embodiments, the one or more amino acid recognition factors may be labeled (e.g., conjugated) with a fluorescent dye that can be excited when the one or more recognition factors bind to or near an amino acid of the polypeptide. Thus, the periodic signal emitted by the fluorescent dye may be characteristic of the binding rate of the amino acid recognition factors.
[0036] In operation 1204, a first series of signal pulses can be detected. For example, the integrated device may include one or more light detection regions that detect light emitted by the reaction chamber, and more specifically, by the fluorescent dye excited therein. As described herein, the fluorescent dye labeling one or more amino acid recognition factors can be excited with excitation light (e.g., from at least one light source, such as a pulsed laser). When excited, electrons of the fluorescent dye absorb energy from the excitation light and move to a higher energy level. After a period of time, the electrons return to a ground state. In returning to the ground state, the electrons release energy in the form of photons. The emitted photons (also referred to herein as emitted light or signal) can be detected by the integrated device described herein. The signal pulses detected by the integrated device include information characteristic of the sample to which the fluorescent dye is bound (e.g., pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, whether an amino acid is recognized). In particular, the signal pulses include information characteristic of the fluorescent dye. Each type of amino acid recognition agent can be conjugated to a unique combination of one or more fluorescent dyes, and thus a signal pulse can be correlated to one or more amino acids and used to obtain information about the sample.
[0037] The first series of signal pulses detected in operation 1204 may indicate a first series of binding events between a first set of one or more dye-labeled amino acid recognition factors and at least one amino acid of the polypeptide. As described herein, the one or more dye-labeled amino acid recognition factors may periodically bind to an amino acid (e.g., a terminal amino acid, an internal amino acid) and may emit a signal when bound to the amino acid. Thus, the first series of signal pulses may indicate a series of binding events between the one or more dye-labeled amino acid recognition factors and an amino acid, in some embodiments, one or more upstream amino acids and / or downstream amino acids.
[0038] In operation 1206, at least one chemical property of the polypeptide may be determined based on at least one property of the first series of signal pulses detected in operation 1204. In some embodiments, determining at least one chemical property of the polypeptide includes determining at least one chemical property of a first set of at least two amino acids of the polypeptide. In some embodiments, the at least two amino acids include an amino acid to which one or more amino acid recognition factors are bound and one or more other amino acids (e.g., one or more upstream and / or downstream amino acids). In some embodiments, the at least two amino acids include a terminal amino acid and one or more downstream amino acids.
[0039] As described herein, at least one characteristic of the series of signal pulses can be determined. Examples of the at least one characteristic include, but are not limited to, intensity, fluorescence lifetime, wavelength, pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, absence of signal pulse, and whether amino acid is recognized. In some embodiments, the at least one characteristic of the series of signal pulses comprises the average characteristic of the series of signal pulses.
[0040] In some embodiments, at least one characteristic of the series of signal pulses includes intensity (e.g., average intensity of the series of signal pulses). The intensity can be determined based on the amount of charge carriers detected in a light detection region that receives the emitted light from the fluorescent label. In some embodiments, the emitted light from a particular fluorescent label can have a characteristic intensity such that analyzing the intensity information of the emitted light can facilitate the identification of one or more chemical properties of the polypeptide.
[0041] In some embodiments, at least one characteristic of the train of signal pulses includes wavelength (e.g., the average wavelength of the train of signal pulses). The wavelength of the emitted light may be determined in any suitable manner, for example, by using one or more optical filters and / or light detection regions disposed at different depths. In some embodiments, the emitted light from a particular fluorescent label may have a characteristic wavelength such that analyzing the wavelength information of the emitted light may facilitate identification of one or more chemical properties of the polypeptide.
[0042] In some embodiments, at least one characteristic of the train of signal pulses includes a fluorescence lifetime (e.g., an average fluorescence lifetime of the train of signal pulses). In some embodiments, the fluorescent label, when excited by incident excitation light, emits fluorescence with a characteristic lifetime (e.g., a characteristic emission decay period) such that analyzing the lifetime information of the emitted light can facilitate identification of one or more chemical characteristics of the polypeptide. Fluorescence lifetime, also referred to simply as "lifetime" herein, is a measure of the time a fluorescent dye spends in an excited state before returning to the ground state and emitting a photon. In some embodiments, the fluorescence lifetime information and / or other timing characteristics described herein may be obtained through techniques for time binning charge carriers generated by photons incident on a light detection region (e.g., a photodiode).
[0043] In some embodiments, at least one characteristic of the series of signal pulses includes a pulse duration (e.g., average pulse duration), also referred to herein as pulse width. Pulse duration refers to the time interval measured across a pulse. In some embodiments, the pulse width is measured at full width at half maximum of the pulse. As described herein, the dye-labeled amino acid recognition factor periodically binds and dissociates from the polypeptide (e.g., to an amino acid of the polypeptide). When bound, the dye-labeled amino acid recognition factor can be excited and emit an emitted light. The average duration of each signal pulse emitted by the dye-labeled amino acid recognition factor includes the pulse duration of the fluorescent label. In certain embodiments, for example, at least one characteristic of the first series of signal pulses includes a first pulse duration, and the first pulse duration includes the average duration of each pulse of the first series of signal pulses.
[0044] In some embodiments, at least one characteristic of the series of signal pulses includes an interpulse duration (e.g., an average interpulse duration). Interpulse duration, also referred to herein as interpulse width, refers to the time interval between adjacent pulses. As described herein, the dye-labeled amino acid recognition factor periodically binds and dissociates from the polypeptide (e.g., to an amino acid of the polypeptide). When bound, the dye-labeled amino acid recognition factor can be excited and emit an emission light. The average duration between signal pulses emitted by the fluorescent label includes the interpulse duration of the fluorescent label. In certain embodiments, for example, at least one characteristic of the first series of signal pulses includes a first interpulse duration, and the first interpulse duration includes the average duration between each pulse of the first series of signal pulses.
[0045] In some embodiments, at least one characteristic of the train of signal pulses includes a recognition segment (RS) duration. A recognition segment generally refers to a train of signal pulses indicative of a train of binding events between a type of amino acid recognition factor (e.g., one or more molecules of that type of amino acid recognition factor) and one or more amino acids of a polypeptide. In some cases, for example, a first recognition segment includes a first train of signal pulses indicative of a train of binding events between a first set of one or more amino acid recognition factors and a first amino acid (in some cases, one or more additional amino acids) of a polypeptide. In some cases, a second recognition segment includes a second train of signal pulses indicative of a train of binding events between a second set of one or more amino acid recognition factors and a second amino acid (in some cases, one or more additional amino acids) of a polypeptide. A recognition segment duration generally refers to the length of time over which a train of signal pulses is received (i.e., the duration of a recognition segment). In some cases, for example, a first recognition segment can have a first recognition segment duration that includes the length of time over which a first train of signal pulses is received. In some cases, the second recognition segment can have a second recognition segment duration that includes the length of time during which the second series of signal pulses is received.
[0046] In some embodiments, at least one characteristic of the series of signal pulses includes an inter-segment duration. Inter-segment duration generally refers to the duration between two recognition segments. In certain embodiments, for example, a first inter-segment duration includes the length of time between a first recognition segment and a second recognition segment. The first and second recognition segments may have different characteristics, which may allow the recognition segments to be distinguished from each other, as described herein.
[0047] In some embodiments, at least one characteristic of the series of signal pulses includes a cleavage rate (e.g., average cleavage rate) and / or a cleavage time. In some embodiments, for example, a terminal amino acid of a polypeptide disposed in a reaction chamber is cleaved from the polypeptide. In certain embodiments, cleavage of the terminal amino acid is performed by exposing the polypeptide to a cleavage agent (e.g., one or more aminopeptidases). In some embodiments, the cleavage agent may be included in the same solution as the amino acid recognition factor. The cleavage of an amino acid (e.g., a terminal amino acid) from a polypeptide may be referred to as a cleavage event. The cleavage rate may refer to the number of cleavage events per unit time. The cleavage time may refer to the length of time between cleavage events. In some embodiments, the cleavage event may be determined by observing a change from one recognition segment to another recognition segment (e.g., based on different properties of the recognition segment), a change from a recognition segment to a non-recognition segment (i.e., a segment where no signal pulse is received), and / or a change from a non-recognition segment to a recognition segment.
[0048] In some embodiments, the at least one characteristic comprises the absence of a signal pulse at one or more reference time points.As an illustrative example, the polypeptide may have a predicted sequence, including a first amino acid that the first amino acid recognition factor preferentially binds to.In certain embodiments, the predicted sequence of binding events between the first amino acid and the first amino acid recognition factor may not occur, and the signal pulse may not be present at one or more reference time points.In some cases, the absence of a signal pulse may indicate the presence of a modification (e.g., a post-translational modification of the first amino acid, a mutation to a wild-type protein, binding to a binding component).
[0049] One or more of the characteristics of the series of signal pulses can be used to determine at least one chemical property of a first set of at least two amino acids of a polypeptide. In some embodiments, the first set of at least two amino acids includes a terminal amino acid and at least one downstream amino acid of the polypeptide. In some embodiments, the first set of at least two amino acids includes an internal amino acid of the polypeptide and at least one upstream and / or downstream amino acid. In some embodiments, the at least two amino acids of the polypeptide include an amino acid to which the dye label recognition factor binds and at least one other amino acid (e.g., one or more upstream and / or downstream amino acids). In some embodiments, the first set of at least two amino acids does not consist of the terminal amino acid and the penultimate amino acid of the polypeptide (i.e., the terminal amino acid and the amino acid immediately adjacent). In some embodiments, the at least two amino acids include at least three amino acids, at least four amino acids, at least five amino acids, at least 10 amino acids, at least 15 amino acids, at least 20 amino acids, at least 50 amino acids, or at least 100 amino acids. In some embodiments, at least two amino acids includes 2 amino acids, 3 amino acids, 4 amino acids, 5 amino acids, 10 amino acids, 15 amino acids, 20 amino acids, 50 amino acids, 100 amino acids, etc. In some embodiments, the at least two amino acids include 2-3 amino acids, 2-4 amino acids, 2-5 amino acids, 2-10 amino acids, 2-15 amino acids, 2-20 amino acids, 2-50 amino acids, 2-100 amino acids, 3-5 amino acids, 3-10 amino acids, 3-15 amino acids, 3-20 amino acids, 3-50 amino acids, 3-100 amino acids, 5-10 amino acids, 5-15 amino acids, 5-20 amino acids, 5-50 amino acids, 5-100 amino acids, 10-15 amino acids, 10-20 amino acids, 10-50 amino acids, 10-100 amino acids, 20-50 amino acids, 20-100 amino acids, or 50-100 amino acids.
[0050] In some embodiments, the at least one chemical property comprises an amino acid identity (e.g., the identity of one or more of the first set of at least two amino acids). In some embodiments, the at least one chemical property can be used to determine the identity of one or more of the first set of at least two amino acids.
[0051] In some embodiments, the at least one chemical property comprises a structural property of an amino acid, including one or more structural properties of the first set of at least two amino acids (e.g., whether the amino acid comprises a modification, what type of modification the amino acid comprises). In some embodiments, the modification comprises a post-translational modification, a non-natural modification, an oxidative modification, a cross-linking modification, and / or a chemical modification. In some embodiments, the modification comprises one or more mutations compared to the wild-type protein. In some embodiments, the modification comprises one or more insertions compared to the wild-type protein. In some embodiments, the modification comprises one or more deletions compared to the wild-type protein. In some embodiments, the modification comprises a covalent or non-covalent bond to a binding entity (e.g., a nucleic acid, a linker, an antibody). In some embodiments, the modification affects at least one property of the first series of signal pulses (e.g., pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, absence of a signal pulse, whether an amino acid is recognized). The effect of the modification on at least one property of the first series of signal pulses allows the modification to be identified based on the first series of signal pulses.
[0052] Thus, in some embodiments, determining at least one chemical property of the polypeptide comprises identifying one or more amino acids of the polypeptide. In some embodiments, identifying the amino acids comprises determining which of the 20 naturally occurring amino acids are present. In some embodiments, the identity of the amino acid is selected from the group consisting of alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine. In some embodiments, determining at least one chemical property of the first set of at least two amino acids of the polypeptide comprises identifying at least one (in some embodiments, each) amino acid of the first set of at least two amino acids.
[0053] In some embodiments, determining at least one chemical property of the polypeptide comprises determining a subset of potential amino acids that may be present in the polypeptide. In some embodiments, this can be accomplished by determining that the amino acid is not one or more specific amino acids (and therefore may be any of the other amino acids). In some embodiments, this can be accomplished by determining which of a particular subset of amino acids (e.g., based on size, charge, hydrophobicity, post-translational modification, binding properties) may be present in the polypeptide (e.g., using a recognition agent that binds to a particular subset of two or more amino acids). In some embodiments, determining at least one chemical property of a first set of at least two amino acids of the polypeptide comprises determining that at least one (in some embodiments, each) amino acid of the first set of at least two amino acids is not one or more specific amino acids.
[0054] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that an amino acid of the polypeptide comprises a post-translational modification. The post-translational modification can affect the series of signals emitted by a dye-labeled amino acid recognition factor bound to the polypeptide (e.g., a terminal amino acid and / or an internal amino acid). In some embodiments, the series of signals emitted by the dye-labeled amino acid recognition factor can be affected by the post-translational modification even if the post-translational modification is for an amino acid that is not bound to the dye-labeled amino acid recognition factor. In some embodiments, the post-translational modification of the amino acid bound by the dye-labeled recognition factor and / or the post-translational modification of one or more upstream or downstream amino acids can change (e.g., increase, decrease) at least one characteristic of the series of signal pulses (e.g., pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, absence of signal pulse) compared to the unmodified amino acid. Non-limiting examples of post-translational modifications include acetylation (e.g., acetylated lysine), ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation (e.g., glycosylated asparagine), O-linked glycosylation (e.g., glycosylated serine, glycosylated threonine), hydroxylation, methylation (e.g., methylated lysine, methylated arginine), myristoylation (e.g., myristoylated glycine), NEDDylation, nitration (e.g., nitrated tyrosine), chlorination (e.g., chlorinated tyrosine), oxidation / reduction (e.g., oxidized cysteine), and the like. , oxidized methionine), carbonylation (e.g., carbonylated lysine, carbonylated proline, carbonylated arginine, carbonylated threonine), palmitoylation (e.g., palmitoylated cysteine), phosphorylation, prenylation (e.g., prenylated cysteine), S-nitrosylation (e.g., S-nitrosylated cysteine, S-nitrosylated methionine), sulfation (e.g., sulfated tyrosine), glycosylation (e.g., glycosylated lysine), SUMOylation (e.g., SUMOylated lysine), and ubiquitination (e.g., ubiquitinated lysine).
[0055] In some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that an arginine residue of the polypeptide comprises a post-translational modification. For example, as described herein, the amino acid recognition agent of the present disclosure can distinguish between different arginine modifications, including symmetric dimethylarginine (SDMA), asymmetric dimethylarginine (ADMA), and citrulline (also referred to as citrullinated arginine). In some embodiments, determining at least one chemical characteristic of the first set of at least two amino acids of the polypeptide comprises determining that at least one amino acid of the first set of at least two amino acids is a post-translational modified arginine (e.g., SDMA, ADMA, citrulline).
[0056] In some embodiments, determining at least one chemical property of the polypeptide comprises determining that an amino acid of the polypeptide comprises a phosphorylated side chain. For example, in some embodiments, determining at least one chemical property of the polypeptide comprises determining that the polypeptide comprises a phosphorylated threonine (e.g., phosphothreonine). In some embodiments, determining at least one chemical property of the polypeptide comprises determining that the polypeptide comprises a phosphorylated tyrosine (e.g., phosphotyrosine). In some embodiments, determining at least one chemical property of the polypeptide comprises determining that the polypeptide comprises a phosphorylated serine (e.g., phosphoserine). In some embodiments, determining at least one chemical property of a first set of at least two amino acids of the polypeptide comprises determining that at least one (in some embodiments, each) amino acid of the first set of at least two amino acids comprises a phosphorylated side chain. In certain embodiments, determining at least one chemical characteristic of the first set of at least two amino acids of the polypeptide comprises determining that at least one (and in some embodiments, each) amino acid of the first set of at least two amino acids is a phosphorylated threonine, a phosphorylated tyrosine, and / or a phosphorylated serine.
[0057] In some embodiments, determining at least one chemical characteristic of the polypeptide includes determining that the polypeptide comprises a chemically modified variant of an amino acid, an unnatural amino acid, and / or a proteinogenic amino acid (e.g., selenocysteine, pyrrolysine). Examples of unnatural amino acids include, but are not limited to, 2-naphthyl-alanine, statins, homoalanine, α-amino acids, β2-amino acids, β3-amino acids, γ-amino acids, 3-pyridyl-alanine, 4-fluorophenyl-alanine, cyclohexyl-alanine, N-alkyl amino acids, peptoid amino acids, homo-cysteine, penicillamine, 3-nitro-tyrosine, homo-phenylalanine, t-leucine, hydroxy-proline, 3-Abz, 5-F-tryptophan, and azabicyclo-[2.2.1]heptane. In certain embodiments, determining at least one chemical property of the first set of at least two amino acids of the polypeptide comprises determining that at least one (and in some embodiments, each) amino acid of the first set of at least two amino acids is a chemically modified variant of an amino acid, an unnatural amino acid, and / or a proteinogenic amino acid.
[0058] In some embodiments, determining at least one chemical characteristic of the polypeptide comprises determining that an amino acid of the polypeptide comprises an oxidative modification. For example, as described herein, the amino acid recognition factor of the present disclosure can distinguish between oxidized methionine and its unmodified variants. In some embodiments, the oxidative modification comprises an oxidatively damaged side chain of the amino acid. In some embodiments, the oxidatively damaged side chain is a cysteine-derived product (e.g., disulfide, sulfinic acid, sulfonic acid, sulfenic acid, S-nitrosocysteine), a tyrosine-derived product (e.g., dityrosine, 3,4-dihydroxyphenylalanine, 3-chlorotyrosine, 3-nitrotyrosine), a histidine-derived product (e.g., 2-oxohistidine, 4-hydroxy-2-oxohistidine, dihistidine, asparagine, aspartic acid, urea), a methionine-derived product, or a combination thereof. Oxidatively damaged amino acids include, but are not limited to, oxidatively damaged amino acids (e.g., sulfoxides, sulfones), tryptophan-derived products (e.g., ditryptophan, N-formylkynurenine, kynurenine, 2-oxo-tryptophan oxyindolylalanine, 6-nitrotryptophan, hydroxytryptophan), phenylalanine-derived products (e.g., meta-tyrosine, ortho-tyrosine), and global side chain products (e.g., alcohols, hydroperoxides, aldehyde / ketone carbonyls). Examples of oxidatively damaged amino acids are known in the art. See, for example, Hawkins, CL, Davies, MJ, "Detection, identification, and quantification of oxidative protein modifications," J Biol Chem. 2019 Dec 20; 294(51):19683-19708. In certain embodiments, determining at least one chemical property of the first set of at least two amino acids of the polypeptide comprises determining that at least one (and in some embodiments, each) amino acid of the first set of at least two amino acids comprises an oxidative modification.
[0059] In some embodiments, determining at least one chemical property of the polypeptide includes determining that an amino acid of the polypeptide comprises a side chain characterized by one or more biochemical properties. For example, the amino acid may comprise a non-polar aliphatic side chain, a positively charged side chain, a negatively charged side chain, a non-polar aromatic side chain, or a polar uncharged side chain at physiological pH. Non-limiting examples of amino acids comprising a non-polar aliphatic side chain include alanine, glycine, valine, leucine, methionine, and isoleucine. Non-limiting examples of amino acids comprising a positively charged side chain include lysine, arginine, and histidine. Non-limiting examples of amino acids comprising a negatively charged side chain include aspartic acid and glutamic acid. Non-limiting examples of amino acids comprising a non-polar aromatic side chain include phenylalanine, tyrosine, and tryptophan. Non-limiting examples of amino acids containing a polar, uncharged side chain include serine, threonine, cysteine, proline, asparagine, and glutamine.
[0060] In some embodiments, determining at least one chemical property of the polypeptide comprises identifying one or more mutations relative to the wild-type protein. Non-limiting examples of mutations include substitutions, insertions, and deletions. In certain embodiments, the one or more mutations comprise 2 or more, 3 or more, 4 or more, 5 or more, 10 or more, 15 or more, or 20 or more mutations. In certain embodiments, the one or more mutations comprise 2-3, 2-4, 2-5, 2-10, 2-15, 2-20, 3-4, 3-5, 3-10, 3-15, 3-20, 5-10, 5-15, 5-20, 10-15, 10-20, or 15-20 mutations. In some cases, determining at least one chemical property of the first set of at least two amino acids of the polypeptide comprises determining that at least one (in some embodiments, each) amino acid of the first set of at least two amino acids is mutated compared to the wild-type protein.
[0061] In some embodiments, determining at least one chemical property of the polypeptide includes determining that at least one amino acid is bound (e.g., via a covalent or non-covalent interaction) to a binding entity. Non-limiting examples of suitable binding entities include nucleic acids (e.g., DNA, RNA), linkers, and antibodies. In some examples, one or more amino acids of the polypeptide may be bound to a nucleic acid via one or more non-covalent interactions. In some examples, one or more amino acids of the polypeptide may be bound to a linker via one or more covalent interactions. In certain embodiments, determining at least one chemical property of the first set of at least two amino acids of the polypeptide includes determining that at least one (in some embodiments, each) amino acid of the first set of at least two amino acids is bound covalently or non-covalently to a binding entity.
[0062] In some embodiments, one or more characteristics of the first series of signal pulses indicative of a first series of binding events between a first set of one or more amino acid recognition factors and a first amino acid (e.g., terminal amino acid, internal amino acid) of the polypeptide can be influenced by one or more chemical properties of the polypeptide. In certain examples, one or more modifications of one or more amino acids (e.g., post-translational modification, mutation, binding to a binding component) can promote covalent or non-covalent interactions (e.g., via electrostatic attraction, π-stacking, hydrogen bond formation, etc.) between one or more amino acid recognition factors and the first amino acid, thereby increasing the pulse duration. In certain examples, one or more modifications of one or more amino acids (e.g., post-translational modification, mutation, presence of a binding component) can disrupt covalent or non-covalent interactions (e.g., via electrostatic repulsion, steric hindrance, etc.) between one or more amino acid recognition factors and the first amino acid, thereby decreasing the pulse duration.
[0063] In some embodiments, determining at least one chemical property of the polypeptide may include comparing at least one property of the series of signal pulses to known properties of known amino acid segments. For example, FIGS. 10A-10G show known pulse durations for various amino acid segments. In the illustrated embodiment, known pulse durations are shown for various tripeptide segments. Using a table, such as the table shown in FIGS. 10A-10G, and the properties (e.g., pulse durations) determined from the series of signal pulses may allow for the identification of amino acid segments (e.g., tripeptide segments, tetrapeptide segments). The table of known properties of amino acid segments may be constructed by theoretical means, by simulation, empirically, or by any combination thereof. The amino acid segments may have any length. In some embodiments, the table includes known properties of amino acid segments having 3 amino acids, 4 amino acids, 5 amino acids, 10 amino acids, 15 amino acids, or 20 amino acids. In some embodiments, the table includes known properties of amino acid segments having 3-4, 3-5, 3-10, 3-15, 3-20, 4-5, 4-10, 4-15, 4-20, 5-10, 5-15, 5-20, 10-15, 10-20, or 15-20 amino acids.
[0064] In some embodiments, the protein from which the polypeptide originates can be identified. For example, as described herein, the technology may include identifying one or more of the amino acids of an amino acid segment (e.g., a tripeptide segment, a tetrapeptide segment). Based on the identified amino acids, the protein from which the polypeptide originates can be identified. For example, identifying the protein from which the polypeptide originates can include comparing the identified amino acids of the amino acid segment with known information. In some embodiments, identifying the protein from which the polypeptide originates can include identifying a pattern in the amino acid segment that is also present in a candidate matching protein. The pattern can be unique to the candidate matching protein compared to other candidate matching proteins. Thus, the technology described herein may allow for identifying the protein from which the polypeptide originates based on identifying only a portion of the amino acids of the polypeptide. In some embodiments, identifying the polypeptide includes identifying the protein from which the polypeptide originates. In some embodiments, identifying the polypeptide includes identifying a pattern of amino acids present in the polypeptide, and identifying a candidate matching polypeptide that includes the pattern of amino acids.
[0065] The polypeptides described herein can be of any type. In some embodiments, the polypeptide comprises a fragment of a protein. In some embodiments, the polypeptide is derived from a biological source. In certain embodiments, the polypeptide is derived from the digestion of one or more proteins present in a biological sample (e.g., a human sample, a non-human animal sample, a plant sample). In some embodiments, the polypeptide is a recombinant polypeptide. In some embodiments, the polypeptide is a synthetic polypeptide.
[0066] In some embodiments, the protein is present in a biological sample (e.g., blood, plasma, tissue, saliva, urine, or other biological source). In certain cases, the biological sample is obtained from a human subject or a non-human animal subject. In certain cases, the biological sample is obtained from a plant, fungus, virus, or bacteria. The protein can be a wild-type or mutant protein. In some embodiments, the protein is a recombinant protein. In some embodiments, the protein is a synthetic protein.
[0067] In some embodiments, the protein is digested (e.g., by an enzyme or chemical reagent) to produce multiple polypeptides. Non-limiting examples of suitable reagents for enzymatic and / or chemical digestion include Lys-C, Arg-C, Asp-N, Lys-N, trypsin, chemotrypsin, BNPS-skatole, CNBr, caspase, formic acid, glutamyl endopeptidase, hydroxylamine, iodosobenzoic acid, neutrophil elastase, pepsin, proline endopeptidase, proteinase K, staphylococcal peptidase I, thermolysin, and thrombin.
[0068] In some embodiments, a solution containing a mixture of polypeptides can be introduced onto the integrated device. In some embodiments, the reaction chamber can receive at least one polypeptide. In some embodiments, the reaction chamber can receive at least two polypeptides, which may be different polypeptides. In some embodiments, the first polypeptide and the second polypeptide are disposed in different reaction chambers. A respective set of signals from each polypeptide can be obtained and used to obtain at least one chemical characteristic of the amino acids described herein.
[0069] It should be understood that determining at least one chemical property of the first set of at least two amino acids may, in some embodiments, be obtained based on multiple series of signal pulses. For example, one or more of the at least two amino acids may be identified and / or otherwise characterized based on a first series of signal pulses (e.g., a series of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and a first amino acid) and at least one additional series of signal pulses (e.g., a series of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and a second amino acid) as described herein. In some embodiments, the first series of signal pulses may be obtained when one or more amino acids are at a first position (e.g., a position other than a terminal position) of the chain of amino acids of the polypeptide, and the second series of signal pulses may be obtained when one or more amino acids are at a second position (e.g., a terminal position) of the chain of amino acids of the polypeptide different from the first position. In some embodiments, one or more additional series of signal pulses may be used to identify and / or otherwise characterize the at least two amino acids. Such techniques for multi-sampling the same amino acid may ensure greater accuracy in identifying and / or otherwise characterizing at least two amino acids.
[0070] In some embodiments, the method for determining a chemical property of a polypeptide (e.g., method 1200) further comprises detecting a second series of signal pulses indicative of a second series of binding events between the second set of one or more amino acid recognition factors and the polypeptide. In some embodiments, the method further comprises determining at least one chemical property of a second set of at least two amino acids of the polypeptide based on at least one property of the second series of signal pulses. In certain embodiments, the method further comprises identifying the polypeptide based on at least one chemical property of the second set of at least two amino acids of the polypeptide. In certain embodiments, the method further comprises identifying the polypeptide based on at least one chemical property of the first set of at least two amino acids and at least one chemical property of the second set of at least two amino acids.
[0071] In certain embodiments, the second set of at least two amino acids of the polypeptide comprises at least one amino acid of the first set of at least two amino acids. As an illustrative example, the first set of at least two amino acids may comprise a first amino acid (e.g., a terminal amino acid), a second amino acid, and a third amino acid, and the second set of at least two amino acids may comprise a second amino acid, a third amino acid, and a fourth amino acid.
[0072] The at least one characteristic of the second series of signal pulses can include any characteristic described herein (e.g., pulse duration, interpulse duration, recognition segment duration, intersegment duration, cleavage rate, cleavage time, whether an amino acid is recognized). In some embodiments, the at least one characteristic of the second series of signal pulses includes a second recognition segment duration (e.g., the length of time the second series of signal pulses are received).
[0073] In some embodiments, at least one characteristic of the second series of signal pulses is based in part on at least one characteristic of the first series of signal pulses. In certain embodiments, at least one characteristic of the second series of signal pulses includes a first inter-segment duration. In some cases, the first inter-segment duration includes a length of time between a first recognition segment where the first series of signal pulses are received and a second recognition segment where the second series of signal pulses are received. In certain embodiments, at least one characteristic of the second series of signal pulses includes an average of the first recognition segment duration and the second recognition segment duration. In certain embodiments, at least one characteristic of the second series of signal pulses includes an average of the first inter-segment duration and the second inter-segment duration. In some examples, the second inter-segment duration includes a length of time during which a third series of signal pulses is received indicative of a third series of binding events between the third set of one or more amino acid recognition factors and the polypeptide between the second recognition segment and a third recognition segment.
[0074] The at least one chemical property of the second set of at least two amino acids may include any chemical property described herein. In certain embodiments, determining the at least one chemical property of the second set of at least two amino acids includes identifying at least one (in some cases, each) amino acid of the second set of at least two amino acids. In certain embodiments, determining the at least one chemical property of the second set of at least two amino acids includes identifying a modification of at least one (in some cases, each) amino acid of the second set of at least two amino acids. In some examples, the modification includes a post-translational modification, a non-natural modification, an oxidative modification, a cross-linking modification, and / or a chemical modification. In some examples, the modification includes one or more mutations compared to the wild-type protein. In some examples, the modification includes a covalent or non-covalent bond between the at least one amino acid and a binding component (e.g., a nucleic acid, a linker, an antibody).
[0075] As described herein, the inventors have recognized that the characteristics of the signal pulse are influenced not only by the amino acid to which the dye-labeled amino acid recognition factor is bound, but also by one or more upstream and / or downstream amino acids that may or may not be bound to the dye-labeled amino acid recognition factor. Thus, in some embodiments, the signals from one or more series of signal pulses can be used to determine the chemical properties of a number of amino acids that is greater than the number of series of signal pulses used. In some embodiments, at least one of the amino acids may not be recognized (e.g., cannot be recognized) by any of the amino acid recognition factors present in the reaction chamber, meaning that no signal pulse would otherwise result from the dye-labeled amino acid recognition factor bound to the amino acid. Although at least one of the amino acids is not recognized, information about the amino acid can still be obtained from other series of signal pulses.
[0076] 12B illustrates an exemplary method 1210 for determining a chemical property of a polypeptide, where one or more amino acids of the polypeptide are unrecognized. The method 1210 may begin at operation 1212, where data is obtained during the degradation process of the polypeptide. In some embodiments, the data may include at least one series of signal pulses indicative of a series of binding events between the polypeptide and one or more amino acid recognition factors. In some embodiments, the series of binding events may be between one or more amino acid recognition factors and at least one amino acid of the polypeptide (e.g., a terminal amino acid exposed at an end of the polypeptide, an internal amino acid). The data may be obtained according to any of the techniques described herein.
[0077] In operation 1214, the acquired data may be analyzed to determine portions of the data. Each of the determined portions of the data may include a recognition segment, as described herein. For example, each of the determined portions of the data may correspond to an amino acid of the polypeptide during the degradation process (e.g., an amino acid exposed at a terminus of the polypeptide during the degradation process). The data may include at least one recognition segment and at least one non-recognition segment (e.g., a period during which a series of pulse segments are expected to be received, but no light is received due to, for example, an amino acid at the terminus of the polypeptide not being recognized). In some embodiments, the data may include a first portion corresponding to a first amino acid of the polypeptide. In certain embodiments, the first portion of the data may include a first plurality of signal pulses indicative of a series of binding events between a first type of amino acid recognition factor and the first amino acid. In some embodiments, the data may include a second portion corresponding to a second amino acid of the polypeptide. In certain embodiments, the second portion of the data may not include a signal pulse indicative of a binding event between any type of amino acid recognition factor and the second amino acid (e.g., due to the second amino acid not being recognized by one or more amino acid recognition factors).
[0078] In operation 1216, at least one chemical property of the first amino acid and / or the second amino acid may be determined based on at least one property of the first portion of data and at least one property of the second portion of data. The at least one property of each portion of data may include any of the properties described herein (e.g., pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, absence of a signal pulse, whether the amino acid is recognized). In some embodiments, the at least one property of the second portion of data includes the duration of the second portion of data (e.g., the duration of the absence of a signal pulse). The at least one chemical property of each amino acid may include any of the chemical properties described herein (e.g., identity, structural modification, presence of a binding entity). In certain embodiments, the at least one chemical property of the first amino acid and / or the second amino acid includes the identity of the first amino acid and / or the second amino acid. In certain embodiments, the at least one chemical property comprises a modification (e.g., post-translational modification, mutation, conjugation to a binding entity) of the first amino acid and / or the second amino acid. In some embodiments, operation 1216 comprises determining at least one chemical property of each of the first amino acid and the second amino acid.
[0079] As described herein, the techniques described herein can be used to identify amino acid characteristics based on known information. FIG. 12C shows an exemplary method 1220 for determining a chemical characteristic of a polypeptide. The method 1220 can begin with operation 1222, in which a first series of signal pulses is detected. The first series of signal pulses can indicate a first series of binding events between a first set of one or more amino acid recognition factors and the polypeptide. In some embodiments, the first series of signal pulses indicates a first series of binding events between a first set of one or more amino acid recognition factors and amino acids (e.g., terminal amino acids, internal amino acids) of the polypeptide.
[0080] At least one characteristic of the first series of signal pulses may be determined in operation 1224. For example, any of the characteristics of the signal pulses described herein (e.g., pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, absence of a signal pulse, whether an amino acid is recognized) may be determined in operation 1224.
[0081] In operation 1226, at least one characteristic of the first series of signal pulses can be compared to known characteristics of a plurality of amino acid segments comprising at least two amino acids. For example, the "heat maps" shown in FIGS. 10A-10G show known pulse durations for different known tripeptide segments. Operation 1226 can be performed using a table, such as the table shown in FIGS. 10A-10G, to compare at least one characteristic of the first series of signal pulses to known characteristics of a plurality of amino acid segments (e.g., tripeptide segments, tetrapeptide segments). The table of known characteristics of a plurality of amino acid segments can be constructed by theoretical means, by simulation, empirically, or by any combination thereof. The amino acid segments of the plurality of amino acid segments can have any suitable length. In some embodiments, one or more (and in some cases, all) of the amino acid segments of the plurality of amino acid segments have a length of at least 3 amino acids, at least 4 amino acids, at least 5 amino acids, at least 10 amino acids, at least 15 amino acids, or at least 20 amino acids. In some embodiments, one or more (and in some cases, all) of the amino acid segments of the plurality of amino acid segments have a length of 3 amino acids, 4 amino acids, 5 amino acids, 10 amino acids, 15 amino acids, or 20 amino acids. In some embodiments, one or more of at least two amino acids of an amino acid segment are contiguous. In some embodiments, one or more of at least two amino acids of an amino acid segment are non-contiguous (e.g., separated by one or more amino acids).
[0082] In operation 1228, at least one chemical property of at least two amino acids of the polypeptide may be determined based on the comparison. For example, any of the chemical properties described herein may be determined (e.g., identity of the amino acid, identity or presence of a modification). In some embodiments, determining the at least one chemical property of the at least two amino acids includes identifying at least one (and in some cases both) of the at least two amino acids. In some embodiments, determining the at least one chemical property of the at least two amino acids includes identifying a modification (e.g., post-translational modification, mutation, binding to a binding entity) of at least one (and in some cases both) of the at least two amino acids. The comparison may be performed according to any of the techniques described herein. For example, the comparison may be performed using an algorithm or by manual comparison. In some embodiments, the method may further include identifying the protein from which the polypeptide is derived based on the determined chemical property.
[0083] In some embodiments, additional series of signal pulses may be obtained. For example, the method may further include detecting a second series of signal pulses indicative of a second series of binding events between a second set of one or more amino acid recognition factors and the polypeptide. In certain embodiments, the second series of signal pulses may be indicative of a second series of binding events between a second set of one or more amino acid recognition factors and a subsequent amino acid of the polypeptide (e.g., a second amino acid that becomes the terminal amino acid after cleaving the first terminal amino acid). In some embodiments, at least one characteristic of the second series of signal pulses may be determined and compared to known characteristics of the plurality of amino acid segments. At least one chemical characteristic of the amino acid may be determined based on one or more characteristics of each of the first series of signal pulses and the second series of signal pulses.
[0084] 12D shows an exemplary method 1230 for identifying a protein from which a polypeptide is derived based on a pulse pattern that includes at least three recognition segments. The method 1230 may begin at 1232, where data is acquired during the degradation process of the polypeptide. The data may include a series of signal pulses (e.g., at least three recognition segments) indicative of respective binding events with at least three amino acids.
[0085] In operation 1234, the data can be analyzed to determine at least three portions of the data. In some embodiments, each portion corresponds to an amino acid of the polypeptide and includes a plurality of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and the amino acid. For example, in certain embodiments, the first portion of the data includes a first recognition segment and corresponds to a first amino acid. In some embodiments, the first portion of the data includes a plurality of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and the first amino acid. In certain embodiments, the second portion of the data includes a second recognition segment and corresponds to a second amino acid. In some embodiments, the second portion of the data includes a plurality of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and the second amino acid. In certain embodiments, the third portion of the data includes a third recognition segment and corresponds to a third amino acid. In some embodiments, the third portion of the data includes a plurality of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and the third amino acid. In some cases, the first amino acid is an amino acid exposed at a terminus of the polypeptide during the degradation process. In some cases, the first amino acid is an internal amino acid.
[0086] In act 1236, one or more characteristics of each of the at least three portions of data may be determined. The one or more characteristics may include any of the characteristics described herein (e.g., pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, absence of a signal pulse, whether an amino acid is recognized, etc.).
[0087] In operation 1238, the protein from which the polypeptide is derived may be identified based on the order of the at least three portions of data and one or more characteristics of each of the at least three portions of data. For example, one or more characteristics of the at least three portions of data may be used to identify at least three amino acids of the polypeptide. The identity of the at least three amino acids may be used, for example, in accordance with the techniques described herein to identify the protein from which the polypeptide is derived. In some embodiments, the at least three portions of data include at least four portions, at least five portions, at least six portions, at least seven portions, at least eight portions, at least nine portions, at least ten portions, at least fifteen portions, at least twenty portions, or at least fifty portions of data. In certain embodiments, the at least three portions of data include 3-4 portions, 3-5 portions, 3-10 portions, 3-15 portions, 3-20 portions, 3-50 portions, 5-10 portions, 5-15 portions, 5-20 portions, 5-50 portions, 10-15 portions, 10-20 portions, 10-50 portions, or 20-50 portions of data.
[0088] 12E illustrates an exemplary method 1240 for characterizing a second amino acid based on a pulse pattern emitted by one or more amino acid recognition factors bound to the first amino acid. The method 1240 may begin with operation 1242, in which a series of signal pulses indicative of a series of binding events between one or more amino acid recognition factors and the first amino acid of the polypeptide is detected. Detecting the series of signal pulses may be performed according to any of the techniques described herein.
[0089] In operation 1244, at least one characteristic of the series of signal pulses can be used to determine at least one chemical characteristic of a second amino acid of the polypeptide. The at least one characteristic of the series of signal pulses can be any characteristic described herein (e.g., pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, absence of signal pulse, whether the amino acid is recognized). The at least one chemical characteristic of the second amino acid can be any chemical characteristic described herein. Thus, the signal obtained based on the binding of one or more amino acid recognition factors to the first amino acid of the polypeptide can be used to identify (or otherwise characterize) the second amino acid of the polypeptide. The inventors have recognized that such techniques are particularly beneficial when the second amino acid cannot be recognized (e.g., by one or more amino acid recognition factors).
[0090] In some embodiments, determining at least one chemical property of the second amino acid comprises identifying the second amino acid, hi some embodiments, determining at least one chemical property of the second amino acid comprises identifying a modification (e.g., post-translational modification, mutation, binding to a binding entity) of the second amino acid.
[0091] A polypeptide can include a chain of amino acids that includes a first and a second amino acid. In some embodiments, the second amino acid is downstream of the first amino acid. In some embodiments, the second amino acid is upstream of the first amino acid. In some embodiments, the second amino acid is contiguous (e.g., adjacent) to the first amino acid. In some embodiments, the second amino acid is at least one amino acid (e.g., a third amino acid) away from the first amino acid in the chain of amino acids. In some embodiments, the second amino acid is at least 5 amino acids, at least 10 amino acids, at least 15 amino acids, or at least 20 amino acids away from the first amino acid. In some embodiments, the second amino acid is no more than 5 amino acids away from the first amino acid.
[0092] 12F illustrates an exemplary method 1250 for determining at least one chemical property of an amino acid of a polypeptide. The method 1250 may begin at operation 1252, where a first series of signal pulses indicative of a first series of binding events between a first set of one or more amino acid recognition factors and a first amino acid of the polypeptide is detected. Detecting the first series of signal pulses may be performed according to any of the techniques described herein.
[0093] A second series of signal pulses indicative of a second series of binding events between a second set of amino acid recognition factors and a second amino acid of the polypeptide can be detected in operation 1254. Detecting the second series of signal pulses can be performed according to any of the techniques described herein.
[0094] In operation 1256, at least a chemical property of the second amino acid may be determined based on at least one property of the first series of signal pulses and at least one property of the second series of signal pulses. As described herein, the inventors have recognized that multiple sampling of signal pulses for an amino acid may be advantageous. For example, multiple sampling may advantageously increase the accuracy of the identification and / or other characterization of the amino acid. In operation 1256, at least one chemical property of the second amino acid may be determined based on at least one property from each of the two signal pulses. The at least one property of each series of signal pulses may be the same in some embodiments and may be different in other embodiments. The property of the series of signal pulses may be any of the properties described herein, including but not limited to pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, absence of signal pulse, whether the amino acid is recognized, or any other property.
[0095] In some embodiments, additional series of signal pulses may be used. For example, in some embodiments, a third series of signal pulses may be detected that indicates a series of binding events between a third set of one or more amino acid recognition factors and a third amino acid of the polypeptide. Determining at least one chemical property of the second amino acid may be based on at least one property of each of the first, second, and third series of signal pulses.
[0096] 12G shows an exemplary method 1260 of identifying a disease or disorder in a subject. In some embodiments, the subject is a human subject. In some embodiments, the subject is a non-human animal subject.
[0097] The method 1260 may begin at operation 1262, where proteins in a sample from a subject may be digested to produce a plurality of polypeptides. The proteins may be any protein. Examples of proteins of interest include, but are not limited to, vimentin and β-amyloid protein. Protein digestion may be performed according to any of the enzymatic and / or chemical techniques described herein.
[0098] In operation 1264, a polypeptide of the plurality of polypeptides may be contacted with one or more amino acid recognition factors and a cleavage agent. The amino acid recognition factor may be any amino acid recognition factor described herein. The cleavage agent may be any cleavage agent described herein.
[0099] In operation 1266, one or more series of signal pulses indicative of a binding event between one or more amino acid recognition factors and the polypeptide are detected as amino acids are progressively cleaved from the end of the polypeptide by the cleavage agent. Detecting the series of signal pulses can be performed according to any of the techniques described herein.
[0100] In operation 1268, at least one characteristic of the one or more series of signal pulses can be used to determine at least one chemical property of the polypeptide. The at least one characteristic of the one or more series of signal pulses may include any property described herein. In some embodiments, the at least one characteristic includes pulse duration, inter-pulse duration, recognition segment duration, inter-segment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, whether an amino acid is recognized. In some embodiments, the at least one characteristic includes the absence of a signal pulse at one or more reference time points.
[0101] The at least one chemical property may include any chemical property described herein. In certain embodiments, the at least one chemical property is indicative of a modification of the protein. In some embodiments, the modification is a post-translational modification. The post-translational modification may be any post-translational modification described herein. In certain embodiments, the post-translational modification includes citrullination of at least one amino acid. In some examples, the at least one amino acid includes arginine. In certain embodiments, the post-translational modification includes methylation (e.g., demethylation) of at least one amino acid. In some examples, the at least one amino acid includes arginine and / or lysine. In certain embodiments, the post-translational modification includes phosphorylation of at least one amino acid. In some examples, the at least one amino acid includes threonine, tyrosine, and / or serine. In certain embodiments, the post-translational modification includes acetylation of at least one amino acid. In some examples, the at least one amino acid includes lysine. In certain embodiments, the post-translational modification includes oxidation of at least one amino acid. In some examples, the at least one amino acid includes methionine and / or cysteine. In some embodiments, the modification comprises one or more mutations compared to the wild-type protein.
[0102] In some embodiments, the modification of the protein is indicative of a disease or disorder in the subject. Non-limiting examples of diseases or disorders include cardiovascular disease, autoimmune disease, cancer, and / or neurodegenerative disease. In certain embodiments, the disease or disorder comprises an autoimmune disease. Non-limiting examples of autoimmune diseases include rheumatoid arthritis, Crohn's disease, lupus, and multiple sclerosis. In certain embodiments, the disease or disorder comprises cancer. Non-limiting examples of cancer include lung cancer, breast cancer, prostate cancer, skin cancer, brain cancer, oral cancer, gastrointestinal cancer, and colorectal cancer. In certain embodiments, the disease or disorder comprises a neurodegenerative disease. A non-limiting example of a neurodegenerative disease is Alzheimer's disease.
[0103] Amino acid recognition factor In some embodiments, the techniques described herein can be implemented using any amino acid recognition factor known in the art. For example, see International Publication No. 2020102741 (A1), filed November 15, 2019, and International Publication No. 2021236983 (A2), filed May 20, 2021, which describe amino acid recognition factors (e.g., recognition molecules) in detail, the relevant contents of which are incorporated by reference in their entirety.
[0104] In some embodiments, the amino acid recognition factor of the present disclosure comprises an amino acid binding protein having an amino acid sequence selected from Table 1. Table 1 herein provides a list of exemplary sequences of amino acid binding proteins. It is understood that these sequences and other examples described herein are meant to be non-limiting and that an amino acid recognition factor according to the present disclosure can include any homolog, variant, or fragment thereof that minimally contains the domain or subdomain involved in amino acid recognition.
[0105] In some embodiments, the amino acid binding protein has an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 1. In some embodiments, the amino acid binding protein has at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or more amino acid sequence identity to an amino acid sequence selected from Table 1. In some embodiments, the amino acid binding protein has 25-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-95%, 95-99%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, or 95-100% amino acid sequence identity to an amino acid sequence selected from Table 1.
[0106] For purposes of comparing two or more amino acid sequences, the percentage of "sequence identity" (also referred to herein as "amino acid identity") between a first amino acid sequence and a second amino acid sequence can be calculated by dividing the number of amino acid residues in the first amino acid sequence that are identical to amino acid residues at corresponding positions in the second amino acid sequence by the total number of amino acid residues in the first amino acid sequence and multiplying by 0100, where each deletion, insertion, substitution or addition of an amino acid residue in the second amino acid sequence compared to the first amino acid sequence is considered to be a difference at a single amino acid residue (position). Alternatively, the degree of sequence identity between two amino acid sequences can be calculated using known computer algorithms (e.g., by the local homology algorithm of Smith and Waterman (1970) Adv. Appl. Math. 2:482c, by the homology alignment algorithm of Needleman and Wunsch, J. Mol. Biol. (1970) 48:443, by the similarity search method of Pearson and Lipman. Proc. Natl. Acad. Sci. USA (1998) 85:2444, or by computerized implementations of algorithms available as Blast, Clustal Omega, or other sequence alignment algorithms), for example, using standard settings. Usually, for the purpose of determining the percentage of "sequence identity" between two amino acid sequences according to the calculation method outlined above, the amino acid sequence with the largest number of amino acid residues is referred to as the "first" amino acid sequence, and the other amino acid sequence is referred to as the "second" amino acid sequence.
[0107] Additionally or alternatively, two or more sequences may be evaluated for identity between sequences. The term "identical" or percent "identity" in the context of two or more amino acid sequences refers to two or more sequences or subsequences that are the same. Two sequences are "substantially identical" if they have a certain percentage of amino acid residues that are the same over a particular region or entire sequence (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical) when compared and aligned for maximum correspondence over a comparison window or designated region, as measured using one of the sequence comparison algorithms described above, or by manual alignment and visual inspection. Optionally, the identity exists over a region that is at least about 25, 50, 75, or 100 amino acids in length, or over a region that is 100-150, 150-200, 100-200, or more than 200 amino acids in length.
[0108] Additionally or alternatively, two or more sequences may be evaluated for alignment between sequences. The term "alignment" or percent "alignment" in the context of two or more amino acid sequences refers to two or more sequences or subsequences that are the same. Two sequences are "substantially aligned" if they have a certain percentage of amino acid residues that are the same over a particular region or entire sequence (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or 99.9% identical) when compared and aligned for maximum correspondence over a comparison window or designated region, as measured using one of the sequence comparison algorithms described above or by manual alignment and visual inspection. Optionally, the alignment exists over a region that is at least about 25, 50, 75, or 100 amino acids in length, or over a region that is 100-150, 150-200, 100-200, or 200 or more amino acids in length.
[0109] In some embodiments, the amino acid binding proteins include modified amino acid binding proteins and include one or more amino acid deletions, additions, or mutations relative to the sequences shown in Table 1. In some embodiments, the modified amino acid binding proteins include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more amino acid deletions, additions, or mutations (which may or may not be contiguous amino acids) relative to the sequences shown in Table 1.
[0110] Devices and Systems The method according to the present disclosure may be implemented in some aspects using a system that allows single molecule analysis. The system may include an integrated device and an instrument configured to interface with the integrated device. The integrated device may include an array of pixels, each pixel including a reaction chamber and at least one photodetector. The reaction chambers of the integrated device may be formed on or through a surface of the integrated device and may be configured to receive a sample disposed on the surface of the integrated device. Collectively, the reaction chambers may be considered an array of reaction chambers. The multiple reaction chambers may have a suitable size and shape such that at least a portion of the reaction chambers receive a single sample (e.g., a single molecule such as a polypeptide). In some embodiments, the number of samples in the reaction chambers may be distributed among the reaction chambers of the integrated device such that some reaction chambers contain one sample while other reaction chambers contain zero, two, or more samples.
[0111] Excitation light is provided to the integrated device from one or more light sources external to the integrated device. Optical components of the integrated device can receive the excitation light from the light source and direct the light to an array of reaction chambers of the integrated device to illuminate an illumination region within the reaction chamber. In some embodiments, the reaction chambers may have a configuration that allows a sample to be held in close proximity to a surface of the reaction chamber, which may facilitate delivery of the excitation light to the sample and detection of emission light from the sample. A sample positioned within the illumination region may emit emission light in response to being illuminated by the excitation light. For example, the sample may be labeled with a fluorescent label that emits light in response to achieving an excited state through illumination with the excitation light. The emission light emitted by the sample may then be detected by one or more photodetectors in pixels that correspond to the reaction chamber with the sample being analyzed. According to some embodiments, multiple samples can be analyzed in parallel when performed across an array of reaction chambers, which may range in number between about 10,000 pixels to 10,000,000 pixels.
[0112] The integrated device may include an optical system for receiving the excitation light and directing the excitation light among the reaction chamber array. The optical system may include one or more grating couplers configured to couple the excitation light to other optical components of the integrated device and direct the excitation light to the other optical components. For example, the optical system may include optical components that direct the excitation light from the grating coupler toward the reaction chamber array. Such optical components may include an optical splitter, an optical combiner, and a waveguide. In some embodiments, the one or more optical splitters may combine the excitation light from the grating coupler and deliver the excitation light to at least one waveguide. According to some embodiments, the optical splitter may have a configuration that allows for a substantially uniform delivery of the excitation light across all the waveguides such that each of the waveguides receives substantially the same amount of excitation light. Such embodiments may improve the performance of the integrated device by improving the uniformity of the excitation light received by the reaction chambers of the integrated device. For example, examples of suitable components for coupling excitation light to the reaction chamber and / or directing emitted light to a photodetector and for inclusion within an integrated device are described in U.S. patent application Ser. No. 14 / 821,688, filed Aug. 7, 2015, entitled "INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES," and U.S. patent application Ser. No. 14 / 543,865, filed Nov. 17, 2014, entitled "INTEGRATED DEVICE WITH EXTERNAL LIGHT SOURCE FOR PROBING, DETECTING, AND ANALYZING MOLECULES," both of which are incorporated by reference in their entireties. Examples of suitable grating couplers and waveguides that may be implemented in an integrated device are described in U.S. patent application Ser. No. 15 / 844,403, filed Dec. 15, 2017, entitled “OPTICAL COUPLER AND WAVEGUIDE SYSTEM,” which is incorporated by reference in its entirety.
[0113] Additional photonic structures may be disposed between the reaction chamber and the photodetector and configured to reduce or prevent excitation light from reaching the photodetector, which may otherwise contribute to signal noise when detecting the emission light. In some embodiments, the metal layer, which may act as a circuit for the integrated device, may also act as a spatial filter. Examples of suitable photonic structures may include spectral filters, polarizing filters, and spatial filters, and are described in U.S. Patent Application No. 16 / 042,968, filed July 23, 2018, entitled "OPTICAL REJECTION PHOTONIC STRUCTURES," and U.S. Provisional Patent Application No. 63 / 124,655, filed December 11, 2020, entitled "INTEGRATED CIRCUIT WITH IMPROVED CHARGE TRANSFER EFFICIENCY AND ASSOCIATED TECHNIQUES," both of which are incorporated by reference in their entirety.
[0114] Components located remotely from the integrated device can be used to position and align the excitation source with respect to the integrated device. Such components can include optical components including lenses, mirrors, prisms, windows, apertures, attenuators, and / or optical fibers. Additional mechanical components can be included in the instrument to allow control of one or more alignment components. Such mechanical components can include actuators, stepper motors, and / or knobs. Examples of suitable excitation sources and alignment mechanisms are described in U.S. Patent Application No. 15 / 161,088, filed May 20, 2016, entitled "PULSED LASER AND SYSTEM," which is incorporated by reference in its entirety. Another example of a beam steering module is described in U.S. Patent Application No. 15 / 842,720, filed December 14, 2017, entitled "COMPACT BEAM SHAPING AND STEERING ASSEMBLY," which is incorporated by reference herein. Additional examples of suitable excitation sources are described in U.S. patent application Ser. No. 14 / 821,688, filed Aug. 7, 2015, entitled “INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES,” which is incorporated by reference in its entirety.
[0115] The photodetector(s) disposed with each pixel of the integrated device may be configured and arranged to detect light emission from the corresponding reaction chamber of the pixel. Examples of suitable photodetectors are described in U.S. Patent Application No. 14 / 821,656, entitled "INTEGRATED DEVICE FOR TEMPORAL BINNING OF RECEIVED PHOTONS," filed August 7, 2015, and incorporated by reference in its entirety. In some embodiments, the reaction chambers and their respective photodetectors may be aligned along a common axis. In this manner, the photodetectors may overlap the reaction chambers within the pixel.
[0116] Characteristics of the detected emitted light may provide an indication for identifying a label associated with the emitted light. Such characteristics may include any suitable type of characteristics, including the arrival time of a photon detected by a photodetector, the amount of photons accumulated over time by a photodetector, and / or the distribution of photons across two or more photodetectors. In some embodiments, such characteristics may be any one or a combination of two or more of luminescence lifetime, luminescence intensity, brightness, absorption spectrum, emission spectrum, luminescence quantum yield, wavelength (e.g., peak wavelength), and signal characteristics (e.g., pulse duration, inter-pulse duration, change in signal intensity).
[0117] In some embodiments, the photodetector may have a configuration that allows for detection of one or more timing characteristics (e.g., luminescence lifetime) associated with the emission of the sample. The photodetector may detect a distribution of photon arrival times after a pulse of excitation light propagates through the integrated device, and the distribution of arrival times may provide an indication of the timing characteristics of the emission light of the sample (e.g., a proxy for the luminescence lifetime). In some embodiments, the one or more photodetectors provide an indication of the probability of emission light being emitted by the label (e.g., luminescence intensity). In some embodiments, multiple photodetectors may be sized and positioned to capture the spatial distribution of the emission light. The output signal from the one or more photodetectors may then be used to distinguish one label among multiple labels, and multiple labels may be used to identify the sample within the sample. In some embodiments, the sample may be excited by multiple excitation energies, and the emission light and / or timing characteristics of the emission light emitted by the sample in response to the multiple excitation energies may distinguish one label from multiple labels.
[0118] In operation, parallel analysis of samples in the reaction chambers is performed by exciting some or all of the samples in the chambers using excitation light and detecting signals from the sample emissions using photodetectors. Emission light from the samples may be detected by corresponding photodetectors and converted into at least one electrical signal. The electrical signal may be transmitted along conductive lines in the circuitry of the integrated device that may be connected to an instrument interfaced with the integrated device. The electrical signal may then be processed and / or analyzed. Processing or analysis of the electrical signal may be performed on a suitable computing device located either on the instrument or off the instrument.
[0119] The instrument may include a user interface for controlling the operation of the instrument and / or the integrated device. The user interface may be configured to allow a user to input information into the instrument, such as commands and / or settings used to control the instrument's functions. In some embodiments, the user interface may include buttons, switches, dials, and a microphone for voice commands. The user interface may allow a user to receive feedback regarding the instrument and / or the performance of the integrated device, such as information obtained by proper alignment and / or readout signals from a photodetector on the integrated device. In some embodiments, the user interface may provide feedback using a speaker to provide audible feedback. In some embodiments, the user interface may include indicator lights and / or a display screen to provide visual feedback to the user.
[0120] In some embodiments, the instrument may include a computer interface configured to connect with a computing device. The computer interface may be a USB interface, a FireWire interface, or any other suitable computer interface. The computing device may be any general-purpose computer, such as a laptop or desktop computer. In some embodiments, the computing device may be a server (e.g., a cloud-based server) accessible over a wireless network via a suitable computer interface. The computer interface may facilitate communication of information between the instrument and the computing device. Input information for controlling and / or configuring the instrument may be provided to the computing device and transmitted to the instrument via the computer interface. Output information generated by the instrument may be received by the computing device via the computer interface. The output information may include feedback regarding the performance of the instrument, the performance of the integrated device, and / or data generated from the readout signal of the photodetector.
[0121] In some embodiments, the instrument may include a processing device configured to analyze data received from the one or more photodetectors of the integrated device and / or transmit control signals to the excitation source. In some embodiments, the processing device may comprise a general-purpose processor, a specially adapted processor (e.g., a central processing unit (CPU), such as one or more microprocessors or microcontroller cores, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a custom integrated circuit, a digital signal processor (DSP), or a combination thereof). In some embodiments, the processing of data from the one or more photodetectors may be performed by both the processing device of the instrument and an external computing device. In other embodiments, the external computing device may be omitted and the processing of data from the one or more photodetectors may be performed only by the processing device of the integrated device.
[0122] According to some embodiments, an instrument configured to analyze a sample based on luminescence emission characteristics may detect differences in luminescence lifetimes and / or intensities between different luminescent molecules (e.g., fluorescent molecules) and / or differences in lifetimes and / or intensities of the same luminescent molecule in different environments. The inventors have recognized and appreciated that differences in luminescence emission lifetimes may be used to distinguish the presence or absence of different luminescent molecules and / or to distinguish different environments or conditions to which the luminescent molecules are exposed. In some cases, aspects of the system may be simplified by distinguishing luminescent molecules based on lifetimes (e.g., rather than emission wavelengths). As an example, wavelength-discriminating optics (e.g., wavelength filters, dedicated detectors for each wavelength, dedicated pulsed light sources at different wavelengths, and / or diffractive optics) may be reduced in number or eliminated when distinguishing luminescent molecules based on lifetimes. In some cases, a single pulsed light source operating at a single characteristic wavelength may be used to excite different luminescent molecules that emit within the same wavelength region of the optical spectrum but have measurably different lifetimes. Analytical systems that use a single pulsed light source, rather than multiple light sources operating at different wavelengths, to excite and distinguish different luminescent molecules that emit in the same wavelength region can be less complex to operate and maintain, more compact, and can be manufactured at lower cost.
[0123] Although an analysis system based on luminescence lifetime analysis may have certain advantages, the amount of information obtained by the analysis system and / or the detection accuracy may be increased by enabling additional detection techniques. For example, some embodiments of the system may be further configured to identify one or more characteristics of the sample based on the emission wavelength and / or the emission intensity. In some implementations, the luminescence intensity may additionally or alternatively be used to distinguish different luminescence labels. For example, some luminescence labels may emit at significantly different intensities or have significant differences in the probability of excitation (e.g., at least about 35% difference) even if their decay rates are similar. By referencing the binned signal to the measured excitation light, it may be possible to distinguish different luminescence labels based on intensity levels.
[0124] According to some embodiments, the different luminescence lifetimes can be distinguished by a photodetector configured to time-bin the luminescence emission events following excitation of the luminescent labels. The time binning can occur during a single charge accumulation cycle for the photodetector. A charge accumulation cycle is the interval between readout events during which photogenerated carriers accumulate in bins of the time binning photodetector. An example of a time binning photodetector is described in U.S. Patent Application No. 14 / 821,656, filed August 7, 2015, entitled "INTEGRATED DEVICE FOR TEMPORAL BINNING OF RECEIVED PHOTONS," which is incorporated herein by reference. In some embodiments, the time binning photodetector can generate charge carriers in a photon absorption / carrier generation region and directly transfer the charge carriers to charge carrier storage bins in a charge carrier storage region. In such embodiments, the time binning photodetector may not include a carrier transfer / capture region. Such a time binning photodetector may be referred to as a "direct binning pixel." An example of a time binning photodetector including a direct binning pixel is described in U.S. patent application Ser. No. 15 / 852,571, entitled “INTEGRATED PHOTODETECTOR WITH DIRECT BINNING PIXEL,” filed Dec. 22, 2017, which is incorporated herein by reference.
[0125] In some embodiments, different numbers of fluorophores of the same type can be attached to different reagents in a sample, so that each reagent can be identified based on emission intensity. For example, two fluorophores can be attached to a first labeled recognition molecule, and four or more fluorophores can be attached to a second labeled recognition molecule. Due to the different numbers of fluorophores, there can be different excitation and fluorophore emission probabilities associated with different recognition molecules. For example, there can be more emission events for the second labeled recognition molecule during a signal accumulation interval, so that the apparent intensity of the bin is significantly higher than that of the first labeled recognition molecule.
[0126] The inventors have recognized and appreciated that distinguishing biological or chemical samples based on fluorophore decay rates and / or fluorophore intensity may allow for simplification of optical excitation and detection systems. For example, optical excitation may be performed using a single wavelength source (e.g., a light source that produces one characteristic wavelength rather than multiple light sources, or a light source that operates at multiple different characteristic wavelengths). Furthermore, wavelength discrimination optics and filters may not be required in the detection system. Also, a single photodetector may be used for each reaction chamber to detect emissions from different fluorophores. The phrase "characteristic wavelength" or "wavelength" is used to refer to a central or dominant wavelength within a limited emission bandwidth (e.g., a central or peak wavelength within a 20 nm bandwidth output by a pulsed light source). In some cases, "characteristic wavelength" or "wavelength" may be used to refer to a peak wavelength within the full bandwidth of emission output by a source.
[0127] According to one aspect of the present disclosure, an exemplary integrated device may be configured to perform single molecule analysis in combination with the instruments described above. It should be understood that the exemplary integrated device described herein is intended to be exemplary, and that other integrated device configurations may be configured to perform any or all of the techniques described herein.
[0128] 13 shows a cross-sectional view of pixel 1-112 of integrated device 1-102. Pixel 1-112 includes a photodetection region, which may be a pinned photodiode (PPD), and a charge storage region, which may be a storage diode (SDO). In some embodiments, the photodetection region and the charge storage region may be formed in the semiconductor material of the pixel by doping a region of the semiconductor material. For example, the photodetection region and the charge storage region may be formed with the same conductivity type (e.g., n-type doping or p-type doping).
[0129] During operation of the pixel 1-112, excitation light can illuminate the reaction chamber 1-108 to cause incident photons, including fluorescent emission from the sample, to flow along the optical axis to the photodetection region PPD. As shown in FIG. 13, the pixel 1-112 can include a waveguide 1-220 configured to optically couple (e.g., evanescently couple) the excitation light from a grating coupler of an integrated device (not shown) to the reaction chamber 1-108. In response, the sample in the reaction chamber 1-108 can emit fluorescent light toward the photodetection region PPD. In some embodiments, the pixel 1-112 can also include one or more photonic structures 1-230, which can include one or more light removal structures, such as a spectral filter, a polarizing filter, and / or a spatial filter. For example, the photonic structures 1-230 can be configured to reduce the amount of excitation light reaching the photodetection region PPD and / or increase the amount of fluorescent emission reaching the photodetection region PPD. Also, as shown in pixel 1-112, pixel 1-112 may include one or more metal layers 1-240, which may be configured as filters and / or carry control signals from control circuitry configured to control the transfer gates, as described further herein.
[0130] In some embodiments, the pixel 1-112 may include one or more transfer gates configured to control the operation of the pixel 1-112 by applying an electrical bias to one or more semiconductor regions of the pixel 1-112 in response to one or more control signals. For example, when the transfer gate ST0 induces a first electrical bias in the semiconductor region between the photodetection region PPD and the storage region SD0, a transfer path (e.g., a charge transfer channel) may be formed in the semiconductor region. Charge carriers (e.g., photoelectrons) generated in the photodetection region PPD by incident photons flow along the transfer path to the storage region SD0. In some embodiments, the first electrical bias may be applied during a collection period in which charge carriers from the sample are selectively directed to the storage region SD0. Alternatively, when the transfer gate ST0 applies a second electrical bias to the semiconductor region between the photodetection region PPD and the storage region SD0, charge carriers from the photodetection region PPD may be blocked from reaching the storage region SD0 along the transfer path. In some embodiments, the drain gate REJ can provide a channel to the drain D to pull noise charge carriers generated in the photodetection region PPD by the excitation light away from the photodetection region PPD and the storage region SD0, such as during a rejection period before fluorescent emission photons from the sample reach the photodetection region PPD. In some embodiments, during a readout period, the transfer gate ST0 can provide a second electrical bias, and the transfer gate TX0 can provide an electrical bias to flow the charge carriers stored in the storage region SD0 to the readout region, which can be a floating diffusion (FD) region, for processing.
[0131] It should be appreciated that, according to various embodiments, the transfer gates described herein may comprise semiconductor materials and / or metals, and may include the gate of a field effect transistor (FET), the base of a bipolar junction transistor (BJT), or the like.
[0132] In some embodiments, the operation of the pixels 1-112 may include one or more collection sequences, each of which includes one or more rejection (e.g., drain) periods and one or more collection periods. In one example, a collection sequence performed according to one or more pulses of an excitation light source may begin with a rejection period, such as to discard charge carriers generated in the pixels 1-112 (e.g., in the photodetection region PD) in response to excitation photons from the light source. For example, excitation photons may arrive at the pixels 1-112 prior to the arrival of fluorescent emission photons from the reaction chamber. A transfer gate for the charge accumulation region may be biased to have low conductivity in a charge transfer channel that couples the charge accumulation region to the photodetection region, preventing transfer and accumulation of charge carriers in the charge accumulation region. A drain gate for the drain region may be biased to have high conductivity in a drain channel between the photodetection region and the drain region, facilitating drainage of charge carriers from the photodetection region to the drain region. A transfer gate for any charge storage region coupled to a photodetection region may be biased to have low conductivity between the photodetection region and the charge storage region, thereby preventing charge carriers from being transferred or stored in the charge storage region during the rejection period.
[0133] Following the rejection period, a collection period may occur in which charge carriers generated in response to incident photons are transferred to one or more charge accumulation regions. During the collection period, the incident photons may include fluorescent emission photons, resulting in the accumulation of fluorescent emission charge carriers in the charge accumulation region(s). For example, a transfer gate for one of the charge accumulation regions may be biased to have a high conductivity between the photodetection region and the charge accumulation region, facilitating the accumulation of charge carriers in the charge accumulation region. Any drain gate coupled to the photodetection region may be biased to have a low conductivity between the photodetection region and the drain region, such that charge carriers are not discarded during the collection period.
[0134] Some embodiments may include multiple rejection and / or collection periods in a collection sequence, such as a first rejection and collection period followed by a second rejection and collection period, with each pair of rejection and collection periods occurring in response to a pulse of excitation light. In one example, charge carriers generated in the light detection region during each collection period of a collection sequence (e.g., in response to multiple pulses of excitation light) may be collected in a single charge accumulation region. In some embodiments, the charge carriers collected in the charge accumulation region may be read out for processing before the next collection sequence. Alternatively or additionally, in some embodiments, the charge carriers collected in the first charge accumulation region during the first collection sequence may be transferred to a second charge accumulation region sequentially coupled to the first charge accumulation region and read out simultaneously with the next collection sequence. In some embodiments, the processing circuitry configured to read out the charge carriers from one or more pixels may be configured to determine one or more of luminescence intensity information, luminescence lifetime information, luminescence spectrum information, and / or any other mode of luminescence information associated with performing the techniques described herein.
[0135] In some embodiments, the first collection sequence may include transferring charge carriers generated in the light detection response to the excitation pulse to the charge accumulation region at a first time following each excitation pulse, and the second collection sequence may include transferring charge carriers generated in the light detection response to the excitation pulse to the charge accumulation region at a second time following each excitation pulse. For example, the number of charge carriers collected after the first and second times may indicate luminance lifetime information of the received light.
[0136] As further described herein, the pixels of the integrated circuit can be controlled to perform one or more collection sequences using one or more control signals from the control circuitry of the integrated circuit, for example, by providing control signals to the drain and / or transfer gate of the pixels of the integrated circuit. In some embodiments, the charge carriers can be read out from the FD region of each pixel during a readout pixel associated with each pixel and / or row or column of pixels for processing. In some embodiments, the FD region of the pixel can be read out using a correlated double sampling (CDS) technique.
[0137] Sequence information Table 1: Non-limiting example sequences of amino acid binding proteins
[0138] [Table 1-1]
[0139] [Table 1-2]
[0140] [Table 1-3]
[0141] [Table 1-4]
[0142] [Table 1-5]
[0143] [Table 1-6]
[0144] [Table 1-7]
[0145]
Table 1-8
[0146]
Table 1-9
[0147]
Table 1-10
[0148]
Table 1-11
[0149]
Table 1-12
[0150]
Table 1-13
[0151]
Table 1-14
[0152]
Table 1-15
[0153]
Table 1-16
[0154]
Table 1-17
[0155]
Table 1-18
[0156]
Table 1-19
[0157]
Table 1-20
[0158]
Table 1-21
[0159]
Table 1-22
[0160]
Table 1-23
[0161]
Table 1-24
[0162]
Table 1-25
[0163]
Table 1-26
[0164]
Table 1-27
[0165] [Table 1-28]
[0166] [Table 1-29]
[0167] [Table 1-30]
[0168] [Table 1-31]
[0169] [Table 1-32]
[0170] [Table 1-33] EXAMPLES
[0171] Example 1. Real-time dynamic single molecule protein sequencing on an integrated semiconductor device In this example, we demonstrate a dynamic sequencing-by-degradation approach in which single surface-immobilized peptide molecules are probed in real time by a mixture of dye-labeled N-terminal amino acid recognition factors. By measuring the fluorescence intensity, lifetime, and intermolecular dynamics of the recognition factors on the semiconductor chip, we demonstrate the ability to annotate amino acids and collectively identify peptide sequences. By exploiting the dynamics of binding, each recognition factor is able to uniquely identify multiple amino acids. Principles and processes for expanding the number of recognizable amino acids are also described herein. Furthermore, we show that the method is compatible with both synthetic peptides and native peptides isolated from recombinant human proteins, and is capable of detecting single amino acid changes and post-translational modifications. The results demonstrate a robust core technology that can serve as a highly accurate, sensitive, and scalable next-generation sequencing platform for proteins.
[0172] Measurement of the proteome provides deep and informative insights into biological processes. However, more sensitive methods are needed to fully understand the complex and dynamic state of the proteome in cells and the proteomic changes that occur in disease states and to make this information more accessible. The complex nature of the proteome and the chemical properties of proteins present several fundamental challenges to achieving the same comprehensive sensitivity, throughput, and adoption as DNA sequencing techniques. These challenges include the large number of different proteins (>10,000) and even larger numbers of proteoforms per cell; the very wide dynamic range of protein abundance in cells and biofluids and the lack of correlation with transcript levels; the cost and high detection limits of current mass spectrometry methods; and the inability to copy or amplify proteins. Methods that directly sequence single protein molecules have the potential to provide the greatest possible detection sensitivity, allowing single-cell input, digital quantification based on read counts, detection of post-translational modifications (PTMs) and low-abundance or aberrant proteoforms, and cost and throughput levels favorable for widespread adoption.
[0173] Here, we demonstrate a single-molecule protein sequencing approach and integrated system for massively parallel proteomics studies. In this approach, peptides are immobilized in nanoscale reaction chambers on a semiconductor chip and N-terminal amino acids (NAAs) are detected in real time with dye-labeled NAA recognition factors. Aminopeptidases sequentially remove individual NAA to expose subsequent amino acids for recognition, eliminating the need for complex chemicals and fluidics (Figure 1). We constructed a benchtop device with a 532 nm pulsed laser source for fluorescence excitation and electronics for signal processing (Figure 6A). The semiconductor chip uses intensity and fluorescence lifetime rather than emission wavelength for discrimination of dye labels. The recognition factors detect one or more types of NAA and provide information for peptide identification based on the temporal order of NAA recognition and on-off binding kinetics.
[0174] Using CMOS fabrication techniques, we constructed a custom time-domain-sensitive semiconductor chip with nanosecond precision, with fully integrated components for single-molecule detection, including photosensors, optical waveguide circuits, and a reaction chamber for biomolecule immobilization (Figure 1). An observation volume of less than 5 attoliters was achieved through evanescent illumination at the bottom of the reaction chamber from a nearby waveguide, enabling highly sensitive single-molecule detection in the presence of high freely diffusing dye concentrations (>1 μM).
[0175] The semiconductor chip uses a filterless system that rejects excitation light based on photon arrival time, achieving over 10,000-fold attenuation of the incident excitation light. Elimination of the need for an integrated optical filter layer improves the efficiency of fluorescence collection and allows for scalable manufacturing of the chip. To allow discrimination of fluorochrome labels attached to NAA recognition agents by their fluorescence lifetime and intensity, the chip rapidly alternates between early and late signal collection windows associated with each laser pulse, thereby collecting different parts of the exponential fluorescence lifetime decay curve. The relative signals in these collection windows (called "bin ratios") provide a reliable indication of the fluorescence lifetime (Figure 6B-F, and Materials and Methods).
[0176] For an NAA-binding protein to function as a recognition factor in this approach, the average lifetime of the bound recognition factor-peptide complex should be long enough (typically >120 ms) to generate detectable single molecule binding events. Proteins from the N-end rule adaptor family ClpS, which naturally bind to N-terminal phenylalanine, tyrosine, and tryptophan, were evaluated. Using PS610, a recognition factor derived from ClpS2 from A. tumefaciens, it was confirmed that this recognition factor detectably binds to immobilized peptides bearing these NAAs. Importantly, it was also determined that the kinetics of binding differed for each NAA. To demonstrate these properties, immobilized peptides containing the first N-terminal sequence FAA, YAA, or WAA were incubated with PS610 on separate chips and data were collected for 10 hours (Methods). NAA recognition was observed by PS610 and was characterized by continuous on-off binding during the incubation period, with different pulse durations (PD) for each peptide (Figure 2A). The median PD values were 2.51, 0.73, and 0.31 s for FAA, YAA, and WAA, respectively. These values reflect the differences in binding affinity driven by different dissociation rates for each type of protein-NAA interaction (Figures 7A-B).
[0177] To expand the set of recognizable NAAs, N-end rule pathway proteins were explored as a source of additional recognition factors. In a comprehensive screen of diverse ClpS family proteins, a group of ClpS proteins from the bacterial phylum Planctomycetes was discovered that have native binding to N-terminal leucine, isoleucine, and valine. Directed evolution techniques were applied to generate a Planctomycetes ClpS mutant, PS961, with submicromolar affinity for N-terminal leucine, isoleucine, and valine, demonstrating recognition of these NAAs (Figure 2B). The median PDs for binding to peptides with N-terminal LAA, IAA, and VAA were 1.21, 0.28, and 0.21 s, respectively, consistent with bulk characterization (Figure 7C).
[0178] In another screen, we investigated a diverse set of UBR box domains from the UBR family of ubiquitin ligases that naturally bind N-terminal arginine, lysine, and histidine. The UBR box domain from the yeast K. lactis UBR1 protein showed the highest affinity for N-terminal arginine, and this protein was used to generate the arginine recognition factor PS691. PS691 recognized arginine in peptides with an N-terminal RLA with a median PD of 0.23 seconds (Figure 2C). The low affinity binding to N-terminal lysine and histidine (Figure 7D-E) was insufficient for single molecule detection.
[0179] To demonstrate that amino acids in a single peptide molecule can be sequentially exposed by aminopeptidases and recognized in real time with distinguishable kinetics, an immobilized peptide containing the initial sequence FAAWAAYAA (sequence number 832) was incubated with PS610 for 15 min, followed by the addition of PhTET3 (an aminopeptidase from P. horikoshii). The collected trace consisted of regions of distinct pulsing, termed recognition segments (RS), separated by regions lacking recognition pulsing (non-recognition segments, NRS). Analysis software was developed to automatically identify pulsing regions and transition points within the trace (Methods). The trace began with recognition of phenylalanine with a median PD of 2.36 s (Figure 2D), consistent with the PD observed for FAA in the recognition-only assay. This pattern was terminated after addition of aminopeptidase (on average 11 min after addition) and was followed by the orderly appearance of two RSs with median PDs of 0.25 and 0.49 s, corresponding to the short and moderate PDs obtained in assays with YAA and WAA recognition alone (Figure 2D). Thus, introduction of aminopeptidase activity into the reaction led to the sequential appearance of distinct RSs in the correct order and with the expected kinetic properties.
[0180] To demonstrate dynamic sequencing with two NAA recognition factors, PS610 and PS961 were labeled with distinguishable dyes atto-Rho6G and Cy3, respectively, and an immobilized peptide of sequence LAQFASIAAYASDDD (SEQ ID NO: 793) was exposed to a solution containing both recognition factors. After 15 min, two P. horikoshii aminopeptidases, PhTET2 and PhTET3, with complementary activities covering all 20 amino acids, were added. The collected traces showed distinct segments of alternating pulsing between PS961 and PS610 according to the order of recognizable amino acids in the peptide sequence (Figure 2E). The average bin ratios and average PDs associated with each RS easily distinguished the two dye labels and the four types of recognized NAA (Figure 2F). The median PDs were 2.70, 1.43, 0.25, and 0.66 s for N-terminal LAQ, FAS, IAA, and YAS, respectively (Figure 2G).
[0181] NAA-binding ClpS and UBR proteins also contact residues at positions 2 (P2) and 3 (P3) from the N-terminus, which affect binding affinity. These effects are reflected in the modulation of PD depending on downstream P2 and P3 residues, as observed above for LAA (1.21 s) compared to LAQ (2.70 s). We found that these effects on PD vary within informatively favorable ranges and can be empirically determined or estimated in silico to model peptide sequencing behavior a priori (Figure 7F-H). A powerful feature of this recognition behavior for peptide identification is that each RS contains information about potential downstream P2 and P3 residues or PTMs, regardless of whether these positions are targets of NAA recognition factors.
[0182] To evaluate the kinetic principles of dynamic sequencing when applied to diverse sequences, we characterized the synthetic peptide DQQRLIFAG (SEQ ID NO: 794), which corresponds to a segment of human ubiquitin (Figure 3A-D). Sequencing reactions were performed using three differentially labeled recognition factors, PS610, PS961, and PS691, in combination with two aminopeptidases, PhTET2 and PhTET3 (Materials and Methods). The example trace in Figure 3A begins with an NRS corresponding to the time interval in which the residues in the first DQQ motif are present at the N-terminus. The first RS begins at 120 min, when the N-terminal arginine is exposed for recognition by PS691. Subsequent cleavage events sequentially expose the N-terminal leucine, isoleucine, and phenylalanine to their corresponding recognition factors, with a rapid transition (average <10 s) from one RS to the next. The transition from leucine recognition to isoleucine recognition by PS961 is easily identified as an abrupt change in the average PD. This overall pattern is reproduced across many instances of sequencing the same peptide, with similar PD statistics across traces (Figures 3B-3C), because each peptide molecule follows the same reaction pathway over the course of the sequencing run. Due to the stochastic timing of the cleavage events, each trace shows distinct onset and duration times for each RS (Figure 3C).
[0183] This technique reports the binding kinetics at each recognizable amino acid position and the kinetics of aminopeptidase cleavage along the peptide sequence. Each RS typically contains tens to hundreds of on-off binding events, resulting in a distribution of PD and inter-pulse duration (IPD) measurements that can be statistically analyzed, so high-precision kinetic information on binding is obtained from a single trace. Repetitive probing of each NAA also provides high-precision recognition factor calls, because the call is not based on the error-prone detection of a single event associated with one fluorophore molecule (Figure 6F). The recognition factor concentration defines the IPD of each RS. Higher recognition factor concentrations result in shorter average IPDs and faster pulse rates (Figures 8A-8B). However, higher recognition factor concentrations can increase the fluorescence background from freely diffusing recognition factors, resulting in lower pulse signal-to-noise, and competing with aminopeptidases for N-terminal access. In practice, an IPD in the range of about 2-10 seconds provides a favorable balance between these factors.
[0184] The distribution of RS durations over the ensemble of repeated traces defines the cleavage rate of each recognizable NAA. For the DQQRLIFAG (SEQ ID NO: 794) peptide, mean cleavage times of 31, 54, 39, and 86 min were observed for the N-terminal arginine, leucine, isoleucine, and phenylalanine, respectively, with approximate monoexponential decay statistics for each position (Figure 3D, Figure 8C). The distribution of NRS durations reports the cleavage rates of one or more unrecognized NAA runs. The mean NRS duration for the first DQQ motif was 153 min (Figure 3D). The mean cleavage rate is a critical parameter and is controlled by the aminopeptidase concentration in the assay (Figure 8D-Figure 8E). Considering the exponential behavior, a mean RS duration of 10-40 min was targeted to provide sufficient time for pulse data collection, avoiding loss of RS due to rapid cleavage, and minimizing excessively long RS durations. We found it useful to visualize peptide sequencing profiles as kinetic signature plots, simplified trace-like representations of the complete peptide sequencing time course, including the median PD for each RS and the average duration of each RS and NRS (Figure 3E). These highly distinctive features provide a wealth of sequence-dependent information for mapping peptide-derived traces to their proteins of origin.
[0185] To demonstrate that this core methodology and its kinetic principles apply to a wide range of peptide sequences, synthetic peptides DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 795), RLAFSALGAADDD (SEQ ID NO: 796), and EFIAWLV (SEQ ID NO: 797) (a segment of human GLP-1) were sequenced under the same sequencing conditions used for DQQRLIFAG (SEQ ID NO: 794) (Figure 3F). Each peptide produced a characteristic kinetic signature according to its sequence (Figure 3G). Readouts were obtained up to position 18 (the furthest amino acid residue recognizable) of peptide DQQIASSRLAASFAAQQYPDDD (SEQ ID NO: 795), demonstrating that the method is compatible with long peptides and can deeply access sequence information for peptides of the length found in typical protein digests.
[0186] To demonstrate the extent to which the kinetic parameters obtained from sequencing are sensitive to changes in sequence composition, sequencing was performed with a set of three peptides that differ only at a single position located directly downstream of the PS961 N-terminal target leucine: RLAFAYPDDD (SEQ ID NO: 798), RLIFAYPDDD (SEQ ID NO: 799), and RLVFAYPDDD (SEQ ID NO: 800). Each type of amino acid at this position had a different effect on the PD obtained during recognition of the N-terminal leucine by PS961. Median PDs of 1.29 s, 2.22 s, and 4.21 s were observed for LAF, LIF, and LVF, respectively (Figure 4B). In addition to the differences in PD for leucine, each peptide showed a characteristic RS or NRS in the interval between leucine and phenylalanine recognition (Figure 4A, Figure 9A). These results demonstrate the sensitivity of the sequencing readout to variations at a single position and indicate that both the directly recognized NAA and the adjacent residues can affect the complete kinetic signature obtained from sequencing.
[0187] Since the aminoacyl-proline bond of the YP motif in peptides such as RLIFAYPDDD (SEQ ID NO: 799) cannot be cleaved by PhTET aminopeptidase, the observation of YP pulsing at the end of the trace confirms that cleavage proceeded completely from the first to the last recognizable amino acid. Thus, the sequencing output from RLIFAYPDDD (SEQ ID NO: 799) provided a convenient data set to investigate biochemical sources of non-ideal behavior that may lead to errors in peptide identification. The main contributors to incomplete information in the trace were the loss of expected RS due to the stochastic occurrence of rapid successive cleavage events (Figure 9B) and premature termination of the read resulting from photodamage or surface detachment (Figure 9C).
[0188] In addition to changes in amino acid sequence composition, the sequencing readout is sensitive to changes due to PTMs. As an example, methionine oxidation was examined. The thioether moiety of the methionine side chain is susceptible to oxidation during peptide synthesis and sequencing. PS961 has a K DIt was determined that PS961 binds peptides with an N-terminal methionine at P2 (Figure 9D), and oxidation resulting in a polar methionine sulfoxide side chain was hypothesized to eliminate binding and reduce NAA binding affinity when located at P2. It was computationally determined that methionine sulfoxide is highly unfavorable in the PS961 NAA binding pocket, and non-polar residues are preferred at P2 (Figure 9E). The synthetic peptide RLMFAYPDDD (SEQ ID NO: 801) was sequenced, and two populations of traces with distinct kinetic signatures were observed, the first population containing leucine recognition with a median PD of 0.86 seconds, and the second population with a median PD of 0.35 seconds (Figure 4C). Traces from the first population also showed methionine recognition with a short PD in the time interval between leucine and phenylalanine recognition (Figure 4E). Methionine recognition was absent in traces from the second population (Figure 4D), indicating that the methionine side chain in these peptides could not be recognized by PS961. When methionine was fully oxidized by preincubation with hydrogen peroxide (Materials and Methods), as expected, elimination of both methionine-recognition and leucine-recognition clusters was observed with a long median PD (Figure 4E). These results demonstrate the ability of extremely sensitive detection of PTMs due to their kinetic effects on recognition.
[0189] Proteomics applications require the identification of peptides in mixtures derived from biological sources. To extend the results to peptide mixtures and biologically derived peptides, two experiments were performed. First, DQQRLIFAG (SEQ ID NO: 794) and RLAFSALGAADDD (SEQ ID NO: 796) peptides were mixed, immobilized on the same chip, and sequencing experiments were performed. Data analysis (Materials and Methods) identified two populations of traces corresponding to each peptide, with kinetic signatures that were nearly identical to those identified in runs with individual peptides (Figure 5A, Figure 9F). Second, to demonstrate that the method extends to biologically derived peptides, sequencing experiments were performed with peptide libraries generated using a simple workflow from recombinant human ubiquitin (76 amino acids) and GLP-1 (37 amino acids) proteins digested with AspN / LysC and trypsin, respectively (Materials and Methods). For both libraries, data analysis readily identified traces matching the expected recognition patterns for the protease cleavage products DQQRLIFAGK (SEQ ID NO: 802) and EFIAWLVK (SEQ ID NO: 803) of ubiquitin and GLP-1, respectively, and generated kinetic signatures consistent with synthetic versions of these peptides (Figure 5B, Figure 9G). Matches to the kinetic signature of the ubiquitin peptide DQQRLIFAGK (SEQ ID NO: 802) were identified across the human proteome using simple sequence constraints provided by the kinetic information (Materials and Methods). Only one protein other than ubiquitin was found to contain a peptide that could potentially match this signature. Thus, even the short signatures are consistent with 10 4 These results demonstrate the potential of full kinetic output from sequencing to enable digital mapping of peptides to their protein of origin.
[0190] Consideration The simple real-time kinetic approach is significantly different from other recently described single molecule approaches that rely on stepwise Edman chemistry or complex iterative methods involving hundreds of cycles of epitope probing. Nanopore approaches offer the potential for real-time readout and simplicity, but face significant challenges related to the size and biophysical complexity of polypeptides. The sequencing technology described herein is easily expanded in its capabilities, and multiple areas for improvement exist. Expanded proteome coverage can be achieved through directed evolution and engineering of recognition factors. The NAA targets demonstrated here comprise approximately 35.6% of the human proteome, but lower affinity NAA targets require longer PD to enable detection in all sequence contexts.
[0191] Recognition factors for new amino acids or PTMs can be evolved from current recognition factors or identified in screening other types of NAA or other scaffolds such as PTM-binding proteins or aptamers. In general, expansion to detection of all 20 natural amino acids and multiple PTMs is feasible for de novo sequencing. However, partial sequences are sufficient for most proteomics applications that rely on mapping to a predefined set of candidate proteins. Aminopeptidases can be engineered to optimize cleavage rates and minimize RS deletions from rapid sequential cleavage. It is envisioned that the dynamic range of the sample and the most suitable application for the system will tend to scale with the number of reaction chambers on the chip, and compression of the dynamic range will be necessary for certain applications.
[0192] The sequencing technologies demonstrated herein promise to increase the availability of proteomic testing, enable new discoveries in biological and clinical research, and help power a new generation of precision medicine.
[0193] Materials and Methods Semiconductor device operation and bin ratio calculation Experiments were performed on a pre-fabricated semiconductor chip with 296K active wells, allowing for some losses leading to flow cell blockage of the sensor array. The dual chamber flow cell allows two independent samples to be sequenced in parallel, each utilizing 148K active wells. The initial production device will have 2M active wells, scaling to tens of millions of active wells using standard CMOS processing for the first product line. Pulsed 532 nm excitation light from a 67 MHz mode-locked laser is coupled into a grating coupler at the edge of the semiconductor chip. The use of a single laser wavelength combined with fluorochrome discrimination by fluorescence intensity and lifetime reduces size, cost, and complexity, contributing to the scalability of the platform. A network of optical waveguides splits the excitation light and routes it to the sensor array to illuminate each reaction chamber. Each CMOS pixel contains a single light-sensitive photodiode with two high-speed global shutters (rejection and collection gates) that discard and collect photoelectrons (the chip photonic structure reduces pixel-to-pixel crosstalk to less than 2%). Control waveforms are applied to the collection and reject gates in synchronization with the incident pulsed light source (Figure 6B). Approximately 1 ns prior to the excitation pulse, the reject gate is charged to >3 volts and the collection gate is discharged to <1 volt. Scattered 532 nm excitation photons generate photoelectrons within the photodiode. The photoelectrons are rapidly transferred to a high voltage drain by a built-in potential field within the photodiode and the reject gate potential. Between 1 and 3 ns after excitation, the collection gate is charged to >3 volts and the reject gate is discharged to <1 volt. Photoelectrons generated from emission photons that reach the photodiode after the collection gate is opened are transferred to a storage node within each pixel. Photoelectrons are configurably accumulated for 7.5 to 30 ms within each pixel over approximately 500,000 to 2,000,000 laser pulses (Figure 6B). The charge stored in the storage node is measured using standard transfer gates, floating diffusions, source followers, row selects, and on-chip analog-to-digital converters common to all CMOS image sensors, allowing scalability to large array sizes with small pixels.Fluorescence lifetime information is obtained by alternating the timing of the collection and rejection gating waveforms between subsequent measurements. In the first measurement, only emission photoelectrons arriving >3 ns after the excitation pulse are collected (bin 0). In the second measurement, emission photoelectrons arriving >1 ns after the excitation pulse are collected (bin 1). The signal measured from the pixel as the phase relationship between the excitation source and the gating waveform is adjusted through the entire excitation cycle demonstrates that the pixel transitions from 100% collection of photons during the collection phase to greater than 99.99% disappearance of photons during the rejection phase in less than 1 ns (Figure 6C). The ratio of these two measurements (bin ratio) provides an estimate of the fluorescence lifetime (Figure 6D). We have demonstrated the ability to distinguish between multiple dyes based on the bin ratio alone (Figure 6E).
[0194] Peptide synthesis and labeling Peptides were synthesized on rink amide resin on a PurePrep® Chorus solid-phase peptide synthesizer (Gyros Protein Technology) using standard Fmoc chemistry. All synthetic peptides contained a C-terminal Fmoc-azidolysine. The resin was deprotected in a mixture of TFA / TIPS / H2O (2.5% / 2.5% / 95%) for 1.5 h at room temperature. The deprotection mixture was concentrated under a stream of argon. Peptides were precipitated from cold diethyl ether, resuspended in 1:1 water-acetonitrile, and purified by reverse-phase HPLC (X-bridge C18, Waters) using a gradient of 10-70% acetonitrile (0.05% TFA) over 20 min. The residue was dried under high vacuum to produce a white pellet. To a solution of DBCO-DNA-biotin (2 nmol in 100 μL PBS) was added peptide stock solution (4 μL, 5 mM) at room temperature. The progress of the reaction was monitored by LC-MS (Thermo UlTiMate 3000 Executive Plus). After the reaction was complete, the mixture was conjugated to excess streptavidin. The peptide-DNA-streptavidin complex was purified by ion-exchange HPLC (DNAPac 200, Thermo). Gradient, Buffer A, 20 mM sodium phosphate buffer, pH 8.5, Buffer B, 1 M NaBr, 20 mM sodium phosphate buffer, pH 8.5, 20-60% B over 15 min. The purified complex was buffer exchanged on a 30K MWCO spin filter into a solution containing 50 mM MOPS (pH 8.0) and 60 mM potassium acetate before use. Peptides containing fully oxidized methionine were prepared by mixing 3% hydrogen peroxide with the methionine peptide in 1:1 water-methanol for 20 min at room temperature. The product was immediately purified by reversed-phase HPLC using the same peptide purification method described above, the purity was confirmed by reversed-phase HPLC (Thermo UlTiMate 3000) on an analytical column (Zorbax® SB-Aq, 5 μm, 4.6 × 250 mm) and the exact mass of the oxidation product was confirmed by LC-MS (Agilent LC-MSD-iQ, positive mode).
[0195] Protein digestion and labeling GLP-1 7-37, GLP-2, and ubiquitin (1-76) recombinant proteins were purchased as lyophilized powders from RnD Systems. Each protein was reconstituted to a final concentration of 200 μM in 100 mM HEPES, pH 8.0 (20% acetonitrile). When necessary, cysteines were reduced and alkylated using TCEP (2 mM) and iodoacetamide (10 mM). GLP1 and GLP2 were digested overnight at 37° C. using 1 μg of trypsin (LCMS grade, Pierce). Ubiquitin was digested using 1 μg of LysC (LCMS grade, Pierce) and 1 μg of rAspN (LCMS grade, Promega). After protease digestion, the pH of the peptide mixture was adjusted to pH 10.5 using potassium carbonate (57 mM) and lysines were converted to azidolysines using imidazole-1-sulfonyl azide (ISA, 2 mM) and copper sulfate catalyst (0.5 mM). ISA was quenched using amine-functionalized polyurethane beads (Oligo Factory). The mixture was then filtered and adjusted to pH 7-8 using 1 M acetic acid. The solution was diluted with 50% (v / v) 10 mM MOPS, 10 mM KOAc, pH 7.5, added to the DNA-streptavidin-DBCO complex and incubated at 37 °C for 12-16 h. When required, the detergent, cetrimonium bromide, was added to the reaction at a final concentration of 0.25 mM.
[0196] Recognition factor purification, labeling, and characterization Expression vectors (with pET30a+ backbone) for the recognition factors and biotin ligase were co-transformed into BL21(DE3) chemically competent E. coli cells. Transformed cells were plated on Luria agar plates containing carbenicillin (50 μg / mL) and kanamycin (25 μg / mL) and incubated overnight at 37° C. to obtain single colonies. Starter liquid cultures inoculated with colonies were grown in Luria broth containing ampicillin (50 μg / mL) and kanamycin (25 μg / mL) to inoculate large-scale cultures at a starting optical density (OD600) of approximately 0.01. Expression cultures were incubated at 37° C. and 230 rpm until the OD600 approached approximately 0.7. Cultures were then induced with 4 mM IPTG. The expressed recognition factors were biotinylated in vivo by adding 8 mM biotin simultaneously with IPTG. Approximately 12 hours after expression, cells were harvested by centrifugation at 10,000g at 4°C, and the cell pellet was washed with 1x PBS buffer pH 7.4. Cells were resuspended in Bug buster® HT (Thermo Fisher Scientific) and incubated at room temperature on a magnetic stirrer for 30 minutes. The cell suspension was then diluted with an equal volume of 2x lysis buffer (100mM Tris-HCl pH 7.5, 10% glycerol, 0.5M NaCl) and incubated at room temperature on a magnetic stirrer for 30 minutes. The lysate was centrifuged at 21000g at 4°C to remove cell debris. The supernatant was collected and loaded onto a nickel NTA resin (Cytiva) affinity column pre-equilibrated with buffer A (50mM Tris-HCl pH 7.5, 10% glycerol, 0.5M NaCl) on an AKTA Pure (Cytiva) system. The column was washed with at least 10 column volumes of buffer containing 10 mM imidazole. Elution was performed with a 10-300 mM imidazole gradient. The eluted fraction was dialyzed overnight at 4 °C in a 10 kDa cassette against 4 L of dialysis buffer (50 mM Tris-HCl pH 7.5, 0.2 M NaCl, 50% glycerol).
[0197] For labeling of the recognition factor, equal volumes of recognition factor and DNA-dye-streptavidin complex were mixed at a molar ratio of 5:1 (recognition factor:DNA-dye-SV). The mixture was incubated on ice for 30 min and dialyzed overnight against SEC buffer (25 mM HEPES pH 8.0, 150 mM KCl). The recognition factor-dye conjugate was recovered from dialysis and centrifuged at 10,000 g at 4° C. The supernatant was collected and concentrated using a 10 kDa cut-off concentrator. The concentrated conjugate was purified on an Agilent 1260 Infinity HPLC system using a size exclusion column (BioSEC-3 300 Å, 3 μm).
[0198] Binding affinity was measured by polarization using labeled peptides. Polarization response and total intensity measurements were performed at 20°C in a microplate fluorometer with 480 nm excitation and 530 nm emission. Interaction of recognition factor with labeled peptide (XAKLDEESILKQK-FITC (SEQ ID NO: 833)) containing the target N-terminal residue was performed in PBS buffer at pH 7.4 and readings were collected after 30 min. To obtain titration curves, multiple analyses were performed with increasing concentrations of recognition factor at a fixed concentration of target peptide. The equilibrium polarization response at each concentration was plotted and fitted to obtain the K D was calculated.
[0199] Using a stopped-flow apparatus, the off-rates (k off ) was measured. Labeled peptide (50 nM) was mixed with PS610 in PBS buffer (pH 7.4) containing 0.01% Tween-20 and incubated at 30 °C. After 30 min of incubation, the recognition factor:peptide complex was rapidly mixed with a 10- to 20-fold molar excess of unlabeled trap peptide, and the reaction was followed in real time by measuring the fluorescence intensity. At least three time-lapse traces were averaged and fitted to an exponential equation.
[0200] Aminopeptidase purification Expression vectors (with pET30a+ backbone) for the aminopeptidases PhTET2 and PhTET3 were transformed into BL21(DE3) chemically competent E. coli cells. The transformed cells were plated on Luria agar plates containing kanamycin (25 μg / mL) and incubated overnight at 37° C. to obtain single colonies. Starter liquid cultures inoculated with the colonies were grown in Luria Broth (LB) containing kanamycin (25 μg / mL) and inoculated into large-scale cultures at an initial optical density (OD600) of approximately 0.01. The expression cultures were incubated at 37° C. and 230 rpm until the OD600 approached approximately 0.7. The cultures were then induced with 0.4 mM IPTG. The expressed aminopeptidases were purified as described above for the recognition factors. For conditioning, the aminopeptidase protein was dialyzed against 50 mM MOPS (pH 8.0) / 60 mM potassium acetate and then exposed to a final concentration of 400 μM cobalt acetate for 1–1.5 h at 65°C to form the active dodecamer complex. The conditioned aminopeptidase preparation was further dialyzed against 50 mM MOPS pH 8.0 / 60 mM potassium acetate, aliquoted, and flash frozen.
[0201] Peptide loading, recognition, and dynamic sequencing The semiconductor chip was placed in the sequencing device and a chip check was performed to test electronic circuit function and optimize laser coupling alignment. The chip was then removed from the device socket and the chip was washed twice with 50 μL of 70% isopropanol, followed by four washes with 30 μL of wash buffer (50 mM MOPS pH 8.0, 60 mM potassium acetate, 50 mM glucose, 20 mM magnesium acetate, and surfactant mix) through a flow cell attached to the chip. A second chip check was then performed. The laser was then shut off and peptide complexes were added to a final concentration of 1-10 nM, mixed thoroughly, and the chip was incubated for 15 min, via an integrated software-controlled shutter. The chip was then washed six times with wash buffer, followed by the addition of imaging solution (wash buffer containing 5 mM Trolox and an oxygen scavenging and removal system). The laser was unblocked and the occupancy percentage (target 10-30%, Poisson distribution) was recorded by acquiring the photobleaching signal from the fluorophore bound to the peptide complex during 5 min of laser irradiation. For NAA recognition only assays, after peptide loading, labeled recognition factors were added to a final concentration of 50 nM PS610, 100 nM PS691, or 250 nM PS961 (as indicated according to the experiment) and data were recorded for 10 h. For kinetic sequencing assays, after peptide loading, a mixture of labeled recognition factors was added to give a final concentration of 50 nM PS610, 100 nM PS691, and 250 nM PS961. Data were recorded for 15 min. The laser was then briefly blocked and aminopeptidase was added to the sequencing reaction via the flow cell and mixed thoroughly (final concentration 2-8 μM PhTET2 and / or 20-80 μM PhTET3 as indicated according to the experiment). The laser was then unblocked and data was recorded for 10 h. For all runs, 30 μL of mineral oil was added to the fluid reservoirs of each port of the flow cell to prevent evaporation during the run.
[0202] Signal Processing and Trace Segmentation The signal measured on-chip contains various noise components, the most dominant being due to the fluorescent emission from the diffusing recognition factor in the reaction chamber. The pulse caller algorithm for a given reaction chamber starts by estimating the statistical properties of this background noise component. Once an estimate within certain error limits is established, the algorithm operates in an on-line manner observing new frames of data as they are generated. At each point in time, the algorithm maintains a state indicating whether the signal is due to background components only or whether a pulse from recognition factor-NAA interaction is being observed. The background-to-pulse state transition is triggered using an edge detection test where a shift in the signal is expected to be significant relative to the statistical distribution of the background components. The pulse-to-background state transition is triggered when a small window of the most recent frames of signal appears to again follow the distribution of the background components. The algorithm maintains an updated model of the background components as new background frames are observed. This provides robustness against drifts in the signal intensity, along with a feedback control loop that maintains stable optical coupling of the laser to the chip based on any such detected drifts. Because detected pulses may be due to true recognition agent-dipeptide interaction events as well as other incidental transient noise spikes, a downstream filter layer is used to test the significance of pulse events based on their duration, intensity, and noise pattern within the context of the entire timeline of the run and the entire data set of the reaction chamber.
[0203] Initial regions are determined by performing a sliding window calculation of pulse rates along the time dimension of the series of pulses. Regions with a mean pulse rate >1 pulse / min are then subdivided according to a greedy bisection approach, where the left and right pulses of each potential split are evaluated for statistically significant deviations in any of four separate pulse characteristics (intensity, bin ratio, pulse duration, and interpulse duration) using a Mann-Whitney U test. To define the RS, the split point with the lowest p-value for any of the four characteristics is used to subdivide the region, with a p-value <10 in any comparison. -5 The process continues until there are no more regions with candidate division points with . In this way, the transition from one RS to the next in a region of successive pulsing is determined a priori based on the change in the fluorescent properties of the pulsing kinetics. The resulting region is called the recognition segment (RS).
[0204] Recognition Segment Classification RS classification for reactions containing a single synthetic peptide was performed using an unsupervised clustering algorithm. A Gaussian mixture model (GMM) was pre-trained using a subset of RSs including those with an average signal-to-noise ratio of the constituent pulses of ≥ 3 to identify approximate centroids for each of the N classes of recognition, where N is equal to the number of expected recognizable peptide states with F, Y, W, L, I, V, or R at the N-terminus. Identified clusters were assigned to recognizable peptide states by matching the predominant order of observed cluster sequences to the expected amino acid sequence and by using prior knowledge of the dye properties to identify binders that are active in each RS. Subsequent rounds of GMM fitting were performed on all RSs that matched the expected order of these events to refine the GMM model until no further sequences appeared in the expected order. The final model was then applied to all RSs in a given reaction.
[0205] RS classification for reactions containing library prepared peptides and mixtures of peptides was performed using a random forest classifier pre-trained on annotated RS pulse features from previous synthetic peptide experiments. Unless otherwise stated, figures and statistics generated from classified RS are derived from reaction chambers with the expected sequence of RS.
[0206] Molecular dynamics and binding energy calculations A homology model of PS961 complexed to the peptide was generated using the internal crystal structure, mutations were applied, and optimized using protCAD prior to molecular dynamics. AMBER20 implicit solvent molecular dynamics simulations with a general Born solvation potential were performed using the ff19SB force field with no interatomic distance cutoff. Minimization was performed using steepest descent, followed by conjugate gradient minimization. Langevin dynamics and 3ps -1 The system was thermalized from 0 to 300 K using a collision frequency of 0.5 μm. Molecular dynamics simulations of the equilibrated recognition factor-peptide complex, the free recognition factor, and the free peptide were performed independently at 300 K for 5 ns, and binding energy calculations were performed using MMPBSA, where 125 frames, each containing 10,000 2-femtosecond steps, were used for the calculations from the three simulations. The binding energy and the decomposition of all residues contributing to the binding energy were calculated at 0.15 M salt concentration.
[0207] Example 2. Peptide identification using modeled proteome-wide kinetic signatures Using sequencing and biochemical data, predicted pulse durations were determined for recognition factors binding to all possible tripeptide targets. Figures 10A-10C show heat maps of predicted pulse durations for PS961-bound tripeptide targets with leucine (Figure 10A), isoleucine (Figure 10B), or valine (Figure 10C) at the N-terminal position. Figures 10D-10F show heat maps of predicted pulse durations for PS610-bound tripeptide targets with phenylalanine (Figure 10D), tyrosine (Figure 10E), or tryptophan (Figure 10F) at the N-terminal position. Figure 10G shows heat maps of predicted pulse durations for PS1122-bound tripeptide targets with arginine at the N-terminal position. The predicted pulse durations correlated highly with the actual pulse durations from on-chip experimental results for PS961 (Figure 10H, left plot) and PS610 (Figure 10H, right plot).
[0208] This database of predicted tripeptide pulse durations can be used to model the expected kinetic signatures of all peptides in the human proteome, which can provide improved understanding and utilization of the ability to identify proteins from sequencing output. Kinetic signatures are average representations of peptide-on-chip sequencing behavior, as detailed in Example 1 above. The information in kinetic signatures derived from single molecule traces dramatically improves the ability to map sequencing data to proteomes (e.g., compared to methods based on alignment of text strings, such as in DNA sequencing). Kinetic information can include, for example, pulse duration, inter-pulse duration, and recognition segment (RS) duration.
[0209] To prepare a model that demonstrates the ability to uniquely map peptides to the human proteome (using recognition factors PS961, PS610, and PS1122), we performed an in silico digestion of the proteome with AspN / LysC, followed by selection of all peptides that terminated with lysine (used for on-chip immobilization) and were greater than seven amino acids in length. The results are shown below.
[0210] [Table 2]
[0211] A predicted pulse duration was assigned to all visible amino acids in the set of 273112 peptides (positions with predicted mean PD less than 0.18 seconds were treated as invisible). The distribution of predicted RS in the first 15 residues is shown in Figure 10I (left plot). 82068 peptides contained 4 or more RSs (and were therefore considered potentially informative). A kinetic signature was generated for each of these peptides.
[0212] The kinetic signature contains the expected binder and average PD at each visible position, and gaps representing runs of one or more invisible amino acids. For each peptide, the number of peptides with the same kinetic signature was then determined (signatures were considered identical if they had the same order of RS and gaps, and the predicted PD at each RS was somewhat similar (if in any pairwise comparison the shorter PD was more than half of the longer PD)). According to this analysis, 38849 of the 82068 peptides yielded unique kinetic signatures with no other matches in the human proteome. An additional 10571 peptides had only one other match. The distribution of kinetic matches per peptide is shown in Figure 10I (middle plot). 14167 proteins (69% of all proteins) contained at least one uniquely mappable peptide. On average, there were 2.5 uniquely mappable peptides per protein. The distribution of uniquely mappable peptides per protein is shown in Figure 10I (right plot).
[0213] To further illustrate this data and how it can be used to model protein behavior, results with the IL6 protein are shown in Figure 10J (for simplicity, the residue immediately preceding the C-terminal lysine was treated as invisible and the XP motif was treated as cleavable). As shown in Figure 10J, two peptides contain at least four RSs. As shown in Figure 10K, one of these peptides uniquely maps to IL6, and the other peptides match the kinetic signatures of eight different peptides from eight proteins.
[0214] To provide an illustrative example of using a smaller proteome, the E. coli proteome (containing only 4392 proteins) was analyzed as described above for the human proteome. The results are shown below.
[0215] [Table 3]
[0216] The distribution of predicted RSs in the first 15 residues is shown in Figure 10L (left plot). 9925 peptides contained 4 or more RSs (and were therefore considered potentially informative). A kinetic signature was generated for each of these peptides. For each peptide, the number of peptides with the same kinetic signature was determined. According to this analysis, 7740 of the 9925 peptides yielded unique kinetic signatures with no other matches in the E. coli proteome. The distribution of kinetic matches per peptide is shown in Figure 10L (middle plot). 3187 proteins contained at least one uniquely mappable peptide. On average, there were 2.4 uniquely mappable peptides per protein. The distribution of uniquely mappable peptides per protein is shown in Figure 10L (right plot). To illustrate this data and how it can be used to model protein behavior, results using a protein from E. coli containing 6 uniquely mappable peptides are shown in Figure 10M.
[0217] These results demonstrate the utility of a kinetics-centric view of peptide identification, which also provides the ability to model with high accuracy the informative effects of changes to reaction conditions such as adding new recognition factors, increasing recognition factor pulse durations, changing frame rates, and adding new dye labels.
[0218] Example 3. Direct identification of arginine post-translational modifications Proteins undergo diverse post-translational modifications (PTMs) on their amino acid side chains that can strongly affect protein function and mediate complex cellular events. Measuring the diversity, dynamics, and functional consequences of protein PTM states across the proteome is essential to understand the role of proteins in health and disease. However, discovery and detection of PTMs and routine measurement of complex PTM states remain extremely challenging, and the diverse proteoforms in the human proteome remain largely unmapped. New methods enabling sensitive detection of PTMs would greatly aid biomarker discovery, drug discovery, and the development of approaches to precision personalized medicine.
[0219] Modification of arginine side chain is of particular biomedical interest. Methylation and citrullination of arginine residues in many human proteins have been shown to play important roles in disease states such as cardiovascular disease, autoimmune disease, and cancer. In this example, aspects of the technology described herein are applied to the detection of arginine methylation and citrullination with single molecule resolution and high sensitivity.
[0220] Arginine plays an important role in protein structure and function due to the unique properties of the guanidinium group that terminates its side chain (Figure 11A). This group is positively charged and can form an extended hydrogen bond network as well as cation-π interactions with other amino acids and nucleic acids. Thus, arginine often mediates important interactions between protein binding partners or between proteins and DNA.
[0221] The two most common arginine PTMs, dimethylation and citrullination, modify the arginine side chain, changing its properties (Figure 11A), potentially resulting in important downstream effects on cellular processes. Dimethylation retains the positive charge of arginine but increases its size and hydrophobicity, blocking hydrogen bond formation. Citrullination eliminates the positive charge of arginine, resulting in a neutral side chain with altered properties, which can profoundly affect protein conformation and function.
[0222] Arginine dimethylation and citrullination are carried out by enzymes and may be part of the normal regulation of cellular processes or may be involved in disease states. Arginine dimethylation is catalyzed by protein arginine methyltransferase (PRMT). PRMT asymmetrically transfers two methyl groups onto the same nitrogen atom resulting in asymmetric dimethylarginine (ADMA) or onto opposite nitrogen atoms resulting in symmetric dimethylarginine (SDMA). These modifications increase size and hydrophobicity and block hydrogen bonds. Arginine citrullination is catalyzed by protein arginine deiminase (PAD). PAD performs hydrolysis of the positively charged guanidinium group of arginine to produce a neutral ureido group. This conversion results in a negligible mass increase of 0.9840 Da, but the loss of the positive charge can dramatically change the conformation and function of the protein. Figure 11A shows the structures of SDMA, ADMA, standard arginine, and citrulline.
[0223] Arginine PTMs have emerged as important targets in biomedical research. Methylated arginine residues and their respective PRMTs have been implicated in important diseases such as cardiovascular disease and cancer. Critical involvement of arginine citrullination in immune system function, skin keratinization, myelination, and regulating gene expression has also been demonstrated. Notably, removal of the positive charge of arginine can, in some cases, cause proteins to activate the immune system and contribute to autoimmune diseases.
[0224] Difficulties with detecting arginine PTMs Research into these arginine PTMs has been particularly challenging because they are difficult to detect and distinguish with current proteomic methods. Mass spectrometry is the most frequently utilized tool for detecting protein PTMs. However, ADMA and SDMA are structural isomers with identical masses and cannot be easily distinguished by mass spectrometry. Similarly, deimination of arginine to citrulline results in a negligible mass increase of 0.9840 Da. This mass difference is 13 It can easily be confused with the C isotope or can be misinterpreted as deamidation of nearby asparagine or glutamine residues. In addition, mass spectrometry techniques for arginine PTM detection require highly specialized knowledge and training as well as sophisticated analytical methods.
[0225] Another common method of PTM detection, enzyme-linked immunosorbent assay (ELISA), uses specifically generated antibodies to detect modified proteins of interest. Although arginine PTMs are presumed to be widespread in human cells, commercially available antibodies against arginine PTMs are limited to specific sites on a few highly tested proteins. The need to generate new antibodies, along with the complex workflow, expense, antibody reproducibility, and other difficulties associated with ELISA assay development, likely hinder the discovery and further testing of novel arginine PTM sites.
[0226] Continued development towards novel methods is needed to facilitate the direct detection of arginine PTMs in proteins. Single molecule protein sequencing offers an alternative approach to the detection of ADMA, SDMA, and citrulline that is not based on mass-to-charge ratio or antibody specificity, but rather on the kinetic signature of binding between the recognition factor and the N-terminal amino acid (NAA).
[0227] Embodiments of the technology described herein overcome current technology gaps and provide direct detection of arginine PTMs, gaining insight into these PTMs with single molecule resolution. Methodology and Workflow PTM detection involved isolating peptides and subjecting them to a real-time single-molecule protein sequencing reaction. Proteins were first digested into peptide fragments and conjugated at the C-terminus to a polymeric linker. The peptide complexes were immobilized at the bottom of nanoscale wells on a semiconductor chip, yielding single peptide molecules with exposed N-termini ready for sequencing. During the sequencing reaction, the surface-immobilized peptides were exposed to a solution containing dye-labeled NAA recognition factors that bind on and off to cognate NAAs with characteristic kinetic properties. Aminopeptidases in the solution successively removed individual NAAs to expose subsequent amino acids for recognition. Fluorescence lifetime, intensity, and kinetic data were collected in real time and analyzed to determine amino acid sequence and PTM content.
[0228] The trace-level output contained distinct pulsing regions called recognition segments (RS). Each RS corresponded to the period between aminopeptidase cleavage events during which an NAA recognition factor binds on and off to its exposed target NAA. Chemical modifications to the target NAA or nearby downstream amino acids can modulate recognition factor affinity, resulting in characteristic changes in the average pulse duration (PD) in the RS compared to the unmodified peptide. These modifications can also affect the rate of aminopeptidase cleavage of the NAA, resulting in characteristic changes in the average duration of the corresponding RS.
[0229] An overview of the workflow for peptide sequencing and PTM detection is shown in Figure 11B. Results and Discussion Detection of arginine dimethylation First, we demonstrate the detection and differentiation of arginine, ADMA, and SDMA by single molecule protein sequencing, focusing on a critical segment of the signaling protein P38MAPKα. Dimethylation of arginine residue 70 of P38MAPKα in myoblasts by PRMT7 is a key regulatory step in the activation of myoblast differentiation in humans.
[0230] A synthetic peptide corresponding to residues 69-76 of p38MAPKα was generated in three versions containing either arginine, ADMA, or SDMA at position 2: YRELRLLK (SEQ ID NO: 834), YR ADMA ELRLLK (SEQ ID NO: 835), and YR SDMA ELRLLK (SEQ ID NO: 836). Each peptide was sequenced using three recognition factors, PS610 (F, Y, W), PS961 (L, I, V), and PS621 (R), and the data was analyzed to identify the RS, determine the average PD of each RS, and characterize the kinetic signature of each peptide. Each peptide showed a distinct pattern due to the distinct kinetic effects of arginine, ADMA, and SDMA on recognition factor binding (see example traces in Figure 11C-A).
[0231] Arginine and ADMA residues showed binding to the recognition factor PS621 with similar PDs, whereas SDMA showed no binding (Figure 11C-A, 11C-B). This result indicated that symmetric dimethylation of arginine, in contrast to asymmetric dimethylation, reduced the affinity of PS621 for N-terminal arginine, providing a clear kinetic difference between these isomeric arginine PTMs. The NAA recognition factors used in this example contact residues 2 and 3 from the N-terminus when they bind to their target NAA. Thus, modification of these downstream residues may affect recognition factor binding affinity. A strong effect of arginine dimethylation on the recognition of upstream tyrosine residues in these peptides by PS610 was observed (Figure 11C). The median pulse duration of tyrosine recognition ranged from 0.69 s for YRE to 0.59 s for YR ADMA E and YR SDMA In addition, the median interpulse duration (IPD) of arginine recognition by PS621 decreased from 10.05 s for unmodified arginine to 5.82 s for ADMA (Fig. 11C-C).
[0232] The effect of these dimethylated arginine residues on the recognition of the preceding NAA serves as a powerful feature for protein sequencing with single molecule sensitivity and accuracy. These results demonstrate the capabilities of unprecedented sensitivity in the detection of arginine dimethylation using aspects of the technology described herein.
[0233] Detection of arginine citrullination We next demonstrated that citrullinated arginine residues can be rapidly distinguished from natural arginine residues using differential binding kinetics. Two synthetic peptide sequences, LRLAFAYPDDDK (SEQ ID NO: 817) and LCitLAFAYPDDDK (SEQ ID NO: 839), containing either arginine or citrulline at position 2, were generated and sequenced using the three recognition agents as described above. Each peptide showed highly distinct kinetic signatures due to the influence of the different arginine and citrulline side chains on recognition (Figures 11D-A, 11D-B). Citrullination eliminated N-terminal arginine recognition by PS621 (see example trace in Figure 11D-A). Citrullination at position 2 also resulted in a large increase in the median PD of recognition of the N-terminal leucine located at the preceding position by PS961. The median PD was 0.43 seconds for LRL and increased to 0.78 seconds for LCitL (Figure 11D-B). These results demonstrate the ability to detect and digitally quantify arginine citrullination.
[0234] conclusion In this example, arginine PTMs were directly detected. Arginine PTMs play an important role in human health and disease, but have been difficult to test. Current proteomics methods (e.g., mass spectrometry and ELISA) can only indirectly identify these arginine PTMs using highly specialized techniques, or are limited to a small set of specific proteins based on antibody availability and other difficulties. The ability to directly detect PTMs offers great potential for accelerating biomedical research, as well as a wide range of commercial applications in drug discovery and biomarker development.
[0235] Example 4. Identification of threonine post-translational modifications Sequencing reactions with the recognition factors PS691 (R), PS610 (FYW), and PS961 (LIV) were performed separately for peptides RLTFIAYPDDD (SEQ ID NO: 821) and RLpTFIAYPDDD (SEQ ID NO: 822) (pT is phosphothreonine). Recognition of N-terminal leucine preceding threonine or phosphothreonine by PS961 was observed, with distinct pulse durations for threonine following leucine (RS average PD = 1.2 sec; Figure 14A) compared to phosphothreonine following leucine (RS average PD = 0.3 sec; Figure 14B). Furthermore, the recognition segment (RS) duration for leucine recognition was longer when phosphothreonine followed leucine (RS average duration = 130 min; Figure 14C, right panel) compared to threonine (RS average duration = 8.1 min; Figure 14C, left panel). These data demonstrate the ability to discriminate between unmodified and post-translationally modified threonine side chains.
[0236] Example 5. Identification of tyrosine post-translational modifications Sequencing reactions with the recognition agents PS691 (R), PS610 (FYW) and PS961 (LIV) were performed separately for peptides RLYFIAYPDDD (SEQ ID NO: 823) and RLpYFIAYPDDD (SEQ ID NO: 824) (pY is phosphotyrosine). Recognition of N-terminal arginine and leucine residues preceding tyrosine or phosphotyrosine by PS691 and PS961, respectively, was observed with distinct pulse durations depending on whether the peptide contained tyrosine (FIG. 15A) or phosphotyrosine (FIG. 15B). Recognition of N-terminal arginine occurred with an RS average PD of 0.9 s for RLY and 0.45 s for RLpY. Recognition of N-terminal leucine occurred with an RS average PD of 2.45 s for LYF and 3.4 s for LpYF. Furthermore, the trace from peptide RLpYFIAYPDD (SEQ ID NO: 840) contained a consensus gap between L and F as pY was not recognized by PS610, whereas the trace from peptide RLYFIAYPDDD (SEQ ID NO: 823) contained Y recognition by PS610 in this interval. These data demonstrate the ability to discriminate between unmodified and post-translationally modified tyrosine side chains.
[0237] Example 6. Identification of lysine post-translational modifications Sequencing reactions with the recognition agents PS691(R), PS610(FYW), PS961(LIV) and PS1165(A) were performed separately for the peptides RLYFKAYPDDD (SEQ ID NO: 825) and RLK{acetyl}FIAYPDDD (SEQ ID NO: 826) (K{acetyl} is an acetylated lysine). Recognition of N-terminal phenylalanine and alanine residues preceding lysine or acetyl-lysine by PS610 ...
Claims
[Claim 1] The invention described in this specification.