Determination of Protein Information by Recoding Amino Acid Polymers into DNA Polymers
By recoding amino acid sequences into DNA polymers for high-throughput analysis, the method addresses the limitations of current proteome characterization tools, enabling sensitive and economical detection of protein abnormalities for early disease diagnosis and treatment.
Patent Information
- Application Number
- JP2025501695
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-19
- Filing Date
- 2023-07-12
- Publication Date
- 2025-07-17
AI Technical Summary
Current tools and techniques for sensitive, accurate, and economical characterization of the proteome are lacking, hindering early detection of abnormal protein sequences and concentrations, which are crucial for diagnosing and treating diseases like cancer.
A method involving recoding amino acid sequences of peptides into DNA polymers for high-throughput analysis, using chemically reactive conjugates to immobilize and sequence amino acids, followed by DNA sequencing to determine identity and positional information.
Enables highly parallelized and cost-effective proteome analysis, overcoming limitations of existing methods by providing detailed sequence and presence information of proteins, facilitating early-stage disease detection and therapeutic discovery.
Smart Images

Figure 2025523087000001_ABST
Abstract
Description
Technical Field
[0001] Cross-reference This application claims the benefit of U.S. Provisional Application No. 63 / 388,317, filed Jul. 22, 2022; U.S. Provisional Application No. 63 / 399,294, filed Aug. 19, 2022; U.S. Provisional Application No. 63 / 439,523, filed Jan. 17, 2023; and U.S. Provisional Application No. 63 / 467,729, filed May 19, 2023, the disclosures of which are hereby incorporated by reference in their entirety.
[0002] Sequence Listing This application contains a Sequence Listing that has been electronically submitted in XML format and is hereby incorporated by reference in its entirety. The name of the XML copy created on Jul. 11, 2023 is 062954-501001US_SL.xml, and the size is 118,986 bytes.
[0003] Field The present disclosure relates to compositions, methods, and systems for analyzing polymeric macromolecules, including polymeric macromolecules such as peptides, polypeptides, and proteins.
Background Art
[0004] Background Proteins are fundamental to cell function. Thus, the sequences and their concentrations of the thousands of proteins within each cell are important indicators of cell health. Abnormal protein sequences or concentrations may indicate a disease state. However, tools and techniques for the sensitive, accurate, economical, and unbiased characterization of the proteome are currently lacking. Early detection of abnormal sequences and / or concentrations is important for the diagnosis and treatment of many diseases, such as cancer. For these and other reasons, better tools for assessing the sequences and concentrations of proteins and peptides in biological samples should be developed.
[0005] When such tools become available, the discovery of new biomarkers, the accurate determination of the concentration of even the lowest-abundance proteins, the discovery of important post-translational modifications, and the monitoring of proteome dynamics are some of the first steps towards improving healthcare. These initial steps towards a deeper understanding and earlier detection of important signatures of cancer and other health conditions enable early-stage diagnosis, facilitate therapeutic discovery, and inform the course of treatment, thereby having a beneficial impact on patient care.
[0006] Accordingly, there is a need in the art for compositions, methods, and systems of materials for highly parallelized, accurate, sensitive, high-throughput proteome analysis. The present disclosure addresses this need and other needs. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0007] Abstract The present disclosure relates to compositions, methods, and systems for analyzing polymeric macromolecules, including peptides, polypeptides, and proteins, in a highly parallel and high-throughput manner by recoding their sequences into DNA polymers.
[0008] In some embodiments, a method for determining the identity and positional information of amino acid residues of a peptide coupled to a solid support is disclosed herein. The method comprises: (a) providing a peptide to the solid support, wherein the peptide is coupled to the solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) providing a chemically reactive conjugate, wherein the chemically reactive conjugate comprises: (x) a cycle tag comprising a cyclic nucleic acid associated with the number of cycles; (y) a reactive moiety for binding to the N-terminal amino acid residue of the peptide; and (z) an immobilization moiety for immobilization to the solid support; (c) contacting the peptide with the chemically reactive conjugate, thereby coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex to the solid support via the immobilization moiety; (e) cleaving the N-terminal amino acid residue and separating it from the peptide, thereby providing an immobilized amino acid complex, wherein the immobilized amino acid complex comprises the cleaved and separated N-terminal amino acid residue; (f) contacting the immobilized amino acid complex with a binder, wherein the binder comprises a binding moiety for preferentially binding to the immobilized amino acid complex and a recoding tag comprising a recoding nucleic acid corresponding to the binder, thereby forming an affinity complex, wherein the affinity complex comprises the immobilized amino acid complex and the binder, thereby bringing the cycle tag into proximity to the recoding tag within the affinity complex; (g) transmitting the information of the recoding nucleic acid to the cyclic nucleic acid of the immobilized conjugate complex to generate a recoding block; (j) obtaining the sequence information of the recoding block; and (k) determining the identity and positional information of the amino acid residues of the peptide based on the obtained sequence information. In some embodiments, cleaving the N-terminal amino acid residue from the peptide exposes the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide. In some embodiments, the reactive moiety of the chemically reactive conjugate cleaves the N-terminal amino acid residue from the peptide.Some embodiments include repeating steps (b) to (k) for each subsequent amino acid of the peptide. Some embodiments include washing the immobilized amino acid complex before contacting the immobilized amino acid complex with the binder. Some embodiments include determining a possible three-dimensional structure of the peptide based on sequence information. In some embodiments, the recoded nucleic acid comprises DNA or RNA. In some embodiments, the cyclic nucleic acid comprises DNA or RNA. In some embodiments, obtaining sequence information of the recoded block comprises performing sequencing. In some embodiments, the binding moiety comprises a peptide, an antibody, an antibody fragment, or an antibody derivative. In some embodiments, the binding moiety comprises an aptamer. In some embodiments, the binding moiety binds to a natural amino acid, a post-translationally modified amino acid, a derivatized version of an amino acid, a derivatized or stabilized version of a post-translationally modified amino acid, a synthetic amino acid, an amino acid having a specific side chain, an amino acid having a phosphorylated side chain, an amino acid having a glycosylated side chain, an amino acid having a methylation modification, or a D-amino acid of an amino acid, or a combination thereof. In some embodiments, the solid support comprises beads, plates, or chips. In some embodiments, the solid support comprises a glass slide, silica, resin, gel, membrane, polystyrene, metal, nitrocellulose, mineral, plastic, polyacrylamide, latex, or ceramic. In some embodiments, the peptide comprises a hormone, a neurotransmitter, an enzyme, an antibody, a viral protein, a bacterial protein, a synthetic peptide, a bioactive peptide, a peptide hormone, an oligopeptide, a polypeptide, a fusion protein, a cyclic peptide, a branched peptide, a recombinant protein, a tumor marker, a therapeutic peptide, an antigenic peptide, or a signaling peptide. In some embodiments, the peptide is derived from a cell lysate, a blood sample, a plasma sample, a serum sample, a tissue biopsy material, a saliva sample, a urine sample, a cerebrospinal fluid sample, a sweat sample, a synovial fluid sample, a fecal sample, an intestinal microbiota sample, an environmental water sample, a soil sample, a bacterial culture, a viral culture, an organoid, a tumor biopsy material, a sputum sample, or a hair sample. In some embodiments, the peptide is associated with a disease.In some embodiments, transmitting the information comprises performing nucleic acid amplification, enzymatic ligation, splint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a splint oligonucleotide, use of a catalyst, use of a crosslinking molecule, use of a condensing agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-stranded binding protein, a click chemistry reaction, phosphodiester bond formation, or peptide nucleic acid-mediated ligation. In some embodiments, the information of the recoded nucleic acid comprises the sequence of the recoded nucleic acid or the reverse complement of the sequence of the recoded nucleic acid. In some embodiments, transmitting the information comprises joining the recoded nucleic acid or the reverse complement of the recoded nucleic acid to a cycled nucleic acid.
[0009] In some embodiments, methods are disclosed herein for determining the identity and positional information of a plurality of amino acid residues of a peptide, the peptide comprising n amino acid residues, the method comprising: (a) coupling the peptide to a solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support but is exposed to reaction conditions; (b) providing a chemically reactive conjugate, the chemically reactive conjugate comprising: (x) a cycle tag comprising a cyclic nucleic acid associated with the number of cycles; (y) a reactive moiety for binding to and cleaving the N-terminal amino acid residue of the peptide to expose the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide; and (z) an immobilization moiety for immobilization to the solid support; (c) contacting the peptide with the chemically reactive conjugate, thereby coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex to the solid support via the immobilization moiety; (e) cleaving the N-terminal amino acid residue and separating it from the peptide, thereby exposing the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide, providing an immobilized amino acid complex, the immobilized amino acid complex comprising the cleaved and separated N-terminal amino acid residue; (f) repeating (b) to (e) n - 1 times to assemble n - 1 additional immobilized amino acid complexes, each additional immobilized amino acid complex comprising a nucleic acid associated with cycles 2 to n, respectively; (g) contacting the immobilized amino acid complex with a binder, the binder comprising a binding moiety for preferentially binding to one or a subset of the immobilized amino acid complexes and a recoding tag comprising a recoding nucleic acid corresponding to the binder, thereby forming one or more affinity complexes, each affinity complex comprising the immobilized amino acid complex and the binder, thereby bringing the cycle tag into proximity with the recoding tag in each of the formed affinity complexes; (h) joining the cycle tag or its reverse complement to the recoding tag within each of the formed affinity complexes to form a recoding block, thereby creating a plurality of recoding blocks, each recoding block corresponding to the formed affinity complex,Creating a plurality of recoding blocks, (i) joining two or members of the plurality of recoding blocks to form a memory oligonucleotide, (j) obtaining sequence information of the memory oligonucleotide, and (k) based on the obtained sequence information, determining the identity and position information of a plurality of amino acid residues of a peptide. In some embodiments, (g)-(h) are repeated 2, 3, 4 times or more. In some embodiments, n is an integer greater than or equal to 2. In some embodiments, each binder comprises a recoding tag having a unique nucleic acid sequence. In some embodiments, the plurality of binders comprise recoding tags having the same nucleic acid sequence. In some embodiments, the binder comprises a recoding tag having a unique sequence portion and a common sequence portion. Some embodiments include deprotecting the cycle tag between (f) and (g). Some embodiments include washing the immobilized amino acid complex before contacting the immobilized amino acid complex with the binder. Some embodiments include determining a possible three-dimensional structure of the peptide based on the sequence information. In some embodiments, the recoding nucleic acid comprises DNA. In some embodiments, the cycle nucleic acid comprises DNA. Some embodiments in which obtaining the sequence information of the memory oligonucleotide includes performing sequencing. In some embodiments, the binding moiety comprises an antibody or a fragment thereof. In some embodiments, the binding moiety binds to a natural amino acid, a derivatized amino acid, a synthetic amino acid, or a D-amino acid. In some embodiments, the binding moiety binds to a post-translationally modified amino acid. In some embodiments, the solid support comprises beads, plates, or chips. In some embodiments, the solid support comprises a glass slide, silica, resin, gel, membrane, polystyrene, metal, nitrocellulose, mineral, plastic, polyacrylamide, latex, or ceramic. In some embodiments, determining the identity and position information of a plurality of amino acid residues of the peptide includes determining the identity and position information of all amino acid residues of the peptide. In some embodiments,Determining the identity and position information of multiple amino acid residues of a peptide includes determining the identity and position information of amino acid residues of only a subset of the peptide. Some embodiments include identifying the peptide by comparing the identity and position information of the multiple amino acid residues to a database.,
[0010] In some embodiments, a chemically reactive conjugate (CRC) is disclosed herein that includes (A) a nucleic acid sequence tag, (B) a reactive moiety for attaching and cleaving an N-terminal amino acid residue to a peptide, and (C) an immobilization moiety for immobilizing on a solid support. Some embodiments include Formula I:
Chemical formula
Chemical formula
Chemical formula
[0011] In some embodiments, kits for determining the identity and positional information of amino acid residues of a peptide are disclosed herein. The kits include (a) nucleic acid sequence tags, and (b) a chemically reactive conjugate including a reactive moiety that couples to the N-terminal amino acid residue of the peptide, thereby forming a conjugate complex including a chemically reactive conjugate coupled to the N-terminal amino acid of the peptide, a binding agent including a binding moiety for preferentially binding to the conjugate complex, and a recoding tag including a recoding nucleic acid corresponding to the binding agent, and reagents for transmitting the information of the recoding nucleic acid to the cyclic nucleic acid of the conjugate complex to generate a recoding block.
[0012] In some embodiments, methods for sequencing a subset of nucleotides of an oligonucleotide are disclosed herein. The methods include providing, in a nucleic acid sequencing reaction, a combination of reversibly terminated nucleotides and non-reversibly terminated nucleotides, wherein the nucleotides of the nucleic acid to be sequenced corresponding to the non-reversibly terminated nucleotides are not sequenced. Some embodiments include identifying the nucleotides of the nucleic acid to be sequenced corresponding to the reversibly terminated nucleotides. In some embodiments, the nucleic acid to be sequenced includes a region including only a subset of nucleotides selected from A, C, G, and T, and the subset of nucleotides is not sequenced. In some embodiments, the subset of nucleotides selected from A, C, G, and T includes two nucleotides selected from A, C, G, and T. In some embodiments, the subset of nucleotides selected from A, C, G, and T includes three nucleotides selected from A, C, G, and T. In some embodiments, the region includes a primer sequence. In some embodiments, the region does not include a barcode sequence, a recoding nucleic acid sequence or a part thereof, or a cyclic nucleic acid sequence or a part thereof.
[0013] In some embodiments, a method is disclosed herein, the method comprising providing a conjugate comprising a reactive molecule coupled to a protected oligonucleotide; contacting the reactive moiety with a terminal amino acid of a peptide, thereby binding the reactive moiety to the terminal amino acid and optionally cleaving the terminal amino acid from the peptide; deprotecting the oligonucleotide; and contacting the deprotected oligonucleotide with an enzyme or reagent for ligation or polymerization. Some embodiments include reprotecting the oligonucleotide. In some embodiments, the reactive moiety cleaves the terminal amino acid from the peptide to expose the next terminal amino acid, and the method further comprises contacting the next amino acid with another conjugate after reprotecting the oligonucleotide. In some embodiments, the terminal amino acid is an N-terminus. In some embodiments, the peptide is immobilized on a solid support. In some embodiments, the conjugate comprises an organic small molecule. In some embodiments, the conjugate comprises a chemically reactive conjugate (CRC), the chemically reactive conjugate (CRC) comprising (A) an oligonucleotide, (B) a reactive moiety, and (C) an immobilizing moiety. In some embodiments, the oligonucleotide comprises a cyclic nucleic acid.
[0014] In some embodiments, a method is disclosed herein, the method comprising providing a conjugate comprising a peptide coupled to a protected oligonucleotide, contacting a terminal amino acid of the peptide, thereby attaching a reactive moiety to the terminal amino acid, optionally cleaving the terminal amino acid from the peptide, deprotecting the oligonucleotide, and contacting the deprotected oligonucleotide with an enzyme or reagent for ligation or polymerization. Some embodiments include reprotecting the oligonucleotide. In some embodiments, the reactive moiety cleaves the terminal amino acid from the peptide to expose the next terminal amino acid, and the method further comprises contacting the next amino acid with another conjugate after reprotecting the oligonucleotide. In some embodiments, the terminal amino acid is the N-terminus. In some embodiments, the peptide is immobilized on a solid support. In some embodiments, the conjugate comprises an organic small molecule.
[0015] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Other features, details, utilities, and advantages of the claimed subject matter will become apparent from the following detailed description, which includes the embodiments shown in the accompanying drawings and aspects defined in the appended claims. These aspects of the disclosure, as well as other features and advantages, will be described in more detail below.
Brief Description of the Drawings
[0016] To enable a more particular understanding of the above-described features of the present disclosure, a more specific description of the present disclosure, briefly summarized above, may be obtained by reference to the embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the accompanying drawings illustrate only exemplary embodiments and should not be regarded as limiting the scope thereof, and other equally effective embodiments may be recognized.
[0017] Accordingly, the foregoing and other features and advantages of the present disclosure will be more fully understood from the following detailed description of the exemplary embodiments, taken in conjunction with the accompanying drawings.
[0018]
Figure 1
[0019]
Figure 2
[0020]
Figure 3-1
Figure 3-2
Figure 3-3
Figure 3-4
Figure 3-5
Figure 3-6
[0021]
Figure 4
[0022]
Figure 5
[0023]
Figure 6
[0024]
Figure 7
[0025]
Figure 8-1
Figure 8-2
Figure 8-3
[0026]
Figure 9
[0027]
Figure 10
[0028]
Figure 11
[0029]
Figure 12-1
Figure 12-2
[0030]
Figure 13
[0031]
Figure 14-1
Figure 14-2
Figure 14-3
Figure 14-4
Figure 14-5
Figure 14-6
[0032]
Figure 15
[0033]
Figure 16
[0034]
Figure 17-1
Figure 17-2
[0035]
Figure 18
[0036]
Figure 19
[0037]
Figure 20A
Figure 20B
[0038]
Figure 21
[0039]
Figure 22
[0040]
Figure 23
[0041]
Figure 24
[0042]
Figure 25
[0043]
Figure 26-1
Figure 26-2
[0044]
Figure 27
[0045]
Figure 28
[0046]
Figure 29
[0047]
Figure 30
[0048]
Figure 31
[0049]
Figure 32A
Figure 32B-1
Figure 32B-2
Figure 32B-3
[0050]
Figure 33
[0051]
Figure 34
[0052]
Figure 35
[0053]
Figure 36A-D
[0054]
Figure 37
[0055] It should be understood that the drawings are not necessarily to scale and that like reference numerals refer to like features. Elements and features of one embodiment may be beneficially incorporated into other embodiments without further recitation.
BRIEF DESCRIPTION OF THE DRAWINGS
[0056] DETAILED DESCRIPTION The methods and compositions described herein may be useful for determining the identity and positional information of amino acid residues of a peptide. The peptide may be coupled to a solid support, the N-terminal amino acid of the peptide may be cleaved, and contacted with a chemically reactive conjugate that couples the N-terminal amino acid to the solid support with a cycle tag. This may then be contacted with a binding agent, such as a binding agent specific for the N-terminal amino acid. The binding agent may include a recode tag. The cycle tag and the recode tag may include nucleic acid information that can be sequenced to obtain the identity and positional information of the N-terminal amino acid. This process may be repeated for various amino acids of the peptide. Thus, the position and information of the amino acid residues of a protein can be recoded using nucleic acids and obtained by sequencing the nucleic acids.
[0057] In some embodiments, a method for determining the identity and positional information of amino acid residues of a peptide coupled to a solid support is disclosed herein. The method comprises: (a) providing a peptide to the solid support, wherein the peptide is coupled to the solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) providing a chemically reactive conjugate, wherein the chemically reactive conjugate comprises: (x) a cycle tag comprising a cyclic nucleic acid associated with the number of cycles; (y) a reactive moiety for binding to the N-terminal amino acid residue of the peptide; and (z) an immobilization moiety for immobilization to the solid support; (c) contacting the peptide with the chemically reactive conjugate, thereby coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex to the solid support via the immobilization moiety; (e) cleaving the N-terminal amino acid residue and separating it from the peptide, thereby providing an immobilized amino acid complex, wherein the immobilized amino acid complex comprises the cleaved and separated N-terminal amino acid residue; (f) contacting the immobilized amino acid complex with a binder, wherein the binder comprises a binding moiety for preferentially binding to the immobilized amino acid complex and a recoding tag comprising a recoding nucleic acid corresponding to the binder, thereby forming an affinity complex, wherein the affinity complex comprises the immobilized amino acid complex and the binder, thereby bringing the cycle tag into proximity to the recoding tag within the affinity complex; (g) transferring the information of the recoding nucleic acid to the cycle nucleic acid of the immobilized conjugate complex to generate a recoding block; (j) obtaining the sequence information of the recoding block; and (k) determining the identity and positional information of the amino acid residues of the peptide based on the obtained sequence information. Some embodiments include repeating any or all of steps (b) to (k) for each subsequent amino acid of the peptide.
[0058] In some embodiments, methods are disclosed herein for determining the identity and positional information of amino acid residues of a peptide coupled to a solid support. The method can include providing the peptide to the solid support. In some embodiments, the peptide is coupled to the solid support such that, for example, the N-terminal amino acid residue of the peptide is not directly coupled to the solid support or is exposed to reaction conditions. The method can include providing a chemically reactive conjugate. The chemically reactive conjugate can include a cycle tag. The cycle tag can include a cyclic nucleic acid associated with the number of cycles. The chemically reactive conjugate can include a reactive moiety. The reactive moiety can be useful for binding the N-terminal amino acid residue of the peptide. The chemically reactive conjugate can include an immobilization moiety. The immobilization moiety can be useful for immobilization to the solid support. The method can include contacting the peptide with the chemically reactive conjugate. By contacting the peptide with the chemically reactive conjugate, the chemically reactive conjugate can be coupled to the N-terminal amino acid of the peptide to form a conjugate complex. The method can include immobilizing the conjugate complex to the solid support, for example, via the immobilization moiety. The method can include cleaving or separating the N-terminal amino acid residue from the peptide. By cleaving or separating the N-terminal amino acid residue from the peptide, an immobilized amino acid complex can be provided. The immobilized amino acid complex can include the cleaved and separated N-terminal amino acid residue. The method can include contacting the immobilized amino acid complex with a binder. The binder can include a binding moiety. The binding moiety can be useful for preferential binding to the immobilized amino acid complex. The binder can include a recoding tag. The recoding tag can include a recoding nucleic acid corresponding to the binder. By contacting the immobilized amino acid complex with the binder, an affinity complex can be formed. The affinity complex can include the immobilized amino acid complex. The affinity complex can include the binder. By contacting the immobilized amino acid complex with the binder, the cycle tag can be brought into proximity to the recoding tag, for example, within the affinity complex. The method can include transmitting the information of the recoding nucleic acid to the cycle nucleic acid.Thereby, a recoded block can be generated. The recoded block can be assembled into a memory oligonucleotide. The method can include joining one or more recoded blocks made from one or more amino acid residues. The method can include a step of obtaining sequence information of the recoded block. The method can include obtaining sequence information of the memory oligonucleotide. The method can include determining information on amino acid residues of the peptide based on the obtained sequence information. The information can include identification information. The information can include position information. In some embodiments, cleaving the N-terminal amino acid residue from the peptide exposes the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide. In some embodiments, the reactive portion of the chemical reactive conjugate cleaves the N-terminal amino acid residue from the peptide. Some embodiments include repeating any of the foregoing steps for each subsequent amino acid of the peptide. In some embodiments, the immobilized portion includes an alkyne which is a chemically activatable moiety. Some embodiments include joining the chemical moiety to a solid support. In some embodiments, cleaving the N-terminal amino acid residue from the peptide exposes the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide. In some embodiments, the reactive portion of the chemical reactive conjugate cleaves the N-terminal amino acid residue from the peptide. Some embodiments include washing away the chemically reactive conjugate not joined to the solid support before contacting the next N-terminal amino acid of the peptide with the chemical reactive complex. Some embodiments include contacting the immobilized amino acid complex with a binder to form an affinity complex. Some embodiments include washing the immobilized amino acid complex before contacting the immobilized amino acid complex with the binder. Some embodiments include washing the immobilized amino acid affinity complex after contacting the affinity complex with one or a set of binders.
[0059] In some embodiments, a method for determining the identity and positional information of a plurality of amino acid residues of a peptide, wherein the peptide comprises n amino acid residues, comprising: (a) coupling the peptide to a solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) providing a chemically reactive conjugate, the chemically reactive conjugate comprising: (x) a cycle tag comprising a cyclic nucleic acid associated with the number of cycles; (y) a reactive moiety for binding to and cleaving the N-terminal amino acid residue of the peptide to expose the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide; and (z) an immobilization moiety for immobilization on the solid support; (c) contacting the peptide with the chemically reactive conjugate, thereby coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex on the solid support via the immobilization moiety; (e) cleaving the N-terminal amino acid residue and separating it from the peptide, thereby exposing the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide and providing an immobilized amino acid complex, the immobilized amino acid complex comprising the cleaved and separated N-terminal amino acid residue; (f) repeating (b) to (e) n - 1 times to assemble n - 1 additional immobilized amino acid complexes, each additional immobilized amino acid complex comprising a nucleic acid associated with cycles 2 to n accordingly; (g) contacting the immobilized amino acid complexes with a binder or a set of binders, each binder comprising a binding moiety for preferentially binding to one or a subset of the immobilized amino acid complexes and a recoding tag comprising a recoding nucleic acid corresponding to the binder, thereby forming one or more affinity complexes, each affinity complex comprising an immobilized amino acid complex and a binder, thereby bringing the cycle tag into proximity to the recoding tag in each formed affinity complex; (h) in each formed affinity complex, ligating the cycle tag or its reverse complement to the recoding tag to form a recoding block, orOtherwise, transfer the information of the recoded nucleic acid to the cyclic nucleic acid of the immobilized conjugate complex, thereby creating a plurality of recoded blocks, each recoded block corresponding to the formed affinity complex, creating a plurality of recoded blocks, and (i) joining two or more members of the plurality of recoded blocks to form a memory oligonucleotide; (j) obtaining the sequence information of the memory oligonucleotide; and (k) determining the identity and position information of a plurality of amino acid residues of the peptide based on the obtained sequence information. A method is disclosed herein.
[0060] In some embodiments, methods are disclosed herein for determining the identity and positional information of a plurality of amino acid residues of a peptide. The peptide can include n amino acid residues. The method can include coupling the peptide to a solid support. The coupling can be such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support or is exposed to the reaction conditions. The method can include providing a chemically reactive conjugate. The chemically reactive conjugate can include a cycle tag that includes a cyclic nucleic acid related to the number of cycles. The chemically reactive conjugate can include a reactive moiety. The reactive moiety can bind and / or cleave to the N-terminal amino acid residue of the peptide. The reactive moiety can expose the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide. The chemically reactive conjugate can include an immobilization moiety for immobilization to the solid support. The method can include contacting the peptide with the chemically reactive conjugate. Such contact can couple the chemically reactive conjugate to the N-terminal amino acid of the peptide and form a conjugate complex. The method can include immobilizing the conjugate complex to the solid support. The immobilization can be via the immobilization moiety. The method can include cleaving and thereby separating the N-terminal amino acid residue. The cleavage can expose the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide. The method can include providing an immobilized amino acid complex. The immobilized amino acid complex can include the cleaved and separated N-terminal amino acid residue. The method can include repeating the steps n - 1 times to assemble n - 1 additional immobilized amino acid complexes. The additional immobilized amino acid complexes can include nucleic acids related to cycles 2 to n. The method can include contacting the immobilized amino acid complex with one or a set of binders. The binder can include a binding moiety for preferentially binding to one or a subset of the immobilized amino acid complexes. The binder can include a recoding tag. The recoding tag can include a recoding nucleic acid corresponding to the binder. Contacting the immobilized amino acid complex with one or more binders can form one or more affinity complexes.The affinity complex can include an immobilized amino acid complex and a binder. By contacting the immobilized amino acid complex with the binder, a cycle tag can be brought into proximity to a recode tag within the formed affinity complex. The method can include joining a cycle tag or its reverse complement to a recode tag within each formed affinity complex. The joining can form a recode block. The joining or method can include creating a plurality of recode blocks. Each recode block can correspond to a formed affinity complex. The method can include joining two or more members of the plurality of recode blocks to form a memory oligonucleotide. The method can include obtaining sequence information of the memory oligonucleotide. The method can include determining the identity and position information of a plurality of amino acid residues of a peptide based on the obtained sequence information. In some embodiments, n is an integer greater than or equal to 2. In some embodiments, each binder includes a recode tag having a unique nucleic acid sequence. In some embodiments, the plurality of binders include recode tags having the same nucleic acid sequence. In some embodiments, the binder includes a recode tag that can have a unique sequence portion and a common sequence portion.
[0061] In some embodiments, disclosed herein is a chemically reactive conjugate comprising (a) a nucleic acid sequence tag, (b) a reactive moiety for binding and cleaving an N-terminal amino acid residue from a peptide, and (c) an immobilization moiety for immobilization to a solid support.
[0062] In some embodiments, disclosed herein is a chemically reactive conjugate. The chemically reactive conjugate can include a nucleic acid sequence tag. The chemically reactive conjugate can include a reactive moiety. The reactive moiety can be useful for binding an N-terminal amino acid residue. The reactive moiety can be useful for cleaving an N-terminal amino acid residue from a peptide. The chemically reactive conjugate can include an immobilization moiety. The immobilization moiety can be useful for immobilization to a solid support. Also disclosed is a kit containing any of the components described herein.
[0063] Introduction The sequences and concentrations of cells and secreted proteins are useful indicators of cell health. Abnormal sequences or concentrations may indicate a disease state. However, tools and techniques for the sensitive, accurate, economical, and unbiased characterization of the proteome are currently lacking. Early detection of abnormal sequences and / or concentrations is important for the diagnosis and treatment of many diseases, such as cancer. For these and other reasons, better tools for assessing the sequences and concentrations of proteins and peptides in biological samples must be developed.
[0064] Next-generation sequencing (NGS) of DNA and RNA polymers has transformed diagnostic, clinical, and research approaches by enabling clinicians and researchers to analyze billions of DNA sequences with high throughput and low cost. However, the ability to detect and quantify proteins and peptides has lagged that of nucleic acids because there is no equivalent of the polymerase chain reaction (PCR) for amino acid polymers. New tools for sensitively quantifying proteins and evaluating their sequences can, like NGS, help in understanding cell processes, continue to transform research, diagnostic, and clinical approaches, and facilitate precision medicine.
[0065] Current state-of-the-art proteomics toolkits include: 1) Edman degradation followed by conventional chromatography; 2) fragmentation followed by high-resolution separation and mass spectrometry (MS) techniques; and 3) common approaches for the recognition of proteins via affinity molecules. These methods provide very useful information. However, none of these approaches generates information at the scale, throughput, reproducibility, accessibility, or cost required to unlock transformative applications in research, diagnosis, or treatment.
[0066] Peptide sequencing based on Edman degradation was first proposed and automated in 1950 by Pehr Edman. This process is similar to Sanger sequencing. Briefly, a series of chemical reactions and downstream HPLC analysis are used to collect peptide sequence information by stepwise degradation of the N-terminal amino acid on the peptide. First, the N-terminal amino acid is reacted with phenylisothiocyanate (PITC) under basic conditions (typically NMP / methanol / water) to form a phenylthiocarbamoyl (PTC) derivative. In the second step, the PTC-modified amino group is treated with an acid (typically TFA anhydride) to obtain an ATZ-modified (2-anilino-5(4)-thiazolinone) amino acid, separate the amino acid from the polymer, and create the next N-terminus on the polypeptide. The cyclic ATZ-amino acid is converted to a PTH-amino acid derivative and analyzed by chromatography. These steps are then sequentially repeated to determine the peptide sequence. This is effective, but has high prior protein sample requirements, and the process lacks throughput and cost for supporting large-scale discovery.
[0067] More recently, multiplexed methods and devices for Edman degradation-based peptide sequencing of trace amounts of protein have been developed. See, for example, Chharbra U.S. Patent No. 7,611,834 B2. However, such methods and devices are still not suitable for highly parallelized high-throughput proteome analysis.
[0068] Over the past 20 years, peptide analysis by fragmentation and analysis by mass spectrometry (here, LC / MS) have been increasingly used to quantify protein abundance and determine sequences. Additionally, in certain applications, recognition-based proteomics has been used. In this approach, affinity molecules such as antibodies or antibody fragments, aptamers, RNAs, or modified proteins are generally engineered to recognize the tertiary structure of the analyte. Often, these are linked to molecular beacons that emit fluorescence or provide other means of detecting binding events such as ELISA assays. However, like previous methods, fragmentation and recognition-based methods lack the throughput and efficiency to support large-scale discovery.
[0069] The present disclosure provides methods for analyzing polymeric macromolecules such as peptides, polypeptides, and proteins. Accordingly, aspects of the present disclosure relate to the field of proteomics.
[0070] Figure 1 shows the segmentation of the field of proteomics by technology. As noted above, the current state of proteome analysis includes: 1) Edman degradation followed by conventional chromatography; 2) fragmentation followed by advanced separation and mass spectrometry techniques; and 3) the general approach of protein recognition via affinity molecules. These (and other) approaches can provide useful information to researchers, but they do not provide such information at the scale, throughput, or cost necessary to unlock transformative applications in research, diagnosis, or therapy. Some more specific challenges associated with current approaches (e.g., Edman, LC / MS, and affinity approaches) include the following: (a) Protein folding is dynamic and proteins can lose their characteristic shape. Then, recognition-based methods become inaccurate. This can occur for unstable proteins or in cases of uncontrolled sample handling prior to analysis. (b)Knowledge-based methods do not provide information regarding whether a protein sequence is a catalytically inactive variant, as is often the case in cancer biology. (c)The biomarker of interest is likely to be present at concentrations below fM, below the detection limit of most available tools used to quantify the absolute abundance of multiple proteins. (d)The universe of protein molecules is vast. Due to the additional diversity introduced by post-translational modifications (PTMs), it is far more complex than the RNA transcriptome. (e)Intracellular proteins change dynamically (in terms of expression level and modification state) in response to the environment, physiological state, and disease state. Thus, proteins contain a vast amount of relevant information that has been scarcely investigated. (f)Generating an effective ensemble of affinity agents with low cross-reactivity between off-target macromolecules can be time-consuming. (g)Multiplexing the readout of an ensemble of affinity agents that minimizes cross-reactivity between the affinity agent and off-target macromolecules is difficult. (h)The automation around existing methods and current approaches is slow, costly, and in the case of the Edman method, limited to processing only a few peptides per day. (i)LC / MS has drawbacks including high instrument costs, requirements for sophisticated users, insufficient quantification ability, and insufficient dynamic range. Since proteins ionize with different efficiencies, absolute and even relative quantification between samples is difficult. (j)Since LC / MS analyzes more abundant species, complex pre-sample preparation, such as using nanoparticle coronas, is required, and characterizing low-abundant proteins is difficult. (k)The throughput of LC / MS samples is typically limited to thousands of peptides per run. (l) Recent attempts to develop methodologies have suffered from insufficient discrimination of N-terminal or C-terminal AAs and are confounded by the "neighboring effect." Still other single-molecule methodologies under development are unable to amplify the analyte and require expensive equipment because they have to detect a small number of photons, electrons, or detection elements.
[0071] The present disclosure addresses the above and other needs by providing methods, systems, and compositions for analyzing polymeric macromolecules via recoding of sequences into DNA polymers for subsequent DNA sequencing and analysis. Referring to FIG. 1, certain embodiments of the present disclosure are included in the "Proteomics" >> "Cutting Edge Research" >> "Chemistry" >> "Sequence-Based" segments. Numerous applications of the present disclosure include peptide sequence and quantification determination in synthetic and biologically derived samples containing multiple protein complexes, proteins, and / or polypeptide components.
[0072] Some embodiments of the methods described herein include: 1) a step of binding to a substrate; 2) a step of functionalized PITC conjugation to an amino acid; 3) a step of immobilizing the PITC conjugate to a hydrogel substrate; 4) a step of cleaving the amino acid by Edman degradation; 4a) nucleotide deprotection; 5) a step of constructing a recoding block with a binder; 6) a memory oligo assembly step; and 7) a step of releasing the oligo for sequencing.
[0073] Improved methods for determining the sequence and presence of proteins Figure 2 shows a simplified block diagram of an exemplary workflow 200 for analyzing polymer macromolecules according to an embodiment of the present disclosure. More specifically, workflow 200 includes a high-level overview of the various methods herein and how such methods are synergistically compatible with DNA sequencing technology. As shown, a sample of macromolecules, such as proteins and peptides, is prepared and immobilized on a solid support (box 1). While immobilized, the amino acid sequence of the macromolecule is converted to a DNA sequence (e.g., “recoded”) (box 2), and the DNA sequence is amplified into a library for NGS sequencing (box 3). The DNA library is then sequenced (box 4) and analyzed (box 5) by high-throughput and high-precision methods, thereby enabling low cost.
[0074] Figure 3 schematically shows the various operations of the workflow of Figure 2 according to an embodiment of the present disclosure. More specifically, Figure 3 shows the main stages of the “recoding” operation of Figure 2 as process 300. As illustrated, the recoding process 300 has three separate and separable stages, each stage being shown in a row of operations.
[0075] In the first stage (operations 1-4 of Figure 3), periodic information is captured. In operation 1, the surface of the solid support is prepared for attachment of a polymeric analyte, a set of universal primers, and a trifunctional chemically reactive conjugate. This can be achieved by providing three (or more) orthogonal chemistries on the surface of the support, shown as aldehyde-hydrazine, azide-alkyne, and thiol in Figure 3. It should be noted that multiple conjugation chemistries are possible, including alternative chemical functional groups for immobilizing the primers, polymeric analyte, and chemically reactive conjugate to the solid support, as described below. Using at least one of the orthogonal reaction modes of the support, multiple polymeric analytes, such as proteins, protein fragments (i.e., peptides), or other polymers, are immobilized on the support surface.
[0076] Next, perform operations 2-4 to immobilize the trifunctional chemical reactive conjugate (as a conjugate-AA-cycle tag complex). As shown, in operation 2, the N-terminal amino acid of the immobilized analyte is contacted with a chemical reactive conjugate containing a reactive group for the amino terminus, such as an Edman reagent (phenylisothiocyanate (PITC)), an orthogonal reactive group for the support, and a nucleic acid molecule carrying information regarding the cycle when the conjugate is contacted with the analyte. Under basic conditions, the PITC conjugate reacts with the N-terminal amino acid to form a phenylthiocarbamoyl-amino acid (PTC) conjugate. By stringent washing, unreacted PITC conjugate is removed, and then, in operation 3, activation of the orthogonal chemistry used to tether the conjugate to the support is initiated, and the PTC conjugate is immobilized in proximity to the anchor point of the relevant analyte. For example, by changing the redox conditions to induce dithiol formation or adding stabilizers and redox components to induce a click reaction, a PTC-thiol conjugate or a PTC-alkyne complex can be immobilized on the solid support. Following immobilization of the conjugate to the solid support, a conjugate reactive scavenger can be added to cap the reactivity of the bound conjugate not washed away in the previous step and render it inert to future N-terminal amino acid reactions. In operation 4, peptide bond cleavage targeting the N-terminal amino acid of the peptide is induced. In an example using Edman degradation chemistry, this is facilitated by a change in pH from basic to acidic conditions. Then, operations 1-4 can be repeated for n cycles to generate a lawn of n cycle-tagged conjugates localized on the solid support. 2+ 、stabilizers and redox components can be added to induce a click reaction to immobilize the PTC-thiol conjugate or the PTC-alkyne complex on the solid support. Following immobilization of the conjugate to the solid support, a conjugate reactive scavenger can be added to cap the reactivity of the bound conjugate not washed away in the previous step and render it inert to future N-terminal amino acid reactions. In operation 4, peptide bond cleavage targeting the N-terminal amino acid of the peptide is induced. In an example using Edman degradation chemistry, this is facilitated by a change in pH from basic to acidic conditions. Then, operations 1-4 can be repeated for n cycles to generate a lawn of n cycle-tagged conjugates localized on the solid support.
[0077] In fact, the first iteration of operations 2-4 (i.e., the first cycle) provides information about the terminal monomer of the immobilized polymer analyte. The second cycle thereof provides information about the next monomer of the immobilized polymer analyte, etc. By repeating steps 2-4 for n cycles, a spatially localized conjugate lawn that holds cycle information is created. By placing an appropriate spacing between the anchor points of the immobilized polymer analyte, conjugates related to a single analyte are placed in the same location and isolated from conjugates of other analytes.
[0078] The second row of Figure 3 shows the operations of the iterative process in which the recoding block is constructed. In this operation, amino acid information and cycle information are associated. Briefly, a plurality of binders that recognize the immobilized conjugate-AA-cycle tag complex are introduced in operation 5a and bound to their cognate targets in operation 5b. The binders are engineered to preferentially recognize specific conjugates based on the differences in cognate amino acids of the immobilized conjugate. Thereby, a drug that holds both cognate AA and cognate cycle information instructs ligation of the AA information to the cycle tag of the corresponding conjugate complex (operations 5c and 5d). By repeating binding, washing, and ligation, multiple attempts are possible to find cognate partners and transfer information to each immobilized conjugate-AA-cycle tag complex to construct the recoding block. Thus, using multiple consecutive binding cycles, the yield of information transfer from the binder to the immobilized conjugate can be driven to arbitrarily high completion levels.
[0079] In the third row of FIG. 3, the formed recoded blocks are assembled into memory oligos (e.g., combined into a single memory oligo). This oligo can be amplified on a solid support or in solution and then analyzed using DNA sequencing to determine the sequence and / or abundance of the immobilized analyte. Briefly, in operation 6, the co-localized recoded blocks interact based on their complementary DNA sequences to assemble DNA oligonucleotides that represent the sequence of the original macromolecule. This process is similar to the g-block assembly of gene products. The assembly can be facilitated by a polymerase extension-ligation process or by a ligation process.
[0080] Gaps in connectivity between co-localized conjugates can exist, for example, due to a) incomplete information accumulation during sequential degradation of peptides and immobilization of PTC-AA-cycle tag conjugate complexes, b) incomplete information transfer from recode tags to cycle tags during recoded block assembly, or c) simple incomplete ligation of available existing recoded block information during memory oligo assembly. To improve these gaps and enable high-yield assembly of information into a single oligo that can be analyzed using DNA sequencing, a ligation step using a generic print oligo can be performed. Thus, in operation 7, incomplete assembly of co-localized recoded blocks and / or memory oligos is corrected by adding a generic print that can replace recoded block sequences not created in operation 5. If recoded block information is missing, amino acid information related to incorrect cycles is lost, but substantial recoded block information is assembled into the memory oligo. In operation 8, the tether of the recoded block is released and an amplification product is generated via polymerase extension. Optionally, the solid support surface can be restored by cleaving the conjugate from the surface.
[0081] Figure 4 schematically shows an exemplary solid support for spatially assisting macromolecular analytes according to an embodiment of the present disclosure. As shown, the solid support is coated with a hydrogel that supports orthogonal chemistries. The orthogonal chemistries shown are aldehyde-hydrazine, azide-alkyne, and thiol. Depending on the selected immobilization scheme, either thiol or click chemistry can be activated for the attachment of trifunctional chemically reactive conjugates. Aldehyde-hydrazine conjugation is an exemplary chemistry that can provide specific and orthogonal immobilization of macromolecular analytes. On the surface, macromolecular analytes are seeded such that the reactants, which are mainly spatially separated and interact with one macromolecule, do not interact with another macromolecule. The volume element is defined by the radius circumscribed by the length of the polymer polymer, as well as the lengths of the linkers of the conjugate complex and the binder.
[0082] Figure 5 schematically shows the interaction between a chemically reactive conjugate and the terminal amino acid of an immobilized peptide during operation 2 of the recoding process 300 of FIG. 3 according to an embodiment of the present disclosure. Generally, the conjugate 1) binds to the terminal amino acid and cleaves the peptide bond between the terminal amino acid and the next amino acid in the polymer (in the case of an N-terminal reaction, this is equivalent to the classical function of an Edman reagent); 2) immobilizes the conjugate on the solid support; 3) Note the 1:1 relationship between the immobilized peptide and the conjugate in a given cycle having three functions, carrying a cycle tag oligo. The conjugate that reacts and binds to the terminal amino acid in operation 2 of the recoding process 300 is shown as a black triangle in FIG. 5. Under basic conditions, the PITC conjugate reacts with the N-terminal amino acid to form a phenylthiocarbamoyl-amino acid (PTC) conjugate. The unreacted conjugate is shown as a white triangle. The unreacted PITC conjugate can be washed from the surface of the solid support before causing a chemical reaction that joins the PTC conjugate to the surface.
[0083] FIG. 6 schematically shows the immobilization of a chemically reactive conjugate onto a solid support during operation 3 of the recoding process 300 of FIG. 3, according to an embodiment of the present disclosure. Generally, conjugate immobilization reactions can be induced by light, addition of a catalyst, or by changing the properties or temperature of a buffer to control the reaction rate. For example, by lowering the redox potential, formation of stable dithiol bonds becomes possible. Stringent washing removes unreacted PITC conjugate, and then the activation of orthogonal chemistry used to join the conjugate to the solid support is initiated to immobilize the PTC conjugate in proximity to the anchor point of the associating peptide. For example, by changing redox conditions to induce dithiol formation or by adding Cu 2+ , stabilizers, and redox components to induce a click reaction, the PTC-thiol conjugate or PTC-alkyne complex can be immobilized onto the solid support. Note that the length of the peptide defines the volume element around the anchor point with the support, and the conjugate associated with a particular peptide co-localizes at that anchor point. Following the immobilization of the conjugate onto the solid support, a conjugate reactivity scavenger may be added to cap the reactivity of residual conjugate that does not react with the N-terminal amino acid, is incompletely washed, and adheres to the solid support. In this way, incomplete removal of the PITC conjugate is remedied by introducing an amino acid mimic. The unreacted PITC conjugate can bind to the surface, but its future reactivity towards amino acids is quenched, leaving spectator conjugate on the surface that cannot participate in downstream processes.
[0084] FIG. 7 schematically shows the cleavage of a terminal amino acid (e.g., cleavage of a peptide bond) in operation 4 of the recoding process 300 of FIG. 3 according to an embodiment of the present disclosure. In an example using Edman's degradation chemistry, this is achieved by a change in pH from basic to harsh acidic conditions, sometimes in an organic solvent. Thus, the hydrogel and conjugation reactions are designed to withstand the peptide bond cleavage conditions. Also, for this reason, the cycle-tag nucleic acids immobilized on the solid support and any other nucleic acids contain protecting groups that prevent degradation of the amine or other reactive moieties of the nucleic acids. Cleavage of the terminal amino acid results in the release of the peptide and provides a new terminal amino acid on the immobilized peptide analyte. In FIG. 7, the immobilized PTC-AA cycle-tag conjugate complex is located near the anchor point of the peptide analyte.
[0085] FIG. 8 schematically shows the results of repeatedly iterating the operations of FIGS. 5-7 according to an embodiment of the present disclosure. More specifically, FIG. 8 shows two to four iterations of operation 2 of the recoding process 300. As shown, a series of co-localized conjugates each having a cycle-tag carrying information regarding the relative position of the amino acids in one immobilized peptide analyte are spatially isolated from the conjugates of other peptide analytes. The details of the preparation of each conjugate are irrelevant to the information carried by the conjugate. Thus, immobilized conjugates having information derived from the carboxy-terminal chemistry and immobilized conjugates having information derived from the amine-terminal chemistry may be combined in downstream steps.
[0086] Figure 9 schematically shows the assembly of recoded blocks, such as the above-described operations 5a to 5b, according to an embodiment of the present disclosure. In this process of the recoding process 300, amino acid identity information is aggregated with cycle information. Homologous binders interact with the immobilized conjugate as shown in the upper panel of Figure 9. The binder is engineered to preferentially recognize a specific conjugate based on the differences in the homologous amino acids of the immobilized conjugate. The binding energy of the binder is a combination of (a) the binding energy between the affinity-binding moiety and the conjugate, and (b) the hybridization energy between the cycle tag oligo of the conjugate and the recode tag oligo of the binder. A binder possessing both homologous AA and homologous cycle information directs the ligation of the AA information to the cycle tag, as shown in the lower panel. In practice, components for the recognition of all AA conjugates of all cycles are present simultaneously, and recoded blocks are created simultaneously. Discrimination can be enhanced under "competitive" conditions in cooperation with slow annealing to find the overall maximum binding energy of the combined affinity-binding moiety and nucleic acid. Note that the recognition of the immobilized PTC conjugate on the solid support avoids the proximity effect from the amino acids that were adjacent on the original peptide.
[0087] Figure 10 schematically shows a preparatory operation of an exemplary process for assembling the recoding block of FIG. 9 according to an embodiment of the present disclosure. The lower panel of FIG. 10 shows a binder including a joining portion and a recoding tag, and a conjugate having a cycle tag. Since binders for recognizing all AA conjugates of all cycles are present simultaneously, there are several possible interactions that may exist between the binder and the immobilized conjugate. They can be classified as (a) correct cognate AA, correct cognate nucleotide; (b) correct cognate AA, incorrect cognate nucleotide; (c) incorrect cognate AA, correct cognate nucleotide; (d) incorrect cognate AA, incorrect cognate nucleotide; and (e) non-specific binding. Stringent washing conditions remove weakly bound binders from the surface. These can be due to cross-reactive binding of the joining portion to non-cognate PTC-AA-cycle tag conjugate complexes and include interactions classified as either (c), (d) or (e). The interaction classified as (a) is productive during the next step of oligoligation where information is transferred from the recoding tag to the cycle tag to form the recoding block. The interaction classified as (b) does not occur during the next step of oligoligation. k off is the characteristic time (1 / k off ) which is the off-rate of the cognate binder may exceed the time to effectively wash the solid support.
[0088] Figure 11 schematically shows the transfer of amino acid identity information from the recoding tag of a binder to the cycle tag of an immobilized conjugate for forming a recoding block via ligation, for example operations 5c-5d of the recoding process 300, according to an embodiment of the present disclosure. As shown, complementary ligation oligos and ligase are added in a suitable buffer to assist ligation and information transfer is received only when cognate amino acid and complementary nucleic acid conditions are met. A binder containing a recoding tag that is cognate to the amino acid but non-complementary to the cycle tag does not receive information transfer. Similarly, a ligation oligo that is not complementary to the recoding tag of the binder does not receive information transfer.
[0089] Figure 12 schematically shows the iterative implementation of operations 5a - 5d of the recoding process 300 for assembling recoding blocks according to an embodiment of the present disclosure. During its implementation, a binder for recognizing all AA conjugates is present simultaneously for all cycles. Thus, the efficiency of correct binding of homologous pairs in any one trial can be low. Slow annealing helps to distinguish interactions with similar binding energies and promotes the binding of homologous pairs. However, this may not improve the efficiency to the desired level. Furthermore, steric hindrance due to the co - localization of immobilized conjugates can limit the access of the binder to one or more conjugates in any given trial. To drive a high percentage of recoding block assembly, multiple trials of binding, washing, and ligation can be used. In trials where non - productive binding events occur, ligation does not occur. Stringent washing to remove non - homologous binders creates a new opportunity to find and anneal homologous agents in the next trial. In principle, assuming no systematic effects, repeating the trials drives the recoding block assembly until completion.
[0090] Figure 13 schematically shows the relative sizes of various components of the recoding process 300 according to an embodiment of the present disclosure. As shown, the relative sizes of the various components emphasize the need to provide linkers / spacers that allow sufficient freedom for the components to interact while maintaining co - localized isolation of each immobilized polymer analyte on the solid support.
[0091] Figure 14 schematically shows the separation of incompatible chemical operations during the recoding process 300 according to an embodiment of the present disclosure. As illustrated, the recoding process 300 is suitable for separating these steps so that it is not necessary to switch the chemistry to complete the cycles and / or reversible chemistry.
[0092] FIG. 15 schematically shows the assembly of memory oligos for subsequent DNA sequencing analysis in operations 6-8 of the recoding process 300 according to an embodiment of the present disclosure. As shown, the overlapping and complementary sequences of the co-localized recoding blocks facilitate their assembly into a single oligo (memory oligo) that serves as a seed for analysis using DNA sequencing techniques. Several molecular biology methods may be useful during the assembly. For example, memory oligos can be assembled using polymerase extension followed by ligation, or simply by using the ligation method. In the case of assembly by ligation, the addition of a single-stranded 5'-phosphorylated DNA oligo complementary to the AA tag of the recoding block facilitates the assembly. Direct ligation to primer sequences immobilized on a solid support, such as the P5 and P7 sequences shown in FIG. 15, using a chimeric primer having sequences complementary to the recoding block and the P5 or P7 sequence, can promote memory oligo amplification. Following memory oligo assembly, the recoding block tether to the solid support may be cleaved if necessary, and polymerase extension from the 3' end of the immobilized P5 or P7 may initiate cluster generation. Alternatively, the assembled memory oligos can undergo end repair, A-tailing, sequencing adapter ligation, and amplification either in situ or in solution.
[0093] FIG. 16 schematically shows the repair of an incomplete recoding block in a memory oligo assembly according to an embodiment of the present disclosure. Thus, FIG. 16 shows operation 7 of the recoding process 300. As shown, gaps in the connectivity between co-localized conjugates can be due to, for example, a) incomplete information accumulation during sequential degradation of the peptide and immobilization of the PTC-AA-cycle tag conjugate, b) incomplete information transfer from the recoding tag to the cycle tag during recoding block assembly, or c) incomplete ligation of available existing recoding block information in the memory oligo assembly. To improve these gaps and enable high-yield assembly of information into a single oligo that can be analyzed using DNA sequencing, a ligation step using a generic print oligo can be performed. The repair can be achieved simultaneously for all cycles by using a pool containing a sprint that can assemble any non-ligated recoding block with any other non-ligated recoding block. Alternatively, the correction may be achieved stepwise using a subset of the described pools. The "..." in FIG. 16 indicates the completion of the series and represents intervening ligation oligos not explicitly shown. C1 indicates the cycle tag sequence (or its complementary sequence), and n indicates the total number of cycles.
[0094] FIG. 17 schematically shows various oligonucleotide components within a sample volume during recoding block assembly. By considering interaction and reaction conditions, accurate and complete assembly during the recoding process 300 is facilitated. Within any given volume element surrounding an anchor point for a protein or peptide, there are immobilized PTC-AA-cycle tag conjugate complexes that do not have complexes with (a) the same AA, different cycle information, and (b) different AAs, different cycle information, but (c) different AAs, the same cycle information, or (d) the same AA, the same cycle information. In FIG. 17, there are "group 1" components to assist in the assembly of cycle 1 information. "Group 2" components are present to assist in the assembly of cycle 2 information, etc., from group 3 to group "n". Interactions within and between groups are cataloged at the top of each column. The total number of oligo components is shown for each type of component. Weak interactions are possible due to hybridization of shortmer oligos. These are overridden by relatively strong interactions where the binding moiety directs the oligo for assembly. The heavy line indicates the desired interaction assumed for a given recoding block AA1-cycle 1 and represents the total binding energy of the interaction. The light lines indicate exemplary possible oligo interactions. The Tm of these interactions is low, and thus, incorrect ligation leading to incorrect recoding block information is controlled. The recoding block is shown with various tether sites to, for example, the 5', 3', and internal nucleosides. The "..." in FIG. 17 indicates the completion of a series and represents intervening cycle tags, ligation oligos, or recoding tags not explicitly shown. C1 indicates the cycle tag sequence (or its complementary sequence), AA indicates the amino acid recoding sequence, and "n" indicates the total number of cycles.
[0095] FIG. 18 schematically shows various oligonucleotide components within a sample volume during memory oligo assembly according to an embodiment of the present disclosure. In FIG. 18, the effective concentration of the components is high for co-localization within a volume element defined by the length of the polymeric analyte and the length of the linker of the associated recoding block. However, the complexity of the oligos is not high. Also, since the cycle code (C1, C2, ... Cn) and the amino acid code (AA1, AA2, ... AAn) are defined using a schema based on communication theory, mismatch ligation is unlikely to occur. Note that since the cycle information is adjacent to the amino acid information, even an "erroneous assembly" resulting from mismatch ligation generates an oligo with useful polymeric analyte sequence information. The continuous information blocks within the memory oligo are redundant when determining the sequence of the peptide analyte. The "..." indicates the completion of the series and represents intervening recoding blocks or AA' complements not explicitly shown. n indicates the total number of cycles.
[0096] FIG. 19 schematically shows the release of memory oligos and conjugate complexes from a solid support in operation 8 of the recoding process 300 according to an embodiment of the present disclosure. An exemplary memory oligo having p7 and P5 adapters is shown in FIG. 19. The memory oligo may also include a sample index, UMI, CRISPR PAM or spacer sequence, or other identifying nucleic acid sequences that can be incorporated during NGS library preparation. Cleaving the tether (or a subset of the tethers) from the solid support is an optional step to improve the efficiency of PCR extension containing the memory oligo. Removal of the conjugate from the surface is an optional process for purifying the solid support before use in downstream processes such as cluster generation and NGS sequencing. FIG. 19 shows the reduction of disulfide bonds that can be mediated by the addition of dithiothreitol to the solution contacting the support surface.
[0097] In certain embodiments, the recoding block includes sequences, such as CRISPR PAM or spacer sequences, that facilitate the assembly of memory oligos and / or facilitate target enrichment, target depletion, and / or sequencing sample preparation (e.g., NGS sample preparation). For example, approximately 90% of the protein content in human plasma is albumin. It would be advantageous to deplete albumin in plasma to improve the sensitivity for detecting low abundance proteins that interact with albumin therein. Thus, depletion by DNA methods of enrichment or depletion after recoding can provide a less biased sample preparation than depletion or enrichment of protein samples by conventional recognition-based methods of protein enrichment or depletion. Thus, oligo design for cycle tags, recoding tags, recoding blocks and / or memory oligos can preferentially deplete recoded albumin peptide sequences via cleavage of memory oligo amplicons by CRISPR nucleases or other enzymes, for albumin, such as NGG, C1-AAtag Met -C2-AAtag Lys and may include CRISPR PAM and spacer sequences (or others) specific to.
[0098] Figures 20A - 20B show the fluorescence values obtained through the execution of steps 1 - 4 of Figure 3. The relative fluorescence units (RFU) of the fluorescent oligonucleotides complementary to the cycle tags indicate the progress of the steps of the present method. In Figure 20A, each bar shows the measured fluorescence value in the progress step. Bar 1 shows the minimum autofluorescence of the peptides and solid support used in the study. Bar 3 shows the capture of the fluorescent oligonucleotide by the CRC immobilized on the solid support through the reaction of their reactive part (PITC) with the N - terminal amino acid of the immobilized peptide. The low signal of Bar 2 supports that the signal does not relate to the unbound fluorescent oligonucleotide in solution and coincides with the signal arising from the fluorescent oligo captured by the CRC reacted with the immobilized peptide on the solid support. Bar 4 shows the signal from the fluorescent oligo released from the surface when the surface is exposed to mild chemical conditions that promote the de - hybridization of the oligonucleotide. The relative values of Bars 3 and 4 can be explained by the difference in volume during measurement. Bar 5 corroborates the de - hybridization of the fluorescent oligonucleotide from the surface. During the measurement of the bars, 5 and 6 CRCs of sample B were immobilized on the surface by the Cu - catalyzed Huisgen cycloaddition reaction. Also, during the measurement of Bars 5 and 6, the surface was subjected to anhydrous acid under conditions that assist in the cleavage of the N - terminal amino acid, exposing the next amino acid residue as the N - terminal amino acid residue on the cleaved peptide. Bar 6 shows the progress of contacting the surface with a second CRC having a different cycle - tag sequence through the reactive part (PITC). The CRC is reactive towards the newly exposed N - terminal amino acid of the immobilized peptide after the cleavage of the first N - terminal amino acid by acid. Bar 8 demonstrates the capture of the new fluorescent oligo by the CRC immobilized on the solid support through the reaction of its reactive part (PITC) with the new terminal amino acid of the immobilized peptide. The low signal of Bar 7 supports that the signal does not relate to the unbound fluorescent oligonucleotide in solution and coincides with the signal arising from the fluorescent oligo captured by the CRC reacted with the immobilized peptide on the solid support. Bar 9 shows the signal from the fluorescent oligo released from the surface when the surface is exposed to mild chemical conditions that promote the de - hybridization of the oligonucleotide.The relative values of bars 8 and 9 can be explained by the difference in volume during measurement. Bar 10 corroborates the dehybridization of the fluorescent oligonucleotide from the surface. The progression of the fluorescent signal is used to confirm the reaction, capture, and cleavage of the N-terminal amino acid residue of the peptide using the substances and methods disclosed herein. The strong signals of the bars in steps 3, 4, 8, and 9 confirm the function of the CRC for performing steps 2-4 of FIG. 3. In FIG. 20B, each bar shows the fluorescence in the progress of the method. The conditions and conclusions are the same as those for bar B in FIG. 20A, except that the starting azide-functionalized silica surface was supplied by a commercial source.
[0099] Figure 21 schematically shows how the efficiency of memory oligo assembly can be adjusted according to the method described herein. As shown, the large spheres in Figure 21 represent the volume defined by the length of an analyte polymer, such as an amino acid polymer. Inside the large spheres are many small spheres. Each of these smaller spheres may represent the volume defined by the binders and conjugates utilized during the recoding process, more specifically, the binders and conjugates utilized during operation 5 described above. Such volume mainly depends on the linker lengths of both the binder and the conjugate. Thus, to facilitate the association of recoding blocks in memory oligo assembly, the polymer (represented by the larger sphere) can be degraded via a known polymer degradation mechanism such as that described in Leonid Lonov, Hydrogel-based actuators: possibilities and limitations, Materials Today, 17, 10, 494 (2014), which is hereby incorporated by reference in its entirety. Alternatively, the binder and conjugate (represented by the smaller sphere) can expand, for example, by utilizing in-silico swellable spacers, linking oligos, and / or deconvolution of rare events, as described elsewhere herein, thereby facilitating communication between adjacent recoding blocks. Note that the recoding blocks can be linked in any order to create a memory oligo. The swellable spacer can include a molecule containing multiple thiol groups. When a disulfide bond is formed, the range of the spacer becomes shorter, and when the crosslinker decreases, for example, by the addition of DTT, the range of the spacer becomes wider.
[0100] FIG. 22 shows the use of a universal array to facilitate the concatenation of recoding blocks in memory oligo assembly, regardless of a particular order. As described above, the recoding blocks may be concatenated in any order to create a memory oligo. This is because cycle information is directly adjacent to amino acid information in the assembled recoding blocks, whether the recoding blocks are contiguous or non - contiguous within the memory oligo. Assembly of the recoding blocks in the correct order of the analyte can be efficient, but the adjacent nature of the cycles and the amino acid information within the recoding blocks can cause redundancy. Thus, to avoid these redundancies while relaxing the criteria for memory oligo assembly, the recoding blocks can be assembled in a random order.
[0101] To facilitate the assembly of memory oligos regardless of a particular order, a universal assembly array may be utilized during the recoding process. Such a universal array can be attached to the 5' and / or 3' ends of the cycle tags and / or recoding tags before introducing these tags to the immobilized analyte. By attaching complementary universal arrays to two or more cycle tags and / or recoding tags, random concatenation (e.g., ligation) of the recoding blocks obtained in memory oligo assembly is facilitated, regardless of the sequential order, and the correct macromolecular analyte sequence can be assigned during post - sequencing analysis.
[0102] Figure 23 schematically shows the transfer of information from a positional oligo to a recoding block. In certain embodiments, during the recoding process, the peptide is attached to a solid support via a positional linker that may include any molecule configured to attach the peptide to the solid support and further configured to bind to a nucleic acid. The nucleic acid may include any suitable type of nucleic acid sequence having code information regarding the position of immobilized PTC conjugate isolated on the solid support. The nucleic acid may be directly conjugated to a hydrogel. This nucleic acid may be referred to as a "positional oligo". The positional oligo may be attached to the site-specific linker before or after binding of the peptide to the solid support and / or before or after immobilization of the peptide to the solid support. During the recoding process, after transferring the information of the recoding tag to the immobilized conjugate complex to generate a recoding block, a PCR-like thermal cycling process may be performed to sequentially transfer the positional oligo information to a plurality of proximal recoding blocks via polymerase extension. Use of unnatural nucleotides (shown as circles in the figure) in the synthetic nucleic acid and polymerase extension using only A, G, C, T, and iC in the reaction solution can control unwanted polymerase extension.
[0103] In short, the above recoding process avoids important issues related to 1) incompatible chemistry / protecting chemistry, 2) reversible chemistry, and 3) binder molecule specificity. Regarding incompatible chemistry and protecting chemistry: The harsh chemical conditions associated with peptide bond cleavage are carried out in a single block of a process that can maintain the integrity of nucleic acids using protecting groups. Regarding reversible chemistry: Since the information is aggregated into blocks, switching between chemical properties, blocking and deblocking of labile chemical moieties, and other complexities is avoided. The operations can be carried out in parallel instead of accumulating information sequentially, so reversible chemistry is not required. This greatly expands the potential chemical universe that can be deployed within the workflow. Regarding the specificity of the binder molecule: The specificity of the binder molecule for a single amino acid is achieved by isolating the recognition event for each individual amino acid from the influence of adjacent amino acids in the peptide by recognizing the amino acid within the isolated context of the immobilized PTC conjugate. Amino acid identity is recoded separately from the cycle information (position within the polypeptide chain), providing flexibility and simplicity to the workflow / process and reducing the complexity of the amino acid recognition event. The DNA library recoded from the peptide sequence can be amplified directly on a solid support, or the nucleic acid library can be released from the solid support and amplified using standard NGS library preparation reagent kits, or amplified via standard molecular biology techniques. Analysis using any high-throughput NGS method yields millions of reads per run, translating into millions of peptides sequenced in a single run.
[0104] FIG. 24 schematically shows an example of an alternative event during the implementation of the recoding method described herein according to an embodiment of the present disclosure. More specifically, FIG. 24 shows an inaccurate association of cycle tag information to monomers of an analyte caused by two conjugates immobilized in proximity to each other on a solid support. Referring back to operation 5a of the recoding process, when a binder is introduced to the immobilized conjugate-AA-cycle tag complex, the binder having both homologous AA and homologous cycle information should recognize and bind to their target conjugate. Thereafter, the binder should direct the ligation of its AA information, in the form of a recode tag, to the cycle tag of the conjugate complex. However, the proximity of the immobilized conjugate complex on the solid support may rarely bind a binder having homologous AA but no homologous cycle information to the conjugate complex, which results in an alternative recode tag cycle tag ligation, and thus an inaccurate association of cycle tag information to monomers of the analyte.
[0105] In Figure 24, the binder "AA1:C12" is shown as binding to the conjugate-AA-cycle tag complex "C1:AA1". In this example, the binder AA1:C12 correctly recognizes the cognate amino acid of CA:AA1 (e.g., AA1 is recognized). However, since the cycle tag of the complex C1:AA1 is non-cognate, the binder should not bind to C1:AA1. Nevertheless, the binder AA1:C12 recognizes the cycle tag C12 of the neighboring conjugate-AA-cycle tag complex "C12:AA3" and promotes an "alternative" binding of the binder AA1:C12 to the complex C1:AA1. Thus, this binding is promoted by the binding activity of the binding moiety of the binder AA1:C12 to AA1 in addition to the hybridization energy of the C12:AA3 complex in proximity to the CA:AA1 complex. As a result of this binding, the amino acid 3 (AA3) is "alternatively" assigned to cycle 12 (C12), but the correct assignment in this example would be AA1 to C1. Note that in this example, the binder and the neighboring conjugate complex must retain the same cycle information to enable alternative events.
[0106] The alternative assignment in Figure 24 cannot be repaired by sequentially introducing binders to the immobilized conjugate-AA-cycle tag complexes in the order of AA 1-n :C1, AA 1-n :C2, AA 1-n :C3, etc. Similarly, for all amino acid binding moieties, the assignment cannot be improved by continuously introducing binders to the immobilized conjugate-AA-cycle tag complexes in the order of AA1:C1, AA2:C2, AA3:C3, etc. Rather, spatial separation of the conjugate-AA-cycle tag complexes can be promoted to reduce or eliminate the type of interaction in Figure 24.
[0107] Conjugate AA cycle tag complexes and short spacers for binders can be used during the recoding block assembly process (e.g., operations 5a - 5d above) to effectively avoid these alternative events. However, since such assembly is facilitated by the interaction of recoding blocks, such spacers may have an adverse effect on the assembly of memory oligos. To overcome these opposing spatial constraints, spacer molecules that can be controllably lengthened or extended can be used. For example, cysteine may be incorporated at both ends of the spacer molecule via a disulfide bridge, thereby facilitating a shortened linker during recoding block assembly (e.g., the operations shown in FIG. 24). This spacer can be extended by reducing the disulfide bond during memory oligo assembly. Alternatively, a controllably inflatable polymer can be utilized. For example, a two - component hydrogel configured to disintegrate or expand based on solution / solvent conditions, or a polymer with reactivity enabling expansion can be incorporated into the solid support. During recoding block assembly, the polymer can be relaxed, thereby increasing the distance between the immobilized conjugate - AA - cycle tag complexes, but during other steps of the recoding process, the polymer can be disintegrated.
[0108] In yet another embodiment, to alleviate the inability of recoding blocks to access each other, connecting oligos with cross - linking ability can be utilized (see, e.g., FIG. 22). For example, as shown in FIG. 22, a long connecting oligo can be used to bridge a gap through the recoding blocks via an extension: ligation approach. In other embodiments, events such as those shown in FIG. 24 are rare and depend on the proximity of the conjugate AA - cycle tag complexes, so prior information and probabilities can be used to improve the accuracy of identification. AI and other computer - based methods can be useful for recognizing these events such that they can be corrected in silico.
[0109] Figures 25A - 25C include an exemplary method that may include separation, assignment, and assembly. Figure 25A shows an example of isolation: The N - terminal amino acid can be successively removed from the peptide in a series of cycles using a trifunctional molecule, each of which results in the immobilization of one amino acid complex adjacent to the anchor point of its cognate peptide. Multiple cycles can create a spatially localized lawn of complexes that hold circular DNA, as shown in the top - most geometry panel, where the large spheres represent the localization of the protein and the small spheres represent the localization of its isolated amino acids. The cycles may be known, but the amino acid identities have not yet been determined. Figure 25B shows an example of assignment: After removing the protecting group from the isolated complex and transitioning from an anhydrous environment to an aqueous environment, amino acid identities can be added to the isolated complex through recognition by an affinity construct that brings the circular DNA into proximity with the circular DNA form of the identity information. The "identity" and "circular" DNAs can be ligated in a high - fidelity reaction. Figure 25C shows an example of assembly: As shown in the bottom - most geometry panel, extension - ligation of the region DNA into a long construct that reflects the original peptide information, which can be analyzed using NGS sequencing.
[0110] Figure 26 shows a trifunctional molecule of approximately 1 kd having the following: (1) a base structure: phenylisothiocyanate at the oligo position to simplify the analytical characterization of NNN - (propargyl - PEG2)(6 - oxo - 6 - (dibenzo[b,f]azacyclooct - 4 - yl) - caproic acid)(PEG3 - 1 - acetamido - 4 - iso - thiocyanato - benzene), (2) propargyl, and (3) model vanillin. The molecular structure was confirmed using LC - ESI - MS and its function was tested. HPLC analysis showed the formation of a high - yield product and the functional activity of the important reactive isothiocyanate moiety.
[0111] Figure 27 shows an agarose electrophoresis gel demonstrating the effective in situ ligation of cycle tags and recode tag oligos. Lane 1 has a dsDNA ladder (cat#10597012 from Invitrogen) with the brightest band appearing at 100 base pairs. Lane 2 is the product from the ligation of a 30mer oligo by a tether arm having 45mer ligation oligos at both the 3’ (Sys#001 LO2, 30, SEQ ID NO:85) and 5’ (Sys#001, LO1, 30, SEQ ID NO:84) termini. Three bands are seen: the product where both 30mer oligos are ligated, a faint band indicating that one or the other 30mer oligo was ligated, and a faint band indicating the unligated 45mer oligo. Lanes 3 and 4 show the ligation products when only one or the other of the 30mer ligation oligos was added to the reaction, and thus shorter products are produced. Lane 5 shows the ligation mixture without added ligase, with the primary band being the 45mer oligo band. Lane 6 shows the ligation product with the “tetherless” version of the 45mer oligo, and the three bands are similar to those in Lane 2, indicating the presence of the double ligation product, single ligation product, and unligated 45mer product.
[0112] Figure 28 shows a block diagram of the steps of a periodic protection and deprotection workflow.
[0113] Figure 29 shows a reaction scheme for the stepwise assembly of an immobilized CRC complex by reacting an N-terminal amino acid with an amine-reactive molecule having a second reactive functional group (e.g., tetrazine). A trifunctional construct having a nucleic acid cycle tag and a surface immobilization moiety and a trans-cyclooctene can be reacted with the tetrazine of the amine-reactive molecule to form the immobilized CRC complex.
[0114] Figure 30 shows the reaction scheme for the stepwise assembly of an immobilized CRC complex by reacting the N-terminal amino acid with an amine-reactive molecule having a second reactive functional group (e.g., trans-cyclooctene). A trifunctional construct having a nucleic acid cycle tag and a surface immobilization moiety can be reacted with the tetrazine functional group of the amine-reactive molecule to form an immobilized CRC complex.
[0115] Figure 31 shows the reaction scheme for the stepwise assembly of an immobilized CRC complex, where a tetrazine-labeled oligo is reacted with a functional group (e.g., trans-cyclooctene) of a trifunctional construct having reactive moieties for binding and cleaving the N-terminal amino acid residue of a peptide and a surface immobilization moiety.
[0116] Figures 32A - 32B show the CRC synthesis process and intermediate molecules. Figure 32A is a block diagram showing the process of synthesizing PPO starting from PDA. This can be converted to PDON-tBOC, deprotected to form PDON, then converted to PDO, and subsequently to PPO. Figure 32B includes the chemical structures of PPO and the intermediates.
[0117] Figure 33 shows the function of PPO. It shows the relative fluorescence units (RFU) of PPO immobilized on an azide-modified surface by Cu-catalyzed Huisgen cycloaddition followed by reaction with amine-labeled fluorescein. Multiple fractions of purified PPO function similarly. A strong signal above background confirms the function of both the alkyne and ITC chemical reactivity elements of CRC.
[0118] Figure 34 shows the function of PPO. It shows the relative fluorescence units of a fluorescent oligo complementary to the oligo on PPO immobilized on an azide-modified surface via Cu-catalyzed Huisgen cycloaddition. Multiple fractions of purified PPO function similarly. A strong signal above background confirms the function of both the alkyne and oligo elements of CRC.
[0119] Figure 35 shows the function of PPO. It shows the relative fluorescence units (RFU) of PPO immobilized on an amine-modified surface via a reactive ITC moiety and subsequently subjected to Cu-catalyzed Huisgen cycloaddition to an azide-labeled fluorescein reagent. Multiple fractions of purified PPO function similarly. A strong signal above background confirms the function of the alkyne, the ITC chemical reactivity element of CRC, and the ability to use CRC in multiple embodiments of the solid support.
[0120] Figures 36A - 36D show exemplary simulations and the binding kinetics of a commercial antibody (Sigma, SAB5200015) to an immobilized phosphotyrosine - PTH - ligand. Figure 36D shows representative data of a strongly reproducible binding curve generated using the Nicoya SPR system.
[0121] Figure 37 shows the PCR data of the ligated recoded block. Amplification of the ligated oligos with and without the tether shows the amplification of the ligated recoded block with the tether, and thus demonstrates the ability to generate amplicons from the tethered recoded block for subsequent acquisition of sequence information of the memory oligonucleotide or recoded block.
[0122] In certain embodiments, a method for analyzing one or more peptides from a sample comprising a plurality of peptides, proteins, and / or protein complexes, the method comprising: (a) providing a peptide of mer length n = 2 to 2000 bound to a solid support; (b) providing a first chemically reactive conjugate, such as a PITC conjugate, the first chemically reactive conjugate comprising a cycle tag (e.g., "cycleTag") having identification information regarding the workflow cycle of the method, a reactive moiety that can bind to and cleave the terminal amino acid of the peptide, and a reactive moiety that facilitates immobilization to the solid support; (c) contacting the peptide with the first chemically reactive conjugate, the first chemically reactive conjugate binding to the terminal amino acid or modified terminal portion of the peptide to form a conjugate complex, such as a PTC-AA-cycleTag conjugate complex; (d) immobilizing the conjugate complex to the solid support; (e) cleaving the terminal amino acid from the peptide, thereby providing the immobilized conjugate complex and a new terminal amino acid of the peptide bound to the solid support of (a); (f) contacting the immobilized conjugate complex with a first binder capable of binding to the immobilized conjugate complex, the first binder comprising a binding moiety and a first recode tag (e.g., "recodeTag") having identification information regarding the first binder; (g) transmitting the information of the first recode tag associated with the first binder to the cycle tag of the immobilized complex to generate a first recode block (e.g., "recodeBlock"); (h) optionally repeating steps (b)-(g) to assemble a second recode block having recoding information for the new terminal amino acid of the peptide; (i) optionally repeating step (h) for additional iteration cycles to create additional recode blocks for additional amino acids of the immobilized peptide of (a); (j) optionally deprotecting the nucleic acids of the first recode block, the second recode block, and the additional recode blocks;(k) Under conditions that allow extension-ligation to assemble the recoding block into a memory oligonucleotide (e.g., "memoryOligo"), contacting the recoding block with polymerase, nucleotides, ligase, and buffer, and (l) analyzing the memory oligonucleotide, a method is disclosed herein.
[0123] In some embodiments, one or more operations of the method are repeated one or more times to increase the process yield of the method. For example, in certain embodiments, operations (e), (f), and / or (g) are repeated one or more times to increase the process yield.
[0124] In some embodiments, the method further comprises contacting the immobilized conjugate complex with an indiscriminate binder that can bind to the immobilized conjugate complex regardless of the identity of the amino acid (AA) within the conjugate complex, between operations (h) and (j) and / or after operation (k). The indiscriminate binder comprises a binding moiety that associates with the immobilized conjugate regardless of the AA. The indiscriminate binder is capable of hybridizing with specific cycle information, or any cycle tag (or subset of cycle tags), and may carry an indiscriminate recoding tag (e.g., inosine base) that conveys identification information regarding the indiscriminate binder. This provides robustness to the binding recognition operation and may be repeated one or more times to increase the process yield. In such embodiments, operation (k) may be repeated after contacting the immobilized conjugate complex with the indiscriminate binder.
[0125] In some embodiments, the peptide comprises any suitable polymeric polymer, including proteins, peptides, complex carbohydrates, etc. In such embodiments, the monomer units of the polymeric polymer can include amino acids, carbohydrates, and / or any monomer moiety that can be incorporated into the polymer.
[0126] In some embodiments, the conjugate complex includes zero, one or more reactive moieties (e.g., moieties used to conjugate the complex to a solid support), and the reaction includes activatable chemistry. In some embodiments, the conjugate complex includes zero, one or more reactive moieties (e.g., moieties used to conjugate the complex to a solid support), and the reaction includes activatable chemistry. In some embodiments, the conjugate complex includes zero, one or more reactive moieties (e.g., moieties used to conjugate the complex to a solid support), and the reaction includes reversible chemistry and activatable chemistry.
[0127] In some embodiments, the recoding tag linked to the binder is a nucleic acid having a sequence corresponding to the (n - 1)th cycle tag or the (n + / - i)th cycle tag, an amino acid (AA) tag (e.g., "AAtag"), and the nth cycle tag. Optionally, the recoding tag linked to the binder is a nucleic acid having a universal sequence for amplification or assembly, a sequence complementary to a cycle tag (e.g., "cycle tag complementary sequence"), and an amino acid (AA) tag (e.g., "AAtag").
[0128] In some embodiments, operation (k) comprises contacting the recoding block with a ligase, an AA tag oligonucleotide complement, and a buffer under conditions that allow assembly of the recoding block and the AA tag oligonucleotide complement onto the memory oligo by ligation or creation of a fragment of the memory oligo.
[0129] In certain embodiments, a method for analyzing one or more peptides from a sample comprising a plurality of peptides, proteins, and / or protein complexes, the method comprising: (a) providing a peptide of mer length n = 2 to 2000 attached to a solid support; (b) providing a first chemically reactive conjugate (e.g., a PITC conjugate), the conjugate comprising a cycle tag having identification information regarding the workflow cycle of the method, a reactive moiety capable of binding and cleaving to the terminal amino acid of the peptide, and a reactive moiety facilitating immobilization to the solid support; (c) contacting the peptide with the first chemically reactive conjugate, the first chemically reactive conjugate binding to the terminal amino acid or modified terminal portion of the peptide to form a first conjugate complex, e.g., a PIT-AA-cycle tag conjugate complex; (d) immobilizing the first conjugate complex to the solid support; (e) cleaving the terminal amino acid from the peptide, thereby providing a first immobilized conjugate complex and a new terminal amino acid of the peptide conjugated to the solid support of (a); (f) optionally repeating (b) - (e) to assemble a second immobilized conjugate complex having cycle information for the new terminal amino acid of the peptide; (g) optionally repeating (f) for additional iteration cycles to create additional immobilized conjugate complexes for additional amino acids of the peptide of step (a); (h) optionally deprotecting the nucleic acid of the conjugate complex and / or any protected nucleic acid associated with the solid support; (i) contacting the first immobilized conjugate complex with a first binder capable of binding to the first immobilized conjugate complex, the first binder comprising a binding moiety and a recoding tag having identification information regarding the first binder; (j) transmitting the information of the recoding tag associated with the first binder to the cycle tag of the first immobilized complex to generate a first recoding block; (k) optionally repeating (i) and (j) using a second binder comprising a binding moiety and a recoding tag having identification information regarding the second binder,When the information of the recoding tag associated with the second binder is transmitted to the second immobilized conjugate complex to generate a second recoding block, (l) optionally, repeating (k) for additional cycles to create a recoding block for additional amino acids of the peptide of step (a); and (m) contacting the recoding block with a polymerase, nucleotides, ligase, and buffer under conditions that allow extension-ligation to assemble the recoding block into a memory oligo or create fragments of the memory oligo; and (n) analyzing the memory oligo. A method is disclosed herein.
[0130] In some embodiments, one or more operations of the method are repeated one or more times to increase the process yield of the method. For example, in certain embodiments, operation (e), (i), and / or (j) are repeated one or more times to increase the process yield.
[0131] In some embodiments, the method further comprises, after operation (m), contacting the first immobilized conjugate complex with a promiscuous binder that can bind to the first immobilized conjugate complex regardless of the identity of the amino acids within the conjugate complex, the promiscuous binder comprising a binding moiety that associates with the immobilized conjugate regardless of the amino acid. The promiscuous binder can hybridize with specific cycle information, or any cycle tag (or subset of cycle tags), and can carry a promiscuous recoding tag (e.g., inosine base) that conveys identification information regarding the promiscuous binder. This provides robustness to the binding recognition operation and may be repeated one or more times to increase the process yield. In such embodiments, operation (m) may be repeated after contacting the immobilized conjugate complex with the promiscuous binder.
[0132] In some embodiments, the assembly (e.g., ligation) of the recoded blocks is facilitated by the use of a permissive polymerase such as polymerase theta (PolΘ), or by the use of proteins involved in a blunt-end DNA ligation process similar to non-homologous end joining (NHEJ). See, for example, Poplawski T et al., Postepy Biochem 2009;55(1):36-45; Davis AJ, Chen DJ, Transl Cancer Res.2013 June;2(3):130-143.
[0133] In some embodiments, the peptide comprises any suitable polymeric polymer including proteins, peptides, polypeptides, and the like. In such embodiments, the monomer units of the polymeric polymer can include amino acids, carbohydrates, and / or any monomer moiety that can be incorporated into the polymer.
[0134] In certain embodiments, a method for analyzing one or more peptides from a sample comprising a plurality of peptides, proteins, and / or protein complexes, the method comprising: (a) providing a peptide of mer length n = 2-2000 conjugated to a solid support using a positional linker, the positional linker being bound to a location oligo (e.g., "locationOligo"); (b) providing a first chemically reactive conjugate (e.g., a PITC conjugate), the conjugate comprising a cycle tag having identification information regarding the workflow cycle of the method, a reactive moiety that binds to and can cleave the terminal amino acid of the peptide, and a reactive moiety that facilitates immobilization to the solid support; (c) contacting the peptide with the first chemically reactive conjugate, the first chemically reactive conjugate binding to the terminal amino acid or modified terminal portion of the peptide to form a first conjugate complex, e.g., a PIT-AA-cycle tag conjugate complex; (d) immobilizing the first conjugate complex to the solid support; (e) cleaving the terminal amino acid from the peptide, thereby providing a first immobilized conjugate complex and a new terminal amino acid of the peptide conjugated to the solid support of (a); (f) optionally repeating (b)-(e) to assemble a second immobilized conjugate complex having cycle information for the new terminal amino acid of the peptide; (g) optionally repeating (f) for additional iterative cycles to create additional immobilized conjugate complexes for additional amino acids of the peptide of step (a); (h) optionally deprotecting the nucleic acid of the conjugate complex and / or any protected nucleic acid associated with the solid support; (i) contacting the first immobilized conjugate complex with a first binder capable of binding to the first immobilized conjugate complex, the first binder comprising a binding moiety and a recode tag having identification information regarding the first binder; (j) transmitting the information of the recode tag associated with the first binder to the cycle tag of the first immobilized complex to generate a first recode block; (k) optionally,Using a second binder comprising a binding moiety and a recoding tag having identification information regarding the second binder, steps (i) and (j) are repeated to transfer information of the recoding tag associated with the second binder to a second immobilized conjugate complex to generate a second recoding block. Then, (l) optionally, step (k) is repeated for additional cycles to create a recoding block for additional amino acids of the peptide of step (a); (m) at least the first recoding block and the corresponding position oligo are contacted with a polymerase, nucleotides, and a buffer under conditions that allow the extension to transfer information from the position oligo to the first recoding block, thereby creating a memory oligo; (n) optionally, step (m) is repeated to transfer information from the position oligo to additional proximal recoding blocks of the position oligo; (o) the memory oligo is released from the solid support via tether cleavage, hydrogel dissociation, polymerization, or another means; (p) optionally, the memory oligo is assembled (in situ) into a longer memory oligo; and (q) the memory oligo is analyzed. A method is disclosed herein that includes these steps.
[0135] In some embodiments, one or more operations of the method are repeated one or more times to increase the process yield of the method. For example, in certain embodiments, operation (e), (i), and / or (j) are repeated one or more times to increase the process yield.
[0136] In some embodiments, the peptide comprises any suitable polymeric polymer, including proteins, peptides, complex carbohydrates, etc. In such embodiments, the monomer units of the polymeric polymer can include amino acids, carbohydrates, and / or any monomer moiety that can be incorporated into the polymer.
[0137] In some embodiments, the recoding tag linked to the binder is a nucleic acid having a sequence corresponding to the (n-1)th cycle tag or the (n+ / -i)th cycle tag, an amino acid (AA) tag (e.g., "AAtag"), and the nth cycle tag. Optionally, the recoding tag linked to the binder is a nucleic acid having a universal sequence for amplification or assembly, a sequence complementary to a cycle tag (e.g., "cycle tag complementary sequence"), and an amino acid (AA) tag (e.g., "AAtag").
[0138] In some embodiments, in some embodiments, information is transferred from the position oligo to the recoding block using a ligase.
[0139] In some embodiments, each individual memory oligo is analyzed by itself or randomly assembled with other memory oligos from the same or different analytes of the sample. This approach can facilitate the rationalization of the recoding process and enable more efficient analysis.
[0140] In some embodiments, position oligos can be used to determine the spatial position within a histological tissue section and, in combination with specific data in silico, enable the spatial resolution of individual protein molecules. Determining the spatial position of protein molecules within a histological tissue section enables spatial multi-omic analysis. Spatial multi-omics is the study of gene / RNA expression and protein abundance in a spatial context to elucidate functional biology. Integrating different scale analyses from spatial multi-omics can facilitate the improvement of the understanding of the tissue and cellular microenvironment.
[0141] In some embodiments, the conjugate complex contains zero, one or more reactive moieties (e.g., used to conjugate the complex to a solid support), and the reaction contains an activatable chemistry.
[0142] In some embodiments, the conjugate complex comprises zero, one or more reactive moieties (e.g., those used to conjugate the complex to a solid support), and the reaction comprises reversible chemistry.
[0143] In some embodiments, the conjugate complex comprises zero, one or more reactive moieties (e.g., those used to conjugate the complex to a solid support), and the reaction comprises activatable and reversible chemistry.
[0144] In some embodiments, regardless of which amino acid (or monomer) of the cycle is identified, one or more amino acids (or monomer subunits) are removed from the immobilized peptide (or macromolecular analyte). These “skipped” amino acid cycles are recorded in silico, and the analysis algorithm accounts for the known translation of the skipped information during alignment with the reference sequence. In the case of peptides, this can be achieved by performing one or more iterations of Operations 2-4 described below, where PITC is replaced with a chemically reactive conjugate (e.g., a PITC conjugate), as needed. This can be referred to as “strobe” leading or “strobe” sequencing. One advantage of this embodiment is that protein isoforms can be readily determined by the leading segments of the protein that are not adjacent to each other to achieve long-range information. This can save time and cost in obtaining intervening or redundant information contained in the peptide, or in combination with genomic information related to the peptide. For example, this embodiment can include five cycles of peptide cleavage using a chemically reactive conjugate, followed by 30 cycles using PITC or enzymatic cleavage, then another five cycles using a chemically reactive conjugate, and so on.
[0145] In some embodiments, the use of a predetermined subset of binders enables the identification of a subset of amino acids of a peptide, polypeptide, protein, or protein complex. Given that the site of interest (e.g., post-translational modification (PTM) or splice site) can vary across different proteins in a mixed population, this embodiment eliminates the need to measure / determine the identity for all single amino acids in a sample for each cycle, a task that would otherwise require significantly more sequencing.
[0146] In some embodiments, the subset of amino acids identified by a subset of binders is modified with a post-translational modification. By doing so, the information density of the subset of amino acids can be significantly increased during analysis.
[0147] In some embodiments, an aminopeptidase (e.g., CAS number: 37288-67-8) or a similar agent / construct is used to remove one or more amino acids (or monomer subunits) from the immobilized peptide (or macromolecular analyte), regardless of the amino acid (or monomer) identified in that cycle. Additionally, this technique can be applied to prepare the N-terminus of a protein or peptide protected by acylation for treatment with a chemically reactive conjugate. Further, this method can be used in some instances to "strobe" through amino acids such as proline that may not be effectively cleaved under chemical conditions using a chemically reactive conjugate.
[0148] In some embodiments, one or more operations of the method are performed simultaneously. For example, in certain embodiments, operations (i)-(l) are performed simultaneously.
[0149] In some embodiments, operation (m) comprises contacting a recoding block with a ligase, an AA tag oligonucleotide complement, and a buffer under conditions that enable the ligation to assemble the recoding block and the AA tag oligonucleotide complement onto the memory oligonucleotide.
[0150] In some embodiments, the memory oligo, cycle tag, recode block, AA tag complement, and / or ligation oligo or component can include DNA molecules, RNA molecules, other types of nucleic acid molecules, DNA molecules having pseudo-complementary bases (e.g., inosine), or combinations or chimeras thereof.
[0151] In some embodiments, the memory oligo or ligation component includes a universal priming site, which can include a priming site for amplification, a priming site for sequencing, or both.
[0152] In some embodiments, the memory oligo includes a sample index, a spacer, a unique molecular identifier (UMI), a universal priming site, a CRISPR protospacer adjacent motif (PAM) sequence, or any combination thereof.
[0153] In some embodiments, the memory oligo and / or the chemically reactive conjugate includes a spacer having a length of 0.1 nm to 500 nm attached to its 3' end, 5' end, or to a modified nucleotide base.
[0154] In some embodiments, the memory oligo is related to a unique molecular identifier (UMI) or barcode.
[0155] In some embodiments, the solid supports described herein include solid beads, porous beads, solid planar supports, porous planar supports, patterned or unpatterned surfaces, nanoparticles, or inorganic or polymeric microspheres. In some embodiments, the supports can include glass slides or wafers, silicon slides or wafers, PC, PTC, PE, HDPE, or other plastic surfaces, Teflon®, nylon, nitrocellulose, or other membranes, and the particles / beads can be polystyrene, cross-linked polystyrene, agarose, or acrylamide.
[0156] In some embodiments, the beads or nanoparticles are magnetic or paramagnetic.
[0157] In some embodiments, the solid support may be passivated with glass, silicon oxide, tantalum pentoxide, DLC diamond-like carbon, or other passivating agents, or the solid support may include a passivated or activated membrane via, for example, corona or other plasma treatment methods.
[0158] In some embodiments, the solid support may or may not be assembled with other components to facilitate fluid transport and / or detection (e.g., flow cells, biochips, microtiter plates).
[0159] In some embodiments, the solid support is composed of a hydrogel that supports the conjugation of components for macromolecular recoding and / or analysis workflows.
[0160] In some embodiments, the hydrogel is formed from synthetic polymers, natural polymers, and / or hybrid polymers. The monomers can include one or more of acrylamide, dihydroxymethacrylate, methacrylic acid, etc., in linear, branched, and / or cross-linked configurations, block copolymer configurations, or other configurations useful for sequencing macromolecules.
[0161] In some embodiments, the hydrogel comprises at least three orthogonal conjugation chemistry modalities.
[0162] In some embodiments, macromolecules (e.g., proteins, peptides) and / or universal primer sequences are covalently conjugated to a solid support.
[0163] In some embodiments, the binder comprises a polypeptide or protein, such as an antibody or a portion thereof (e.g., single-chain variable fragment (scFv), fragment antigen-binding (Fab) region, Fab2 region), nanobody, DNA aptamer, RNA aptamer, modified aptamer, photoactive or non-photoactive cage compound, oligopeptide permease (Opp), aminoacyl-tRNA synthetase (aaRS), periplasmic binding protein (PBP), dipeptide permease (Dpp), proton-dependent oligopeptide transporter (POT), modified aminopeptidase, modified aminoacyl-tRNA synthetase, modified anticalin, or modified Clp protease adapter protein (ClpS).
[0164] In some embodiments, the binder can selectively bind to an immobilized conjugate complex depending on the AA that is part of the complex.
[0165] In some embodiments, the binder comprises a binding moiety and a recoding tag.
[0166] In some embodiments, the recoding tag comprises a sequence representing AA information, and the recoding block comprises a sequence representing both workflow cycle and amino acid (or monomer identity) information.
[0167] In some embodiments, the binding moiety and the recoding tag are joined by a linker having a length of 0.1 nm to 500 nm.
[0168] In some embodiments, the chemically reactive conjugate and / or conjugate complex further comprises a spacer, a workflow cycle-specific sequence, a unique molecular identifier, a universal priming site, a restriction endonuclease cleavage sequence, or any combination thereof.
[0169] In some embodiments, the chemically reactive conjugate and / or conjugate complex comprises a spacer associated with a reactive moiety used for immobilization of the chemically reactive conjugate complex to a hydrogel surface, the spacer comprising a restriction endonuclease cleavage sequence that can release the PITC-AA moiety and / or cycle tag from the conjugate complex.
[0170] In some embodiments, the chemically reactive conjugate and / or conjugate complex comprises a spacer associated with a reactive moiety used for binding and cleavage of terminal amino acids, the spacer comprising a restriction endonuclease cleavage sequence that can release the cycle tag and / or reactive moiety used for immobilization from the conjugate complex.
[0171] In some embodiments, the chemically reactive conjugate can be in a pro-form, which means that through addition, activation, cleavage reactions or other operations, it can perform functions such as cycle identification (e.g., cycle tag), binding and cleavage of amino acids (e.g., PITC), and reactions to surfaces such as hydrogel-coated surfaces.
[0172] In some embodiments, transmission of the recode tag information to the recode block is mediated by DNA ligase and ligation oligos.
[0173] In some embodiments, transmission of the recode tag information to the recode block is mediated by DNA polymerase or by a combination of DNA polymerase and ligase.
[0174] In some embodiments, the transfer of recoding tag information to a recoding block is mediated by chemical ligation.
[0175] In some embodiments, a plurality of macromolecules and associated conjugate complexes are conjugated to a solid support.
[0176] In some embodiments, a plurality of pools having different combinations or compositions of binders with completely distinct or distinct but overlapping affinities can be introduced onto the surface of an immobilized chemically reactive conjugate. By using different pools with distinct binding characteristics, a more comprehensive and accurate characterization of the immobilized peptides can be achieved.
[0177] In some embodiments, the plurality of macromolecules are spaced on the solid support at an average distance exceeding 100 nm.
[0178] In some embodiments, the reactivity of residual chemically reactive conjugates (e.g., conjugates that are unreacted with amino acids but are immobilized on the surface due to insufficient removal by washing prior to initiating the immobilization reaction) is quenched by amino acids or amino acid mimics so as to be a bystander in future cycles.
[0179] In some embodiments, the modification of the terminal amino acids of a peptide prior to contacting it with a first chemically reactive conjugate increases the reactivity of the chemically active conjugate towards the modified amino acids compared to the unmodified amino acids. For example, the activation of the C-terminal amino acid by acetic anhydride prior to contacting with trimethylsilyl isothiocyanate has been described. Bailey, J.M., Shenoy, N.R., Ronk, M., & Shively, J.E., 1992, Protein Sci. 1, 68 - 80.
[0180] In some embodiments, the methods described herein comprise contacting a recoding block with polymerase, nucleotides, ligase, and / or buffer under conditions that allow elongation-ligation or ligation to assemble the recoding block into a memory oligo, and then contacting a plurality of incompletely ligated memory oligos with a linking oligo, polymerase, nucleotides, ligase, and / or buffer under conditions that allow elongation-ligation or ligation to assemble the incompletely ligated memory oligos into a memory oligo. Thus, the yield during memory oligo assembly can be increased.
[0181] In some embodiments, the methods described herein comprise contacting a recoding block with polymerase, nucleotides, ligase, and / or buffer under conditions that allow elongation-ligation or ligation to assemble the recoding block into a memory oligo, and then contacting a plurality of incompletely ligated memory oligo fragments and / or recoding blocks with a linking oligo, ligase, and buffer under conditions that facilitate ligation of the recoding block and the memory oligo fragments. Thus, the yield during memory oligo assembly can be increased.
[0182] In some embodiments, the linking oligo comprises a sequence complementary to the sequence of the recoding block, thereby facilitating ligation of any recoding blocks that were not ligated during contact with polymerase, nucleotides, ligase, and buffer.
[0183] In some embodiments, the linking oligo comprises additional nucleotide sequences that are encoded to carry information related to the sample or process and / or to assist with ligation or elongation-ligation.
[0184] In some embodiments, the memory oligos are amplified prior to analysis, for example, by bridge amplification, ExAmp NGS clustering, isothermal clustering, solution-based PCR amplification, A-tailing to add primer sequences prior to solution-based amplification, or any suitable DNA amplification method.
[0185] In some embodiments, the memory oligos optionally include a sample index, a spacer, a unique molecular identifier (UMI), a universal priming site, a CRISPR protospacer adjacent motif (PAM) sequence, or any combination thereof.
[0186] In some embodiments, a plurality of memory oligos are concentrated prior to analysis via a depletion process or a normalization process, for example, to remove or reduce the fraction of oligos associated with abundant proteins, peptides, or macromolecules. In some embodiments, concentration or depletion can be performed via a commercially available kit such as Agilent SureSelect, or via a custom concentration or depletion method using oligonucleotides that are partially complementary to the memory oligo sequences, for example, complementary to the AA tag sequence of the target memory oligo.
[0187] In some embodiments, a plurality of memory oligos representing a plurality of macromolecules are analyzed in parallel.
[0188] In some embodiments, the analysis of the memory oligos includes a nucleic acid sequencing method.
[0189] In some embodiments, analyzing the memory oligos includes analysis by multiplex PCR.
[0190] In some embodiments, the nucleic acid sequencing method includes sequencing by synthesis, sequencing by ligation, sequencing by hybridization, or pyrosequencing.
[0191] In some embodiments, the nucleic acid sequencing method includes single-molecule microscopy sequencing or nanopore sequencing.
[0192] In some embodiments, the memory oligos are configured to be analyzed using commercially available NGS technologies such as NGS methods exemplified by Illumina, Element Bio, and Singular Genomics.
[0193] In some embodiments, the chemically reactive conjugate and / or conjugate complex includes a cleavable group adjacent to a matching unique molecular identifier (UMI) within the cycle tag to facilitate cleavage of the memory oligo at a specified location. In these embodiments, one or more restriction endonuclease sequences carried by one or more cycle tag sequences assembled to the memory oligo are cleaved to create one or more oligonucleotides (memory oligos). The oligonucleotides are short enough to be fully read using short-read DNA sequencing technologies, including short-chain DNA sequencing methods and devices commercially available from Illumina, Element Bio, and Singular Genomics.
[0194] In some embodiments, helicase can be utilized during the assembly of the memory oligo. The use or strobe of helicase during one or more assembly processes can, in some instances, improve access to DNA blocks and facilitate longer memory oligo assembly.
[0195] In some embodiments, the memory oligo or its recoding block is configured to be analyzed using a decode-based methodology. Further information regarding decode-based technologies can be found in Gunderson et al., Decoding Randomly Ordered DNA Arrays, Genome Res., 2004 May;14(5):870-7, which is hereby incorporated by reference in its entirety for all purposes.
[0196] In some embodiments, a set of any such spatially limited constructs that include memory oligos or fragments of recode blocks, or sequences and identity information related to a given peptide, protein, protein complex or polymer are analyzed using a decode-based methodology. See Gunderson et al.
[0197] In some embodiments, the identification components are selected from UMIs, sample indices, recode tags, recode blocks, ligation oligos, AA tags, their complements, or any combination thereof.
[0198] In some embodiments, the N-terminal AA of the peptide is removed by chemical cleavage in place of Edman cleavage.
[0199] In some embodiments, one or more chemically reactive conjugates bind to the terminal amino acid residue of the peptide.
[0200] In some embodiments, one or more binders bind to the conjugate complex.
[0201] In some embodiments, the conjugate complex includes post-translationally modified amino acids.
[0202] In some embodiments, the identification components of the recode tag, recode block, or both include error detection and / or correction bits.
[0203] In some embodiments, the error detection / correction sequences are derived from Hamming distance theory, or other up-to-date digital code space theories (e.g., Lee, Levenshtein-Tenengolts, Reed-Solomon, or others).
[0204] In some embodiments, the components of the recode tag, recode block, or both include 2, 3, 4, 5, 6 or more different types of nucleotides.
[0205] In some embodiments, the code (e.g., an array) associated with a recode tag or recode block via analysis of a memory oligo is derived from 2, 3, 4, 5, 6 or more types of nucleotides.
[0206] In some embodiments, the number of different types of nucleotides used to create the code for recoding is not equal to the number of nucleotide types including the recode tag, cycle tag, or both.
[0207] In some embodiments, macromolecule, fragment or peptide activation includes functional moieties such as NHS groups, aldehyde groups, azide groups, alkyne groups, maleimide groups, thiol groups, tetrazine and trans - cyclooctene.
[0208] In some embodiments, the immobilized peptide is linearized (denatured) using detergents, surfactants, chaotropic agents, reducing agents, and / or alkylating agents.
[0209] In some embodiments, a chemically reactive conjugate reacts to cleave from the C - terminus rather than the N - terminus of the peptide to create a recode block that can be assembled using any of the methods described herein.
[0210] In some embodiments, "paired - end read" information is used to create a recode block that can be assembled using the methods described herein by using chemically reactive conjugates that operate sequentially or in parallel at both the N - terminus and C - terminus of a given protein complex, protein, or peptide to create a recode block that can be assembled using the methods described herein, and the recode block can be collected from the immobilized protein complex, protein, or peptide.
[0211] In certain embodiments, a method is provided for obtaining pre-defined code information via sequencing a subset of nucleotide types in an oligonucleotide or oligonucleotide cluster. This is particularly useful when considering the reading of information stored in DNA (e.g., DNA data storage information technology reading).
[0212] In some aspects, information re-coded into a memory oligo is obtained via sequencing a subset of nucleotide types in the memory oligo. For example, in a sequencing read by introducing non-fluorescent non-reversible terminating nucleotides into an SBS sequencing reagent mix, a subset of nucleotide types may be identified, and a subset of nucleotide types may not be identified. In certain embodiments, the subset is two of the four natural nucleotides.
[0213] In certain embodiments, a method for preparing a plurality of peptides of mer length n = 2 to 2000 conjugated to a solid support, the method comprising: (a) fragmenting one or more peptides, proteins, and / or protein complexes in a sample; (b) activating zero, one, two, or more moieties of each fragmented peptide, protein, and / or protein complex; (c) optionally, binding a sample-specific nucleotide index sequence to the activated peptide, protein, and / or protein complex; and (d) conjugating the peptide to the solid support.
[0214] In some aspects, one or more operations of the method are performed in any suitable sequential order or simultaneously.
[0215] In some embodiments, subunits of a given protein are co-immobilized either directly or through interaction with native subunits on the surface. Subsequently, one or more subunits can be simultaneously recoded within the same local region by processes (b)-(m) including alternative embodiments related to the method. The information of the memory oligos can include a mixture of subunits (protein and native) that can be deconvolved in silico.
[0216] In certain embodiments, a method for preparing an interaction peptide or peptides conjugated to a solid support, comprising: (a) crosslinking crosslinking peptides, proteins, and / or protein complexes in one or more samples (e.g., using homo-bifunctional, hetero-bifunctional, or photoreactive methods as described in Kluger, et al., (2004) Bioorganic Chemistry v32:6,451); (b) activating zero, one, two, or more portions of each crosslinked peptide, protein, and / or protein complex for immobilization to the solid support; (c) optionally, conjugating a sample-specific nucleotide index sequence to the activated peptide, protein, and / or protein complex; and (d) conjugating the complex to the solid support. In some embodiments, one or more operations of the method are performed in any suitable sequential order or simultaneously. Generally, this method enables the analysis of in vivo-related proteins and their interactions, thus facilitating the discovery, identification, and investigation of protein intermediates.
[0217] In certain embodiments, a method for preparing interacting DNA - peptide or multiple interacting DNA - peptide complexes linked to a solid support, the method comprising: (a) cross - linking a peptide, protein, and / or protein complex with native DNA to which the protein has associated in a biological context for one or more samples (e.g., using formaldehyde or other methods known in the art); (b) activating zero, one, two, or more portions of each cross - linked peptide - DNA, protein, and / or protein complex - DNA complex; (c) optionally, conjugating a sample - specific nucleotide index sequence to the activated peptide - DNA and / or protein - DNA complex; and (d) conjugating the complex to a solid support. In some embodiments, one or more operations of the method are performed in any suitable sequential order or simultaneously. Generally, the method provides an analysis of in vivo interactions between proteins and DNA.
[0218] In some embodiments, fragmentation includes physical shearing, endopeptidase activity, modified endopeptidase activity, proteases, metalloproteases, and / or other suitable fragmentation methods.
[0219] In some embodiments, the peptide includes any suitable macromolecular polymer including proteins, peptides, etc. In such embodiments, the monomer units of the macromolecular polymer can include amino acids, carbohydrates, and / or any monomer moiety that can be incorporated into the polymer.
[0220] In some embodiments, the method further comprises depleting one or more abundant proteins from the sample prior to any of operations (a), (b), (c), and / or (d).
[0221] In certain embodiments, the use of a chemically reactive conjugate with a cleavable spacer enables rejuvenation of the surface of the substrate for a second recoding. For example, in certain embodiments, a method for analyzing one or more residual immobilized analytes from a surface having a plurality of peptides, proteins, and / or protein complexes, the method comprising: (a) providing a surface that is rejuvenated by cleaving a spacer of a first chemically reactive conjugate and that is used in a previous round prior to the recoding operations (b)-(d) described below; (b) providing a second chemically reactive conjugate (e.g., a PITC conjugate), the conjugate comprising a cycle tag having identification information regarding the workflow cycle of the method, a reactive moiety that can bind and cleave to the terminal amino acid of the peptide, and a reactive moiety that facilitates immobilization to a solid support; (c) contacting the peptide with the second chemically reactive conjugate, the second chemically reactive conjugate binding to the terminal amino acid or modified terminal portion of the peptide to form a second conjugate complex, e.g., a PIT-AA-cycle tag conjugate complex; (d) immobilizing the second conjugate complex to the solid support; (e) cleaving the terminal amino acid from the peptide, thereby providing a second immobilized conjugate complex and a new terminal amino acid of the peptide conjugated to the solid support of (a); (f) optionally repeating (b)-(e) to assemble a second immobilized conjugate complex having cycle information for the new terminal amino acid of the peptide; (g) optionally repeating (f) for additional iteration cycles to create additional immobilized conjugate complexes for additional amino acids of the peptide of step (a); (h) optionally deprotecting the nucleic acid of the conjugate complex and / or any protected nucleic acid associated with the solid support; (i) contacting the second immobilized conjugate complex with a binder that can bind to the second immobilized conjugate complex, the binder comprising a binding moiety and a recoding tag having identification information regarding the first binder.(j) To transfer information of the recoding tag associated with the binder to the cycle tag of the second immobilized complex to generate a second recoding block; and (k) Optionally, repeating (i) and (j) using a second binder including a binding portion and a recoding tag having identification information regarding the second binder to transfer information of the recoding tag associated with the second binder to the second immobilized conjugate complex to generate a second recoding block; (l) Optionally, repeating (k) for additional cycles to create a recoding block for additional amino acids of the peptide of step (a); and (m) Contacting the recoding block with a polymerase, nucleotides, ligase, and buffer under conditions that allow extension-ligation to assemble the recoding block into a memory oligo or create fragments of the memory oligo; and (n) Analyzing the memory oligo.
[0222] In some embodiments, the foregoing embodiments associated with the first operation round are applied to the second operation round.
[0223] In some embodiments, one or more operations of the method are performed in any suitable sequential order or simultaneously.
[0224] In some embodiments, the recovery process is repeated one or more times.
[0225] In some embodiments, only a portion of the chemically reactive conjugate is cleaved from the surface because it may be desirable to retain a portion of the recoding block to facilitate in silico mapping and assembly over iterative cycles of memory oligo assembly.
[0226] In some embodiments, surface recovery may include "stropping" the protein using either a chemical (e.g., phenylisothiocyanate (PITC)) or biological (e.g., aminopeptidase) method.
[0227] In some embodiments, the amine groups of the remaining non-cleaved recoded block nucleobases are protected by reaction with fluorenylmethyloxycarbonyl (FMOC) or other standard protection chemistries.
[0228] In some embodiments, following process (m) of the method, a plurality of assembly oligos containing all or a portion of the possible assembly oligos are hybridized to memory oligos, ligated, and dehybridized to form solution-phase memory oligos.
[0229] In some embodiments, a method for analyzing one or more peptides from a sample comprising a plurality of peptides, proteins, and / or protein complexes, the method comprising: (a) providing a peptide of mer length n = 2 to 2000 attached to a solid support; (b) providing a first chemically reactive conjugate, the conjugate comprising a cycle tag, a reactive moiety that can bind to and cleave the terminal amino acid of the peptide, and a reactive moiety that facilitates immobilization to the solid support; (c) contacting the peptide with the first chemically reactive conjugate, wherein the first chemically reactive conjugate binds to the terminal amino acid or modified terminal portion of the peptide to form a first conjugate complex; (d) immobilizing the first conjugate complex to the solid support; (e) cleaving the terminal amino acid from the peptide, thereby providing a first immobilized conjugate complex and a new terminal amino acid of the peptide attached to the solid support of (a); (f) optionally repeating steps (b)-(e) to assemble a second immobilized conjugate complex having cycle information about the new terminal amino acid of the peptide; (g) optionally repeating (f) for additional iterative cycles to generate additional immobilized conjugate complexes for additional amino acids of the peptide of step (a); (h) optionally deprotecting nucleic acids of the conjugate complex and / or any protected nucleic acids associated with the solid support; (i) contacting the first immobilized conjugate complex with a first binder capable of binding to the first immobilized conjugate complex, the first binder comprising a binding moiety and a recoding tag having identification information regarding the first binder; (j) transmitting information of the recoding tag associated with the first binder to the cycle tag of the first immobilized complex to generate a first recoding block; (k) optionally repeating (i) and (j) using a second binder comprising a binding moiety and a recoding tag having identification information regarding the second binder to transmit information of the recoding tag associated with the second binder to the second immobilized conjugate complex to generate a second recoding block; (l) optionally,Repeating (k) for additional cycles to create a recoding block for additional amino acids of the peptide of step (a), and (m) contacting the recoding block with a polymerase, nucleotides, ligase, and buffer under conditions that allow extension-ligation to assemble the recoding block into a memory oligo or create fragments of the memory oligo, and (n) analyzing the memory oligo. Any of the foregoing method steps may be used alone or in combination with other steps or methods described herein. In some embodiments, (e), (i), and (j) are repeated one or more times to increase the process yield of the method. Some embodiments include, after (m) and / or (l), contacting a first immobilized conjugate complex with an indiscriminate binder that can bind to the first immobilized conjugate complex regardless of the identity of the amino acid within the conjugate complex, the indiscriminate binder including a binding moiety that associates with the immobilized conjugate regardless of the amino acid and an indiscriminate recoding tag that can hybridize with any cycle tag and carries identification information regarding the indiscriminate binder. In some embodiments, the conjugate complex includes zero, one, or more reactive moieties, and the reaction includes activatable chemistry and / or reversible chemistry. In some embodiments, the recoding tag associated with the first binder is a nucleic acid having a sequence corresponding to the (n-1)th cycle tag, amino acid (AA) tag, and nth cycle tag. In some embodiments, (i)-(l) are performed simultaneously. In some embodiments, operation (m) includes contacting the recoding block with a ligase, an AA tag oligonucleotide complement, and a buffer under conditions that allow ligation to assemble the recoding block and the AA tag oligonucleotide complement into the memory oligo. In some embodiments, the memory oligo, cycle tag, and recoding block each include a nucleic acid molecule. In some embodiments, the memory oligo includes a universal priming site, and the universal priming site is a priming site for amplification or a priming site for sequencing,or both. In some embodiments, the binder comprises a polypeptide or a protein.,
[0230] In some embodiments, a method for determining the identity and positional information of amino acid residues of a peptide coupled to a solid support, the method comprising: (a) providing the peptide to the solid support, wherein the peptide is coupled to the solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) providing a chemically reactive conjugate, the chemically reactive conjugate comprising: (x) a cycle tag comprising a cyclic nucleic acid associated with the number of cycles; (y) a reactive moiety for binding to and cleaving the N-terminal amino acid residue of the peptide to expose the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide; and (z) an immobilization moiety for immobilizing on the solid support; (c) contacting the peptide with the chemically reactive conjugate, thereby coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex on the solid support via the immobilization moiety; (e) cleaving the N-terminal amino acid residue and separating it from the peptide, thereby providing an immobilized amino acid complex, the immobilized amino acid complex comprising the cleaved and separated N-terminal amino acid residue; (f) contacting the immobilized amino acid complex with a binder, the binder comprising a binding moiety for preferentially binding to the immobilized amino acid complex and a recoding tag comprising a recoding nucleic acid corresponding to the binder, thereby forming an affinity complex, the affinity complex comprising the immobilized amino acid complex and the binder, and thereby bringing the cycle tag formed in each affinity complex into proximity with the recoding tag in the affinity complex; (g) transmitting the information of the nucleic acid recoding tag associated with the first binder to the cycle tag of the first immobilized complex to generate a first recoding block; (j) obtaining the sequence information of the recoding block; and (k) determining the identity and positional information of the amino acid residues of the peptide based on the obtained sequence information. In some embodiments, the immobilized amino acid complex is washed before contacting with the binder.In some embodiments, the sequence information is used to determine the possible three-dimensional structure of the peptide. Some embodiments include repeating steps (b)-(k) for each successive amino acid in the peptide.
[0231] In some embodiments, a method for determining the identity and positional information of amino acid residues of a peptide coupled to a solid support, the method comprising: (a) providing the peptide to the solid support, wherein the peptide is coupled to the solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) providing a chemically reactive conjugate, the chemically reactive conjugate comprising: (x) a cycle tag comprising a cyclic nucleic acid associated with the number of cycles; (y) a reactive moiety for binding to and cleaving the N-terminal amino acid residue of the peptide to expose the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide; and (z) an immobilization moiety for immobilizing to the solid support; (c) contacting the peptide with the chemically reactive conjugate, thereby coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex to the solid support via the immobilization moiety; (e) cleaving the N-terminal amino acid residue and separating it from the peptide, thereby providing an immobilized amino acid complex, the immobilized amino acid complex comprising the cleaved and separated N-terminal amino acid residue; (f) contacting the immobilized amino acid complex with a binding agent, the binding agent comprising a binding moiety for preferentially binding to the immobilized amino acid complex and a recoding tag comprising a recoding nucleic acid corresponding to the binding agent, thereby forming an affinity complex, the affinity complex comprising the immobilized amino acid complex and the binding agent, thereby bringing the cycle tag into proximity to the recoding tag within the affinity complex; (g) ligating the recoding nucleic acid or the sequence of the recoding nucleic acid to the cycle nucleic acid or the sequence of the cycle nucleic acid to generate a recoding block; (j) obtaining the sequence information of the recoding block; and (k) determining the identity and positional information of the amino acid residues of the peptide based on the obtained sequence information. Some embodiments include repeating steps (b)-(k) for the next amino acid of the peptide. Some embodiments include repeating steps (b)-(k) for each subsequent amino acid of the peptide.Some embodiments include washing the immobilized amino acid complex before contacting the immobilized amino acid complex with a binder. Some embodiments include determining a potential three-dimensional structure of a peptide based on sequence information.
[0232] In some embodiments, a method for determining the identity and positional information of a plurality of amino acid residues of a peptide, wherein the peptide comprises n amino acid residues, comprising: (a) coupling the peptide to a solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) providing a chemically reactive conjugate, the chemically reactive conjugate comprising: (x) a cycle tag comprising a cycle nucleic acid associated with the number of cycles; (y) a reactive moiety for binding to and cleaving the N-terminal amino acid residue of the peptide to expose the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide; and (z) an immobilization moiety for immobilizing on the solid support; (c) contacting the peptide with the chemically reactive conjugate, thereby coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex on the solid support via the immobilization moiety; (e) cleaving the N-terminal amino acid residue, thereby separating it from the peptide and exposing the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide, providing an immobilized amino acid complex, the immobilized amino acid complex comprising the cleaved and separated N-terminal amino acid residue; (f) repeating (b) to (e) n - 1 times to assemble n - 1 additional immobilized amino acid complexes, each additional immobilized amino acid complex comprising a nucleic acid associated with cycles 2 to n accordingly; (g) contacting the immobilized amino acid complex with a binder, each binder comprising a binding moiety for preferentially binding to one or a subset of the immobilized amino acid complexes and a recoding tag comprising a recoding nucleic acid corresponding to the binder, thereby forming one or more affinity complexes, each affinity complex comprising the immobilized amino acid complex and the binder, thereby bringing the cycle tag into proximity to the recoding tag in each formed affinity complex; (h) joining the cycle tag to the recoding tag in each formed affinity complex to form a recoding block, thereby creating a plurality of recoding blocks, each recoding block corresponding to the formed affinity complex, creating a plurality of recoding blocks;(i) joining two or members of a plurality of recoding blocks to form a memory oligonucleotide; (j) obtaining sequence information of the memory oligonucleotide; and (k) determining the identity and position information of a plurality of amino acid residues of a peptide based on the obtained sequence information. A method is disclosed herein. In some embodiments, n is an integer greater than or equal to 2. In some embodiments, each binder comprises a recoding tag having a unique nucleic acid sequence. In some embodiments, the plurality of binders comprises recoding tags having the same nucleic acid sequence. In some embodiments, the binder comprises a recoding tag that may have a unique sequence portion and a common sequence portion.,
[0233] In some embodiments, determining the identity and position information of a plurality of amino acid residues of a peptide includes determining the identity and position information of all amino acid residues of the peptide. In some embodiments, determining the identity and position information of a plurality of amino acid residues of a peptide includes determining the identity and position information of only a subset of amino acid residues of the peptide. Some embodiments include, for example, identifying the peptide by comparing the identity and position information of the plurality of amino acid residues to a database.
[0234] In some embodiments, a method for determining the identity and positional information of amino acid residues of a peptide coupled to a solid support, the method comprising: (a) providing a peptide to the solid support, wherein the peptide is coupled to the solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) providing a chemically reactive conjugate, the chemically reactive conjugate comprising: (x) a cycle tag associated with the number of cycles; (y) a reactive moiety for binding to and cleaving the N-terminal amino acid residue of the peptide to expose the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide; and (z) an immobilization moiety for immobilizing to the solid support; (c) contacting the peptide with the chemically reactive conjugate, thereby coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex; (d) immobilizing the conjugate complex to the solid support via the immobilization moiety; (e) cleaving the N-terminal amino acid residue, thereby separating it from the peptide, thereby providing an immobilized amino acid complex, the immobilized amino acid complex comprising the cleaved and separated N-terminal amino acid residue; (f) contacting the immobilized amino acid complex with a binder, the binder comprising a binding moiety for preferentially binding to the immobilized amino acid complex and a recoding tag comprising a recoding nucleic acid corresponding to the binder, thereby forming an affinity complex, the affinity complex comprising the immobilized amino acid complex and the binder, thereby bringing the cycle tag into proximity to the recoding tag within the affinity complex; (g) transmitting the information of the recoding nucleic acid associated with the binder to the cycle tag of the immobilized complex to generate a recoding block; (j) obtaining the sequence information of the recoding block; and (k) determining the identity and positional information of the amino acid residues of the peptide based on the obtained sequence information. A method is disclosed herein.
[0235] Recoding tag In this specification, in some embodiments, recoding tags are disclosed. The recoding tags may be part of a binder. The recoding tags may correspond to a binder. For example, the recoding tags may convey information regarding the molecule (e.g., amino acid or PTM) to which the binder binds. The recoding tags may include nucleic acids such as recoding nucleic acids. In some embodiments, the recoding nucleic acids include DNA or RNA. In some embodiments, the recoding tags are DNA sequences. In some embodiments, the recoding tags are RNA sequences. The recoding nucleic acids may be useful for encoding amino acid information in nucleic acids. The recoding tags may be used in the methods described herein, for example, methods for determining protein information such as the position or identity of amino acids.
[0236] Recoding block In this specification, in some embodiments, recoding blocks are disclosed. The recoding blocks may include a cycle tag and a recoding tag or its reverse complement. The recoding blocks may include a cycle tag or its reverse complement and a recoding tag. The recoding blocks may include a cycle tag or its reverse complement and a recoding tag or its reverse complement. The recoding blocks may include a cycle tag and a recoding tag, or information corresponding to a cycle tag and a recoding tag. For example, the recoding blocks may include a cyclic nucleic acid, a cyclic nucleic acid sequence, or their reverse complements, and may include a recoding nucleic acid, a recoding nucleic acid sequence, or their reverse complements. The recoding blocks may be useful for joining to a memory oligonucleotide, and any of the memory oligonucleotides may convey information regarding the amino acid position and identity within a protein. The recoding blocks may be used in the methods described herein, for example, methods for determining protein information such as the position or identity of amino acids.
[0237] In some embodiments, the recoding block comprises a recoded nucleic acid, the sequence of the recoded nucleic acid, or the reverse complement of the sequence of the recoded nucleic acid conjugated or combined with a cycled nucleic acid, the sequence of the cycled nucleic acid, or the reverse complement of the sequence of the cycled nucleic acid. In some embodiments, the recoding block comprises the reverse complement of the sequence of the recoded nucleic acid conjugated with a recoded nucleic acid or a cycling tag. In some embodiments, the recoding block comprises a recoded nucleic acid, the sequence of the recoded nucleic acid, or the reverse complement of the sequence of the recoded nucleic acid. In some embodiments, the recoding block comprises a cycled nucleic acid, the sequence of the cycled nucleic acid, or the reverse complement of the sequence of the cycled nucleic acid.
[0238] Transfer of information In some embodiments, methods are disclosed herein that include transferring information. For example, the method can include transferring the information of a recoded nucleic acid to a cycled nucleic acid of an immobilized conjugate complex to generate a recoding block. Transfer of information may form a recoding block or may be used to form a memory oligonucleotide. Transfer of information can be included in the methods described herein, such as methods for determining protein information such as the position or identity of an amino acid.
[0239] In some embodiments, communicating the information includes performing nucleic acid sequence-based amplification, for example, to generate a sequence of a recoded nucleic acid or a sequence of a cyclic nucleic acid. In some embodiments, communicating the information includes performing a polymerase chain reaction (PCR) to generate a sequence of a recoded nucleic acid or a sequence of a cyclic nucleic acid. In some embodiments, PCR includes real-time PCR, digital PCR, multiplex PCR, nested PCR, hot-start PCR, touchdown PCR, or quantitative PCR. In some embodiments, communicating the information includes performing or carrying out, for example, a ligase chain reaction, helicase-dependent amplification, strand displacement amplification, loop-mediated isothermal amplification, rolling circle amplification, recombinase polymerase amplification, nicking enzyme amplification reaction, whole genome amplification, transcription-mediated amplification, multiple displacement amplification, or multiple annealing and loop-based amplification cycles to generate a sequence of a recoded nucleic acid or a sequence of a cyclic nucleic acid. The amplification or other procedure can be to generate a sequence of a recoded nucleic acid, a sequence of a cyclic nucleic acid, a reverse complement, or a combination thereof. In some embodiments, the information of the recoded nucleic acid includes a sequence of the recoded nucleic acid or a reverse complement of the sequence of the recoded nucleic acid.
[0240] In some embodiments, the transmission of information involves a polymerase chain reaction. In some embodiments, the transmission of information involves a reverse transcription polymerase chain reaction. In some embodiments, the transmission of information includes a real-time polymerase chain reaction. In some embodiments, the transmission of information includes a digital polymerase chain reaction. In some embodiments, the transmission of information includes a multiplex polymerase chain reaction. In some embodiments, the transmission of information includes a nested polymerase chain reaction. In some embodiments, the transmission of information includes a hot start polymerase chain reaction. In some embodiments, the transmission of information includes a touchdown polymerase chain reaction. In some embodiments, the transmission of information involves a quantitative polymerase chain reaction. In some embodiments, the transmission of information involves a ligase chain reaction. In some embodiments, the transmission of information includes a helicase-dependent amplification. In some embodiments, the transmission of information includes a strand displacement amplification. In some embodiments, the transmission of information includes a loop-mediated isothermal amplification. In some embodiments, the transmission of information includes a rolling circle amplification. In some embodiments, the transmission of information involves a recombinase polymerase amplification. In some embodiments, the transmission of information includes a nicking enzyme amplification reaction. In some embodiments, the transmission of information includes a whole genome amplification. In some embodiments, the transmission of information includes a transcription-mediated amplification. In some embodiments, the transmission of information includes a multiple displacement amplification. In some embodiments, the transmission of information includes a plurality of annealing and loop-based amplification cycles. In some embodiments, the transmission of information includes a nucleic acid sequence-based amplification.
[0241] In some embodiments, transmitting the information includes conjugating a recoded nucleic acid or a reverse complement of the recoded nucleic acid to a cyclic nucleic acid.
[0242] Conjugation In some embodiments, methods involving ligation are disclosed herein. For example, a recoded nucleic acid or its reverse complement may be ligated to a cyclic nucleic acid or its reverse complement. The ligation may form a recoding block or may be used to form a memory oligonucleotide. The ligation may be included in the methods described herein, such as methods for determining protein information such as amino acid position or identity.
[0243] In some embodiments, the ligation includes enzymatic ligation. In some embodiments, the ligation includes splint ligation. In some embodiments, the ligation includes chemical ligation. In some embodiments, the ligation includes template-assisted ligation. In some embodiments, ligating includes the use of a ligase enzyme. In some embodiments, the ligation includes the use of a splint oligonucleotide. In some embodiments, the ligation includes the use of a catalyst. In some embodiments, the ligation includes the use of a crosslinking molecule. In some embodiments, the ligation includes the use of a condensing agent. In some embodiments, the ligation includes the use of a coupling reagent. In some embodiments, the ligation includes the use of a polymerase enzyme. In some embodiments, the ligation includes the use of a complementary nucleic acid sequence. In some embodiments, the ligation includes the use of a nicking enzyme. In some embodiments, the ligation includes the use of a nucleic acid modifying enzyme. In some embodiments, the ligation includes the use of a recombinase. In some embodiments, the ligation includes the use of a strand-displacing polymerase. In some embodiments, the ligation includes the use of a single-strand binding protein. In some embodiments, the ligation includes a click chemistry reaction. In some embodiments, the ligation includes phosphodiester bond formation. In some embodiments, the ligation includes peptide nucleic acid-mediated ligation. In some embodiments, each binder includes a recoding tag having a unique nucleic acid sequence. In some embodiments, a plurality of binders includes a recoding tag having the same nucleic acid sequence. In some embodiments, the binder includes a recoding tag that may have a unique sequence portion and a common sequence portion.
[0244] In some embodiments, joining a recoding nucleic acid or a sequence of a recoding nucleic acid to a cycling nucleic acid or a sequence of a cycling nucleic acid to generate a recoded block includes (i) joining a recoding nucleic acid to a cycling nucleic acid, (ii) joining a recoding nucleic acid to a sequence of a cycling nucleic acid, (iii) joining a sequence of a recoding nucleic acid to a cycling nucleic acid, or (iv) joining a sequence of a recoding nucleic acid to a sequence of a cycling nucleic acid. Some embodiments include performing nucleic acid sequence-based amplification to generate a sequence of a recoding nucleic acid or a sequence of a cycling nucleic acid. Some embodiments include performing polymerase chain reaction (PCR) to generate a sequence of a recoding nucleic acid or a sequence of a cycling nucleic acid. In some embodiments, PCR includes real-time PCR, digital PCR, multiplex PCR, nested PCR, hot start PCR, touchdown PCR, or quantitative PCR. Some embodiments include performing ligase chain reaction, helicase-dependent amplification, strand displacement amplification, loop-mediated isothermal amplification, rolling circle amplification, recombinase polymerase amplification, nicking enzyme amplification reaction, whole genome amplification, transcription-mediated amplification, multiple displacement amplification, or multiple annealing and loop-based amplification cycles to generate a sequence of a recoding nucleic acid or a sequence of a cycling nucleic acid.
[0245] In some embodiments, joining includes performing enzymatic ligation, sprint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a sprint oligonucleotide, use of a catalyst, use of a crosslinking molecule, use of a condensing agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand displacement polymerase, use of a single-stranded binding protein, click chemistry reaction, phosphodiester bond formation, or peptide nucleic acid-mediated ligation.
[0246] Some embodiments include contacting an additional immobilized amino acid complex with a second binder. In some embodiments, the binder and the second binder include distinct recoding tags having different recoded nucleic acids from each other. In some embodiments, the binder and the second binder include recoding tags having the same recoded nucleic acid as each other. In some embodiments, the binder and the second binder include distinct recoding tags having different sequences from each other and having a portion of the recoded nucleic acid that is the same recoded nucleic acid.
[0247] In some embodiments, communicating the information includes joining or combining a recoded nucleic acid, a sequence of a recoded nucleic acid, or a reverse complement of a sequence of a recoded nucleic acid, with a cycled nucleic acid, a sequence of a cycled nucleic acid, or a reverse complement of a sequence of a cycled nucleic acid, to generate a recoding block.
[0248] Memory oligo readout In some embodiments, methods are disclosed herein that include memory oligonucleotides. The memory oligonucleotides can include a plurality of recoding blocks, reverse complements of a plurality of recoding blocks, or one or more recoding blocks and one or more reverse complements of recoding blocks. The memory oligonucleotides can be used in the methods described herein, such as methods for determining protein information such as amino acid position or identity.
[0249] In some embodiments, obtaining array information of recoding blocks includes performing array determination. In some embodiments, obtaining array information of memory oligonucleotides includes performing array determination. Memory oligonucleotides may include a recoding block or multiple recoding blocks. In some embodiments, sequencing includes Sanger sequencing. In some embodiments, sequencing includes next-generation sequencing. In some embodiments, sequencing includes pyrosequencing, sequencing by synthesis, sequencing by ligation, Illumina sequencing, Ion Torrent sequencing, Pacific Biosciences sequencing, Oxford Nanopore sequencing, SOLiD sequencing, nanopore sequencing, single molecule real-time (SMRT) sequencing, 454 sequencing, Complete Genomics sequencing, Helicos sequencing, MinION sequencing, direct RNA sequencing, Linked-Read sequencing, mate pair sequencing, or target gene sequencing.
[0250] In some embodiments, the sequence information of the memory oligonucleotide is obtained by sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by Sanger sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by next-generation sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by pyrosequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by synthetic sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by ligation-based sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by Illumina sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by Ion Torrent sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by Pacific Biosciences sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by Oxford Nanopore sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by SOLiD sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by nanopore sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by single-molecule real-time (SMRT) sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by 454 sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by complete genomics sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by Helicos sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by MinION sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by direct RNA sequencing.In some embodiments, the sequence information of the memory oligonucleotide is obtained by Linked-Read Sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by mate pair sequencing. In some embodiments, the sequence information of the memory oligonucleotide is obtained by targeted gene sequencing.
[0251] Some embodiments include the aggregation of information from only a subset of the cycles. Some embodiments include, for example, the analysis of peptide information that does not include all amino acids of a peptide using sequencing information generated by a recoding process that does not include all amino acids of the peptide (e.g., from memory oligonucleotides formed from the sequences of recoding tags and cycle tags). In some embodiments, only some amino acids of a protein are recoded into a recoding block. A memory oligo may include a recoding block corresponding to all or only some of the amino acids of a peptide. Missing amino acid information may be considered when reconstructing or identifying a peptide. Some memory oligonucleotides include a recoding block having a recoding tag sequence and a cycle tag sequence.
[0252] Binder In some embodiments, a binder is disclosed herein. The binder may include a recoding tag and a binding moiety. The recoding tag may include a recoding nucleic acid. The binder may be used in the methods described herein, for example, in methods for determining protein information such as the position or identity of an amino acid.
[0253] In some embodiments, the binding moiety comprises a peptide. In some embodiments, the binding moiety comprises an antibody. In some embodiments, the antibody comprises a monoclonal antibody, a polyclonal antibody, an antibody fragment, an antibody derivative, a bispecific antibody, a nanobody, or a single-domain antibody. In some embodiments, the antibody comprises an antibody fragment such as Fab, F(ab’)2, or scFv. In some embodiments, the binding moiety comprises an antibody derivative such as an antibody-drug conjugate, a synthetic antibody, an antibody mimetic, an engineered protein binder such as a DARPin or an Affibody, an aptamer, a ligand for a peptide receptor, a small molecule, a lectin, an enzyme substrate, an RNA molecule, or a DNA molecule.
[0254] In some embodiments, the binder comprises an antibody. In some embodiments, the binder comprises a monoclonal antibody. In some embodiments, the binder comprises a polyclonal antibody. In some embodiments, the binder comprises an antibody fragment such as Fab, F(ab’)2, or scFv. In some embodiments, the binder comprises an antibody derivative such as an antibody-drug conjugate. In some embodiments, the binder comprises a bispecific antibody. In some embodiments, the binder comprises a synthetic antibody or an antibody mimetic. In some embodiments, the binder comprises an aptamer. In some embodiments, the binder comprises a nanobody or a single-domain antibody. In some embodiments, the binder comprises an engineered protein binder such as DARPins or Affibodies. In some embodiments, the binder comprises a peptide. In some embodiments, the binder comprises a ligand for a peptide receptor. In some embodiments, the binder comprises a small molecule. In some embodiments, the binder comprises a lectin. In some embodiments, the binder comprises an enzyme substrate. In some embodiments, the binder comprises an RNA molecule. In some embodiments, the binder comprises a DNA molecule.
[0255] In some embodiments, the binder further comprises a second tag. In some embodiments, the second tag comprises a fluorescent tag for visualization, a biotin tag for interacting with streptavidin, a radioactive tag for detection, a quantum dot for visualization, a detection tag based on mass spectrometry, a chromogenic tag for visualization, a chemiluminescent tag for detection, a photoacoustic imaging tag, a single molecule imaging tag, or a dual modality imaging tag.
[0256] In some embodiments, the binder is labeled with a second tag for visualization. In some embodiments, the binder is labeled with a fluorescent tag for visualization. In some embodiments, the binder is labeled with a biotin tag for subsequent interaction with streptavidin. In some embodiments, the binder is labeled with a radioactive tag for detection. In some embodiments, the binder is labeled with a quantum dot for visualization. In some embodiments, the binder is labeled with a second tag for detection based on mass spectrometry. In some embodiments, the binder is labeled with a chromogenic tag for visualization. In some embodiments, the binder is labeled with a chemiluminescent tag for detection. In some embodiments, the binder is labeled with a second tag for photoacoustic imaging. In some embodiments, the binder is labeled with a second tag for single molecule imaging. In some embodiments, the binder is labeled with a second tag for dual modality imaging.
[0257] In some embodiments, the linking moiety binds to any of the following amino acids: Ala, Arg, Asn, Asp, Cys, Gln, Glu, Gly, His, Ile, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, or Val. In some embodiments, the linking moiety binds to Ala. In some embodiments, the linking moiety binds to Arg. In some embodiments, the linking moiety binds to Asn. In some embodiments, the linking moiety binds to Asp. In some embodiments, the linking moiety binds to Cys. In some embodiments, the linking moiety binds to Gln. In some embodiments, the linking moiety binds to Glu. In some embodiments, the linking moiety binds to Gly. In some embodiments, the linking moiety binds to His. In some embodiments, the linking moiety binds to Ile. In some embodiments, the linking moiety binds to Leu. In some embodiments, the linking moiety binds to Lys. In some embodiments, the linking moiety binds to Met. In some embodiments, the linking moiety binds to Phe. In some embodiments, the linking moiety binds to Pro. In some embodiments, the linking moiety binds to Ser. In some embodiments, the linking moiety binds to Thr. In some embodiments, the linking moiety binds to Trp. In some embodiments, the linking moiety binds to Tyr. In some embodiments, the linking moiety binds to Val. In some embodiments, the linking moiety binds to any combination of the aforementioned amino acids. Multiple binding agents can be used, along with various binding agents having linking moieties that bind to distinct amino acids and having different recoding tags corresponding to the distinct amino acids. Multiple binding agents may be used, and the various binding agents have linking moieties that bind to multiple amino acids, groups of amino acids, or that bind slightly preferentially to some amino acids over others. The multiple binding agents can be used with binding agents having a combination of properties including some binding to distinct amino acids and other bindings to groups of amino acids.
[0258] In some embodiments, the binding moiety binds to a dipeptide. In some embodiments, the binding moiety binds to a tripeptide. In some embodiments, the binding moiety binds to a natural amino acid, a post-translationally modified (PTM) amino acid, a derivatized version of an amino acid, a derivatized or stabilized version of a post-translationally modified amino acid, a synthetic amino acid, an amino acid having a specific side chain, an amino acid having a phosphorylated side chain, an amino acid having a glycosylated side chain, an amino acid having a methylation modification, or a D-amino acid. In some embodiments, the binding moiety binds to any combination of the aforementioned amino acids. In some embodiments, the binding moiety binds to a group of amino acids. For example, the binding moiety may bind to a plurality of a number of amino acids, such as all positively charged or phosphorylated PTMS. In some embodiments, the binding moiety is weakly specific for an amino acid or a group of amino acids. For example, in some embodiments, the binding moiety only moderately prefers one amino acid or group of amino acids over another amino acid or group of amino acids. In some embodiments, PTMs such as phosphotyrosine, phosphothreonine, or phosphoserine are recognized. The binding moiety may bind to a phosphorylated amino acid. The binding moiety may bind to a glycosylated amino acid. The binding moiety may bind to a methylated amino acid. The binding moiety may bind to a ubiquitinated amino acid. A plurality of different binding moieties may be used for a plurality of binders, and each binder may include a recoding tag corresponding to each of the plurality of different binding moieties. The binding moiety may bind to a derivatized or stabilized version of an amino acid of other natural or synthetic amino acids, a post-translationally modified amino acid. The binding moiety may bind to an amino acid that has been subjected to sumoylation, prenylation, nitrosylation, sulfation, ADP-ribosylation, palmitoylation, myristoylation, carboxylation, hydroxylation, or other modifications. The binding moiety may bind to a group or class of amino acids having such modifications or similar modifications. For example, the binding moiety may bind to a group of any amino acids having a specific PTM, such as all phosphorylated amino acids.
[0259] Solid support In some embodiments, a solid support is disclosed herein. A peptide may be coupled to the solid support. The chemical reactive conjugate may bind to the solid support. The solid support may be used in the methods described herein, such as methods for determining protein information such as the position or identity of an amino acid.
[0260] In some embodiments, the solid support includes beads, plates, or chips. In some embodiments, the solid support includes glass slides, silica, resin, gel, membrane, polystyrene, metal, nitrocellulose, minerals, plastics, polyacrylamide, latex, or ceramics. In some embodiments, the solid support includes magnetic beads, glass slides, microarray chips, nanoparticles, silica gel, resin, polystyrene beads, gold plates, silicon chips, nitrocellulose membranes, quartz slides, multi-well plates, cellulose paper, agarose beads, plastic beads, polyacrylamide gels, magnetic nanoparticles, latex beads or ceramic beads. In some embodiments, the solid support is included within a flow cell or within a well plate.
[0261] In some embodiments, the solid support is a bead, plate, or chip. In some embodiments, the solid support is a magnetic bead. In some embodiments, the solid support is a glass slide. In some embodiments, the solid support is a microarray chip. In some embodiments, the solid support is a nanoparticle. In some embodiments, the solid support is silica gel. In some embodiments, the solid support is a resin. In some embodiments, the solid support is a polystyrene bead. In some embodiments, the solid support is a gold plate. In some embodiments, the solid support is a silicon chip. In some embodiments, the solid support is a nitrocellulose membrane. In some embodiments, the solid support is a quartz slide. In some embodiments, the solid support is a multiwell plate. In some embodiments, the solid support is cellulose paper. In some embodiments, the solid support is an agarose bead. In some embodiments, the solid support is a plastic bead. In some embodiments, the solid support is a polyacrylamide gel. In some embodiments, the solid support is a magnetic nanoparticle. In some embodiments, the solid support is a latex bead. In some embodiments, the solid support is a ceramic bead. In some embodiments, the solid support is contained within a flow cell. In some embodiments, the solid support is contained within a well plate.
[0262] In some embodiments, the solid support comprises a bead, plate, chip, polymer, metal, or glass. In some embodiments, the solid support is a bead. In some embodiments, the solid support is a plate. In some embodiments, the solid support is a chip. In some embodiments, the solid support is composed of a polymer. In some embodiments, the solid support is composed of a metal. In some embodiments, the solid support is composed of glass.
[0263] Peptide In some embodiments, the peptide is disclosed herein. The peptide can be the subject of a method for obtaining information about the peptide, such as information regarding the identity or position of one or more amino acids of the peptide. The peptide can be included in the methods described herein, such as methods for determining protein information such as amino acid position or identity.
[0264] In some embodiments, the peptide comprises a polypeptide or a protein. In some embodiments, the peptide comprises a hormone, neurotransmitter, enzyme, antibody, viral protein, bacterial protein, synthetic peptide, bioactive peptide, peptide hormone, oligopeptide, polypeptide, fusion protein, cyclic peptide, branched peptide, recombinant protein, tumor marker, therapeutic peptide, antigenic peptide, or signaling peptide.
[0265] In some embodiments, the peptide is a polypeptide or a protein. In some embodiments, the peptide is a hormone. In some embodiments, the peptide is a neurotransmitter. In some embodiments, the peptide is an enzyme. In some embodiments, the peptide is an antibody. In some embodiments, the peptide is a viral protein. In some embodiments, the peptide is a bacterial protein. In some embodiments, the peptide is a synthetic peptide. In some embodiments, the peptide is a bioactive peptide. In some embodiments, the peptide is a peptide hormone. In some embodiments, the peptide is an oligopeptide. In some embodiments, the peptide is a polypeptide. In some embodiments, the peptide is a fusion protein. In some embodiments, the peptide is a cyclic peptide. In some embodiments, the peptide is a branched peptide. In some embodiments, the peptide is a recombinant protein. In some embodiments, the peptide is a tumor marker. In some embodiments, the peptide is a therapeutic peptide. In some embodiments, the peptide is an antigenic peptide. In some embodiments, the peptide is a signaling peptide.
[0266] In some embodiments, peptides coupled to a solid support are disclosed herein. In some embodiments, the peptide is coupled to the solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support. For example, the peptide can be directly coupled to the solid support by a C-terminal amino acid residue or by an internal (e.g., non-N-terminal and non-C-terminal) amino acid residue. In some embodiments, the N-terminus of the peptide is indirectly linked or coupled to the solid support via a chain of other amino acids of the peptide.
[0267] In some embodiments, the peptide is coupled to a solid support such that the N-terminal amino acid residue is exposed to the reaction conditions. For example, the N-terminal amino acid residue can be on the outside of the peptide. In some embodiments, the N-terminal amino acid residue exposed to the reaction conditions is exposed to a solvent.
[0268] In some embodiments, the peptide is derived from a human, plant, bacterium, fungus, animal, virus, mammal, bird, marine organism, insect, reptile, amphibian, synthetic source, protist, yeast, primate, cell culture, parasite, patient sample, environmental sample, or genetically modified organism.
[0269] In some embodiments, the peptide is derived from a cell lysate, blood sample, plasma sample, serum sample, tissue biopsy material, saliva sample, urine sample, cerebrospinal fluid sample, sweat sample, synovial fluid sample, fecal sample, gut microbiota sample, environmental water sample, soil sample, bacterial culture, virus culture, organoid, tumor biopsy material, sputum sample, or hair sample.
[0270] In some embodiments, the peptide is of human origin. In some embodiments, the peptide is of plant origin. In some embodiments, the peptide is of bacterial origin. In some embodiments, the peptide is of fungal origin. In some embodiments, the peptide is of animal origin. In some embodiments, the peptide is of viral origin. In some embodiments, the peptide is of mammalian origin. In some embodiments, the peptide is of avian origin. In some embodiments, the peptide is of marine organism origin. In some embodiments, the peptide is of insect origin. In some embodiments, the peptide is of reptilian origin. In some embodiments, the peptide is of amphibian origin. In some embodiments, the peptide is of synthetic origin. In some embodiments, the peptide is of protist origin. In some embodiments, the peptide is of yeast origin. In some embodiments, the peptide is of primate origin. In some embodiments, the peptide is of cell culture origin. In some embodiments, the peptide is of parasite origin. In some embodiments, the peptide is of patient sample origin. In some embodiments, the peptide is of environmental sample origin. In some embodiments, the peptide is of genetically modified organism origin.
[0271] In some embodiments, the peptide is derived from a cell lysate. In some embodiments, the peptide is derived from a plasma sample. In some embodiments, the peptide is derived from a tissue biopsy material. In some embodiments, the peptide is derived from a serum sample. In some embodiments, the peptide is derived from a saliva sample. In some embodiments, the peptide is derived from a urine sample. In some embodiments, the peptide is derived from a cerebrospinal fluid sample. In some embodiments, the peptide is derived from a sweat sample. In some embodiments, the peptide is derived from a synovial fluid sample. In some embodiments, the peptide is derived from a fecal sample. In some embodiments, the peptide is derived from a gut microbiota sample. In some embodiments, the peptide is derived from an environmental water sample. In some embodiments, the peptide is derived from a soil sample. In some embodiments, the peptide is derived from a bacterial culture. In some embodiments, the peptide is derived from a viral culture. In some embodiments, the peptide is derived from an organoid. In some embodiments, the peptide is derived from a tumor biopsy material. In some embodiments, the peptide is derived from a sputum sample. In some embodiments, the peptide is derived from a hair sample.
[0272] In some embodiments, the peptide is associated with a disease state. In some embodiments, the peptide is associated with a cancerous disease state, an autoimmune disease state, a neurodegenerative disease state, a cardiovascular disease state, a metabolic disease state, a genetic disease state, a viral infection, a bacterial infection, a fungal infection, a parasitic infection, an inflammatory condition, an endocrine disorder, an immunodeficiency, a respiratory disorder, a skin disorder, a gastrointestinal disorder, a psychiatric disorder, an aging process, a muscle disorder, or a kidney disorder.
[0273] In some embodiments, the peptide is associated with a particular disease state. In some embodiments, the peptide is associated with a cancerous disease state. In some embodiments, the peptide is associated with an autoimmune disease state. In some embodiments, the peptide is associated with a neurodegenerative disease state. In some embodiments, the peptide is associated with a cardiovascular disease state. In some embodiments, the peptide is associated with a metabolic disease state. In some embodiments, the peptide is associated with a genetic disease state. In some embodiments, the peptide is associated with a viral infection. In some embodiments, the peptide is associated with a bacterial infection. In some embodiments, the peptide is associated with a fungal infection. In some embodiments, the peptide is associated with a parasitic infection. In some embodiments, the peptide is associated with inflammatory symptoms. In some embodiments, the peptide is associated with an endocrine disorder. In some embodiments, the peptide is associated with immunodeficiency. In some embodiments, the peptide is associated with a respiratory disorder. In some embodiments, the peptide is associated with a skin disorder. In some embodiments, the peptide is associated with a gastrointestinal disorder. In some embodiments, the peptide is associated with a mental disorder. In some embodiments, the peptide is associated with the aging process. In some embodiments, the peptide is associated with a muscle disorder. In some embodiments, the peptide is associated with a kidney disorder.
[0274] In some embodiments, the peptide is a biomarker of a disease or condition, a drug target of a disease or condition, an antigen for the development of a vaccine, used in the manufacture of a biosimilar or generic drug, used to evaluate the efficacy of a drug treatment, used in personalized medicine for a particular disease or condition, used in immuno-oncology research, used in the validation of a diagnostic test, used in the development of a peptide-based therapeutic agent, a therapeutic agent for a disease or condition, used in structure-activity relationship studies, used in the development of immunoassays, used in the study of protein-protein interactions, used in the design of drug delivery systems, used in high-throughput screening assays, used in pharmacokinetic studies, used in the formulation of nutraceuticals, used in the development of probiotic products, or used in proteomics research, and is a component of a cell signaling pathway.
[0275] In some embodiments, the peptide is a biomarker for a disease or condition. In some embodiments, the peptide is a drug target for a specific disease or condition. In some embodiments, the peptide is an antigen for vaccine development. In some embodiments, the peptide is used for patient stratification in clinical trials. In some embodiments, the peptide is a therapeutic agent for a specific disease or condition. In some embodiments, the peptide is used in the manufacture of biosimilars or generic drugs. In some embodiments, the peptide is used to evaluate the effectiveness of drug treatment. In some embodiments, the peptide is used for personalized medicine for a specific disease or condition. In some embodiments, the peptide is used in immuno-oncology research. In some embodiments, the peptide is used for the validation of diagnostic tests. In some embodiments, the peptide is used in the development of peptide-based therapeutic agents. In some embodiments, the peptide is a component of a cell signaling pathway. In some embodiments, the peptide is used in structure-activity relationship studies. In some embodiments, the peptide is used in the development of immunoassays. In some embodiments, the peptide is used in the study of protein-protein interactions. In some embodiments, the peptide is used in the design of drug delivery systems. In some embodiments, the peptide is used in high-throughput screening assays. In some embodiments, the peptide is used in pharmacokinetic studies. In some embodiments, the peptide is used in the formulation of nutraceuticals. In some embodiments, the peptide is used in the development of probiotic products. In some embodiments, the peptide is used in proteomics research.
[0276] Deprotection and Reprotection of Oligonucleotides In some embodiments, methods are disclosed herein that include protection and / or deprotection. For example, some aspects include any or all of the embodiments shown in FIG. 28. Some embodiments include sequential and repeated deprotection and reprotection of oligonucleotides during protein sequencing methods to minimize the impact of peptide cleavage chemical conditions on the molecular structure of the oligonucleotides. Protection, deprotection, or reprotection can be used in the methods described herein, such as methods for determining protein information such as amino acid sequence, identity, or position.
[0277] Some embodiments include methods that involve continuously protecting and deprotecting oligonucleotides. Continuous protection and deprotection can mitigate DNA damage. Some embodiments include methods of cyclically protecting and deprotecting oligonucleotides that are directly or indirectly bound to a solid support in the presence of a peptide that is also directly or indirectly bound to the solid support. This can be useful for mitigating DNA damage during cyclic N-terminal cleavage of the peptide and the biochemistry within each subsequent cycle. Some embodiments include methods of cyclically protecting and deprotecting oligonucleotides in peptide sequencing methods where the nucleic acid is not directly or indirectly bound to a solid support.
[0278] Any or all of the following steps can be included in the peptide sequencing methods described herein: (1) Deprotecting an oligonucleotide related to the identity of a cyclic amino acid or cyclic peptide to enable polymerization, ligation, or DNA manipulation by an enzyme known in the art to modify, extend, amplify, convert, or ligate DNA (2) Reprotecting the oligonucleotide (3) Cleaving the terminal amino acid (4) Repeating.
[0279] Cleavage can be performed using a chemically-reactive conjugate (CRC). In some embodiments, the sequential protection and deprotection of oligonucleotides can be carried out in relation to a protein sequencing protocol, for example, within a protein sequencing method or within a barcoding and / or detection method.
[0280] In some embodiments, the protection and deprotection steps can be repeated. The cycle tags may be deprotected. In some embodiments, the positional oligos can be protected, deprotected, and / or reprotected.
[0281] Oligonucleotides are developed for phosphoramidite oligonucleotide synthesis and can be protected using the protection chemistry utilized during phosphoramidite oligonucleotide synthesis. These protecting groups can withstand the TCA anhydride that is central to the synthesis. For example, N(6)-benzoyl A, N(4)-benzoyl C, and N(2)-isobutyryl G can be employed during DNA synthesis and can be suitable for protection within a protein sequencing method. Also, protecting groups removable under mild alkaline conditions, such as phenoxyacetyl (Pac)-protected dA and 4-isopropyl-phenoxyacetyl (iPr-Pac)-protected dG, may be used together with acetyl-protected dC. As a non-limiting example, the protection of the individual bases A, G, and C can be achieved by acylation reactions with appropriate acid chlorides. The specific acid chlorides used can be benzoyl chloride for adenine and cytosine and isobutyryl chloride for guanine. A solution of benzoyl chloride in a solvent such as dimethylformamide (DMF) and isobutyl chloride in DMF can be prepared and applied to reprotect the oligonucleotide bound to a solid support. In some embodiments, thymine is not protected, but if necessary, it can be protected using, for example, diphenylcarbamoyl chloride.
[0282] In some embodiments, a method is disclosed herein that includes: (a) protecting an oligonucleotide of a binding or reactive molecule; (b) contacting the molecule with the N-terminus of a peptide bound to a solid support; (c) cleaving one or more amino acid residues from the peptide; (d) deprotecting the oligonucleotide of the binding or reactive molecule; and (e) contacting the deprotected oligo with a reagent to convey information by enzymatic ligation, polymerase extension, or chemical ligation. Some embodiments include repeating any of the foregoing steps. The chemically reactive species can include the chemically reactive conjugates described herein.
[0283] In some embodiments, a method is disclosed herein that includes: (a) protecting an oligonucleotide conjugated to a peptide; (b) contacting the N-terminus of the peptide with a reagent to cleave one or more amino acid residues from the peptide; (c) deprotecting the oligonucleotide conjugated to the peptide; and (d) contacting the deprotected oligonucleotide with a reagent to convey information by enzymatic ligation, polymerase extension, or chemical ligation. Some embodiments include repeating any of the foregoing steps. The chemically reactive species can include the chemically reactive conjugates described herein.
[0284] In some embodiments, a method is disclosed herein that includes: (a) protecting an oligonucleotide related to the position or identity of a peptide; (b) contacting the N-terminus of the peptide with a reagent to cleave one or more amino acid residues from the peptide; (c) deprotecting the oligonucleotide conjugated to the peptide; and (d) contacting the deprotected oligonucleotide with a reagent to convey information by enzymatic ligation, polymerase extension, or chemical ligation. Some embodiments include repeating any of the foregoing steps. The chemically reactive species can include the chemically reactive conjugates described herein.
[0285] In some embodiments, methods are disclosed herein that include: (a) protecting an oligonucleotide coupled to a solid support; (b) attaching a chemically reactive species to the terminal amino acid of a peptide coupled to the solid support; (c) deprotecting the oligonucleotide; (d) reacting a reagent with the oligonucleotide; and (e) reprotecting the oligonucleotide. Some embodiments include cleaving the terminal amino acid of the peptide after reprotecting the oligonucleotide. Some embodiments include deprotecting the oligonucleotide after cleaving the terminal amino acid of the peptide and then reacting a second reagent with the oligonucleotide. Some examples include a washing step before or after (a), (b), (c), (d), or (e). Washing can include exchanging solutions and removing excess reagents or solutions. Any of the foregoing steps (e.g., step (e)) or combinations of such steps can be optional in some embodiments.
[0286] In some embodiments, methods are disclosed herein that include: (a) protecting an oligonucleotide coupled to a solid support; (b) cleaving the terminal amino acid of a peptide coupled to the solid support; (c) deprotecting the oligonucleotide; (d) reacting a reagent with the oligonucleotide; and (e) reprotecting the oligonucleotide. Some embodiments include attaching a chemically reactive species to the terminal amino acid of the peptide after reprotecting the oligonucleotide. Some embodiments include deprotecting the oligonucleotide after attaching a chemically reactive species to the terminal amino acid of the peptide and then reacting a second reagent with the oligonucleotide. Some examples include a washing step before or after (a), (b), (c), (d), or (e). Washing can include exchanging solutions and removing excess reagents or solutions. Any of the foregoing steps (e.g., step (e)) or combinations of such steps can be optional in some embodiments.
[0287] Some embodiments relate to a method. The method may include providing a conjugate comprising a reactive molecule coupled to a protected oligonucleotide. The method may include contacting the reactive moiety with the terminal amino acid of a peptide, thereby, for example, coupling the reactive moiety to the terminal amino acid. The method may include cleaving the terminal amino acid from the peptide, if necessary. The method may include deprotecting the oligonucleotide. The method may include contacting the deprotected oligonucleotide with an enzyme or reagent for ligation or polymerization. In some embodiments, a method is disclosed herein that includes providing a conjugate comprising a reactive molecule coupled to a protected oligonucleotide; contacting the reactive moiety with the terminal amino acid of a peptide, thereby coupling the reactive moiety to the terminal amino acid and, if necessary, cleaving the terminal amino acid from the peptide; deprotecting the oligonucleotide; and contacting the deprotected oligonucleotide with an enzyme or reagent for ligation or polymerization. Some embodiments include reprotecting the oligonucleotide. In some embodiments, the reactive moiety cleaves the terminal amino acid from the peptide to expose the next terminal amino acid, and the method further includes contacting the next amino acid with another conjugate after reprotecting the oligonucleotide. In some embodiments, the terminal amino acid is an N-terminus. In some embodiments, the peptide is immobilized on a solid support. In some embodiments, the conjugate comprises an organic small molecule. In some embodiments, the conjugate comprises a chemically reactive conjugate (CRC) comprising (A) an oligonucleotide, (B) a reactive moiety, and (C) an immobilization moiety. In some embodiments, the oligonucleotide comprises a cyclic nucleic acid.
[0288] Some embodiments relate to a method. The method may include providing a conjugate comprising a peptide coupled to a protected oligonucleotide. The method may include contacting a terminal amino acid of the peptide, e.g., thereby attaching a reactive moiety to the terminal amino acid. The method may include cleaving the terminal amino acid from the peptide, if necessary. The method may include deprotecting the oligonucleotide. The method may include contacting the deprotected oligonucleotide with an enzyme or reagent for ligation or polymerization. In some embodiments, a method is disclosed herein that includes providing a conjugate comprising a peptide coupled to a protected oligonucleotide, contacting a terminal amino acid of the peptide, thereby attaching a reactive moiety to the terminal amino acid, cleaving the terminal amino acid from the peptide, if necessary, deprotecting the oligonucleotide, and contacting the deprotected oligonucleotide with an enzyme or reagent for ligation or polymerization. Some embodiments include reprotecting the oligonucleotide. In some embodiments, the reactive moiety cleaves the terminal amino acid from the peptide to expose the next terminal amino acid, and the method further includes contacting the next amino acid with another conjugate after reprotecting the oligonucleotide. In some embodiments, the terminal amino acid is an N-terminus. In some embodiments, the peptide is immobilized on a solid support. In some embodiments, the conjugate includes an organic small molecule.
[0289] Subset sequencing Disclosed herein are methods for sequencing a nucleotide or subset of nucleotides or excluding a nucleotide or subset of nucleotides from a sequencing. Methods for sequencing a subset of nucleotides can be included as part of a method for determining protein information such as amino acid sequence, identity, or position. The method can be useful in a separate method including DNA sequencing. In some embodiments, only a subset of nucleotides is sequenced. In some embodiments, some nucleotides are not sequenced. For example, in some embodiments, only two nucleotides (such as A and C) of a sequence are sequenced and other nucleotides are not sequenced. This can reduce the sequencing cost by reducing the need for sequencing of reagents.
[0290] Subset sequencing can be particularly useful when oligonucleotides function during physicochemical activities such as primers for PCR or spacer oligos and require a function to preserve information. In some embodiments, the nucleotides of a sequence that are functional during the physicochemical activity provide redundant conserved information. Aspects such as barcoded nucleic acids or recoded nucleic acids can include nucleotides such as A, G, C, and T, but the information content of a physiochemically functional sequence can be represented by a subset of nucleotides (e.g., A and C, or T and G). In some embodiments, recoding tags, cycle tags, and / or recoding block nucleic acids include sequences useful to obtain. In some aspects, this information can be obtained by sequencing a subset of nucleotides including the nucleic acid. When an oligonucleotide containing redundant information is sequenced, a subset of nucleotides can be skipped during the sequencing.
[0291] In some embodiments, methods are disclosed herein for sequencing a subset of nucleotides of an oligonucleotide. The method can include (a) providing, in a nucleic acid sequencing reaction, a combination of reversibly terminated nucleotides and non-reversibly terminated nucleotides. In some embodiments, the reversibly terminated nucleotides are fluorescent. In some embodiments, the non-reversibly terminated nucleotides are fluorescent. In some embodiments, the nucleotides of the nucleic acid being sequenced corresponding to the non-reversibly terminated nucleotides are not sequenced. In some embodiments, only a subset of the nucleotides of the nucleic acid is sequenced. In some embodiments, a subset of the nucleotides of the nucleic acid is excluded from the sequencing. The method can include providing, in a nucleic acid sequencing reaction, a combination of reversibly terminated nucleotides and non-reversibly terminated nucleotides, wherein the nucleotides of the nucleic acid being sequenced corresponding to the non-reversibly terminated nucleotides are not sequenced. The method can include identifying the nucleotides of the nucleic acid being sequenced corresponding to the reversibly terminated nucleotides. In some embodiments, the nucleic acid being sequenced includes a region that includes only a subset of nucleotides selected from A, C, G, and T, and the subset of nucleotides is not sequenced. In some embodiments, the subset of nucleotides selected from A, C, G, and T includes two nucleotides selected from A, C, G, and T. In some embodiments, the subset of nucleotides selected from A, C, G, and T includes three nucleotides selected from A, C, G, and T. In some embodiments, the region includes a primer sequence. In some embodiments, the region does not include a barcode sequence, a recoded nucleic acid sequence or a part thereof, or a cyclic nucleic acid sequence or a part thereof. The region not sequenced can include 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 750, 1000 or more nucleotides, or a range of nucleotides defined by any of two or more of the foregoing integers.The portion to be sequenced may include 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 500, 750, 1000 or more nucleotides, or a range of nucleotides defined by any two or more of the aforementioned integers.
[0292] In some embodiments, the subset includes a combination of A, G, C, or T. In some embodiments, the subset of nucleotide components identified by DNA sequencing is two of the four natural nucleotides (e.g., two of A, G, C, and T). The subset can include A and G, A and C, A and T, G and C, G and T, or C and T. In some embodiments, the subset of nucleotides identified by DNA sequencing is A and C.
[0293] In some embodiments, the subset to be sequenced includes all four natural nucleotides, non-natural nucleotides are incorporated and not sequenced, and are skipped by non-reversible terminal nucleotides.
[0294] In some embodiments, the subset of nucleotide components identified by DNA sequencing is three of the four natural nucleotides (e.g., three of A, G, C, and T). The subset can include A, G, and C; A, G, and T; A, C, and T; or G, C, and T. The subset can exclude A, G, and C; A, G, and T; A, C, and T; or G, C, and T.
[0295] A subset of nucleotides can be sequenced by the use of modified nucleotides (e.g., dideoxy (ddNTP) that can be used in Sanger sequencing, etc.). The modified nucleotides can include reversible terminator chemistry. The modified nucleotides can include dyes or tags such as fluorescent dyes or tags. The modified nucleotides can be provided in the sequencing reaction. In some embodiments, other nucleotides not included in the subset are not sequenced (e.g., skipped). Nucleotides not included in the subset can exclude modifications. For example, unmodified nucleotides corresponding to nucleotides that are skipped or not included in the subset can be used in the sequencing reaction mix.
[0296] In some embodiments, a method may include sequencing a subset of nucleotides of an oligonucleotide molecule, the method comprising: (a) providing a solution comprising the oligonucleotide to be sequenced; (b) providing a sequencing reagent comprising one or more nucleotides primarily as reversibly terminated nucleotides and one or more nucleotides primarily as non-terminated nucleotides; (c) preparing (a) for sequencing according to a protocol for a sequencing system; (d) using the sequencing reagent of (b) as at least one component of the sequencing reagent to sequence the prepared solution of (a) for at least one cycle of DNA sequencing; and (e) obtaining the sequence order of a subset of nucleotides in the original oligonucleotide sequence. In some embodiments, the oligonucleotide is designed to contain information regarding the composition of a peptide or amino acids derived from a peptide. In some embodiments, the oligonucleotide is a memory oligo, a recoding tag, a recoding block, or a cycle tag. In some embodiments, the oligonucleotide is derived from a protein sequencing method that creates barcoded nucleic acid information representing a protein sequence and / or protein identity. In some embodiments, the oligonucleotide is any nucleic acid sequence that embodies information regarding the sequence or composition of a peptide or amino acid. In some embodiments, the information of the memory oligo is obtained via DNA sequencing of a subset of nucleotides comprising the memory oligo. In some embodiments, any suitable subset of nucleotides is identified by the DNA sequencing process. In some embodiments, the DNA sequencing method is next-generation sequencing (NGS). In some embodiments, the DNA sequencing is a sequencing-by-synthesis approach using an Illumina sequencer or a PacBio sequencer. In some embodiments, the DNA sequencing is by a ligation approach, a sequence hybridization approach, and / or a ligation-based approach.In some embodiments, the subset of nucleotides identified by DNA sequencing is A and C. In some embodiments, the subset of nucleotide components identified by DNA sequencing is 2 out of the 4 natural nucleotides. In some embodiments, the subset is one of the combinations of A, G, C, or T. Some embodiments include introducing non-fluorescent non-reversible terminating nucleotides into the NGS sequencing reagent mixture. In some embodiments, the nucleotides in the oligonucleotide are natural nucleotides (e.g., A, C, G, and / or T). In some embodiments, the nucleotides in the oligonucleotide include non-natural nucleotides.
[0297] In some embodiments, a method for analyzing one or more peptides from a sample comprising a plurality of peptides, proteins, and / or protein complexes includes: (a) designing an oligonucleotide that includes two, three, four, five, six, or more different types of nucleotide components and uses a subset of the nucleotide components to represent cycle, amino acid, position, and / or protein information; (b) utilizing the physicochemical properties of the oligonucleotide designed within a protein sequencing method as described herein; (c) collecting DNA sequence information of nucleotides that represent protein information; and (d) analyzing the DNA sequence information of a subset of nucleotides to infer protein information. In some embodiments, the oligonucleotide is a memory oligo, a recoding tag, a recoding block, or a cycle tag. In some embodiments, the oligonucleotide is derived from a protein sequencing method that creates barcoded nucleic acid information representing a protein sequence and / or protein identity. In some embodiments, the oligonucleotide is any nucleic acid sequence that embodies information regarding the sequence or composition of a peptide or amino acid. In some embodiments, the information of the memory oligo is obtained via DNA sequencing of a subset of nucleotides that includes the memory oligo. In some embodiments, the DNA sequencing method is NGS. In some embodiments, the DNA sequencing is a synthetic sequencing approach using an Illumina sequencer or a PacBio sequencer. In some embodiments, the DNA sequencing is by a ligation approach, a sequence hybridization approach, and / or a ligation-based approach. In some embodiments, the subset of nucleotides identified by DNA sequencing is A and C. In some embodiments, the subset of nucleotide components identified by DNA sequencing is two of the four natural nucleotides. In some embodiments, the subset is one of the combinations of A, G, C, or T.In some embodiments, any suitable subset of nucleotides is identified by a DNA sequencing process. In some embodiments, the method includes introducing non-fluorescent non-reversible terminator nucleotides into an NGS sequencing reagent mixture.
[0298] In some embodiments, SBS sequencing reagent mixes are disclosed herein. Some aspects include SBS sequencing reagent mixes that include one or more nucleotides primarily as reversibly terminated nucleotides and one or more nucleotides primarily as non-terminated nucleotides.
[0299] Chemically reactive conjugate In some embodiments, a chemically reactive conjugate (CRC) is disclosed herein. The CRC can be used in the methods described herein, such as methods for determining protein information such as amino acid sequences, identities, or positions. The chemically reactive conjugate (CRC) can include a nucleic acid sequence tag. The chemically reactive conjugate can include a reactive moiety. The reactive moiety can bind to a peptide and cleave the N-terminal amino acid residue from the peptide. The chemically reactive conjugate can include an immobilization moiety. The immobilization moiety can bind to a solid support and can thus be useful for immobilization to a solid support. The chemically reactive conjugate can include (A) a cycle tag, (B) a reactive moiety for binding and cleaving the N-terminal amino acid residue from a peptide, and (C) an immobilization moiety for immobilization to a solid support.
[0300] The CRC is:
Chem.
Chem.
Chemical Structure
[0301] The chemically reactive conjugate may include a central moiety. The central moiety may be a central carbon or may include a central carbon. The central carbon may be attached to other carbons, for example, three other carbons, or may be linked to the arm of the chemically reactive conjugate. The central moiety may include a heterocyclic ring, a carbocyclic ring, or trivalent nitrogen. The trivalent nitrogen may include an amine. The amine may include a tertiary amine. The central moiety may be trivalent boron, phosphorus of trivalent or higher valence, tetravalent silicon, polyhedral oligomeric silsesquioxane (POSS), siloxane, branched siloxane, polyether, phosphazene, phosphonium, ammonium, imidazolium, methane, propane, butane, pentane, hexane, C1-C24 alkyl, benzene, toluene, xylene, phenol, N,N-disubstituted aniline, anisole, trihydroxybenzene, benzenetricarboxylic acid, phthalic acid, trimesic acid, cyclopropane, glycol, glycerol, ethylene glycol, oligoethylene glycol, branched oligoethylene glycol, multi-arm oligoethylene glycol, dendrimer, propylene glycol, oligopropylene glycol, trimethylolpropane, pentaerythritol, dipentaerythritol, sugar, glycoside, saccharide, glucose, fructose, furanose, galactose, mannose, cyclohexane, cyclooctane, cycloheptane, cyclopentane, cyclobutene, cyclononane, cyclohexene, cyclobutene, cyclopentene, cyclooctene, cyclononane, adamantane, naphthalene, anthracene, pyrene, annulene, pyridine, N-substituted piperazine, N,It may contain N-disubstituted piperazine, thiophene, indole, pyrazine, isoquinoline, pyran, furan, pyrimidine, purine, oxazole, benzofuran, carbazole, xanthene, coumarin, oxazine, benzothiophene, benzoxazole, acridine, dibenzofuran, fluorene, N-substituted azepine, N-substituted azocine, thiocane, N-substituted azonane, spiro compound, indolizine, benzimidazole, isoindole, azoindole, cyclotrisiloxane, cyclotetrasiloxane, polycyclic aromatic hydrocarbon, alkene, biphenyl, terphenyl, triphenylmethane, decalin, phenanthrene, phosphonate, trisubstituted phosphine, phosphonic acid, phosphite, borate, norbornane, oxanorborene, norbornene, oxanorborene, dioxane, di-tertiary amine, tri-tertiary amine, tetra-tertiary amine, amide, N,N-dialkylamide, sulfonamide, phosphonamide, phthalimide, gallate, ether, thioether, thioamide, mesitylene, carboxylic acid functional molecule, diene, cyanurate, guanidine, urea, substituted urea, thiourea, hydrazone, oxime, dibenzocyclooctene, triazole or ester. The central portion may join the A, B, and C elements of the chemically reactive conjugate.,
[0302] In some embodiments, the chemically reactive conjugate is prepared by organic synthesis methods. Some examples of multi-component reaction schemes are shown in FIGS. 29 to 32B.
[0303] In some embodiments, a chemically reactive conjugate comprising (A) a cycle tag, (B) a reactive moiety, and (C) an immobilization moiety is disclosed herein. In some embodiments, (A), (B), and (C) are linearly oriented relative to each other. In some embodiments, (A), (B), and (C) are oriented in any of the following orders: (A)-(B)-(C) (similar to Formula II), (A)-(C)-(B), or (B)-(A)-(C). In some embodiments, (A), (B), and (C) are linearly similar to Formula II and include any linker between (A), (B), and (C), but in the following order: (A)-(C)-(B). In some embodiments, (A), (B), and (C) are linearly similar to Formula II and include any linker between (A), (B), and (C), but in the following order: (B)-(A)-(C). In some embodiments, each of (A), (B), and (C) is on an independent arm relative to each other.
[0304] In some embodiments, the CRC is linear in the order of (A)-(B)-(C). In some embodiments, the CRC is linear in the order of (A)-(C)-(B). In some embodiments, the CRC is linear in the order of (B)-(A)-(C). In some embodiments, each of the CRCs of (A), (B), and (C) is on an independent arm.
[0305] Some embodiments include cleavable groups between (A) and (B), between (B) and (C), between (A) and (C), between (A) and (C), between (A) and (B + C), between (B) and (A + C), between (C) and (A + B), or any combination thereof. Some embodiments include a cleavable group between (A) and (B). Some embodiments include a cleavable group between (B) and (C). Some embodiments include a cleavable group between (A) and (C). Some embodiments include a cleavable group between (A) and (B + C). Some embodiments include a cleavable group between (C) and (A + C). Some embodiments include a cleavable group between (C) and (A + B).
[0306] Some embodiments include a non-nucleic acid label (e.g., Element A). In some embodiments, the detectable label includes a fluorophore, a radioactive label, an isotope label, a mass tag, a chemiluminescent tag, or an imaging tag. Some embodiments include a detectable label. In some embodiments, the detectable label is a fluorophore. In some embodiments, the detectable label is a radioactive label.
[0307] In some embodiments, the CRC includes a pre-nucleic acid sequence tag that includes a group for attaching a nucleic acid sequence. In some embodiments, the group for attaching a nucleic acid sequence includes an oxyamine group, tetrazine, azide, alkyne, alkene, trans-cyclooctene, DBCO, bicyclononine, norbornene, strained alkyne, strained alkene, or a derivative thereof. In some embodiments, the group for attaching a nucleic acid sequence is subsequently used to attach a nucleic acid sequence. In some embodiments, the nucleic acid sequence tag is generated by conjugating a nucleic acid sequence to a group for attaching a nucleic acid sequence that includes an oxyamine group, tetrazine, azide, alkyne, alkene, trans-cyclooctene, DBCO, bicyclononine, norbornene, strained alkyne, or strained alkene, or a derivative thereof. In some embodiments, the nucleic acid sequence tag is generated by conjugating a nucleic acid sequence to a group for attaching a nucleic acid sequence that includes a protected oxyamine group, a protected thiol, a protected amine, a protected hydrazine, tetrazine, azide, alkyne, alkene, trans-cyclooctene, DBCO, bicyclononine, norbornene, strained alkyne, or strained alkene, or a derivative thereof. In some embodiments, the conjugation is performed prior to the peptide sequencing step. In some embodiments, the conjugation occurs after the CRC has reacted with the N-terminal amino acid. In some embodiments, the conjugation occurs after the CRC has reacted with the N-terminal amino acid and has then been cleaved from the N-terminal amino acid, but before the start of the next cycle.
[0308] In some embodiments, the CRC comprises a pre-reactive moiety that includes a group (e.g., as Element B) for joining the reactive moiety. In some embodiments, the pre-reactive moiety for attaching the reactive moiety includes tetrazine, azide, alkene, alkyne, trans-cyclooctene, DBCO, bicyclononine, norbornene, strained alkyne, strained alkene, or derivatives thereof. In some embodiments, the group for attaching the reactive moiety is then used to couple a reactive moiety for attaching and cleaving the N-terminal amino acid. In some embodiments, the group for attaching the reactive moiety is used to join the CRC to a reactive moiety attached to the N-terminal amino acid.
[0309] Some examples of chemically reactive conjugates are shown in Table 1. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10] [Table 1-11]
Table 1-12
[0310] Cycle tag In some embodiments, a cycle tag is disclosed herein. The cycle tag may be associated with the number of cycles. The number of cycles may correspond to the amino acid number, for example, the amino acid number of a peptide when numbered from N to C. The cycle tag may be part of a chemically reactive conjugate.
[0311] The cycle tag may include a cyclic nucleic acid. In some embodiments, the cyclic nucleic acid includes DNA or RNA. In some embodiments, the cycle tag nucleic acid includes RNA, peptide, synthetic small molecule, or peptide nucleic acid. In some embodiments, the cycle tag is a fluorescent tag.
[0312] In some embodiments, the cycle tag includes a peptide. In some embodiments, the cycle tag includes a peptide nucleic acid. In some embodiments, the cycle tag includes a fluorescent tag. In some embodiments, the cycle tag includes a small molecule. In some embodiments, the cycle tag includes a nucleic acid. In some embodiments, the cycle tag is synthetic.
[0313] In some embodiments, a nucleic acid tag is disclosed herein. The nucleic acid tag may be included within a chemically reactive conjugate. The nucleic acid tag of the chemically reactive conjugate may be referred to or included as an example of a cyclic nucleic acid tag. In some embodiments, the nucleic acid sequence tag includes a DNA or RNA sequence. In some embodiments, the nucleic acid sequence tag includes at least 10 nucleotides. In some embodiments, the nucleic acid sequence tag is ligated or bound to an additional oligonucleotide.
[0314] In some embodiments, the nucleic acid sequence tag is a DNA sequence. In some embodiments, the nucleic acid sequence tag is an RNA sequence. In some embodiments, the nucleic acid sequence tag is a sequence of at least 10 nucleotides. In some embodiments, the nucleic acid sequence tag is a site for ligating or binding additional oligonucleotides and may not contain the nucleic acid itself.
[0315] Reactive moiety In some embodiments, the reactive moiety is disclosed herein. The reactive moiety can be included as part of a chemically reactive conjugate.
[0316] In some embodiments, the reactive moiety includes an Edman degradation reagent. In some embodiments, the reactive moiety includes phenylisothiocyanate (PITC). In some embodiments, the reactive moiety includes isothiocyanate (ITC) or some of its derivatives. In some embodiments, the reactive moiety includes dansyl chloride or some of its derivatives. In some embodiments, the reactive moiety includes dinitrofluorobenzene (DNFB) or some of its derivatives.
[0317] In some embodiments, the reactive moiety comprises an enzyme or a peptide. In some embodiments, the reactive moiety is an enzyme. In some embodiments, the reactive moiety is a peptide. In some embodiments, the reactive moiety specifically cleaves at a particular amino acid. In some embodiments, the reactive moiety specifically cleaves at a particular amino acid that is not the N-terminus. In some embodiments, the reactive moiety specifically cleaves at a particular amino acid that is not an N-terminal acid. In some embodiments, the enzyme or peptide has aminopeptidase activity. In some embodiments, the enzyme or peptide is a modified aminopeptidase. In some embodiments, the reactive moiety cleaves more than one amino acid. In some embodiments, the reactive moiety cleaves 2, 3, 4, 5 or more amino acids. In some embodiments, the reactive moiety cleaves amino acids at a particular motif. In some embodiments, the motif is on the carboxyl side of lysine (K) and arginine (R) amino acid residues, unless the next residue is proline. In some embodiments, the reactive moiety binds to and cleaves the C-terminal amino acid. In some embodiments, the reactive moiety that binds to and cleaves the C-terminal amino acid comprises a modified carboxypeptidase. In some embodiments, the reactive moiety cleaves more than one amino acid. Examples of reactive moieties that can bind to and cleave more than one amino acid can include peptidyl dipeptidase, or a modified peptidyl dipeptidase, such as a modified angiotensin converting enzyme (ACE). The reactive moiety can comprise ACE or a modified ACE.
[0318] Some embodiments include C-terminal peptide cleavage, for example, according to the alkylated thiohydantoin method described by DuPont et al. DuPont DR, Bozzini M, Boyd VL. Alkylated thiohydantoin method for C-terminal sequence analysis. EXS. 2000;88:119-31. https: / / doi.org / 10.1007 / 978-3-0348-8458-7_8. The C-terminal carboxyl can be converted to thiohydantoin by treatment with acetic anhydride followed by thiocyanate ion under acidic conditions. Optionally, the C-terminal can be converted to thiohydantoin via reaction with diphenyl phosphoroisothiocyanatidate (DPP-ITC). Alkylation of thiohydantoin can be achieved by reaction with an alkyl halide functional chemical reactive conjugate under basic conditions, resulting in alkylation at the sulfur of thiohydantoin. This is useful for linking the C-terminal to CRC. Cleavage of the C-terminal amino acid conjugate can be achieved using thiocyanate ion under acidic conditions.
[0319] In some embodiments, the reactive moiety comprises a group on CRC for attachment to a cleavable derivatized N-terminal amino acid, including tetrazine, azide, alkene, alkyne, trans-cyclooctene, DBCO, bicyclononine, norbornene, strained alkyne, or strained alkene, or derivatives thereof.
[0320] Immobilized moiety In some embodiments, an immobilized moiety is disclosed herein. The immobilized moiety can be included as part of a chemical reactive conjugate.
[0321] In some embodiments, the immobilization moiety comprises a thiol group, an amine group, or a carboxyl group. In some embodiments, the immobilization moiety comprises a protected thiol group, a protected amine group, or a carboxyl group, an azide, an alkyne, an alkene, an arylboronic acid, a halogenated aryl, a haloalkyne, a silyl alkyne, an Si-H group, a protected or photo-protected reactive group, or a photo-activated reactive group. In some embodiments, the immobilization moiety comprises an azide, an alkyne, an alkene, an arylboronic acid, a halogenated aryl, a haloalkyne, a silyl alkyne, an Si-H group, a protected or photo-protected reactive group, or a photo-activated reactive group. The immobilization moiety may comprise a thiol. The immobilization moiety may comprise an amine. The immobilization moiety may comprise an alkyne. The immobilization moiety may comprise an azide. The immobilization moiety may comprise an alkene. The immobilization moiety may comprise an arylboronic acid. The immobilization moiety may comprise a halogenated aryl. The immobilization moiety may comprise a haloalkyne. The immobilization moiety may comprise a silyl alkyne. The immobilization moiety may comprise an Si-H group. The immobilization moiety may comprise a protected or photo-protected reactive group (e.g., pyridyldisulfide, phenylacyl-protected thiol, nitrobenzyl-protected thiol, photo-caged DBCO). The immobilization moiety may comprise a photo-activated reactive group (e.g., azirine, tetrazole, sydnone, 3-hydroxynaphthalen-2-ol).
[0322] In some embodiments, the immobilization moiety is a thiol group. In some embodiments, the immobilization moiety is an amine group. In some embodiments, the immobilization moiety is a carboxyl group. In some embodiments, the moiety comprises a protected amine, a protected oxyamine, a protected hydrazine or a blocked isocyanate.
[0323] Linker Any of the components of the CRC may be linked. The linkage may be via a linker. The components may have the same or different linkers. When the CRC comprises the structure of Formula I, L A 、L Bor L C may include a linker. L A may contain a linker. L B may contain a linker. When CRC contains the structure of Formula II, L AB or L BC . L A may contain a linker. L AB may contain a linker. L BC may contain a linker. In some embodiments, CRC is L A , L B and / or L C and contains a linker located at
[0324] In some embodiments, the linker includes polyethylene glycol (PEG), hydrocarbon, ether, carboxyl, amine, amide, azide, thiol, azide-thiol, alkylene, heteroalkylene, cyclic group, phenyl, or a combination thereof. The linker may include polyethylene glycol (PEG). PEG may be PEG 1~20 such as PEG n and may include.
[0325] In some cases, the linker includes alkylene. In some examples, the alkylene is C1-C20 alkylene or a derivative thereof. In some examples, the C1-C20 alkylene may be a substituted variant thereof as needed. In some examples, the alkylene is C1-C10 alkylene or a derivative thereof. In some cases, the linker includes heteroalkylene. In some examples, the heteroalkylene is PEG 1-nincluding, where n is any suitable integer. In some examples, n is an integer from 2 to 100. In some examples, n is an integer from 2 to 50. In some examples, n is an integer from 2 to 25. In some examples, n is in the form of an integer from 2 to 20. In some examples, the heteroalkylene includes PEG1-20 (e.g., polyethylene glycol of 1-20 units) or a derivative thereof. In some examples, PEG1-20 can be its substitution variant as needed. The linker can include oligoethylene glycol, peptide, oligopropylene glycol, oligoamide, oligosaccharide, siloxane, fully alkylated polyamine, polyol, oligomeric polyester, nucleic acid, or oligomeric poly(tetramethylene oxide). In some embodiments, the linker can be modified with, for example, 1 or more heterocycles, carbocycles, thioesters, ethers, thioethers, tertiary amines, amides, carbamates, sulfonamides, dibenzocyclooctene, triazoles, thioamides, oximes, hydrazones, ureas, thioureas, carbonyls (such as esters or amides), or carbonates. The number of PEG units in the PEG linker or carbon atoms in the alkylene linker can be decreased or increased as needed. Changing the number of PEGs or carbon atoms in the linker can change the reaching effect of the chemical reaction arm. For example, a longer PEG arm can be useful to allow for greater flexibility or promiscuity, while a shorter PEG arm can provide greater rigidity or specificity.
[0326] The linker may include -C(O)-, -O-, -S-, -S(O)-, -C(O)O-, -C(O)C1-C10 alkyl, -C(O)C1-C10 alkyl-O-, -C(O)C1-C10 alkyl-CO2-, -C(O)C1-C10 alkyl-S-, -C(O)C1-10 alkyl-NH-C(O)-, -C1-C10 alkyl-, -C1-C10 alkyl-O-, -C1-C10 alkyl-CO2-, -C1-C10 alkyl-S-, -C1-C10 alkyl-NH-C(O)-, -CH2CH2SO2-C1-C10 alkyl-, CH2C(O)-C1-C1-10 alkyl-, =N-(O or N)-C1-C10 alkyl-O-, =N-(O or N)-C1-C10 alkyl-CO2-, =N-(O or N)-C1-C10 alkyl-S-.
Chemical formula
[0327] The linker may be included between the cycle tag and the reactive moiety (e.g., of the linear version of CRC), and the linker may be cleavable. The linker may be included between the cycle tag and the immobilization moiety (e.g., of the linear version of CRC), and the linker may be cleavable. The linker may be included between the reactive moiety and the immobilization moiety (e.g., of the linear version of CRC), and the linker may be cleavable. Any combination of the aforementioned linkers may be used.
[0328] In some embodiments, one or more linkers are cleavable. In some embodiments, one or more cleavable linkers include a disulfide. The linker may include a cleavable moiety. In some aspects, the cleavable moiety is cleaved by light, an enzyme, or a combination thereof. In some aspects, the light includes UV light, visible light, IR light, a laser, or a combination thereof. In some aspects, the cleavable moiety includes a photocleavable moiety. In some aspects, the photocleavable moiety includes an o-nitrobenzyloxy group, an o-nitrobenzylamino group, an o-nitrobenzyl group, an o-nitroveratryl group, a phenacyl group, a p-alkoxyphenacyl group, a benzoin group, or a pivaloyl group. In some aspects, the photocleavable moiety includes an o-nitrobenzyl group. In some aspects, the o-nitrobenzyl group is substituted with a methoxy group or an ethoxy group.
[0329] The cleavable moiety can be cleaved by light, under acidic conditions, under basic conditions, by an enzyme, or a combination thereof. In some cases, the light can include UV light, visible light, IR light, a laser, or a combination thereof. In such cases, the cleavable moiety can be a photocleavable moiety. The photocleavable moiety can include an electron-withdrawing group, such as, but not limited to, a nitro group or a halide group. Alternatively, the cleavable moiety can be an enzymatically cleavable moiety.
[0330] The cleavable moiety can include a pH-sensitive cleavable bond that is cleaved under acidic or basic conditions. In some non-limiting examples, the cleavable moiety can include a pH-sensitive cleavable bond that is cleaved by acidifying the solution. In some non-limiting examples, the cleavable moiety can include a pH-sensitive cleavable bond that is cleaved by making the solution basic. The pH-sensitive cleavable bond is advantageous because it can deliver a molecule but does not react until it is in a slightly acidified environment that can be beneficial for methods of protein sequencing.
[0331] The cleavable moiety may include a disulfide bond. The disulfide bond may be formed chemically or enzymatically. The disulfide bond may be cleaved by a reducing agent. The disulfide bond may be enzymatically cleavable. The cleavable moiety may include a protein or peptide sequence that is recognized and cleaved by an enzyme. For example, the cleavable moiety may include the peptide sequence ENLYFQ*S, where * indicates the cleavage site. The disulfide bond may be included as part of the peptide.
[0332] The enzyme that cleaves the cleavable moiety may include an enzyme that cleaves a disulfide bond. Some examples of enzymes that can cleave a disulfide bond include thioredoxin or glutaredoxin. The enzyme may include trypsin. The enzyme may include a virus that cleaves a specific peptide sequence. For example, a tobacco etch virus (TEV) protein that specifically cleaves the peptide sequence ENLYFQ*S (* represents the cleavage site) may be used. This or another peptide sequence may be present between the central portion and one (or either) of the arms. After ligation and concentration, the bond may be cleaved, thereby releasing the molecule of interest.
[0333] The photocleavable moiety can be cleaved by UV light. The UV light can have a wavelength in the range of about 100 nm to about 400 nm, about 200 nm to about 400 nm, about 250 nm to about 400 nm, about 280 nm to about 400 nm, about 100 nm to about 370 nm, about 200 nm to about 370 nm, about 250 nm to about 370 nm, or about 280 nm to about 370 nm. In some examples, the photocleavable moiety includes a nitrobenzyloxy group, a nitrobenzylamino group, a nitrobenzyl group, a nitroveratryl group, a phenacyl group, an alkoxyphenacyl group, a benzoin group, or a pivaloyl group. In some examples, the nitro group can be ortho to the benzyl, veratryl, phenacyl, benzoin, or pivaloyl group relative to the cleavage site (e.g., o-nitrobenzyloxy group, o-nitrobenzylamino group, o-nitrobenzyl group, o-nitroveratryl group). In some examples, the alkoxy group can be para to the benzyl, veratryl, phenacyl, benzoin, or pivaloyl group relative to the cleavage site (e.g., p-alkoxyphenacyl group). In one aspect, the photocleavable moiety includes a nitrobenzyl group. The nitro group may be ortho to the benzyl group relative to the cleavage site (o-nitrobenzyl group). The o-nitrobenzyl group may be substituted with methoxy or ethoxy. In some cases, the methoxy or ethoxy may be substituted para to the nitro of the o-nitrobenzyl group. In a further example, the o-nitrobenzyl group may include a linker that further connects to a central portion, such as those described herein. The linker may be meta to the nitro group. The linker may include an ester, an ether, an amine, an amide, a carbamate, -O-C1-C10 alkyl-, or any other linker described herein. In some examples, the photocleavable moiety has the formula:
Chemical formula
[0334] L A , L B , LC , L AB or L BC Any or all of the linkers such as etc. may independently include, or be selected from, any of the aforementioned cleavable linkers or non-cleavable linkers, or a combination of a cleavable linker and a non-cleavable linker.
[0335] Kit In some embodiments, kits are disclosed herein. The kit may include any component in this specification or any aspect described. The kit may be useful for analyzing polymeric macromolecules including polymeric macromolecules such as peptides, polypeptides, and proteins.
[0336] Some embodiments include instructions such as instructions for use. For example, the kit may include instructions for use in a method for determining the identity and positional information of amino acid residues of a peptide.
[0337] In some embodiments, the kit includes a chemically reactive conjugate.
[0338] In some embodiments, the kit includes a binder.
[0339] In some embodiments, the kit includes reagents for transferring the information of a recoded nucleic acid to the cyclic nucleic acid of a conjugate complex to generate a recoded block.
[0340] Some embodiments include those for analyzing polymeric macromolecules such as peptide, polypeptide, or protein polymeric macromolecules, and include (a) a nucleic acid sequence tag, and (b) a reactive moiety that couples to the N-terminal amino acid residue of a peptide, thereby forming a conjugate complex that includes a chemically reactive conjugate coupled to the N-terminal amino acid of the peptide, a binding agent that includes a binding moiety for preferentially binding to the conjugate complex, and a recoding tag that includes a recoding nucleic acid corresponding to the binding agent, and a reagent for transmitting information of the recoding nucleic acid to the cyclic nucleic acid of the conjugate complex to generate a recoding block.
[0341] In some embodiments, the kit comprises any or all of the following aspects: (a) a solid support for coupling a peptide to the solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support and is exposed to reaction conditions; (b) one or more reagents having a chemically reactive conjugate, wherein the chemically reactive conjugate comprises (x) a cycle tag comprising a cyclic nucleic acid related to the number of cycles, (y) a reactive moiety for binding to the N-terminal amino acid residue of the peptide, and (z) an immobilization moiety for immobilizing to the solid support, one or more reagents having a chemically reactive conjugate; (c) a reagent for coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex when the peptide is contacted with the chemically reactive conjugate; (d) one or more reagents for immobilizing the complex to the solid support via the immobilization moiety; (e) one or more reagents for cleaving and thereby separating the N-terminal amino acid residue from the peptide, thereby providing an immobilized amino acid complex, the immobilized amino acid complex comprising the cleaved and separated N-terminal amino acid residue, one or more reagents; (f) one or more reagents having one or more binders comprising (i) a binding moiety for preferentially binding to the immobilized amino acid complex and (ii) a recoding tag comprising a recoding nucleic acid corresponding to the binder, wherein when the immobilized amino acid complex contacts the binder, the immobilized amino acid complex and the binder form an affinity complex, the affinity complex comprising the immobilized amino acid complex and the binder, one or more reagents; (g) one or more reagents for transferring the information of the recoding nucleic acid to the cyclic nucleic acid of the immobilized conjugate complex to generate a recoding block; one or more reagents for joining two or members of a plurality of recoding blocks to form a memory oligonucleotide; and / or (j) one or more sequencing reagents for obtaining the sequence information of the recoding block.
[0342] The kit can be used to sequence a subset of nucleotides of an oligonucleotide and can include one or more reagents for sequencing a subset of nucleotides of an oligonucleotide. Some embodiments include an SBS sequencing reagent mix that includes one or more nucleotides as primarily reversibly terminated nucleotides and one or more nucleotides as primarily non-terminated nucleotides.
[0343] The kit can include any of the reagents or embodiments described herein.
[0344] Definitions As used in this disclosure, the term "amino acid" and the notation "AA" refer to natural d-, l-, unnatural, and post-translationally modified amino acids. "N-terminal amino acid" refers to an amino acid that has a free amine group and is linked only to one other amino acid of a peptide via an amide bond. Similarly, "C-terminal amino acid" refers to an amino acid that has a free carboxyl group and is linked only to one other amino acid of a peptide via an amide bond.
[0345] The term "AA tag" refers to a nucleic acid molecule of any length, typically in the range of 5-20 bases, that includes a sequence defined to represent a particular amino acid or class of amino acids that share structural or functional similarity. When recoding a polymer that does not contain amino acids, the AA tag sequence can be defined to represent a particular monomer or class of monomers that share structural or functional similarity. It may also refer to any construct that enables subsequent identification methods of cycle information such as mass tags.
[0346] The terms "analyze" and "analyzing" refer to assigning a sequence, and / or quantification, and / or identity to a portion of a macromolecule or macromolecular analyte.
[0347] The term "assembly" (e.g., assemblyOligo) refers to a nucleic acid that can hybridize to a memory oligo tethered to a solid support and / or hydrogel. Assembly oligos can be used to facilitate ligation assembly of complementary DNA strands to memory oligos tethered to the hydrogel surface and / or solid support as a template. Ligation assembly of complementary strands avoids the need for polymerase extension through the tethered nucleic acid to create a solution-phase nucleic acid representing the analyte sequence. Assembly oligos contain sequences complementary to cycle tag sequences and sequences complementary to amino acid sequences.
[0348] The term "binding agent" refers to an entity consisting of a binding moiety conjugated to a recode tag. The binding moiety and the recode tag can be conjugated by a linker.
[0349] The term "binding moiety" refers to a molecule or macromolecule that recognizes and binds to a target analyte or a characteristic of a target analyte. Exemplary binding moieties include antibodies, F(ab’)2, Fab and scFv regions, nanobodies, DNA aptamers, RNA aptamers, modified aptamers, photoactive or non-photoactive cage compounds, oligopeptide permeases (Opp), amino-acyl t-RNA synthetases (aaRS), periplasmic binding proteins (PBP), dipeptide permeases (Dpp), proton-dependent oligopeptide transporters (POT), modified aminopeptidases, modified aminoacyl tRNA synthetases, modified anticalins, modified ClpS, lectins, or clathrates. The binding moiety can form a covalent or non-covalent association with the target analyte, and the target analyte includes an immobilized conjugate complex, e.g., an immobilized PTC-AA-cycle tag conjugate complex. The binding moiety can exhibit preferential binding to one conjugate complex over another depending on the amino acids of the complex. The binding moiety can preferentially bind to a class of amino acids that are structurally or functionally similar within the conjugate complex.
[0350] In addition to caged drugs and bioactive small molecules, amino acids and derivatized amino acids offer many possibilities for caging. For example, amines, carboxylates, and amino acid side chains provide several functional groups that are readily caged.More specifically, caged serine, threonine, tyrosine, cysteine, methionine, aspartic acid, glutamic acid, and lysine have all been reported (see Pirrung et al., Synthesis of photodeprotectable serine derivatives-caged serine, Bioorg. Med. Chem. Lett. 2, 1489-1492 (1992); Tatsu et al., Solid-phase synthesis of caged peptides using tyrosine modified with a photocleavable protecting group, Biochem. Biophys. Res. Comm. 227, 688-693 (1996); Gee, K.R., Carpenter, B.K., and Hess, G.P., Synthesis, photochemistry, and biological characterization of photolabile protecting groups for carboxylic acids and neurotransmitters, Met. Enz. 291, 30-50 (1998); Tatsu et al., Synthesis of caged peptides using caged lysine: Application to the synthesis of caged AIP, a highly specific inhibitor of calmodulin-dependent protein kinase II, Bioorg. Med. Chem. Lett. 9, 1093-1096 (1999); Okuno, T., Hirota, S., and Yamauchi, Ol., Folding character of cytochrome c studied by onitrobenzyl modification of methionine 65 and subsequent ultraviolet light irradiation, Biochem. 39, 7538-7545 (2000)).
[0351] The terms "biochip" and "microarray" refer to consumable devices that assist in fluid handling and further support recoding workflows. In some embodiments, these may include flow cells that are directly used by NGS sequencing devices in the DNA sequencing process.
[0352] The term "biologically or synthetically derived sample" refers to a sample of macromolecules having its origin from a biological process such as a cell lysate solution, or having its origin from a sample produced using synthetic biology techniques, or a sample of macromolecules produced purely using chemical synthesis, such as a solution of synthetic peptides, synthetic nucleic acids, or chemically synthesized polymers.
[0353] The term "chemically reactive conjugate" refers to a conjugate comprising (a) a reactive moiety that can bind to and cleave a terminal amino acid, (b) a reactive moiety that enables immobilization to a solid support, and (c) a cycle tag having identification information regarding the cycles of the workflow.
[0354] The term "codespace" refers to a region of code that is associated with cycle tags and AA tags and is used to represent the identification information of the cycles and monomers of the workflow, respectively. The codespace provides a practical separation distance between codes and is defined by a set of rules that improve fidelity and accuracy while reading information. For example, applying Hamming distance theory or other up-to-date digital code space theories (e.g., Lee, Levenshtein-Tenengolts, Reed-Solomon, or others) to assign codes enables error detection and error correction capabilities, and can account for a combination of errors that may occur during 1) NGS sequencing errors during analysis, 2) errors in oligonucleotide synthesis, 3) errors in reagents used in the recoding process, 3) errors occurring during the assembly of recoding blocks, 4) errors occurring during the assembly of memory oligos, or any step in the determination of protein sequence and protein abundance by recoding an amino acid polymer into a DNA polymer and analyzing it.
[0355] The term "isotopic binder" refers to a binder designed to bind with high relative affinity to an isotopic target analyte or a feature or part of an isotopic target analyte. This is not designed to bind to a non-isotopic target analyte or a feature or part of a non-isotopic target analyte, and thus is in contrast to a "non-isotopic binder" that interacts with a non-isotopic target analyte or a feature or part of a non-isotopic target analyte with low relative affinity. As a result, the non-isotopic binder does not effectively transfer re-encoding tag information to the re-encoding block under conditions appropriate for re-encoding block assembly by the isotopic binder.
[0356] The terms "conjugate complex" and "immobilized conjugate complex" refer to a chemically reactive conjugate that is joined, as needed in context, to an amino acid (e.g., a monomer of a macromolecular analyte), a peptide, a linker, a solid support, and / or a cycle tag.
[0357] The term "complementary" refers to Watson-Crick base pairing between nucleotides, specifically to nucleotide hydrogens that are joined to each other by a thymine or uracil residue linked to an adenine residue by two hydrogen bonds and a cytosine and guanine residue linked by three hydrogen bonds. Generally, a nucleic acid contains a nucleotide sequence that is described as having a "percent complementarity" or "percent homology" to a designated second nucleotide sequence. For example, a nucleotide sequence can have 80%, 90% or 100% complementarity to a designated second nucleotide sequence, which indicates that 8 out of 10 nucleotides, 9 out of 10 nucleotides or 10 out of 10 nucleotides of the sequence are complementary to the designated second nucleotide sequence.
[0358] The term "cycle tag" (e.g., "cycleTag") refers to a nucleic acid molecule of any length, typically in the range of 5 to 20 bases, having an array defined to represent a specific cycle of a recoding workflow. The length of the cycle tag can vary for different cycles of the workflow. The cycle tag can optionally include additional nucleic acid sequences that direct the assembly of memory oligos in a later step, such as a universal assembly sequence that facilitates recoding block assembly regardless of the order of assembly. In a specific example, the cycle tag can optionally include a restriction endonuclease sequence. The term "cycle tag" may also refer to any construct that enables subsequent identification of cycle information such as mass tags.
[0359] The term "deprotection" refers to the removal of a protecting moiety that maintains the integrity of a functional group during conditions that would otherwise react to change the functional group and exposure to potential reactants. Exemplary protecting agents for nucleic acids include FMOC, acetyl (Ac), benzoyl (Bz), dimethylformamidine (DMFA), and phenoxyacetyl (PAC). See Radhakrishnan P. Iyer, Current Protocols in Nucleic Acid Chemistry.
[0360] The terms "homology" or "identity" or "similarity" refer to sequence similarity between two peptides or between two nucleic acid molecules.
[0361] The term "hydrogel" refers to synthetic polymers, natural polymers, and / or hybrid polymers. Exemplary monomers that can form hydrogels include acrylamide, acrylate, vinyl pyridine, dihydroxy methacrylate, other methacrylates, HEMA, PHEMA, PVA, HPMC, PLGA, PEG, etc., having linear, branched and crosslinked configurations, block copolymer configurations, or other configurations that promote sequence-defined macromolecules. See Faisal Raza, Hajra Zafar, Ying Zhu, Yuan Ren, Aftab-Ullah, Asif Ullah Khan, Xinyi He, Han Han, Md Aquib, Kofi Oti Boakye-Yiadom and Liang Ge, A Review on Recent Advances in Stabilizing Peptides / Proteins upon Fabrication in Hydrogels from Biodegradable Polymers, Pharmaceutics 2018, 10, 16. Hydrogels can associate with a solid support via covalent or non-covalent interactions. Hydrogels can further include orthogonal conjugation chemistry modalities to assist in the recoding workflow.
[0362] Terms such as "the i-th", "(i-1)-th", etc. refer to any position in the polymeric analyte and its immediate vicinity.
[0363] The term "ligation oligo" (e.g., "ligationOligo") refers to a nucleic acid that becomes ligated to the cycle tag of an immobilized conjugate complex when appropriately directed by a cognate binder via hybridization to the recode tag of the cognate binder. In certain embodiments, the ligation oligo may carry information regarding amino acids and workflow cycle assembly and be complementary to the recode tag of the cognate binder. It is also recognized that the ligation oligo may be in another molecular format that is not a nucleic acid and may recode cycle information of amino acids and workflows that can be joined to the cycle tag via a chemical reaction. In certain embodiments, the ligation oligo may optionally include sequences that facilitate ligation of a recode block to another recode block, extension: ligation, or chemical ligation, regardless of the order of assembly. For example, by including 3' and / or 5' universal assembly sequences in multiple recode blocks such that at least two recode blocks share the same universal assembly sequence, assembly into a memory oligo in any given order of such recode blocks becomes possible.
[0364] The term "linker" or "spacer" refers to a molecule used to join two or more molecules. The composition of the molecule may be a polymer, a monomer, or a combination of both. The linker may further include reactive elements that facilitate covalent and / or non-covalent conjugation between molecules. Exemplary linkers include those used to join a binder to a recode tag, or a cycle tag to another element of the conjugate complex, such as a molecule having an NHS-ester at one end and an azide at the other end of a PEG molecule, or a molecule having a biotin at one end and a maleimide moiety at the other end of a nucleic acid.
[0365] The term "linking oligo" (e.g., "linkingOligo") refers to a nucleic acid that can facilitate ligation between a recoding block associated with a given workflow cycle and a second recoding block associated with any other workflow cycle of the recoding process. Linking oligos can replace, for example, errors in upstream processes that have resulted in incomplete or unexpected recoding block sequences in one or more workflow cycles, the absence of recoding block assembly in one or more workflow cycles, or steric effects that prevent interactions between recoding blocks and the assembly of recoding blocks, and are useful for completing the assembly of memory oligos. Linking oligos can optionally include sequences complementary to the cycle tag sequence of one workflow cycle and the cycle tag sequence of any other workflow cycle. Ligation of recoding blocks via linking oligos can result in a lack of information regarding recoding blocks skipped in the assembly of memory oligos. In this case, it is recognized that memory oligos can still be valuable for the analysis of macromolecular information because information can be inferred during analysis that an unknown (or multiple unknown) monomer separates the positions of known monomers, and mapping to a reference sequence enables macromolecular sequence and identity information. In certain embodiments, linking oligos can optionally include sequences for facilitating ligation between a recoding block associated with a workflow cycle and a second recoding block associated with another workflow cycle of the recoding process. For example, such ligation can be facilitated through complementarity between cycle tags and / or universal assembly sequences of recoding tags.
[0366] The term "position linker" refers to any molecule configured to attach a peptide to a solid support and further configured to bind to a nucleic acid. In some examples, a position linker refers to a molecule having three or more functional elements that facilitate the attachment of a peptide, nucleic acid, and solid support. In some examples, the nucleic acid can be a UMI that carries code information regarding the isolation position of an isolated immobilized PTC conjugate.
[0367] The term "position oligo" (e.g., "locationOligo") refers to a nucleic acid of any suitable length, but typically in the range of 10 to 40 bases, that contains a sequence representing the x, y, and z coordinates of an immobilized macromolecular analyte and is held in proximity to the macromolecule via a position linker. Position oligos are useful for transmitting position information to spatially adjacent immobilized recoding blocks.
[0368] The terms "macromolecule" and "macromolecular polymer" refer to high molecular weight molecules composed of subunits. Examples of macromolecules include protein complexes such as photosynthetic reaction center antenna complexes, multi-subunit proteins such as photosynthetic reaction centers or pore proteins, single-subunit proteins such as cytochrome-c, protein fragments, peptides, polypeptides, nucleic acids, carbohydrates, and polymers such as urethane or acrylamide, but are not limited thereto. "Macromolecule" also describes natural and synthetic combinations of two or more macromolecular types such as peptides covalently bound to nucleic acids or lectins bound to carbohydrates via electrostatic, van der Waals forces, or any non-covalent binding force.
[0369] The term "memoryOligo" (e.g., "memoryOligo") refers to a construct that includes positional information, monomer relative position information, and / or monomer identity information. This is typically assembled by aggregating information from recoding blocks. Typically, a memoryOligo contains information about one related macromolecular analyte. However, it is recognized that there are embodiments where a memoryOligo contains identification information for one or more macromolecular analytes. Optionally, a memoryOligo may further include a sample index, UMI, universal priming site, linker, and other identifiers of macromolecular origin. The length of a memoryOligo is typically 25 to 25,000 base pairs. When fully assembled, the length of a memoryOligo is equal to the sum of the length of the provenance identifier, the cycle tag, and the AA tag sequence, multiplied by the number of cycles in the workflow. It is recognized that the length of the cycle tag may vary for each cycle of the workflow. Incomplete assembly of the recoding blocks results in a length that is shorter or longer than that of a fully assembled memoryOligo, and since cycle and amino acid (e.g., monomer) information is transmitted to adjacent registers of the memoryOligo, it should be noted that useful memoryOligos for macromolecular analysis can be generated. Further, it is recognized that a sequential assembly that recodes recoding block information into a memoryOligo is not necessary to provide an analytical memoryOligo useful for macromolecular analyte analysis.
[0370] The term "n" refers to the length of the target macromolecular analyte or the number of cycles in the workflow. It can also refer to the terminal subunit of the macromolecular analyte, e.g., the nth subunit. Thus, the next subunit is shown as n-1, then n-2, and so on for the length of the peptide. These labels can be assigned starting from the N-terminus or C-terminus of the macromolecule.
[0371] The terms "n-1", "n-2", etc. refer to cycles such as the cycle before the last cycle. It can also refer to the subunits of the macromolecular analyte that are closest and next closest to the terminal subunit.
[0372] The terms "polynucleic acid" or "polynucleotide" refer to a polymer of deoxyribonucleotides linked by 3'-5' phosphodiester bonds. This includes polymers having nucleotide analogs and unnatural nucleotides, such as Iso-G and Iso-C. This also includes nucleotides linked by phosphorothioate bonds or peptidyl bonds, such as PNA. This also encompasses RNA and polymers having modified ribose moieties, such as LNA, XNA, or BNA.
[0373] The terms "nucleic acid sequencing", "NGS", or "next-generation sequencing" refer to high-throughput methods for determining the sequence of a nucleic acid polymer. These methods are exemplified by products commercially available from Illumina, Pacific Biosciences, and Oxford Nanopore.
[0374] The terms "peptide" or "polypeptide" refer to a chain of two or more amino acids, and the distinction with respect to length is not implied by the terms peptide, polypeptide, or protein. Similarly, no distinction or limitation is implied with respect to l-, d-, unnatural, or post-translationally modified amino acid monomers that include peptides.
[0375] The term "PITC conjugate" refers to a chemically reactive conjugate that has not reacted with an amino acid or a solid support. The modifier "PITC" is recognized as a representative term for describing any number of molecules (or sets of molecules) that can similarly function to bind to the N-terminal or C-terminal amino acid and cleave the terminal subunit.
[0376] The terms "conjugate complex", "PTC-conjugate", and "PTC-AA-cycle tag conjugate complex" refer to chemically reactive conjugates that have reacted with amino acids but are not necessarily immobilized on a solid support. The modifier "PTC" is recognized as a representative term for describing any number of alternative molecules (or sets of molecules) that can similarly function to bind to the N-terminal or C-terminal amino acid and cleave the terminal subunit. The terms "immobilized conjugate complex", "immobilized PTC-conjugate", and "immobilized PTC-AA-cycle tag conjugate complex" refer to chemically reactive conjugates that have reacted with amino acids immobilized on a solid support. The modifier "PTC" is recognized as a representative term for describing any number of alternative molecules (or sets of molecules) that can similarly function to bind to the N-terminal or C-terminal amino acid and cleave the terminal subunit.
[0377] The term "post-translational modification" refers to any modification, either biologically or synthetically, of l-, d-, or unnatural amino acids. Modifications can occur at the terminal amine, terminal carboxyl, or any reactive moiety of the peptide. Examples include, but are not limited to, phosphorylation, glycosylation, glycanation, methylation, acetylation, ubiquitination, carboxylation, hydroxylation, biotinylation, pegylation, and succinylation. Further information regarding post-translational modifications can be found in DOI: 10.1021 / acs.biochem.7b00861 Biochemistry 2018, 57, 177 - 185, which is hereby incorporated by reference in its entirety.
[0378] The term "recode block" (e.g., "recodeBlock") refers to a construct created by the interaction between the cycle tag of an immobilized conjugate complex and the recode tag of a cognate binder. Typically, a recode block is a chimeric nucleic acid molecule that includes information regarding the cycle of a workflow and the composition of an amino acid or class of amino acids that includes the conjugate complex. Further, the recode block holds information for instructing the assembly of memory oligos and / or amplifies the recode block. A recode block can be formed by utilizing an extension-ligation method to transfer information from a recode tag to a recode block or via a ligation reaction under appropriate conditions in the presence of a ligase and ligation oligos. The format of a recode block does not necessarily have to be nucleic acid. It may also take the form of a mass tag that can be used to assign the identity of the cognate conjugate complex and cycle of amino acids, or other modalities that represent information of the immobilized conjugate complex and are suitable for grouping that information for analysis.
[0379] The term "recode tag" (e.g., "recodeTag") refers to a nucleic acid molecule of any length, typically in the range of 15 to 60 bases, having an array consisting of the i-th cycle tag complement, the AA tag complement, and the (i - 1)-th cycle tag complement. It provides for specifying amino acid (or monomer subunit) information about its associated binder. It may uniquely specify one amino acid or may specify a class of amino acids having structural and / or functional similarity. The recode tag provides a probabilistic estimate regarding the identity of the amino acid component of the immobilized PTC-AA cycle tag conjugate complex, thereby providing sufficient information for analysis. In certain embodiments, the recode tag may optionally include the i-th cycle tag complement, the AA tag complement, and / or a universal assembly sequence or the complement of a universal assembly sequence, which aids in the assembly of the memory oligo. In certain embodiments, the recode tag may optionally include universal assembly sequences at both the 3' and 5' ends to facilitate memory oligo assembly, regardless of the assembly order of the components of the recode block. In a further embodiment, the recode tag may include a sequence that facilitates amplification of the recode block.
[0380] The term "sample index" refers to an identifier that is incorporated during the preparation after recoding of a DNA library for NGS analysis, or that can be ligated as a component of a memory oligo during its assembly and can be used during NGS analysis to identify the origin of the oligonucleotides in the DNA library.
[0381] The term "solid support" or "surface" refers to any solid material substrate in a planar form, spherical form, or combination of forms, including but not limited to solid beads, porous beads, solid planar materials, porous planar materials, patterned or unpatterned solid materials, nanoparticles, or inorganic or polymeric microspheres, or capillaries. For example, the solid support can include a glass slide or wafer, silicon slide or wafer, PC, PTC, polyethylene (PE), high-density polyethylene (HDPE), or other plastic slides, Teflon®, nylon, nitrocellulose membrane, or borosilicate capillary. Particles and beads can be formed from polystyrene, cross-linked polystyrene, agarose, or acrylamide. Beads or nanoparticles can be magnetic or paramagnetic to assist in a separation or purification process. The solid support may be passivated with glass, silicon oxide, tantalum pentoxide, DLC diamond-like carbon, or other passivating agents. A "solid support" including a membrane can be passivated or activated by corona or other plasma treatment methods. The solid support may be further assembled with other components (e.g., flow cell, biochip, microtiter plate) to facilitate fluid transport and / or detection. The solid support can include an associated hydrogel that aids in the joining of components for macromolecule recoding and / or analysis workflows. In certain examples, the term "solid support" can include any of the above solid supports further associated with a hydrogel.
[0382] The term "sprint" refers to a nucleic acid having complementarity to the 5' end of one nucleic acid and the 3' end of another nucleic acid, such that hybridization of the sprint to both nucleic acids brings the 5' and 3' ends into proximity to facilitate either chemical or biological ligation.
[0383] The term "strobe array determination" refers to a method of sequencing (e.g., nucleic acids, peptides, and other polymers) in which reads with short gaps or scattered subreads are generated from contiguous fragments rather than a single contiguous read. Such subreads are referred to as "strobes" or "strobe" reads.
[0384] As used in this disclosure, the term "unique molecular identifier" or "UMI" refers to a nucleic acid molecule 10 to 40 bases in length that can be assembled, for example, into memory oligos and provides unique identification for in silico deconvolution of NGS sequencing data for a particular memory oligo.
[0385] The term "universal priming site" or "universal primer" refers to a nucleic acid molecule that can be used during library amplification and / or NGS. Exemplary universal priming sequences can include P5, P7, P5’, P7’, SBS Read 1, and SBS Read 2 primers.
[0386] The term "universal sequence" or "universal assembly sequence" or "universal amplification sequence" refers to a common complementary polynucleotide sequence that can be added to the 3’ and / or 5’ ends of tags, such as recoding tags, to facilitate amplification by common primers or assembly into oligos, such as memory oligos. In certain embodiments, the universal sequence includes repetitive sequences, such as dinucleotide repeat sequences n such as (GT), or other relatively short nucleotide motifs. The universal sequence is silent during sequencing of the oligo and can facilitate efficient detection and analysis of the assembled components of the oligo.
[0387] The term "workflow cycle" or "cycle" refers to the number of repetitions of any one of the operations of the process flow or method described herein.
[0388] Several references to oligonucleotides may be used, or may be included in or named in Table 2. [Table 2-1] [Table 2-2]
[0389] It should be noted that, as used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, a reference to "an oligo" refers to one or more oligos and the like. Further, terms such as "left", "right", "top", "bottom", "front", "rear", "side", "height", "length", "width", "upper", "lower", "inner", "outer", "inside", "outside", etc., as used herein, merely describe a reference point and are not necessarily intended to limit the embodiments of the present disclosure to a particular orientation or configuration. Further, terms such as "first", "second", "third", etc., merely identify one of a number of parts, components, steps, operations, functions, and / or reference points disclosed herein, and likewise are not necessarily intended to limit the embodiments of the present disclosure to a particular configuration or orientation.
[0390] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. All publications mentioned herein are incorporated by reference for the purpose of describing and disclosing devices, methods, and cell populations that may be used in connection with the disclosure described herein.
[0391] When a range of values is provided, it is understood that each intervening value between the upper and lower limits of that range, and any other stated value or intervening value within the stated range, is encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are encompassed within the disclosure subject to any specifically excluded limits within the stated range. When the stated range includes one or both of the limits, ranges excluding one or both of those included limits are also included in the disclosure.
[0392] In this description, numerous specific details are set forth in order to provide a more thorough understanding of the disclosure. However, it will be apparent to one of ordinary skill in the art that the disclosure may be practiced without one or more of these specific details. In other instances, well-known features and procedures have not been described in order to avoid obscuring the disclosure.
[0393] The functions described in connection with one embodiment are intended to be applicable to additional embodiments described herein, unless explicitly stated otherwise, or the feature or function is incompatible with an additional embodiment. For example, if a given feature or function is explicitly described in connection with one embodiment but not explicitly mentioned in connection with an alternative embodiment, it is understood that the feature or function may be developed, utilized, or implemented in connection with the alternative embodiment, so long as the feature or function is not incompatible with the alternative embodiment.
[0394] The practice of the technology described herein, unless otherwise specified, can use conventional techniques and descriptions of organic chemistry, polymer technology, molecular biology (including recombinant techniques), cell biology, proteomics, biochemistry, and sequencing technology, which are within the skill of those in the art. Such conventional techniques include polymer array synthesis, hybridization and ligation of polynucleotides and other polymers, and detection of hybridization using labels. Specific illustrations of appropriate techniques can be obtained by referring to the examples in this specification. However, of course, other equivalent conventional procedures can also be used.Such conventional techniques and descriptions can be found in standard laboratory manuals such as Green et al., Eds. (1999), Genome Analysis: A Laboratory Manual Series (Vols. I-IV); Weiner, Gabriel, Stephens, Eds. (2007), Genetic Variation: A Laboratory Manual; Dieffenbach, Dveksler, Eds. (2003), PCR Primer: A Laboratory Manual; Mount (2004), Bioinformatics: Sequence and Genome Analysis; Sambrook and Russell (2006), Condensed Protocols from Molecular Cloning: A Laboratory Manual; and Sambrook and Russell (2002), Molecular Cloning: A Laboratory Manual (all from Cold Spring Harbor Laboratory Press); Stryer, L. (1995) Biochemistry (4th Ed.) W.H. Freeman, New York N.Y.; Gait, ‘‘Oligonucleotide Synthesis: A Practical Approach’’ 1984, IRL Press, London; Nelson and Cox (2000), Lehninger, Principles of Biochemistry 3rd Ed., W.H. Freeman Pub., New York, N.Y.; Berg et al. (2002) Biochemistry, 5th Ed., W.H. Freeman Pub., New York, N.Y.; etc., all of which are hereby incorporated by reference in their entirety for all purposes.
Example
[0395] The following examples are described to provide a complete disclosure and description of how to make and use the present disclosure to one of ordinary skill in the art, and are not intended to limit the scope that the inventors regard as their disclosure, nor are they intended to represent or imply that the following experiments are all or the only experiments that have been conducted. One of ordinary skill in the art will understand that numerous variations and / or modifications can be made to the present disclosure as shown in the specific embodiments without departing from the spirit or scope of the present disclosure as broadly described. Accordingly, this embodiment should be considered exemplary in all respects and not limiting.
[0396] Example 1: Characterization and Verification of the Trifunctional Chemical Reactive Conjugate (TCRC) Function A chemical reactive conjugate (CRC) can include (x) a cycle tag (or a moiety for covalently attaching to a cycle tag such as an aminooxy group in this example), (y) a reactive moiety (such as PITC in this example) for binding and cleaving the N-terminal amino acid residue of a peptide and exposing the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide, and (z) an immobilization moiety (such as propargyl in this example) for immobilizing on a solid support. The ability to synthesize a trifunctional molecule, bind its reactive moiety to the N-terminal amino acid of an immobilized peptide, cleave the N-terminal amino acid, hybridize to the cycle tag, ligate the cycle tag, and bind the CRC to a solid support via the immobilization moiety was demonstrated using PPO (an example of a CRC compound as shown in FIG. 32B).
[0397] Accordingly, the examples shown to be functional here are exemplary CRCs shown to be functional here. PPO (propargyl-PITC-oligo): 1-(1-deoxyribonucleotide-indol-3-yl)-N-(12-(4-(3-(4-isothiocyanatophenyl)-3,9-dihydro-8H-dibenzo[b,f][1,2,3]triazolo[4,5-d]azocine-8-yl)-4-oxobutanoyl)-3,6,9,15,18-pentaoxa-12-azahenicos-20-in-1-yl)-3,6,9,12,15-pentaoxa-2-azaoctadec-1-en-18-amide.
[0398] The chemical names of the intermediates that can be formed during synthesis, such as the synthesis shown in FIGS. 32A to 32B, can be as follows. · PDA: N-(propargyl-PEG2)-DBCO-PEG3-amine (Broadpharm catalog number 29932) · PDON: N-(12-(4-(11,12-didehydrodibenzo[b,f]azocine-5(6H)-yl)-4-oxobutanoyl)-3,6,9,15,18-pentaoxa-12-azahenicos-20-in-1-yl)-2,5,8,11,14-pentaoxa-1-azapentadecan-17-amide · PDON-tBOC: PDON tert-butyloxycarbonyl · PDO: 1-(1-deoxyribonucleotide-indol-3-yl)-N-(12-(4-(11,12-didehydrodibenzo[b,f]azocine-5(6H)-yl)-4-oxobutanoyl)-3,6,9,15,18-pentaoxa-12-azahenicos-20-in-1-yl)-3,6,9,12,15-pentaoxa-2-azaoctadec-1-en-18-amide
[0399] As a preliminary test, in Figure 26, a model trifunctional molecule of approximately 1 kd containing vanillin instead of oligonucleotide was generated using the following: (1) base structure: phenyl isothiocyanate at the oligo position to simplify the analytical characterization of NNN-(propargyl-PEG2)(6-oxo-6-(dibenzo[b,f]azacyclooct-4-yl)-caproic acid)(PEG3-1-acetamido-4-iso-thiocyanato-benzene), (2) propargyl, and (3) model vanillin. The molecular structure was confirmed using LC-ESI-MS and its function was tested. HPLC analysis showed the formation of a high-yield product and the functional activity of the important reactive isothiocyanate moiety. A modular design was used to create the trifunctional molecular base such that each component and linkage could be exchanged for alternative structures as needed. The composition was designed for stability under cyclic Edman conditions while retaining downstream functionality: · PEG is inert to acid and base degradation. · The peptide bond used to connect the modular components is similar to the internal peptide bonds of proteins that are little affected during the Edman degradation process. · 1,2,3-triazole is thought to be stable to anhydrous acid and basic conditions useful for protein sequencing described herein. · The same oligonucleotide protecting groups used during phosphoramidite synthesis may be used. Exposure to trichloroacetic acid (TCA) during the synthesis of long oligos exceeds the expected chemical stress of the protein sequencing steps herein.
[0400] Synthesis of PPO, trifunctional CRC Synthesis of PDON-tBOC: N-(Propargyl-PEG2)-DBCO-PEG3-amine, TFA salt (PDA, Broadpharm catalog number 29932, 4.56 mg, 0.0063 mmol) was dissolved in 200 μL of 100 mM pH 8.65 phosphate buffer and mixed with 15.8 μL of 400 mM carbonate buffer pH 9.6. t-Boc-aminooxy-PEG4-NHS ester (broadpharm catalog number 24429, 10 mg, 0.021 mmol) was dissolved in 100 μL of DMSO. The solutions were combined, mixed with a pipette, 200 μL of dimethyl sulfoxide (DMSO) was added, and the reaction was incubated at room temperature (RT) for 18 h. The product was purified using high performance liquid chromatography (HPLC). m / z = 969 (positive mode) [M+H] + The electrospray ionization-mass spectrometry (ESI-MS) peak at + indicated the success of the synthesis.
[0401] Synthesis of PDON: PDON-tBOC was evaporated at 45 °C for 3 h under reduced pressure and then redissolved in dichloromethane (100 μl). Trifluoroacetic acid was added (30 μL), the mixture was incubated at room temperature (RT) for 1.5 h, neutralized by adding 180 μL of 4.1 M imidazole in acetonitrile / methanol (2:3 volume:volume), and purified using HPLC. The success of the synthesis of the intermediate PDON was confirmed by ESI-MS analysis. Specifically, the observation of a peak at m / z = 869 [M+H]+ indicated the success of the synthesis of PDON.
[0402] Synthesis of PDO: PDON was partially evaporated at 45 °C for 3 hours under reduced pressure. The concentration was quantified using optical density measurement (OD = 45, 3.4 mM) at 310 nm. Sys3 SOC Oligonucleotide ( / 5Phos / TGAAGGG / iFormInd / TGACCTAGCAATGGTGAAGTTAATGCAGGTAGTTAAG (SEQ ID NO: 108), Integrated DNA Technology, 178.8 nmol, where iFormInd represents a formyl indole modification for subsequent tethering)) was resuspended in 100 μL of 1x SSPE buffer (Sigma), and 10 μL of the oligo solution was added to 10 μL of 11.3 mM 5-aminoindole and 20 μL of 390 mM pH 5.5 acetate buffer. 15 μL of the PDON solution was added to this mixture, the solution was mixed by pipetting, and the reaction was incubated at room temperature for 18 hours. After the reaction, the product was purified using high-performance liquid chromatography (HPLC) and then dried under reduced pressure at a temperature of 45 °C for 4 hours. Electrospray ionization time-of-flight mass spectrometry (ESI-TOF-MS) analysis was performed on the product, and a peak occurred at m / z = 14920 [M+H]+, indicating the successful synthesis of compound PDO. In the control experiment, the mass spectrum of the Sys3 SOC oligonucleotide alone was found to have a peak at m / z = 14069 [M+H]+.
[0403] Synthesis of PPO: PDO was resuspended in 28 μL of Milli-Q water (OD260 = 20, 43 μM). 4-Azidophenyl isothiocyanate (N3PITC, 1.30 mg, 0.0074 mmol) was dissolved in 1 mL of DMSO to form a 7.4 mM solution. 90 μL of the N3PITC solution was added to the PDO solution and mixed with a pipette. The reaction was incubated at room temperature for 3 hours. After this incubation period, the product was purified using high-performance liquid chromatography (HPLC), and two prominent peaks were obtained at 14 minutes and 18 minutes. These product peaks were further analyzed using quadrupole time-of-flight mass spectrometry (QTOF MS). This analysis resulted in a peak at m / z = 15094 [M]+ for the product corresponding to the 14-minute mark of the HPLC analysis. For the product corresponding to the 18-minute time point, a peak of m / z = 15095 [M + H]+ was observed, indicating the success of the synthesis process. The two peaks can be assumed to be isomers (e.g., positional isomers ...
Claims
1. A method for determining the identity and positional information of amino acid residues of a peptide coupled to a solid support, said method comprising: (a) providing said peptide to said solid support, wherein said peptide is coupled to said solid support such that the N-terminal amino acid residue of said peptide is not directly coupled to said solid support and is exposed to reaction conditions; (b) providing a chemically reactive conjugate, said chemically reactive conjugate comprising (x) a cycle tag comprising a cyclic nucleic acid associated with the number of cycles, (y) a reactive moiety for binding to said N-terminal amino acid residue of said peptide, and (z) an immobilization moiety for immobilizing on said solid support; (c) contacting said peptide with said chemically reactive conjugate, thereby coupling said chemically reactive conjugate to said N-terminal amino acid of said peptide to form a conjugate complex; (d) immobilizing said conjugate complex on said solid support via said immobilization moiety; (e) cleaving said N-terminal amino acid residue and separating it from said peptide, thereby providing an immobilized amino acid complex, said immobilized amino acid complex comprising said cleaved and separated N-terminal amino acid residue; (f) contacting said immobilized amino acid complex with a binder, said binder comprising a binding moiety for preferentially binding to said immobilized amino acid complex and a recoding tag comprising a recoding nucleic acid corresponding to said binder, thereby forming an affinity complex, said affinity complex comprising said immobilized amino acid complex and said binder, thereby bringing said cycle tag into proximity to said recoding tag within said affinity complex; (g) transmitting the information of said recoding nucleic acid to said cyclic nucleic acid of said immobilized conjugate complex to generate a recoding block; (j) obtaining the sequence information of said recoding block; (k) determining the identity and positional information of the amino acid residues of said peptide based on said obtained sequence information. A method as described above.
2. The method according to claim 1, wherein cleaving said N-terminal amino acid residue from said peptide exposes the next amino acid residue as the N-terminal amino acid residue on said cleaved peptide.
3. The method according to claim 2, wherein the reactive moiety of the chemically reactive conjugate cleaves the N-terminal amino acid residue from the peptide.
4. The method according to claim 2, further comprising repeating steps (b) to (k) for each subsequent amino acid of the peptide.
5. The method according to claim 1, further comprising washing the immobilized amino acid complex before contacting the immobilized amino acid complex with the binder.
6. The method according to claim 1, further comprising determining a possible three-dimensional structure of the peptide based on the sequence information.
7. The method according to claim 1, wherein the recoded nucleic acid comprises DNA or RNA.
8. The method according to claim 1, wherein the cyclic nucleic acid comprises DNA or RNA.
9. The method according to claim 1, wherein obtaining the sequence information of the recoding block comprises performing sequencing.
10. The method according to claim 1, wherein the binding moiety comprises a peptide.
11. antibody
12. antibody fragment, antibody derivative, or aptamer.
13. The method according to claim 1, wherein the binding moiety binds to a natural amino acid, a post-translationally modified amino acid, a derivatized version of an amino acid, a derivatized or stabilized version of a post-translationally modified amino acid, a synthetic amino acid, an amino acid having a specific side chain, an amino acid having a phosphorylated side chain, an amino acid having a glycosylated side chain, an amino acid having a methylation modification, or a D-amino acid of an amino acid, a phenylthiohydantoin (PTH) derivative or an anilinothiazolinone (ATZ) derivative, or a combination thereof.
14. The method according to claim 1, wherein the solid support comprises beads, plates, or chips.
15. The method according to claim 1, wherein the solid support comprises a glass slide, silica, resin, gel, hydrogel, membrane, polystyrene, metal, nitrocellulose, mineral, plastic, polyacrylamide, latex, or ceramic.
16. The method according to claim 1, wherein the peptide comprises a hormone, neurotransmitter, enzyme, antibody, viral protein, bacterial protein, synthetic peptide, bioactive peptide, peptide hormone, oligopeptide, polypeptide, fusion protein, cyclic peptide, branched peptide, recombinant protein, tumor marker, therapeutic peptide, antigenic peptide, or signal transduction peptide.
17. The method according to claim 1, wherein the peptide is derived from a cell lysate, blood sample, plasma sample, serum sample, tissue biopsy material, saliva sample, urine sample, cerebrospinal fluid sample, sweat sample, synovial fluid sample, fecal sample, gut microbiota sample, environmental water sample, soil sample, bacterial culture, viral culture, organoid, tumor biopsy material, sputum sample, or hair sample.
18. The method according to claim 1, wherein the peptide is associated with a disease.
19. The method according to claim 1, wherein transmitting the information comprises performing nucleic acid amplification.
20. Enzyme ligation, splint ligation, chemical ligation, template-assisted ligation, use of ligase enzyme, use of splint oligonucleotide, use of catalyst, use of cross-linking molecule, use of condensing agent, use of coupling reagent, use of polymerase enzyme, use of complementary nucleic acid sequence, use of nicking enzyme, use of nucleic acid modifying enzyme, use of recombinase, use of strand displacement polymerase, use of single-stranded binding protein, click chemistry reaction, phosphodiester bond formation, or peptide nucleic acid-mediated ligation.
21. The method according to claim 1, wherein the information of the recoded nucleic acid comprises the sequence of the recoded nucleic acid or the reverse complement of the sequence of the recoded nucleic acid.
22. The method according to claim 1, wherein transmitting the information comprises conjugating the recoded nucleic acid or the reverse complement of the recoded nucleic acid with the cyclic nucleic acid.
23. A method for determining the identity and position information of a plurality of amino acid residues of a peptide, wherein the peptide comprises n amino acid residues, and the method comprises (a) coupling the peptide to a solid support such that the N-terminal amino acid residue of the peptide is not directly coupled to the solid support but is exposed to reaction conditions. to provide a chemically reactive conjugate, wherein the chemically reactive conjugate comprises: (x) a cycle tag comprising a cyclic nucleic acid related to the number of cycles; (y) a reactive moiety for binding to and cleaving the N-terminal amino acid residue of the peptide to expose the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide; and (z) an immobilization moiety for immobilizing on the solid support to contact the peptide with the chemically reactive conjugate, thereby coupling the chemically reactive conjugate to the N-terminal amino acid of the peptide to form a conjugate complex to immobilize the conjugate complex on the solid support via the immobilization moiety to cleave the N-terminal amino acid residue, thereby separating it from the peptide and exposing the next amino acid residue as the N-terminal amino acid residue on the cleaved peptide, to provide an immobilized amino acid complex, wherein the immobilized amino acid complex comprises the cleaved and separated N-terminal amino acid residue to repeat steps (b) to (e) n - 1 times to assemble n - 1 additional immobilized amino acid complexes, wherein each additional immobilized amino acid complex comprises a nucleic acid related to cycles 2 to n accordingly to contact the immobilized amino acid complex with a binder, wherein the binder comprises a binding moiety for preferentially binding to one or a subset of the immobilized amino acid complexes and a recoding tag comprising a recoding nucleic acid corresponding to the binder, thereby forming one or more affinity complexes, each affinity complex comprising an immobilized amino acid complex and the binder, thereby bringing the cycle tag into proximity to the recoding tag in each formed affinity complex in each formed affinity complex, to ligate the cycle tag or its reverse complement to the recoding tag to form a recoding block, or otherwise transfer the information of the recoding nucleic acid to the cycle nucleic acid of the immobilized conjugate complex, thereby creating a plurality of recoding blocks, each recoding block corresponding to the formed affinity complex, to create a plurality of recoding blocks (i) joining two or more members of the plurality of recoding blocks to form a memory oligonucleotide; (j) obtaining sequence information of the memory oligonucleotide; (k) determining the identity and position information of a plurality of amino acid residues of the peptide based on the obtained sequence information. A method comprising:
24. The method according to claim 23, wherein (g) to (h) are repeated 2, 3, 4 times or more.
25. The method according to claim 23, wherein n is an integer greater than or equal to 2.
26. The method according to claim 23, wherein each binder comprises a recoding tag having a unique nucleic acid sequence.
27. The method according to claim 23, wherein a plurality of binders comprise recoding tags having the same nucleic acid sequence.
28. The method according to claim 23, wherein the binder comprises a recoding tag having a unique sequence portion and a common sequence portion.
29. The method according to claim 23, further comprising washing the immobilized amino acid complex before contacting the immobilized amino acid complex with the binder.
30. The method according to claim 23, further comprising determining a possible three-dimensional structure of the peptide based on the sequence information.
31. The method according to claim 23, wherein the recoding nucleic acid comprises DNA.
32. The method according to claim 23, wherein the cyclic nucleic acid comprises DNA.
33. The method according to claim 23, wherein obtaining the sequence information of the memory oligonucleotide comprises performing sequencing.
34. The method according to claim 23, wherein the binding moiety comprises an antibody or a fragment thereof, or an aptamer.
35. The method according to claim 23, wherein the binding moiety binds to a natural amino acid, a derivatized amino acid, a synthetic amino acid, or a D-amino acid.
36. The method according to claim 23, wherein the binding moiety binds to a post-translational modification.
37. The method according to claim 23, wherein the solid support comprises beads, plates, or chips.
38. The method according to claim 23, wherein the solid support comprises a glass slide, silica, resin, gel, hydrogel, membrane, polystyrene, metal, nitrocellulose, mineral, plastic, polyacrylamide, latex, or ceramic.
39. The method of claim 23, further comprising deprotecting the cycle tag between (f) and (g).
40. The method of claim 23, wherein transmitting the information comprises performing nucleic acid amplification, enzymatic ligation, sprint ligation, chemical ligation, template-assisted ligation, use of a ligase enzyme, use of a sprint oligonucleotide, use of a catalyst, use of a cross-linking molecule, use of a condensing agent, use of a coupling reagent, use of a polymerase enzyme, use of a complementary nucleic acid sequence, use of a nicking enzyme, use of a nucleic acid modifying enzyme, use of a recombinase, use of a strand-displacing polymerase, use of a single-stranded binding protein, a click chemistry reaction, phosphodiester bond formation, or peptide nucleic acid-mediated ligation.
41. The method of claim 23, wherein the information of the recoded nucleic acid comprises the sequence of the recoded nucleic acid or the reverse complement of the sequence of the recoded nucleic acid.
42. The method of claim 23, wherein transmitting the information comprises conjugating the recoded nucleic acid or the reverse complement of the recoded nucleic acid to the cycle nucleic acid.
43. The method of claim 23, wherein determining the identity and positional information of the plurality of amino acid residues of the peptide comprises determining the identity and positional information of all amino acid residues of the peptide.
44. The method of claim 23, wherein determining the identity and positional information of the plurality of amino acid residues of the peptide comprises determining the identity and positional information of amino acid residues of only a subset of the peptide.
45. The method of claim 23, further comprising identifying the peptide by comparing the identity and positional information of the plurality of amino acid residues to a database.
46. A chemically reactive conjugate (CRC) comprising: (A) a nucleic acid sequence tag; (B) a reactive moiety for binding and cleaving an N-terminal amino acid residue to a peptide or a cleavable derivative thereof; and (C) an immobilization moiety for immobilization to a solid support.
47. Formula I: 【Chemical Formula 9】 (wherein A includes a cycle tag, B includes a reactive moiety, C includes an immobilization moiety, L A includes any linker, L B includes any linker, L C includes any linker, 【Chemical 10】 is a central moiety) represented CRC.
48. Formula II: 【Chemical 11】 (wherein A includes a cycle tag, B includes a reactive moiety, C includes an immobilized moiety, and L AB includes any linker, and L BC includes any linker) represented by CRC.
49. The CRC according to claim 46 or 47, comprising a cleavable group between (A) and (B), between (B) and (C), between (A) and (C), between (A) and (C), between (A) and (B + C), between (B) and (A + C), between (C) and (A + B), or any combination thereof.
50. The CRC according to claim 48, comprising a cleavable group between (A) and (B), between (B) and (C), or a combination thereof.
51. The CRC according to claim 46, wherein (A), (B), and (C) are linearly oriented relative to each other in any of the following orders: (A)-(B)-(C), (A)-(C)-(B), or (B)-(A)-(C).
52. The CRC according to any one of claims 46 to 48, wherein the reactive moiety comprises phenyl isothiocyanate (PITC), isothiocyanate (ITC), dansyl chloride, dinitrofluorobenzene (DNFB), an enzyme or a peptide, or a combination or derivative thereof.
53. The CRC according to any one of claims 46 to 48, wherein the reactive moiety specifically cleaves at a specific amino acid.
54. The CRC according to any one of claims 46 to 48, wherein the reactive moiety cleaves more than one amino acid or motif.
55. The CRC according to any one of claims 46 to 48, wherein the immobilized moiety comprises a protected thiol group, a protected amine group, or a carboxyl group, azide, alkyne, alkene, arylboronic acid, halogenated aryl, haloalkyne, silylalkyne, Si-H group, a protected or photo-protected reactive group, or a photo-activated reactive group.
56. The CRC according to any one of claims 46 to 48, wherein the nucleic acid sequence tag is generated by conjugating the nucleic acid sequence to a group for attaching a nucleic acid sequence comprising a protected oxyamine group, a protected thiol, a protected amine, a protected hydrazine, tetrazine, azide, alkyne, alkene, trans-cyclooctene, DBCO, bicyclononine, norbornene, strained alkyne, or strained alkene, or a derivative thereof.
57. The CRC according to any one of claims 46 to 48, wherein the reactive moiety comprises a group on the CRC for attachment to a cleavable derivatized N-terminal amino acid comprising tetrazine, azide, alkene, alkyne, trans-cyclooctene, DBCO, bicyclononine, norbornene, strained alkyne, or strained alkene, or a derivative thereof.
58. A kit for determining the identity and position information of amino acid residues of a peptide, comprising: (a) a nucleic acid sequence tag; and (b) a chemical reactive conjugate comprising a reactive moiety that couples to the N-terminal amino acid residue of the peptide, thereby forming a conjugate complex comprising the chemical reactive conjugate coupled to the N-terminal amino acid of the peptide. A binder comprising a binding moiety for preferentially binding to the conjugate complex, a recoding tag comprising a recoding nucleic acid corresponding to the binder, and A kit comprising a reagent for transmitting information of the recoding nucleic acid to the cyclic nucleic acid of the conjugate complex to generate a recoding block.
59. A method for sequencing a subset of nucleotides of an oligonucleotide, comprising: Providing a combination of reversibly terminated nucleotides and non-reversibly terminated nucleotides in a nucleic acid sequencing reaction, wherein the nucleotides of the nucleic acid to be sequenced corresponding to the non-reversibly terminated nucleotides are not sequenced.
60. The method according to claim 59, further comprising identifying the nucleotides of the nucleic acid to be sequenced corresponding to the reversibly terminated nucleotides.
61. The method according to claim 59, wherein the nucleic acid to be sequenced comprises a region comprising only a subset of nucleotides selected from A, C, G, and T, and the subset of nucleotides is not sequenced.
62. The method according to claim 61, wherein the subset of nucleotides selected from A, C, G, and T comprises two nucleotides selected from A, C, G, and T.
63. The method according to claim 61, wherein the subset of nucleotides selected from A, C, G, and T comprises three nucleotides selected from A, C, G, and T.
64. The method according to claim 61, wherein the region comprises a primer sequence.
65. The method according to claim 61, wherein the region does not include a barcode array, a recoded nucleic acid sequence or a part thereof, or a cyclic nucleic acid sequence or a part thereof.
66. A method comprising: providing a conjugate comprising a reactive molecule coupled to a protected oligonucleotide; contacting the reactive moiety with a terminal amino acid of the peptide, thereby binding the reactive moiety to the terminal amino acid and, optionally, cleaving the terminal amino acid from the peptide; deprotecting the oligonucleotide; contacting the deprotected oligonucleotide with an enzyme or reagent for ligation or polymerization.
67. The method according to claim 66, further comprising reprotecting the oligonucleotide.
68. The method according to claim 67, wherein the reactive moiety cleaves the terminal amino acid from the peptide to expose the next terminal amino acid, and the method further comprises contacting the next amino acid with another conjugate after reprotecting the oligonucleotide.
69. The method according to claim 66, wherein the terminal amino acid is an N-terminal.
70. The method according to claim 66, wherein the peptide is immobilized on a solid support.
71. The method according to claim 66, wherein the conjugate comprises an organic small molecule.
72. The method according to claim 66, wherein the conjugate comprises a chemically reactive conjugate (CRC), and the chemically reactive conjugate (CRC) comprises (A) the oligonucleotide, (B) the reactive moiety, and (C) an immobilized moiety.
73. The method according to claim 66, wherein the oligonucleotide comprises a cyclic nucleic acid.
74. A method comprising: providing a conjugate comprising a peptide coupled to a protected oligonucleotide; contacting the terminal amino acid of the peptide, thereby binding a reactive moiety to the terminal amino acid and, optionally, cleaving the terminal amino acid from the peptide; deprotecting the oligonucleotide; contacting the deprotected oligonucleotide with an enzyme or reagent for ligation or polymerization.
75. The method according to claim 74, further comprising reprotecting the oligonucleotide. **Claim 76** The method according to claim 75, wherein the reactive moiety cleaves the terminal amino acid from the peptide to expose the next terminal amino acid, and the method further comprises contacting the next amino acid with another conjugate after re - protecting the oligonucleotide. **Claim 77** The method according to claim 74, wherein the terminal amino acid is an N - terminal amino acid. **Claim 78** The method according to claim 74, wherein the peptide is immobilized on a solid support. **Claim 79** The method according to claim 74, wherein the conjugate comprises a small organic molecule.