Improved library preparation for polypeptide sequencing
The method of protein digestion, derivatization, and conjugation to an immobilization complex, along with signal monitoring, addresses the limitations of conventional protein characterization methods by enhancing polypeptide identification and sequencing, especially for proteins with post-translational modifications.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- QUANTUM SI INC
- Filing Date
- 2025-11-26
- Publication Date
- 2026-06-04
AI Technical Summary
Conventional methods for protein characterization, such as mass spectrometry and affinity-based methods, face challenges in identifying unknown proteins and differentiating unmodified proteins from those with post-translational modifications, limiting the effectiveness of proteomics in understanding cellular processes and disease progression.
A method involving protein digestion, derivatization, and conjugation to an immobilization complex, combined with the use of amino acid recognizers and cleaving agents, to prepare polypeptides for sequencing, enabling the determination of chemical characteristics through signal monitoring.
Enhances the identification and sequencing of polypeptides, particularly those with post-translational modifications, improving the accuracy and coverage of protein analysis.
Smart Images

Figure US2025057343_04062026_PF_FP_ABST
Abstract
Description
[0001] IMPROVED LIBRARY PREPARATION FOR POLYPEPTIDE SEQUENCING
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003]
[0001] The present application claims the benefit of priority of U. S. Provisional Application No.
[0004] 63 / 726,142, filed November 27, 2024, the entire contents of which are incorporated herein by reference.
[0005] BACKGROUND OF THE INVENTION
[0006]
[0002] Measurements of the proteome provide deep and valuable insight into key biological processes. Protein characterization has a number of important applications, including determination of the presence or absence of a protein (e.g., a disease-relevant protein) in a biological sample, identification of an unknown protein in a biological sample, and identification of a protein responsible for biological activity in an isolated protein fraction. However, conventional methods of characterizing proteins, such as mass spectrometry and affinity-based methods, often face substantial challenges, including the inability to identify unknown proteins and / or differentiate unmodified proteins from proteins with post-translational modifications (PTMs).
[0007]
[0003] Protein sequencing can provide insights into cellular processes and response patterns, which lead to improved diagnostic and therapeutic strategies. In adjacent fields, such as genomics, advances in DNA sequencing technology have proven extremely valuable in improving understanding of the progression of complex human disease. However, applying similar approaches to proteomics has been challenging for a number of reasons, including the large number of different proteins and proteoforms, the wide dynamic range of protein abundance in cells and biological fluids, and the inability to copy or amplify proteins. Accordingly, improved approaches for the preparation of protein samples for sequencing are needed.
[0008] SUMMARY OF THE INVENTION
[0009]
[0004] The present disclosure provides methods, articles, kits, and / or systems for the preparation and / or analysis of a polypeptide. Through the use of methods, articles, kits, and / or systems of the present disclosure, polypeptides may be more readily prepared for sequencing and / or may be more readily sequenced.
[0010]
[0005] Accordingly, in one aspect, provided herein is a method of preparing a polypeptide sample from a protein for analysis, comprising:
[0011] exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample; derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides; and
[0012] conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0013]
[0006] In some embodiments, the method further comprises denaturing the protein before exposing the protein to the protein digestion agent. In some embodiments, the method further comprises exposing the R0708.70180WO00 / R0708.70180US01 1 / 122
[0014] #14644680vl protein to a reducing agent before exposing the protein to the protein digestion agent. In some embodiments, the method further comprises exposing the protein to an amino acid side chain capping agent before exposing the protein to the protein digestion agent. In some embodiments, the method does not comprise a buffer exchange step before denaturing the protein. In some embodiments, the method does not comprise a buffer exchange step before exposing the protein to the reducing agent.
[0015]
[0007] In another aspect, provided herein is a method of preparing a polypeptide sample from a protein for analysis, comprising:
[0016] exposing the protein to a reducing agent;
[0017] exposing the protein to an amino acid side chain capping agent;
[0018] exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample.
[0019]
[0008] In some embodiments, the method further comprises derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides. In some embodiments, the method further comprises conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0020]
[0009] In another aspect, provided herein is a method of preparing a polypeptide sample from a protein for analysis, comprising:
[0021] exposing the protein to a reducing agent;
[0022] exposing the protein to an amino acid side chain capping agent;
[0023] exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample; derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides; and
[0024] conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0025]
[0010] In some embodiments, the method further comprises:
[0026] contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents and one or more amino acid recognizers;
[0027] monitoring a signal for signal pulses corresponding to interactions between one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides of the polypeptide sample; and
[0028] determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal.
[0029]
[0011] In another aspect, provided herein is a method of protein analysis, comprising:
[0030] exposing the protein to a reducing agent;
[0031] exposing the protein to an amino acid side chain capping agent;
[0032] exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample; derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized R0708.70180WO00 / R0708.70180US01 2 / 122
[0033] #14644680vl polypeptide sample comprising one or more derivatized polypeptides;
[0034] conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0035] contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents and one or more amino acid recognizers;
[0036] monitoring a signal for signal pulses corresponding to interactions between one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides of the polypeptide sample; and
[0037] determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal.
[0038]
[0012] In some embodiments, the method further comprises denaturing the protein before exposing the protein to the reducing agent. In some embodiments, the method does not comprise a buffer exchange step before denaturing the protein. In some embodiments, the method does not comprise a buffer exchange step before exposing the protein to the reducing agent.
[0039]
[0013] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III):
[0040]
[0041] or a salt thereof, wherein L1and R1are defined herein.
[0042]
[0014] In some embodiments, the protein digestion agent induces proteolysis of the protein to form one or more capped polypeptides, thereby forming the digested polypeptide sample. In some embodiments, the derivatizing comprises derivatizing an amino acid side chain of the one or more capped polypeptides using a derivatization agent to form an unquenched mixture comprising one or more derivatized polypeptides. In some embodiments, the unquenched mixture further comprises excess derivatization agent. In some embodiments, the method further comprises quenching the unquenched mixture to form a quenched mixture by removing at least some of the excess derivatization agent.
[0043]
[0015] In some embodiments, one or more capped polypeptides are of Formula (I):
[0044] S'U- / RN
[0045] N
[0046]
[0047] (I),
[0048] or a salt thereof, wherein L1, R1, Rc, and RNare defined herein.
[0049]
[0016] In some embodiments, one or more derivatized polypeptides are of Formula (IV):
[0050] / RN
[0051] N
[0052]
[0053] (IV),
[0054] or a salt thereof, wherein L1, Rc, and RNare defined herein.
[0055] R0708.70180WO00 / R0708.70180US01 3 / 122
[0056] #14644680vl
[0017] In some embodiments, the method further comprises binding the protein to a solid substrate before, at the same time as, or after exposing the protein to the protein digestion agent. In some embodiments, the method further comprises purifying the digested polypeptide sample to form a purified digested polypeptide sample.
[0057]
[0018] In some embodiments, the method further comprises purifying the derivatized polypeptide sample to form a purified derivatized polypeptide sample.
[0058]
[0019] In another aspect, provided herein is a compound of Formula (I):
[0059]
[0060] (I),
[0061] or a salt thereof, wherein L1, R1, Rc, and RNare defined herein.
[0062]
[0020] In another aspect, provided herein is a method of preparing a compound of Formula (I):
[0063] RCY^N-RN
[0064]
[0065] 0 H(1),
[0066] or a salt thereof, comprising contacting a compound of Formula (II):
[0067]
[0068] (II),
[0069] or a salt thereof, with a compound of Formula (III):
[0070] Y^L1^N(R1)2
[0071]
[0072] (III),
[0073] or a salt thereof, to obtain the compound of Formula (I), or a salt thereof, wherein Y, L1, R1, Rc, and RNare defined herein.
[0074]
[0021] In another aspect, provided herein is a method of protein digestion, comprising:
[0075] contacting a compound of Formula (II):
[0076]
[0077] (II),
[0078] or a salt thereof, with a compound of Formula (III):
[0079] Y^L1^N(R1)2
[0080]
[0081] (III),
[0082] R0708.70180WO00 / R0708.70180US01 4 / 122
[0083] #14644680vl or a salt thereof, to obtain a compound of Formula (I):
[0084] , S^L1^N(R’)2
[0085] RV N-RN
[0086]
[0087] 0 H(I),
[0088] or a salt thereof; and
[0089] exposing the compound of Formula (I), or salt thereof, to a protein digestion agent, thereby forming a digested polypeptide sample, wherein Y, L1, R1, Rc, and RNare defined herein.
[0090]
[0022] In another aspect, provided herein is a method of protein analysis, comprising:
[0091] contacting a compound of Formula (II):
[0092]
[0093] (II),
[0094] or a salt thereof, with a compound of Formula (III):
[0095] ^N(R1)2
[0096]
[0097] L(HI),
[0098] or a salt thereof, to obtain a compound of Formula (I):
[0099] S^L,„N(R1)2
[0100] / RN
[0101] N
[0102] H
[0103]
[0104] (I),
[0105] or a salt thereof;
[0106] exposing the compound of Formula (I), or salt thereof, to a protein digestion agent, thereby forming a digested polypeptide sample;
[0107] derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides;
[0108] conjugating the one or more derivatized polypeptides to an immobilization complex to form a polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0109] contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents and one or more amino acid recognizers;
[0110] monitoring a signal for signal pulses corresponding to interactions between one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides of the polypeptide sample; and
[0111] determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal;
[0112] wherein Y, L1, R1, Rc, and RNare defined herein.
[0113] R0708.70180WO00 / R0708.70180US01 5 / 122
[0114] #14644680vl
[0023] In some embodiments, the protein digestion agent is an enzymatic protein digestion agent. In some embodiments, the protein digestion agent is Lys-C.
[0115]
[0024] It should be appreciated that the foregoing concepts, and the additional concepts discussed below, may be arranged in any suitable combination, as the present disclosure is not limited in this respect. Further, other advantages and novel features of the present disclosure will become apparent from the following detailed description of various non-limiting embodiments when considered in conjunction with the accompanying drawings.
[0116] BRIEF DESCRIPTION OF THE DRAWINGS
[0117]
[0025] FIG. 1 shows an example overview of real-time dynamic protein sequencing. Protein samples are digested into peptide fragments, immobilized in nanoscale reaction chambers, and incubated with a mixture of freely-diffusing N-terminal amino acid (NAA) recognizers and aminopeptidases that carry out the sequencing process. The labeled recognizers bind on and off to the peptide when one of their cognate NAAs is exposed at the N-terminus, thereby producing characteristic pulsing patterns. The NAA is cleaved by an aminopeptidase, exposing the next amino acid for recognition. The temporal order of NAA recognition and the kinetics of binding enable peptide identification and are sensitive to features that modulate binding kinetics, such as post-translational modifications (PTMs).
[0118]
[0026] FIG. 2 shows a schematic of a sample preparation process. The starting protein is reduced and alkylated, subjected to enzymatic digestion, derivatized (functionalized), and conjugated to a linker for subsequent analysis (e.g., sequencing).
[0119]
[0027] FIG. 3 shows V2 Library Prep Workflow (SOP, central box) and improvements to the process (boxes on right hand side).
[0120]
[0028] FIG. 4A shows that bromoethylamine (BEA) converts cysteines into pseudo-lysine sites for Lys-C.
[0121] FIG. 4B shows that magnetic SP3 beads (Single Pot, Solid Phase enhanced Sample Preparation) can be used for protein clean up, on bead digestion, and peptide clean up.
[0122]
[0029] FIG. 5 shows that exposure to a post digestion C18 resin benefits some proteins. Post digestion C18 resin showed equal or better Inference Precision (IF) for 4 / 5 tested proteins.
[0123]
[0030] FIGs, 6A-6D show library preparation improvement with the combination of modifications of: (1) bromoethylamine (BEA) instead of chloroacetamide (CAA); (2) SP3 magnetic beads for protein clean up, on bead digestion, and peptide clean up; and (3) post digestion C18 resin. As a proof of concept, BEA / SP3 beads / C18 post-diazotransfer cleanup workflow demonstrated better sequencing performance (Alignment / Rl) than SOP for proteins MFN2, TPM1 and TMLH.
[0124]
[0031] FIG. 7 shows library preparation improvement with nickel (II) acetate vs. SOP (CUSO4). Proteins on the list have different size and different number of cysteines.
[0125]
[0032] FIG. 8 shows percent change alignments of individual peptides (percent change in nickel compared to copper).
[0126]
[0033] FIG. 9 shows the molecular weights of proteins studied in Example 1.
[0127]
[0034] FIG. 10 shows results for protein sequencing when the protein samples were prepared according to R0708.70180WO00 / R0708.70180US01 6 / 122
[0128] #14644680vl Example 1.
[0129]
[0035] FIG. 11 shows comparisons of the alignment ratio (top) and number of peptides ratio (bottom) between the sample preparation of Example 1 and the previous method (V2).
[0130]
[0036] FIG. 12 shows the inference ranks for sequencing of samples from the sample preparation method of Example 1 with or without buffer exchange, compared to those from a previous method (V2).
[0131]
[0037] FIG. 13 shows a direct comparison of sequencing results obtained from the sample preparation method of Example 1 (V3) and the previous sample preparation method (V2).
[0132]
[0038] FIGs. 14A-14B show a comparison of standard and BEA chemistry workflows for protein sequencing. FIG. 14A shows standard workflow: Following reduction with TCEP, cysteines are alkylated with chloroacetamide (CAA) to prevent disulfide bond reformation. Lys-C digestion cleaves only at native lysine residues, followed by peptide derivatization and NGPS analysis. This approach provides limited coverage for lysine-poor proteins. FIG. 14B shows BEA workflow: Following reduction, cysteines are aminoethylated with 2-bromoethylamine (BEA) to generate pseudolysine residues. These pseudolysines serve as additional Lys-C cleavage sites, producing more peptides of suitable length for derivatization and sequencing, thereby enhancing sequence coverage.
[0133]
[0039] FIGs. 15A-15F show BEA chemistry achieves conversion efficiency comparable to native lysine. Two synthetic peptides differing only in the substitution of lysine with cysteine at internal and terminal position (FIGs. 15A-15B) were processed using either standard chemistry or BEA derivatization. Sequencing results showed equivalent alignment counts between the two methods (FIGs. 15C-15D), with comparable sequence coverage demonstrated by kinetic coverage maps (FIGs. 15E-15F).
[0134]
[0040] FIG. 16 shows BEA-mediated conversion of cysteines in pseudolysines generated appropriately sized peptides for NGPS. SDS-PAGE comparison of Lys-C digestion products from CD6 and LRC32 under standard and BEA chemistry workflows. High molecular weight bands present with standard chemistry (CD6: ~50 kDa; LRC32: 25-75 kDa) are eliminated with BEA treatment, demonstrating effective reduction of fragment size through increased cleavage site density.
[0135]
[0041] FIGs. 17A-17B show BEA chemistry enables confident identification of lysine-poor CD6 protein.
[0136] FIG. 17A shows standard chemistry yields no confidently identified peptides (FDR < 10%) despite native lysine residues (K) present in the CD6 sequence. FIG. 17B shows BEA chemistry converts cysteine residues to pseudolysines, creating additional Lys-C cleavage sites that led to identification of four high-confidence peptides (indicated with *, FDR < 10%). These newly generated and identified peptides enabled correct inference of CD6 as the top-ranked protein.
[0137]
[0042] FIGs. 18A-18B show BEA chemistry enables confident identification of lysine-poor LRC32 protein. FIG. 18A shows standard chemistry yields no confidently identified peptides (FDR < 10%) despite native lysine residues (K) present in the LRC32 sequence. FIG. 18B shows BEA chemistry converts cysteine residues to pseudolysines, creating additional Lys-C cleavage sites that identified five high-confidence peptides (indicated with *, FDR < 10%). This expansion of the peptides in the library enables correct inference of LRC32 as the top-ranked protein.
[0138]
[0043] FIGs. 19A-19E show the effect of surfactant condition (i.e., identity and amount) on protein R0708.70180WO00 / R0708.70180US01 7 / 122
[0139] #14644680vl sequencing. FIG. 19A shows alignments of sequencing runs, grouped by condition with a line connecting the median values. FIG. 19B shows alignments / RRLl+ of sequencing runs, grouped by condition with a line connecting the median values. FIG. 19C shows number of peptides identified for sequencing runs, grouped by condition with a line connecting the median values. FIG. 19D shows mean number of alignments for sequencing runs of each surfactant condition divided by the mean number of alignments for all SOP (no surfactant) runs. FIG. 19E shows mean number of alignments for each protein (grouped by condition) divided by the mean number of alignments for all SOP (no surfactant) runs for the same protein.
[0140] DEFINITIONS
[0141]
[0044] Definitions of specific functional groups and chemical terms are described in more detail below. The chemical elements are identified in accordance with the Periodic Table of the Elements, CAS version, Handbook of Chemistry and Physics, 75thEd., inside cover, and specific functional groups are generally defined as described therein. Additionally, general principles of organic chemistry, as well as specific functional moieties and reactivity, are described in Thomas Sorrell, Organic Chemistry, University Science Books, Sausalito, 1999; Michael B. Smith, March ’s Advanced Organic Chemistry, 7thEdition, John Wiley & Sons, Inc., New York, 2013; Richard C. Larock, Comprehensive Organic Transformations, John Wiley & Sons, Inc., New York, 2018; and Carruthers, Some Modem Methods of Organic Synthesis, 3rdEdition, Cambridge University Press, Cambridge, 1987.
[0142]
[0045] Compounds described herein can comprise one or more asymmetric centers, and thus can exist in various stereoisomeric forms, e.g., enantiomers and / or diastereomers. For example, the compounds described herein can be in the form of an individual enantiomer, diastereomer or geometric isomer, or can be in the form of a mixture of stereoisomers, including racemic mixtures and mixtures enriched in one or more stereoisomer. Isomers can be isolated from mixtures by methods known to those skilled in the art, including chiral high pressure liquid chromatography (HPLC) and the formation and crystallization of chiral salts; or preferred isomers can be prepared by asymmetric syntheses. See, for example, Jacques et al., Enantiomers, Racemates and Resolutions (Wiley Interscience, New York, 1981); Wilen et al., Tetrahedron 33:2725 (1977); Eliel, E. L. Stereochemistry of Carbon Compounds (McGraw-Hill, NY, 1962); and Wilen, S. H., Tables of Resolving Agents and Optical Resolutions p. 268 (E. L. Eliel, Ed., Univ, of Notre Dame Press, Notre Dame, IN 1972). The present disclosure additionally encompasses compounds as individual isomers substantially free of other isomers, and alternatively, as mixtures of various isomers.
[0143]
[0046] Unless otherwise provided, formulae and structures depicted herein include compounds that do not include isotopically enriched atoms, and also include compounds that include isotopically enriched atoms. For example, compounds having the present structures except for the replacement of hydrogen by deuterium or tritium, replacement of19F with18F, or the replacement of a carbon by a13C- or deenriched carbon are within the scope of the disclosure. Such compounds are useful, for example, as analytical tools or probes in biological assays.
[0144] R0708.70180WO00 / R0708.70180US01 8 / 122
[0145] #14644680vl
[0047] When a range of values (“range”) is listed, it encompasses each value and sub-range within the range. A range is inclusive of the values at the two ends of the range unless otherwise provided. For example “Ci.6alkyl” encompasses, C C2, C3, C4, C5, C6, Ci_6, C1-5, Ci^, C1-3, C1-2, C2-6, C2-5, C2 4. C2-3, C3-6, C3-5, C34. C4 4. C4-5, and Cs alkyl.
[0146]
[0048] The term “aliphatic” refers to alkyl, alkenyl, alkynyl, and carbocyclic groups. Likewise, the term “heteroaliphatic” refers to heteroalkyl, heteroalkenyl, heteroalkynyl, and heterocyclic groups.
[0147]
[0049] The term “alkyl” refers to a radical of a straight-chain or branched saturated hydrocarbon group having from 1 to 20 carbon atoms (“Ci-2o alkyl”). In some embodiments, an alkyl group has 1 to 12 carbon atoms (“C1-12 alkyl”). In some embodiments, an alkyl group has 1 to 10 carbon atoms (“C1-10 alkyl”). In some embodiments, an alkyl group has 1 to 9 carbon atoms (“C1-9 alkyl”). In some embodiments, an alkyl group has 1 to 8 carbon atoms (“Ci-s alkyl”). In some embodiments, an alkyl group has 1 to 7 carbon atoms (“C1-7 alkyl”). In some embodiments, an alkyl group has 1 to 6 carbon atoms (“C1-6 alkyl”). In some embodiments, an alkyl group has 1 to 5 carbon atoms (“C1-5 alkyl”). In some embodiments, an alkyl group has 1 to 4 carbon atoms (“Ci^ alkyl”). In some embodiments, an alkyl group has 1 to 3 carbon atoms (“C1-3 alkyl”). In some embodiments, an alkyl group has 1 to 2 carbon atoms (“C1-2 alkyl”). In some embodiments, an alkyl group has 1 carbon atom (“Ci alkyl”). In some embodiments, an alkyl group has 2 to 6 carbon atoms (“C2.6 alkyl”). Examples of C, <> alkyl groups include methyl (Ci), ethyl (C2), propyl (C3) (e.g., w-propyl, isopropyl), butyl (C4) (e.g., w-butyl, tert-butyl, sec-butyl, isobutyl), pentyl (C5) (e.g., w-pentyl, 3-pentanyl, amyl, neopentyl, 3-methyl-2-butanyl, tertamyl), and hexyl (Cg) (e.g., w-hexyl). Additional examples of alkyl groups include w-heptyl (C7), w-octyl (C8), w-dodecyl (Ci2), and the like. Unless otherwise specified, each instance of an alkyl group is independently unsubstituted (an “unsubstituted alkyl”) or substituted (a “substituted alkyl”) with one or more substituents (e.g., halogen, such as F). In some embodiments, the alkyl group is an unsubstituted Ci-i2 alkyl (such as unsubstituted C 1,, alkyl, e.g., -CH3(Me), unsubstituted ethyl (Et), unsubstituted propyl (Pr, e.g., unsubstituted n-propyl (n-Pr), unsubstituted isopropyl (z-Pr)), unsubstituted butyl (Bu, e.g., unsubstituted w-butyl (w-Bu). unsubstituted tert-butyl (tert-Bu or / -Bu). unsubstituted sec-butyl (sec-Bu or s- Bn). unsubstituted isobutyl (z-Bu)). In some embodiments, the alkyl group is a substituted Ci-i2alkyl (such as substituted Ci „ alkyl, e g., -CH2F, -CHF2, -CF3, -CH2CH2F, -CH2CHF2, -CH2CF3, or benzyl (Bn)).
[0148]
[0050] The term “heteroalkyl” refers to an alkyl group, which further includes at least one heteroatom (e.g., 1, 2, 3, or 4 heteroatoms) selected from oxygen, nitrogen, or sulfur within (e.g., inserted between adjacent carbon atoms of) and / or placed at one or more terminal position(s) of the parent chain. In some embodiments, a heteroalkyl group refers to a saturated group having from 1 to 20 carbon atoms and 1 or more heteroatoms within the parent chain (“heteroCi-2o alkyl”). In some embodiments, a heteroalkyl group refers to a saturated group having from 1 to 12 carbon atoms and 1 or more heteroatoms within the parent chain (“heteroCi-i2alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 11 carbon atoms and 1 or more heteroatoms within the parent chain (“heteroCi-n alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 10 carbon atoms and 1 or more R0708.70180WO00 / R0708.70180US01 9 / 122
[0149] #14644680vl heteroatoms within the parent chain (“heteroCi-10 alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 9 carbon atoms and 1 or more heteroatoms within the parent chain (“heteroCi-9 alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 8 carbon atoms and 1 or more heteroatoms within the parent chain (“heteroCi-s alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 7 carbon atoms and 1 or more heteroatoms within the parent chain (“heteroCi-7 alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 6 carbon atoms and 1 or more heteroatoms within the parent chain (“heteroCi^ alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 5 carbon atoms and 1 or 2 heteroatoms within the parent chain (“heteroCi-5 alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 4 carbon atoms and lor 2 heteroatoms within the parent chain (“heteroCi^ alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 3 carbon atoms and 1 heteroatom within the parent chain (“heteroCi-s alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 to 2 carbon atoms and 1 heteroatom within the parent chain (“heteroCi-2 alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 1 carbon atom and 1 heteroatom (“heteroCi alkyl”). In some embodiments, a heteroalkyl group is a saturated group having 2 to 6 carbon atoms and 1 or 2 heteroatoms within the parent chain (“heteroC2-6 alkyl”). Unless otherwise specified, each instance of a heteroalkyl group is independently unsubstituted (an “unsubstituted heteroalkyl”) or substituted (a “substituted heteroalkyl”) with one or more substituents. In some embodiments, the heteroalkyl group is an unsubstituted heteroCi-12 alkyl. In some embodiments, the heteroalkyl group is a substituted heteroCi-12 alkyl.
[0150]
[0051] The term “alkenyl” refers to a radical of a straight-chain or branched hydrocarbon group having from 1 to 20 carbon atoms and one or more carbon-carbon double bonds (e.g., 1, 2, 3, or 4 double bonds). In some embodiments, an alkenyl group has 1 to 20 carbon atoms (“C1-20 alkenyl”). In some embodiments, an alkenyl group has 1 to 12 carbon atoms (“C1-12 alkenyl”). In some embodiments, an alkenyl group has 1 to 11 carbon atoms (“Ci-n alkenyl”). In some embodiments, an alkenyl group has 1 to 10 carbon atoms (“C1-10 alkenyl”). In some embodiments, an alkenyl group has 1 to 9 carbon atoms (“C1-9 alkenyl”). In some embodiments, an alkenyl group has 1 to 8 carbon atoms (“Ci-s alkenyl”). In some embodiments, an alkenyl group has 1 to 7 carbon atoms (“C1-7 alkenyl”). In some embodiments, an alkenyl group has 1 to 6 carbon atoms (“C1-6 alkenyl”). In some embodiments, an alkenyl group has 1 to 5 carbon atoms (“C1-5 alkenyl”). In some embodiments, an alkenyl group has 1 to 4 carbon atoms (“Ci^ alkenyl”). In some embodiments, an alkenyl group has 1 to 3 carbon atoms (“C1-3 alkenyl”). In some embodiments, an alkenyl group has 1 to 2 carbon atoms (“C1-2 alkenyl”). In some embodiments, an alkenyl group has 1 carbon atom (“Ci alkenyl”). The one or more carbon-carbon double bonds can be internal (such as in 2-butenyl) or terminal (such as in 1-butenyl). Examples of Ci^ alkenyl groups include methylidenyl (Ci), ethenyl (C2), 1-propenyl (C3), 2-propenyl (C3), 1-butenyl (C4), 2-butenyl (C4), butadienyl (C4), and the like. Examples of C1-6 alkenyl groups include the aforementioned C2-4 alkenyl groups as well as pentenyl (C5), pentadienyl (C5), hexenyl (Ce), and the like. Additional examples of alkenyl include heptenyl (C7), octenyl (Cs), octatrienyl (Cs), and the like. Unless otherwise specified, R0708.70180WO00 / R0708.70180US01 10 / 122
[0151] #14644680vl each instance of an alkenyl group is independently unsubstituted (an “unsubstituted alkenyl”) or substituted (a “substituted alkenyl”) with one or more substituents. In some embodiments, the alkenyl group is an unsubstituted C1.20 alkenyl. In some embodiments, the alkenyl group is a substituted C1.20 alkenyl. In an alkenyl group, a C=C double bond for which the stereochemistry is not specified (e.g., -CH=CHCH3 or
[0152]
[0153] ) may be in the ( / ■.')- or (^-configuration.
[0154]
[0052] The term “alkynyl” refers to a radical of a straight-chain or branched hydrocarbon group having from 1 to 20 carbon atoms and one or more carbon-carbon triple bonds (e.g., 1, 2, 3, or 4 triple bonds) (“C1-20 alkynyl”). In some embodiments, an alkynyl group has 1 to 10 carbon atoms (“C1-10 alkynyl”). In some embodiments, an alkynyl group has 1 to 9 carbon atoms (“C1-9 alkynyl”). In some embodiments, an alkynyl group has 1 to 8 carbon atoms (“Cns alkynyl”). In some embodiments, an alkynyl group has 1 to 7 carbon atoms (“C1-7 alkynyl”). In some embodiments, an alkynyl group has 1 to 6 carbon atoms (“C1-6 alkynyl”). In some embodiments, an alkynyl group has 1 to 5 carbon atoms (“C1-5 alkynyl”). In some embodiments, an alkynyl group has 1 to 4 carbon atoms (“C1.4 alkynyl”). In some embodiments, an alkynyl group has 1 to 3 carbon atoms (“C1.3 alkynyl”). In some embodiments, an alkynyl group has 1 to 2 carbon atoms (“C1.2 alkynyl”). In some embodiments, an alkynyl group has 1 carbon atom (“Ci alkynyl”). The one or more carbon-carbon triple bonds can be internal (such as in 2-butynyl) or terminal (such as in 1-butynyl). Examples of C1.4 alkynyl groups include, without limitation, methylidynyl (Ci), ethynyl (C2), 1-propynyl (C3), 2-propynyl (C3), 1-butynyl (C4), 2-butynyl (C4), and the like. Examples of C1-6 alkenyl groups include the aforementioned C2-4 alkynyl groups as well as pentynyl (C5), hexynyl (Ce), and the like. Additional examples of alkynyl include heptynyl (C7), octynyl (Cs), and the like. Unless otherwise specified, each instance of an alkynyl group is independently unsubstituted (an “unsubstituted alkynyl”) or substituted (a “substituted alkynyl”) with one or more substituents. In some embodiments, the alkynyl group is an unsubstituted C1-20 alkynyl. In some embodiments, the alkynyl group is a substituted C1-20 alkynyl.
[0155]
[0053] The term “carbocyclyl” or “carbocyclic” refers to a radical of a non-aromatic cyclic hydrocarbon group having from 3 to 14 ring carbon atoms (“C3-14 carbocyclyl”) and zero heteroatoms in the non-aromatic ring system. In some embodiments, a carbocyclyl group has 3 to 14 ring carbon atoms (“C3-14 carbocyclyl”). In some embodiments, a carbocyclyl group has 3 to 13 ring carbon atoms (“C3-13 carbocyclyl”). In some embodiments, a carbocyclyl group has 3 to 12 ring carbon atoms (“C3-12 carbocyclyl”). In some embodiments, a carbocyclyl group has 3 to 11 ring carbon atoms (“C3-11 carbocyclyl”). In some embodiments, a carbocyclyl group has 3 to 10 ring carbon atoms (“C3-10 carbocyclyl”). In some embodiments, a carbocyclyl group has 3 to 8 ring carbon atoms (“C3-8 carbocyclyl”). In some embodiments, a carbocyclyl group has 3 to 7 ring carbon atoms (“C3-7 carbocyclyl”). In some embodiments, a carbocyclyl group has 3 to 6 ring carbon atoms (“C3-6 carbocyclyl”). In some embodiments, a carbocyclyl group has 4 to 6 ring carbon atoms (“C4-6 carbocyclyl”). In some embodiments, a carbocyclyl group has 5 to 6 ring carbon atoms (“C5-6 carbocyclyl”). In some embodiments, a carbocyclyl group has 5 to 10 ring carbon atoms (“C5-10 carbocyclyl”). Exemplary C3-6 carbocyclyl groups include cyclopropyl (C3), cyclopropenyl (C3), R0708.70180WO00 / R0708.70180US01 11 / 122
[0156] #14644680vl cyclobutyl (C4), cyclobutenyl (C4), cyclopentyl (C5), cyclopentenyl (C5), cyclohexyl (Ce), cyclohexenyl (Ce), cyclohexadienyl (Ce), and the like. Exemplary C3-8 carbocyclyl groups include the aforementioned C3-6 carbocyclyl groups as well as cycloheptyl (C7), cycloheptenyl (C7), cycloheptadienyl (C7), cycloheptatrienyl (C7), cyclooctyl (Cs), cyclooctenyl (Cs), bicyclo[2.2.1]heptanyl (C7), bicyclo[2.2.2]octanyl (Cs), and the like. Exemplary C3-10 carbocyclyl groups include the aforementioned C3-8 carbocyclyl groups as well as cyclononyl (C>), cyclononenyl (C>), cyclodecyl (C10), cyclodecenyl (C10), octahydro- IH-indenyl (C>), decahydronaphthalenyl (C10), spiro[4.5]decanyl (C10), and the like. Exemplary C3-8 carbocyclyl groups include the aforementioned C3-10 carbocyclyl groups as well as cycloundecyl (Cn), spiro[5.5]undecanyl (Cn), cyclododecyl (C12), cyclododecenyl (C12), cyclotridecane (C13), cyclotetradecane (C14), and the like. As the foregoing examples illustrate, In some embodiments, the carbocyclyl group is either monocyclic (“monocyclic carbocyclyl”) or polycyclic (e.g., containing a fused, bridged or spiro ring system such as a bicyclic system (“bicyclic carbocyclyl”) or tricyclic system (“tricyclic carbocyclyl”)) and can be saturated or can contain one or more carbon-carbon double or triple bonds. “Carbocyclyl” also includes ring systems wherein the carbocyclyl ring, as defined above, is fused with one or more aryl or heteroaryl groups wherein the point of attachment is on the carbocyclyl ring, and in such instances, the number of carbons continue to designate the number of carbons in the carbocyclic ring system. Unless otherwise specified, each instance of a carbocyclyl group is independently unsubstituted (an “unsubstituted carbocyclyl”) or substituted (a “substituted carbocyclyl”) with one or more substituents. In some embodiments, the carbocyclyl group is an unsubstituted C3-14 carbocyclyl. In some embodiments, the carbocyclyl group is a substituted C3-14 carbocyclyl.
[0157]
[0054] In some embodiments, “carbocyclyl” is a monocyclic, saturated carbocyclyl group having from 3 to 14 ring carbon atoms (“C3-14 cycloalkyl”). In some embodiments, a cycloalkyl group has 3 to 10 ring carbon atoms (“C3-10 cycloalkyl”). In some embodiments, a cycloalkyl group has 3 to 8 ring carbon atoms (“C3-8 cycloalkyl”). In some embodiments, a cycloalkyl group has 3 to 6 ring carbon atoms (“C3-6 cycloalkyl”). In some embodiments, a cycloalkyl group has 4 to 6 ring carbon atoms (“C4-6 cycloalkyl”). In some embodiments, a cycloalkyl group has 5 to 6 ring carbon atoms (“C5-6 cycloalkyl”). In some embodiments, a cycloalkyl group has 5 to 10 ring carbon atoms (“C5-10 cycloalkyl”). Examples of C5-6 cycloalkyl groups include cyclopentyl (C5) and cyclohexyl (C5). Examples of C3-6 cycloalkyl groups include the aforementioned C5-6 cycloalkyl groups as well as cyclopropyl (C3) and cyclobutyl (C4). Examples of C3-8 cycloalkyl groups include the aforementioned C3-6 cycloalkyl groups as well as cycloheptyl (C7) and cyclooctyl (Cs). Unless otherwise specified, each instance of a cycloalkyl group is independently unsubstituted (an “unsubstituted cycloalkyl”) or substituted (a “substituted cycloalkyl”) with one or more substituents. In some embodiments, the cycloalkyl group is an unsubstituted C3-14 cycloalkyl. In some embodiments, the cycloalkyl group is a substituted C3-14 cycloalkyl. In some embodiments, the carbocyclyl includes 0, 1, or 2 C=C double bonds in the carbocyclic ring system, as valency permits.
[0158]
[0055] The term “heterocyclyl” or “heterocyclic” refers to a radical of a 3- to 14-membered non-aromatic ring system having ring carbon atoms and 1 to 4 ring heteroatoms, wherein each heteroatom is R0708.70180WO00 / R0708.70180US01 12 / 122
[0159] #14644680vl independently selected from nitrogen, oxygen, and sulfur (“3-14 membered heterocyclyl”). In heterocyclyl groups that contain one or more nitrogen atoms, the point of attachment can be a carbon or nitrogen atom, as valency permits. A heterocyclyl group can either be monocyclic (“monocyclic heterocyclyl”) or polycyclic (e.g., a fused, bridged or spiro ring system such as a bicyclic system (“bicyclic heterocyclyl”) or tricyclic system (“tricyclic heterocyclyl”)), and can be saturated or can contain one or more carbon-carbon double or triple bonds. Heterocyclyl polycyclic ring systems can include one or more heteroatoms in one or both rings. “Heterocyclyl” also includes ring systems wherein the heterocyclyl ring, as defined above, is fused with one or more carbocyclyl groups wherein the point of attachment is either on the carbocyclyl or heterocyclyl ring, or ring systems wherein the heterocyclyl ring, as defined above, is fused with one or more aryl or heteroaryl groups, wherein the point of attachment is on the heterocyclyl ring, and in such instances, the number of ring members continue to designate the number of ring members in the heterocyclyl ring system. Unless otherwise specified, each instance of heterocyclyl is independently unsubstituted (an “unsubstituted heterocyclyl”) or substituted (a “substituted heterocyclyl”) with one or more substituents. In some embodiments, the heterocyclyl group is an unsubstituted 3-14 membered heterocyclyl. In some embodiments, the heterocyclyl group is a substituted 3-14 membered heterocyclyl. In some embodiments, the heterocyclyl is substituted or unsubstituted, 3- to 7-membered, monocyclic heterocyclyl, wherein 1, 2, or 3 atoms in the heterocyclic ring system are independently oxygen, nitrogen, or sulfur, as valency permits.
[0160]
[0056] In some embodiments, a heterocyclyl group is a 5-10 membered non-aromatic ring system having ring carbon atoms and 1-4 ring heteroatoms, wherein each heteroatom is independently selected from nitrogen, oxygen, and sulfur (“5-10 membered heterocyclyl”). In some embodiments, a heterocyclyl group is a 5-8 membered non-aromatic ring system having ring carbon atoms and 1-4 ring heteroatoms, wherein each heteroatom is independently selected from nitrogen, oxygen, and sulfur (“5-8 membered heterocyclyl”). In some embodiments, a heterocyclyl group is a 5-6 membered non-aromatic ring system having ring carbon atoms and 1-4 ring heteroatoms, wherein each heteroatom is independently selected from nitrogen, oxygen, and sulfur (“5-6 membered heterocyclyl”). In some embodiments, the 5-6 membered heterocyclyl has 1-3 ring heteroatoms selected from nitrogen, oxygen, and sulfur. In some embodiments, the 5-6 membered heterocyclyl has 1-2 ring heteroatoms selected from nitrogen, oxygen, and sulfur. In some embodiments, the 5-6 membered heterocyclyl has 1 ring heteroatom selected from nitrogen, oxygen, and sulfur.
[0161]
[0057] Exemplary 3 -membered heterocyclyl groups containing 1 heteroatom include azirdinyl, oxiranyl, and thiiranyl. Exemplary 4-membered heterocyclyl groups containing 1 heteroatom include azetidinyl, oxetanyl, and thietanyl. Exemplary 5 -membered heterocyclyl groups containing 1 heteroatom include tetrahydrofuranyl, dihydrofuranyl, tetrahydrothiophenyl, dihydrothiophenyl, pyrrolidinyl, dihydropyrrolyl, and pyrrolyl-2,5-dione. Exemplary 5 -membered heterocyclyl groups containing 2 heteroatoms include dioxolanyl, oxathiolanyl and dithiolanyl. Exemplary 5 -membered heterocyclyl groups containing 3 heteroatoms include triazolinyl, oxadiazolinyl, and thiadiazolinyl. Exemplary 6-membered heterocyclyl groups containing 1 heteroatom include piperidinyl, tetrahydropyranyl, R0708.70180WO00 / R0708.70180US01 13 / 122
[0162] #14644680vl dihydropyridinyl, and thianyl. Exemplary 6-membered heterocyclyl groups containing 2 heteroatoms include piperazinyl, morpholinyl, dithianyl, and dioxanyl. Exemplary 6-membered heterocyclyl groups containing 3 heteroatoms include triazinyl. Exemplary 7-membered heterocyclyl groups containing 1 heteroatom include azepanyl, oxepanyl and thiepanyl. Exemplary 8-membered heterocyclyl groups containing 1 heteroatom include azocanyl, oxecanyl and thiocanyl. Exemplary bicyclic heterocyclyl groups include indolinyl, isoindolinyl, dihydrobenzofuranyl, dihydrobenzothienyl, tetrahydrobenzothienyl, tetrahydrobenzofuranyl, tetrahydroindolyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, decahydroquinolinyl, decahydroisoquinolinyl, octahydrochromenyl, octahydroisochromenyl, decahydronaphthyridinyl, decahydro- 1,8-naphthyridinyl, octahydropyrrolo[3,2-b]pyrrole, indolinyl, phthalimidyl, naphthalimidyl, chromanyl, chromenyl, lH-benzo[e][l,4]diazepinyl, 1,4,5,7-tetrahydro-pyrano[3,4-b]pyrrolyl, 5,6-dihydro-4H-furo[3,2-b]pyrrolyl, 6,7-dihydro-5H-furo[3,2-b]pyranyl, 5,7-dihydro-4H-thieno[2,3-c]pyranyl, 2,3-dihydro-lH-pyrrolo[2,3-b]pyridinyl, 2,3-dihydrofuro[2,3-b]pyridinyl, 4,5,6,7-tetrahydro-lH-pyrrolo[2,3-b]pyridinyl, 4,5,6,7-tetrahydrofuro[3,2-c]pyridinyl, 4,5,6,7-tetrahydrothieno[3,2-b]pyridinyl, l,2,3,4-tetrahydro-l,6-naphthyridinyl, and the like.
[0163]
[0058] The term “aryl” refers to a radical of a monocyclic or polycyclic (e.g., bicyclic or tricyclic) 4n+2 aromatic ring system (e.g., having 6, 10, or 1471 electrons shared in a cyclic array) having 6-14 ring carbon atoms and zero heteroatoms provided in the aromatic ring system (“Ce-i4 aryl”). In some embodiments, an aryl group has 6 ring carbon atoms (“Cg aryl”; e.g., phenyl). In some embodiments, an aryl group has 10 ring carbon atoms (“Cio aryl”; e.g., naphthyl such as 1-naphthyl and 2-naphthyl). In some embodiments, an aryl group has 14 ring carbon atoms (“C14 aryl”; e.g., anthracyl). “Aryl” also includes ring systems wherein the aryl ring, as defined above, is fused with one or more carbocyclyl or heterocyclyl groups wherein the radical or point of attachment is on the aryl ring, and in such instances, the number of carbon atoms continue to designate the number of carbon atoms in the aryl ring system. Unless otherwise specified, each instance of an aryl group is independently unsubstituted (an “unsubstituted aryl”) or substituted (a “substituted aryl”) with one or more substituents. In some embodiments, the aryl group is an unsubstituted Cg-i4 aryl. In some embodiments, the aryl group is a substituted Ce-i4 aryl.
[0164]
[0059] The term “heteroaryl” refers to a radical of a 5-14 membered monocyclic or polycyclic (e.g., bicyclic, tricyclic) 4n+2 aromatic ring system (e.g., having 6, 10, or 14 > electrons shared in a cyclic array) having ring carbon atoms and 1-4 ring heteroatoms provided in the aromatic ring system, wherein each heteroatom is independently selected from nitrogen, oxygen, and sulfur (“5-14 membered heteroaryl”). In heteroaryl groups that contain one or more nitrogen atoms, the point of attachment can be a carbon or nitrogen atom, as valency permits. Heteroaryl polycyclic ring systems can include one or more heteroatoms in one or both rings. “Heteroaryl” includes ring systems wherein the heteroaryl ring, as defined above, is fused with one or more carbocyclyl or heterocyclyl groups wherein the point of attachment is on the heteroaryl ring, and in such instances, the number of ring members continue to designate the number of ring members in the heteroaryl ring system. “Heteroaryl” also includes ring systems wherein the heteroaryl ring, as defined above, is fused with one or more aryl groups wherein the R0708.70180WO00 / R0708.70180US01 14 / 122
[0165] #14644680vl point of attachment is either on the aryl or heteroaryl ring, and in such instances, the number of ring members designates the number of ring members in the fused polycyclic (aryl / heteroaryl) ring system. Polycyclic heteroaryl groups wherein one ring does not contain a heteroatom (e.g., indolyl, quinolinyl, carbazolyl, and the like) the point of attachment can be on either ring, e.g., either the ring bearing a heteroatom (e.g., 2-indolyl) or the ring that does not contain a heteroatom (e.g., 5 -indolyl). In some embodiments, the heteroaryl is substituted or unsubstituted, 5- or 6-membered, monocyclic heteroaryl, wherein 1, 2, 3, or 4 atoms in the heteroaryl ring system are independently oxygen, nitrogen, or sulfur. In some embodiments, the heteroaryl is substituted or unsubstituted, 9- or 10-membered, bicyclic heteroaryl, wherein 1, 2, 3, or 4 atoms in the heteroaryl ring system are independently oxygen, nitrogen, or sulfur.
[0166]
[0060] In some embodiments, a heteroaryl group is a 5-10 membered aromatic ring system having ring carbon atoms and 1-4 ring heteroatoms provided in the aromatic ring system, wherein each heteroatom is independently selected from nitrogen, oxygen, and sulfur (“5-10 membered heteroaryl”). In some embodiments, a heteroaryl group is a 5-8 membered aromatic ring system having ring carbon atoms and 1-4 ring heteroatoms provided in the aromatic ring system, wherein each heteroatom is independently selected from nitrogen, oxygen, and sulfur (“5-8 membered heteroaryl”). In some embodiments, a heteroaryl group is a 5-6 membered aromatic ring system having ring carbon atoms and 1-4 ring heteroatoms provided in the aromatic ring system, wherein each heteroatom is independently selected from nitrogen, oxygen, and sulfur (“5-6 membered heteroaryl”). In some embodiments, the 5-6 membered heteroaryl has 1-3 ring heteroatoms selected from nitrogen, oxygen, and sulfur. In some embodiments, the 5-6 membered heteroaryl has 1-2 ring heteroatoms selected from nitrogen, oxygen, and sulfur. In some embodiments, the 5-6 membered heteroaryl has 1 ring heteroatom selected from nitrogen, oxygen, and sulfur. Unless otherwise specified, each instance of a heteroaryl group is independently unsubstituted (an “unsubstituted heteroaryl”) or substituted (a “substituted heteroaryl”) with one or more substituents. In some embodiments, the heteroaryl group is an unsubstituted 5-14 membered heteroaryl. In some embodiments, the heteroaryl group is a substituted 5-14 membered heteroaryl.
[0167]
[0061] Exemplary 5-membered heteroaryl groups containing 1 heteroatom include pyrrolyl, furanyl, and thiophenyl. Exemplary 5 -membered heteroaryl groups containing 2 heteroatoms include imidazolyl, pyrazolyl, oxazolyl, isoxazolyl, thiazolyl, and isothiazolyl. Exemplary 5 -membered heteroaryl groups containing 3 heteroatoms include triazolyl, oxadiazolyl, and thiadiazolyl. Exemplary 5 -membered heteroaryl groups containing 4 heteroatoms include tetrazolyl. Exemplary 6-membered heteroaryl groups containing 1 heteroatom include pyridinyl. Exemplary 6-membered heteroaryl groups containing 2 heteroatoms include pyridazinyl, pyrimidinyl, and pyrazinyl. Exemplary 6-membered heteroaryl groups containing 3 or 4 heteroatoms include triazinyl and tetrazinyl, respectively. Exemplary 7-membered heteroaryl groups containing 1 heteroatom include azepinyl, oxepinyl, and thiepinyl. Exemplary 5,6-bicyclic heteroaryl groups include indolyl, isoindolyl, indazolyl, benzotriazolyl, benzothiophenyl, isobenzothiophenyl, benzofuranyl, benzoisofuranyl, benzimidazolyl, benzoxazolyl, benzisoxazolyl, R0708.70180WO00 / R0708.70180US01 15 / 122
[0168] #14644680vl benzoxadiazolyl, benzthiazolyl, benzisothiazolyl, benzthiadiazolyl, indolizinyl, and purinyl. Exemplary 6,6-bicyclic heteroaryl groups include naphthyridinyl, pteridinyl, quinolinyl, isoquinolinyl, cinnolinyl, quinoxalinyl, phthalazinyl, and quinazolinyl. Exemplary tricyclic heteroaryl groups include phenanthridinyl, dibenzofuranyl, carbazolyl, acridinyl, phenothiazinyl, phenoxazinyl, and phenazinyl.
[0169]
[0062] The term “unsaturated bond” refers to a double or triple bond.
[0170]
[0063] The term “unsaturated” or “partially unsaturated” refers to a moiety that includes at least one double or triple bond.
[0171]
[0064] The term “saturated” or “fully saturated” refers to a moiety that does not contain a double or triple bond, e.g., the moiety only contains single bonds.
[0172]
[0065] Affixing the suffix “-ene” to a group indicates the group is a divalent moiety, e.g., alkylene is the divalent moiety of alkyl, alkenylene is the divalent moiety of alkenyl, alkynylene is the divalent moiety of alkynyl, heteroalkylene is the divalent moiety of heteroalkyl, heteroalkenylene is the divalent moiety of heteroalkenyl, heteroalkynylene is the divalent moiety of heteroalkynyl, carbocyclylene is the divalent moiety of carbocyclyl, heterocyclylene is the divalent moiety of heterocyclyl, arylene is the divalent moiety of aryl, and heteroarylene is the divalent moiety of heteroaryl.
[0173]
[0066] A group is optionally substituted unless expressly provided otherwise. The term “optionally substituted” refers to being substituted or unsubstituted. In some embodiments, alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl groups are optionally substituted. “Optionally substituted” refers to a group which is substituted or unsubstituted (e.g., “substituted” or “unsubstituted” alkyl, “substituted” or “unsubstituted” alkenyl, “substituted” or “unsubstituted” alkynyl, “substituted” or “unsubstituted” heteroalkyl, “substituted” or “unsubstituted” heteroalkenyl, “substituted” or “unsubstituted” heteroalkynyl, “substituted” or “unsubstituted” carbocyclyl, “substituted” or “unsubstituted” heterocyclyl, “substituted” or “unsubstituted” aryl or “substituted” or “unsubstituted” heteroaryl group). In general, the term “substituted” means that at least one hydrogen present on a group is replaced with a permissible substituent, e.g., a substituent which upon substitution results in a stable compound, e.g., a compound which does not spontaneously undergo transformation such as by rearrangement, cyclization, elimination, or other reaction. Unless otherwise indicated, a “substituted” group has a substituent at one or more substitutable positions of the group, and when more than one position in any given structure is substituted, the substituent is either the same or different at each position. The term “substituted” is contemplated to include substitution with all permissible substituents of organic compounds, and includes any of the substituents described herein that results in the formation of a stable compound. The present disclosure contemplates any and all such combinations in order to arrive at a stable compound. For purposes of this disclosure, heteroatoms such as nitrogen may have hydrogen substituents and / or any suitable substituent as described herein which satisfy the valencies of the heteroatoms and results in the formation of a stable moiety. The disclosure is not limited in any manner by the exemplary substituents described herein.
[0174]
[0067] Exemplary carbon atom substituents include halogen, -CN, -NO2, “Ns, -SO2H, -SO3H, -OH, -ORaa, -ON(Rbb)2, -N(Rbb)2, -N(Rbb)3+X, -N(ORcc)Rbb, -SH, -SRaa, -SSRCC, -C(=O)Raa, -CO2H, R0708.70180WO00 / R0708.70180US01 16 / 122
[0175] #14644680vl -CHO, -C(ORCC)2, -CO2Raa, -OC(=O)Raa, -OCO2Raa, -C(=O)N(Rbb)2, -OC(=O)N(Rbb)2, -NRbbC(=O)Raa, -NRbbCO2Raa, -NRbbC(=O)N(Rbb)2, -C(=NRbb)Raa, -C(=NRbb)ORaa, -OC(=NRbb)Raa, -OC(=NRbb)ORaa, -C(=NRbb)N(Rbb)2, -OC(=NRbb)N(Rbb)2, -NRbbC(=NRbb)N(Rbb)2, -C(=O)NRbbSO2Raa, -NRbbSO2Raa, -SO2N(Rbb)2, -SO2Raa, -SO2ORaa, -OSO2Raa, -S(=O)Raa, -OS(=O)Raa, — Si(Raa)3, -OSi(Raa)3-C(=S)N(Rbb)2, -C(=O)SRaa, -C(=S)SRaa, -SC(=S)SRaa, -SC(=O)SRaa, -OC(=O)SRaa, -SC(=O)ORaa, -SC(=O)Raa, -P(=O)(Raa)2, -P(=O)(ORCC)2, -OP(=O)(Raa)2, -OP(=O)(ORCC)2, -P(=O)(N(Rbb)2)2, -OP(=O)(N(Rbb)2)2, -NRbbP(=O)(Raa)2, -NRbbP(=O)(ORcc)2, -NRbbP(=O)(N(Rbb)2)2, -P(RCC)2, -P(ORCC)2, -P(RCC)3+X, -P(ORCC)3+X, -P(RCC)4, -P(ORCC)4,-OP(RCC)2, -OP(RCC)3+X, -OP(ORCC)2, -OP(ORCC)3X, -OP(RCC)4, -OP(ORcc)4,-B(Raa)2, -B(ORCC)2, -BRaa(ORcc), Ci-2o alkyl, C1-20 perhaloalkyl, C1-20 alkenyl, C1-20 alkynyl, heteroCi-2o alkyl, heteroCi-2o alkenyl, heteroCi-2o alkynyl, C3-10 carbocyclyl, 3-14 membered heterocyclyl, Ce-i4aryl, and 5-14 membered heteroaryl, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 Rddgroups; wherein X is a counterion;
[0176] or two geminal hydrogens on a carbon atom are replaced with the group =0, =S, =NN(Rbb)2, =NNRbbC(=0)Raa, =NNRbbC(=0)0Raa, =NNRbbS(=0)2Raa, =NRbb, or =NORCC;
[0177] wherein:
[0178] each instance of Raais, independently, selected from Ci_2o alkyl, Ci_2o perhaloalkyl, Ci_2o alkenyl, Ci_2o alkynyl, hctcroCj _20alkyl, hctcroC _20alkenyl, hctcroCj _20alkynyl, C3-10 carbocyclyl, 3-14 membered heterocyclyl, Ce-i4aryl, and 5-14 membered heteroaryl, or two Raagroups are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each of the alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 Rddgroups;
[0179] each instance of Rbbis, independently, selected from hydrogen, -OH, -ORaa, -N(RCC)2, -CN, -C(=O)Raa, -C(=0)N(RCC)2, -CO2Raa, -SO2Raa, -C(=NRcc)0Raa, -C(=NRCC)N(RCC)2, -SO2N(RCC)2, -SO2RCC, -SO2ORCC, -SORaa, -C(=S)N(RCC)2, -C(=O)SRCC, -C(=S)SRCC, -P(=O)(Raa)2, -P(=O)(ORCC)2, -P(=O)(N(RCC)2)2, CI-20alkyl, Ci_20perhaloalkyl, Ci_20alkenyl, Ci-2o alkynyl, heteroCi-2oalkyl, heteroCi-2oalkenyl, heteroCi-2oalkynyl, C3-10 carbocyclyl, 3-14 membered heterocyclyl, Ce-i4 aryl, and 5-14 membered heteroaryl, or two Rbbgroups are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 Rddgroups;
[0180] each instance of Rccis, independently, selected from hydrogen, C1-20 alkyl, C1-20 perhaloalkyl, Ci_2o alkenyl, Ci_2o alkynyl, hctcroCj _20alkyl, hctcroCj _20alkenyl, hctcroCj _20alkynyl, C3-10 carbocyclyl, 3-14 membered heterocyclyl, Ce-i4aryl, and 5-14 membered heteroaryl, or two Rccgroups are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, R0708.70180WO00 / R0708.70180US01 17 / 122
[0181] #14644680vl heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 Rddgroups;
[0182] each instance of Rddis, independently, selected from halogen, -CN, -NO2, -N3, -SO2H, -SO3H, -OH, -ORee, -ON(Rff)2, -N(Rff)2, -N(R'')3X. -N(ORee)Rff, -SH, -SRee, -SSRee, -C(=O)Ree, -CO2H, -CO2Ree, -OC(=O)Ree, -OCO2Ree, -C(=O)N(Rff)2, -OC(=O)N(Rff)2, -NRffC(=O)Ree, -NRffCO2Ree, -NRffC(=O)N(Rff)2, -C(=NRff)ORee, -OC(=NRff)Ree, -OC(=NRff)ORee, -C(=NRff)N(Rff)2, -OC(=NRff)N(Rff)2, -NRffC(=NRff)N(Rff)2, -NRffSO2Ree, -SO2N(Rff)2, -SO2Ree, -SO2ORee, -OSO2Ree, -S(=O)Ree, -Si(Ree)3, -OSi(Ree)3, -C(=S)N(Rff)2, -C(=O)SRee, -C(=S)SRee, -SC(=S)SRee, -P(=O)(ORee)2, -P(=O)(Ree)2, -OP(=O)(Ree)2, -OP(=O)(ORee)2, Ci-10 alkyl, C1-10 perhaloalkyl, C1-10 alkenyl, C1-10 alkynyl, heteroCi-ioalkyl, heteroCi-ioalkenyl, heteroCi-ioalkynyl, C3-10 carbocyclyl, 3-10 membered heterocyclyl, Ce-io aryl, and 5-10 membered heteroaryl, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R88groups, or two geminal Rddsubstituents are joined to form =0 or =S; wherein X is a counterion;
[0183] each instance of Reeis, independently, selected from C1-10 alkyl, C1-10 perhaloalkyl, C1-10 alkenyl, Ci-w alkynyl, heteroCi-10 alkyl, heteroCi-10 alkenyl, heteroCi-10 alkynyl, C3-10 carbocyclyl, Cg-io aryl, 3-10 membered heterocyclyl, and 3-10 membered heteroaryl, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R88groups;
[0184] each instance of Rffis, independently, selected from hydrogen, Ci-io alkyl, Ci-io perhaloalkyl, C1-10 alkenyl, C1-10 alkynyl, heteroCi-10 alkyl, heteroCi-10 alkenyl, heteroCi-10 alkynyl, C3-10 carbocyclyl, 3-10 membered heterocyclyl, Ce-io aryl, and 5-10 membered heteroaryl, or two Rffgroups are joined to form a 3-10 membered heterocyclyl or 5-10 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 R88groups;
[0185] each instance of R88is, independently, halogen, -CN, -NO2, -N3, -SO2H, -SO3H, -OH, -OC1-6 alkyl, -ON(CI-6 alkyl)2, -N(CI-6 alkyl)2, -N(CI-6 alkyl)3X. -NH(CI-6 alkyl)2X.
[0186] -NH2(CI -6alkyl)+X ", -NH3X. -N(OCI_6alkyl)(Ci-6 alkyl), -N(OH)(Ci^ alkyl), -NH(OH), -SH, -SCi_6 alkyl, -SS(Ci_6alkyl), -C(=O)(CI„ alkyl), -CO2H, -CO2(CI_6 alkyl), -OC(=O)(C1 6 alkyl), -OCO2(CI„ alkyl), -C(=O)NH2, -C(=O)N(CI„ alkyl)2, -OC(=O)NH(Ci6alkyl), -NHC(=O)( CKalkyl), -N(Ci^ alkyl)C(=O)( C,, alkyl), -NHCO2(CI„ alkyl), -NHC(=O)N(Ci _6 alkyl)2, -NHC(=O)NH(CI„ alkyl), -NHC(=0)NH2, -C(=NH)O(Ci6alkyl), -OC(=NH)(Cj „ alkyl), -OC(=NH)OCj „ alkyl, -C(=NH)N(Cj „ alkyl)2, -C(=NH)NH(Cj „ alkyl), -C(=NH)NH2, -OC(=NH)N(CI6alkyl)2, -OC(NH)NH(Ci6alkyl), -OC(NH)NH2, -NHC(NH)N(Cj „ alkyl)2, -NHC(=NH)NH2, -NHSO2(Cj „ alkyl), -SO2N(Cj „ alkyl)2, -SO2NH(Cj „ alkyl), -SO2NH2, -SO2C,,, alkyl, -SO2OCi6alkyl, -OSO2Ci6alkyl, -SOCj „ R0708.70180WO00 / R0708.70180US01 18 / 122
[0187] #14644680vl alkyl, -Si(Ci^ alkyl)3, -OSi(Ci^ alkyl)3-C(=S)N(Ci^ alkyl)2, C(=S)NH(Cj alkyl), C(=S)NH2, C(=O)S(C alkyl), -C(=S)SCI„ alkyl, SC(=S )SC,, alkyl, P(=O)(OC,, alkyl)2, -P(=O)(C1-6alkyl)2, -OP(=O)(C1-6alkyl)2, -OP(=O)(OC1-6alkyl)2, CMO alkyl, Ci_10perhaloalkyl, Ci-io alkenyl, Ci-io alkynyl, heteroCuo alkyl, heteroCi-10 alkenyl, heteroCi-10 alkynyl, C3-10 carbocyclyl, Ce-io aryl, 3-10 membered heterocyclyl, or 5-10 membered heteroaryl; or two geminal R88substituents can be joined to form =0 or =S; and
[0188] each X is a counterion.
[0189]
[0068] In some embodiments, each carbon atom substituent is independently halogen, substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-6 alkyl, -ORaa, -SRaa, -N(Rbb)2, -CN, -SCN, -NO2, -C(=O)Raa, -CO2Raa, -C(=O)N(Rbb)2, -OC(=O)Raa, -OCO2Raa, -OC(=O)N(Rbb)2, -NRbbC(=O)Raa, -NRbbCO2Raa, or -NRbbC(=O)N(Rbb)2. In some embodiments, each carbon atom substituent is independently halogen, substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-w alkyl, -ORaa, -SRaa, -N(Rbb)2, -CN, -SCN, -NO2, -C(=O)Raa, -CO2Raa, -C(=O)N(Rbb)2, -OC(=O)Raa, -OCO2Raa, -OC(=O)N(Rbb)2, -NRbbC(=O)Raa, -NRbbCO2Raa, or -NRbbC(=O)N(Rbb)2, wherein Raais hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-w alkyl, an oxygen protecting group (e.g., silyl, TBDPS, TBDMS, TIPS, TES, TMS, MOM, THP, t-Bu, Bn, allyl, acetyl, pivaloyl, or benzoyl) when attached to an oxygen atom, or a sulfur protecting group (e.g., acetamidomethyl, t-Bu, 3 -nitro-2 -pyridine sulfenyl, 2-pyridine-sulfenyl, or triphenylmethyl) when attached to a sulfur atom; and each Rbbis independently hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-io alkyl, or a nitrogen protecting group (e.g., Bn, Boc, Cbz, Fmoc, trifluoroacetyl, triphenylmethyl, acetyl, or Ts). In some embodiments, each carbon atom substituent is independently halogen, substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-6 alkyl, -ORaa, -SRaa, -N(Rbb)2, -CN, -SCN, or -NO2. In some embodiments, each carbon atom substituent is independently halogen, substituted (e.g., substituted with one or more halogen moieties) or unsubstituted Ci-w alkyl, -ORaa, -SRaa, -N(Rbb)2, -CN, -SCN, or -NO2, wherein Raais hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-io alkyl, an oxygen protecting group (e.g., silyl, TBDPS, TBDMS, TIPS, TES, TMS, MOM, THP, t-Bu, Bn, allyl, acetyl, pivaloyl, or benzoyl) when attached to an oxygen atom, or a sulfur protecting group (e.g., acetamidomethyl, t-Bu, 3-nitro-2-pyridine sulfenyl, 2-pyridine-sulfenyl, or triphenylmethyl) when attached to a sulfur atom; and each Rbbis independently hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-io alkyl, or a nitrogen protecting group (e.g., Bn, Boc, Cbz, Fmoc, trifluoroacetyl, triphenylmethyl, acetyl, or Ts).
[0190]
[0069] In some embodiments, the molecular weight of a carbon atom substituent is lower than 250, lower than 200, lower than 150, lower than 100, or lower than 50 g / mol. In some embodiments, a carbon atom substituent consists of carbon, hydrogen, fluorine, chlorine, bromine, iodine, oxygen, sulfur, nitrogen, and / or silicon atoms. In some embodiments, a carbon atom substituent consists of carbon, hydrogen, fluorine, chlorine, bromine, iodine, oxygen, sulfur, and / or nitrogen atoms. In some embodiments, a carbon atom substituent consists of carbon, hydrogen, fluorine, chlorine, bromine, and / or iodine atoms. R0708.70180WO00 / R0708.70180US01 19 / 122
[0191] #14644680vl In some embodiments, a carbon atom substituent consists of carbon, hydrogen, fluorine, and / or chlorine atoms.
[0192]
[0070] The term “halo” or “halogen” refers to fluorine (fluoro, -F), chlorine (chloro, -Cl), bromine (bromo, -Br), or iodine (iodo, -I).
[0193]
[0071] The term “hydroxyl” or “hydroxy” refers to the group -OH. The term “substituted hydroxyl” or “substituted hydroxy,” by extension, refers to a hydroxyl group wherein the oxygen atom directly attached to the parent molecule is substituted with a group other than hydrogen, and includes groups selected from -ORaa, -0N(Rbb)2, -OC(=O)SRaa, -OC(=O)Raa, -OCO2Raa, -OC(=O)N(Rbb)2, -OC(=NRbb)Raa, -OC(=NRbb)ORaa, -OC(=NRbb)N(Rbb)2, -OS(=O)Raa, -OSO2Raa, -OSi(Raa)3, -OP(RCC)2, -OP(RCC)3+X, -OP(ORCC)2, -OP(ORCC)3X, -OP(=O)(Raa)2, -OP(=O)(ORCC)2, and -OP(=O)(N(Rbb))2, wherein X, Raa, Rbb, and Rccare as defined herein.
[0194]
[0072] The term “amino” refers to the group -NH2. The term “substituted amino,” by extension, refers to a monosubstituted amino, a disubstituted amino, or a trisubstituted amino. In some embodiments, the “substituted amino” is a monosubstituted amino or a disubstituted amino group.
[0195]
[0073] The term “monosubstituted amino” refers to an amino group wherein the nitrogen atom directly attached to the parent molecule is substituted with one hydrogen and one group other than hydrogen, and includes groups selected from -NH(Rbb), -NHC(=O)Raa, -NHCO2Raa, -NHC(=O)N(Rbb)2, -NHC(=NRbb)N(Rbb)2, -NHSO2Raa, -NHP(=O)(ORCC)2, and -NHP(=O)(N(Rbb)2)2, wherein Raa, Rbband Rccare as defined herein, and wherein Rbbof the group -NH(Rbb) is not hydrogen.
[0196]
[0074] The term “disubstituted amino” refers to an amino group wherein the nitrogen atom directly attached to the parent molecule is substituted with two groups other than hydrogen, and includes groups selected from -N(Rbb)2, -NRbbC(=O)Raa, -NRbbCO2Raa, -NRbbC(=O)N(Rbb)2, -NRbbC(=NRbb)N(Rbb)2, -NRbbSO2Raa, -NRbbP(=O)(ORcc)2, and -NRbbP(=O)(N(Rbb)2)2, wherein Raa, Rbb, and Rccare as defined herein, with the proviso that the nitrogen atom directly attached to the parent molecule is not substituted with hydrogen.
[0197]
[0075] The term “trisubstituted amino” refers to an amino group wherein the nitrogen atom directly attached to the parent molecule is substituted with three groups, and includes groups selected from -N(Rbb)3and -N(Rhh)3X. wherein Rbband X are as defined herein.
[0198]
[0076] The term “acyl” refers to a group having the general formula -C(=O)RX1, -C(=O)ORX1, -C(=O)-O-C(=O)RX1, -C(=O)SRX1, -C(=O)N(RX1)2, -C(=S)RX1, -C(=S)N(RX1)2, and -C(=S)S(RX1), -C(=NRX1)RX1, -C(=NRX1)ORX1, -C(=NRX1)SRX1, and -C(=NRX1)N(RX1)2, wherein RX1is hydrogen; halogen; substituted or unsubstituted hydroxyl; substituted or unsubstituted thiol; substituted or unsubstituted amino; substituted or unsubstituted acyl, cyclic or acyclic, substituted or unsubstituted, branched or unbranched aliphatic; cyclic or acyclic, substituted or unsubstituted, branched or unbranched heteroaliphatic; cyclic or acyclic, substituted or unsubstituted, branched or unbranched alkyl; cyclic or acyclic, substituted or unsubstituted, branched or unbranched alkenyl; substituted or unsubstituted alkynyl; substituted or unsubstituted aryl, substituted or unsubstituted heteroaryl, aliphaticoxy, heteroaliphaticoxy, alkyloxy, heteroalkyloxy, aryloxy, heteroaryloxy, aliphaticthioxy,
[0199] R0708.70180WO00 / R0708.70180US01 20 / 122
[0200] #14644680vl heteroaliphaticthioxy, alkylthioxy, heteroalkylthioxy, arylthioxy, heteroarylthioxy, mono- or di-aliphaticamino, mono- or di- heteroaliphaticamino, mono- or di- alkylamino, mono- or diheteroalkylamino, mono- or di-arylamino, or mono- or di-heteroarylamino; or two RX1groups taken together form a 5- to 6-membered heterocyclic ring. Exemplary acyl groups include aldehydes (-CHO), carboxylic acids (-CO2H), ketones, acyl halides, esters, amides, imines, carbonates, carbamates, and ureas. Acyl substituents include, but are not limited to, any of the substituents described herein, that result in the formation of a stable moiety (e.g., aliphatic, alkyl, alkenyl, alkynyl, heteroaliphatic, heterocyclic, aryl, heteroaryl, acyl, oxo, imino, thiooxo, cyano, isocyano, amino, azido, nitro, hydroxyl, thiol, halo, aliphaticamino, heteroaliphaticamino, alkylamino, heteroalkylamino, arylamino, heteroarylamino, alkylaryl, arylalkyl, aliphaticoxy, heteroaliphaticoxy, alkyloxy, heteroalkyloxy, aryloxy, heteroaryloxy, aliphaticthioxy, heteroaliphaticthioxy, alkylthioxy, heteroalkylthioxy, arylthioxy, heteroarylthioxy, acyloxy, and the like, each of which may or may not be further substituted).
[0201]
[0077] The term “carbonyl” refers to a group wherein the carbon directly attached to the parent molecule is sp2hybridized, and is substituted with an oxygen, nitrogen or sulfur atom, e.g., a group selected from ketones (-C(=O)Raa), carboxylic acids (-CO2H), aldehydes (-CHO), esters (-CCER13, -C(=O)SRaa, -C(=S)SRaa), amides (-C(=O)N(Rbb)2, -C(=O)NRbbSO2Raa, -C(=S)N(Rbb)2), and imines (-C(=NRbb)Raa, -C(=NRbb)ORaa), -C(=NRbb)N(Rbb)2), wherein Raaand Rbbare as defined herein.
[0202]
[0078] Nitrogen atoms can be substituted or unsubstituted as valency permits, and include primary, secondary, tertiary, and quaternary nitrogen atoms. Exemplary nitrogen atom substituents include hydrogen, -OH, -ORaa, -N(RCC)2, -CN, -C(=O)Raa, -C(=O)N(RCC)2, -CO2Raa, -SO2Raa, -C(=NRbb)Raa, -C(=NRcc)ORaa, -C(=NRCC)N(RCC)2, -SO2N(RCC)2, -SO2RCC, -SO2ORCC, -SORaa, -C(=S)N(RCC)2, -C(=O)SRCC, -C(=S)SRCC, -P(=O)(ORCC)2, -P(=O)(Raa)2, -P(=O)(N(RCC)2)2, Ci 20 alkyl, Ci 20 perhaloalkyl, C1-20 alkenyl, C1-20 alkynyl, hetero C1-20 alkyl, hetero C1-20 alkenyl, hetero C1-20 alkynyl, C3-10 carbocyclyl, 3-14 membered heterocyclyl, Ce-i4 aryl, and 5-14 membered heteroaryl, or two Rccgroups attached to an N atom are joined to form a 3-14 membered heterocyclyl or 5-14 membered heteroaryl ring, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 Rddgroups, and wherein Raa, Rbb, Rccand Rddare as defined above.
[0203]
[0079] In some embodiments, each nitrogen atom substituent is independently substituted (e.g., substituted with one or more halogen) or unsubstituted C1-6 alkyl, -C(=O)Raa, -CC R”, -C(=O)N(Rbb)2, or a nitrogen protecting group. In some embodiments, each nitrogen atom substituent is independently substituted (e.g., substituted with one or more halogen) or unsubstituted C1-10 alkyl, -C(=O)Raa, -CC>2Raa, -C(=O)N(Rbb)2, or a nitrogen protecting group, wherein Raais hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted C1-10 alkyl, or an oxygen protecting group when attached to an oxygen atom; and each Rbbis independently hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted CMO alkyl, or a nitrogen protecting group. In some embodiments, each nitrogen atom substituent is independently substituted (e.g., substituted with one or more halogen) or unsubstituted C1.6 alkyl or a nitrogen protecting group.
[0204] R0708.70180WO00 / R0708.70180US01 21 / 122
[0205] #14644680vl
[0080] In some embodiments, the substituent present on the nitrogen atom is a nitrogen protecting group (also referred to herein as an “amino protecting group”). Nitrogen protecting groups include -OH, -ORaa, -N(RCC)2, -C(=O)Raa, -C(=O)N(RCC)2, -CO2Raa, -SO2Raa, -C(=NRcc)Raa, -C(=NRcc)ORaa, -C(=NRCC)N(RCC)2, -SO2N(RCC)2, -SO2RCC, -SO2ORCC, -SORaa, -C(=S)N(RCC)2, -C(=O)SRCC, -C(=S)SRCC, Ci-io alkyl (e.g., aralkyl, heteroaralkyl), C1-20 alkenyl, C1-20 alkynyl, hetero C1-20 alkyl, hetero C1-20 alkenyl, hetero C1-20 alkynyl, C3-10 carbocyclyl, 3-14 membered heterocyclyl, Ce-i4 aryl, and 5-14 membered heteroaryl groups, wherein each alkyl, alkenyl, alkynyl, heteroalkyl, heteroalkenyl, heteroalkynyl, carbocyclyl, heterocyclyl, aralkyl, aryl, and heteroaryl is independently substituted with 0, 1, 2, 3, 4, or 5 Rddgroups, and wherein Raa, Rbb, Rccand Rddare as defined herein. Nitrogen protecting groups are well known in the art and include those described in detail in Protecting Groups in Organic Synthesis, T. W. Greene and P. G. M. Wuts, 3rdedition, John Wiley & Sons, 1999, incorporated herein by reference.
[0206]
[0081] For example, In some embodiments, at least one nitrogen protecting group is an amide group (e.g., a moiety that include the nitrogen atom to which the nitrogen protecting groups (e.g., -C(=O)Raa) is directly attached). In certain such embodiments, each nitrogen protecting group, together with the nitrogen atom to which the nitrogen protecting group is attached, is independently selected from the group consisting of formamide, acetamide, chloroacetamide, trichloroacetamide, trifluoroacetamide, phenylacetamide, 3-phenylpropanamide, picolinamide, 3 -pyridylcarboxamide, JV-benzoylphenylalanyl derivatives, benzamide, p-phenylbenzamide, o-nitophenylacetamide, o-nitrophenoxyacetamide, acetoacetamide, (N’ -dithiobenzyloxyacylamino)acetamide, 3-(p-hydroxyphenyl)propanamide, 3-(o-nitrophenyl)propanamide, 2-methyl-2-(o-nitrophenoxy)propanamide, 2-methyl-2-(o-phenylazophenoxy)propanamide, 4-chlorobutanamide, 3-methyl-3-nitrobutanamide, o-nitrocinnamide, / V-acetylmethionine derivatives, o-nitrobenzamide, and o-(benzoyloxymethyl)benzamide.
[0207]
[0082] In some embodiments, at least one nitrogen protecting group is a carbamate group (e.g., a moiety that include the nitrogen atom to which the nitrogen protecting groups (e.g., -C(=O)ORaa) is directly attached). In certain such embodiments, each nitrogen protecting group, together with the nitrogen atom to which the nitrogen protecting group is attached, is independently selected from the group consisting of methyl carbamate, ethyl carbamate, 9-fluorenylmethyl carbamate (Fmoc), 9-(2-sulfo)fluorenyhnethyl carbamate, 9-(2,7-dibromo)fluoroenyhnethyl carbamate, 2,7-di-t-butyl-[9-(10,10-dioxo-10,10,10,10-tetrahydrothioxanthyl)]methyl carbamate (DBD-Tmoc), 4-methoxyphenacyl carbamate (Phenoc), 2,2,2-trichloroethyl carbamate (Troc), 2-trimethylsilylethyl carbamate (Teoc), 2-phenylethyl carbamate (hZ), l-(l-adamantyl)-l -methylethyl carbamate (Adpoc), 1,1 -dimethyl -2 -haloethyl carbamate, 1,1-dimethyl-2,2-dibromoethyl carbamate (DB-t-BOC), 1,1 -dimethyl -2, 2, 2-trichloroethyl carbamate (TCBOC), 1-methyl-l-(4-biphenylyl)ethyl carbamate (Bpoc), l-(3,5-di-t-butylphenyl)-l-methylethyl carbamate (t-Bumeoc), 2-(2'- and 4'-pyridyl)ethyl carbamate (Pyoc), 2-(N, N-dicyclohexylcarboxamido)ethyl carbamate, / -butyl carbamate (BOC or Boc), 1-adamantyl carbamate (Adoc), vinyl carbamate (Voc), allyl carbamate (Alloc), 1 -isopropylallyl carbamate (Ipaoc), cinnamyl carbamate (Coc), 4-nitrocinnamyl carbamate (Noc), 8-quinolyl carbamate, N-hydroxypiperidinyl carbamate, alkyldithio carbamate, benzyl R0708.70180WO00 / R0708.70180US01 22 / 122
[0208] #14644680vl carbamate (Cbz), p-methoxybenzyl carbamate (Moz), p-nitobenzyl carbamate, p-bromobenzyl carbamate, p-chlorobenzyl carbamate, 2,4-dichlorobenzyl carbamate, 4-methylsulfinylbenzyl carbamate (Msz), 9-anthrylmethyl carbamate, diphenylmethyl carbamate, 2-methylthioethyl carbamate, 2-methylsulfonylethyl carbamate, 2-(p-toluenesulfonyl)ethyl carbamate, [2-(l,3-dithianyl)]methyl carbamate (Dmoc), 4-methylthiophenyl carbamate (Mtpc), 2,4-dimethylthiophenyl carbamate (Bmpc), 2-phosphonioethyl carbamate (Peoc), 2-triphenylphosphonioisopropyl carbamate (Ppoc), 1,1 -dimethyl -2-cyanoethyl carbamate, m-chloro-p-acyloxybenzyl carbamate, p-(dihydroxyboryl)benzyl carbamate, 5-benzisoxazolylmethyl carbamate, 2-(trifluoromethyl)-6-chromonylmethyl carbamate (Tcroc), m-nitrophenyl carbamate, 3, 5 -dimethoxybenzyl carbamate, o-nitrobenzyl carbamate, 3,4-dimethoxy-6-nitrobenzyl carbamate, phenyl(o-nitrophenyl)methyl carbamate, / -am l carbamate,. S'-bcnzyl thiocarbamate, p-cyanobenzyl carbamate, cyclobutyl carbamate, cyclohexyl carbamate, cyclopentyl carbamate, cyclopropylmethyl carbamate, p-decyloxybenzyl carbamate, 2,2-dimethoxyacylvinyl carbamate, o-(A'. A'-dimcthylcarboxamido)bcnzyl carbamate, I. I -dimcthy l-3-(AC / V-dimethylcarboxamido)propyl carbamate, 1,1-dimethylpropynyl carbamate, di(2-pyridyl)methyl carbamate, 2-furanylmethyl carbamate, 2-iodoethyl carbamate, isoborynl carbamate, isobutyl carbamate, isonicotinyl carbamate, p-(p ’-methoxyphenylazo)benzyl carbamate, 1 -methylcyclobutyl carbamate, 1-methylcyclohexyl carbamate, 1 -methyl- 1 -cyclopropylmethyl carbamate, 1 -methyl- 1 -(3,5-dimethoxyphenyl)ethyl carbamate, 1 -methyl- l-(p-phenylazophenyl)ethyl carbamate, 1-methyl-l-phenylethyl carbamate, 1 -methyl- l-(4-pyridyl)ethyl carbamate, phenyl carbamate, p-(phenylazo)benzyl carbamate, 2.4.6-tri - / -bnty Iphcny I carbamate, 4-(trimethylammonium)benzyl carbamate, and 2,4,6-trimethylbenzyl carbamate.
[0209]
[0083] In some embodiments, at least one nitrogen protecting group is a sulfonamide group (e.g., a moiety that include the nitrogen atom to which the nitrogen protecting groups (e.g., -S(=O)2Raa) is directly attached). In certain such embodiments, each nitrogen protecting group, together with the nitrogen atom to which the nitrogen protecting group is attached, is independently selected from the group consisting of p-toluenesulfonamide (Ts), benzenesulfonamide, 2,3,6-trimethyl-4-methoxybenzenesulfonamide (Mtr), 2,4,6-trimethoxybenzenesulfonamide (Mtb), 2,6-dimethyl-4-methoxybenzenesulfonamide (Pme), 2, 3,5,6-tetramethyl-4-methoxybenzenesulfonamide (Mte), 4-methoxybenzenesulfonamide (Mbs), 2,4,6-trimethylbenzenesulfonamide (Mts), 2,6-dimethoxy-4-methylbenzenesulfonamide (iMds), 2, 2, 5,7,8-pentamethylchroman-6-sulfonamide (Pmc), methane sulfonamide (Ms), (3-trimethylsilylethanesulfonamide (SES), 9-anthracenesulfonamide, 4-(4',8'-dimethoxynaphthylmethyl)benzenesulfonamide (DNMBS), benzylsulfonamide, trifluoromethylsulfonamide, and phenacylsulfonamide.
[0210]
[0084] In some embodiments, each nitrogen protecting group, together with the nitrogen atom to which the nitrogen protecting group is attached, is independently selected from the group consisting of phenothiazinyl-(10)-acyl derivatives, / V’-p-toluenesulfonylaminoacyl derivatives, N’-phenylaminothioacyl derivatives, JV-benzoylphenylalanyl derivatives, / V-acetylmethionine derivatives, 4,5-diphenyl-3-oxazolin-2-one, / V-phthalimide, / V-dithiasuccinimide (Dts), / V-2,3-diphenylmaleimide, N- R0708.70180WO00 / R0708.70180US01 23 / 122
[0211] #14644680vl 2,5 -dimethylpyrrole, A-l,l,4,4-tetramethyldisilylazacyclopentane adduct (STABASE), 5-substituted 1,3-dimethyl-l,3,5-triazacyclohexan-2-one, 5-substituted l,3-dibenzyl-l,3,5-triazacyclohexan-2-one, 1-substituted 3,5-dinitro-4-pyridone, A-methylamine, A-allylamine, A-| 2-(trimethylsilyl)ethoxy]methylamine (SEM), N-3 -acetoxypropylamine, A-(l -isopropyl -4-nitro-2-oxo-3-pyroolin-3-yl)amine, quaternary ammonium salts, A-benzylamine, A-di(4-methoxyphenyl)methylamine, A-5-dibenzosuberylamine, A-tri phenyl methylamine (Tr), A- 1 (4-mcthoxy phcny I )di phenyl methyl ] amine (MMTr), A-9-phenylfluorenylamine (PhF), A-2,7-dichloro-9-fluorenylmethyleneamine, A-ferrocenyhnethylamino (Fem), A-2 -picolylamino A ’-oxide, A- 1,1 -dimethylthiomethyleneamine, A-benzylideneamine, A-p-methoxybenzylideneamine, A-diphenylmethyleneamine, A-[(2-pyridyl)mesityl]methyleneamine, A-(A’, A’-dimethylaminomethylene)amine, A-p-nitrobenzylideneamine, A-salicylideneamine, A-5-chlorosalicylideneamine, A-(5-chloro-2-hydroxyphenyl)phenylmethyleneamine, A-cyclohexylideneamine, A-(5,5 -dimethyl-3 -oxo- 1 -cyclohexenyl)amine, A-borane derivatives, A-diphenylborinic acid derivatives, A-[phenyl(pentaacylchromium- or tungsten)acyl]amine, A-copper chelate, A-zinc chelate, A-nitroamine, A-nitrosoamine, amine A-oxide, diphenylphosphinamide (Dpp), dimethylthiophosphinamide (Mpt), diphenylthiophosphinamide (Ppt), dialkyl phosphoramidates, dibenzyl phosphoramidate, diphenyl phosphoramidate, benzene sulfenamide, o-nitrobenzenesulfenamide (Nps), 2,4-dinitrobenzenesulfenamide, pentachlorobenzenesulfenamide, 2-nitro-4-methoxybenzenesulfenamide, triphenylmethylsulfenamide, and 3-nitropyridinesulfenamide (Npys). In some embodiments, two instances of a nitrogen protecting group together with the nitrogen atoms to which the nitrogen protecting groups are attached are A, A5-isopropylidenediamine.
[0212]
[0085] In some embodiments, at least one nitrogen protecting group is Bn, Boc, Cbz, Fmoc, trifluoroacetyl, triphenylmethyl, acetyl, or Ts.
[0213]
[0086] In some embodiments, each oxygen atom substituent is independently substituted (e.g., substituted with one or more halogen) or unsubstituted Cnio alkyl, -C(=O)Raa, -CC>2Raa, -C(=0)N(Rbb)2, or an oxygen protecting group. In some embodiments, each oxygen atom substituents is independently substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-6 alkyl, -C(=O)Raa, -CC R”, -C(=0)N(Rbb)2, or an oxygen protecting group, wherein Raais hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted Cnio alkyl, or an oxygen protecting group when attached to an oxygen atom; and each Rbbis independently hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted Cnio alkyl, or a nitrogen protecting group. In some embodiments, each oxygen atom substituent is independently substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-6 alkyl or an oxygen protecting group.
[0214]
[0087] In some embodiments, the substituent present on an oxygen atom is an oxygen protecting group (also referred to herein as an “hydroxyl protecting group”). Oxygen protecting groups include -Raa, -N(Rbb)2, -C(=O)SRaa, -C(=O)Raa, -CO2Raa, -C(=O)N(Rbb)2, -C(=NRbb)Raa, -C(=NRbb)ORaa, -C(=NRbb)N(Rbb)2, -S(=O)Raa, -SO2Raa, -Si(Raa)3, -P(RCC)2, -P(RCC)3+X,-P(ORcc)2, -P(ORCC)3+X, -P(=O)(Raa)2, -P(=O)(ORCC)2, and -P(=O)(N(Rbb)2)2, wherein X, Raa, Rbb, and Rccare as defined herein. R0708.70180WO00 / R0708.70180US01 24 / 122
[0215] #14644680vl Oxygen protecting groups are well known in the art and include those described in detail in Protecting Groups in Organic Synthesis, T. W. Greene and P. G. M. Wuts, 3rdedition, John Wiley & Sons, 1999, incorporated herein by reference.
[0216]
[0088] In some embodiments, each oxygen protecting group, together with the oxygen atom to which the oxygen protecting group is attached, is selected from the group consisting of methyl, methoxymethyl (MOM), methylthiomethyl (MTM), / -butylthiomethyl, (phenyldimethylsilyl)methoxymethyl (SMOM), benzyloxymethyl (BOM),p-methoxybenzyloxymethyl (PMBM), (4-methoxyphenoxy)methyl (p-AOM), guaiacolmethyl (GUM), / -butoxymcthyl. 4-pentenyloxymethyl (POM), siloxymethyl, 2-methoxyethoxymethyl (MEM), 2,2,2-trichloroethoxymethyl, bis(2-chloroethoxy)methyl, 2-(trimethylsilyl)ethoxymethyl (SEMOR), tetrahydropyranyl (THP), 3-bromotetrahydropyranyl, tetrahydrothiopyranyl, 1 -methoxy cyclohexyl, 4-methoxytetrahydropyranyl (MTHP), 4-methoxytetrahydrothiopyranyl, 4-methoxytetrahydrothiopyranyl S, S-dioxide, l-[(2-chloro-4-methyl)phenyl]-4-methoxypiperidin-4-yl (CTMP), l,4-dioxan-2-yl, tetrahydrofuranyl, tetrahydrothiofuranyl, 2,3,3a,4,5,6,7,7a-octahydro-7,8,8-trimethyl-4,7-methanobenzofuran-2-yl, 1-ethoxyethyl, l-(2-chloroethoxy)ethyl, 1 -methyl- 1 -methoxyethyl, 1 -methyl- 1 -benzyloxyethyl, 1 -methyl- 1-benzyloxy-2-fluoroethyl, 2,2,2-trichloroethyl, 2-trimethylsilylethyl, 2-(phenylselenyl)ethyl, / -butyl, allyl, p-chlorophenyl, p-methoxyphenyl, 2,4-dinitrophenyl, benzyl (Bn), p-methoxybenzyl (PMB), 3,4-dimethoxybenzyl, o-nitrobenzyl, p-nitrobenzy 1, p-halobenzy 1, 2,6-dichlorobenzyl, p-cyanobenzy 1, p-phenylbenzyl, 2-picolyl, 4-picolyl, 3 -methyl -2 -picolyl A-oxido, diphenylmethyl, p,p ’-dinitrobenzhydryl, 5-dibenzosuberyl, triphenylmethyl, 4,4'-dimethoxytrityl (4,4'-dimethoxytriphenyhnethyl or DMT), a-naphthyldiphenylmethyl, p-methoxyphenyldipheny Imethy 1, di(p-methoxypheny 1 )pheny Imethyl, t ri (p-methoxyphenyl)methyl, 4-(4’-bromophenacyloxyphenyl)diphenylmethyl, 4,4',4"-tris(4,5-dichlorophthalimidophenyl)methyl, 4,4',4"-tris(levulinoyloxyphenyl)methyl, 4, 4', 4"-tris(benzoyloxyphenyl)methyl, 4,4’-Dimethoxy-3"‘-[N-(imidazolylmethyl) ]trityl Ether (IDTr-OR), 4,4’-Dimethoxy-3"‘-[N-(imidazolylethyl)carbamoyl]trityl Ether (lETr-OR), l,l-bis(4-methoxyphenyl)-l'-pyrenylmethyl, 9-anthryl, 9-(9-phenyl)xanthenyl, 9-(9-phenyl-10-oxo)anthryl, l,3-benzodithiolan-2-yl, benzisothiazolyl AS'-dioxido. trimethylsilyl (TMS), triethylsilyl (TES), triisopropylsilyl (TIPS), dimethylisopropylsilyl (IPDMS), diethylisopropylsilyl (DEIPS), dimethylthexylsilyl, / -bnty Idimcthy I sily 1 (TBDMS), / -bnty Idi phenyl si lyl (TBDPS), tribenzylsilyl, tri-p-xylylsilyl, triphenylsilyl, diphenylmethylsilyl (DPMS), t-butylmethoxyphenylsilyl (TBMPS), formate, benzoylformate, acetate, chloroacetate, dichloroacetate, trichloroacetate, trifluoroacetate, methoxyacetate, triphenylmethoxyacetate, phenoxyacetate, p-chlorophenoxyacetate, 3 -phenylpropionate, 4-oxopentanoate (levulinate), 4,4-(ethylenedithio)pentanoate (levulinoyldithioacetal), pivaloate, adamantoate, crotonate, 4-methoxycrotonate, benzoate, p-phenylbenzoate, 2,4,6-trimethylbenzoate (mesitoate), methyl carbonate, 9-fluorenylmethyl carbonate (Fmoc), ethyl carbonate, 2,2,2-trichloroethyl carbonate (Troc), 2-(trimethylsilyl)ethyl carbonate (TMSEC), 2-(phenylsulfonyl) ethyl carbonate (Psec), 2-(triphenylphosphonio) ethyl carbonate (Peoc), isobutyl carbonate, vinyl carbonate, allyl carbonate, / -butyl carbonate (BOC or Boc), p-nitrophenyl carbonate, benzyl carbonate, p-methoxybenzyl carbonate, 3,4- R0708.70180WO00 / R0708.70180US01 25 / 122
[0217] #14644680vl dimethoxybenzyl carbonate, o-nitrobenzyl carbonate, p-nitrobenzy 1 carbonate,. S'-bcnzy I thiocarbonate, 4-ethoxy-l-napththyl carbonate, methyl dithiocarbonate, 2-iodobenzoate, 4-azidobutyrate, 4-nitro-4-methylpentanoate, o-(dibromomethyl)benzoate, 2-formylbenzenesulfonate, 2-(methylthiomethoxy)ethyl carbonate (MTMEC-OR), 4-(methylthiomethoxy)butyrate, 2-(methylthiomethoxymethyl)benzoate, 2,6-dichloro-4-methylphenoxyacetate, 2,6-dichloro-4-(l, l,3,3-tetramethylbutyl)phenoxyacetate, 2,4-bis(l, 1-dimethylpropyl)phenoxyacetate, chlorodiphenylacetate, isobutyrate, monosuccinoate, (A)-2-mcthyl-2-butenoate, o-(methoxyacyl)benzoate, a-naphthoate, nitrate, alkyl N, N, N’, N’-tetramethylphosphorodiamidate, alkyl / V-phenylcarbamate, borate, dimethylphosphinothioyl, alkyl 2,4-dinitrophenylsulfenate, sulfate, methane sulfonate (mesylate), benzylsulfonate, and tosylate (Ts).
[0218]
[0089] In some embodiments, at least one oxygen protecting group is silyl, TBDPS, TBDMS, TIPS, TES, TMS, MOM, THP, / -Bn. Bn, allyl, acetyl, pivaloyl, or benzoyl.
[0219]
[0090] In some embodiments, each sulfur atom substituent is independently substituted (e.g., substituted with one or more halogen) or unsubstituted C O alkyl, -C(=O)Raa, -CO2Raa, -C(=O)N(Rbb)2, or a sulfur protecting group. In some embodiments, each sulfur atom substituent is independently substituted (e.g., substituted with one or more halogen) or unsubstituted CMO alkyl, -C(=O)Raa, -CO2Raa, -C(=O)N(Rbb)2, or a sulfur protecting group, wherein Raais hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted CMO alkyl, or an oxygen protecting group when attached to an oxygen atom; and each Rbbis independently hydrogen, substituted (e.g., substituted with one or more halogen) or unsubstituted C O alkyl, or a nitrogen protecting group. In some embodiments, each sulfur atom substituent is independently substituted (e.g., substituted with one or more halogen) or unsubstituted Ci-6 alkyl or a sulfur protecting group.
[0220]
[0091] In some embodiments, the substituent present on a sulfur atom is a sulfur protecting group (also referred to as a “thiol protecting group”). In some embodiments, each sulfur protecting group is selected from the group consisting of-Raa, -N(Rbb)2, -C(=O)SRaa, -C(=O)Raa, -CO2Raa, -C(=O)N(Rbb)2, -C(=NRbb)Raa, -C(=NRbb)ORaa, -C(=NRbb)N(Rbb)2, -S(=O)Raa, -SO2Raa, -Si(Raa)3, -P(RCC)2, -P(RCC)3+X, -P(ORCC)2, -P(ORCC)3+X, -P(=O)(Raa)2, -P(=O)(ORCC)2, and -P(=O)(N(Rbb) 2)2, wherein Raa, Rbb, and Rccare as defined herein. Sulfur protecting groups are well known in the art and include those described in detail in Protecting Groups in Organic Synthesis, T. W. Greene and P. G. M. Wuts, 3rdedition, John Wiley & Sons, 1999, incorporated herein by reference.
[0221]
[0092] In some embodiments, the molecular weight of a substituent is lower than 250, lower than 200, lower than 150, lower than 100, or lower than 50 g / mol. In some embodiments, a substituent consists of carbon, hydrogen, fluorine, chlorine, bromine, iodine, oxygen, sulfur, nitrogen, and / or silicon atoms. In some embodiments, a substituent consists of carbon, hydrogen, fluorine, chlorine, bromine, iodine, oxygen, sulfur, and / or nitrogen atoms. In some embodiments, a substituent consists of carbon, hydrogen, fluorine, chlorine, bromine, and / or iodine atoms. In some embodiments, a substituent consists of carbon, hydrogen, fluorine, and / or chlorine atoms. In some embodiments, a substituent comprises 0, 1, 2, or 3 hydrogen bond donors. In some embodiments, a substituent comprises 0, 1, 2, or 3 hydrogen bond acceptors.
[0222] R0708.70180WO00 / R0708.70180US01 26 / 122
[0223] #14644680vl
[0093] A “counterion” or “anionic counterion” is a negatively charged group associated with a positively charged group in order to maintain electronic neutrality. An anionic counterion may be monovalent (e.g., including one formal negative charge). An anionic counterion may also be multivalent (e.g., including more than one formal negative charge), such as divalent or trivalent. Exemplary counterions include halide ions (e.g., F, Cl” Br ", I"), NOs. CIO4”, OH", H2PO4. HCO3, HSO4. sulfonate ions (e.g., methanesulfonate, trifluoromethanesulfonate (triflate), p-toluenesulfonate, benzene sulfonate, 10-camphor sulfonate, naphthalene-2-sulfonate, naphthalene- 1 -sulfonic acid-5-sulfonate, ethan-1-sulfonic acid-2-sulfonate, and the like), carboxylate ions (e.g., acetate, propanoate, benzoate, glycerate, lactate, tartrate, glycolate, gluconate, and the like), BF4, PF4”, PFg”, AsFe", SbFe", B 13.5-(CF3)2C. H, | 1 - B(CeF5)4, BPI14", A1(OC(CF3)3)4 ”, and carborane anions (e.g., CB11H12” or (HCBnMesBre) ). Exemplary counterions which may be multivalent include CO32, HPO42, PO43, B4O72, SO42, S2O32, carboxylate anions (e.g., tartrate, citrate, fumarate, maleate, malate, malonate, gluconate, succinate, glutarate, adipate, pimelate, suberate, azelate, sebacate, salicylate, phthalates, aspartate, glutamate, and the like), and carboranes.
[0224]
[0094] A “leaving group” (LG) is an art-understood term referring to an atomic or molecular fragment that departs with a pair of electrons in heterolytic bond cleavage, wherein the molecular fragment is an anion or neutral molecule. In some embodiments, a leaving group is an atom or a group capable of being displaced by a nucleophile. See e.g., Smith, March Advanced Organic Chemistry 6th ed. (501-502). Exemplary leaving groups include, but are not limited to, halo (e.g., fluoro, chloro, bromo, iodo) and activated substituted hydroxyl groups (e.g., -OC(=O)SRaa, -OC(=O)Raa, -OCO2Raa, -OC(=O)N(Rbb)2, -OC(=NRbb)Raa, -OC(=NRbb)ORaa, -OC(=NRbb)N(Rbb)2, -OS(=O)Raa, -OSO2Raa, -OP(RCC)2, -OP(RCC)3, -OP(=O)2Raa, -OP(=O)(Raa)2, -OP(=O)(ORCC)2, -OP(=O)2N(Rbb)2, and -OP(=O)(NRbb)2, wherein Raa, Rbb, and Rccare as defined herein). Additional examples of suitable leaving groups include, but are not limited to, halogen alkoxycarbonyloxy, aryloxycarbonyloxy, alkane sulfonyloxy, arenesulfonyloxy, alkylcarbonyloxy (e.g., acetoxy), arylcarbonyloxy, aryloxy, methoxy, N, O-dimethylhydroxylamino, pixyl, and haloformates. In some embodiments, the leaving group is a sulfonic acid ester, such as toluene sulfonate (tosylate, -OTs), methanesulfonate (mesylate, -OMs),p-bromobenzenesulfonyloxy (brosylate, -OBs), -OS(=O)2(CF2)3CF3 (nonaflate, -ONf), or trifluoromethanesulfonate (triflate, -OTf). In some embodiments, the leaving group is a brosylate, such as p-bromobenzenesulfonyloxy. In some embodiments, the leaving group is a nosylate, such as 2-nitrobenzenesulfonyloxy. In some embodiments, the leaving group is a sulfonate-containing group. In some embodiments, the leaving group is a tosylate group. In some embodiments, the leaving group is a phosphineoxide (e.g., formed during a Mitsunobu reaction) or an internal leaving group such as an epoxide or cyclic sulfate. Other non-limiting examples of leaving groups are water, ammonia, alcohols, ether moieties, thioether moieties, zinc halides, magnesium moieties, diazonium salts, and copper moieties.
[0225]
[0095] Use of the phrase “at least one instance” refers to 1, 2, 3, 4, or more instances, but also encompasses a range, e.g., for example, from 1 to 4, from 1 to 3, from 1 to 2, from 2 to 4, from 2 to 3, or from 3 to 4 instances, inclusive.
[0226] R0708.70180WO00 / R0708.70180US01 27 / 122
[0227] #14644680vl
[0096] It is also to be understood that compounds that have the same molecular formula but differ in the nature or sequence of bonding of their atoms or the arrangement of their atoms in space are termed “isomers”. Isomers that differ in the arrangement of their atoms in space are termed “stereoisomers”.
[0228]
[0097] Stereoisomers that are not mirror images of one another are termed “diastereomers” and those that are non-superimposable mirror images of each other are termed “enantiomers”. When a compound has an asymmetric center, for example, it is bonded to four different groups, a pair of enantiomers is possible. An enantiomer can be characterized by the absolute configuration of its asymmetric center and is described by the R- and S-sequencing rules of Cahn and Prelog, or by the manner in which the molecule rotates the plane of polarized light and designated as dextrorotatory or levorotatory (i.e., as (+) or (-)-isomers respectively). A chiral compound can exist as either individual enantiomer or as a mixture thereof. A mixture containing equal proportions of the enantiomers is called a “racemic mixture”.
[0229]
[0098] These and other exemplary substituents are described in more detail in the Detailed Description, Examples, and Claims. The present disclosure is not limited in any manner by the above exemplary listing of substituents.
[0230]
[0099] As used herein, the term “salt” refers to any and all salts, and encompasses pharmaceutically acceptable salts. Salts include ionic compounds that result from the neutralization reaction of an acid and a base. A salt is composed of one or more cations (positively charged ions) and one or more anions (negative ions) so that the salt is electrically neutral (without a net charge). Salts of the compounds of the present disclosure include those derived from inorganic and organic acids and bases. Examples of acid addition salts are salts of an amino group formed with inorganic acids, such as hydrochloric acid, hydrobromic acid, phosphoric acid, sulfuric acid, and perchloric acid, or with organic acids, such as acetic acid, oxalic acid, maleic acid, tartaric acid, citric acid, succinic acid, or malonic acid or by using other methods known in the art such as ion exchange. Other salts include adipate, alginate, ascorbate, aspartate, benzenesulfonate, benzoate, bisulfate, borate, butyrate, camphorate, camphorsulfonate, citrate, cyclopentanepropionate, digluconate, dodecylsulfate, ethanesulfonate, formate, fumarate, glucoheptonate, glycerophosphate, gluconate, hemisulfate, heptanoate, hexanoate, hydroiodide, 2-hydroxy-ethanesulfonate, lactobionate, lactate, laurate, lauryl sulfate, malate, maleate, malonate, methanesulfonate, 2-naphthalene sulfonate, nicotinate, nitrate, oleate, oxalate, palmitate, pamoate, pectinate, persulfate, 3-phenylpropionate, phosphate, picrate, pivalate, propionate, stearate, succinate, sulfate, tartrate, thiocyanate, p-toluenesulfonate, undecanoate, valerate, hippurate, and the like. Salts derived from appropriate bases include alkali metal, alkaline earth metal, ammonium and N (C’i 4 alky 1 )4salts. Representative alkali or alkaline earth metal salts include sodium, lithium, potassium, calcium, magnesium, and the like. Further salts include ammonium, quaternary ammonium, and amine cations formed using counterions such as halide, hydroxide, carboxylate, sulfate, phosphate, nitrate, lower alkyl sulfonate, and aryl sulfonate.
[0231]
[0100] As used herein, the term “about X,” or “approximately X,” where X is a number or percentage, refers to a number or percentage that is between 99.5% and 100.5%, between 99% and 101%, between 98% and 102%, between 97% and 103%, between 96% and 104%, between 95% and 105%, between R0708.70180WO00 / R0708.70180US01 28 / 122
[0232] #14644680vl 92% and 108%, or between 90% and 110%, inclusive, ofX.
[0233]
[0101] The terms “polynucleotide”, “nucleotide sequence”, “nucleic acid”, “nucleic acid molecule”, “nucleic acid sequence”, and “oligonucleotide” refer to a series of nucleotide bases (also called “nucleotides”) in DNA and RNA, and mean any chain of two or more nucleotides. The polynucleotides can be chimeric mixtures or derivatives or modified versions thereof, single -stranded or double-stranded. The oligonucleotide can be modified at the base moiety, sugar moiety, or phosphate backbone, for example, to improve stability of the molecule, its hybridization parameters, etc. The antisense oligonuculeotide may comprise a modified base moiety which is selected from the group including, but not limited to, 5 -fluorouracil, 5 -bromouracil, 5 -chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5 -(carboxyhydroxylmethyl) uracil, 5 -carboxymethylaminomethyl -2 -thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1 -methylinosine, 2,2- dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5- methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5- methoxyaminomethyl-2 -thiouracil, beta-D-mannosylqueosine, 5’-methoxycarboxymethyluracil, 5 -methoxyuracil, 2-methylthio-N6-isopentenyladenine, wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5 -methyl -2 -thiouracil, 2-thiouracil, 4-thiouracil, 5 -methyluracil, uracil- 5-oxyacetic acid methylester, uracil-5 -oxyacetic acid, 5-methyl-2- thiouracil, 3-(3-amino-3-N-2-carboxypropyl) uracil, a thio-guanine, and 2,6-diaminopurine. A nucleotide sequence typically carries genetic information, including the information used by cellular machinery to make proteins and enzymes. These terms include double- or single-stranded genomic and cDNA, RNA, any synthetic and genetically manipulated polynucleotide, and both sense and antisense polynucleotides. This includes single- and double -stranded molecules, i.e., DNA-DNA, DNA-RNA and RNA-RNA hybrids, as well as “protein nucleic acids” (PNAs) formed by conjugating bases to an amino acid backbone. This also includes nucleic acids containing carbohydrate or lipids. Exemplary DNAs include single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), plasmid DNA (pDNA), genomic DNA (gDNA), complementary DNA (cDNA), antisense DNA, chloroplast DNA (ctDNA or cpDNA), microsatellite DNA, mitochondrial DNA (mtDNA or mDNA), kinetoplast DNA (kDNA), provirus, lysogen, repetitive DNA, satellite DNA, and viral DNA. Exemplary RNAs include single-stranded RNA (ssRNA), doublestranded RNA (dsRNA), small interfering RNA (siRNA), messenger RNA (mRNA), precursor messenger RNA (pre-mRNA), small hairpin RNA or short hairpin RNA (shRNA), microRNA (miRNA), guide RNA (gRNA), transfer RNA (tRNA), antisense RNA (asRNA), heterogeneous nuclear RNA (hnRNA), coding RNA, non-coding RNA (ncRNA), long non-coding RNA (long ncRNA or IncRNA), satellite RNA, viral satellite RNA, signal recognition particle RNA, small cytoplasmic RNA, small nuclear RNA (snRNA), ribosomal RNA (rRNA), Piwi-interacting RNA (piRNA), polyinosinic acid, ribozyme, flexizyme, small nucleolar RNA (snoRNA), spliced leader RNA, viral RNA, and viral satellite RNA.
[0234]
[0102] Polynucleotides described herein may be synthesized by standard methods known in the art, e.g., by use of an automated DNA synthesizer (such as those that are commercially available from Biosearch, R0708.70180WO00 / R0708.70180US01 29 / 122
[0235] #14644680vl Applied Biosystems, etc.). As examples, phosphorothioate oligonucleotides may be synthesized by the method of Stein et al., Nucl. Acids Res., 16, 3209, (1988), methylphosphonate oligonucleotides can be prepared by use of controlled pore glass polymer supports (Sarin et al., Proc. Natl. Acad. Sci. U. S. A. 85, 7448-7451, (1988)). A number of methods have been developed for delivering antisense DNA or RNA to cells, e.g., antisense molecules can be injected directly into the tissue site, or modified antisense molecules, designed to target the desired cells (antisense linked to peptides or antibodies that specifically bind receptors or antigens expressed on the target cell surface) can be administered systemically.
[0236] Alternatively, RNA molecules may be generated by in vitro and in vivo transcription of DNA sequences encoding the antisense RNA molecule. Such DNA sequences may be incorporated into a wide variety of vectors that incorporate suitable RNA polymerase promoters such as the T7 or SP6 polymerase promoters. Alternatively, antisense cDNA constructs that synthesize antisense RNA constitutively or inducibly, depending on the promoter used, can be introduced stably into cell lines. However, it is often difficult to achieve intracellular concentrations of the antisense sufficient to suppress translation of endogenous mRNAs. Therefore a preferred approach utilizes a recombinant DNA construct in which the antisense oligonucleotide is placed under the control of a strong promoter. The use of such a construct to transfect target cells in the patient will result in the transcription of sufficient amounts of single stranded RNAs that will form complementary base pairs with the endogenous target gene transcripts and thereby prevent translation of the target gene mRNA. For example, a vector can be introduced in vivo such that it is taken up by a cell and directs the transcription of an antisense RNA. Such a vector can remain episomal or become chromosomally integrated, as long as it can be transcribed to produce the desired antisense RNA. Such vectors can be constructed by recombinant DNA technology methods standard in the art. Vectors can be plasmid, viral, or others known in the art, used for replication and expression in mammalian cells. Expression of the sequence encoding the antisense RNA can be by any promoter known in the art to act in mammalian, preferably human, cells. Such promoters can be inducible or constitutive. Any type of plasmid, cosmid, yeast artificial chromosome, or viral vector can be used to prepare the recombinant DNA construct that can be introduced directly into the tissue site.
[0237]
[0103] The polynucleotides may be flanked by natural regulatory (expression control) sequences or may be associated with heterologous sequences, including promoters, internal ribosome entry sites (IRES) and other ribosome binding site sequences, enhancers, response elements, suppressors, signal sequences, polyadenylation sequences, introns, 5'- and 3 '-non-coding regions, and the like. The nucleic acids may also be modified by many means known in the art. Non-limiting examples of such modifications include methylation, “caps”, substitution of one or more of the naturally occurring nucleotides with an analog, and intemucleotide modifications, such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoroamidates, carbamates, etc.) and with charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.). Polynucleotides may contain one or more additional covalently linked moieties, such as, for example, proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), intercalators (e.g., acridine, psoralen, etc.), chelators (e.g., metals, radioactive metals, iron, oxidative metals, etc.), and alkylators. The polynucleotides may be derivatized R0708.70180WO00 / R0708.70180US01 30 / 122
[0238] #14644680vl by formation of a methyl or ethyl phosphotriester or an alkyl phosphoramidate linkage. Furthermore, the polynucleotides herein may also be modified with a label capable of providing a detectable signal, either directly or indirectly. Exemplary labels include radioisotopes, fluorescent molecules, isotopes (e.g., radioactive isotopes), biotin, and the like.
[0239]
[0104] A “protein,” “peptide,” or “polypeptide” comprises a polymer of amino acid residues linked together by peptide bonds. The term refers to proteins, polypeptides, and peptides of any size, structure, or function. Typically, a protein will be at least three amino acids long. A protein may refer to an individual protein or a collection of proteins. Inventive proteins preferably contain only natural amino acids, although non-natural amino acids (z. e., compounds that do not occur in nature but that can be incorporated into a polypeptide chain) and / or amino acid analogs as are known in the art may alternatively be employed. Also, one or more of the amino acids in a protein may be modified, for example, by the addition of a chemical entity such as a carbohydrate group, a hydroxyl group, a phosphate group, a famesyl group, an isofamesyl group, a fatty acid group, a linker for conjugation or functionalization, or other modification. A protein may also be a single molecule or may be a multi-molecular complex. A protein may be a fragment of a naturally occurring protein or peptide. A protein may be naturally occurring, recombinant, synthetic, or any combination of these.
[0240]
[0105] Amino acid residues may be indicated by their corresponding single letter codes, e.g., R (arginine), H (histidine), K (lysine), D (aspartic acid), E (glutamic acid), S (serine), T (threonine), N (asparagine), Q (glutamine), C (cysteine), G (glycine), P (proline), A (alanine), V (valine), I (isoleucine), L (leucine), M (methionine), F (phenylalanine), Y (tyrosine), W (tryptophan).
[0241]
[0106] In some embodiments, the terminology includes identifying one or more amino acids of a polypeptide. As used herein, in some embodiments, “identifying,” “determining the identity,” and like terms, in reference to an amino acid, include determination of an express identity of an amino acid as well as determination of a probability of an express identity of an amino acid. For example, in some embodiments, an amino acid is identified by determining a probability (e.g., from 0% to 100%) that the amino acid is of a specific type, or by determining a probability for each of a plurality of specific types. Accordingly, in some embodiments, the terms “amino acid sequence,” “polypeptide sequence,” and “protein sequence” as used herein may refer to the polypeptide or protein material itself and is not restricted to the specific sequence information (e.g., the succession of letters representing the order of amino acids from one terminus to another terminus) that biochemically characterizes a specific polypeptide or protein.
[0242]
[0107] For the purposes of comparing two or more amino acid sequences, the percentage of “sequence identity” between a first amino acid sequence and a second amino acid sequence (also referred to herein as “amino acid identity”) may be calculated by: dividing [the number of amino acid residues in the first amino acid sequence that are identical to the amino acid residues at the corresponding positions in the second amino acid sequence] by [the total number of amino acid residues in the first amino acid sequence] and multiplying by
[0100] , in which each deletion, insertion, substitution or addition of an amino acid residue in the second amino acid sequence compared to the first amino acid sequence is R0708.70180WO00 / R0708.70180US01 31 / 122
[0243] #14644680vl considered as a difference at a single amino acid residue (position).
[0244]
[0108] Alternatively, the degree of sequence identity between two amino acid sequences may be calculated using a known computer algorithm (e.g., by the local homology algorithm of Smith and Waterman (1970) Adv. Appl. Math. 2:482c, by the homology alignment algorithm of Needleman and Wunsch, J. Mol. Biol. (1970) 48:443, by the search for similarity method of Pearson and Lipman. Proc. Natl. Acad. Sci. USA (1998) 85:2444, or by computerized implementations of algorithms available as Blast, Clustal Omega, or other sequence alignment algorithms) and, for example, using standard settings. Usually, for the purpose of determining the percentage of “sequence identity” between two amino acid sequences in accordance with the calculation method outlined hereinabove, the amino acid sequence with the greatest number of amino acid residues will be taken as the “first” amino acid sequence, and the other amino acid sequence will be taken as the “second” amino acid sequence.
[0245]
[0109] Additionally, or alternatively, two or more sequences may be assessed for the identity between the sequences. The terms “identical” or percent “identity” in the context of two or more amino acid sequences, refer to two or more sequences or subsequences that are the same. Two sequences are “substantially identical” if two sequences have a specified percentage of amino acid residues that are the same (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identical) over a specified region or over the entire sequence, when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the above sequence comparison algorithms or by manual alignment and visual inspection. Optionally, the identity exists over a region that is at least about 25, 50, 75, or 100 amino acids in length, or over a region that is 100 to 150, 150 to 200, 100 to 200, or 200 or more, amino acids in length.
[0246]
[0110] Additionally, or alternatively, two or more sequences may be assessed for the alignment between the sequences. The terms “alignment” or “percent alignment” in the context of two or more amino acid sequences, refer to two or more sequences or subsequences that are the same. Two sequences are “substantially aligned” if two sequences have a specified percentage of amino acid residues that are the same (e.g., at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or 99.9% identical) over a specified region or over the entire sequence, when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the above sequence comparison algorithms or by manual alignment and visual inspection. Optionally, the alignment exists over a region that is at least about 25, 50, 75, or 100 amino acids in length, or over a region that is 100 to 150, 150 to 200, 100 to 200, or 200 or more amino acids in length.
[0247]
[0111] A “peptidase,” “protease,” or “proteinase” is an enzyme that catalyzes the hydrolysis of a peptide bond. Peptidases digest polypeptides into shorter fragments and may be generally classified into endopeptidases and exopeptidases, which cleave a polypeptide chain internally and terminally, respectively. An exopeptidase in accordance with the application may be an “aminopeptidase” or a “carboxypeptidase,” which cleaves a single amino acid from an amino- or a carboxy-terminus, respectively. A peptidase (e.g., an aminopeptidase) may also be referred to as a “cutter” or a “cleaving agent.”
[0248] R0708.70180WO00 / R0708.70180US01 32 / 122
[0249] #14644680vl
[0112] The term “avidin protein” refers to a biotin-binding protein, generally having a biotin binding site at each of four subunits of the avidin protein. Avidin proteins include, for example, avidin, streptavidin, traptavidin, tamavidin, bradavidin, xenavidin, and homologs and variants thereof. In some cases, the monomeric, dimeric, or tetrameric form of the avidin protein can be used. In some embodiments, the avidin protein of an avidin protein complex is streptavidin in a tetrameric form (e.g., a homotetramer).
[0250]
[0113] As used herein, in some embodiments, the term “bond” or “bonds” refers to any non-covalent interaction (e.g., a hydrogen bond, a van der Waals interaction, an aromatic interaction, an electrostatic interaction) or covalent interaction between specified binding components or any plurality thereof, and the terms “bind,” “binding,” “bound,” and like terms refer to the formation and / or existence of any such bonds. As an illustrative example, a binding event between an amino acid recognizer and an amino acid may comprise the formation of one or more non-covalent or covalent interactions between the amino acid recognizer and the amino acid.
[0251]
[0114] The term “click chemistry” refers to a chemical synthesis technique introduced by K. Barry Sharpless of The Scripps Research Institute, describing chemistry tailored to generate covalent bonds quickly and reliably by joining small units comprising reactive groups together. See, e.g., Kolb, Finn and Sharpless Angewandte Chemie International Edition (2001) 40: 2004-2021; Evans, Australian Journal of Chemistry (2007) 60: 384-395). Exemplary coupling reactions (some of which may be classified as “click chemistry”) include, but are not limited to, formation of esters, thioesters, amides (e.g., such as peptide coupling) from activated acids or acyl halides; nucleophilic displacement reactions (e.g., such as nucleophilic displacement of a halide or ring opening of strained ring systems); azide-alkyne Huisgen cycloaddition; thiol-yne addition; imine formation; Michael additions (e.g., maleimide addition); and Diels- Alder reactions (e.g., tetrazine [4 + 2] cycloaddition). Exemplary click chemistry reactions include, but are not limited to, azide-alkyne Huisgen cycloaddition; and Diels- Alder reactions (e.g., tetrazine [4 + 2] cycloaddition). In some embodiments, click chemistry reactions are modular, wide in scope, give high chemical yields, generate inoffensive byproducts, are stereospecific, exhibit a large thermodynamic driving force > 84 kJ / mol to favor a reaction with a single reaction product, and / or can be carried out under physiological conditions. In some embodiments, a click chemistry reaction exhibits high atom economy, can be carried out under simple reaction conditions, use readily available starting materials and reagents, uses no toxic solvents or use a solvent that is benign or easily removed (preferably water), and / or provides simple product isolation by non-chromatographic methods (crystallization or distillation).
[0252]
[0115] The term “click chemistry handle,” as used herein, refers to a reactant, or a reactive group, that can partake in a click chemistry reaction. For example, a strained alkyne, e.g., a cyclooctyne, is a click chemistry handle, since it can partake in a strain-promoted cycloaddition (see, e.g., Table 1). In general, click chemistry reactions require at least two molecules comprising click chemistry handles that can react with each other. Such click chemistry handle pairs that are reactive with each other are sometimes referred to herein as partner click chemistry handles. For example, an azide is a partner click chemistry handle to a cyclooctyne or any other alkyne. Exemplary click chemistry handles suitable for use R0708.70180WO00 / R0708.70180US01 33 / 122
[0253] #14644680vl according to some aspects of this invention are described herein, for example, in Tables 1 and 2. In some embodiments, click chemistry handles are used that can react to form covalent bonds in the presence of a metal catalyst, e.g., copper (II). In some embodiments, click chemistry handles are used that can react to form covalent bonds in the absence of a metal catalyst. Additional suitable click chemistry handles are well known to those of skill in the art, and such click chemistry handles include, but are not limited to, the click chemistry reaction partners, groups, and handles described in Becer, Hoogenboom, and Schubert, Click Chemistry beyond Metal-Catalyzed Cycloaddition, Angewandte Chemie International Edition (2009) 48: 4900 - 4908 and PCT / US2012 / 044584 and references therein, which references are incorporated herein by reference for click chemistry handles and methodology.
[0254] Table 1: Exemplary click chemistry handles and reactions.
[0255]
[0256] 1,3-dipolar cycloaddition terminal alkyne azide
[0257]
[0258] strained alkyne
[0259]
[0260] Table 2: Exemplary click chemistry handles and reactions (from Becer, Hoogenboom, and Schubert, Click Chemistry Beyond Metal-Catalyzed Cycloaddition, Angewandte Chemie International Edition (2009) 48: 4900 - 4908.).
[0261] Reagent Reagent B Mechanism Notes on reaction131Reference A
[0262] 0 azide alkyne Cu-catalyzed [3+2] 2 h at 60°C in H₂O [9]
[0263] azide -alkyne
[0264] cycloaddition (CuAAC)
[0265] 1 azide cyclooctyne strain-promoted [3+2] 1 h at RT [6- azide-alkyne 8,10,11] cycloaddition (SPAAC)
[0266] 2 azide activated [3+2] Huisgen 4 h at 50°C
[0012]
[0267] alkyne cycloaddition
[0268] R0708.70180WO00 / R0708.70180US01 34 / 122
[0269] #14644680vl Reagent Reagent B Mechanism Notes on reaction131Reference A
[0270] 3 azide electron- [3+2] cycloaddition 12 h at RT in H₂O
[0013]
[0271] deficient
[0272] alkyne
[0273] 4 azide aryne [3+2] cycloaddition 4 h at RT in THF with [14,15] crown ether or 24 h at
[0274] RT in CHUN
[0275] 5 tetrazine alkene Diels-Alder retro- [4+2] 40 min at 25 °C (100% [36-38] cycloaddition yield)
[0276] N2 is the only by-product
[0277] 6 tetrazole alkene 1,3 -dipolar cycloaddition few min UV irradiation [39,40] (photoclick) and then overnight at
[0278] 4°C
[0279] 7 dithioester diene hetero-Diels-Alder 10 min at RT
[0043]
[0280] cycloaddition
[0281] g anthracene maleimide [4+2] Diels-Alder 2 days at reflux in
[0041]
[0282] reaction toluene
[0283] 9 thiol alkene radical addition 30 min UV (quantitative [19-23] (thio click) conv.) or
[0284] 24 h UV irradiation
[0285] (>96%)
[0286] 10 thiol enone Michael addition 24 h at RT in CH3CN
[0027] 11 thiol maleimide Michael addition 1 h at 40°C in THF or [24-26]
[0287] 16 h at RT in dioxane
[0288] 12 thiol para-fluoro nucleophilic substitution overnight at RT in DMF
[0032]
[0289] or
[0290] 60 min at 40°C in DMF
[0291] 13 amine para-fluoro nucleophilic substitution 20 min MW at 95°C in
[0030]
[0292] NMP as solvent
[0293] [a ]RT=room temperature, DMF=N,N-dimethylformamide. NMP=N-methylpyrrolidone, THF=tetrahydrofuran, CH₃CN=acetonitrile
[0294] DETAILED DESCRIPTION OF CERTAIN EMBODIMENTS
[0295]
[0116] Aspects of the disclosure relate to methods of preparing polypeptide samples from a protein for analysis (e.g., protein sequencing). Embodiments described herein may provide various benefits during the preparation of polypeptide samples, including reduction of protein loss, improvement of protein digestion efficiency, improvement in protein functionalization (e.g., derivatization), and / or improvement in conjugation to a functionalized solid substrate. These embodiments may also provide advantages for the analysis of the polypeptide samples (e.g., for protein sequencing). For example, these embodiments may allow lower concentrations of protein to be used, improve the likelihood of correctly identifying a protein, and / or enable the identification of proteins not readily identified by existing methods.
[0296]
[0117] The aspects described herein are not limited to specific embodiments, systems, compositions, methods, or configurations, and as such can, of course, vary. The terminology used herein is for the purpose of describing particular aspects only and, unless specifically defined herein, is not intended to be limiting.
[0297] R0708.70180WO00 / R0708.70180US01 35 / 122
[0298] #14644680vl Methods of Preparing a Polypeptide Sample and Methods of Protein Analysis
[0299]
[0118] In one aspect, provided herein is a method of preparing a polypeptide sample from a protein for analysis, comprising:
[0300] exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample; derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides; and
[0301] conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0302]
[0119] In some embodiments, the method further comprises denaturing the protein before exposing the protein to the protein digestion agent. In some embodiments, the method further comprises exposing the protein to a reducing agent before exposing the protein to the protein digestion agent. In some embodiments, the method further comprises exposing the protein to an amino acid side chain capping agent before exposing the protein to the protein digestion agent. In some embodiments, the method further comprises exposing the protein to a surfactant before exposing the protein to the protein digestion agent.
[0303]
[0120] In some embodiments, the method further comprises denaturing the protein before exposing the protein to the protein digestion agent and exposing the protein to a reducing agent before exposing the protein to the protein digestion agent. In some embodiments, the method further comprises denaturing the protein before exposing the protein to the protein digestion agent and exposing the protein to an amino acid side chain capping agent before exposing the protein to the protein digestion agent. In some embodiments, the method further comprises exposing the protein to a reducing agent before exposing the protein to the protein digestion agent and exposing the protein to an amino acid side chain capping agent before exposing the protein to the protein digestion agent.
[0304]
[0121] In another aspect, provided herein is a method of preparing a polypeptide sample from a protein for analysis, comprising:
[0305] exposing the protein to a reducing agent;
[0306] exposing the protein to an amino acid side chain capping agent;
[0307] exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample.
[0308]
[0122] In some embodiments, the method further comprises derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides. In some embodiments, the method further comprises conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0309]
[0123] In some embodiments, the method further comprises derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides and conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more R0708.70180WO00 / R0708.70180US01 36 / 122
[0310] #14644680vl immobilization complex-conjugated polypeptides.
[0311]
[0124] In another aspect, provided herein is a method of preparing a polypeptide sample from a protein for analysis, comprising:
[0312] exposing the protein to a reducing agent;
[0313] exposing the protein to an amino acid side chain capping agent;
[0314] exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample; derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides; and
[0315] conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0316]
[0125] In some embodiments, the method further comprises:
[0317] contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents and one or more amino acid recognizers;
[0318] monitoring a signal for signal pulses corresponding to interactions between one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides of the polypeptide sample; and
[0319] determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal.
[0320]
[0126] In some embodiments, the analysis is polypeptide sequencing.
[0321]
[0127] In another aspect, provided herein is a method of protein analysis, comprising:
[0322] exposing the protein to a reducing agent;
[0323] exposing the protein to an amino acid side chain capping agent;
[0324] exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample; derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides;
[0325] conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0326] contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents and one or more amino acid recognizers;
[0327] monitoring a signal for signal pulses corresponding to interactions between one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides of the polypeptide sample; and
[0328] determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal.
[0329]
[0128] In some embodiments, the method of protein analysis is a method of polypeptide sequencing.
[0330]
[0129] In some embodiments, the method further comprises denaturing the protein before exposing the R0708.70180WO00 / R0708.70180US01 37 / 122
[0331] #14644680vl protein to the reducing agent. In some embodiments, the method does not comprise a buffer exchange step before denaturing the protein. In some embodiments, the method does not comprise a buffer exchange step before exposing the protein to the reducing agent.
[0332] Protein Sample
[0333]
[0130] In some embodiments, the protein is from a protein sample.
[0334]
[0131] In some embodiments, the protein sample comprises a biological sample. In some embodiments, a protein sample comprises blood, saliva, sputum, feces, urine or buccal swab sample. In some embodiments, a biological sample is from a human, a non-human primate, a rodent, a dog, a cat, a horse, or any other mammal. In some embodiments, a biological sample is from a bacterial cell culture (e.g., an E. coli bacterial cell culture). A bacterial cell culture may comprise gram positive bacterial cells and / or gram negative bacterial cells. In some embodiments, a sample is a purified sample proteins that have been previously extracted. A blood sample may be a freshly drawn blood sample from a subject (e.g., a human subject) or a dried blood sample (e.g., preserved on solid media (e.g., Guthrie cards)). A blood sample may comprise whole blood, serum, plasma, red blood cells, and / or white blood cells.
[0335]
[0132] In some embodiments, a protein sample (e.g., a sample comprising cells or tissue), may be prepared, e.g., lysed (e.g., disrupted, degraded and / or otherwise digested) in a process in accordance with the instant disclosure. In some embodiments, a protein sample to be prepared, e.g., lysed, comprises cultured cells, tissue samples from biopsies (e.g., tumor biopsies from a cancer patient, e.g., a human cancer patient), or any other clinical sample. In some embodiments, a protein sample comprising cells or tissue is lysed using any one of known physical or chemical methodologies to release a target molecule (e.g., a target protein) from said cells or tissues. In some embodiments, a protein sample may be lysed using an electrolytic method, an enzymatic method, a detergent-based method, and / or mechanical homogenization. In some embodiments, a protein sample (e.g., complex tissues, gram positive or gram negative bacteria) may require multiple lysis methods performed in series. In some embodiments, if a protein sample does not comprise cells or tissue (e.g., a protein sample comprising purified protein), a lysis step may be omitted. In some embodiments, lysis of a protein sample is performed to isolate target protein(s). In some embodiments, a lysis method further includes use of a mill to grind a protein sample, sonication, surface acoustic waves (SAW), freeze-thaw cycles, heating, addition of detergents, addition of protein degradants (e.g., enzymes such as hydrolases or proteases), and / or addition of cell wall digesting enzymes (e.g., lysozyme or zymolase). Exemplary detergents (e.g., non-ionic detergents) for lysis include polyoxyethylene fatty alcohol ethers, polyoxyethylene alkylphenyl ethers, polyoxyethylene -polyoxypropylene block copolymers, polysorbates and alkylphenol ethoxylates, preferably nonylphenol ethoxylates, alkylglucosides and / or polyoxyethylene alkyl phenyl ethers. In some embodiments, lysis methods involve heating a protein sample for at least 1-30 min, 1-25 min, 5-25 min, 5-20 min, 10-30 min, 5-10 min, 10-20 min, or at least 5 min at a desired temperature (e.g., at least 60° C, at least 70° C, at least 80° C, at least 90° C, or at least 95° C).
[0336]
[0133] In some embodiments, a protein sample is prepared, e.g., lysed, in the presence of a buffer system. This buffer system may be used to make a slurry of the protein sample, to suspend the protein sample, R0708.70180WO00 / R0708.70180US01 38 / 122
[0337] #14644680vl and / or to stabilize the protein sample during any known lysis methodology, including those methods described herein. In some embodiments, a protein sample is prepared, e.g., lysed, in the presence of RIPA buffer, GCI buffer that comprises Guanidine-HCl buffer, Gly-NP40 buffer, a TRIS buffer, a HEPES buffer, or any other known buffering solution.
[0338]
[0134] Many of the lysis methods described herein allow for the protein sample to be lysed by mechanically homogenizing the protein sample such that the cell walls of the protein sample break down. For example, methods that cause lysis by mechanical homogenization include, but are not limited to bead-beating, heating (e.g., to high temperatures sufficient to disrupt cell walls, e.g., greater than 50° C, 60° C, 70° C, 80° C, 90° C, or 95° C), syringe / needle / microchannel passage (to cause shearing), sonication, or maceration with a grinder. In some embodiments, any lysis methodology may be combined with any other lysis methodology. For example, any lysis methodology may be combined with heating and / or sonication and / or syringe / needle / microchannel passage to quicken the rate of lysis.
[0339]
[0135] In some embodiments, protein sample preparation comprises cell disruption (i.e., subsequent removal of unwanted cell and tissue elements following lysis). In some embodiments, cell disruption involves protein precipitation. In some embodiments, following precipitation, the lysed and disrupted protein sample is subjected to centrifugation. In some embodiments, following centrifugation, the supernatant is discarded. Precipitation can be accomplished through multiple processes, including but not limited to those methods described in Winter, D. and H. Steen (2011). " Optimization of cell lysis and protein digestion protocols for the analysis ofHeLa S3 cells by LC-MS / MS." PROTEOMICS 11(24): 4726-4730. In some embodiments, proteins or peptides are immunoprecipitated. In some embodiments, centrifugation of precipitated proteins is followed by discarding of the supernatant and subsequent washing of the pellet fraction (e.g., washing using chloroform / methanol or trichloroacetic acid).
[0340]
[0136] In some embodiments, a protein sample (e.g., a protein sample comprising a target protein) may be purified, e.g., following lysis, in a process in accordance with the instant disclosure. In some embodiments, a protein sample may be purified using chromatography (e.g., affinity chromatography that selectively binds the protein sample) or electrophoresis. In some embodiments, a protein sample may be purified in the presence of precipitating agents. In some embodiments, after a purification step or method, a protein sample may be washed and / or released from a purification matrix (e.g., affinity chromatography matrix) using an elution buffer. In some embodiments, a purification step or method may comprise the use of a reversibly switchable polymer, such as an electroactive polymer. In some embodiments, a protein sample may be initially purified by electrophoretic passage of a protein sample through a porous matrix (e.g., cellulose acetate, agarose, acrylamide).
[0341]
[0137] In some embodiments, exposing the protein to the protein digestion agent comprises exposing the protein sample to the protein digestion agent.
[0342]
[0138] In some embodiments, the method further comprises exposing the protein to a surfactant before exposing the protein to the protein digestion agent. In some embodiments, the surfactant is selected from RapiGest (e.g., RapiGest SF (Waters)), sodium dodecyl sulfate (SDS), sodium deoxycholate, Sarkosyl, n-octyl glucoside (OG), Triton X-100 (TX-100), Triton X-l 14 (IX- 114), 3-[(3- R0708.70180WO00 / R0708.70180US01 39 / 122
[0343] #14644680vl cholamidopropyl)dimethylammonio] - 1 -propanesulfonate (CHAPS), 3 - [(3 -cholamidopropyl)dimethylammonio] -2 -hydroxy- 1 -propanesulfonate (CHAPSO), dodecyl maltoside (DDM), ProteaseMAX (Promega), Tergitol-type NP-40 (NP-40), polysorbate 20 (e.g., Tween 20), polysorbate 80 (e.g., Tween 80), Brij 35, and Brij 58. In some embodiments, the surfactant comprises a non-ionic detergent. In some embodiments, the surfactant comprises a non-ionic detergent selected from n-octyl glucoside (OG), Triton X-100 (TX-100), Triton X-l 14 (TX-114), dodecyl maltoside (DDM), Tergitol-type NP-40 (NP-40), polysorbate 20 (e.g., Tween 20), polysorbate 80 (e.g., Tween 80), Brij 35, and Brij 58. In some embodiments, the surfactant comprises an ionic detergent. In some embodiments, the surfactant comprises an ionic detergent selected from RapiGest (e.g., RapiGest SF (Waters)), sodium dodecyl sulfate (SDS), sodium deoxycholate, Sarkosyl, and ProteaseMAX (Promega). In some embodiments, the surfactant comprises a zwitterionic detergent. In some embodiments, the surfactant comprises a zwitterionic detergent selected from 3-[(3-cholamidopropyl)dimethylammonio]-l-propane sulfonate (CHAPS) and 3-[(3-cholamidopropyl)dimethylammonio]-2-hydroxy-l-propane sulfonate (CHAPSO). In some embodiments, the surfactant comprises RapiGest (e.g., RapiGest SF (Waters)). In some embodiments, the surfactant comprises sodium dodecyl sulfate (SDS). In some embodiments, the surfactant comprises sodium deoxycholate. In some embodiments, the surfactant comprises Sarkosyl. In some embodiments, the surfactant comprises n-octyl glucoside (OG). In some embodiments, the surfactant comprises Triton X-100 (TX-100). In some embodiments, the surfactant comprises Triton X-l 14 (TX-114). In some embodiments, the surfactant comprises 3-[(3-cholamidopropyl)dimethylammonio]-l -propane sulfonate (CHAPS). In some embodiments, the surfactant comprises 3- [(3 -cholamidopropyl)dimethylammonio] -2 -hydroxy- 1 -propanesulfonate (CHAPSO). In some embodiments, the surfactant comprises dodecyl maltoside (DDM). In some embodiments, the surfactant comprises ProteaseMAX (Promega). In some embodiments, the surfactant comprises Tergitol-type NP-40 (NP-40). In some embodiments, the surfactant comprises polysorbate 20 (e.g., Tween 20). In some embodiments, the surfactant comprises polysorbate 80 (e.g., Tween 80). In some embodiments, the surfactant comprises Brij 35. In some embodiments, the surfactant comprises Brij 58.
[0344]
[0139] In some embodiments described herein, protein samples are buffered to maintain pH within particular ranges. For instance, in some embodiments, protein samples are buffered to maintain pH greater than or equal to 6, greater than or equal to 7, greater than or equal to 8, greater than or equal to 9, greater than or equal to 10, and / or greater at room temperature. In some embodiments, protein samples are buffered to maintain pH less than or equal to 11, less than or equal to 10, less than or equal to 9, less than or equal to 8, less than or equal to 7, and / or less at room temperature. Combinations of these ranges are possible. For example, in some embodiments, protein samples are buffered to maintain a pH of between 6 and 9.
[0345]
[0140] In some embodiments described herein, a protein sample may be buffered to a first pH range for a first step, and buffered to a second pH range for a second step. For example, in some embodiments, a protein sample is buffered to a pH of 6 to 9 during incubation, and is then buffered to a pH of between 10 and 11 for a derivatization step. In some embodiments, the protein sample is buffered to a desirable pH R0708.70180WO00 / R0708.70180US01 40 / 122
[0346] #14644680vl range for three, for four, for five, for six, for seven, for eight, for nine, and / or for ten or more steps. For example, in some embodiments, a protein sample is buffered to a pH of 6 to 9 during incubation, and is then buffered to a pH of between 10 and 11 for a derivatization step, before being buffered to a pH of 7-8 for an immobilization complex forming step and a purification step.
[0347]
[0141] Protein samples may be buffered with any buffers suitable to the desired pH range of a protein sample. For instance, in some embodiments it may be desirable to maintain a pH of between 6 and 9 for a protein sample. Exemplary buffers appropriate to such pH ranges may comprise: HEPES buffer, phosphate buffers (e.g., PBS), Tris, Bis-Tris, carbonate buffers (e.g., buffers comprising: carbonates, such as sodium or potassium carbonate; and / or bicarbonates, such as sodium bicarbonate), which may be used separately or in combination to stabilize pH within a desired range. In some embodiments, a buffer appropriate to such pH ranges comprises: HEPES buffer, phosphate buffers (e.g., PBS), and / or carbonate buffers (e.g., buffers comprising: carbonates, such as sodium or potassium carbonate; and / or bicarbonates, such as sodium bicarbonate). One of ordinary skill in the art would be familiar with these and many other buffer systems, and the use of un-listed buffer systems is contemplated here.
[0348]
[0142] Protein samples may be diluted prior to digestion (e.g., prior to exposing the protein to a protein digestion agent). In some embodiments, the protein sample is diluted with a buffer. In some embodiments, the protein sample is diluted with a surfactant solution. In some embodiments, the surfactant solution comprises a surfactant selected from RapiGest (e.g., RapiGest SF (Waters)), sodium dodecyl sulfate (SDS), sodium deoxycholate, Sarkosyl, n-octyl glucoside (OG), Triton X-100 (TX-100), Triton X-114 (TX-114), 3-[(3-cholamidopropyl)dimethylammonio]-l-propanesulfonate (CHAPS), 3-[(3-cholamidopropyl)dimethylammonio] -2 -hydroxy- 1 -propanesulfonate (CHAPSO), dodecyl maltoside (DDM), ProteaseMAX (Promega), Tergitol-type NP-40 (NP-40), polysorbate 20 (e.g., Tween 20), polysorbate 80 (e.g., Tween 80), Brij 35, and Brij 58. In some embodiments, the surfactant solution comprises a non-ionic detergent. In some embodiments, the surfactant solution comprises a non-ionic detergent selected from n-octyl glucoside (OG), Triton X-100 (TX-100), Triton X-l 14 (TX-114), dodecyl maltoside (DDM), Tergitol-type NP-40 (NP-40), polysorbate 20 (e.g., Tween 20), polysorbate 80 (e.g., Tween 80), Brij 35, and Brij 58. In some embodiments, the surfactant solution comprises an ionic detergent. In some embodiments, the surfactant solution comprises an ionic detergent selected from RapiGest (e.g., RapiGest SF (Waters)), sodium dodecyl sulfate (SDS), sodium deoxycholate, Sarkosyl, and ProteaseMAX (Promega). In some embodiments, the surfactant solution comprises a zwitterionic detergent. In some embodiments, the surfactant solution comprises a zwitterionic detergent selected from 3-[(3-cholamidopropyl)dimethylammonio]-l-propanesulfonate (CHAPS) and 3-[(3-cholamidopropyl)dimethylammonio]-2-hydroxy-l-propanesulfonate (CHAPSO). In some embodiments, the surfactant solution comprises RapiGest (e.g., RapiGest SF (Waters)). In some embodiments, the surfactant solution comprises sodium dodecyl sulfate (SDS). In some embodiments, the surfactant solution comprises sodium deoxycholate. In some embodiments, the surfactant solution comprises Sarkosyl. In some embodiments, the surfactant solution comprises n-octyl glucoside (OG). In some embodiments, the surfactant solution comprises Triton X-100 (TX-100). In some embodiments, the R0708.70180WO00 / R0708.70180US01 41 / 122
[0349] #14644680vl surfactant solution comprises Triton X-l 14 (TX-114). In some embodiments, the surfactant solution comprises 3- [(3 -cholamidopropyl)dimethylammonio]-l -propanesulfonate (CHAPS). In some embodiments, the surfactant solution comprises 3-[(3-cholamidopropyl)dimethylammonio]-2-hydroxy-l-propane sulfonate (CHAPSO). In some embodiments, the surfactant solution comprises dodecyl maltoside (DDM). In some embodiments, the surfactant solution comprises ProteaseMAX (Promega). In some embodiments, the surfactant solution comprises Tergitol-type NP-40 (NP-40). In some embodiments, the surfactant solution comprises polysorbate 20 (e.g., Tween 20). In some embodiments, the surfactant solution comprises polysorbate 80 (e.g., Tween 80). In some embodiments, the surfactant solution comprises Brij 35. In some embodiments, the surfactant solution comprises Brij 58.
[0350] Amino Acid Side Chain Reduction, Amino Acid Side Chain Capping, and Protein Digestion
[0351]
[0143] In some embodiments, the method comprises exposing the protein to a reducing agent.
[0352]
[0144] In some embodiments, the reducing agent reduces an amino acid side chain of the protein to form a reduced amino acid side chain of the protein. Any suitable reducing agent may be used to reduce a protein.
[0353]
[0145] In some embodiments, the reducing agent is suitable for reducing a disulfide bond. In some embodiments, the reducing agent may reversibly reduce a disulfide bond. Suitable reversable reducing agents may comprise compounds such as dithiothreitol (DTT), (3-mercaptoethanol (BME), and / or Glutathione (GSH).
[0354]
[0146] In some embodiments, the reducing agent may irreversibly reduce a disulfide bond. Suitable irreversible reducing agents may comprise compounds such as tris(2-carboxyethyl)phosphine (TCEP) or tris(hydroxypropyl)phosphine (THP). In some specific embodiments, the reducing agent comprises tris(2-carboxyethyl)phosphine (TCEP). In some embodiments, the reducing agent comprises tris(hydroxypropyl)phosphine (THP). In some embodiments, the reducing agent does not comprise tris(2-carboxyethyl)phosphine (TCEP). In some embodiments, the reducing agent comprises tris(hydroxypropyl)phosphine (THP) and does not comprise tris(2-carboxyethyl)phosphine (TCEP).
[0355]
[0147] In some embodiments, the method comprises exposing the protein to an amino acid side chain capping agent.
[0356]
[0148] In some embodiments, the amino acid side chain capping agent forms a covalent bond with the reduced amino acid side chain to form a capped amino acid side chain of the protein. Any suitable amino acid side chain capping agent may be used to cap amino acid side chains of a protein within a polypeptide sample. In some embodiments, the protein comprises the capped amino acid side chain.
[0357]
[0149] In some embodiments, the amino acid side chain capping agent prevents the formation of disulfide bonds. In some embodiments, the amino acid side chain capping agent prevents the amino acid side chain from undergoing further reactivity such as nucleophile / electrophile or redox reactivity. In some embodiments, the amino acid side chain capping agent comprises a cysteine alkylation agent. In some embodiments, the amino acid side chain capping agent is a sulfhydryl -reactive alkylating reagent (e.g., a cysteine alkylation agent). For instance, in some embodiments, the amino acid side chain capping agent comprises a haloacetamide (e.g., chloroacetamide, iodoacetamide) or a haloacetate / haloacetic acid (e.g., R0708.70180WO00 / R0708.70180US01 42 / 122
[0358] #14644680vl chloroacetate / chloroacetic acid, iodoacetate / iodoacetic acid). In some embodiments, the amino acid side chain capping agent comprises iodoacetamide (IAA) and / or chloroacetamide (CAA). In some embodiments, the amino acid side chain capping agent is an aromatic benzyl halide. For example, the amino acid side chain capping agent may be an aromatic benzyl halide derivative based on a benzene aromatic group, a pyridine aromatic group, a pyrazine aromatic group, and the like. Other examples of suitable cysteine alkylating agents include 4-vinylpyridine, acrylamide, and methanethiosulfonate.
[0359]
[0150] In some embodiments, the amino acid side chain capping agent comprises an alkyl halide.
[0360]
[0151] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III):
[0361]
[0362] or a salt thereof, wherein:
[0363] Y is a leaving group;
[0364] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; and each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group.
[0365]
[0152] As generally described herein, Y is a leaving group.
[0366]
[0153] In some embodiments, Y is halo (e.g., -F, -Cl, -Br, -I), an activated substituted hydroxyl group (e.g., -OC(=O)SRaa, -OC(=O)Raa, -OCO2Raa, -OC(=O)N(Rbb)2, -OC(=NRbb)Raa, -OC(=NRbb)ORaa, -OC(=NRbb)N(Rbb)2, -OS(=O)Raa, -OSO2Raa, -OP(RCC)2, -OP(RCC)3, -OP(=O)2Raa, -OP(=O)(Raa)2, -OP(=O)(ORCC)2, -OP(=O)2N(Rbb)2, -OP(=O)(NRbb)2, wherein Raa, Rbb, and Rccare as defined herein), or a sulfonic acid ester (e.g., toluenesulfonate (tosylate, -OTs), methane sulfonate (mesylate, -OMs), p-bromobenzenesulfonyloxy (brosylate, -OBs), -OS(=O)2(CF2)3CF3 (nonaflate, -ONf), trifluoromethanesulfonate (triflate, -OTf)).
[0367]
[0154] In some embodiments, Y is halo (e.g., -F, -Cl, -Br, -I). In some embodiments, Y is -Cl, -Br, or-I. In some embodiments, Y is -Cl. In some embodiments, Y is -Br. In some embodiments, Y is
[0368] -I.
[0369]
[0155] As generally described herein, L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene.
[0370]
[0156] In some embodiments, L1is optionally substituted Ci-6 alkylene. In some embodiments, L1is substituted Ci-6 alkylene. In some embodiments, L1is unsubstituted Ci-6 alkylene. In some embodiments, L1is optionally substituted C1-3 alkylene. In some embodiments, L1is substituted C1-3 alkylene. In some embodiments, L1is unsubstituted C1-3 alkylene. In some embodiments, L1is methylene, ethylene, or n-propylene. In some embodiments, L1is ethylene.
[0371]
[0157] In some embodiments, L1is optionally substituted C1-6 heteroalkylene. In some embodiments, L1is substituted C1-6 heteroalkylene. In some embodiments, L1is unsubstituted C1-6 heteroalkylene. In some embodiments, L1is optionally substituted C1-3 heteroalkylene. In some embodiments, L1is substituted Ci-3 heteroalkylene. In some embodiments, L1is unsubstituted C1-3 heteroalkylene.
[0372]
[0158] As generally described herein, each instance of R1is independently hydrogen, optionally R0708.70180WO00 / R0708.70180US01 43 / 122
[0373] #14644680vl substituted aliphatic, or a nitrogen protecting group.
[0374]
[0159] In some embodiments, at least one instance of R1is hydrogen. In some embodiments, each instance of R1is independently hydrogen.
[0375]
[0160] In some embodiments, at least one instance of R1is optionally substituted aliphatic. In some embodiments, at least one instance of R1is optionally substituted alkyl. In some embodiments, at least one instance of R1is a nitrogen protecting group.
[0376]
[0161] In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-a):
[0377]
[0378] (Ill-a),
[0379] or a salt thereof, wherein:
[0380] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; and each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group.
[0381]
[0162] In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-a), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene or unsubstituted Ci-6 heteroalkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III -a), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-a), or a salt thereof, wherein L1is unsubstituted C1-3 alkylene.
[0382]
[0163] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-b):
[0383] KL1-N(R1)2
[0384]
[0385] (III-b),
[0386] or a salt thereof, wherein:
[0387] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; and each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group.
[0388]
[0164] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-b), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene or unsubstituted Ci-6 heteroalkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-b), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-b), or a salt thereof, wherein L1is unsubstituted C1-3 alkylene.
[0389]
[0165] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-c):
[0390] Y'
[0391]
[0392] X / XN(R’)2 aII<)
[0393] or a salt thereof, wherein:
[0394] Y is a leaving group; and
[0395] each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group.
[0396] R0708.70180WO00 / R0708.70180US01 44 / 122
[0397] #14644680vl
[0166] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-c), or a salt thereof, wherein Y is halo. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-c), or a salt thereof, wherein Y is -Br. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-c), or a salt thereof, wherein Y is -I.
[0398]
[0167] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d):
[0399] Y^L1 / NH2
[0400]
[0401] (III-d),
[0402] or a salt thereof, wherein:
[0403] Y is a leaving group; and
[0404] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene.
[0405]
[0168] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d), or a salt thereof, wherein Y is halo. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d), or a salt thereof, wherein Y is -Br. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d), or a salt thereof, wherein Y is -I. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene or unsubstituted Ci-6 heteroalkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d), or a salt thereof, wherein L1is unsubstituted C1-3 alkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d), or a salt thereof, wherein Y is halo; and L1is unsubstituted C1-6 alkylene or unsubstituted C1-6 heteroalkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d), or a salt thereof, wherein Y is -Br or -I; and L1is unsubstituted C1-6 alkylene.
[0406]
[0169] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-e):
[0407]
[0408] (III-e),
[0409] or a salt thereof, wherein Y is a leaving group.
[0410]
[0170] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-e), or a salt thereof, wherein Y is halo.
[0411]
[0171] In some embodiments, the amino acid side chain capping agent is a compound of formula:
[0412] Br'^'NH
[0413]
[0414] 2 orl'^xNH2
[0415] or a salt thereof.
[0416]
[0172] In some embodiments, the amino acid side chain capping agent comprises bromoethylamine (BEA). In some embodiments, the amino acid side chain capping agent is bromoethylamine (BEA).
[0417]
[0173] In some embodiments, the amino acid side chain capping agent comprises iodoethylamine (IEA). In some embodiments, the amino acid side chain capping agent is iodoethylamine (IEA).
[0418]
[0174] In some embodiments, the protein comprises fewer than 5% lysine residues. In some embodiments, the protein comprises a continuous sequence of amino acids that does not comprise a lysine residue. In R0708.70180WO00 / R0708.70180US01 45 / 122
[0419] #14644680vl some embodiments, the continuous sequence of amino acids that does not comprise a lysine residue comprises at least 25, 50, 75, 100, 125, 150, 175, or 200 amino acid residues. In some embodiments, the continuous sequence of amino acids that does not comprise a lysine residue comprises at least 150 amino acid residues.
[0420]
[0175] In some embodiments, the method comprises exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample.
[0421]
[0176] In some embodiments, the protein digestion agent induces proteolysis of the protein to form one or more capped polypeptides, thereby forming the digested polypeptide sample. In general, protein digestion can be conducted using any known method, but typically will involve a non-enzymatic or an enzymatic method.
[0422]
[0177] In some embodiments, the protein digestion agent is a non-enzymatic protein digestion agent. Approaches for non-enzymatic digestion include, but are not limited to, acid hydrolysis and / or cleavage using a non-enzymatic protein digestion agent such as cyanogen bromide, hydroxylamine, iodosobenzoic acid, dimethyl sulfoxide-hydrochloric acid, BNPS-skatole [2-(2 -nitrophenylsulfenyl) -3-methylindole], or 2-nitro-5 -thiocyanobenzoic acid. Electro-physical digestion methods may be employed as well, including electrochemical oxidation and / or digestion in conjunction with microwaves.
[0423]
[0178] In some embodiments, the protein digestion agent is an enzymatic protein digestion agent. For example, in some embodiments, the protein digestion agent comprises a protease, which can fragment a protein into component peptides. In some embodiments, the protease comprises trypsin, chymotrypsin, Lys-C, Lys-N, Asp-N, Glu-C, and / or Arg-C. In some embodiments, the protease comprises trypsin, Lys-C, Asp-N, and / or Glu-C. In some embodiments, the protease comprises trypsin. In some embodiments, the protease comprises Lys-C. In some embodiments, the protease comprises Asp-N. In some embodiments, the protease comprises Glu-C. In some embodiments, the protease is Lys-C. Enzymatic fragmentation / digestion methods may be selected and adjusted for ease of use, speed, automation and / or effectiveness. In some embodiments, enzymatic methods include enzyme immobilization on solid substrates. An enzymatic digestion may utilize any number or combination of enzymes and may further comprise any of the known non-enzymatic methods.
[0424]
[0179] In some embodiments, one or more capped polypeptides are of Formula (I):
[0425] , S^L,_N(R1)2
[0426] RYVRN
[0427]
[0428] 0 H(I),
[0429] or a salt thereof, wherein:
[0430] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene;
[0431] each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group;
[0432] Rcis -OH, an amino acid moiety, or a peptide; and
[0433] RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0434] R0708.70180WO00 / R0708.70180US01 46 / 122
[0435] #14644680vl
[0180] As generally described herein, L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene.
[0436]
[0181] In some embodiments, L1is optionally substituted Ci-6 alkylene. In some embodiments, L1is substituted Ci-6 alkylene. In some embodiments, L1is unsubstituted Ci-6 alkylene. In some embodiments, L1is optionally substituted C1-3 alkylene. In some embodiments, L1is substituted C1-3 alkylene. In some embodiments, L1is unsubstituted C1-3 alkylene. In some embodiments, L1is methylene, ethylene, or n-propylene. In some embodiments, L1is ethylene.
[0437]
[0182] In some embodiments, L1is optionally substituted Ci-6 heteroalkylene. In some embodiments, L1is substituted Ci-6 heteroalkylene. In some embodiments, L1is unsubstituted Ci-6 heteroalkylene. In some embodiments, L1is optionally substituted C1-3 heteroalkylene. In some embodiments, L1is substituted Ci-3 heteroalkylene. In some embodiments, L1is unsubstituted C1-3 heteroalkylene.
[0438]
[0183] As generally described herein, each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group.
[0439]
[0184] In some embodiments, at least one instance of R1is hydrogen. In some embodiments, each instance of R1is independently hydrogen.
[0440]
[0185] In some embodiments, at least one instance of R1is optionally substituted aliphatic. In some embodiments, at least one instance of R1is optionally substituted alkyl. In some embodiments, at least one instance of R1is a nitrogen protecting group.
[0441]
[0186] In some embodiments, Rcis -OH. In some embodiments, Rcis an amino acid moiety or a peptide. In some embodiments, Rcis an amino acid moiety. In some embodiments, Rcis a peptide. In some embodiments, RNis hydrogen. In some embodiments, RNis a nitrogen protecting group. In some embodiments, RNis an amino acid moiety or a peptide. In some embodiments, RNis an amino acid moiety. In some embodiments, RNis a peptide. In some embodiments, Rcis -OH, and RNis an amino acid moiety or a peptide. In some embodiments, Rcis -OH, and RNis an amino acid moiety. In some embodiments, Rcis -OH, and RNis a peptide. In some embodiments, Rcis an amino acid moiety or a peptide, and RNis hydrogen. In some embodiments, Rcis an amino acid moiety, and RNis hydrogen. In some embodiments, Rcis a peptide, and RNis hydrogen. In some embodiments, Rcis an amino acid moiety or a peptide; and RNis an amino acid moiety or a peptide. In some embodiments, Rcis an amino acid moiety; and RNis an amino acid moiety. In some embodiments, Rcis an amino acid moiety; and RNis a peptide. In some embodiments, Rcis a peptide; and RNis an amino acid moiety. In some embodiments, Rcis a peptide; and RNis a peptide.
[0442]
[0187] In some embodiments, RNis a peptide comprising fewer than 5% lysine residues. In some embodiments, Rcis a peptide comprising fewer than 5% lysine residues. In some embodiments, RNis a peptide comprising fewer than 5% lysine residues and / or Rcis a peptide comprising fewer than 5% lysine residues. In some embodiments, RNis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue. In some embodiments, Rcis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue. In some embodiments, RNis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue and / or Rcis a R0708.70180WO00 / R0708.70180US01 47 / 122
[0443] #14644680vl peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue. In some embodiments, RNis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue. In some embodiments, Rcis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue. In some embodiments, RNis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue and / or Rcis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue.
[0444]
[0188] In some embodiments, one or more capped polypeptides are of Formula (I-a):
[0445] , S^L1^NH2
[0446] RcJL ^RN
[0447] n N
[0448] AH
[0449]
[0450] 0(I-a),
[0451] or a salt thereof, wherein:
[0452] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene;
[0453] Rcis -OH, an amino acid moiety, or a peptide; and
[0454] RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0455]
[0189] In some embodiments, one or more capped polypeptides are of Formula (I-a), or salt thereof, wherein Rcis an amino acid moiety or a peptide. In some embodiments, one or more capped polypeptides are of Formula (I-a), or salt thereof, wherein RNis an amino acid moiety or a peptide. In some embodiments, one or more capped polypeptides are of Formula (I-a), or salt thereof, wherein Rcis an amino acid moiety or a peptide; and RNis an amino acid moiety or a peptide.
[0456]
[0190] In some embodiments, one or more capped polypeptides are of Formula (I-b):
[0457] <S^ NH2
[0458] DC 1 RN
[0459]
[0460] 0 H(I-b).
[0461] or a salt thereof, wherein:
[0462] Rcis -OH, an amino acid moiety, or a peptide; and
[0463] RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0464]
[0191] In some embodiments, one or more capped polypeptides are of Formula (I-b), or salt thereof, wherein Rcis an amino acid moiety or a peptide. In some embodiments, one or more capped polypeptides are of Formula (I-b), or salt thereof, wherein RNis an amino acid moiety or a peptide. In some embodiments, one or more capped polypeptides are of Formula (I-b), or salt thereof, wherein Rcis an amino acid moiety or a peptide; and RNis an amino acid moiety or a peptide.
[0465]
[0192] In some embodiments, the method further comprises binding the protein to a solid substrate. In some embodiments, the method further comprises binding the protein to a solid substrate before, at the same time as, or after exposing the protein to the protein digestion agent. In some embodiments, the method further comprises binding the protein to a solid substrate before exposing the protein to the R0708.70180WO00 / R0708.70180US01 48 / 122
[0466] #14644680vl protein digestion agent. In some embodiments, the method further comprises binding the protein to a solid substrate after exposing the protein to the protein digestion agent. In some embodiments, the method further comprises binding the protein or the one or more capped polypeptides to a solid substrate after exposing the protein to the protein digestion agent. In some embodiments, the solid substrate is a bead. In some embodiments, the solid substrate is or comprises a polymeric material (e.g., a bead comprising a polymeric material). In some embodiments, the solid substrate comprises polystyrene. In some embodiments, the solid substrate is a resin. In some embodiments, the bead is paramagnetic. In some embodiments, the bead comprises magnetite. In some embodiments, the solid substrate comprises polystyrene and magnetite. In some embodiments, the binding comprises anoncovalent interaction between functional groups on a surface of the solid substrate and the protein or the one or more capped polypeptides. Advantageously, a resin or bead may be removed by filtration.
[0467]
[0193] In some embodiments, the method further comprises purifying the digested polypeptide sample to form a purified digested polypeptide sample. In some embodiments, purifying the digested polypeptide sample to form the digested polypeptide sample comprises removing at least some of any salts, detergents, or chaotropes present in the digested polypeptide sample. In some embodiments, purifying the digested polypeptide sample to form the digested polypeptide sample comprises exposure of the digested polypeptide sample to a Cl 8 matrix. In some embodiments, purifying the digested polypeptide sample to form the digested polypeptide sample comprises passing the digested polypeptide sample through a Cl 8 column.
[0468]
[0194] In some embodiments, the method comprises an incubation step, wherein:
[0469] the reducing agent reduces an amino acid side chain of the protein to form a reduced amino acid side chain of the protein;
[0470] the amino acid side chain capping agent forms a covalent bond with the reduced amino acid side chain to form a capped amino acid side chain of the protein; and
[0471] the protein digestion agent induces proteolysis of the protein to form one or more capped polypeptides, thereby forming the digested polypeptide sample.
[0472]
[0195] In some embodiments, the incubating step comprises maintaining the protein and / or the one or more capped polypeptides at a temperature greater than or equal to 20°C, greater than or equal to 25°C, greater than or equal to 30°C, greater than or equal to 35 °C, or greater than or equal to 37°C. In some embodiments, an incubating step comprises maintaining the protein and / or the one or more capped polypeptides at a temperature less than or equal to 70°C, less than or equal to 50°C, less than or equal to 37°C, less than or equal to 35°C, or less than or equal to 30°C. Combinations of these ranges are possible. For example, an incubating step may comprise maintaining the protein and / or the one or more capped polypeptides at a temperature greater than or equal to 20°C and less than or equal to 70°C. In some embodiments, an incubating step comprises maintaining the protein and / or the one or more capped polypeptides at a temperature within the above-mentioned ranges (e.g., 37°C) for at least 1 minute, at least 2 minutes, at least 5 minutes, at least 10 minutes, at least 15 minutes, at least 20 minutes, at least 25 minutes, at least 30 minutes, at least 45 minutes, at least 1 hour, at least 2 hours, at least 3 hours, at least R0708.70180WO00 / R0708.70180US01 49 / 122
[0473] #14644680vl 4 hours, at least 5 hours, at least 6 hours, or greater. In some embodiments, an incubating step comprises maintaining the protein and / or the one or more capped polypeptides at a temperature within the above-mentioned ranges (e.g., 37°C) for less than or equal to 20 hours, less than or equal to 15 hours, less than or equal to 10 hours, or less. Combinations (e.g., maintaining an above-mentioned temperature for at least one minute and less than or equal to 20 hours, at least 6 hours and less than or equal to 10 hours) are possible.
[0474] Polypeptide Derivatization
[0475]
[0196] In some embodiments, the method comprises derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides.
[0476]
[0197] In some embodiments, the derivatizing comprises derivatizing an amino acid side chain of the one or more capped polypeptides using a derivatization agent to form an unquenched mixture comprising one or more derivatized polypeptides. In some embodiments, the derivatizing comprises derivatizing a terminal end (e.g., its N-terminal end or its C-terminal end) of the one or more capped polypeptides using a derivatization agent to form an unquenched mixture comprising one or more derivatized polypeptides.
[0477]
[0198] In some embodiments, a derivatization agent is used to derivatize an amino acid side chain (e.g., by one or more of the methods described below).
[0478]
[0199] In some embodiments, the derivatization agent is or comprises an azide transfer agent (e.g., imidazole -1 -sulfonyl azide (ISA), benzene sulfonyl azide). In some embodiments, the azide transfer agent comprises imidazole- 1 -sulfonyl azide (ISA). In some embodiments, the azide transfer agent is imidazole- 1-sulfonyl azide (ISA). In some embodiments, the azide transfer agent is imidazole- 1 -sulfonyl azide tetrafluoroborate.
[0479]
[0200] In some embodiments, the derivatizing comprises exposing the one or more capped polypeptides to the derivatization agent and one or more derivatization reagents.
[0480]
[0201] In some embodiments, the one or more derivatization reagents comprises a pH adjusting reagent. In the context of the present disclosure, a pH adjusting reagent may comprise any chemical suitable for adjusting the pH of a solution to a desired value for a chemical reaction. In some embodiments, a pH adjusting reagent comprises a base (e.g., a strong base, a weak base). In some embodiments, a pH adjusting reagent comprises an acid (e.g., strong acid, a weak acid). In some embodiments, a pH adjusting reagent comprises a buffer. In some embodiments, the pH adjusting reagent is potassium carbonate (K2CO3).
[0481]
[0202] In some embodiments, the one or more derivatization reagents comprises a catalyst for a derivatization reaction between the amino acid side chain of the one or more capped polypeptides and the derivatization agent. In some embodiments, the one or more derivatization reagents comprises a source of Cu2+. In some embodiments, the source of Cu2+is CuCl2, CuBr2, Cu(OH)2, or CuSO4. In some embodiments, the source of Cu2+is copper (II) sulfate (CuSO4). In some embodiments, the one or more derivatization reagents comprises a source of Ni2+. In some embodiments, the source of Ni2+is nickel (II) acetate (Ni(OAc)2).
[0482] R0708.70180WO00 / R0708.70180US01 50 / 122
[0483] #14644680vl
[0203] In some embodiments, the derivatizing comprises exposing the one or more capped polypeptides to, in order: a pH adjusting reagent, one or more derivatization reagents, and the derivatization agent. In some embodiments, the derivatizing comprises exposing the one or more capped polypeptides to, in order: a pH adjusting reagent, a source of Cu2+, and an azide transfer agent. In some embodiments, the derivatizing comprises exposing the one or more capped polypeptides to, in order: potassium carbonate (K2CO3), nickel (II) acetate (Ni(OAc)2), and imidazole- 1 -sulfonyl azide (ISA). In some embodiments, the derivatizing comprises exposing the one or more capped polypeptides to, in order: a pH adjusting reagent, a source of Ni2+, and an azide transfer agent. In some embodiments, the derivatizing comprises exposing the one or more capped polypeptides to, in order: potassium carbonate (K2CO3), nickel (II) acetate (Ni(OAc)2), and imidazole- 1 -sulfonyl azide (ISA).
[0484]
[0204] In some embodiments, the derivatizing comprises a temperature of about 20-30°C, e.g., 20-25°C, 22-27°C, 25-30°C, 20°C, 21°C, 22°C, 23°C, 24°C, 25°C, 26°C, 27°C, 28°C, 29°C, or 30°C. In some embodiments, the derivatizing comprises a temperature of about 30-60 minutes, e.g., 30-35 minutes, 35-40 minutes, 40-45 minutes, 45-50 minutes, 50-55 minutes, or 55-60 minutes. In a particular embodiment, the derivatizing comprises reaction at ambient temperature (e.g., about 25 °C) for about 60 minutes.
[0485]
[0205] In some embodiments, the N-terminal selectivity of the diazo transfer reaction is at least about 90%.
[0486]
[0206] In some embodiments, one or more derivatized polypeptides are of Formula (IV):
[0487] RVVRN
[0488]
[0489] sH
[0490] or a salt thereof, wherein:
[0491] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene;
[0492] Rcis -OH, an amino acid moiety, or a peptide; and
[0493] RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0494]
[0207] In some embodiments, one or more derivatized polypeptides are of Formula (IV-a):
[0495] <S^ N3
[0496] □C 1 RN
[0497]
[0498] 0 H(IV-a),
[0499] or a salt thereof, wherein:
[0500] Rcis -OH, an amino acid moiety, or a peptide; and
[0501] RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0502]
[0208] In some embodiments, one or more capped polypeptides are of Formula (IV-a), or salt thereof, wherein Rcis an amino acid moiety or a peptide. In some embodiments, one or more capped polypeptides are of Formula (IV-a), or salt thereof, wherein RNis an amino acid moiety or a peptide. In some embodiments, one or more capped polypeptides are of Formula (IV-a), or salt thereof, wherein RcR0708.70180WO00 / R0708.70180US01 51 / 122
[0503] #14644680vl is an amino acid moiety or a peptide; and RNis an amino acid moiety or a peptide.
[0504]
[0209] In some embodiments, one or more capped polypeptides are of Formula (I):
[0505]
[0506] or a salt thereof, and one or more derivatized polypeptides are of Formula (IV):
[0507]
[0508] (IV),
[0509] or a salt thereof, wherein:
[0510] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group;
[0511] each instance of Rcis independently -OH, an amino acid moiety, or a peptide; and
[0512] each instance of RNis independently hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0513]
[0210] In some embodiments, one or more capped polypeptides are of formula:
[0514]
[0515] (I-b),
[0516] or a salt thereof, and one or more derivatized polypeptides are of formula:
[0517]
[0518] or a salt thereof, wherein:
[0519] each instance of Rcis independently -OH, an amino acid moiety, or a peptide; and
[0520] each instance of RNis independently hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0521]
[0211] In some embodiments, RNis a peptide comprising fewer than 5% lysine residues. In some embodiments, Rcis a peptide comprising fewer than 5% lysine residues. In some embodiments, RNis a peptide comprising fewer than 5% lysine residues and / or Rcis a peptide comprising fewer than 5% lysine residues. In some embodiments, RNis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue. In some embodiments, Rcis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue. In some embodiments, RNis a peptide R0708.70180WO00 / R0708.70180US01 52 / 122
[0522] #14644680vl comprising a continuous sequence of amino acids that does not comprise a lysine residue and / or Rcis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue. In some embodiments, RNis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue. In some embodiments, Rcis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue. In some embodiments, RNis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue and / or Rcis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue.
[0523]
[0212] In some embodiments, the unquenched mixture further comprises excess derivatization agent.
[0524]
[0213] In some embodiments, the method further comprises quenching (i.e., neutralizing) the unquenched mixture to form a quenched mixture by removing at least some of the excess derivatization agent. In some embodiments, the method further comprises quenching ( / . e., neutralizing) the unquenched mixture to form a quenched mixture by removing at least some of the excess derivatization agent, wherein the derivatization agent is an azide transfer agent (e.g., imidazole- 1 -sulfonyl azide (ISA)).
[0525]
[0214] In some embodiments, the quenching comprises reacting the at least some of the excess derivatization agent with functional groups on a surface of a solid substrate. In some embodiments, the solid substrate is a bead. In some embodiments, the solid substrate is or comprises a polymeric material (e.g., a bead comprising a polymeric material). In some embodiments, the solid substrate is a polystyrene bead. In some embodiments, the functional groups of the solid substrate comprise amine groups. In some embodiments, the solid substrate is a polystyrene polyamine bead. In some embodiments, the solid substrate is a resin. Advantageously, a resin or bead may be removed by fdtration.
[0526]
[0215] Alternatively, in some embodiments, the quenching comprises exposing the unquenched mixture comprising the excess derivatization agent to a C 18 matrix. Alternatively, in some embodiments, the quenching comprises passing the unquenched mixture comprising the excess derivatization agent through a Cl 8 column.
[0527]
[0216] In some embodiments, the method further comprises purifying the derivatized polypeptide sample to form a purified derivatized polypeptide sample. In some embodiments, purifying the derivatized polypeptide sample to form the purified derivatized polypeptide sample comprises removing at least some of any remaining non-derivatized polypeptides of the derivatized polypeptide sample.
[0528]
[0217] In some embodiments, purifying the derivatized polypeptide sample to form the purified derivatized polypeptide sample comprises passing the derivatized polypeptide sample through a size exclusion medium. In some embodiments, the size exclusion medium may be a column. The column may be a desalting column. In some embodiments, the column is a Zeba column (e.g., a Zeba 7 kDa or a Zeba 40 kDa column). In some embodiments, the size exclusion medium is part of a fluidic device. In some embodiments, the size exclusion medium is part of a system, but is not part of a fluidic device of that system.
[0529]
[0218] In some embodiments, purifying a protein comprises purification via immunoprecipitation. In some embodiments, immunoprecipitation comprises precipitating a target protein out of sample (e.g., a sample R0708.70180WO00 / R0708.70180US01 53 / 122
[0530] #14644680vl before or after derivatization) using an antibody that specifically binds to the target protein. Conjugation to Immobilization Complex
[0531]
[0219] In some embodiments, the method comprises conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0532]
[0220] In some embodiments, conjugating the one or more derivatized polypeptides to the immobilization complex comprises exposing the one or more derivatized polypeptides to cetrimonium bromide (CTAB).
[0533]
[0221] In some embodiments, conjugating the one or more derivatized polypeptides to the immobilization complex comprises exposing the one or more derivatized polypeptides to ethylenediaminetetraacetic acid (EDTA). In some embodiments, the EDTA is present in a concentration ranging from 1 mM to 10 mM. In some embodiments, the EDTA is present in a concentration of about 5 mM.
[0534]
[0222] In some embodiments, conjugating the one or more derivatized polypeptides to the immobilization complex comprises a click chemistry reaction. In some embodiments, conjugating the one or more derivatized polypeptides to the immobilization complex comprises an azide-alkyne cycloaddition. In some embodiments, the one or more derivatized polypeptides comprise an azide.
[0535]
[0223] In some embodiments, one or more derivatized polypeptides comprise an azide moiety. In some embodiments, the method further comprises contacting the one or more derivatized polypeptides with the immobilization complex, such that the azide moiety of the one or more derivatized polypeptides reacts with the immobilization complex via click chemistry to provide one or more immobilization complex-conjugated polypeptides comprising a triazole.
[0536]
[0224] In some embodiments, the immobilization complex comprises a strained alkyne. In some embodiments, the immobilization complex comprises a cyclooctyne or an azacyclooctyne. In some embodiments, the immobilization complex comprises a cyclooctyne. In some embodiments, the immobilization complex comprises an azacyclooctyne. In some embodiments, the immobilization complex comprises dibenzoazacyclooctyne (DIBAC or DBCO), biarylazacyclooctynone (BARAC), dibenzocyclooctyne (DIBO), difluorinated cyclooctyne (DIFO), bicyclononyne (BCN), dimethoxyazacyclooctyne (DIMAC), monofluorinated cyclooctyne (MOFO), cyclooctyne (OCT), and / or aryl-less cyclooctyne (ALO). In some embodiments, the immobilization complex comprises dibenzoazacyclooctyne (DIBAC or DBCO). In some embodiments, the immobilization complex comprises biarylazacyclooctynone (BARAC). In some embodiments, the immobilization complex comprises dibenzocyclooctyne (DIBO). In some embodiments, the immobilization complex comprises difluorinated cyclooctyne (DIFO). In some embodiments, the immobilization complex comprises bicyclononyne (BCN). In some embodiments, the immobilization complex comprises dimethoxyazacyclooctyne (DIMAC). In some embodiments, the immobilization complex comprises monofluorinated cyclooctyne (MOFO). In some embodiments, the immobilization complex comprises cyclooctyne (OCT). In some embodiments, the immobilization complex comprises aryl -less cyclooctyne (ALO).
[0537]
[0225] In some embodiments, the immobilization complex comprises polyethylene glycol (PEG). In some R0708.70180WO00 / R0708.70180US01 54 / 122
[0538] #14644680vl embodiments, the immobilization complex comprises an oligonucleotide. In some embodiments, the immobilization complex comprises single-stranded DNA. In some embodiments, the immobilization complex comprises double-stranded DNA. In some embodiments, when the immobilization complex comprises single-stranded DNA, the method further comprises hybridizing a complementary DNA strand to the single-stranded DNA to obtain a compound wherein the immobilization complex comprises double -stranded DNA. In some embodiments, the single-stranded DNA is Q24 and the complementary DNA strand is Cy3B. In some embodiments, the immobilization complex further comprises biotin (e.g., bisbiotin).
[0539]
[0226] In some embodiments, the immobilization complex comprises TCO, single-stranded DNA, and biotin (e.g., bisbiotin). In some embodiments, the immobilization complex is Q24-BisBt-BCN. In some embodiments, the immobilization complex is Q24- BisBt-DBCO. In some embodiments, the immobilization complex is Q24- BisBt-TCO. Generally, the immobilization complex may comprise a branching moiety (e.g., a 1, 3, 5 -tricarboxylate moiety), wherein two branches are direct or indirect attachments to biotin moieties, and the third branch is an attachment to the water soluble moiety (e.g., a polynucleotide such as Q24). In some embodiments, the immobilization complex comprises a triazole moiety derived from the click-coupling of fragments comprising (i) a bisbiotin-azide functionalized linker and (ii) an alkyne (e.g., BCN) -functionalized polynucleotide (e.g., Q24). The click-coupled product may be derivatized to introduce a further click handle, such as BCN or DBCO.
[0540]
[0227] In another embodiment, the immobilization complex comprises DBCO, single-stranded DNA, and streptavidin (SV). In certain particular embodiments, the immobilization complex is DBCO-Q24-SV.
[0541]
[0228] In some embodiments, when the immobilization complex comprises biotin (e.g., bisbiotin), the method further comprises contacting the biotin (e.g., bisbiotin) with streptavidin to obtain a compound wherein the immobilization complex comprises biotin (e.g., bisbiotin) and streptavidin. In other embodiments, when the immobilization complex comprises streptavidin, the method further comprises contacting the streptavidin with biotin (e.g., bisbiotin) to obtain a compound wherein the immobilization complex comprises streptavidin and biotin (e.g., bisbiotin). In some embodiments, the immobilization complex is a streptavidin-bearing immobilization complex.
[0542]
[0229] In some embodiments, the immobilization complex comprises a linking group. In some embodiments, the linking group comprises a polypeptidyl group. In some embodiments, the polypeptidyl group comprises at least 5 amino acid residues, at least 10 amino acid residues, at least 15 amino acid residues, or at least 20 amino acid residues. In some embodiments, the polypeptidyl group comprises between 5 and 10 amino acid residues, between 5 and 15 amino acid residues, between 5 and 20 amino acid residues, between 10 and 15 amino acid residues, between 10 and 20 amino acid residues, or between 15 and 20 amino acid residues. In some embodiments, the polypeptidyl group comprises between 5 and 15 amino acid residues.
[0543]
[0230] In some embodiments, the polypeptidyl group has a length of at least about 20 A, 25 A, 30 A, 35 A, 40 A, 45 A, 50 A, 55 A, 60 A, 65 A, 70 A, or 75 A. In some embodiments, the polypeptidyl group has a length in a range from 20 A to 30 A, 20 A to 35 A, 20 A to 40 A, 20 A to 45 A, 20 A to 50 A, 20 A to R0708.70180WO00 / R0708.70180US01 55 / 122
[0544] #14644680vl 55 A, 20 A to 60 A, 20 A to 65 A, 20 A to 70 A, 20 A to 75 A, 30 A to 40 A, 30 A to 45 A, 30 A to 50 A, 30 A to 55 A, 30 A to 60 A, 30 A to 65 A, 30 A to 70 A, 30 A to 75 A, 40 A to 50 A, 40 A to 55 A, 40 A to 60 A, 40 A to 65 A, 40 A to 70 A, 40 A to 75 A, 50 A to 60 A, 50 A to 65 A, 50 A to 70 A, 50 A to 75 A, 60 A to 70 A, or 60 A to 75 A.
[0545]
[0231] In some embodiments, the polypeptidyl group comprises at least 1 negatively charged moiety at physiological pH. In some embodiments, the polypeptidyl group comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 negatively charged moieties at physiological pH. In some embodiments, the polypeptidyl group comprises between 1 and 2, 1 and 3, 1 and 4, 1 and 5, 1 and 6, 1 and 7, 1 and 8, 1 and 9, 1 and 10, 1 and 11, 1 and 12, 1 and 13, 1 and 14, 1 and 15, 2 and 3, 2 and 4, 2 and 5, 2 and 6, 2 and 7, 2 and 8, 2 and 9, 2 and 10, 2 and 11, 2 and 12, 2 and 13, 2 and 14, 2 and 15, 3 and 4, 3 and 5, 3 and 6, 3 and 7, 3 and 8, 3 and 9, 3 and 10, 3 and 11, 3 and 12, 3 and 13, 3 and 14, 3 and 15, 4 and 5, 4 and 6, 4 and 7, 4 and 8, 4 and 9, 4 and 10, 4 and 11, 4 and 12, 4 and 13, 4 and 14, 4 and 15, 5 and 6, 5 and 7, 5 and 8, 5 and 9, 5 and 10, 5 and 11, 5 and 12, 5 and 13, 5 and 14, 5 and 15, 6 and 10, 6 and 15, 7 and 10, 7 and 15, 8 and 10, 8 and 15, 9 and 10, 9 and 15, or 10 and 15 negatively charged moieties at physiological pH. In some embodiments, the polypeptidyl group comprises between 1 and 10 negatively charged moieties at physiological pH.
[0546]
[0232] In some embodiments, the polypeptidyl group comprises at least 1 aspartate residue. In some embodiments, the polypeptidyl group comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 aspartate residues. In some embodiments, the polypeptidyl group comprises between 1 and 2, 1 and 3, 1 and 4, 1 and 5, 1 and 6, 1 and 7, 1 and 8, 1 and 9, 1 and 10, 1 and 11, 1 and 12, 1 and 13, 1 and 14, 1 and 15, 2 and 3, 2 and 4, 2 and 5, 2 and 6, 2 and 7, 2 and 8, 2 and 9, 2 and 10, 2 and 11, 2 and 12, 2 and 13, 2 and 14, 2 and 15, 3 and 4, 3 and 5, 3 and 6, 3 and 7, 3 and 8, 3 and 9, 3 and 10, 3 and 11, 3 and 12, 3 and 13, 3 and 14, 3 and 15, 4 and 5, 4 and 6, 4 and 7, 4 and 8, 4 and 9, 4 and 10, 4 and 11, 4 and 12, 4 and 13, 4 and 14, 4 and 15, 5 and 6, 5 and 7, 5 and 8, 5 and 9, 5 and 10, 5 and 11, 5 and 12, 5 and 13, 5 and 14, 5 and 15, 6 and 10, 6 and 15, 7 and 10, 7 and 15, 8 and 10, 8 and 15, 9 and 10, 9 and 15, or 10 and 15 aspartate residues. In some embodiments, the polypeptidyl group comprises between 1 and 10 aspartate residues.
[0547]
[0233] In some embodiments, the polypeptidyl group comprises at least 1 phenylalanine residue. In some embodiments, the polypeptidyl group comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 phenylalanine residues. In some embodiments, the polypeptidyl group comprises between 1 and 2, 1 and 3, 1 and 4, 1 and 5, 1 and 6, 1 and 7, 1 and 8, 1 and 9, 1 and 10, 1 and 11, 1 and 12, 1 and 13, 1 and 14, 1 and 15, 2 and 3, 2 and 4, 2 and 5, 2 and 6, 2 and 7, 2 and 8, 2 and 9, 2 and 10, 2 and 11, 2 and 12, 2 and 13, 2 and 14, 2 and 15, 3 and 4, 3 and 5, 3 and 6, 3 and 7, 3 and 8, 3 and 9, 3 and 10, 3 and 11, 3 and 12, 3 and 13, 3 and 14, 3 and 15, 4 and 5, 4 and 6, 4 and 7, 4 and 8, 4 and 9, 4 and 10, 4 and 11, 4 and 12, 4 and 13, 4 and 14, 4 and 15, 5 and 6, 5 and 7, 5 and 8, 5 and 9, 5 and 10, 5 and 11, 5 and 12, 5 and 13, 5 R0708.70180WO00 / R0708.70180US01 56 / 122
[0548] #14644680vl and 14, 5 and 15, 6 and 10, 6 and 15, 7 and 10, 7 and 15, 8 and 10, 8 and 15, 9 and 10, 9 and 15, or 10 and 15 phenylalanine residues.
[0549]
[0234] In some embodiments, the polypeptidyl group comprises at least 1 glycine residue. In some embodiments, the polypeptidyl group comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 glycine residues. In some embodiments, the polypeptidyl group comprises between 1 and 2, 1 and 3, 1 and 4, 1 and 5, 1 and 6, 1 and 7, 1 and 8, 1 and 9, 1 and 10, 1 and 11, 1 and 12, 1 and 13, 1 and 14, 1 and 15, 2 and 3, 2 and 4, 2 and 5, 2 and 6, 2 and 7, 2 and 8, 2 and 9, 2 and 10, 2 and 11, 2 and 12, 2 and 13, 2 and 14, 2 and 15, 3 and 4, 3 and 5, 3 and 6, 3 and 7, 3 and 8, 3 and 9, 3 and 10, 3 and 11, 3 and 12, 3 and 13, 3 and 14, 3 and 15, 4 and 5, 4 and 6, 4 and 7, 4 and 8, 4 and 9, 4 and 10, 4 and 11, 4 and 12, 4 and 13, 4 and 14, 4 and 15, 5 and 6, 5 and 7, 5 and 8, 5 and 9, 5 and 10, 5 and 11, 5 and 12, 5 and 13, 5 and 14, 5 and 15, 6 and 10, 6 and 15, 7 and 10, 7 and 15, 8 and 10, 8 and 15, 9 and 10, 9 and 15, or 10 and 15 glycine residues.
[0550]
[0235] In some embodiments, the polypeptidyl group comprises at least 1 proline residue. In some embodiments, the polypeptidyl group comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 proline residues. In some embodiments, the polypeptidyl group comprises between 1 and 2, 1 and 3, 1 and 4, 1 and 5, 1 and 6, 1 and 7, 1 and 8, 1 and 9, 1 and 10, 1 and 11, 1 and 12, 1 and 13, 1 and 14, 1 and 15, 2 and 3, 2 and 4, 2 and 5, 2 and 6, 2 and 7, 2 and 8, 2 and 9, 2 and 10, 2 and 11, 2 and 12, 2 and 13, 2 and 14, 2 and 15, 3 and 4, 3 and 5, 3 and 6, 3 and 7, 3 and 8, 3 and 9, 3 and 10, 3 and 11, 3 and 12, 3 and 13, 3 and 14, 3 and 15, 4 and 5, 4 and 6, 4 and 7, 4 and 8, 4 and 9, 4 and 10, 4 and 11, 4 and 12, 4 and 13, 4 and 14, 4 and 15, 5 and 6, 5 and 7, 5 and 8, 5 and 9, 5 and 10, 5 and 11, 5 and 12, 5 and 13, 5 and 14, 5 and 15, 6 and 10, 6 and 15, 7 and 10, 7 and 15, 8 and 10, 8 and 15, 9 and 10, 9 and 15, or 10 and 15 proline residues.
[0551]
[0236] In some embodiments, the polypeptidyl group comprises at least 1 DD repeat, GG repeat, FF repeat, DDD repeat, GGG, and / or FFF repeat. In some embodiments, the polypeptidyl group comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 DD repeats, GG repeats, FF repeats, DDD repeats, GGG, and / or FFF repeats. In some embodiments, the polypeptidyl group comprises between 1 and 2, 1 and 3, 1 and 4, 1 and 5, 1 and 6, 1 and 7, 1 and 8, 1 and 9, 1 and 10, 1 and 11, 1 and 12, 1 and 13, 1 and 14, 1 and 15, 2 and 3, 2 and 4, 2 and 5, 2 and 6, 2 and 7, 2 and 8, 2 and 9, 2 and 10, 2 and 11, 2 and 12, 2 and 13, 2 and 14, 2 and 15, 3 and 4, 3 and 5, 3 and 6, 3 and 7, 3 and 8, 3 and 9, 3 and 10, 3 and 11, 3 and 12, 3 and 13, 3 and 14, 3 and 15, 4 and 5, 4 and 6, 4 and 7, 4 and 8, 4 and 9, 4 and 10, 4 and 11, 4 and 12, 4 and 13, 4 and 14, 4 and 15, 5 and 6, 5 and 7, 5 and 8, 5 and 9, 5 and 10, 5 and 11, 5 and 12, 5 and 13, 5 and 14, 5 and 15, 6 and 10, 6 and 15, 7 and 10, 7 and 15, 8 and 10, 8 and 15, 9 and 10, 9 and 15, or 10 and 15 DD repeats, GG repeats, FF repeats, DDD repeats, GGG, and / or FFF repeats.
[0552]
[0237] In some embodiments, the polypeptidyl group comprises a sequence selected from the group consisting of GPPPPPPPPG (SEQ ID NO: 34), isoEGWRW (SEQ ID NO: 35), DDGGGDDDFF (SEQ R0708.70180WO00 / R0708.70180US01 57 / 122
[0553] #14644680vl ID NO: 36), GGSSSGSGNDEEFQ (SEQ ID NO: 37), GGGGGDPDPDFF (SEQ ID NO: 38), GDGDGDGDGDFF (SEQ ID NO: 39), NNGGGNNNFF (SEQ ID NO: 40), and DDGGGCyCyCyFF (SEQ ID NO: 41), or a salt thereof, wherein Cy is a cysteic acid. In some embodiments, the polypeptidyl group comprises DDGGGDDDFF (SEQ ID NO: 36). In some embodiments, the oligonucleotide has a length of at least 25 nucleotides, and the polypeptidyl group comprises DDGGGDDDFF (SEQ ID NO: 36).
[0554]
[0238] In some embodiments, the linking group comprises an oligonucleotide. In some embodiments, the oligonucleotide is a single-stranded oligonucleotide. In some embodiments, the oligonucleotide is a double -stranded oligonucleotide. In some embodiments, the oligonucleotide has a length of at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 nucleotides. In some embodiments, the oligonucleotide has a length in a range from 15 to 20, 15 to 25, 15 to 30, 15 to 35, 15 to 40, 15 to 45, 15 to 50, 20 to 25, 20 to 30, 20 to 35, 20 to 40, 20 to 45, 20 to 50, 25 to 30, 25 to 35, 25 to 40, 25 to 45, 25 to 50, 30 to 35, 30 to 40, 30 to 45, 30 to 50, 35 to 40, 35 to 45, 35 to 50, 40 to 45, 40 to 50, or 45 to 50 nucleotides. In some embodiments, the oligonucleotide has a length of at least 25 nucleotides.
[0555]
[0239] In certain embodiments, at least one strand of the oligonucleotide has a sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to 5'-CCACGCGTGGAACCCTTGGGATCCA-3' (SEQ ID NO: 32). In some embodiments, at least one strand of the oligonucleotide has a sequence that is at least 80% identical to 5'-CCACGCGTGGAACCCTTGGGATCCA-3' (SEQ ID NO: 32). In certain embodiments, at least one strand of the oligonucleotide has a sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to 5'-TGG AGT CAA GGT CCT CTG ATG CCA T-3’ (SEQ ID NO: 33).
[0556]
[0240] In some embodiments, the linking group further comprises at least one of optionally substituted alkylene, optionally substituted alkenylene, optionally substituted alkynylene, optionally substituted heteroalkylene, optionally substituted heteroalkenylene, optionally substituted heteroalkynylene, optionally substituted heterocyclylene, optionally substituted carbocyclylene, optionally substituted arylene, optionally substituted heteroarylene, or a combination thereof.
[0557]
[0241] In some embodiments, the immobilization complex comprises a binding group. In some embodiments, the binding group comprises a biotin moiety. In some embodiments, the biotin moiety is a bis-biotin moiety.
[0558]
[0242] In some embodiments, the binding group comprises at least one tag sequence. In some embodiments, the at least one tag sequence comprises at least one biotin ligase recognition sequence that permits biotinylation of the immobilization complex (e.g., incorporation of one or more biotin moieties, including biotin and bis-biotin moieties). In some embodiments, the at least one tag sequence comprises two biotin ligase recognition sequences oriented in tandem. In some cases, a biotin ligase recognition sequence refers to an amino acid sequence that is recognized by a biotin ligase, which catalyzes a R0708.70180WO00 / R0708.70180US01 58 / 122
[0559] #14644680vl covalent linkage between the sequence and a biotin molecule. Each biotin ligase recognition sequence of a tag sequence can be covalently linked to a biotin moiety, such that a tag sequence having multiple biotin ligase recognition sequences can be covalently linked to multiple biotin molecules. A region of a tag sequence having one or more biotin ligase recognition sequences can be generally referred to as a biotinylation tag or a biotinylation sequence. In some embodiments, a bis-biotin or bis-biotin moiety can refer to two biotins bound to two biotin ligase recognition sequences oriented in tandem. In some embodiments, the binding group comprises at least one biotin ligase recognition sequence having a biotin moiety attached thereto or at least two biotin ligase recognition sequences, each having a biotin moiety attached thereto.
[0560]
[0243] In some embodiments, the binding group comprises or is conjugated to an avidin protein. In some embodiments, the biotin moiety comprises an avidin protein. In some embodiments, the biotin moiety is conjugated to an avidin protein. The term “avidin protein” refers to a biotin-binding protein, generally having a biotin binding site at each of four subunits of the avidin protein. Non-limiting examples of avidin proteins include avidin, streptavidin, traptavidin, tamavidin, bradavidin, xenavidin, and homologs and variants thereof. In some cases, the avidin protein may have a monomeric, dimeric, or tetrameric form. In some embodiments, the avidin protein is streptavidin in a tetrameric form (e.g., a homotetramer). In some embodiments, the streptavidin in a tetrameric form may be bound to one component (e.g., a first component comprising a first mono-biotin moiety or a first bis-biotin moiety), two components (e.g., a first component comprising a first mono-biotin moiety or a first bis-biotin moiety and a second component comprising a second mono-biotin moiety or a second bis-biotin moiety), three components (e.g., a first component comprising a first bis-biotin moiety, a second component comprising a first mono-biotin moiety, and a third component comprising a second mono-biotin moiety), or four components (e.g., four components, each comprising a mono-biotin moiety).
[0561]
[0244] In some embodiments, conjugating the one or more derivatized polypeptides to the immobilization complex comprises maintaining the one or more derivatized polypeptides and / or the one or more immobilization complex-conjugated polypeptides at a temperature greater than or equal to 20°C, greater than or equal to 25 °C, greater than or equal to 30°C, greater than or equal to 35 °C, or greater than or equal to 37°C.
[0562] Polypeptide Sample Analysis and Protein Sequencing
[0563]
[0245] Aspects of the instant disclosure also involve methods of protein sequencing and identification, methods of protein sequencing and identification, methods of amino acid identification, and compositions, systems, and devices for performing such methods. In some aspects, methods of determining the sequence of a target protein are described. In some embodiments, the target protein is enriched (e.g., enriched using electrophoretic methods, e.g., affinity SCODA) prior to determining the sequence of the target protein. In some aspects, methods of determining the sequences of a plurality of proteins (e.g., at least 2, 3, 4, 5, 10, 15, 20, 30, 50, or more) present in a sample (e.g., a purified sample, a cell lysate, a single-cell, a population of cells, or a tissue) are described. In some embodiments, a sample is prepared as described herein (e.g., digested, lysed, purified, fragmented, and / or enriched for a target R0708.70180WO00 / R0708.70180US01 59 / 122
[0564] #14644680vl protein) prior to determining the sequence of a target protein or a plurality of proteins present in a sample.
[0565]
[0246] For example, FIG. 1 shows an example of a dynamic peptide sequencing reaction in which individual on-off binding events give rise to signal pulses of a signal output. As shown at left, a protein sample may be fragmented into peptides, which are immobilized in reaction chambers and exposed to a mixture of amino acid recognizers and cleaving agents. As shown at right, amino acid recognizers reversibly bind to the peptide, producing a series of changes in signal output (e.g., signal pulses) as amino acids are progressively cleaved from the peptide terminus. The temporal order of recognition and the kinetics of binding and / or cleaving can be used to determine structural information for the peptide.
[0566]
[0247] Compositions, systems, and methods for performing dynamic polypeptide sequencing and analyzing data obtained therefrom are described in PCT International Publication No.
[0567] W02020102741A1, filed November 15, 2019, PCT International Publication No. WO2021236983A2, filed May 20, 2021, PCT International Publication No. WO2022 / 159495A1, filed January 19, 2022, and PCT International Publication No. WO2023 / 122769A2, filed December 22, 2022, each of which is incorporated by reference in its entirety.
[0568]
[0248] The methods may facilitate obtaining information regarding multiple amino acids. For example, a polypeptide comprising a chain of amino acids may be used with the techniques described herein. The chain of amino acids may comprise at least one amino acid to which a dye-labeled recognizer binds. In some embodiments, the chain of amino acids comprises a terminal amino acid and one or more downstream amino acids (e.g., amino acids at position 1, 2, 3, 4, and / or 5 relative to the polypeptide terminus). In some embodiments, one or more amino acid recognizers may bind to the terminal amino acid. In some embodiments, the one or more amino acid recognizers may bind to one or more amino acids downstream of the terminal amino acid in addition to the terminal amino acid of the peptide. In some embodiments, the one or more amino acid recognizers may bind to an internal amino acid and one or more amino acids upstream or downstream of the internal amino acid.
[0569]
[0249] The polypeptide may comprise any number of amino acids. In some embodiments, the polypeptide comprises at least 5 amino acids, at least 10 amino acids, at least 15 amino acids, at least 20 amino acids, at least 50 amino acids, or at least 100 amino acids. In some embodiments, the polypeptide comprises 5-10, 5-15, 5-20, 5-50, 5-100, 10-15, 10-20, 10-50, 10-100, 15-20, 15-50, 15-100, 20-50, 20-100, or 50-100 amino acids.
[0570]
[0250] In some embodiments, to obtain information regarding the chain of amino acids, a sample comprising at least a portion (e.g., all or a fragment thereof) of the polypeptide may be loaded onto an integrated device. In particular, the polypeptide may be loaded into a reaction chamber of the integrated device. In some cases, the polypeptide may be bound to the surface of the chamber via a covalent or non-covalent bond (e.g., a streptavidin-biotin bond, a click chemistry bond) which immobilizes the polypeptide in the chamber, forming one or more immobilization complex-conjugated polypeptides.
[0571]
[0251] In some embodiments, multiple polypeptides may be loaded onto the integrated device and multiple chambers of the integrated device may receive one or more of the polypeptides. The techniques R0708.70180WO00 / R0708.70180US01 60 / 122
[0572] #14644680vl described herein for obtaining information regarding polypeptides may be performed in a parallel manner (e.g., concurrently, simultaneously).
[0573]
[0252] In some embodiments, the method further comprises contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents and one or more amino acid recognizers. In some embodiments, the method further comprises contacting the polypeptide sample with a reaction mixture comprising one or more aminopeptidases and one or more amino acid recognizers.
[0574]
[0253] In some embodiments, the method further comprises monitoring a signal for signal pulses corresponding to interactions between one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides of the polypeptide sample.
[0575]
[0254] In some embodiments, the method further comprises determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal.
[0576]
[0255] In some embodiments, the method further comprises:
[0577] contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents and one or more amino acid recognizers;
[0578] monitoring a signal for signal pulses corresponding to interactions between one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides of the polypeptide sample; and
[0579] determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal.
[0580]
[0256] In some embodiments, the method further comprises outputting an amino acid sequence representative of the polypeptide.
[0581] Reaction Mixture
[0582]
[0257] In some embodiments, a polypeptide sequencing reaction in accordance with the disclosure is performed under conditions in which recognition and cleavage of amino acids can occur simultaneously in a single reaction mixture. For example, in some embodiments, a polypeptide sequencing reaction is performed in a reaction mixture having a pH at which association events and cleavage events can occur. Accordingly, in some embodiments, a reaction mixture has a pH of between about 6.5 and about 9.0. In some embodiments, a reaction mixture has a pH of between about 7.0 and about 8.5 (e.g., between about 7.0 and about 8.0, between about 7.5 and about 8.5, between about 7.5 and about 8.0, or between about 8.0 and about 8.5).
[0583]
[0258] In some embodiments, a polypeptide sequencing reaction is performed in a reaction mixture comprising one or more buffering agents. In some embodiments, a reaction mixture comprises a buffering agent in a concentration of at least 10 mM (e.g., at least 20 mM and up to 250 mM, at least 50 mM, 10-250 mM, 10-100 mM, 20-100 mM, 50-100 mM, or 100-200 mM). In some embodiments, a reaction mixture comprises a buffering agent in a concentration of between about 10 mM and about 50 mM (e.g., between about 10 mM and about 25 mM, between about 25 mM and about 50 mM, or between about 20 mM and about 40 mM). Examples of buffering agents include, without limitation, HEPES (4-R0708.70180WO00 / R0708.70180US01 61 / 122
[0584] #14644680vl (2 -hydroxyethyl)- 1 -piperazineethanesulfonic acid), Tris (tris(hydroxymethyl)aminomethane), and MOPS (3 -(N -morpholino)propane sulfonic acid).
[0585]
[0259] In some embodiments, a polypeptide sequencing reaction is performed in a reaction mixture comprising salt in a concentration of at least 10 mM. In some embodiments, a reaction mixture comprises salt in a concentration of at least 10 mM (e.g., at least 20 mM, at least 50 mM, at least 100 mM, or more). In some embodiments, a reaction mixture comprises salt in a concentration of between about 10 mM and about 250 mM (e.g., between about 20 mM and about 200 mM, between about 50 mM and about 150 mM, between about 10 mM and about 50 mM, or between about 10 mM and about 100 mM). Examples of salts include, without limitation, sodium salts, potassium salts, and acetates, such as sodium chloride (NaCl), sodium acetate (NaOAc), and potassium acetate (KO Ac).
[0586]
[0260] Additional examples of components for use in a reaction mixture include divalent cations (e.g., Mg2+, Co2+) and surfactants (e.g., polysorbate 20). In some embodiments, a reaction mixture comprises a divalent cation in a concentration of between about 0.1 mM and about 50 mM (e.g., between about 10 mM and about 50 mM, between about 0.1 mM and about 10 mM, or between about 1 mM and about 20 mM). In some embodiments, a reaction mixture comprises a surfactant in a concentration of at least 0.01% (e.g., between about 0.01% and about 0.10%). In some embodiments, a reaction mixture comprises one or more components useful in single-molecule analysis, such as an oxygen-scavenging system (e.g., a PCA / PCD system or a Pyranose oxidase / Catalase / glucose system) and / or one or more triplet state quenchers (e.g., trolox, COT, and NBA).
[0587]
[0261] In some embodiments, a polypeptide sequencing reaction is performed at a temperature at which association events and cleavage events can occur. In some embodiments, a polypeptide sequencing reaction is performed at a temperature of at least 10 °C. In some embodiments, a polypeptide sequencing reaction is performed at a temperature of between about 10 °C and about 50 °C (e.g., 15-45 °C, 20-40 °C, at or around 25 °C, at or around 30 °C, at or around 35 °C, at or around 37 °C). In some embodiments, a polypeptide sequencing reaction is performed at or around room temperature.
[0588]
[0262] As detailed above, a real-time sequencing process as illustrated by FIG. 1 can generally involve cycles of amino acid recognition and terminal amino acid cleavage. In some embodiments, the relative occurrence of recognition and cleavage can be controlled by a concentration differential between one or more amino acid recognizers and at least one cleaving agent. In some embodiments, the concentration differential can be optimized such that the number of signal pulses detected during recognition of an individual amino acid provides a desired confidence interval for identification. For example, if an initial sequencing reaction provides signal data with too few signal pulses between cleavage events to permit determination of characteristic patterns with a desired confidence interval, the sequencing reaction can be repeated using a decreased concentration of non-specific exopeptidase relative to recognition molecule.
[0589]
[0263] In some embodiments, polypeptide analysis in accordance with the disclosure may be carried out by contacting a polypeptide with a reaction mixture comprising one or more amino acid recognizers and one or more cleaving agents (e.g., aminopeptidases). In some embodiments, a reaction mixture comprises an amino acid recognizer at a concentration of between about 10 nM and about 10 pM. In some R0708.70180WO00 / R0708.70180US01 62 / 122
[0590] #14644680vl embodiments, a reaction mixture comprises a cleaving agent at a concentration of between about 500 nM and about 500 pM.
[0591]
[0264] In some embodiments, a reaction mixture comprises an amino acid recognizer at a concentration of between about 100 nM and about 10 pM, between about 250 nM and about 10 pM, between about 100 nM and about 1 pM, between about 250 nM and about 1 pM, between about 250 nM and about 750 nM, or between about 500 nM and about 1 pM. In some embodiments, a reaction mixture comprises an amino acid recognizer at a concentration of about 100 nM, about 250 nM, about 500 nM, about 750 nM, or about 1 pM. In some embodiments, a reaction mixture comprises a cleaving agent at a concentration of between about 500 nM and about 250 pM, between about 500 nM and about 100 pM, between about 1 pM and about 100 pM, between about 500 nM and about 50 pM, between about 1 pM and about 100 pM, between about 10 pM and about 200 pM, or between about 10 pM and about 100 pM. In some embodiments, a reaction mixture comprises a cleaving agent at a concentration of about 1 pM, about 5 pM, about 10 pM, about 30 pM, about 50 pM, about 70 pM, or about 100 pM.
[0592]
[0265] In some embodiments, a reaction mixture comprises an amino acid recognizer at a concentration of between about 10 nM and about 10 pM, and a cleaving agent at a concentration of between about 500 nM and about 500 pM. In some embodiments, a reaction mixture comprises an amino acid recognizer at a concentration of between about 100 nM and about 1 pM, and a cleaving agent at a concentration of between about 1 pM and about 100 pM. In some embodiments, a reaction mixture comprises an amino acid recognizer at a concentration of between about 250 nM and about 1 pM, and a cleaving agent at a concentration of between about 10 pM and about 100 pM. In some embodiments, a reaction mixture comprises an amino acid recognizer at a concentration of about 500 nM, and a cleaving agent at a concentration of between about 25 pM and about 75 pM. In some embodiments, the concentration of an amino acid recognizer and / or the concentration of a cleaving agent in a reaction mixture is as described elsewhere herein.
[0593]
[0266] In some embodiments, a reaction mixture comprises an amino acid recognizer and a cleaving agent in a molar ratio of about 500: 1, about 400: 1, about 300: 1, about 200: 1, about 100:1, about 75:1, about 50:1, about 25:1, about 10:1, about 5:1, about 2: 1, or about 1: 1. In some embodiments, a reaction mixture comprises an amino acid recognizer and a cleaving agent in a molar ratio of between about 10:1 and about 200: 1. In some embodiments, a reaction mixture comprises an amino acid recognizer and a cleaving agent in a molar ratio of between about 50:1 and about 150: 1. In some embodiments, the molar ratio of an amino acid recognizer to a cleaving agent in a reaction mixture is between about 1: 1,000 and about 1:1 or between about 1:1 and about 100:1 (e.g., 1:1,000, about 1:500, about 1:200, about 1:100, about 1:10, about 1:5, about 1:2, about 1:1, about 5:1, about 10:1, about 50:1, about 100:1). In some embodiments, the molar ratio of an amino acid recognizer to a cleaving agent in a reaction mixture is between about 1:100 and about 1:1 or between about 1:1 and about 10:1. In some embodiments, the molar ratio of an amino acid recognizer to a cleaving agent in a reaction mixture is as described elsewhere herein.
[0594]
[0267] In some embodiments, a reaction mixture comprises one or more amino acid recognizers and one R0708.70180WO00 / R0708.70180US01 63 / 122
[0595] #14644680vl or more cleaving agents described herein. In some embodiments, a reaction mixture comprises at least three amino acid recognizers and at least one cleaving agent. In some embodiments, the reaction mixture comprises two or more cleaving agents. In some embodiments, the reaction mixture comprises at least one and up to ten cleaving agents (e.g., 1-3 cleaving agents, 2-10 cleaving agents, 1-5 cleaving agents, 3-10 cleaving agents). In some embodiments, the reaction mixture comprises at least three and up to thirty amino acid recognizers (e.g., between 3 and 25, between 3 and 20, between 3 and 10, between 3 and 5, between 5 and 30, between 5 and 20, between 5 and 10, or between 10 and 20, amino acid recognizers).
[0596]
[0268] In some embodiments, a reaction mixture comprises more than one amino acid recognizer and / or more than one cleaving agent. In some embodiments, a reaction mixture described as comprising more than one amino acid recognizer or cleaving agent refers to the mixture as having more than one type of amino acid recognizer or cleaving agent. For example, in some embodiments, a reaction mixture comprises two or more cleaving agents, where the two or more cleaving agents refer to two or more types of aminopeptidases. In some embodiments, one type of aminopeptidase has an amino acid sequence that is different from another type of aminopeptidase in the reaction mixture. In some embodiments, one type of cleaving agent cleaves an amino acid or subset of amino acids that is different from an amino acid or subset of amino acids cleaved by another type of cleaving agent in the reaction mixture.
[0597] Cleaving Agents
[0598]
[0269] In some embodiments, the method further comprises contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents. In some embodiments, the method further comprises contacting the polypeptide sample with a reaction mixture comprising one or more aminopeptidases.
[0599]
[0270] In some embodiments, the cleaving agent comprises an aminopeptidase having an amino acid sequence selected from Table 3. It should be appreciated that the example sequences in Table 3 and other examples described herein are meant to be non-limiting, and aminopeptidases in accordance with the disclosure can include any homologs, variants, or fragments thereof minimally containing domains or subdomains responsible for amino acid cleavage.
[0600]
[0271] In some embodiments, a cleaving agent comprises an aminopeptidase having an amino acid sequence that is at least 80% identical to an amino acid sequence selected from Table 3. In some embodiments, an aminopeptidase has at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 98%, or higher, amino acid sequence identity to an amino acid sequence selected from Table 3. In some embodiments, an aminopeptidase has 25-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-95%, 92-99%, 94-99%, 95-99%, 40-100%, 50-100%, 60-100%, 70-100%, 80-100%, 90-100%, 92-100%, 94-100%, 95-100%, 96-100%, or 100% amino acid sequence identity to an amino acid sequence selected from Table 3.
[0601]
[0272] In some embodiments, a cleaving agent comprises a synthetic or recombinant aminopeptidase. In some embodiments, a cleaving agent comprises a monomeric aminopeptidase. In some embodiments, a cleaving agent comprises a multimeric aminopeptidase (e.g., a multimeric complex of monomeric subunits, which may be the same or different).
[0602] R0708.70180WO00 / R0708.70180US01 64 / 122
[0603] #14644680vl
[0273] In some embodiments, a cleaving agent comprises an aminopeptidase obtained or derived from a particular source (e.g., organism). As described herein, in some embodiments, an aminopeptidase identified as being from a particular organism does not impart a requirement that the aminopeptidase have an amino acid sequence that is 100% identical to a naturally-occurring aminopeptidase from the organism, although it may in some embodiments.
[0604]
[0274] For example, in some embodiments, a cleaving agent comprises an aminopeptidase from Pyrococcus horikoshii (e.g., Pyrococcus horikoshii TET Aminopeptidase II, Pyrococcus horikoshii TET Aminopeptidase III). In some embodiments, an aminopeptidase from Pyrococcus horikoshii is at least 80%, at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to a naturally-occurring aminopeptidase from Pyrococcus horikoshii (e.g., Pyrococcus horikoshii TET Aminopeptidase II, Pyrococcus horikoshii TET Aminopeptidase III).
[0605]
[0275] In some embodiments, a cleaving agent comprises an aminopeptidase from Yersinia pestis (e.g., Yersinia pestis Xaa-Prolyl Aminopeptidase). In some embodiments, an aminopeptidase from Yersinia pestis is at least 80%, at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to a naturally-occurring aminopeptidase from Yersinia pestis (e.g., Yersinia pestis Xaa-Prolyl Aminopeptidase).
[0606]
[0276] In some embodiments, a cleaving agent comprises an aminopeptidase from Pyrococcus furiosus (e.g., Pyrococcus furiosus Aminopeptidase I). In some embodiments, an aminopeptidase from Pyrococcus furiosus is at least 80%, at least 85%, at least 90%, at least 95%, 80-100%, 85-95%, 90-99%, 95-99%, or 100% identical to a naturally-occurring aminopeptidase from Pyrococcus furiosus (e.g., Pyrococcus furiosus Aminopeptidase I).
[0607] Table 3. Non-limiting example sequences of aminopeptidases
[0608] Name Sequence
[0609] Pyrococcus MEVRNMVDYELLKKVVEAPGVSGYEFLGIRDVVIEEIKDYVDEVKVDKLGNVIAHKKGEGPKVMI horikoshii TET II AAHMDQIGLMVTHIEKNGFLRVAPIGGVDPKTLIAQRFKVWIDKGKFIYGVGASVPPHIQKPEDR Aminopeptidase KKAPDWDQIFIDIGAESKEEAEDMGVKIGTVITWDGRLERLGKHRFVSIAFDDRIAVYTILEVAK (hTET II) QLKDAKADVYFVATVQEEVGLRGARTSAFGIEPDYGFAIDVTIAADIPGTPEHKQVTHLGKGTAI KIMDRSVICHPTIVRWLEELAKKHEIPYQLEILLGGGTDAGAIHLTKAGVPTGALSVPARYIHSN TEVVDERDVDATVELMTKALENIHELKI (SEQ ID NO: 1)
[0610] AP30 MEVRNMVDYELLKKVVEAPGVSGYEFLGIRDVVIEEIKDYVDEVKVDKLGNVIAHKKGEGPKVMI AAHMDQIGLMVTHIEKNGFLRVAPIGGVDPKTLIAQRFKVWIDKGKFIYGVGASVPPHIQKPEDR KKAPDWDQIFIDIGAESKEEAEDMGVKIGTVITWDGRLERLGKHRFVSIAFDDRIAVYTILEVAK QLKDAKADVYFVATVQEEVGLRGARTSAFGIEPDYGFAIDVTIAADIPGTPEHKQVTHLGKGTAI KIMDRSVICHPTIVRWLEELAKKHEIPYQLEILLGGGTDAGAIHLTKAGVPTGALSVPARYIHSN TEVVDERDVDATVELMTKALENIHELKIGGSHHHHHHHHHHGGGSGGGSGGGSGLNDFFEAQKIE WHEGGGSGGGSGGGSGLNDFFEAQKIEWHE (SEQ ID NO: 2)
[0611] Pyrococcus MDLKGGESMVDWKLMQEIIEAPGVSGYEHLGIRDIVVDVLKEVADEVKVDKLGNVIAHFKGSSPR horikoshii TET III IMVAAHMDKIGVMVNHIDKDGYLHIVPIGGVLPETLVAQRIRFFTEKGERYGVVGVLPPHLRRGQ Aminopeptidase EDKGSKIDWDQIVVDVGASSKEEAEEMGFRVGTVGEFAPNFTRLNEHRFATPYLDDRICLYAMIE (hTET III) AARQLGDHEADIYIVGSVQEEVGLRGARVASYAINPEVGIAMDVTFAKQPHDKGKIVPELGKGPV MDVGPNINPKLRAFADEVAKKYEIPLQVEPSPRPTGTDANMQINREGVATAVLSIPIRYMHSQVE LADARDVDNTIKLAKALLEELKPMDFTP (SEQ ID NO: 3)
[0612] AP37 MDLKGGESMVDWKLMQEIIEAPGVSGYEHLGIRDIVVDVLKEVADEVKVDKLGNVIAHFKGSSPR
[0613]
[0614] IMVAAHMDKIGVMVNHIDKDGYLHIVPIGGVLPETLVAQRIRFFTEKGERYGVVGVLPPHLRRGQ R0708.70180WO00 / R0708.70180US01 65 / 122
[0615] #14644680vl EDKGSKIDWDQIVVDVGASSKEEAEEMGFRVGTVGEFAPNFTRLNEHRFATPYLDDRICLYAMIE AARQLGDHEADIYIVGSVQEEVGLRGARVASYAINPEVGIAMDVTFAKQPHDKGKIVPELGKGPV MDVGPNINPKLRAFADEVAKKYEIPLQVEPSPRPTGTDANMQINREGVATAVLSIPIRYMHSQVE LADARDVDNTIKLAKALLEELKPMDFTPGHHHHHHHHHH (SEQ ID NO: 4) Yersinia pestis Xaa- MTQQEYQNRRQALLAKMAPGSAAIIFAAPEATRSADSEYPYRQNSDFSYLTGFNEPEAVLILVKS Prolyl DETHNHSVLFNRIRDLTAEIWFGRRLGQEAAPTKLAVDRALPFDEINEQLYLLLNRLDVIYHAQG aminopeptidase QYAYADNIVFAALEKLRHGFRKNLRAPATLTDWRPWLHEMRLFKSAEEIAVLRRAGEISALAHTR (yPIP) AMEKCRPGMFEYQLEGEILHEFTRHGARYPAYNTIVGGGENGCILHYTENECELRDGDLVLIDAG CEYRGYAGDITRTFPVNGKFTPAQRAVYDIVLAAINKSLTLFRPGTSIREVTEEVVRIMVVGLVE LGILKGDIEQLIAEQAHRPFFMHGLSHWLGMDVHDVGDYGSSDRGRILEPGMVLTVEPGLYIAPD ADVPPQYRGIGIRIEDDIVITATGNENLTASVVKDPDDIEALMALNHAGENLYFQLE (SEQ ID NO: 5)
[0616] yPIP-6x His MTQQEYQNRRQALLAKMAPGSAAIIFAAPEATRSADSEYPYRQNSDFSYLTGFNEPEAVLILVKS DETHNHSVLFNRIRDLTAEIWFGRRLGQEAAPTKLAVDRALPFDEINEQLYLLLNRLDVIYHAQG QYAYADNIVFAALEKLRHGFRKNLRAPATLTDWRPWLHEMRLFKSAEEIAVLRRAGEISALAHTR AMEKCRPGMFEYQLEGEILHEFTRHGARYPAYNTIVGGGENGCILHYTENECELRDGDLVLIDAG CEYRGYAGDITRTFPVNGKFTPAQRAVYDIVLAAINKSLTLFRPGTSIREVTEEVVRIMVVGLVE LGILKGDIEQLIAEQAHRPFFMHGLSHWLGMDVHDVGDYGSSDRGRILEPGMVLTVEPGLYIAPD ADVPPQYRGIGIRIEDDIVITATGNENLTASVVKDPDDIEALMALNHAGENLYFQLEHHHHHH
[0617] (SEQ ID NO: 6)
[0618] yPIP (truncated) MTQQEYQNRRQALLAKMAPGSAAIIFAAPEATRSADSEYPYRQNSDFSYLTGFNEPEAVLILVKS DETHNHSVLFNRIRDLTAEIWFGRRLGQEAAPTKLAVDRALPFDEINEQLYLLLNRLDVIYHAQG QYAYADNIVFAALEKLRHGFRKNLRAPATLTDWRPWLHEMRLFKSAEEIAVLRRAGEISALAHTR AMEKCRPGMFEYQLEGEILHEFTRHGARYPAYNTIVGGGENGCILHYTENECELRDGDLVLIDAG CEYRGYAGDITRTFPVNGKFTPAQRAVYDIVLAAINKSLTLFRPGTSIREVTEEVVRIMVVGLVE LGILKGDIEQLIAEQAHRPFFMHGLSHWLGMDVHDVGDYGSSDRGRILEPGMVLTVEPGLYIAPD ADVPPQYRGIGIRIEDDIVITATGNENLTASVVKDPDDIEALMALNHAGENLYFQ (SEQ ID NO: 7)
[0619] AP70 MTQQEYQNRRQALLAKMAPGSAAIIFAAPEATRSADSEYPYRQNSDFSYLTGFNEPEAVLILVKS DETHNHSVLFNRIRDLTAEIWFGRRLGQEAAPTKLAVDRALPFDEINEQLYLLLNRLDVIYHAQG QYAYADNIVFAALEKLRHGFRKNLRAPATLTDWRPWLHEMRLFKSAEEIAVLRRAGEISALAHTR AMEKCRPGMFEYQLEGEILHEFTRHGARYPAYNTIVGGGENGCILHYTENECELRDGDLVLIDAG CEYRGYAGDITRTFPVNGKFTPAQRAVYDIVLAAINKSLTLFRPGTSIREVTEEVVRIMVVGLVE LGILKGDIEQLIAEQAHRPFFMHGLSHWLGMDVHDVGDYGSSDRGRILEPGMVLTVEPGLYIAPD ADVPPQYRGIGIRIEDDIVITATGNENLTASVVKDPDDIEALMALNHAGENLYFQGGSHHHHHH
[0620] (SEQ ID NO: 8)
[0621] L. pneumophila Ml MMVKQGVFMKTDQSKVKKLSDYKSLDYFVIHVDLQIDLSKKPVESKARLTVVPNLNVDSHSNDLV Aminopeptidase LDGENMTLVSLQMNDNLLKENEYELTKDSLIIKNIPQNTPFTIEMTSLLGENTDLFGLYETEGVA (Glu / Asp Specific) LVKAESEGLRRVFYLPDRPDNLATYKTTIIANQEDYPVLLSNGVLIEKKELPLGLHSVTWLDDVP KPSYLFALVAGNLQRSVTYYQTKSGRELPIEFYVPPSATSKCDFAKEVLKEAMAWDERTFNLECA LRQHMVAGVDKYASGASEPTGLNLFNTENLFASPETKTDLGILRVLEWAHEFFHYWSGDRVTIR DWFNLPLKEGLTTFRAAMFREELFGTDLIRLLDGKNLDERAPRQSAYTAVRSLYTAAAYEKSADI FRMMMLFIGKEPFIEAVAKFFKDNDGGAVTLEDFIESISNSSGKDLRSFLSWFTESGIPELIVTD ELNPDTKQYFLKIKTVNGRNRPIPILMGLLDSSGAEIVADKLLIVDQEEIEFQFENIQTRPIPSL LRSFSAPVHMKYEYSYQDLLLLMQFDTNLYNRCEAAKQLISALINDFCIGKKIELSPQFFAVYKA LLSDNSLNEWMLAELITLPSLEELIENQDKPDFEKLNEGRQLIQNALANELKTDFYNLLFRIQIS GDDDKQKLKGFDLKQAGLRRLKSVCFSYLLNVDFEKTKEKLILQFEDALGKNMTETALALSMLCE INCEEADVALEDYYHYWKNDPGAVNNWFSIQALAHSPDVIERVKKLMRHGDFDLSNPNKVYALLG SFIKNPFGFHSVTGEGYQLVADAIFDLDKINPTLAANLTEKFTYWDKYDVNRQAMMISTLKIIYS NATSSDVRTMAKKGLDKVKEDLPLPIHLTFHGGSTMQDRTAQLIADGNKENAYQLH (SEQ ID NO: 9)
[0622] E. coli methionine MGTAISIKTPEDIEKMRVAGRLAAEVLEMIEPYVKPGVSTGELDRICNDYIVNEQHAVSACLGYH aminopeptidase GYPKSVCISINEVVCHGIPDDAKLLKDGDIVNIDVTVIKDGFHGDTSKMFIVGKPTIMGERLCRI (Met specific) TQESLYLALRMVKPGINLREIGAAIQKFVEAEGFSVVREYCGHGIGRGFHEEPQVLHYDSRETNV
[0623] VLKPGMTFTIEPMVNAGKKEIRTMKDGWTVKTKDRSLSAQYEHTIVVTDNGCEILTLRKDDTIPA
[0624]
[0625] IISHD (SEQ ID NO: 10)
[0626] R0708.70180WO00 / R0708.70180US01 66 / 122
[0627] #14644680vl M. smegmatis MGTLEANTNGPGSMLSRMPVSSRTVPFGDHETWVQVTTPENAQPHALPLIVLHGGPGMAHNYVAN Proline IAALADETGRTVIHYDQVGCGNSTHLPDAPADFWTPQLFVDEFHAVCTALGIERYHVLGQSWGGM iminopeptidase LGAEIAVRQPSGLVSLAICNSPASMRLWSEAAGDLRAQLPAETRAALDRHEAAGTITHPDYLQAA (Pro specific) AEFYRRHVCRWPTPQDFADSVAQMEAEPTVYHTMNGPNEFHVVGTLGDWSVIDRLPDVTAPVLV IAGEHDEATPKTWQPFVDHIPDVRSHVFPGTSHCTHLEKPEEFRAVVAQFLHQHDLAADARV
[0628] (SEQ ID NO: 11)
[0629] P. furiosus MDTEKLMKAGEIAKKVREKAIKLARPGMLLLELAESIEKMIMELGGKPAFPVNLSINEIAAHYTP methionine YKGDTTVLKEGDYLKIDVGVHIDGFIADTAVTVRVGMEEDELMEAAKEALNAAISVARAGVEIKE aminopeptidase LGKAIENEIRKRGFKPIVNLSGHKIERYKLHAGISIPNIYRPHDNYVLKEGDVFAIEPFATIGAG QVIEVPPTLIYMYVRDVPVRVAQARFLLAKIKREYGTLPFAYRWLQNDMPEGQLKLALKTLEKAG AIYGYPVLKEIRNGIVAQFEHTIIVEKDSVIVTQDMINKSTLE (SEQ ID NO: 12) Aeromonas sobria HMSSPLHYVLDGIHCEPHFFTVPLDHQQPDDEETITLFGRTLCRKDRLDDELPWLLYLQGGPGFG Proline APRPSANGGWIKRALQEFRVLLLDQRGTGHSTPIHAELLAHLNPRQQADYLSHFRADSIVRDAEL aminopeptidase IREQLSPDHPWSLLGQSFGGFCSLTYLSLFPDSLHEVYLTGGVAPIGRSADEVYRATYQRVADKN RAFFARFPHAQAIANRLATHLQRHDVRLPNGQRLTVEQLQQQGLDLGASGAFEELYYLLEDAFIG EKLNPAFLYQVQAMQPFNTNPVFAILHELIYCEGAASHWAAERVRGEFPALAWAQGKDFAFTGEM IFPWMFEQFRELIPLKEAAHLLAEKADWGPLYDPVQLARNKVPVACAVYAEDMYVEFDYSRETLK GLSNSRAWITNEYEHNGLRVDGEQILDRLIRLNRDCLE (SEQ ID NO: 13) Pyrococcus furiosus MKERLEKLVKFMDENSIDRVFIAKPVNVYYFSGTSPLGGGYIIVDGDEATLYVPELEYEMAKEES Proline KLPVVKFKKFDEIYEILKNTETLGIEGTLSYSMVENFKEKSNVKEFKKIDDVIKDLRIIKTKEEI Aminopeptidase (X- EIIEKACEIADKAVMAAIEEITEGKREREVAAKVEYLMKMNGAEKPAFDTIIASGHRSALPHGVA / -Pro) SDKRIERGDLWIDLGALYNHYNSDITRTIVVGSPNEKQREIYEIVLEAQKRAVEAAKPGMTAKE LDSIAREIIKEYGYGDYFIHSLGHGVGLEIHEWPRISQYDETVLKEGMVITIEPGIYIPKLGGVR IEDTVLITENGAKRLTKTERELL (SEQ ID NO: 14)
[0630] Elizabethkingia MIPITTPVGNFKVWTKRFGTNPKIKVLLLHGGPAMTHEYMECFETFFQREGFEFYEYDQLGSYYS meningoseptica DQPTDEKLWNIDRFVDEVEQVRKAIHADKENFYVLGNSWGGILAMEYALKYQQNLKGLIVANMMA Proline SAPEYVKYAEVLSKQMKPEVLAEVRAIEAKKDYANPRYTELLFPNYYAQHICRLKEWPDALNRSL aminopeptidase KHVNSTVYTLMQGPSELGMSSDARLAKWDIKNRLHEIATPTLMIGARYDTMDPKAMEEQSKLVQK GRYLYCPNGSHLAMWDDQKVFMDGVIKFIKDVDTKSFN (SEQ ID NO: 15)
[0631] N. gonorrhoeae MYEIKQPFHSGYLQVSEIHQIYWEESGNPDGVPVIFLHGGPGAGASPECRGFFNPDVFRIVIIDQ Proline RGCGRSHPYACAEDNTTWDLVADIEKVREMLGIGKWLVFGGSWGSTLSLAYAQTHPERVKGLVLR Iminopeptidase GIFLCRPSETAWLNEAGGVSRIYPEQWQKFVAPIAENRRNRLIEAYHGLLFHQDEEVCLSAAKAW ADWESYLIRFEPEGVDEDAYASLAIARLENHYFVNGGWLQGDKAILNNIGKIRHIPTVIVQGRYD LCTPMQSAWELSKAFPEAELRWQAGHCAFDPPLADALVQAVEDILPRLL (SEQ ID NO: 16)
[0632] E. coli MTQQPQAKYRHDYRAPDYQITDIDLTFDLDAQKTVVTAVSQAVRHGASDAPLRLNGEDLKLVSVH Aminopeptidase N INDEPWTAWKEEEGALVISNLPERFTLKIINEISPAANTALEGLYQSGDALCTQCEAEGFRHITY (Zinc YLDRPDVLARFTTKIIADKIKYPFLLSNGNRVAQGELENGRHWVQWQDPFPKPCYLFALVAGDFD Metalloprotease) VLRDTFTTRSGREVALELYVDRGNLDRAPWAMTSLKNSMKWDEERFGLEYDLDIYMIVAVDFFNM GAMENKGLNIFNSKYVLARTDTATDKDYLDIERVIGHEYFHNWTGNRVTCRDWFQLSLKEGLTVF RDQEFSSDLGSRAVNRINNVRTMRGLQFAEDASPMAHPIRPDMVIEMNNFYTLTVYEKGAEVIRM IHTLLGEENFQKGMQLYFERHDGSAATCDDFVQAMEDASNVDLSHFRRWYSQSGTPIVTVKDDYN PETEQYTLTISQRTPATPDQAEKQPLHIPFAIELYDNEGKVIPLQKGGHPVNSVLNVTQAEQTFV FDNVYFQPVPALLCEFSAPVKLEYKWSDQQLTFLMRHARNDFSRWDAAQSLLATYIKLNVARHQQ GQPLSLPVHVADAFRAVLLDEKIDPALAAEILTLPSVNEMAELFDIIDPIAIAEVREALTRTLAT ELADELLAIYNANYQSEYRVEHEDIAKRTLRNACLRFLAFGETHLADVLVSKQFHEANNMTDALA ALSAAVAAQLPCRDALMQEYDDKWHQNGLVMDKWFILQATSPAANVLETVRGLLQHRSFTMSNPN RIRSLIGAFAGSNPAAFHAEDGSGYLFLVEMLTDLNSRNPQVASRLIEPLIRLKRYDAKRQEKMR AALEQLKGLENLSGDLYEKITKALA (SEQ ID NO: 17)
[0633] P. falciparum Ml PKIHYRKDYKPSGFIINQVTLNINIHDQETIVRSVLDMDISKHNVGEDLVFDGVGLKINEISINN aminopeptidase KKLVEGEEYTYDNEFLTIFSKFVPKSKFAFSSEVIIHPETNYALTGLYKSKNIIVSQCEATGFRR ITFFIDRPDMMAKYDVTVTADKEKYPVLLSNGDKVNEFEIPGGRHGARFNDPPLKPCYLFAVVAG DLKHLSATYITKYTKKKVELYVFSEEKYVSKLQWALECLKKSMAFDEDYFGLEYDLSRLNLVAVS DFNVGAMENKGLNIFNANSLLASKKNSIDFSYARILTVVGHEYFHQYTGNRVTLRDWFQLTLKEG LTVHRENLFSEEMTKTVTTRLSHVDLLRSVQFLEDSSPLSHPIRPESYVSMENFYTTTVYDKGSE VMRMYLTILGEEYYKKGFDIYIKKNDGNTATCEDFNYAMEQAYKMKKADNSANLNQYLLWFSQSG
[0634]
[0635] TPHVSFKYNYDAEKKQYSIHVNQYTKPDENQKEKKPLFIPISVGLINPENGKEMISQTTLELTKE R0708.70180WO00 / R0708.70180US01 67 / 122
[0636] #14644680vl SDTFVFNNIAVKPIPSLFRGFSAPVYIEDQLTDEERILLLKYDSDAFVRYNSCTNIYMKQILMNY NEFLKAKNEKLESFQLTPVNAQFIDAIKYLLEDPHADAGFKSYIVSLPQDRYIINFVSNLDTDVL ADTKEYIYKQIGDKLNDVYYKMFKSLEAKADDLTYFNDESHVDFDQMNMRTLRNTLLSLLSKAQY PNILNEIIEHSKSPYPSNWLTSLSVSAYFDKYFELYDKTYKLSKDDELLLQEWLKTVSRSDRKDI YEILKKLENEVLKDSKNPNDIRAVYLPFTNNLRRFHDISGKGYKLIAEVITKTDKFNPMVATQLC EPFKLWNKLDTKRQELMLNEMNTMLQEPQISNNLKEYLLRLTNK (SEQ ID NO: 18) Puromycin-sensitive MWLAAAAPSLARRLLFLGPPPPPLLLLVFSRSSRRRLHSLGLAAMPEKRPFERLPADVSPINYSL aminopeptidase CLKPDLLDFTFEGKLEAAAQVRQATNQIVMNCADIDIITASYAPEGDEEIHATGFNYQNEDEKVT (NPEPPS) LSFPSTLQTGTGTLKIDFVGELNDKMKGFYRSKYTTPSGEVRYAAVTQFEATDARRAFPCWDEPA IKATFDISLVVPKDRVALSNMNVIDRKPYPDDENLVEVKFARTPVMSTYLVAFVVGEYDFVETRS KDGVCVRVYTPVGKAEQGKFALEVAAKTLPFYKDYFNVPYPLPKIDLIAIADFAAGAMENWGLVT YRETALLIDPKNSCSSSRQWVALVVGHELAHQWFGNLVTMEWWTHLWLNEGFASWIEYLCVDHCF PEYDIWTQFVSADYTRAQELDALDNSHPIEVSVGHPSEVDEIFDAISYSKGASVIRMLHDYIGDK DFKKGMNMYLTKFQQKNAATEDLWESLENASGKPIAAVMNTWTKQMGFPLIYVEAEQVEDDRLLR LSQKKFCAGGSYVGEDCPQWMVPITISTSEDPNQAKLKILMDKPEMNWLKNVKPDQWVKLNLGT VGFYRTQYSSAMLESLLPGIRDLSLPPVDRLGLQNDLFSLARAGIISTVEVLKVMEAFVNEPNYT VWSDLSCNLGILSTLLSHTDFYEEIQEFVKDVFSPIGERLGWDPKPGEGHLDALLRGLVLGKLGK AGHKATLEEARRRFKDHVEGKQILSADLRSPVYLTVLKHGDGTTLDIMLKLHKQADMQEEKNRIE RVLGATLLPDLIQKVLTFALSEEVRPQDTVSVIGGVAGGSKHGRKAAWKFIKDNWEELYNRYQGG FLISRLIKLSVEGFAVDKMAGEVKAFFESHPAPSAERTIQQCCENILLNAAWLKRDAESIHQYLL QRKASPPTV (SEQ ID NO: 19)
[0637] NPEPPS E366V MWLAAAAPSLARRLLFLGPPPPPLLLLVFSRSSRRRLHSLGLAAMPEKRPFERLPADVSPINYSL CLKPDLLDFTFEGKLEAAAQVRQATNQIVMNCADIDIITASYAPEGDEEIHATGFNYQNEDEKVT LSFPSTLQTGTGTLKIDFVGELNDKMKGFYRSKYTTPSGEVRYAAVTQFEATDARRAFPCWDEPA IKATFDISLVVPKDRVALSNMNVIDRKPYPDDENLVEVKFARTPVMSTYLVAFVVGEYDFVETRS KDGVCVRVYTPVGKAEQGKFALEVAAKTLPFYKDYFNVPYPLPKIDLIAIADFAAGAMENWGLVT YRETALLIDPKNSCSSSRQWVALVVGHVLAHQWFGNLVTMEWWTHLWLNEGFASWIEYLCVDHCF PEYDIWTQFVSADYTRAQELDALDNSHPIEVSVGHPSEVDEIFDAISYSKGASVIRMLHDYIGDK DFKKGMNMYLTKFQQKNAATEDLWESLENASGKPIAAVMNTWTKQMGFPLIYVEAEQVEDDRLLR LSQKKFCAGGSYVGEDCPQWMVPITISTSEDPNQAKLKILMDKPEMNWLKNVKPDQWVKLNLGT VGFYRTQYSSAMLESLLPGIRDLSLPPVDRLGLQNDLFSLARAGIISTVEVLKVMEAFVNEPNYT VWSDLSCNLGILSTLLSHTDFYEEIQEFVKDVFSPIGERLGWDPKPGEGHLDALLRGLVLGKLGK AGHKATLEEARRRFKDHVEGKQILSADLRSPVYLTVLKHGDGTTLDIMLKLHKQADMQEEKNRIE RVLGATLLPDLIQKVLTFALSEEVRPQDTVSVIGGVAGGSKHGRKAAWKFIKDNWEELYNRYQGG FLISRLIKLSVEGFAVDKMAGEVKAFFESHPAPSAERTIQQCCENILLNAAWLKRDAESIHQYLL QRKASPPTV (SEQ ID NO: 20)
[0638] Francisella MIYEFVMTDPKIKYLKDYKPSNYLIDETHLIFELDESKTRVTANLYIVANRENRENNTLVLDGVE tularensis LKLLSIKLNNKHLSPAEFAVNENQLIINNVPEKFVLQTVVEINPSANTSLEGLYKSGDVFSTQCE Aminopeptidase N ATGFRKITYYLDRPDVMAAFTVKIIADKKKYPIILSNGDKIDSGDISDNQHFAVWKDPFKKPCYL FALVAGDLASIKDTYITKSQRKVSLEIYAFKQDIDKCHYAMQAVKDSMKWDEDRFGLEYDLDTFM IVAVPDFNAGAMENKGLNIFNTKYIMASNKTATDKDFELVQSVVGHEYFHNWTGDRVTCRDWFQL SLKEGLTVFRDQEFTSDLNSRDVKRIDDVRIIRSAQFAEDASPMSHPIRPESYIEMNNFYTVTVY NKGAEIIRMIHTLLGEEGFQKGMKLYFERHDGQAVTCDDFVNAMADANNRDFSLFKRWYAQSGTP NIKVSENYDASSQTYSLTLEQTTLPTADQKEKQALHIPVKMGLINPEGKNIAEQVIELKEQKQTY TFENIAAKPVASLFRDFSAPVKVEHKRSEKDLLHIVKYDNNAFNRWDSLQQIATNIILNNADLND EFLNAFKSILHDKDLDKALISNALLIPIESTIAEAMRVIMVDDIVLSRKNVVNQLADKLKDDWLA VYQQCNDNKPYSLSAEQIAKRKLKGVCLSYLMNASDQKVGTDLAQQLFDNADNMTDQQTAFTELL KSNDKQVRDNAINEFYNRWRHEDLVVNKWLLSQAQISHESALDIVKGLVNHPAYNPKNPNKVYSL IGGFGANFLQYHCKDGLGYAFMADTVLALDKFNHQVAARMARNLMSWKRYDSDRQAMMKNALEKI KASNPSKNVFEIVSKSLES (SEQ ID NO: 21)
[0639] T. aquations MDAFTENLNKLAELAIRVGLNLEEGQEIVATAPIEAVDFVRLLAEKAYENGASLFTVLYGDNLIA Aminopeptidase T RKRLALVPEAHLDRAPAWLYEGMAKAFHEGAARLAVSGNDPKALEGLPPERVGRAQQAQSRAYRP TLSAITEFVTNWTIVPFAHPGWAKAVFPGLPEEEAVQRLWQAIFQATRVDQEDPVAAWEAHNRVL HAKVAFLNEKRFHALHFQGPGTDLTVGLAEGHLWQGGATPTKKGRLCNPNLPTEEVFTAPHRERV EGVVRASRPLALSGQLVEGLWARFEGGVAVEVGAEKGEEVLKKLLDTDEGARRLGEVALVPADNP IAKTGLVFFDTLFDENAASHIAFGQAYAENLEGRPSGEEFRRRGGNESMVHVDWMIGSEEVDVDG
[0640]
[0641] LLEDGTRVPLMRRGRWVI (SEQ ID NO: 22)
[0642] R0708.70180WO00 / R0708.70180US01 68 / 122
[0643] #14644680vl Bacillus MAKLDETLTMLKALTDAKGVPGNEREARDVMKTYIAPYADEVTTDGLGSLIAKKEGKSGGPKVMI stearothermophilus AGHLDEVGFMVTQIDDKGFIRFQTLGGWWSQVMLAQRVTIVTKKGDITGVIGSKPPHILPSEARK Peptidase M28 KPVEIKDMFIDIGATSREEAMEWGVRPGDMIVPYFEFTVLNNEKMLLAKAWDNRIGCAVAIDVLK QLKGVDHPNTVYGVGTVQEEVGLRGARTAAQFIQPDIAFAVDVGIAGDTPGVSEKEAMGKLGAGP HIVLYDATMVSHRGLREFVIEVAEELNIPHHFDAMPGVGTDAGAIHLTGIGVPSLTIAIPTRYIH SHAAILHRDDYENTVKLLVEVIKRLDADKVKQLTFDE (SEQ ID NO: 23)
[0644] Vibrio cholera MEDKVWISMGADAVGSLNPALSESLLPHSFASGSQVWIGEVAIDELAELSHTMHEQHNRCGGYMV Aminopeptidase HTSAQGAMAALMMPESIANFTIPAPSQQDLVNAWLPQVSADQITNTIRALSSFNNRFYTTTSGAQ ASDWLANEWRSLISSLPGSRIEQIKHSGYNQKSVVLTIQGSEKPDEWVIVGGHLDSTLGSHTNEQ SIAPGADDDASGIASLSEIIRVLRDNNFRPKRSVALMAYAAEEVGLRGSQDLANQYKAQGKKVVS VLQLDMTNYRGSAEDIVFITDYTDSNLTQFLTTLIDEYLPELTYGYDRCGYACSDHASWHKAGFS AAMPFESKFKDYNPKIHTSQDTLANSDPTGNHAVKFTKLGLAYVIEMANAGSSQVPDDSVLQDGT AKINLSGARGTQKRFTFELSQSKPLTIQTYGGSGDVDLYVKYGSAPSKSNWDCRPYQNGNRETCS FNNAQPGIYHVMLDGYTNYNDVALKASTQ (SEQ ID NO: 24)
[0645] Photobacterium MEDKVWISIGSDASQTVKSVMQSNARSLLPESLASNGPVWVGQVDYSQLAELSHHMHEDHQRCGG halotolerans YMVHSSPESAIAASNMPQSLVAFSIPEISQQDTVNAWLPQVNSQAITGTITSLTSFINRFYTTTS Aminopeptidase GAQASDWLANEWRSLSASLPNASVRQVSHFGYNQKSWLTITGSEKPDEWIVLGGHLDSTIGSHT NEQSVAPGADDDASGIASVTEIIRVLSENNFQPKRSIAFMAYAAEEVGLRGSQDLANQYKAEGKQ VISALQLDMTNYKGSVEDIVFITDYTDSNLTTFLSQLVDEYLPSLTYGFDTCGYACSDHASWHKA GFSAAMPFEAKFNDYNPMIHTPNDTLQNSDPTASHAVKFTKLGLAYAIEMASTTGGTPPPTGNVL KDGVPVNGLSGATGSQVHYSFELPAQKNLQISTAGGSGDVDLYVSFGSEATKQNWDCRPYRNGNN EVCTFAGATPGTYSIMLDGYRQFSGVTLKASTQ (SEQ ID NO: 25) Yersinia pestis MTQQPQAKYRHDYRAPDYTITDIDLDFALDAQKTTVTAVSKVKRQGTDVTPLILNGEDLTLISVS Aminopeptidase N VDGQAWPHYRQQDNTLVIEQLPADFTLTIVNDIHPATNSALEGLYLSGEALCTQCEAEGFRHITY YLDRPDVLARFTTRIVADKSRYPYLLSNGNRVGQGELDDGRHWVKWEDPFPKPSYLFALVAGDFD VLQDKFITRSGREVALEIFVDRGNLDRADWAMTSLKNSMKWDETRFGLEYDLDIYMIVAVDFFNM GAMENKGLNVFNSKYVLAKAETATDKDYLNIEAVIGHEYFHNWTGNRVTCRDWFQLSLKEGLTVF RDQEFSSDLGSRSVNRIENVRVMRAAQFAEDASPMAHAIRPDKVIEMNNFYTLTVYEKGSEVIRM MHTLLGEQQFQAGMRLYFERHDGSAATCDDFVQAMEDVSNVDLSLFRRWYSQSGTPLLTVHDDYD VEKQQYHLFVSQKTLPTADQPEKLPLHIPLDIELYDSKGNVIPLQHNGLPVHHVLNVTEAEQTFT FDNVAQKPIPSLLREFSAPVKLDYPYSDQQLTFLMQHARNEFSRWDAAQSLLATYIKLNVAKYQQ QQPLSLPAHVADAFRAILLDEHLDPALAAQILTLPSENEMAELFTTIDPQAISTVHEAITRCLAQ ELSDELLAVYVANMTPVYRIEHGDIAKRALRNTCLNYLAFGDEEFANKLVSLQYHQADNMTDSLA ALAAAVAAQLPCRDELLAAFDVRWNHDGLVMDKWFALQATSPAANVLVQVRTLLKHPAFSLSNPN RTRSLIGSFASGNPAAFHAADGSGYQFLVEILSDLNTRNPQVAARLIEPLIRLKRYDAGRQALMR KALEQLKTLDNLSGDLYEKITKALAA (SEQ ID NO: 26)
[0646] Vibrio anguillarum MEEKVWISIGGDATQTALRSGAQSLLPENLINQTSVWVGQVPVSELATLSHEMHENHQRCGGYMV Aminopeptidase HPSAQSAMSVSAMPLNLNAFSAPEITQQTTVNAWLPSVSAQQITSTITTLTQFKNRFYTTSTGAQ ASNWIADHWRSLSASLPASKVEQITHSGYNQKSVMLTITGSEKPDEWWIGGHLDSTLGSRTNES SIAPGADDDASGIAGVTEIIRLLSEQNFRPKRSIAFMAYAAEEVGLRGSQDLANRFKAEGKKVMS VMQLDMTNYQGSREDIVFITDYTDSNFTQYLTQLLDEYLPSLTYGFDTCGYACSDHASWHAVGYP AAMPFESKFNDYNPNIHSPQDTLQNSDPTGFHAVKFTKLGLAYVVEMGNASTPPTPSNQLKNGVP VNGLSASRNSKTWYQFELQEAGNLSIVLSGGSGDADLYVKYQTDADLQQYDCRPYRSGNNETCQF SNAQPGRYSILLHGYNNYSNASLVANAQ (SEQ ID NO: 27)
[0647] Salinivibrio MEDKKVWISIGADAQQTALSSGAQPLLAQSVAHNGQAWIGEVSESELAALSHEMHENHHRCGGYI spYCSC6 VHSSAQSAMAASNMPLSRASFIAPAISQQALVTPWISQIDSALIVNTIDRLTDFPNRFYTTTSGA Aminopeptidase QASDWIKQRWQSLSAGLAGASVTQISHSGYNQASVMLTIEGSESPDEWVVVGGHLDSTIGSRTNE QSIAPGADDDASGIAAVTEVIRVLAQNNFQPKRSIAFVAYAAEEVGLRGSQDVANQFKQAGKDVR GVLQLDMTNYQGSAEDIVFITDYTDNQLTQYLTQLLDEYLPTLNYGFDTCGYACSDHASWHQVGY PAAMPFEAKFNDYNPNIHTPQDTLANSDSEGAHAAKFTKLGLAYTVELANADSSPNPGNELKLGE PINGLSGARGNEKYFNYRLDQSGELVIRTYGGSGDVDLYVKANGDVSTGNWDCRPYRSGNDEVCR FDNATPGNYAVMLRGYRTYDNVSLIVE (SEQ ID NO: 28)
[0648] Vibrio proteolyticus MPPITQQATVTAWLPQVDASQITGTISSLESFTNRFYTTTSGAQASDWIASEWQALSASLPNASV Aminopeptidase I KQVSHSGYNQKSVVMTITGSEAPDEWIVIGGHLDSTIGSHTNEQSVAPGADDDASGIAAVTEVIR VLSENNFQPKRSIAFMAYAAEEVGLRGSQDLANQYKSEGKNVVSALQLDMTNYKGSAQDVVFITD YTDSNFTQYLTQLMDEYLPSLTYGFDTCGYACSDHASWHNAGYPAAMPFESKFNDYNPRIHTTQD
[0649]
[0650] TLANSDPTGSHAKKFTQLGLAYAIEMGSATGDTPTPGNQLE (SEQ ID NO: 29) R0708.70180WO00 / R0708.70180US01 69 / 122
[0651] #14644680vl Vibrio proteolyticus MPPITQQATVTAWLPQVDASQITGTISSLESFTNRFYTTTSGAQASDWIASEWQFLSASLPNASV Aminopeptidase I KQVSHSGYNQKSVVMTITGSEAPDEWIVIGGHLDSTIGSHTNEQSVAPGADDDASGIAAVTEVIR (A55F) VLSENNFQPKRSIAFMAYAAEEVGLRGSQDLANQYKSEGKNVVSALQLDMTNYKGSAQDVVFITD YTDSNFTQYLTQLMDEYLPSLTYGFDTCGYACSDHASWHNAGYPAAMPFESKFNDYNPRIHTTQD TLANSDPTGSHAKKFTQLGLAYAIEMGSATGDTPTPGNQLE (SEQ ID NO: 30)
[0652] P. furiosus MVDWELMKKIIESPGVSGYEHLGIRDLVVDILKDVADEVKIDKLGNVIAHFKGSAPKVMVAAHMD Aminopeptidase I KIGLMVNHIDKDGYLRVVPIGGVLPETLIAQKIRFFTEKGERYGVVGVLPPHLRREAKDQGGKID WDSIIVDVGASSREEAEEMGFRIGTIGEFAPNFTRLSEHRFATPYLDDRICLYAMIEAARQLGEH EADIYIVASVQEEIGLRGARVASFAIDPEVGIAMDVTFAKQPNDKGKIVPELGKGPVMDVGPNIN PKLRQFADEVAKKYEIPLQVEPSPRPTGTDANVMQINREGVATAVLSIPIRYMHSQVELADARDV
[0653]
[0654] DNTIKLAKALLEELKPMDFTPLE (SEQ ID NO: 31)
[0655] Amino Acid Recognizers
[0656]
[0277] In some embodiments, the method further comprises contacting the polypeptide sample with a reaction mixture comprising one or more amino acid recognizers (e.g., one or more amino acid binding proteins not having peptide cleavage activity). In some embodiments, the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0657]
[0278] In some embodiments, an amino acid recognizer comprises an amino acid binding protein, such as a ClpS protein (e.g., Planctomycetia bacterium ClpS protein), a UBR protein (e.g., Kluyveromyces marxianus UBR protein), an Ntaql protein (e.g., Scleropages formosus Ntaql protein), or a variant or homolog thereof. In some embodiments, an amino acid recognizer comprises a label (e.g., a detectable label, such as a luminescent label). Examples of amino acid recognizers (e.g., recognition molecules) are described in detail in PCT International Publication No. W02020 / 102741A1, filed November 15, 2019, PCT International Publication No. WO2021 / 236983A2, filed May 20, 2021, and co-pending U. S. Serial No. 63 / 395,328, filed August 4, 2022, which describe amino acid recognizers (e.g., recognition molecules) in detail, the relevant content of each of which is incorporated by reference in its entirety.
[0658]
[0279] As described herein, the polypeptide sample may comprise a plurality of amino acids. For instance, the one or more immobilization complex-conjugated polypeptides (also referred to as “a polypeptide” or “the polypeptide”) may comprise a plurality of amino acids. In some embodiments, the polypeptide comprises a first amino acid (e.g., a terminal amino acid, an internal amino acid) to which the one or more amino acid recognizers may bind, and at least one other (e.g., upstream, downstream) amino acid (e.g., a second amino acid).
[0659]
[0280] The one or more amino acid recognizers may comprise a first set of one or more amino acid recognizers that bind to the polypeptide. In some embodiments, the first set of one or more amino acid recognizers may bind to an amino acid of the polypeptide and, in some embodiments, to one or more additional amino acids. In some embodiments, the first set of one or more amino acid recognizers may bind to a terminal amino acid of the polypeptide and, in some embodiments, to one or more downstream amino acids. In some embodiments, the first set of one or more amino acid recognizers may bind to an internal amino acid of the polypeptide and, in some embodiments, to one or more upstream and / or downstream amino acids. At least one (and, in some embodiments, each) of the one or more amino acid recognizers may be labeled with a fluorescent dye that emits emission light when excited with excitation R0708.70180WO00 / R0708.70180US01 70 / 122
[0660] #14644680vl light, as described herein. In some cases, the one or more amino acid recognizers comprise a plurality of types of amino acid recognizers. In certain cases, each type of amino acid recognizer may only bind to certain amino acids. As an illustrative example, a first type of amino acid recognizer may preferentially bind to leucine, isoleucine, and valine. As another illustrative example, a second type of amino acid recognizer may preferentially bind to phenylalanine, tyrosine, and tryptophan. As another illustrative example, a third type of amino acid recognizer may preferentially bind to arginine. In some embodiments, the first set of one or more amino acid recognizers comprises one type of amino acid recognizer. In some embodiments, the first set of one or more amino acid recognizers comprises two or more types of amino acid recognizers. In some embodiments, each type of recognizer is labeled with a unique dye and / or a unique number of dyes. Accordingly, the emission light from the dye-labeled amino acid recognizers may be used to obtain information about (e.g., identify) the amino acid to which the dye-labeled amino acid recognizer is bound, and in some embodiments, about one or more additional amino acids.
[0661] Polypeptide Analysis
[0662]
[0281] In some embodiments, the method further comprises monitoring a signal for signal pulses corresponding to interactions between one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides (also referred to as “a polypeptide” or “the polypeptide”) of the polypeptide sample.
[0663]
[0282] In some embodiments, the method further comprises determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal.
[0664]
[0283] In some embodiments, one or more characteristics of a first series of signal pulses indicative of a first series of binding events between a first set of one or more amino acid recognizers and a first amino acid of a polypeptide (e.g., a terminal amino acid, an internal amino acid) may be impacted by one or more chemical characteristics of the polypeptide. In certain instances, one or more modifications of one or more amino acids (e.g., post-translational modifications, mutations, bonds to binding components) may promote a covalent or non-covalent interaction between one or more amino acid recognizers and the first amino acid (e.g., through electrostatic attraction, pi stacking, hydrogen bond formation, etc.), thereby increasing pulse duration. In certain instances, one or more modifications of one or more amino acids (e.g., post-translational modifications, mutations, presence of binding components) may discourage a covalent or non-covalent interaction between one or more amino acid recognizers and the first amino acid (e.g., through electrostatic repulsion, steric hindrance, etc.), thereby decreasing pulse duration.
[0665]
[0284] In some embodiments, determining at least one chemical characteristic of the polypeptide may comprise comparing at least one characteristic of the series of signal pulses with known characteristics of known amino acid segments. In some embodiments, a protein from which the polypeptide originated may be identified. For example, as described herein, the techniques may include identifying one of more of the amino acids of an amino acid segment (e.g., a tripeptide segment, a tetrapeptide segment). Based on the identified amino acids, a protein from which the polypeptide originated may be identified. For R0708.70180WO00 / R0708.70180US01 71 / 122
[0666] #14644680vl example, identifying the protein from which the polypeptide originated may comprise comparing the identified amino acids of the amino acid segment to known information. In some embodiments, identifying a protein from which the polypeptide originated may comprise identifying a pattern in the amino acid segment(s) also present in a candidate matching protein. The pattern may be unique to the candidate matching protein relative to other candidate matching proteins. Accordingly, the techniques described herein may allow for identifying a protein from which a polypeptide originated based on identifying only a portion of the amino acids of the polypeptide. In some embodiments, identifying a polypeptide comprises identifying a protein from which a polypeptide originated. In some embodiments, identifying a polypeptide comprises identifying a pattern of amino acids present in the polypeptide and identifying a candidate matching polypeptide comprising the pattern of amino acids.
[0667]
[0285] In some embodiments, contacting the polypeptide with the one or more amino acid recognizers comprises introducing the one or more amino acid recognizers onto a device (e.g., by loading a solution comprising the one or more amino acid recognizers onto an integrated device comprising the polypeptide). The one or more amino acid recognizers may periodically bind to the polypeptide (e.g., to at least one amino acid of the polypeptide). The rate at which the one or more amino acid recognizers bind to the polypeptide is referred to herein as the binding rate. In some embodiments, the one or more amino acid recognizers may be labeled with (e.g., conjugated to) fluorescent dyes that may become excited when the one or more recognizers are bound to or in the vicinity of an amino acid of the polypeptide. Therefore, periodic signals emitted by the fluorescent dyes may be characteristic of the binding rate of the amino acid recognizers.
[0668]
[0286] In some embodiments, the monitoring comprises detecting a series of signal pulses while the polypeptide is being degraded, wherein a characteristic pattern in the series of signal pulses is indicative of the at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides.
[0669]
[0287] In some embodiments, determining at least one chemical characteristic of the polypeptide comprises identifying one or more amino acids present in the polypeptide. In some embodiments, identifying an amino acid comprises determining which of the naturally-occurring 20 amino acids is present. In some embodiments, the identity of an amino acid is selected from the group consisting of alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, and valine.
[0670]
[0288] In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining a subset of potential amino acids that can be present in the polypeptide. In some embodiments, this can be accomplished by determining that an amino acid is not one or more specific amino acids (and therefore could be any of the other amino acids). In some embodiments, this can be accomplished by determining which of a specified subset of amino acids (e.g., based on size, charge, hydrophobicity, post-translational modification, binding properties) could be in the polypeptide (e.g., using a recognizer that binds to a specified subset of two or more amino acids).
[0671] R0708.70180WO00 / R0708.70180US01 72 / 122
[0672] #14644680vl
[0289] In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises a post-translational modification. The post-translational modification may affect the series of signals emitted by a dye-labeled amino acid recognizer bound to the polypeptide (e.g., to a terminal amino acid and / or an internal amino acid). In some embodiments, the series of signals emitted by the dye-labeled amino acid recognizer may be impacted by the post-translational modification even if the post-translational modification is to an amino acid which does not bind to the dye-labeled amino acid recognizer. In some embodiments, a post-translational modification of an amino acid to which a dye-labeled recognizer binds and / or a post-translational modification of one or more upstream or downstream amino acids may cause at least one characteristic of a series of signal pulses (e.g., pulse duration, interpulse duration, recognition segment duration, intersegment duration, cleavage rate, cleavage time, intensity, wavelength, fluorescence lifetime, absence of signal pulses) to change (e.g., increase, decrease) relative to an unmodified amino acid. Non-limiting examples of post-translational modifications include acetylation (e.g., acetylated lysine), ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation (e.g., glycosylated asparagine), O-linked glycosylation (e.g., glycosylated serine, glycosylated threonine), hydroxylation, methylation (e.g., methylated lysine, methylated arginine), myristoylation (e.g., myristoylated glycine), neddylation, nitration (e.g., nitrated tyrosine), chlorination (e.g., chlorinated tyrosine), oxidation / reduction (e.g., oxidized cysteine, oxidized methionine), carbonylation (e.g., carbonylated lysine, carbonylated proline, carbonylated arginine, carbonylated threonine), palmitoylation (e.g., palmitoylated cysteine), phosphorylation, prenylation (e.g., prenylated cysteine), S-nitrosylation (e.g., S-nitrosylated cysteine, S-nitrosylated methionine), sulfation, glycation (e.g., glycated lysine), sumoylation (e.g., sumoylated lysine), and ubiquitination (e.g., ubiquitinated lysine).
[0673]
[0290] In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises a phosphorylated side chain. For example, in some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises phosphorylated threonine (e.g., phospho-threonine). In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises phosphorylated tyrosine (e.g., phospho-tyrosine). In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises phosphorylated serine (e.g., phospho-serine).
[0674]
[0291] In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises a chemically modified variant, an unnatural amino acid, or a proteinogenic amino acid such as selenocysteine and pyrrolysine. Examples of unnatural amino acids include, without limitation, 2-naphthyl-alanine, statine, homoalanine, a-amino acid, [32-amino acid, [33-amino acid, y-amino acid, 3-pyridyl-alanine, 4-fluorophenyl-alanine, cyclohexyl-alanine, N-alkyl amino acid, peptoid amino acid, homo-cysteine, penicillamine, 3-nitro-tyrosine, homo-phenyl-alanine, / -leucine, hydroxy-proline, 3-Abz, 5 -F -tryptophan, and azabicyclo- [2.2. l]heptane.
[0675]
[0292] In some embodiments, determining at least one chemical characteristic of a polypeptide comprises R0708.70180WO00 / R0708.70180US01 73 / 122
[0676] #14644680vl determining that an amino acid comprises an oxidative modification. In some embodiments, the oxidative modification comprises an oxidatively-damaged side chain of an amino acid. In some embodiments, the oxidatively-damaged side chain comprises a cysteine-derived product (e.g., disulfide, sulfinic acid, sulfonic acid, sulfenic acid, S-nitrosocysteine), a tyrosine-derived product (e.g., di-tyrosine, 3,4-dihydroxyphenylalanine, 3 -chlorotyrosine, 3-nitrotyrosine), a histidine-derived product (e.g., 2-oxohistidine, 4-hydroxy-2-oxohistidine, di-histidine, asparagine, aspartic acid, urea), a methioninederived product (e.g., sulfoxide, sulfone), a tryptophan-derived product (e.g., di-tryptophan, N-formylkynurenine, kynurenine, 2-oxo-tryptophan oxindolylalanine, 6-nitrotryptophan, hydroxytryptophan), a phenylalanine-derived product (e.g., meta-tyrosine, o / 7 / ?o-tyrosinc). or a generic side-chain product (e.g., alcohol, hydroperoxide, aldehyde / ketone carbonyl). Examples of oxidatively damaged amino acids are known in the art, see, e.g., Hawkins, C. L., Davies, M. J. Detection, identification, and quantification of oxidative protein modifications. J Biol Chem. 2019 Dec 20;294(51): 19683-19708.
[0677]
[0293] In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that an amino acid comprises a side chain characterized by one or more biochemical properties. For example, an amino acid may comprise a nonpolar aliphatic side chain, a positively charged side chain, a negatively charged side chain, a nonpolar aromatic side chain, or a polar uncharged side chain. Non-limiting examples of an amino acid comprising a nonpolar aliphatic side chain include alanine, glycine, valine, leucine, methionine, and isoleucine. Non-limiting examples of an amino acid comprising a positively charged side chain includes lysine, arginine, and histidine. Non-limiting examples of an amino acid comprising a negatively charged side chain include aspartate and glutamate. Non-limiting examples of an amino acid comprising a nonpolar, aromatic side chain include phenylalanine, tyrosine, and tryptophan. Non-limiting examples of an amino acid comprising a polar uncharged side chain include serine, threonine, cysteine, proline, asparagine, and glutamine.
[0678]
[0294] In some embodiments, determining at least one chemical characteristic of a polypeptide comprises determining that at least one amino acid is bound (e.g., via a covalent or non-covalent interaction) to a binding component. Non-limiting examples of suitable binding components include a nucleic acid (e.g., DNA, RNA), a linker, and an antibody. In some instances, one or more amino acids of a polypeptide may be bound to a nucleic acid via one or more non-covalent interactions. In some instances, one or more amino acids of a polypeptide may be bound to a linker via one or more covalent interactions.
[0679]
[0295] In some embodiments, a protein or polypeptide can be digested into a plurality of smaller polypeptides and chemical characteristics can be determined for one or more of these smaller polypeptides. In some embodiments, a first terminus (e.g., N or C terminus) of a polypeptide is immobilized and the other terminus (e.g., the C or N terminus) is analyzed as described herein.
[0680]
[0296] A non-limiting example of polypeptide structure analysis by detecting single molecule binding interactions during a polypeptide degradation process is illustrated in FIG. 1. An example signal trace is shown depicting different association (e.g., binding) events at times corresponding to changes in the signal. As shown, an association event between an amino acid recognizer and a terminal end of a R0708.70180WO00 / R0708.70180US01 74 / 122
[0681] #14644680vl polypeptide produces a change in magnitude of the signal that persists for a duration of time. Different association events are illustrated for different amino acids exposed at the terminal end of the polypeptide. As described herein, an amino acid that is “exposed” at the terminus of a polypeptide is an amino acid that is still attached to the polypeptide and that becomes the terminal amino acid upon removal of the prior terminal amino acid during degradation (e.g., either alone or along with one or more additional amino acids).
[0682]
[0297] As generically depicted, the association events between amino acid recognizers and different types of amino acids at the terminal end of the polypeptide produce distinctive changes in the signal, referred to herein as a characteristic pattern, which may be used to determine chemical characteristics of the polypeptide. In some embodiments, a characteristic pattern corresponding to one type of terminal amino acid can be used to determine structural information for the terminal amino acid and one or more amino acids contiguous to the terminal amino acid. Accordingly, in some embodiments, a characteristic pattern corresponding to one type of terminal amino acid can be used to determine structural information for at least two (e.g., at least three, at least four, at least five, two, three, four, or between two and five) amino acids of a polypeptide.
[0683]
[0298] In some embodiments, a transition from one characteristic pattern to another is indicative of amino acid cleavage. As used herein, in some embodiments, amino acid cleavage refers to the removal of at least one amino acid from a terminus of a polypeptide (e.g., the removal of at least one terminal amino acid from the polypeptide). In some embodiments, amino acid cleavage is determined by inference based on a time duration between characteristic patterns. In some embodiments, amino acid cleavage is determined by detecting a change in signal produced by association of a labeled cleaving agent with an amino acid at the terminus of the polypeptide. As amino acids are sequentially cleaved from the terminus of the polypeptide during degradation, a series of changes in magnitude, or a series of signal pulses, is detected.
[0684]
[0299] In some embodiments, signal data can be analyzed to extract signal pulse information by applying threshold levels to one or more parameters of the signal data. For example, in some embodiments, a threshold magnitude level may be applied to the signal data of a signal trace. In some embodiments, the threshold magnitude level is a minimum difference between a signal detected at a point in time and a baseline determined for a given set of data. In some embodiments, a signal pulse is assigned to each portion of the data that is indicative of a change in magnitude exceeding the threshold magnitude level and persisting for a duration of time. In some embodiments, a threshold time duration may be applied to a portion of the data that satisfies the threshold magnitude level to determine whether a signal pulse is assigned to that portion. For example, experimental artifacts may give rise to a change in magnitude exceeding the threshold magnitude level but that does not persist for a duration of time sufficient to assign a signal pulse with a desired confidence (e.g., transient association events which could be non-discriminatory for amino acid type, non-specific detection events such as diffusion into an observation region or reagent sticking within an observation region). Accordingly, in some embodiments, a signal pulse is extracted from signal data based on a threshold magnitude level and a threshold time duration. R0708.70180WO00 / R0708.70180US01 75 / 122
[0685] #14644680vl
[0300] In some embodiments, a peak in magnitude of a signal pulse is determined by averaging the magnitude detected over a duration of time that persists above the threshold magnitude level. It should be appreciated that, in some embodiments, a “signal pulse” as used herein can refer to a change in signal data that persists for a duration of time above a baseline (e.g., raw signal data), or to signal pulse information extracted therefrom (e.g., processed signal data).
[0686]
[0301] In some embodiments, signal pulse information can be analyzed to identify different types of amino acids in a polypeptide based on different characteristic patterns in a series of signal pulses. For example, as shown in FIG. 1, the signal pulse information is indicative of different types of amino acids at a terminal end of a polypeptide (e.g., arginine, leucine, isoleucine, phenylalanine). By way of example, the signal pulses detected at the earliest time points provide information indicative of (at least) arginine at the terminus of the polypeptide based on a first characteristic pattern, and the signal pulses detected at the latest time points provide information indicative of at least phenylalanine at the terminus of the polypeptide based on a second characteristic pattern.
[0687]
[0302] In some embodiments, each signal pulse of a characteristic pattern comprises a pulse duration corresponding to an association event between an amino acid recognizer and an amino acid ligand. In some embodiments, the pulse duration is characteristic of a dissociation rate of binding. In some embodiments, each signal pulse of a characteristic pattern is separated from another signal pulse of the characteristic pattern by an interpulse duration. In some embodiments, the interpulse duration is characteristic of an association rate of binding. In some embodiments, a change in magnitude in a signal can be determined for a signal pulse based on a difference between baseline and the peak of a signal pulse. In some embodiments, a characteristic pattern is determined based on pulse duration. In some embodiments, a characteristic pattern is determined based on pulse duration and interpulse duration. In some embodiments, a characteristic pattern is determined based on any one or more of pulse duration, interpulse duration, and change in magnitude.
[0688]
[0303] Accordingly, as illustrated by FIG. 1, in some embodiments, polypeptide analysis is performed by detecting a series of signal pulses indicative of association of one or more amino acid recognizers with successive amino acids exposed at the terminus of a polypeptide in an ongoing degradation reaction. The series of signal pulses can be analyzed to determine characteristic patterns in the series of signal pulses, and the time course of characteristic patterns can be used to determine chemical characteristics throughout an amino acid sequence of the polypeptide.
[0689]
[0304] As described herein, signal pulse information may be used to identify an amino acid based on a characteristic pattern in a series of signal pulses. In some embodiments, a characteristic pattern comprises a plurality of signal pulses, each signal pulse comprising a pulse duration. In some embodiments, the plurality of signal pulses may be characterized by a summary statistic (e.g., mean, median, time decay constant) of the distribution of pulse durations in a characteristic pattern. In some embodiments, the mean pulse duration of a characteristic pattern is between about 1 millisecond and about 10 seconds (e.g., between about 1 ms and about 1 s, between about 1 ms and about 100 ms, between about 1 ms and about 10 ms, between about 10 ms and about 10 s, between about 100 ms and about 10 s, between about 1 s R0708.70180WO00 / R0708.70180US01 76 / 122
[0690] #14644680vl and about 10 s, between about 10 ms and about 100 ms, or between about 100 ms and about 500 ms). In some embodiments, the mean pulse duration is between about 50 milliseconds and about 2 seconds, between about 50 milliseconds and about 500 milliseconds, or between about 500 milliseconds and about 2 seconds.
[0691]
[0305] In some embodiments, different characteristic patterns corresponding to different types of amino acids in a single polypeptide may be distinguished from one another based on a statistically significant difference in the summary statistic. For example, in some embodiments, one characteristic pattern may be distinguishable from another characteristic pattern based on a difference in mean pulse duration of at least 10 milliseconds (e.g., between about 10 ms and about 10 s, between about 10 ms and about 1 s, between about 10 ms and about 100 ms, between about 100 ms and about 10 s, between about 1 s and about 10 s, or between about 100 ms and about 1 s). In some embodiments, the difference in mean pulse duration is at least 50 ms, at least 100 ms, at least 250 ms, at least 500 ms, or more. In some embodiments, the difference in mean pulse duration is between about 50 ms and about 1 s, between about 50 ms and about 500 ms, between about 50 ms and about 250 ms, between about 100 ms and about 500 ms, between about 250 ms and about 500 ms, or between about 500 ms and about 1 s. In some embodiments, the mean pulse duration of one characteristic pattern is different from the mean pulse duration of another characteristic pattern by about 10-25%, 25-50%, 50-75%, 75-100%, or more than 100%, for example by about 2-fold, 3 -fold, 4-fold, 5 -fold, or more. It should be appreciated that, in some embodiments, smaller differences in mean pulse duration between different characteristic patterns may require a greater number of pulse durations within each characteristic pattern to distinguish one from another with statistical confidence.
[0692]
[0306] In some embodiments, a characteristic pattern generally refers to a plurality of association events between an amino acid of a polypeptide and a means for binding the amino acid (e.g., an amino acid recognition molecule). In some embodiments, a characteristic pattern comprises at least 10 association events (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000, or more, association events). In some embodiments, a characteristic pattern comprises between about 10 and about 1,000 association events (e.g., between about 10 and about 500 association events, between about 10 and about 250 association events, between about 10 and about 100 association events, or between about 50 and about 500 association events). In some embodiments, the plurality of association events is detected as a plurality of signal pulses.
[0693]
[0307] In some embodiments, a characteristic pattern refers to a plurality of signal pulses which may be characterized by a summary statistic as described herein. In some embodiments, a characteristic pattern comprises at least 10 signal pulses (e.g., at least 25, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1,000, or more, signal pulses). In some embodiments, a characteristic pattern comprises between about 10 and about 1,000 signal pulses (e.g., between about 10 and about 500 signal pulses, between about 10 and about 250 signal pulses, between about 10 and about 100 signal pulses, or between about 50 and about 500 signal pulses).
[0694]
[0308] In some embodiments, a characteristic pattern refers to a plurality of association events between an R0708.70180WO00 / R0708.70180US01 77 / 122
[0695] #14644680vl amino acid recognition molecule and an amino acid of a polypeptide occurring over a time interval prior to removal of the amino acid (e.g., a cleavage event). In some embodiments, a characteristic pattern refers to a plurality of association events occurring over a time interval between two cleavage events (e.g., prior to removal of the amino acid and after removal of an amino acid previously exposed at the terminus). In some embodiments, the time interval of a characteristic pattern is between about 1 minute and about 30 minutes (e.g., between about 1 minute and about 20 minutes, between about 1 minute and 10 minutes, between about 5 minutes and about 20 minutes, between about 5 minutes and about 15 minutes, or between about 5 minutes and about 10 minutes).
[0696]
[0309] In some embodiments, the series of signal pulses comprises a series of changes in magnitude of an optical signal overtime. In some embodiments, the series of changes in the optical signal comprises a series of changes in luminescence produced during association events. In some embodiments, luminescence is produced by a detectable label associated with one or more reagents of a sequencing reaction. For example, in some embodiments, each of the one or more amino acid recognizers comprises a luminescent label. In some embodiments, a cleaving agent comprises a luminescent label. Examples of luminescent labels and their use in accordance with the disclosure are provided herein.
[0697]
[0310] In some embodiments, the series of signal pulses comprises a series of changes in magnitude of an electrical signal overtime. In some embodiments, the series of changes in the electrical signal comprises a series of changes in conductance produced during association events. In some embodiments, conductivity is produced by a detectable label associated with one or more reagents of a sequencing reaction. For example, in some embodiments, each of the one or more amino acid recognizers comprises a conductivity label. Examples of conductivity labels and their use in accordance with the disclosure are provided elsewhere herein. Methods for identifying single molecules using conductivity labels have been described (see, e.g., U. S. Patent Publication No. 2017 / 0037462).
[0698]
[0311] In some embodiments, the series of changes in conductance comprises a series of changes in conductance through a nanopore. For example, methods of evaluating receptor-ligand interactions using nanopores have been described (see, e.g., Thakur, A. K. & Movileanu, L. (2019) Nature Biotechnology 37(1)). The inventors have recognized and appreciated that such nanopores may be used to monitor polypeptide sequencing reactions in accordance with the disclosure. Accordingly, in some embodiments, the disclosure provides methods of polypeptide analysis comprising contacting a single polypeptide molecule with one or more amino acid recognizers described herein, where the single polypeptide molecule is immobilized to a nanopore. In some embodiments, the methods further comprise detecting a series of changes in conductance through the nanopore indicative of association of the one or more amino acid recognizers with successive amino acids exposed at a terminus of the single polypeptide while the single polypeptide is being degraded.
[0699]
[0312] As used herein, sequencing a polypeptide refers to determining sequence information for a polypeptide. In some embodiments, this can involve determining the identity of each sequential amino acid for a portion (or all) of the polypeptide. However, in some embodiments, this can involve assessing the identity of a subset of amino acids within the polypeptide (e.g., and determining the relative position R0708.70180WO00 / R0708.70180US01 78 / 122
[0700] #14644680vl of one or more amino acid types without determining the identity of each amino acid in the polypeptide). However, in some embodiments, amino acid content information can be obtained from a polypeptide without directly determining the relative position of different types of amino acids in the polypeptide. The amino acid content alone may be used to infer the identity of the polypeptide that is present (e.g., by comparing the amino acid content to a database of polypeptide information and determining which polypeptide(s) have the same amino acid content).
[0701]
[0313] In some embodiments, sequence information for a plurality of polypeptide products obtained from a longer polypeptide or protein (e.g., via enzymatic and / or chemical cleavage) can be analyzed to reconstruct or infer the sequence of the longer polypeptide or protein.
[0702]
[0314] In some aspects, the polypeptide analysis described herein generates data indicating how a polypeptide interacts with a binding means while the polypeptide is being degraded by a cleaving means. As discussed above, the data can include a series of characteristic patterns corresponding to association events at a terminus of a polypeptide in between cleavage events at the terminus. In some embodiments, methods of polypeptide analysis described herein comprise contacting a single polypeptide molecule with a binding means and a cleaving means, where the binding means and the cleaving means are configured to achieve at least 10 association events prior to a cleavage event. In some embodiments, the means are configured to achieve the at least 10 association events between two cleavage events.
[0703]
[0315] In some embodiments, a plurality of single -molecule sequencing reactions are performed in parallel in an array of sample wells. In some embodiments, an array comprises between about 10,000 and about 1,000,000 sample wells. The volume of a sample well may be between about 10'21liters and about 10'15liters, in some implementations. Because the sample well has a small volume, detection of singlemolecule events may be possible as only about one polypeptide may be within a sample well at any given time. Statistically, some sample wells may not contain a single -molecule sequencing reaction and some may contain more than one single polypeptide molecule. However, an appreciable number of sample wells may each contain a single-molecule reaction (e.g., at least 30% in some embodiments), so that single-molecule analysis can be carried out in parallel for a large number of sample wells. In some embodiments, the binding means and the cleaving means are configured to achieve at least 10 association events prior to a cleavage event in at least 10% (e.g., 10-50%, more than 50%, 25-75%, at least 80%, or more) of the sample wells in which a single-molecule reaction is occurring. In some embodiments, the binding means and the cleaving means are configured to achieve at least 10 association events prior to a cleavage event for at least 50% (e.g., more than 50%, 50-75%, at least 80%, or more) of the amino acids of a polypeptide in a single-molecule reaction.
[0704] Devices and Systems
[0705]
[0316] Methods in accordance with the disclosure, in some aspects, may be performed using a system that permits single-molecule analysis. The system may include an integrated device and an instrument configured to interface with the integrated device. The integrated device may include an array of pixels, where individual pixels include a sample well and at least one photodetector. The sample wells of the integrated device may be formed on or through a surface of the integrated device and be configured to R0708.70180WO00 / R0708.70180US01 79 / 122
[0706] #14644680vl receive a sample placed on the surface of the integrated device. Collectively, the sample wells may be considered as an array of sample wells. The plurality of sample wells may have a suitable size and shape such that at least a portion of the sample wells receive a single sample (e.g., a single molecule, such as a polypeptide). In some embodiments, the number of samples within a sample well may be distributed among the sample wells of the integrated device such that some sample wells contain one sample while others contain zero, two or more samples.
[0707]
[0317] Excitation light is provided to the integrated device from one or more light source external to the integrated device. Optical components of the integrated device may receive the excitation light from the light source and direct the light towards the array of sample wells of the integrated device and illuminate an illumination region within the sample well. In some embodiments, a sample well may have a configuration that allows for the sample to be retained in proximity to a surface of the sample well, which may ease delivery of excitation light to the sample and detection of emission light from the sample. A sample positioned within the illumination region may emit emission light in response to being illuminated by the excitation light. For example, the sample may be labeled with a fluorescent label, which emits light in response to achieving an excited state through the illumination of excitation light. Emission light emitted by a sample may then be detected by one or more photodetectors within a pixel corresponding to the sample well with the sample being analyzed. When performed across the array of sample wells, which may range in number between approximately 10,000 pixels to 1,000,000 pixels according to some embodiments, multiple samples can be analyzed in parallel.
[0708]
[0318] The integrated device may include an optical system for receiving excitation light and directing the excitation light among the sample well array. Examples of suitable components, e.g., for coupling excitation light to a sample well and / or directing emission light to a photodetector, to include in an integrated device are described in U. S. Patent Application No. 14 / 821,688, fded August 7, 2015, titled “INTEGRATED DEVICE FOR PROBING, DETECTING AND ANALYZING MOLECULES,” and U. S. Patent Application No. 14 / 543,865, fded November 17, 2014, titled “INTEGRATED DEVICE WITH EXTERNAL LIGHT SOURCE FOR PROBING, DETECTING, AND ANALYZING MOLECULES,” both of which are incorporated by reference in their entirety. Examples of suitable grating couplers and waveguides that may be implemented in the integrated device are described in U. S. Patent Application No. 15 / 844,403, fded December 15, 2017, titled “OPTICAL COUPLERAND WAVEGUIDE SYSTEM,” which is incorporated by reference in its entirety.
[0709]
[0319] Additional photonic structures may be positioned between the sample wells and the photodetectors and configured to reduce or prevent excitation light from reaching the photodetectors, which may otherwise contribute to signal noise in detecting emission light. Examples of suitable photonic structures may include spectral filters, a polarization filters, and spatial filters and are described in U. S. Patent Application No. 16 / 042,968, fded July 23, 2018, titled “OPTICAL REJECTION PHOTONIC STRUCTURES,” and U. S. Provisional Patent Application No. 63 / 124,655, fded December 11, 2020, titled “INTEGRATED CIRCUIT WITH IMPROVED CHARGE TRANSFER EFFICIENCY AND ASSOCIATED TECHNIQUES,” both of which are incorporated by reference in their entirety.
[0710] R0708.70180WO00 / R0708.70180US01 80 / 122
[0711] #14644680vl
[0320] Characteristics of the detected emission light may provide an indication for identifying the label associated with the emission light. Such characteristics may include any suitable type of characteristic, including an arrival time of photons detected by a photodetector, an amount of photons accumulated over time by a photodetector, and / or a distribution of photons across two or more photodetectors. In some embodiments, such characteristics can be any one or a combination of two or more of luminescence lifetime, luminescence intensity, brightness, absorption spectra, emission spectra, luminescence quantum yield, wavelength (e.g., peak wavelength), and signal characteristics (e.g., pulse duration, interpulse durations, change in signal magnitude).
[0712]
[0321] In some embodiments, a photodetector may have a configuration that allows for the detection of one or more timing characteristics associated with a sample’s emission light (e.g., luminescence lifetime). The photodetector may detect a distribution of photon arrival times after a pulse of excitation light propagates through the integrated device, and the distribution of arrival times may provide an indication of a timing characteristic of the sample’s emission light (e.g., a proxy for luminescence lifetime). In some embodiments, the one or more photodetectors provide an indication of the probability of emission light emitted by the label (e.g., luminescence intensity). In operation, parallel analyses of samples within the sample wells are carried out by exciting some or all of the samples within the wells using excitation light and detecting signals from sample emission with the photodetectors.
[0713]
[0322] The instrument may include a user interface for controlling operation of the instrument and / or the integrated device. In some embodiments, the instrument may include a computer interface configured to connect with a computing device. In some embodiments, the instrument may include a processing device configured to analyze data received from one or more photodetectors of the integrated device and / or transmit control signals to the excitation source(s).
[0714]
[0323] According to some embodiments, the instrument that is configured to analyze samples based on luminescence emission characteristics may detect differences in luminescence lifetimes and / or intensities between different luminescent molecules, and / or differences between lifetimes and / or intensities of the same luminescent molecules in different environments. The inventors have recognized and appreciated that differences in luminescence emission lifetimes can be used to discern between the presence or absence of different luminescent molecules and / or to discern between different environments or conditions to which a luminescent molecule is subjected. Although analytic systems based on luminescence lifetime analysis may have certain benefits, the amount of information obtained by an analytic system and / or detection accuracy may be increased by allowing for additional detection techniques. For example, some embodiments of the systems may additionally be configured to discern one or more properties of a sample based on luminescence wavelength and / or luminescence intensity. In some embodiments, different numbers of fluorophores of the same type may be linked to different reagents in a sample, so that each reagent may be identified based on luminescence intensity.
[0715]
[0324] The inventors have recognized and appreciated that distinguishing biological or chemical samples based on fluorophore decay rates and / or fluorophore intensities may enable a simplification of the optical excitation and detection systems. For example, optical excitation may be performed with a single- R0708.70180WO00 / R0708.70180US01 81 / 122
[0716] #14644680vl wavelength source (e.g., a source producing one characteristic wavelength rather than multiple sources or a source operating at multiple different characteristic wavelengths). Additionally, wavelength discriminating optics and filters may not be needed in the detection system. Also, a single photodetector may be used for each sample well to detect emission from different fluorophores. The phrase “characteristic wavelength” or “wavelength” is used to refer to a central or predominant wavelength within a limited bandwidth of radiation (e.g., a central or peak wavelength within a 20 nm bandwidth output by a pulsed optical source). In some cases, “characteristic wavelength” or “wavelength” may be used to refer to a peak wavelength within a total bandwidth of radiation output by a source.
[0717]
[0325] The methods described herein may be implemented by a system. For example, in some embodiments, the system comprises at least one non-transitory computer-readable medium having instructions encoded thereon that, when executed, cause a processor to perform one or more of the methods described herein. In some embodiments, the system further comprises the processor. The system may comprise any of the components of the integrated device described herein.
[0718] Compounds and Additional Methods
[0719]
[0326] In another aspect, provided herein is a compound of Formula (I):
[0720] , S^L,, N(R1j2
[0721] □C I pN
[0722] RV" N
[0723] nH
[0724]
[0725] 0(I),
[0726] or a salt thereof, wherein:
[0727] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group;
[0728] Rcis -OH, an amino acid moiety, or a peptide; and
[0729] RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0730]
[0327] Compounds of Formula (I), or salts thereof, may act as pseudo-lysine sites, facilitating protein digestion, e.g., enzymatic protein digestion with Lys-C. Accordingly, in some embodiments, the compound of Formula (I), or salt thereof, is for use in protein sequencing.
[0731]
[0328] In another aspect, provided herein is a method of preparing a compound of Formula (I):
[0732] . S^L1^N(R1)2
[0733] RCy-VRN
[0734]
[0735] 0 H(I),
[0736] or a salt thereof, comprising contacting a compound of Formula (II):
[0737] R0708.70180WO00 / R0708.70180US01 82 / 122
[0738] #14644680vl SH
[0739]
[0740] or a salt thereof, with a compound of Formula (III):
[0741] Y^L1^N(R1)2
[0742]
[0743] (III),
[0744] or a salt thereof, to obtain the compound of Formula (I), or a salt thereof, wherein:
[0745] Y is a leaving group;
[0746] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group;
[0747] Rcis -OH, an amino acid moiety, or a peptide; and
[0748] RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0749]
[0329] In another aspect, provided herein is a method of protein digestion, comprising:
[0750] contacting a compound of Formula (II):
[0751]
[0752] or a salt thereof, with a compound of Formula (III):
[0753] Y^L1^N(R1)2
[0754]
[0755] (III),
[0756] or a salt thereof, to obtain a compound of Formula (I):
[0757] . S^L1^N(R1)2
[0758] RCy-VRN
[0759]
[0760] 0 H(I),
[0761] or a salt thereof; and
[0762] exposing the compound of Formula (I), or salt thereof, to a protein digestion agent, thereby forming a digested polypeptide sample, wherein:
[0763] Y is a leaving group;
[0764] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group;
[0765] Rcis -OH, an amino acid moiety, or a peptide; and
[0766] RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0767]
[0330] In some embodiments, the protein digestion agent is an enzymatic protein digestion agent. In some embodiments, the protein digestion agent is Lys-C.
[0768] R0708.70180WO00 / R0708.70180US01 83 / 122
[0769] #14644680vl
[0331] In another aspect, provided herein is a method of protein analysis, comprising:
[0770] contacting a compound of Formula (II):
[0771]
[0772] (II),
[0773] or a salt thereof, with a compound of Formula (III):
[0774] Y^L1^N(R1)2
[0775]
[0776] (III),
[0777] or a salt thereof, to obtain a compound of Formula (I):
[0778]
[0779] or a salt thereof;
[0780] exposing the compound of Formula (I), or salt thereof, to a protein digestion agent, thereby forming a digested polypeptide sample;
[0781] derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides;
[0782] conjugating the one or more derivatized polypeptides to an immobilization complex to form a polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
[0783] contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents and one or more amino acid recognizers;
[0784] monitoring a signal for signal pulses corresponding to interactions between one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides of the polypeptide sample; and
[0785] determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal;
[0786] wherein:
[0787] Y is a leaving group;
[0788] L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group;
[0789] Rcis -OH, an amino acid moiety, or a peptide; and
[0790] RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
[0791]
[0332] In some embodiments, the protein digestion agent is an enzymatic protein digestion agent. In some embodiments, the protein digestion agent is Lys-C.
[0792] R0708.70180WO00 / R0708.70180US01 84 / 122
[0793] #14644680vl
[0333] As generally described herein, Y is a leaving group.
[0794]
[0334] In some embodiments, Y is halo (e.g., -F, -Cl, -Br, -I), an activated substituted hydroxyl group (e.g., -OC(=O)SRaa, -OC(=O)Raa, -OCO2Raa, -OC(=O)N(Rbb)2, -OC(=NRbb)Raa, -OC(=NRbb)ORaa, -OC(=NRbb)N(Rbb)2, -OS(=O)Raa, -OSO2Raa, -OP(RCC)2, -OP(RCC)3, -OP(=O)2Raa, -OP(=O)(Raa)2, -OP(=O)(ORCC)2, -OP(=O)2N(Rbb)2, -OP(=O)(NRbb)2, wherein Raa, Rbb, and Rccare as defined herein), or a a sulfonic acid ester (e.g., toluenesulfonate (tosylate, -OTs), methanesulfonate (mesylate, -OMs), p-bromobenzenesulfonyloxy (brosylate, -OBs), -OS(=O)2(CF2)3CF3 (nonaflate, -ONf), trifluoromethanesulfonate (triflate, -OTf)).
[0795]
[0335] In some embodiments, Y is halo (e.g., -F, -Cl, -Br, -I). In some embodiments, Y is -Cl, -Br, or-I. In some embodiments, Y is -Cl. In some embodiments, Y is -Br. In some embodiments, Y is -I.
[0796]
[0336] As generally described herein, L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene.
[0797]
[0337] In some embodiments, L1is optionally substituted Ci-6 alkylene. In some embodiments, L1is substituted Ci-6 alkylene. In some embodiments, L1is unsubstituted Ci-6 alkylene. In some embodiments, L1is optionally substituted C1-3 alkylene. In some embodiments, L1is substituted C1-3 alkylene. In some embodiments, L1is unsubstituted C1-3 alkylene. In some embodiments, L1is methylene, ethylene, or n-propylene. In some embodiments, L1is ethylene.
[0798]
[0338] In some embodiments, L1is optionally substituted Ci-6 heteroalkylene. In some embodiments, L1is substituted C1.6 heteroalkylene. In some embodiments, L1is unsubstituted Ci-6 heteroalkylene. In some embodiments, L1is optionally substituted C1.3 heteroalkylene. In some embodiments, L1is substituted Ci.
[0799] 3 heteroalkylene. In some embodiments, L1is unsubstituted C1.3 heteroalkylene.
[0800]
[0339] As generally described herein, each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group.
[0801]
[0340] In some embodiments, at least one instance of R1is hydrogen. In some embodiments, each instance of R1is independently hydrogen.
[0802]
[0341] In some embodiments, at least one instance of R1is optionally substituted aliphatic. In some embodiments, at least one instance of R1is optionally substituted alkyl. In some embodiments, at least one instance of R1is a nitrogen protecting group.
[0803]
[0342] In some embodiments, Rcis -OH. In some embodiments, Rcis an amino acid moiety or a peptide. In some embodiments, Rcis an amino acid moiety. In some embodiments, Rcis a peptide. In some embodiments, RNis hydrogen. In some embodiments, RNis a nitrogen protecting group. In some embodiments, RNis an amino acid moiety or a peptide. In some embodiments, RNis an amino acid moiety. In some embodiments, RNis a peptide. In some embodiments, Rcis -OH, and RNis an amino acid moiety or a peptide. In some embodiments, Rcis -OH, and RNis an amino acid moiety. In some embodiments, Rcis -OH, and RNis a peptide. In some embodiments, Rcis an amino acid moiety or a peptide, and RNis hydrogen. In some embodiments, Rcis an amino acid moiety, and RNis hydrogen. In some embodiments, Rcis a peptide, and RNis hydrogen. In some embodiments, Rcis an amino acid moiety or a peptide; and RNis an amino acid moiety or a peptide. In some embodiments, Rcis an amino R0708.70180WO00 / R0708.70180US01 85 / 122
[0804] #14644680vl acid moiety; and RNis an amino acid moiety. In some embodiments, Rcis an amino acid moiety; and RNis a peptide. In some embodiments, Rcis a peptide; and RNis an amino acid moiety. In some embodiments, Rcis a peptide; and RNis a peptide.
[0805]
[0343] In some embodiments, RNis a peptide comprising fewer than 5% lysine residues. In some embodiments, Rcis a peptide comprising fewer than 5% lysine residues. In some embodiments, RNis a peptide comprising fewer than 5% lysine residues and / or Rcis a peptide comprising fewer than 5% lysine residues. In some embodiments, RNis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue. In some embodiments, Rcis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue. In some embodiments, RNis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue and / or Rcis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue. In some embodiments, RNis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue. In some embodiments, Rcis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue. In some embodiments, RNis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue and / or Rcis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue.
[0806]
[0344] In some embodiments, the compound of Formula (I) is of Formula (I-a):
[0807]
[0808] (I-a),
[0809] or a salt thereof.
[0810]
[0345] In some embodiments of Formula (I-a), Rcis an amino acid moiety or a peptide. In some embodiments of Formula (I-a), RNis an amino acid moiety or a peptide. In some embodiments of Formula (I-a), Rcis an amino acid moiety or a peptide; and RNis an amino acid moiety or a peptide.
[0811]
[0346] In some embodiments, the compound of Formula (I) is of Formula (I-b):
[0812]
[0813] (I-b),
[0814] or a salt thereof.
[0815]
[0347] In some embodiments of Formula (I-b), Rcis an amino acid moiety or a peptide. In some embodiments of Formula (I-b), RNis an amino acid moiety or a peptide. In some embodiments of Formula (I-b), Rcis an amino acid moiety or a peptide; and RNis an amino acid moiety or a peptide.
[0816]
[0348] In some embodiments of Formula (II), Rcis an amino acid moiety or a peptide. In some embodiments of Formula (II), RNis an amino acid moiety or a peptide. In some embodiments of Formula (II), Rcis an amino acid moiety or a peptide; and RNis an amino acid moiety or a peptide.
[0817] R0708.70180WO00 / R0708.70180US01 86 / 122
[0818] #14644680vl
[0349] In some embodiments, the compound of Formula (III) is of Formula (III-a):
[0819] Br^ 1-N(R1)2
[0820]
[0821] (III-a),
[0822] or a salt thereof.
[0823]
[0350] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-a), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene or unsubstituted Ci-6 heteroalkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III -a), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-a), or a salt thereof, wherein L1is unsubstituted C1-3 alkylene.
[0824]
[0351] In some embodiments, the compound of Formula (III) is of Formula (III-b):
[0825] L1-N(R1)2
[0826]
[0827] or a salt thereof.
[0828]
[0352] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-b), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene or unsubstituted Ci-6 heteroalkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-b), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-b), or a salt thereof, wherein L1is unsubstituted C1-3 alkylene.
[0829]
[0353] In some embodiments, the compound of Formula (III) is of Formula (III-c):
[0830] Y'"'"~'N(R1)2 (iii.c)
[0831]
[0832] or a salt thereof.
[0833]
[0354] In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-c), or a salt thereof, wherein Y is halo. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-c), or a salt thereof, wherein Y is -Br. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-c), or a salt thereof, wherein Y is -I.
[0834]
[0355] In some embodiments, the compound of Formula (III) is of Formula (Ill-d):
[0835] Y^L1 / NH2
[0836]
[0837] (III-d),
[0838] or a salt thereof.
[0839]
[0356] In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-d), or a salt thereof, wherein Y is halo. In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-d), or a salt thereof, wherein Y is -Br. In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-d), or a salt thereof, wherein Y is -I. In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-d), or a salt thereof, wherein L1is unsubstituted Ci-6 alkylene or unsubstituted Ci-6 heteroalkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-d), or a salt R0708.70180WO00 / R0708.70180US01 87 / 122
[0840] #14644680vl thereof, wherein L1is unsubstituted Ci-6 alkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-d), or a salt thereof, wherein L1is unsubstituted C1.3 alkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (III-d), or a salt thereof, wherein Y is halo; and L1is unsubstituted C1-6 alkylene or unsubstituted C1-6 heteroalkylene. In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-d), or a salt thereof, wherein Y is -Br or -I; and L1is unsubstituted C1-6 alkylene.
[0841]
[0357] In some embodiments, the compound of Formula (III) is of Formula (Ill-e):
[0842]
[0843] or a salt thereof.
[0844]
[0358] In some embodiments, the amino acid side chain capping agent is a compound of Formula (Ill-e), or a salt thereof, wherein Y is halo.
[0845]
[0359] In some embodiments, the compound of Formula (III) is of formula:
[0846] B
[0847]
[0848] r'^'NH2 or1NH2
[0849] or a salt thereof.
[0850] EXAMPLES
[0851]
[0360] In order that the present disclosure may be more fully understood, the following examples are set forth. The synthetic and biological examples described in this application are offered to illustrate the methods and compounds provided herein and are not to be construed in any way as limiting in their scope.
[0852] Example 1: Exemplary Sample Preparation Protocol.
[0853]
[0361] This Example describes an exemplary sample preparation protocol (e.g., method of preparing a polypeptide sample).
[0854]
[0362] Library preparation begins by reducing, alkylating, and enzymatically digesting intact proteins containing lysine residues into peptides. C-terminal lysine residues on each peptide are then activated with azido groups and then conjugated with a macromolecule linker which enables immobilization on the Sequencing Chip.
[0855] Materials
[0856] Table 4. Components Used in Library Preparation
[0857] Component Storage Component Storage Temperature Temperature Lys-C (Endopeptidase Lys-C) -20°C K2CO3 (Potassium 4°C Carbonate)
[0858] Lys-C Buffer -20°C CuSO4(Copper(II) 4°C
[0859] Sulfate)
[0860] TCEP (Tris(2- -20°C EDTA 4°C
[0861] carboxyethyl)phosphine
[0862]
[0863] hydrochloride)
[0864] R0708.70180WO00 / R0708.70180US01 88 / 122
[0865] #14644680vl CAA (Chloroacetamide) -20°C Acetic Acid 4°C ISA (Imidazole- 1 -sulfonyl Azide) -20°C Polyamine Beads 4°C
[0866] K-Linker Complex -20°C Sample Buffer 4°C
[0867] CTAB (Cetyltrimethyl ammonium -20°C Empty Spin Column 4°C
[0868]
[0869] bromide)
[0870]
[0363] All reagents should be discarded after 4 freeze-thaw cycles. Frozen reagents should be thawed on ice for 15-20 min before use, unless otherwise specified. Reagents should be kept on ice until use and returned to freeze immediately after use. Reagents should be briefly vortexed to homogenize before use, except Lys-C. All reagent tubes should be briefly spun in a minicentrifuge to collect reagents at the bottom of the tube and to prevent any bubbles or liquid on the side walls. If using a microcentrifuge, it should not exceed 2,680 x g. Remaining rehydrated Lys-C should be immediately stored at -80°C as 4 pL aliquots in 0.5-1.5 mL tubes to avoid freeze-thaw cycles.
[0871]
[0364] Sample considerations. Sample preparation and sequencing are optimized for protein samples with the following attributes: protein molecular weight 10-80 kDa; volume 100 pL, concentration 5 pM, and up to 10 proteins in a mixture. Protein concentrations of 5 pM in 100 pL can be used for this protocol. The sample should be diluted with Sample Buffer if concentration is higher than 5 pM. It is not recommended to use less than 5 pM of protein.
[0872] Digestion
[0873]
[0365] Protein digestion begins with cysteine reduction using TCEP (Tris(2-carboxyethyl)phosphine hydrochloride) to disrupt the disulfide bridges that have formed between cysteine amino acid residues. Next, exposed thiol groups on cysteine side chains are alkylated with CAA (Chloroacetamide) to prevent the disulfide bridges from reforming. Finally, the endopeptidase Lys-C is added to digest the protein into peptides with a lysine (K) residue at the C-terminus.
[0874]
[0366] The following reagents are briefly vortexed and spun before use: TCEP (thawed on ice), CAA (thawed on ice), and Lys-C buffer (thawed on ice). Lys-C is rehydrated in 40 pL of Lys-C Buffer and pipette mixed 10-20 times to ensure Lys-C is fully dissolved. Aliquots of 4 pL Lys-C are prepared in 0.5- 1.5 mL Protein Lo-Bind tubes for any unused reactions, and are immediately stored at -80°C. Unused Lys-C is discarded after a single use of each aliquot to avoid freeze-thaws.
[0875]
[0367] Reduction. 2 pL TCEP is added to the Sample tube containing 5 pM of protein and mixed well by vortexing and briefly spinning down in a minicentrifuge. The Sample Tube is incubated in a heat block at 37°C for 30 minutes.
[0876]
[0368] Alkylation. 2 pL CAA is added to the Sample tube and mixed well by vortexing and briefly spinning down in a ...
Claims
1. CLAIMS2.What is claimed is:
1. A method of preparing a polypeptide sample from a protein for analysis, comprising:4.exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample;5.derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides; and conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
2. The method of claim 1 further comprising denaturing the protein before exposing the protein to the protein digestion agent.
3. The method of claim 1 or 2 further comprising exposing the protein to a reducing agent before exposing the protein to the protein digestion agent.
4. The method of any one of claims 1-3 further comprising exposing the protein to an amino acid side chain capping agent before exposing the protein to the protein digestion agent.
5. The method of any one of claims 1-4 further comprising exposing the protein to a surfactant before exposing the protein to the protein digestion agent.
6. A method of preparing a polypeptide sample from a protein for analysis, comprising:11.exposing the protein to a reducing agent;12.exposing the protein to an amino acid side chain capping agent;13.exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample.
7. The method of claim 6 further comprising derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides.15.R0708.70180WO00 / R0708.70180US01 109 / 12216.#14644680vl 8. The method of claim 6 or 7 further comprising conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
9. A method of preparing a polypeptide sample from a protein for analysis, comprising:18.exposing the protein to a reducing agent;19.exposing the protein to an amino acid side chain capping agent;20.exposing the protein to a protein digestion agent, thereby forming a digested polypeptide sample;21.derivatizing one or more polypeptides of the digested polypeptide sample to form a derivatized polypeptide sample comprising one or more derivatized polypeptides; and conjugating the one or more derivatized polypeptides to an immobilization complex to form the polypeptide sample, wherein the polypeptide sample comprises one or more immobilization complex-conjugated polypeptides.
10. The method of any one of claims 3-9 further comprising exposing the protein to a surfactant before exposing the protein to the reducing agent.
11. The method of claim 10, wherein the surfactant is selected from RapiGest (e.g., RapiGest SF (Waters)), sodium dodecyl sulfate (SDS), sodium deoxycholate, Sarkosyl, n-octyl glucoside (OG), Triton X-100 (TX-100), Triton X-114 (TX-114), 3-[(3-cholamidopropyl)dimethylammonio]-l-propane sulfonate (CHAPS), 3 - [(3 -cholamidopropyl)dimethylammonio] -2 -hydroxy- 1 -propane sulfonate (CHAPSO), dodecyl maltoside (DDM), ProteaseMAX (Promega), Tergitol-type NP-40 (NP-40), polysorbate 20 (e.g., Tween 20), polysorbate 80 (e.g., Tween 80), Brij 35, and Brij 58.
12. The method of any one of claims 3-11 further comprising denaturing the protein before exposing the protein to the reducing agent.
13. The method of claim 12, wherein the method does not comprise a buffer exchange step before denaturing the protein.
14. The method of any one of claims 3-13, wherein the method does not comprise a buffer exchange step before exposing the protein to the reducing agent.
15. The method of any one of claims 3-14 further comprising exposing the protein to a surfactant before exposing the protein to the protein digestion agent.28.R0708.70180WO00 / R0708.70180US01 110 / 12229.#14644680vl 16. The method of any one of claims 3-14, wherein the reducing agent reduces an amino acid side chain of the protein to form a reduced amino acid side chain of the protein.
17. The method of any one of claims 3-15, wherein the reducing agent comprises tris(2-carboxyethyl)phosphine (TCEP).
18. The method of any one of claims 3-17, wherein the reducing agent comprises tris(hydroxypropyl)phosphine (THP).
19. The method of any one of claims 4-18, wherein the amino acid side chain capping agent forms a covalent bond with the reduced amino acid side chain to form a capped amino acid side chain of the protein.
20. The method of any one of claims 4-19, wherein the amino acid side chain capping agent comprises a cysteine alkylation agent.
21. The method of any one of claims 4-20, wherein the amino acid side chain capping agent comprises iodoacetamide (IAA) and / or chloroacetamide (CAA).
22. The method of any one of claims 4-20, wherein the amino acid side chain capping agent is a compound of Formula (III):36.Y^L,_N(R1)238.
39. (III),40.or a salt thereof, wherein:41.Y is a leaving group;42.L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; and each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group.
23. The method of claim 22, wherein Y is halogen.
24. The method of claim 22 or 23, wherein the amino acid side chain capping agent is a compound of Formula (III -a) or (III-b):45.Br^L,_N(R’)2^L1--N(R1)247.
48. (Ill-a) or (III-b), or a salt thereof.49.R0708.70180WO00 / R0708.70180US01 111 / 12250.#14644680vl 25. The method of any one of claims 22-24, wherein L1is unsubstituted Ci-6 alkylene.
26. The method of any one of claims 22, 23, and 25, wherein the amino acid side chain capping agent is a compound of Formula (III-c):
53.
54. N(R1)2 (III_C),55.or a salt thereof.
27. The method of any one of claims 22-26, wherein at least one instance of R1is hydrogen.
28. The method of any one of claims 22, 23, 25, and 27, wherein the amino acid side chain capping agent is a compound of Formula (Ill-d):58.^NH260. 61.L(III-d),62.or a salt thereof.
29. The method of any one of claims 22, 23, and 25-28, wherein the amino acid side chain capping agent is a compound of Formula (Ill-e):
65.
66. NH2(Ill-e),67.or a salt thereof.
30. The method of any one of claims 22-29, wherein the amino acid side chain capping agent is a compound of formula:
70. 72.or a salt thereof.
31. The method of any one of claims 4-30, wherein the amino acid side chain capping agent comprises bromoethylamine (BEA).
32. The method of any one of claims 4-31, wherein the amino acid side chain capping agent comprises iodoethylamine (IEA).
33. The method of any one of claims 1-32, wherein the protein comprises fewer than 5% lysine residues.76.R0708.70180WO00 / R0708.70180US01 112 / 12277.#14644680vl 34. The method of any one of claims 1-33, wherein the protein comprises a continuous sequence of amino acids that does not comprise a lysine residue.
35. The method of claim 34, wherein the continuous sequence of amino acids that does not comprise a lysine residue comprises at least 150 amino acid residues.
36. The method of any one of claims 1-35, wherein the protein digestion agent induces proteolysis of the protein to form one or more capped polypeptides, thereby forming the digested polypeptide sample.
37. The method of any one of claims 19-36, wherein the protein comprises the capped amino acid side chain.
38. The method of claim 36 or 37, wherein one or more capped polypeptides are of Formula (I):82.. S^LI„N(R')283.pC 1 f? N84.AH86. 87.0(i),88.or a salt thereof, wherein:89.L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group;90.Rcis -OH, an amino acid moiety, or a peptide; and91.RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
39. The method of any one of claims 36-38, wherein one or more capped polypeptides are of Formula (I-b):
94. 95.0(I-b),96.or a salt thereof.
40. The method of any one of claims 1-39, wherein the protein digestion agent comprises a protease.
41. The method of claim 40, wherein the protease comprises trypsin, Lys-C, Asp-N, and / or Glu-C.99.R0708.70180WO00 / R0708.70180US01 113 / 122100.#14644680vl 42. The method of claim 40 or 41, wherein the protease comprises Lys-C.
43. The method of any one of claims 1-42 further comprising binding the protein to a solid substrate before exposing the protein to the protein digestion agent.
44. The method of any one of claims 36-43 further comprising binding the protein or the one or more capped polypeptides to a solid substrate after exposing the protein to the protein digestion agent.
45. The method of claim 44, wherein the solid substrate comprises a bead.
46. The method of claim 45, wherein the bead is paramagnetic.
47. The method of any one of claims 43-46, wherein the binding comprises a noncovalent interaction between functional groups on a surface of the solid substrate and the protein or the one or more capped polypeptides.
48. The method of any one of claims 1-47, comprising an incubation step, wherein:107.the reducing agent reduces an amino acid side chain of the protein to form a reduced amino acid side chain of the protein;108.the amino acid side chain capping agent forms a covalent bond with the reduced amino acid side chain to form a capped amino acid side chain of the protein; and109.the protein digestion agent induces proteolysis of the protein to form one or more capped polypeptides, thereby forming the digested polypeptide sample.
49. The method of claim 48, wherein the incubating step comprises maintaining the protein and / or the one or more capped polypeptides at a temperature greater than or equal to 20°C, greater than or equal to 25 °C, greater than or equal to 30°C, greater than or equal to 35 °C, or greater than or equal to 37°C.
50. The method of any one of claims 1-49 further comprising purifying the digested polypeptide sample to form a purified digested polypeptide sample.
51. The method of claim 50, wherein purifying the digested polypeptide sample to form the digested polypeptide sample comprises removing at least some of any salts, detergents, or chaotropes present in the digested polypeptide sample.
52. The method of claim 50 or 51, wherein purifying the digested polypeptide sample to form the digested polypeptide sample comprises exposure of the digested polypeptide sample to a Cl 8 matrix.114.R0708.70180WO00 / R0708.70180US01 114 / 122115.#14644680vl 53. The method of any one of claims 1-52, wherein the derivatizing comprises derivatizing an amino acid side chain of the one or more capped polypeptides using a derivatization agent to form an unquenched mixture comprising one or more derivatized polypeptides.
54. The method of claim 53, wherein the derivatization agent comprises an azide transfer agent.
55. The method of claim 54, wherein the azide transfer agent comprises imidazole- 1 -sulfonyl azide (ISA).
56. The method of any one of claims 53-55, wherein the derivatizing comprises exposing the one or more capped polypeptides to the derivatization agent and one or more derivatization reagents.
57. The method of claim 56, wherein the one or more derivatization reagents comprises a pH adjusting reagent.
58. The method of claim 57, wherein the pH adjusting reagent is potassium carbonate (K2CO3).
59. The method of any one of claims 56-58, wherein the one or more derivatization reagents comprises a catalyst for a derivatization reaction between the amino acid side chain of the one or more capped polypeptides and the derivatization agent.
60. The method of any one of claims 56-59, wherein the one or more derivatization reagents comprises a source of Cu2+.
61. The method of claim 60, wherein the source of Cu2+is copper (II) sulfate (CUSO4).
62. The method of any one of claims 56-61, wherein the one or more derivatization reagents comprises a source of Ni2+.
63. The method of claim 62, wherein the source of Ni2+is nickel (II) acetate (Ni(OAc)2).
64. The method of any one of claims 36-63, wherein the derivatizing comprises exposing the one or more capped polypeptides to, in order: a pH adjusting reagent, one or more derivatization reagents, and the derivatization agent.127.R0708.70180WO00 / R0708.70180US01 115 / 122128.#14644680vl 65. The method of any one of claims 36-64, wherein the derivatizing comprises exposing the one or more capped polypeptides to, in order: a pH adjusting reagent, a source of Cu2+, and an azide transfer agent.
66. The method of any one of claims 36-65, wherein the derivatizing comprises exposing the one or more capped polypeptides to, in order: potassium carbonate (K2CO3), copper (II) sulfate (CuSO4), and imidazole -1 -sulfonyl azide (ISA).
67. The method of any one of claims 36-63, wherein the derivatizing comprises exposing the one or more capped polypeptides to, in order: a pH adjusting reagent, a source of Ni2+, and an azide transfer agent.
68. The method of any one of claims 36-63 and 67, wherein the derivatizing comprises exposing the one or more capped polypeptides to, in order: potassium carbonate (K2CO3), nickel (II) acetate (Ni(0Ac)2), and imidazole -1 -sulfonyl azide (ISA).
69. The method of any one of claims 1-68, wherein one or more capped polypeptides are of Formula (I):133.. S^L1^N(R1)2134.pC I f? N136. 137.0(I),138.or a salt thereof, and one or more derivatized polypeptides are of Formula (IV):
139. / SX1-N3140.RVVRN142.
143. sH144.or a salt thereof, wherein:145.L1is optionally substituted C1-6 alkylene or optionally substituted C1-6 heteroalkylene; each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group;146.Rcis -OH, an amino acid moiety, or a peptide; and147.RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
70. The method of claim 69, wherein one or more capped polypeptides are of formula:149.R0708.70180WO00 / R0708.70180US01 116 / 122150.#14644680vl151. 153.or a salt thereof, and one or more derivatized polypeptides are of formula:
155.
156. (IV-a),157.or a salt thereof.
71. The method of any one of claims 53-70, wherein the unquenched mixture further comprises excess derivatization agent.
72. The method of claim 71 further comprising quenching the unquenched mixture to form a quenched mixture by removing at least some of the excess derivatization agent.
73. The method of claim 71 or 72, wherein the quenching comprises reacting the at least some of the excess derivatization agent with functional groups on a surface of a solid substrate.
74. The method of claim 73, wherein the solid substrate comprises a bead.
75. The method of claim 73 or 74, wherein the functional groups of the solid substrate comprise amine groups.
76. The method of any one of claims 72-75, wherein the quenching comprises exposure of the unquenched mixture to a C 18 matrix.
77. The method of any one of claims 1-5 and 7-76 further comprising purifying the derivatized polypeptide sample to form a purified derivatized polypeptide sample.
78. The method of claim 77, wherein purifying the derivatized polypeptide sample to form the purified derivatized polypeptide sample comprises removing at least some of any remaining non-derivatized polypeptides of the derivatized polypeptide sample.166.R0708.70180WO00 / R0708.70180US01 117 / 122167.#14644680vl 79. The method of claim 77 or 78, wherein purifying the derivatized polypeptide sample to form the purified derivatized polypeptide sample comprises passing the derivatized polypeptide sample through a size exclusion medium.
80. The method of any one of claims 1-5 and 8-76, wherein conjugating the one or more derivatized polypeptides to the immobilization complex comprises exposing the one or more derivatized polypeptides to cetrimonium bromide (CTAB).
81. The method of any one of claims 1-5, 8-76, and 80, wherein conjugating the one or more derivatized polypeptides to the immobilization complex comprises exposing the one or more derivatized polypeptides to ethylenediaminetetraacetic acid (EDTA).
82. The method of any one of claims 1-5, 8-76, 80, and 81, wherein conjugating the one or more derivatized polypeptides to the immobilization complex comprises a click chemistry reaction.
83. The method of any one of claims 1-5, 8-76, and 80-82, wherein the immobilization complex is a streptavidin-bearing immobilization complex.
84. The method of any one of claims 1-5, 8-76, and 80-83, wherein conjugating the one or more derivatized polypeptides to the immobilization complex comprises maintaining the one or more derivatized polypeptides and / or the one or more immobilization complex-conjugated polypeptides at a temperature greater than or equal to 20°C, greater than or equal to 25 °C, greater than or equal to 30°C, greater than or equal to 35°C, or greater than or equal to 37°C.
85. The method of any one of claims 1-84 further comprising:174.contacting the polypeptide sample with a reaction mixture comprising one or more cleaving agents and one or more amino acid recognizers;175.monitoring a signal for signal pulses corresponding to interactions between the one or more amino acid recognizers and the one or more immobilization complex-conjugated polypeptides of the polypeptide sample; and176.determining at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides based on a characteristic pattern in the signal.
86. The method of claim 85, wherein the monitoring comprises detecting a series of signal pulses while the one or more immobilization complex-conjugated polypeptides is being degraded, wherein a characteristic pattern in the series of signal pulses is indicative of the at least one chemical characteristic of the one or more immobilization complex-conjugated polypeptides.178.R0708.70180WO00 / R0708.70180US01 118 / 122179.#14644680vl 87. The method of claim 85 or 86, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the one or more immobilization complex-conjugated polypeptides as a naturally occurring amino acid, an unnatural amino acid, or a modified variant thereof.
88. The method of any one of claims 85-87, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the one or more immobilization complex-conjugated polypeptides as having a side chain that is negatively charged, positively charged, uncharged, polar, non-polar, hydrophobic, aromatic, or a combination thereof.
89. The method of any one of claims 85-88, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the one or more immobilization complex-conjugated polypeptides as one type selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, selenocysteine, serine, threonine, tryptophan, tyrosine, and valine.
90. The method of any one of claims 85-89, wherein determining the at least one chemical characteristic comprises identifying at least one amino acid in the one or more immobilization complex-conjugated polypeptides as having a post-translational modification.
91. The method of claim 90, wherein the post-translational modification is selected from acetylation, ADP-ribosylation, caspase cleavage, citrullination, formylation, N-linked glycosylation, O-linked glycosylation, hydroxylation, methylation, myristoylation, neddylation, nitration, chlorination, oxidation / reduction, carbonylation, palmitoylation, phosphorylation, prenylation, S-nitrosylation, sulfation, glycation, sumoylation, and ubiquitination.
92. A method of preparing a compound of Formula (I):185., S^L1„N(R1)2186.RVVRN188.
189. OH(I),190.or a salt thereof, comprising contacting a compound of Formula (II):191.. SH192.DC I f? N193.nH195. 196.0(II),197.or a salt thereof, with a compound of Formula (III):198.R0708.70180WO00 / R0708.70180US01 119 / 122199.#14644680vl Y^L1_N(R1)2201.
202. (Ill),203.or a salt thereof, to obtain the compound of Formula (VIII), or a salt thereof, wherein:204.Y is a leaving group;205.L1is optionally substituted Ci-6 alkylene or optionally substituted Ci-6 heteroalkylene; each instance of R1is independently hydrogen, optionally substituted aliphatic, or a nitrogen protecting group;206.Rcis -OH, an amino acid moiety, or a peptide; and207.RNis hydrogen, a nitrogen protecting group, an amino acid moiety, or a peptide.
93. The method of claim 92, wherein Y is halogen.
94. The method of claim 92 or 93, wherein the compound of Formula (III) is of formula:210.Br^ ^N(R1)2^L1-N(R1)2212.
213. L or214.or a salt thereof.
95. The method of any one of claims 92-94, wherein L1is unsubstituted Ci-6 alkylene.
96. The method of any one of claims 92, 93, and 95, wherein the compound of Formula (III) is of formula:
218. 219.Y'X / ^N(R1)2220.or a salt thereof.
97. The method of any one of claims 92-96, wherein at least one instance of R1is hydrogen.
98. The method of any one of claims 92, 93, 95, and 97, wherein the compound of Formula (III) is of formula:223.Y^L1^NH2224.or a salt thereof.
99. The method of any one of claims 92, 93, and 95-98, wherein the compound of Formula (III) is of formula:
227. 229.or a salt thereof.230.R0708.70180WO00 / R0708.70180US01 120 / 122231.#14644680vl 100. The method of any one of claims 92-99, wherein the compound of Formula (III) is of formula:
233. 235.or a salt thereof.
101. The method of any one of claims 92-100, wherein the compound of Formula (I) is of formula:
238. 240.or a salt thereof.
102. The method of any one of claims 92-101, wherein:242.Rcis a peptide comprising fewer than 5% lysine residues; and / or243.RNis a peptide comprising fewer than 5% lysine residues.
103. The method of any one of claims 92-102, wherein:245.Rcis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue; and / or246.RNis a peptide comprising a continuous sequence of amino acids that does not comprise a lysine residue.
104. The method of any one of claims 92-103, wherein:248.Rcis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue; and / or249.RNis a peptide comprising a continuous sequence of at least 150 amino acids that does not comprise a lysine residue.250.R0708.70180WO00 / R0708.70180US01 121 / 122251.#14644680vl