Peptide sequencer

By combining fluorescence lifetime imaging and Edman degradation, the problems of low sensitivity and high cost of peptide sequencing in existing technologies are solved, and high-confidence peptide sequence prediction and accurate sequencing are achieved, which is suitable for multiple amino acid reads.

CN120752535APending Publication Date: 2025-10-03OREGON HEALTH & SCI UNIV
View PDF 32 Cites 0 Cited by

Patent Information

Application Number
CN202480014566.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-18
Filing Date
2024-01-04
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing de novo protein sequencing technologies have the disadvantages of low sensitivity, high cost, limited dynamic range, and ambiguity in the assignment of peptide or peptide fragment sequences, especially for peptides or peptide fragments with the same mass-to-charge ratio, and are unable to interpret the sensitivity of post-translational modifications with high fidelity.

Method used

Fluorescence lifetime imaging (FLIM) was used to measure the fluorescence lifetime of single-molecule fluorophores, combined with cyclic Edman degradation chemistry. The N-terminal amino acid of the peptide was functionalized by a universal docking DNA oligonucleotide functionalized with phenyl isothiocyanate (PITC). Single-molecule fluorescence lifetime measurements were performed using a library of fluorophores and imager chains, and the peptide sequence was reconstructed using a machine learning algorithm.

Benefits of technology

It achieves high-confidence peptide sequence prediction, avoids the difficulty of developing amino acid-specific binders, enables efficient and economical peptide sequencing, is suitable for multiple amino acid reads, and improves sequencing accuracy and throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752535A_ABST
    Figure CN120752535A_ABST
Patent Text Reader

Abstract

The present invention describes a peptide sequencing method in which the non-linked terminus of a surface end immobilized peptide is functionalized with a universal docking chain (DS) DNA oligonucleotide. Each sequential terminal amino acid of the peptide is characterized using a library of signal molecules, such as fluorophores, each of the signal molecules being conjugated to an imaging strand (IS) oligonucleotide complementary to the DS oligonucleotide. Each amino acid is identified using computer-aided analysis based on the difference of the measured signal caused by the proximity of the signal molecule to each terminal amino acid.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of earlier applications of U.S. Provisional Application No. 63 / 478,661, filed on January 5, 2023, and U.S. Provisional Application No. 63 / 485,904, filed on February 18, 2023, both of which are incorporated herein by reference in their entireties.

[0003] Incorporation by Reference into the Sequence Listing

[0004] A computer-readable text file with a file size of 9,295 bytes and titled "O046-0072PCT_ST26.xml," created on or about January 3, 2024, contains the sequence listing of the present application and is hereby incorporated by reference in its entirety. Technical Field

[0005] The present disclosure relates generally to de novo protein and peptide sequencing. Further, the present disclosure relates to methods for identifying the terminal amino acid of a peptide using its interaction with a signaling molecule. Background Art

[0006] The flow of genetic information in a cell can be described by three basic transformations: DNA to DNA (replication), DNA to RNA (transcription), and RNA to protein (translation). Functionally, DNA and RNA determine the structure of proteins, but the ultimate function of a cell (or dysfunction in disease) is determined by the proteins expressed in the cell. Given how DNA / RNA sequencing platforms have enabled insights into general biology, including cancer and other diseases and conditions, knowing the exact protein profile of a sample may help to further this understanding. Therefore, de novo protein sequencing platforms are needed in disease research to discover differentiating factors in disease progression and / or novel disease biomarkers.

[0007] Currently, untargeted proteomics relies primarily on digesting intact proteins into small peptides and reading their sequences using liquid chromatography-mass spectrometry (LC-MS). However, LC-MS has several limitations: low sensitivity, high cost, limited dynamic range, and ambiguity in assigning sequences to peptides or peptide fragments, especially those with identical mass-to-charge ratios.

[0008] De novo peptide sequencing technology is being developed. Although the final readout method is different, various companies are using N-terminal amino acid (NTAA) specific binders to read out digested proteins (peptides). As discussed in US2021 / 0396762, progress has been made in the field of digital analysis of proteins by end sequencing (DAPES). For example, in one method, a modified Edman degradation step is used to directly sequence surface-bound peptides, followed by detection, such as with a labeled antibody (WO2010 / 065531). A modification of DAPES is disclosed in which single-molecule sequencing of peptides is achieved by contacting the peptide with a fluorescently labeled N-terminal amino acid binding protein (NAAB), detecting the fluorescence of the NAAB bound to the amino acid, identifying the N-terminal amino acid based on the detected fluorescence, removing the NAAB from the peptide, and repeating with NAAB bound to a different N-terminal amino acid (WO2014 / 0273004). The N-terminal amino acid is cleaved from the polypeptide by Edman degradation, and the procedure is repeated for each newly exposed N-terminal amino acid. Other teachings use labeled N-terminal amino acid complexes for sequencing polypeptides, followed by cycles of Edman degradation or aminopeptidase cleavage (WO2010 / 065322); or methods for peptide analysis using multi-component detectors, for example, comprising a first detector and a second detector that are capable of producing a detectable signal when brought into proximity (US2021 / 0396762). Other techniques for characterizing proteins include those disclosed in US2003 / 0138831, US2014 / 0349860, and WO2013 / 112745.

[0009] While some of these approaches may be promising, each faces significant challenges. Developing specific binders for each amino acid is an extremely challenging task. Using covalent dyes attached to amino acids only allows for the reading of a small subset of amino acids in a peptide sequence. One peptide sequencing approach under development uses nanopores for the detection of amino acids, but suffers from scalability and reliability issues. Whole protein profiling has also been attempted, but the approach relies on libraries of affinity reagents that do not yet exist. Therefore, all previously available de novo peptide sequencing technologies are in significant need of additional development, or even the discovery of biochemical and / or technical methods that can compete with LC-MS. Summary of the Invention

[0010] The advent of next-generation sequencing has dramatically accelerated clinical and translational discovery, but genomic sequencing incompletely characterizes the proteomic landscape of biological systems. Similarly, future de novo protein sequencing methods promise to revolutionize nearly all areas of biological research and medicine. The recent global push for next-generation protein sequencing has yielded powerful mass spectrometry- and fluorescence-based methods; however, these technologies cannot fully sequence proteins de novo, have low throughput, and are unable to account for sensitivity to post-translational modifications with high fidelity.

[0011] Fluorescence lifetime imaging (FLIM) measures the single-molecule fluorescence of individual fluorophores to determine the time spent in the excited state before photon relaxation and emission. For some fluorophores, the excited-state lifetime can be extremely sensitive to local and global environmental changes. In one embodiment, a single-molecule peptide sequencing method is described herein that uses cyclic Edman degradation-based chemistry with an optical readout of fluorescence lifetime measurements.

[0012] This paper describes a new peptide sequencing method that can be implemented with readily available reagents and equipment. A representative design workflow is shown in Figures 1A-1G It begins with enzymatic digestion of proteins in the sample to obtain unmodified peptides, just like the LC-MS based workflow. The peptide is then attached to a solid phase substrate from its C-terminus. The N-terminal amino acid of the immobilized peptide is functionalized with a universal docking DNA oligonucleotide conjugated to phenyl isothiocyanate (PITC) or a functionalizing equivalent. Here, PITC serves two purposes: (i) to mediate the conjugation of the docking oligonucleotide; and (ii) to implement the cleavage of the N-terminal amino acid (NTAA) when the next amino acid needs to be read out using Edman degradation.

[0013] A library of fluorophores conjugated to imaging strand (IS) oligonucleotides complementary to the docking strand (DS) oligonucleotides is used to determine the N-terminal amino acid. The readout begins with the introduction of the first fluorophore (conjugated to the IS) for docking (to the DS) and a single molecule fluorescence lifetime measurement is performed for each peptide. The fluorescence lifetime of the fluorophore will be different from its free form because it interacts with the NTAA. In addition, designed sequence differences in the IS can (optionally) further tune the fluorescence lifetime readout. The single molecule fluorescence lifetime measurement can be repeated for one or more additional combinations of IS and fluorophores in the library. The N-terminal amino acid is then cleaved ("read"), for example using Edman degradation, making the construct ready for the next cycle - analysis of the next amino acid now at the N-terminus of the peptide. Therefore, each analysis cycle begins with DS oligonucleotide conjugation and ends with Edman degradation.

[0014] Reading out a series of amino acids yields a set of fluorescence lifetime measurements for each amino acid, corresponding to each different fluorophore and IS, for as many amino acids as are used in the analysis. This data can be fed into a machine learning-based prediction algorithm that generates the sequence of all peptides.

[0015] The embodiments of the provided sequence methods have numerous benefits. For example, in the provided embodiments, there is no need to develop new amino acid specific binders, which are generally difficult to identify (or generate) and verify. All measurements can be completed using standard fluorophores. The amino acids in each embodiment can be read more than once with different fluorophores and / or DNA imaging strands. This makes it possible to predict sequences with high confidence based on machine learning, which uses multiple different data inputs for each amino acid position. Different imaging oligonucleotides and fluorescent dye designs can be used to further expand the variance of fluorescence lifetime readouts.

[0016] In the exemplary embodiment provided, the peptide was conjugated to a glass surface via the C-terminus, leaving a primary amine at the N-terminus for covalent attachment of a phenyl-isothiocyanate-functionalized oligonucleotide. A complementary imager strand was hybridized to the docking oligonucleotide to bring the fluorophore into close proximity with the N-terminal amino acid and imaged by two-photon FLIM to interrogate the fluorescence lifetime of AF488. Other imager strands conjugated to BODIPY, KU530-6, or KU530-R-4 were also used and imaged to interrogate their lifetimes before removal of the N-terminus by Edman degradation. Among the amino acids tested, tryptophan, arginine, phenylalanine, serine, glutamine, glutamate, and phosphoserine showed significant differences in fluorescence lifetime with the displayed fluorophores. Additionally, amino acids at positions N-1 and N-2 were shown to contribute to the lifetime variation.

[0017] In one embodiment, a method for sequencing a peptide having an initial NTAA is provided, the method comprising sequentially interrogating the initial NTAA using a library of at least two different combinations of ssDNA DS and ssDNA IS, wherein a fluorophore is conjugated to the IS to generate a set of fluorescence lifetime data having a characteristic fingerprint for each different amino acid of the peptide, wherein each pair of IS and DS is at least partially complementary in sequence. In an example of this embodiment, the sequential interrogation comprises detecting and / or measuring the interaction between the fluorophore and the amino acid side chain at or near the NTAA by detecting the fluorescence lifetime data for each pair of IS and DS in the library, for example using FLIM single-molecule fluorescence measurement. Optionally, these sequencing methods can also include removing the initial NTAA peptide by Edman degradation, Edman degradation enzyme reaction, or similar processes.

[0018] Also contemplated are methods for sequencing a peptide, wherein the method is repeated for each subsequent amino acid in the peptide to generate a matrix of fluorescence lifetime data. Optionally, the data is input into a machine learning algorithm to reconstruct the polypeptide sequence.

[0019] Yet another embodiment provides a method for identifying a terminal amino acid (TAA) of a peptide having an N-terminal amino acid (NTAA) and a C-terminal amino acid (CTAA), the method comprising: binding the NTAA of the peptide or the CTAA of the peptide to a solid surface to produce a bound TAA; attaching a ssDNA docking strand (DS) to the unbound TAA of the peptide; hybridizing a first ssDNA imaging strand (IS) to the DS, the first IS comprising a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, the second IS comprising a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; and identifying the original TAA of the peptide based on the detected fluorescence lifetime data of the first fluorophore and the second fluorophore.

[0020] In any of the method embodiments, the library of ISs can include a plurality of ssDNA oligonucleotides that are varied such that the spatial positioning and / or degrees of freedom of the attached fluorophore are modulated to adjust the interaction with the NTAA side chain and thereby modulate the measured fluorescence lifetime. For example, in some instances, the library of ISs includes a plurality of ssDNA oligonucleotides that are varied by one or more of: including modified nucleotides, including non-natural nucleotides, including a 5' IS overhang relative to the cognate DS, or including a 5' IS without an overhang relative to the cognate DS.

[0021] In various embodiments, the fluorophore is conjugated at the end of the IS. Alternatively, the peptide is conjugated to a modified nucleotide within the DS, and the fluorophore is conjugated to a modified nucleotide within the IS ( FIG. 10 ).

[0022] In the examples provided, removal of NTAA is performed under conditions such that the remaining peptide has a new N-terminal amino acid.

[0023] Optionally, the peptide to be sequenced is immobilized on a solid support.

[0024] Yet another embodiment provides a method for identifying the NTAA of a peptide, the method comprising: binding the C-terminal amino acid of the peptide to a solid surface; attaching a ssDNA DS to the NTAA of the peptide; hybridizing a first ssDNA IS to the DS, the first IS comprising a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, the second IS comprising a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; and identifying the initial NTAA of the peptide based on the detected fluorescence lifetime data of the first fluorophore and the second fluorophore. Optionally, the method may further comprise cleaving the initial NTAA from the peptide to leave the next NTAA of the peptide. Optionally, the method further comprises repeating the method multiple times to identify the sequence of the peptide.

[0025] In examples of these method embodiments, cleaving the initial NTAA comprises an Edman degradation reaction, enzymatic cleavage or digestion, or the like.

[0026] Another embodiment is a method for sequencing a peptide, the method comprising: attaching a peptide to be sequenced via its C-terminus to a solid phase substrate; functionalizing the initial N-terminal amino acid of the immobilized peptide with a universal DS ssDNA oligonucleotide; contacting the DS with an IS oligonucleotide complementary to the DS oligonucleotide, the IS conjugated to a first fluorophore; obtaining single-molecule FLIM measurements of the first fluorophore for each peptide; optionally, repeating the single-molecule FLIM measurements for one or more additional combinations of IS and fluorophores in a library; and cleaving the initial N-terminal amino acid from the peptide to reveal a second N-terminal amino acid; and optionally, performing another analysis cycle on the second N-terminal amino acid.

[0027] Also provided are methods of sequencing peptides substantially as described herein.

[0028] Another embodiment is a kit for performing any of the described method embodiments, the kit comprising at least one pair of IS and DS. In an example of the kit embodiment, the kit comprises at least two pairs of IS and DS, wherein the two pairs differ in the fluorophore, sequence, or both contained in the IS.

[0029] Also provided is a database containing matrices of fluorescence lifetime data generated by any of the methods described.

[0030] Additional embodiments include a method for sequencing a peptide having an initial terminal amino acid (TAA), the method comprising: interrogating the initial TAA using a ssDNA DS linked to the initial TAA and a ssDNA IS conjugated to a signal molecule to generate a measurement result of spectral characteristic data having a characteristic fingerprint of the initial TAA of the peptide; wherein the IS and the DS are at least partially complementary in sequence. Optionally, such a method may further comprise: sequentially interrogating the initial TAA using a library of at least two different combinations of ssDNA DS and ssDNA imaging strands (IS), wherein a signal molecule is conjugated to the IS to generate a set of spectral characteristic data having a characteristic fingerprint of the initial TAA of the peptide. In various instances, the initial TAA is: the NTAA of the peptide; or the carboxyl terminal amino acid (CTAA) of the peptide. Optionally, any of these method embodiments may be performed in parallel on multiple peptides. Optionally, the signal molecule may be a fluorophore (such as Alexa Fluor®). 488 (AF488), BODIPY-FL, BODIPY-TR or TAMRA), and optionally, the spectral characteristic comprises fluorescence lifetime.

[0031] Yet another described embodiment is a method of sequencing a peptide having an initial NTAA, the method comprising sequentially interrogating the initial NTAA using a library of at least two different combinations of ssDNA DS and ssDNA IS, wherein a fluorophore is conjugated to the IS, to generate a set of fluorescence lifetime data having characteristic measurements for each combination of DS, IS, and fluorophore, wherein each pair of IS and DS is at least partially complementary in sequence.

[0032] In any of the method embodiments, the DS may be a general DS.

[0033] Also provided are peptide analysis (e.g., sequencing) methods, wherein the sequential interrogation comprises detecting and / or measuring an interaction between a fluorophore and an amino acid side chain at or near the CTAA or NTAA by detecting fluorescence lifetime data for each of a plurality of IS / DS pairs in the library. Optionally, in any embodiment of the method, detecting or measuring the interaction comprises obtaining FLIM single-molecule fluorescence measurements for each of a plurality of IS / DS pairs in the library. Any of methods 1 to 8 may further comprise removing the original CTAA or NTAA of the peptide by Edman degradation, enzymatic digestion, or the like.

[0034] The described method can optionally be repeated for each subsequent amino acid in the peptide, thereby generating a matrix of fluorescence lifetime data. In various embodiments, the data is input into a machine learning algorithm to reconstruct the polypeptide sequence.

[0035] Also provided herein are methods for peptide analysis, wherein a library of ISs comprises a plurality of ssDNA oligonucleotides that are varied such that the spatial positioning and / or degrees of freedom of the attached fluorophore are varied to modulate the interaction with the CTAA side chain or NTAA side chain and thereby modulate the measured fluorescence lifetime. For example, the library of ISs can comprise a plurality of ssDNA oligonucleotides that are varied by one or more of the following: including modified nucleotides, including non-natural nucleotides, including 5' IS overhangs relative to the cognate DS, or including 5' IS non-overhangs relative to the cognate DS. Further, the interaction between the CTAA or NTAA in the embodiments of the method can be further affected by one or more of the DS position, degrees of freedom, or another variable described herein.

[0036] In any of the method embodiments, a signaling molecule (which may optionally be a fluorophore) is conjugated to a nucleotide (eg, a modified nucleotide) of an IS. The nucleotide may optionally be located at either end (5' or 3') of the IS or somewhere within the IS.

[0037] In any of the method embodiments, the peptide is conjugated to a modified nucleotide within the DS (ie, not at a terminus), and the fluorophore is conjugated to a modified nucleotide within the IS (eg, Figure 11 shown).

[0038] In any of the method embodiments, removal of CTAA or NTAA can be performed under conditions such that the remaining peptides have a new terminal amino acid available for another cycle of analysis.

[0039] In any of the method embodiments, the peptide can be immobilized on a solid support.

[0040] Also provided is a database containing a matrix of signal molecule spectral property data prepared using any of the methods described herein. In some instances, this data includes measurements of the spectral lifetimes of a plurality of different signal molecules, as such lifetimes are affected by the proximity of the side chains of the different terminal amino acids of the peptide being analyzed.

[0041] Yet another embodiment provides a method for identifying an NTAA of a peptide, the method comprising: binding the C-terminal amino acid of the peptide to a solid surface; attaching a ssDNA DS to the NTAA of the peptide; hybridizing a first ssDNA IS to the DS, the first IS comprising a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, the second IS comprising a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; and identifying the initial NTAA of the peptide based on the detected fluorescence lifetime data of the first and second fluorophores. Optionally, the method further comprises: cleaving the initial TAA from the peptide to leave the next TAA of the peptide. For example, cleaving the initial NTAA can comprise an Edman degradation reaction, an Edman degradation enzyme reaction, or a similar process. Optionally, the method is repeated multiple times to identify the sequence of the peptide.

[0042] Another embodiment provides a method for identifying a CTAA of a peptide, the method comprising: binding the N-terminal amino acid of the peptide to a solid surface; attaching a ssDNA docking strand (DS) to the CTAA of the peptide; hybridizing a first ssDNA IS to the DS, the first IS comprising a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, the second IS comprising a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; and identifying the initial NTAA of the peptide based on the detected fluorescence lifetime data of the first and second fluorophores. Optionally, the method further comprises cleaving the initial TAA from the peptide to leave the next TAA of the peptide. For example, cleaving the initial NTAA can comprise an Edman degradation reaction, an Edman degradation enzyme reaction, or a similar process. Optionally, the method is repeated multiple times to identify the sequence of the peptide.

[0043] Also provided is a method for sequencing a peptide, the method comprising: attaching a peptide to be sequenced via its C-terminus to a solid phase substrate; functionalizing an initial N-terminal amino acid of the immobilized peptide with a universal DS ssDNA oligonucleotide; contacting the DS with an IS oligonucleotide complementary to the DS oligonucleotide, the IS conjugated to a first fluorophore; obtaining a single molecule FLIM measurement of the first fluorophore for each peptide; optionally, repeating the single molecule FLIM measurement for one or more additional combinations of IS and fluorophores in a library; and cleaving the initial N-terminal amino acid from the peptide to reveal a second N-terminal amino acid; and optionally, performing another analysis cycle on the second N-terminal amino acid.

[0044] Also provided is a method for sequencing a peptide, the method comprising: attaching a peptide to be sequenced via its N-terminus to a solid phase substrate; functionalizing an initial C-terminal amino acid of the immobilized peptide with a universal DS ssDNA oligonucleotide; contacting the DS with an IS oligonucleotide complementary to the DS oligonucleotide, the IS conjugated to a first fluorophore; obtaining a single molecule FLIM measurement of the first fluorophore for each peptide; optionally, repeating the single molecule FLIM measurement for one or more additional combinations of IS and fluorophores in a library; and cleaving the initial C-terminal amino acid from the peptide to reveal a second C-terminal amino acid; and optionally, performing another analysis cycle on the second C-terminal amino acid.

[0045] Another embodiment is a method of sequencing a peptide substantially as described herein. The method contemplated in this embodiment comprises detecting at least one spectral characteristic of a signaling molecule, wherein the spectral characteristic is not fluorescence lifetime.

[0046] Yet another provided embodiment is a kit for practicing the method of any provided embodiment, the kit comprising at least one pair of IS and DS. For example, an example of such a kit comprises at least two pairs of IS and DS, wherein the two pairs differ in the signaling molecule (e.g., fluorophore), or sequence, or both, attached to the IS.

[0047] Also described herein are compounds having formula (I):

[0048]

[0049] or a salt or solvate thereof, wherein: x is 0, 1 or 2; each R is independently selected from the group consisting of: C1-C6 alkyl, -NO2, halogen, -C=OR, -C=SR, -C=ONR, -C=OOR, -SO3 or any other common electron-withdrawing group; R 1 and R 2 are independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, —O—(C1-C6 alkyl), C1-C6 alkyl, hydroxy, halogen, —O-alkyl, —S-alkyl, —OC(═O)R, —N—(C═O)—R, —OC(═O)OR, —NC(═S)NR, —N—(C═O)—OR, or any other common electron-donating group; y is 0, 1, 2, or 3; and each R 3 are independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group. In an embodiment, x is 0; and / or y is 0; and / or R1 or R 2 The C1-C6 alkyl group is a methyl group, and R 1 or R 2 -O-(C1-C6 alkyl) is methoxy.

[0050] Additionally provided are compounds having the formula (II)

[0051]

[0052] or a salt or solvate thereof, wherein:

[0053] x is 0, 1 or 2; each R is independently selected from the group consisting of: C1-C6 alkyl, -NO2, halogen, -C=OR, -C=SR, -C=ONR, -C=OOR, -SO3 or any other common electron withdrawing group; R 1 and R 2 are independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, —O—(C1-C6 alkyl), C1-C6 alkyl, hydroxy, halogen, —O-alkyl, —S-alkyl, —OC(═O)R, —N—(C═O)—R, —OC(═O)OR, —NC(═S)NR, —N—(C═O)—OR, or any other common electron-donating group; y is 0, 1, 2, or 3; and each R 3 are independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group. In an embodiment, x is 0; and / or y is 0; and / or R 1 or R 2 The C1-C6 alkyl group is a methyl group, and R 1 or R 2 -O-(C1-C6 alkyl) is methoxy.

[0054] Additionally provided are compound examples having the structure:

[0055]

[0056] wherein R1 and R2 are each selected from the group consisting of: H, CH3, OH, and OCH3; provided that R1 and R2 are the same; or a salt or solvate thereof. For example, an exemplary compound according to Example 40 is (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)phenyl)carbamothioic acid pivaloic acid thioanhydride; or a salt or solvate thereof.

[0057] Also provided is a method for preparing a compound of formula (I)

[0058]

[0059] or its salt or solvate, the method comprising the steps of:

[0060]

[0061] Converted into a compound of formula (II) or a salt or solvate thereof

[0062]

[0063] and thereafter converting the compound of formula (II) or its salt or solvate into the compound of formula (I) or its salt or solvate, wherein: x is 0, 1 or 2; each R is independently selected from the group consisting of: C1-C6 alkyl, -NO2, halogen, -C=OR, -C=SR, -C=ONR, -C=OOR, -SO3 or any other common electron-withdrawing group; R 1 and R 2 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, -O-(C1-C6 alkyl), halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group; y is 0, 1, 2, or 3; and each R 3 is independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR or any other common electron-donating group. In an example of this embodiment, the compound of formula (III) or a salt or solvate thereof is first converted to a compound of formula (IV) or a solvate thereof,

[0064]

[0065] Subsequently, the compound of formula (IV) or its salt or solvate is converted into the compound of formula (II) or its solvate. In the example of these methods or embodiments, the conversion of the compound of formula (III) or its salt or solvate to the compound of formula (IV) or its salt or solvate is by reacting carbon disulfide (CS2) with the compound of formula (III). For example, the reaction is carried out in the presence of a base (such as (C1-C6 alkyl)3N or more specifically triethylamine).

[0066] Also provided are examples of embodiments of this method, wherein the conversion of the compound of formula (IV) or a salt or solvate thereof to the compound of formula (II) or a salt or solvate thereof is by reacting the compound of formula (IV) or a salt or solvate thereof with di-tert-butyl carbonate (O-(C(=O)-OC(CH3)2)2). For example, the reaction can occur in the presence of one or more bases, such as the one or more bases including dimethylaminopyridine (DMAP) and triethylamine.

[0067] Also described are method embodiments wherein the compound has the structure:

[0068]

[0069] wherein R1 and R2 are each selected from the group consisting of: H, CH3, OH, and OCH3; provided that R1 and R2 are the same; or a salt or solvate thereof. For example, in some instances, the compound is (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)phenyl)carbamothioic acid pivaloic acid thioanhydride; or a salt or solvate thereof.

[0070] Also provided is the use of any of the described compounds in any of the method embodiments of peptide analysis as described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figures 1A-1G Schematic diagram of an example peptide sequencing workflow based on fluorescence lifetime imaging (FLIM). The presented method allows sequential interrogation of N-terminal amino acids (NTAAs) using a library of different combinations of ssDNA imager strands (IS) and fluorophores to generate a set of fluorescence lifetime data, with each distinct amino acid having a characteristic fingerprint. The NTAA is removed by Edman degradation, and the method is repeated for one or more subsequent amino acids to generate a matrix of fluorescence lifetime data. This data can be input into a machine learning algorithm to reconstruct the peptide sequence.

[0072] Figure 2 Single-stranded DNA (ssDNA) docking strand and imaging strand sequences and schematic arrangements are shown. In various embodiments, the ssDNA imaging strand (IS) can be modified in various ways, such as including modified or non-natural nucleotides, including a 5' IS overhang relative to the docking strand (DS), or including a 5' IS non-overhang relative to the DS. These variables adjust the spatial position and / or degrees of freedom of the attached fluorophore and thereby the interaction with the NTAA side chain, which adjusts the measured fluorescence lifetime. This enables the identification of each NTAA. Figure 2Imaged in the bottom panel are DS1 (SEQ ID NO: 1), IS1 (SEQ ID NO: 2), IS2 (SEQ ID NO: 3), and IS3 (SEQ ID NO: 4).

[0073] Figure 3 Examples of the chemical structures of the bifunctional linker precursor (top left) and product, known as maleimidophenyl isothiocyanate (MPITC) (IUPAC: 1-(4-isothiocyanatophenyl)-1H-pyrrole-2,5-dione) (bottom left); a DNA DS-peptide conjugate (center); and a DNA IS-fluorescent dye conjugate (right) are shown. In the example shown, the DNA DS has a 3' propylthiol modification, enabling conjugation to the maleimido group of the MPITC linker. The peptide is conjugated to the isothiocyanate group of the MPITC linker. In the example shown, the DNA IS has a 5' hexylamine modification for conjugation to a fluorophore modified with a chemical crosslinker such as a maleimido group.

[0074] Figure 4 A bar graph showing the fluorescence lifetimes of AF488 and BODIPY-FL conjugated to imager strand 1 ("IS1") (format: N-terminal-AA1-AA2-AA3) near various synthetic peptides containing different N-terminal amino acids as indicated on the X-axis. This figure demonstrates that the use of different fluorophores enables discrimination of sequential amino acids, as shown for FGG, SGG, and RGG, for example. [*, p < 0.05; ***, p < 0.005; two-way ANOVA with Tukey post hoc; N = multiple fields within 3-8 samples].

[0075] Figure 5 is a bar graph showing the average normalized fluorescence intensity measured by various parts of the embodiment (e.g., various assemblies of the complete workflow). Three substructures (silane-PEG; silane-PEG-peptide+MPITC+DS; silane-PEG+DS+IS) and the complete embodiment were analyzed. A peptide containing a serine at the N-terminus and imaging strand 1 ("IS1") conjugated to AF488 was used in the analysis. These data in the figure show that the collected fluorescence lifetime data is mainly dominated by the fluorophore and indicate that other components of the construct do not produce significant fluorescence. Data are normalized to the blank sample [*, p>0.05; ***, p<0.005; ****, p<0.0001; one-way ANOVA with Tukey's post hoc test; N=3 samples].

[0076] Figure 6is a bar graph showing queries of N-terminal amino acids with post-translational modifications (PTMs). Four individual fluorophores (Alexa Fluor 500, ... Fluorescence lifetimes of 488, BODIPY-FL, BODIPY-TR, and TAMRA (N = 1 sample).

[0077] Figures 7A-7C Shown are the normalized fluorescence lifetimes of a fluorophore conjugated to IS1 before and after removal of the NTAA (position "N") to expose the next amino acid in the polypeptide chain (position "N-1"). In the examples shown, Edman degradation was used to remove the NTAA. Figure 7A Fluorescence lifetime measured by imager strand 1 (IS1) conjugated to Alexa Fluor 488 (IS1-AF488) before and after a single Edman degradation of a peptide to remove the N-terminal amino acid at position "N" and expose the next amino acid in the polypeptide chain ("N-1"). The cleaved amino acid in the sequence is indicated in parentheses. [N = multiple fields of view within 3 samples]. Figure 7B Fluorescence lifetime measured by imager strand 1 (IS1) conjugated to BODIPY-FL (IS1-BODIPY-FL) before and after a single Edman degradation of the peptide to remove the N-terminal amino acid at position "N" and expose the next amino acid in the polypeptide chain ("N-1"). The cleaved amino acid in the sequence is indicated in parentheses. [*, p < 0.05; one-way ANOVA with Tukey's post hoc test; N = 3 samples within multiple fields]. Brackets in the N-terminal sequence (X-axis labels) indicate the amino acid removed by Edman degradation. Data were normalized to a free (unattached) IS1 fluorophore control. Figure 7C A complete cycle of terminal amino acid analysis, including reset to the second terminal amino acid, is shown; the different fluorophores are indicated by the displayed stars.

[0078] Figures 8A-8H is a series of bar graphs showing the normalized fluorescence lifetimes of fluorophores conjugated to IS1 before and after multiple cycles of Edman degradation to sequentially remove and collect lifetime information of amino acids in a polypeptide chain. IS1-AF488 is shown in Figures 8A-8D and IS1-BODIPY-FL is shown in Figures 8E-8H from the nucleotide sequence containing tryptophan at the second ("N-1") position and the third ("N-2") position. Figure 8A and 8B as well as Figure 8E and 8F ) and arginine ( Figure 8C and 8Das well as Figure 8G and 8H ) of the peptide adjacent to the fluorophore to collect the fluorescence lifetime. Figure 8A Fluorescence lifetime measured by imager strands conjugated to Alexa Fluor 488 for a peptide with a tryptophan at the second position (“N-1”) along the peptide before and after multiple cycles of Edman degradation [***, p<0.005; one-way ANOVA with Tukey’s post hoc test; N=3 fields within samples]. Figure 8B Fluorescence lifetime measured by imager strands conjugated to AlexaFluor 488 for a peptide with a tryptophan at the third position (“N-2”) along the peptide before and after multiple cycles of Edman degradation [*, p < 0.05; **, p < 0.01; one-way ANOVA with Tukey’s post hoc test; N = multiple fields within 3 samples]. Figure 8C Fluorescence lifetime measured by an imager strand conjugated to Alexa Fluor 488 for a peptide with an arginine at the second position along the peptide (“N-1”) before and after multiple cycles of Edman degradation [N=multiple fields of view within 3 samples]. Figure 8D Fluorescence lifetime measured by imager strands conjugated to Alexa Fluor 488 for a peptide with an arginine at the third position (“N-2”) along the peptide before and after multiple cycles of Edman degradation. [*, p<0.05; one-way ANOVA with Tukey’s post hoc test; N=multiple fields within 3 samples]. Figure 8E Fluorescence lifetime measured by imager strands conjugated to BODIPY-FL for a peptide with a tryptophan at the second position (“N-1”) along the peptide before and after multiple cycles of Edman degradation [*, p < 0.05; one-way ANOVA with Tukey’s post hoc test; N = multiple fields within 3 samples]. Figure 8F Fluorescence lifetime measured by imager strands conjugated to BODIPY-FL for a peptide with a tryptophan at the third position (“N-2”) along the peptide before and after multiple cycles of Edman degradation [*, p<0.05; one-way ANOVA with Tukey’s post hoc test; N=3 fields within samples]. Figure 8G Fluorescence lifetime measured by imager strands conjugated to BODIPY-FL for a peptide with an arginine at the second position along the peptide ("N-1") before and after multiple cycles of Edman degradation [*, p < 0.05; one-way ANOVA with Tukey's post hoc test; N = multiple fields within 3 samples]. Figure 8HFluorescence lifetime measured by imager strands conjugated to BODIPY-FL for a peptide with an arginine at the third position ("N-2") along the peptide before and after multiple cycles of Edman degradation [*, p < 0.05; **, p < 0.01, one-way ANOVA with Tukey's post hoc test; N = multiple fields within 3 samples].

[0079] Figure 9 is a series of bar graphs showing the dependence of fluorescence lifetime on the combination of DNA IS sequence and fluorescent dye selection. The data shown uses the same peptide but with different imager chains (IS1 (SEQ ID NO: 2), IS2 (SEQ ID NO: 3), and IS3 (SEQ ID NO: 4)) and fluorophores (Alexa Fluor in the top panel). 488; BODIPY at the bottom). IS1 and IS2 have higher melting temperatures, and IS3 has a single base overhang. Data are normalized to the fluorescence lifetime of the corresponding free fluorescent dye. [N = 1 sample]

[0080] Figures 10A-10C The generation of unique peptide fingerprints using the presented method is demonstrated. Figure 10A is a two-dimensional graph showing the fluorescence lifetime fingerprints of different N-terminal amino acids. Figure 10B Demonstrating that fluorophore cycling generates robust amino acid fingerprints. Figure 10C It is shown that complex multivariate results improve the accuracy of neural network predictions. In the case of a theoretical neural network approach for amino acid identification, the fluorescence lifetime fingerprints of different N-terminal amino acids are illustrated.

[0081] Figure 11 An alternative peptide sequencing system embodiment using a DNA major groove design is presented. The immobilized peptide is conjugated to a modified nucleotide within the DS (i.e., not directly near either end of the DS), and the fluorophore is conjugated to a modified nucleotide within the IS (i.e., not directly near either end of the IS). This provides additional adjustable control over the interaction between the fluorescent dye and the structural components of the DNA docking strand: imaging strand (DS:IS) complex. In the current embodiment, there is an interaction between the fluorescent dye and the blunt-ended nucleobases of the DS:IS complex, whereas in this alternative embodiment, the fluorescent dye is conjugated internally and therefore cannot reach the blunt end of the complex. This has been verified by computer-simulated molecular dynamics simulations. The peptide is conjugated to the modified nucleotide within the docking strand, and the fluorophore is conjugated to the modified nucleotide within the imaging strand.

[0082] Figure 12 is a bar graph depicting the native fluorescence lifetime measured by a fluorophore-conjugated imager strand (IS1) in water. [ *, p<0.05; **, p<0.01; ****, p<0.001; one-way ANOVA with Tukey's post hoc test; N = 3 samples within multiple fields].

[0083] Figure 13 Bar graph showing the various fluorescence lifetimes measured by various imager strands ("IS1") containing the long-lived fluorophores KU530-6 and KU530-R-4 to identify the N-terminal amino acid [*, p < 0.01; ****, p < 0.001; two-way ANOVA with Tukey's post hoc test; N = multiple fields within 3-4 samples].

[0084] Figure 14 Fluorescence lifetimes measured by imager strands conjugated to AF488, BODIPY-FL, or KU530-6 for peptides with common post-translational modifications at the N-terminus, such as phosphorylation of serine or acetylation of lysine [*, p < 0.05; one-way ANOVA with Tukey's post hoc test; N = multiple fields within 3 samples].

[0085] Figure 15 Fluorescence lifetimes reported by IS1 conjugated to AF488 in SGG and post-translationally modified serine (PhosSGG) control samples. Phosphorylation was removed from FLIM of PhosSGG and IS1-AF488 using phosphatase, and lifetimes similar to those of the control were reported. [*, p < 0.05; one-way ANOVA with Tukey's post hoc test; N = 3 samples within multiple plots].

[0086] Figure 16 Normalized lifetime “heatmaps” from sequencing data as reported by 4 separate fluorophore-conjugated imager strands and peptide screening experiments. Each data point was normalized to the GGGS of each fluorophore and patterned based on the corresponding normalized lifetime range.

[0087] Figure 17 In the case of peptides with tryptophan or arginine at the second and third positions along the peptide ("N-2"), the fluorescence lifetime was measured by an imager strand conjugated to Alexa Fluor 488 to determine whether the amino acid at N-1 or N-2 affects the lifetime of the fluorophore. [N = multiple fields of view within 3 samples].

[0088] FIG1. ​​Fluorescence lifetimes measured by imager strands conjugated to BODIPY-FL for peptides with tryptophan or arginine at the second and third positions (“N-2”) along the peptide to determine whether the amino acid at N-1 or N-2 affects the lifetime of the fluorophore. [N = multiple fields of view within 3 samples].

[0089] Figure 19 Fluorescence lifetime measured by imager strands conjugated to KU530-6 for peptides with tryptophan or arginine at the second and third positions ("N-2") along the peptide to determine whether the amino acid at N-1 or N-2 affects the lifetime of the fluorophore [*, p < 0.05; ***, p < 0.005, one-way ANOVA with Tukey's post hoc test; N = multiple fields within 3 samples].

[0090] Figure 20 Fluorescence lifetime measured by imager strand 1 (IS1) conjugated to KU530-6 before and after a single Edman degradation of the peptide to remove the N-terminal amino acid at position "N" and expose the next amino acid in the polypeptide chain ("N-1"). The cleaved amino acid in the sequence is indicated in parentheses. [**, p < 0.01; one-way ANOVA with Tukey's post hoc test; N = multiple fields within 3 samples].

[0091] Figures 21A-21C From a more detailed Figure 21A ; SEQ ID NO: 5)、RGWSGGSDC( Figure 21B ; SEQ ID NO: 6) and WRGSGGSDC ( Figure 21C Sequencing data for a synthetic peptide with the full sequence of SEQ ID NO: 7). A complete workflow has been completed, including imager strand cycling and multiple Edman degradation cycles. As shown, at each step of the workflow, the fluorescence lifetime was measured by an imager strand conjugated to Alexa Fluor 488. Brackets indicate the amino acids that were cleaved during the Edman degradation process. Figure 21A : ** , p < 0.01; ***, p < 0.005; ****, p < 0.001, one-way ANOVA, Tukey post hoc test; Figure 21B : * , p < 0.05; **, p < 0.01; ****, p < 0.001, one-way ANOVA, Tukey post hoc test; Figure 21C : * , p < 0.05; ****, p < 0.001, one-way ANOVA, Tukey's post hoc test; for all panels, N = 3 fields within a sample].

[0092] Figures 22A-22C From a more detailed Figure 22A ; SEQ ID NO: 5)、RGWSGGSDC( Figure 22B ; SEQ ID NO: 6) and WRGSGGSDC ( Figure 22CSequencing data for a synthetic peptide with the full sequence of SEQ ID NO: 7). A complete workflow has been completed, including imager strand cycling and multiple Edman degradation cycles. As shown, at each step of the workflow, the fluorescence lifetime was measured by the imager strand conjugated to BODIPY-FL. Brackets indicate the amino acids that were cleaved during the Edman degradation process. Figure 22B : * , p < 0.05; ****, p < 0.001, one-way ANOVA, Tukey's post hoc test; for all panels, N = 3 fields within a sample].

[0093] Figures 23A-23C From a more detailed Figure 23A ; SEQ ID NO: 5)、RGWSGGSDC( Figure 23B ; SEQ ID NO: 6) and WRGSGGSDC ( Figure 23C Sequencing data for a synthetic peptide with the full sequence of SEQ ID NO: 7). A complete workflow has been completed, including imager strand cycling and multiple Edman degradation cycles. As shown, at each step of the workflow, the fluorescence lifetime was measured by the imager strand conjugated to KU530-6. Brackets indicate the amino acids that were cleaved during the Edman degradation process. Figure 23A : * , p < 0.05; **, p < 0.01; ****, p < 0.001, one-way ANOVA, Tukey post hoc test; Figure 23B : * , p<0.05; **, p<0.01, one-way ANOVA, Tukey's post hoc test; for all panels, N=3 multiple fields within samples].

[0094] Figures 24A-24C From a more detailed Figure 24A ; SEQ ID NO: 5)、RGWSGGSDC( Figure 24B ; SEQ ID NO: 6) and WRGSGGSDC ( Figure 24C Sequencing data for a synthetic peptide with the full sequence of (SEQ ID NO: 7). A complete workflow has been completed, including imaging strand cycles and multiple Edman degradation cycles. At each step of the workflow, fluorescence lifetime was measured by strands conjugated to KU530-R-4, as shown. Brackets indicate amino acids that were cleaved during the Edman degradation process. [For all panels, N = multiple fields within 3 samples].

[0095] Figure 25Normalized lifetime "heatmap" from sequencing data as reported by four separate fluorophore-conjugated imager strands. Each data point is normalized to the GGGS of the corresponding fluorophore and patterned based on the corresponding normalized lifetime range. Peptides are shown in SEQ ID NO: 5 (top), SEQ ID NO: 6 (middle), and SEQ ID NO: 7 (bottom).

[0096] Figure 26 Fluorescence lifetimes measured by various KU dyes in the presence of different peptides. [*, p < 0.05; ***, p < 0.005; ****, p < 0.001; one-way ANOVA with Tukey post hoc test; N = 3 samples within multiple fields] Reference sequence

[0097] Nucleic acid and / or amino acid sequences described herein are shown using standard letter abbreviations as defined in 37 CFR §1.822.Only one strand of each nucleic acid sequence is shown, but the complementary strand is understood to be included in appropriate embodiments.

[0098] SEQ ID NO: 1 shows the nucleic acid sequence of an exemplary docking strand DS1: ATCTACATATCTC.

[0099] SEQ ID NO: 2 shows the nucleic acid sequence of the first exemplary imager strand IS1: TAGATGTATAGAG.

[0100] SEQ ID NO: 3 shows the nucleic acid sequence of the second exemplary imaging strand IS2: T L A L GATGTATAGAG (where "L" indicates the nucleic acid is locked).

[0101] SEQ ID NO: 4 shows the nucleic acid sequence of the third exemplary imager strand IS3: TT L A L GATGTATAGAG (where "L" indicates the nucleic acid is locked).

[0102] SEQ ID NO: 5 shows the amino acid sequence of the synthetic peptide: WGRSGGSDC

[0103] SEQ ID NO: 6 shows the amino acid sequence of the synthetic peptide: RGWSGGSDC

[0104] SEQ ID NO: 7 shows the amino acid sequence of the synthetic peptide: WRGSGGSDC

[0105] SEQ ID NO:8 shows the amino acid sequence shared by SEQ ID NO:5-7: (XXX)SGGSDC DETAILED DESCRIPTION

[0106] The advent of next-generation sequencing has greatly accelerated clinical and translational discovery, but genomic sequencing incompletely depicts the proteomic landscape of biological systems. Similarly, future de novo protein sequencing methods will revolutionize many areas of biological research and medicine. The recent global push for next-generation protein sequencing has yielded powerful mass spectrometry- and fluorescence-based methods; however, these technologies cannot fully sequence proteins de novo, have low throughput, and are unable to account for sensitivity to post-translational modifications with high fidelity.

[0107] In the present disclosure, methods for determining the terminal amino acid of a peptide are described. In various embodiments, this is achieved by distinguishing the change in fluorescence lifetime of a common fluorophore upon interaction with each terminal amino acid. Fluorophores are attached to DNA oligonucleotides using the described techniques, the fluorophores are strategically placed in close proximity to the amino acids, and fluorescence lifetimes based on the fluorophore-amino acid interactions are collected. Additionally, the use of oligonucleotides expands the characterization of amino acids within a sequence, as various fluorophores can be attached to collect different fluorescence lifetime information. Additionally, changes in the characteristics of the imaging strand oligonucleotide can be included to alter the fluorophore-amino acid interactions, which can be helpful, for example, in characterizing post-translationally modified (PTM) amino acids. Multiple cycles of terminal amino acid sequence removal (e.g., Edman degradation cycles) can be performed with the current workflow described to achieve sequencing of peptides.

[0108] This paper describes a single-molecule peptide sequencing method, embodiments of which use cyclic Edman degradation-based chemistry with an optical readout, such as fluorescence lifetime measurements. Fluorescence lifetime imaging (FLIM) measures the single-molecule fluorescence of individual fluorophores to determine the time spent in an excited state before photon relaxation and emission. For some fluorophores, the excited-state lifetime can be extremely sensitive to local and global (but molecular-level) environmental changes.

[0109] In the exemplary workflow described, peptides are bound to a glass surface via the C-terminus, leaving a primary amine at the N-terminus for covalent attachment of a phenyl-isothiocyanate functionalized oligonucleotide (docking strand, DS). A complementary imaging strand (IS) is hybridized to the docking DS oligonucleotide so that the fluorophore (attached to the IS) is in close proximity to the N-terminal amino acid, and the system is imaged by two-photon FLIM to interrogate the fluorescence lifetime of the fluorophore (e.g., AF488). The lifetime of the IS labeled with a second fluorophore (specifically, the BODIPY-FL imaging strand) is also measured before removal of the N-terminus by Edman degradation. Among the amino acids tested, tryptophan, arginine, phenylalanine, serine, and phosphoserine showed significant differences in fluorescence lifetime. In addition, it has been demonstrated that the amino acids in positions N-1 and N-2 contribute to the lifetime variation. This technology enables a highly sensitive, high-throughput method for sequencing peptides.

[0110] Various aspects of the present disclosure are now described with reference to the following additional details and options: (I) Overview of polypeptide sequencing methods; (II) Selection and preparation of peptides for analysis; (III) Blocking of undesired chemical reactivity with protecting groups; (IV) Support surfaces for immobilizing peptides; (V) Support surface conjugation of peptides; (VI) Modification of the free end of the immobilized peptide; (VII) Signaling molecules; (VIII) Imager strands and their attachment to signaling molecules; (IX) Docking strands and their attachment to the immobilized peptide; (X) Binding of labeled imager strands to the docking strand on the immobilized peptide; (XI) Detection of imager strand signal; (XII) Cleavage of terminal amino acids and release of docking strands; (XIII) Cycling methods and their applications; (XIV) Sequence assembly - comparison with databases; (XV) Kits and articles of manufacture; (XVI) Representative definitions; (XVII) Exemplary examples; (XVIII) Experimental examples; (XV) Selected references; and (XVI) Concluding paragraphs. These headings do not limit the interpretation of the present disclosure and are provided for organizational purposes only.

[0111] (I) Overview of Peptide Sequencing Methods

[0112] The present disclosure proposes a novel peptide sequencing method that can be implemented using readily available reagents and equipment as described herein. The design workflow of one embodiment can be shown in Figures 1A-1G It begins with obtaining unmodified peptides that have been enzymatically digested from whole proteins (as in an LC-MS based peptide sequencing workflow). Enzymatic digestion cleaves the protein. These digested (or otherwise fragmented) peptides are attached from their c-termini to a solid phase substrate (e.g., via thiol-maleimide click chemistry between the thiol of a cysteine ​​residue at the c-terminus and the maleimide of a silane-PEG-maleimide linker that has been previously conjugated to a glass surface). Figure 1A ).

[0113] The n-terminus of the immobilized peptide was functionalized with a docking strand DNA oligonucleotide conjugated to phenyl isothiocyanate (PITC), which is optionally a universal docking strand, as Figure 1B Here, PITC serves two different purposes: 1) mediating the conjugation of the docking oligonucleotide; and 2) implementing cleavage of the N-terminal amino acid when the next amino acid is to be read out by using Edman degradation (Smith, in Encyclopedia of Life Sciences. 2001, MacMillan Publishers Ltd, Nature Publishing Group, available online at els.net.). A fluorophore (or another signaling molecule) is conjugated to a complementary oligonucleotide, called an imaging strand (IS), and is used to place the fluorophore in close proximity to the N-terminal amino acid for direct interaction ( Figure 1C In this example, a library of fluorophores can be attached to imager strands to measure the lifetimes (or other signaling molecules and other spectral properties) of various fluorophores and their subsequent differential interactions with the n-terminal amino acid. Additionally, the imager strands can be varied to alter the distance between the fluorophore and the n-terminal amino acid, ultimately changing the fluorescence lifetime readout and providing more data to evaluate each amino acid individually.

[0114] Readout begins with docking the first fluorophore (attached to the IS) to the DS and collecting single-molecule fluorescence lifetime measurements for each peptide ( Figure 1D As depicted in 1E, the fluorescence lifetime of the fluorophore will differ from its free form due to its interaction with the n-terminal amino acid and varies between different fluorophores (Anju et al., ACS Omega 4(7):12357-12565, 2019; U.S. Patent No. 7,046,661). Single-molecule FLIM can be repeated for other fluorophores in the IS library ( Figures 1C-1E Finally, the N-terminal amino acid is cleaved, for example by Edman degradation, and the immobilized peptide is then ready for the next cycle ( Figure 1F The next cycle begins with conjugation of the DS oligonucleotide to the new n-terminal amino acid ( Figure 1B ), FLIM measurements were performed using various IS ( Figures 1C-1E ), and ends in another (Edman) degradation ( Figure 1F ), releasing N-1TAA (compared with the initial TAA).

[0115] Reading out the amino acids of a peptide (e.g., all the amino acids of a peptide) yields a set of fluorescence lifetime measurements for each amino acid corresponding to different fluorophores. This data can be processed in a machine learning-based prediction algorithm that will generate a peptide sequence ( Figure 1G To improve prediction performance, a training dataset of measurements with known peptide sequences can be provided and referenced. It is expected that this approach will be able to read out millions of peptides, each containing 10-20 or more amino acids, and analyze them in parallel.

[0116] In a representative embodiment, the C-terminal carboxylic acid (or side chain of an amino acid, such as the sulfhydryl group of a cysteine ​​residue) of a natural, synthetic, or modified peptide (or a collection of two or more peptides) is first conjugated to the surface of a support ( Figures 1A-1G In an example, such peptides can be prepared by fragmentation of the protein, such as by enzymatic degradation of the protein by treatment with one or more proteases (e.g., peptidases and / or proteases). Alternatively, the protein can be chemically digested with agents such as cyanogen bromide.

[0117] Before or after enzymatic or chemical digestion, various known chemical reactions can be used to block reactive chemical groups on proteins or peptides, such as the side chain functional groups of lysine (amino), aspartic acid (carboxyl), glutamic acid (carboxyl), cysteine ​​(sulfhydryl), serine (hydroxyl), threonine (hydroxyl), tyrosine (hydroxyl), and arginine (guanidino). This prevents the labile side chains of these amino acids from interfering with subsequent chemical reactions in the workflow.

[0118] Surface conjugation of peptides can be achieved by linking to an azide-functionalized glass surface (or more generally, a functionalized support surface) using a bifunctional DBCO-tetraethylene glycol-maleimide crosslinker. The maleimide group of the crosslinker reacts with the sulfhydryl group of the C-terminal cysteine ​​residue of the synthetic peptide to form a covalent bond, and the DBCO group of the crosslinker reacts with the azide group on the support surface to form a covalent bond. In a second example, a maleimide-functionalized glass surface is prepared, and the sulfhydryl group of the cysteine ​​residue contained in the synthetic peptide is covalently bonded to the maleimide.

[0119] Subsequently, in one example, the N-terminal amine of the peptide is covalently conjugated to the isothiocyanate group of a bifunctional maleimidophenyl isothiocyanate (MPITC) cross-linker ( Figure 3 ). Then, a single-stranded DNA (ssDNA) docking strand (DS) of the selected sequence with modifications on the 3' end of the oligonucleotide was designed and synthesized ( Figure 2) is covalently attached to the N-terminal maleimido group of the surface-conjugated MPITC-peptide construct (Figures 1 and 3). In one example, the modification is an alkylthiol group, such as (3-mercaptopropyl) phosphate. The DS can be universal in that it can be conjugated to any / all peptides in an experiment, regardless of the primary sequence of the peptide.

[0120] For this synthetic ssDNA DS, an ssDNA imaging strand (IS) with a 5' conjugated fluorophore is attached to form a double-stranded DS:IS complex. The designed ssDNA IS of the selected sequence has a modification on the 5' end of the oligonucleotide. In one example, this modification is an alkylamine, such as (6-aminohexyl) phosphate, which is covalently conjugated to the maleimide group of the maleimide-functionalized fluorophore ( Figure 2 and 3 ). Peptide-DS conjugates are combined with IS-fluorophores (including Alexa Fluor) The complete assembly of ISs (fluorophores such as 488 [AF488] and BODIPY-FL) has been modeled and simulated in silico to guide the design of oligonucleotide sequences and individual linkers within the complete assembly. The length of the IS can be influenced by the melting temperature of the DS:IS complex. Designing an IS with a specific guanine and cytosine content and adjusting environmental salt conditions can help produce the desired IS. The minimum length of an IS can be 8-10 base pairs. The maximum length of an IS can be influenced by (e.g., constrained by) an upper limit on the melting temperature of the DS:IS complex.

[0121] Once the DS:IS-fluorophore complex is formed, the fluorophore attached to the 5' end of the IS oligonucleotide is positioned in close spatial proximity to the side chain of the peptide's NTAA (Figures 1 and 2). This close spatial proximity promotes weak molecular interactions between the fluorophore and the NTAA. Based on the chemical structure of the NTAA side chain and the chemical structure of the selected fluorophore, changes in the measured fluorescence lifetime of the fluorophore relative to the fluorescence lifetime of the free fluorophore are observed. In addition, the ssDNA IS sequence can contain various synthetic structural features, including modified or non-natural nucleotides, 5' single-stranded overhangs of one or more nucleotides, or 5' unoverhangs of one or more nucleotides ( Figure 2). In one example, a locked nucleic acid (LNA), also known as a locked nucleotide, can be included at the 5' end of the IS oligonucleotide sequence. LNA exhibits stronger base pairing with its complementary nucleic acid. In this case, the inclusion of LNA increases the stability of the DS:IS-fluorophore complex and limits the degree of freedom of the fluorophore by preventing the terminal base pairing from being temporarily interrupted in a random manner. These synthetic structural features of the IS affect the positioning and degree of freedom of the fluorophore, which changes the relative positioning of the fluorophore and the NTAA side chain and causes the measured fluorescence lifetime of the fluorophore to vary. Therefore, based on the sequence design of the IS and the choice of the conjugated fluorophore, NTAA acquires a unique fluorescence lifetime characteristic.

[0122] The non-covalent bonding of the DS:IS-fluorophore complex is then broken, allowing the IS-fluorophore to be removed and new IS-fluorophore molecules with different ssDNA sequence designs and / or different fluorophores to form new DS:IS-fluorophore complexes. In one example, the double-stranded DNA DS:IS-fluorophore complex can be broken with a chemical denaturant (such as concentrated urea) and / or by heating to raise the ambient temperature above the melting temperature of the complex. After the DS:IS-fluorophore non-covalent bond is broken, the IS-fluorophore is washed off and a new IS-fluorophore combination can be added.

[0123] Different combinations of IS sequences and conjugated fluorophores were sequentially attached to the DS ( Figures 1A-1G ). Fluorescence lifetime measurements are taken for each IS-fluorophore combination, which is then removed and a different combination is attached to the DS. In this way, NTAA is repeatedly interrogated to generate a characteristic fingerprint of fluorescence lifetime data that varies depending on the identity of the NTAA, the choice of IS sequence, and the choice of fluorophore conjugated to the selected IS ( Figure 4-9 The library of different IS-fluorophore combinations can be expanded or narrowed as needed to determine the identity of NTAA, as it has been shown that other components of the embodiment do not dominate or confound the fluorescence ( Figure 5 ).

[0124] Next, the N-terminal amino acid is removed from the peptide by performing Edman degradation. Briefly, the phenylisothiocyanate (PITC) group of the MPITC linker is able to undergo a cyclization reaction with NTAA, which ultimately hydrolyzes the peptide bond between NTAA and the amino acid in the next position within the polypeptide chain (the "N minus 1" (N-1) position). Figure 3 This exposes a new N-terminal amine to which new MPITC and DS can be conjugated ( Figures 1A-1G A library of IS-fluorophore combinations was then used to interrogate this novel NTAA and generate a fingerprint of fluorescence lifetime data based on the identity of this novel NTAA ( Figures 7A-7C).

[0125] The described method is repeated for each subsequent amino acid until some or all amino acids in the polypeptide chain have been interrogated by the library of IS-fluorophores and a three-dimensional matrix of fluorescence lifetime data for each combination of NTAA, IS sequence and fluorophore is generated (Figures 1, 8A-8H and 10A-10C). The length of the peptide can be up to 50 amino acids, for example, ranging from 10-50, 10-40, 10-30, 20-50, 30-50, 30-40, 20-40, 20-30, 10-20, etc. In addition to identifying proteinogenic amino acids, this method can also be used to detect and identify amino acid post-translational modifications (PTMs) and atypical or unnatural amino acids ( Figure 6 and 7A and 7B). This matrix of fluorescence lifetime data is then used as input to a machine learning algorithm (such as a convolutional neural network) to determine the identity of the amino acids in the peptide chain (Figures 1 and 10C). Based on the number of Edman degradation cycles performed, the position of each amino acid within the peptide chain is known. For peptides derived from naturally occurring proteins, after identifying some or all of the amino acids in the polypeptide chain, the sequence can be aligned with the proteome of the source organism. In the scientific literature, it has been shown that only a subset of amino acids in enzymatically digested proteins need to be identified with positional accuracy to determine the identity of the protein by proteomic alignment (Swaminthan et al., PLOS Comp Biol., 2015; doi.org / 10.1371 / journal.pcbi.1004080).

[0126] In an alternative embodiment (e.g. Figure 11 In one embodiment, the peptide and fluorophore are linked to an internally modified nucleotide within the DS oligonucleotide sequence, rather than to the terminus of the DS oligonucleotide. Similarly, the fluorophore is linked to an internally modified nucleotide within the IS oligonucleotide sequence, rather than to the terminus of the IS oligonucleotide. Computer-generated molecular dynamics simulations have shown that, for this embodiment, the peptide and fluorophore can be positioned within the major groove of the double-stranded DNA to minimize interactions between the fluorophore and the nucleobases of the DS:IS complex. In another embodiment, the terminal nucleobase of the blunt end of the DS:IS complex may interact with the fluorophore conjugated to the IS and affect the modulation of the fluorescence lifetime. Thus, in various embodiments, DNA bases or grooves may also interact with the dye and / or TAA to restrict degrees of freedom or otherwise affect lifetime measurements.

[0127] (II) Selection and preparation of peptides for analysis

[0128] This paper provides methods for analyzing proteins, polypeptides and peptides. These methods are applicable to any polypeptide molecule, no matter how its source. In each embodiment, polypeptide is advantageously processed or prepared before analysis.

[0129] In certain embodiments, the protein, polypeptide or peptide to be analyzed is obtained from a biological sample. For example, the sample can contain mammalian (e.g., human) cells, plant cells, fungal (e.g., yeast) cells and / or prokaryotic (e.g., bacterial) cells. In certain embodiments, the sample contains cells from a sample obtained from a multicellular organism. For example, a sample can be isolated from an individual (also referred to as a subject, or in some cases, a patient). The sample can contain a single cell type or multiple cell types. The sample can include two or more cells.

[0130] The sample can be obtained from a mammalian organism or a human, for example, by puncture or other collection or sampling procedures, such as those known in the art.

[0131] The peptides can be composed of L-amino acids, D-amino acids, or both. The peptides, polypeptides, proteins, or protein complexes can contain one or more of the following: standard, naturally occurring amino acids, modified amino acids (e.g., post-translationally modified), amino acid analogs or mimetics, or any combination thereof. In some embodiments, the polypeptides to be analyzed are naturally occurring, synthetically produced, or recombinantly expressed.

[0132] The standard naturally occurring amino acids include alanine (A or Ala), cysteine ​​(C or Cys), aspartic acid (D or Asp), glutamate (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or He), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or GIn), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp), and tyrosine (Y or Tyr). Non-standard amino acids include selenocysteine, pyrrolysine and N-formylmethionine, β-amino acids, homologous amino acids, proline and pyruvate derivatives, 3-substituted alanine derivatives, glycine derivatives, cyclo-substituted phenylalanine and tyrosine derivatives, linear core amino acids, and N-methyl amino acids.

[0133] In any of the foregoing embodiments, the peptide, polypeptide, protein or polypeptide complex may further include one or more post-translational modifications. The post-translational modification (PTM) of the peptide, polypeptide or protein may be a covalent modification or an enzymatic modification. Examples of PTMs include acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, C-terminal amidation, deamidation, deimination, diphtheria amide formation, disulfide bridge formation, elimination, farnesylation, flavin linkage, formylation, γ-carboxylation, glutamylation, glycylation, glycosylation (including C-linked, N-linked, O-linked glycosylation and phosphoglycosylation), glycosylation. Phosphatidylinositylation, heme C linkage, hydroxylation, hydroxyputrescine lysine formation, iodination, prenylation, lipidation, lipoylation, malonylation, methylation, myristoylation, oxidation, palmitoylation, PEGylation, phosphopantetheinylation, phosphorylation, prenylation, propionylation, retinylidene Schiff base formation, S-glutathionylation, S-nitrosylation, S-sulfenylation, selenoylation, succinylation, sulfinylation, and ubiquitination.

[0134] PTM includes modification of the amino terminus and / or carboxyl terminus of a peptide, polypeptide or protein. The modification of the terminal amino group includes amino, N-low alkyl, N-di-low alkyl and N-acyl modifications. The modification of the terminal carboxyl group includes amide, lower alkyl amide, dialkyl amide and lower alkyl ester modification (for example, wherein the lower alkyl is C1-C4 alkyl). Post-translational modification also includes modification of the amino acid between the amino terminus and the carboxyl terminus of a peptide, polypeptide or protein, such as those mentioned above. Post-translational modification can affect the characteristics and / or function of intracellular proteins, such as their activity, structure, stability or positioning. For example, phosphorylation plays an important role in the regulation of some proteins, particularly in cell signaling (Prabakaran et al., Wiley Interdisciplinary Review-System Biology and Medicine (Wiley Interdiscip Rev Syst Biol Med) 4:565-583,2012). Similarly, adding sugar (for example, glycosylation) to a protein is considered to promote protein folding, improve stability and change regulatory function; and the connection of lipids can target proteins to the cell membrane.

[0135] Post-translational modifications can also include modifications of the peptide, polypeptide, or protein performed by experimental or scientific procedures, such as attachment of detectable labels, linkers, and the like.

[0136] Provided herein is a method for determining (e.g., sequencing) polypeptides, proteins and / or peptides. The method also allows for simultaneous analysis of multiple different peptides (two or more peptides), such as multiplexing. As used herein, simultaneously refers to analyzing (e.g., sequencing) multiple peptides with different sequences in the same assay. The multiple peptides analyzed can be present in the same sample (e.g., biological sample) or in different samples. The multiple polypeptides can be derived from the same subject or different subjects. In certain embodiments, the method is performed on multiple separated polypeptides from a sample. In some aspects, the identity of the polypeptide is unknown. The multiple polypeptides analyzed can be different polypeptides, or the same polypeptide derived from different samples. A plurality of polypeptides includes 2 or more polypeptides, 5 or more polypeptides, 10 or more polypeptides, 50 or more polypeptides, 100 or more polypeptides, 500 or more polypeptides, 1000 or more polypeptides, 5,000 or more polypeptides, 10,000 or more polypeptides, 50,000 or more polypeptides, 100,000 or more polypeptides, 500,000 or more polypeptides, or 1,000,000 or more polypeptides.

[0137] It is also contemplated that different peptide fragments of a single (or multiple) polypeptide are analyzed simultaneously, for example, by fragmenting the polypeptide in a certain manner before analysis. A plurality of peptide fragments can comprise, in various embodiments, 2 or more peptides, 5 or more peptides, 10 or more peptides, 50 or more peptides, 100 or more peptides, 500 or more peptides, 1000 or more peptides, 5,000 or more peptides, 10,000 or more peptides, 50,000 or more peptides, 100,000 or more peptides, 500,000 or more peptides, or 1,000,000 or more peptides.

[0138] Therefore, in certain embodiments, peptide, polypeptide or protein can be fragmentation.Peptide, polypeptide or protein can be carried out fragmentation by any means known in the art, comprise that carry out fragmentation and chemical or physical fragmentation by proteolytic enzyme or endopeptidase.Fragmentation can be carried out by targeting using specific protease or endopeptidase, and described specific protease or endopeptidase combine and cut at specific consensus sequence place.In other embodiments, by using non-specific protease or endopeptidase, fragmentation is non-targeted or random.Non-specific protease can combine and cut at specific amino acid residue rather than consensus sequence place. The following proteases and endopeptidases, as well as others known in the art (e.g., Granvogl et al., Anal Bioanal Chem 389:991-1002, 2007), can be used to cleave proteins or polypeptides into peptide fragments: Proteinase K (a non-specific serine protease), TEV protease (cleaves at a specific consensus sequence), Trypsin, Chymotrypsin, Pepsin, Thermolysin, Thrombin, Factor Xa, Furin, Endopeptidase, Papain, Pepsin, Subtilisin, Elastase, Enterokinase, Genenase TM I, intracellular protease LysC, intracellular protease AspN, intracellular protease GluC and the like. Engineered proteases can also be used to fragment polypeptides before analysis, such as heat-labile versions of proteinase K that can be rapidly inactivated (see, for example, WO2019 / 17089). Proteinase K is also known to be stable in denaturing agents (such as urea and SDS), which enables digestion of partially or completely denatured proteins. Technicians can select proteases from the database based on the desired properties of protease (including specificity for a particular amino acid or amino acid sequence), which are referred to as protease substrates. Processed proteolytic databases known in the art can include the MEROPS database (available at: ebi.ac.uk / merops / ), the PANTHER database (available at: pantherdb.org), the BRENDA database (available at: brenda-enzymes.org), the TopFIND database (available at: topfind.clip.msl.ubc.ca) and the UniProt database (available at: uniprot.org).

[0139] Polypeptides can also be fragmented using chemical reagents. Chemical reagents for fragmenting polypeptides or proteins into smaller peptides are known in the art and include cyanogen bromide (CNBr; which hydrolyzes peptide bonds at the C-terminus of methionine residues), hydroxylamine, hydrazine, formic acid, BNPS-skatole [2-(2-nitrophenylsulfenyl)-3-methylindole], iodobenzoic acid, NTCB+Ni (2-nitro-5-thiocyanatobenzoic acid), and the like.

[0140] In certain embodiments, after enzymatic or chemical cleavage, the length of the resulting peptide fragment is about the same desired length, for example, 10 to 100 amino acids, 10 to 80 amino acids, 10 to 60 amino acids, 10 to 40 amino acids, 10 to 30 amino acids, 20 to 100 amino acids, 20 to 80 amino acids, 20 to 60 amino acids, 20 to 40 amino acids, 20 to 30 amino acids, 30 to 70 amino acids, 30 to 60 amino acids, 30 to 50 amino acids, or 15 to 40 amino acids. The cleavage reaction can be monitored in real time, for example, by incorporating a short test fluorescence resonance energy transfer (FRET) peptide containing a protease or endopeptidase cleavage site into a protein or polypeptide sample. A fluorescent group and a quencher group are attached to either end of the FRET peptide sequence comprising the cleavage site, and FRET between the quencher and the fluorophore results in low fluorescence. Once the test peptide is cleaved at the included cleavage site (e.g., by a protease or endopeptidase), the quencher and fluorophore separate, resulting in a measurable (and quantifiable) increase in fluorescence. This enables the cleavage reaction to be stopped at a certain fluorescence intensity, which provides a reproducible cleavage endpoint.

[0141] In some aspects, samples can be fractionated, wherein proteins or peptides are separated by one or more properties (such as cellular location, molecular weight, hydrophobicity, isoelectric point, or protein enrichment method) to reduce the complexity of the sample to be analyzed. Optionally, a subset of macromolecules (e.g., proteins) within the sample is fractionated so that the subset of macromolecules is separated from the rest of the sample. The sample can be fractionated before being attached to a support.

[0142] Alternatively or additionally, protein enrichment methods can be used to select specific proteins or peptides (see, e.g., Whiteaker et al., Anal. Biochem. 362:44-54, 2007) or to select specific post-translational modifications (see, e.g., Huang et al., J. Chromatogr. A 1372:1-17, 2014). Alternatively, affinity enrichment or selection can be performed on one or more specific classes of proteins for analysis—e.g., by exploiting the binding properties of such proteins. One class is immunoglobulins or specific immunoglobulin (Ig) isotypes. In the case of immunoglobulin molecules, analysis of the sequence and abundance or frequency of hypervariable sequences involved in affinity binding is of particular interest, particularly because they change in response to disease progression or are associated with health, immunity, and / or disease phenotypes.

[0143] Standard methods can also be used, including, for example, immunoaffinity methods, to remove overabundant proteins from samples. For plasma samples where more than 80% of the protein components are albumin and immunoglobulins, it may be useful to remove abundant proteins. Several commercial products are available for removing the overabundant proteins of plasma samples, including spin columns (Pierce, Agilent), or PROTIA and PROT20 (Sigma-Aldrich) that remove the top 2-20 plasma proteins.

[0144] Proteins, polypeptides, or peptides to be analyzed according to the methods described herein can be enriched prior to analysis. Methods for enriching the polypeptide of interest can include removing the polypeptide of interest from the sample (direct or positive enrichment) or removing or deducting other polypeptides from the sample (indirect or negative enrichment or depletion), or both. Enrichment can increase the efficiency of the disclosed methods, improve the dynamic range, and / or improve the ability to detect low-abundance polypeptides in complex samples. Enrichment methods can include removal of bulk species (not the intended target of the analysis), such as albumin; enrichment by specific targeting of a particular protein (e.g., by antibody or other affinity capture) (or deduction of non-targets by such capture); enrichment using one or more general properties of proteins (e.g., size, pI, hydrophobicity, etc.) (or deduction of non-targets using these properties); enrichment by targeting several classes of polypeptides (e.g., by post-translational modifications, such as phosphorylated and glycosylated proteins) (or deduction of non-targets); enrichment by the ability to bind certain molecules (e.g., DNA binding proteins); ATP binding proteins; enrichment / deduction by subcellular localization (e.g., nuclear, mitochondrial, Golgi apparatus, endoplasmic reticulum, etc.); enrichment by cell populations (e.g., T cells, B cells, etc.) that produce the target polypeptide, where the cells can be identified, sorted, or otherwise captured (e.g., by cell surface markers). Art-recognized methods and techniques for enrichment include centrifugation, chromatography, electrophoresis, binding, filtration, precipitation, and degradation. The dynamic range of a sample can also be adjusted by fractionating the sample using standard fractionation methods such as electrophoresis and liquid chromatography (Zhou et al., Anal Chem 84(2):720-734, 2012).

[0145] (III) Blocking undesirable chemical reactivity with protecting groups

[0146] Before or after enzymatic or chemical digestion of a polypeptide, a protecting group can be used to block reactive side chains or chemical groups to prevent side reactions, such as in a subsequent conjugation reaction. Such reactive groups include amino, carboxyl, sulfhydryl, hydroxyl, and guanidinyl side chains. A protecting group (PG) is a chemical moiety that protects or masks the reactive portion of a molecule to prevent side reactions in the reactive portion of the molecule while manipulating or reacting different portions of the molecule. After the manipulation or reaction is complete, the protecting group can be removed without degrading or decomposing the remainder of the molecule, i.e., the protected reactive portion of the molecule is "deprotected." Protecting groups (PGs) and their reactions are well known and can be selected by one skilled in the art based on compatibility with downstream chemistry (Isidro-Llobet et al., Chem Ref. 109(6):2455-2504, 2019; doi.org / 10.1021 / cr800323s; Protecting Groups 3rd ed. ISBN-13:978-1588902351). Many conventional protecting groups are known in the art, for example as described in McOmie, JFW, ed., Protective Groups in Organic Chemistry, Planum Press, 1973, Greene, TW and Wuts, PGM, Protective Groups in Organic Synthesis, John Wiley & Sons, 3rd ed., 1999, and Kocienski, P. Protecting Groups, 3rd ed., 2003, Georg Thieme Verlag (Americas). Examples of protecting groups include t-Boc, C1-6 acyl, Ac, Ts, Ms, silyl ethers such as TMS, TBDMS, TBDPS, Tf, Ns, Bn, Fmoc, dimethoxytrityl, methoxyethoxymethyl ether, methoxymethyl ether, pivaloyl, p-methoxybenzyl ether, tetrahydropyranyl, trityl, ethoxyethyl ether, carbobenzyloxy, benzoyl, etc. For example, the protecting group is an amine protecting group in some cases.

[0147] Protective groups can be added prior to enzymatic cleavage (this may require sacrificing the N-terminal and C-terminal component peptides, as the N-terminal amine and C-terminal carboxylic acid of the protein may also be protected) - this will render it inert to downstream reactions of the provided sequencing protocol. Protective groups can be added after enzymatic digestion, using specific protecting agents that will not protect the N-terminal functional group or the C-terminal functional group of the component peptide. Examples of such protecting groups are known in the art and are exemplified herein.

[0148] For example, in various embodiments, the docking chain is attached to NTAA via a terminal amine and isothiocyanate (ITC) group. The ITC can react with the alpha amine on NTAA, as well as the epsilon amino group on lysine and, to a lesser extent, the thiol on cysteine. These residues within the peptide to be analyzed can be capped / protected. For amines, NTAA and lysine amines can be capped / protected, and the NTAA cap can then be removed by Edman degradation to reveal the new amine. Lysine remains capped because the amide bond is not cleaved under Edman conditions.

[0149] Advantageously, the protecting group can be selected / designed such that, when interrogated as a TAA, the interaction between the fluorophore and the protected side chain is altered compared to the unprotected side chain. This can have a significant impact on the lifetime change, which can be taken into account when processing the signals used to train the database and, therefore, identify amino acids in test assays. Thus, embodiments are contemplated in which the analysis of a target polypeptide includes analysis with differently modified (e.g., protected) side chains and comparing the resulting changes in a spectral analysis.

[0150] (IV) Support surface for peptide immobilization

[0151] In the described peptide analysis methods, polypeptides, proteins, or peptides are immobilized on a surface (support surface) via one terminus (amino terminus or carboxyl terminus), and any peptide in a "run" (e.g., all peptides) are attached via the same terminus. Attachment to the support surface enables reliable interrogation of each individual feature (i.e., the location at which each peptide is immobilized), including through multiple steps or analysis cycles.

[0152] The polypeptide, protein or peptide may be directly or indirectly attached (covalently attached) to the support surface by any means known in the art. In some cases, it is desirable to use a support with a large loading capacity to immobilize a large number of (different) polypeptides.

[0153] Solid support surfaces to which proteins, peptides, and polypeptides can be attached are known in the art, as described, for example, in U.S. Patent Nos. 7,972,827, 10,852,305, 11,105,812, and 11,268,963, and published patent applications US2022 / 0155316, US2021 / 0396762, WO 2010 / 065531, and WO2016 / 069124.

[0154] The carrier surface can include any substrate of any size (e.g., glass, quartz, plastic, silicon, silicon oxide, ceramic, metal, metal oxide, alloy, or semiconductor) on which the biological sample (e.g., containing polypeptides or peptides) is placed (or arrayed) for analysis (thus also an "analytical substrate"). In various embodiments, the carrier surface can be a microscope slide, such as a standard 3" x 1" slide or a standard 75mm x 25mm slide. Additional examples of substrates include substrates used to assist in sample analysis, such as mass spectrometry platforms, such as SELDI and MALDI chips. The proteins, peptides, and polypeptides described herein can be applied to any type of analytical substrate typically used to detect the type of signal employed. In some embodiments, the carrier is a planar substrate.

[0155] In some embodiments, a three-dimensional carrier (e.g., a porous matrix or beads) is used to immobilize the polypeptide. In other embodiments, a carrier compatible with the signal detection method, sensor, and / or device to be used in the analysis is used to immobilize the polypeptide. Solid or semi-solid phase carriers can include surfaces such as glass, plastic, ceramic, and / or metal; particles such as nano-, micro-, or millimeter-scale particles composed of materials such as polystyrene, iron oxide, tentagel, glass, ceramic, and / or plastic; and substances of other shapes and forms. In certain embodiments, the carrier is a bead (or a collection of beads), such as polystyrene beads, polymer beads, polyacrylate beads, agarose beads, cellulose beads, dextran beads, acrylamide beads, solid beads, porous beads, paramagnetic beads, glass beads, silica-based beads, or controlled pore beads, or any combination thereof. In an embodiment, the carrier is a bead array. Although it is contemplated that beads can be used as a carrier surface, porous beads may allow for undesirable crosstalk between peptides. The loading density can be optimized to minimize this problem.

[0156] (V) Peptide Conjugation to Carrier Surface

[0157] A variety of reactions can be used to link the polypeptide to the surface of a support (eg, a solid support or a porous support). The polypeptide can be linked to the support directly or indirectly (eg, via a linker).

[0158] Various methods for attaching proteins to the surface of a solid support are known in the art. See, for example, Chan et al. (PloS One, 2(11):e1164, 2007, doi:10.1371 / magazine.pone.001165), Camarero and Kwon (IntJPeptide Res Therap. 14:351-357, 2008); Camarero et al. (JAm Chem Soc. 126(45):14730-14731, 2024); Kwon et al. (Angewandte Chem. 1996:1066-1074, 2008). Chemie), 45(11):1725-1729, 2006, doi.org / 10.1002 / anie.200503475); U.S. Patent Nos. 7,972,827, 10,852,305, 11,105,812, and 11,268,963; and published patent applications US 2022 / 0155316, US 2021 / 0396762, WO 2010 / 065531, and WO 2016 / 069124.

[0159] Exemplary reactions include the copper-catalyzed reaction of azides and alkynes to form triazoles (Huisgen 1,3-dipolar cycloaddition), strain-promoted azide-alkyne cycloaddition (SPAAC), reactions of dienes and dienophiles (Diels-Alder), strain-promoted alkyne-nitrone cycloaddition, reactions of strained alkenes with azides, tetrazines, or tetrazoles, [3+2] cycloadditions of alkenes and azides, inverse electron demand Diels-Alder (IEDDA) reactions of alkenes and tetrazines (e.g., m-tetrazine (mTet) or phenyltetrazine (pTet) and trans-cyclooctene (TCO)); or pTet and alkenes), photoreactions of alkenes and tetrazoles, Staudinger ligations of azides and phosphines, and various displacement reactions, such as displacement of leaving groups by nucleophilic attack on electrophilic atoms (Horisawa Frontiers in Physiology). Physiol.) 5:457, 2014, doi:10.3389 / fphys.2014.00457; Knall et al., Tetrahedron Lett. 55(34):4763-4766 2014, doi:10.1016 / j.tetlet.2014.07.002).

[0160] Exemplary displacement reactions include reactions of amines with activated esters; N-hydroxysuccinimide esters; isocyanates; isothiocyanates, aldehydes, epoxides, and the like. In some embodiments, iEDDA click chemistry is used to immobilize polypeptides onto supports because it is rapid and provides high yields at low input concentrations. In another embodiment, m-tetrazine is used in the iEDDA click chemistry reaction instead of tetrazine because m-tetrazine has improved bond stability. In another embodiment, phenyltetrazine (pTet) is used in the iEDDA click chemistry reaction. In one case, a polypeptide is labeled with a bifunctional click chemistry reagent, such as an alkyne-NHS ester (acetylene-PEG-NETS ester) reagent or an alkyne-benzophenone to generate an alkyne-tagged polypeptide. In some embodiments, the alkyne can also be a strained alkyne, such as cyclooctyne, including dibenzocyclooctyl (DBCO), and the like.

[0161] To interrogate multiple peptides on an immobilized surface, in some embodiments, the minimum distance between peptides is 2 nanometers or greater. The minimum distance may be affected by the resolution of the optical instrument used to read the fluorescence lifetime (or other signals from the IS, and by the local molecular environment). Methods for measuring the required minimum, optimized, or optimal distance between peptides on an immobilized surface to achieve optical detection of individual features and how to separate reads from these features (e.g., with the assistance of computer analysis of the signals) are well known in the art.

[0162] The spacing will be limited in particular by the resolution of the fluorescence detection. On widefield or confocal microscopes, the resolution is limited by the diffraction of the photons. This is proportional to the wavelength (λ) of the photons and the numerical aperture (NA) of the objective: resolution ≈ (λ) / (2xNA). When using super-resolution methods such as DNA-PAINT or dSTORM, this resolution can vary but will typically not exceed 10 nm. This is the minimum separation that can be achieved between peptides in the example of the array. This is also related to The FRET effect is on the same order of magnitude, where the donor and acceptor fluorophores or quenchers tend to interact when the separation between them is about 10 nm or less. Contributions from amino acid side chains on adjacent molecules may contribute to energy transfer (such as quenching or FRET) at separations less than 10 nm. However, if the distal contributions are sufficiently small (empirical evidence does not exist), then it may be possible to deconvolute or denoise the proximal (desired amino acid) and distal (neighboring amino acid crosstalk) effects of amino acids on fluorescence lifetime. In this case, peptide / oligonucleotide barcoding, substrate patterning or sparse / random labeling and / or fluorophore emission can support separations much smaller than 10 nm. In theory, both of these separation limitations can be addressed by using nanoscale wells or pits with a radius less than 5 nm, as only a single complex can occupy these wells or pits. In this case, the physical boundaries of the wells can suppress crosstalk, and single sensor-based detection at the base of each pit or well can address the diffraction-limited resolution problem. Nanoneedles or DNA origami may also be included. These can also provide a spacing of less than 10nm and can be mixed with discrete detection sensors similar to Pac Bio. Another possibility is to use an AFM cantilever to scan the sample or connect the molecule to a hollow nano-pyramid / pillar array and use a near-field fluorescence method. For AFM, NSOM flux may be lower and may reduce the spacing to about 50nm. See also Pan et al., Optics Communications (OpticsComm.) 445: 273-276, 2019, doi.org / 10.1016 / j.optcom.2019.04.053) and the description of the near-field scanning optical microscope provided by Olympus Life Sciences (Olympus Lifesciences) (available online at olympus-lifescience.com / en / microscope-resource / primer / techniques / nearfield / nearfieldintro / ).

[0163] Therefore, in certain embodiments where multiple polypeptides are immobilized on the same carrier, the polypeptides may be appropriately spaced apart to accommodate the method for performing the binding reaction and any downstream detection and / or analysis steps for evaluating the polypeptides. For example, for a signal detection step, it may be advantageous to optimally space the molecules apart. In some cases, the appropriate spacing depends on the type of signal generated and the detection method or sensor used to detect the signal. In some cases, the spacing of the targets on the carrier is determined based on the consideration that the signal generated in association with one polypeptide may be blurred or indistinguishable from the signal generated with an adjacent molecule. In some embodiments, the polypeptides are immobilized on the carrier and spaced apart at an optically resolvable distance.

[0164] In some embodiments, the surface of the support is blocked—that is, the surface is treated with a layer of material. Methods for blocking the surface include standard methods originally developed for fluorescent single molecule analysis, including blocking the surface with polymers such as polyethylene glycol (PEG) (Pan et al., Phys. Biol. 12:045006, 2015), polysiloxanes (e.g., Pluronic F-127), star polymers (e.g., star PEG) (Groll et al., Methods of Enzymology 12:045006, 2015), and the like.

[0010] Examples of such blocking agents include hydrophobic dichlorodimethylsilane (DDS) + self-assembled Tween-20 (Hua et al., Nat. Methods 11: 1233-1236, 2014), diamond-like carbon (DLC), DLC + PEG (Stavis et al., Proc. Natl. Acad. Sci. USA 108: 983-988, 2011), and zwitterionic moieties (e.g., US 2006 / 0183863). In addition to covalent surface modification, a variety of blocking agents can also be used, including surfactants such as Tween-20, polysiloxanes in solution (Pluronic series), polyvinyl alcohol (PVA), and proteins such as BSA and casein.

[0165] In various embodiments, when a protein, polypeptide, or peptide is immobilized on a solid substrate, the density of the protein, polypeptide, or peptide on the surface or within the bulk of the solid substrate can be titrated by incorporating competitors or "pseudo" reactive molecules.

[0166] The spacing of the immobilized polypeptides on the support can also be controlled by varying (titrating) the density of functional coupling groups used to attach the polypeptides (e.g., TCO or carboxyl (COOH)) to the substrate surface. In some embodiments, a plurality of molecules are spaced apart on the surface or within the volume of the support (e.g., a porous support) such that adjacent molecules are spaced apart by a distance of 50 nm to 500 nm, or 50 nm to 400 nm, or 50 nm to 300 nm, or 50 nm to 200 nm, or 50 nm to 100 nm. In some embodiments, a plurality of molecules are spaced apart on the surface of the support by an average distance of at least 50 nm, at least 60 nm, at least 70 nm, at least 80 nm, at least 90 nm, at least 100 nm, at least 150 nm, at least 200 nm, at least 250 nm, at least 300 nm, at least 350 nm, at least 400 nm, at least 450 nm, or at least 500 nm.

[0167] Thus, in some embodiments, appropriate spacing on the support is achieved by titrating the ratio of available linker molecules on the substrate surface. In some instances, the substrate surface (e.g., bead surface) is functionalized with carboxyl groups (COOH) and then treated with an activating agent (e.g., EDC and sulfo-NHS). In some instances, the substrate surface (e.g., bead surface) includes NHS moieties. In some embodiments, mPEG is added to the activated beads. n -NH2 and NH2-PEG n -mTet mixture (where n is any number, such as 1-100). mPEG3-NH2 (not available for coupling) can be titrated with NH2-PEG 24 In some embodiments, the ratio of the coupling moieties (e.g., NH2-PEG4-mTet) to the coupling moieties (e.g., NH2-PEG4-mTet) on the solid surface is at least 50 nm, at least 100 nm, at least 250 nm, or at least 500 nm. In some embodiments, the spacing of the polypeptides on the support is achieved by controlling the concentration and / or number of available COOH or other functional groups on the support.

[0168] (VI) Modification of the Free End of the Immobilized Peptide

[0169] Optionally, in some embodiments, the terminal amino acid of the polypeptide can be derivatized prior to conjugation of the docking strand (DS) to the terminal amino acid to achieve or enable conjugation. For example, in one embodiment, the terminal amino acid is NTAA, and the NTAA is derivatized with an Edman reagent, such as phenyl isothiocyanate (PITC).

[0170] Also contemplated are embodiments using an isoselenocyanate group in place of an isothiocyanate; this can be substituted to provide schemes similar to those provided herein.

[0171] 1-(4-Isothiocyanatophenyl)-1H-pyrrole-2,5-dione (CAS Reg. No. 60283-89-8, also known as maleimidophenyl isothiocyanate or MPITC; Keana et al., J Am Chem Soc. 108:7947-7963, 1986; commercially available from Chemieliva Biotech Co. Limited, Chongqing, China) described and produced herein offers beneficial properties in terms of allowing docking chain conjugation and Edman degradation capability (via maleimide and isothiocyanate groups, respectively). Other bifunctional linkers can only perform one of these functions, limiting their usefulness and requiring multiple linkers to replicate the functionality of MPITC.

[0172] The final product is obtained from the commercially available precursor 1-(4-aminophenyl)-1H-pyrrole-2,5-dione (CAS Reg. No. 29753-26-2). The synthesis envisions reacting the precursor's aniline group with carbon disulfide in the presence of DMAP and ET3N to produce an intermediate. The reaction is completed by reacting the high-energy intermediate with Boc2O under anhydrous conditions.

[0173] A representative reaction scheme is as follows:

[0174]

[0175]

[0176] The penultimate intermediate compound in the above scheme can be referred to as (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)phenyl)carbamothioic acid pivaloic acid thioanhydride. Additional intermediate compounds used in this process include, but are not limited to, (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)-2,6-dimethylphenyl)carbamothioic acid pivaloic acid thioanhydride; (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)-2,6-dihydroxyphenyl)carbamothioic acid pivaloic acid thioanhydride; and (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)-2,6-dimethoxyphenyl)carbamothioic acid pivaloic acid thioanhydride.

[0177] The central phenyl structure provides rigidity and thus fewer degrees of freedom between the conjugated modules, ensuring consistent interactions, a useful feature for kinetically sensitive signaling such as fluorescence and chemiluminescence.

[0178] Maleimides allow conjugation to a range of thiol- and amine-containing molecules, thus expanding their use beyond a single type of system to include other potential biomarkers such as antibodies, modified oligonucleotides, various peptides, proteins, etc. At the same time, the mild nature of the synthesis allows other coupling groups to be attached to the linker, including DBCO, NHS esters, TCO, azide, etc., without interfering with ITC formation.

[0179] The MPITC linker can optionally be modified for FLIM use by adding electron-donating groups at the ortho position of the ITC. The resulting enhanced electron density may allow for faster Edman degradation, thereby increasing the efficiency of the workflow. A scheme with several examples is as follows:

[0180]

[0181] a) 1-(4-isothiocyanatophenyl)-1H-pyrrole-2,5-dione;

[0182] b) 1-(4-isothiocyanato-3,5-dimethylphenyl)-1H-pyrrole-2,5-dione;

[0183] c) 1-(3,5-dihydroxy-4-isothiocyanatophenyl)-1H-pyrrole-2,5-dione; and

[0184] d) 1-(4-Isothiocyanato-3,5-dimethoxyphenyl)-1H-pyrrole-2,5-dione.

[0185] In another embodiment, peptide analysis (including peptide sequencing) described herein can be performed with C-terminal attachment of an oligonucleotide docking strand (DS) and removal of amino acids by C-terminal degradation. First, the N-terminus of the peptide is attached to a substrate using standard amine coupling reagents / protocols.

[0186] A C-terminal thiohydantoin moiety is formed and then S-alkylated using a DS carrying a leaving group, including but not limited to acyl halides (e.g., chlorides, bromides, iodides, etc.) and tosylate. This alkylation constitutes the addition of a DS to the C-terminus, which enables the application of a FLIM workflow for interrogating the C-terminal amino acid (CTAA) using a sequential imaging strand (IS)-fluorophore. Interrogation of the CTAA side chain is performed in the same manner as interrogation of NTAA using an N-terminally attached DS. The IS-fluorophore conjugate is attached so that the fluorophore is in proximity to the CTAA or "C-1" amino acid side chain, as determined by the linker composition and length; the fluorescence lifetime of the fluorophore is measured; the IS-fluorophore is removed by heating, chemical denaturation, or a combination of both; another IS-fluorophore in which the IS, fluorophore, or both are different is attached; and the interrogation workflow is cycled until a sufficient amount of modulated fluorescence lifetime data for the CTAA or "C-1" amino acid is acquired.

[0187] To remove CTAA and form the C-terminus from the "C-1" position amino acid, Schlack-Kumpf degradation was performed (protocol provided below; see also Li and Liang, Analytical Biochemistry 302(1):108-0113, doi.org / 10.1006 / abio.2001.5505). This cleavage reaction reforms the peptide-thiohydantoin at the C-terminus of the N-terminally immobilized peptide. After the original "C" position amino acid CTAA is cleaved from the peptide, the "C-1" position amino acid constitutes the newly formed C-terminus and can then be referred to as CTAA.

[0188] In this method, the C-terminus is first activated and converted into a thiohydantoin moiety. The C-terminal carboxylic acid is converted into an anhydride group by combining with acetic anhydride at 50°C to 80°C for 5 minutes before adding 0.5M triphenylgermanium isothiocyanate (Ph3Ge-ITC) or a similar highly substituted analog in acetonitrile. The peptide-thiohydantoin is alkalized with reagents that may include triethylamine, sodium bicarbonate, and sodium borate, and then combined with an oligonucleotide docking strand (DS) modified with a leaving group (such as an acyl chloride or other). The leaving group can be located at the 5' end, the 3' end, or within the DS, depending on which interrogation scheme is to be used. This results in conjugation of the DS to the C-terminus of the peptide via the thiolate moiety (Boyd et al., Analytical Biochemistry 206(2):344-352, 1992, doi.org / 10.1016 / 0003-2697(92)90376-i).

[0189] After fluorescence lifetime data is acquired, C-terminal degradation is performed under acidic conditions by adding hydrogen isothiocyanate (or isothiocyanate anion), which can be generated by donors including (trimethylsilyl)isothiocyanate. This results in cleavage of the CTAA-DS conjugate from the peptide. This cleavage reaction reforms the thiohydantoin at the C-terminus of the peptide from the original "C-1" position amino acid (see scheme below).

[0190]

[0191] Example reaction scheme for the attachment of an oligonucleotide docking strand (DS) to the C-terminus via a thiolate moiety and subsequent C-terminal degradation and removal of CTAA. Attachment of the DS to the peptide can be at the 3' end, the 5' end, or internally. Once cleavage of the CTAA is performed, the thiohydantoin moiety is reformed, and the DS can be attached to the newly formed C-terminus via an alkylation reaction without the need for reactivation of the C-terminus. This schematic does not show attachment of the peptide to the substrate via its N-terminal amine or repeated interrogation (via IS-fluorophore bonding and fluorescence lifetime data acquisition).

[0192] While methods using isothiocyanates (ITC, RN=C=S) are exemplified herein, ITC analogs such as isoselenocyanates (ISC, RN=C=Se) are also contemplated. Thus, variants of the peptide-to-docking linker are contemplated that include an ISC group in place of the ITC, and Edman degradation can be performed by the ISC rather than the ITC.

[0193] In embodiments of peptide analysis / sequencing workflows (exemplified by FLIM), including embodiments using N-terminal (or C-terminal) degradation methods, isothiocyanate analogs can be used instead. These analogs include isoselenocyanates (ISCs) (Maeda et al., Heterocycles 82(2):2010doi.org / 10.3987 / com-10-s(e)116; Iskierko et al., Ann Univ MariaeCurie Sklodowska Med (in Boer), 3169-76, 1976). Using such analogs can be beneficial because they exhibit different levels of reactivity and can therefore be designed to vary under different reaction conditions used throughout the FLIM workflow. This promotes the use of milder conditions, which can include shorter reaction times, lower temperatures, and lower extreme pH values. Milder conditions can enhance the relative stabilities of structures (peptides, etc.) and attachments (peptide to surface, etc.) used in the workflow, which can improve data fidelity.

[0194] Also provided herein are compounds of formula (I):

[0195]

[0196] or a salt or solvate thereof, wherein:

[0197] x is 0, 1, or 2;

[0198] Each R is independently selected from the group consisting of: C1-C6 alkyl, -NO2, halogen, -C=OR, -C=SR, -C=ONR, -C=OOR, -SO3 or any other common electron withdrawing group;

[0199] R 1 and R 2 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, —O—(C1-C6 alkyl), C1-C6 alkyl, hydroxy, halogen, —O-alkyl, —S-alkyl, —OC(═O)R, —N—(C═O)—R, —OC(═O)OR, —NC(═S)NR, —N—(C═O)—OR, or any other common electron-donating group;

[0200] y is 0, 1, 2, or 3; and

[0201] Each R 3 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group.

[0202] Also provided are compounds of formula (II):

[0203]

[0204] or a salt or solvate thereof, wherein:

[0205] x is 0, 1, or 2;

[0206] Each R is independently selected from the group consisting of: C1-C6 alkyl, -NO2, halogen, -C=OR, -C=SR, -C=ONR, -C=OOR, -SO3 or any other common electron withdrawing group;

[0207] R 1 and R 2independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, —O—(C1-C6 alkyl), C1-C6 alkyl, hydroxy, halogen, —O-alkyl, —S-alkyl, —OC(═O)R, —N—(C═O)—R, —OC(═O)OR, —NC(═S)NR, —N—(C═O)—OR, or any other common electron-donating group;

[0208] y is 0, 1, 2, or 3; and

[0209] Each R 3 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group.

[0210] Also contemplated are compounds of formula (I) or formula (II) in which the six-membered ring (as shown below) may be replaced by other aromatic groups (e.g., naphthyl, other benzo-fused rings), heteroaryl groups, heterocyclyl groups, or alkyl groups, each of which may be unsubstituted or substituted.

[0211]

[0212] (VII) Signaling molecules

[0213] In the provided polypeptide analysis methods, individual terminal amino acids are distinguished based on their effect on the local environment (at the molecular level). Based on this effect, the effect is detected and distinguished by using a signal compound / molecule that exhibits a detectable / measurable change in a measurable property. For example, the signal compound has one or more spectral properties that are affected by proximity to one or more amino acid side groups—and the effect is detected by analyzing and comparing the spectral properties.

[0214] In one embodiment, the protein analysis method includes comparing one or more spectral characteristics of a signal molecule near the terminal amino acid (e.g., NTAA or CTAA) of an immobilized polypeptide with a set of reference spectral characteristics of the interaction of the signal molecule when near the known terminal amino acid. In one embodiment, the spectral characteristics include a fluorescent signal from a fluorophore, luminescence from a luminescent molecule, and phosphorescence from a phosphorescent molecule. Luminescent molecules do not require heat (i.e., photons) for excitation, and the lifetime of phosphorescent molecules is typically longer than that of fluorescent molecules (1 μs to several seconds). In the current photoelectron case, the sensitivity of FLIM instruments is about 0.05-0.3 ns, which limits the detection of small differences in amino acids in this embodiment. Molecules with longer lifetimes may increase sensitivity; however, this may increase the overall measurement time.

[0215] In one embodiment, suitable signal compounds exhibit different spectral properties when approaching different N-terminal amino acids. For example, it is shown herein that various fluorescent dyes exhibit predictable and variable spectral properties (represented herein by fluorescence lifetime) when approaching different amino acid residues.

[0216] As specific examples, it is shown herein that the following fluorophores (each of which is commercially available) are able to "detect" differences in proximal amino acids based on changes in their fluorescence lifetimes as measured by FLIM.

[0217] AF488 (lifetime = 4.1ns)

[0218]

[0219] BODIPY FL (lifetime = 5.87ns)

[0220]

[0221] BODIPY TR (lifetime = 5.7ns)

[0222]

[0223] TAMRA (lifetime = 5.7ns)

[0224]

[0225] KU530-6 (lifespan = approximately 24ns)

[0226]

[0227] KU530-R-4 (lifespan = approximately 24ns)

[0228]

[0229] KU560-6 (lifetime = about 20ns)

[0230]

[0231] KU560-R-4 (lifespan = approximately 20ns)

[0232]

[0233] Thus, the signal molecule may be a fluorescent dye, and the spectral characteristic may be the fluorescence lifetime of the dye, and the change in proximity to different peptide-terminal amino acids. Fluorescent dyes can be detected in real time with high resolution, and many fluorescent dyes can have different excitation and emission wavelengths. A panel of fluorescent dyes can be selected so that more than one dye is detected simultaneously in the same reaction, for example as a way of controlling for complete washing (removal) of a previously interrogated IS. For example, the following dyes can be detected and distinguished simultaneously: Cy3, Cy5, FAM, JOE, TAMRA, ROX, dR110, dR6G, dTAMRA, and dRox. Any of these dyes can be used alone or in any combination to practice the embodiments herein.

[0234] Dyes can allow single molecule detection. A large number of fluorescent dyes have been synthesized and are commercially available in different forms (see, for example, compounds available from Invitrogen). This can include fluorescent dyes having a linker region and a hydrazine group that allows coupling to nucleic acids in a reaction with a dialdehyde group.

[0235] In some embodiments, 2-3 fluorophores are sequentially conjugated to the 5' end of the IS and act as a FRET pair. Due to the fixed rigid distance (10 angstroms to 10 nm) between each donor and acceptor, more than one amino acid residue along the peptide sequence is interrogated and the fluorescence lifetime of all fluorophores is collected. When performing FLIM or another peptide sequencing method provided herein, the lifetime can vary based on the proximal amino acid and the nearby FRET donor / acceptor, thus providing more information about the environment.

[0236] The present disclosure is not limited to the use of specific fluorescent dyes, but different dyes can be applied to achieve the same effect. In fact, this paper shows a library of two or more different fluorescent molecules to better characterize amino acids. This paper demonstrates this by sequentially interrogating the same terminal amino acid with more than one signal molecule (e.g., more than one fluorophore), each of which is connected to an IS that is coupled in series to a DS connected to an immobilized peptide. A different readout for each interaction can be used to improve the fidelity of amino acid identification.

[0237] Non-limiting examples of signal molecules can include 5-FAM (also known as 5-carboxyfluorescein; 6-carboxy-4',5'-dimethylfluorescein (also known as spiro[isobenzofuran-1(3H),9'-(9H)xanthene]-5-carboxylic acid, 3',6'-dihydroxy-3-oxo-6-carboxyfluorescein; Cdmfda); 5-hexachloro-fluorescein; ([4,7,2',4',5',7'-hexachloro-(3',6'-dipivaloyl-fluoresceinyl)-6- carboxylic acid]); 6-hexachloro-fluorescein; ([4,7,2',4',5',7'-hexachloro-(3',6'-dipivaloylfluoresceinyl)-5-carboxylic acid]); 5-tetrachloro-fluorescein; ([4,7,2',7'-tetrachloro-(3',6'-dipivaloylfluoresceinyl)-5-carboxylic acid]); 6-tetrachloro-fluorescein; ([4,7,2',7'-tetrachloro-(3',6'-dipivaloylfluoresceinyl)-6-carboxylic acid]); 5-TAMRA (5-carboxytetramethylrhodamine); 9-(2,4-dicarboxyphenyl)-3,6-bis(dimethylamino)xanthene cation; 6-TAMRA (6-carboxytetramethylrhodamine); 9-(2,5-dicarboxyphenyl)-3,6-bis(dimethylamino); EDANS (5-((2-aminoethyl)amino)naphthalene-1-sulfonic acid); 1,5-IAEDANS (5-((((2-iodoacetyl)amino)ethyl)amino)naphthalene-1-sulfonic acid); Cy5 (indoledicarbocyanine-5); Cy3 (indole-dicarbocyanine-3); and BODIPY FL (2,6-dibromo-4,4-difluoro-5,7-dimethyl-4-bora-3a,4a-diaza-s-dicyclopentaphenacene-3-propionic acid); Quasar TM -670 dye (Biosearch Technologies); Cal Fluor TM Orange dye (Biosearch Technologies); Rox dye; Max dye (Integrated DNA Technologies), tetrachlorofluorescein (TET), 4,7,2'-trichloro-7'-phenyl-6-carboxyfluorescein (VIC), HEX, Cy3, Cy 3.5, Cy 5, Cy 5.5, Cy7, tetramethylrhodamine, ROX and JOE and suitable derivatives thereof. The label can be Alexa Fluor Dyes, such as Alexa 350, 405, 430, 488, 532, 546, 555, 568, 594, 633, 647, 660, 680, 700 and 750. The markers can be Cascade Blue, Marina Blue, Oregon Green 500, Oregon Green 514, Oregon Green 488, Oregon Green 488-X, Pacific Blue, Rhodamine Green, Rhodol Green, Rhodamine Green-X, Rhodamine Red-X and Texas Red-X. The markers can be a group of longer-lived dyes, such as those from KU TM Dye family (KU450, KU470, KU483, KU500, KU510, KU530, KU542, KU560, KU600, KU600, KU625, KU-P and KU-T) (KU dyes, Department of Chemistry, University of Copenhagen, Denmark). The label can be located at the 5' end of the probe, at the 3' end of the probe, at both the 5' and 3' ends of the probe, or inside the probe. A distinguishable (e.g., unique) label can be used to detect each different locus in the experiment, such as the two ends of a target polynucleotide (e.g., mRNA).

[0238] Non-limiting examples of dye-hydrazides that can be used as signaling molecules include Alexa Fluor TM -Hydrazide and its salts, 1-pyrenebutyric acid-hydrazide, 7-diethylaminocoumarin-3-carboxylic acid-hydrazide (DCCH), Cascade Blue TM Hydrazide and its salts, biocytin-hydrazide, 2-acetamido-4-mercaptobutyric acid-hydrazide (AMBH), BODIPY TM FL-hydrazide, biotin-hydrazide, TexasRed TM -hydrazide, biocytin-hydrazide, luminol (3-aminophthalhydrazide), and Marina Blue TM Hydrazide. Non-limiting examples of dyes that can be used for labeling include 5-dimethylaminonaphthalene-1-(N-(2-aminoethyl))sulfonamide (dansylethylenediamine), Cascade Blue TMEthylenediamine and its salts, N-(2-aminoethyl)-4-amino-3,6-disulfo-1,8-naphthalimide (fluorescein yellow ethylenediamine) and its salts, N-(biotinyl)-N'-(iodoacetyl)ethylenediamine, N-(-2-aminoethyl) biotinamide, hydrobromide (biotin ethylenediamine), 4,4-difluoro-5,7-dimethyl-4-borane-3a,4a-diaza-s-dicyclopentaphenone-3-propionylethylenediamine and its salts (BODIPY TM FL EDA), Lissamine TM Rhodamine B ethylenediamine and DSB-X TM Biotin ethylenediamine (dethiobiotin-X ethylenediamine, hydrochloride).

[0239] Non-limiting examples of dye-cadaverines that can be used as signal molecules include 5-dimethylaminonaphthalene-1-(N-(5-aminopentyl))sulfonamide (dansylcadaverine), 5-(and-6)-((N-(5-aminopentyl)amino)carbonyl)tetramethylrhodamine (tetramethylrhodaminecadaverine), N-(5-aminopentyl)-4-amino-3,6-disulfo-1,8-naphthalene dicarboximide and its salts (fluorescein cadaverine), N-(5-aminopentyl)biotinamide and its salts (biotin cadaverine), biotin-X cadaverine (5-((N-(biotinyl)amino)hexanoyl)aminopentylamine and its salts, Texas Red TM Texas Red TM C5), 5-(((4-(4,4-difluoro-5-(2-thienyl)-4-boran-3a,4a-diaza-s-dicyclopenta-benzophen-3-yl)phenoxy)acetyl)amino)pentylamine and its salts (BODIPY TM TR cadaverine), Oregon Green TM Cadaverine, Alexa Cadaverine and 5-(5-aminopentyl)thioureido)fluorescein and its salts (fluorescein cadaverine).

[0240] Alternative spectroscopic techniques such as anisotropy or fluorescence polarization can be used to enhance the differentiation of amino acid residues along a peptide. Amino acids in close proximity to a fluorophore may alter the spatial orientation of the fluorophore on the IS. This can affect the emission direction from being uniform and isotropic to being directionally dependent or anisotropic. These spectroscopic techniques can be combined with existing FLIM techniques discussed in the art by polarizing the excitation light and increasing the number of detectors within the current optical path. Thus, by collecting both spatial and temporal information (i.e., anisotropy and fluorescence lifetime) from the fluorophore, the sensitivity of determining amino acid differences can be increased, particularly when determining post-translational modifications.

[0241] (VIII) Imaging chain and its connection with signaling molecules

[0242] In the provided peptide sequencing methods, an ssDNA imager strand (IS) with attached signal molecules is used, wherein the IS is ligated to a single-stranded docking site (DS) to form a double-stranded DS:IS complex. The IS carries one or more signal molecules, which, depending on the embodiment, are covalently attached to an IS oligonucleotide at, near, or within one or both ends.

[0243] The designed ssDNA IS of the selected sequence has, for example, a signal tag attachment modification on the 5' end of the oligonucleotide. The signal tag attachment modification is customized based on the type of signal molecule to be attached to the IS to achieve the attachment chemistry. In one example, this modification is an alkylamine, such as (6-aminohexyl) phosphate, covalently conjugated to the maleimide group of a maleimide-functionalized fluorophore; this is for example in Figure 2 and 3 Shown in.

[0244] The length of the IS may be affected by the melting temperature of the DS:IS complex. Designing an IS with a certain guanine and cytosine content and adjusting the environmental salt conditions can help produce the desired IS. In various embodiments, the minimum length of the IS can be 8-10 base pairs. The maximum length of the IS may be affected by the upper limit of the melting temperature of the DS:IS complex (e.g., a constraint). The DNA hybridization melting temperature (T m ) are affected by the salt concentration in the reaction buffer and wash solution. If PAINT technology is used, lower salt, higher temperature, and / or cosolvents can modulate these dynamics. Similarly, dyes with low water solubility can benefit from the use of certain solvents (DMSO, DMF, PEG). Washing to remove imager strands can be improved by using detergents (anionic, zwitterionic, cationic).

[0245] In FLIM, many different dyes (signal molecules) with different spectral properties (excitation / emission wavelengths, lifetimes, etc.) can be used. The lifetime of each dye may vary depending on the TAA (e.g., NTAA) close to it. Therefore, a set of dyes can provide a unique signal for each NTAA. By limiting the movement or freedom of the dye around the NTAA, better sensitivity and resolution are achieved. Using internally modified bases to connect the docking chain and using imaging chains to which the dye is connected through internal base modifications can help to position the dye and NTAA along the major or minor groove of the DNA. This may result in higher reproducibility or more consistent lifetime effects.

[0246] Molecular dynamics simulations were performed to model the interaction of the fluorophore of the DNA imaging strand (IS)-fluorophore conjugate with the side chain of the N-terminal amino acid (NTAA). Based on the propensity of the NTAA side chain to interact with the fluorophore, the length and chemical composition of the various linkers involved in the workflow were optimized ( Figure 3 Similarly, variations in the IS sequence were modeled, which could result in overhangs or non-overhangs relative to the DNA docking strand (DS). Figure 2 For some NTAAs, the change in fluorescence lifetime of a specific fluorophore can be predicted in silico based on the proportion of time the amino acid side chain interacts with the fluorophore. For example, Alexa 488 (AF488) interacts with the side chain of the N-terminal tryptophan group more than with arginine or glycine in the simulation, so those skilled in the art can predict that for the modeled combination of IS and fluorophore, the N-terminal tryptophan residue is expected to suppress the fluorescence lifetime of AF488 more than the N-terminal arginine or glycine residue.

[0247] Also contemplated are libraries of imaging strands that can contain, for example: a set of ISs having the same primary nucleotide sequence, but differing in that each IS includes a different signaling molecule (the signaling molecules can all be of one type, e.g., all fluorophores where FLIM is used as a measure of the altered proximal environment; or of different types or detectable signals; etc.); a set of ISs that are diverse in terms of primary nucleotide sequence (e.g., different lengths, different primary sequences, different sequences but with equal purine / pyrimidine content ratios, one or more modified nucleotides, different overhangs or no overhangs compared to the homologous DS, etc.), but each IS binds the same signaling molecule; etc.

[0248] We also considered using unlabeled IS chains with dye-labeled minor groove binders (MGBs). MGBs have a sequence-specific (primarily A:T) design with DS linked to NTAA, and the surrounding sequence context allows for the positioning of the dye-MGB, allowing for understanding of lifetime changes. In this example, the dye is non-covalently linked to the IS.

[0249] Using DNA docking strands and imaging strands, the sequencing methods described herein are based on super-resolution imaging (changes in one or more spectral properties of a signal molecule, which are affected based on the proximity to different terminal amino acids of an immobilized peptide). It may be advantageous to design an IS that can exchange or dehybridize at a slower rate than fluorescence lifetime measurement to reduce or minimize signal attenuation caused by photobleaching. DNA-PAINT (Civitci et al., Nature Comm. 11, Article No. 4339, 2020, doi.org / 10.1038 / s41467-020-18181-6) can be used with the provided peptide sequence methods.

[0250] Methods for preparing both natural and modified (non-naturally occurring) oligonucleotides are well known in the art, and the oligonucleotides can be used to prepare the imaging strand oligonucleotides used in the methods disclosed herein. Oligonucleotides are connected by internucleotide bonds, which refer to the chemical bonds between the two nucleoside moieties. Modification of the phosphate backbone of a DNA or RNA oligonucleotide can increase binding affinity or stability of the oligonucleotide, or reduce the susceptibility of the oligonucleotide to nuclease digestion. Cationic modifications include, but are not limited to, diethyl-ethylenediamide (DEED) or dimethyl-aminopropylamine (DMAP), which may be particularly useful due to reducing the electrostatic repulsion between the oligonucleotide and the target. Modification of the phosphate backbone can also include a sulfur atom replacing one of the non-bridge oxygens in the phosphodiester bond. This substitution produces a thiophosphate internucleoside bond instead of a phosphodiester bond. Oligonucleotides containing thiophosphate internucleoside bonds have been shown to be more stable in vivo.

[0251] Examples of modified nucleotides with reduced charge include modified internucleotide linkages such as phosphate analogs with achiral and uncharged intersubunit linkages (e.g., Sterchak et al., Organic Chem., 52:4202, (1987)), and uncharged morpholino-based polymers with achiral intersubunit linkages (see, e.g., U.S. Patent No. 5,034,506). Some internucleotide linkage analogs include morpholine, acetal, and polyamide-linked heterocycles.

[0252] In another embodiment, the oligonucleotide is composed of a locked nucleic acid, or may contain at least one locked nucleic acid. Locked nucleic acids (LNA) are modified RNA nucleotides (see, for example, Braasch et al., Chem. Biol., 8(1): 1-7, 2001). LNA forms hybrids with DNA that are more stable than DNA / DNA hybrids and have properties similar to those of peptide nucleic acid (PNA) / DNA hybrids. Therefore, LNA can be used like PNA molecules. In some embodiments, LNA binding efficiency can be increased by adding a positive charge thereto. Commercial nucleic acid synthesizers and standard phosphoramidite chemistry can be used to prepare LNA.

[0253] In certain embodiments, oligonucleotide includes peptide nucleic acid. Peptide nucleic acid (PNA) is a synthetic DNA mimic, and the phosphate backbone of oligonucleotide is (for example, integrally) replaced by repeating N-(2-aminoethyl)-glycine unit, and phosphodiester bond is usually replaced by peptide bond. Various heterocyclic bases are connected to the backbone by methylene carbonyl bond. PNA maintains the interval similar to conventional DNA oligonucleotide of heterocyclic base, but is achiral and neutrally charged molecule. Peptide nucleic acid comprises peptide nucleic acid monomer.

[0254] Other main chain modifications include peptide and amino acid variations and modifications. Therefore, the main chain components of oligonucleotides such as PNAs can be peptide bonds, or alternatively they can be non-peptide peptide bonds. Examples include acetyl caps, such as 8-amino-3,6-dioxaoctanoic acid and other amino spacers (referred to herein as O-linkers). If a positive charge is required in PNAs, amino acids such as lysine are particularly useful. Methods for chemically assembling PNAs are well known. See, for example, U.S. Patents No. 5,539,082, No. 5,527,675, No. 5,623,049, No. 5,714,331, No. 5,736,336, No. 5,773,571 and No. 5,786,571.

[0255] Oligonucleotide optionally includes one or more terminal residues or modifications at either or both ends to increase the stability and / or affinity of the oligonucleotide to its target. Commonly used positively charged moieties include amino acids lysine and arginine, although other positively charged moieties may also be useful. Propylamine groups may be used to further modify the oligonucleotide to add caps, thereby preventing degradation. The procedures for 3' or 5' capping oligonucleotides are well known in the art.

[0256] (IX) Docking Chain, and Connection of the Docking Chain to the Immobilized Peptide

[0257] In an embodiment, the docking strand (DS) is a synthetic single-stranded DNA (ssDNA) having a selected sequence and a modification (e.g., at the 3' end of the oligonucleotide) that enables the ssDNA to be covalently linked to the free terminal amino acid of the immobilized protein, polypeptide, or peptide. In one example, the modification is an alkylthiol group, such as (3-mercaptopropyl) phosphate, which is capable of covalently linking to the N-terminal maleimidyl group of the surface-conjugated MPITC peptide. Although examples are provided, embodiments using other DNA oligonucleotide modifications are also contemplated, including embodiments in which modifications are made at the 5' end, within the oligonucleotide, or at the 3' end. Internal modifications are generally limited to dT, abasic, or spacers. Recognized and commercially available types of oligonucleotide modifications are available for use; see, for example, information provided online by Integrated DNA Technologies (see, for example, resources available at idtdna.com / pages / products / custom-dna-rna / oligo-modifications).

[0258] The DS used in the provided methods can be "universal" in that the oligonucleotide can bind to any / all peptides in an experiment, regardless of the primary sequence of the peptide. This versatility is beneficial because it simplifies the analytical system, as different docking chains (or other binding moieties) do not need to be associated / labeled with the immobilized peptide. The universal DS is independent of the primary sequence of the peptide to be analyzed and can bind (covalently link) to any terminal amino acid present on each immobilized peptide.

[0259] The DS is (chemically) linked to the TAA of a peptide / polypeptide immobilized on the support surface (or one DS is linked to each TAA of an array of peptides / polypeptides). The single-stranded region of the DS is available to allow binding to the IS so that the signaling molecule (e.g., fluorophore) linked to the IS is accessible to the TAA and its side chain.

[0260] The DS sequence can be the same between each cleavage cycle, or can change to a different / partially different sequence. The DS can be used in FLIM measurements to fingerprint amino acids.

[0261] Provided herein are examples of joints that can be used to connect docking chains to immobilized peptides for use in protein / peptide analysis (sequence) methods. Typically, these joints can be imagined as three-part structures, with a "DNA connection" function on one end and a "peptide connection" function on the other end, as well as some chemical bridges or other parts that connect the two functional elements together. The "DNA connection" portion is characterized in that it has a chemical structure that can (covalently) be connected to a DNA molecule (particularly ssDNA DS), while the "peptide connection" portion has a chemical structure that can (covalently) be connected to the end of a peptide molecule. According to an embodiment, the "peptide connection" portion of the joint can be connected to the N-terminus of the peptide; starting from the C-terminus of the peptide. Similarly, according to an embodiment, the "DNA connection" portion of the joint can be connected to the 5' end of the oligonucleotide; the 3' end of the oligonucleotide; or a site within the internal sequence of the oligonucleotide. Additional details and options are provided herein.

[0262] Sample barcodes, or UMIs, can be incorporated into docking strands and decoded by sequential hybridization. Therefore, DS can also be used to encode other information, such as sample or spatial barcodes. For sample barcodes, DS is added to the peptide before immobilization on the surface. Different DS sequences are attached to different samples. They are then mixed together and immobilized on the surface. Labeled imaging strands bind to complementary DSs and report the location of individual peptides and the sample ID. If the number of samples is greater than the number of dye colors, the DS can have multiple IS binding sequences, and the sequential binding of isA and isB can be demultiplexed. 3 colors, 2 IS sites = 3^2 = 9 sample barcodes, similar to MER-fish. Ideally, the barcode DS would be cleaved and then ligated to a FLIM DS for sequencing. Spatial barcodes are similar to sample barcodes, but have more combinations. 4^8 = 65K barcodes. Unlabeled ISs can be "colors." True UMIs can be difficult to decode because there are so many molecules.

[0263] Methods of preparing both natural and modified (non-naturally occurring) oligonucleotides are well known in the art and can be used to prepare docking strand oligonucleotides for use in the methods of the present disclosure.Exemplary methods are provided herein.

[0264] (X) Binding of labeled imager strands to docking strands on immobilized peptides

[0265] The binding of the IS oligonucleotide (which acts as an ssDNA molecule) to the DS oligonucleotide (which acts as an ssDNA molecule) relies on conventional DNA double-strand hydrogen bonding. Therefore, binding occurs under conditions recognized in the art that allow the two ssDNA strands to pair, such that hydrogen bonds form between the bases adenine and thymine to form an AT base pair, and between the bases guanine and cytosine to form a GC base pair.

[0266] One of ordinary skill in the art will appreciate that conditions can be varied to affect bonding affinity, including altering salt and other buffer conditions, temperature, etc. The primary sequence of the two strands, including modifications and overhangs or lack thereof, can also affect bonding.

[0267] The melting temperature of the DS:IS complex will be affected by the length of the IS and DS oligonucleotides as well as the guanine and cytosine content and the environmental salt and other buffer conditions used in the association (and dissociation). m ) are affected by the salt concentration in the reaction buffer and wash solution. If PAINT technology is used, lower salt, higher temperature, and / or cosolvents can modulate these dynamics. Similarly, dyes with low water solubility can benefit from the use of certain solvents (DMSO, DMF, PEG). Washing to remove imager strands can be improved by using detergents (anionic, zwitterionic, cationic).

[0268] It is also known that lifetime effects can be influenced by solvent conditions; see, for example, Boens et al. (Analytical Chemistry 79(5):2137-2149, 2007, doi:10.1021 / ac062160k). If DNA-PAINT technology is used, lower salt, higher temperature, and / or cosolvents can modulate these dynamics. Similarly, dyes with low water solubility can benefit from the use of some solvents (DMSO, DMF, PEG). Washing can be improved by using detergents (anionic, zwitterionic, cationic) to remove imager strands.

[0269] (XI) Detection of imaging chain signals

[0270] Various statistical methods known in the art can be used to compare spectra of signal molecules near different amino acids in order to identify the closest match and thereby identify the terminal amino acid (such as the N-terminal amino acid) of the immobilized polypeptide.

[0271] In one embodiment, a suitable method generates a quantitative measure of the similarity or difference between a measured spectrum (e.g., fluorescence lifetime) and a reference spectrum (e.g., fluorescence lifetime), such as a spectrum obtained using the same signal molecule at a known amino acid near the terminal position of the immobilized polypeptide.

[0272] In various embodiments, the method further comprises generating a statistical measure or probability score that the spectrum indicates the presence of a particular terminal amino acid (e.g., NTAA or CTAA, depending on the embodiment and method) adjacent to the signal molecule. In one embodiment, the method used herein for comparing the measured spectral characteristics of a signal molecule near a test terminal amino acid to a reference / control terminal amino acid uses one or more probabilistic algorithms. For example, a probabilistic algorithm can be trained to identify different N-terminal amino acids by correlating specific spectra with specific known N-terminal amino acids. As demonstrated herein, the N-1 and N-2 amino acids may also affect the spectrum, and therefore this can be taken into account when developing reference / control measurements.

[0273] In various embodiments herein, lifetime measurements are described with respect to measuring the fluorescence lifetime of fluorophores on an IS. However, other marker systems are contemplated in which the signal "lifetime" of different markers can be measured. For example, the imaging chain can be configured to measure luminescence or phosphorescence lifetime.

[0274] Systems that can perform single-photon or multi-photon lifetime imaging (including time domain and frequency domain), including commercially available systems, may be capable of performing these measurements and analyses employed in the methods described herein. Commercial systems include systems from Becker & Hickl, PicoQuant, Olympus, Leica-microsystems, ISS, Nikon, and Zeiss. In-house systems require pulsed laser sources, fast electronics, and detectors (photomultiplier tubes, avalanche photodiodes, metal oxide semiconductor sensors), and microscopy equipment to perform similar measurements discussed in the art.

[0275] It should also be noted that different signal molecules (including, for example, fluorophores) can exhibit different spectral properties, depending on the buffer and salt conditions under which the spectral properties are measured. Therefore, additional discriminating information can be obtained by changing the conditions under which the spectral measurements are performed. Fluorescence lifetime is extremely sensitive to changes in the immediate molecular environment. Factors that affect fluorescence lifetime include pH, ion concentration, viscosity, hydrophobicity, and oxygen concentration. Oxygen can damage the dye, so this can be mitigated by using anti-photobleaching agents (commercially available solutions).

[0276] Alternative spectroscopic techniques such as anisotropy or fluorescence polarization can be used to enhance the discrimination of amino acid residues along the peptide.

[0277] DNA-PAINT for nanoscale topography imaging is a super-resolution technique that allows for image resolutions below the diffraction limit of light. In FLIM, IS chains are stably bound to DS, so all peptide positions have a signal and must be spatially resolvable to be properly distinguished. This creates a density, or spacing, of peptides on the support surface. In embodiments of peptide sequencing using PAINT, IS and DS bind transiently, so each peptide position "blinks" randomly. Because fractions of peptides are associated at any given time, the spacing density can be increased to below the normal resolution limit. Another benefit of PAINT is that it exchanges IS, thereby alleviating photobleaching or dye issues. For example, a disadvantage of using PAINT is that only a fraction of the peptide is detected at any given time, so the accumulation time increases to 5-15 minutes. In such combined embodiments, the association / dissociation or blinking rate of PAINT is slower than the fluorescence lifetime (or any other spectral property being measured).

[0278] Cleavage events can also be measured by re-interrogation, which provides error correction. Errors are tracked during data collection and corrected during analysis. Accuracy is important in sequencing. Tracking NTAA cleavage will inform phasing. If DS / IS hybridization is negative after cleavage, the spot signal is monitored longitudinally. If the spot dims in subsequent cycles, either cleavage failed and the DS was damaged, or new DS conjugation failed. If the spot reappears in subsequent cycles, either cleavage occurred or DS conjugation was successful. It is unknown which error occurred, but it will be known that a gapped alignment model was run at that position. If DS / IS hybridization is intact after cleavage, then several approaches may be useful: one approach is to hybridize the IS after cleavage to determine which DSs remain—indicating a cleavage failure. These peptides are lagging. In practice, if cleavage is efficient, this can be difficult, and the number of positive spots will be low. This can pose problems for image registration. Fiducial markers can help resolve this issue, but may not if image distortion is present. Spots that remain dim after new DS conjugation indicate failed conjugation. Another approach is to use a different DS and a different set of ISs in subsequent cycles. Only the sequence has changed and the dye panel remains the same. Missing peptide signals relative to the previous image may be due to failed cleavage or failed DS conjugation. To avoid doubling the amount of reagents, the DS can have a second binding region specifically for error checking. The IS binding region remains unchanged. DS1 and DS2 are alternated between cycles, with IS1' and IS2' - preferably different colors for detection - to track the cleavage / conjugation results of each cycle. In this method, most spots (positions) will emit light; those that remain dark indicate failed DS binding. Tracking peptide signals between cycles will tell you the type of error and will affect the alignment algorithm that can be used to reconstruct the sequence.

[0279] Those skilled in the art will recognize that buffer conditions, such as pH, co-solvents, salts, metals, etc., may also affect the lifetime changes of TAA-dependent spectral characteristics. Varying the buffer / solvent conditions can be used to modulate the interaction between the TAA side chains and the signaling molecule on the IS. Therefore, additional data can be collected by interrogating the same peptide / DS / IS / signaling molecule combination under different buffer conditions. For example, salts and pH may affect ionic interactions between groups; solvents, surfactants, etc. may affect other interactions (such as weak interactions, van der Wahls forces, hydrophobicity, aromaticity, ionicity, etc.).

[0280] (XII) Cleavage of the terminal amino acid and release of the docking chain

[0281] In certain embodiments involving analysis (e.g., sequencing) of peptides, after binding to the terminal amino acid (N-terminus or C-terminus) by a binding agent and detecting the signal generated by binding of a labeled IS to a DS, the terminal amino acid is removed or cleaved from the peptide to expose a new terminal amino acid. Optionally, the labeled IS is dissociated from the DS and washed away prior to such cleavage.

[0282] In some embodiments, the terminal amino acid is NTAA. In other embodiments, the terminal amino acid is CTAA.

[0283] Cleavage of the terminal amino acid can be accomplished by a number of known techniques, including chemical cleavage and enzymatic cleavage or digestion. In some embodiments, a catalytic engineered enzyme or reagent that promotes the removal of PITC-derived or otherwise labeled N-terminal amino acids is used. In some embodiments, the terminal amino acid is removed or eliminated using any of the methods described in US2020 / 0348307, WO2020 / 223133, or WO2020 / 198264. In some embodiments, cleavage of the terminal amino group uses a carboxypeptidase, an aminopeptidase, a dipeptidyl peptidase, a dipeptidyl aminopeptidase, or a variant, mutant, or modified protein thereof; a hydrolase, or a variant, mutant, or modified protein thereof; a mild Edman degradation reagent; an Edman enzyme; anhydrous TFA, a base; or any combination thereof. In some embodiments, mild Edman degradation uses dichloro or monochloro acid; mild Edman degradation uses TFA, TCA, or DCA; or mild Edman degradation uses triethylamine, triethanolamine, or triethylammonium acetate (Et3NHOAc). In some cases, the reagent used to remove the amino acid includes a base. In some embodiments, the base is a hydroxide, an alkylated amine, a cyclic amine, a carbonate buffer, a trisodium phosphate buffer, or a metal salt. The use of exopeptidases or endopeptidases for digestion is also contemplated.

[0284] In some embodiments, the chemical reagent used to remove a portion of the polypeptide is selected from phenyl isothiocyanate (PITC), nitro-PITC, sulfo-PITC, phenyl isocyanate (PIC), nitro-PIC, sulfo-PIC, Cbz-Cl (benzyl chloroformate) or Cbz-OSu (benzyloxycarbonyl N-succinimide), anhydride, 1-fluoro-2,4-dinitrobenzene (Sanger's reagent, DNFB), dansyl chloride (DNS-Cl or 1-dimethylaminonaphthalene-5-sulfonyl chloride), 4-sulfonyl-2-nitrofluorobenzene (SNFB), 2-pyridinecarboxaldehyde, 2-formylphenylboronic acid, 2-acetylphenylboronic acid. Acid, 1-fluoro-2,4-dinitrobenzene, 4-chloro-7-nitrobenzofuran, pentafluorophenyl isothiocyanate, 4-(trifluoromethoxy)-phenyl isothiocyanate, 4-(trifluoromethyl)-phenyl isothiocyanate, 3-(carboxylic acid)-phenyl isothiocyanate, 3-(trifluoromethyl)-phenyl isothiocyanate, 1-naphthyl isothiocyanate, N-nitroimidazole-1-carboximidamide, N,N'-bis(pivaloyl)-1H-pyrazole-1-carboximidamide, N,N'-bis(benzyloxycarbonyl)-1H-pyrazole-1-carboximidamide, acetylating reagent, guanylating reagent, thioacylation reagent, thioacetylation reagent, thiobenzylation reagent and diheterocyclic methylimine reagent or its derivatives.

[0285] The enzymatic cleavage of NTAA can be completed by peptidases, such as carboxypeptidases, aminopeptidases or dipeptidyl peptidases, dipeptidyl aminopeptidases or their variants, mutants or modified proteins. Aminopeptidases exist naturally in monomeric and multimeric enzyme forms and can be metal or ATP dependent. Aminopeptidases are enzymes that cut amino acids from the N-terminus of proteins or peptides. Natural aminopeptidases have limited specificity and typically cut the N-terminal amino acid in a continuous manner, cutting amino acids one by one (Kishor et al., Analytical Biochemistry 488: 6-8, 2015). However, residue-specific aminopeptidases have been identified (Eriquez et al., J. Clin. Microbiol., 12:667-71, 1980; Wilce et al., Proc. Natl. Acad. Sci. USA 95:3472-3477, 1998; Liao et al., Prot. Sci. 13:1802-10, 2004).

[0286] For the methods described herein, aminopeptidases (e.g., metalloenzyme aminopeptidases) can be engineered to have specific binding or catalytic activity to NTAA only when modified with an N-terminal tag. For example, if the aminopeptidase is modified with a group such as PTC, modified PTC, Cbz, DNP, SNP, acetyl, guanidinyl, diheterocyclic methylamine, the aminopeptidase can be engineered to only cut the N-terminal amino acid. In this way, the aminopeptidase only cuts one amino acid from the N-terminus at a time and allows for controlled degradation cycles. In some embodiments, the modified aminopeptidase is non-selective for amino acid residue identity and selective for N-terminal tags. In other embodiments, the modified aminopeptidase is selective for both amino acid residue identity and N-terminal tags.

[0287] For embodiments involving CTAA analysis, methods for cleaving CTAA from peptides are also known in the art. For example, U.S. Patent No. 6,046,053 discloses a method for reacting a peptide or protein with an alkyl anhydride to convert the carboxyl terminus to an oxazolone, followed by release of the C-terminal amino acid by reaction with an acid and alcohol or an ester. Enzymatic cleavage of CTAA can also be accomplished by carboxypeptidases.

[0288] In various embodiments, the methods described herein include cleaving the N-terminal amino acid or N-terminal amino acid derivative of a polypeptide using Edman chemistry or related chemical degradation. In one embodiment, the methods described herein include enzymatic cleavage of the N-terminal amino acid or N-terminal amino acid derivative using a protease (e.g., an aminopeptidase).

[0289] Edman degradation generally involves two steps, a coupling step and a cleavage step. These steps can be iteratively repeated, each time removing the exposed N-terminal amino acid residue of the polypeptide. In one embodiment, Edman degradation is performed by contacting the polypeptide with a suitable Edman reagent (such as PITC or an analog containing ITC) at elevated pH to form an N-terminal thiocarbamoyl derivative. Lowering the pH, such as by adding trifluoroacetic acid, results in cleavage of the N-terminal amino acid thiocarbamoyl derivative from the polypeptide to form a free anilinothiazolinone (ATZ) derivative. Optionally, this ATZ derivative can be washed away from the sample. In one embodiment, the pH of the sample is adjusted to control the reactions of the coupling and cleavage steps.

[0290] In some embodiments, the N-terminal amino acid is contacted with a suitable Edman reagent (e.g., PITC or an analog containing ITC) at elevated pH prior to contacting the immobilized polypeptide with a plurality of probes that selectively bind to the N-terminal amino acid derivative. Optionally, the cleavage step comprises lowering the pH to cleave the N-terminal amino acid derivative.

[0291] In conventional Edman degradation, peptides are sequenced by degradation from their N-termini using the Edman reagent phenylisothiocyanate (PITC). The process uses two steps: coupling and cleavage. In the first step (coupling), the N-terminal amino group of the peptide is reacted with phenylisothiocyanate to form thiourea. In the second step, treatment of the thiourea with anhydrous acid (e.g., trifluoroacetic acid) results in cleavage of the peptide bond between the first and second amino acids. The N-terminal amino acid is released as a thiazolinone derivative.

[0292] In various embodiments, an Edman degrading enzyme (such as those discussed in U.S. Patent No. 10,852,305, which can be used to cleave the N-terminal amino acid of a peptide or polypeptide) can be used to remove the terminal amino acid from the peptide being analyzed. Such enzymes can catalyze the cleavage step of Edman degradation in aqueous buffer and neutral pH, thereby providing an alternative to the harsh chemical conditions typically employed in conventional Edman degradation. An example of an Edman degrading enzyme can be a modified cruzain enzyme ("cruzipain"), where cruzain is a cysteine ​​protease from the protozoan Trypanosoma cruzi. See, for example, Santos et al. (Sci. Reports 11:18231, 2021).

[0293] (XIII) Cyclic method and its application

[0294] In the methods provided herein, after one cycle of contacting the docking chain attached to the immobilized polypeptide with a labeled IS and performing signal detection, these steps can be repeated one or more times sequentially. Two types of cycles are envisioned: interrogating the same terminal amino acid with a series of two or more different signal molecules on a sequentially bound IS (which sequentially binds to the same DS), and performing signal detection on each amino acid; and removing the terminal amino acid (and its attached DS), subsequently connecting the DS to the newly exposed (next) terminal amino acid, binding to a labeled IS, and performing signal detection. Optionally, these two cycles can be combined together - wherein each amino acid in the immobilized polypeptide is interrogated with a series of two or more different signal molecules (each attached to an IS), and then the terminal amino acid is removed, and the next (newly exposed) terminal amino acid is interrogated with a series of two or more different signal molecules (each attached to an IS), and so on. Each sequential terminal amino acid can be interrogated with the same set of different signal molecules or different sets; and the order of such interrogation can be the same or different.

[0295] In some embodiments, the polypeptide method includes removing a portion of the polypeptide. In some embodiments, the method includes removing the terminal amino acid (and the DS to which it is attached) from the peptide, thereby generating a newly exposed terminal amino acid. This newly exposed terminal amino acid can be connected to a new DS, which is then contacted with an IS labeled with a detection agent, and a signal is detected (the signal is affected by the local environment, just as it is affected by the newly exposed terminal amino acid) - and thus the sequence can be repeated on each newly exposed terminal amino acid. Removing a portion of a polypeptide, such as a terminal amino acid, such as NTAA, can be achieved by any number of known techniques, including chemical techniques and enzymatic techniques (including those described and illustrated herein). In some embodiments, the repetitive steps for analyzing the newly exposed NTAA are substantially similar to the first cycle, including connecting the docking chain to the newly exposed NTAA, binding at least one labeled IS to the newly attached DS, and detecting the signal (such as spectral characteristics, such as lifetime measurement) generated by the signal molecule label in the proximal environment, which is formed when the signal molecule approaches the newly exposed NTAA. In some cases, it may be beneficial to wash the polypeptide with, for example, a suitable buffer (e.g., by washing the solid surface on which the polypeptide is fixed) to remove and / or dissociate components between steps.

[0296] Thus, in various embodiments, the NTAA of the polypeptide is cleaved (and the C-terminus of the polypeptide is immobilized on a support). Chopping off the initial NTAA exposes the N-terminal amino group of the adjacent (penultimate) amino acid on the polypeptide, thereby making the adjacent amino acid available for reaction with DS - and thereby characterizing the identity of the amino acid. Optionally, the polypeptide is sequentially cleaved (each repetition of which can be considered a cycle) until the last amino acid in the polypeptide (the C-terminal amino acid) is reached. However, less than all amino acids in the immobilized polypeptide can optionally be analyzed.

[0297] In other embodiments, the CTAA of the polypeptide is cleaved (and the N-terminus of the polypeptide is immobilized on a support). Cleavage of the initial CTAA exposes the C-terminal carboxyl group of the adjacent (penultimate) amino acid on the polypeptide, thereby making the adjacent amino acid available for reaction with DS - and thereby characterizing the identity of the amino acid. Optionally, the polypeptide is cleaved sequentially (each repetition of which can be considered a cycle) until the last amino acid in the polypeptide (the N-terminal amino acid) is reached. However, less than all amino acids in the immobilized polypeptide can optionally be analyzed.

[0298] (XIV) Assembly of polypeptide sequences - comparison with databases;

[0299] Although the identification of amino acids as described herein employs new and inventive methods, compositions, and techniques, the process of assembling the resulting peptide sequences into proteins and their identification based on comparison with protein sequence databases, etc., is generally conventional. For example, methods for assembling peptide sequences based on other peptide analyses, such as mass spectrometry, are well known; see, for example, Zhang et al. (Curr. Protoc. Mol. Biol. 108: 10.12.1-10.21.30, 2014, doi.org / 10.1002 / 0471142727.mb1021s108).

[0300] A widely used method does not attempt to extract peptide sequence information directly from the spectrum. Instead, the method uses an algorithm to match experimental data with theoretical spectra, which are calculated for all peptides in the database. A score is assigned to each match to determine the confidence of the match. This method is highly automated and is most suitable for high-throughput proteomic analysis of complex samples (Link et al., Nat. Biotechnol. 17(7): 676-682, 1999). A variety of algorithms have been developed for this method (Sadygov et al., Analytical Chemistry 76(6): 1664-1671, 2004; Kapp et al., Proteomics, 5(13): 3475-3490, 2005). The most widely used algorithms are MASCOT (Perkins et al., Electrophoresis. 20(18):3551-3567, 1999), SEQUEST (Link et al., Nature Biotechnology 17(7):676-682, 1999), X! Tandem (Craig and Beavis, Bioinformatics. 20(9):1466-1467, 2004), and OMSSA (Geer et al., J Proteome Res. 3(5):958-965, 2004).

[0301] In one embodiment of the present specification, the method includes comparing the sequence information obtained for each polypeptide molecule with a reference protein sequence database. In some embodiments, small fragments of 10-40 or fewer sequenced continuous or gapped amino acid residues can be used to detect the identity of the polypeptide in the sample. It has been shown that protein sequencing can be accomplished by identifying only a subset of amino acids within the sequence and then comparing partial or incomplete alignments with a database (e.g., sparse sequencing). See, for example, Swaminathan et al., Public Library of Science Computational Biology 11(2):31004080, 2015, doi:10.1371 / journal.pcbi.1004080; and Swaminathan et al., Nature Biology 36:1075-1082, 2018).

[0302] Applications such as Proteome Discoverer (ThermoFisher) can align peptide fragments from the above search algorithms with putative proteins.

[0303] As will be appreciated by one of ordinary skill in the art, aspects of the present disclosure may be embodied as systems, methods, or computer program products. Thus, the various embodiments of the present disclosure may be manifested as complete hardware embodiments, complete software embodiments (including firmware, resident software, microcode, etc.), or embodiments combining both software and hardware aspects. These may all be generally referred to herein as "circuits," "engines," "modules," or "systems." Additionally, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable media, having computer-readable program code instantiated on the one or more computer-readable media.

[0304] Aspects of the present disclosure may be implemented using one or more analog and / or digital electrical or electronic components, and may include microprocessors, microcontrollers, application specific integrated circuits, field programmable gate arrays, programmable logic and / or other analog and / or digital circuit elements configured to perform the various input / output, control, analysis, and other functions described herein, such as by executing instructions of a computer program product.

[0305] A system for generating a database for diagnosing and treating systemic inflammatory conditions may include a corresponding computer device, a computer readable medium, a network, and a remote device.

[0306] The computing device may include, but is not limited to, a processor, memory, input and / or output devices, and a display device. The memory may include, but is not limited to, one or more databases, applications for operating the mass spectrometer and analyzing mass spectrometer data, and client-oriented applications. The computing device may be accessed by a remote device via a network.

[0307] In some cases, software tools can be used to store, analyze and / or determine the information from the provided method. The software can utilize the information of the binding properties of each binding agent. The software can also utilize a list of some or all of the spatial positions of the signal generated or not generated by a detectable label. In certain embodiments, the software can include a database. The database can contain the sequence of the known protein in the species from which the sample is obtained, or also include related species (e.g., homologues). In some cases, if the species of the sample is unknown, a database of some or all protein sequences can be used. The database can also contain the characteristics and / or sequence of any known protein variant and its mutant protein.

[0308] In some embodiments, the software may include one or more algorithms, such as machine learning, deep learning, statistical learning, supervised learning, unsupervised learning, clustering, expectation maximization, maximum likelihood estimation, Bayesian inference, linear regression, logistic regression, binary classification, multinomial classification, or other pattern recognition algorithms. For example, the software may execute the one or more algorithms to analyze information about: (i) the spectral characteristics of each signal molecule used, (ii) information from a protein database, and / or (iii) a list of observed positions (included in different cycles) to generate or assign a possible identity and / or confidence level (e.g., confidence level and / or confidence interval) for each signal detected.

[0309] (XV) Kits and Products

[0310] Also provided herein are kits and articles comprising components for carrying out peptide sequencing analysis using one of the methods described herein. In certain embodiments, the kits further contain other reagents for processing and analyzing proteins, polypeptides, or peptides. Kits and articles may include any one or more reagents and components used in the methods provided. In certain embodiments, the kits include one or more of the following: a carrier surface, a docking strand oligonucleotide (optionally functionalized to be connected to the peptide to be analyzed), an imaging strand oligonucleotide (optionally modified by connecting a signal molecule), a process or reaction compound or solution for a peptide sequence method.

[0311] For example, an exemplary peptide sequencing method kit includes a reagent package comprising a combination of two or more of the following: a DS library and an IS library, a coupling buffer, a hybridization buffer, a wash buffer, and a cleavage buffer. A flow cell kit may include a flow cell / substrate for immobilizing the peptide to be analyzed, a terminal activation reagent (C-terminal or N-terminal, depending on the type of analysis), and immobilization and wash / blocking buffers. The kit may optionally include components that can be used for polypeptide fragmentation, but commercially available peptide fragmentation kits and systems may also be used.

[0312] In certain embodiments, the test kit further includes reagents for preparing proteins or polypeptides. Any combination of protein classification, enrichment and deduction methods can be performed. For example, reagents can be used for fragmentation or digestion of proteins. In some cases, the test kit includes reagents and components for classifying, separating, deducting and / or enriching the protein (or peptide) to be analyzed. In some instances, the test kit further includes a protease. In certain embodiments, the test kit includes a carrier surface on which one or more polypeptides in the polypeptide can be fixed, and one or more reagents for fixing the polypeptide (or peptide) on a carrier.

[0313] In certain embodiments, the test kit also includes one or more buffers or reaction fluids used or necessary for any reaction. Buffers such as washing buffer, reaction buffer, binding buffer, elution buffer are known to those skilled in the art or those of ordinary skill. In certain embodiments, the test kit further includes buffer and one or more other components to accompany other reagents described herein. Reagent, buffer and other components can be provided in bottles (such as sealed vials), vessels, ampoules, bottles, jars, flexible packaging (for example, sealed polyester film or plastic bags) etc. Any component of the test kit can be sterilized and / or sealed.

[0314] In addition to the components described above, the subject kits can further include instructions for practicing the subject methods using the components of the kit, such as instructions for sample preparation, sequence acquisition, and / or analysis of data obtained from the methods. The kits described herein can also include other materials that may be considered desirable from a commercial and user perspective, including other buffers, diluents, filters, syringes, and / or package inserts with instructions for performing at least one of the methods described herein.

[0315] Any of the above kit components, as well as any molecule, molecular complex or conjugate, reagent (e.g., chemical or biological), agent, structure (e.g., support, surface, particle or bead), reaction intermediate, reaction product, binding complex or any other article of manufacture disclosed and / or used in the exemplary kits and methods can be provided alone or in any suitable combination to form a kit.

[0316] Also contemplated are devices for applying chemicals or compositions or washes / solutions to the immobilized support, as well as devices for exposing the immobilized peptides to components of the provided methods (e.g., washes, buffers, reaction solutions, etc.), including flow-through fluidics devices and microfluidics devices.

[0317] Also contemplated are devices for detecting and measuring the spectroscopic properties of the terminal amino acid being analyzed, such as devices for detecting fluorescence lifetime.

[0318] Further embodiments are analysis software and amino acid deconvolution databases prepared using or intended for use with any of the described peptide analysis / sequencing methods.

[0319] (XVI) Representative Definition

[0320] For ease of understanding, a number of terms are defined below. The terms used herein (unless otherwise indicated) have the meanings commonly understood by those of ordinary skill in the art related to the present disclosure. The terms herein are used to describe specific embodiments, but their use is not intended to be limiting, except as outlined in the claims.

[0321] The term "alkyl" refers to a straight or branched chain hydrocarbon. For example, an alkyl group can have from 1 to 6 carbon atoms (i.e., a C1-C6 alkyl group or a C 1-6 alkyl), 1 to 4 carbon atoms (i.e., C1-C4 alkyl or C 1-4 alkyl) or 1 to 3 carbon atoms (i.e., C1-C3 alkyl or C 1-3Examples of suitable alkyl groups include, but are not limited to, methyl (Me, -CH3), ethyl (Et, -CH2CH3), 1-propyl (n-Pr, n-propyl, -CH2CH2CH3), 2-propyl (iso-Pr, isopropyl, -CH(CH3)2), 1-butyl (n-Bu, n-butyl, -CH2CH2CH2CH3), 2-methyl-1-propyl (iso-Bu, isobutyl, -CH2CH(CH3)2), 2-butyl (sec-Bu, sec-butyl, -CH(CH3)2), 2-methyl-1-propyl (iso-Bu, isobutyl, -CH2CH(CH3)2), 2-butyl (sec-Bu, sec-butyl, -CH(CH3)2), 2-methyl-2-propyl (iso-Bu, iso-butyl, -CH2CH(CH3)2), 2-methyl-3 ... )CH2CH3), 2-methyl-2-propyl (tert-Bu, tert-butyl, -C(CH3)3), 1-pentyl (n-pentyl, -CH2CH2CH2CH2CH3), 2-pentyl (-CH(CH3)CH2CH2CH3), 3-pentyl (-CH(CH2CH3)2), 2-methyl-2-butyl (-C(CH3)2CH2CH3), 3-methyl-2-butyl (-CH(CH3)CH(CH3)2), 3-methyl-1-butyl (-CH2CH2CH(CH3)2), 2-methyl-1-butyl (-CH2CH(CH3)CH2CH3), 1-hexyl (-CH2CH2CH2CH2CH2CH3), 2-hexyl (-CH(CH3)CH2CH2CH2CH3), 3-hexyl (-CH(CH2CH3)(CH2CH2CH3)), 2-methyl-2-pentyl (-C(CH3)2CH2CH2CH3), 3-methyl-2-pentyl (-C H(CH3)CH(CH3)CH2CH3), 4-methyl-2-pentyl (-CH(CH3)CH2CH(CH3)2), 3-methyl-3-pentyl (-C(CH3)(CH2CH3)2), 2-methyl-3-pentyl (-CH(CH2CH3)CH(CH3)2), 2,3-dimethyl-2-butyl (-C(CH3)2CH(CH3)2) and 3,3-dimethyl-2-butyl (-CH(CH3)C(CH3)3.

[0322] The term "amino acid" generally refers to an organic compound containing at least one amino group (-NH2) and one carboxyl group (-COOH), wherein the carboxylic acid is deprotonated at neutral pH and has the basic formula NH2CHRCOOH. Amino acids, and therefore peptides, have an N (amino) terminal residue region and a C (carboxyl) terminal residue region. The term "terminus" (referred to as the singular terminus and the plural terminus) refers to the terminal position of a peptide or protein. The "N-terminus" or N-terminal amino acid is the amino acid found at the amino terminus of a peptide, while the "C-terminus" or C-terminal amino acid is the amino acid found at the carboxyl terminus. The phrase "N-terminal amino acid" refers to an amino acid that has a free amine group and is linked to another amino acid only by a peptide amide bond in a polypeptide. Optionally, an "N-terminal amino acid" can be an "N-terminal amino acid derivative." As used herein, an "N-terminal amino acid derivative" refers to an N-terminal amino acid residue that has been chemically modified, for example, by Edman's reagent or other chemical substances in vitro or in cells through natural post-translational modification (e.g., phosphorylation) mechanisms.

[0323] Amino acids include 20 standard naturally occurring or typical amino acids and non-standard amino acids. Standard naturally occurring amino acids include alanine (A or Ala), cysteine ​​(C or Cys), aspartic acid (D or Asp), glutamic acid (E or Glu), phenylalanine (F or Phe), glycine (G or Gly), histidine (H or His), isoleucine (I or Ile), lysine (K or Lys), leucine (L or Leu), methionine (M or Met), asparagine (N or Asn), proline (P or Pro), glutamine (Q or Gln), arginine (R or Arg), serine (S or Ser), threonine (T or Thr), valine (V or Val), tryptophan (W or Trp) and tyrosine (Y or Tyr). Amino acids can be L-amino acids or D-amino acids. Non-standard amino acids can be naturally occurring or chemically synthesized modified amino acids, amino acid analogs, amino acid mimetics, non-standard protein amino acids or non-protein amino acids. Examples of non-standard amino acids include selenocysteine, pyrrolysine and N-formylmethionine, β-amino acids, homo-amino acids, proline and pyruvate derivatives, 3-substituted alanine derivatives, glycine derivatives, ring-substituted phenylalanine and tyrosine derivatives, linear core amino acids, N-methyl amino acids.

[0324] The terms "amino acid sequence", "peptide", "peptide sequence", "polypeptide" and "polypeptide sequence" are used interchangeably herein to refer to at least two amino acids or amino acid analogs covalently linked by a peptide (amide) bond or a peptide bond analog. The term peptide includes oligomers and polymers of amino acids or amino acid analogs. The term peptide also includes molecules commonly referred to as peptides, which typically contain two (2) to twenty (20) amino acids. The term peptide also includes molecules commonly referred to as polypeptides, which typically contain twenty (20) to fifty amino acids (50). The term peptide also includes molecules commonly referred to as proteins, which typically contain fifty (50) to three thousand (3000) amino acids. The amino acids of a peptide can be L-amino acids or D-amino acids. A peptide, polypeptide or protein can be synthetic, recombinant or naturally occurring. A synthetic peptide is a peptide produced artificially in vitro.

[0325] As used herein, " analysis " polypeptide means that all or part of the components of polypeptide are identified, detected, quantitatively, characterized, distinguished or its combination. For example, analyzing peptides, polypeptides or proteins includes determining all or part of the amino acid sequence (continuous or discontinuous) of peptides. Analyzing polypeptides also includes the partial identification of the components of polypeptides. For example, the amino acid in the partial identification polypeptide protein sequence can be used to identify the amino acid in the protein as belonging to a possible amino acid subset. Analysis usually starts from analyzing n NTAA, and then proceeds to the next amino acid (that is, n-1, n-2, n-3 etc.) of the peptide. This is achieved by eliminating nNTAA, thereby converting the n-1 amino acid of the peptide into N-terminal amino acid (referred to as " n-1NTAA " in this article).

[0326] Analysis of the peptide may also include determining the presence, identity, and / or frequency of post-translational modifications on the peptide, and may optionally include information regarding the order of post-translational modifications on the peptide, protein, or polypeptide.

[0327] Analyzing peptides can also include determining the presence and frequency of recognized characteristics of proteins, such as recognized domains and / or functional domains affected by the primary (or secondary) sequence of the peptide, which may or may not include information about the precedence or position of domains within the polypeptide or peptide. Domains can include, for example, epitopes in a peptide, which may or may not include information about the precedence or position of epitopes within the peptide. Analyzing peptides can include combining different types of analyses, such as obtaining amino acid sequence information and post-translational modification information, or identification of primary amino acid sequence and domains.

[0328] As used herein, the term "barcode" refers to a molecule that provides a unique identifier tag or source information for a polypeptide, a binding agent, a set of binding agents from a binding cycle, a sample polypeptide, a set of samples, a polypeptide within a compartment (e.g., a droplet, a bead, or a separate location), a polypeptide within a set of compartments, a polypeptide fraction, a set of polypeptide fractions, a spatial region or a set of spatial regions, a polypeptide library, or a binding agent library. A "nucleic acid barcode" refers to a nucleic acid molecule having 2 to 30 bases (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 bases). "Peptide barcode" or "amino acid barcode" refers to an amino acid sequence that can be, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, 50, 75, or 100 amino acids in length. A particular peptide barcode can be distinguished from other peptide barcodes by having a different length, sequence, or other physical property (e.g., hydrophobicity). A barcode can be an artificial sequence or a naturally occurring sequence. In certain embodiments, each barcode within a population of barcodes is different. In other embodiments, a portion of the barcodes in a barcode population are different, for example, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% of the barcodes in the barcode population are different. The number of barcodes can be randomly generated or non-randomly generated. In certain embodiments, the barcode population is error-correcting or error-tolerant barcodes. The barcodes can be used to computationally deconvolute multiplexed sequencing data and identify sequence reads from individual polypeptides, samples, libraries, etc.

[0329] The term "nucleic acid molecule" or "polynucleotide" refers to a single-stranded or double-stranded polynucleotide and polynucleotide analogs containing deoxyribonucleotides or ribonucleotides linked by 3'-5' phosphodiester bonds. Nucleic acid molecules include DNA, RNA and cDNA. Polynucleotide analogs can have a backbone other than the standard phosphodiester linkages found in natural polynucleotides, and optionally have one or more modified sugar moieties other than ribose or deoxyribose. Polynucleotide analogs contain bases that can hydrogen bond with standard polynucleotide bases by Watson-Crick base pairing, wherein the analog backbone presents the bases in a manner that allows such hydrogen bonding to occur between the oligonucleotide analog molecule and the standard polynucleotide bases in a sequence-specific manner. Examples of polynucleotide analogs include xenogeneic nucleic acids (XNA), bridged nucleic acids (BNA), glycol nucleic acids (GNA), peptide nucleic acids (PNA), γPNA, morpholino polynucleotides, locked nucleic acids (LNA), threose nucleic acids (TNA), 2'-O-methyl polynucleotides, 2'-O-alkylribosyl substituted polynucleotides, phosphorothioate polynucleotides, and borophosphate polynucleotides. Polynucleotide analogs can have purine or pyrimidine analogs, including, for example, 7-deazapurine analogs, 8-halogenated purine analogs, 5-halogenated pyrimidine analogs, or universal base analogs that can pair with any base, including hypoxanthine, nitroazoles, isoquinolone (isocarbostyril) analogs, azole carboxamides, and aromatic triazole analogs, or base analogs with additional functionality, such as biotin moieties for affinity binding. In some embodiments, the nucleic acid molecule or oligonucleotide is a modified oligonucleotide. In some embodiments, the nucleic acid molecule or oligonucleotide is a DNA with pseudo-complementary bases, a DNA with protected bases, an RNA molecule, a BNA molecule, an XNA molecule, an LNA molecule, a PNA molecule, a γPNA molecule, or a morpholino DNA, or a combination thereof. In some embodiments, the nucleic acid molecule or oligonucleotide is backbone-modified, sugar-modified, or nucleobase-modified. In some embodiments, the nucleic acid molecule or oligonucleotide has a nucleobase protecting group (e.g., Alloc), an electrophilic protecting group (e.g., sulfane), an acetyl protecting group, a nitrobenzyl protecting group, a sulfonate protecting group, or a traditional base instability protecting group.

[0330] The phrase "detectable label" refers to a substance that can indicate the presence of another substance when associated with the other substance. A detectable label can be a substance that is connected to or incorporated into the substance to be detected. In certain embodiments, the detectable label is suitable for allowing detection and also quantitative, for example, a detectable label that emits a detectable and measurable signal. Detectable labels include any label that can be used and is compatible with the provided polypeptide analysis assay format, and include bioluminescent labels, biotin / avidin labels, chemiluminescent labels, chromophores, coenzymes, dyes, electroactive groups, electrochemiluminescent labels, enzymatic labels, fluorescent labels, latex particles, magnetic particles, metals, metal chelates, phosphorescent dyes, protein labels, radioactive elements or moieties and stable free radicals. In the embodiments provided, fluorescent labels are preferred.

[0331] Direct and indirect connections (e.g., detectable labels directly and indirectly connected to oligonucleotides or other substances) can include covalent bonds or non-covalent interactions. Covalent bonds include the sharing of electrons in chemical bonds. Non-covalent interactions include dispersed electromagnetic interactions, such as hydrogen bonds (e.g., occurring between paired nucleic acid chains), ionic bonds, van der Waals interactions, and hydrophobic bonds.

[0332] "Fluorescence" refers to the visible light emitted by substances that absorb light of different wavelengths. In certain embodiments, fluorescence provides a non-destructive way to track and / or analyze biomolecules based on the fluorescence emission of a specific wavelength. Proteins (including antibodies), peptides, nucleic acids, oligonucleotides (including single-stranded and double-stranded primers), etc. can be "labeled" with any of the various exogenous fluorescent molecules called fluorophores. Isothiocyanate derivatives of fluorescein, such as carboxyfluorescein, are examples of fluorophores that can be conjugated to proteins (such as antibodies for immunohistochemistry) or nucleic acids. In certain embodiments, fluorescein can be conjugated to nucleoside triphosphates and incorporated into nucleic acid probes (such as "fluorescence-conjugated primers") to perform in situ hybridization.

[0333] The term "individual" or "subject" includes birds (e.g., chickens, ducks, geese, turkeys, quail, songbirds, etc.), other non-mammalian vertebrates (e.g., fish), and mammals (e.g., mice, rats, rabbits, and other rodents; cats and other felines; dogs and other canines; other domestic animals; pigs, cows, cattle, sheep, goats, horses, and other livestock; monkeys and other non-human primates). In various embodiments, the individual or subject is a human.

[0334] The term "linker" refers to one or more nucleotides, nucleotide analogs, amino acids, peptides, polypeptides, polymers, or non-nucleotide chemical moieties used to connect two molecules to each other. Linkers can be used to connect nucleic acids (e.g., DS) to polypeptides, polypeptides to vectors, detection agents to nucleic acids (e.g., IS), etc. In certain embodiments, linkers connect two molecules via an enzymatic reaction or a chemical reaction (e.g., click chemistry).

[0335] As used herein, "next generation sequencing" refers to a high-throughput sequencing method that allows millions to billions of molecules to be sequenced in parallel. Examples of next generation sequencing methods include sequencing by synthesis, sequencing by ligation, sequencing by hybridization, polymerase colony sequencing (polony sequencing), ion semiconductor sequencing, and pyrophosphate sequencing. By connecting primers to a solid substrate and connecting complementary sequences to nucleic acid molecules, nucleic acid molecules can hybridize with the solid substrate by primers, and then utilize polymerase to produce multiple copies in discrete regions on the solid substrate for amplification (these groupings are sometimes referred to as polymerase colonies or clones). Therefore, during the sequencing process, the nucleotides at a specific position can be sequenced multiple times (e.g., hundreds or thousands of times) - this coverage depth is referred to as "deep sequencing." Examples of high-throughput nucleic acid sequencing technologies include platforms offered by Illumina, BGI, Qiagen, Thermo-Fisher, and Roche, including formats such as parallel bead arrays, sequencing by synthesis, sequencing by ligation, capillary electrophoresis, electronic microchips, "biochips," microarrays, parallel microchips, and single-molecule arrays (see, e.g., Service, Science 311: 1544-1546, 2006).

[0336] As used herein, "single molecule sequencing" or "third generation sequencing" refers to a next generation sequencing method, in which the reads from a single molecule sequencing instrument are generated by sequencing a single molecule, typically a molecule of DNA. Unlike the next generation sequencing method that relies on amplification and clones many DNA molecules in parallel in a staged manner for sequencing, single molecule sequencing interrogates a single molecule (e.g., minute of DNA) and does not require amplification or synchronization. Single molecule sequencing includes methods that require pausing the sequencing reaction after each base is incorporated ("wash and scan" cycles) and methods that do not require stopping between reading steps. Examples of single molecule sequencing methods include single molecule real-time sequencing (Pacific Biosciences), nanopore-based sequencing (Oxford Nanopore), duplex interruption nanopore sequencing, and direct imaging of DNA using advanced microscopes.

[0337] The term "sample" refers to anything that may contain an analyte (e.g., a protein or peptide) for which an analyte assay (e.g., detection, quantification, and / or sequencing) is desired. The term "sample" may include a solution, a suspension, a liquid, a powder, a paste, any of aqueous or non-aqueous, or any combination thereof. The sample may be a biological sample, such as a biological fluid or a biological tissue or an individual cell. Examples of biological fluids include urine, blood, plasma, serum, saliva, semen, feces, sputum, cerebrospinal fluid, tears, mucus, amniotic fluid, and the like. Biological tissue is an aggregate of cells, typically cells of a particular species (or a mixture of two or more species), together with their intercellular material, that form one of the structural materials of a human, animal, plant, bacterial, fungal, or viral structure, including connective tissue, epithelial tissue, muscle tissue, and nervous tissue. Examples of biological tissue also include organs, tumors, lymph nodes, and arteries. In some embodiments, the sample can be derived from a tissue or body fluid, such as connective tissue, epithelial tissue, muscle tissue, or neural tissue; a tissue selected from the group consisting of brain, lung, liver, spleen, bone marrow, thymus, heart, lymph, blood, bone, cartilage, pancreas, kidney, gall bladder, stomach, intestine, testicle, ovary, uterus, rectum, nervous system, glands, and internal blood vessels; or a body fluid selected from the group consisting of blood, urine, saliva, bone marrow, semen, ascites, and subfractions thereof, such as serum or plasma.

[0338] As used herein, the term "post-translational modification" refers to modifications that occur on a peptide after its translation (eg, by ribosomal translation) is complete. Post-translational modifications can be covalent chemical modifications or enzymatic modifications. Examples of post-translational modifications include acylation, acetylation, alkylation (including methylation), biotinylation, butyrylation, carbamylation, carbonylation, C-terminal amidation, deamidation, deimination, diphtheria amide formation, disulfide bridge formation, elimination, farnesylation, flavin attachment, formylation, γ-carboxylation, glutamylation, glycylation, glycosylation, glycosylphosphatidylinositolation, heme C attachment, hydroxylation, hydroxyputrescine lysine formation, iodination, prenylation, lipidation, fatty acylation, malonylation, methylation, myristoylation, oxidation, palmitoylation, PEGylation, phosphopantetheinylation, phosphorylation, prenylation, propionylation, retinylidene Schiff base formation, S-glutathionylation, S-nitrosylation, S-sulfinylation, selenoylation, succinylation, sulfinylation, and ubiquitination. Post-translational modification includes modification of the amino terminal and / or carboxyl terminal of the peptide. Modification of the terminal amino group includes deamination, N-low alkyl, N-di-low alkyl and N-acyl modification. Modification of the terminal carboxyl group includes amide, low alkyl amide, dialkyl amide and low alkyl ester modification (for example, wherein low alkyl is C1-C4 alkyl). Post-translational modification also includes modification of the amino acid falling between the amino and carboxyl terminal, such as those mentioned above. The term post-translational modification can also include peptide modifications comprising one or more detectable labels.

[0339] The term "proteome" includes a whole set of proteins, polypeptides or peptides (including conjugates or complexes thereof) expressed at a specific time by the genome, cells, tissues or organisms of any organism. On the one hand, it is a collection of proteins expressed in a given type of cell or organism under given conditions at a given time. For example, a "cellular proteome" may include a collection of proteins found in a specific cell type under a specific set of environmental conditions (e.g., exposed to hormone stimulation). The complete proteome of an organism may include a complete collection of proteins from all various cellular proteomes. A proteome may also include a collection of proteins in certain subcellular biological systems. For example, all proteins in a virus may be referred to as a viral proteome. As used herein, the term "proteome" includes subsets of the proteome, including the kinome; the secretome; the receptorome (e.g., GPCRome); the immune proteome; the nutritional proteome; subsets of the proteome defined by post-translational modification (e.g., phosphorylation, ubiquitination, methylation, acetylation, glycosylation, oxidation, lipidation and / or nitrosylation), such as the phosphotyrosine proteome (e.g., the phosphotyrosine proteome, the tyrosine proteome and the tyrosine phosphorylated proteome), the glycoproteome, etc.; subsets of the proteome associated with a tissue or organ, a developmental stage or a physiological or pathological condition; subsets of the proteome associated with a cellular process (e.g., cell cycle, differentiation (or dedifferentiation), cell death, senescence, cell migration, transformation or metastasis); or any combination thereof.

[0340] Proteomics is the study of the proteome. Therefore, the term "proteomics" encompasses the quantitative analysis of the proteome within cells, tissues, and body fluids, as well as the corresponding spatial distribution of the proteome within cells and tissues. Furthermore, proteomic studies encompass the dynamic state of the proteome, which continuously changes over time in response to biological and defined biological or chemical stimuli.

[0341] In the embodiments provided herein, the term "sample" includes any material containing one or more polypeptides. A sample can be a biological sample, such as an animal or plant tissue, biopsy, organ, cell, membrane vesicle, plasma membrane, organelle, cell extract, secretion, urine or mucus or other secretion, tissue extract or other biological sample of both natural or synthetic origin. The term sample also includes a single cell, organelle or intracellular material isolated from a biological sample, or a virus, prion, bacteria, fungus or an isolate thereof. A sample can also be an environmental sample, such as a water sample or soil sample, or a sample of any artificial or natural material containing one or more polypeptides.

[0342] The term "side chain" or "R" (or R group) refers to the unique structure attached to the alpha carbon (linking the amine and carboxylic acid groups of an amino acid) that makes each amino acid unique. R groups have a variety of shapes, sizes, charges, and reactivities, including positively or negatively charged polar side chains such as lysine (+), arginine (+), histidine (+), aspartic acid (-), and glutamic acid (-). Amino acids can also be basic, such as lysine, or acidic, such as glutamic acid. Uncharged polar side chains have hydroxyl, amide, or thiol groups, such as cysteine, which has a chemically reactive side chain, i.e., a thiol group that can form a bond with another cysteine ​​with a hydroxyl R side chain of varying size; serine (Ser) and threonine (Thr); asparagine (Asn), glutamine (Gln), and tyrosine (Tyr). Nonpolar, hydrophobic amino acid side chains include the amino acids glycine; alanine, valine, leucine, and isoleucine, whose aliphatic hydrocarbon side chains range in size from the methyl group of alanine to the isomeric butyl groups of leucine and isoleucine. Methionine (Met) has a thiol ether side chain, and proline (Pro) has a cyclic pyrrolidine side group. Phenylalanine (with its phenyl moiety) (Phe) and tryptophan (Trp) (with its indole group) contain aromatic side groups and are characterized by being blocky and nonpolar.

[0343] As used herein, the terms "solid support," "solid surface," or "solid substrate" or "sequencing substrate" refer to any solid material, including porous and non-porous materials, that can associate directly or indirectly with a polypeptide by any means known in the art, including covalent and non-covalent interactions, or any combination thereof. A solid support can be two-dimensional (e.g., a planar surface) or three-dimensional (e.g., a gel matrix or beads). A solid support can be any carrier surface, including beads, microbeads, arrays, glass surfaces, silicon surfaces, plastic surfaces, filters, membranes, PTFE membranes, PTFE membranes, nitrocellulose membranes, nitrocellulose-based polymer surfaces, nylon, silicon wafer chips, flow-through chips, flow cells, biochips including signal transduction electronic devices, channels, microtiter wells, ELISA plates, interferometry turntables, nitrocellulose membranes, nitrocellulose-based polymer surfaces, polymer matrices, nanoparticles, or microspheres. The material of solid phase carrier comprises acrylamide, agarose, cellulose, dextran, nitrocellulose, glass, gold, quartz, polystyrene, polyethylene vinyl acetate, polypropylene, polyester, polymethacrylate, polyacrylate, polyethylene, polyethylene oxide, polysilicate, polycarbonate, polyvinyl alcohol (PVA), Teflon, fluorocarbon, nylon, silicone rubber, polyanhydride, polyglycolic acid, polyvinyl chloride, polylactic acid, polyorthoester, functionalized silane, polypropyl fumarate, collagen, glycosaminoglycan, polyamino acid, dextran or its any combination.Solid phase carrier further comprises film, membrane, bottle, dish, fiber, woven fiber, shaped polymer, as tube, granule, bead, microsphere, microparticle or its any combination.

[0344] For example, when the solid surface is a bead, the bead can include ceramic beads, polystyrene beads, polymer beads, polyacrylate beads, methylstyrene beads, agarose beads, cellulose beads, dextran beads, acrylamide beads, solid beads, porous beads, paramagnetic beads, glass beads, controlled pore beads, silica-based beads or any combination thereof. The bead can be spherical or irregularly shaped. The bead or carrier can be porous. The size of the bead can be in the range of nanometers (e.g., 100nm) to millimeters (e.g., 1mm). In certain embodiments, the size of the bead is in the range of 0.2 microns to 200 microns or 0.5 microns to 5 microns. In certain embodiments, the diameter of the bead can be about 1, 1.5, 2, 2.5, 2.8, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 15 or 20 μm. In certain embodiments, " bead " solid phase carrier can refer to independent bead or a plurality of beads.In certain embodiments, solid surface is nanoparticle.In certain embodiments, the size range of nanoparticle is from diameter about 1nm to about 500nm, for example, diameter is between 1nm and 20nm, between 1nm and 50nm, between 1nm and 100nm, between 10nm and 50nm, between 10nm and 100nm, between 10nm and 200nm, between 50nm and 100nm, between 50nm and 150, between 50nm and 200nm, between 100nm and 200nm or between 200nm and 500.In certain embodiments, the diameter of nanoparticle can be 10nm, 50nm, 100nm, 150nm, 200nm, 300nm or 500nm.In certain embodiments, the diameter of nanoparticle is less than about 200nm.

[0345] The most widely used reaction for the sequential analysis of the N-terminal residues of peptides is the Edman degradation method (Edman et al., Acta Chem. Scand. 4:283-293, 1950). Edman degradation is a method for sequencing the amino acids in a peptide (or protein) in which the amino-terminal residue is labeled and cleaved from the peptide without disrupting the peptide bonds between other amino acid residues. In the Edman procedure, phenyl isothiocyanate (PITC) reacts quantitatively with the free amino groups of a peptide to produce the corresponding phenylthiocarbamoyl peptide. Upon treatment with anhydrous acid, the N-terminal residue is cleaved into the phenylthiocarbamoyl amino acid; this leaves the remainder of the peptide chain intact. One aspect of the Edman degradation method is that the remainder of the peptide chain (after removal of the N-terminal amino acid) remains intact during further cycles of this procedure; thus, the Edman method can be used in a sequential, iterative manner to identify multiple consecutive amino acid residues starting from the N-terminus of the peptide to be analyzed.

[0346] As used herein, the phrase "universal" docking strand refers to a single-stranded DNA oligonucleotide that can bind to any peptide, regardless of the primary sequence of the peptide.

[0347] The following exemplary embodiments and examples are included to illustrate various embodiments of the present disclosure. Those skilled in the art should, in light of the present disclosure, appreciate that many changes can be made to the specific embodiments disclosed herein and still obtain a like or similar result without departing from the spirit and scope of the present disclosure.

[0348] (XVII) Exemplary Embodiments

[0349] First embodiment group

[0350] 1. A method for sequencing a peptide having a primary N-terminal amino acid (NTAA), the method comprising sequentially interrogating the primary NTAA using a library of at least two different combinations of a ssDNA docking strand (DS) and a ssDNA imaging strand (IS), wherein a fluorophore is conjugated to the IS to generate a set of fluorescence lifetime data having a characteristic fingerprint for each different amino acid of the peptide, wherein each pair of IS and DS is at least partially complementary in sequence.

[0351] 2. The method of embodiment 1, wherein the sequential interrogation comprises detecting and / or measuring the interaction between the fluorophore and the amino acid side chain at or near the NTAA by detecting fluorescence lifetime data for each pair of IS and DS in the library, for example using fluorescence lifetime imaging (FLIM) single-photon fluorescence measurement.

[0352] 3. The method of embodiment 1 or embodiment 2, further comprising removing the initial NTAA peptide by Edman degradation reaction, Edman degradation enzyme reaction, or a similar process.

[0353] 4. The method of any one of embodiments 1 to 3, wherein the method is repeated for each subsequent amino acid in the peptide to generate a matrix of fluorescence lifetime data.

[0354] 5. The method of embodiment 4, wherein the data is input into a machine learning algorithm to reconstruct the polypeptide sequence.

[0355] 6. The method of any one of embodiments 1 to 6, wherein the library of IS comprises a plurality of ssDNA oligonucleotides that are varied such that the spatial positioning and / or degree of freedom of the attached fluorophore can be modulated to tune the interaction with the NTAA side chain and thereby modulate the measured fluorescence lifetime.

[0356] 7. The method of embodiment 6, wherein the library of IS comprises a plurality of ssDNA oligonucleotides varied by one or more of: including modified nucleotides, including non-natural nucleotides, including 5' IS overhangs relative to a cognate DS, or including 5' IS non-overhangs relative to a cognate DS.

[0357] 8. The method of any one of embodiments 1 to 7, wherein the fluorophore comprises Alexa Fluor 88(AF488), BODIPY-FL, BODIPY-TR or TAMRA.

[0358] 9. The method of any one of embodiments 1 to 8, wherein the fluorophore is conjugated at the end of the IS.

[0359] 10. The method of any one of embodiments 1 to 8, wherein the peptide is conjugated to a modified nucleotide within the DS, and the fluorophore is conjugated to a modified nucleotide within the IS. Figure 11 ).

[0360] 11. The method of any one of embodiments 1 to 10, wherein the removal of the NNTAA is performed under conditions such that the remaining peptide has a new N-terminal amino acid.

[0361] 12. The method according to any one of embodiments 1 to 11, wherein the peptide is immobilized on a solid support.

[0362] 13. A database comprising the matrix of the fluorescence lifetime data of embodiment 4.

[0363] 14. A method for identifying the N-terminal amino acid (NTAA) of a peptide, the method comprising: binding the C-terminal amino acid of the peptide to a solid surface; connecting a ssDNA docking strand (DS) to the NTAA of the peptide; hybridizing a first ssDNA imaging strand (IS) to the DS, the first IS comprising a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, the second IS comprising a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; and identifying the initial NTAA of the peptide based on the detected fluorescence lifetime data of the first fluorophore and the second fluorophore.

[0364] 15. The method of embodiment 14, further comprising: cleaving the initial NTAA from the peptide to leave the next NTAA of the peptide.

[0365] 16. The method of embodiment 15, wherein cleaving the initial NTAA comprises an Edman degradation reaction, an Edman degradation enzyme reaction, or a similar process.

[0366] 17. The method of any one of embodiments 14 to 16, comprising repeating the method multiple times to identify the sequence of the peptide.

[0367] 18. A method for sequencing peptides, the method comprising: attaching a peptide to be sequenced via its C-terminus to a solid phase substrate; functionalizing the initial N-terminal amino acid of the immobilized peptide with a universal docking strand (DS) ssDNA oligonucleotide; contacting the DS with an imaging strand (IS) oligonucleotide complementary to the DS oligonucleotide, the IS conjugated to a first fluorophore; obtaining single-molecule fluorescence lifetime (FLIM) measurements of the first fluorophore for each peptide; optionally, repeating the single-molecule FLIM measurements for one or more additional combinations of IS and fluorophores in a library; and cleaving the initial N-terminal amino acid from the peptide to reveal a second N-terminal amino acid; and optionally, performing another analysis cycle on the second N-terminal amino acid.

[0368] 19. A method of sequencing a peptide substantially as described herein.

[0369] 20. A kit for performing the method according to any one of embodiments 1 to 18, comprising at least one pair of IS and DS.

[0370] 21. The kit of embodiment 20, comprising at least two pairs of IS and DS, wherein the two pairs differ in the fluorophore, sequence, or both contained in the IS.

[0371] Second embodiment group

[0372] 1. A method for identifying a terminal amino acid (TAA) of a peptide having an N-terminal amino acid (NTAA) and a C-terminal amino acid (CTAA), the method comprising: binding the NTAA of the peptide or the CTAA of the peptide to a solid surface to generate a bound TAA; attaching a ssDNA docking strand (DS) to the unbound TAA of the peptide; hybridizing a first ssDNA imaging strand (IS) to the DS, the first IS comprising a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, the second IS comprising a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; and identifying the original TAA of the peptide based on the detected fluorescence lifetime data of the first fluorophore and the second fluorophore.

[0373] 2. A method for sequencing a peptide having an initial terminal amino acid (TAA), the method comprising: interrogating the initial TAA using a single-stranded DNA (ssDNA) docking strand (DS) connected to the initial TAA and an ssDNA imaging strand (IS) conjugated to a signal molecule to generate measurement results of spectral characteristic data having a characteristic fingerprint of the initial TAA of the peptide; wherein the IS and the DS are at least partially complementary in sequence.

[0374] 3. The method according to Example 2 further comprises: sequentially interrogating the initial TAA using a library of at least two different combinations of single-stranded DNA (ssDNA) docking strands (DS) and ssDNA imaging strands (IS), wherein a signal molecule is conjugated to the IS to generate a set of spectral characteristic data having a characteristic fingerprint of the initial TAA of the peptide.

[0375] 4. The method of embodiment 2, wherein the initial TAA is: the amino-terminal (N-terminal) amino acid (NTAA) of the peptide; or the carboxyl (C-terminal) amino acid (CTAA) of the peptide.

[0376] 5. The method of embodiment 2, wherein the method is performed on multiple peptides in parallel.

[0377] 6. The method of embodiment 2, wherein the signaling molecule comprises a fluorophore and the spectral characteristic comprises a measure of fluorescence.

[0378] 7. The method of embodiment 6, wherein the spectral characteristic comprises fluorescence lifetime.

[0379] 8. A method for sequencing a peptide having an initial N-terminal amino acid (NTAA) and a C-terminal amino acid (CTAA), the method comprising: sequentially interrogating the initial NTAA using a library of at least two different combinations of a ssDNA docking strand (DS) and a ssDNA imaging strand (IS), wherein a fluorophore is conjugated to the IS to generate a set of fluorescence lifetime data having characteristic measurements for each combination of DS, IS, and fluorophore, wherein each pair of IS and DS is at least partially complementary in sequence.

[0380] 9. The method of any one of embodiments 1 to 8, wherein the DS is a universal DS.

[0381] 10. The method of any one of embodiments 1 to 8, wherein the interrogation or the sequential interrogation comprises detecting and / or measuring the interaction between a fluorophore and an amino acid side chain at or near the CTAA or the NTAA by detecting fluorescence lifetime data for each of a plurality of IS / DS pairs in the library.

[0382] 11. The method of embodiment 10, wherein detecting or measuring the interaction comprises obtaining fluorescence lifetime imaging (FLIM) single molecule fluorescence measurements of each of the plurality of IS / DS pairs in the library.

[0383] 12. The method of embodiment 1 or embodiment 4, further comprising removing the original CTAA or NTAA of the peptide by Edman degradation, enzymatic digestion, or the like.

[0384] 13. The method of any one of embodiments 1 to 8, wherein the method is repeated for at least two subsequent amino acids in the peptide to generate a matrix of fluorescence lifetime data.

[0385] 14. The method of embodiment 13, wherein the method is performed for each subsequent amino acid in the peptide to generate a matrix of fluorescence lifetime data.

[0386] 15. The method of embodiment 13, wherein the fluorescence lifetime data is input into a machine learning algorithm to reconstruct the polypeptide sequence.

[0387] 16. The method of embodiment 14, wherein the fluorescence lifetime data is input into a machine learning algorithm to reconstruct the polypeptide sequence.

[0388] 17. The method of embodiment 3 or embodiment 8, wherein the library of IS comprises a plurality of ssDNA oligonucleotides, wherein the ssDNA oligonucleotides are varied such that the spatial positioning and / or degree of freedom of the attached fluorophore are varied to modulate the interaction with the CTAA side chain or the NTAA side chain and thereby modulate the measured fluorescence lifetime.

[0389] 18. The method of embodiment 17, wherein the library of IS comprises a plurality of ssDNA oligonucleotides varied by one or more of: including modified nucleotides, including non-natural nucleotides, including 5' IS overhangs relative to a cognate DS, or including 5' IS non-overhangs relative to a cognate DS.

[0390] 19. The method of embodiment 10, wherein the interaction between the CTAA or the NTAA is further influenced by one or more of DS positioning, degrees of freedom, or another variable described herein.

[0391] 20. The method of any one of embodiments 1, 6, or 8, wherein the fluorophore comprises Alexa Fluor. 488 (AF488), BODIPY-FL, BODIPY-TR, TAMRA, or KU dyes.

[0392] 21. The method of embodiment 20, wherein the fluorophore is conjugated at the end of the IS.

[0393] 22. The method of any one of embodiments 1 to 8, wherein: the peptide is conjugated to a modified nucleotide within the DS and the fluorophore is conjugated to a modified nucleotide within the IS; or the peptide is conjugated to a modified nucleotide at or near the end of the IS and the fluorophore is conjugated to a modified nucleotide within the IS; or the peptide is conjugated to a modified nucleotide within the DS and the fluorophore is conjugated to a modified nucleotide at or near the end of the IS; or the peptide is conjugated to a modified nucleotide at or near the end of the DS and the fluorophore is conjugated to a modified nucleotide at or near the end of the IS.

[0394] 23. The method of embodiment 12, wherein the removal of CTAA or NTAA is performed under conditions such that the remaining peptides have a new terminal amino acid that can be used in another analysis cycle.

[0395] 24. The method of any one of embodiments 1 to 8, wherein the or each peptide is immobilized on a solid support.

[0396] 25. A database comprising a matrix of the fluorescence lifetime data of embodiment 14.

[0397] 26. A method for identifying the NTAA of a peptide having an N-terminal amino acid (NTAA) and a C-terminal amino acid (CTAA), the method comprising: binding the CTAA of the peptide to a solid surface; connecting an ssDNA docking strand (DS) to the NTAA of the peptide; hybridizing a first ssDNA imaging strand (IS) to the DS, the first IS comprising a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, the second IS comprising a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; and identifying the initial NTAA of the peptide based on the detected fluorescence lifetime data of the first fluorophore and the second fluorophore.

[0398] 27. A method for identifying the CTAA of a peptide having a C-terminal amino acid (CTAA) and an N-terminal amino acid (NTAA), the method comprising: binding the NTAA of the peptide to a solid surface; connecting an ssDNA docking strand (DS) to the CTAA of the peptide; hybridizing a first ssDNA imaging strand (IS) to the DS, the first IS comprising a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, the second IS comprising a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; and identifying the initial NTAA of the peptide based on the detected fluorescence lifetime data of the first fluorophore and the second fluorophore.

[0399] 28. The method of embodiment 26 or embodiment 27, further comprising: cleaving the initial TAA from the peptide to leave the next TAA of the peptide.

[0400] 29. The method of embodiment 28, wherein cleaving the initial NTAA comprises an Edman degradation reaction, an Edman degradation enzyme reaction, or a similar process.

[0401] 30. The method of embodiment 27 or embodiment 27, comprising repeating the method multiple times to identify the sequence of the peptide.

[0402] 31. A method for sequencing peptides, each of which has a C-terminal amino acid (CTAA) and an N-terminal amino acid (NTAA), the method comprising: attaching the peptide to be sequenced via its C-terminus to a solid substrate to form an immobilized peptide; functionalizing the initial N-terminal amino acid of the immobilized peptide with a universal docking strand (DS) ssDNA oligonucleotide; contacting the DS with an imaging strand (IS) oligonucleotide complementary to the DS oligonucleotide, the IS conjugated to a first fluorophore; obtaining single-molecule fluorescence lifetime (FLIM) measurements of the first fluorophore for each peptide; optionally, repeating the single-molecule FLIM measurements for one or more additional combinations of IS and fluorophores in a library; cleaving the initial N-terminal amino acid from the peptide to reveal a second N-terminal amino acid; and optionally, performing another analysis cycle on the second N-terminal amino acid.

[0403] 32. A method for sequencing peptides, each of which has a C-terminal amino acid (CTAA) and an N-terminal amino acid (NTAA), the method comprising: attaching the peptide to be sequenced via its N-terminus to a solid substrate to form an immobilized peptide; functionalizing the initial C-terminal amino acid of the immobilized peptide with a universal docking strand (DS) ssDNA oligonucleotide; contacting the DS with an imaging strand (IS) oligonucleotide complementary to the DS oligonucleotide, the IS conjugated to a first fluorophore; obtaining single-molecule fluorescence lifetime (FLIM) measurements of the first fluorophore of each peptide; optionally, repeating the single-molecule FLIM measurements for one or more additional combinations of IS and fluorophores in a library; cleaving the initial C-terminal amino acid from the peptide to reveal a second C-terminal amino acid; and optionally, performing another analysis cycle on the second C-terminal amino acid.

[0404] 33. A method of sequencing a peptide substantially as described herein.

[0405] 34. The method of embodiment 33, wherein the method comprises detecting at least one spectral characteristic of the signaling molecule, wherein the spectral characteristic is not fluorescence lifetime.

[0406] 35. A kit for performing the method according to any one of embodiments 1 to 34, comprising at least one pair of an IS and a DS.

[0407] 36. The kit of embodiment 35, comprising at least two pairs of IS and DS, wherein the two pairs differ in the fluorophore, sequence, or both contained in the IS.

[0408] 37. A compound of formula (II)

[0409]

[0410] or a salt or solvate thereof, wherein: x is 0, 1 or 2; each R is independently selected from the group consisting of: C1-C6 alkyl, -NO2, halogen, -C=OR, -C=SR, -C=ONR, -C=OOR, -SO3 or any other common electron-withdrawing group; R 1 and R 2 are independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, —O—(C1-C6 alkyl), C1-C6 alkyl, hydroxy, halogen, —O-alkyl, —S-alkyl, —OC(═O)R, —N—(C═O)—R, —OC(═O)OR, —NC(═S)NR, —N—(C═O)—OR, or any other common electron-donating group; y is 0, 1, 2, or 3; and each R 3independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group.

[0411] 38. The compound of embodiment 37, or a salt or solvate thereof, wherein x is 0.

[0412] 39. The compound of embodiment 37, or a salt or solvate thereof, wherein y is 0.

[0413] 40. The compound of embodiment 38, or a salt or solvate thereof, wherein y is 0.

[0414] 41. The compound according to any one of embodiments 37 to 40, wherein: R 1 or R 2 The C1-C6 alkyl group is a methyl group, and R 1 or R 2 -O-(C1-C6 alkyl) is methoxy.

[0415] 42. The compound of embodiment 37 having the structure:

[0416]

[0417] wherein R1 and R2 are each selected from the group consisting of: H, CH3, OH, and OCH3; provided that R1 and R2 are the same; or a salt or solvate thereof.

[0418] 43. The compound of embodiment 42 which is (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)phenyl)carbamylthioic acid pivalic acid thioanhydride; or a salt or solvate thereof.

[0419] 44. A method for preparing a compound of formula (I)

[0420]

[0421] or its salt or solvate, the method comprising the steps of:

[0422]

[0423] Converted into a compound of formula (II) or a salt or solvate thereof

[0424]

[0425] and thereafter converting the compound of formula (II) or its salt or solvate into the compound of formula (I) or its salt or solvate, wherein: x is 0, 1 or 2; each R is independently selected from the group consisting of: C1-C6 alkyl, -NO2, halogen, -C=OR, -C=SR, -C=ONR, -C=OOR, -SO3 or any other common electron-withdrawing group; R 1 and R 2 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, -O-(C1-C6 alkyl), halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group; y is 0, 1, 2, or 3; and each R 3 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group.

[0426] 45. The method according to embodiment 44, wherein the compound of formula (III) or a salt or solvate thereof is first converted to a compound of formula (IV) or a solvate thereof,

[0427]

[0428] The compound of formula (IV) or a salt or solvate thereof is then converted to the compound of formula (II) or a solvate thereof.

[0429] 46. ​​The method of embodiment 45, wherein the conversion of the compound of formula (III) or a salt or solvate thereof to the compound of formula (IV) or a salt or solvate thereof occurs by reacting carbon disulfide (CS2) with the compound of formula (III).

[0430] 47. The method of embodiment 46, wherein the reacting occurs in the presence of a base.

[0431] 48. The method of embodiment 47, wherein the base is (C1-C6 alkyl)3N.

[0432] 49. The method of embodiment 48, wherein the (C1-C6 alkyl)3N is triethylamine.

[0433] 50. A method according to embodiment 45, wherein the conversion of the compound of formula (IV) or its salt or solvate to the compound of formula (II) or its salt or solvate occurs by reacting the compound of formula (IV) or its salt or solvate with di-tert-butyl carbonate (O-(C(=O)-OC(CH3)2)2).

[0434] 51. The method of embodiment 50, wherein the reacting occurs in the presence of one or more bases.

[0435] 52. The method of embodiment 51, wherein the one or more bases comprise dimethylaminopyridine (DMAP) and triethylamine.

[0436] 53. The method of embodiment 45, wherein the compound has the structure:

[0437]

[0438] wherein R1 and R2 are each selected from the group consisting of: H, CH3, OH, and OCH3; provided that R1 and R2 are the same; or a salt or solvate thereof.

[0439] 54. The method of embodiment 53, wherein the compound is (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)phenyl)carbamylthioic acid pivalic acid thioanhydride; or a salt or solvate thereof.

[0440] 55. Use of a compound according to any one of embodiments 37 to 54 in a peptide assay as described herein.

[0441] (XVIII) Experimental Examples

[0442] Example 1: Peptide sequencing by detecting fluorescence lifetime perturbations

[0443] Fluorescence lifetime-based peptide sequencing ( Figures 1A-1G) has been demonstrated, demonstrating that fluorescence lifetime can be influenced by the selected sequence of the DNA imaging strand (IS), including synthetic modifications to the nucleobase or sugar-phosphate backbone; the selected fluorophore; and the identity of the N-terminal amino acid (NTAA), including intrinsic chemical variations or post-translational modifications (PTMs) of the NTAA side chain. This approach is independent of the chemical identity of the NTAA and, therefore, can detect and identify atypical or unnatural amino acids in addition to proteinogenic amino acids and their associated PTMs. The NTAA is repeatedly interrogated with different IS-fluorophore conjugates, where either or both the IS sequence and the fluorophore can be varied, to generate a set of fluorescence lifetime data. Edman degradation is performed to demonstrate that the NTAA can be removed to generate a new NTAA, which can then be interrogated to generate fluorescence lifetime data. This new NTAA can be repeatedly interrogated with different IS-fluorophore conjugates to generate a set of fluorescence lifetime data. This workflow can be repeated until all or a subset of amino acids within the polypeptide chain have been interrogated. Edman degradation can be performed sequentially to sequentially expose amino acids at the N-terminal position of the polypeptide chain. All or a subset of these novel NTAAs can be interrogated with IS-fluorophore conjugates to generate data corresponding to the identity of the NTAA and the location of each amino acid within the peptide chain based on the number of Edman degradation cycles performed.

[0444] In the examples presented, synthetic peptides are used as substitutes for naturally occurring peptides or component peptides of enzymatically or chemically digested proteins. These synthetic peptides are attached to the glass surface via the sulfhydryl group of the C-terminal cysteine, but the C-terminus can also be covalently attached to the surface via the C-terminal carboxyl group. This can be achieved by preparing an amine-functionalized surface, such as by silanization with (3-aminopropyl) triethoxysilane (APTES, CAS Reg. No.: 919-30-2, IUPAC: 3-(triethoxysilyl)propan-1-amine) and using a coupling agent such as a carbodiimide (e.g., N,N'-diisopropylcarbodiimide (DIC) or 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC)) or N-hydroxysuccinimide (NHS) to covalently attach the C-terminal carboxyl group of the peptide to the surface. In order to consistently execute this strategy and minimize side reactions, the side chains containing amino, hydroxyl, and carboxyl groups can be first protected. In addition, in the case of a DNA docking strand (DS) with or without conjugation to the maleimido group of a maleimidophenylisothiocyanate (MPITC) linker, the N-terminal amino group can also be reversibly protected or otherwise conjugated to the isothiocyanate group of the linker.

[0445] In one experiment, eleven different synthetic peptides that differed only in the identity of the N-terminal amino acid (NTAA) were interrogated with IS1-AF488 and IS1-BODIPY-FL. Figure 4 ). Unique differential responses were detected for tryptophan, arginine, phenylalanine, serine, and phosphoserine NTAA peptides. In addition, for serine, arginine, and phenylalanine NTAA peptides, the responses measured for IS1-AF488 and IS1-BODIPY-FL were different.

[0446] Each example was tested in different assemblies to determine if any of the components of the example would confound the fluorescence signal. Each assembly was made in a separate well of an 8-well chamber slide according to the current protocol. Fluorescence microscopy was performed and fluorescence intensity was collected from each sample under the same excitation and emission parameters. Figure 5 As shown, the fluorescence intensity is higher when the imager strand is attached to the construct. This indicates that the collected fluorescence lifetime data is mainly dominated by the fluorophore and suggests that other components of the construct do not produce significant fluorescence.

[0447] To demonstrate the effect of fluorescence lifetime changes of different fluorophores conjugated to the same IS1 imaging strand, serine and phosphoserine NTAA peptides were conjugated to glass and used to compare the lifetimes of AF488, BODIPY-FL, BODIPY-TR, and TAMRA (carboxytetramethylrhodamine). In the experiments, the NTAA peptide remained unchanged while the IS1-conjugated fluorophore complex was cycled. The results showed that the fluorophore could be successfully and completely removed, as when changing from a green fluorophore (such as AF488 or BODIPY-FL) to a red fluorophore (such as BODIPY-TR or TAMRA), the signal of the previous fluorophore in the green channel was minimal (not shown). Additionally, the fluorescence lifetime results suggest that some fluorophores may be more sensitive to certain NTAAs than others ( Figure 6 ).

[0448] To test whether alternative imager strands alter the fluorescence lifetime of the fluorophore, three different synthetic peptides were sequentially interrogated with different DNA imager strands (IS) and fluorophore conjugates (IS-fluorophore). The three peptides differed only in the identity of the N-terminal amino acid (NTAA): the first peptide contained an N-terminal glycine residue (G), the second peptide contained an N-terminal arginine residue (R), and the third peptide contained an N-terminal tryptophan residue (W). These peptides were sequentially interrogated with the following IS-fluorophore conjugates (IS sequences as shown in ( Figure 2 ) shown): IS1-AF488, IS2-AF488, IS3-AF488 and IS3-BODIPY. For each different IS-fluorophore conjugate, the fluorescence lifetime was recorded. The data show that both the sequence of the IS and the choice of fluorophore affect the measured fluorescence lifetime ( Figure 9). In short, for each IS-AF488 combination, the measured fluorescence lifetime of the tryptophan NTAA peptide was suppressed compared to the fluorescence lifetimes of the glycine and arginine NTAA peptides. However, for IS3-BODIPY, the measured fluorescence lifetime was less suppressed in the case of the tryptophan-containing peptide than in the case of the other two peptides. In addition, for IS1-AF488, the measured fluorescence lifetime of the arginine NTAA peptide was longer than that of the glycine NTAA peptide, but this trend was reversed for IS2-AF488 and IS3-AF488.

[0449] To thoroughly test the complete workflow, multiple synthetic peptides were interrogated with IS1-AF488 and IS1-BODIPY-FL before and after multiple cycles of Edman degradation (Figures 7 and 8). Prior to Edman degradation, in the case of IS1-AF488 and IS1-BODIPY-FL, the arginine N-terminal amino acid (NTAA) peptides exhibited higher normalized fluorescence lifetime measurements than the glycine NTAA peptides, and the tryptophan NTAA peptides exhibited lower fluorescence lifetime measurements than the glycine NTAA peptides. Additional measurements showed that phenylalanine, tyrosine, histidine, and methionine had similar fluorescence lifetimes. Upon Edman degradation to remove the NTAA from each synthetic peptide and expose the N-terminal glycine residue of each peptide, the fluorescence lifetimes of IS1-AF488 and IS1-BODIPY-FL converged to similar values ​​measured for the original glycine NTAA peptide as predicted by theory ( Figures 7A-7B ). Additional experiments were performed in which multiple Edman cycles were performed on synthetic peptides containing tryptophan or arginine in the second ("N-1") or third ("N-2") position along the peptide. Figures 8A-8H As shown, the fluorescence lifetimes of IS1-AF488 as well as IS1-BODIPY-FL fluctuate after each Edman degradation cycle, indicating that the new N-terminal amino acid after Edman is changing the fluorescence lifetime compared to the pre-Edman state.

[0450] These experiments demonstrate the effectiveness of the workflow (e.g. Figures 1A-1G ), demonstrating that fluorescence lifetime measurements vary depending on the choice of DNA imaging strand (IS) sequence, the choice of fluorophore conjugated to the IS, and the identity of the peptide's N-terminal amino acid (NTAA). Furthermore, proof-of-concept experiments using Edman degradation demonstrate that the workflow can be repeated for each subsequent amino acid within the peptide sequence until some or all of the amino acids have been interrogated.

[0451] method

[0452] Molecular Dynamics Simulations: Modeling and simulations were performed to predict the optimal linker length and functional dye-N-terminal amino acid combinations. The DNA docking strand (DS) and imaging strand (IS)-fluorophore oligonucleotide were modeled using a combination of homology modeling and energy minimization, followed by equilibration in all-atom molecular dynamics (MD) simulations in water. The Molecular Operating Environment (MOE) software was primarily used for modeling and initial energy minimization.

[0453] All-atom MD simulations were performed on an Exacloud cluster at Oregon Health & Science University (Portland, OR) using the GRONINGEN MAchine for Computer Simulations (GROMACS-2018).

[0454] All-atom AMBER-type force field parameters were generated for the system subjected to all-atom MD simulations. After ensuring the correct protonation state of the molecules, the AMBER tool was used to calculate the partial charges at HF / 6-31G (AM1-BCC), and the General AMBER Force Field (GAFF) was used for bonding and van der Waals interactions.

[0455] MD simulations were performed in an aqueous environment. The TIP3P water model was used to simulate the water with an appropriate number of counterions (Na + or Cl - ) in an aqueous environment to ensure charge neutrality. The complex was centered using a 3D periodic box, at least 1.0 nm from the edge and accounting for >2 nm of the solvent buffer. 5 ns equilibrium and 100 ns generation runs were run in the NPT ensemble using a V-rescale thermostat and a Parrinello-Rahman barostat, respectively, with the temperature maintained at 300 K and the pressure maintained at 1 bar. The MD simulations incorporated a leapfrog algorithm with a 2 fs time step to integrate the equations of motion. Long-range electrostatic interactions were calculated using the particle mesh Ewald (PME) algorithm with a real space cutoff of 1.2 nm. LJ interactions were also cut off at 1.2 nm. The LINCS algorithm was used to constrain the motion of hydrogen atoms bonded to heavy atoms. The coordinates of the protein molecules were stored every 1 ps for additional analysis. In order to monitor the system reaching equilibrium, the root mean square deviation (RMSD) of the complex structure was calculated over time.

[0456] Surface functionalization and synthetic peptide attachment: Glass coverslips were etched with 6 mM KOH at room temperature for 20 minutes. The surface was then functionalized with a 1:1 solution of silane-PEG:silane-PEG-maleimide to evenly distribute the maleimide functionalized groups on the glass surface for the addition of peptide linkers via click chemistry. The silane linkers were dissolved in a solution of 95% ethanol, 1% acetic acid, and the remainder ultrapure water, with a final concentration of 3.2 mM for each linker. The glass coverslips were incubated with the silane linker mixture at room temperature for 30 minutes. After incubation, the coverslips were rinsed with ultrapure water. Each synthetic peptide containing a C-terminal cysteine ​​residue was dissolved in ultrapure water at a final concentration of 1 μM, added to the glass surface, and incubated at room temperature for 4 hours to covalently attach the thiol group of the cysteine ​​side chain to the maleimide group on the functionalized surface.

[0457] Synthesis of Maleimidophenyl Isothiocyanate (MPITC) Linker: To synthesize the MPITC linker, 0.2 mmol of the starting material (CAS Registry Number: 29753-26-2, IUPAC: 1-(4-aminophenyl)-1H-pyrrole-2,5-dione) was dissolved in 5 ml of anhydrous ethanol in a round-bottom flask. 10 molar equivalents of carbon disulfide (5 M in tetrahydrofuran) were added to the flask. Triethylamine was added to the starting material in a 1:1 molar ratio, and the reaction mixture was stirred at 100 rpm at room temperature for 30 minutes using a stir bar. The flask was transferred to an ice bath. A catalytic amount (3 mol% of the starting material) of 4-dimethylaminopyridine (DMAP) was dissolved in anhydrous ethanol. With stirring, the DMAP solution and 98 mol% (relative to the starting material) of di-tert-butyl decarbonate were added to the flask simultaneously. The reaction mixture was incubated in an ice bath with stirring for 5 minutes and then transferred to room temperature and stirred at 200 rpm overnight. The solvent was removed by rotary evaporation at 40°C and 175 mbar until visibly dry, and then at 40°C and 0 mbar for 5 minutes to give a white crystalline solid.

[0458] The products were confirmed using nuclear magnetic resonance spectroscopy (NMR) and Fourier transform infrared spectroscopy (FTIR). Proton and carbon NMR of the products were performed in deuterated chloroform on a 400 MHz Avance NEO NanoBay spectrometer (Bruker, Inc.). FTIR was performed on a Nicolet iS5 KBR window FTIR spectrometer (Thermo Fisher Scientific) with an iD7 anti-reflection diamond crystal attenuated total reflectance (ATR) module. 2 μl of 7.4 mM MPITC or 2 μl of 7.4 mM starting material in anhydrous ethanol was added directly to the ATR crystal. Each sample was dried in a clean, dry air stream at 2 cm -1 Resolution from 4000cm-1 Up to 400cm -1 The next scan is 256 times.

[0459] The peptide was conjugated to the DNA docking strand using the MPITC linker: the N-terminal amine of the peptide was first covalently linked to the isothiocyanate group of MPITC ( Figure 3 ). Anhydrous ethanol containing 7.46 mM MPITC was incubated on the peptide-functionalized glass coverslip for 5 hours to complete the conjugation. The MPITC-peptide-functionalized glass coverslip was washed with ultrapure water. Next, a DNA docking strand (DS) oligonucleotide with 3' (3-mercaptopropyl) phosphate was covalently attached to the maleimide group of MPITC ( Figure 3 A 1 nmol amount of a 100 μM DS solution was diluted to 10 μM in DNA buffer containing 1 μl of tris(2-carboxyethyl)phosphine (TCEP) to reduce 3' modifications and expose free sulfhydryl groups. This mixture was added to the MPITC-peptide-functionalized surface and incubated at room temperature for 3 hours to obtain a DS-MPITC-peptide-functionalized surface.

[0460] Preparation of DNA imaging strand-fluorophore conjugates: To conjugate each fluorophore to each DNA imaging strand (IS), the maleimide-functionalized fluorophore was dissolved in pure DMSO at a concentration of approximately 20 mM, and the IS with a 5' (6-aminohexyl) phosphate modification was dissolved in ultrapure nuclease-free water to 1 mM. For each fluorophore and IS combination, the fluorophore and IS were combined together in ultrapure water containing 10% 1M NaHCO3 at a dilution of 1:10 and 1:20, respectively, with a pH of approximately 8.0. The fluorophore and IS were reacted at room temperature on a ThermoMixer (Eppendorf GmbH) at 500 rpm to form a covalent bond between the 5' amino group of the IS and the maleimide group of the fluorophore ( Figure 3 ).

[0461] To purify the IS-fluorophore conjugate from excess fluorophore, the solution was mixed with 3M sodium acetate in a 1:1 ratio. This solution was diluted 1:5 in absolute ethanol and incubated at -80°C overnight to precipitate the IS-fluorophore conjugate. The precipitated IS-fluorophore conjugate was pelleted by centrifugation at approximately 20,000 × g for 30 minutes at 2°C, the supernatant removed, and the pellet resuspended in absolute ethanol at -20°C; these centrifugation wash steps were repeated three times. The pellet was resuspended in 50 μl of ultrapure water and diluted 1:1 with 3M sodium acetate. Absolute ethanol was added to a concentration of 80% v / v, and the solution was incubated at -80°C overnight. The precipitated IS-fluorophore conjugate was pelleted and resuspended three times as described, and the pellet was then lyophilized for 10 minutes or until dry. The lyophilized conjugate was resuspended in ultrapure water. The concentration of IS was determined by measuring the absorbance at 260 nm based on the absorption properties and extinction coefficient of the fluorophore, and the absorbance was measured to calculate the dye concentration. The IS-fluorophore conjugate was stored at -20°C.

[0462] DNA docking strand: imager strand hybridization and imager strand removal: To attach the imager strand (IS)-fluorophore conjugate to the peptide-conjugated DNA docking strand (DS), 1 nmol of 10 μM IS-fluorophore conjugate in DNA buffer was added to the DS-MPITC-peptide functionalized surface and incubated at room temperature for 1-5 minutes for hybridization. After hybridization and thorough washing with 10 mM HEPES buffer to remove excess unhybridized IS-fluorophore, fluorescence lifetime measurements were performed.

[0463] To perform multiple interrogations of the N-terminal amino acid (NTAA), the IS-fluorophore conjugate can be removed and an IS-fluorophore conjugate with a different oligonucleotide sequence and / or fluorophore can be attached to the DS for additional fluorescence lifetime measurements. The IS-fluorophore conjugate is removed from the DS by incubation with a chemical denaturant (such as 8 M urea) for 1-5 minutes, followed by extensive washing with 10 mM HEPES buffer to remove free IS-fluorophore conjugate.

[0464] Fluorescence lifetime measurements: Fluorescence lifetime imaging (FLIM) of all synthetic peptides was performed on a Zeiss LSM 880 scanning confocal microscope equipped with a Chameleon Ti:Sapphire (Coherent) multiphoton source operated at a pulse repetition rate of 80 MHz. 2Multiphoton excitation of the sample was achieved by scanning the excitation beam over an area of ​​​​(100 nm). After passing through a 1.4NA63x magnification objective and a 640nm longpass filter, single-photon fluorescence events were captured on a Big.2 gallium arsenide phosphide photomultiplier tube (GaAsP PMT) to construct a 512 x 512 image with 290nm pixels. The area was scanned approximately 100 times to fill the fluorescence lifetime distribution for each pixel. Time-correlated single-photon counting (TCSPC) was performed using Becker & Hickel TCSPC electronics and SPCM software (B&H).

[0465] FLIM data analysis was performed using the open-source software tool FLIMfit (available online at flimfit.org / ) using Matlab Compiler Runtime R2016b. To characterize the system, the instrument response function (IRF) was determined by second harmonic generation imaging of dried urea crystals on a microscope slide. Using nonlinear least squares, the fluorophore lifetime measurements at each pixel in the image were fitted with the following expression:

[0466]

[0467] Among them I 背景 is the average background intensity measured by ambient light and detector noise. i and τ i are the amplitude and lifetime according to the exponential fit, respectively. By evaluating the mean χ 2 To evaluate the fit of the life curve. 2 A value less than 1.2 and greater than 0.88 is determined to be a good fit for the decay. The average lifetime is calculated from the fitting equation:

[0468]

[0469] where α i is the fractional amplitude of component i in the exponential decay fit, and the sum of the fractional amplitudes of all components equals 1. For samples containing a known single synthetic peptide, the calculated mean lifetime is averaged over all pixels of the image.

[0470] Edman degradation to remove the N-terminal amino acid (NTAA): Edman degradation was performed to remove the N-terminal amino acid (NTAA) and expose the N-terminal amine of the next amino acid in the polypeptide chain. A 0.5% v / v (pH 2) aqueous solution of TFA was added to the DS-MPITC-peptide or IS:DS-MPITC-peptide-functionalized glass surface and incubated at 55°C for 40 minutes. The surface was then washed extensively with 10 mM HEPES buffer to discard any removed NTAA-linked oligonucleotide and prepare for the next reaction with the MPITC linker.

[0471] Microfluidic chip preparation: For some experiments, a microfluidic chip was assembled around a peptide-functionalized glass coverslip, and subsequent chemical reactions were performed within the microfluidic chip.

[0472] The microfluidics contains polydimethylsiloxane (PDMS) guides for inlet and outlet microfluidic tubing. To produce the PDMS for these guides, Sylgard 184 silicone elastomer (Dow, Inc.) was used. The components were combined according to the kit instructions (10:1 dimethylsiloxane to siloxane and silicone), mixed thoroughly, and then removed under vacuum to remove bubbles. The PDMS mixture was poured onto a clean silicon wafer and baked in an oven at 80°C for 90 minutes. The cured PDMS was cut to size with a scalpel, and a 1 mm diameter biopsy punch was used to make the tube channels. The punched PDMS was thoroughly washed in IPA and ultrapure water. Holes to accommodate the microfluidic tubing were drilled on a glass slide using a diamond-reinforced drill bit. The punched PDMS and the drilled glass slide were treated with oxygen plasma under argon carrier gas. The punched PDMS and the plasma-treated surfaces of the drilled glass slide were aligned with an inspection microscope and then baked in a hotplate or oven at 100-105° C. for 30 minutes to adhere the plasma-treated surfaces.

[0473] Next, double-sided adhesive, such as 127 μ m thick 200MP 468MP type acrylic adhesive (3M company) is cut into required geometric shape, such as oval, and is applied around the borehole of slide glass. Then, the peptide-functionalized glass cover glass prepared as described herein is adhered to adhesive tape to form closed microchannel. For fluid handling, thin-walled Teflon (PTFE) tubing, such as TT-26 (Weico Wire & Cable, Inc.) with an inner diameter of 0.018 in and a wall thickness of 0.009 in, is inserted through a PDMS guide, passed through the borehole in the slide glass and entered into the microchannel of the microfluidic chip. All chemical reactions can be carried out by injecting reaction components into the microchannel by a syringe.

[0474] Example 2: Identifying amino acids using machine learning algorithms

[0475] Empirical and Monte Carlo simulated time-correlated single photon counting (TCSPC) lifetime data can be used to train a convolutional neural network (CNN) for amino acid calling. To this end, fluorescence lifetimes (e.g., generated using the method described in Example 1) can be abstracted into an array of intensity pixels weighted by the average lifetime of a single molecule. The resulting image corresponds to a unique amino acid fingerprint that can be interrogated by a deep learning network.

[0476] The recursive nature of the described method enables the probing of individual residues and the construction of a large M × N intensity matrix, where the fields M and N are "fluorophore" and "imager strand," respectively. To minimize residual errors in the fit, these conditions can be combined with other spectroscopic techniques to construct an n-dimensional tensor for input to a CNN. For example, a CNN can receive a 3-dimensional tensor with the fields M, N, and L, where L is the average value derived from the fluorescence autocorrelation function.

[0477] Example 3

[0478] This example demonstrates that the use of alternative imager strand designs can be used to increase the flexibility of the provided peptide sequencing methods and systems.

[0479] The sequences and schematic arrangement of the single-stranded DNA (ssDNA) docking strand and imaging strand are shown in Figure 2 The ssDNA imaging strand (IS) can be modified in various ways, such as including modified or non-natural nucleotides, including a 5' IS overhang relative to the docking strand (DS), or including a 5' IS non-overhang relative to the DS. These variables adjust the spatial position and / or degree of freedom of the attached fluorophore and, thereby, the interaction with the NTAA side chain, which modulates the measured fluorescence lifetime. This enables the identification of each NTAA.

[0480] The same peptide was used but with different imager strands (IS1 (SEQ ID NO: 2), IS2 (SEQ ID NO: 3), and IS3 (SEQ ID NO: 4)) and fluorophores (Alexa Fluor in the top panel). 488; the data obtained in the case of BODIPY at the bottom are shown in Figure 9 Data are normalized to the fluorescence lifetime of the corresponding free fluorescent dye. IS1 and IS2 have higher melting temperatures, and IS3 has a single-base overhang. These data demonstrate that alternative IS designs can be used to increase the flexibility of peptide sequencing analyses.

[0481] Example 4

[0482] This example demonstrates that the peptide analysis methods provided herein can be used to generate unique peptide identification fingerprints.

[0483] The fluorescence lifetimes of AF488 and BODIPY near the N-terminus of the peptide were plotted according to the workflow ( Figure 10A The marker radius represents the standard deviation of the bulk distribution of lifetimes measured for each dye-AA. The W and G lines are drawn to demonstrate the unique fingerprint recognition between each amino acid.

[0484] Using the empirically derived means and standard deviations of the fluorescence lifetime measurements observed in the workflow, simulated 2-dimensional Gaussian distributions of dye-AA lifetimes for AF488 and BODIPY were generated. Figure 10B The two-dimensional distributions in are plotted as scatter histograms to demonstrate how multiple dyes can be used to name (identify) amino acids. Alternative fluorescence lifetime fingerprinting highlights the prediction of different N-terminal amino acids using two separate fluorophores attached to IS1. The point spread fluorescence lifetimes of AF488 and BODIPY-FL are depicted for tryptophan, glycine, and arginine. As shown, predicting the fluorescence lifetime differences between the three amino acids using only one dye can be challenging; however, separation of the species can be observed when two dyes are used. Therefore, Figure 10B Demonstrating that fluorophore cycling generates robust amino acid fingerprints.

[0485] Complex multivariate results improve the accuracy of neural network predictions. The dye-AA lifetimes generated in the workflow are used to generate 3D arrays ( Figure 10C The intensity of the pixel blocks represents lifetime, the Y-axis represents dye measurements for different imager strands, and the X-axis represents unknown amino acids. These image arrays generate fingerprints that can be used in convolutional neural networks for amino acid prediction and sequencing. Using a theoretical neural network approach for amino acid identification, the fluorescence lifetime fingerprints of different N-terminal amino acids are illustrated.

[0486] Example 5

[0487] This example provides evidence that the provided methods can also operate when the free N-terminus of the peptide is linked to a modified nucleotide within the DS oligonucleotide sequence rather than at (or near) the terminus of the DS oligonucleotide. As demonstrated in this example, the fluorophore is linked to a modified nucleotide within the IS oligonucleotide sequence rather than at (or near) the terminus of the IS oligonucleotide.

[0488] Figure 11An alternative embodiment of the peptide sequencing system provided using a DNA major groove design is presented. This embodiment provides additional adjustable control for the interaction between the fluorescent dye and the structural components of the DNA docking strand: imaging strand (DS:IS) complex. In contrast to the embodiment in which the fluorescent dye interacts with the blunt-ended nucleobases of the DS:IS complex, in this alternative embodiment, the fluorescent dye is conjugated inside the IS and is therefore unable to reach the blunt end of the complex. This has been verified by computer-simulated molecular dynamics simulations. The immobilized peptide is conjugated to the modified nucleotides within the DS (i.e., not directly near either end of the DS), and the fluorophore is conjugated to the modified nucleotides within the IS (i.e., not directly near either end of the IS). Computer-simulated molecular dynamics simulations have shown that for this embodiment, the peptide and fluorophore can be positioned within the major groove of the double-stranded DNA to maximize the interaction between the fluorophore and the N-terminal amino acid side chain of the peptide. Figure 11 One such configuration generated by MD simulations described in the Methods section is shown. In this configuration, the N-terminus of the peptide is linked to a thiol-modified dT nucleobase on a DS oligonucleotide using a maleimidophenyl isothiocyanate bifunctional linker. The NHS fluorophore is linked to an amine-modified dT nucleobase on an IS oligonucleotide. 3D modeling revealed that the attachment of the peptide and fluorophore at the fifth position of the nucleobase positions the peptide and fluorophore in the major groove of the DNA double helix and on the 3' side of its corresponding oligonucleotide. For the linker described here, MD simulations showed that the interaction between the fluorophore and peptide is maximized when they are kept four to six bases apart. Furthermore, positioning the fluorophore away from the end of the oligonucleotide minimizes the interaction of the fluorophore with the nucleobases of both the DS and IS oligonucleotides.

[0489] Example 6

[0490] This example demonstrates the change in the fluorescence lifetime of a fluorophore conjugated to IS1.

[0491] Using methods substantially similar to those described in Example 1, various fluorophores were conjugated to IS1, and each IS was suspended in water. The fluorescence lifetime ( Figure 12 Longer-lived fluorophores exhibit greater lifetime differences, with KU560-6 having a significantly higher lifetime among the KU dyes. These larger lifetime differences may improve sensitivity when measuring lifetime differences between AAs with small differences, such as those measured by AF488 and BODIPY-FL.

[0492] Example 7

[0493] This example demonstrates that the peptide analysis method provided herein can be used to analyze KU with a longer lifespan. TM The dye generates a unique fingerprint for peptide identification.

[0494] Figure 13 is a bar graph showing various fluorescence lifetimes measured by various imager strands containing the long-lived fluorophores KU530-6 and KU530-R-4 to identify the N-terminal amino acid.

[0495] To screen peptides, the fluorescence lifetime was measured by the long-lived fluorophores KU530-6 and KU530-R-4 when close to W, F, Y, G, H, M, Q, E, S, R at the N-terminus in an established workflow ( Figure 13 When comparing the lifetimes of specific AAs reported by each dye, there were discrepancies in the measurements for W, Y, and H. In addition, KU530-6 showed significant variations in lifetimes between several amino acids. When combined with the previously tested AF488 and BODIPY-FL, the additional dyes helped define a characteristic lifetime fingerprint for each amino acid that could be used to distinguish terminal AAs.

[0496] Example 8

[0497] This example demonstrates that the peptide analysis methods provided herein can be used to identify post-translational modifications on NTAA.

[0498] Fluorescence lifetime was measured by imager strands (IS1) conjugated to Alexa Fluor 488, BODIPY-FL, or KU530-6 hybridized to surface-bound peptides with common post-translational modifications of phosphorylation of serine or acetylation of lysine at the N-terminus. In the case of acetylated lysine and phosphorylated serine, the lifetimes measured by AF488 did not differ; however, KU530-6 reported different lifetimes when comparing the two PTMs ( Figure 14 This highlights the potential of dyes with longer lifetimes to discern some common post-translationally modified amino acids. When comparing the lifetimes of phosphorylated serine, the lifetime reported by BODIPY-FL was significantly higher than that reported by AF488, suggesting that using multiple dyes may be helpful in elucidating PTMs. [*, p < 0.05; one-way ANOVA with Tukey's post hoc test; multiple fields within N = 3 samples]. This suggests that each dye has different interactions with various amino acids and, therefore, different lifetimes, which could ultimately be used to sequence peptides.

[0499] To further elucidate the molecules responsible for the lifetime differences, fluorescence lifetimes were collected from IS1 conjugated to AF488 in SGG and control samples of post-translationally modified serine (PhosSGG). After FLIM was performed, IS1-AF488 was removed and phosphatase was added to the PTM sample to remove the phosphorylation group (PhosSGG) on the modified serine. IS1-AF488 was reintroduced to hybridize with DS and FLIM was performed. Figure 15 In the figure, the bar graph depicts the lifespan measured by unmodified serine (SGG), phosphorylated serine PTM (PhosSGG), and dephosphorylated PTM (DePhosSGG). The results show that the lifespan reported by dephosphorylated serine (DePhos) is not significantly different from that of the unmodified control (SGG), and that phosphorylated serine PTM is significantly lower than that of the control [*, p < 0.05; one-way ANOVA with Tukey's post hoc test; multiple regions within N = 3 samples]. This suggests that phosphorylation of modified serine is altering lifespan. In addition, AF488 conjugated to IS1 may be a good candidate for detecting differences between PTMs and unmodified peptides.

[0500] Example 9

[0501] This example provides a visual depiction of sequencing data to interpret the data and identify variations in the collected lifespan data, and further demonstrates the power of this sequencing approach.

[0502] Normalized lifetime “heat maps” from sequencing data as reported by four separate fluorophore-conjugated imager strands as well as peptide screening experiments ( Figure 4 and 13 ). Each data point was normalized to the measured lifetime of GGGS of each fluorophore-conjugated IS and patterned based on the corresponding normalized lifetime range. Following the established workflow, peptides ending in G, W, F, Y, H, M, Q, E, S, or R were attached to the solid substrate, with each group in a separate well. After DS ligation, IS1-AF488 was hybridized and lifetime data for each individual peptide sequence was collected. Dehybridization of IS1-AF488 was performed, followed by hybridization of different IS1-fluorophore conjugates. Lifetimes were collected. Figure 4 and 13 This cycle was repeated for all IS1-fluorophore conjugates listed in (BODIPY-FL, KU530-6, and KU530-R-4). The lifetime information collected from each peptide and fluorophore was normalized relative to the GGGS of each fluorophore. The data were compiled based on their normalized lifetimes and presented in a heatmap ( Figure 16). Each amino acid had a different lifetime pattern with the four dyes, with no two showing identical response patterns. This suggests that by using additional dyes, even for a given fluorophore-IS combination, amino acids with similar lifetimes can be resolved. Similar testing and mapping of post-translational modifications on terminal amino acids was performed for phosphorylated serine and acetylated lysine using AF488, BODIPY-FL, and KU530-6 conjugated to IS1. The normalized lifetime for each fluorophore had similar characteristics to the modified amino acid and did not respond consistently with any other characteristic, indicating that PTMs can also be successfully identified using this approach.

[0503] Example 10

[0504] This example determines the effect of certain additional amino acids along the peptide on the measured fluorescence lifetime. The amino acids at the second position ("N-1") and the third position ("N-2") of the N-terminus are modified with glycines surrounding the amino acids. Lifetimes are collected using a standard workflow.

[0505] Tryptophan and arginine were incorporated at the second ("N-1") and third ("N-2") positions along the otherwise identical peptide to determine whether the amino acid at N-1 or N-2 affected the fluorescence lifetime ( Figure 17 AF488 was conjugated to IS1 and used with each peptide. Peptides terminated with glycine (W) or R at N-1 or N-2 showed no statistically significant differences in lifetime compared to the control GGGS, indicating that AF488 is only sensitive to the terminal AA of the peptide. [N = multiple fields within 3 samples]. However, a slight increase in lifetime was observed for R at N-2, suggesting that IS1-AF488 may be sensitive to arginine at another position along the peptide sequence.

[0506] Tryptophan and arginine were incorporated at the second ("N-1") and third ("N-2") positions along the otherwise identical peptide to determine whether the amino acid at N-1 or N-2 affected the fluorescence lifetime ( Figure 18 ). BODIPY-FL was conjugated to IS1 and used with each peptide. Peptides terminated with glycine or R at N-1 or N-2 showed no statistically significant differences in lifetime when compared to the control GGGS, indicating that BODIPY-FL was only sensitive to the terminal AA of the peptide. However, a slight increase in the lifetime of R was observed at N-1, and an increase in the lifetime of R was observed at N-2 compared to GGGS, suggesting that IS1-BODIPY-FL may be sensitive to arginine at other positions along the peptide sequence. This IS1-fluorophore can be used to further sequence amino acids along the peptide sequence without Edman degradation with the help of a trained machine model.

[0507] exist Figure 19 In this study, fluorescence lifetime was measured with imager strands conjugated to KU530-6 in the presence of peptides with tryptophan or arginine at the second and third positions along the peptide to determine whether the amino acid at N-1 or N-2 affected the lifetime of the fluorophore. KU530-6 exhibited minimal sensitivity to amino acids located elsewhere along the sequence, ultimately highlighting its specificity in interacting with terminal amino acids and manipulating dye lifetime. [*, p < 0.05; ***, p < 0.005, one-way ANOVA with Tukey's post hoc test; N = 3 samples within multiple fields of view].

[0508] Example 11

[0509] This example demonstrates the complete workflow for this embodiment. This includes binding the construct to a glass substrate, performing FLIM, imaging the strands, and performing cyclic Edman degradation to further determine the peptide sequence. This example focuses on the Edman degradation process and the resulting lifetime changes based on the novel terminal amino acid.

[0510] Following an established workflow, peptides with sequences of different NTAAs (G, W, R, F, Y, H, and M) and identical remaining sequences were attached to a glass substrate. An imaging strand (IS1) containing KU530-6 was hybridized to the conjugated DS on the peptide. FLIM was performed on each set and then Edman degradation was performed to remove the NTAA as well as the DS, linker, and IS. A new linker was attached after DS conjugation. The new IS1-KU530-6 was hybridized and FLIM was performed on each set (Edman-cleaved amino acids are shown in brackets). The collected data were normalized to GGGS ( Figure 20 Differences in lifetimes were observed for tryptophan and arginine as terminal amino acids compared to their Edman-cleaved counterparts, indicating that Edman degradation is complete with lifetimes stabilizing to values ​​similar to those of GGGS [**, p < 0.01; one-way ANOVA with Tukey post hoc test; N = 3 samples across multiple fields]. Furthermore, after the first Edman degradation cycle in each group, lifetimes were similar between groups, demonstrating the power of the existing workflow to perform Edman degradation cycles and the similar values ​​reported for each group for IS1-KU530-6.

[0511] The amino acids in the longer peptides were analyzed sequentially to show Figures 1A-1GFeasibility of the complete workflow presented. A solid glass surface was etched with potassium hydroxide to provide -OH groups on its surface. Silane-PEG-maleimide was added to bind the silane to the -OH groups. Synthetic peptides with the sequence WGRSGGDC, RGWSGGDC, or WRGSGGSDC were immobilized on a glass substrate via a cysteine-maleimide reaction for testing. Each peptide was conjugated with an MPITC linker before covalently attaching it to a DNA docking strand (DS) oligonucleotide. AF488 was conjugated to IS1 separately, then added to the DS-MPITC-peptide and bonded to hybridize the DS and IS. The fluorescence lifetime of the fluorophore was measured. The IS1-AF488 was then dehybridized and IS1-BODIPY was added to hybridize with the DS. FLIM of the IS was performed. Each IS1-conjugated fluorophore was cycled until all four fluorophores were used. Edman degradation was then performed to cleave the N-terminal AA. The new NTAA was then exposed to the same linker and DS ligation process, followed by hybridization with the IS-fluorophore construct and measurement of FLIM, IS-fluorophore recycling and Edman degradation.

[0512] This complete process was completed three times and the sequential lifetime measurements are presented in Figures 21A-21C Using BODIPY-FL( Figures 22A-22C )、KU530-6( Figures 23A-23C ) and KU530-R-4( Figures 24A-24C ) completed the same workflow. The data from these experiments are Figure 25 The results are summarized in a heatmap, where the fluorescence lifetimes for each analysis step are normalized to GGGS. Each peptide in the peptide library exhibits a unique range of lifetimes for each dye, demonstrating that the entire workflow can be successfully implemented with existing fluorophores to sequence intact peptides. Furthermore, the results suggest that amino acids in the N-1 and N-2 positions may also influence the fluorescence lifetime of each dye; however, the use of machine learning algorithms will help further elucidate these differences and ultimately enable the sequencing of intact peptides.

[0513] exist Figure 25 In, for example Figures 23A-24CThe measured lifetimes reported in were normalized and combined to create a lifetime "heat map" depicting the sequencing data for the four individual fluorophore-conjugated imaging strands. The heat map shows the average lifetime values ​​normalized to GGGS for each respective fluorophore and patterned based on the corresponding normalized lifetime range. Each amino acid produced a different lifetime pattern with the four dyes and for each peptide sequence. This suggests that by using the selected dyes, amino acids with similar lifetimes can be resolved for a given fluorophore-IS combination. In addition, this heat map highlights the sensitivity of some fluorophores to amino acids at N-1 and N-2, such as IS1-AF488 and WGRSGGSDC (SEQ ID NO:5). After three cycles of Edman degradation, the sequences of all peptides contained the same sequence (XXX)SGGSDC (SEQ ID NO:8). The lifetimes measured by this sequence did not differ significantly between peptides for BODIPY-FL, KU530-6, and KU530-R-4; however, for IS1-AF488, the reported lifetime for (RGW)SGGSDC (SEQ ID NO: 6) was significantly lower than that reported for (WGR)SGGSDC (SEQ ID NO: 5) and (WRG)SGGSDC (SEQ ID NO: 7). This suggests that the sample was not completely washed, as the lifetime of the following BODIPY-FL-conjugated IS1 varied across all peptide sequences. These data demonstrate a comprehensive workflow for sequencing peptides of up to 4 amino acids within a sequence using 4 imager strands, resulting in different lifetimes that can be further elucidated by machine learning algorithms.

[0514] Example 12

[0515] This example highlights the variability in the measured lifetimes of various longer-lived dyes in the current workflow. Candidates used in most studies were selected.

[0516] exist Figure 26In this study, a preliminary screen was conducted using commercially available dyes with longer lifetimes: KU530-6, KU530-R-4, KU560-6, and KU560-R-4. These dyes were conjugated to IS1 and hybridized to DS covalently linked to surface-immobilized peptides containing various NTAAs, tryptophan, arginine, and glycine. Longer lifetimes were reported compared to those of AF488 and BODIPY-FL. The three selected peptides significantly altered the lifetime of KU530-6. The lifetime of arginine was significantly higher than that of WGG or KU560-6. [*, p < 0.05; ***, p < 0.005; ****, p < 0.001; one-way ANOVA with Tukey's post hoc test; N = 3 samples, multiple fields of view]. KU530-6 was selected for this example due to its demonstrated sensitivity to amino acids. KU530-R-4 was chosen because it has the same fluorophore as KU530-6 but has a rigid linker that may interact differently with amino acids and appears to report an inverse relationship in lifetimes compared to KU530-6.

[0517] (XV) Selected References

[0518] U.S. Patent No. 8,373,115 Method and apparatus for identifying proteins in a mixture

[0519] U.S. Patent No. 10,545,153 for single-molecule peptide sequencing

[0520] U.S. Patent No. 10,852,230 for molecules and methods for iterative polypeptide analysis and processing

[0521] U.S. Patent No. 11,001,875 for methods for nucleic acid sequencing

[0522] U.S. Patent No. 11,162,952 for single-molecule peptide sequencing

[0523] U.S. Patent No. 11,268,963 for protein sequencing methods and reagents

[0524] U.S. Patent Application Publication No. 2014 / 0273004 for Molecules and Methods for Iterative Polypeptide Analysis and Processing

[0525] U.S. Patent Application Publication No. 2019 / 0145982 for Nucleic Acid-Encoded Macromolecular Analysis

[0526] U.S. Application Publication No. 2021 / 0396762 for Peptide Analysis Methods Employing Multicomponent Detection Reagents and Related Kits

[0527] International Patent Publication No. WO 2010 / 065531 for Single-Molecule Protein Screening

[0528] International Patent Publication No. WO 2016069124 for Improved Single-Molecule Peptide Sequencing

[0529] An Zhu et al., ACS Omega 4:12357-12365, 2019

[0530] Swaminathan et al., Nature Biotechnology DOI: 10.1038 / Nbt.4278.2018

[0531] Timp & Timp, Sci Adv. 6:eaax8978 (16 pages), 2020

[0532] (XVI) Concluding paragraph

[0533] As will be understood by one of ordinary skill in the art, each embodiment disclosed herein may include, consist essentially of, or consist of the elements, steps, ingredients, or components specifically recited therein. Thus, the term "include" or "including" should be interpreted as stating: "include, consist of, or consist essentially of." The transitional term "comprise" or "comprises" means having, but not limited to, and permits the inclusion of unspecified elements, steps, ingredients, or components, even if they constitute a major amount. The transitional phrase "consisting of" does not include any unspecified elements, steps, ingredients, or components. The transitional phrase "consisting essentially of" limits the scope of the embodiment to the specified elements, steps, ingredients, or components, and to elements, steps, ingredients, or components that do not materially affect the embodiment.

[0534] Unless otherwise indicated, all numbers expressing quantities of ingredients, properties (such as molecular weight), reaction conditions, and the like used in the specification and claims should be understood as being modified in all instances by the term "about." Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and appended claims are approximations that may vary depending on the desired properties sought to be obtained by the present invention. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. Where further clarification is needed, the term “about” has the meaning reasonably given to it by one skilled in the art when used in conjunction with a stated value or range, i.e., to mean slightly more or slightly less than the stated value or range, within ±20% of the stated value; within ±19% of the stated value; within ±18% of the stated value; within ±17% of the stated value; within ±16% of the stated value; within ±15% of the stated value; within ±14% of the stated value; within ±13% of the stated value; within ±12% of the stated value; within ±11% of the stated value; within ±10% of the stated value; within ±9% of the stated value; within ±8% of the stated value; within ±7% of the stated value; within ±6% of the stated value; within ±5% of the stated value; within ±4% of the stated value; within ±3% of the stated value; within ±2% of the stated value; or within ±1% of the stated value.

[0535] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the invention are approximations, the numerical values ​​set forth in the specific examples are reported as precisely as possible. Any numerical value, however, inherently contains certain errors necessarily resulting from the standard deviation found in their respective testing measurements.

[0536] Unless otherwise indicated herein or obviously contradictory to the context, the terms "a / an" and "the" and similar references used in the context of describing the present invention (particularly in the context of the following claims) should be interpreted as covering both the singular and the plural. The description of the range of values ​​herein is only intended to serve as a simplified method of referring to each individual value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into this specification as if each individual value were individually quoted herein. Unless otherwise indicated herein or obviously contradictory to the context, all methods described herein can be performed in any suitable order. Unless otherwise required, the use of any and all examples or exemplary languages ​​(e.g., "such as") provided herein is only intended to better illustrate the present invention and is not intended to limit the scope of the present invention. Any language in the specification should not be interpreted as indicating that any unclaimed element is necessary for practicing the present invention.

[0537] Throughout this disclosure, various aspects or variables are presented in range format. It should be understood that descriptions in range format are for convenience and brevity only and should not be construed as hard limits on the range. Therefore, the description of a range will be considered to have specifically disclosed all possible subranges within the range, as well as individual numerical values. For example, a description of a range such as 1 to 6 is intended to be considered to have explicitly disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 1 to 7, 2 to 4, 2 to 6, 3 to 6, etc.; and individual numbers within the range, such as specifically, 1, 2, 3, 4, 5, and 6. This applies regardless of the width of the range and is limited to integer amounts required by the context.

[0538] The grouping of alternative elements or embodiments of the present invention disclosed herein should not be construed as limiting. Each group member can be cited and protected individually or in any combination with other members in the group or other elements found herein. For convenience and / or patentability reasons, it is estimated that one or more members in the group may be included in the group or deleted from the group. When any such inclusion or deletion is performed, the specification is considered to contain the modified group, thereby meeting the written description of all Markush groups used in the appended claims.

[0539] Certain embodiments of the present invention are described herein, including the best modes for carrying out the present invention known to the inventors. Of course, by reading the foregoing description, variations of these described embodiments will become clear to those of ordinary skill in the art. The inventors expect that technicians will adopt these variations when appropriate, and the intention of the inventors is to practice the present invention in a manner different from that specifically described herein. Therefore, where permitted by applicable law, the present invention includes all modifications and equivalents to the subject matter recited in the appended claims. In addition, unless otherwise specified herein or otherwise clearly contradictory to the context, the present invention encompasses any combination of the above-mentioned elements with all possible variations thereof.

[0540] In addition, numerous references are made in this specification to patents, printed publications, journal articles, other written texts, and website content (referenced herein). Each referenced material is individually incorporated herein by reference in its entirety for teaching purposes as of the filing date of the first application in the priority chain that includes the specific reference. For example, with respect to chemical compounds, nucleic acid, and amino acid sequences cited herein that are available in public databases, the information in the database entries is incorporated herein by reference as of the filing date in the priority chain where the database identifier for the compound or sequence is first included in the text.

[0541] It should be understood that the embodiments of the present invention disclosed herein are illustrative of the principles of the present invention. Other modifications that may be employed are also within the scope of the present invention. Thus, by way of example, and not limitation, alternative configurations of the present invention may be utilized in accordance with the teachings herein. Therefore, the present invention is not limited to exactly as shown and described.

[0542] The details shown herein are merely examples and are intended only for illustrative discussion of preferred embodiments of the present invention and are presented to provide what is believed to be the most useful and easily understood description of the principles and conceptual aspects of the various embodiments of the present invention. In this regard, no attempt is made to illustrate the structural details of the present invention in more detail than is necessary for a basic understanding of the present invention, and the description in conjunction with the accompanying drawings and / or examples makes it clear to those skilled in the art how several forms of the present invention may be embodied in practice.

[0543] Unless explicitly modified in an example, or when the application of meaning makes any construction meaningless or essentially meaningless, the definitions and explanations used in this disclosure are intended to and are intended to control any future construction. If the construction of a term would make it meaningless or essentially meaningless, the definition should be taken from Webster's Dictionary, 11th edition, or a related dictionary known to those of ordinary skill in the art, such as Oxford Dictionary of Biochemistry and Molecular Biology, 2nd edition (edited by Anthony Smith, Oxford University Press, Oxford, 2006) and / or Dictionary of Chemistry, 8th edition (edited by J. Law and R. Rennie, Oxford University Press, 2020).

Claims

1. A method for identifying a terminal amino acid (TAA) of a peptide having an N-terminal amino acid (NTAA) and a C-terminal amino acid (CTAA), the method comprising: binding the NTAA of the peptide or the CTAA of the peptide to a solid surface to produce a bound TAA; Attaching a ssDNA docking strand (DS) to the unbound TAA of the peptide; hybridizing a first ssDNA imager strand (IS) to the DS, wherein the first IS comprises a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, wherein the second IS comprises a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; as well as The original TAA of the peptide is identified based on the detected fluorescence lifetime data of the first fluorophore and the second fluorophore.

2. A method for sequencing a peptide having an initial terminal amino acid (TAA), the method comprising: interrogating the initial TAA using a single-stranded DNA (ssDNA) docking strand (DS) attached to the initial TAA and an ssDNA imaging strand (IS) conjugated to a signal molecule to generate measurements of spectral characteristic data having a characteristic fingerprint of the initial TAA of the peptide; wherein the IS and the DS are at least partially complementary in sequence.

3. The method according to claim 2, further comprising: The initial TAA is sequentially interrogated using a library of at least two different combinations of single-stranded DNA (ssDNA) docking strands (DS) and ssDNA imaging strands (IS), wherein a signal molecule is conjugated to the IS, to generate a set of spectral characteristic data having a characteristic fingerprint of the initial TAA of the peptide.

4. The method of claim 2, wherein the initial TAA is: the N-terminal amino acid (NTAA) of the peptide; or The C-terminal amino acid (CTAA) of the peptide. The method according to claim 2 , wherein the method is performed on multiple peptides in parallel. The method of claim 2 , wherein the signaling molecule comprises a fluorophore and the spectral characteristic comprises a measure of fluorescence. The method of claim 6 , wherein the spectral characteristic comprises fluorescence lifetime.

8. A method for sequencing a peptide having an initial N-terminal amino acid (NTAA) and a C-terminal amino acid (CTAA), the method comprising: sequentially interrogating the initial NTAA using a library of at least two different combinations of a ssDNA docking strand (DS) and a ssDNA imaging strand (IS), wherein a fluorophore is conjugated to the IS, to generate a set of fluorescence lifetime data having characteristic measurements for each combination of DS, IS, and fluorophore, Each pair of IS and DS is at least partially complementary in sequence.

9. The method according to any one of claims 1 to 8, wherein the DS is a general DS.

10. The method of any one of claims 1 to 8, wherein the interrogation or the sequential interrogation comprises detecting and / or measuring the interaction between a fluorophore and an amino acid side chain at or near the CTAA or the NTAA by detecting fluorescence lifetime data for each of a plurality of IS / DS pairs in the library.

11. The method of claim 10, wherein detecting or measuring the interaction comprises obtaining fluorescence lifetime imaging (FLIM) single molecule fluorescence measurements for each of the plurality of IS / DS pairs in the library.

12. The method of claim 1 or claim 4, further comprising removing the original CTAA or NTAA of the peptide by Edman degradation reaction, enzymatic digestion, or a similar process.

13. The method of any one of claims 1 to 8, wherein the method is repeated for at least two subsequent amino acids in the peptide to generate a matrix of fluorescence lifetime data.

14. The method of claim 13, wherein the method is performed for each subsequent amino acid in the peptide to produce a matrix of fluorescence lifetime data.

15. The method of claim 13, wherein the fluorescence lifetime data is input into a machine learning algorithm to reconstruct the polypeptide sequence.

16. The method of claim 14, wherein the fluorescence lifetime data is input into a machine learning algorithm to reconstruct the polypeptide sequence.

17. The method of claim 3 or claim 8, wherein the library of IS comprises a plurality of ssDNA oligonucleotides, wherein the plurality of ssDNA oligonucleotides are varied such that the spatial positioning and / or degree of freedom of the attached fluorophore are varied to modulate the interaction with the CTAA side chain or the NTAA side chain and thereby modulate the measured fluorescence lifetime.

18. The method of claim 17, wherein the library of IS comprises a plurality of ssDNA oligonucleotides varied by one or more of: including modified nucleotides, including unnatural nucleotides, Include a 5' IS overhang relative to the cognate DS, or Include a 5' IS without overhang relative to the cognate DS.

19. The method of claim 10, wherein the interaction between the CTAA or the NTAA is further influenced by one or more of DS positioning, degrees of freedom, or another variable described herein.

20. The method of any one of claims 1, 6, or 8, wherein the fluorophore comprises Alexa Fluor 488 (AF488), BODIPY-FL, BODIPY-TR, TAMRA, or KU dyes. The method of claim 20 , wherein the fluorophore is conjugated at the end of the IS.

22. The method according to any one of claims 1 to 8, wherein: The peptide is conjugated to a modified nucleotide within the DS, and the fluorophore is conjugated to a modified nucleotide within the IS; or The peptide is conjugated to a modified nucleotide at or near a terminus of the IS, and the fluorophore is conjugated to a modified nucleotide within the IS; or The peptide is conjugated to a modified nucleotide within the DS, and the fluorophore is conjugated to a modified nucleotide at or near the end of the IS; or The peptide is conjugated to a modified nucleotide at or near the end of the DS, and the fluorophore is conjugated to a modified nucleotide at or near the end of the IS.

23. The method of claim 12, wherein the removal of CTAA or NTAA is performed under conditions such that the remaining peptides have a new terminal amino acid that can be used in another analysis cycle.

24. A method according to any one of claims 1 to 8, wherein the or each peptide is immobilised on a solid support.

25. A database comprising a matrix of the fluorescence lifetime data of claim 14.

26. A method for identifying the NTAA of a peptide having an N-terminal amino acid (NTAA) and a C-terminal amino acid (CTAA), the method comprising: binding the CTAA of the peptide to a solid surface; ligating a ssDNA docking strand (DS) to the NTAA of the peptide; hybridizing a first ssDNA imager strand (IS) to the DS, wherein the first IS comprises a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, wherein the second IS includes a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; as well as The original NTAA of the peptide is identified based on the detected fluorescence lifetime data of the first fluorophore and the second fluorophore.

27. A method for identifying a CTAA of a peptide having a C-terminal amino acid (CTAA) and an N-terminal amino acid (NTAA), the method comprising: binding the NTAA of the peptide to a solid surface; ligating a ssDNA docking strand (DS) to the CTAA of the peptide; hybridizing a first ssDNA imager strand (IS) to the DS, wherein the first IS comprises a first fluorophore; detecting fluorescence lifetime data of the first fluorophore; dissociating the first IS from the DS; hybridizing a second ssDNA IS to the DS, wherein the second IS includes a second fluorophore; detecting fluorescence lifetime data of the second fluorophore; as well as The original NTAA of the peptide is identified based on the detected fluorescence lifetime data of the first fluorophore and the second fluorophore.

28. The method of claim 26 or claim 27, further comprising: The initial TAA is cleaved from the peptide to leave the next TAA of the peptide.

29. The method of claim 28, wherein cleaving the initial NTAA comprises an Edman degradation reaction, an Edman degradation enzyme reaction, or a similar process.

30. The method of claim 27 or claim 27, comprising repeating the method a plurality of times to identify the sequence of the peptide.

31. A method of sequencing peptides, each of the peptides having a C-terminal amino acid (CTAA) and an N-terminal amino acid (NTAA), the method comprising: The peptide to be sequenced is linked to a solid substrate via its C-terminus to form an immobilized peptide; functionalizing the initial N-terminal amino acid of the immobilized peptide with a universal docking strand (DS) ssDNA oligonucleotide; contacting the DS with an imager strand (IS) oligonucleotide complementary to the DS oligonucleotide, the IS conjugated to a first fluorophore; obtaining a single molecule fluorescence lifetime (FLIM) measurement of the first fluorophore for each peptide; Optionally, repeating the single-molecule FLIM measurement for one or more additional combinations of IS and fluorophores in the library; cleaving the initial N-terminal amino acid from the peptide to reveal a second N-terminal amino acid; and Optionally, another analysis cycle is performed on the second N-terminal amino acid.

32. A method of sequencing peptides, each of the peptides having a C-terminal amino acid (CTAA) and an N-terminal amino acid (NTAA), the method comprising: The peptide to be sequenced is linked to a solid phase substrate via its N-terminus to form an immobilized peptide; functionalizing the initial C-terminal amino acid of the immobilized peptide with a universal docking strand (DS) ssDNA oligonucleotide; contacting the DS with an imager strand (IS) oligonucleotide complementary to the DS oligonucleotide, the IS conjugated to a first fluorophore; obtaining a single molecule fluorescence lifetime (FLIM) measurement of the first fluorophore for each peptide; Optionally, repeating the single-molecule FLIM measurement for one or more additional combinations of IS and fluorophores in the library; cleaving the initial C-terminal amino acid from the peptide to reveal a second C-terminal amino acid; and Optionally, another analysis cycle is performed on the second C-terminal amino acid.

33. A method of sequencing a peptide substantially as herein described.

34. The method of claim 33, wherein the method comprises detecting at least one spectral characteristic of a signal molecule, wherein the spectral characteristic is not fluorescence lifetime.

35. A kit for performing the method according to any one of claims 1 to 34, comprising at least one pair of IS and DS.

36. The kit of claim 35, comprising at least two pairs of IS and DS, wherein the two pairs differ in the fluorophore, sequence, or both contained in the IS.

37. A compound of formula (II) or a salt or solvate thereof, wherein: x is 0, 1, or 2; Each R is independently selected from the group consisting of: C1-C6 alkyl, -NO2, halogen, -C=OR, -C=SR, -C=ONR, -C=OOR, -SO3 or any other common electron withdrawing group; R 1 and R 2 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, -O-(C1-C6 alkyl), C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR or any other common electron-donating groups; y is 0, 1, 2, or 3; and Each R 3 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group.

38. The compound according to claim 37, or a salt or solvate thereof, wherein x is 0.

39. The compound according to claim 37, or a salt or solvate thereof, wherein y is 0.

40. The compound according to claim 38, or a salt or solvate thereof, wherein y is 0.

41. The compound of any one of claims 37 to 40, wherein: R 1 or R 2 The C1-C6 alkyl group is a methyl group, and R 1 or R 2 The -O-(C1-C6 alkyl) is a methoxy group.

42. The compound of claim 37 having the structure: wherein R1 and R2 are each selected from the group consisting of: H, CH3, OH, and OCH3; provided that R1 and R2 are the same; or a salt or solvate thereof.

43. The compound according to claim 42, which is (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)phenyl)carbamylthioic acid pivalic acid thioanhydride; or a salt or solvate thereof.

44. A method for preparing a compound of formula (I) or a salt or solvate thereof, the method comprising: Converted into a compound of formula (II) or a salt or solvate thereof and thereafter converting the compound of formula (II) or its salt or solvate into the compound of formula (I) or its salt or solvate, wherein: x is 0, 1, or 2; Each R is independently selected from the group consisting of: C1-C6 alkyl, -NO2, halogen, -C=OR, -C=SR, -C=ONR, -C=OOR, -SO3 or any other common electron withdrawing group; R 1 and R 2 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, —O—(C1-C6 alkyl), halogen, —O-alkyl, —S-alkyl, —OC(═O)R, —N—(C═O)—R, —OC(═O)OR, —NC(═S)NR, —N—(C═O)—OR, or any other common electron-donating group; y is 0, 1, 2, or 3; and Each R 3 independently selected from the group consisting of hydrogen, C1-C6 alkyl, hydroxy, halogen, -O-alkyl, -S-alkyl, -OC(=O)R, -N-(C=O)-R, -OC(=O)OR, -NC(=S)NR, -N-(C=O)-OR, or any other common electron-donating group.

45. The method according to claim 44, wherein the compound of formula (III) or a salt or solvate thereof is first converted into a compound of formula (IV) or a solvate thereof, The compound of formula (IV) or a salt or solvate thereof is then converted to the compound of formula (II) or a solvate thereof.

46. ​​The process of claim 45, wherein the conversion of the compound of formula (III) or a salt or solvate thereof to the compound of formula (IV) or a salt or solvate thereof occurs by reacting carbon disulfide (CS2) with the compound of formula (III).

47. The method of claim 46, wherein the reaction occurs in the presence of a base.

48. The method of claim 47, wherein the base is (C1-C6 alkyl)3N.

49. The method of claim 48, wherein the (C1-C6 alkyl)3N is triethylamine.

50. The process of claim 45, wherein the conversion of the compound of formula (IV) or its salt or solvate to the compound of formula (II) or its salt or solvate occurs by reacting the compound of formula (IV) or its salt or solvate with di-tert-butyl carbonate (O-(C(=O)-OC(CH3)2)2).

51. The method of claim 50, wherein the reacting occurs in the presence of one or more bases.

52. The method of claim 51, wherein the one or more bases comprise dimethylaminopyridine (DMAP) and triethylamine.

53. The method of claim 45, wherein the compound has the structure: wherein R1 and R2 are each selected from the group consisting of: H, CH3, OH, and OCH3; provided that R1 and R2 are the same; or a salt or solvate thereof.

54. The method of claim 53, wherein the compound is (4-(2,5-dioxo-2,5-dihydro-1H-pyrrol-1-yl)phenyl)carbamylthioic acid pivalic acid thioanhydride; or a salt or solvate thereof.

55. Use of a compound according to any one of claims 37 to 54 in a peptide assay as described herein.

Citation Information

Patent Citations

  • Single molecule peptide sequencing

    US10545153B2

  • Sensor characterization through forward voltage measurements

    US10852230B1

  • Molecules and methods for iterative polypeptide analysis and processing

    US10852305B2

  • Methods for nucleic acid sequencing

    US11001875B2

  • Identifying peptides at the single molecule level

    US11105812B2