Bi-functionalized peptide-oligonucleotide conjugate and process for the preparation thereof
A process for preparing bi-functionalized peptide-oligonucleotide conjugates via click reactions at peptide termini allows nanopore sequencing of peptides and PTM detection, addressing the limitations of existing methods by enabling analysis of complex peptide mixtures from biological proteins.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- WAGENINGEN UNIVERSITEIT
- Filing Date
- 2025-11-03
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for nanopore sequencing of peptides and post-translational modifications rely on pre-designed synthetic constructs, lacking a general approach for preparing bi-functionalized peptide-oligonucleotide conjugates from biological proteins suitable for this technique.
A process for preparing bi-functionalized peptide-oligonucleotide conjugates involves modifying the C- and N-termini of peptides with click tags and reacting them with template DNA or oligonucleotides using click reactions, enabling nanopore sequencing of peptides and detection of PTMs.
Enables the determination of amino acid sequences and associated post-translational modifications of peptides through nanopore sequencing, facilitating the analysis of complex peptide mixtures derived from proteins.
Smart Images

Figure EP2025081720_07052026_PF_FP_ABST
Abstract
Description
[0001] Bi-functionalized peptide-oligonucleotide conjugate and process for the preparation thereof
[0002] Field of the invention
[0003] The present invention is generally in the field of peptide-oligonucleotide conjugates (POCs). More specifically, the invention provides a process for the preparation of a bi-functionalized conjugate, wherein the conjugate is a peptide-oligonucleotide or a peptide-oligopeptide conjugate, from a peptide, from a protein, or from a mixture of peptides obtained from a digested mixture of proteins. The amino acid sequence of the peptide-moiety of especially these bi-functionalized POCs may be determined by e.g. nanopore sequencing.
[0004] Background of the invention
[0005] Nanopore protein sequencing is a rapidly developing technique that provides for single-molecule measurements of peptides with high accuracy on individual amino acid substitutions and associated chemical modifications (Restrepo-Perez et al., Nat. Nanotechnol. 2018, 13 (9), 786-796; Schmid et al., Essays in Biochemistry 2021, 65 (1), 93-107; both incorporated by reference). This technique detects the change in ionic current as a solitary molecule moves through a membrane's nanopore. Nanopores have been employed in the past to sequence individual molecules of DNA. For label-based sequencing, a DNA-translocating motor enzyme called Hel308 helicase is used to translocate and read DNA via the biological nanopore (MspA) step- by-step (Butler et al., Proc. Natl. Acad. Sci. 2008, 105 (52), 20647-20652; Derrington et al., Nat. Biotechnol. 2015, 33 (10), 1073-1075; both incorporated by reference).
[0006] This method has recently been applied for peptide sequencing and post-translational modification (PTM) identification and localization.
[0007] A proof of concept of nanopore sequencing of peptides with either a conjugated N-terminus or a conjugated C-terminus is described by Yan et al. (Nano Lett., 2021, 21 (15), 6703-6710, incorporated by reference). By constructing peptide-oligonucleotide conjugates and subjecting them to measurements with nanopore-induced phase-shift sequencing (NIPSS), direct observation of the ratcheting motion of peptides was achieved and a sequence-dependent change in ion current as the peptide translocated through the pore was observed. The results suggested a proof of concept of nanopore sequencing of a peptide and may be useful for peptide fingerprinting.
[0008] In another proof of concept a peptide-oligonucleotide conjugate was pulled through an MspA nanopore by Hel308 helicase, and discrimination of individual amino acid substitutions on single peptides was obtained (Brinkerhoff et al., Science 2021, 374 (6574), 1509-1513, incorporated by reference). Recently, evidence emerged that this technique enables discrimination of PTMs that can otherwise only be discriminated with great difficulty (Chen et al., ACS Nano 2024, 18 (42), 28999-29007, incorporated by reference). Importantly, all published examples relied on pre-designed synthetic constructs containing the required features for nanopore analysis.
[0009] There is a need in the art for a general approach that enables the preparation of bi-functionalized conjugates, particularly peptide oligonucleotide conjugates (POCs), especially derived from proteins of biological origin, to render them suitable for nanopore sequencing.
[0010] Summary of the invention
[0011] In a first aspect, the invention relates to a process for the preparation of a bi-functionalized conjugate, wherein the conjugate is a peptide-oligopeptide or a peptide-oligonucleotide conjugate (POC), the process comprising the steps of:
[0012] (i) providing a peptide according to structure (1): wherein:
[0013] G is a peptide comprising 1 - 50 amino acids; and
[0014] AA1and AA2are independently selected from the group of amino acid side chains;
[0015] AA3is a tyrosine, tryptophan, arginine or lysine side chain;
[0016] (ii) modifying the C-terminus of the peptide by:
[0017] (ii-al) attaching a click tag F1via AA3; or
[0018] (ii-a2) converting AA3into a 1,2-quinone click tag F1via oxidation, with the proviso that AA3is a tyrosine side chain; and
[0019] (ii-b) reacting click tag F1with D1-(L1)n-Q1, wherein Q1is a click tag that is capable of reacting with F1in a click reaction, L1is a linker and n is 0 or 1, and D1is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer;
[0020] (iii) modifying the N-terminus of the peptide by: (iii-al) formation of an imidazolidinone via imine condensation of the terminal N-atom with a 2-pyridinecarboxaldehyde and cyclization to form the imidazolidinone, wherein the 2-pyridinecarboxaldehyde comprises a click tag F2, and with the proviso that AA2is not a proline side chain; or
[0021] (iii-a2) attaching a click tag F2via AA1; and
[0022] (iii-b) reacting click tag F2with D2-(L2)n-Q2, wherein Q2is a click tag that is capable of reacting with F2in a click reaction, L2is a linker and n is 0 or 1, and D2is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threating oligopeptide and a positively charged polymer; wherein a template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides; a template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids; a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged; a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively or positively charged; and with the proviso that one of D1and D2is a template DNA oligonucleotide or a template oligopeptide.
[0023] In this process it is preferred that in step (i) a mixture of two or more peptides according to structure (1) is provided.
[0024] The present invention also relates to the bi-functionalized conjugates, wherein the conjugates are peptide-oligonucleotide or peptide-oligopeptide conjugates, obtainable by the process according to the invention, and to the mixture of bi-functionalized conjugates, wherein the conjugates are peptide-oligonucleotide or peptide-oligopeptide conjugates, obtainable by the process when in step (i) a mixture of two or more peptides according to structure (1) is provided.
[0025] The present invention also relates to a bi-functionalized conjugate, wherein the conjugate is a peptide-oligonucleotide or a peptide-oligopeptide conjugate, according to structure (34):
[0026] wherein:
[0027] AA1, AA2, G, D1, D2, L1, L2and n are as defined above; m is 0 or 1;
[0028] L3is a linker;
[0029] R1is independently selected from the group consisting of hydrogen and linear or branched C1 - C10 alkyl groups; and
[0030] Z1and Z2are connecting groups.
[0031] In another aspect the invention relates to a bi-functionalized conjugate, wherein the conjugate is a peptide-oligonucleotide or a peptide-oligopeptide conjugate, according to structure (36): wherein:
[0032] AA1, AA2, G, D1, F2, L1, R1, L3and m are as defined above.
[0033] The bi-functionalized conjugates, particularly the POCs, prepared by the process of the present invention may be used for the determination of the amino acid sequence and, if present, associated post-translational modifications (PTMs) of the (POC) peptide moiety, for example by nanopore sequencing. In order to be able to determine the amino acid sequence and associated post-translational modification via nanopore sequencing, one of the functionalizations of the bi-functionalized conjugate, particularly a bi-functionalized POC, should be a DNA oligonucleotide or oligopeptide that can be read by the nanopore, in other words said DNA oligonucleotide or oligopeptide should be recognized by a nanopore motor protein. Such a DNA oligonucleotide or oligopeptide is herein referred to as a template DNA oligonucleotide or template oligopeptide. The nanopore reads the template DNA oligonucleotide or template oligopeptide first, followed by sensing of the linker (if present) and reading of the peptide sequence and, if present, associated post-translational modifications (Laszlo et al., Nat. Biotechnol., 2014, 32 (8), 829-833, incorporated by reference). At the other terminus of the bi-functionalized conjugate, particularly the bi-functionalized POC, a charge is needed in order to cause the conjugate to insert into the nanopore. Typically, a negative charge is employed, for example provided by a negatively charged polymer, a chain of glutamate or aspartate residues, or a negatively charged second oligonucleotide which is also referred to as a threading DNA oligonucleotide (Chen et al., Chem. Sci., 2021, 12 (47), 15750-15756, incorporated by reference). Alternatively, in some embodiments a positive charge may facilitate insertion, for example through the use of a positively charged polymer, a chain of arginine or lysine residues, or a negatively charged oligopeptide.
[0034] The present invention provides a general approach for the preparation of bi-functionalized conjugates, wherein the conjugates are peptide-oligonucleotide or peptide-oligopeptide conjugates, typically from one or more peptides, one or more proteins, or a mixture of peptides obtained from a digested mixture of proteins.
[0035] Brief description of the Figures
[0036] Figure 1 shows a particularly preferred embodiment of the process according to the invention. A C-terminal tyrosine side chain is in situ oxidized to an ortho-quinone using l,3-dihydro-l-hydroxy-3-oxo-l,2-benziodoxole-4-carboxylic acid 1-oxide (mIBX), followed by SPOCQ with BCN-DNA. The obtained POC is then / V-terminally modified with an azide functionalized pyridine carboxaldehyde that subsequently is clicked to another BCN-DNA via SPAAC. AA1and AA2indicate amino acid side chains (AA2cannot be a proline residue), R represents any spacer between the pyridine and azide group.
[0037] In Figure 2 several combinations of compatible click tags are shown, as well as several connecting groups Z that result from click reaction between different click tags Q and F. The nature of Z depends on the type of click reaction, F, and Q. When the skilled person selects a click reaction between a certain click tag F and a compatible click tag Q, the skilled person knows what the corresponding connecting group Z looks like.
[0038] Figure 3 (Example 8) shows the chemical reaction of a peptide containing a C-terminal tyrosine with l,3-dihydro-l-hydroxy-3-oxo-l,2-benziodoxole-4-carboxylic acid 1-oxide (mIBX) and BCN-Cg-threading DNA (threading DNA) via SPOCQ. mIBX oxidizes the tyrosine residue to an ortho-quinone that is subsequently clicked with Threading DNA resulting in a peptide-oligonucleotide conjugate (POC).
[0039] Figure 4 (Example 8) shows the chemical structures and deconvolution of HESI mass spectra of SPOCQ products 88 (peptide 85 + threading DNA; Figure 4A), 89 (peptide 86 + threading DNA; Figure 4B), and 90 (peptide 87 + threading DNA; Figure 4C).
[0040] Figure 5 (Example 9) shows the chemical reaction of a POC with an azide functionalized 2-PCA resulting in the / V-terminal imidazolidinone formation of the POC. Figure 6 (Example 9) shows the deconvoluted mass spectra and chemical structures of POCs modified with 6-AM-2-PCA (79) resulting in products 91 (from POC 88), 92 (from POC 89), and 93 (from POC 90).
[0041] Figure 7 (Example 9) shows the chemical structures of POCs modified with 6-AAPM-2-PCA (83) resulting in products 94 (from POC 88), 95 (from POC 89), and 96 (from POC 90) and the corresponding deconvoluted mass spectra.
[0042] Figure 8 (Example 9) shows the chemical structures of POCs modified with (D) 6-A-PEG4-PM-2-PCA (84) resulting in products 97 (from POC 88), 22 (from POC 98), and 99 (from POC 90) and the deconvoluted mass spectra.
[0043] Figure 9 (Example 10) shows the chemical reaction of a / V-terminal modified POC with template DNA via SPAAC.
[0044] Figure 10 shows several products (100, 101 and 102) that were obtained in the reaction of Figure 9.
[0045] Figure 11 (Example 10) the optimization of product purification after SPAAC click reaction is shown. Here, POC 92 (Figure 6) was taken and resulted in purified POC 102. Excess of template DNA was removed by reacting with biotin-PEGs-methyltetrazine, followed by treatment with streptavidin-agarose. Template DNA that reacted with traces of azide functionalized 2-PCA was removed by reacting with biotin-PEG4- oxyamine, followed by treatment with streptavidin-agarose. After agarose treatment, the product was purified by 30 kDa MWCO spin filters, resulting in a pure POC.
[0046] Figure 12 shows the deconvolution of HESI mass spectra of POCs 100 and 101, and the combined POCs 100, 101 and 102 in a one-pot SPAAC reaction.
[0047] Figure 13 (Example 11) shows the digestion of the peptide human Angiotensin 2 into two fragments of which the one with a C-terminal tyrosine was converted into a POC by the process according to the invention. Figure 13A shows the structure of the native peptide human Angiotensin 2. Figure 13B shows the fragments that were obtained after human angiotensin was treated with chymotrypsin. Figure 13C shows the conjugation products that were formed when the mixture that is shown in Figure 13B was processed into POCs according to the invention. Only the tyrosine fragment is converted into a POC, the peptide with the sequence lle-His-Pro-Phe is not converted. Through the process according to the invention, the peptide with the sequence Aps-Arg-Val-Tyr is first converted into POC 103, then converted into 2-PCA-functionalized POC 104, and subsequently into the final POC 105. In this final product, the peptide sequence is sandwiched in between two different DNA strands.
[0048] Figure 14 (Example 11) shows the digestion of the protein human Lysozyme into peptide fragments of which several ones with a C-terminal tyrosine were converted into a POC by the process according to the invention. Figure 14A shows the primary structure of human lysozyme (abbreviated as "LysoH"), and the HPLC chromatogram of the mixture that was obtained after digestion of LysoH into peptides. In this mixture, the five peptides have been identified explicitly and are indicated with the numbers 1-5. Figure 14B shows the deconvoluted MS spectrum of the peptides that were obtained after chymotrypsin digestion of LysoH after treatment during the first steps of POC synthesis by the process according to the invention. The numbers in the left panel of Figure 14B correspond to the numbers of the peptides in Figure 14A but now containing a C-terminal DNA strand (numbers 1, 2, 3, and 5) or a C-terminal DNA strand, and an N-terminal 2-PCA modification (numbers l'-5'). The numbers in the right panel of Figure 14B correspond to the five final sandwich POCs that were formed by the process according to the invention. These are compatible with nanopore analysis. Figure 14C shows the terminal amino acids of the fragments of LysoH that were identified by LC-MS, including the associated sequence, and the molecular weight of the peptide fragment ('Mass' column), C-terminal DNA-conjugate ('SPOCQ' column), 2-PCA-functionalized N-termini ('PCA (')' column), and the final POC that is obtained at the end (the 'SPAAC' column).
[0049] Detailed Description of the invention
[0050] General Definitions
[0051] In this document and in its claims, the verb "to comprise" and its conjugations are used in their non-limiting sense to mean that items following the word are included, but items not specifically mentioned are not excluded. In addition, reference to an element by the indefinite article "a" or "an" does not exclude the possibility that more than one of the element is present, unless the context clearly requires that there be one and only one of the elements. The indefinite article "a" or "an" thus usually means "at least one". The word "about" or "approximately" when used in association with a numerical value (e.g. about 10) preferably means that the value may be the given value plus or minus 5% of the value.
[0052] The present invention is described herein with reference to a number of exemplary embodiments. Modifications and alternative implementations of some parts or elements are possible and are included in the scope of protection as defined in the appended claims. All citations of literature and patent documents are hereby incorporated by reference.
[0053] The compounds disclosed in this description and in the claims may comprise one or more asymmetric centres, and different diastereomers and / or enantiomers may exist of the compounds. The description of any compound in this description and in the claims is meant to include all diastereomers, and mixtures thereof, unless stated otherwise. In addition, the description of any compound in this description and in the claims is meant to include both the individual enantiomers, as well as any mixture, racemic or otherwise, of the enantiomers, unless stated otherwise. When the structure of a compound is depicted as a specific enantiomer, it is to be understood that the invention of the present application is not limited to that specific enantiomer.
[0054] The compounds may occur in different tautomeric forms. The compounds according to the invention are meant to include all tautomeric forms, unless stated otherwise. When the structure of a compound is depicted as a specific tautomer, it is to be understood that the invention of the present application is not limited to that specific tautomer. The compounds disclosed in this description and in the claims may further exist as R and S stereoisomers. Unless stated otherwise, the description of any compound in the description and in the claims is meant to include both the individual R and the individual S stereoisomers of a compound, as well as mixtures thereof. When the structure of a compound is depicted as a specific S or R stereoisomer, it is to be understood that the invention of the present application is not limited to that specific S or R stereoisomer.
[0055] The compounds disclosed in this description and in the claims may further exist as exo and endo diastereoisomers. Unless stated otherwise, the description of any compound in the description and in the claims is meant to include both the individual exo and the individual endo diastereoisomers of a compound, as well as mixtures thereof. When the structure of a compound is depicted as a specific endo or exo diastereomer, it is to be understood that the invention of the present application is not limited to that specific endo or exo diastereomer.
[0056] The compounds according to the invention may exist in salt form, which are also covered by the present invention. The term "salt thereof" means a compound formed when an acidic proton, typically a proton of an acid, is replaced by a cation, such as a metal cation or an organic cation and the like. For example, in a salt of a compound the compound may be protonated by an inorganic or organic acid to form a cation, with the conjugate base of the inorganic or organic acid as the anionic component of the salt.
[0057] The term "protein" is herein used in its normal scientific meaning. Herein, polypeptides comprising about 100 or more amino acids are considered proteins. A protein may comprise natural amino acids, post- translationally modified amino acids, but also unnatural amino acids.
[0058] A "linker" is herein defined as a moiety that connects two or more elements of a compound. A linker may comprise one or more linkers and spacer-moieties that connect various moieties within the linker.
[0059] A "spacer" or "spacer-moiety" is herein defined as a moiety that spaces (i.e. provides distance between) and covalently links together two (or more) parts of a linker.
[0060] The terms "tyrosinase" and "(poly)phenol oxidase" refer to an enzyme that is capable of catalysing the ortho-hydroxylation of a monophenol moiety to an ortho-dihydroxybenzene (catechol) moiety, followed by further oxidation of the ortho-dihydroxybenzene moiety to produce an ortho-quinone (1,2-quinone) moiety.
[0061] The term "click tag" refers to a functional moiety that is capable of undergoing a click reaction, i.e. two compatible click tags mutually undergo a click reaction such that they are covalently linked in the resulting product. Click reactions are known in the art and include cycloaddition reactions such as the [4+2] cycloaddition (e.g. the Diels-Alder reaction, inverse electron-demand Diels-Alder reaction), [4+1] cycloaddition (e.g. between isocyanides and tetrazines) and the [3+2] cycloaddition (e.g. 1,3-dipolar cycloaddition between an azide and alkyne). Typical examples of click reactions include the strain-promoted alkyne-azide cycloaddition (SPAAC), strain-promoted alkyne-nitrone cycloaddition (SPANC), strain-promoted alkyne-nitrile oxide cycloaddition (SPANOC), strain-promoted oxidation-controlled 1,2-quinone cycloaddition (SPOCQ) to a cycloalkene or cycloalkyne, sulfur(VI) fluoride exchange reaction (SuFEx), thiol-ene and thiolyne reactions. Compatible tags for click reactions are known in the art. An example of a compatible click tags is a (hetero)cycloalkyne and an azide. Other examples of compatible click tags include a (hetero)cycloalkene and an azide, a (hetero)cycloalkyne and a nitrone, a (hetero)cycloalkyne and a 1,2-quinone, a (hetero)cycloalkene and a 1,2-quinone, a tetrazine and a cyclooctyne, a tetrazine and a trans-cyclooctyne, etc. In the context of the present invention, click tag F1is capable of reacting with click tag Q1and thus click tags F1and Q1are compatible click tags. Similarly, click tag F2and Q2are compatible click tags. A comprehensive review of click chemistry is provided by Nguyen et al., Nature Rev. Chem. 2020, 4, 476-489 (incorporated by reference).
[0062] The term "(hetero)cycloalkyne" refers to a cycloalkyne as well as to a heterocycloalkyne. Similarly the term "(hetero)cycloalkene" refers to a cycloalkene as well as to a heterocycloalkane, etc.
[0063] The invention relates to a process for the preparation of a bi-functionalized conjugate, wherein the conjugate is a peptide-oligopeptide or a peptide-oligonucleotide conjugate (POC), the process comprising the steps of:
[0064] (i) providing a peptide according to structure (1): wherein:
[0065] G is a peptide comprising 1-50 amino acids; and
[0066] AA1and AA2are independently selected from the group of amino acid side chains;
[0067] AA3is a tyrosine, tryptophan, arginine or lysine side chain;
[0068] (ii) modifying the C-terminus of the peptide by:
[0069] (ii-al) attaching a click tag F1via AA3; or
[0070] (ii-a2) converting AA3into a 1,2-quinone click tag F1via oxidation, with the proviso that AA3is a tyrosine side chain; and
[0071] (ii-b) reacting click tag F1with D1-(L1)n-Q1, wherein Q1is a click tag that is capable of reacting with F1in a click reaction, L1is a linker and n is 0 or 1, and D1is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer;
[0072] (iii) modifying the N-terminus of the peptide by:
[0073] (iii-al) formation of an imidazolidinone via imine condensation of the terminal N-atom with a 2-pyridinecarboxaldehyde and cyclization to form the imidazolidinone, wherein the 2-pyridinecarboxaldehyde comprises a click tag F2, and with the proviso that AA2is not a proline side chain; or
[0074] (iii-a2) attaching a click tag F2via AA1; and
[0075] (iii-b) reacting click tag F2with D2-(L2)n-Q2, wherein Q2is a click tag that is capable of reacting with F2in a click reaction, L2is a linker and n is 0 or 1, and D2is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer; wherein a template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides; a template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids; a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged; a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively or positively charged; and with the proviso that one of D1and D2is a template DNA oligonucleotide or a template oligopeptide.
[0076] The present invention also relates to the bi-functionalized conjugates, wherein the conjugates are peptide-oligonucleotide or peptide-oligopeptide conjugates, obtainable by the process according to the invention.
[0077] A peptide-oligopeptide conjugate is a molecule wherein a peptide moiety is covalently linked, via one or more bonds, to an oligopeptide moiety. A peptide-oligonucleotide conjugate (POC) is a molecule wherein a peptide moiety is covalently linked, via one or more bonds, to an oligonucleotide moiety. The POCs, and peptide-oligopeptide conjugates, according to the invention are bi-functionalized, wherein the term "bifunctionalized" means that both the N- and the C-terminus of the peptide are functionalized, i.e. modified. Since both ends of the peptide are modified, these bi-functionalized POCs may also be referred to as "sandwich-POCs". In addition to the functionalization at the N- and the C-terminus of the peptide other functionalizations may be present in the bi-functionalized conjugates (e.g., POCs), e.g. post-translational modifications (PTMs). Typically these additional functionalizations, if any, are present in the peptide moiety.
[0078] As described above, the bi-functionalized conjugates, especially the POCs, may be used for the determination of the amino acid sequence in the peptide moiety of the conjugates, preferably by nanopore sequencing. Therefore either the N-terminus or the C-terminus of the bi-functionalized conjugates, particularly the POCs, is functionalized with a template DNA oligonucleotide or a template oligopeptide, wherein a template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides, and wherein a template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids. Template DNA oligonucleotides and template oligopeptides are described in more detail below.
[0079] Step (i) Provision of a peptide or a mixture of two or more peptides
[0080] In step (i) of the process according to the invention a peptide according to structure (1), or a mixture of two or more peptides according to structure (1), is provided.
[0081] G is a peptide comprising 1-50 amino acids. In a preferred embodiment, G is a peptide comprising 1-40, more preferably 1-30, and even more preferably 2-25 amino acids.
[0082] AA1and AA2are independently selected from the group of amino acid side chains. Preferably, AA1and AA2are independently selected from the group of amino acid side chains of natural amino acids. The skilled person knows which amino acids are natural amino acids. A common definition of natural amino acids is that these are the amino acids that proteins are made up of. These amino acids are alanine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, proline, histidine, lysine, arginine, aspartic acid, glutamic acid, serine, threonine, cysteine, tyrosine, asparagine, glutamine, glycine, selenocysteine and pyrrolysine.
[0083] As is known in the art, an amino acid side chain is the side chain that is attached to the alpha-carbon atom of an amino acid. The amino acid side chain is specific for each amino acid. For example in glycine the side chain is -H, in alanine the amino acid side chain is -CH3, in valine -CH(CH3h, in cysteine -CH2SH, in serine -CH2OH, in lysine -(CI- hNI- , in tyrosine -CI- fCsF OH) etc. In proline the side chain attached to the alphacarbon atom forms a five-membered ring with the proline nitrogen atom via a -(CHzh- moiety.
[0084] AA3is a tyrosine, tryptophan, arginine or lysine side chain, preferably a tyrosine, tryptophan or arginine side chain. In a particularly preferred embodiment AA3is a tyrosine side chain. In a preferred embodiment of the process of the invention, a mixture of two or more peptides according to structure (1) is provided in step (i). Such peptide mixture, or peptide according to structure (1), may be obtained by hydrolysis of one or more peptides, one or more proteins, a mixture of peptides obtained from a digested mixture of proteins, or a mixture of proteins. Hydrolysis methods are known in the art. The hydrolysis may be a chemical hydrolysis or an enzymatic hydrolysis. Selective chemical hydrolysis methods wherein cleavage occurs at a specific site in a protein are known in the art, e.g. CNBr assisted cleavage at the C-terminus of methionine. In a preferred embodiment the hydrolysis occurs via enzymatic hydrolysis with a proteolytic enzyme. Proteolytic enzymes, also referred to as proteases, proteinases or peptidases, are known in the art (see e.g. Rao et al., Microbiol. Mol. Biol. Rev. 1998, 62, 597 - 635, incorporated by reference). The type of enzyme used for the hydrolysis of the protein or mixture of proteins depends on the nature of the peptide bond that is to be cleaved. Suitable enzymes with a preferential cleavage at specific amino acids are known in the art.
[0085] When for example cleavage of a protein at the C-terminus of tyrosine is desired to obtain a peptide of structure (1) wherein AA3is a tyrosine side chain, an endopeptidase, e.g. a serine endopeptidase (enzyme class 3.4.21), aspartate endopeptidase (enzyme class 3.4.23) or metalloendopeptidase (enzyme class 3.4.24), with known preferential cleavage at the C-terminus of tyrosine may be used. Examples of such serine endopeptidases include chymotrypsin (3.4.21.1), including a- and p-chymotrypsin from different sources, metridin (3.4.21.3), cathepsin G (3.4.21.20), chymase (3.4.21.39) and semenogelase (3.4.21.77). Examples of such aspartate endopeptidases include pepsin A (3.4.23.1), gastricsin (3.4.23.4), and examples of such metalloendopeptidases include neprilysin (3.4.24.11).
[0086] When cleavage of a protein at the C-terminus of tryptophan is desired to obtain a peptide according to structure (1) wherein AA3is a tryptophan side chain, a peptidase with known preferential cleavage at the C-terminus of tryptophan may be used. An example is chymotrypsin (3.4.21.1), including a- and - chymotrypsin from different sources.
[0087] When for example cleavage of a protein at the C-terminus of arginine is desired to obtain a peptide wherein AA3is an arginine side chain, a peptidase with known preferential cleavage at the C-terminus of arginine may be used. Examples include trypsin (3.4.21.4), thrombin (3.4.21.5), plasmin (3.4.21.7), acrosin (3.4.21.10), tryptase (3.4.21.59), peptidase K (3.4.21.64) and hepsin (3.4.21.106).
[0088] When, for example, cleavage of a protein at the C-terminus of lysine is desired to obtain a peptide wherein AA3is a lysine side chain, a peptidase with known preferential cleavage at the C-terminus of lysine may be used. Examples include trypsin (3.4.21.4), plasmin (3.4.21.7), acrosin (3.4.21.10), lysyl endopeptidase (3.4.21.50), tryptase (3.4.21.59) and peptidase K (3.4.21.64).
[0089] When in step (i) a mixture of two or more peptides of structure (1) is provided, said peptides typically differ at least in peptide G. Amino acid side chains AA1, AA2and AA3may be the same or different. As is clear to the skilled person, the nature of the peptide mixture will depend on the method of preparing the peptide mixture from a protein or mixture of proteins (e.g. the enzyme used for hydrolysis) and on the protein or mixture of proteins itself.
[0090] Also when in step (i) a mixture of two or more peptides according to structure (1) is provided, the process of the invention may be performed as a one-pot process with peptide mixtures, wherein the peptide mixtures are obtained via hydrolysis of a protein or a protein mixture.
[0091] In steps (ii) and (iii) of the process according to the invention the C-terminus and the N-terminus of the peptide, or mixture of peptides, provided in step (i) are modified. If in step (i) a mixture of two or more peptides is provided, the C-terminus and the N-terminus of the peptides in this mixture are modified in steps (ii) and (iii). The modification of the C-terminus and the N-terminus of the peptide or mixture of peptides may be performed in any order. In one embodiment of the process, the C-terminus of the peptide or mixture of peptides is modified before modification of the N-terminus of the peptide or mixture of peptides, i.e. step
[0092] (ii) of the process is performed before step (iii). In this embodiment the order of the process steps is (i), (ii),
[0093] (iii). In another embodiment of the process the N-terminus of the peptide or mixture of peptides is modified before modification of the C-terminus of the peptide or mixture of peptides, i.e. step (iii) of the process is performed before step (ii). In this embodiment the order of the process steps is (i), (iii), (ii). It is preferred that the C-terminus of the peptide or mixture of peptides is modified before the N-terminus, i.e. in a preferred embodiment step (ii) takes place before step (iii) and the order of the process steps is (i), (ii), (iii).
[0094] Step (ii): Modification of the C-terminus of peptide (1)
[0095] In step (ii) of the process according to the invention, the C-terminus of the peptide according to structure (1), or at least part, such as all, of the C-termini of peptides in the mixture of two or more peptides of structure (1), is modified. Said modification comprises the following steps:
[0096] (ii-al) attaching a click tag F1via AA3; or
[0097] (ii-a2) converting AA3into a 1,2-quinone click tag F1via oxidation, with the proviso that AA3is a tyrosine side chain; and
[0098] (ii-b) reacting click tag F1with D1-(L1)n-Q1, wherein Q1is a click tag that is capable of reacting with F1in a click reaction, L1is a linker and n is 0 or 1, and D1is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer, wherein a template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides; a template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids; a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged; and a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively charged or positively charged.
[0099] The template DNA oligonucleotide, template oligopeptide, threading DNA oligonucleotide, and threading oligopeptide are described in more detail below.
[0100] In a preferred embodiment steps (ii-a), i.e. either (ii-al) or (ii-a2), and (ii-b) are performed in a one- pot process without the need to isolate the intermediate product of step (ii-a).
[0101] As defined above, the term "click tag" refers to a functional moiety that is capable of undergoing a click reaction, i.e. two compatible click tags mutually undergo a click reaction such that they are covalently linked in the resulting product. Click tags and compatible click tags for click reactions are known in the art.
[0102] In step (ii-al) of the modification of the C-terminus of the peptide, a click tag F1is attached to the peptide via amino acid side chain AA3. As described above, AA3is a tyrosine, tryptophan, arginine or lysine side chain, preferably a tyrosine, tryptophan or arginine side chain, more preferably a tyrosine side chain. Suitable methods for the attachment of F1depend on the nature of AA3. The skilled person will be able to determine which methods are suitable for the introduction of a particular click tag F1via a particular amino acid side chain AA3. Several examples are provided below.
[0103] Step (ii-a2) may be employed when AA3is a tyrosine side chain. This tyrosine side chain may be converted into a 1,2-quinone click tag F1via oxidation. The oxidation of the phenol in the tyrosine side chain may be accomplished in situ, for example via reaction with an o-iodoxybenzoic acid (IBX), e.g. 1,3-dihydro-l- hydroxy-3-oxo-l,2-benziodoxole-4-carboxylic acid 1-oxide (mIBX). The oxidation of the phenol in the tyrosine side chain may also be accomplished, preferably in situ, via enzymatic oxidation by the action of an oxidative enzyme capable of oxidizing tyrosine. Such enzymes are known in the art and are preferably selected from tyrosinases (e.g. mushroom tyrosinase), phenol oxidases and polyphenol oxidases. A compatible click probe Q1for a 1,2-quinone is a click probe comprising a strained C=C bond or a strained CEC bond.
[0104] When AA3is a tyrosine side chain, there are many ways to employ step (ii-al) to attach a click tag F1via the tyrosine side chain, and the skilled person will know which methods are suitable for the introduction of a particular click tag F1. Methods for the attachment of a click tag to a tyrosine chain are reviewed in Szjika et al., Org. Biomol. Chem. 2020, 18, 9018 - 9028 (incorporated by reference).
[0105] Click tag F1may for example be attached to the tyrosine side chain via a Mannich-type reaction between tyrosine, an aldehyde and an aniline, wherein the aniline comprises a click tag F1(e.g. an azide or an alkyne, but many others are possible) that can react with a compatible click tag Q1(e.g. an alkyne or an azide). As shown below, a moiety comprising the click tag F1is introduced at the ortho position of the tyrosine phenol -OH group. The reaction shown below is e.g. performed at room temperature and neutral pH, in a suitable buffer.
[0106] Further examples of step (ii-al) when AA3is tyrosine include the attachment of F1via reaction with a diazonium salt as shown below. Said reaction is e.g. performed at a temperature in the range of 0-40 °C, at a pH in the range of 8-10 and in a suitable buffer.
[0107] Click tag F1may also be introduced via reaction with an azodicarbonyl carbonyl compound such as a 1, 2, 4-triazoline-3, 5-dione, as shown below. This reaction is e.g. performed at a temperature in the range of 0-40 °C, at a pH in the range of 6-9 in a suitable buffer.
[0108] Another example of the attachment of a click tag F1via step (ii-al) when AA3is tyrosine is reaction with / V-methyl luminol (NML) using laccase, horseradish peroxidase (HRP), hemin, DNAzyme, or electrochemistry, as shown below. Reaction conditions are e.g. 0-40 °C, a pH in the range of 6-9 and a suitable buffer as solvent.
[0109] When AA3is a tryptophan side chain, click tag F1may for example be introduced via reaction with ABNO (oxyl radical of 9-azabicyclo[3.3.1]nonan-3-one 9-oxoammonium cation) in the presence of sodium nitrite (NaNOj). Reaction conditions are e.g. 0-40 °C, a pH range of 6-9 and a suitable buffer as solvent. The carbonyl (C=O) of the attached ABNO can subsequently be functionalized with F1by means of established methods known to the skilled person.
[0110] When AA3is an arginine side chain, a click tag F1may for example be introduced via a glyoxal-based reagent comprising click tag F1, for example an aryl glyoxal reagent wherein the aryl group is further substituted with a click tag F1. Selective functionalization occurs via reaction between the glyoxal group and the guanidine group that is present in arginine side chain AA3. An example of such glyoxal-based reagent wherein F1is an azide is 4-azidophenyl glyoxal (APG). Different click tags F1may be attached to the arginine side chain AA3in a similar way.
[0111] In a preferred embodiment of step (ii-a), AA3is tyrosine. In this embodiment it is further preferred that click tag F1is introduced via step ( i i-a 2), i.e. via converting AA3into a 1,2-quinone click tag F1by oxidation, as described in more detail above.
[0112] In the process according to the invention, preferred options for click tag F1are selected from the group consisting of (hetero)cycloalkynes, (hetero)cycloalkenes, azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones.
[0113] In step (ii-b) of the modification of the C-terminus of the peptide, click tag F1is reacted with D1- ( Jn-Q1, wherein Q1is a click tag that is capable of reacting with F1in a click reaction, L1is a linker and n is 0 or 1, and D1is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a negatively charged moiety, and a positively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, wherein the positively charged moiety is selected from a threating oligopeptide and a positively charged polymer, wherein a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged and a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively or positively charged, and wherein the negatively or positively charged polymer may be of synthetic or biological origin and typically contains sufficient charge for nanopore insertion (i.e., suitable for nanopore insertion).
[0114] As described above, step (ii-a), i.e. the introduction of click tag F1via step (ii-al) or (ii-a2), and step (ii-b), i.e. the reaction of F1with D1-(L1)n-Q1, may be performed in a one-pot process, without the necessity to isolate the intermediate product of step (ii-a) comprising click tag F1.
[0115] In D1-(L1)n-Q1, Q1is a click tag that is capable of reacting with F1in a click reaction, i.e. Q1is a click tag that is compatible with F1. The skilled person knows which click tags are compatible, in other words, the skilled person knows which click tag Q1may be used to react with a particular click tag F1. For a particular click tag F1there may be more compatible click tags Q1, and vice versa. For example, when F1is a strained alkyne, e.g. a (hetero)cycloalkyne, Q1may be, amongst others, an azide, a nitrone, a 1,2-quinone, a nitrile oxide or a tetrazine.
[0116] In the process according to the invention, preferred options for click tag Q1are selected from the group consisting of (hetero)cycloalkynes, (hetero)cycloalkenes, azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones, wherein click tag Q1is compatible with click tag F1.
[0117] The click reaction of F1with Q1and F2with Q2, and preferred embodiments thereof, are described in more detail below.
[0118] L1is a linker. The presence of L1is optional (n is 0 or 1). A linker is herein defined as a moiety that connects two or more elements of a compound. In D^L^n-Q1, D1and Q1are covalently connected to each other via linker L1, if present. Linkers are known in the art. L1may for example be selected from the group consisting of linear or branched Ci -C200 alkylene groups, C2 - C200 alkenylene groups, C2 - C200 alkynylene groups, C3 - C200 cycloalkylene groups, Cg - C200 cycloalkenylene groups, Cg - C200 cycloalkynylene groups, C7- C200 alkylarylene groups, C7- C200 arylalkylene groups, Cg - C200 arylalkenylene groups and Cg - C200 arylalkynylene groups. Optionally the alkylene groups, alkenylene groups, alkynylene groups, cycloalkylene groups, cycloalkenylene groups, cycloalkynylene groups, alkylarylene groups, arylalkylene groups, arylalkenylene groups and arylalkynylene groups may be substituted, and optionally said groups may be interrupted by one or more heteroatoms, preferably 1 to 100 heteroatoms, said heteroatoms preferably being selected from the group consisting of O, S(O)y and NR2, wherein y is 0, 1 or 2, preferably y = 2, and R2 is independently selected from the group consisting of hydrogen, halogen, Ci - C24 alkyl groups, Cg - C24 (hetero)aryl groups, C7- C24 alkyl(hetero)aryl groups and C7- C24 (hetero)arylalkyl groups. The optional substituents may be selected from polar groups, such as oxo groups, (poly)ethylene glycol diamines, (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains, carboxylic acid groups, carbonate groups, carbamate groups, cyclodextrins, crown ethers, saccharides (e.g. monosaccharides, oligosaccharides), phosphates or esters thereof, phosphonic acid or ester, phosphinic acid or ester, sulfoxides, sulfones, sulfonic acid or ester, sulfinic acid, or sulfenic acid. Said polar groups may also be present in the chain of L1. Preferred polar group include (poly)ethylene glycol diamines (e.g. 1,8-diamino- 3,6-dioxaoctane or equivalents comprising longer ethylene glycol chains), (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains and l,y' -diaminoalkanes (wherein y' is the number of carbon atoms in the alkane, preferably y' is 1-10).
[0119] Preferably, L1is selected from the group consisting of linear or branched, preferably linear, Ci - Cwo alkylene groups, C2 - Cwo alkenylene groups, C2 - Cwo alkynylene groups, C3 - C100 cycloalkylene groups, Cg - C100 cycloalkenylene groups, Cg - C100 cycloalkynylene groups, C7- Cwo alkylarylene groups, C7- Cioo arylalkylene groups, Cg- Cioo arylalkenylene groups and Cg- Cwo arylalkynylene groups, more preferably from the group consisting of linear or branched, preferably linear, Ci -C5o alkylene groups, C2 - C5o alkenylene groups, C2 - C5o alkynylene groups, C3 - C5o cycloalkylene groups, Cg - C5o cycloalkenylene groups, Cg - C5o cycloalkynylene groups, C7- C5o alkylarylene groups, C7- C5o arylalkylene groups, Cg - C5o arylalkenylene groups and Cg - C5o arylalkynylene groups, and even more preferably from the group consisting of linear or branched, preferably linear, Ci - C30 alkylene groups, C2 - C30 alkenylene groups, C2 - C30 alkynylene groups, C3 - C30 cycloalkylene groups, Cg - C30 cycloalkenylene groups, Cg - C30 cycloalkynylene groups, C7- C30 alkylarylene groups, C7- C30 arylalkylene groups, Cg - C30 arylalkenylene groups and Cg - C30 arylalkynylene groups, wherein optionally the alkylene groups, alkenylene groups, alkynylene groups, cycloalkylene groups, cycloalkenylene groups, cycloalkynylene groups, alkylarylene groups, arylalkylene groups, arylalkenylene groups and arylalkynylene groups may be substituted, and optionally said groups may be interrupted by one or more heteroatoms, said heteroatoms being selected from the group consisting of O, S(O)y and NR2, wherein y is 0, 1 or 2, preferably y = 2, and R2is independently selected from the group consisting of hydrogen, halogen, Ci - C12 alkyl groups, Cg - C12 (hetero)aryl groups, C7- C12 alkyl(hetero)aryl groups and C7- C12 (hetero)arylalkyl groups. As above, the optional substituents may be selected from polar groups, such as oxo groups, (poly)ethylene glycol diamines, (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains, carboxylic acid groups, carbonate groups, carbamate groups, cyclodextrins, crown ethers, saccharides (e.g. monosaccharides, oligosaccharides), phosphates or esters thereof, phosphonic acid or ester, phosphinic acid or ester, sulfoxides, sulfones, sulfonic acid or ester, sulfinic acid, or sulfenic acid. Said polar groups may also be present in the chain of L1. Preferred polar group include (poly)ethylene glycol diamines (e.g. l,8-diamino-3,6-dioxaoctane or equivalents comprising longer ethylene glycol chains), (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains and l,y' -diaminoalkanes (wherein y' is the number of carbon atoms in the alkane, preferably y' is 1-10).
[0120] D1is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide and a negatively or positively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer. The template DNA oligonucleotide, template oligopeptide, threading DNA oligonucleotide and threading oligopeptide are described in more detail below.
[0121] As described above, the bi-functionalized conjugates, particularly the POCs, according to the present invention, may be used for the determination of the amino acid sequence of the conjugate peptide moiety, for example by nanopore sequencing. In order to be able to determine the amino acid sequence via nanopore sequencing, one of the functionalizations D1or D2of the bi-functionalized conjugate (e.g., POC) should be a DNA oligonucleotide or oligopeptide that can be recognized by the nanopore motor protein. Such DNA oligonucleotide or oligopeptide is herein referred to as a template DNA oligonucleotide or template oligopeptide, respectively. A template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides. A template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids.
[0122] The required template depends on the type of nanopore that is used. Depending on the particular nanopore the template is a template DNA oligopeptide having a particular nucleotide sequence, or a template oligopeptide having a particular amino acid sequence. Some nanopores have more than one template that may be used. The skilled person knows which template to select for a particular nanopore. For example, when the nanopore is an MspA-Hel308 nanopore, the template is a DNA template oligonucleotide and the preferred sequence of the template DNA oligonucleotide is according to SEQ ID NO:1. Another template DNA oligonucleotide for the MspA-Hel308 nanopore is according to SEQ ID NO:2. When the nanopore is a CsgG nanopore, the template is a template oligopeptide and the preferred sequence of the template oligopeptide is according to SEQ ID NO:2. D1may also be a negatively charged moiety, also referred to as a threading moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide and a threading polymer, wherein a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged; a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively charged; and a threading polymer is defined as a negatively charged polymer.
[0123] D1may also be a positively charged moiety, also referred to as a threading moiety, wherein the positively charged moiety is selected from a threading oligopeptide and a threading polymer, wherein a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is positively charged; and a threading polymer is defined as a positively charged polymer.
[0124] In order to be able to determine the amino acid sequence via nanopore sequencing, one of D1and D2in the bi-functionalized conjugate (e.g., POC) should be a negatively charged moiety or a positively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a threading polymer, and wherein the positively charge moiety is selected from a threading oligopeptide and a threading polymer. This threading moiety provides an electrophoretic force which is needed to cause the conjugate to enter into the nanopore. The threading moiety may be any anionic or cationic molecule that enables threading through a nanopore, such as a MspA or a CsgG nanopore. The threading moiety should have a density of negative or positive charges that is sufficient for the conjugate to enter the nanopore.
[0125] Preferably, the threading moiety is a monodisperse oligonucleotide or oligopeptide comprising 10- 100 nucleotides or 2-100 amino acids, respectively, wherein the oligonucleotide or oligopeptide is negatively charged, or wherein the oligopeptide is positively charged.
[0126] A preferred threading DNA oligonucleotide is a monodisperse oligonucleotide of thymine nucleotides (T)x, wherein x is 10-100. Preferably, x is 20-80, more preferably 25-70 and even more preferably 30-60. Examples of several representative threading moieties have a sequence according to SEQ ID NO: 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14 and 15. Particularly preferred threading DNA oligonucleotide (T)xare (T)4o, (T)45, (T)5o, (T)S5 and (T)60, corresponding to SEQ ID NO: 8, SEQ ID NO: 9, SEQ ID NO: 10, SEQ ID NO: 11 and SEQ ID NO: 12, respectively.
[0127] Another preferred threading moiety is a monodisperse oligopeptide of glutamic acid (Glu)y, or, since glutamic acid has the one-letter code E, (E)y, wherein y is 3-40. Preferably, y is 4-30, more preferably 5-25. Particularly preferred threading oligopeptides are (Glu)?, (Glu)io, (Glu)iz, (Glu)i4, (Glu)is, (Glu)is and (Glu)ig, corresponding to SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23 and SEQ ID NO: 24, respectively. Yet another preferred threading moiety is a monodisperse oligopeptide of aspartic acid (Asp)z, or, since aspartic acid has the one-letter code D, (D)z, wherein z is 3-40. Preferably, z is 4-30, more preferably 5-25. Particularly preferred threading oligopeptides are (Asp)7, (Asp)io, (Asp)u, (Asp)i4, (Asp)i5, (Asp)ig and (Asp)i8, corresponding to SEQ ID NO: 31, SEQ ID NO: 32 SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 35, SEQ ID NO: 36 and SEQ ID NO: 37, respectively.
[0128] Yet another preferred threading moiety is a monodisperse oligopeptide of arginine (Arp)z, or, since arginine has the one-letter code R, (R)z, wherein z is 3-40. Preferably, z is 4-30, more preferably 5-25. Particularly preferred threading oligopeptides are (Arg)7, (Arg)io, (Arg)i2, (Arg)M, (Arg)i5, (Arg)i6and (Arg)i8, corresponding to SEQ ID NO: 44, SEQ ID NO: 45 SEQ ID NO: 46, SEQ ID NO: 47, SEQ ID NO: 48, SEQ ID NO: 49 and SEQ ID NO: 50, respectively.
[0129] Yet another preferred threading moiety is a monodisperse oligopeptide of lysine (Lys)z, or, since lysine has the one-letter code K, (K)z, wherein z is 3-40. Preferably, z is 4-30, more preferably 5-25. Particularly preferred threading oligopeptides are (l_ys)7, (Lys)io, (l_ys)i2, (l_ys)i4, (l_ys)i5, (l_ys)i6and (Lys)i8, corresponding to SEQ ID NO: 57, SEQ ID NO: 58 SEQ ID NO: 59, SEQ ID NO: 60, SEQ ID NO: 61, SEQ ID NO: 62 and SEQ ID NO: 63, respectively.
[0130]
[0131] In the bi-functionalized conjugates (e.g., POCs) according to the invention, one of D1and D2should be a threading moiety, i.e., a negatively charged moiety that is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a threading polymer as defined above, or a positively charged moiety that is selected from a threading oligopeptide and a threading polymer as defined above. When D1is a threading moiety then D2is a template DNA oligonucleotide or a template oligopeptide, and when D2is a threading moiety, then D1is a template DNA oligonucleotide or a template oligopeptide, wherein the template DNA oligonucleotide and the template oligopeptide are as defined above.
[0132] In a preferred embodiment of the process according to the invention, the modification of the C-terminus of the peptide of structure (1), or of the mixture of peptides according to structure (1), AA3is tyrosine and the tyrosine side chain is converted into a 1,2-quinone click tag F1via oxidation (step (ii-a2)„ followed by reaction of said 1,2-quinone click probe F1with D^L^n-Q1, wherein Q1is a click tag comprising a strained C=C bond or a strained CEC bond. In this embodiment it is further preferred that Q1is a strained (hetero)cyclooctyne or a strained (hetero)cycloalkene, more preferably a (hetero)cycloalkyne according to structure (3)-(20) or a (hetero)cycloalkene according to structure (21)-(33), as defined in more detail below. In a particularly preferred embodiment, Q1is a strained cyclooctyne according to structure (3). In this embodiment the oxidation of the tyrosine phenol group occurs preferably via reaction with an o- iodoxybenzoic acid ( I BX), preferably l,3-dihydro-l-hydroxy-3-oxo-l,2-benziodoxole-4-carboxylic acid 1-oxide (mIBX).
[0133] When step (ii-a2) is performed via oxidation with IBX, e.g. mIBX, a suitable solvent is a buffer solution such as for example phosphate, buffered saline (e.g. phosphate-buffered saline, Bis-Tris-buffered saline), citrate, HEPES, Bis-Tris and glycine. Suitable buffers are known in the art. Preferably, the buffer solution is phosphate- buffered saline (PBS) or Bis-Tris buffer. Said reaction is preferably performed at a temperature in the range of about 4 °C to about 50 °C, more preferably in the range of about 10 °C to about 45 °C, even more preferably in the range of about 20 °C to about 40 °C, and even preferably in the range of about 20 °C to about 30 °C. The pH is preferably in the range of about 5 to about 9, more preferably in the range of about 5.5 to about 8.5, yet more preferably in the range of about 6 to about 8. Even more preferably, step (ii) is performed at a pH in the range of about 7 to about 8.
[0134] When step (ii-a2) is performed via enzymatic oxidation with e.g. mushroom tyrosinase, the solvent is a suitable buffer solution, preferably a phosphate-buffered saline (PBS). Said reaction is preferably performed at a temperature of about 2-10 °C, preferably of about 3-6 °C and more preferably at a temperature of about 4 °C. The pH is preferably in the range of about 4 to about 7, more preferably in the range of about 5 to about 6, yet more preferably at a pH of about 5.5.
[0135] Step (Hi): Modification of the N-terminus of peptide (1)
[0136] In step (iii) of the process according to the invention, the N-terminus of the peptide according to structure (1), or the N-terminus of peptides in the mixture of two or more peptides of structure (1), is modified. Said modification comprises the following steps:
[0137] (iii-al) formation of an imidazolidinone via imine condensation of the terminal N-atom with a 2- pyridinecarboxaldehyde and cyclization to form the imidazolidinone, wherein the 2-pyridinecarboxaldehyde comprises a click tag F2, and with the proviso that AA2is not a proline side chain; or
[0138] (iii-a2) attaching a click tag F2via AA1, and
[0139] (iii-b) reacting click tag F2with D2-(L2)n-Q2, wherein Q2is a click tag that is capable of reacting with F2in a click reaction, L2is a linker and n is 0 or 1, and D2is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer; wherein a template DNA oligonucleotide, a template oligopeptide, a threading DNA oligonucleotide and a threading oligopeptide are as defined above.
[0140] In step (iii-al) an imidazolidinone is formed via imine condensation of the terminal N-atom with a 2-pyridinecarboxaldehyde (a 2-PCA), followed by cyclization to form the imidazolidinone. The 2-pyridine- carboxaldehyde comprises a click tag F2.
[0141] Preferably, the 2-pyridinecarboxaldehyde (2-PCA) in step (iiia-1) is according to structure (2): wherein:
[0142] F2is a click tag; m is 0 or 1;
[0143] L3is a linker; and
[0144] R1is independently selected from the group consisting of hydrogen and linear or branched Ci - Cio alkyl groups.
[0145] R1is preferably independently selected from the group consisting of hydrogen and linear or branched Ci - C4alkyl groups, more preferably from the group consisting of hydrogen and linear or branched Ci - C3 alkyl groups, yet more preferable from the group consisting of hydrogen and Ci - C? alkyl groups. Even more preferably R1is hydrogen or -CH3 and most preferably R1is hydrogen.
[0146] L3is a linker. L3is selected independently from L1and L2, in other words, i.e. L1, L2and L3may be the same or different. The presence of L3is optional (m is 0 or 1). As described above, linkers are known in the art. L3may for example be selected from the group consisting of linear or branched (preferably linear) Ci - C200 alkylene groups, C2 - C200 alkenylene groups, C2 - C2ooalkynylene groups, C3 - C200 cycloalkylene groups, Cg - C200 cycloalkenylene groups, Cg - C200 cycloalkynylene groups, C7 - C200 alkylarylene groups, C7- C200 arylalkylene groups, Cg - C200 arylalkenylene groups and Cg - C200 arylalkynylene groups. Optionally the alkylene groups, alkenylene groups, alkynylene groups, cycloalkylene groups, cycloalkenylene groups, cycloalkynylene groups, alkylarylene groups, arylalkylene groups, arylalkenylene groups and arylalkynylene groups may be substituted, and optionally said groups may be interrupted by one or more heteroatoms, preferably 1 to 100 heteroatoms, said heteroatoms preferably being selected from the group consisting of O, S(O)y and NR2, wherein y is 0, 1 or 2, preferably y = 2, and R2is independently selected from the group consisting of hydrogen, halogen, Ci - C24 alkyl groups, Cg - C24 (hetero)aryl groups, C7- C24 alkyl(hetero)aryl groups and C7- C24 (hetero)arylalkyl groups. The optional substituents may be selected from polar groups, such as oxo groups, (poly)ethylene glycol diamines, (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains, carboxylic acid groups, carbonate groups, carbamate groups, cyclodextrins, crown ethers, saccharides (e.g. monosaccharides, oligosaccharides), phosphates or esters thereof, phosphonic acid or ester, phosphinic acid or ester, sulfoxides, sulfones, sulfonic acid or ester, sulfinic acid, or sulfenic acid. Said polar groups may also be present in the chain of L1. Preferred polar group include (poly)ethylene glycol diamines (e.g. l,8-diamino-3,6-dioxaoctane or equivalents comprising longer ethylene glycol chains), (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains and l,y' -diaminoalkanes (wherein y' is the number of carbon atoms in the alkane, preferably y' is 1-10).
[0147] Preferably, L3is selected from the group consisting of linear or branched, preferably linear, Ci - Cwo alkylene groups, C2 - C100 alkenylene groups, C2 - Cwo alkynylene groups, C3 - C100 cycloalkylene groups, Cg - C100 cycloalkenylene groups, Cg - C100 cycloalkynylene groups, C7- C100 alkylarylene groups, C7- C100 arylalkylene groups, Cg - C100 arylalkenylene groups and Cg - C100 arylalkynylene groups, more preferably from the group consisting of linear or branched, preferably linear, Ci -C5o alkylene groups, C2 - C5o alkenylene groups, C2 - C5o alkynylene groups, C3 - C5o cycloalkylene groups, Cg - C5o cycloalkenylene groups, Cg - C5o cycloalkynylene groups, C7- C5o alkylarylene groups, C7- C5o arylalkylene groups, Cg - C5o arylalkenylene groups and Cg - C5o arylalkynylene groups, and even more preferably from the group consisting of linear or branched, preferably linear, Ci -C30 alkylene groups, C2 - C30 alkenylene groups, C2 - C30 alkynylene groups, C3 - C30 cycloalkylene groups, Cg - C30 cycloalkenylene groups, Cg - C30 cycloalkynylene groups, C7- C30 alkylarylene groups, C7- C30 arylalkylene groups, Cg- Cso arylalkenylene groups and Cg - Cso arylalkynylene groups, wherein optionally the alkylene groups, alkenylene groups, alkynylene groups, cycloalkylene groups, cycloalkenylene groups, cycloalkynylene groups, alkylarylene groups, arylalkylene groups, arylalkenylene groups and arylalkynylene groups may be substituted, and optionally said groups may be interrupted by one or more heteroatoms, said heteroatoms being selected from the group consisting of O, S(O)y and NR2, wherein y is 0, 1 or 2, preferably y = 2, and R2is independently selected from the group consisting of hydrogen, halogen, Ci - C12 alkyl groups, Cg - C12 (hetero)aryl groups, C7- C12 alkyl(hetero)aryl groups and C7- C12 (hetero)arylalkyl groups. As above, the optional substituents may be selected from polar groups, such as oxo groups, (poly)ethylene glycol diamines, (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains, carboxylic acid groups, carbonate groups, carbamate groups, cyclodextrins, crown ethers, saccharides (e.g. monosaccharides, oligosaccharides), phosphates or esters thereof, phosphonic acid or ester, phosphinic acid or ester, sulfoxides, sulfones, sulfonic acid or ester, sulfinic acid, or sulfenic acid. Said polar groups may also be present in the chain of L3. Preferred polar group include (poly)ethylene glycol diamines (e.g. l,8-diamino-3,6-dioxaoctane or equivalents comprising longer ethylene glycol chains), (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains and l,y' -diaminoalkanes (wherein y' is the number of carbon atoms in the alkane, preferably y' is 1-10).
[0148] Step (iii-al) of the process according to the invention is shown below.
[0149] In 2-pyridinecarboxaldehyde (2) it is preferred that R1is H and m is 1. Preferably, L3is a linear linker.
[0150] Preferred structures for L3include (49) - (51):
[0151] (49) (50) wherein: o is 1 to 60, preferably 1-40, more preferably 1-20 and even more preferably 1-10; and p is 1 to 20, preferably 1-15, more preferably 1-10 and even more preferably 1-5.
[0152] In step (iii-a2) of the modification of the N-terminus of the peptide, a click tag F2is attached to the peptide via amino acid side chain AA1. The skilled person knows which methods are suitable for the introduction of a particular click tag F2via a particular amino acid side chain AA1.
[0153] Several examples of such modifications are described above, in the context of step (ii-al), for the amino acid side chains of tyrosine, tryptophan and arginine. When AA1is a serine side chain, this side chain can easily be converted into e.g. a nitrone or nitrile oxide.
[0154] Kjaersgaard et al., ChemBioChem. 2022, 23, e202200245 (incorporated by reference) reviewed the introduction of a click tag via the side chain of numerous amino acids.
[0155] In a preferred embodiment click tag F2is selected from the group consisting of (hetero)cycloalkynes, (hetero)cycloalkenes, azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones.
[0156] In step (iii-b) click tag F2is reacted with D2-(L2)n-Q2, wherein Q2is a click tag that is capable of reacting with F2in a click reaction, L2is a linker and n is 0 or 1, and D2is selected from the group consisting of a template DNA oligonucleotide or template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer. Preferably, the negatively charged moiety is a threading DNA oligonucleotide or a threading oligopeptide. Preferably, the positively charged moiety is a threading oligopeptide. A template DNA oligonucleotide, template polypeptide, threading DNA oligonucleotide, threading oligopeptide, negatively charged polymer, and positively charged polymer are as defined above for step (ii-b).
[0157] In a preferred embodiment click tag Q2is selected from the group consisting of of (hetero)cycloalkynes, (hetero)cycloalkenes, azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones.
[0158] The click reaction of F1with Q1and F2with Q2, and preferred embodiments thereof, are described in more detail below.
[0159] L2is a linker. L2is selected independently from L1, in other words, i.e. L1and L2may be the same or different. The presence of L2is optional (m is 0 or 1). In D2-(L2)m-Q2, D2and Q2are covalently connected to each other via linker L2, if present. As described above, linkers are known in the art. L2may for example be selected from the group consisting of linear or branched Ci -C200 alkylene groups, C2 - C200 alkenylene groups, C2 - C200 alkynylene groups, C3 - C200 cycloalkylene groups, Cg - C200 cycloalkenylene groups, Cg - C200 cycloalkynylene groups, C7- C200 alkylarylene groups, C7- C200 arylalkylene groups, Cg - C200 arylalkenylene groups and Cg - C200 arylalkynylene groups. Optionally the alkylene groups, alkenylene groups, alkynylene groups, cycloalkylene groups, cycloalkenylene groups, cycloalkynylene groups, alkylarylene groups, arylalkylene groups, arylalkenylene groups and arylalkynylene groups may be substituted, and optionally said groups may be interrupted by one or more heteroatoms, preferably 1 to 100 heteroatoms, said heteroatoms preferably being selected from the group consisting of O, S(O)y and NR2, wherein y is 0, 1 or 2, preferably y = 2, and R2is independently selected from the group consisting of hydrogen, halogen, Ci - C24 alkyl groups, Cg - C24 (hetero)aryl groups, C7- C24 alkyl(hetero)aryl groups and C7- C24 (hetero)arylalkyl groups. The optional substituents may be selected from polar groups, such as oxo groups, (poly)ethylene glycol diamines, (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains, carboxylic acid groups, carbonate groups, carbamate groups, cyclodextrins, crown ethers, saccharides (e.g. monosaccharides, oligosaccharides), phosphates or esters thereof, phosphonic acid or ester, phosphinic acid or ester, sulfoxides, sulfones, sulfonic acid or ester, sulfinic acid, or sulfenic acid. Said polar groups may also be present in the chain of L1. Preferred polar group include (poly)ethylene glycol diamines (e.g. 1,8-diamino- 3,6-dioxaoctane or equivalents comprising longer ethylene glycol chains), (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains and l,y' -diaminoalkanes (wherein y' is the number of carbon atoms in the alkane, preferably y' is 1-10). Preferably, L2is selected from the group consisting of linear or branched, preferably linear, Ci - Cioo alkylene groups, C2 - Cioo alkenylene groups, C2 - Cioo alkynylene groups, C3 - Cioo cycloalkylene groups, Cg - Cioo cycloalkenylene groups, Cg - Cioo cycloalkynylene groups, C7- Cioo alkylarylene groups, C7 - C100 arylalkylene groups, Cg- Cioo arylalkenylene groups and Cg- Cioo arylalkynylene groups, more preferably from the group consisting of linear or branched, preferably linear, Ci -C5o alkylene groups, C2 - C5o alkenylene groups, C2 - C5o alkynylene groups, C3 - C5o cycloalkylene groups, Cg - C5o cycloalkenylene groups, Cg - C5o cycloalkynylene groups, C7- C5o alkylarylene groups, C7- C5o arylalkylene groups, Cg - C5o arylalkenylene groups and Cg - C5o arylalkynylene groups, and even more preferably from the group consisting of linear or branched, preferably linear, Ci -C30 alkylene groups, C2 - C30 alkenylene groups, C2 - Csoalkynylene groups, C3
[0160] - C30 cycloalkylene groups, Cg - C30 cycloalkenylene groups, Cg - C30 cycloalkynylene groups, C7- C30 alkylarylene groups, C7- C30 arylalkylene groups, Cg - C30 arylalkenylene groups and Cg - C30 arylalkynylene groups, wherein optionally the alkylene groups, alkenylene groups, alkynylene groups, cycloalkylene groups, cycloalkenylene groups, cycloalkynylene groups, alkylarylene groups, arylalkylene groups, arylalkenylene groups and arylalkynylene groups may be substituted, and optionally said groups may be interrupted by one or more heteroatoms, said heteroatoms being selected from the group consisting of O, S(O)y and NR2, wherein y is 0, 1 or 2, preferably y = 2, and R2is independently selected from the group consisting of hydrogen, halogen, Ci - C12 alkyl groups, Cg - C12 (hetero)aryl groups, C7- C12 alkyl(hetero)aryl groups and C7
[0161] - C12 (hetero)arylalkyl groups. As above, the optional substituents may be selected from polar groups, such as oxo groups, (poly)ethylene glycol diamines, (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains, carboxylic acid groups, carbonate groups, carbamate groups, cyclodextrins, crown ethers, saccharides (e.g. monosaccharides, oligosaccharides), phosphates or esters thereof, phosphonic acid or ester, phosphinic acid or ester, sulfoxides, sulfones, sulfonic acid or ester, sulfinic acid, or sulfenic acid. Said polar groups may also be present in the chain of L2. Preferred polar group include (poly)ethylene glycol diamines (e.g. l,8-diamino-3,6-dioxaoctane or equivalents comprising longer ethylene glycol chains), (poly)ethylene glycol or (poly)ethylene oxide chains, (poly)propylene glycol or (poly)propylene oxide chains and l,y' -diaminoalkanes (wherein y' is the number of carbon atoms in the alkane, preferably y' is 1-10).
[0162] In the bi-functionalized conjugate (e.g., POC) according to the invention one of D1and D2is a template DNA oligonucleotide or a template oligopeptide. In other words, either D1or D2, but not both, is a template DNA oligonucleotide or a template oligopeptide. In one embodiment of the bi-functionalized conjugate (e.g., POC), D1is a template DNA oligonucleotide or a template oligopeptide and D2 is a negatively charged moiety or a positively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer. In another embodiment, D1is a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, a negatively charged polymer, and D2is a template DNA oligonucleotide or a template oligopeptide. In a further preferred embodiment, the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide and a negatively charged polymer, and D2is a template DNA oligonucleotide or template oligopeptide. In yet another embodiment, D1is a positively charged moiety, wherein the positively charged moiety is selected from a threating oligopeptide and a positively charged polymer, and D2is a template DNA oligonucleotide or a template oligopeptide.
[0163] Step (iii) is performed in a suitable solvent, preferably a suitable buffer solution such as for example phosphate, buffered saline (e.g. phosphate-buffered saline, tris-buffered saline), citrate, HEPES, tris and glycine. Suitable buffers are known in the art. Preferably, the buffer solution is phosphate-buffered saline (PBS) or tris buffer. Step (iii) is preferably performed at a temperature in the range of about 4 °C to about 50 °C, more preferably in the range of about 10 °C to about 45 °C, even more preferably in the range of about 20 °C to about 40 °C, and most preferably in the range of about 20 °C to about 30 °C. The pH is preferably in the range of about 5 to about 9, more preferably in the range of about 5.5 to about 8.5, yet more preferably in the range of about 6 to about 8. Even more preferably, step (iii) is performed at a pH in the range of about 7 to about 8.
[0164] As described above, the order of step (ii) and step (iii) may be reversed, and the modification of the N-terminus of peptide (1) of step (iii) may be performed before the modification of the C-terminus of peptide (1) of step (ii). Preferably, the C-terminus of the peptide or mixture of peptides is performed before the N- terminus; the order of the steps of the process according to the invention is thus preferably (i), (ii), (iii).
[0165] In steps (ii) and (iii) preferably an excess of reagents is used to drive the reaction to completion. After step (ii) it is preferred to remove excess reagents remaining in the reaction mixture before step (iii-al) or (iii- a2) is performed. Similarly, it is preferred to remove excess reagents after step (iii-al) or (iii-a2) before step (iii-b) is performed. Also when modification of the N-terminus of peptide (1) or a mixture of two or more peptides (1) occurs before the C-terminus modification, i.e. step (iii) is performed prior to step (ii), it is preferred that any excess reagents is removed in between the process steps.
[0166] Any reactants that may remain in the reaction mixture after step (ii), step (iii-a) and step (iii-b) may be removed via purification methods known in the art. For example, if an excess of 2-pyridinecarboxaldehyde (2-PCA) is used in step (iii-al) for the installation of click tag F2at the N-terminus of peptide (1), said excess may be removed. These purification methods are known in the art and the skilled person is able to select suitable methods based on the specific circumstances of the process. As an example, after step (iii-al) peptide-fragments that have not been connected to DNA by means of the SPOCQ. reaction may be removed, by spin-filtration, for example over 10 kDa MWCO filter. After step (iii-b) unreacted Q2-(L2)n-D2may be removed by streptavidin capture approaches, e.g. biotin-streptavidin bead fishing using MeTz-biotin. Unreacted PCA of step (iii-a)-l may be caught and removed by biotin-oxyamine. In the process according to the present invention, a click reaction between click tags F1and Q1takes place in step (ii-b) and a click reaction between click tags F2and Q2takes place in step (iii-b). Because these click reactions take place in separate steps of the process there is no need for these click reactions to differ from each other. Both click reactions may be the same, and compatible click tags F1and Q1may be the same as compatible click tags F2and Q2. However, both click reactions may also be different and compatible click tags F1and Q1may differ from compatible click tags F2and Q2.
[0167] In a preferred embodiment, the conjugation reaction between Q1and F1is a cycloaddition, preferably a 1,3-dipolar cycloaddition or a [4+2] cycloaddition. The cycloaddition is preferably a metal-free strain-promoted cycloaddition.
[0168] A typical [4+2] cycloaddition is the (hetero)-Diels-Alder reaction, wherein Q1is a diene or a dienophile. As appreciated by the skilled person, the term "diene" in the context of the Diels-Alder reaction refers to l,3-(hetero)dienes, and includes conjugated dienes (R2C=CR— CR=CR2), imines (e.g. R2C=CR-N=CR2or R2C=CR— CR=NR, R2C=N— N=CR2) and carbonyls (e.g. R2C=CR— CR=O or O=CR-CR=O). Hetero-Diels-Alder reactions with N- and O-containing dienes are known in the art. Any diene known in the art to be suitable for [4+2] cycloadditions may be used as reactive group Q1. Preferred dienes include tetrazines, 1,2-quinones and triazines. Although any dienophile known in the art to be suitable for [4+2] cycloadditions may be used as reactive group Q1, the dienophile is preferably an alkene or alkyne group as described above, even more preferably an alkyne group. For conjugation via a [4+2] cycloaddition, it is preferred that Q1is a dienophile (and F1is a diene), more preferably Q1is or comprises an alkynyl group.
[0169] For a 1,3-dipolar cycloaddition, Q1is a 1,3-dipole or a dipolarophile. Any 1,3-dipole known in the art to be suitable for 1,3-dipolar cycloadditions may be used as reactive group Q1. Preferred 1,3-dipoles include azido groups, nitrone groups, nitrile oxide groups, nitrile imine groups and diazo groups. Although any dipolarophile known in the art to be suitable for 1,3-dipolar cycloadditions may be used as reactive groups Q1, the dipolarophile is preferably an alkene or alkyne group, more preferably an alkyne group. For conjugation via a 1,3-dipolar cycloaddition, it is preferred that Q1is a dipolarophile (and F1is a 1,3-dipole), more preferably Q1is or comprises an alkynyl group.
[0170] Thus, in a preferred embodiment, Q1is selected from dipolarophiles and dienophiles and F1is selected from dipoles and dienes.
[0171] The skilled person is capable to determine which combination of Q1and F1is suitable for a proper conjugation reaction, i.e. which Q1and F1are compatible. Preferred options for F1are selected from an azide, tetrazine, triazine, nitrone, nitrile oxide, nitrile imine, diazo compound, 1,2-quinone, dioxothiophene and sydnone. Preferably F1is an azide or a 1,2-quinone.
[0172] The above equally applies to the conjugation reaction between Q2and F2. In a preferred embodiment of the process of the invention, F1and F2are independently selected from the group consisting of (hetero)cycloalkynes, (hetero)cycloalkenes, azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones.
[0173] In another preferred embodiment of the process of the invention Q1and Q2are independently selected from the group consisting of (hetero)cycloalkynes, (hetero)cycloalkenes, azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones.
[0174] It is further preferred that F1or Q1, and / or F2or Q2, is a (hetero)cycloalkyne or a (hetero)cycloalkene.
[0175] More preferably F1or Q1, and / or F2or Q2, preferably Q1and / or Q2, comprises a (hetero)cycloalkynyl group according to structure (52): wherein:
[0176] R8is independently selected from the group consisting of hydrogen, halogen, -OR9, -NO2, -CN, -S(O)2R9, -S(O)3( ), Ci - C24 alkyl groups, Cg - C24 (hetero)aryl groups, C7- C24 alkyl(hetero)aryl groups and C7- C24 (hetero)arylalkyl groups and wherein the alkyl groups, (hetero)aryl groups, alkyl(hetero)aryl groups and (hetero)arylalkyl groups are optionally substituted, wherein two substituents R8may be linked together to form an optionally substituted annulated cycloalkyl or an optionally substituted annulated (hetero)arene substituent, and wherein R9is independently selected from the group consisting of hydrogen, halogen, Ci - C24 alkyl groups, Cg - C24 (hetero)aryl groups, C7- C24 alkyl(hetero)aryl groups and C7- C24 (hetero)arylalkyl groups; u is 0, 1, 2, 3, 4 or 5; u' is 0, 1, 2, 3, 4 or 5, wherein u + u' = 4, 5, 6, 7 or 8; v = an integer in the range of 8-16.
[0177] In a preferred embodiment, u + u' = 4, 5 or 6, more preferably u + u' = 5. Typically, v = (u + u') x 2 or [(u + u') x 2] - 1. In a preferred embodiment, v = 8, 9 or 10, more preferably v = 9 or 10, even more preferably v = 10.
[0178] In a particularly preferred embodiment, F1or Q1and / or F2or Q2comprises a (hetero)cycloalkynyl group according to structure (3) - (20) or a (hetero)cycloalkenyl group according to structure (21) - (33):
[0179] wherein R3in (27) and (28) is alkyl or aryl.
[0180] More preferably, Q1and / or Q2is a (hetero)cycloalkyne or a (hetero)cycloalkene, and yet more preferably Q1and / or Q2is a (hetero)cycloalkyne according to structure (3)-(20) or a (hetero)cycloalkene according to structure (21)-(33). In an especially preferred embodiment, F1or Q1and / or F2or Q2comprises a cyclooctynyl group according to structure (53): wherein:
[0181] R10is independently selected from the group consisting of hydrogen, halogen, -OR9, -NO2, -CN, -S(O)2R9, -S(O)3( ), CI - C24 alkyl groups, C5 - C24 (hetero)aryl groups, C7- C24 alkyl(hetero)aryl groups and C7- C24 (hetero)arylalkyl groups and wherein the alkyl groups, (hetero)aryl groups, alkyl(hetero)aryl groups and (hetero)arylalkyl groups are optionally substituted, wherein two substituents R15may be linked together to form an optionally substituted annulated cycloalkyl or an optionally substituted annulated (hetero)arene substituent, and wherein R9is independently selected from the group consisting of hydrogen, halogen, Ci - C24 alkyl groups, Cg - C24 (hetero)aryl groups, C7
[0182] - C24 alkyl(hetero)aryl groups and C7- C24 (hetero)arylalkyl groups;
[0183] R11is independently selected from the group consisting of hydrogen, halogen, Ci - C24 alkyl groups, Cg
[0184] - C24 (hetero)aryl groups, C7- C24 alkyl(hetero)aryl groups and C7- C24 (hetero)arylalkyl groups;
[0185] R12is selected from the group consisting of hydrogen, halogen, Ci - C24 alkyl groups, Cg - C24 (hetero)aryl groups, C7- C24 alkyl(hetero)aryl groups and C7- C24 (hetero)arylalkyl groups, the alkyl groups optionally being interrupted by one of more hetero-atoms selected from the group consisting of O, N and S, wherein the alkyl groups, (hetero)aryl groups, alkyl(hetero)aryl groups and (hetero)arylalkyl groups are independently optionally substituted, or R12is a second occurrence of Q1or Q2connected via a spacer moiety; and
[0186] I is an integer in the range of 0 to 10.
[0187] More preferably, Q1and / or Q2is a cycloalkyl group according to structure (53).
[0188] In a preferred embodiment of the reactive group according to structure (53), R10is independently selected from the group consisting of hydrogen, halogen, -OR9, Ci - Cg alkyl groups, C5- Cg (hetero)aryl groups, wherein R9is hydrogen or Ci - Cg alkyl, more preferably R10is independently selected from the group consisting of hydrogen and Ci - Cg alkyl, even more preferably all R10are H. In a preferred embodiment of the reactive group according to structure (53), R11is independently selected from the group consisting of hydrogen, Ci - Cg alkyl groups, more preferably both R11are H. In a preferred embodiment of the reactive group according to structure (53), R12is H. In a preferred embodiment of the reactive group according to structure (53), I is 0 or 1, more preferably I is 1. Since F1and Q1are compatible click tags, and F2and Q2are compatible click tags, it is also preferred that F1and / or F2is a click probe that is compatible with a (hetero)cycloalkyne or a (hetero)cycloalkene. Therefore, F1and / or F2are independently selected from the group consisting of azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones.
[0189] It is further preferred that F1and / or F2are independently according to structure (37)-(48): wherein:
[0190] R4is selected from Ci - C4alkyl or aryl;
[0191] R5is selected from hydrogen, alkyl, aryl, acyl and sulfonyl;
[0192] R6is selected from hydrogen, Ci - C24 alkyl groups, C2 - C24 acyl groups, C3 - C24 cycloalkyl groups, C2 - C24 (hetero)aryl groups, C3 - C24 alkyl(hetero)aryl groups, C3 - C24 (hetero)arylalkyl groups and Ci - C24 sulfonyl groups, each of which (apart from hydrogen) may optionally be substituted and optionally interrupted by one or more heteroatoms selected from O, S and NR7wherein R7is independently selected from the group consisting of hydrogen and Ci - C4alkyl groups.
[0193] When a tyrosine side chain is converted into a 1,2-quinone in step (ii-a2), in other words when F1is a 1,2-quinone, the click reaction with Q1in step (ii-b) preferably occurs via a Strain-Promoted Oxidation-Controlled 1,2-Quinone Cycloaddition (SPOCQ) to a (hetero)cycloalkyne or a (hetero)cycloalkene. Preferred (hetero)cycloalkynes are according to structures (3)-(20), (52) and (53), and preferred (hetero)- cycloalkenes are according to structures (21)-(33).
[0194] In another preferred embodiment F1is an azide. A preferred click reaction when F1is an azide is a Strain-Promoted Azide-Alkyne Cycloaddition (SPAAC). In this reaction Q1is preferably a (hetero)cycloalkyne or a (hetero)cycloalkene. Preferred (hetero)cycloalkynes are according to structures (3)-(20), (52) and (53), and preferred (hetero)cycloalkenes are according to structures (21)-(33).
[0195] As described above, the amino acid sequence of the peptide moiety of the bi-functionalized conjugate (e.g., POC) according to the process of the invention may be determined. The present invention thus also relates to a process for the preparation of a bi-functionalized conjugate, wherein the conjugate is a peptideoligonucleotide or peptide-oligopeptide conjugate, as defined above, the process further comprising a step of determining the amino acid sequence of the obtained bi-functionalized conjugate, or, when in step (i) a mixture or more two or more peptides of structure (1) was provided, wherein the process further comprises a step of determining the amino acid sequence of at least one of the obtained bi-functionalized conjugates.
[0196] Preferably, the peptide amino acid sequence is determined via nanopore analysis (sequencing). It is further preferred that the peptide amino acid sequence is determined, including post-translational modifications (PTMs) that may be present. The present invention thus also relates to a process for the preparation of a bi-functionalized conjugate, wherein the conjugate is a peptide-oligonucleotide or a peptide-oligopeptide conjugate, as defined above, the process further comprising a step of determining the amino acid sequence of the obtained bi-functionalized conjugate, or, when in step (i) a mixture or more two or more peptides of structure (1) was provided, wherein the process further comprises a step of determining the amino acid sequence of at least one of the obtained bi-functionalized conjugates, wherein the amino acid sequence is determined via nanopore sequencing, and wherein the peptide amino acid sequence is determined including post-translational modifications (PTMs) that may be present.
[0197] A particularly preferred embodiment of the process according to the present invention is shown in Figure 1. A C-terminal tyrosine side chain is oxidized to an ortho-quinone using l,3-dihydro-l-hydroxy-3-oxo- l,2-benziodoxole-4-carboxylic acid 1-oxide (mIBX), followed by SPOCQ. with BCN-DNA. The obtained POC is then / V-te rm inally modified with an azide functionalized pyridine carboxaldehyde that subsequently is clicked to another BCN-DNA via SPAAC. AA1and AA2indicate amino acid side chains (AA2cannot be a proline residue), R represents any spacer between the pyridine and azide group.
[0198] A preferred embodiment of the process according to the invention to synthesize POCs from proteins that subsequently can be analyzed by nanopore sequencing is shown in Figure 1. The process according to the invention involves bioconjugations on both the C- and the N-terminus of the peptide according to structure (1), resulting in a bi-functionalized conjugate (e.g., POC). In the process shown in Figure 1, a C-terminal tyrosine residue is required in order to install the electrophoretic linker via click chemistry. C- terminal tyrosine may be obtained in a peptide mixture from protein digestion by C-terminal cleavage of tyrosine residues, e.g. by the digestive enzyme chymotrypsin or chymase.
[0199] Obtained peptide mixtures of two or more peptides of structure (1) from protein digestion were subjected to ligation to D1and D2, wherein one of D1and D2is a template DNA oligonucleotide or a template oligopeptide, and the other of D1and D2is a threading moiety, which is a negatively or positively charged moiety. The threading moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, a negatively charged polymer, and a positively charged polymer. Herein, a template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides; a template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids; a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged; a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively charged or positively charged; a negatively or positively charged polymer typically is a biological or synthetic polymer.
[0200] The C-terminal tyrosine residue was subjected to Strain-Promoted Oxidation-Controlled Cyclooctyne-1,2-Quinone Cycloaddition (SPOCQ) chemistry to introduce a threading moiety D1. For this, the tyrosine residue is oxidized using l,3-dihydro-l-hydroxy-3-oxo-l,2-benziodoxole-4-carboxylic acid 1-oxide (mIBX) to generate ortho-quinone in situ and is immediately conjugated to a threading moiety, e.g. a DNA strand, that contains a strained cyclic alkyne following the inverse electron-demand Diels-Alder (IEDDA) mechanism. SPOCQ chemistry benefits from fast reaction rates, with particularly rapid kinetics for strained alkyne bicyclo[6.1.0]non-4-yne (BCN)-OH (53).
[0201] Introduction of the template moiety to the obtained C-terminal modified conjugate (e.g., POC) requires N-terminal modification. In Figure 1 a general approach for N-terminal protein modification under mild conditions is shown, presented by 2-pyridinecarboxaldehyde (2-PCA). In this method, the N-terminal amine first undergoes imine condensation, followed by cyclization via nucleophilic attack of the a-amide nitrogen from the neighboring amino acid to form a stable imidazolidinone. The imidazolidinone can only form at the N-terminus, as the e-amine of a lysine side chain lacks the nearby amide group, making the 2-PCA N-terminal modification highly selective. Restrictions are the presence of a proline residue as the second amino acid, i.e. AA2is not proline. The 2-PCA platform has been enriched with orthogonal functionalities to enable a bioconjugation follow-up reaction. For example, an azide group was installed on 2-PCA and the N-terminal modified peptide was conjugated to a template moiety, e.g. a template DNA oligonucleotide as defined above, via SPAAC click chemistry. SPAAC is a catalyst-free click reaction that proceeds under physiological conditions.
[0202] Each modification and bioconjugation step in the process shown in Figure 1 was explored using dummy peptides, and an optimized protocol was developed and applied to enzymatically digested peptides and protein. The process according to the invention is applicable to all kind of peptides and proteins, which also may contain PTMs such as sulfotyrosine and / or phosphotyrosine residues. The amino acid sequence of the bi-functionalized conjugates (e.g., POCs) according to the invention may be determined by use of nanopore protein sequencing.
[0203] Bi-functionalized peptide-oligonucleotide or peptide-oligopeptide conjugate
[0204] In a second aspect the present invention relates to the bi-functionalized conjugate, wherein the conjugate are peptide-oligonucleotide or peptide-oligopeptide conjugate, obtainable by the process according to the invention, and to a mixture of bi-functionalized conjugates, wherein the conjugates are peptide-oligonucleotide or peptide-oligopeptide conjugates, obtainable by the process when in step (i) a mixture of two or more peptides according to structure (1) is provided.
[0205] In addition, the invention relates to a bi-functionalized conjugate, wherein the conjugate is a peptideoligonucleotide or a peptide-oligopeptide conjugate, according to structure (34): wherein:
[0206] G is a peptide comprising 1 - 50 amino acids;
[0207] AA1and AA2are independently selected from the group of amino acid side chains, with the proviso that AA2is not proline;
[0208] L1, L2and L3are linkers; m is 0 or 1; n is independently selected from 0 or 1;
[0209] R1is independently selected from the group consisting of hydrogen and linear or branched Ci - Cio alkyl groups;
[0210] Z1is a connecting group;
[0211] Z2is a connecting group;
[0212] D1and D2are selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a negatively charged moiety, and a positively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer; wherein a template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides; a template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids; a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged; a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively charged or positively charged; and with the proviso that one of D1and D2is a template DNA oligonucleotide or a template oligopeptide.
[0213] As defined above, AA1and AA2are independently selected from the group of amino acid side chains. Preferably, AA1and AA2are independently selected from the group of amino acid side chains of natural amino acids, typically those described herein.
[0214] Z1and Z2are connecting groups, which covalently connect D1and D2, via L1and L2if present, with the peptide moiety of the bifunctional conjugate, wherein the conjugate is a peptide-oligonucleotide or a peptide-oligopeptide conjugate, according to the invention. The term "connecting group" herein refers to the structural element resulting from a reaction, here between click tag Q. and click tag F, connecting one part of the conjugate with another part of the same conjugate. Z1is formed by a click reaction between click tags Q1and F1. Likewise, Z2is formed by a click reaction between click tag Q2and F2. Herein, Z refers to Z1and Z2, Q. refers to Q1and Q2, and F refers to F1and F2.
[0215] The exact nature of connecting group Z depends on the exact structure of the click tags Q. and F from which it is formed, as will be understood by the person skilled in the art. The skilled person is aware of complementary click tags which are reactive towards each other and form suitable reaction partners Q7F1and Q2 / F2. For example, when F comprises or is an alkynyl group, complementary groups Q. include azido groups, and when F comprises or is an azido group, complementary groups Q. include alkynyl groups. For example, when F comprises or is a cyclopropenyl group, a trans-cyclooctene group, a cycloheptyne or a cyclooctyne group, complementary groups Q. include tetrazinyl groups. In these particular cases, Z is only an intermediate structure and will expel N2, thereby generating a dihydropyridazine (from the reaction with alkene) or pyridazine (from the reaction with alkyne) as shown in Figure 2.
[0216] Connecting groups Z are obtained by a cycloaddition reaction, preferably wherein the cycloaddition is a [4+2] cycloaddition or a 1,3-dipolar cycloaddition. Click reactions via cycloadditions are known to the skilled person and are described above in more detail. As also described above, preferred cycloadditions are a [4+2]- cycloaddition (e.g. a Diels-Alder reaction) or a [3+2]-cycloaddition (e.g. a 1,3-dipolar cycloaddition). Preferably, the cycloaddition is a Diels-Alder reaction or a 1,3-dipolar cycloaddition. The preferred Diels-Alder reaction is the inverse electron-demand Diels-Alder cycloaddition (IEDDA). In another preferred embodiment, the 1,3-dipolar cycloaddition is used, more preferably the alkyne-azide cycloaddition. Cycloadditions, such as Diels-Alder reactions and 1,3-dipolar cycloadditions are known in the art, and the skilled person knows how to perform them.
[0217] The bi-functionalized conjugate according to structure (34) is obtainable by the process according to the invention, which is described in detail above. Connecting group Z1is obtained via the click reaction of F1and Q1in step (ii) of the process, and connecting group Z2is obtained via the click reaction of F2and Q2in step (iii) of the process. When F1and F2are independently selected from the group consisting of azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones, complementary click tags Q1and Q2include (hetero)cycloalkynes and (hetero)cycloalkenes. Examples of connecting groups Z for several of these complementary click tags are shown in Figure 2. The connecting groups Z shown in Figure 2 are preferred connecting groups in the present invention. Of note, the pyridine or pyridazine connecting group is the product of the rearrangement of the tetrazabicyclo[2.2.2]octane connecting group that is formed upon reaction of triazine or tetrazine with alkyne (but not alkene), respectively, with loss of N?.
[0218] Preferably, Z1in the bi-functionalized conjugate (e.g., POC) according to structure (34) is formed via a click reaction wherein F1is 1,2-quinone and Q1a (hetero)cycloalkyne or (hetero)cycloalkene. Consequently Z1is preferably according to structure (55): wherein the bond depicted as — is a single bond or a double bond, Z is obtained via a cycloaddition, and ring size Z is defined by the ring size of (hetero)cycloalkyne or (hetero)cycloalkene click tag Q1.
[0219] Template DNA oligonucleotides, template polypeptides, threading DNA oligopeptides and threading oligopeptides, and preferred embodiments thereof, are described in more detail above.
[0220] R1is preferably independently selected from the group consisting of hydrogen and linear or branched Ci - C4alkyl groups, more preferably from the group consisting of hydrogen and linear or branched Ci - C3 alkyl groups, yet more preferable from the group consisting of hydrogen and Ci - C2alkyl groups. Even more preferably R1is hydrogen or -CH3 and most preferably R1is hydrogen.
[0221] Preferred embodiments of connecting groups Z1and Z2are according to structures (56) - (76) as shown below.
[0222]
[0223] In a particularly preferred embodiment, the bi-functionalized conjugate (e.g., POC) is according to structure (35): wherein: AA1, AA2, G, D1, D2, L1, L2, R1, L1, m and n are as defined above, and with the proviso that one of D1and
[0224] D2is a template DNA oligonucleotide or a template oligopeptide.
[0225] Preferred embodiments of AA1, AA2, G, D1, D2, L1, L2, R1and L1are as described above. The invention further relates to a bi-functionalized conjugate, wherein the conjugate is a peptideoligonucleotide or a peptide-oligopeptide conjugate, according to structure (36): wherein:
[0226] AA1, AA2, G, D1, F2, L1, R1, L3and m are as defined above.
[0227] The bi-functionalized conjugate (e.g., POC) according to structure (36) may be obtained via the process according to the invention, by performing the C-terminal modification of a peptide of structure (1) via step (ii-a2) followed by (ii-b), and then introducing a click tag F2via step (iii-al) of said process. Bi-functionalized conjugate (e.g., POC) (36) is thus an intermediate of the process according to the invention.
[0228] The amino acid sequence of the peptide part of the bi-functionalised conjugate (e.g., POC) obtainable by the process according to the invention, and the bi-functionalised conjugate (e.g., POC) according to structure (34) and (35), may be determined by nanopore sequencing. In these bi-functionalised conjugates either D1or D2is a template DNA oligonucleotide or a template oligopeptide, and the other is a threading moiety, i.e. a negatively charged moiety selected from a threading DNA oligonucleotide, a threading oligopeptide and a negatively charged polymer, preferably selected from a threading DNA oligonucleotide and a threading oligopeptide, or a positively charged moiety selected from a threading oligopeptide and a positively charged polymer. As described above, for a particular type of nanopore, a particular type of template, i.e. a particular template DNA oligonucleotide or template oligopeptide, is required. The template DNA oligonucleotide or template oligopeptide is recognized by a nanopore. The threading moiety, which is negatively charged or positively charged, provides the required electrophoretic force that is needed to cause the conjugate, particularly the POC, to enter into the nanopore.
[0229] Since the process according to the invention tolerates labile sulfotyrosine as a post-translational modification (PTM), it is expected that the orthogonal click approach, in particular the SPOCQ / SPAAC click approach, to synthesize sandwich POCs will be useful for a variety of protein mixtures, revealing new information on structure and modifications associated to biologically relevant proteins.
[0230] Examples
[0231] General procedures Starting materials, reagents, and solvents were purchased from commercial vendors and used as received unless stated otherwise. Triethylammonium acetate (IM solution) was purchased from Honeywell / Fluka. 6-(bromomethyl)-2-pyridinemethanol, pyridinium chlorochromate (PCC), iodoacetamide (IAA), and tetrabutylammonium hydrogen sulfate (TBAHS) were purchased from TCI Europe. 1,3-dihydro-l- hydroxy-3-oxo-l,2-benziodoxole-4-carboxylic acid 1-oxide (mIBX), agarose, lysozyme human recombinant expressed in rice (LysoH), insulin chain B oxidized from bovine pancreas (InsuB), sodium azide (NaNs), 6-(l- piperazinylmethyl)-2-pyridinecarboxaldehyde bis-tosylate salt, and urea were purchased from Sigma-Aldrich. Modified amino acids Fmoc-Tyr(SO2ONp)-OH and Fmoc-Tyr(PO(OBzl)OH)-OH, angiotensin II human (Angio2), biotin-dPEGs-oxyamine HCI, and Tris hydrochloride (Tris-HCI) were purchased from Merck Life Science. Chymotrypsin Sequencing Grade was manufactured by Roche Diagnostics GmbH, and purchased from Merck Life Science. Recombinant Macrovipera lebetina Chymotrypsin-like protease VLCTLP was purchased from Cusabio. BCN-Cg-DNA oligonucleotides were purchased from Thermo Fisher Life Sciences (USA). Dithiothreitol (DTT), Pierce™ High Capacity Streptavidin Agarose, Amicon™ Ultra-0.5 Centrifugal Filter Units 10 kDa and 30 kDa MWCO, Pierce™ Peptide Desalting Spin Columns, Pierce™ Centrifuge Columns 0.8 mL, Tris-borate-EDTA buffer (TBE buffer lOx), 6x DNA loading dye, and SYBR™ Gold Nucleic Acid Gel Stain were purchased from Thermo Fisher Scientific, ss-20 DNA Ladder was purchased from Simplex Sciences. Azidoacetic acid NHS ester, azido-PEG4-NHS ester, biotin-PEG4-methyltetrazine were purchased from Broadpharm.
[0232] Reactions were monitored by thin-layer chromatography (TLC) using Merck aluminum sheets (Silica gel 60 F254) with detection by UV absorption (254 nm), by spraying with a solution of KMnO4(lO g / L) and K2CO3 (50 g / L) in water, followed by charring at about 150 °C. Organic solvents were removed under reduced pressure at 40 °C. Flash column chromatography was performed using Silia Flash® P60 silica gel (particle size of 40-63 pm, pore diameter of 60 A) with the indicated eluents.3H and13C NMR spectra were recorded using a Bruker AV-400 (400 and 101 MHz, respectively) spectrometer in the given deuterated solvent. Chemical shifts are given in ppm (6) relative to the residual solvent peak or tetramethyl silane (0 ppm) as internal standard and coupling constants are given in Hz. Multiplicity is reported as s: singlet, d: doublet, dd: doublet of doublets, td: triplet of doublets, t: triplet, m: multiplet. Assignments were made by standard COSY and HSQC analysis. High-resolution mass spectrometry (HRMS) analysis was performed with a Q Exactive Focus Mass Spectrometer (Thermo Fisher), equipped with a heated electrospray ion source (HESI) in positive mode. MS-grade methanol was used as eluent. The high-resolution mass spectrometer was calibrated prior to measurements with a calibration mixture (Thermo Finnigan).
[0233] Synthesized peptides and peptide-oligonucleotide conjugates (POCs) were analyzed by ultra-high performance liquid chromatography (UHPLC) coupled to mass spectrometry (HESI-MS, measuring in negative mode, calibrated with a Thermo Finnigan calibration mixture) using a Q. Exactive Focus Agilent 1290 Infinity UHPLC-MS system. Generally, for peptide analysis 10 pL of a 1 mg / mL solution is injected, and for POC analysis 10 pL of a 50 pM solution is injected. The UHPLC system is equipped with a diode array detector (DAD G4212A, at 215 nm for peptide detection and 260 nm for DNA detection) and a Dr. Maisch ReproSil Gold 120 C18, 3 pm, 200 x 3 mm column containing a 10 mm guard with a flow rate of 0.4 mL / min. The eluent contained for buffer A: 10 mM triethylammonium acetate in M ill i-Q (deionized water, produced with a Milli- Q Integral 3 system; Millipore, Molsheim / France) and for buffer B: acetonitrile (MeCN). A gradient of 5^5^75^75^5^5%, percentage buffer B (0^5^25^28^29^35 min) was applied. Mass spectrometry data is deconvoluted with the use of UniDec software.41
[0234] For each step of the process according to the invention, an optimized protocol was developed.
[0235] Example 1. General procedure for Solid-Phase Peptide Synthesis
[0236] Peptides were synthesized following Fmoc / tBu Solid-Phase Peptide Synthesis (SPPS) strategy. Chain elongation was initiated from Fmoc-Tyr(tBu)-Wang resin 100-200 mesh Novabiochem®, a p-alkoxy-benzyl alcohol polymer-bound (polystyerene-1%, DVB) amino acid (loading capacity 0.66 mmol / g). SPPS was performed in an automatic peptide synthesizer (CS336X Peptide Synthesizer CS BIO Co.). In general, amino acids were added as follows: the resin was pre-swollen with dichloromethane (DCM). The Fmoc-protecting group was removed using 20% piperidine in / V, / V-dimethylformamide (DMF, 2 x 8 min). The resin was then washed with DMF (3 x 2 min). The next amino acid was treated with 2-(lH-benzotriazole-l-yl)-l, 1,3,3- tetramethyluronium hexafluorophosphate (HBTU), 1-hydroxy-benzotriazole (HOBt), and N,N- diisopropylethylamine (DIPEA) in DMF for 2 min before being added to the resin. The reaction mixture was allowed to couple for 3 h. The resin was then washed with DMF (3 x 2 min), deprotected by 20% piperidine in DMF (2 x 8 min), and washed again with DMF (3 x 2 min). The same steps were repeated until the desired sequence was obtained. After last Fmoc removal step, the peptides were cleaved from the resin by treatment with a cocktail of 95% trifluoroacetic acid (TFA), 2.5% triisopropylsilane (TIS), and 2.5% Milli-Q for 2 h (10 mL TFA cocktail / 1 g initial resin). The cleaved peptides were precipitated by dropwise addition of the TFA cocktail-peptide mixture to ice-cold diethyl ether (1 : 1 of ether : hexane, lOx initial cocktail volume) and the cleaved peptide resin was washed once with a small amount of fresh cleavage cocktail. The precipitate was centrifuged for 10 min at 6000 rpm. The supernatant was discarded and the precipitate was washed with ice- cold diethyl ether and again centrifuged for 10 min at 6000 rpm. This washing step was repeated once more. The resulting precipitate was then dried in a light stream of Nj, redissolved in MeCN : Milli-Q (1 : 1), snap frozen with liquid N2and then lyophilized (Labconco FreeZone lyophilizer, 2.5 L, -84 °C, connected to a 35j xDS Edwards Oil-Free Dry Scroll Pump). The crude peptides containing a sulfated tyrosine residue were treated with 2M NH4OAc to remove the neopentyl protecting group at 45 °C for 40 h. Obtained (deprotected) peptides were purified by (semi)preparative reverse phase HPLC (Agilent 1260 Preparative HPLC with a DAD G7115A and MSD) using either a preparative Grace Alltima column (C18, 22 x 250 mm, 5-Micron) with a flow rate of 20 mL / min or a semi-preparative Zorbax Eclipse column (XDB-C18, 9.4 x 250 mm, 5-Micron) with a flow rate of 10 mL / min. Peptides without a sulfated tyrosine residue were purified with buffers containing 0.1% formic acid (FA). Buffer A: 95% Milli-Q, 5% MeCN, 0.1% FA; buffer B: 95% MeCN, 5% M illi-Q, 0.1% FA. A gradient of 5^5^95^95^5^5% percentage buffer B (0^5^25^30^35^40 min) was used. For peptides with sulfated tyrosine residue(s) the buffer contained 10 mM triethylammonium acetate (TEAA). Buffer A: 10 mM TEAA in Milli-Q; buffer B: MeCN. A gradient of 5^5^60^60^5^5% percentage buffer B (0^5^25^30^35^40 min) was used. The purified peptide fractions were subsequently lyophilized.
[0237] Example 2. N-terminal modification of peptides by 2-PCA azide derivatives
[0238] Synthesized 2-PCA azide derivatives were aliquoted and stored as 100 mM solution in DMSO at -20 °C. From a 2 mM stock solution of peptide, 2 pL (4 nmol, final concentration 100 pM) was taken and added to 32 pL of 50 mM phosphate buffer at pH 7.5. To this solution 4 pL from a 100 mM stock solution of 2-PCA azide derivative (400 nmol, final concentration 10 mM) was added. The reaction was stirred overnight at room temperature and then analyzed by UHPLC-MS.
[0239] Example 3. General procedure for Peptide-Oligonucleotide Conjugation
[0240] 1 mM of peptide or peptide mixture from digestion (20 nmol, from a 2 mM stock solution in Milli-Q), 250 pM of BCN-Cg-threading DNA (5 nmol, from a 1 mM stock solution in Milli-Q), and 25 mM mIBX (500 nmol, from a 100 mM stock solution in PBS buffer; 10 mM NajHPO^ 1.8 mM NaHjPO^ 137 mM NaCI, 2.7 mM KCI, pH 7.3) were added to a 0.5 mL microcentrifuge tube and reacted at room temperature for 4 h. (To analyze the SPOCQ reaction, 10 pL of a 50 pM DNA solution was injected in the UHPLC-MS system.) Then, the solution was diluted with 40 pL phosphate buffer (50 mM, pH 7.5) and 20 pL of azide functionalized 2- PCA compound (2 pmol, from a 100 mM stock solution in DMSO, 25 mM) was added. The reaction mixture was stirred overnight before being purified with Amicon® ultra-spin filtration units with 10 kDa molecular weight cut-off (MWCO) using phosphate buffer. Spin filter was first pre-rinsed with 450 pL phosphate buffer and centrifuged at 14000 g for 5 min. Samples were diluted with phosphate buffer to 450 pL and loaded to the pre-rinsed spin filter and centrifuged at 14000 g for 15 min. The reaction tube was washed 3x with 150 pL phosphate buffer and loaded into the spin filter. The spin filter was again centrifuged at 14000 g for 15 min. Subsequently, the spin-filter was washed with 400 pL phosphate buffer and centrifuged at 14000 g for 15 min. The washing step was repeated two more times, while for the last washing step the centrifugation lasted 20 min. Product was then recovered by reversed filtration of spin filter at 5000 g for 1 min. Spin filter was then washed once with 20 pL phosphate buffer and again reversed filtrated at 5000 g for 1 min. (To analyze the PCA-modified POC, 10 pL of a 50 pM DNA solution was injected in the UHPLC-MS system.) Obtained / V-terminal modified POC was then reacted overnight with BCN-Cg-template DNA (10 nmol, from a 1 mM stock solution in Milli-Q, about 143 pM) at room temperature. Subsequently, 2 pL biotin-PEG4-methyltetrazine (20 nmol, from 10 mM stock in DMSO) and 2 pL biotin-PEGs-oxyamine (20 nmol, from 10 mM stock in DMSO) were added. The reaction mixture was allowed to stir overnight. To a 0.8 mL centrifuge column 530 pL of High Capacity Streptavidin Agarose Resin was added and centrifuged at 5000 g for 1 min to remove the slurry. The column was then equilibrated by washing with PBS buffer (3 x 300 pL) and subsequently centrifuged at 5000 g for 1 min. Sample was then loaded to the centrifuge column containing streptavidin agarose and reaction tube was once washed with 100 pL PBS buffer. The centrifuge tube was incubated for 15 min while shaking on a thermoshaker at 500 rpm. After incubation, the centrifuge column was centrifuged at 5000 g for 1 min and washed once with 200 pL PBS buffer and centrifuged again at 5000 g for 1 min. Combined filter fractions were then further purified with Amicon® ultra-spin filtration units with 30 kDa MWCO using phosphate buffer, similarly as described above. Spin filters were pre-rinsed, loaded with sample, loaded with wash sample from reaction tube, and washed 3x. All centrifugation steps were executed at 14000 g for 10 min. The final washing step lasted for 15 min. Product was recovered by reversed spin filtration at 5000 g for 1 min and spin filter was washed once with 20 pL phosphate buffer and again reversed filtrated at 5000 g for 1 min. Obtained template DNA-PCA-peptide-threading DNA construct was then analyzed with UHPLC-MS (10 pL of recovered product was injected in the UHPLC-MS system).
[0241] Example 4. General procedure for peptide digestion with chymotrypsin
[0242] Peptides were enzymatically hydrolyzed by enzyme in 50 mM NH4CO3, pH 8.3, at 37 °C. In a typical experiment, 20 pL of 50 mM NH4CO3 solution of substrate (5 mg / mL) in an Eppendorf tube was thermally equilibrated to 37 °C in the Thermoshaker at 550 rpm. The reaction was started by adding 5 pL enzyme solution (1 mg / mL in 1 mM HCI) to obtain a 1 : 20 enzyme : substrate w / w ratio. The mixture was diluted to a 100 pg / mL enzyme concentration and is reacted for 18 h. The reaction was quenched by diluting the mixture with 0.1% TFA, and was subsequently treated with Pierce™ Peptide Desalting Spin Columns. Spin columns were conditioned, sample was bound, washed and eluted as desalted peptides. The procedure followed is described in the User Guide found at ThermoFisher Scientific. The obtained eluate was lyophilized and re-suspended in PBS buffer to a 2 mM stock solution. Obtained fragments were analyzed by UHPLC-MS and further used in the procedure for peptide-oligonucleotide conjugation.
[0243] Example s. General procedure for protein digestion with proteases
[0244] Example 5a. Human lysozyme digestion with chymotrypsin
[0245] Human lysozyme (LysoH), recombinant expressed in rice, was dissolved in a minimal amount of 6 M urea and further diluted with Tris buffer (100 mM Tris-HCI, 10 mM CaCL, pH 8.0) to obtain a 1 mM stock solution of LysoH (max. 20% urea). Disulfide bonds were reduced with DTT. Typically, 20 pL of 1 mM LysoH was reacted with 4 pL of 30 mM DTT (in Tris buffer) at 55 °C for 20 min (5 mM DTT concentration in reaction). Free cysteine residues were then alkylated with IAA. To the reduced mixture, 4 pL of 105 mM IAA (in Tris buffer) was added and reacted for 15 min in the dark (15 mM IAA concentration in reaction). After alkylation, the reaction mixture was further diluted to 50 .L total volume (urea concentration is now 0.6 M). The mixture was thermally equilibrated to 37 °C in the Thermoshaker at 550 rpm. The reaction was started by adding 10 pL enzyme solution (1 mg / mL in 1 mM HCI) to obtain a 1 : 20 enzyme : substrate w / w ratio. The mixture was diluted to a 100 pg / mL enzyme concentration and is reacted for 18 h. The reaction was quenched by diluting the mixture with 0.1% TFA, and was subsequently treated with Pierce™ Peptide Desalting Spin Columns. Spin columns were conditioned, sample was bound, washed and eluted as desalted peptides. The procedure followed is described in the User Guide found at ThermoFisher Scientific. The obtained eluate was lyophilized and re-suspended in PBS buffer to a 2 mM stock solution. Obtained fragments were analyzed by UHPLC-MS and further used in the procedure for peptide-oligonucleotide conjugation.
[0246] Example 5b. Human Lysozyme digestion with chymase
[0247] Human lysozyme (LysoH), recombinant expressed in rice, was dissolved in a minimal amount of 6 M urea and further diluted with Tris buffer (100 mM Tris-HCI, 10 mM CaCL, pH 8.0) to obtain a 1 mM stock solution of LysoH (max. 20% urea). Disulfide bonds were reduced with DTT. Typically, 20 pL of 1 mM LysoH was reacted with 4 pL of 30 mM DTT (in Tris buffer) at 55 °C for 20 min (5 mM DTT concentration in reaction). Free cysteine residues were then alkylated with IAA. To the reduced mixture, 4 pL of 105 mM IAA (in Tris buffer) was added and reacted for 15 min in the dark (15 mM IAA concentration in reaction). After alkylation, the reaction mixture was further diluted to 50 pL total volume (urea concentration is now 0.6 M). The mixture was thermally equilibrated to 37 °C in the Thermoshaker at 550 rpm. The reaction was started by adding 50 pL enzyme solution (250 pg / mL in 1 mM HCI) to obtain a 1 : 40 enzyme : substrate w / w ratio. The mixture was diluted to a 100 pg / mL enzyme concentration with PBS pH 7.3, and is reacted for 18 h at 37 °C. The reaction was quenched by diluting the mixture with 0.1% TFA, and was subsequently treated with Pierce™ Peptide Desalting Spin Columns. Spin columns were conditioned, sample was bound, washed and eluted as desalted peptides. The procedure followed is described in the User Guide found at ThermoFisher Scientific. The obtained eluate was lyophilized and re-suspended in PBS buffer to a 2 mM stock solution. Obtained fragments were analyzed by UHPLC-MS.
[0248] Example 6. Agarose gel electrophoresis for SPAAC purification optimization
[0249] Agarose (2.5 g) was added to TBE buffer (lx, 50 mL) and dissolved by heating in a microwave. Upon fully dissolving, the mixture was left to cool to 60 °C and SYBR™ Gold Nucleic Acid Gel Stain (lOOOOx, 5 pL) was added and mixed by swirling through the solution. Subsequently, the gel was immediately cast and left to solidify (5% agarose gel). Samples were prepared after concentration determination by a Nanodrop system. With the known extinction coefficient (template DNA: 706740 M^-crn1, template DNA + threading DNA combined: 1126740 M^-crn1), the concentration of the samples were calculated by using the obtained A260 values. Samples were prepared in M ill i-Q. using 6x DNA loading dye. The gel casket was filled with lx TBE buffer and the gel was placed in the casket. Wells were loaded with the prepared samples; for both template DNA and POC samples 200 ng was loaded. For the marker, ss-20 DNA Ladder (600 ng) was loaded. The gel was run with a Bio-Rad Horizontal Electrophoresis System at 150 V for 60 min. Afterwards, the gel was scanned with a GelDoc Imaging System using ImageLab software.
[0250] Example 7 Synthesis of 2-pyridinecarboxaldehyde azide derivatives
[0251] Example 7.1. 6-(azidomethyl)-2-pyridinemethanol (78)
[0252] The synthesis was performed as described previously.366-(bromomethyl)-2-pyridinemethanol (77, 85 mg, 0.42 mmol) was dissolved in 750 pL THF. NaNa (55 mg, 0.84 mmol) was dissolved in 750 pL H2O and subsequently added to the red solution. TBAHS (15 mg, 0.042 mmol) was added portion wise. The biphasic mixture was stirred vigorously at room temperature in the dark for 3 h. After reaction completion, the layers were separated and the organic layer was concentrated under reduced pressure. The crude product was purified by silica gel column chromatography (30% ethyl acetate in hexane) to afford the title compound 78 as a white solid (40 mg, 0.24 mmol, 57%).3H NMR (400 MHz, CDCI3) 6 7.73 (t, J = 7.7 Hz, 1H), 7.25 (d, J = 7.2 Hz, 1H), 7.22 (d, J = 7.8 Hz, 1H), 4.77 (s, 2H), 4.48 (s, 2H), 3.88 (s, 1H).13C NMR (101 MHz, CDCI3) 6 159.3, 154.9, 138.1, 120.7, 112.0, 64.0, 55.2. HRMS (ESI): m / z = [M+Na]+calc for C7H8N4ONa 187.0590, found 187.0589.
[0253] Example 7.2. 6-(azidomethyl)-2-pyridinecarboxaldehyde (6-AM-2-PCA, 79)
[0254] The synthesis was performed as described previously.366-(azidomethyl)-2-pyridinemethanol (78, 40 mg, 0.24 mmol) was dissolved in 1 mL dichloromethane (DCM) and cooled to 0 °C. After ten minutes, PCC (52 mg, 0.24 mmol) was added portion wise. The reaction mixture was stirred for 1 h, filtered, and concentrated under reduced pressure. The crude product was purified by silica gel column chromatography (hexane 20% ethyl acetate in hexane) to obtain the final product 79 as a colorless oil (25 mg, 0.15 mmol, 63%).3H NMR (400 MHz, DMSO) 6 9.98 (d, J = 0.7 Hz, 1H), 8.10 (td, J = 7.7, 0.8 Hz, 1H), 7.90 (dd, J = 7.7, 1.1 Hz, 1H), 7.75 (dd, J = 7.7, 1.1 Hz, 1H), 4.69 (s, 2H).13C NMR (101 MHz, DMSO) 6 193.3, 156.9, 152.1, 138.9, 127.0, 121.0, 54.0. HRMS (ESI): m / z = [M+Na]+calc for C7H6N4ONa 185.0434, found 185.0435. Example 7.3. azidoacetyl-piperazin-2-pyridinecarboxaldehyde (6-AAPM-2-PCA, 83)
[0255] 6-(l-Piperazinylmethyl)-2-pyridinecarboxaldehyde bis-tosylate salt (82, 40 mg, 73 pmol) was dissolved in / V, / V-dimethylformamide (DMF, 500 pL). Azidoacetic acid NHS ester (80, 17.3 mg, 87 pmol) and triethylamine (Et3N, 50.7 pL, 364 pmol) were added. The reaction mixture was stirred for 2 h. The reaction was monitored by TLC (10% methanol in ethyl acetate). Upon reaction completion, the crude mixture was concentrated under reduced pressure and subsequently purified by silica gel column chromatography (DCM 70% ethyl acetate in DCM) to afford 6-AAPM-2-PCA 83 as a colorless oil (15.7 mg, 54 pmol, 75%).3H NMR (400 MHz, DMSO) 6 9.97 (d, J = 0.7 Hz, 1H), 8.04 (t, J = 7.7 Hz, 1H), 7.84 (dd, J = 7.6, 1.2 Hz, 1H), 7.77 (dd, J = 7.7, 1.1 Hz, 1H), 4.14 (s, 2H), 3.75 (s, 2H), 3.49 (d, J = 5.9 Hz, 2H), 3.35 (d, J = 4.9 Hz, 2H), 2.59 (s, 2H), 2.46 (s, 2H).13C NMR (101 MHz, DMSO) 6 193.7, 172.8, 165.8, 151.8, 138.2, 127.6, 120.4, 63.0, 54.9, 52.6, 52.2, 49.6, 25.2. HRMS (ESI): m / z = [M+Na]+calc for Ci3Hi6N6O2Na 311.1227, found 311.1226.
[0256] Example 7.4. azido-PEG4-acetyl-piperazin-2-pyridinecarboxaldehyde (6-A-PEG4-PM-2-PCA, 84)
[0257] 6-(l-Piperazinylmethyl)-2-pyridine carboxaldehyde bis-tosylate salt (82, 40 mg, 73 pmol) was dissolved in / V, / V-dimethylformamide (DMF, 500 pL). Azido-PEG4-NHS ester (81, 34 mg, 87 pmol) and triethylamine (Et3N, 50.7 pL, 364 pmol) were added. The reaction mixture was stirred for 2 h. The reaction was monitored by TLC (10% methanol in ethyl acetate). Upon reaction completion, the crude mixture was concentrated under reduced pressure and subsequently purified by silica gel column chromatography (DCM ethyl acetate 5% methanol in ethyl acetate) to afford 6-A-PEG4-PM-2-PCA 84 as a colorless oil (27.3 mg, 57 pmol, 78%).TH NMR (400 MHz, DMSO) 6 9.96 (d, J = 0.8 Hz, 1H), 8.04 (td, J = 7.7, 0.7 Hz, 1H), 7.83 (dd, J = 7.6, 1.1 Hz, 1H), 7.77 (dd, J = 7.8, 1.2 Hz, 1H), 3.74 (s, 2H), 3.62-3.46 (m, 19H), 3.38 (dd, J = 5.6, 4.3 Hz, 2H), 2.59 (s, 1H), 2.55 (t, J = 6.7 Hz, 2H), 2.45 (t, J = 4.8 Hz, 2H), 2.40 (t, J = 5.2 Hz, 2H). 13C NMR (101 MHz, DMSO) 6 193.7, 172.8, 168.7, 151.8, 138.1, 127.5, 120.3, 69.83, 69.79, 69.77, 69.69, 69.67, 69.3, 66.8, 63.1, 54.9, 53.0, 52.5, 50.0, 45.0, 32.8, 25.2. HRMS (ESI): m / z = [M+Na]+calc for C22H34N6O6Na 501.2432, found 501.2413.
[0258] Example 8. C-terminal SPOCQ of dummy peptides For the conjugation of DNA to C-terminally positioned tyrosine residues, a one-pot SPOCQ approach was developed. Peptide and threading DNA (containing a BCN group) are added together with mIBX. The latter oxidizes tyrosine to generate an ortho-quinone which induces SPOCQ with the threading DNA in situ (Figure 3). The optimal reaction conditions were determined to be 1 mM peptide, 250 pM of threading DNA, and 25 mM mIBX in PBS buffer at pH 7.3 and room temperature. SPOCQ products 88-90 were obtained from respectively dummy peptides 85-87 (Figure 4). Reactions were monitored by UHPLC-MS and obtained MS spectra were deconvoluted using UniDec software (Marty et al., Anal. Chem., 2015, 87 (8), 4370-4376). Deconvoluted masses were in agreement with calculated masses (see ESI). Reaction optimization demonstrated that 4 hours of reaction time was sufficient to obtain full conversion of the SPOCQ click reaction. Moreover, it was found that purification after mIBX mediated oxidation-inducible SPOCQ reaction was not required before addition of the 2-PCA azide derivative for the / V-terminal modification reaction.
[0259] Figure 3 shows the chemical reaction of a peptide containing a C-terminal tyrosine with l,3-dihydro-l-hydroxy-3-oxo-l,2-benziodoxole-4-carboxylic acid 1-oxide (mIBX) and BCN-Cg-threading DNA (threading DNA) via SPOCQ. mIBX oxidizes the tyrosine residue to an ortho-quinone and is subsequently clicked with Threading DNA resulting in a peptide-oligonucleotide conjugate (POC).
[0260] Figure 4 (Example 8) shows the chemical structures and deconvolution of HESI mass spectra of SPOCQ products 88 (peptide 85 + threading DNA; Figure 4A), 89 (peptide 86 + threading DNA; Figure 4B), and 90 (peptide 87 + threading DNA; Figure 4C).
[0261] Example 9. N-terminal modification of POCs (after C-terminal SPOCQ)
[0262] After installation of the treading DNA at the C-terminus of the dummy peptides by means of SPOCQ, / V-terminal modification was achieved by addition of a 2-PCA molecule to form an / V-terminal imidazolidinone containing an azide click handle for down-stream conjugation (Figure 5). To minimize intermediate purification steps, we diluted the reaction mixture resulting from the SPOCQ reaction with 50 mM phosphate buffer at pH 7.5, and added azide functionalized 2-PCA compound (79, 83, or 84) (25 mM, 25% DMSO present in reaction). After overnight incubation and dilution of the sample with phosphate buffer, / V-terminally modified POCs were purified by 10 kDa MWCO spin filters. POCs 88-90 were treated with 2-PCA compound 79 to produce POCs 91-93 (Figure 6), with 2-PCA compound 83 to provide POCs 94-96 (Figure 7), and with 2-PCA compound 84 to yield POCs 97-99 (Figure 8).
[0263] Figure 5 shows the chemical reaction of a POC with an azide functionalized 2-PCA resulting in the / V-terminal imidazolidinone formation of the POC.
[0264] Figure 6 shows the deconvoluted mass spectra and chemical structures of POCs modified with 6-AM-2-PCA (79) resulting in products 91 (from POC 88), 92 (from POC 89), and 93 (from POC 90).
[0265] Figure 7 shows the deconvoluted mass spectra and chemical structures of POCs modified with 6-AAPM-2-PCA (83) resulting in products 94 (from POC 88), 95 (from POC 89), and 96 (from POC 90). Figure 8 shows the deconvoluted mass spectra and chemical structures of POCs modified with (D) 6-A-PEG4-PM-2-PCA (84) resulting in products 97 (from POC 88), 22 (from POC 98), and 99 (from POC 90).
[0266] Example 10. SPAAC with N-terminal modified POCs
[0267] The / V-terminal modified POCs recovered from 10 kDa spin filtration membranes were further reacted with template DNA in SPAAC click chemistry to obtain the final DNA-peptide-DNA sandwich POCs (Figure 9). For this, an excess of BCN-functionalized template DNA was added to ensure efficient SPAAC ligation. After this, it was found that optimization steps were required to purify the desired sandwich POC product from starting materials present in this step. For this, POC 93 was used for initial test reactions and after purification with 30 kDa MWCO spin filters, some starting materials and side product remained in the mixture, although desired product POC 102 was also abundantly present (Figure 5.5B). We also found that traces of unmodified 2-PCA-free POC 90 were present, but as this POC does not contain template DNA it only translocate through the nanopore in an uncontrolled manner.
[0268] As remaining template DNA could block Hel308, it was preferred to remove this from the mixture. An initial attempt to achieve this was by using azide- or methyltetrazine-functionalized agarose beads (de Bever et al., Bioconjugate Chem. 2023, 34 (3), 538-548). As this was only partially successful, we explored a two-step snatch-and-catch procedure by first reacting BCN-functionalized template DNA with biotin-PEGa- methyltetrazine (MeTz) followed by capturing with high capacity streptavidin agarose beads (Figure 5.5B). Incubation of biotin-captured template DNA with streptavidin agarose resulted in full removal of unreacted template DNA.
[0269] As can be seen in Figure 11, panel upper right corner, a side produce with a mass of 21925 Da, that corresponds to 2-PCA-functionalized template DNA (DNA-2-PCA), remained present. In order to remove this, we used biotin-PEG4-oxyamine (OA) to react with the aldehyde of template DNA-2-PCA under formation of an oxime, followed by capturing of the biotin-functionalized template DNA strand with streptavidin-agarose and centrifugal column purification (Timmers et al., Chem. Commun. 2023, 59 (76), 11397-11400). Afterwards, the obtained diluted solution was further concentrated and purified from small fragments with the use of 30 kDa MWCO spin filters, resulting in a clean sandwich POC that contains an / V-terminal template DNA and a C-terminal threading DNA (Figure 11). The other 6-AM-2-PCA (79) / V-terminally modified POCs (91 and 92) were treated according to the optimized procedure for SPAAC click and purification, and afforded sandwich POCs 100, 101 and 102, respectively (Figure 10). Similarly, SPAAC with 94, 95 and 96 afforded S10- S12 (not shown), and resulted in POCs S13-S15 with 97, 98, 99 (not shown), respectively.
[0270] The clean-up of unreacted materials and side-products was further evaluated with agarose gel electrophoresis. On a 5% agarose gel, product 102 was loaded without treatment, with only MeTz biotinstreptavidin treatment, with only OA biotin-streptavidin treatment, and with both MeTz and OA biotinstreptavidin treatment. The band intensity at the height of template DNA (about 70 nt's) was strongly reduced upon treatment with MeTz biotin-streptavidin. The combination of MeTz and OA biotin-streptavidin clean-up resulted in most intensified band for the product (between 120-140 nt's). The same samples were measured with UHPLC-MS. LC traces and HR-HESI mass spectra deconvolutions shows the necessity of the combination treatment with both bifunctional biotinylated compounds with subsequent streptavidin capturing to ensure fully purified sandwich POCs. Lastly, using the one-pot procedure with a single-sample mixture containing POCs 91, 92, 93, reacted with 2-PCA derivative 79, resulted in the expected mixture of POCs 100, 101, 102 (Figure 12). Similarly, a mixture of sandwich POC products S10-S12 and S13-S15 were obtained from POC mixtures 94, 95, 96 and 97, 98, 99 (not shown).
[0271] Example 11. Conversion of peptides obtained from a digested peptide or protein into POCs
[0272] 1 mM of peptide or peptide mixture from digestion (20 nmol, from a 2 mM stock solution in Milli-Q), 250 pM of BCN-Cg-threading DNA (5 nmol, from a 1 mM stock solution in Milli-Q), and 25 mM mIBX (500 nmol, from a 100 mM stock solution in PBS buffer; 10 mM Na2HPO4, 1.8 mM NaHzPO^ 137 mM NaCI, 2.7 mM KCI, pH 7.3) were added to a 0.5 mL microcentrifuge tube and reacted at room temperature for 4 h. Then, the solution was diluted with 40 pL phosphate buffer (50 mM, pH 7.5) and 20 pL of azide functionalized 2- PCA compound (2 pmol, from a 100 mM stock solution in DMSO, 25 mM) was added. The reaction mixture was stirred overnight before being purified with Amicon® ultra-spin filtration units with 10 kDa molecular weight cut-off (MWCO) using phosphate buffer. Spin filter was first pre-rinsed with 450 pL phosphate buffer and centrifuged at 14000 g for 5 min. Samples were diluted with phosphate buffer to 450 pL and loaded to the pre-rinsed spin filter and centrifuged at 14000 g for 15 min. The reaction tube was washed 3x with 150 pL phosphate buffer and loaded into the spin filter. The spin filter was again centrifuged at 14000 g for 15 min. Subsequently, the spin-filter was washed with 400 pL phosphate buffer and centrifuged at 14000 g for 15 min. The washing step was repeated two more times, while for the last washing step the centrifugation lasted 20 min. Product was then recovered by reversed filtration of spin filter at 5000 g for 1 min. Spin filter was then washed once with 20 pL phosphate buffer and again reversed filtrated at 5000 g for 1 min. Obtained / V-terminal modified POC was then reacted overnight with BCN-Cg-template DNA (10 nmol, from a 1 mM stock solution in Milli-Q, ~143 pM) at room temperature. Subsequently, 2 pL biotin-PEG4-methyltetrazine (20 nmol, from 10 mM stock in DMSO) and 2 pL biotin-PEGs-oxyamine (20 nmol, from 10 mM stock in DMSO) were added. The reaction mixture was allowed to stir overnight. To a 0.8 mL centrifuge column 530 pL of High Capacity Streptavidin Agarose Resin was added and centrifuged at 5000 g for 1 min to remove the slurry. The column was then equilibrated by washing with PBS buffer (3 x 300 pL) and subsequently centrifuged at 5000 g for 1 min. Sample was then loaded to the centrifuge column containing streptavidin agarose and reaction tube was once washed with 100 pL PBS buffer. The centrifuge tube was incubated for 15 min while shaking on a thermoshaker at 500 rpm. After incubation, the centrifuge column was centrifuged at 5000 g for 1 min and washed once with 200 pL PBS buffer and centrifuged again at 5000 g for 1 min. Combined filter fractions were then further purified with Amicon® ultra-spin filtration units with 30 kDa MWCO using phosphate buffer, similarly as described above. Spin filters were pre-rinsed, loaded with sample, loaded with wash sample from reaction tube, and washed 3x. All centrifugation steps were executed at 14000 g for 10 min. The final washing step lasted for 15 min. Product was recovered by reversed spin filtration at 5000 g for 1 min and spin filter was washed once with 20 pL phosphate buffer and again reversed filtrated at 5000 g for
[0273] 1 min. Obtained template DNA-PCA-peptide-threading DNA construct was then analyzed with UHPLC-MS (10 pL of recovered product was injected in the UHPLC-MS system).
Claims
1. Claims1. A process for the preparation of a bi-functionalized conjugate, wherein the conjugate is a peptideoligonucleotide or a peptide-oligopeptide conjugate, the process comprising the steps of:(i) providing a peptide according to structure (1):wherein:G is a peptide comprising 1-50 amino acids; andAA1and AA2are independently selected from the group of amino acid side chains;AA3is a tyrosine, tryptophan, arginine or lysine side chain;(ii) modifying the C-terminus of the peptide by:(ii-al) attaching a click tag F1via AA3; and(ii-b) reacting click tag F1with D1-(L1)n-Q1, wherein Q1is a click tag that is capable of reacting with F1in a click reaction, L1is a linker and n is 0 or 1, and D1is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer;(iii) modifying the N-terminus of the peptide by:(iii-al) formation of an imidazolidinone via imine condensation of the terminal N-atom with a 2-pyridinecarboxaldehyde and cyclization to form the imidazolidinone, wherein the 2-pyridinecarboxaldehyde comprises a click tag F2, and with the proviso that AA2is not a proline side chain; or(iii-a2) attaching a click tag F2via AA1, and(iii-b) reacting click tag F2with D2-(L2)n-Q2, wherein Q2is a click tag that is capable of reacting with F2in a click reaction, L2is a linker and n is 0 or 1, and D2is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negativelycharged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer; wherein a template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides; a template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids; a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged; a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively charged or positively charged; and with the proviso that one of D3and D2is a template DNA oligonucleotide or a template oligopeptide.
2. A process for the preparation of a bi-functionalized conjugate, wherein the conjugate is a peptideoligonucleotide or a peptide-oligopeptide conjugate, the process comprising the steps of:(i) providing a peptide according to structure (1):wherein:G is a peptide comprising 1-50 amino acids; andAA1and AA2are independently selected from the group of amino acid side chains;AA3is a tyrosine, tryptophan, arginine or lysine side chain;(ii) modifying the C-terminus of the peptide by:(ii-a2) converting AA3into a 1,2-quinone click tag F1via oxidation, with the proviso that AA3is a tyrosine side chain; and(ii-b) reacting click tag F1with D1-(L1)n-Q1, wherein Q1is a click tag that is capable of reacting with F1in a click reaction, L1is a linker and n is 0 or 1, and D1is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged55polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer;(iii) modifying the N-terminus of the peptide by:(iii-al) formation of an imidazolidinone via imine condensation of the terminal N-atom with a 2-pyridinecarboxaldehyde and cyclization to form the imidazolidinone, wherein the 2-pyridinecarboxaldehyde comprises a click tag F2, and with the proviso that AA2is not a proline side chain; or(iii-a2) attaching a click tag F2via AA1, and(iii-b) reacting click tag F2with D2-(L2)n-Q2, wherein Q2is a click tag that is capable of reacting with F2in a click reaction, L2is a linker and n is 0 or 1, and D2is selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negatively charged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer; wherein a template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides; a template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids; a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged; a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively charged or positively charged; and with the proviso that one of D1and D2is a template DNA oligonucleotide or a template oligopeptide.
3. Process according to claim 1 or 2, wherein AA3is a tyrosine, tryptophan or arginine side chain.
4. Process according to any one of claims 1-3, wherein AA3is a tyrosine side chain.
5. Process according to any one of claims 1-4, wherein in step (i) a mixture of two or more peptides according to structure (1) is provided.
6. Process according to claim 5, wherein the peptide mixture of step (i) is obtained by digestion of a protein or a mixture of proteins with a proteolytic enzyme.
7. Process according to any of claims 1-6, wherein F1and F2are independently selected from the group consisting of (hetero)cycloalkynes, (hetero)cycloalkenes, azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones.
8. Process according to any one of claims 1-7, wherein Q1and Q2are independently selected from the group consisting of (hetero)cycloalkynes, (hetero)cycloalkenes, azides, nitrones, nitrile oxides, tetrazines, triazines, 1,2-quinones, thiols, nitrile imines, diazo compounds, dioxothiophenes and sydnones.
9. Process according to any one of claims 1-8, wherein the 2-pyridinecarboxaldehyde in step (iii-al) is according to structure (2):wherein:F2is as defined in claim 1; m is 0 or 1;L3is a linker; andR1is independently selected from the group consisting of hydrogen and linear or branched Ci - Cio alkyl groups.
10. Process according to any of claims 1-9, wherein click tag F1or Q1and / or F2or Q2is a (hetero)cycloalkyne according to structure (3)-(20) or a (hetero)cycloalkene according to structure (21)-(33):57wherein:Y3in (23) is selected from C(R13, NR13and O, wherein R13is individually hydrogen or Ci - C6alkyl; and R3in (27) and (28) is alkyl or aryl.
11. Process according to any of claims 1-10, wherein AA3is a tyrosine side chain, F1is a 1,2-quinone andQ1is a (hetero)cycloalkyne or a (hetero)cycloalkene.
12. Process according to any of claims 1-11, wherein F2is an azide and Q2is a (hetero)cycloalkyne or a (hetero)cycloalkene.
13. Process according to any one of claims 1-11, wherein F2is a (hetero)cycloalkyne or a(hetero)cycloalkene and Q2is an azide.
14. Process according to any of claims 1-13, wherein the process is followed by a step of determining the amino acid sequence of the obtained bi-functionalized conjugate, or, when in step (i) a mixture or more two or more peptides of structure (1) was provided, wherein the process is followed by a step of determining the amino acid sequence of at least one of the obtained bi-functionalized conjugates.
15. Process according to claim 14, wherein the peptide amino acid sequence is determined via nanopore analysis.
16. Bi-functionalized conjugate, wherein the conjugate is a peptide-oligonucleotide or a peptide-oligopeptide conjugate, obtainable by the process according to any one of claims 1-12.
17. Mixture of bi-functionalized conjugates, wherein the conjugates are peptide-oligonucleotide or peptide-oligopeptide conjugates, obtainable by the process according to any one of claims 5-12.
18. Bi-functionalized conjugate, wherein the conjugate is a peptide-oligonucleotide or a peptide-oligopeptide conjugate, according to structure (34):G is a peptide comprising 1-50 amino acids;AA1and AA2are independently selected from the group of amino acid side chains, with the proviso that AA2is not proline;L1, L2and L3are linkers; m is 0 or 1; n is independently selected from 0 or 1;R1is independently selected from the group consisting of hydrogen and linear or branched Ci - Cio alkyl groups;Z1is a connecting group;Z2is a connecting group;D1and D2are selected from the group consisting of a template DNA oligonucleotide, a template oligopeptide, a positively charged moiety, and a negatively charged moiety, wherein the negativelycharged moiety is selected from a threading DNA oligonucleotide, a threading oligopeptide, and a negatively charged polymer, and wherein the positively charged moiety is selected from a threading oligopeptide and a positively charged polymer; wherein a template DNA oligonucleotide is defined as a DNA oligonucleotide that is recognized by a nanopore motor protein, said DNA oligonucleotide comprising 10 to 100 nucleotides; a template oligopeptide is defined as an oligopeptide that is recognized by a nanopore motor protein, said oligopeptide comprising 10 to 100 amino acids; a threading DNA oligonucleotide is defined as a DNA oligonucleotide comprising 10 to 100 nucleotides, wherein the oligonucleotide is negatively charged; a threading oligopeptide is defined as an oligopeptide comprising 2 to 100 amino acids, wherein the oligopeptide is negatively charged or positively charged; and with the proviso that one of D1and D2is a template DNA oligonucleotide or a template oligopeptide; and a connecting group is defined as the structural element resulting from a click reaction between click tag Q. and click tag F.
19. Bi-functionalized conjugate, wherein the conjugate is a peptide-oligonucleotide or a peptide-oligopeptide conjugate, according to claim 18, wherein said conjugate is according to structure (35):wherein:AA1, AA2, G, D1, D2, L1, L2and n are as defined in claim 1 or 2; and R1, LI and m are as defined in claim 9.
20. A bi-functionalized conjugate, wherein the conjugate is a peptide-oligonucleotide or a peptide-oligopeptide conjugate, according to structure (36):wherein:AA1, AA2, G, D1, F2, L1and n are as defined in claim 1 or 2; and R1, L3and m are as defined in claim 9.61
Citation Information
Patent Citations
Method of characterising a target polypeptide using a nanopore
US20230024319A1
Protein and Peptide Fingerprinting and Sequencing by Nanopore Translocation of Peptide-Oligonucleotide Complexes
US20230039783A1
Protein sequencing via coupling of polymerizable molecules
US20240337660A1
AU2022422583A1