Nucleic acid-polypeptide-nucleic acid ternary complexes and their use in polypeptide nanopore sequencing
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HUADA GENE INST
- Filing Date
- 2023-12-21
- Publication Date
- 2026-06-23
AI Technical Summary
There is a phenomenon of sequencing signal lag in the existing polypeptide nanopore sequencing technology, which leads to the generation of invalid current signals and reduces the capture efficiency of the polypeptide signal.
The nucleic acid-polypeptide-nucleic acid ternary complex is used to extend the nucleic acid sequence in the single-stranded nucleic acid-polypeptide-single-stranded nucleic acid complex to reduce sequencing signal lag and increase the number of effective polypeptide sequencing signals per unit time.
It effectively reduces the phenomenon of electrical signal lag in polypeptide nanopore sequencing, increases the number of effective polypeptide sequencing signals per unit time, and improves the sequencing quality.
Smart Images

Figure CN122270685A_ABST
Abstract
Description
Nucleic acid-peptide-nucleic acid ternary complex and its application in peptide nanopore sequencing Technical Field
[0001] The present invention relates to the technical field of protein sequencing, and in particular to a nucleic acid-polypeptide-nucleic acid ternary complex and its application in polypeptide nanopore sequencing. Background Art
[0002] As one of the most important life science research topics after the Human Genome Project, proteomics can effectively conduct high-throughput detection and analysis of protein expression levels, epigenetic information, and protein-protein interactions in different species, thereby analyzing the structural information, physiological functions, and regulatory mechanisms of various proteins, thereby better revealing the causes of species differences in nature and the impact of various proteins on individual physiological functions at various stages of an individual's life cycle. However, the foundation of proteomics is the reading of the amino acid sequences that make up proteins. Therefore, in order to obtain protein sequence information more accurately and rapidly, the development of reliable protein sequencing technology is particularly important.
[0003] Nanopore-based peptide sequencing, a recently emerging technology, works by passing a protein under electric field through a solid-state nanopore or biological nanopore. Electrodes are used to capture the current flowing through the nanopore, and by analyzing the changes in current, the protein's type and sequence are determined. This technology allows for rapid sequencing of peptide sequences at the single-molecule level.
[0004] Currently, there are mainly the following peptide sequencing technology solutions based on biological nanopores:
[0005] (1) Huang Shuo's laboratory used phi29 DNA polymerase as a motor protein in combination with the MspA nanopore protein to achieve relatively stable control of the perforation rate of specific negatively charged polypeptide chains and reading of electrical signals (Yan, Shuanghong, et al. "Single molecule ratcheting motion of peptides in a Mycobacterium smegmatis Porin A (MspA) nanopore." Nano letters 21.15(2021):6703-6710.).
[0006] (2) Oxford Nanopore Technologies Limited (ONT) patent document (WO2021 / 111125A1) discloses a sequencing scheme based on oligonucleotide control protein rate control. Its design is to first synthesize the adaptor-peptide-dsDNA tail complex, then anneal a portion of the single-stranded DNA to form double-stranded DNA, and then bind it to the oligonucleotide control protein, and then read the electrical signal through the nanopore.
[0007] (3) The Bai Jingwei research group published a protein sequencing scheme based on MTA helicase speed control similar to the above-mentioned ONT patent document. Its design is also to first synthesize a ssDNA-peptide-ssDNA (polyT) complex, bind to the MTA helicase without annealing to form a double strand, and drive the polypeptide fragment through the MspA-M2 nanopore protein through the regular sliding of the MTA helicase on the ssDNA chain to measure the sequencing electrical signal (Chen, Zhijie, et al. "Controlled movement of ssDNA conjugated peptide through Mycobacterium smegmatis porin A (MspA) nanopore by a helicase motor for peptide sequencing application." Chemical science 12.47 (2021): 15750-15756.).
[0008] (4) Dekker and Cees' laboratory also published a method to use Hel308 helicase to control the speed of peptides to achieve repeated reading, thereby reducing the error rate to 10 -6(Brinkerhoff, Henry, et al. "Multiple rereads of single proteins at single–amino acid resolution using nanopores." Science 374.6574(2021):1509-1513.). Similar to the protein sequencing scheme based on MTA helicase speed control proposed by Bai Jingwei's research group, the Dekker research group used Hel308 helicase to achieve speed control of the ssDNA-peptide complex. Because the ssDNA fragments they designed can be captured by multiple Hel308 helicases at the same time, when one of the Hel308 helicases pulls the complex through the nanopore to obtain an electrical signal, as it dissociates from the sequence, the sequencing complex will be pulled back into the pore under the action of the electric field force, and the next molecule of Hel308 helicase will start pulling and lifting the ssDNA-peptide again, thereby achieving the purpose of repeated sequencing.
[0009] However, the above-mentioned prior art does not disclose or report the phenomenon that sequencing signal jamming and invalid current signals occur during the actual polypeptide nanopore sequencing process, and therefore there is currently no effective solution to this problem.
[0010] Summary of the Invention
[0011] The main purpose of the present invention is to provide a nucleic acid-polypeptide-nucleic acid ternary complex and its application in polypeptide nanopore sequencing to improve the sequencing signal jamming that occurs in the existing polypeptide nanopore sequencing process.
[0012] To achieve the above object, according to one aspect of the present invention, a nucleic acid-polypeptide-nucleic acid ternary complex is provided. The ternary complex has the structural formula shown in formula (1): Single-stranded nucleic acid 1-polypeptide-single-stranded nucleic acid 2 Formula (1)
[0013] The single-stranded nucleic acid 1, the polypeptide and the single-stranded nucleic acid 2 are covalently linked in sequence, wherein the single-stranded nucleic acid 1 is used to connect to the sequencing adapter, and the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2 each independently further include an extension chain, which is arranged at any end or in the middle of the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2. The extension chain includes a chain formed by a random arrangement of one or more of the following molecules: nucleosides, nucleotides or organic linkers, and the extension chain located on the single-stranded nucleic acid 2 is not a polymononucleotide sequence.
[0014] Furthermore, when the extended chain contains nucleosides and / or nucleotides, the extended chain is located at the 5' end or 3' end of the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2. Preferably, the length of the extended chain is 5 to 100 bp, preferably 10 to 30 bp, and more preferably 10 to 20 bp; preferably, the nucleosides are selected from modified nucleosides and / or unmodified nucleosides; preferably, the unmodified nucleosides are selected from deoxyribonucleosides and / or ribonucleosides; more preferably, the deoxyribonucleosides are selected from any one or more of A, T, C or G; more preferably Preferably, the ribonucleoside is selected from any one or more of A, U, C or G; preferably, the modified nucleoside is selected from any one or more of the following: 2-aminopurine nucleoside, 5-bromodeoxyuridine, dideoxynucleoside, 5-methylcytosine deoxynucleoside, 5-hydroxymethylcytosine deoxynucleoside, N6-methyladenosine, deoxyinosine or 5-aza-2-deoxycytidine or G-quadruplex; more preferably, the dideoxynucleoside is selected from any one or more of the following: ddA, ddT, ddC or ddG.
[0015] Furthermore, the nucleotides are selected from modified nucleotides and / or unmodified nucleotides, preferably, the unmodified nucleotides are selected from any one or more of the following: DNA or RNA composed of a random base sequence; preferably, the modified nucleotides are selected from any one or more of the following: deoxynucleotides without side chains, dT inverted nucleotides, dG inverted nucleotides, G-quadruplexes or nucleotides with any one or more of the following modifications: phosphorylation, amination, carboxylation, aldehydeation, azidation, alkynylation, acrylamidation, maleamidation, DBCOation, BCNation, sulfhydrylation, dithiolation, biotinylation, desthiobiotinylation, sterylation or fluorescent group; preferably, the fluorescent group is selected from any one of the following: Cy3, Cy5, FAM, Alexa Fluor488 or Texas Red.
[0016] Furthermore, the organic linker is selected from any one or more of the following: Spacer C3, Spacer C6, Spacer 9, Spacer C12 or Spacer 18; preferably, the organic linker is located at the 5' end, 3' end or the middle of the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2, more preferably in the middle.
[0017] Furthermore, when the extended chain consists only of organic linkers, the number of organic linkers is greater than or equal to 2; when the extended chain includes organic linkers in addition to nucleosides and / or nucleotides, the number of organic linkers is at least 1, and the organic linkers are located at any position of the nucleosides and / or nucleotides; more preferably, the organic linkers are located in the middle position of the nucleosides and / or nucleotides.
[0018] Furthermore, the C-terminus or N-terminus of the polypeptide is modified to achieve connection with the single-stranded nucleic acid 2; preferably, the C-terminus or N-terminus of the polypeptide is azide-modified.
[0019] Furthermore, the single-stranded nucleic acid 1, the polypeptide and the single-stranded nucleic acid 2 are covalently linked by any one or more of the following methods: peptide bond linkage, ester bond linkage, ether bond linkage, thiol-maleimide linkage, carbonyl-hydroxylamine compound oxime linkage, carbonyl-hydrazine compound hydrazone linkage, carbonyl-urea structure compound urea linkage, disulfide bond linkage, thioether bond linkage, halogen-nucleophile substitution linkage, 1,3 dipolar cycloaddition reaction linkage, copper-catalyzed azide-alkynyl cycloaddition reaction linkage, ruthenium-catalyzed azide-alkynyl cycloaddition reaction linkage, azide compound-alkynyl compound click chemistry Reaction connection or natural chemical connection; preferably, the azide compound-alkynyl compound click chemistry reaction connection includes azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection, and azide-cyclooctene click chemistry reaction connection.
[0020] Further, the combination of single-stranded nucleic acid 1 and single-stranded nucleic acid 2 is selected from any one of the following groups: 1) DNA11 and DNA12; 2) DNA13 and DNA14; 3) DNA15 and DNA16; wherein, DNA11 is selected from the sequence shown in SEQ ID NO: 11: AAAAAAAAAAAGCTTCTCGTG, wherein the 5' end has a phosphorylation group and the 3' end has a DBCO group; DNA12 is selected from the sequence shown in SEQ ID NO: 12: GCTGTCTTCTGTCGTCGTTTCCTTCTCTGCAAAAAAAAAAA, wherein the 5' end has a maleinamide group; DNA13 is selected from the sequence shown in SEQ ID NO: 13: GCTTCTCGTGAGAGAGGCGG, wherein the 5' end has a phosphorylation group and the 3' end has a DBCO group; DNA14 is selected from the sequence shown in SEQ ID NO: 14: GGCGGAGAGAGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC, wherein the 5' end has a maleinamide group; DNA15 is selected from the sequence shown in SEQ ID The sequence shown in NO:15 is: GCTTCTCGTGGTCGAAAAAGAGAGAGGCGG, wherein the 5' end has a phosphorylation group and the 3' end has a DBCO group; DNA16 is selected from the sequence shown in SEQ ID NO:16: GGCGGAGAGAAAAAAGCTGGGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC, wherein the 5' end has a malein amidation group.
[0021] Furthermore, the connection direction of single-stranded nucleic acid 1 to polypeptide and single-stranded nucleic acid 2 is: 5'-single-stranded nucleic acid 1-3'-N-polypeptide-C-5'-single-stranded nucleic acid 2-3' or 5'-single-stranded nucleic acid 1-3'-C-polypeptide-N-5'-single-stranded nucleic acid 2-3'.
[0022] According to a second aspect of the present application, a polypeptide nanopore sequencing library is provided, comprising any one of the above-mentioned nucleic acid-polypeptide-nucleic acid ternary complexes.
[0023] Furthermore, the polypeptide nanopore sequencing library includes: a) a double-stranded annealing complex; and / or b) a linker complex covalently linked to the double-stranded annealing complex; wherein the double-stranded annealing complex includes: single-stranded nucleic acid 1-polypeptide-single-stranded nucleic acid 2 shown in formula (1), complementary fragment 1 complementary to the single-stranded nucleic acid 1, and complementary fragment 2 complementary to the single-stranded nucleic acid 2; the linker complex includes: linker sequence 1 and linker sequence 2 and a motor protein, linker sequence 1 includes a first segment and a second segment connected in sequence from the 5' end to the 3' end, wherein the first segment is not complementary to the linker sequence 2, and the second segment is complementary to the linker sequence 2, and the motor protein is movably bound to the first segment of the linker sequence 1; wherein the linker sequence 1 is covalently linked to the 5' end of the single-stranded nucleic acid 1; and the linker sequence 2 is covalently linked to the 3' end of the complementary fragment 1 of the single-stranded nucleic acid 1.
[0024] Furthermore, the combination of complementary fragment 1 and complementary fragment 2 is selected from any one of the following groups:
[0025] 1) DNA 17 represented by SEQ ID NO: 17 and DNA 18 represented by SEQ ID NO: 18;
[0026] 2) DNA 19 represented by SEQ ID NO: 19 and DNA 20 represented by SEQ ID NO: 20;
[0027] 3) DNA21 represented by SEQ ID NO: 21 and DNA22 represented by SEQ ID NO: 17;
[0028] 4) DNA23 shown in SEQ ID NO: 23 and DNA24 shown in SEQ ID NO: 24;
[0029] 5) DNA29 represented by SEQ ID NO: 29 and DNA30 represented by SEQ ID NO: 30;
[0030] 6) DNA31 represented by SEQ ID NO: 31 and DNA32 represented by SEQ ID NO: 32;
[0031] SEQ ID NO: 17:TTTTTTTTTTTCACGAGAAGC;
[0032] SEQ ID NO: 18:
[0033] GCAGAGAAGGAAACGACGACAGAAGACAGCTTTTTTTTTTTT;
[0034] SEQ ID NO: 19: CACGAGAAGCTTTTTTTTTTT;
[0035] SEQ ID NO: 20:
[0036] TTTTTTTTTTTGCAGAGAAGGAAACGACGACAGAAGACAGC;
[0037] SEQ ID NO: 21: CCGCCTTCTCCACGAGAAGC;
[0038] SEQ ID NO: 22:
[0039] GCAGAGAAGGAAACGACGACAGAAGACAGCTCTCTCCGCC;
[0040] SEQ ID NO: 23: CCGCCTCTCTCTTTTTCGACCACGAGAAGC;
[0041] SEQ ID NO: 24:
[0042] GCAGAGAAGGAAACGACGACAGAAGACAGCCCAGCTTTTTTCTCTCCGCC;
[0043] SEQ ID NO: 29:AAAAAAAAAAAACACGAGAAGC;
[0044] SEQ ID NO: 30:
[0045] GCAGAGAAGGAAACGACGACAGAAGACAGCAAAAAAAAAAAA;
[0046] SEQ ID NO: 31: CACGAAAAAAAAAAAAAGAAGC;
[0047] SEQ ID NO: 32:
[0048] GCAGAGAAGGAAACGAAAAAAAAAAAAACGACAGAAGACAGC;
[0049] Preferably, the sequence of linker sequence 1 is as shown in SEQ ID NO: 5:
[0050] 5'-XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXTTTTTTTTTTYYYYGGTTGTTTCTGTTGGTGCTGATATTGCT-3', where X=SpC3, Y=iSp18;
[0051] Preferably, the sequence of linker sequence 2 is as shown in SEQ ID NO: 6:
[0052] 5′-phosphorylated-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3′.
[0053] Furthermore, the motor protein is selected from any one of the following: phi29 polymerase, Hel308 helicase, MTA helicase or DDA helicase.
[0054] According to the third aspect of the present application, a polypeptide nanopore sequencing kit is provided, which comprises: single-stranded nucleic acid 1 and single-stranded nucleic acid 2 in any of the aforementioned nucleic acid-polypeptide-nucleic acid ternary complexes, and any one or more of the following optional components: a nanopore and the complementary fragment 1, complementary fragment 2, linker sequence 1, linker sequence 2 and motor protein in the aforementioned polypeptide nanopore sequencing library.
[0055] Furthermore, the linker sequence 1, the linker sequence 2 and the motor protein exist in the form of a linker complex; preferably, the nanopore is a protein nanopore or a solid-state nanopore; preferably, the protein nanopore is selected from the nanopore of any one of the following proteins or their variants: α-hemolysin, Aerolysin, MspA, CsgG or FraC; preferably, the solid-state nanopore is selected from any one of the following: graphene nanopore, gold nanopore, silicon nitride nanopore, silica nanopore or alumina nanopore.
[0056] According to the fourth aspect of the present application, a method for constructing a polypeptide nanopore sequencing library is provided, which construction method comprises: preparing the polypeptide to be tested into any one of the above-mentioned nucleic acid-polypeptide-nucleic acid ternary complexes; annealing the nucleic acid-polypeptide-nucleic acid ternary complex with the complementary fragment 1 of the single-stranded nucleic acid 1 and the complementary fragment 2 of the single-stranded nucleic acid 2 to form a double-stranded annealing complex; and connecting the double-stranded annealing complex with a linker complex containing a motor protein by a ligase to form a polypeptide nanopore sequencing library.
[0057] Furthermore, preparing the polypeptide to be tested into a nucleic acid-polypeptide-nucleic acid ternary complex includes: covalently linking single-stranded nucleic acid 2 to the polypeptide to be tested to obtain a polypeptide-single-stranded nucleic acid 2 complex; covalently linking single-stranded nucleic acid 1 to the polypeptide-single-stranded nucleic acid 2 complex to obtain a nucleic acid-polypeptide-nucleic acid ternary complex; preferably, covalently linking single-stranded nucleic acid 2 with a maleimide group at the 5' end to a polypeptide with an azide group at the N-terminus or C-terminus through a thiol-maleimide addition reaction to obtain a polypeptide-single-stranded nucleic acid 2 complex; and covalently linking single-stranded nucleic acid 1 with a DBCO group at the 3' end to the azide group on the polypeptide through a click chemistry reaction to obtain a nucleic acid-polypeptide-nucleic acid ternary complex.
[0058] Further, the linker complex includes: a linker sequence 1, a linker sequence 2 that is complementary to the 3' end of the linker sequence 1 but not complementary to the 5' end, and a motor protein located on the linker sequence 1; preferably, the sequence of the linker sequence 1 is as shown in SEQ ID NO: 5: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXTTTTTTTTTYYYYGGTTGTTTCTGTTGGTGCTGATATTGCT, wherein X=SpC3, Y=iSp18; preferably, the sequence of the linker sequence 2 is as shown in SEQ ID NO: 6: 5'-phosphorylated-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3'; preferably, the motor protein is selected from any one of the following: phi29 polymerase, Hel308 helicase, MTA helicase or DDA helicase.
[0059] Furthermore, the combination of complementary fragment 1 and complementary fragment 2 is selected from any one of the following groups:
[0060] 1) DNA 17 represented by SEQ ID NO: 17 and DNA 18 represented by SEQ ID NO: 18;
[0061] 2) DNA 19 represented by SEQ ID NO: 19 and DNA 20 represented by SEQ ID NO: 20;
[0062] 3) DNA21 represented by SEQ ID NO: 21 and DNA22 represented by SEQ ID NO: 17;
[0063] 4) DNA23 shown in SEQ ID NO: 23 and DNA24 shown in SEQ ID NO: 24;
[0064] 5) DNA29 represented by SEQ ID NO: 29 and DNA30 represented by SEQ ID NO: 30;
[0065] 6) DNA31 represented by SEQ ID NO: 31 and DNA32 represented by SEQ ID NO: 32;
[0066] SEQ ID NO: 17:TTTTTTTTTTTCACGAGAAGC;
[0067] SEQ ID NO: 18: GCAGAGAAGGAAACGACGACAGAAGACAGCTTTTTTTTTTTT;
[0068] SEQ ID NO: 19: CACGAGAAGCTTTTTTTTTTT;
[0069] SEQ ID NO: 20: TTTTTTTTTTTGCAGAGAAGGAAACGACGACAGAAGACAGC;
[0070] SEQ ID NO: 21: CCGCCTTCTCCACGAGAAGC;
[0071] SEQ ID NO: 22: GCAGAGAAGGAAACGACGACAGAAGACAGCTCTCTCCGCC;
[0072] SEQ ID NO: 23: CCGCCTCTCTCTTTTTCGACCACGAGAAGC;
[0073] SEQ ID NO: 24:
[0074] GCAGAGAAGGAAACGACGACAGAAGACAGCCCAGCTTTTTTCTCTCCGCC;
[0075] SEQ ID NO: 29:AAAAAAAAAAAACACGAGAAGC;
[0076] SEQ ID NO: 30:
[0077] GCAGAGAAGGAAACGACGACAGAAGACAGCAAAAAAAAAAAA;
[0078] SEQ ID NO: 31: CACGAAAAAAAAAAAAAGAAGC;
[0079] SEQ ID NO: 32:
[0080] GCAGAGAAGGAAACGAAAAAAAAAAAAACGACAGAAGACAGC.
[0081] To achieve the above objectives, according to one aspect of the present invention, a polypeptide nanopore sequencing method is provided, comprising: co-incubating a polypeptide nanopore sequencing library constructed by any of the above construction methods with an anchor sequence to obtain an incubation complex; adding the incubation complex to a solution compartment of a sequencing chip, and under the action of an electric field force, controlling the test polypeptide to pass through the nanopore through the reaction of the motor protein unwinding the double-stranded DNA, thereby obtaining an electrical signal corresponding to the test polypeptide; and decoding the electrical signal to obtain an amino acid sequence corresponding to the test polypeptide.
[0082] Furthermore, one end of the anchor sequence is complementary to the end of the linker sequence 2 away from the complementary fragment 1, and the other end carries an anchor group; preferably, the anchor group is selected from any one of lipids, carbon nanotubes, polypeptides, proteins and / or amino acids; preferably, the lipid is selected from any one of fatty acids, sterols, cholesterol, palmitate or tocopherol; preferably, the anchor sequence is as shown in SEQ ID NO:7: 5'-Chol-TEG-TTYYYYTTGACCGCTCGCCTC-3', wherein Y=iSp18, Chol-TEG represents cholesterol-polyethylene glycol.
[0083] Furthermore, the nanopore is a protein nanopore or a solid-state nanopore; preferably, the protein nanopore is selected from the nanopore of any one of the following proteins or their variants: α-hemolysin, Aerolysin, MspA, CsgG or FraC; preferably, the solid-state nanopore is selected from any one of the following: graphene nanopore, gold nanopore, silicon nitride nanopore, silica nanopore or alumina nanopore.
[0084] The present invention reduces the jamming phenomenon in polypeptide nanopore sequencing by extending the nucleic acid sequence in the single-stranded nucleic acid-polypeptide-single-stranded nucleic acid complex, while increasing the number of effective polypeptide sequencing signals per unit time. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0086] FIG1 is a schematic diagram of the structure of OPO synthesized by ligation reaction in one embodiment of the present invention;
[0087] FIG2 is a schematic diagram of the library structure through annealing and ligase coupling in one embodiment of the present invention;
[0088] FIG3 is a schematic diagram of a protein sequencing process according to an embodiment of the present invention;
[0089] Figure 4 is a schematic diagram of three main signals detected during protein sequencing in one embodiment of the present invention, wherein (a) is a protein sequencing signal, (b) is a protein sequencing freeze signal, and (c) is a platform freeze signal;
[0090] Figure 5 is a distribution diagram of nanopore sequencing signal types of four OPOs in one embodiment of the present invention, wherein (a) corresponds to DNA1-Peptide1-DNA2; (b) corresponds to DNA1-Peptide2-DNA2; (c) corresponds to DNA8-Peptide1-DNA2; and (d) corresponds to DNA8-Peptide1-DNA2.
[0091] Figure 6 is a distribution diagram of nanopore sequencing signal types of six optimized OPOs in one embodiment of the present invention, including (a) D9P1D10; (b) D11P1D12; (c) D13P1D14; (d) D15P1D16; (f) D25P1D26; and (g) D27P1D28. DETAILED DESCRIPTION
[0092] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the embodiments.
[0093] As mentioned in the background art, the prior art discloses a variety of nanopore polypeptide sequencing technology solutions. However, the inventors of this application found that: in all the above-mentioned published solutions, only the effective sequencing current fingerprints containing polypeptide signals are displayed, and there is no mention of whether there are other types of invalid signals, and the actual effective signal detection efficiency, that is, the number of effective sequencing current maps per unit time. In the actual nanopore sequencing process, when non-DNA structures pass through the motor protein, there is a certain probability that the sequencing signal will be stuck, resulting in the generation of invalid current signals. The increase in the number of invalid signals will significantly reduce the capture efficiency of polypeptide signals. At the same time, none of the above solutions provide a feasible technical solution to reduce the number of invalid signals and improve the efficiency of polypeptide signal capture.
[0094] Therefore, in order to improve this situation, the inventors of the present application provide a nanopore polypeptide sequencing technology based on motor protein speed control. More specifically, the inventors of the present application provide a library design method for quickly reading polypeptide signals. By designing different connection structures, the jamming phenomenon during sequencing is significantly reduced, the quality of polypeptide sequencing is improved, and the number of effective polypeptide sequencing signals per unit time is increased.
[0095] To reduce the large number of invalid stutter signals observed during nanopore sequencing, the present invention proposes to increase the number of valid polypeptide sequencing signals per unit time by extending the nucleic acid sequence in the single-stranded nucleic acid-polypeptide-single-stranded nucleic acid complex. The specific implementation steps are as follows:
[0096] (1) Construction of OPO: A ternary complex structure of "nucleic acid-polypeptide-nucleic acid" (OPO) was constructed through a thiol-maleimide addition reaction and an azide-DBCO click chemistry reaction, and the pure product was purified by HPLC to be used as a substrate for peptide nanopore sequencing. The structure and synthesis process of OPO are shown in Figure 1, where DNA1 (DNA8) is a DNA fragment connected to the linker and serves to guide the peptide sequence into the pore under the action of the electric field; DNA2 is the DNA fragment to be unwound, which is used to unwind the DNA double helix under the control of the motor protein, thereby controlling the speed of peptide perforation.
[0097] (2) Construction of sequencing library: The schematic diagram of the sequencing library structure is shown in Figure 2. Specifically, the "nucleic acid-polypeptide-nucleic acid" complex described in (1) is annealed with DNA1 (DNA8) and DNA2's complementary fragments DNA3 and DNA4 to form a double-stranded structure, and then the double-stranded annealed complex is connected to the prefabricated linker complex (a complex formed by DNA5, DNA6 and motor protein) by ligase to form a polypeptide sequencing library.
[0098] (3) Sequencing: The schematic diagram of the sequencing process is shown in Figure 3. Specifically, the library described in (2) is co-incubated with DNA7 and then added to the solution chamber of the sequencing chip. Under the action of the electric field force, the DNA double strands are continuously unwound to stably control the polypeptide to pass through the nanopore and obtain the electrical signal.
[0099] By counting the number of occurrences of perforation signal types, the effects of different nucleic acid sequences on the number of effective nanopore electrical signals in the "nucleic acid-polypeptide-nucleic acid" library were compared, proving that the scheme proposed in the present invention is effective in increasing the number of effective polypeptide sequencing signals per unit time.
[0100] Based on the research results, the inventors of this application have proposed a series of protection schemes for this application. In a first typical embodiment, a nucleic acid-polypeptide-nucleic acid ternary complex is provided, having the structural formula shown in formula (1):
[0101] Single-stranded nucleic acid 1-polypeptide-single-stranded nucleic acid 2 formula (1)
[0102] The single-stranded nucleic acid 1, the polypeptide and the single-stranded nucleic acid 2 are covalently linked in sequence, wherein the single-stranded nucleic acid 1 is used to connect to the sequencing adapter, and the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2 each independently further include an extension chain, which is arranged at any end or in the middle of the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2. The extension chain includes a chain formed by a random arrangement of one or more of the following molecules: nucleosides, nucleotides or organic linkers, and the extension chain is not a polymononucleotide sequence when located on the single-stranded nucleic acid 2.
[0103] As previously mentioned, this application has demonstrated through a series of experiments that extending the length of the single-stranded nucleic acid sequence at at least one end of a polypeptide can help reduce the stuttering phenomenon that occurs during polypeptide nanopore sequencing. This also increases the number of effective polypeptide sequencing signals per unit time.
[0104] In the nucleic acid-polypeptide-nucleic acid ternary complex of the present application, the specific composition of the extended chain can be a chain formed by randomly arranging at least one of the above-mentioned nucleosides, nucleotides and organic linkers. Including but not limited to deoxyribonucleosides (A / T / G / C), ribonucleosides (A / U / G / C), DNA / RNA composed of random base sequences, abaciated spacer / dspacer, organic linkers (such as 5'C3Spacer, 5'C6Spacer, 5'Spacer 9, 5'Spacer 18, 3'C3Spacer, 3'C6Spacer, 3'Spacer 9, 3'Spacer 18, int Spacer C3, int Spacer 18, etc.), modified nucleic acids (e.g., phosphorylated, aminated, carboxylated, aldehyded, azidated, alkynylated, acrylamidated, maleamidated, DBCO-allylated, BCN-allylated, sulfhydrylated, dithiolated, biotinylated, dethiobiotinylated, sterylated, fluorescently modified nucleosides, etc.), special nucleotides (i.e., nucleotides containing modified bases, modified ribose, or modified deoxyribose, such as 2-aminopurine nucleoside, 5-bromodeoxyuridine nucleoside, dideoxynucleosides (ddA, ddT, ddC, ddG), 5-methylcytosine deoxynucleoside, 5-hydroxymethylcytosine deoxynucleoside, N6-methyladenosine, deoxyinosine, 5-aza-2-deoxycytidine, etc.), dT / dG inverted nucleotides, G-quadruplexes, etc.
[0105] The length of the extended chain, measured in nucleosides, can be 5 to 100 bp. Preferably, it is 10 to 30 bp, and more preferably, it is 10 to 20 bp. The position of the extended chain can be designed to be located at any position of the single-stranded nucleic acid sequence at either end, including the beginning, any position in the middle, and the end of the sequence. When the extended chain contains nucleosides and / or nucleotides, it is preferably located at the 5' end or the 3' end of the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2.
[0106] It should be noted that the above-mentioned nucleosides are selected from modified nucleosides and / or unmodified nucleosides; preferably, the unmodified nucleosides are selected from deoxyribonucleosides and / or ribonucleosides; more preferably, the deoxyribonucleosides are selected from any one or more of A, T, C or G; more preferably, the ribonucleosides are selected from any one or more of A, U, C or G. Preferably, the modified nucleosides are selected from any one or more of the following: 2-aminopurine nucleoside, 5-bromodeoxyuridine, dideoxynucleoside, 5-methylcytosine deoxynucleoside, 5-hydroxymethylcytosine deoxynucleoside, N6-methyladenosine, deoxyinosine or 5-aza-2-deoxycytidine or G-quadruplex; more preferably, the dideoxynucleoside is selected from any one or more of the following: ddA, ddT, ddC or ddG.
[0107] Similarly, the above-mentioned nucleotides are selected from modified nucleotides and / or unmodified nucleotides, preferably, the unmodified nucleotides are selected from any one or more of the following: DNA or RNA composed of a random base sequence; preferably, the modified nucleotides are selected from any one or more of the following: deoxynucleotides without side chains, dT inverted nucleotides, dG inverted nucleotides, G-quadruplexes or nucleotides with any one or more of the following modifications: phosphorylation, amination, carboxylation, aldehydeation, azidation, alkynylation, acrylamidation, maleamidation, DBCOation, BCNation, sulfhydrylation, dithiolation, biotinylation, desthiobiotinylation, sterylation or fluorescent group; preferably, the fluorescent group (most of which is modified on the base of the nucleotide) is selected from any one of the following: Cy3, Cy5, FAM, Alexa Fluor 488 or Texas Red.
[0108] In some preferred embodiments of the present application, the organic linker is selected from any one or more of the following: Spacer C3, Spacer C6, Spacer 9, Spacer C12 or Spacer 18.
[0109] Spacer refers to an inter-arm modification, and most Spacers are composed of straight carbon chains or ethylene glycol. Spacers are introduced into oligonucleotides, usually to establish a distance between oligonucleotides or between oligonucleotides and other functional groups, in order to avoid steric hindrance, reduce adverse interactions between groups, increase flexibility, etc. Different Spacers have different numbers of atoms, and the required spatial distance can be achieved by adjusting the number and type of inserted Spacers. Common Spacers include hydrophobic Spacer C3, C6, and C12 (C3 refers to 3 CH2, C6 refers to 6 CH2, and C12 refers to 12 CH2), and hydrophilic Spacer 9 and Spacer18 (Spacer 9 is a linker composed of 3 consecutive ethylene glycols, and Spacer 18 is a linker composed of 6 consecutive ethylene glycols). These spacers can be connected to any position of the single-stranded nucleic acid 1 or 2, such as the 5' end, 3' end or any position in the middle, such as 5' end Spacer C3, 5' end Spacer C6, 5' end Spacer 9, 5' end Spacer 18, 3' end Spacer C3, 3' end Spacer C6, 3' end Spacer 9, 3' end Spacer 18, int Spacer C3 (i.e., spacer C3 in the middle position) or int Spacer18 (i.e., spacer 18 in the middle position).
[0110] In some embodiments, the extended chain is solely composed of an organic linker, in which case the number of organic linkers is greater than or equal to 2. In other embodiments, when the extended chain comprises nucleosides and / or nucleotides and further comprises an organic linker, the number of organic linkers is at least 1, and the organic linker is located at any position of the nucleosides and / or nucleotides; more preferably, the organic linker is located at the middle position of the nucleosides and / or nucleotides.
[0111] In the aforementioned nucleic acid-polypeptide-nucleic acid ternary complex, the polypeptide is the polypeptide to be sequenced, derived from the sample to be sequenced. The C-terminus or N-terminus of the polypeptide is modified to facilitate attachment to the single-stranded nucleic acid 2. Preferably, the C-terminus or N-terminus of the polypeptide is azide-modified, with the amino group replaced by an azide group, such as azide-modified lysine (an unnatural amino acid), having the following structural formula:
[0112] In the above-mentioned nucleic acid-polypeptide-nucleic acid ternary complex, the covalent connection mode between the single-stranded nucleic acid 1, the polypeptide and the single-stranded nucleic acid 2 includes but is not limited to any one or more of the following: peptide bond connection, ester bond connection, ether bond connection, thiol-maleimide connection, carbonyl-hydroxylamine compound oxime connection, carbonyl-hydrazine compound hydrazone connection, carbonyl-urea structure compound urea connection, disulfide bond connection, thioether bond connection, halogen-nucleophile substitution connection, 1,3 dipolar cycloaddition reaction connection (Huisgen cycloaddition), copper-catalyzed azide-alkynyl cycloaddition reaction connection, ruthenium-catalyzed azide-alkynyl cycloaddition reaction connection, azide Nitrogen compound-alkynyl compound click chemistry reaction connection or natural chemical connection; preferably, azide compound-alkynyl compound click chemistry reaction connection includes azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection, azide-cyclooctene click chemistry reaction connection. The schematic diagram of the above molecular structure is as follows:
[0113] In some preferred embodiments, in the nucleic acid-polypeptide-nucleic acid ternary complex of the present application, the combination of single-stranded nucleic acid 1 and single-stranded nucleic acid 2 is selected from any one of the following groups: 1) DNA11 and DNA12; 2) DNA13 and DNA14; 3) DNA15 and DNA16; wherein, DNA11 is selected from the sequence AAAAAAAAAAAGCTTCTCGTG shown in SEQ ID NO: 11, wherein the 5' and 3' ends are modified by groups, the 5' end modification group is a phosphorylation group, and the 3' end modification group is a DBCO group; DNA12 is selected from the sequence GCTGTCTTCTGTCGTCGTTTCCTTCTCTGCAAAAAAAAAAA shown in SEQ ID NO: 12, wherein the 5' end is modified by groups, the 5' end modification group is a malein amidation group; DNA13 is selected from the sequence GCTTCTCGTGAGAGAGGCGG shown in SEQ ID NO: 13, wherein the 5' and 3' ends are modified by groups, the 5' end modification group is a phosphorylation group, and the 3' end modification group is a DBCO group; DNA14 is selected from the sequence SEQ ID The sequence shown in NO:14 is GGCGGAGAGAGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC, wherein the 5' end carries a modification group, preferably, the modification group at the 5' end is a maleimide group; DNA15 is selected from the sequence shown in SEQ ID NO:15: GCTTCTCGTGGTCGAAAAAGAGAGAGGCGG, wherein the 5' end and 3' end carry modification groups, the modification group at the 5' end is a phosphorylation group, and the modification group at the 3' end is a DBCO group; DNA16 is selected from the sequence shown in SEQ ID NO:16: GGCGGAGAGAAAAAAGCTGGGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC, wherein the 5' end carries a modification group, and the modification group at the 5' end is a maleimide group. In the above combination, the 5' end of the single-stranded nucleic acid 1 carries a phosphorylation group for subsequent connection to the linker, while the DBCO group at the 3' end and the maleimide group at the 5' end of the single-stranded nucleic acid 2 are both for achieving covalent coupling with the intermediate polypeptide.
[0114] In the nucleic acid-polypeptide-nucleic acid ternary complex of the present application, the orientation of the polypeptide in the two single-stranded nucleic acids is not particularly limited, and can be either 5'-single-stranded nucleic acid 1-3'-N-polypeptide-C-5'-single-stranded nucleic acid 2-3' (for example, the N-terminus of the polypeptide is cysteine C, and the C-terminus is LYS(N3), i.e., azide-modified lysine), or 5'-single-stranded nucleic acid 1-3'-C-polypeptide-N-5'-single-stranded nucleic acid 2-3' (for example, the N-terminus of the polypeptide is LYS(N3), i.e., azide-modified lysine, and the C-terminus is cysteine C). It should also be noted here that the amino acids at the junctions between the two ends of the polypeptide sequence of the present application and the single-stranded nucleic acid can be designed as any amino acid, including modified amino acids.
[0115] In a second exemplary embodiment of the present application, a polypeptide nanopore sequencing library is provided, comprising any of the aforementioned nucleic acid-polypeptide-nucleic acid ternary complexes. A polypeptide nanopore library containing the aforementioned nucleic acid-polypeptide-nucleic acid ternary complexes can reduce sequencing lags and increase the number of effective polypeptide sequencing signals per unit time during sequencing.
[0116] In some preferred embodiments, the polypeptide nanopore sequencing library comprises: a) a double-stranded annealing complex; and / or b) a linker complex covalently linked to the double-stranded annealing complex; wherein the double-stranded annealing complex comprises: single-stranded nucleic acid 1-polypeptide-single-stranded nucleic acid 2 shown in the aforementioned formula (1), complementary fragment 1 complementary to the single-stranded nucleic acid 1, and complementary fragment 2 complementary to the single-stranded nucleic acid 2; the linker complex comprises: linker sequence 1 and linker sequence 2 and a motor protein, linker sequence 1 comprises a first segment and a second segment connected sequentially from the 5' end to the 3' end, wherein the first segment is not complementary to the linker sequence 2, and the second segment is complementary to the linker sequence 2, and the motor protein is movably bound to the first segment of the linker sequence 1; wherein the linker sequence 1 is covalently linked to the 5' end of the single-stranded nucleic acid 1; and the linker sequence 2 is covalently linked to the 3' end of the complementary fragment 1 of the single-stranded nucleic acid 1.
[0117] In some more specific embodiments, the combination of complementary fragment 1 and complementary fragment 2 is selected from any one of the following groups:
[0118] 1) DNA 17 represented by SEQ ID NO: 17 and DNA 18 represented by SEQ ID NO: 18;
[0119] 2) DNA 19 represented by SEQ ID NO: 19 and DNA 20 represented by SEQ ID NO: 20;
[0120] 3) DNA21 represented by SEQ ID NO: 21 and DNA22 represented by SEQ ID NO: 17;
[0121] 4) DNA23 shown in SEQ ID NO: 23 and DNA24 shown in SEQ ID NO: 24;
[0122] 5) DNA29 represented by SEQ ID NO: 29 and DNA30 represented by SEQ ID NO: 30;
[0123] 6) DNA31 represented by SEQ ID NO: 31 and DNA32 represented by SEQ ID NO: 32;
[0124] SEQ ID NO: 17:TTTTTTTTTTTCACGAGAAGC;
[0125] SEQ ID NO: 18:
[0126] GCAGAGAAGGAAACGACGACAGAAGACAGCTTTTTTTTTTTT;
[0127] SEQ ID NO: 19: CACGAGAAGCTTTTTTTTTTT;
[0128] SEQ ID NO: 20:
[0129] TTTTTTTTTTTGCAGAGAAGGAAACGACGACAGAAGACAGC;
[0130] SEQ ID NO: 21: CCGCCTTCTCCACGAGAAGC;
[0131] SEQ ID NO: 22:
[0132] GCAGAGAAGGAAACGACGACAGAAGACAGCTCTCTCCGCC;
[0133] SEQ ID NO: 23: CCGCCTCTCTCTTTTTCGACCACGAGAAGC;
[0134] SEQ ID NO: 24:
[0135] GCAGAGAAGGAAACGACGACAGAAGACAGCCCAGCTTTTTTCTCTCCGCC;
[0136] SEQ ID NO: 29:AAAAAAAAAAAACACGAGAAGC;
[0137] SEQ ID NO: 30:
[0138] GCAGAGAAGGAAACGACGACAGAAGACAGCAAAAAAAAAAAA;
[0139] SEQ ID NO: 31: CACGAAAAAAAAAAAAAGAAGC;
[0140] SEQ ID NO: 32:
[0141] GCAGAGAAGGAAACGAAAAAAAAAAAAACGACAGAAGACAGC;
[0142] Preferably, the sequence of linker sequence 1 is as shown in SEQ ID NO: 5:
[0143] 5'-XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXTTTTTTTTTTYYYYGGTTGTTTCTGTTGGTGCTGATATTGCT-3', where X=SpC3, Y=iSp18;
[0144] Preferably, the sequence of linker sequence 2 is as shown in SEQ ID NO: 6:
[0145] 5′-phosphorylated-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3′.
[0146] In some embodiments of the present application, the applicable motor protein is selected from any one of the following: phi29 polymerase, Hel308 helicase, MTA helicase or DDA helicase.
[0147] In a third typical embodiment of the present application, a polypeptide nanopore sequencing kit is provided, which comprises: single-stranded nucleic acid 1 and single-stranded nucleic acid 2 in any of the aforementioned nucleic acid-polypeptide-nucleic acid ternary complexes, and any one or more of the following optional components (as needed): a nanopore and complementary fragment 1, complementary fragment 2, linker sequence 1, linker sequence 2 and motor protein in the aforementioned polypeptide nanopore sequencing library.
[0148] In some embodiments, the adapter sequence 1, adapter sequence 2, and motor protein in the above kit exist in the form of an adapter complex.
[0149] In the above kit, the nanopore can be any of the following protein nanopores or solid-state nanopores; wherein the protein nanopore is selected from the nanopore of any of the following proteins or their variants: α-hemolysin, Aerolysin, MspA, CsgG or FraC; the solid-state nanopore is selected from any of the following: graphene nanopore, gold nanopore, silicon nitride nanopore, silica nanopore or alumina nanopore.
[0150] In a fourth typical embodiment of the present application, a method for constructing a polypeptide nanopore sequencing library is provided, which construction method includes: preparing the polypeptide to be tested into the aforementioned nucleic acid-polypeptide-nucleic acid ternary complex; annealing the nucleic acid-polypeptide-nucleic acid ternary complex with the complementary fragment 1 of the single-stranded nucleic acid 1 and the complementary fragment 2 of the single-stranded nucleic acid 2 to form a double-stranded annealing complex; and connecting the double-stranded annealing complex with a linker complex containing a motor protein through a ligase to form a polypeptide sequencing library.
[0151] This construction method adds a step of forming a nucleic acid-polypeptide-nucleic acid ternary complex containing an extended chain of the present application. On this basis, the polypeptide nanopore sequencing library of the present application can be obtained by adopting steps similar to the prior art of annealing the corresponding complementary fragments to form a double-stranded complex and connecting it with the adapter complex.
[0152] In some preferred embodiments, preparing the test polypeptide into a nucleic acid-polypeptide-nucleic acid ternary complex comprises: covalently linking single-stranded nucleic acid 2 to the test polypeptide to form a polypeptide-single-stranded nucleic acid 2 complex; and covalently linking single-stranded nucleic acid 1 to the polypeptide-single-stranded nucleic acid 2 complex to form a nucleic acid-polypeptide-nucleic acid ternary complex. This method, by first covalently linking the test polypeptide to single-stranded nucleic acid 2 and then to single-stranded nucleic acid 1, offers advantages over other linking methods, such as directly covalently linking nucleic acid 1, nucleic acid 2, and polypeptide, resulting in higher reaction efficiency, higher conversion rate of double-end specific modification, and easier purification.
[0153] The specific covalent connection method can be any of the following methods: peptide bond connection, ester bond connection, ether bond connection, thiol-maleimide connection, carbonyl-hydroxylamine compound oxime connection, carbonyl-hydrazine compound hydrazone connection, carbonyl-urea structure compound urea connection, disulfide bond connection, thioether bond connection, halogen-nucleophile substitution connection, 1,3 dipolar cycloaddition connection (Huisgen cycloaddition), copper-catalyzed azide-alkynyl cycloaddition connection, ruthenium-catalyzed azide-alkynyl cycloaddition connection, azide compound-alkynyl compound click chemistry reaction connection or natural chemical connection; preferably, the azide compound-alkynyl compound click chemistry reaction connection includes azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection, azide-cyclooctene click chemistry reaction connection.
[0154] In some preferred embodiments, the steps for preparing the nucleic acid-polypeptide-nucleic acid ternary complex, as shown in Figure 1, include: first, covalently linking a single-stranded nucleic acid 2 with a maleimide group at the 5' end to a polypeptide with an azide group at the N-terminus or C-terminus through a thiol-maleimide addition reaction to obtain a polypeptide-single-stranded nucleic acid 2 complex; then, covalently linking a single-stranded nucleic acid 1 with a DBCO group at the 3' end to the azide group on the polypeptide through a click chemistry reaction to obtain a nucleic acid-polypeptide-nucleic acid ternary complex.
[0155] In the above construction method, the linker complex includes: a linker sequence 1, a linker sequence 2 that is complementary to the 3' end of the linker sequence 1 but not complementary to the 5' end, and a motor protein located on a sequence on the linker sequence 1 that is not complementary to the linker sequence 2; preferably, the sequence of the linker sequence 1 is as shown in SEQ ID NO: 5 (i.e., DNA5 in the present application): XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXTTTTTTTTYYYYGGTTGTTTCTGTTGGTGCTGATATTGCT, wherein X=SpC3, Y=iSp18; preferably, the sequence of the linker sequence 2 is as shown in SEQ ID NO: 6 (i.e., DNA6 in the present application): 5'-phosphorylated-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3'; preferably, the motor protein is selected from any one of the following: phi29 polymerase, Hel308 helicase, MTA helicase or DDA helicase.
[0156] In the preferred embodiment described above, the motor protein's binding site is located on a 10-nt T sequence preceding the "YYYY" region in the DNA5 sequence (SEQ ID NO: 5). Prior to sequencing, the motor protein must bind before the complementary double-stranded region; otherwise, it will automatically unwind the double strands, preventing the ability to initiate sequencing after entering the nanopore.
[0157] In some more preferred embodiments, the combination of the complementary fragment 1 and the complementary fragment 2 is selected from any one of the following groups:
[0158] 1) DNA 17 represented by SEQ ID NO: 17 and DNA 18 represented by SEQ ID NO: 18;
[0159] 2) DNA 19 represented by SEQ ID NO: 19 and DNA 20 represented by SEQ ID NO: 20;
[0160] 3) DNA21 represented by SEQ ID NO: 21 and DNA22 represented by SEQ ID NO: 17;
[0161] 4) DNA23 shown in SEQ ID NO: 23 and DNA24 shown in SEQ ID NO: 24;
[0162] 5) DNA29 represented by SEQ ID NO: 29 and DNA30 represented by SEQ ID NO: 30;
[0163] 6) DNA31 represented by SEQ ID NO: 31 and DNA32 represented by SEQ ID NO: 32;
[0164] SEQ ID NO: 17:TTTTTTTTTTTCACGAGAAGC;
[0165] SEQ ID NO: 18: GCAGAGAAGGAAACGACGACAGAAGACAGCTTTTTTTTTTTT;
[0166] SEQ ID NO: 19: CACGAGAAGCTTTTTTTTTTT;
[0167] SEQ ID NO: 20: TTTTTTTTTTTGCAGAGAAGGAAACGACGACAGAAGACAGC;
[0168] SEQ ID NO: 21: CCGCCTTCTCCACGAGAAGC;
[0169] SEQ ID NO: 22: GCAGAGAAGGAAACGACGACAGAAGACAGCTCTCTCCGCC;
[0170] SEQ ID NO: 23: CCGCCTCTCTCTTTTTCGACCACGAGAAGC;
[0171] SEQ ID NO: 24:
[0172] GCAGAGAAGGAAACGACGACAGAAGACAGCCCAGCTTTTTTCTCTCCGCC;
[0173] DNA29 (SEQ ID NO: 29):AAAAAAAAAAAACACGAGAAGC;
[0174] DNA30 (SEQ ID NO: 30):
[0175] GCAGAGAAGGAAACGACGACAGAAGACAGCAAAAAAAAAAAA;
[0176] DNA31 (SEQ ID NO: 31): CACGAAAAAAAAAAAAAGAAGC;
[0177] DNA32 (SEQ ID NO: 32):
[0178] GCAGAGAAGGAAACGAAAAAAAAAAAAACGACAGAAGACAGC.
[0179] As demonstrated in the examples of the present application, the above combination can be used in conjunction with the preferred combination of single-stranded nucleic acid 1 and single-stranded nucleic acid 2 of the present application to significantly reduce the jamming phenomenon that occurs in polypeptide nanopore sequencing.
[0180] In a fifth typical embodiment of the present application, a polypeptide nanopore sequencing method is provided, which includes: co-incubating the polypeptide nanopore sequencing library constructed by the aforementioned construction method with an anchor sequence to obtain an incubation complex; adding the incubation complex to a solution compartment of a sequencing chip, and under the action of an electric field force, controlling the test polypeptide to pass through the nanopore through the reaction of the motor protein unwinding the double-stranded DNA, thereby obtaining an electrical signal corresponding to the test polypeptide; and decoding the electrical signal to obtain an amino acid sequence corresponding to the test polypeptide.
[0181] The above-mentioned step of co-incubating the library with the anchor sequence is to anchor the library near the nanopore protein to facilitate the subsequent polypeptide to pass through the nanopore for sequencing. Therefore, the anchor sequence needs to be anchored to the linker sequence at one end and anchored to the nanopore membrane structure (such as the phospholipid bilayer) at the other end. In a preferred embodiment of the present application, one end of the anchor sequence is complementary to the end of the linker sequence 2 away from the complementary fragment 1, and the other end carries an anchor group. Preferably, the anchor group is selected from any one of lipids, carbon nanotubes, polypeptides, proteins and / or amino acids; preferably, the lipid is selected from any one of fatty acids, sterols, cholesterol, palmitate or tocopherol; preferably, the anchor sequence is as shown in SEQ ID NO: 7 5'-Chol-TEG-TTYYYYTTGACCGCTCGCCTC-3', wherein Y = iSp18, and Chol-TEG represents cholesterol-polyethylene glycol.
[0182] In the nanopore sequencing method of the present application, any of the following nanopore protein sequencing technologies can be used: α-hemolysin nanopore, Aerolysin nanopore, MspA nanopore, CsgG nanopore, FraC nanopore, phi29 nanopore, corresponding variants and mutations of the above nanopores, and other solid-state nanopore structures (such as graphene nanopore, gold nanopore, silicon nitride nanopore, silica nanopore, alumina nanopore, etc.).
[0183] The sequencing method of the present invention reduces the electrical signal jamming phenomenon that occurs during polypeptide nanopore sequencing by extending and optimizing the sequence structure of the nucleic acid-polypeptide-nucleic acid complex, increases the number of effective polypeptide sequencing signals per unit time, and achieves the purpose of improving the effective polypeptide nanopore electrical signal capture rate.
[0184] The beneficial effects of the present application will be further explained in detail below with reference to specific embodiments.
[0185] The DNA used in the following examples was synthesized by BGI Liuhe (sequences are expressed in the 5' end → 3' end direction), and the polypeptides were synthesized by GenScript (sequences are expressed in the N-terminus → C-terminus direction). The sequence information is as follows:
[0186] DNA1 (SEQ ID NO: 1):Phosphorylation-GCTTCTCGTG-DBCO;
[0187] DNA2 (SEQ ID NO: 2): Maleimide-GCTGTCTTCTGTCGTCGTTTCCTTCTCTGC;
[0188] DNA3(SEQ ID NO:3):CACGAGAAGCA;
[0189] DNA4(SEQ ID NO:4):GCAGAGAAGGAACGACGAC;
[0190] DNA5(SEQ ID NO:5):
[0191] XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXTTTTTTTTTTTYYYYGGTTGTTTCTGTTGGTGCTGATATTGCT(XSpC3SYSiSp18)
[0192] DNA6(SEQ ID NO:6):
[0193] Phosphorylation-GCAATATCAGCACCAACAAACAACCTTTGAGGCGAGCGGTCAA.
[0194] DNA7(SEQ ID NO:7):Chol-TEG / TTYYYYTTGACCGCTCCGCCTC(Y / iSp18),Chol-TEG cleavage-protein fragment.
[0195] DNA8(SEQ ID NO:8):Phosphorylation-GCTTCTCGTG-BCN1
[0196] DNA9(SEQ ID NO:9):Phosphorylation-GCTTCTCGTGAAAAAAAAAAA-DBCO;
[0197] DNA10(SEQ ID NO:10):
[0198] Maleimide-AAAAAAAAAAAGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC;
[0199] DNA11(SEQ ID NO:11):Phosphorylation-AAAAAAAAAAAGCTTCTCGTG-DBCO;
[0200] DNA12(SEQ ID NO:12):
[0201] Maleimide-GCTGTCTTCTGTCGTCGTTTCCTTCTCTGCAAAAAAAAAAAA;
[0202] DNA13(SEQ ID NO:13):Phosphorylation-GCTTCTCGTGAGAGAGGCGG-DBCO;
[0203] DNA14(SEQ ID NO:14):
[0204] Maleimide-GGCGGAGAGAGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC;
[0205] DNA15(SEQ ID NO:15):
[0206] Phosphorylation-GCTTCTCGTGGTCGAAAAAGAGAGAGGCGG-DBCO;
[0207] DNA16(SEQ ID NO:16):
[0208] Maleimide-GGCGGAGAGAAAAAAGCTGGGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC;
[0209] DNA17(SEQ ID NO:17):TTTTTTTTTTTCACGAGAAGC;
[0210] DNA18(SEQ ID NO:18):
[0211] GCAGAGAAGGAAACGACGACAGAAGACAGCTTTTTTTTTTT;
[0212] DNA19(SEQ ID NO:19):CACGAGAAGCTTTTTTTTTTT;
[0213] DNA20(SEQ ID NO:20):
[0214] TTTTTTTTTTTGCAGAGAAGGAAACGACGACAGAAGACAGC;
[0215] DNA21(SEQ ID NO:21):CCGCCTCTCTCACGAGAAGC;
[0216] DNA22(SEQ ID NO:22):
[0217] GCAGAGAAGGAAACGACGACAGAAGACAGCTCTCTCCGCC;
[0218] DNA23(SEQ ID NO:23):CCGCCTCTCTCTTTTTCGACCACGAGAAGC;
[0219] DNA24(SEQ ID NO:24):
[0220] GCAGAGAAGGAAACGACGACAGAAGACAGCCCAGCTTTTTTCTCTCCGCC;
[0221] DNA25(SEQ ID NO:25):Phosphorylation-GCTTCTCGTGYYYY-DBCO,Y=iSp 18;
[0222] DNA26(SEQ ID NO:26):
[0223] Maleimide-YYYYGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC,Y=iSp 18;
[0224] DNA27(SEQ ID NO:27):Phosphorylation-GCTTCYYYYTCGTG-DBCO,Y=iSp 18;
[0225] DNA28(SEQ ID NO:28):
[0226] Maleimide-GCTGTCTTCTGTCGTYYYYCGTTTCCTTCTCTGC,Y=iSp 18;
[0227] DNA29(SEQ ID NO:29):AAAAAAAAAAAACACGAGAAGC;
[0228] DNA30(SEQ ID NO:30):
[0229] GCAGAGAAGGAAACGACGACAGAAGACAGCAAAAAAAAAAAA;
[0230] DNA31(SEQ ID NO:31):CACGAAAAAAAAAAAAAGAAGC;
[0231] DNA32(SEQ ID NO:32):
[0232] GCAGAGAAGGAAACGAAAAAAAAAAAAACGACAGAAGACAGC;
[0233] Peptide1 (SEQ ID NO: 33):CGSGDDGSG{LYS(N3)};
[0234] Peptide2 (SEQ ID NO: 34):CGSGRRRGSG{LYS(N3)}.
[0235] The meanings of the modifications present in the above sequences are summarized as follows:
[0236] DBCO: dibenzocyclooctyne, used in copper-free azide-alkyne cycloaddition (SPAAC) reactions.
[0237] SpC3: Spacer C3, the sequence contains Spacer C3 located in the middle position and the 5' end position.
[0238] iSp18: Int Spacer 18, indicating Spacer 18 located in the middle position.
[0239] Chol-TEG: stands for cholesterol-polyethylene glycol.
[0240] Phosphorylation: phosphorylation.
[0241] BCN: A commonly used triple-bond compound used for click chemistry reactions with N3. It also has a carbon chain structure connected to the 3' end of DNA. The specific structural formula is as follows:
[0242] Maleimide: Maleimide modification refers to the modification of nucleic acids with maleimide functional groups. There is also a carbon chain structure connected between the maleimide group and the 5' end of DNA. The specific structural formula is as follows:
[0243] LYS(N3): Also known as Lys(N3), azide-modified lysine is an unnatural amino acid with an azide group replacing the original amino group. Its structural formula is as follows:
[0244] Example 1: Construction of Nucleic Acid-Polypeptide-Nucleic Acid Complex (OPO Ligation Product)
[0245] 1. Experimental Methods
[0246] (1) Two oligo-peptide (OP) ligation products were prepared: (a) Peptide 1-DNA 2 and (b) Peptide 2-DNA 2. DNA 2 is shown in FIG1 , and its sequence is SEQ ID NO: 2. Peptide 1 and Peptide 2 are the polypeptides shown in FIG1 , and their sequences are SEQ ID NO: 33 and SEQ ID NO: 34. The preparation methods for each are as follows:
[0247] DNA2 powder was fully dissolved in pure water, and the concentration was quantified using the Qubit ssDNA Assay Kit (ThermoFisher). Peptide powder was fully dissolved in pure water to 10 mg / mL. To an EP tube, 1 μL of 10× ligation reaction buffer (1 M HEPES (pH 7.2), 50 mM EDTA) was added, along with 80 nmol of peptide and 200 mM TCEP solution to a 1:1 molar ratio of peptide to TCEP. The reaction was diluted with pure water, vortexed, and incubated in a metal bath at 25°C for 10 minutes. After 10 minutes, 5 nmol of DNA1 was added to a final reaction volume of 10 μL. After vortexing, the reaction was incubated in a metal bath at 25°C for 4 hours. Following the reaction, the DNA1-peptide (OP) ligation product fraction was purified using an Agilent 1260 Infinity II HPLC system and lyophilized overnight for the next OPO ligation reaction.
[0248] (2) Four oligo-peptide-oligo (OPO) conjugates were prepared: (a) DNA1-Peptide1-DNA2, (b) DNA1-Peptide2-DNA2, (c) DNA8-Peptide1-DNA2, and (d) DNA8-Peptide2-DNA2. DNA1 and DNA8 are shown in FIG1 , and their sequences are SEQ ID NO: 1 and SEQ ID NO: 8, respectively. The preparation methods for each conjugate are as follows:
[0249] DNA1 / DNA8 powder was fully dissolved in pure water, and its concentration was quantified using the Qubit ssDNA Assay Kit (ThermoFisher). The DNA1 / DNA8 solution was added to the lyophilized OP ligation product powder at a molar ratio of 1:7.5 (OP ligation product: DNA1 / DNA8). 10 μL of 10× ligation buffer (1 M HEPES (pH = 7.2), 50 mM EDTA) was added, and the final reaction volume was adjusted to 100 μL with pure water. After vortexing, the reaction was incubated in a metal bath at 25°C for 16 hours. After completion of the reaction, the DNA1 / DNA8-peptide-DNA2 (DPD) ligation product fraction was purified using a 1260 Infinity II HPLC (Agilent). The concentration of the fraction was quantified using the Qubit ssDNA Assay Kit (ThermoFisher). The samples were aliquoted, lyophilized overnight, and stored at −80°C.
[0250] 2. Experimental Results
[0251] FIG1 shows the OPO synthesis process described in this example. The four OPO ligation products were purified and tested by ESI-MS. The results showed that the molecular weights matched the theoretical molecular weights, indicating that they were correct products.
[0252] Example 2: Nanopore sequencing of OPO
[0253] 1. Experimental Methods
[0254] DNA3 (SEQ ID NO: 3) and DNA4 (SEQ ID NO: 4) powders were fully dissolved in pure water, and their concentrations were quantified using the Qubit ssDNA Assay Kit (ThermoFisher). To the EP tubes containing the four OPO ligation products, aqueous solutions of DNA3 (SEQ ID NO: 3) and DNA4 (SEQ ID NO: 4) were added at a molar ratio of 1:1:1. The mixture was heated at 65°C for 5 minutes, then slowly cooled to 25°C and held for 30 minutes to complete annealing.
[0255] The annealed complex was mixed with a pre-prepared adapter complex (a complex formed by DNA5 (SEQ ID NO: 5), DNA6 (SEQ ID NO: 6) and motor protein), T4 ligase (NEB), and T4 ligase buffer (NEB) and incubated at room temperature for 30 minutes. The concentration of the OPO annealing complex in the mixture was 0.4 μM. Then, 2 μL of the incubated mixture was mixed with DNA7 (SEQ ID NO: 7) and sequencing buffer (0.5 M KCl, 10 mM HEPES, 0.5 mM ATP, 1 mM MgCl2, pH 8) as the sequencing sample. The final concentration of the OPO annealing complex was 2.67 nM.
[0256] This example uses a patch clamp amplifier (other electrical signal amplifiers can also be used) to collect current signals. A single-channel nanopore detection system based on patch clamp and signal amplifier was constructed according to the method disclosed in the literature (Ji Z, Guo P. Channel from bacterial virus T7 DNA packaging motor for the differentiation of peptides composed of a mixture of acidic and basic amino acids. Biomaterials. 2019 May 21; 214: 119-222). A planar 1,2-diphytanoyl-sn-glycero-3-phosphocholine (DPhPC, Avanti Polar Lipids) phospholipid bilayer membrane was used to divide the electrolytic cell into two chambers: the cis chamber and the trans chamber. A pair of Ag / AgCl electrodes was placed in each chamber. The nanopore protein CsgG was added to the bilayer membrane, and a voltage of 180 mV was applied to promote the insertion of the pore protein into the phospholipid bilayer membrane, forming a single nanopore channel. After the single nanopore protein was inserted into the phospholipid membrane, sequencing buffer (0.5 M KCl, 10 mM HEPES, 0.5 mM ATP, 1 mM MgCl2, pH 8) was pushed in to remove excess pore protein. The above-mentioned co-incubation mixture was then added to the cis chamber and incubated at 25°C for 10 min. Finally, 180 mV was applied, and nanopore current data were recorded at a frequency of 5 kHz.
[0257] 2. Experimental Results
[0258] Figure 4 shows the three main signal types detected when performing nanopore sequencing on a sequencing library containing peptide fragments.
[0259] (a) Represents the complete protein sequencing signal, where the area enclosed by the dotted box is the nanopore signal of the polypeptide, the area before the dotted box is the signal of DNA5-DNA1, and the area after the dotted box is the signal of DNA2.
[0260] (b) represents a protein sequencing stuck signal, that is, only a partial signal of DNA5-DNA1 is detected, and then the electrical signal is stuck at a lower current and cannot return to the open pore current on its own. The signals of polypeptide and DNA2 are missing, which is an invalid sequencing signal.
[0261] (c) represents a platform stutter signal, that is, only 1-2 obvious electrical signal stutter platforms can be detected, and no signal of any DNA or polypeptide segment can be observed. It is also an invalid sequencing signal.
[0262] Example 3: Nanopore signal type statistics
[0263] 1. Experimental Methods
[0264] Each sample was run in triplicate, and the frequency of each of the three signal types was counted within 20–30 minutes of applying the 180 mV sequencing voltage. The frequency of each signal type was then plotted using GraphPad Prism.
[0265] 2. Experimental Results
[0266] Figure 5 shows the distribution of nanopore sequencing signal types for the four OPOs in Example 1. (a), (b), (c), and (d) correspond to the statistical results of the four OPO sequencing signals of DNA1-Peptide1-DNA2, DNA1-Peptide2-DNA2, DNA8-Peptide1-DNA2, and DNA8-Peptide2-DNA2, respectively. It can be clearly seen from the figure that the complete protein sequencing signal is the least abundant signal type, accounting for less than 20% of the total signal; the protein sequencing jam signal is the second most abundant signal type, accounting for between 20% and 40%; the most abundant of the three signals is the platform jam signal, accounting for more than 50%. At the same time, when changing the polypeptide sequence (mainly the electrical properties, Peptide1 is electronegative and Peptide2 is electropositive) and the chemical linking group on the guide DNA sequence (DBCO for DNA1 and BCN for DNA8), there is no significant effect on the proportion of the three signals, and the capture rate of the complete protein sequencing signal is 1.2-1.4 signals / minute.
[0267] Example 4: Construction of optimized nucleic acid-polypeptide-nucleic acid complex
[0268] 1. Experimental Methods
[0269] Referring to the method of Example 1, the DNA fragments used were replaced to synthesize the following sequence:
[0270] (1) DNA9-Peptide1-DNA10 (using DNA9 instead of DNA1 / DNA8, and DNA10 instead of DNA2);
[0271] (2) DNA11-Peptide1-DNA12 (using DNA11 instead of DNA1 / DNA8 and DNA12 instead of DNA2);
[0272] (3) DNA13-Peptide1-DNA14 (using DNA13 instead of DNA1 / DNA8 and DNA14 instead of DNA2);
[0273] (4) DNA15-Peptide1-DNA16 (using DNA15 instead of DNA1 / DNA8 and DNA16 instead of DNA2);
[0274] (5) DNA25-Peptide1-DNA26 (using DNA25 instead of DNA1 / DNA8, and DNA26 instead of DNA2);
[0275] (6) DNA27-Peptide1-DNA28 (DNA27 is used instead of DNA1 / DNA8, and DNA28 is used instead of DNA2).
[0276] 2. Experimental Results
[0277] After purification, the four OPO ligation products were detected by ESI-MS, and the results showed that the molecular weights matched the theoretical molecular weights, confirming that they were correct products.
[0278] Example 5: Nanopore sequencing and nanopore signal type statistics of optimized OPO
[0279] 1. Experimental Methods
[0280] Referring to the method of Example 2, the OPO and DNA sequences were replaced accordingly during annealing, as follows:
[0281] (1) DNA9-Peptide1-DNA10 (D9P1D10): DNA17 replaces DNA3, and DNA18 replaces DNA4;
[0282] (2) DNA11-Peptide1-DNA12 (D11P1D12): DNA19 replaces DNA3, and DNA20 replaces DNA4;
[0283] (3) DNA13-Peptide1-DNA14 (D13P1D14): DNA21 replaces DNA3, and DNA22 replaces DNA4;
[0284] (4) DNA15-Peptide1-DNA16 (D15P1D16): DNA23 replaces DNA3, and DNA24 replaces DNA4;
[0285] (5) DNA25-Peptide1-DNA26 (D25P1D26): DNA29 replaces DNA3, and DNA30 replaces DNA4;
[0286] (6) DNA27-Peptide1-DNA28 (D27P1D28): DNA31 replaces DNA3, and DNA32 replaces DNA4.
[0287] The methods of library construction and nanopore sequencing were consistent with those described in Example 2.
[0288] Referring to the method of Example 3, statistics of nanopore signal types were performed on the six libraries constructed above.
[0289] 2. Experimental results:
[0290] Figure 6 shows the distribution of nanopore sequencing signal types for the six sequencing libraries in this example. Figures (a), (b), (c), (d), (e), and (f) correspond to the statistical results of six OPO sequencing signals: DNA9-Peptide1-DNA10, DNA11-Peptide1-DNA12, DNA13-Peptide1-DNA14, DNA15-Peptide1-DNA16, DNA25-Peptide1-DNA26, and DNA27-Peptide1-DNA28, respectively. The following conclusions can be drawn from these figures:
[0291] (1) As can be seen from Figure (a), compared with the signal type ratios shown in Figure 5, extending the DNA sequences at both ends of the peptide, the 3' end of DNA9 and the 5' end of DNA10 by 11 nt respectively, can significantly increase the proportion of complete peptide sequencing signals to more than 60%, while greatly reducing the two types of jamming signals;
[0292] (2) As can be seen from Figure (b), extending the DNA sequences at both ends of the polypeptide, namely the 5' end of DNA11 and the 3' end of DNA12, by 11 nt respectively, can also achieve similar effects as described in (1). This shows that as long as the DNA sequence coupled to both ends of the polypeptide is extended, no matter which end is extended, the purpose of reducing jamming and increasing the proportion of correct signals can be achieved;
[0293] (3) As can be seen from Figure (c), the extended polypeptide end sequences, regardless of whether the extended sequence consists of a poly base sequence, can achieve the purpose of reducing jams and increasing the correct signal ratio. After replacing the poly base with a random base sequence, the correct sequencing signal ratio is further improved (the correct sequencing signal ratio is close to 80%).
[0294] (4) As can be seen from Figure (d), further extending the nucleic acid sequence to a random sequence of 20 nt can also significantly reduce the jamming and increase the proportion of correct signals. After the sequence is extended even longer, the proportion of correct sequencing signals is close to 90%, and the platform jamming almost disappears, and the optimization effect is more obvious.
[0295] (5) As can be seen from Figure (e), after the extended fragment was replaced by nucleotides with organic linkers, a significant reduction in jamming was observed, which increased the proportion of correct signals, with the correct sequencing signal proportion reaching 80%.
[0296] (6) As can be seen from Figure (f), when the organic linker used to extend the DNA sequence is placed in the middle of the original DNA sequence, similar results to (e) can be observed. Compared with before optimization, the jamming is significantly reduced, and the proportion of correct signals is significantly increased, with the proportion of correct sequencing signals reaching 80%.
[0297] In summary, after optimization, the capture rate of complete protein sequencing signals can reach about 5-10 signals per minute.
[0298] From the above description, it can be seen that the above embodiments of the present invention achieve the following technical effects:
[0299] 1) For the first time, various signal types that may occur during nanopore protein sequencing using motor protein speed control were demonstrated, including the first description of the problem of signal stuttering during nanopore sequencing.
[0300] 2) The improved nucleic acid sequences at both ends of the extended polypeptide of the present invention can not only reduce jamming and invalid current signals, thereby improving sequencing quality, but also increase the proportion of correct polypeptide signals by introducing special groups or sequences into the nucleic acid sequences at both ends.
[0301] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A nucleic acid-polypeptide-nucleic acid ternary complex, characterized in that: It has the structural formula shown in formula (1): Single-stranded nucleic acid 1-polypeptide-single-stranded nucleic acid 2 Formula (1) The single-stranded nucleic acid 1, the polypeptide and the single-stranded nucleic acid 2 are covalently linked in sequence, wherein the single-stranded nucleic acid 1 is used to connect to the sequencing adapter. The single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2 each independently further includes an extension chain, which is arranged at any end or in the middle of the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2, and the extension chain includes a chain formed by a random arrangement of one or more of the following molecules: nucleosides, nucleotides or organic linkers, and the extension chain located on the single-stranded nucleic acid 2 is not a polymononucleotide sequence.
2. The nucleic acid-polypeptide-nucleic acid ternary complex according to claim 1, characterized in that: When the extended chain contains the nucleoside and / or the nucleotide, the extended chain is located at the 5' end or the 3' end of the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2, Preferably, the length of the extended chain is 5 to 100 bp, preferably 10 to 30 bp, more preferably 10 to 20 bp; Preferably, the nucleoside is selected from modified nucleosides and / or unmodified nucleosides; Preferably, the unmodified nucleosides are selected from deoxyribonucleosides and / or ribonucleosides; More preferably, the deoxyribonucleoside is selected from any one or more of A, T, C or G; More preferably, the ribonucleoside is selected from any one or more of A, U, C or G; Preferably, the modified nucleoside is selected from any one or more of the following: 2-aminopurine nucleoside, 5-bromodeoxyuridine, dideoxynucleoside, 5-methylcytosine deoxynucleoside, 5-hydroxymethylcytosine deoxynucleoside, N6-methyladenosine, deoxyinosine or 5-nitro-2-deoxycytidine nucleoside or G-quadruplex; More preferably, the dideoxynucleoside is selected from any one or more of the following: ddA, ddT, ddC or ddG.
3. The nucleic acid-polypeptide-nucleic acid ternary complex according to claim 1, characterized in that: The nucleotides are selected from modified nucleotides and / or unmodified nucleotides, Preferably, the unmodified nucleotides are selected from any one or more of the following: DNA or RNA composed of random base sequences; Preferably, the modified nucleotide is selected from any one or more of the following: deoxynucleotides without side chains, dT inverted nucleotides, dG inverted nucleotides, G-quadruplexes, or nucleotides with any one or more of the following modifications: phosphorylation, amination, carboxylation, aldehydeation, azidation, alkynylation, acrylamidation, maleamidation, DBCO-ylation, BCN-ylation, sulfhydrylation, dithiolation, biotinylation, dethiobiotinylation, sterolization, or a fluorescent group; Preferably, the fluorescent group is selected from any one of the following: Cy3, Cy5, FAM, Alexa Fluor 488 Or Texas Red.
4. The nucleic acid-polypeptide-nucleic acid ternary complex according to claim 1, characterized in that: The organic linker is selected from any one or more of the following: Spacer C3, Spacer C6, Spacer 9, Spacer C12 or Spacer 18; Preferably, the organic linker is located at the 5' end, 3' end or the middle of the single-stranded nucleic acid 1 and / or the single-stranded nucleic acid 2, more preferably in the middle.
5. The nucleic acid-polypeptide-nucleic acid ternary complex according to claim 1 or 4, characterized in that: When the extended chain is only the organic linker, the number of the organic linkers is greater than or equal to 2; When the extended chain further comprises the organic linker in addition to the nucleoside and / or the nucleotide, the number of the organic linker is at least 1, and the organic linker is located at any position of the nucleoside and / or the nucleotide; More preferably, the organic linker is located in the middle position of the nucleoside and / or the nucleotide.
6. The nucleic acid-polypeptide-nucleic acid ternary complex according to claim 1, characterized in that: The C-terminus or N-terminus of the polypeptide is modified to achieve connection with the single-stranded nucleic acid 2; Preferably, the C-terminus or N-terminus of the polypeptide has an azide modification.
7. The nucleic acid-polypeptide-nucleic acid ternary complex according to any one of claims 1 to 6, characterized in that The single-stranded nucleic acid 1, the polypeptide and the single-stranded nucleic acid 2 are covalently linked by any one or more of the following methods: peptide bond linkage, ester bond linkage, ether bond linkage, thiol-maleimide linkage, carbonyl-hydroxylamine compound oxime linkage, carbonyl-hydrazone linkage, carbonyl-urea structure compound urea linkage, disulfide bond linkage, thioether bond linkage, halogen-nucleophile substitution linkage, 1,3 dipolar cycloaddition reaction linkage, copper-catalyzed azide-alkynyl cycloaddition reaction linkage, ruthenium-catalyzed azide-alkynyl cycloaddition reaction linkage, azide compound-alkynyl compound click chemistry reaction linkage or natural chemical linkage; Preferably, the azide compound-alkynyl compound click chemistry reaction connection includes azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection, and azide-cyclooctene click chemistry reaction connection.
8. The nucleic acid-polypeptide-nucleic acid ternary complex according to claim 7, characterized in that: The combination of the single-stranded nucleic acid 1 and the single-stranded nucleic acid 2 is selected from any one of the following groups: 1) DNA11 and DNA12; 2) DNA13 and DNA14; 3) DNA15 and DNA16; Wherein, the DNA 11 is selected from the sequence shown in SEQ ID NO: 11: AAAAAAAAAAAGCTTCTCGTG, wherein the 5' end has a phosphorylation group and the 3' end has a DBCO group; The DNA 12 is selected from the sequence shown in SEQ ID NO: 12: GCTGTCTTCTGTCGTCGTTTCCTTCTCTGCAAAAAAAAAAA, wherein the 5' end has a maleamidated group; The DNA 13 is selected from the sequence shown in SEQ ID NO: 13: GCTTCTCGTGAGAGAGGCGG, wherein the 5' end has a phosphorylation group and the 3' end has a DBCO group; The DNA 14 is selected from the sequence shown in SEQ ID NO: 14: GGCGGAGAGAGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC, wherein the 5' end has a maleamidated group; The DNA 15 is selected from the sequence shown in SEQ ID NO: 15: GCTTCTCGTGGTCGAAAAAGAGAGAGGCGG, in which the 5' end has a phosphorylation group and the 3' end has a DBCO group; The DNA 16 is selected from the sequence shown in SEQ ID NO: 16: GGCGGAGAGAAAAAAGCTGGGCTGTCTTCTGTCGTCGTTTCCTTCTCTGC, wherein the 5' end has a malein amidation group.
9. The nucleic acid-polypeptide-nucleic acid ternary complex according to any one of claims 1 to 7, characterized in that The connection direction of the single-stranded nucleic acid 1, the polypeptide and the single-stranded nucleic acid 2 is: 5'-single-stranded nucleic acid 1-3'-N-polypeptide-C-5'-single-stranded nucleic acid 2-3' or 5'-single-stranded nucleic acid 1-3'-C-polypeptide-N-5'-single-stranded nucleic acid 2-3'.
10. A polypeptide nanopore sequencing library, characterized in that: The method comprises the nucleic acid-polypeptide-nucleic acid ternary complex according to any one of claims 1 to 9.
11. The polypeptide nanopore sequencing library according to claim 10, wherein The polypeptide nanopore sequencing library comprises: a) double-stranded annealing complex; and / or b) a linker complex covalently linked to the double-stranded annealing complex; The double-stranded annealing complex comprises: single-stranded nucleic acid 1-polypeptide-single-stranded nucleic acid 2 as shown in formula (1), complementary fragment 1 complementary to the single-stranded nucleic acid 1, and complementary fragment 2 complementary to the single-stranded nucleic acid 2; The linker complex comprises: a linker sequence 1, a linker sequence 2, and a motor protein, wherein the linker sequence 1 comprises a first segment and a second segment connected sequentially from the 5' end to the 3' end, wherein the first segment is not complementary to the linker sequence 2, and the second segment is complementary to the linker sequence 2, and the motor protein is movably bound to the linker sequence 1. On the first paragraph; Wherein, the linker sequence 1 is covalently linked to the 5' end of the single-stranded nucleic acid 1; The linker sequence 2 is covalently linked to the 3′ end of the complementary fragment 1 of the single-stranded nucleic acid 1 .
12. The polypeptide nanopore sequencing library according to claim 11, wherein The combination of the complementary fragment 1 and the complementary fragment 2 is selected from any one of the following groups: 1) DNA 17 represented by SEQ ID NO: 17 and DNA 18 represented by SEQ ID NO: 18; 2) DNA 19 represented by SEQ ID NO: 19 and DNA 20 represented by SEQ ID NO: 20; 3) DNA21 represented by SEQ ID NO: 21 and DNA22 represented by SEQ ID NO: 17; 4) DNA23 shown in SEQ ID NO: 23 and DNA24 shown in SEQ ID NO: 24; 5) DNA29 represented by SEQ ID NO: 29 and DNA30 represented by SEQ ID NO: 30; 6) DNA31 represented by SEQ ID NO: 31 and DNA32 represented by SEQ ID NO: 32; SEQ ID NO: 17:TTTTTTTTTTTCACGAGAAGC; SEQ ID NO: 18: GCAGAGAAGGAAACGACGACAGAAGACAGCTTTTTTTTTTTT; SEQ ID NO: 19: CACGAGAAGCTTTTTTTTTTT; SEQ ID NO: 20: TTTTTTTTTTTGCAGAGAAGGAAACGACGACAGAAGACAGC; SEQ ID NO: 21: CCGCCTTCTCCACGAGAAGC; SEQ ID NO: 22: GCAGAGAAGGAAACGACGACAGAAGACAGCTCTCTCCGCC; SEQ ID NO: 23: CCGCCTCTCTCTTTTTCGACCACGAGAAGC; SEQ ID NO: 24: GCAGAGAAGGAAACGACGACAGAAGACAGCCCAGCTTTTTTCTCTCCGCC; DNA29 (SEQ ID NO: 29):AAAAAAAAAAAACACGAGAAGC; DNA30 (SEQ ID NO: 30): GCAGAGAAGGAAACGACGACAGAAGACAGCAAAAAAAAAAAA; DNA31 (SEQ ID NO: 31): CACGAAAAAAAAAAAAAGAAGC; DNA32 (SEQ ID NO: 32): GCAGAGAAGGAAACGAAAAAAAAAAAAACGACAGAAGACAGC; Preferably, the sequence of the linker sequence 1 is shown in SEQ ID NO: 5: 5'-XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXTTTTTTTTTTYYYYGGTTGTTTCTGTTGGTGCTGATATTGCT-3', where X=SpC3, Y=iSp18; Preferably, the sequence of the linker sequence 2 is as shown in SEQ ID NO: 6: 5′-phosphorylated-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3′.
13. The polypeptide nanopore sequencing library according to claim 11, wherein The motor protein is selected from any one of the following: phi29 polymerase, Hel308 helicase, MTA helicase or DDA helicase.
14. A polypeptide nanopore sequencing kit, characterized in that: The kit comprises: the single-stranded nucleic acid 1 and the single-stranded nucleic acid 2 in the nucleic acid-polypeptide-nucleic acid ternary complex according to any one of claims 1 to 9, and any one or more of the following optional components: A nanopore and the complementary fragment 1, complementary fragment 2, adapter sequence 1, adapter sequence 2 and motor protein in the polypeptide nanopore sequencing library according to claim 11.
15. The kit according to claim 14, characterized in that The linker sequence 1, the linker sequence 2, and the motor protein exist in the form of a linker complex; Preferably, the nanopore is a protein nanopore or a solid-state nanopore; Preferably, the protein nanopore is selected from the group consisting of nanopores of any one of the following proteins or variants thereof: α-hemolysi, Aerolysin, MspA, CsgG, or FraC; Preferably, the solid-state nanopore is selected from any one of the following: graphene nanopore, gold nanopore, silicon nitride nanopore, silicon dioxide nanopore or aluminum oxide nanopore.
16. A method for constructing a polypeptide nanopore sequencing library, characterized in that: The construction method comprises: preparing the polypeptide to be tested into a nucleic acid-polypeptide-nucleic acid ternary complex according to any one of claims 1 to 9; Annealing the nucleic acid-polypeptide-nucleic acid ternary complex with the complementary fragment 1 of the single-stranded nucleic acid 1 and the complementary fragment 2 of the single-stranded nucleic acid 2 to form a double-stranded annealing complex; The double-stranded annealing complex is connected to a linker complex containing a motor protein by a ligase to form the polypeptide nanopore sequencing library.
17. The construction method according to claim 16, characterized in that: The preparation of the polypeptide to be tested into the nucleic acid-polypeptide-nucleic acid ternary complex comprises: Covalently linking the single-stranded nucleic acid 2 to the polypeptide to be tested to obtain a polypeptide-single-stranded nucleic acid 2 complex; Covalently linking the single-stranded nucleic acid 1 to the polypeptide-single-stranded nucleic acid 2 complex to obtain the nucleic acid-polypeptide-nucleic acid ternary complex; Preferably, the single-stranded nucleic acid 2 having a maleimide group at the 5' end is covalently linked to the polypeptide having an azide group at the N-terminus or C-terminus through a thiol-maleimide addition reaction to obtain the polypeptide-single-stranded nucleic acid 2 complex; The single-stranded nucleic acid 1 with a DBCO group at the 3' end is covalently linked to the azide group on the polypeptide through a click chemistry reaction to obtain the nucleic acid-polypeptide-nucleic acid ternary complex.
18. The construction method according to claim 16, characterized in that: The linker complex comprises: a linker sequence 1, a linker sequence 2 that is complementary to the 3' end of the linker sequence 1 but not complementary to the 5' end, and the motor protein located on the linker sequence 1; Preferably, the sequence of the linker sequence 1 is shown in SEQ ID NO: 5: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXTTTTTTTTTYYYYGGTTGTTTCTGTTGGTGCTGATATTGCT, where, X=SpC3, Y=iSp18; Preferably, the sequence of the linker sequence 2 is as shown in SEQ ID NO: 6: 5′-phosphorylated-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3′; Preferably, the motor protein is selected from any one of the following: phi29 polymerase, Hel308 helicase, MTA helicase or DDA helicase.
19. The construction method according to claim 16, characterized in that: The combination of the complementary fragment 1 and the complementary fragment 2 is selected from any one of the following groups: 1) DNA 17 represented by SEQ ID NO: 17 and DNA 18 represented by SEQ ID NO: 18; 2) DNA 19 represented by SEQ ID NO: 19 and DNA 20 represented by SEQ ID NO: 20; 3) DNA21 represented by SEQ ID NO: 21 and DNA22 represented by SEQ ID NO: 17; 4) DNA23 shown in SEQ ID NO: 23 and DNA24 shown in SEQ ID NO: 24; 5) DNA29 represented by SEQ ID NO: 29 and DNA30 represented by SEQ ID NO: 30; 6) DNA31 represented by SEQ ID NO: 31 and DNA32 represented by SEQ ID NO: 32; SEQ ID NO: 17:TTTTTTTTTTTCACGAGAAGC; SEQ ID NO: 18: GCAGAGAAGGAAACGACGACAGAAGACAGCTTTTTTTTTTTT; SEQ ID NO: 19: CACGAGAAGCTTTTTTTTTTT; SEQ ID NO: 20: TTTTTTTTTTTGCAGAGAAGGAAACGACGACAGAAGACAGC; SEQ ID NO: 21: CCGCCTTCTCCACGAGAAGC; SEQ ID NO: 22: GCAGAGAAGGAAACGACGACAGAAGACAGCTCTCTCCGCC; SEQ ID NO: 23: CCGCCTCTCTCTTTTTCGACCACGAGAAGC; SEQ ID NO: 24: GCAGAGAAGGAAACGACGACAGAAGACAGCCCAGCTTTTTTCTCTCCGCC; SEQ ID NO: 29:AAAAAAAAAAAACACGAGAAGC; SEQ ID NO: 30: GCAGAGAAGGAAACGACGACAGAAGACAGCAAAAAAAAAAAA; SEQ ID NO: 31: CACGAAAAAAAAAAAAAGAAGC; SEQ ID NO: 32: GCAGAGAAGGAAACGAAAAAAAAAAAAACGACAGAAGACAGC.
20. A polypeptide nanopore sequencing method, characterized in that: The sequencing method comprises: Co-incubating the polypeptide nanopore sequencing library constructed by the construction method of any one of claims 16 to 19 with the anchor sequence to obtain an incubation complex; The incubation complex is added to the solution chamber of the sequencing chip, and under the action of the electric field force, the motor protein unwinds the double-stranded DNA to control the test polypeptide to pass through the nanopore, thereby obtaining an electrical signal corresponding to the test polypeptide; The electrical signal is decoded to obtain the amino acid sequence corresponding to the polypeptide to be tested.
21. The sequencing method according to claim 20, characterized in that One end of the anchor sequence is complementary to the end of the linker sequence 2 in the linker complex that is away from the complementary fragment 1, and the other end carries an anchor group; Preferably, the anchoring group is selected from any one of lipids, carbon nanotubes, polypeptides, proteins and / or amino acids; Preferably, the lipid is selected from any one of fatty acids, sterols, cholesterol, palmitate or tocopherol; Preferably, the anchor sequence is shown in SEQ ID NO: 7: 5'-Chol-TEG-TTYYYYTTGACCGCTCGCCTC-3', where Y=iSp18, Chol-TEG represents Cholesterol-polyethylene glycol.
22. The sequencing method according to claim 20, characterized in that The nanopore is a protein nanopore or a solid-state nanopore; Preferably, the protein nanopore is selected from the group consisting of nanopores of any one of the following proteins or variants thereof: α-hemolysin, Aerolysin, MspA, CsgG, or FraC; Preferably, the solid-state nanopore is selected from any one of the following: graphene nanopore, gold nanopore, silicon nitride nanopore, silicon dioxide nanopore or aluminum oxide nanopore.