Binary complex containing chaperone molecule and use thereof in polypeptide sequencing

By using chaperone molecules to covalently link peptides and motor protein control in nanopore peptide sequencing, the problem of low resolution in nanopore peptide sequencing has been solved, and the accuracy and stability of peptide sequencing have been improved.

WO2026050906A1PCT designated stage Publication Date: 2026-03-12SHENZHEN HUADA GENE INST
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing nanopore peptide sequencing technologies suffer from low resolution and difficulty in accurately resolving amino acid composition, mainly due to the limited spatial resolution of nanopores and the low electrical signal resolution caused by the difference in molecular scale between peptides and nucleic acids.

Method used

A binary complex containing a chaperone molecule is used to covalently link the peptide and the chaperone molecule to form a double-chain annealed complex. Motor proteins are then used to control the passage of the complex through nanopores to obtain a stable electrical signal.

Benefits of technology

This improved the resolution of peptide sequencing, obtained reproducible and consistent electrical signals, and enhanced the accuracy of nanopore peptide sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116688_12032026_PF_FP_ABST
    Figure CN2024116688_12032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a binary complex containing a chaperone molecule and the use thereof in polypeptide sequencing. The binary complex comprises a chaperone molecule and a polypeptide to be detected that are covalently linked, wherein the chaperone molecule is selected from any one or more of the following: PEG, a spacer, a deoxyribose phosphate, a ribose phosphate, a nucleotide, a deoxynucleotide, a peptide nucleic acid or a locked nucleotide; the number of constituent units of the chaperone molecule is ≥1; and the number of amino acids of said polypeptide is ≥2. The above development and use are conducive to improving the resolution of nanopore polypeptide sequencing, and are of great significance for the advancement of nanopore polypeptide sequencing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Binary complex containing chaperone molecules and its application in polypeptide sequencing TECHNICAL FIELD

[0001] The present application relates to the technical field of protein molecule sequencing, in particular, it relates to a binary complex containing chaperone molecules and its application in polypeptide sequencing. BACKGROUND

[0002] The related research of polypeptide sequencing by nanopore has been very common, and there are two common technical means: 1) the polypeptide is directly passed through the nanopore under the action of electric field force, and the polypeptide is analyzed by using the pore signal; 2) the phosphoribosyl backbone of nucleic acid is connected with polypeptide to form a nucleic acid polypeptide complex, which is stably passed through the nanopore under the traction of motor protein such as helicase or polymerase, and the signal is analyzed. The technical solutions of the prior art include the following four kinds:

[0003] (1) The denatured polypeptide directly passes through the pore. For example, Yu et al. denatured the protein by using guanidine hydrochloride, so that the full-length protein was unfolded and translocated through the nanopore. By analyzing the current change generated when passing through the pore, three kinds of proteins can be distinguished (DOI: https: / / doi.org / 10.1038 / s41587-022-01598-3).

[0004] (2) The protein is enzymatically digested into multiple peptide segments and directly passed through the pore. For example, Lucas et al. used trypsin to digest the measured protein to obtain polypeptide fragments passing through the nanopore, and collected the electrical signals during translocation, mainly the degree of current signal blockage and blockage time. At the same time, the electrical signals of the corresponding model polypeptide were also collected. By comparing the polypeptide signals obtained by enzymatic digestion with the electrical signals of the model polypeptide, the measured polypeptide can be identified. (DOI: https: / / doi.org / 10.1038 / s41467-021-26046-9) (Figure 1).

[0005] (3) The polypeptide is coupled with nucleic acid and passed through the pore by polymerase traction. For example, the laboratory of Huang Shuo linked polypeptide with DNA, and used the polymerization of single-stranded DNA by polymerase phi29 to pull the polypeptide chain to move in the nanopore, so as to sequence the polypeptide chain. (DOI: https: / / doi.org / 10.1021 / acs.nanolett.1c02371).

[0006] (4) Coupling the polypeptide with nucleic acid, and pulling through the pore by helicase. For example, the Whitehead Institute used MTA helicase to control the sequencing of single-stranded "nucleic acid-polypeptide-nucleic acid" (DOI: 10.1039 / D1SC04342K). From the sequencing signal, first, the relatively flat step signal of single-stranded DNA appears, then the amplitude suddenly increases, the polypeptide sequencing signal appears, then it decreases and gradually tends to be flat, and finally the nucleic acid signal appears. The last piece of nucleic acid signal is relatively flat and has no obvious step signal because this piece of DNA sequence is polyT (Figure 2). For another example, patent application WO2021 / 111125A1 discloses connecting polypeptide to nucleic acid at both ends, then annealing single-stranded DNA to form double-stranded DNA, and combining with helicase to control the translocation speed of DNA through the helicase, and then indirectly control the translocation speed of polypeptide, so as to collect the nanopore signal of polypeptide.

[0007] The above-mentioned schemes (1) and (2) belong to the method of directly passing the whole polypeptide or polypeptide fragment through the pore. When the polypeptide fragment is directly passed through the pore, the speed is too fast, the time of staying in the nanopore is very short, and therefore the current signal information that can be obtained is less, generally only three information of current blockage degree, current fluctuation degree and pore passing time, which makes the resolution of the sequencing signal lower. In addition, the protein is digested into polypeptide fragments, and these polypeptide fragments cannot be assembled, which makes it difficult to assemble the sequence information of the whole protein. The above-mentioned schemes (3) and (4) belong to the method of passing the polypeptide through the pore under the control of motor protein molecules (such as polymerase or helicase, etc.) after coupling the polypeptide with nucleic acid. There is a natural difference in molecular size between polypeptide and nucleic acid, and in these technical solutions, the nucleic acid molecules have a better interaction with the pore, and there is a clear step signal when passing through the pore. However, since the polypeptide is smaller than the nucleic acid, it cannot have a stable interaction with the nanopore, and therefore the resolution of the electric signal measured by the polypeptide part is lower, and the step of the electric signal is not obvious, which is not conducive to the sequencing of the polypeptide.

[0008] In summary, at present there is no method that can accurately analyze the composition of amino acids in polypeptides. The main reasons are as follows: first, the spatial resolution of the nanopore itself is limited, and the molecular radius of the amino acid is much smaller than the diameter of the nanopore, and only a weak current block signal can be generated when passing through the nanopore, and many times it does not interact with the nanopore. Second, common nanopore nucleic acid sequencing only needs to analyze the signals of 4 bases, while there are as many as 20 kinds of common natural amino acids and numerous post-translational modification sites, and the arrangement and combination of these different types make the analysis of the sequencing signal impossible. Therefore, most of the existing methods are to comprehensively analyze the fingerprint information of the polypeptide, which is far from the level of sequencing. How to improve the resolution of nanopore polypeptide sequencing has become a key technical problem to be solved at present.

[0009] SUMMARY

[0010] The main purpose of the present application is to provide a binary complex containing a partner molecule and its application in polypeptide sequencing, so as to solve the problem of low resolution in nanopore polypeptide sequencing in the prior art.

[0011] In order to achieve the above-mentioned purpose, according to the first aspect of the present application, a binary complex is provided, comprising a covalently linked partner molecule and a polypeptide to be detected, wherein the partner molecule is selected from any one or more of the following: PEG, spacer, deoxyribose phosphate, ribose phosphate, nucleotide, deoxynucleotide, peptide nucleic acid or locked nucleotide; wherein the number of constituent units of the partner molecule is ≥1; the number of amino acids of the polypeptide to be detected is ≥2.

[0012] Further, the covalently linked form of the partner molecule and the polypeptide to be detected includes any one or more of any one of the following groups:

[0013] 1) one end of a polypeptide to be detected and one end of a partner molecule are single-end bonded;

[0014] 2) the same end of a polypeptide to be detected is single-end bonded to two ends of a partner molecule, one end of a polypeptide to be detected is single-end bonded to the middle position of a partner molecule, or the same end of a partner molecule is single-end bonded to one end of another partner molecule and one end of a polypeptide to be detected;

[0015] 3) the same end of a partner molecule is single-end bonded to two ends of a polypeptide to be detected, one end of a partner molecule is single-end bonded to the middle position of a polypeptide to be detected, or the same end of a polypeptide to be detected is single-end bonded to one end of a partner molecule and one end of a polypeptide to be detected;

[0016] 4) the same end of a complex element 1 is single-end bonded to two ends of a partner molecule, a complex element 1 is single-end bonded to the middle position of a partner molecule, or the same end of a partner molecule is single-end bonded to one end of a complex element 1 and one end of a partner molecule;

[0017] 5) the same end of a complex element 1 is single-end bonded to one end of a partner molecule and one end of a polypeptide to be detected, the same end of a partner molecule is single-end bonded to one end of a complex element 1 and one end of a polypeptide to be detected, or one end of a complex element 2 is single-end bonded to one end of a polypeptide to be detected;

[0018] 6) a polypeptide to be detected is two-end bonded to a partner molecule;

[0019] 7) one end of a complex element 3 is single-end bonded to any one end of a partner molecule;

[0020] 8) one end of a complex element 3 is single-end bonded to any one end of a polypeptide to be tested;

[0021] 9) any one end of a complex element 3 is single-end bonded to any one end of a partner molecule;

[0022] 10) any one end of a complex element 3 is single-end bonded to any one end of a polypeptide to be tested;

[0023] 11) one end of a complex element 1 is double-end bonded to a partner molecule;

[0024] 12) one end of a complex element 4 is single-end bonded to any one end of a partner molecule;

[0025] 13) any one end of a complex element 4 is single-end bonded to any one end of a partner molecule;

[0026] 14) one end of a complex element 4 is single-end bonded to any one end of a polypeptide to be tested;

[0027] 15) any one end of a complex element 4 is single-end bonded to any one end of a polypeptide to be tested;

[0028] wherein the partner molecules at different positions in any one form of binary complex are the same or different, and the polypeptides to be tested at different positions in any one form of binary complex are the same or different;

[0029] a complex element 1 is a complex in which one end of a polypeptide to be tested is single-end bonded to any one end of a partner molecule;

[0030] a complex element 2 is a complex in which any one end of a complex element 1 is single-end bonded to any one end of a partner molecule;

[0031] a complex element 3 is a complex in which a polypeptide to be tested is double-end bonded to a partner molecule;

[0032] a complex element 4 is a complex in which a complex element 1 is double-end bonded to a partner molecule.

[0033] Further, the chaperone molecule and the polypeptide to be detected are covalently connected by any one or more of the following ways: peptide bond connection, ester bond connection, ether bond connection, thiol-maleimide connection, oxime formation of carbonyl-hydroxylamine-containing compound, hydrazone formation of carbonyl-hydrazine-containing compound, urea formation of carbonyl-urea-containing compound, disulfide bond connection, sulfide bond connection, halogen-nucleophile substitution connection, 1,3 dipole cycloaddition reaction connection, copper-catalyzed azide-alkynyl cycloaddition reaction connection, ruthenium-catalyzed azide-alkynyl cycloaddition reaction connection, azide compound-alkynyl compound click chemistry reaction connection or natural chemical connection.

[0034] Preferably, the azide compound-alkynyl compound click chemistry reaction connection includes any one or more of the following: azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection or azide-cyclooctene click chemistry reaction connection.

[0035] Further, the spacer is selected from any one or more of the following: Spacer C3, Spacer C6, Spacer 9, Spacer C12 or Spacer 18.

[0036] Further, the average molecular weight of the PEG is 300-20000; preferably, the PEG is selected from any one or more of the following: PEG-300, PEG-400, PEG-800, PEG-1000, PEG-1500, PEG-2000, PEG-3000, PEG-4000, PEG-6000, PEG-8000 or PEG-20000.

[0037] In order to achieve the above-mentioned purpose, according to a second aspect of the present application, a double-stranded annealing complex is provided, which includes: a single-stranded nucleic acid 1 and a complementary fragment 1 which is at least partially complementary to the single-stranded nucleic acid 1, wherein the middle position of the single-stranded nucleic acid 1 is covalently connected with the above-mentioned binary complex.

[0038] Further, the first segment of the single-stranded nucleic acid 1 and the second segment of the single-stranded nucleic acid 1 are respectively covalently connected with two ends of the binary complex; wherein, the covalent connection mode is selected from any one or more of the following: peptide bond connection, ester bond connection, ether bond connection, thiol-maleimide connection, oxime formation of carbonyl-hydroxylamine-containing compound connection, hydrazone formation of carbonyl-hydrazine-containing compound connection, urea formation of carbonyl-urea structure-containing compound connection, disulfide bond connection, sulfide bond connection, halogen-nucleophile substitution connection, 1,3 dipole cycloaddition reaction connection, copper-catalyzed azide-alkynyl cycloaddition reaction connection, ruthenium-catalyzed azide-alkynyl cycloaddition reaction connection, azide compound-alkynyl compound click chemistry reaction connection or natural chemical connection;

[0039] Preferably, the covalent connection mode is diester bond connection; more preferably, the covalent connection mode is phosphodiester bond connection.

[0040] Preferably, the azide compound-alkynyl compound click chemistry reaction connection includes azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection, azide-cyclooctene click chemistry reaction connection.

[0041] Further, the structural formula of the single-stranded nucleic acid 1 is: the first segment of the single-stranded nucleic acid 1-int DBCO modified deoxyribonucleotide-binary complex-int DBCO modified deoxyribonucleotide-the second segment of the single-stranded nucleic acid 1.

[0042] Further, the length of the first segment of the single-stranded nucleic acid 1 is ≥1 nt, the length of the second segment of the single-stranded nucleic acid 1 is ≥1 nt; the length of the complementary fragment 1 is ≥1 nt; preferably, the length of the first segment of the single-stranded nucleic acid 1 is 5-500 nt; the length of the second segment of the single-stranded nucleic acid 1 is 5-500 nt; the length of the complementary fragment 1 is 5-1200 nt.

[0043] Further, the 5' end of the first segment of the single-stranded nucleic acid 1 has a phosphorylation group, and the 3' end of the second segment of the single-stranded nucleic acid 1 has any one or more of the following dideoxynucleotides: ddA, ddT, ddC or ddG.

[0044] Further, the first segment of the single-stranded nucleic acid 1 is selected from the sequence shown in SEQ ID NO: 1: 5'-GCTTCTCGTGTTTTTTTTCTCTC-3'; the second segment of the single-stranded nucleic acid 1 is selected from the sequence shown in SEQ ID NO: 2: 5'-CCCTTTTTTTTTTGCTGTCTTCTGTCGTCGTTT-3'; and the complementary fragment 1 is selected from the sequence shown in SEQ ID NO: 5: 5'-GAAACGACGACAGAAGACAGCAAAAAAAAAAGGGATTTTTTAGAGAGAAAAAAAAC GAGAAGCA-3'.

[0045] To achieve the above object, according to a third aspect of the present application, there is provided a polypeptide library comprising the above-mentioned binary complex or the above-mentioned double-stranded annealed complex.

[0046] Further, the polypeptide library comprises the double-stranded annealed complex and a linker complex covalently linked to the double-stranded annealed complex;

[0047] The linker complex comprises: a linker sequence 1 and a linker sequence 2, and a motor protein, the linker sequence 1 comprises, in the direction from the 5' end to the 3' end, a first segment and a second segment connected in sequence, wherein the first segment of the linker sequence 1 is not complementary to the linker sequence 2, the second segment of the linker sequence 1 is complementary to the linker sequence 2, and the motor protein is movably bound to the first segment of the linker sequence 1.

[0048] The linker sequence 1 is covalently linked to the 5' end of the single-stranded nucleic acid 1; and the linker sequence 2 is covalently linked to the 3' end of the complementary fragment 1 which is at least partially complementary to the single-stranded nucleic acid 1.

[0049] Further, the length of the linker sequence 1 is ≥1 nt, and the length of the linker sequence 2 is ≥1 nt; preferably, the length of the linker sequence 1 is 10-100 nt, and the length of the linker sequence 2 is 10-100 nt.

[0050] Further, the linker sequence 1 is selected from the sequence shown in 5'-XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX-SEQ ID NO: 6-YYYY-SEQ ID NO: 3-3', wherein X = iSpC3, the sequence of SEQ ID NO: 6 is TTTTTTTTTT, and Y = iSp18, the sequence of SEQ ID NO: 3 is GGTTGTTTCTGTTGGTGCTGATATTGCT; and the linker sequence 2 is selected from the sequence shown in SEQ ID NO: 7: 5'-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3', wherein the 5' end is phosphorylated.

[0051] Further, the motor protein is selected from a polymerase or a helicase; preferably, the polymerase is selected from any one of Bst DNA polymerase, SD DNA polymerase, phi29 DNA polymerase, Bsu Large Fragment DNA polymerase, Klenow Fragment DNA polymerase, T3 RNA polymerase, T7 RNA polymerase, SP6 RNA polymerase or E. coli RNA polymerase; preferably, the helicase is selected from any one of Dda, Hel308, RecD, UvrD, Rep, RecQ, PcrA, eIF4A, NS3, gp41, T7 gp4 or BCH105.

[0052] To achieve the above object, according to a fourth aspect of the present application, there is provided a polypeptide sequencing kit, the kit comprising: the chaperone molecule in the binary complex, and any one or more of the following optional components: a membrane, a nanopore, a single-stranded nucleic acid 1 in the polypeptide library, a complementary fragment 1, a linker sequence 1, a linker sequence 2 and a motor protein.

[0053] Further, the linker sequence 1, the linker sequence 2 and the motor protein are in the form of a linker complex; preferably, the nanopore is a protein nanopore or a solid-state nanopore;

[0054] More preferably, the protein nanopore is selected from any one or more of the following combinations: a-haemolysin, haemolysin, leukocidin, Mycobacterium smegmatis pore protein A (MspA), MspB, MspC, MspD, a-Haemolysin, CsgG, Aerolysin, cytolysin, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, Neisseria self-transporting lipoprotein (NalP), WZA, GspD, BCP34 and BCP58;

[0055] More preferably, the solid-state nanopore is selected from any one or more of the following combinations: graphene nanopore, gold nanopore, silicon nitride nanopore, silicon dioxide nanopore or aluminum oxide nanopore;

[0056] Preferably, the nanopore is located on and penetrates the membrane; preferably, the membrane is selected from any one or more of the following: a phospholipid membrane, a polymer membrane or a solid-state thin film.

[0057] To achieve the above object, according to a fifth aspect of the present application, there is provided a method for constructing a polypeptide library, the method comprising:

[0058] Preparation of the polypeptide to be tested into the binary complex;

[0059] Linking the binary complex with the double-stranded nucleic acid to obtain a double-stranded complex;

[0060] connecting the double-stranded complex with the double-stranded linker complex containing the motor protein to obtain the polypeptide library;

[0061] The double-stranded linker complex containing the motor protein comprises: a linker sequence 1, a linker sequence 2 complementary to the 3' end of the linker sequence 1 and non-complementary to the 5' end, and a motor protein located on the linker sequence 1.

[0062] Further, the double-stranded nucleic acid comprises a first segment of the single-stranded nucleic acid 1, a second segment of the single-stranded nucleic acid 1, and a complementary segment 1 complementary to at least part of the single-stranded nucleic acid 1.

[0063] Further, the construction method comprises: covalently connecting the two ends of the binary complex to the 3' end of the first segment of the single-stranded nucleic acid 1 and the 5' end of the second segment of the single-stranded nucleic acid 1, respectively; annealing the single-stranded nucleic acid 1 with the binary complex to the complementary segment 1 to form a double-stranded annealing complex; and connecting the double-stranded annealing complex with the linker complex containing the motor protein by a ligase to form the polypeptide library.

[0064] Further, the construction method comprises: annealing the single-stranded nucleic acid 1 with the complementary segment 1 to form a double-stranded nucleic acid; covalently connecting the two ends of the binary complex to the 3' end of the first segment of the single-stranded nucleic acid 1 and the 5' end of the second segment of the single-stranded nucleic acid 1 in the double-stranded nucleic acid to obtain a double-stranded annealing complex; and connecting the double-stranded annealing complex with the linker complex containing the motor protein by a ligase to form the polypeptide library.

[0065] Further, the preparation of the polypeptide to be tested into a binary complex comprises: covalently connecting a partner molecule and the polypeptide to be tested to obtain a binary complex; wherein the partner molecule is selected from any one or more of the following: PEG, spacer, deoxyribose phosphate, ribose phosphate, nucleotide, deoxynucleotide, peptide nucleic acid, or locked nucleotide; preferably, the spacer is selected from any one or more of the following: Spacer C3, Spacer C6, Spacer 9, Spacer C12, or Spacer 18; preferably, the average molecular weight of the PEG is 300-20000.

[0066] Further, the covalent connection of the partner molecule and the polypeptide to be tested comprises any one or more of any one of the following groups:

[0067] 1) one end of a polypeptide to be tested and one end of a partner molecule are single-end bonded;

[0068] 2) one end of a test polypeptide simultaneously binds to one end of two partner molecules, one end of a test polypeptide binds to a middle position of one partner molecule, or one end of a partner molecule simultaneously binds to one end of another partner molecule and one end of a test polypeptide;

[0069] 3) one end of a partner molecule simultaneously binds to one end of two test polypeptides, one end of a partner molecule binds to a middle position of one test polypeptide, or one end of a test polypeptide simultaneously binds to one end of a partner molecule and one end of a test polypeptide;

[0070] 4) one end of a complex element 1 simultaneously binds to one end of two partner molecules, one end of a complex element 1 binds to a middle position of one partner molecule, or one end of a partner molecule simultaneously binds to one end of a complex element 1 and one end of a partner molecule;

[0071] 5) one end of a complex element 1 simultaneously binds to one end of a partner molecule and one end of a test polypeptide, one end of a partner molecule simultaneously binds to one end of a complex element 1 and one end of a test polypeptide, or one end of a complex element 2 binds to one end of a test polypeptide;

[0072] 6) one test polypeptide binds to one partner molecule at both ends;

[0073] 7) both ends of a complex element 3 each bind to one end of one partner molecule;

[0074] 8) both ends of a complex element 3 each bind to one end of one test polypeptide;

[0075] 9) one end of a complex element 3 binds to one end of one partner molecule;

[0076] 10) one end of a complex element 3 binds to one end of one test polypeptide;

[0077] 11) one complex element 1 binds to one partner molecule at both ends;

[0078] 12) both ends of a complex element 4 each bind to one end of one partner molecule;

[0079] 13) one end of a complex element 4 binds to one end of one partner molecule;

[0080] 14) one complex element 4 is single-end bonded to either end of any one of the polypeptides to be tested;

[0081] 15) one complex element 4 is single-end bonded to either end of any one of the polypeptides to be tested;

[0082] wherein the partner molecules at different positions in any one of the binary complexes are the same or different, and the polypeptides to be tested at different positions in any one of the binary complexes are the same or different;

[0083] complex element 1 is a complex of two ends of a polypeptide to be tested each single-end bonded to either end of a partner molecule;

[0084] complex element 2 is a complex of either end of a complex element 1 single-end bonded to either end of a partner molecule;

[0085] complex element 3 is a complex of a polypeptide to be tested double-end bonded to a partner molecule;

[0086] complex element 4 is a complex of a complex element 1 double-end bonded to a partner molecule.

[0087] Further, the partner molecules and the polypeptides to be tested are covalently linked by any one or more of the following: peptide bond linkage, ester bond linkage, ether bond linkage, thiol-maleimide linkage, oxime formation of carbonyl-hydroxylamine-containing compounds, hydrazone formation of carbonyl-hydrazine-containing compounds, urea formation of carbonyl-urea- containing compounds, disulfide bond linkage, sulfide bond linkage, halogen-nucleophile substitution linkage, 1,3 dipolar cycloaddition reaction linkage, copper-catalyzed azide-alkynyl cycloaddition reaction linkage, ruthenium-catalyzed azide-alkynyl cycloaddition reaction linkage, azide compound-alkynyl compound click chemistry reaction linkage, or natural chemical linkage;

[0088] Preferably, the azide compound-alkynyl compound click chemistry reaction linkage comprises any one or more of the following: azide-DBCO click chemistry reaction linkage, azide-OCT click chemistry reaction linkage, azide-DIBO click chemistry reaction linkage, azide-BARAC click chemistry reaction linkage, azide-ALO click chemistry reaction linkage, azide-DIFO click chemistry reaction linkage, azide-MOFO click chemistry reaction linkage, azide-DIBAC click chemistry reaction linkage, azide-DIMAC click chemistry reaction linkage, or azide-cyclooctene click chemistry reaction linkage.

[0089] In order to achieve the above-mentioned purpose, according to a sixth aspect of the present application, a polypeptide sequencing method is provided, the sequencing method comprising: co-incubating the polypeptide library constructed by the construction method or the polypeptide library with an anchor sequence to obtain an incubation complex; adding the incubation complex into a sequencing solution tank, under the action of an electric field force, controlling the binary complex containing the polypeptide to be sequenced to pass through the nanopore by the motor protein, so as to obtain the corresponding electric signal of the polypeptide to be sequenced; decoding the electric signal to determine the amino acid sequence of the polypeptide to be sequenced.

[0090] Further, one end of the anchor sequence is complementarily paired with the end of the linker sequence 2 in the linker complex away from the complementary fragment 1, and the other end is provided with an anchor group; preferably, the anchor group is selected from any one of a lipid, a carbon nanotube, a polypeptide, a protein and / or an amino acid; preferably, the lipid is selected from any one of a fatty acid, a sterol, a palmitate or a tocopherol; preferably, the anchor sequence is 5'-Chol-TEG-TT-YYYY-SEQ ID NO:8, wherein Y = iSp18, Chol-TEG represents cholesterol-polyethylene glycol, and the sequence of SEQ ID NO:8 is 5'-TTGACCGCTCGCCTC-3'.

[0091] Further, the nanopore is a protein nanopore or a solid-state nanopore; preferably, the protein nanopore is selected from any one or a combination of the following: alpha-hemolysin, hemolysin, leukocidin, Mycobacterium smegmatis porin A (MspA), MspB, MspC, MspD, alpha-Haemolysin, CsgG, Aerolysin, cytolysin, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, Neisseria self-transporting lipoprotein (NalP), WZA, GspD, BCP34 and BCP58; preferably, the solid-state nanopore is selected from any one or a combination of the following: a graphene nanopore, a gold nanopore, a silicon nitride nanopore, a silicon dioxide nanopore or an aluminum oxide nanopore.

[0092] By using the technical solution of the present application, the double (multiple) molecules are adapted to the nanopore channel in the spatial scale by the method of sequencing by combining the polypeptide molecules with the chaperone molecules, which can greatly improve the interaction probability between the amino acid residues in the polypeptide and the nanopore, and on the other hand, under the assistance of the chaperone molecules, the polypeptide can pass through the nanopore at a stable speed, so that a sequencing signal with good reproducibility and consistency is obtained, and the resolution of the polypeptide nanopore sequencing is greatly improved. The development and application of the present application are helpful to improve the resolution of the nanopore polypeptide sequencing, and have important significance for the development of the nanopore polypeptide sequencing technology. BRIEF DESCRIPTION OF DRAWINGS

[0093] The accompanying drawings, which form a part of this specification, are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification. The embodiments of the application, and their

[0094] Figure 1 shows a comparison of polypeptide signals obtained by passing polypeptide fragments resulting from enzymatic digestion of a test protein with trypsin through a nanopore with the electrical signals of model polypeptides.

[0095] Figure 2 shows a sequencing signal diagram for sequencing a single strand of "nucleic acid-polypeptide-nucleic acid" using MTA helicase by the Whitehead Institute for Biomedical Research.

[0096] Figure 3 shows a schematic diagram of a single end ligation of either end of a test polypeptide with either end of a partner molecule.

[0097] Figure 4 shows a schematic diagram of a single end ligation of either end of a test polypeptide with either end of a partner molecule, a single end ligation of either end of a test polypeptide with a middle position of a partner molecule, or a single end ligation of either end of a partner molecule with either end of another partner molecule and either end of a test polypeptide.

[0098] Figure 5 shows a schematic diagram of a single end ligation of either end of a partner molecule with either end of a test polypeptide, a single end ligation of either end of a partner molecule with a middle position of a test polypeptide, or a single end ligation of either end of a test polypeptide with either end of a partner molecule and either end of a test polypeptide.

[0099] Figure 6 shows a schematic diagram of a single end ligation of either end of a complex element 1 with either end of a partner molecule, a single end ligation of a complex element 1 with a middle position of a partner molecule, or a single end ligation of either end of a partner molecule with either end of a complex element 1 and either end of a partner molecule, wherein the complex element 1 is a complex of a test polypeptide with either end of a partner molecule.

[0100] Figure 7 shows a schematic diagram of a single end ligation of either end of a complex element 1 with either end of a partner molecule and either end of a test polypeptide, a single end ligation of either end of a partner molecule with either end of a complex element 1 and either end of a test polypeptide, or a single end ligation of either end of a complex element 2 with either end of a test polypeptide, wherein the complex element 1 is as in Figure 6 and the complex element 2 is a complex of either end of a complex element 1 with either end of a partner molecule.

[0101] Figure 8 shows a schematic diagram of two-end bonding of a polypeptide to be tested to a partner molecule according to the present application.

[0102] Figure 9 shows a schematic diagram of single-end bonding of each of the two ends of a complex element 3 to any one end of a partner molecule according to the present application, wherein the complex element 3 is a complex of two-end bonding of a polypeptide to be tested to a partner molecule.

[0103] Figure 10 shows a schematic diagram of single-end bonding of each of the two ends of a complex element 3 to any one end of a polypeptide to be tested according to the present application, wherein the complex element 3 is the same as that of Figure 9.

[0104] Figure 11 shows a schematic diagram of single-end bonding of any one end of a complex element 3 to any one end of a partner molecule according to the present application, wherein the complex element 3 is the same as that of Figure 9.

[0105] Figure 12 shows a schematic diagram of single-end bonding of any one end of a complex element 3 to any one end of a polypeptide to be tested according to the present application, wherein the complex element 3 is the same as that of Figure 9.

[0106] Figure 13 shows a schematic diagram of two-end bonding of a complex element 1 to a partner molecule according to the present application, wherein the complex element 1 is the same as that of Figure 6.

[0107] Figure 14 shows a schematic diagram of single-end bonding of each of the two ends of a complex element 4 to any one end of a partner molecule according to the present application, wherein the complex element 4 is a complex of two-end bonding of a complex element 1 to a partner molecule, and the complex element 1 is the same as that of Figure 6.

[0108] Figure 15 shows a schematic diagram of single-end bonding of any one end of a complex element 4 to any one end of a partner molecule according to the present application, wherein the complex element 4 is the same as that of Figure 14.

[0109] Figure 16 shows a schematic diagram of single-end bonding of each of the two ends of a complex element 4 to any one end of a polypeptide to be tested according to the present application, wherein the complex element 4 is the same as that of Figure 14.

[0110] Figure 17 shows a schematic diagram of single-end bonding of any one end of a complex element 4 to any one end of a polypeptide to be tested according to the present application, wherein the complex element 4 is the same as that of Figure 14.

[0111] Figure 18 shows a schematic diagram of a binary complex passing through a nanopore according to the first preferred embodiment of the present application, wherein any one end of a polypeptide to be tested and any one end of a partner molecule are single-end bonded to form the binary complex.

[0112] Figure 19 shows a schematic diagram of the way a binary complex passes through a nanopore in a second preferred embodiment of the application, wherein the same end of one polypeptide under test simultaneously binds to either end of two partner molecules, either end of one polypeptide under test binds to the middle of one partner molecule, or the same end of one partner molecule simultaneously binds to either end of another partner molecule and either end of one polypeptide under test to form the binary complex.

[0113] Figure 20 shows a schematic diagram of the way a binary complex passes through a nanopore in a third preferred embodiment of the application, wherein the same end of one partner molecule simultaneously binds to either end of two polypeptides under test, either end of one partner molecule binds to the middle of one polypeptide under test, or the same end of one polypeptide under test simultaneously binds to either end of one partner molecule and either end of one polypeptide under test to form the binary complex.

[0114] Figure 21 shows a schematic diagram of the way a binary complex passes through a nanopore in a fourth preferred embodiment of the application, wherein the same end of one complexing element 1 simultaneously binds to either end of two partner molecules, either end of one complexing element 1 binds to the middle of one partner molecule, or the same end of one partner molecule simultaneously binds to either end of one complexing element 1 and either end of one partner molecule to form the binary complex, wherein complexing element 1 is as in Figure 6.

[0115] Figure 22 shows a schematic diagram of the way a binary complex passes through a nanopore in a fifth preferred embodiment of the application, wherein the same end of one complexing element 1 simultaneously binds to either end of one partner molecule and either end of one polypeptide under test, the same end of one partner molecule simultaneously binds to either end of one complexing element 1 and either end of one polypeptide under test, or either end of one complexing element 2 binds to either end of one polypeptide under test to form the binary complex, wherein complexing element 1 is as in Figure 6 and complexing element 2 is as in Figure 7.

[0116] Figure 23 shows a schematic diagram of the way a binary complex passes through a nanopore in a sixth preferred embodiment of the application, wherein one polypeptide under test binds to one partner molecule at both ends to form the binary complex.

[0117] Figure 24 shows a schematic diagram of the way a binary complex passes through a nanopore in a seventh preferred embodiment of the application, wherein both ends of one complexing element 3 each bind to either end of one partner molecule to form the binary complex, wherein complexing element 3 is as in Figure 9.

[0118] Figure 25 shows a schematic diagram of the way a binary complex passes through a nanopore in an eighth preferred embodiment of the present application, in which one end of each of two complexing elements 3 is single-end bonded to either end of a polypeptide to be tested to form the binary complex, in which the complexing element 3 is as in Figure 9.

[0119] Figure 26 shows a schematic diagram of the way a binary complex passes through a nanopore in a ninth preferred embodiment of the present application, in which one end of a complexing element 3 is single-end bonded to either end of a partner molecule to form the binary complex, in which the complexing element 3 is as in Figure 9.

[0120] Figure 27 shows a schematic diagram of the way a binary complex passes through a nanopore in a tenth preferred embodiment of the present application, in which one end of a complexing element 3 is single-end bonded to either end of a polypeptide to be tested to form the binary complex, in which the complexing element 3 is as in Figure 9.

[0121] Figure 28 shows a schematic diagram of the way a binary complex passes through a nanopore in an eleventh preferred embodiment of the present application, in which one end of a complexing element 1 is two-end bonded to a partner molecule to form the binary complex, in which the complexing element 1 is as in Figure 6.

[0122] Figure 29 shows a schematic diagram of the way a binary complex passes through a nanopore in a twelfth preferred embodiment of the present application, in which one end of each of two complexing elements 4 is single-end bonded to either end of a partner molecule to form the binary complex, in which the complexing element 4 is as in Figure 14.

[0123] Figure 30 shows a schematic diagram of the way a binary complex passes through a nanopore in a thirteenth preferred embodiment of the present application, in which one end of a complexing element 4 is single-end bonded to either end of a partner molecule to form the binary complex, in which the complexing element 4 is as in Figure 14.

[0124] Figure 31 shows a schematic diagram of the way a binary complex passes through a nanopore in a fourteenth preferred embodiment of the present application, in which one end of each of two complexing elements 4 is single-end bonded to either end of a polypeptide to be tested to form the binary complex, in which the complexing element 4 is as in Figure 14.

[0125] Figure 32 shows a schematic diagram of the way a binary complex passes through a nanopore in a fifteenth preferred embodiment of the present application, in which one end of a complexing element 4 is single-end bonded to either end of a polypeptide to be tested to form the binary complex, in which the complexing element 4 is as in Figure 14.

[0126] Figure 33 shows a schematic diagram of a nanopore used in polypeptide nanopore sequencing using the binary complex of the present application, wherein A is a biological nanopore (such as a protein nanopore) and B is a solid state nanopore.

[0127] Figure 34 shows a schematic diagram of the interaction forces that can occur between the polypeptide to be sequenced and the partner molecule of the present application.

[0128] Figure 35 shows a schematic diagram of a partner molecule of the present application.

[0129] Figure 36 shows a schematic diagram of a polypeptide library constructed using a binary complex of the present application.

[0130] Figure 37 shows sequencing signal graphs for four blank control libraries (i.e. DNA double stranded templates without coupled polypeptides) of Example 5 of the present application, wherein a represents a partner molecule double end ligated DNA double stranded template, b represents a partner molecule 5' end ligated DNA double stranded template, c represents a partner molecule 3' end ligated DNA double stranded template, and d represents a DNA double stranded template not ligated to a partner molecule.

[0131] Figure 38 shows sequencing signal graphs for (a) a blank library of DNA1-DNA5 partner molecule templates (i.e. containing only partner molecules, without coupled polypeptides), (b) a D1P1-D5 library, (c) a D1P2-D5 library, and (d) a control library of the prior art conventional "DNA1-Peptide-DNA2" sandwich structure of Example 8 of the present application.

[0132] Figure 39 shows sequencing signal graphs for (a) D9P1-D10 and (b) D9P2-D10 of Example 9 of the present application.

[0133] Figure 40 shows a schematic diagram of a single-stranded nucleic acid 1 and a complementary fragment 1 that are completely complementary to each other in one embodiment of the present application.

[0134] Figure 41 shows a schematic diagram of a complementary fragment 1 that is completely complementary to a first segment of a single-stranded nucleic acid 1 and partially complementary to a second segment of the single-stranded nucleic acid 1 to form a single overhang in one embodiment of the present application.

[0135] Figure 42 shows a schematic diagram of a complementary fragment 1 that is completely complementary to a first segment of a single-stranded nucleic acid 1 and partially complementary to a second segment of the single-stranded nucleic acid 1 to form two overhangs in one embodiment of the present application.

[0136] Figure 43 shows a schematic diagram of a complementary fragment 1 that is completely complementary to a first segment of a single-stranded nucleic acid 1 and completely non-complementary to a second segment of the single-stranded nucleic acid 1 to form a single overhang in one embodiment of the present application.

[0137] Figure 44 shows a schematic diagram of the complete complementary pairing of the complementary fragment 1 with the first segment of the single-stranded nucleic acid 1 and the complete non-complementary pairing of the complementary fragment 1 with the second segment of the single-stranded nucleic acid 1, resulting in two protruding ends in one embodiment of the present application. DETAILED DESCRIPTION

[0138] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the embodiments.

[0139] Terminology:

[0140] The constituent unit of the partner molecule: the constituent unit of the partner molecule in the present application refers to the component that constitutes the partner molecule. It can be a plurality of identical single components, such as 6 Spacer C3, the number of constituent units is 6; such as 8 abasic deoxynucleotides, the number of constituent units is 8; or a plurality of different components, such as 4 Spacer C3+2 Spacer C6, the number of constituent units is 6.

[0141] As mentioned in the background, the resolution of various nanopore polypeptide sequencing technical solutions disclosed in the prior art is low. In order to improve this situation, the present application designs a partner molecule assisted protein nanopore sequencing technology. By covalently linking the polypeptide signal and the partner molecule to form a binary complex, and then fixing the binary complex on the single strand of the DNA double strand, the simultaneous translocation of the polypeptide and the partner molecule through the nanopore and the signal capture are realized. The specific implementation steps are as follows:

[0142] (1) Construct the sequence to be sequenced containing both the partner molecule and the polypeptide molecule by covalent linkage, and pair it with the complementary sequence to form a complex that can be used to prepare a nanopore polypeptide sequencing library.

[0143] (2) Link the complex in (1) with the pre-prepared linker complex through the ligase to form a polypeptide nanopore sequencing library.

[0144] (3) After co-incubation of the polypeptide nanopore sequencing library in (2) with the anchor sequence, add the solution bin of the sequencing chip. Under the action of the electric field force, through the continuous unwinding reaction of the DNA double strand, the simultaneous translocation of the polypeptide and the partner molecule through the nanopore is stably controlled, and the electrical signal is obtained.

[0145] (4) Algorithm analysis is performed on the obtained electrical signal.

[0146] The inventors found that the resolution of amino acids was significantly improved by analyzing the electrical signals of different amino acids passing through the nanopore. Based on this research result, it is expected to realize the accurate analysis of the composition of amino acids in the polypeptide by using the technical solution of the present application.

[0147] Based on the above research results, the applicant of the present application proposes a series of protection schemes. In a first typical embodiment of the present application, a binary complex is provided, comprising a covalently linked chaperone molecule and a polypeptide to be tested, wherein the chaperone molecule (see Figure 35) is selected from any one or more of the following: PEG, spacer, deoxyribose phosphate, ribose phosphate, nucleotide, deoxynucleotide, peptide nucleic acid or locked nucleotide; wherein the number of constituent units of the chaperone molecule is ≥ 1; and the number of amino acids of the polypeptide to be tested is ≥ 2. Those skilled in the art can understand that the number of constituent units of the chaperone molecule and the number of amino acids of the polypeptide are generally integers.

[0148] The present application forms a binary complex by covalently linking a chaperone molecule of different spatial size and length with a polypeptide to be tested, so that it approaches or reaches the matching with the nanopore at the molecular scale. The size refers to the size of each unit constituting the chaperone molecule, and the length refers to the number of constituent units of the chaperone molecule. The chaperone molecule can be a nucleic acid backbone or a high-molecular structure with negative electricity, which can form a stable interaction with the motor protein, so that the speed control of the motor protein is stable and continuous. The use of the binary complex of the present application for polypeptide nanopore sequencing can make the amino acid side chain have a greater probability of interacting with the nanopore channel, thereby improving the resolution of the signal.

[0149] According to the differences in the number and connection position of the chaperone molecule and the polypeptide to be tested, the covalent connection form of the chaperone molecule and the polypeptide to be tested includes but is not limited to any one or more of any one group as follows:

[0150] 1) Single-end bonding of any one end of a polypeptide to be tested and any one end of a chaperone molecule (see Figure 3);

[0151] 2) Single-end bonding of the same end of a polypeptide to be tested with any one end of two chaperone molecules, single-end bonding of any one end of a polypeptide to the middle position of a chaperone molecule, or single-end bonding of the same end of a chaperone molecule with any one end of another chaperone molecule and any one end of a polypeptide to be tested (see Figure 4);

[0152] 3) Single-end bonding of the same end of a chaperone molecule with any one end of two polypeptides to be tested, single-end bonding of any one end of a chaperone molecule to the middle position of a polypeptide to be tested, or single-end bonding of the same end of a polypeptide to be tested with any one end of a chaperone molecule and any one end of a polypeptide to be tested (see Figure 5);

[0153] 4) one end of one complex element 1 is simultaneously single-end bonded to any one end of two partner molecules, one complex element 1 is single-end bonded to the middle of one partner molecule, or one end of one partner molecule is simultaneously single-end bonded to any one end of one complex element 1 and any one end of one partner molecule (see Figure 6);

[0154] 5) one end of one complex element 1 is simultaneously single-end bonded to any one end of one partner molecule and any one end of one polypeptide to be tested, one end of one partner molecule is simultaneously single-end bonded to any one end of one complex element 1 and any one end of one polypeptide to be tested, or any one end of one complex element 2 is single-end bonded to any one end of one polypeptide to be tested (see Figure 7);

[0155] 6) one polypeptide to be tested is double-end bonded to one partner molecule (see Figure 8);

[0156] 7) two ends of one complex element 3 are respectively single-end bonded to any one end of one partner molecule (see Figure 9);

[0157] 8) two ends of one complex element 3 are respectively single-end bonded to any one end of one polypeptide to be tested (see Figure 10);

[0158] 9) any one end of one complex element 3 is single-end bonded to any one end of one partner molecule (see Figure 11);

[0159] 10) any one end of one complex element 3 is single-end bonded to any one end of one polypeptide to be tested (see Figure 12);

[0160] 11) one complex element 1 is double-end bonded to one partner molecule (see Figure 13);

[0161] 12) two ends of one complex element 4 are respectively single-end bonded to any one end of one partner molecule (see Figure 14);

[0162] 13) any one end of one complex element 4 is single-end bonded to any one end of one partner molecule (see Figure 15);

[0163] 14) two ends of one complex element 4 are respectively single-end bonded to any one end of one polypeptide to be tested (see Figure 16);

[0164] 15) any one end of one complex element 4 is single-end bonded to any one end of one polypeptide to be tested (see Figure 17);

[0165] wherein the partner molecules at different positions in any form of binary complex are the same or different, and the polypeptides to be tested at different positions in any form of binary complex are the same or different;

[0166] The complex element 1 is a complex of the test polypeptide and the partner molecule after single-end bonding of the test polypeptide and the partner molecule at one end of the test polypeptide and one end of the partner molecule respectively;

[0167] The complex element 2 is a complex of the complex element 1 and the partner molecule after single-end bonding of one end of the complex element 1 and one end of the partner molecule;

[0168] The complex element 3 is a complex of the test polypeptide and the partner molecule after two-end bonding of the test polypeptide and the partner molecule;

[0169] The complex element 4 is a complex of the complex element 1 and the partner molecule after two-end bonding of the complex element 1 and the partner molecule.

[0170] In the above-mentioned various binary complexes, the test polypeptide and the partner molecule can be connected in various ways, including but not limited to covalent connection in any one or more of the following ways: peptide bond connection, ester bond connection, ether bond connection, thiol-maleimide connection, oxime formation of carbonyl-hydroxylamine-containing compounds, hydrazone formation of carbonyl-hydrazine-containing compounds, urea formation of carbonyl-urea-containing compounds, disulfide bond connection, sulfide bond connection, halogen-nucleophile substitution connection, 1,3 dipolar cycloaddition reaction connection, copper-catalyzed azide-alkynyl cycloaddition reaction connection, ruthenium-catalyzed azide-alkynyl cycloaddition reaction connection, azide compound-alkynyl compound click chemistry reaction connection, or natural chemical connection.

[0171] In some preferred embodiments, the azide compound-alkynyl compound click chemistry reaction connection includes but is not limited to any one or more of the following: azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection, or azide-cyclooctene click chemistry reaction connection. This connection method has the advantages of short reaction time, high coupling efficiency, high product conversion rate, and stable chemical structure after connection.

[0172] In a preferred embodiment of the present application, when the partner molecule contains a spacer, the spacer is selected from any one or more of the following: Spacer C3, Spacer C6, Spacer 9, Spacer C12, or Spacer 18. Using a spacer as a partner molecule has the advantages of easy synthesis, providing sufficient length and spatial size for the test polypeptide, and facilitating smooth penetration of the binary complex.

[0173] In another preferred embodiment of the present application, the chaperone molecule contains PEG, and the average molecular weight of the PEG is 300-20000; preferably, the PEG is selected from any one or more of PEG-300, PEG-400, PEG-800, PEG-1000, PEG-1500, PEG-2000, PEG-3000, PEG-4000, PEG-6000, PEG-8000 or PEG-20000. The use of PEG as the chaperone molecule helps to provide sufficient length and spatial size for the polypeptide to be detected, and to enhance the affinity of the chaperone molecule to the polypeptide to be detected.

[0174] In a second typical embodiment of the present application, a double-stranded annealing complex is provided, which comprises: a single-stranded nucleic acid 1 and a complementary fragment 1 which is at least partially complementary to the single-stranded nucleic acid 1, wherein the middle position of the single-stranded nucleic acid 1 is covalently linked to the above-mentioned binary complex. The use of this double-stranded annealing complex to construct a nucleic acid-polypeptide library for nanopore sequencing has the advantages of facilitating connection with a double-stranded sequencing adapter containing a motor protein, and providing stable speed control for the double-stranded annealing complex during sequencing.

[0175] It should be noted that the single-stranded nucleic acid 1 and the complementary fragment 1 can be completely complementary (see Figure 40), or partially complementary, including one protruding end (see Figures 41 and 43) or two protruding ends (see Figures 42 and 44).

[0176] The first segment of the single-stranded nucleic acid 1 and the second segment of the single-stranded nucleic acid 1 are respectively covalently linked to the two ends of the binary complex; the covalent linkage can be achieved in various ways, wherein the covalent linkage is selected from any one or more of peptide bond linkage, ester bond linkage, ether bond linkage, thiol-maleimide linkage, oxime formation of carbonyl-hydroxylamine-containing compounds, hydrazone formation of carbonyl-hydrazine-containing compounds, urea formation of carbonyl-urea-containing compounds, disulfide bond linkage, sulfide bond linkage, halogen-nucleophile substitution linkage, 1,3 dipolar cycloaddition reaction linkage, copper-catalyzed azide-alkynyl cycloaddition reaction linkage, ruthenium-catalyzed azide-alkynyl cycloaddition reaction linkage, azide compound-alkynyl compound click chemistry reaction linkage or natural chemical linkage; preferably, the covalent linkage is diester bond linkage; more preferably, the covalent linkage is phosphodiester bond linkage. The use of this preferred covalent linkage has the advantages of low synthesis difficulty and easy availability.

[0177] In some preferred embodiments, the azide compound-alkynyl compound click chemistry reaction connection includes, but is not limited to: azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection, azide-cyclooctene click chemistry reaction connection. The above-mentioned covalent connection has the advantages of short reaction time, high coupling efficiency, high product conversion rate, and stable chemical structure after connection.

[0178] In a preferred embodiment of the present application, the structural formula of the single-stranded nucleic acid 1 is: first segment of single-stranded nucleic acid 1-int DBCO modified deoxyribonucleotide-binary complex-int DBCO modified deoxyribonucleotide-second segment of single-stranded nucleic acid 1.

[0179] The length of the first segment of the single-stranded nucleic acid 1 is ≥1 nt, the length of the second segment of the single-stranded nucleic acid 1 is ≥1 nt, and the length of the complementary fragment 1 is ≥1 nt. Preferably, the length of the first segment of the single-stranded nucleic acid 1 is 5-500 nt, the length of the second segment of the single-stranded nucleic acid 1 is 5-500 nt, and the length of the complementary fragment 1 is 5-1200 nt. The single-stranded nucleic acid with the above-mentioned length has the advantages of significant sequencing signal characteristics, appropriate sequencing read length, and easy data analysis.

[0180] The 5' end of the first segment of the single-stranded nucleic acid 1 is provided with a phosphorylation group, and the 3' end of the second segment of the single-stranded nucleic acid 1 is provided with any one or more of the following dideoxynucleotides: ddA, ddT, ddC or ddG. The dideoxynucleotide has better chemical stability, which can reduce error signals and interference in the sequencing process, thereby improving the quality and reliability of the data. In addition, the dideoxynucleotide can also reduce base modification and base mismatch in the sequencing process, which helps to improve the accuracy of sequencing.

[0181] Further, the first segment of the single-stranded nucleic acid 1 is selected from the sequence shown in SEQ ID NO: 1: 5'-GCTTCTCGTGTTTTTTTTCTCTC-3'; the second segment of the single-stranded nucleic acid 1 is selected from the sequence shown in SEQ ID NO: 2: 5'-CCCTTTTTTTTTTGCTGTCTTCTGTCGTCGTTT-3'; and the complementary fragment 1 is selected from the sequence shown in SEQ ID NO: 5: 5'-GAAACGACGACAGAAGACAGCAAAAAAAAAAGGGATTTTTTAGAGAGAAAAAAAAC GAGAAGCA-3'.

[0182] In a third typical embodiment of the present application, a polypeptide library is provided, which comprises the above-mentioned binary complex or the above-mentioned double-stranded annealed complex.

[0183] In a preferred embodiment of the present application, the polypeptide library comprises the double-stranded annealed complex and a linker complex covalently linked to the double-stranded annealed complex;

[0184] The linker complex comprises: a linker sequence 1 and a linker sequence 2, and a motor protein, the linker sequence 1 comprises a first segment and a second segment connected in sequence from 5' end to 3' end direction, wherein the first segment of the linker sequence 1 is not complementary to the linker sequence 2, the second segment of the linker sequence 1 is complementary to the linker sequence 2, and the motor protein is movably bound to the first segment of the linker sequence 1; the linker sequence 1 is covalently linked to the 5' end of the single-stranded nucleic acid 1; and the linker sequence 2 is covalently linked to the 3' end of the complementary fragment 1 which is at least partially complementary to the single-stranded nucleic acid 1 (see FIG. 36).

[0185] The length of the linker sequence 1 is ≥1 nt, and the length of the linker sequence 2 is ≥1 nt; preferably, the length of the linker sequence 1 is 10-100 nt, and the length of the linker sequence 2 is 10-100 nt.

[0186] The linker sequence 1 is selected from the sequence shown in 5'-XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX-SEQ ID NO: 6-YYYY-SEQ ID NO: 3-3', wherein X = iSpC3, the sequence of SEQ ID NO: 6 is TTTTTTTTTT, and Y = iSp18, the sequence of SEQ ID NO: 3 is GGTTGTTTCTGTTGGTGCTGATATTGCT; and the linker sequence 2 is selected from the sequence shown in SEQ ID NO: 7: 5'-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3', wherein the 5' end is modified with phosphorylation.

[0187] The motor protein is selected from a polymerase or a helicase; preferably, the polymerase is selected from any one of Bst DNA polymerase, SD DNA polymerase, phi29 DNA polymerase, Bsu Large Fragment DNA polymerase, Klenow Fragment DNA polymerase, T3 RNA polymerase, T7 RNA polymerase, SP6 RNA polymerase, or E. coli RNA polymerase; and preferably, the helicase is selected from any one of Dda, Hel308, RecD, UvrD, Rep, RecQ, PcrA, eIF4A, NS3, gp41, T7gp4, or BCH105.

[0188] In a fourth exemplary embodiment of the present application, a polypeptide sequencing kit is provided, which comprises the chaperone molecule in the binary complex described above, and any one or more of the following optional components: a membrane, a nanopore, a single-stranded nucleic acid 1, a complementary fragment 1, a linker sequence 1, a linker sequence 2, and a motor protein in the polypeptide library described above.

[0189] In the kit described above, the linker sequence 1, the linker sequence 2, and the motor protein are preferably present in the form of a linker complex, which is more convenient for connection.

[0190] The nanopore in the kit described above is a biological nanopore (such as a protein nanopore) or a solid-state nanopore.

[0191] The specific protein nanopore is selected from any one or more combinations of the following: a-haemolysin, haemolysin, leukocidin, Mycobacterium smegmatis porin A (MspA), MspB, MspC, MspD, a-Haemolysin, CsgG, Aerolysin, cytolysin, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, Neisseria self-transporting lipoprotein (NalP), WZA, GspD, BCP34, and BCP58.

[0192] The solid-state nanopore is selected from any one or more combinations of the following: a graphene nanopore, a gold nanopore, a silicon nitride nanopore, a silicon dioxide nanopore, or an aluminum oxide nanopore.

[0193] In some preferred embodiments, the nanopore is located on and penetrates the membrane; more preferably, the membrane is selected from any one or more of the following: a phospholipid membrane, a polymer membrane, or a solid-state thin film.

[0194] Polypeptide sequencing using the kit of the present application has the beneficial effects of simple operation and high resolution of individual amino acids in the polypeptide to be tested.

[0195] In a fifth exemplary embodiment of the present application, a method for constructing a polypeptide library is provided, which comprises: preparing a polypeptide to be tested into a binary complex described above; connecting the binary complex with a double-stranded nucleic acid to obtain a double-stranded complex; connecting the double-stranded complex with a double-stranded linker complex containing a motor protein to obtain a polypeptide library; wherein the double-stranded linker complex containing the motor protein comprises: a linker sequence 1, a linker sequence 2 complementary to the 3' end of the linker sequence 1 and non-complementary to the 5' end, and a motor protein located on the linker sequence 1.

[0196] The application provides a polypeptide library construction method, which comprises the following steps: preparing a binary complex of a polypeptide to be detected, connecting the binary complex to a double-stranded nucleic acid, and connecting a sequencing adaptor to the double-stranded nucleic acid, so that the polypeptide to be detected in the binary complex can better contact the wall of a nanopore, the interaction between the amino acid in the polypeptide to be detected and the nanopore constriction is enhanced, the electrical signal characteristics of a single polypeptide are improved, and the resolution of a single amino acid is improved.

[0197] In a preferred embodiment of the application, the double-stranded nucleic acid comprises a first segment of the single-stranded nucleic acid 1, a second segment of the single-stranded nucleic acid 1, and a complementary fragment 1 which is at least partially complementary to the single-stranded nucleic acid 1.

[0198] The polypeptide library can be prepared by annealing (the single-stranded nucleic acid 1 is at least partially complementary to the complementary fragment 1) first and then connecting (the binary complex is connected to the double-stranded nucleic acid), or by connecting (the binary complex is connected to the single-stranded nucleic acid 1) first and then annealing (the single-stranded nucleic acid 1 is at least partially complementary to the complementary fragment 1).

[0199] In a preferred embodiment of the application, the construction method comprises the following steps: covalently connecting the two ends of the binary complex to the 3' end of the first segment of the single-stranded nucleic acid 1 and the 5' end of the second segment of the single-stranded nucleic acid 1, respectively; annealing the single-stranded nucleic acid 1 connected with the binary complex and the complementary fragment 1 to form a double-stranded annealed complex; and connecting the double-stranded annealed complex and the adaptor complex comprising the motor protein by using a ligase to form the polypeptide library.

[0200] In another preferred embodiment of the application, the construction method comprises the following steps: annealing the single-stranded nucleic acid 1 and the complementary fragment 1 to form a double-stranded nucleic acid; covalently connecting the two ends of the binary complex to the 3' end of the first segment of the single-stranded nucleic acid 1 and the 5' end of the second segment of the single-stranded nucleic acid 1 in the double-stranded nucleic acid to obtain a double-stranded annealed complex; and connecting the double-stranded annealed complex and the adaptor complex comprising the motor protein by using a ligase to form the polypeptide library.

[0201] The preparation of the binary complex of the polypeptide to be detected comprises the following steps: covalently combining a chaperone molecule and the polypeptide to be detected to obtain the binary complex; and the chaperone molecule is selected from any one or more of the following: PEG, a spacer, deoxyribose phosphate, ribose phosphate, a nucleotide, a deoxyribonucleotide, a peptide nucleic acid or a locked nucleotide; preferably, the spacer is selected from any one or more of the following: Spacer C3, Spacer C6, Spacer 9, Spacer C12 or Spacer 18; and preferably, the average molecular weight of the PEG is 300-20,000.

[0202] The covalent connection of the chaperone molecule and the polypeptide to be detected can be in any one or more of any one of the following groups:

[0203] 1) one end of a test polypeptide and one end of a partner molecule are single-end bonded;

[0204] 2) the same end of a test polypeptide is single-end bonded to either end of two partner molecules, one end of a test polypeptide is single-end bonded to a middle position of a partner molecule, or the same end of a partner molecule is single-end bonded to either end of another partner molecule and one end of a test polypeptide;

[0205] 3) the same end of a partner molecule is single-end bonded to either end of two test polypeptides, one end of a partner molecule is single-end bonded to a middle position of a test polypeptide, or the same end of a test polypeptide is single-end bonded to either end of a partner molecule and one end of a test polypeptide;

[0206] 4) the same end of a complex element 1 is single-end bonded to either end of two partner molecules, a complex element 1 is single-end bonded to a middle position of a partner molecule, or the same end of a partner molecule is single-end bonded to either end of a complex element 1 and one end of a partner molecule;

[0207] 5) the same end of a complex element 1 is single-end bonded to either end of a partner molecule and one end of a test polypeptide, the same end of a partner molecule is single-end bonded to either end of a complex element 1 and one end of a test polypeptide, or one end of a complex element 2 is single-end bonded to one end of a test polypeptide;

[0208] 6) a test polypeptide is two-end bonded to a partner molecule;

[0209] 7) both ends of a complex element 3 are single-end bonded to either end of a partner molecule;

[0210] 8) both ends of a complex element 3 are single-end bonded to either end of a test polypeptide;

[0211] 9) one end of a complex element 3 is single-end bonded to either end of a partner molecule;

[0212] 10) one end of a complex element 3 is single-end bonded to either end of a test polypeptide;

[0213] 11) a complex element 1 is two-end bonded to a partner molecule;

[0214] 12) both ends of a complex element 4 are single-end bonded to either end of a partner molecule;

[0215] 13) one end of a complexing element 4 is single-end bonded with one end of a partner molecule;

[0216] 14) two ends of a complexing element 4 are each single-end bonded with one end of a polypeptide to be tested;

[0217] 15) one end of a complexing element 4 is single-end bonded with one end of a polypeptide to be tested;

[0218] wherein the partner molecules at different positions in any one form of the binary complex are the same or different, and the polypeptides to be tested at different positions in any one form of the binary complex are the same or different;

[0219] The complexing element 1 is a complex in which two ends of a polypeptide to be tested are each single-end bonded with one end of a partner molecule;

[0220] The complexing element 2 is a complex in which one end of a complexing element 1 is single-end bonded with one end of a partner molecule;

[0221] The complexing element 3 is a complex in which a polypeptide to be tested is two-end bonded with a partner molecule;

[0222] The complexing element 4 is a complex in which a complexing element 1 is two-end bonded with a partner molecule.

[0223] Different binary complexes are different in the way of passing through the nanopore. In a first preferred embodiment of the present application, the binary complex passes through the nanopore as shown in Fig. 18, in which one end of a polypeptide to be tested is single-end bonded with one end of a partner molecule to form the binary complex.

[0224] In a second preferred embodiment of the present application, the binary complex passes through the nanopore as shown in Fig. 19, in which the same end of a polypeptide to be tested is simultaneously single-end bonded with one end of two partner molecules, one end of a polypeptide to be tested is single-end bonded with a middle position of a partner molecule, or the same end of a partner molecule is simultaneously single-end bonded with one end of another partner molecule and one end of a polypeptide to be tested to form the binary complex.

[0225] In a third preferred embodiment of the present application, the binary complex passes through the nanopore as shown in Fig. 20, in which the same end of a partner molecule is simultaneously single-end bonded with one end of two polypeptides to be tested, one end of a partner molecule is single-end bonded with a middle position of a polypeptide to be tested, or the same end of a polypeptide to be tested is simultaneously single-end bonded with one end of a partner molecule and one end of a polypeptide to be tested to form the binary complex.

[0226] In a fourth preferred embodiment of the present application, a binary complex is formed by nano-pore as shown in Figure 21, wherein one end of a complexing element 1 is simultaneously single-end bonded to either end of two partner molecules, one end of a complexing element 1 is single-end bonded to a middle position of one partner molecule, or one end of a partner molecule is simultaneously single-end bonded to either end of one complexing element 1 and either end of one partner molecule, wherein the complexing element 1 is the same as that shown in Figure 6.

[0227] In a fifth preferred embodiment of the present application, a binary complex is formed by nano-pore as shown in Figure 22, wherein one end of a complexing element 1 is simultaneously single-end bonded to either end of one partner molecule and either end of one polypeptide to be detected, one end of one partner molecule is simultaneously single-end bonded to either end of one complexing element 1 and either end of one polypeptide to be detected, or either end of one complexing element 2 is single-end bonded to either end of one polypeptide to be detected, wherein the complexing element 1 is the same as that shown in Figure 6, and the complexing element 2 is the same as that shown in Figure 7.

[0228] In a sixth preferred embodiment of the present application, a binary complex is formed by nano-pore as shown in Figure 23, wherein one polypeptide to be detected is double-end bonded to one partner molecule.

[0229] In a seventh preferred embodiment of the present application, a binary complex is formed by nano-pore as shown in Figure 24, wherein two ends of one complexing element 3 are respectively single-end bonded to either end of one partner molecule, wherein the complexing element 3 is the same as that shown in Figure 9.

[0230] In an eighth preferred embodiment of the present application, a binary complex is formed by nano-pore as shown in Figure 25, wherein two ends of one complexing element 3 are respectively single-end bonded to either end of one polypeptide to be detected, wherein the complexing element 3 is the same as that shown in Figure 9.

[0231] In a ninth preferred embodiment of the present application, a binary complex is formed by nano-pore as shown in Figure 26, wherein either end of one complexing element 3 is single-end bonded to either end of one partner molecule, wherein the complexing element 3 is the same as that shown in Figure 9.

[0232] In a tenth preferred embodiment of the present application, a binary complex is formed by nano-pore as shown in Figure 27, wherein either end of one complexing element 3 is single-end bonded to either end of one polypeptide to be detected, wherein the complexing element 3 is the same as that shown in Figure 9.

[0233] In a eleventh preferred embodiment of the present application, the binary complex is formed by a nano-pore method as shown in FIG. 28, wherein one complexing element 1 is two-end bonded with one partner molecule to form the binary complex, wherein the complexing element 1 is the same as FIG. 6.

[0234] In a twelfth preferred embodiment of the present application, the binary complex is formed by a nano-pore method as shown in FIG. 29, wherein one complexing element 4 is single-end bonded with either end of one partner molecule to form the binary complex, wherein the complexing element 4 is the same as FIG. 14.

[0235] In a thirteenth preferred embodiment of the present application, the binary complex is formed by a nano-pore method as shown in FIG. 30, wherein one complexing element 4 is single-end bonded with either end of one partner molecule to form the binary complex, wherein the complexing element 4 is the same as FIG. 14.

[0236] In a fourteenth preferred embodiment of the present application, the binary complex is formed by a nano-pore method as shown in FIG. 31, wherein one complexing element 4 is single-end bonded with either end of one polypeptide to be tested to form the binary complex, wherein the complexing element 4 is the same as FIG. 14.

[0237] In a fifteenth preferred embodiment of the present application, the binary complex is formed by a nano-pore method as shown in FIG. 32, wherein one complexing element 4 is single-end bonded with either end of one polypeptide to be tested to form the binary complex, wherein the complexing element 4 is the same as FIG. 14.

[0238] The partner molecule and the polypeptide to be tested are covalently linked by any one or more of the following: peptide bond, ester bond, ether bond, thiol-maleimide bond, oxime bond of carbonyl-hydroxylamine-containing compound, hydrazone bond of carbonyl-hydrazine-containing compound, urea bond of carbonyl-urea-structured compound, disulfide bond, thioether bond, halogen-nucleophile substitution bond, 1,3 dipolar cycloaddition reaction bond, copper-catalyzed azide-alkynyl cycloaddition reaction bond, ruthenium-catalyzed azide-alkynyl cycloaddition reaction bond, azide compound-alkynyl compound click chemistry reaction bond, or natural chemical bond;

[0239] Preferably, the azide compound-alkynyl compound click chemistry reaction bond includes any one or more of the following: azide-DBCO click chemistry reaction bond, azide-OCT click chemistry reaction bond, azide-DIBO click chemistry reaction bond, azide-BARAC click chemistry reaction bond, azide-ALO click chemistry reaction bond, azide-DIFO click chemistry reaction bond, azide-MOFO click chemistry reaction bond, azide-DIBAC click chemistry reaction bond, azide-DIMAC click chemistry reaction bond, or azide-cyclooctene click chemistry reaction bond.

[0240] It should be noted that when the polypeptide to be tested and the partner molecule are covalently linked, the polypeptide to be tested and the partner molecule are close to each other, and may generate interaction forces such as hydrogen bonds (as shown in FIG. 34).

[0241] In a sixth typical embodiment of the present application, a polypeptide sequencing method is provided, which comprises: co-incubating the polypeptide library constructed by the construction method or the polypeptide library with an anchor sequence to obtain an incubation complex; adding the incubation complex into a sequencing solution tank, and under the action of an electric field force, controlling the binary complex containing the polypeptide to be tested to pass through a nanopore by a motor protein, so as to obtain an electric signal corresponding to the polypeptide to be tested; and decoding the electric signal to determine the amino acid sequence of the polypeptide to be tested.

[0242] One end of the anchor sequence is complementary to one end of the linker sequence 2 away from the complementary fragment 1 in the linker complex, and the other end is provided with an anchor group; preferably, the anchor group is selected from any one of a lipid, a carbon nanotube, a polypeptide, a protein and / or an amino acid; preferably, the lipid is selected from any one of a fatty acid, a sterol, a palmitate or a tocopherol; preferably, the anchor sequence is 5'-Chol-TEG-TT-YYYY-SEQ ID NO:8, wherein Y = iSp18, Chol-TEG represents cholesteryl-PEG, and the sequence of SEQ ID NO:8 is 5'-TTGACCGCTCGCCTC-3'.

[0243] The nanopore is a protein nanopore or a solid-state nanopore (FIG. 33); preferably, the protein nanopore is selected from any one or a combination of the following: a-haemolysin, haemolysin, leukocidin, Mycobacterium smegmatis porin A (MspA), MspB, MspC, MspD, a-Haemolysin, CsgG, Aerolysin, cytolysin, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, Neisseria autotransporter lipoprotein (NalP), WZA, GspD, BCP34 and BCP58; preferably, the solid-state nanopore is selected from any one or a combination of the following: a graphene nanopore, a gold nanopore, a silicon nitride nanopore, a silicon dioxide nanopore or an aluminum oxide nanopore.

[0244] The beneficial effects of the present application will be further explained in detail below in combination with specific examples.

[0245] The DNA used in the following examples is synthesized by Huada Lihe (the sequence is represented as 5'→3'), and the polypeptide is synthesized by Jin Sui (the sequence is represented as N-terminal→C-terminal), and the sequence information is as follows:

[0246] DNA1: SEQ ID NO: 1 -int DBCO dT-XXXXXX-int DBCO dT-SEQ ID NO: 2, wherein the sequence of SEQ ID NO: 1 is 5'-GCTTCTCGTGTTTTTTTTCTCTC-3'; the sequence of SEQ ID NO: 2 is 5'- CCCTTTTTTTTTTGCTGTCTTCTGTCGTCGTTT-3', with a ddC modification at the 3' end;

[0247] DNA2: SEQ ID NO: 1 -int DBCO dT-XXXXXX-SEQ ID NO: 4, wherein the sequence of SEQ ID NO: 1 is 5'-GCTTCTCGTGTTTTTTTTCTCTC-3'; the sequence of SEQ ID NO: 4 is 5'- TCCCTTTTTTTTTTGCTGTCTTCTGTCGTCGTTT-3', with a ddC modification at the 3' end;

[0248] DNA3: SEQ ID NO: 9-XXXXXX-int DBCO dT-SEQ ID NO: 2, wherein the sequence of SEQ ID NO: 9 is 5'-GCTTCTCGTGTTTTTTTTCTCTCT-3'; the sequence of SEQ ID NO: 2 is 5'- CCCTTTTTTTTTTGCTGTCTTCTGTCGTCGTTT-3', with a ddC modification at the 3' end;

[0249] DNA4: SEQ ID NO: 9-XXXXXX-SEQ ID NO: 4, wherein the sequence of SEQ ID NO: 9 is 5'- GCTTCTCGTGTTTTTTTTCTCTCT-3'; the sequence of SEQ ID NO: 4 is 5'-TCCCTTTTTTTTTTGCTGTCTTCTGTCGTCGTTT-3', with a ddC modification at the 3' end;

[0250] DNA5 has a sequence as set forth in SEQ ID NO: 5: 5'-GAAACGACGACAGAAGACAGCAAAAAAAAAAGGGATTTTTTAGAGAGAAAAAAAAC GAGAAGCA-3';

[0251] DNA6: XXXXXXXXXXXXXXXXXXXXXXXXXXXXX- SEQ ID NO: 6- YYYY- SEQ ID NO: 3; wherein, X = iSpC3, the sequence of SEQ ID NO: 6 is 5'- TTTTTTTTTT -3', Y = iSp18, the sequence of SEQ ID NO: 3 is 5'- GGTTGTTTCTGTTGGTGCTGATATTGCT -3';

[0252] DNA7 has the sequence as shown in SEQ ID NO: 7: 5'- GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA -3', wherein, the 5' end is with phosphorylation modification;

[0253] DNA8: Chol-TEG-TT- YYYY- SEQ ID NO: 8, wherein, Chol-TEG represents cholesteryl- polyethylene glycol, Y = iSp18, the sequence of SEQ ID NO: 8 is 5'- TTGACCGCTCGCCTC -3';

[0254] DBCO represents: dibenzocyclooctyne, for azide-alkyne cycloaddition (SPAAC) reaction without copper ion catalysis;

[0255] The structural formula of Int DBCO dT is shown in the following figure:

[0256] X in DNA1, DNA2, DNA3 and DNA4 is deoxynucleoside without base (i.e. phosphate deoxyribose), and its structural formula is:

[0257] ddC is double deoxy cytosine nucleoside, and its structural formula is:

[0258] iSp18: Int Spacer 18, which represents the spacer 18 located in the middle position.

[0259] Peptide1: N3- SEQ ID NO: 10- AznL, wherein, the sequence of SEQ ID NO: 10 is DALTAIE;

[0260] Peptide2: N3- SEQ ID NO: 11- AznL, wherein, the sequence of SEQ ID NO: 11 is ERVEE;

[0261] N3: This modification is the replacement of NH2on the lysine side chain residue with N3, which is itself a non-natural amino acid, and is introduced directly into the sequence during polypeptide synthesis. AznL is the abbreviation for 6-azido-L-norleucine, i.e. the structure after the lysine side chain amino group has been replaced with an azido functional group through a modification reaction:

[0262] Example 1: Construction of DNA double-stranded templates containing only a chaperone molecule, without coupled polypeptide

[0263] Four kinds of DNA double-stranded templates were prepared: (a) double-end coupling template, (b) 5' end coupling template, (c) 3' end coupling template and (d) control template. The preparation method of each is as follows:

[0264] 1) Dissolve DNA1, DNA2, DNA3, DNA4, DNA5 powders in pure water to obtain a stock solution with a final concentration of 100 μM.

[0265] 2) Mix DNA1, DNA2, DNA3 or DNA4 with DNA5 in a 1:1 molar ratio in 1x PBS solution, and then use a PCR instrument to perform gradient cooling annealing (heat at 95°C for 5 minutes, then slowly cool to 25°C at a rate of 0.1°C / s, and then hold for 30 minutes to complete annealing), to obtain a DNA double-stranded template with a final concentration of 10 μM, which is denoted as (a) DNA1-DNA5, (b) DNA2-DNA5, (c) DNA3-DNA5, (d) DNA4-DNA5, respectively.

[0266] Example 2: Preparation of sequencing adapters

[0267] 1) Dissolve DNA6 and DNA7 powders in TE buffer (pH 8) to obtain a stock solution with a final concentration of 100 μM.

[0268] 2) Then add 10 μL of DNA6 / DNA7 stock solution to 40 μL of TE buffer (pH 8) to dilute to a working solution with a final concentration of 20 μM.

[0269] 3) Next, mix 30 μL of the working solution of DNA6 / DNA7 obtained by dilution in the previous step together in a 1:1 molar ratio, and then use a PCR instrument to perform gradient cooling annealing (heat at 70°C for 10 minutes, then slowly cool to 25°C at a rate of 0.1°C / s, and then hold for 30 minutes to complete annealing), to obtain an adapter solution with a final concentration of 10 μM.

[0270] Example 3: Cloning, expression and purification of motor protein (Dda helicase)

[0271] 1) Cloning and expression of Dda helicase

[0272] The full-length cDNA sequence of Dda helicase was ordered from Shengwo Biological and ligated into pET28a(+) plasmid using Ndel and Xhol as double enzyme digestion sites, so that the N-terminal of the expressed Dda protein has a 6xHis tag and a thrombin cleavage site. The cloned pET28a(+)-Dda plasmid was transformed into ArcticExpress(DE3) competent bacteria (Tolo Biotech., 96183-02) or its derivative bacteria. Single colonies were picked and inoculated into 5 mL of LB medium containing kanamycin and incubated at 37°C overnight. Then, 1 L of LB (containing kanamycin) was inoculated and incubated at 37°C until OD600=0.6-0.8, then cooled to 16°C, and 500 μM of IPTG was added to induce Dda expression overnight. Five kinds of buffers were prepared according to the following formula:

[0273] Buffer A: 20 mM Tris-HCl pH 7.5, 250 mM NaCl, 20 mM imidazole;

[0274] Buffer B: 20 mM Tris-HCl pH 7.5, 250 mM NaCl, 300 mM imidazole;

[0275] Buffer C: 20 mM Tris-HCl pH 7.5, 50 mM NaCl;

[0276] Buffer D: 20 mM Tris-HCl pH 7.5, 1000 mM NaCl;

[0277] Buffer E: 20 mM Tris-HCl pH 7.5, 100 mM NaCl.

[0278] The DNA sequence of Dda helicase (SEQ ID NO: 12) is as follows:

[0279] The amino acid sequence of Dda helicase (SEQ ID NO: 13) is as follows:

[0280] 2) Purification of Dda helicase

[0281] The bacteria expressing Dda were collected, resuspended in buffer A, and lysed with a cell disrupter. The supernatant was collected after centrifugation. The supernatant was mixed with Ni-NTA resin previously equilibrated with buffer A, and allowed to bind for 1 h. The resin was collected and washed extensively with buffer A until no more contaminant proteins were washed out. Then, Dda was eluted from the resin by adding buffer B. The eluted Dda was desalted by buffer exchange using a desalting column equilibrated with buffer C. Then, an appropriate amount of thrombin (Yeasen, 20402ES05) was added, and the mixture was added to ssDNA cellulose (Sigma, D8273-10G) resin equilibrated with buffer C, and allowed to digest and bind overnight at 4°C. The ssDNA cellulose resin was collected and washed with buffer C for 3-4 times, and then eluted with buffer D. The purified protein on ssDNA cellulose was concentrated and loaded onto a Superdex 200 column (Sigma, GE28-9909-44) equilibrated with buffer E. The protein of interest was collected, concentrated, and stored at -80°C. The concentration of the purified protein was quantified using a Nanodrop. The purity of the protein was also determined using HPLC and SDS-PAGE.

[0282] Example 4: Preparation of sequencing adaptor complex

[0283] 1) To a volumetric flask, 100 mL of 1M Tris-HCl (pH 7.5) buffer and 100 mL of 1M KCl solution were added, and the volume was made up to 1 L with ultrapure water to prepare 2x binding buffer.

[0284] 2) The mixture solution of helicase Dda (prepared in Example 3) and adaptor was prepared according to the following Table 1 on ice, and then incubated at 30°C for 1 h.

[0285] Table 1. Preparation of sequencing adaptor complex

[0286] 3) The sequencing adaptor was quantified using a Qubit DNA HS kit, and the product was stored at 4°C after the concentration was labeled.

[0287] Example 5: Preparation of sequencing library containing only partner molecules without coupled polypeptides

[0288] As shown in Table 2, four DNA double-stranded templates obtained in Example 1 were mixed with pre-prepared linker-motor protein complex (complex formed by DNA6, DNA7 and motor protein), T4 ligase (NEB) and T4 ligase buffer (NEB) respectively, and incubated at room temperature for 30 minutes. Then 2 μL of the mixed solution after incubation was mixed with 3 μL of 1 μM DNA8 and 295 μL of sequencing buffer (0.5 M KCI, 10 mM HEPES, 0.5 mM ATP, 1 mM MgCl2, pH 8) to obtain a sequencing library containing only a partner molecule and no coupled polypeptide, and the final concentration of the library was 2.67 nM.

[0289] Table 2. Construction of sequencing library

[0290] Example 6: Nanopore sequencing using a library containing only a partner molecule and no coupled polypeptide

[0291] This example uses a patch-clamp amplifier to collect current signals. A planar 1,2-diphytanoyl-sn-glycero-3-phosphocholine (DPhPC, Avanti Polar Lipids) phospholipid bilayer lipid membrane is used to divide the electrolytic cell into two chambers: the cis chamber and the trans chamber; and each chamber is placed with a pair of Ag / AgCl electrodes; the nanopore protein CsgG is added to the double-molecular phospholipid membrane, and a voltage of 180 mV is applied to promote the embedding of the pore protein CsgG into the phospholipid bilayer lipid membrane to form a single nanopore channel; after the single nanopore protein CsgG is inserted into the phospholipid membrane, the sequencing buffer (0.5 M KCI, 10 mM HEPES, 0.5 mM ATP, 1 mM MgCl2, pH 8) is pushed in to remove excess pore proteins; then the four libraries mentioned in Example 5 are added to the cis chamber of each independent chip and incubated at 25°C for 10 min; finally, a voltage of 180 mV is applied, and the nanopore current data is recorded at a frequency of 5 kHz.

[0292] Figure 37 shows the sequencing signals of the blank control library (i.e. DNA double-stranded template without coupled polypeptide) of the four nanopore sequencing obtained in Example 5. It can be seen that when Int DBCO dT is introduced into the DNA chain, an electrical signal with up and down fluctuation characteristics will be generated, which is represented by solid and dashed boxes in Figure 37. The positions of these two characteristic signals can help determine the end and start points of the electrical signals at the front and back ends of the partner molecule during sequencing, and thus determine the start and end points of the polypeptide signal.

[0293] Example 7: Preparation of a sequencing library containing a partner molecule and coupled polypeptide

[0294] 1) Dissolve Peptide 1 and Peptide 2 powder in 1x PBS respectively to get the polypeptide stock solution with final concentration of 1 mM.

[0295] 2) Mix the DNA 1 stock solution with Peptide 1 / Peptide 2 stock solution according to the following Table 3 with a molar ratio of 1:19, the final concentration of DNA 1 is 20 mM, and the conjugation product DNA 1-Peptide 1 and DNA 1-Peptide 2 are obtained after 24 h reaction at room temperature.

[0296] Table 3. Conjugation reaction of DNA with chaperone molecule and polypeptide

[0297] 3) Mix the two conjugation products obtained above with DNA 5 in 1x PBS solution according to a molar ratio of 1:1, and then use PCR instrument to perform gradient cooling annealing (heat at 95 °C for 5 min, then slowly cool to 25 °C at a rate of 0.1 °C / s, and keep for 30 min to complete annealing), to obtain polypeptide conjugation double-stranded product with a final concentration of 10 mM DNA, which are denoted as (b) D1P1-D5; (c) D1P2-D5.

[0298] In addition, according to the above steps, a blank library (a) DNA 1-DNA 5 chaperone molecule template is prepared synchronously, which only contains chaperone molecules and is not conjugated with polypeptides. According to the existing literature method (Chen, Zhijie, et al. "Controlled movement of ssDNA conjugated peptide through Mycobacterium smegmatis porin A (MspA) nanopore by a helicase motor for peptide sequencing application." Chemical science 12.47 (2021): 15750-15756.), a control library (d) is prepared synchronously, which is a conventional "DNA-Peptide-DNA" sandwich structure library of the prior art, and the polypeptide is conjugated with SEQ ID NO: 1 and SEQ ID NO: 2 sequences at both ends.

[0299] 4) According to the method described in Example 5, the two polypeptide conjugation double-stranded products obtained in the above step are connected with the linker-motor protein complex, and then mixed with DNA 8 and sequencing buffer to form a sequencing library containing chaperone molecules and conjugated polypeptides.

[0300] Example 8: Nanopore sequencing using library containing chaperone molecules and conjugated polypeptides

[0301] The two libraries of conjugated polypeptides obtained in Example 7, as well as the blank library and the control library, were subjected to nanopore sequencing and signal collection according to the method described in Example 6.

[0302] Figure 38 shows the sequencing signal graphs of (a) the blank library of DNA1-DNA5 chaperone molecule templates (i.e. containing only chaperone molecules, without conjugated polypeptides), (b) the D1P1-D5 library, (c) the D1P2-D5 library, and (d) the control library of the conventional "DNA1-Peptide-DNA2" sandwich structure of the prior art.

[0303] The signal range of the polypeptides was determined in (a) and the polypeptide signal interval was extracted in (b) and (c) according to the method described in Example 6. As can be seen by comparing (b) and (c), when polypeptides with the same number of charges but different lengths are synchronized through the nanopore with the chaperone molecules, the signal trends are similar, but there are still significant differences in the details. As can be seen by comparing (c) and (d), compared with the conventional nanopore sequencing scheme that does not rely on the assistance of molecular chaperones, when polypeptides are synchronized through the nanopore with chaperone molecules, the sequencing signal is more stepped, the signal stability is significantly improved, and the fluctuations are significantly reduced, indicating that the method proposed in the present application can effectively improve the signal quality of protein nanopore sequencing.

[0304] Example 9: Construction of nucleic acid chaperone molecule library and sequencing signal

[0305] In addition to the library structure using deoxynucleosides without bases as chaperone molecules as shown in Examples 6 and 8, this example will show the sequencing signal after replacing the chaperone molecules with deoxynucleotide sequences. At this time, the DNA1 sequence used is changed to DNA9: 5'-SEQ ID NO: 1-int DBCO dT-TTTTTTTTT-int DBCO dT-SEQ ID NO: 2-3', with a ddC modification at the 3' end. The complementary sequence of DNA9 is changed to DNA10, which has the sequence shown in SEQ ID NO: 14: GAAACGACGACAGAAGACAGCAAAAAAAAAAGGGAAAAAAAAAAAGAGAGAAAAAA AACACGAGAAGCA.

[0306] According to the methods described in Examples 7 and 8, two libraries of D9P1-D10 and D9P2-D10 were prepared, and their sequencing signals were obtained.

[0307] Figure 39 shows the sequencing signal diagrams of (a) D9P1-D10 and (b) D9P2-D10. It can be seen that pure deoxynucleotide sequences can also produce nanopore electrical signals as partner molecules together with polypeptides. By comparing Figure 38, it can be found that when the same polypeptide is threaded through the nanopore using different partner molecules, the overall trend has some similarities, but there are also not small differences. In addition, unlike the molecular partner shown in Figure 39, the signal trend of the polypeptide with the same number of charges but different lengths is similar, and after replacing the deoxynucleotide sequence in Figure 38 as a molecular partner, the two polypeptides with the same number of charges but different lengths show more obvious signal differences. This not only shows that the partner molecule assisted threading of the present application has universality in the composition of the partner molecule (without specific molecular species requirements), but also shows the huge practical potential of the present application scheme.

[0308] From the above description, it can be seen that the above-mentioned embodiments of the present application achieve the following technical effects: the present application provides a brand-new sequencing idea of synchronously threading polypeptides and partner molecules through nanopores, and by using the combination of "partner molecule-polypeptide", the complex for preparing a nanopore polypeptide sequencing library can be constructed by coupling first (with a single chain) and then annealing (with a complementary chain), or by annealing first and then coupling.

[0309] By using the combination of "partner molecule-polypeptide" to simulate the interaction between nucleotides threading through the nanopore and the constriction of the nanopore in the DNA nanopore sequencing process, the smaller volume of amino acid sequences is more suitable for the existing nanopore channel diameter, the stability and repeatability of the electrical signal in the polypeptide sequencing process are improved, the interaction between the polypeptide and the nanopore is enhanced, the signal characteristics are improved, and the existing nanopore suitable for DNA sequencing is better adapted to polypeptide sequencing. The present application can better push the amino acid residues on the polypeptide towards the nanopore wall, enhance the interaction between the amino acid residues on the polypeptide to be tested and the amino acid residues at the constriction of the nanopore, enhance the electrical signal characteristics of a single polypeptide, and further improve the resolution of a single amino acid.

[0310] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A binary composite, characterized in that, The binary complex comprises a covalently linked partner molecule and a polypeptide to be tested, wherein the partner molecule is selected from any one or more of PEG, spacer, deoxyribose phosphate, ribose phosphate, nucleotide, deoxynucleotide, peptide nucleic acid or locked nucleotide; wherein the number of constituent units of the partner molecule is ≥1; and the number of amino acids of the polypeptide to be tested is ≥2.

2. The binary composite of claim 1, wherein, The covalently linked form of the partner molecule and the polypeptide to be tested comprises any one or more of any one of the following groups: 1) single-end bonding of any one end of the polypeptide to be tested to any one end of the partner molecule; 2) single-end bonding of the same end of the polypeptide to be tested to any one end of two partner molecules, single-end bonding of any one end of the polypeptide to the middle position of the partner molecule, or single-end bonding of the same end of one partner molecule to any one end of another partner molecule and any one end of the polypeptide to be tested; 3) single-end bonding of the same end of one partner molecule to any one end of two polypeptides to be tested, single-end bonding of any one end of the partner molecule to the middle position of the polypeptide to be tested, or single-end bonding of the same end of one polypeptide to be tested to any one end of one partner molecule and any one end of the polypeptide to be tested; 4) single-end bonding of the same end of one complex element 1 to any one end of two partner molecules, single-end bonding of one complex element 1 to the middle position of the partner molecule, or single-end bonding of the same end of one partner molecule to any one end of one complex element 1 and any one end of the partner molecule; 5) single-end bonding of the same end of one complex element 1 to any one end of one partner molecule and any one end of one polypeptide to be tested, single-end bonding of the same end of one partner molecule to any one end of one complex element 1 and any one end of one polypeptide to be tested, or single-end bonding of any one end of one complex element 2 to any one end of one polypeptide to be tested; 6) two-end bonding of one polypeptide to one partner molecule; 7) single-end bonding of each of the two ends of one complex element 3 to any one end of one partner molecule; 8) single-end bonding of each of the two ends of one complex element 3 to any one end of one polypeptide to be tested; 9) single-end bonding of any one end of one complex element 3 to any one end of one partner molecule; 10) single-end bonding of any one end of one complex element 3 to any one end of one polypeptide to be tested; 11) two-end bonding of one complex element 1 to one partner molecule; 12) single-end bonding of each of the two ends of one complex element 4 to any one end of one partner molecule; 13) single-end bonding of any one end of one complex element 4 to any one end of one partner molecule; 14) single-end bonding of each of the two ends of one complex element 4 to any one end of one polypeptide to be tested; 15) one end of a complex element 4 is single-end bonded with one end of any one of said polypeptides to be tested; wherein the partner molecules at different positions in any one form of said binary complex are the same or different, and the polypeptides to be tested at different positions in any one form of said binary complex are the same or different; a complex element 1 is a complex after one end of said polypeptide to be tested is single-end bonded with one end of any one of said partner molecules; a complex element 2 is a complex after one end of said complex element 1 is single-end bonded with one end of any one of said partner molecules; a complex element 3 is a complex after one end of said polypeptide to be tested is single-end bonded with one end of any one of said partner molecules; a complex element 4 is a complex after one end of said complex element 1 is single-end bonded with one end of any one of said partner molecules.

3. The binary composite of claim 2, wherein, said partner molecules and said polypeptides to be tested are covalently connected by any one or more of the following: peptide bond connection, ester bond connection, ether bond connection, thiol-maleimide connection, oxime formation of carbonyl-hydroxylamine-containing compounds, hydrazone formation of carbonyl-hydrazine-containing compounds, urea formation of carbonyl-urea-containing compounds, disulfide bond connection, sulfide bond connection, halogen-nucleophile substitution connection, 1,3 dipolar cycloaddition reaction connection, copper-catalyzed azide-alkynyl cycloaddition reaction connection, ruthenium-catalyzed azide-alkynyl cycloaddition reaction connection, azide compound-alkynyl compound click chemistry reaction connection, or natural chemical connection; Preferably, said azide compound-alkynyl compound click chemistry reaction connection includes any one or more of the following: azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection, or azide-cyclooctene click chemistry reaction connection.

4. The binary composite of claim 1, wherein, said spacer is selected from any one or more of the following: Spacer C3, Spacer C6, Spacer 9, Spacer C12, or Spacer 18.

5. The binary composite of claim 1, wherein, said PEG has an average molecular weight of 300-20000; Preferably, said PEG is selected from any one or more of the following: PEG-300, PEG-400, PEG-800, PEG-1000, PEG-1500, PEG-2000, PEG-3000, PEG-4000, PEG-6000, PEG-8000, or PEG-20000.

6. A double-stranded annealed complex, comprising: said double-stranded annealed complex comprises: a single-stranded nucleic acid 1 and a complementary fragment 1 that is at least partially complementary to said single-stranded nucleic acid 1, wherein the middle position of said single-stranded nucleic acid 1 is covalently connected with the binary complex of any one of claims 1 to 5.

7. The double-stranded annealed complex of claim 6, wherein, the first segment of said single-stranded nucleic acid 1 and the second segment of said single-stranded nucleic acid 1 are respectively covalently connected with the two ends of said binary complex; The covalent connection is selected from any one or more of the following: peptide bond connection, ester bond connection, ether bond connection, thiol-maleimide connection, oxime formation of carbonyl-hydroxylamine-containing compound connection, hydrazone formation of carbonyl-hydrazine-containing compound connection, urea formation of carbonyl-urea structure-containing compound connection, disulfide bond connection, sulfide bond connection, halogen-nucleophile substitution connection, 1,3 dipolar cycloaddition reaction connection, copper-catalyzed azide-alkynyl cycloaddition reaction connection, ruthenium-catalyzed azide-alkynyl cycloaddition reaction connection, azide compound-alkynyl compound click chemistry reaction connection, or natural chemical connection. Preferably, the covalent connection is diester bond connection. More preferably, the covalent connection is phosphodiester bond connection. Preferably, the azide compound-alkynyl compound click chemistry reaction connection includes azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection, azide-cyclooctene click chemistry reaction connection. The structure of the single-stranded nucleic acid 1 is: first segment of single-stranded nucleic acid 1-int DBCO modified deoxyribonucleotide-the binary complex-int DBCO modified deoxyribonucleotide-second segment of single-stranded nucleic acid 1.

8. The duplex annealing complex of claim 7, wherein, The length of the first segment of the single-stranded nucleic acid 1 is ≥1 nt, and the length of the second segment of the single-stranded nucleic acid 1 is ≥1 nt; the length of the complementary fragment 1 is ≥1 nt.

9. The double-stranded annealed complex of claim 8, wherein, Preferably, the length of the first segment of the single-stranded nucleic acid 1 is 5-500 nt, the length of the second segment of the single-stranded nucleic acid 1 is 5-500 nt, and the length of the complementary fragment 1 is 5-1200 nt. The 5' end of the first segment of the single-stranded nucleic acid 1 has a phosphorylation group, and the 3' end of the second segment of the single-stranded nucleic acid 1 has any one or more of the following dideoxynucleotides: ddA, ddT, ddC, or ddG.

10. The duplex annealing complex of any one of claims 6-9, wherein, The first segment of the single-stranded nucleic acid 1 is selected from the sequence shown in SEQ ID NO: 1: 5'-GCTTCTCGTGTTTTTTTTCTCTC-3'; 11. The duplexed annealed complex of claim 10, wherein, The second segment of the single-stranded nucleic acid 1 is selected from the sequence shown in SEQ ID NO: 2: 5'-CCCTTTTTTTTTTGCTGTCTTCTGTCGTCGTTT-3'; The complementary fragment 1 is selected from the sequence shown in SEQ ID NO: 5: 5'-GAAACGACGACAGAAGACAGCAAAAAAAAAAGGGATTTTTTAGAGAGAAAAAAAACACGAGAAGCA-3'. The polypeptide library comprises the binary complex of any one of claims 1 to 5 or the double-stranded annealed complex of any one of claims 6 to 11.

12. A library of polypeptides, characterized in that, ​ 13. The library of claim 12, wherein, The polypeptide library comprises the double-stranded annealed complex and a linker complex covalently linked to the double-stranded annealed complex; The linker complex comprises: a linker sequence 1 and a linker sequence 2, and a motor protein, the linker sequence 1 comprises a first segment and a second segment connected in sequence from 5' end to 3' end direction, wherein the first segment of the linker sequence 1 is not complementary to the linker sequence 2, the second segment of the linker sequence 1 is complementary to the linker sequence 2, and the motor protein is movably bound to the first segment of the linker sequence 1; The linker sequence 1 is covalently linked to the 5' end of the single-stranded nucleic acid 1; The linker sequence 2 is covalently linked to the 3' end of the complementary fragment 1 which is at least partially complementary to the single-stranded nucleic acid 1.

14. The library of claim 13, wherein, The length of the linker sequence 1 is ≥1 nt, and the length of the linker sequence 2 is ≥1 nt; Preferably, the length of the linker sequence 1 is 10-100 nt, and the length of the linker sequence 2 is 10-100 nt.

15. The library of claim 14, wherein, The linker sequence 1 is selected from the sequence shown as follows: 5'-XXXXXXXXXXXXXXXXXXXXXXXXXXXXXX-SEQ ID NO: 6-YYYY-SEQ ID NO: 3-3', wherein X = iSpC3, the sequence of SEQ ID NO: 6 is TTTTTTTTTT, and Y = iSp18, the sequence of SEQ ID NO: 3 is GGTTGTTTCTGTTGGTGCTGATATTGCT; The linker sequence 2 is selected from the sequence shown in SEQ ID NO: 7: 5'-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3', wherein the 5' end is modified with phosphorylation.

16. The library of claim 13, wherein, The motor protein is selected from a polymerase or a helicase; Preferably, the polymerase is selected from any one of the following: Bst DNA polymerase, SD DNA polymerase, phi29 DNA polymerase, Bsu Large Fragment DNA polymerase, Klenow Fragment DNA polymerase, T3 RNA polymerase, T7 RNA polymerase, SP6 RNA polymerase or E. coli RNA polymerase; Preferably, the helicase is selected from any one of the following: Dda, Hel308, RecD, UvrD, Rep, RecQ, PcrA, eIF4A, NS3, gp41, T7gp4 or BCH105.

17. A polypeptide sequencing kit, characterized in that, The kit comprises: the chaperone molecule in the binary complex of any one of claims 1 to 3, and any one or more of the following optional components: a membrane, a nanopore, the single-stranded nucleic acid 1, the complementary fragment 1, the linker sequence 1, the linker sequence 2 and the motor protein in the polypeptide library of claim 13.

18. The kit of claim 17, wherein The linker sequence 1, the linker sequence 2 and the motor protein exist in the form of a linker complex; Preferably, the nanopore is a protein nanopore or a solid-state nanopore. More preferably, the protein nanopore is selected from any one or more combinations of: a-hemolysin, hemolysin, leukocidin, Mycobacterium smegmatis pore protein A (MspA), MspB, MspC, MspD, a-Haemolysin, CsgG, Aerolysin, cytolysin, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, Neisseria autotransporter lipoprotein (NalP), WZA, GspD, BCP34 and BCP58; More preferably, the solid-state nanopore is selected from any one or more combinations of: graphene nanopore, gold nanopore, silicon nitride nanopore, silicon dioxide nanopore or aluminium oxide nanopore; Preferably, the nanopore is located on and traverses the membrane; Preferably, the membrane is selected from any one or more of: a phospholipid membrane, a polymer membrane or a solid-state membrane.

19. A method of constructing a library of polypeptides, characterized by, The construction method comprises: preparing the polypeptide to be tested into the binary complex of any one of claims 1 to 5; linking the binary complex with a double-stranded nucleic acid to obtain a double-stranded complex; linking the double-stranded complex with a double-stranded linker complex containing a motor protein to obtain the polypeptide library; wherein the double-stranded linker complex containing a motor protein comprises: a linker sequence 1, a linker sequence 2 complementary to the 3' end of the linker sequence 1 and non-complementary to the 5' end, and the motor protein located on the linker sequence 1.

20. The method of construction of claim 19, wherein, The double-stranded nucleic acid comprises a first segment of single-stranded nucleic acid 1, a second segment of single-stranded nucleic acid 1, and a complementary fragment 1 complementary to at least part of the single-stranded nucleic acid 1.

21. The method of construction of claim 20, wherein, The construction method comprises: covalently linking the two ends of the binary complex to the 3' end of the first segment of single-stranded nucleic acid 1 and the 5' end of the second segment of single-stranded nucleic acid 1, respectively; annealing the single-stranded nucleic acid 1 with the binary complex to the complementary fragment 1 to form a double-stranded annealing complex; linking the double-stranded annealing complex with the linker complex containing a motor protein by a ligase to form the polypeptide library.

22. The method of construction of claim 20, wherein, The construction method comprises: annealing the single-stranded nucleic acid 1 with the complementary fragment 1 to form the double-stranded nucleic acid; covalently linking the two ends of the binary complex to the 3' end of the first segment of single-stranded nucleic acid 1 and the 5' end of the second segment of single-stranded nucleic acid 1 in the double-stranded nucleic acid to obtain the double-stranded annealing complex; linking the double-stranded annealing complex with the linker complex containing a motor protein by the ligase to form the polypeptide library.

23. The method of construction of claim 19, wherein, Preparation of the polypeptide to be tested into the binary complex comprises: covalently binding the chaperone molecule and the polypeptide to be tested to obtain the binary complex; wherein the chaperone molecule is selected from any one or more of: PEG, spacer, deoxyribose phosphate, ribose phosphate, nucleotide, deoxynucleotide, peptide nucleic acid or locked nucleotide; Preferably, the spacer is selected from any one or more of: Spacer C3, Spacer C6, Spacer 9, Spacer C12 or Spacer 18; Preferably, the average molecular weight of the PEG is 300-20000.

24. The method of construction of claim 19, wherein, The forms of covalent linkage of the chaperonin and the polypeptide to be detected include any one or more of any one of the following groups: 1) single-end linkage of any one end of one polypeptide to be detected and any one end of one chaperonin; 2) single-end linkage of the same end of one polypeptide to be detected to any one end of two chaperonins, single-end linkage of any one end of one polypeptide to be detected to the middle of one chaperonin, or single-end linkage of the same end of one chaperonin to any one end of another chaperonin and any one end of one polypeptide to be detected; 3) single-end linkage of the same end of one chaperonin to any one end of two polypeptides to be detected, single-end linkage of any one end of one chaperonin to the middle of one polypeptide to be detected, or single-end linkage of the same end of one polypeptide to be detected to any one end of one chaperonin and any one end of one polypeptide to be detected; 4) single-end linkage of the same end of one complex element 1 to any one end of two chaperonins, single-end linkage of one complex element 1 to the middle of one chaperonin, or single-end linkage of the same end of one chaperonin to any one end of one complex element 1 and any one end of one chaperonin; 5) single-end linkage of the same end of one complex element 1 to any one end of one chaperonin and any one end of one polypeptide to be detected, single-end linkage of the same end of one chaperonin to any one end of one complex element 1 and any one end of one polypeptide to be detected, or single-end linkage of any one end of one complex element 2 to any one end of one polypeptide to be detected; 6) two-end linkage of one polypeptide to be detected to one chaperonin; 7) single-end linkage of each of the two ends of one complex element 3 to any one end of one chaperonin; 8) single-end linkage of each of the two ends of one complex element 3 to any one end of one polypeptide to be detected; 9) single-end linkage of any one end of one complex element 3 to any one end of one chaperonin; 10) single-end linkage of any one end of one complex element 3 to any one end of one polypeptide to be detected; 11) two-end linkage of one complex element 1 to one chaperonin; 12) single-end linkage of each of the two ends of one complex element 4 to any one end of one chaperonin; 13) single-end linkage of any one end of one complex element 4 to any one end of one chaperonin; 14) single-end linkage of each of the two ends of one complex element 4 to any one end of one polypeptide to be detected; 15) single-end linkage of any one end of one complex element 4 to any one end of one polypeptide to be detected; wherein the chaperonins at different positions in any one form of the binary complex are the same or different, and the polypeptides to be detected at different positions in any one form of the binary complex are the same or different. The complex element 1 is a complex of single-end bonding of each end of the polypeptide to be detected to any one end of the chaperone molecule; The complex element 2 is a complex of single-end bonding of any one end of the complex element 1 to any one end of the chaperone molecule; The complex element 3 is a complex of two-end bonding of the polypeptide to be detected to the chaperone molecule; The complex element 4 is a complex of two-end bonding of the complex element 1 to the chaperone molecule.

25. The method of construction of claim 19, wherein, The covalent connection of the chaperone molecule and the polypeptide to be detected is achieved by any one or more of the following: peptide bond connection, ester bond connection, ether bond connection, thiol-maleimide connection, oxime formation of carbonyl-hydroxylamine-containing compound, hydrazone formation of carbonyl-hydrazine-containing compound, urea formation of carbonyl-urea-containing compound, disulfide bond connection, sulfide bond connection, halogen-nucleophile substitution connection, 1,3 dipolar cycloaddition reaction connection, copper-catalyzed azide-alkynyl cycloaddition reaction connection, ruthenium-catalyzed azide-alkynyl cycloaddition reaction connection, azide compound-alkynyl compound click chemistry reaction connection or natural chemical connection; Preferably, the azide compound-alkynyl compound click chemistry reaction connection includes any one or more of the following: azide-DBCO click chemistry reaction connection, azide-OCT click chemistry reaction connection, azide-DIBO click chemistry reaction connection, azide-BARAC click chemistry reaction connection, azide-ALO click chemistry reaction connection, azide-DIFO click chemistry reaction connection, azide-MOFO click chemistry reaction connection, azide-DIBAC click chemistry reaction connection, azide-DIMAC click chemistry reaction connection or azide-cyclooctene click chemistry reaction connection.

26. A method of sequencing a polypeptide, comprising: The sequencing method comprises: The polypeptide library constructed by the construction method of any one of claims 19-25 or the polypeptide library of any one of claims 12-16 is co-incubated with an anchor sequence to obtain an incubation complex; The incubation complex is added to a sequencing solution bin, and under the action of an electric field force, a binary complex containing the polypeptide to be detected is controlled by a motor protein to pass through a nanopore, so as to obtain an electric signal corresponding to the polypeptide to be detected; The electric signal is decoded to determine the amino acid sequence of the polypeptide to be detected.

27. The sequencing method of claim 26, wherein, One end of the anchor sequence is complementary to the end of the linker sequence 2 in the linker complex away from the complementary fragment 1, and the other end is provided with an anchor group; Preferably, the anchor group is selected from any one of lipids, carbon nanotubes, polypeptides, proteins and / or amino acids; Preferably, the lipid is selected from any one of fatty acids, sterols, palmitate or tocopherol; Preferably, the anchor sequence is 5'-Chol-TEG-TT-YYYY-SEQ ID NO:8, wherein Y=iSp18, Chol-TEG represents cholesteryl-PEG, and the sequence of SEQ ID NO:8 is 5'-TTGACCGCTCGCCTC-3'.

28. The sequencing method of claim 27, wherein, The nanopore is a protein nanopore or a solid-state nanopore. Preferably, the protein nanopore is selected from any one or more combinations of: a-haemolysin, haemolysin, leukocidin, Mycobacterium smegmatis pore protein A (MspA), MspB, MspC, MspD, a-Haemolysin, CsgG, Aerolysin, cytolysin, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A, Neisseria autotransporter lipoprotein (NalP), WZA, GspD, BCP34 and BCP58; Preferably, the solid state nanopore is selected from any one or more combinations of: graphene nanopore, gold nanopore, silicon nitride nanopore, silicon dioxide nanopore or aluminium oxide nanopore.

Citation Information

Patent Citations

  • Method for controlling speed of polypeptide passing through nanopore and application thereof

    CN112147185A

  • DNA-peptide conjugate, ARCFU-like helicase and application of ARCFU-like helicase in detection of peptide fragment

    CN118086286A

  • Nanopore sequencing method and kit

    CN118207308A

  • Polypeptide tagged nucleotides and use thereof in nucleic acid sequencing by nanopore detection

    US20190002968A1

  • Single molecule nanopore sequencing method

    WO2023125605A1