NANO porin and use thereof

By modifying nanoporin monomers to form a dual-sensor structure, the problem of insufficient sequencing accuracy of existing nanoporin sequencing has been solved, achieving high-precision molecular detection and sequencing.

WO2025260333A1PCT designated stage Publication Date: 2025-12-26SHENZHEN HUADA GENE INST
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/100449
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing nanoporous proteins have insufficient sequencing accuracy during sequencing. Due to their simple pore structure, they are difficult to achieve high-precision molecular detection and sequencing.

Method used

A novel pore protein monomer was developed, and its amino acid sequence was modified to form a pore structure with dual receptors, including a central receptor and a cap-gate receptor, to improve sequencing accuracy.

Benefits of technology

It achieves highly accurate molecular detection and sequencing, and significantly improves the ability to recognize nucleic acid and peptide sequences through the dual-sensor structure, thereby improving sequencing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2024100449-FTAPPB-I100001
    Figure PCTCN2024100449-FTAPPB-I100001
  • Figure PCTCN2024100449-FTAPPB-I100002
    Figure PCTCN2024100449-FTAPPB-I100002
  • Figure PCTCN2024100449-FTAPPB-I100003
    Figure PCTCN2024100449-FTAPPB-I100003
Patent Text Reader

Abstract

Provided is a new porin suitable for nanopore detection. The porin is formed by means of the polymerization of porin monomers, and the porin monomers contain a barrel domain, which barrel domain has: a. an amino acid sequence as shown in SEQ ID NO: 1; b. an amino acid sequence in which, compared to the amino acid sequence as shown in SEQ ID NO: 1, one or more amino acids are substituted, deleted, and / or added, wherein the barrel segment has the function of forming a pore channel structure upon polymerization; or c. an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the amino acid sequence as shown in SEQ ID NO: 1, wherein the barrel segment the function of forming a pore channel structure upon polymerization.
Need to check novelty before this filing date? Find Prior Art

Description

Nanopore protein and application thereof TECHNICAL FIELD

[0001] The present application relates to the field of sequencing technology, in particular to a nanopore protein and application thereof. BACKGROUND

[0002] With the continuous development of sequencing technology, nanopore sequencing has become the mainstream method in the third generation sequencing technology due to its long read length and the advantage of not requiring library construction. Currently, there are commercial nanopore sequencing instruments such as MinION, GridION and PromethION of Oxford Nanopore Technologies in the United Kingdom, QNome-3841 of Zymo Research, etc.

[0003] As a sensor molecule, nanopore protein (channel protein) realizes the sequencing of nucleotide sequence by sensing the current disturbance caused by the penetration of nucleotide to generate specific current signal. In the related art, only a few natural proteins such as Mycobacterium smegmatis porin A (MspA) and curli-specific transport channel (CsgG) can meet the requirements of channel protein as a single molecule detector in industry, but the structures of these channel proteins are relatively single, and their sequencing accuracy is still insufficient.

[0004] Therefore, it is urgent to develop new channel proteins with excellent internal structure that can sensitively characterize different nucleotides.

[0005] SUMMARY

[0006] Embodiments of the present application provide a channel protein monomer, a channel protein, a polynucleotide, an expression vector, a cell, a nanopore sensor, a nanopore sequencing device, a sequencing method and a kit.

[0007] The first aspect of the embodiments of the present application provides a channel protein monomer, which comprises a barrel segment, and the barrel segment has:

[0008] a. an amino acid sequence as shown in SEQ ID NO: 1;

[0009] b. an amino acid sequence having one or more substitutions, deletions and / or additions of amino acids compared with the amino acid sequence shown in SEQ ID NO: 1, and the barrel segment has the function of being polymerized to form a channel structure; or

[0010] c. has an amino acid sequence that is at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identical to the amino acid sequence set forth in SEQ ID NO: 1, and the barrel segment has the function of being polymerized to form a pore structure.

[0011] The second aspect of the present application provides a pore protein comprising a polymer of the pore protein monomer according to any one of the first aspect of the present application, which is optionally polymerized from 12-18 of the pore protein monomers.

[0012] The third aspect of the present application provides a polynucleotide encoding the pore protein monomer according to any one of the first aspect of the present application.

[0013] The fourth aspect of the present application provides an expression vector comprising the polynucleotide according to any one of the third aspect of the present application and a controllable element.

[0014] The fifth aspect of the present application provides a cell comprising the polynucleotide according to any one of the third aspect of the present application or the expression vector according to any one of the fourth aspect of the present application or expressing the pore protein monomer according to any one of the first aspect of the present application.

[0015] The sixth aspect of the present application provides a nanopore sensor comprising a membrane layer and the pore protein according to any one of the second aspect of the present application, wherein the pore protein is inserted into the membrane layer to form a pore, and when a voltage is applied across the membrane layer, the pore generates an electric current.

[0016] The seventh aspect of the present application provides a nanopore sequencing device comprising the nanopore sensor according to any one of the sixth aspect of the present application.

[0017] The eighth aspect of the present application provides a sequencing method using the pore protein according to any one of the second aspect of the present application or the nanopore sensor according to any one of the sixth aspect of the present application or the nanopore sequencing device according to any one of the seventh aspect of the present application to determine the composition of a molecule to be detected by detecting and analyzing the electrical signal generated when the molecule to be detected passes through the pore of the pore protein.

[0018] The ninth aspect of the present application provides a kit comprising at least one of the pore monomer according to any one of the first aspect of the present application, the pore according to any one of the second aspect of the present application, the polynucleotide according to any one of the third aspect of the present application, the expression vector according to any one of the fourth aspect of the present application, the cell according to any one of the fifth aspect of the present application, and the nanopore sensor according to any one of the sixth aspect of the present application.

[0019] The tenth aspect of the present application provides the use of the pore monomer according to any one of the first aspect of the present application, the pore according to any one of the second aspect of the present application, the polynucleotide according to any one of the third aspect of the present application, the expression vector according to any one of the fourth aspect of the present application, the cell according to any one of the fifth aspect of the present application, the nanopore sensor according to any one of the sixth aspect of the present application, or the kit according to any one of the ninth aspect of the present application in biological small molecule detection, nucleic acid sequencing, or polypeptide sequencing.

[0020] The technical solution of the present application achieves the following technical effects:

[0021] The novel nanopore protein proposed in the embodiments of the present application has high stability and excellent internal structure, and can sensitively characterize the biological molecules (such as nucleic acids, polypeptides, etc.) in the pore to achieve more accurate molecular detection and nanopore sequencing. In addition, compared with the single-receptor nanopore protein with only one main narrow region such as MspA and CsgG commonly used in current commercial sequencers, the nanopore protein proposed in the embodiments of the present application can be modified to have two narrow regions (i.e., containing double receptors), thereby greatly improving the sequencing accuracy. At the same time, the nanopore protein and the double-receptor-containing nanopore protein based thereon proposed in the embodiments of the present application are directionally modified, and the mutants thereof also have excellent molecular detection and sequencing performance. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0023] FIG. 1 is a schematic diagram of the structure of a pore monomer according to an embodiment of the present application;

[0024] FIG. 2 is a central sensor region and part of the key sites therebetween of a pore according to an embodiment of the present application;

[0025] Figure 3 is the entrance end of a pore protein and some of the sites therebetween, according to embodiments of the present application;

[0026] Figure 4 is the transmembrane region of a pore protein and some of the sites therebetween, according to embodiments of the present application;

[0027] Figure 5 shows the introduction position of a cap-sensor of a pore protein (left) and the structure of a second sensor structure region (right), according to embodiments of the present application;

[0028] Figure 6 is the cap-sensor region of a pore protein and some of the key sites therebetween, according to embodiments of the present application;

[0029] Figure 7 is a side view of a partial structure of a pore protein, according to embodiments of the present application;

[0030] Figure 8 is a top view of a partial structure of a pore protein, according to embodiments of the present application;

[0031] Figure 9 is a bottom view of a partial structure of a pore protein, according to embodiments of the present application;

[0032] Figure 10 is a view of two monomer structures opposite to each other inside a dimer of a pore protein, according to embodiments of the present application;

[0033] Figure 11 is a side view of a partial structure of a pore protein with a cap-sensor grafted thereon, according to embodiments of the present application;

[0034] Figure 12 is a top view of a partial structure of a pore protein with a cap-sensor grafted thereon, according to embodiments of the present application;

[0035] Figure 13 is a view of two monomer structures opposite to each other inside a dimer of a pore protein with a cap-sensor grafted thereon, according to embodiments of the present application;

[0036] Figure 14 is an electropherogram of a pore protein, according to embodiments of the present application;

[0037] Figures 15a-15b are electropherograms of pore protein mutants and grafts, according to embodiments of the present application;

[0038] Figures 16a-16j show the biosensing currents generated by a pore protein and its grafts and mutants at different voltages, according to embodiments of the present application;

[0039] Figure 17 shows the overall current change when a pore protein is used for sequencing, according to embodiments of the present application;

[0040] Figure 18 is a close-up view of the overall current change when a pore protein is used for sequencing, according to embodiments of the present application;

[0041] FIGS. 19a-19b respectively show the overall current change and the local detail view of the overall current change when the pore protein mutant is used for sequencing according to an embodiment of the present application;

[0042] FIGS. 20a-20b respectively show the overall current change and the local detail view of the overall current change when the pore protein mutant is used for sequencing according to an embodiment of the present application;

[0043] FIGS. 21a-21b respectively show the overall current change and the local detail view of the overall current change when the pore protein mutant is used for sequencing according to an embodiment of the present application;

[0044] FIG. 22 shows the molecular sieve purification results, electrophoresis results and electron microscope structure diagram of the pore protein according to an embodiment of the present application;

[0045] FIG. 23 is a negative staining electron microscope image of the pore protein graft according to an embodiment of the present application;

[0046] FIG. 24 is a negative staining electron microscope image of the pore protein graft according to an embodiment of the present application. DETAILED DESCRIPTION

[0047] The present application will be further described in detail below with specific embodiments, and the embodiments given are only to illustrate the present application and not to limit the scope of the present application. The embodiments provided below can serve as a guide for further improvement by those of ordinary skill in the art, and do not constitute any limitation on the present application in any way.

[0048] The present application is based on the following recognition of the inventor:

[0049] A nanopore-based single-molecule sequencer is a detection system highly integrated with multiple disciplines and technologies. The development of such an instrument requires deep cross-disciplinary and collaborative innovation in physics, biology, chemistry, semiconductors, computers, and other disciplines. Starting from the bottom core module, a high-precision single-molecule nanopore sequencing system is constructed. The commercial nanopore sequencers currently available, such as MinION, GridION and PromethION of Oxford Nanopore Technologies in the United Kingdom, QNome-3841 of Zibec, etc., all have great deficiencies in sequencing accuracy, throughput, chip stability and applicable scenarios, and cannot meet the high-precision needs of molecular biology research.

[0050] Nanopore sequencing requires that the sensing region inside the pore protein is sharp enough to have high spatial resolution both laterally and longitudinally. So far, only a few natural proteins such as Mycobacterium smegmatis pore protein A (MspA), curli-specific transport channel (CsgG), etc. can meet the requirements of becoming a single molecule detector in the industry, but the structures of these pore proteins are relatively single, and their sequencing accuracy is still insufficient due to this limitation.

[0051] Based on this, the inventors of the present application have, through many experiments and tests, excavated a new type of pore protein monomer from deep sea metagenome (derived from a sample of the Mariana Trench 11,000 meters deep), and named it BP62. Multiple such monomers can form a pore protein with a nanopore by covalent connection or non-covalent polymerization. Moreover, by artificial modification, such as substitution, deletion and / or addition of one or more amino acids based on the sequence of the pore protein monomer, a mutant of the protein monomer can be obtained. Whether it is the pore protein or its mutant, it has the characteristics of high stability and excellent internal structure, so they can be used as detection proteins and applied to the detection of small molecules such as nucleotides, amino acids, sugars, vitamins, etc., or to DNA, RNA or polypeptide sequencing based on nanopores. The pore protein and its mutant can sensitively characterize the nucleic acid sequence or polypeptide sequence passing through the pore to achieve higher accuracy nanopore sequencing.

[0052] In some embodiments, the pore protein monomer comprises a barrel segment having the function of forming a pore structure by polymerization. In some embodiments, the barrel segment has an amino acid sequence as shown in SEQ ID NO: 1 or an amino acid sequence having 50% and above identity to the amino acid sequence shown in SEQ ID NO: 1, for example, an amino acid sequence having at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.91%, 99.92%, 99.93%, 99.94%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99%, or any value therebetween but less than 100% identity, and the mutated barrel segment can form a pore structure by polymerization.

[0053] MVYIEALIMEVSVDKAFELGIEWSVFKDTNIGNDDAVVGGGFNASSGFSNAAS NLLYSSGGNVGVVSGTLDLTLGGTTVTIPSLGALIKALQTNEDVHILQTPKLL ATDNKEATIKVGKVIPYQTRVSTTDNETYNTYDYKDVGLTLKITPHISEDRTV RLEIFQELSALTGASAEITTTLPTTLNRTVETTVVVEDQNMLVIGGLIDDTYTN KVTKVPLLGSIPLLGRLFRYDSQVGNKTNLYIFITPRVV (SEQ ID NO: 1, barrel segment)

[0054] In embodiments, the barrel segment of the pore-forming protein monomer has an amino acid sequence with one or more substitutions, deletions, and / or additions of amino acids compared to the amino acid sequence set forth in SEQ ID NO: 1, and the barrel segment has the function of polymerizing to form a pore structure. For example, the barrel segment of the pore-forming protein monomer has at least 1-200 amino acid mutations compared to SEQ ID NO: 1. In some embodiments, the barrel segment of the pore-forming protein monomer has at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20 amino acid mutations compared to SEQ ID NO: 1. In embodiments, the barrel segment of the pore-forming protein monomer has at least 21-100 amino acid mutations compared to SEQ ID NO: 1. These barrel segments with mutations also have the function of polymerizing to form a pore structure.

[0055] In embodiments, the barrel segment of the pore-forming protein monomer comprises a first sensor structure region, a transmembrane region, and an exit structure region, wherein the first sensor structure region is configured to form a first narrow pore in the pore structure. Specifically, as shown in FIG. 1, the first sensor structure region can comprise a beta sheet with a curvature configured to form a first narrow pore in the pore structure to sense a molecule to be detected passing through the first narrow pore.

[0056] In the embodiments of the present application, as shown in FIG. 1, the middle part (i.e., barrel section) of the pore monomer has a curved β-sheet extending towards the C-terminal direction, and when a plurality of pore monomers are polymerized to form a pore, the curved β-sheets can enclose a narrow pore (i.e., a first narrow pore), which can serve as a sensor for large molecules such as nucleic acids, and can sense different current changes when different large molecules pass through the pore, so as to distinguish different pore-passing large molecules, thereby achieving high-accuracy nanopore sequencing. In the embodiments of the present application, the first sensor structure region corresponding to the first narrow pore is also referred to as a central sensor.

[0057] In the embodiments of the present application, the outlet structure region of the barrel section of the pore monomer can include an α-helix and a β-sheet. In addition, the barrel section of the pore monomer proposed in the embodiments of the present application further includes a transmembrane region, which is in contact with the membrane layer after the monomers are polymerized to form a pore, so as to achieve stable pore embedding of the pore. In some embodiments, the transmembrane region can be located adjacent to the outlet structure region, for example, on the outside of the outlet structure region.

[0058] In the embodiments of the present application, the barrel section of the pore monomer in any of the above embodiments can also be modified to optimize the performance of the monomer. Specifically, each part of the barrel section can be subjected to directed mutation to optimize the sequencing performance of the pore monomer.

[0059] In the embodiments of the present application, the first sensor structure region of the barrel section can be modified to optimize the sequencing performance of the pore monomer, for example, polar uncharged amino acids or non-polar amino acids such as A and G can be introduced into the first sensor structure region to modify the first sensor structure region, wherein the polar uncharged amino acids include S, T, Y, N, Q and C. In some embodiments, the polar uncharged amino acids and / or non-polar amino acids are introduced into the first sensor structure region by substitution and / or addition (insertion).

[0060] In some embodiments, the polar uncharged amino acids or non-polar amino acids can be introduced into any core site of the first sensor structure region to improve the molecular detection performance of the pore monomer. In some embodiments, the core site can be an amino acid site located at the loop between the curved β-sheets of the first sensor structure region, for example, positions 126 to 140, particularly positions 130 to 135, of the amino acid sequence shown in SEQ ID NO: 1.

[0061] It can be understood that by modifying the amino acids of the core site of the central sensor structure region of the pore monomer to polar uncharged amino acids or non-polar amino acids, the binding and capture of negatively charged and nucleic acids by the central sensor can be effectively improved, and the quality of the current signal can be improved, thereby improving the molecular detection performance of the pore.

[0062] In some embodiments, the transmembrane region of the barrel segment of the pore-forming protein monomer can be modified by substituting one or more of the polar amino acids at positions 221, 227, 20, 21, 23, 25, 84, 86, 88, 90, 65, and 63 of the transmembrane region of the pore-forming protein monomer having the amino acid sequence of SEQ ID NO: 1 with a highly hydrophobic amino acid to improve the detection performance of the pore-forming protein monomer. It can be appreciated that by modifying the polar amino acids and / or non-polar amino acids in the transmembrane inner region and the transmembrane interface region of the transmembrane region with highly hydrophobic amino acids, the distribution of the pore-forming protein on the membrane can be improved, and the binding stability between the pore-forming protein and the membrane can be enhanced, thus effectively improving the detection efficiency and quality of molecules.

[0063] In some embodiments, the transmembrane region of the barrel segment of the pore-forming protein monomer can be modified by substituting one or more of the polar amino acids at positions 221, 227, 20, 21, 23, 25, 84, 86, 88, 90, 65, and 63 of the transmembrane region of the pore-forming protein monomer having the amino acid sequence of SEQ ID NO: 1 with a highly hydrophobic amino acid to improve the detection performance of the pore-forming protein monomer. It can be appreciated that by modifying the polar amino acids and / or non-polar amino acids in the transmembrane inner region and the transmembrane interface region of the transmembrane region with highly hydrophobic amino acids, the distribution of the pore-forming protein on the membrane can be improved, and the binding stability between the pore-forming protein and the membrane can be enhanced, thus effectively improving the detection efficiency and quality of molecules.

[0064] The pore-forming protein monomer according to the embodiments of the present application can further comprise a second sensor structure region (also referred to as a “cap peptide segment”) that can be grafted to the outlet structure region of the barrel segment of the pore-forming protein monomer to form a second narrow pore in the pore structure.

[0065] In some embodiments, the second sensor structure region can have a double alpha helix structure, which can replace or be inserted into the alpha helix of the outlet structure region of the barrel segment of the pore-forming protein monomer (as shown in the left side of FIG. 5) to form a second narrow pore in the pore structure for sensing the molecules to be detected passing through the second narrow pore. In the embodiments of the present application, the second sensor structure region corresponding to the second narrow pore is also referred to as a cap sensor.

[0066] In some embodiments, the second sensor structure region can have a double alpha helix structure and a beta fold structure. For example, the double alpha helix structure and the beta fold structure with curvature are beneficial for enclosing the second narrow pore in the pore structure to sense the molecules to be detected passing through the second narrow pore.

[0067] In some embodiments, the secondary structure of the second sensor structure region is in the form of an extended strip, which accesses the alpha helix of the exit structure region in the horizontal direction, and can be enclosed at a delicate angle in the pore protein formed by the aggregation of multiple pore protein monomers to form a second narrow pore (i.e. cap gate sensor). Like the central sensor, the second narrow pore can serve as a sensor for macromolecules such as nucleic acids or biological small molecules, and can sense different current changes when different molecules pass through the pore, thereby distinguishing different pore molecules and achieving high-accuracy nanopore sequencing or detection.

[0068] In some embodiments, the second sensor structure region can have any of the amino acid sequences shown in SEQ ID NOs: 2-9. The right side of FIG. 5 shows the structure of the second sensor structure region of SEQ ID NOs: 2-9 of the embodiments of the present application.

[0069] Cap 1: IGKIAAGALAARSIPGTPGSTVTDVNGNVTVNPNGNSTRGDYS (SEQ ID NO: 2)

[0070] Cap 2: LAVLLGRLFGSGGDESDPDADADDEVEIGPIEAVQA (SEQ ID NO: 3)

[0071] Cap 3: IGSVAAGLEAAKTTKGERVPEVRAADTGELITAAYTKNDTKGDFS (SEQ ID NO: 4)

[0072] Cap 4: AIGAAAYSGRGYTTKASQSTTGSGDNQNTTNNPDIVTRADYT (SEQ ID NO: 5)

[0073] Cap 5: IGSIAGALQANRGQDTTTTTVIDPDTGVITETSNPNQGDNGQ (SEQ ID NO: 6)

[0074] Cap 6: IGNVMIGLEEAKDTTQTKAVYDTNNNFLRNETTTTKGDYTKL (SEQ ID NO: 7)

[0075] Cap 7: IGTLGAAISAAKPQKGSTVISENGATTINPDTNGDLS (SEQ ID NO: 8)

[0076] Cap 8: IGKVMVGLEEAKDKTVTDSRWNSDTDKYEPYSRTEAGDYS (SEQ ID NO: 9)

[0077] In the embodiments of the present application, the second sensillum structure region can replace or be inserted into, preferably replace, the V38-G85 peptide segment or any peptide segment therebetween, preferably F48-V65 or any peptide segment therebetween (based on SEQ ID NO: 1) of the outlet structure region of the barrel segment of the pore-forming protein monomer as described in any of the above embodiments. In some embodiments, the pore-forming protein monomer grafted with the second sensillum structure region has an amino acid sequence as shown in SEQ ID NO: 10, wherein the second sensillum structure region corresponds to SEQ ID NO: 9:

[0078] MVYIEALIMEVSVDKAFELGIEWSVFKDTNIGNDDAVVGGGFNASSGIGKVMVGLEEAKDKTVTDSRWNSDTDKYEPYSRTEAGDYSVSGTLDLTLGGTTVTIPSLGALIKALQTNEDVHILQTPKLLATDNKEATIKVGKVIPYQTRVSTTDNETYNTYDYKDVGLTLKITPHISEDRTVRLEIFQELSALTGASAEITTTLPTTLNRTVETTVVVEDQNMLVIGGLIDDTYTNKVTKVPLLGSIPLLGRLFRYDSQVGNKTNLYIFITPRVV (SEQ ID NO: 10, wherein the underlined portion is the second sensillum structure region as shown in SEQ ID NO: 9)

[0079] Compared with the single-sensillum pore-forming protein in the related art, the pore-forming protein containing a double sensillum proposed in the embodiments of the present application has greater advantages in resolving repetitive nucleotide sequences or amino acid sequences, because the double sensillum contained therein provides higher nucleotide / amino acid resolution, thereby greatly improving the sequencing accuracy.

[0080] It can be understood that the second sensillum structure region proposed in the embodiments of the present application can also be other sequences, as long as the shape and size of the protein formed by the sequence can form a narrow pore as a molecular sensillum; in addition, in some embodiments, the second sensillum structure region can also be grafted to other regions of the pore-forming protein monomer or its mutants, as long as it does not interfere with the original shape and original function of the monomer after grafting. These other sequences of the second sensillum structure region and other grafting positions also fall within the protection scope of the present application.

[0081] In the embodiments of the present application, the second poreceptor structure region can also be modified to optimize the sequencing performance of the pore monomer. In some embodiments, a polar uncharged amino acid or a non-polar amino acid such as A and G, etc. can be introduced into any core site of the second poreceptor structure region to modify the second poreceptor structure region, wherein the polar uncharged amino acid includes S, T, Y, N, Q and C. In some embodiments, the core site of the pore monomer having the amino acid sequence shown in SEQ ID NO: 10 is the loop region between the curved beta sheets of the second poreceptor structure region, for example, the 68th to 75th of the amino acid sequence shown in SEQ ID NO: 10.

[0082] It can be understood that, after the above modification of the cap poreceptor structure region of the pore monomer, the pore monomer containing the mutant amino acid can effectively improve the binding and capture of the cap poreceptor to the negatively charged nucleic acid and improve the current signal quality, thereby improving the molecular detection and sequencing performance of the pore monomer.

[0083] In the embodiments of the present application, in addition to the barrel segment described in any of the above embodiments, the pore monomer can also comprise an N-terminal segment, and the pore monomer comprises the N-terminal segment and the barrel segment from N-terminal to C-terminal, wherein the N-terminal segment comprises one or more domains selected from N0, N1, N2 and N3 domains, and preferably comprises N3 domain. At this time, the N-terminal segment serves as an entrance structure region of the pore monomer, and is connected to the barrel segment forming the pore structure, and one or more domains of N0, N1, N2 and N3 domains constitute the entrance structure region and play similar or same functions. In some embodiments, the sequence of the N3 domain is shown in SEQ ID NO: 11. In some embodiments, the sequence of the N-terminal segment is shown in SEQ ID NO: 12, which comprises N0, N1, N2 and N3 domains.

[0084] GKIQVYKLNNAEAEEMAQVLQNLQKGTSGSSDAGKEAIFVSSDISITADKATNSLLIMADSADYATIEGIINELDVPKS (SEQ ID NO: 11, N3 domain)

[0085] AHTILNMVDTDITVLIETISAMTGKSFIIDGNVKGKVNVVSPEKISPEEAYRLFESVLEVNGFALVPSGKFIKVVPAAEARTKSIDTRVDTQKLSNTRNPDDRAITQLIHLKYADAKDVKTLFTPLVSKNSIIQAYADTNTLIIYDVESNIKRLLHILTAIDIPGTGREITTIPLAHADADTLSKTLATIFKTSMSEEKDVSQKTLQFVADERTNTIIFVASEDDTLRIKKLVSLLDKDTPKGKGKIQVYKLNNAEAEEMAQVLQNLQKGTSGSSDAGKEAIFVSSDISITADKATNSLLIMADSADYATIEGIINELDVPKS (SEQ ID NO: 12, N-terminal segment, comprising N0, N1, N2 and N3 domains)

[0086] In some embodiments, the pore-forming protein monomer comprises a S segment and / or a signal peptide in addition to the barrel segment. For example, in some embodiments, the pore-forming protein monomer comprises, from N-terminus to C-terminus, a barrel segment and a S segment. In some embodiments, the pore-forming protein monomer comprises, from N-terminus to C-terminus, an N-terminal segment, a barrel segment and a S segment. In this case, the S segment is a C-terminal structural segment of the pore-forming protein, which is connected to the barrel segment that forms the pore structure, so that the pore-forming protein monomer comprises, from N-terminus to C-terminus, an N-terminal segment, a barrel segment and a S segment. In some embodiments, the pore-forming protein monomer comprises, from N-terminus to C-terminus, a signal peptide, an N-terminal segment, a barrel segment and a S segment. In this case, the S segment is a signal peptide that helps the expression of the pore-forming protein, which is connected to the N-terminal segment of the entry structural segment, so that the pore-forming protein monomer comprises, from N-terminus to C-terminus, a signal peptide, an N-terminal segment, a barrel segment and a S segment. In some embodiments, the sequence of the S segment can be as set forth in SEQ ID NO: 13. In some embodiments, the sequence of the signal peptide can be as set forth in SEQ ID NO: 14.

[0087] KNPSEGTRLTDDLQDTLTPEETGIIKLYGPLENSERSMAPNTTEQPSE (SEQ ID NO: 13, S segment)

[0088] MKPTADVVTFMRRHKQREHRPKQKQFIMKRCYMTRLHRFQTTAFRFFPLMLACWLLLSFSFLPGAFAQNKNTTSSESGVASDAN (SEQ ID NO: 14, signal peptide)

[0089] In the embodiments of the present application, the detection performance of the pore-forming protein monomer can also be optimized by performing targeted mutation on each of the above-mentioned portions of the pore-forming protein monomer other than the barrel segment, such as the N-terminal segment, the signal peptide, and / or the S segment.

[0090] In some embodiments, the entrance structure region corresponding to the N-terminal segment of the pore-forming protein monomer can be engineered, such as by performing synonymous mutation and / or sequence truncation (i.e., deletion) of amino acids on the structure region. For example, as shown in FIG. 1, the complete N0-N3 domains in the N-terminal segment can be retained as the entrance structure region, in which case the entrance end is the N0 domain; in other embodiments, the N-terminal segment can be truncated by deletion mutation, such that any three, any two, or even any one of the N0-N3 domains is included as the entrance structure region, and the entrance end can be the N0, N1, N2, or N3 domain. In some embodiments, the N0, N1, and N2 domains can be deleted, and only the N3 domain is retained as the entrance structure region. It can be understood that the entrance structure region above the central sensor structure region can effectively capture the molecule to be detected, so that it enters the entrance lumen composed of the entrance structure region and passes through the central sensor, thereby achieving effective detection of the molecule to be detected.

[0091] In the embodiments of the present application, the detection performance of the pore-forming protein monomer can also be optimized by engineering one or more amino acids in the entrance structure region, such as the entrance end thereof. In some embodiments, the entrance structure region comprises polar charged amino acids, wherein the polar charged amino acids include polar positively charged amino acids and polar negatively charged amino acids, then a polar positively charged amino acid or a polar uncharged amino acid can be introduced into the site of the polar negatively charged amino acid in the entrance structure region, and / or a polar uncharged amino acid can be introduced into the site of the polar positively charged amino acid in the entrance structure region, to engineer the pore-forming protein monomer. In some embodiments, the polar uncharged amino acid is selected from S, T, Y, N, Q, and C, and the polar positively charged amino acid is selected from K, R, and H. In some embodiments, a non-polar amino acid selected from A and G can also be introduced into the site of the polar charged amino acid. In some embodiments, the above-mentioned polar uncharged amino acid and / or non-polar amino acid selected from A and G can be introduced into any one or more of the positions K2, K7, K25, K35, K50, E12, E14, E15, D32, E36, D43, D49, D60, D63, E68, E73, and D75 of the amino acid sequence shown in SEQ ID NO: 11.

[0092] The entrance end of the central sensor structure region of the pore protein monomer and the mutant thereof with the entrance end modified in the embodiment of the present application comprises a high proportion of polar positively charged amino acids, polar uncharged amino acids, amino acid A or amino acid G, thereby effectively capturing the nucleic acid molecules to be tested, realizing high-strength binding of the molecules to be tested and the pore protein, thus improving the sequencing speed and the quality of the current signal generated in sequencing, thereby realizing high-throughput and high-quality sequencing.

[0093] It can be understood that, in addition to the above-mentioned truncation and modification, the pore protein monomer of the embodiment of the present application can also comprise other mutant forms compared with SEQ ID NO: 11-14, for example: the alpha helix and random coil of the S segment are modified to have a truncated S segment or a new alpha helix or other structure, etc.; other alpha helix is added in the N-terminal segment to form N4, N5, N6 and other domains, or other structure similar domains are used as the entrance structure region, or substitution between polar positively charged amino acids and / or substitution between polar uncharged amino acids is performed in the entrance structure region, as long as these mutant forms do not affect the structure and function of the pore protein monomer; in addition, the pore protein monomer of the present application can also comprise silent mutations, non-natural amino acids and / or protein modifications based on the amino acids of SEQ ID NO: 1-14, and these variants are all included in the protection scope of the present application.

[0094] It should be noted that the pore protein in the embodiment of the present application forms a pore structure based on the barrel segment contained therein to play the function of the pore protein, that is, the pore protein containing only the barrel segment (including the barrel segment containing only the first sensor structure region or containing the first and second sensor structure regions) or the mutant thereof can form a pore structure to play the function of the pore protein, and the N-terminal segment, the signal peptide and the S segment and the like are included in the pore protein monomer as additional structures, which will be beneficial to the expression of the pore protein monomer in the cell, the insertion of the pore protein into the membrane layer to form a more stable nanopore sensor, or the interaction of the pore protein with other proteins, etc., but it can be understood that these structures are not necessary. At the same time, the mutation of these additional structures or their replacement forms also fall within the protection scope of the present application.

[0095] In the embodiment of the present application, the pore protein monomer can have an amino acid sequence as shown in SEQ ID NO: 15 or an amino acid sequence having 50% and above identity with the amino acid sequence shown in SEQ ID NO: 15 corresponding to each part as shown in SEQ ID NO: 1 and the like, and the monomer can form a pore structure after polymerization.

[0096] In embodiments of the present application, the term "percent identity" with respect to a nucleic acid or polypeptide sequence is defined as the percentage of nucleotides or amino acid residues in the candidate sequence that are identical with the known polypeptide after aligning the sequences for maximum percent identity and introducing gaps, if necessary, to achieve the maximum percent homology. N-terminal or C-terminal insertions or deletions are not to be construed as affecting homology. Homology or identity at the nucleotide or amino acid sequence level can be determined by BLAST (Basic Local Alignment Search Tool) analysis using the algorithm employed by the programs blastp, blastn, blastx, tblastn and tblastx (Altschul (1997), Nucleic Acids Res 25, 3389-3402 and Karlin (1990), Proc. Natl. Acad. Sci. USA 87, 2264-2268), which are tailored for sequence similarity searching.

[0097] The pore-forming monomer according to embodiments of the present application can comprise an amino acid sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, 99.91%, 99.92%, 99.93%, 99.94%, 99.95%, 99.96%, 99.97%, 99.98%, 99.99%, or any value therebetween, but less than 100% identity compared to SEQ ID NO: 15.

[0098] The pore-forming protein monomer sequence provided in the embodiments of the present application has one or more amino acid mutations compared to SEQ ID NO: 15, wherein the mutations include substitutions, deletions, and / or additions (insertions), which can occur in the barrel segment as shown in SEQ ID NO: 1 and / or in the signal peptide, N-terminal segment, and / or S segment as shown in SEQ ID NO: 11-14, respectively, and the mutated monomer still has the function of forming a pore structure after polymerization. In some embodiments, the pore-forming protein monomer has at least 1-200 amino acid mutations compared to SEQ ID NO: 15. In some embodiments, the pore-forming protein monomer has at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20 amino acid mutations compared to SEQ ID NO: 15. In the embodiments of the present application, the pore-forming protein monomer has at least 21-100 amino acid mutations compared to SEQ ID NO: 15. In the embodiments of the present application, unless otherwise specified, the remaining amino acids of the pore-forming protein mutated monomer are the same as those of SEQ ID NO: 15, except for the indicated mutations.

[0099] FIG. 1 is a schematic diagram of the structure of a pore-forming protein monomer according to an embodiment of the present application. As shown in FIG. 1, the pore-forming protein monomer (e.g., having an amino acid sequence as shown in SEQ ID NO: 15) can sequentially comprise, from the N-terminal to the C-terminal: a signal peptide (not indicated, the sequence of which is shown in SEQ ID NO: 14), an N-terminal segment (the sequence of which is shown in SEQ ID NO: 12), a barrel segment (the sequence of which is shown in SEQ ID NO: 1), and an S segment (the sequence of which is shown in SEQ ID NO: 13), wherein the barrel segments of a plurality of pore-forming protein monomers can form a pore structure through polymerization.

[0100] Further, as shown in FIG. 1, the entry structure region of the pore-forming protein monomer having an amino acid sequence as shown in SEQ ID NO: 15 can comprise N0, N1, N2, and N3 domains. In some embodiments, the N-terminal segment can comprise one or more alpha helices. In addition, as shown in FIG. 1, the C-terminal of the pore-forming protein monomer having an amino acid sequence as shown in SEQ ID NO: 15 comprises an S segment, wherein the S segment can comprise an alpha helix and a random coil (long loop).

[0101] It is understood that the pore monomer based on the sequence shown in SEQ ID NO: 15 can also correspondingly contain the mutation forms occurring in each domain shown in the sequences of SEQ ID NO: 14, 12, 1, 13, etc. described in any of the above embodiments, that is, the substitution, deletion and / or addition of amino acids can correspondingly occur in the signal peptide, N-terminal segment, barrel segment and / or S segment contained in SEQ ID NO: 15 to improve the detection performance of the pore monomer.

[0102] In some embodiments, the barrel segment of the pore monomer based on the amino acid sequence shown in SEQ ID NO: 15 can be modified to optimize the sequencing performance of the pore monomer. Corresponding to the amino acid mutations occurring on the barrel segment sequence shown in SEQ ID NO: 1 described above, each structural region of the barrel segment of the pore monomer shown in SEQ ID NO: 15 (i.e. SEQ ID NO: 1 therein), such as the first sensor structural region, the transmembrane region, etc. can be mutated to optimize the sequencing performance of the pore monomer.

[0103] In some embodiments, the first sensor structural region of the barrel segment can be modified to optimize the sequencing performance of the pore monomer, and the specific mutation direction is described above for the mutation of the first sensor structural region of the barrel segment based on SEQ ID NO: 1, which is not repeated here. In some embodiments, the polar non-charged amino acid or non-polar amino acid can be introduced into any core site of the first sensor structural region to improve the molecular detection performance of the pore monomer. In some embodiments, the core site can be the amino acid site located at the loop between the curved β-pleated sheets of the first sensor structural region, such as position 533, position 546, and positions 537-542 shown in Figure 2 (corresponding to positions 126, 139, and 130-135 of the barrel segment shown in SEQ ID NO: 1).

[0104] In some embodiments, the amino acids at positions 533, 546, and 537-542 of the pore monomer having the amino acid sequence shown in SEQ ID NO: 15 are R533, D546, T537, D538, N539, E540, T541, and Y542, which respectively correspond to R126, D139, T130, D131, N132, E133, T134, and Y135 of the barrel segment shown in SEQ ID NO: 1. For these core sites, based on SEQ ID NO: 15 and SEQ ID NO: 1 respectively, the mutation forms thereof can be one or more of the following:

[0105] T537 (T130) is substituted with any amino acid selected from the group consisting of G, A, S, N, Q and Y, preferably G, and / or inserted with any amino acid selected from the group consisting of G, A, S, N, Q and Y, preferably G, or a combination thereof;

[0106] D538 (D131) is substituted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y, preferably selected from the group consisting of G, T and Q, and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y, preferably selected from the group consisting of G, T and Q, or a combination thereof;

[0107] N539 (N132) is substituted with any amino acid selected from the group consisting of G, A, S, T, Q and Y, preferably selected from the group consisting of A and Y, and / or inserted with any amino acid selected from the group consisting of G, A, S, T, Q and Y, preferably selected from the group consisting of A and Y, or a combination thereof;

[0108] E540 (E133) is substituted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y, preferably selected from the group consisting of G and Q, and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y, preferably selected from the group consisting of G and Q, or a combination thereof;

[0109] T541 (T134) is substituted with any amino acid selected from the group consisting of G, A, S, N, Q and Y, and / or inserted with any amino acid selected from the group consisting of G, A, S, N, Q and Y, or a combination thereof;

[0110] Y542 (Y135) is substituted with any amino acid selected from the group consisting of G, A, S, T, N and Q, preferably A, and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N and Q, preferably A, or a combination thereof;

[0111] R533 (R126) is substituted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y, preferably A, and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y, preferably A, or a combination thereof; and

[0112] D546 (D139) is substituted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y, preferably Q, and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y, preferably Q, or a combination thereof.

[0113] In some embodiments, for the core positions, based on SEQ ID NO: 15 and SEQ ID NO: 1, respectively, a mutated form thereof can be selected from one or more of the following: T537G (T130G), D538G (D131G) / D538T (D131T) / D538Q (D131Q), N539A (N132A) / N539Y (N132Y), E540G (E133G) / E540Q (E133Q), Y542A (Y135A), R533A (R126A), and D546Q (D139Q).

[0114] In some embodiments, the first sensor structure region can have the following mutant forms based on SEQ ID NO: 15 and SEQ ID NO: 1, respectively:

[0115] R533A(R126A), D546Q(D139Q), D538T(D131T), and E540Q(E133Q);

[0116] R533A(R126A), D546Q(D139Q), D538Q(D131Q), and N539Y(N132Y);

[0117] R533A(R126A), D546Q(D139Q), D538Q(D131Q), E540G(E133G), Y542A(Y135A), and T537G(T130G); or

[0118] D538G(D131G) and N539A(N132A).

[0119] It can be understood that, through the above modification of the central sensor structure region of the pore monomer, the pore monomer containing the mutant amino acid can effectively improve the binding and capture of the central sensor to the negatively charged and nucleic acid and improve the current signal quality, thereby improving the molecular detection performance of the pore.

[0120] In some embodiments, the sequencing performance of the pore monomer can be optimized by modifying the transmembrane region of the barrel segment. For specific mutation directions, see the description of the mutation of the transmembrane region of the barrel segment based on SEQ ID NO: 1 in the above embodiments, which will not be repeated here. In some embodiments, as shown in FIG. 4, the high hydrophobic amino acid can be introduced by substituting any one or more of positions 628, 634, 427, 428, 430, 432, 491, 493, 495, 497, 472, and 470 of the transmembrane region of the pore monomer having the amino acid sequence shown in SEQ ID NO: 15, which correspond to positions 221, 227, 20, 21, 23, 25, 84, 86, 88, 90, 65, and 63 of the barrel segment shown in SEQ ID NO: 1, respectively. The non-polar amino acids corresponding to these positions can include L628, L634, G427, I428, W430, V432, L491, A493, I495, A497, V472, and V470, which correspond to L221, L227, G20, I21, W23, V25, L84, A86, I88, A90, V65, and V63 of the barrel segment shown in SEQ ID NO: 1, respectively. For these core positions, the mutant forms based on SEQ ID NO: 15 and SEQ ID NO: 1, respectively, can be one or more of the following:

[0121] L628 (L221) is substituted with any amino acid selected from G, A, V, I, Y, F or W;

[0122] L634 (L227) is substituted with any amino acid selected from G, A, V, I, Y, F or W;

[0123] G427 (G20) is substituted with any amino acid selected from A, V, L, I, Y, F or W;

[0124] I428 (I21) is substituted with any amino acid selected from A, G, V, L, Y, F or W;

[0125] W430 (W23) is substituted with any amino acid selected from A, G, V, L, I, Y or F;

[0126] V432 (V25) is substituted with any amino acid selected from A, G, L, I, Y, F or W;

[0127] L491 (L84) is substituted with any amino acid selected from G, A, V, I, Y, F or W;

[0128] A493 (A86) is substituted with any amino acid selected from G, V, L, I, Y, F or W;

[0129] I495 (I88) is substituted with any amino acid selected from A, G, V, L, Y, F or W;

[0130] A497 (A90) is substituted with any amino acid selected from G, V, L, I, Y, F or W;

[0131] V472 (V65) is substituted with any amino acid selected from A, G, L, I, Y, F or W; and

[0132] V470 (V63) is substituted with any amino acid selected from A, G, L, I, Y, F or W.

[0133] It can be understood that the pore-forming protein monomer and the pore-forming protein mutant with the transmembrane region modified with respect to the polar amino acid and / or the non-polar amino acid in the transmembrane region have less polar charged amino acids in the transmembrane region and have a higher proportion of highly hydrophobic amino acids in the transmembrane inner region and the transmembrane interface region, thereby improving the distribution rate of the pore-forming protein on the membrane and enhancing the binding stability between the pore-forming protein and the membrane, thus effectively promoting the molecular detection efficiency and quality.

[0134] In some embodiments, the N-terminal segment of the pore-forming protein monomer is truncated to contain only the N3 domain, and the entrance end of the N-terminal segment is the N3 domain. As shown in FIG. 3, which is a structural schematic of the entrance structure region according to embodiments of the present application, the N-terminal segment of the pore-forming protein monomer is truncated to contain only the N3 domain, and the entrance end of the N-terminal segment is the N3 domain. In some embodiments, a non-polar amino acid and / or a polar uncharged amino acid can be introduced into the pore-forming protein having the amino acid sequence set forth in SEQ ID NO: 15 at any one or more of positions 330, 335, 353, 363, 378, 340, 342, 343, 360, 364, 371, 377, 388, 391, 396, 401, and 403, which correspond to positions 2, 7, 25, 35, 50, 12, 14, 15, 32, 36, 43, 49, 60, 63, 68, 73, and 75, respectively, of the amino acid sequence set forth in SEQ ID NO: 11. For example, the entrance end contains the polar positively charged amino acids K330 (corresponding to K2 in SEQ ID NO: 11, and the same below), K335 (K7), K353 (K25), K363 (K35), and K378 (K50); and the polar negatively charged amino acids E340 (E12), E342 (E14), E343 (E15), D360 (D32), E364 (E36), D371 (D43), D377 (D49), D388 (D60), D391 (D63), E396 (E68), E401 (E73), and D403 (D75). In some embodiments, the replacement amino acid of K330 (K2), K335 (K7), K353 (K25), K363 (K35), and K378 (K50) is independently selected from the group consisting of N, A, G, S, T, and Q; and the replacement amino acid of E340 (E12), E342 (E14), E343 (E15), D360 (D32), E364 (E36), D371 (D43), D377 (D49), D388 (D60), D391 (D63), E396 (E68), E401 (E73), and D403 (D75) is independently selected from the group consisting of K, R, N, A, G, S, T, and Q.

[0135] The entrance end above the central sensor structure region of the pore monomer and the mutant thereof with the entrance end modified of the embodiments of the present application comprises a high proportion of polar positively charged amino acids, polar uncharged amino acids, amino acid A or amino acid G, thereby effectively capturing the nucleic acid molecules to be tested, realizing high strength binding of the molecules to be tested and the pore, thus improving the sequencing speed and the quality of the current signal generated in sequencing, thereby realizing high-throughput and high-quality sequencing.

[0136] The pore monomer (for example, the sequence shown in SEQ ID NO: 15) of the embodiments of the present application can also comprise a transplantable second sensor structure region. The description of the transplantable second sensor structure region in the above embodiments also applies to this, and thus will not be repeated. In some embodiments, based on the pore monomer shown in SEQ ID NO: 15, the second sensor structure region can replace or be inserted into, preferably replace, the V445-G492 (corresponding to V38-G85 of SEQ ID NO: 1) peptide segment or any peptide segment therebetween, for example, F455-V472 (corresponding to F48-V65 of SEQ ID NO: 1) of the exit structure region of the barrel segment. In some embodiments, the pore monomer with the second sensor structure region transplanted has an amino acid sequence shown in SEQ ID NO: 16, wherein the second sensor structure region corresponds to SEQ ID NO: 9:

[0137] MKPTADVVTFMRRHKQREHRPKQKQFIMKRCYMTRLHRFQTTAFRFFPLMLACWLLLSFSFLPGAFAQNKNTTSSESGVASDANAHTILNMVDTDITVLIETISAMTGKSFIIDGNVKGKVNVVSPEKISPEEAYRLFESVLEVNGFALVPSGKFIKVVPAAEARTKSIDTRVDTQKLSNTRNPDDRAITQLIHLKYADAKDVKTLFTPLVSKNSIIQAYADTNTLIIYDVESNIKRLLHILTAIDIPGTGREITTIPLAHADADTLSKTLATIFKTSMSEEKDVSQKTLQFVADERTNTIIFVASEDDTLRIKKLVSLLDKDTPKGKGKIQVYKLNNAEAEEMAQVLQNLQKGTSGSSDAGKEAIFVSSDISITADKATNSLLIMADSADYATIEGIINELDVPKSMVYIEALIMEVSVDKAFELGIEWSVFKDTNIGNDDAVVGGGFNASSGIGKVMVGLEEAKDKTVTDSRWNSDTDKYEPYSRTEAGDYSVSGTLDLTLGGTTVTIPSLGALIKALQTNEDVHILQTPKLLATDNKEATIKVGKVIPYQTRVSTTDNETYNTYDYKDVGLTLKITPHISEDRTVRLEIFQELSALTGASAEITTTLPTTLNRTVETTVVVEDQNMLVIGGLIDDTYTNKVTKVPLLGSIPLLGRLFRYDSQVGNKTNLYIFITPRVVKNPSEGTRLTDDLQDTLTPEETGIIKLYGPLENSERSMAPNTTEQPSE (SEQ ID NO: 16, wherein the underlined portion is a second sensor structure region as set forth in SEQ ID NO: 9).

[0138] It can be understood that the second sensor structure region shown as SEQ ID NO: 9 in SEQ ID NO: 16 can also be replaced by the second sensor structure region as described in any one of SEQ ID NOs: 2-8 to obtain a porin monomer comprising two sensor structure regions. In other embodiments, the second sensor structure region shown as SEQ ID NO: 9 in SEQ ID NO: 16 can also be replaced by other sequences as long as the shape and size of the protein formed by the sequence can form a narrow pore as a molecular sensor; in addition, in some embodiments, the second sensor structure region can also be transplanted to other regions of the porin monomer or its mutants as long as it does not interfere with the original shape and original function of the monomer after transplantation, and these other sequences of the second sensor structure region and other transplantation positions also fall within the protection scope of the present application.

[0139] In the embodiments of the present application, the second sensor structure region of the porin monomer can also be modified, and the specific mutation direction is described above for the mutation of the second sensor structure region, which will not be repeated here. In some embodiments, the core site of the porin having the amino acid sequence shown as SEQ ID NO: 16 is the loop region between the curved β sheets of the second sensor structure region, for example, the 475th to 482nd positions shown in FIG. 6.

[0140] In some embodiments, the mutation of the 475th to 482nd amino acids of the porin monomer can include: independently replacing each core site with and / or inserting at each site one or more amino acids selected from the group consisting of G, A, S, T, N, Q and Y. In some embodiments, the 475th to 482nd amino acids of the porin monomer can be W475, N476, S477, D478, T479, D480, K481 and Y482, which respectively correspond to W68, N69, S70, D71, T72, D73, K74 and Y75 of the 68th to 75th positions of SEQ ID NO: 10, and then the mutation form for these core sites can be one or more of the following:

[0141] W475 (W68) is substituted with and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y, or a combination thereof;

[0142] N476 (N69) is substituted with and / or inserted with any amino acid selected from the group consisting of G, A, S, T, Q and Y, or a combination thereof;

[0143] S477 (S70) is substituted with and / or inserted with any amino acid selected from the group consisting of G, A, T, N, Q and Y, or a combination thereof;

[0144] D478 (D71) is substituted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y or a combination thereof;

[0145] T72 (T479) is substituted with any amino acid selected from the group consisting of G, A, S, N, Q and Y and / or inserted with any amino acid selected from the group consisting of G, A, S, N, Q and Y or a combination thereof;

[0146] D480 (D73) is substituted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y or a combination thereof;

[0147] K481 (K74) is substituted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N, Q and Y or a combination thereof; and

[0148] Y482 (Y75) is substituted with any amino acid selected from the group consisting of G, A, S, T, N and Q and / or inserted with any amino acid selected from the group consisting of G, A, S, T, N and Q or a combination thereof.

[0149] It can be understood that, through the above modification of the cap-sensor structure region of the pore protein monomer, the pore protein monomer containing the mutant amino acid can effectively improve the binding and capture of the cap-sensor to the negatively charged nucleic acid and improve the current signal quality, thereby improving the molecular detection and sequencing performance of the pore protein.

[0150] The pore protein according to the embodiments of the present application comprises a polymer of the pore protein monomer according to any of the above embodiments, which is optionally polymerized by 12-18 pore protein monomers, for example, polymerized by 15 monomers.

[0151] FIG. 7 is a side view of a partial structure of a pore protein according to the embodiments of the present application, in which the barrel section (about 250 aa) of the C-terminus of each pore protein monomer is shown, and the pore protein is polymerized by 15 BP62 monomers (i.e., 15-mer). FIG. 8 and FIG. 9 are a top view and a bottom view of FIG. 7, respectively.

[0152] As shown in FIGS. 7-9, from N-terminus to C-terminus, the pore-forming protein according to the embodiments of the present application can comprise: an entrance cavity, which is enclosed by the N-terminal segment of the pore-forming protein monomer, for allowing the target molecules to enter; a first sensor, which is enclosed by the first sensor structure region of the pore-forming protein monomer, for sensing the current change when the target molecules pass through the first sensor; and an exit cavity, which is enclosed by the exit structure region of the pore-forming protein monomer, for allowing the target molecules to move out of the pore-forming protein. In addition, as shown in FIGS. 7-9, the lower edge (i.e., near the C-terminus) of the exit cavity of the pore-forming protein also has two β-pleated sheets, which are slightly curved towards the center of the circle, and above the β-pleated sheets, there is a horizontal a-helix.

[0153] In some embodiments, the pore diameter of the first narrow pore of the first sensor (i.e., the central sensor) can be 0.5-2 nm. Preferably, the pore diameter of the first narrow pore of the first sensor (i.e., the central sensor) can be 0.5-1 nm. For example, the pore diameter of the first narrow pore of the first sensor (i.e., the central sensor) can be 0.5-0.8 nm. It can be understood that the middle part of the pore-forming protein monomer has a curved β-pleated sheet extending towards the C-terminus, so that when a plurality of pore-forming protein monomers are aggregated to form a pore-forming protein, these curved β-pleated sheets can enclose a narrow pore (i.e., the first narrow pore), which can serve as a sensor for large molecules such as nucleic acids and other target small molecules, and can sense different current changes when different molecules pass through the pore, so as to distinguish different pore-passing molecules and achieve high-accuracy nanopore sequencing.

[0154] In the embodiments of the present application, the pore-forming protein can also comprise a transplantable second sensor (i.e., a cap-sensor), which is enclosed by the transplantable second sensor structure region of the pore-forming protein monomer, for sensing the current change when the target molecules pass through the second sensor.

[0155] FIG. 11 is a side view of the partial structure of the pore-forming protein transplanted with the second sensor according to the embodiments of the present application, which is aggregated by 15 pore-forming protein monomers (i.e., 15-mer). FIG. 12 is a top view of FIG. 11.

[0156] As shown in FIGS. 11-12, in the pore-forming protein according to the embodiments of the present application, the second sensor structure region forms an a-helix and a β-pleated sheet extending towards the center of the pore, and based on these transplanted structures, a plurality of pore-forming protein monomers enclose a second narrow pore at a delicate angle, which (i.e., the cap-sensor) can also serve as a sensor for nucleic acid and other target molecules, and can sense different current changes when different molecules pass through the pore, so as to distinguish different pore-passing molecules and achieve high-accuracy nanopore sequencing and molecular detection.

[0157] In some embodiments, the second narrow aperture of the second sensor can have an aperture diameter of For example

[0158] The pore protein provided in the embodiments of the present application has high stability and excellent internal structure, and can sensitively characterize the molecules to be detected in the pore. In addition, the BP62 pore protein can have double sensors, thereby greatly improving the detection accuracy.

[0159] It should be noted that the pore protein provided in the embodiments of the present application is based on a pore protein monomer, which can have the pore protein monomer structure or a mutant form thereof described in any of the above embodiments and achieve the same effect, which will not be repeated here.

[0160] The embodiments of the present application also provide a polynucleotide encoding the pore protein monomer in any of the above embodiments, an expression vector containing the polynucleotide (the expression vector can also contain a controllable element), a cell containing the polynucleotide or the expression vector or expressing the pore protein monomer in any of the above embodiments. It can be understood that the polynucleotide encoding the pore protein monomer or its complementary sequence provided in the embodiments of the present application can be inserted into a eukaryotic expression vector or a prokaryotic expression vector for expression, such as an expression vector with T7 as a promoter, such as PET.28a(+), PET.21a(+), PET.32a(+) and the like; and eukaryotic or prokaryotic cells can be used to express the pore protein monomer, such as Escherichia coli, BL21(DE3), BL21 Star(DE3)pLyss, Rossata(DE3), ArcticExpress and the like, and the present application does not limit the type of expression vector and cell.

[0161] The embodiments of the present application also provide a nanopore sensor, which comprises the pore protein and the membrane layer as described in any of the above embodiments, wherein the pore protein is inserted into the membrane layer to form a pore, so as to form a nanopore sensor, and when a voltage is applied across the membrane layer, the pore generates an electric current.

[0162] In the embodiments of the present application, the membrane layer can be selected in various ways. In some embodiments, the membrane layer can include a lipid layer or an artificial polymer membrane, wherein the lipid layer can include, for example, an amphiphilic lipid. In some embodiments, the amphiphilic lipid can be selected from one or more of the following: phospholipids, fatty acids, fatty acyl groups, glycerol lipids, glycerophospholipids, sphingolipids, sterol lipids (such as cholesterol, cholesterol hemisuccinate or derivatives thereof), prenol lipids, glycolipids, polyketides and amphiphilic block copolymers. In some embodiments, the amphiphilic lipid includes a phospholipid bilayer.

[0163] In some embodiments, the lipid layer can comprise a planar membrane layer or a liposome, where the liposome can be, for example, a multilamellar liposome or a unilamellar liposome. In some embodiments, the phospholipid can be selected from the group consisting of phosphatidylethanolamine (PE), phosphatidylinositol (PI), phosphatidylserine (PS), phosphatidylglycerol (PG), and phosphatidylcholine (PC). In some embodiments, the phospholipid is phosphatidylcholine (PC), and the lipid layer can comprise a phospholipid bilayer consisting of diphytanoyl-sn-glycero-3-phosphocholine (DPhPC). In other embodiments, the lipid layer can comprise dioleoylphosphatidylcholine (DOPC), dipalmitoylphosphatidylcholine (DPPC), 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), dioleoyl-phosphatidylethanolamine (DOPEA), 1,2-distearoyl-sn-glycero-3-phosphoethanolamine (DSPE), palmitoyloleoylphosphatidylcholine (POPC), and the like. It is to be understood that the membrane layer can be of any suitable configuration that provides a suitable detection environment for the pore protein and enables molecular detection by the pore protein, and the present application is not limited to any particular configuration of the membrane layer.

[0164] In the embodiments of the present application, when a voltage is applied across the membrane layer, the molecule to be detected passes through the pore in the nanopore sensor and is displaced, and the pore generates a changing current. It is to be understood that the composition of the molecule to be detected can be determined from the change in the current signal caused by the capture of the molecule to be detected. In the embodiments of the present application, the molecule to be detected can be a metal ion, an inorganic salt, a polymer, an amino acid, a peptide, a protein, a nucleotide, a polynucleotide, a polysaccharide, a lipid, a dye, a bleach, and a drug. In some embodiments, the molecule to be detected can be DNA, RNA, and / or a polypeptide, and analogs / derivatives thereof. For example, the molecule to be detected can be modified or unmodified DNA, RNA, and / or a polypeptide, where the modification can occur, for example, at a base, and the DNA and / or RNA to be detected can include any one or more of the following modified bases: 2-thiouracil, 4-thiouracil, 5-methylcytosine, 5-methyluracil, 5-methoxyuracil, 6-methyladenine, 7-methylguanine, and pseudouracil. It is to be understood that the internal structure of the pore protein according to the embodiments of the present application is excellent, and can sensitively perceive various molecules to be detected that pass through the pore, thereby enabling high-accuracy determination of the composition of the molecule to be detected.

[0165] The embodiments of the present application also provide a nanopore sequencing device, which includes the nanopore sensor according to any one of the above embodiments.

[0166] In the embodiments of the present application, the nanopore sequencing device can specifically include: an electrolytic cell containing a sequencing buffer; a nanopore sensor located in the center of the electrolytic cell and dividing the electrolytic cell and the sequencing buffer into a positive electrolyte area and a negative electrolyte area; and a first electrode and a second electrode arranged in the positive electrolyte area and the negative electrolyte area respectively, and the first electrode and the second electrode are connected to a signal processing chip, preferably, the first electrode and the second electrode include metal or composite electrode materials; preferably, the first electrode and the second electrode are different, being silver and silver chloride respectively; or the first electrode and the second electrode are the same, including gold, platinum, graphene or titanium nitride. It can be understood that the nanopore sequencing device proposed in the embodiments of the present application can include various reagents and elements suitable for cooperating with the nanopore sensor to achieve nanopore sequencing, and the present application does not intend to limit the specific composition of the sequencing device.

[0167] The embodiments of the present application also propose a sequencing method, which includes: determining the composition of the test molecule (for example, the sequence of the test molecule) by detecting and analyzing the electrical signal generated when the test molecule passes through the pore of the pore protein, using the pore protein as described in any of the above embodiments, the nanopore sensor as described in any of the above embodiments, or the nanopore sequencing device as described in any of the above embodiments.

[0168] In the embodiments of the present application, the electrical signal can be an electric current.

[0169] In the embodiments of the present application, based on the test molecule being a nucleic acid, the sequencing method includes: i. contacting the nucleic acid with the pore protein and the nucleic acid binding protein as described in any of the above embodiments, so that the nucleic acid binding protein controls the moving speed of the nucleic acid passing through the pore of the pore protein, wherein the nucleic acid binding protein is selected from any one or more of a nuclease, a polymerase, a topoisomerase, a ligase, a helicase or a single-strand binding protein; and ii. applying a voltage across the pore, measuring the electrical signal when the nucleic acid moves through the pore when the nucleic acid moves through the pore, wherein the electrical signals generated by different types of nucleotides passing through the pore are different, so as to determine the sequence information of the nucleic acid based on the electrical signal.

[0170] The embodiments of the present application also propose a kit, which includes the pore protein monomer, the pore protein, the polynucleotide, the vector, the cell and / or the nanopore sensor as described in any of the above embodiments, for example, a kit for nanopore sequencing. It can be understood that the kit proposed in the embodiments of the present application can also include other reagents for nanopore sequencing and instructions for use, etc., and the present application does not limit other reagents for nanopore sequencing.

[0171] The application further provides application of the pore protein monomer, the pore protein, the polynucleotide, the vector, the cell, the nanopore sensor or the kit in biological small molecule detection, nucleic acid sequencing or polypeptide sequencing.

[0172] It should be noted that the explanations of the pore protein monomer and the pore protein in the embodiments of the application also apply to the polynucleotide, the expression vector, the cell, the nanopore sensor, the kit and the application thereof in molecular detection provided in the embodiments of the application, which will not be repeated here.

[0173] The experimental methods in the following examples are all conventional methods, and are performed according to the techniques or conditions described in the literature in the art or according to the product instructions, unless otherwise specified. The materials, reagents and the like used in the following examples can be obtained from commercial channels, unless otherwise specified.

[0174] Unless otherwise specified, the quantitative experiments in the following examples are all set up in triplicate, and the results are averaged.

[0175] Example 1: Acquisition and structure display of pore protein

[0176] Based on the metagenomic analysis of deep-sea samples, the inventors discovered a new type of pore protein that can be used for nanopore sequencing, which is named BP62, and the sequence of the monomer thereof is shown as SEQ ID NO: 15. The structure of the BP62 monomer was predicted using Alphafold, and the results are shown in FIG. 1. As shown in FIG. 1, from the N-terminus to the C-terminus, the structure of the BP62 monomer is four N-terminal domains N0, N1, N2 and N3, followed by a barrel segment that can form a pore structure by polymerization, and the most C-terminal is a S segment composed of a long a helix and a long loop.

[0177] About 252 amino acids of the C-terminal barrel segment of the BP62 monomer were intercepted, and the polymer structure (12-18mer) was predicted by means of AlphaFold 2 software, and the predicted pore structure is shown in FIGS. 7-9. In FIG. 7, the display direction of the BP62 pore protein is N-terminus up and C-terminus down, and only the barrel segment is shown, without the N-terminal segment and the S segment.

[0178] Overall, the BP62 pore protein presents a barrel structure, and there is a natural narrow hole (i.e., the first narrow hole) inside. FIG. 10 shows that the two monomers inside the polymer are opposite, and in the middle part of the monomer, two β folds with curvature extend towards the center, and these β folds form a narrow hole around a circle in the polymer, with a pore diameter of about 0.5 nm. which can act as a receptor for macromolecules such as nucleic acids and polypeptides, and is therefore named central sensor, also known as the first sensor.

[0179] As can be seen from FIGS. 7-9, the central sensor separates the barrel structure of the BP62 pore protein, thereby forming two inner cavities (i.e., an entrance inner cavity and an exit inner cavity) above and below. As can be seen from the top view of FIG. 8, the entrance inner cavity is mainly surrounded by β-strands, and as can be seen from the bottom view of FIG. 9, the exit inner cavity is also mainly surrounded by β-strands. In addition, there are two β-strands at the lower edge (i.e., near the C-terminus) of the exit inner cavity, which slightly bend toward the center, and there is a horizontal a-helix above the β-strands. Therefore, in the entire structure of the BP62 pore protein shown in FIGS. 7-9, in addition to the N0-N3 domains and the S segment shown in FIG. 1 above the entrance inner cavity, there are, in order from top to bottom, the entrance inner cavity, the central sensor, and the exit inner cavity.

[0180] Example 2: Design of Mutants of Pore Protein

[0181] This example is based on the analysis of the BP62 pore protein based on the structure predicted by Alphafold in Example 1, and different mutants are designed based on SEQ ID NO: 15 to optimize the sequencing performance of the pore protein.

[0182] 2.1 Optimization of Central Sensor Region

[0183] The composition of the amino acids in the core region of the central sensor region plays a decisive role in the generation of current signals. This example analyzes the core sites of the central sensor region and proposes an optimization scheme for the core sites of the central sensor region.

[0184] Specifically, some of the core sites of the central sensor region can be position 533, position 546, and positions 537-542 shown in FIG. 2, and the corresponding amino acids are R533, D546, T537, D538, N539, E540, T541, and Y542. For these sites, polar uncharged amino acids are introduced into them by substitution, deletion, insertion, and other mutation forms, so that they are more sensitive to the sensing of negatively charged nucleic acid molecules. The mutation forms (by substitution, deletion, and / or insertion) that can be introduced into each site are independently selected from the following:

[0185] 2.2 Optimization of Entrance Structure Region

[0186] The charge properties of the amino acids in the entrance structure region are extremely important for the capture of the molecules to be tested. This example takes the truncated N-terminal segment as an example, and optimizes the partial sites of the entrance end of the N3 domain as the entrance end of the entrance structure region.

[0187] Specifically, FIG. 3 shows the partial modification sites of the entry end, i.e. K330, K335, K353, K363 and K378 with positive charges and E340, E342, E343, D360, E364, D371, D377, D388, D391, E396, E401 and D403 with negative charges. For these sites, polar uncharged amino acids or polar positively charged amino acids are introduced by substitution, deletion, insertion and the like mutation forms to improve the capture efficiency of the entry end to the nucleic acid molecule with negative charges, thereby enhancing the binding of the pore protein to the test molecule, so as to improve the sequencing speed and sequencing quality. The mutation forms (by substitution, deletion and / or insertion) that can be introduced into each site are independently selected from the following:

[0188] 2.3 For the transmembrane region

[0189] The transmembrane region located in the exit lumen is the region where the pore protein is embedded in the membrane pore, and the amino acids in this region will affect the pore embedding efficiency of the pore protein. This embodiment analyzes the important sites of the transmembrane region and proposes an optimization scheme for the important sites of the transmembrane region.

[0190] Specifically, FIG. 4 shows some amino acid sites located in the transmembrane region, especially the interface outer wall of the transmembrane region (i.e. towards the membrane direction), i.e. non-polar amino acids L628, L634, G427, I428, W430, V432, L491, A493, I495, A497, V472 and V470. Since polar amino acids and some non-polar amino acids can affect the insertion efficiency, for these sites, high hydrophobic amino acids are introduced by substitution, deletion, insertion and the like mutation forms to enhance the stability of the pore protein binding to the membrane. In addition, the polar amino acids on the interface of the transmembrane region can also be mutated to high hydrophobic amino acids to further improve the pore embedding efficiency, thereby improving the sequencing speed and sequencing quality. The mutation forms (by substitution, deletion and / or insertion) that can be introduced into each site are independently selected from the following:

[0191] It should be noted that the mutations of the above-mentioned central sensor region, entry structure region and transmembrane region will not affect the overall configuration of BP62, and thus will not affect the function of BP62. The test molecule can still pass through the central sensor region of the pore protein, and the current change caused by the perforation can still be detected, thereby playing a detection function as a nanopore sensor.

[0192] Embodiment 3: Pore protein grafted with cap gate sensor (double sensor pore protein)

[0193] Based on the metagenomic analysis of deep-sea samples, the inventors discovered several peptide segments (cap peptides) of different lengths and structures that can be transplanted into the BP62 monomer in Example 1 or 2 and modified into a new pore protein with two sensors. The modified new pore protein with two sensors is named BP62 transplants in this example.

[0194] Through analysis, the cap peptide can be transplanted into the V445-G492 position of the outlet structure region of the BP62 monomer. This example uses AlphaFold2 software to predict the structure of each BP62 transplant to determine whether each cap peptide segment fits the corresponding region of the BP62 monomer.

[0195] Through prediction, this example screened a total of 8 transplantable cap peptides, the amino acid sequences of which are shown as SEQ ID NO: 2-9, and the specific structures are shown as the right side of Figure 5 (cap1-cap8). These cap peptides have different shapes but have substantially the same secondary structure, and the overall structure is an elongated strip shape. These cap peptides are inserted or replaced in a horizontal direction on the a-helix near the structural region of the BP62 monomer (as shown in the left side of Figure 5), and at a suitable angle they can form another narrow pore (i.e., a second narrow pore) in the BP62 barrel, thereby playing a similar role to the central sensor. The resulting BP62 transplant has two sensors, namely the central sensor and the cap sensor.

[0196] The sequence of the cap peptide segment is inserted and replaced in part of the sequence in BP62, and the sequence of the formed pore protein transplant is analyzed.

[0197] Table 1 shows the sequences of the cap domains in the 8 BP62 transplants. The sequences of these BP62 transplants are predicted by AlphaFold 2 software, and are set to 12-18mers. As shown in the structure schematic diagrams of Figures 11 and 12, the insertion of the Cap peptide segment in the BP62-9 transplant forms a cap sensor, and it does not substantially affect other domains of BP62-9.

[0198] Table 1

[0199] The BP62-9 transplant in Table 1 is selected as an example for illustration. Its monomer structure is shown in Figure 13. Compared with the BP62 pore protein with only a central sensor, the introduced cap peptide segment and its adjacent amino acids together form an a-helix and a beta sheet extending towards the center of the pore. These newly formed protein secondary structures form a second narrow part of the pore in the lumen of the pore protein, and its diameter is about (in the figure ), which has the potential to be a sensor, and thus the newly formed second constriction was named as cap sensor. In this way, BP62-9 has two sensors, one is the native center sensor and the other is the transplanted cap sensor.

[0200] Example 4: Mutant design of BP62 transplanted body with cap sensor

[0201] This example analyzes the structure of BP62-9 transplanted body in Example 3, and designs different mutant transplanted bodies based on it to optimize the sequencing performance of the pore protein. The sequence of BP62-9 transplanted body is shown in SEQ ID NO: 16.

[0202] As can be seen from Figures 11-13, the entry end, entry lumen, exit lumen and transmembrane region and transmembrane region interface of the BP62-9 transplanted body are basically the same as those of the BP62 monomer with only the center sensor, which shows that the mutation design scheme for different regions of the BP62 monomer described in Example 2 is also applicable to the BP62 transplanted body.

[0203] In addition, different forms of transplanted mutants are designed for the cap sensor region of the BP62 transplanted body to optimize the sequencing performance of the BP62 transplanted body. As can be seen from Figure 6, the newly formed cap sensor is supported by two alpha helices supporting two beta folds. This example screens some key sites of the cap sensor, i.e. positions 475-482, whose specific amino acid types in SEQ ID NO: 16 are W475, N476, S477, D478, T479, D480, K481 and Y482. Similar to the design of the center sensor region, polar uncharged amino acids can be introduced to these sites by substitution, deletion, insertion and other mutation forms, so as to make them more sensitive to the sensing of negatively charged nucleic acid molecules. The mutation forms (by substitution, deletion and / or insertion) that can be introduced to each site are independently selected from the following:

[0204] Example 5: Preparation of BP62 pore protein and its mutants

[0205] 5.1 Cloning and expression of BP62 protein

[0206] A thrombin cleavage site was added after K328 of the BP62 monomer of the amino acid sequence shown in SEQ ID NO: 1, and the polynucleotide sequence corresponding to the monomer was artificially synthesized. The polynucleotide sequence was ligated into the PET.21a(+) plasmid, respectively, and the double enzyme cleavage sites used were Ndel and Xhol, and the obtained positive plasmid was labeled as PET.21a(+)-BP62. The BP62 protein expressed by the plasmid has a thrombin cleavage site in the middle and a Strep II tag at the C-terminus.

[0207] The PET.21a(+)-BP62 plasmid was transformed into E. coli expression bacteria BL21(DE3) or its derivative bacteria. A single colony was picked and inoculated into 20 mL of LB medium containing ampicillin, and cultured at 37°C overnight with shaking. Then, it was inoculated into 2 L of LB medium containing kanamycin, and cultured at 37°C with shaking until the OD 600 = 0.6-0.8, the temperature was lowered to 16°C, and IPTG was added to a final concentration of 500 μM to induce expression overnight, obtaining bacteria with expressed BP62 protein.

[0208] 5.2 Extraction and purification of BP62 protein

[0209] 5.2.1 Buffer preparation

[0210] Buffer A: 20 mM Tris-HCl, 250 mM NaCl, 1% DDM, pH 8.0.

[0211] Buffer B: 20 mM Tris-HCl, 250 mM NaCl, 0.02% DDM, pH 8.0.

[0212] Buffer C: 20 mM Tris-HCl, 250 mM NaCl, 0.02% DDM, 5 mM desthiobiotin, pH 8.0.

[0213] 5.2.2 Protein extraction and purification

[0214] Resuspend the bacteria at a ratio of 1 g of bacteria to 10 mL of Buffer A, and sonicate the cells until the bacterial solution is clear. Then, rotate on a rotator at 4°C overnight. The next day, centrifuge at 18,000 rpm at 4°C for 1 h, and take the supernatant.

[0215] Strep-Tactin beads (IBA Lifesciences) column was equilibrated with Buffer A for 5 column volumes (CV) using AKTA pure chromatograph, then loaded at 2 mL / min. After loading, the column was washed with Buffer B for 20 CV, and the target protein was eluted with Buffer C, and collected.

[0216] The obtained protein was concentrated to 0.5 mL, and then an appropriate amount of thrombin was added for enzyme digestion at 4°C overnight. Finally, the enzyme digestion solution was passed through a Superdex 6 increase 10 / 300 GL (Cytiva) column equilibrated with buffer B, and the target protein was collected and then stored at -80°C.

[0217] The purified target protein was subjected to SDS-PAGE electrophoresis, and the results are shown in Figure 14. Lanes 2 and 3 are electrophoretic bands of the purified sample of BP62 protein (monomer automatically polymerized into 15-mer in a natural state). It is shown that the BP62 protein (polymer) is greater than 1000 KDa, which is consistent with the expected size.

[0218] 5.3 Preparation and purification of BP62 mutant protein The preparation and purification of BP62 mutant protein are the same as steps 5.1-5.2 above.

[0219] Based on the BP62 monomer sequence shown as SEQ ID NO: 15, partial mutant proteins were prepared, which were named BP62-11 to BP62-17, respectively, and the specific mutation information is as follows:

[0220] wherein the sequence of Strep tag II is Trp-Ser-His-Pro-Gln-Phe-Glu-Lys (WSHPQFEK).

[0221] Figure 15a shows the electrophoretogram of the purified partial mutant proteins. As shown in Figure 15a, the prepared mutant proteins have obvious bands, and the band size indicates that they exist in the form of polymers.

[0222] Example 6: Preparation of BP62 graft and its mutant

[0223] The preparation of BP62 transploms and their mutants is the same as the above steps 5.1-5.2. Figure 15b shows the electrophoretic diagram of BP62-2, BP62-3, BP62-4, BP62-5, BP62-7, BP62-8, BP62-9 transploms, wherein wt is a BP62 protein containing only the central sensor. As can be seen from Figure 15b, each BP62 transplom can express (BP62-4 and BP62-7 express more, and others express less), and compared with wt, each BP62 transplom protein has a larger molecular weight, indicating that the cap sensor transplom is successful.

[0224] Example 7: Construction of nanopore biosensor using BP62 pore protein and its transploms

[0225] The current signal is collected using a patch-clamp amplifier or other electrical signal amplifier. A single-channel nanopore detection system based on patch-clamp and signal amplifier is built according to the method disclosed in the literature (Ji Z, Guo P. Channel from bacterial virus T7 DNA packaging motor for the differentiation of peptides composed of a mixture of acidic and basic amino acids. Biomaterials. 2019 May 21; 214: 119222). The Ag / AgCl electrode is immersed in the sequencing buffer and the electrode is located in the cis and trans regions of the electrolytic cell, respectively. The BP62 pore protein and each transplom prepared in Examples 5 and 6 are diluted by a certain multiple using 1x PBS buffer, and under the action of an applied electric field, a single nanopore protein BP62 is inserted into a phospholipid bilayer composed of 1,2-diphytanoyl-sn-glycero-3-phosphocholine (DPhPC) to form a nanopore biosensor. The specific dilution multiple is determined by whether the nanopore protein is embedded in the membrane (i.e., embedded in the pore). Generally, a protein concentration of 0.1 mg / ml is used, and PBS is diluted by 100 times or 50 times or other times for trial. If a certain dilution concentration fails to embed the pore, the dilution multiple needs to be reduced for further trial until the nanopore protein is successfully embedded in the membrane layer. An external voltage is applied to obtain the current amplitude value of a single pore protein.

[0226] Figures 16a-e show the biosensing currents generated by the nanopore of the nanopore protein BP62 and its different transplants when a voltage of 0.02 V, 0.04 V, 0.10 V, 0.14 V and 0.18 V is applied. As can be seen from Figure 16, BP62 (Figure 16a) and its transplants (Figures 16b-e) can stably output currents at different voltages as a nanopore protein, indicating that the new nanopore protein proposed in the embodiments of the present application can be effectively used for nanopore-based molecular detection.

[0227] Example 8: Constructing a nanopore biosensor using a mutant of the nanopore protein BP62

[0228] According to the method in Example 7, the mutant of the nanopore protein BP62 prepared in Example 5 is diluted by a certain number of times, and then a single nanopore protein BP62 mutant is inserted into a phospholipid bilayer composed of DPhPC under the action of an applied electric field to form a nanopore biosensor, and an applied voltage is applied to obtain the current amplitude value of a single pore protein.

[0229] Figures 16f-j show the biosensing currents generated by the nanopore of the different mutants (BP62-11, BP62-12, BP62-13, BP62-14, BP62-17 in turn) of the nanopore protein BP62 when a voltage of 0.02 V, 0.04 V, 0.10 V, 0.14 V and 0.18 V is applied. As can be seen from Figures 16f-g, BP62-11 and BP62-12 can stably output currents at different voltages as a nanopore protein, indicating that partial truncation (for example, truncation of the N-terminal region) and / or addition of a sequence at the end of the sequence (such as the His-tag and Strep tag used in Example 5) to BP62 based on the full-length sequence of BP62 shown in SEQ ID NO: 15 will not affect the activity of BP62, and these mutants can still function as a nanopore protein.

[0230] In addition, as can be seen from Figures 16h-j, the mutants such as BP62-13 obtained by mutating part of the sites of the barrel segment of BP62 also have the activity of BP62, indicating that the mutants of BP62 can also function as a nanopore protein and can also be used for nanopore sequencing.

[0231] Example 9: Sequencing using the nanopore protein BP62

[0232] The sequencing library containing DC sequence (SEQ ID NO: 17) and single-stranded DNA with cholesterol (5'-cholesterol-TTGACCGCTCGCCTC-3', SEQ ID NO: 18, cholesterol is connected to the 5' end of the DNA, which can bind to the phospholipid bilayer, help the nanopore capture sequencing library, and reduce the loading amount of sequencing library) were mixed with sequencing buffer and added to the nanopore biosensor of the above examples; after applying an external voltage of 0.14V or 0.18V, it was observed that the DNA was captured by the nanopore, and a characteristic current amplitude value was generated. And as the DNA moves through the nanopore, the current amplitude value changes. Different DNA sequences produce different current amplitude values.

[0233] FIGS. 17 and 18 are overall current changes and local detail diagrams of the library DNA passing through the nanopore protein BP62 under the action of an external voltage of 0.14V, respectively. As can be seen from FIGS. 17-18, BP62 can clearly output sequencing signals, indicating that the open pore current of the nanopore protein BP62 is about 150pA under an external voltage of 0.14V, and the sequencing amplitude is about 40pA. The current signal is stable and clear, indicating that the BP62 pore protein proposed in the embodiments of the application can be effectively used for nanopore sequencing of DNA molecules.

[0234] Example 10: Sequencing using BP62 pore protein mutants

[0235] According to the method of Example 9, the mutants of the nanopore protein BP62 are used for nanopore sequencing.

[0236] FIGS. 19a-b show the overall current changes and local detail diagrams of the library DNA passing through the nanopore protein mutant BP62-12 under the action of an external voltage of 0.14V. As can be seen from FIGS. 19a-b, BP62-12 as a pore protein not only can output current at different voltages, but also can stably sequence, which indicates that BP62 as a nanopore protein does not need the participation of N1, N2, N3 and other domains, and the truncation of other parts except the barrel segment and / or the addition of sequence will not affect the functionality of BP62 as a nanopore protein, suggesting that BP62 is functionally stable and has great potential for modification, and its mutants can also be effectively used for nanopore sequencing of DNA molecules.

[0237] FIGS. 20a-b and 21a-b show the overall current changes and local detail diagrams of the library DNA passing through the nanopore protein mutants BP62-13 and BP62-16, respectively, under the action of an external voltage of 0.14V. Similarly, these mutants with core site mutations also exhibit stable current responses, indicating that the BP62 mutants can also be used for nanopore sequencing of DNA molecules.

[0238] Example 11: Verification of the pore-forming ability of BP62 and its transplants by cryo-EM and resolution of the structure of BP62

[0239] The gel filtration and chromatographic protein purification (molecular sieve) of the BP62 pore protein prepared in Example 5 were carried out, and the results are shown in Figure 22a, and the corresponding electrophoretogram of the BP62 pore protein is shown in Figure 22b. The obtained BP62 pore protein was subjected to preliminary negative staining and photographed under an electron microscope, and the collected images are shown in Figure 22c. As can be seen from Figure 22c, BP62 can form stable and uniform pore proteins, indicating that BP62 has good potential as a pore protein.

[0240] Further, high-resolution electron density maps were obtained by cryo-EM, as shown in Figure 22d, in which the right side numerical bar represents the resolution, in angstroms. As can be seen from Figure 22d, the resolution of the structure of BP62 presented by cryo-EM is high, and the structure is clear and realistic. Further, an atomic structure model of BP62 was built, as shown in Figure 22e (top view) and Figure 22f (front view), and BP62 presents a very regular circular barrel structure, and its interior presents a circular small hole that can sense the passage of a test molecule, which is an excellent pore protein.

[0241] In addition, the transplants BP62-7 and BP62-4 of BP62 prepared in Example 6 were also subjected to negative staining, and the collected images are shown in Figures 23 and 24, respectively. As can be seen from Figures 23 and 24, the BP62-7 and BP62-4 transplanted with a second sensor can still form pore proteins, suggesting that the transplants of BP62 also have good potential as double-sensor pore proteins, and compared with BP62 with a single sensor, they will have greater advantages in distinguishing repetitive nucleotide sequences or amino acid sequences.

[0242] In the description of the present specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples. In addition, the skilled person in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples, without contradiction.

[0243] Although the embodiments of the present application have been shown and described above, it is understood that the above-described embodiments are exemplary and are not to be construed as limiting the present application, and that changes, modifications, substitutions and variations can be made by those skilled in the art without departing from the scope of the present application.

Claims

1. A pore protein monomer, the pore protein monomer comprising a cylindrical segment, the cylindrical segment being: a. Has the amino acid sequence shown in SEQ ID NO: 1; b. Compared with the amino acid sequence shown in SEQ ID NO: 1, it has an amino acid sequence with one or more amino acid substitutions, deletions, and / or additions, and the cylindrical segment has the function of forming a channel structure through polymerization; or c. An amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the amino acid sequence shown in SEQ ID NO: 1, and the cylindrical segment having the function of forming a channel structure through polymerization.

2. The pore protein monomer according to claim 1, wherein the tube segment comprises a first receptor structural region, a transmembrane region, and an outlet structural region. in, The first receptor structure region is used to form the first narrow orifice in the channel structure; Preferably, the first receptor structure region includes a β-fold with an arc.

3. The pore protein monomer according to claim 2, wherein the substitution and / or addition of the amino acid occurs in the first receptor structural region, the substitution of the amino acid includes replacing the amino acid in the first receptor structural region with a polar uncharged amino acid and / or a nonpolar amino acid, and the addition of the amino acid includes inserting one or more polar uncharged amino acids and / or nonpolar amino acids into the first receptor structural region. Preferably, the polar uncharged amino acid and the nonpolar amino acid are selected from: G, A, S, T, N, Q and Y; Preferably, the substitution and / or addition of the amino acid occurs at positions 126 to 140, preferably positions 130 to 135, and most preferably positions 126, 139, and 130 to 135 of the amino acid sequence shown in SEQ ID NO:

1.

4. The pore protein monomer according to claim 3, wherein the amino acids at positions 126, 139, and 130 to 135 are R126, D139, T130, D131, N132, E133, T134, and Y135, and the substitution and / or addition of the amino acids are selected from one or more of the following: R126 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, T, N, Q and Y, preferably A; D546 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, T, N, Q and Y, preferably Q; T130 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, N, Q and Y, preferably G; D131 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, T, N, Q and Y, preferably selected from G, T and Q; N132 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, T, Q and Y, preferably selected from A and Y; E133 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, T, N, Q and Y, preferably selected from G and Q; T134 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, N, Q, and Y; and Y135 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, T, N, and Q, preferably A. Optionally, the first receptor structural region has the following abrupt change: R126A, D139Q, D131T, and E133Q; R126A, D139Q, D131Q, and N132Y; R126A, D139Q, D131Q, E133G, Y135A, and T130G; or D131G and N132A.

5. The pore protein monomer according to any one of claims 2 to 4, wherein the substitution and / or addition of the amino acid occurs in the transmembrane region, and the substitution of the amino acid comprises replacing the amino acid in the transmembrane region with a highly hydrophobic amino acid; Preferably, the highly hydrophobic amino acid is selected from: A, G, V, L, I, Y, F, and W; Preferably, the substitution of the amino acid occurs at any one or more of positions 221, 227, 20, 21, 23, 25, 84, 86, 88, 90, 65 and 63 of the amino acid sequence shown in SEQ ID NO: 1; Preferably, the amino acids at positions 221, 227, 20, 21, 23, 25, 84, 86, 88, 90, 65, and 63 are L221, L227, G20, I21, W23, V25, L84, A86, I88, A90, V65, and V63, and the substitutions of the amino acids are selected from one or more of the following: L221 is replaced by any amino acid selected from the following: G, A, V, I, Y, F, or W; L227 is replaced by any amino acid selected from the following: G, A, V, I, Y, F, or W; G20 is replaced by any amino acid selected from the following: A, V, L, I, Y, F, or W; I21 is replaced by any amino acid selected from the following: A, G, V, L, Y, F, or W; W23 is replaced by any amino acid selected from the following: A, G, V, L, I, Y, or F; V25 is replaced by any amino acid selected from the following: A, G, L, I, Y, F, or W; L84 is replaced by any amino acid selected from the following: G, A, V, I, Y, F, or W; A86 is replaced by any amino acid selected from the following: G, V, L, I, Y, F, or W; I88 is replaced by any amino acid selected from the following: A, G, V, L, Y, F, or W; A90 is replaced by any amino acid selected from the following: G, V, L, I, Y, F, or W; V65 is replaced by any amino acid selected from the following: A, G, L, I, Y, F, or W; and V63 is replaced by any amino acid selected from the following: A, G, L, I, Y, F, or W.

6. The pore protein monomer according to any one of claims 1 to 5, further comprising: A transplantable second receptor structure region, which is transplantable to the outlet structure region, wherein the second receptor structure region is used to form a second narrow orifice in the channel structure.

7. The pore protein monomer according to claim 6, wherein the outlet structural region contains an α-helix, the second receptor structural region is a polypeptide containing a double α-helix structure, and the second receptor structural region replaces or inserts into the α-helix of the outlet structural region; Preferably, the second receptor structural region is a polypeptide containing a double α-helix structure and a β-sheet structure.

8. The pore protein monomer according to claim 6 or 7, wherein the secondary structure of the second receptor structural region is an extended elongated strip; preferably, the second receptor structural region has any amino acid sequence as shown in SEQ ID NO: 2-9. Preferably, the second receptor structural region replaces or inserts into the V38-G85 peptide segment or any peptide segment between them in the amino acid sequence shown in SEQ ID NO: 1, and more preferably into the F48-V65 peptide segment or any peptide segment between them. Preferably, the pore protein monomer transplanted with the second receptor structural region has the amino acid sequence shown in SEQ ID NO:

10.

9. The pore protein monomer according to any one of claims 6 to 8, wherein the second receptor structural region further comprises substitution, deletion and / or addition of one or more amino acids, the substitution and / or addition of the amino acids comprising replacing the amino acids in the second receptor structural region with polar uncharged amino acids and / or nonpolar amino acids, the addition of the amino acids comprising inserting one or more polar uncharged amino acids and / or nonpolar amino acids into the second receptor structural region. Preferably, the polar uncharged amino acids and nonpolar amino acids are selected from: G, A, S, T, N, Q, and Y. Preferably, the substitution and / or addition of the amino acid occurs at any one or more positions from position 68 to position 75 of the amino acid sequence shown in SEQ ID NO:

10.

10. The pore protein monomer according to claim 9, wherein the amino acids at positions 68 to 75 are W68, N69, S70, D71, T72, D73, K74, and Y75, and the substitution and / or addition of the amino acids are selected from one or more of the following: W68 is substituted with and / or inserted with any of the following amino acids or combinations thereof: G, A, S, T, N, Q and Y; N69 is replaced by any amino acid selected from the following and / or inserted with any amino acid selected from the following or a combination thereof: G, A, S, T, Q and Y; S70 is substituted with and / or inserted with any amino acid selected from the following: G, A, T, N, Q, and Y; D71 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, T, N, Q, and Y; T72 is replaced by any amino acid selected from the following and / or inserted with any amino acid selected from the following or a combination thereof: G, A, S, N, Q and Y; D73 is replaced by any amino acid selected from the following and / or inserted with any amino acid selected from the following or a combination thereof: G, A, S, T, N, Q and Y; K74 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, T, N, Q, and Y; and Y75 is substituted with and / or inserted with any amino acid selected from the following: G, A, S, T, N, and Q.

11. The pore protein monomer according to any one of claims 1 to 10, wherein the pore protein monomer further comprises an N-terminal segment, wherein the pore protein monomer comprises the N-terminal segment and the barrel segment sequentially from the N-terminus to the C-terminus, wherein the N-terminal segment comprises one or more domains selected from N0, N1, N2 and N3 domains, preferably comprising an N3 domain; Optionally, the sequence of the N3 structural domain is as shown in SEQ ID NO: 11; Optionally, the sequence of the N-terminal segment is as shown in SEQ ID NO:

12.

12. The pore protein monomer according to claim 11, wherein the N-terminal segment contains an inlet structural region, the inlet structural region comprising the substitution, addition and / or deletion of amino acids, wherein the substitution of amino acids comprises replacing the amino acids in the inlet structural region with polar positively charged amino acids, polar uncharged amino acids or nonpolar amino acids selected from A or G; Preferably, the polar positively charged amino acid is selected from: K, R, and H; Preferably, the polar, uncharged amino acid is selected from N, S, T, Y, C, and Q.

13. The pore protein monomer according to claim 12, wherein the amino acid substitution occurs at any one or more of positions 2, 7, 25, 35, 50, 12, 14, 15, 32, 36, 43, 49, 60, 63, 68, 73, and 75 of the amino acid sequence shown in SEQ ID NO:

11. Optionally, the substitution of the amino acid is selected from one or more of the following: K2 is replaced by any amino acid selected from the following: N, A, G, S, T, Q; K7 is replaced by any amino acid selected from the following: N, A, G, S, T, Q; K25 is replaced by any amino acid selected from the following: N, A, G, S, T, Q; K35 is replaced by any amino acid selected from the following: N, A, G, S, T, Q; K50 is replaced by any amino acid selected from the following: N, A, G, S, T, Q; E12 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; E14 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; E15 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; D32 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; E36 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; D43 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; D49 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; D60 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; D63 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; E68 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; E73 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q; D75 is replaced by any amino acid selected from the following: K, R, N, A, G, S, T, Q.

14. The pore protein monomer according to any one of claims 1 to 13, further comprising an S segment and / or a signal peptide, Optionally, the pore protein monomer comprises the tube segment and the S segment sequentially from the N-terminus to the C-terminus; Optionally, the pore protein monomer comprises, from the N-terminus to the C-terminus, the N-terminal segment, the tube segment, and the S-terminus. Optionally, the pore protein monomer comprises, from the N-terminus to the C-terminus, the signal peptide, the N-terminal segment, the tube segment, and the S-terminus. Optionally, the sequence of the S segment is as shown in SEQ ID NO: 13; Optionally, the sequence of the signal peptide is shown in SEQ ID NO:

14.

15. The pore protein monomer according to any one of claims 1 to 5, 11-14, wherein the pore protein monomer has the amino acid sequence shown in SEQ ID NO:

15.

16. The pore protein monomer according to claim 15, wherein, based on SEQ ID NO: 15, the pore protein monomer: a. Having one or more amino acid substitutions, deletions, and / or additions, and the pore protein monomer having the function of polymerizing to form a pore structure; or b. An amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity with the amino acid sequence shown in SEQ ID NO: 15, and the pore protein monomer having the function of forming a pore structure by polymerization.

17. The pore protein monomer according to claim 16, wherein the amino acid sequence of the pore protein monomer as shown in SEQ ID NO: 15 comprises, from the N-terminus to the C-terminus: The signal peptide, the sequence of which is shown in SEQ ID NO: 14; The N-terminal segment, the sequence of which is shown in SEQ ID NO: 12; The cylindrical segment, the sequence of which is shown in SEQ ID NO: 1; and The S segment, the sequence of which is shown in SEQ ID NO: 13, The substitution, deletion, and / or addition of the amino acid occurs in the signal peptide, the N-terminal region, the tubular region, and / or the S-region. The substitution, deletion, and / or addition of the amino acid occurring in the N-terminal segment, the cylindrical segment, and / or the S-segment are defined as in any one of claims 3-5 and 11-13, respectively.

18. The pore protein monomer according to any one of claims 6 to 17, wherein the pore protein monomer has an amino acid sequence, wherein the amino acid sequence is based on SEQ ID NO: 15, and the peptide at positions 455 to 472 of the amino acid sequence of SEQ ID NO: 15 is replaced by an amino acid sequence of a second receptor structural region as shown in any one of SEQ ID NO: 2-9.

19. The pore protein monomer according to claim 18, wherein the pore protein monomer has the amino acid sequence shown in SEQ ID NO:

16.

20. The pore protein monomer according to claim 18 or 19, wherein the substitution, deletion and / or addition of said amino acid occurs in the second receptor structural region, wherein the substitution, deletion and / or addition of said amino acid occurring in the second receptor structural region is as defined in any one of claims 6-10.

21. A pore protein comprising a polymer of pore protein monomers as described in any one of claims 1 to 20, said polymer optionally being formed by polymerizing 12 to 18 of said pore protein monomers.

22. The pore protein of claim 21, comprising: The inlet lumen is formed by the N-terminal segment of the pore protein monomer; The first receptor, the first receptor being enclosed by the first receptor structural region of the pore protein monomer; and The outlet cavity is enclosed by the outlet structural region of the pore protein monomer; Preferably, the aperture of the first narrow aperture of the first receptor is [missing information]. More preferably 23. The pore protein according to claim 22, further comprising: A transplantable second receptor, wherein the second receptor is enclosed by a transplantable second receptor structural region of the pore protein monomer; Preferably, the aperture of the second narrow aperture of the second receptor is [missing information]. More preferably 24. A polynucleotide encoding a pore protein monomer as described in any one of claims 1 to 20.

25. An expression vector comprising the polynucleotide and controllable element as described in claim 24.

26. A cell comprising the polynucleotide of claim 24 or the expression vector of claim 25 or the monomer expressing the pore protein of any one of claims 1 to 20.

27. A nanopore sensor comprising a membrane layer and a pore protein as claimed in any one of claims 21 to 23, wherein the pore protein is inserted into the membrane layer to form a pore, and the pore generates a current when a voltage is applied across the membrane layer.

28. The nanopore sensor of claim 27, wherein the membrane layer comprises a lipid layer or a synthetic polymer membrane; Preferably, the lipid layer comprises amphiphiles; Preferably, the amphiphilic compounds comprise a phospholipid bilayer; Preferably, the lipid layer comprises a planar membrane or a liposome; Preferably, the liposomes comprise multilayer liposomes or monolayer liposomes; Preferably, the lipid layer comprises a phospholipid bilayer composed of diphytylphosphatidylcholine.

29. The nanopore sensor according to claim 27 or 28, wherein when a voltage is applied across the membrane layer, the analyte molecule passes through the pores in the nanopore sensor and is displaced, and the pores generate a changing current; Optionally, the analyte is selected from one or more of the following: metal ions, inorganic salts, polymers, amino acids, peptides, proteins, nucleotides, polynucleotides, polysaccharides, lipids, dyes, bleaching agents, and drugs. Preferably, the molecule to be tested includes modified or unmodified DNA, RNA, and / or polypeptides; Preferably, the DNA and / or RNA includes any one or more of the following modified bases: 2-thiouracil, 4-thiouracil, 5-methylcytosine, 5-methyluracil, 5-methoxyuracil, 6-methyladenine, 7-methylguanine, and pseudouracil.

30. A nanopore sequencing device comprising a nanopore sensor as claimed in any one of claims 27 to 29.

31. The nanopore sequencing device according to claim 30, wherein the nanopore sequencing device specifically comprises: An electrolytic cell containing sequencing buffer; A nanopore sensor, located in the center of the electrolytic cell, divides the electrolytic cell and the sequencing buffer into a positive electrolyte region and a negative electrolyte region; and A first electrode and a second electrode are respectively disposed in the positive electrolyte region and the negative electrolyte region, and the first electrode and the second electrode are connected to the signal processing chip. Preferably, the first electrode and the second electrode comprise metal or composite electrode materials; Preferably, the first electrode and the second electrode are different, being silver and silver chloride, respectively; or the first electrode and the second electrode are the same, including gold, platinum, graphene, or titanium nitride.

32. A sequencing method, wherein the sequencing method utilizes a pore protein as described in any one of claims 21 to 23, or a nanopore sensor as described in any one of claims 27 to 29, or a nanopore sequencing device as described in claim 30 or 31 to determine the composition of the analyte molecule by detecting and analyzing the electrical signal generated when the analyte molecule passes through the pores of the pore protein.

33. The sequencing method according to claim 32, wherein the molecule to be tested comprises modified or unmodified DNA, RNA or polypeptide; Preferably, the electrical signal includes current.

34. The sequencing method according to claim 32 or 33, wherein, based on the molecule to be tested being a nucleic acid, the sequencing method comprises: i. Contacting the nucleic acid with the pore protein and nucleic acid-binding protein as described in any one of claims 21 to 23, such that the nucleic acid-binding protein controls the speed at which the nucleic acid moves through the pores of the pore protein, wherein the nucleic acid-binding protein is selected from any one or more of nucleases, polymerases, topoisomerases, ligases, helicases, or single-stranded binding proteins; and ii. Applying a voltage across the pore, measuring the electrical signal as the nucleic acid moves through the pore, wherein different types of nucleotides produce different electrical signals as they pass through the pore, thereby determining the sequence information of the nucleic acid based on the electrical signal.

35. A kit comprising at least one of the following: a pore protein monomer as claimed in any one of claims 1 to 20, a pore protein as claimed in any one of claims 21 to 23, a polynucleotide as claimed in claim 24, an expression vector as claimed in claim 25, a cell as claimed in claim 26, and a nanopore sensor as claimed in any one of claims 27 to 29.

36. The use of the pore protein monomer as described in any one of claims 1 to 20, the pore protein as described in any one of claims 21 to 23, the polynucleotide as described in claim 24, the expression vector as described in claim 25, the cell as described in claim 26, the nanopore sensor as described in any one of claims 27 to 29, or the kit as described in claim 35 in the detection of biological small molecules, nucleic acid sequencing, or peptide sequencing.

Citation Information

Patent Citations

  • Mutant MspA protein monomer and expressed gene and application thereof

    CN105801676A

  • Screening method of protein nanopore amino acid sequence, protein nanopore and application thereof

    CN113470751A

  • Mutant of pore protein monomer, protein pore and application thereof

    CN113735948A

  • Double-portal pore protein, pore protein mutant, nucleotide sequence and application of double-portal pore protein and pore protein mutant

    CN115974984A

  • Novel nanopore protein mutant and application thereof

    CN117384260A