Pore protein monomer, pore protein, mutant of pore protein and application of pore protein monomer and pore protein
Patent Information
- Application Number
- CN202280102109.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2025-07-04
AI Technical Summary
Existing nanopore sequencers have deficiencies in sequencing accuracy, throughput, and chip stability, and cannot meet the ultimate needs of molecular biology research. Moreover, the stability of natural porins is poor, and there is a lack of suitable porins for single-molecule detection.
A new porin monomer BCP20 and its mutants were developed. Through gene mining and computer-aided structure prediction, porins with high stability and suitable pore structure were formed for nanopore sequencing to realize DNA, RNA and peptides. detection.
It has achieved high accuracy and high stability of single-molecule nanopore sequencing, which can effectively detect small molecules such as nucleotides, amino acids, sugars, and vitamins. It is also suitable for sequencing DNA, RNA, and polypeptides, improving the accuracy of sequencing. and efficiency.
Smart Images

Figure CN120265646A_ABST
Abstract
Description
Porin monomer, porin and its mutant and its application Technical Field
[0001] The present invention relates to the field of single-molecule sequencing, and in particular to a porin monomer, a porin and a mutant thereof, and applications thereof. Background Art
[0002] Currently, Oxford Nanopore Technologies in the UK has commercialized a series of nanopore sequencers, including the MinION, GridION, and PromethION, as well as the QNome-3841 nanopore gene sequencer. However, these instruments still have significant shortcomings in sequencing accuracy, throughput, chip stability, and applicable scenarios, failing to meet the ultimate requirements of molecular biology research. Therefore, there is an urgent need to develop a highly accurate, highly integrated, and highly stable single-molecule sequencer. A nanopore-based single-molecule sequencer is a highly integrated detection system that integrates multiple disciplines and technologies. The development of this instrument requires deep cross-disciplinary collaboration and innovation across multiple disciplines, including physics, biology, chemistry, semiconductors, and computer science, to build a high-precision single-molecule nanopore sequencing system from the underlying core modules.
[0003] Nanopore sequencing requires a sufficiently sharp sensing region within the porin to achieve high spatial resolution in both the lateral and longitudinal directions. Currently, only a few naturally occurring porins, such as the Mycobacterium smegmatis porin A (MspA) and the curli-specific transporter (CsgG), meet the industrial requirements for single-molecule detectors. Discovering more high-quality porins suitable for single-molecule sequencing through gene mining remains an unresolved issue.
[0004] Summary of the Invention
[0005] The main purpose of the present invention is to provide a porin monomer, a porin and a mutant thereof and applications thereof, so as to solve the problem of poor pore stability of porin in the prior art.
[0006] In order to achieve the above object, according to a first aspect of the present invention, a porin monomer is provided, which comprises: (a) a protein consisting of the amino acid sequence shown in SEQ ID NO: 1; or (b) a protein mutant, the amino acid sequence of which is shown in SEQ ID NO: 2; At least one of the following positions of the amino acid sequence shown in NO: 1 is substituted, deleted and / or one or more amino acids are added: 63, 64, 65, 66, 67, 68, 69, 103, 107, 108, 109, 113, 116, 117, 123, 126, 153, 156, 167, 169, 201, 206, 209, 213, 216, 218, and the protein mutant has the function of forming a pore structure by polymerization; or (c) a porin monomer that has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity with the protein in (a) or (b) and has the function of forming a pore structure by polymerization.
[0007] Further, in (b), the type of amino acid substituted at each site is independently selected from the following: G63 is mutated into G63A, G63S or G63T; E64 is mutated into E64G, E64A, E64S, E64T, E64N or E64Q; K65 is mutated into K65G, K65A, K65S, K65T, K65N or K65Q; F66 is mutated into F66G, F66A, F66S, F66T, F66N or F66Q; A67 is mutated into A67G, A67S, A67T, A67N or A67Q; N68 is mutated into N68G, N68A, N68S, N68T or N68Q; I69 mutated to I69G, I69A, I69S, I69T, I69N or I69Q; D103 mutated to D103N, D103A, D103G, D103S, D103T, D103Q, D103R or D103K; K107 mutated to K107N, K107A, K107G, K107S, K107T or K107Q; R109 mutated to R109W, R109F, R109Y, R109N, R109Q, R109S, R109T, R109A or R109G; R113 mutates to R113W, R113F, R113Y, R113N, R113Q, R113S, R113T, R113A or R113G; R116 mutates to R116N, R116A, R116G, R116S, R116T or R116Q; E117 mutates to E117N, E117A, E117G, E117S, E117T, E117Q, E117R or E117K; E123 mutates to E1 23N, E123A, E123G, E123S, E123T, E123Q, E123R or E123K; K126 is mutated to K126N, K126Q, K126S, K126T, K126A or K126G; D153 is mutated to D153A, D153G, D153V, D153L, D153I, D153Y, D153F or D153W; R156 is mutated to R156A, R156G, R156N, R156Q, R156S, R156T, R156 D or R156E; R167 mutated to R167N, R167Q, R167S, R167T, R167A or R167G; D169 mutated to D169N, D169Q, D169S, D169T, D169A or D169G; D201 mutated to D201N, D201Q, D201S, D201T, D201A or D201G; R206 mutated to R206N, R206Q, R206S, R206T, R206A, R206G, R206D or R206E;D209 mutated to D209N, D209Q, D209S, D209T, D209A, D209G, D209R or D209K; E213 mutated to E213N, E213Q, E213S, E213T, E213A or E213G; E216 mutated to E216N, E216Q, E216S, E216T, E216A or E216G; E218 mutated to E218N, E218Q, E218S, E218T, E218A or E218G.
[0008] In order to achieve the above object, according to a second aspect of the present invention, a protein construct is provided. The protein construct is composed of two or more porin monomers mentioned above, linked covalently or non-covalently.
[0009] In order to achieve the above object, according to the third aspect of the present invention, a porin is provided, wherein the porin is composed of 7 to 11 porin monomers mentioned above linked covalently or non-covalently.
[0010] Furthermore, the porin is composed of 9 porin monomers linked non-covalently.
[0011] Furthermore, the pore diameter of the porin is 0.5 to 3 nm.
[0012] In order to achieve the above object, according to a fourth aspect of the present invention, a kit is provided, which comprises the above-mentioned porin monomer, or the above-mentioned protein construct, or the above-mentioned porin.
[0013] Furthermore, the kit further comprises a membrane layer, which comprises a lipid layer or an artificial polymer membrane.
[0014] Furthermore, the kit also includes a sequencing buffer and / or cholesterol-linked single-stranded DNA.
[0015] To achieve the above object, according to the fourth aspect of the present invention, an isolated DNA molecule is provided, wherein the DNA molecule has a nucleotide sequence encoding the porin monomer, or encoding the protein construct, or encoding the porin.
[0016] Furthermore, the DNA molecule has the nucleotide sequence shown in SEQ ID NO:8.
[0017] Furthermore, a DNA molecule having 70% or more, preferably 80% or more, more preferably 90% or more, and even more preferably 95% or more identity with the nucleotide sequence shown in SEQ ID NO: 8 and encoding a protein having the same function.
[0018] In order to achieve the above object, according to the fifth aspect of the present invention, a recombinant vector is provided, which comprises the above DNA molecule.
[0019] In order to achieve the above object, according to the sixth aspect of the present invention, a host cell is provided, wherein the host cell is transformed with the above recombinant vector.
[0020] In order to achieve the above-mentioned purpose, according to the seventh aspect of the present invention, a nanopore sensor is provided, which comprises: a membrane layer; and a pore protein inserted in the middle of the membrane layer to form a pore channel, and when a voltage is applied across the membrane layer, the pore channel generates current; wherein the pore protein comprises the above-mentioned pore protein.
[0021] Furthermore, the membrane layer comprises a lipid layer or an artificial polymer membrane; preferably, the lipid layer comprises amphiphilic lipids; preferably, the amphiphilic lipids comprise a phospholipid bilayer; preferably, the lipid layer comprises a planar membrane layer or a liposome; preferably, the liposome comprises a multilayer liposome or a unilamellar liposome; preferably, the lipid layer comprises a phospholipid bilayer composed of diphytylphosphatidylcholine.
[0022] In order to achieve the above-mentioned objective, according to an eighth aspect of the present invention, a nanopore sequencing device is provided, which includes the above-mentioned nanopore sensor.
[0023] Furthermore, the nanopore sequencing device includes: an electrolytic cell containing a sequencing buffer; a nanopore sensor, the nanopore sensor is located in the center of the electrolytic cell and divides the electrolytic cell and the sequencing buffer into a positive electrolyte region and a negative electrolyte region; a first electrode and a second electrode, the first electrode and the second electrode are respectively arranged in the positive electrolyte region and the negative electrolyte region, and the first electrode and the second electrode are connected to the signal processing chip; preferably, the first electrode and the second electrode include metal or composite electrode materials; preferably, the first electrode and the second electrode are different, namely silver and silver chloride, respectively; or the first electrode and the second electrode are the same, each independently selected from gold, platinum, graphene or titanium nitride.
[0024] To achieve the above objectives, according to the ninth aspect of the present invention, a sequencing method is provided, which utilizes the above-mentioned porin, or the above-mentioned nanopore sensor, or the above-mentioned nanopore sequencing device to detect and analyze the electrical signal generated when the biomolecule to be tested passes through the pore of the porin, thereby determining the sequence of the biomolecule to be tested.
[0025] Furthermore, the biomolecule to be detected includes any one of the following modified or unmodified biomolecules: DNA, RNA or polypeptide.
[0026] Furthermore, the biological molecule to be detected is a target nucleic acid sequence, and the sequencing method includes: (a) contacting the target nucleic acid sequence with a nucleic acid binding protein, which controls the movement speed of the target nucleic acid sequence through the pore of the pore protein; (b) when a voltage is applied across the pore, the target nucleic acid sequence moves through the pore, and the electrical signal passing through the pore is measured, wherein different types of nucleotides generate different electrical signals when passing through the pore, thereby determining the sequence information of the target nucleic acid based on the electrical signal; preferably, the nucleic acid binding protein is selected from nucleases, polymerases, topoisomerases, ligases, helicases and single-stranded binding proteins; preferably, the electrical signal includes an electric current.
[0027] In order to achieve the above-mentioned objectives, according to the tenth aspect of the present invention, there is provided the above-mentioned porin monomer, the above-mentioned porin, the above-mentioned kit, the above-mentioned DNA molecule, the above-mentioned recombinant vector, the above-mentioned host cell, the above-mentioned nanopore sensor, or the above-mentioned nanopore sequencing device, for use in small molecule detection, DNA sequencing, RNA sequencing or polypeptide sequencing.
[0028] The present invention provides a new porin monomer that can polymerize to form porin BCP20. The protein and its mutants have good stability and can meet the requirements of single-molecule nanopore porin sequencing, realizing the detection of small molecules such as nucleotides, amino acids, sugars, vitamins, DNA, RNA and polypeptides. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0030] FIG1 shows a side view of the three-dimensional structure of BCP20 predicted according to Example 1 of the present invention.
[0031] FIG2 shows a top view of the three-dimensional structure of BCP20 predicted according to Example 1 of the present invention.
[0032] Figure 3 shows a schematic diagram of the predicted side chain structure of key amino acids in the BCP20 gating region (sensor region) according to Example 1 of the present invention, wherein A shows the distance between key amino acids in the sensor region, and B shows an enlarged view of the key amino acids in the sensor region.
[0033] FIG4 shows an SDS-PAGE image obtained by purification of the BCP20 protein according to Example 4 of the present invention.
[0034] FIG5 shows a schematic diagram of the sequencing library structure according to Example 5 of the present invention, wherein a: sense strand (top strand); b: antisense strand (bottom strand); c: double-stranded target fragment to be tested; d: helicase BCH105.
[0035] FIG6 shows a graph of the pore opening current of BCP20 according to Example 6 of the present invention at different voltages in a phospholipid bilayer.
[0036] FIG. 7 shows a graph showing the current changes when the DNA to be tested passes through the nanopore BCP20 according to Example 7 of the present invention.
[0037] Figure 8 shows a schematic structural diagram of a sequencing library combined with a single-stranded DNA containing cholesterol according to Example 7 of the present invention, wherein a: sense strand (top strand); b: antisense strand (bottom strand); c: double-stranded target fragment to be tested; d: helicase BCH105; e: single-stranded DNA containing cholesterol.
[0038] FIG9 shows the specific structure of iSp18 according to Example 5 of the present invention. DETAILED DESCRIPTION
[0039] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the embodiments.
[0040] As mentioned in the background technology, there are very few existing porins that can be used for single-molecule nanosequencing, and the protein stability is poor. Therefore, in this application, the inventors tried to develop a new porin. By using computer-assisted structure prediction gene mining methods, a new porin monomer was excavated from the deep-sea metagenome (derived from samples of the Mariana Trench at a depth of 11,000 meters). Multiple monomers can be covalently linked or non-covalently polymerized into a porin with a pore. The porin can be used as a detection protein and applied to the detection of small molecules such as nucleotides, amino acids, sugars, vitamins, or in nanopore-based DNA, RNA or peptide sequencing. Therefore, a series of protection schemes of the present application are proposed.
[0041] In a first typical embodiment of the present application, a porin monomer is provided, which comprises: (a) a protein consisting of the amino acid sequence shown in SEQ ID NO: 1; or (b) a protein mutant, the amino acid sequence of the protein mutant being as shown in SEQ ID NO: (c) a porin monomer having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity with the porin monomer in (a) or (b), and having a function of forming a pore structure by polymerization.
[0042] SEQ ID NO: 1:
[0043]
[0044] The porin monomers defined in (a) above can be polymerized to form the porin BCP20 having a pore structure. When applied to nanopore sequencing, they can allow the biomolecules to be tested to pass through the pore one by one, generating a current signal. Based on the sequence in (a), the protein can be mutated, for example, at other positions such as the mutation sites disclosed in (b) or (c), after substitution and / or deletion and / or addition of one or more amino acids, the pore structure and function of the porin can still be maintained. Mutating the porin monomers may affect the stability of the protein and aggregates, the inner diameter of the pore, and the amino acid residues on the inner wall of the pore, thereby affecting its physicochemical properties and the performance of the biomolecules to be tested. However, the conventional operation mode of mutation and the method of screening to obtain proteins with nanopore structure and functional activity are well known to those skilled in the art.
[0045] Identity in this specification refers to the "identity" between amino acid sequences, that is, the total ratio of identical amino acid residues in an amino acid sequence. Amino acid sequence identity can be determined using alignment programs such as BLAST (Basic Local Alignment Search Tool) and FASTA.
[0046] Proteins with 70%, 75%, 80%, 85%, 90%, 95%, 99% or more (such as 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, or even 99.9% or more) identity and the same function, whose active sites, active pockets, active mechanisms, protein structures, etc. are most likely the same as the proteins provided by the sequence in a), are homologous proteins obtained by amino acid mutations.
[0047] As used herein, amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine (Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).
[0048] Substitution and replacement rules generally refer to the fact that amino acids with similar properties will have similar effects when substituted with each other. For example, conservative amino acid substitutions may occur in the homologous proteins mentioned above. "Conservative amino acid substitutions" include, but are not limited to:
[0049] Hydrophobic amino acids (Ala, Cys, Gly, Pro, Met, Val, Ile, Leu) are replaced by other hydrophobic amino acids;
[0050] Substitution of bulky hydrophobic amino acids (Phe, Tyr, Trp) with other bulky hydrophobic amino acids;
[0051] Amino acids with positively charged side chains (Arg, His, Lys) are replaced by other amino acids with positively charged side chains;
[0052] Amino acids with polar and uncharged side chains (Ser, Thr, Asn, Gln) are replaced by other amino acids with polar and uncharged side chains.
[0053] Those skilled in the art may also perform conservative substitutions on amino acids according to amino acid substitution rules well known to those skilled in the art, such as the "blosum62 scoring matrix" in the prior art.
[0054] The "AlphaFold2-Multimer" used in this application is a publicly available artificial intelligence model that can predict the conformation of protein complexes. Its predictions of protein 3D structures are very close to those observed in real experiments using instruments such as cryo-electron microscopy. This allows for the acquisition of relatively realistic protein structures, thus guiding the study of protein structure and activity.
[0055] In a preferred embodiment, in (b), the type of amino acid substituted at each site is independently selected from the following: G63 mutates to G63A, G63S, G63T; E64 mutates to E64G, E64A, E64S, E64T, E64N, E64Q; K65 mutates to K65G, K65A, K65S, K65T, K65N, K65Q; F66 mutates to F66G, F66A, F66S, F66T, F66N, F66Q; A67 mutates to A67G, A67S, A67T, A67N, A67Q; N68 mutates to N68G, N68A, N68S, N68T, N68Q; I69 mutates to I69G, I69A, A, I69S, I69T, I69N, I69Q; D103 mutated to D103N, D103A, D103G, D103S, D103T, D103Q, D103R, D103K; K107 mutated to K107N, K107A, K107G, K107S, K107T, K107Q; R109 mutated to R109W, R109F, R109Y, R109N, R109Q, R109S, R109T, R109A, R109G; R113 mutated to R113W, R113F, R113Y, R113N, R113Q, R113S, R113T, R113A, R113G; R1 E16 mutated to R116N, R116A, R116G, R116S, R116T, R116Q; E117 mutated to E117N, E117A, E117G, E117S, E117T, E117Q, E117R, E117K; E123 mutated to E123N, E123A, E123G, E123S, E123T, E123Q, E123R, E123K; K126 mutated to K126N, K126Q, K126S, K126T, K126A, K126G; D153 mutated to D153A, D153G, D153V, D153L, D153I, D153Y, D153F, D 153W; R156 mutated to R156A, R156G, R156N, R156Q, R156S, R156T, R156D, R156E; R167 mutated to R167N, R167Q, R167S, R167T, R167A, R167G; D169 mutated to D169N, D169Q, D169S, D169T, D169A, D169G; D201 mutated to D201N, D201Q, D201S, D201T, D201A, D201G; R206 mutated to R206N, R206Q, R206S, R206T, R206A, R206G, R206D, R206E;D209 mutated to D209N, D209Q, D209S, D209T, D209A, D209G, D209R, and D209K; E213 mutated to E213N, E213Q, E213S, E213T, E213A, and E213G; E216 mutated to E216N, E216Q, E216S, E216T, E216A, and E216G; and E218 mutated to E218N, E218Q, E218S, E218T, E218A, and E218G. The letters before the numbers represent the original amino acids, and the letters after the numbers represent the mutated amino acids.
[0056] Among the above-mentioned mutation sites, there are mainly three types of sites: mutation sites in the gate region (sensor region), mutation sites at the entrance of the porin, and mutation sites in the transmembrane region of the porin. Among them, the mutation sites in the sensor region mainly determine the opening of the porin, thereby directly determining the opening current. Because the nucleic acid to be tested is negatively charged, by adjusting the amino acids at the mutation sites at the entrance of the porin, including but not limited to mutating uncharged or negatively charged amino acids to positively charged amino acids, or mutating positively charged amino acids to other types of amino acids, the capture rate of the library can be adjusted. Mutations in the transmembrane region of the porin can enhance the insertion stability of the porin on lipid or polymer membranes. In addition, charged amino acids located on the inner wall or outlet of the pore structure of the porin can also affect the perforation of the sample to be tested.
[0057] Among the above mutation sites, G63, N68, I69, E64, K65, F66, and A67 are located in the sensor region. D103, K107, R109, R113, R116, E117, E123, and K126 are located at the entrance of the porin protein. Mutation of their charge properties can play an important role in regulating the capture rate of the library and sequencing noise. D153 is located on the outer wall of the porin transmembrane region. Mutation of it to a hydrophobic amino acid can enhance the stability of the porin pore. R156, R167, D169, D201, E216, and E218 are located on the inner wall of the pore structure. They are charged amino acids on the inner wall of the porin barrel. Mutation of their charge can promote the smooth passage of the DNA molecules to be tested through the porin. R206, D209, and E213 are located in the loop region (loop region) of the pore structure. Mutation of them can promote the exit of the sequenced nucleic acid chain from the porin protein.
[0058] The aforementioned mutation sites and the resulting amino acids can affect the pore structure, throughput, and library capture capabilities of both monomeric and aggregated porin proteins. In particular, G63, N68, and I69, located in the sensor region, are located at the very center of the porin pore and have a significant impact on properties such as the pore diameter, stability, affinity for molecules passing through it, and the ability to generate electrical signals. N68 and I69 are the central amino acids in the sensor region, and mutations in these sites can directly determine the current magnitude and resolution of the analyte being permeated. Mutations at these two sites include, but are not limited to, mutations to polar, uncharged amino acids such as N, Q, S, T, G, or A, or amino acids with small side chains. G63, located below N68 and I69, can be mutated to amino acids with small side chains, including, but not limited to, A, S, or T, to minimize its impact on sequencing current and nucleic acid permeation.
[0059] In a second typical embodiment of the present application, a protein construct is provided. The protein construct is composed of two or more porin monomers mentioned above, linked covalently or non-covalently.
[0060] In a third typical embodiment of the present application, a porin is provided, wherein the porin is composed of 7-11 porin monomers mentioned above linked covalently or non-covalently.
[0061] In a preferred embodiment, the porin is composed of 9 porin monomers linked non-covalently.
[0062] Porin monomers can spontaneously aggregate together through hydrogen bonds, ionic bonds, hydrophobic interactions, and other forces to form porins. Therefore, porin monomers expressed and purified exist as multimers, especially nonamers, under non-denaturing conditions, while denaturing the protein results in the existence of porin monomers.
[0063] In a preferred embodiment, the pore diameter of the porin is 0.5-3 nm.
[0064] Within a certain range, the smaller the pore diameter of the porin, the higher its accuracy when used for sequencing. If the pore diameter of the porin is too large (more than one molecule may pass through the pore at a time), it is difficult to meet the needs of single-molecule sequencing. When the biological molecule to be tested passes through the overly large pore, the current signal generated may be missed or erroneous, resulting in low sequencing accuracy. In single-molecule sequencing, accurate sequencing results are obtained by sequencing the same molecule multiple times. Therefore, the higher the sequencing accuracy, the shorter the number of sequencing times and time required. Using porins with high sequencing accuracy for sequencing can greatly reduce sequencing time and reduce costs. This advantage is particularly evident in high-throughput sequencing.
[0065] In a fourth typical embodiment of the present application, a kit is provided, which includes the above-mentioned porin monomer, or protein construct, or porin.
[0066] To further improve operational convenience, in a preferred embodiment, the kit further comprises a membrane layer, which comprises a lipid layer or an artificial polymer membrane.
[0067] Preferably, the lipid layer comprises amphiphilic lipids; preferably, the amphiphilic lipids comprise a phospholipid bilayer; preferably, the lipid layer comprises a planar membrane layer or a liposome; preferably, the liposome comprises a multilamellar liposome or a unilamellar liposome; preferably, the lipid layer comprises a phospholipid bilayer composed of diphytylphosphatidylcholine.
[0068] Artificial polymer membranes include, but are not limited to, polysiloxanes, polyolefins, perfluoropolyethers, perfluoroalkyl polyethers, polystyrene, polyoxypropylene, polyvinyl acetate, polyoxybutylene, polyisoprene, polybutadiene, polyvinyl chloride, polyalkyl acrylates, polyalkyl methacrylates, polyacrylonitrile, polypropylene, PTHF, polymethacrylates, polyacrylates, polysulfones, polyethylene ethers, poly(propylene oxide) and copolymers thereof, alkyl-substituted C1-C6 alkyl acrylates and methacrylates, acrylamide, methacrylamide, (C1-C6 alkyl) acrylamide and methacrylamide, N,N-dialkyl-acrylamide, ethoxy acrylate and methacrylate, polyethylene glycol monomethacrylate and polyethylene glycol monomethyl ether methacrylate, hydroxy-substituted (C1-C6 alkyl) acrylamide and methacrylamide, hydroxy-substituted C1-C6 alkyl vinyl ether, sodium vinyl sulfonate, sodium styrene sulfonate, 2-acrylamide-2-methyl The invention also includes one or more of ethylenically unsaturated carboxylic acids having a total of 3 to 5 carbon atoms, amino(C1-C6 alkyl)-, mono(C1-C6 alkylamino)(C1-C6 alkyl)- and bis(C1-C6 alkylamino)(C1-C6 alkyl)-acrylates and methacrylates, allyl alcohol, 2-hydroxypropyl 3-trimethylammonium methacrylate chloride, dimethylaminoethyl methacrylate (DMAEMA), dimethylaminoethyl methacrylamide, glycerol methacrylate, N-(1,1-dimethyl-3-oxobutyl)acrylamide, cyclic imino ethers, vinyl ethers, cyclic ethers including epoxy derivatives, cyclic unsaturated ethers, N-substituted ethylenimines, β-lactones and β-lactams, ketene acetals, vinyl acetals or phosphoranes.
[0069] In a preferred embodiment, the kit further comprises a sequencing buffer and / or cholesterol-linked single-stranded DNA.
[0070] Using the porins, membrane layer, and sequencing buffer in the aforementioned kit, one or more porins can be inserted into the membrane layer within the sequencing buffer to form a nanopore sensor. The nanopore assay buffer provides a neutral environment that maintains the stability of the porins and membrane layer, and the metal ions contained in the nanopore assay buffer provide excellent conductivity. A variety of membrane options are available, allowing porins to be inserted into planar membranes, spherical liposomes, or lipid layers of varying compositions to form nanopore sensors.
[0071] Cholesterol present on single-stranded DNA and RNA molecules can bind to the lipid layer or artificial polymer membrane, facilitating nanopore capture of the sequencing library and reducing the amount of sample loaded. In actual use, the cholesterol in the above-mentioned kit can be first bound to the molecule to be tested and then added to the space within the nanopore sensor for sequencing.
[0072] In a fifth typical embodiment of the present application, an isolated DNA molecule is provided, wherein the DNA molecule has a nucleotide sequence encoding the above-mentioned porin monomer, or encoding the above-mentioned protein construct, or encoding the above-mentioned porin.
[0073] In a preferred embodiment, the DNA molecule has the nucleotide sequence shown in SEQ ID NO:8.
[0074] In a preferred embodiment, the present invention relates to a DNA molecule having more than 70%, preferably more than 80%, more preferably more than 90%, and further preferably more than 95% identity with the nucleotide sequence shown in SEQ ID NO: 8 (for example, it can be 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, or even 99.9% or more) and encoding a protein with the same function.
[0075] SEQ ID NO: 8:
[0076]
[0077] The above-mentioned DNA molecule can encode a porin monomer having the structure and function described above in the present application. Based on the sequence shown in SEQ ID NO: 8, the nucleotides are mutated and hybridized with the DNA molecule defined in (a) under stringent conditions without frameshift mutation. If the mutation occurs in the nucleotides encoding the porin pore, it may result in encoding a porin with an altered pore, affecting the pore diameter of the porin and the properties of the amino acid residues on the inner wall of the pore; if the mutation occurs in the nucleotides encoding the non-pore portion of the protein, it may affect the folding mode, three-dimensional structure and other properties of the encoded protein, thereby affecting the physicochemical properties and stability of the protein.
[0078] As used herein, "isolated" means altered "by the hand of man" from its natural state, i.e., if it occurs in nature, it is altered and / or separated from its original environment. For example, a polynucleotide or polypeptide naturally present in a living organism is not "isolated," whereas the same polynucleotide or polypeptide separated from its coexisting components in its natural state is "isolated" (as the term is used herein).
[0079] In a sixth typical embodiment of the present application, a recombinant vector is provided, which comprises the above-mentioned DNA molecule.
[0080] The porin-expressing DNA molecule is inserted into a recombinant vector, and the porin-expressing gene is replicated in large quantities by utilizing the recombinant vector's ability to replicate itself. "Recombinant" here refers to genetically engineered DNA produced by transplanting or splicing a gene from one species into the cells of a host organism of a different species. This DNA becomes part of the host's genetic structure and is replicated.
[0081] In a seventh typical embodiment of the present application, a host cell is provided, wherein the host cell is transformed with the above-mentioned recombinant vector.
[0082] The above-mentioned recombinant vector is transformed into a host cell, and the host cell replicates, transcribes, and translates the porin expression gene on the recombinant vector, thereby producing large quantities of porin. Host cells include commonly used host cells such as Escherichia coli, yeast, mammalian cells, and insect cells. The host cell folds the porin into the correct three-dimensional structure, thereby obtaining a porin with normal structure and function.
[0083] In an eighth typical embodiment of the present application, a nanopore sensor is provided, which includes: a membrane layer; and a pore protein inserted in the middle of the membrane layer to form a pore, and when a voltage is applied across the membrane layer, the pore generates current; wherein the pore protein includes the above-mentioned pore protein.
[0084] The nanopore sensor in this application specifically refers to a membrane layer with porin inserted into it. This type of nanopore sensor incorporates a porin with a pore channel, which can be oriented in a fixed direction. When a voltage is applied across the membrane, the pore diameter is perpendicular to the direction of the electric field. Ions on both sides of the membrane are forced through the porin pore by the electric field, generating a current.
[0085] In a preferred embodiment, the membrane layer comprises a lipid layer or an artificial polymer membrane; preferably, the lipid layer comprises amphiphilic lipids; preferably, the amphiphilic lipids comprise a phospholipid bilayer; preferably, the lipid layer comprises a planar membrane layer or a liposome; preferably, the liposome comprises a multilamellar liposome or a unilamellar liposome; preferably, the lipid layer comprises a phospholipid bilayer composed of diphytanoylphosphatidylcholine (DPhPC, 1,2-diphytanoyl-sn-glycero-3-phosphocholine).
[0086] There are many options for the selection of lipid layers. Porin can be inserted into planar membrane layers or spherical liposomes, or into lipid layers formed of different components to form nanopore sensors.
[0087] When an electric field is applied across the membrane, the biomolecules being tested pass through the pores of the porin under the influence of the voltage. These biomolecules include biomacromolecules that carry genetic information, such as DNA, RNA, polypeptides, or proteins. These biomolecules can also be modified with groups. These groups include, but are not limited to, cholesterol, polyethylene glycols of varying degrees of polymerization, biotin, or fluorescent groups.
[0088] In a ninth typical embodiment of the present application, a nanopore sequencing device is provided, which includes the above-mentioned nanopore sensor.
[0089] In a preferred embodiment, the nanopore sequencing device includes: an electrolytic cell containing a sequencing buffer; a nanopore sensor, the nanopore sensor being located in the center of the electrolytic cell and dividing the electrolytic cell and the sequencing buffer into a positive electrolyte region and a negative electrolyte region; a first electrode and a second electrode, the first electrode and the second electrode being respectively arranged in the positive electrolyte region and the negative electrolyte region, and the first electrode and the second electrode being connected to a signal processing chip; preferably, the first electrode and the second electrode comprise metal or composite electrode materials; preferably, the first electrode and the second electrode are different, being silver and silver chloride, respectively; or the first electrode and the second electrode are the same, and each is independently selected from gold, platinum, graphene, or titanium nitride.
[0090] A nanopore sequencing device includes an electrolytic cell containing an electrolyte, a nanopore sensor, a first electrode, and a second electrode. The nanopore sensor is placed in the center of the electrolytic cell, which is then divided into a positive electrolyte region and a negative electrolyte region. Two electrodes are positioned in each region, creating an electric field across the nanopore sensor. When the biomolecule to be tested passes through the porins on the membrane, a current amplitude is generated. This current amplitude is received and transmitted to a signal processing chip connected to the electrodes. Based on the current amplitude, the signal processing chip, i.e., the nanopore sequencing device containing the signal processing chip, can perform data analysis and determination of the sequence of the biomolecule to be tested.
[0091] In the tenth typical embodiment of the present application, a sequencing method is provided, which utilizes the above-mentioned pore protein, or nanopore sensor, or nanopore sequencing device to detect and analyze the electrical signal generated when the biological molecule to be tested passes through the pore of the pore protein, and determines the sequence of the biological molecule to be tested.
[0092] In a preferred embodiment, the biomolecule to be detected includes any one of the following modified or unmodified biomolecules: DNA, RNA, or polypeptide.
[0093] In a preferred embodiment, the biological molecule to be detected is a target nucleic acid sequence, and the sequencing method includes: (a) contacting the target nucleic acid sequence with a nucleic acid binding protein, which controls the movement speed of the target nucleic acid sequence through the pore of the pore protein; (b) when a voltage is applied across the pore, the target nucleic acid sequence moves through the pore, and the electrical signal passing through the pore is measured, wherein different types of nucleotides generate different electrical signals when passing through the pore, thereby determining the sequence information of the target nucleic acid based on the electrical signal; preferably, the nucleic acid binding protein is selected from nucleases, polymerases, topoisomerases, ligases, helicases and single-stranded binding proteins; preferably, the electrical signal includes an electric current.
[0094] When a voltage is applied across the pore, ions in the solution pass through the central gating region of the porin, generating an electric current. As the target nucleic acid sequence (including DNA or RNA) moves through the pore, the electrical signal passing through the pore is measured. Because different nucleotide types have different sizes, they block the current to varying degrees. Consequently, different nucleotide compositions produce different current signals when passing through the pore. By analyzing this current signal, the sequence of the target nucleic acid can be determined.
[0095] In an eleventh typical embodiment of the present application, a porin monomer, porin, kit, DNA molecule, recombinant vector, host cell, nanopore sensor, or nanopore sequencing device is provided for use in small molecule detection, DNA sequencing, RNA sequencing, or polypeptide sequencing.
[0096] The above-mentioned small molecules include but are not limited to small molecule compounds such as nucleotides, amino acids, polysaccharides or vitamins.
[0097] A major advantage of nanopore sequencing over traditional sequencing is that it avoids the impact of error accumulation, enabling extremely long read lengths. This can compensate for the gaps that inevitably occur when assembling short fragments in traditional sequencing, allowing for the determination of long deletions, duplications, inversions, and translocations within chromosomes. Nanopore sequencing can cover the entire transcriptome, typically several kilobases in length, providing a novel solution for scientific research on genome assembly, structural variation, and alternative splicing.
[0098] Since nanopore sequencing does not require PCR amplification, the original base modification information on the nucleic acid molecule to be tested can be retained, and the type, site and abundance of the modified base can be directly sequenced at one time. Therefore, the pore protein of the present application can also detect several nucleic acid molecules with DNA / RNA modified bases: including 5-methylcytosine (5mC), 6-methyladenine (m6A), 7-methylguanine (m7G), pseudouracil (pseudouridine, Ψ), etc. By performing specific model training and algorithm development on various modified bases, nanopore sequencing can complete the identification and positioning of more modified bases, thereby constructing a more complete genome / transcriptome modification map.
[0099] Furthermore, from a clinical perspective, nanopore sequencing's long read length, high portability, fast sequencing speed, and real-time readout make it highly suitable for critical epidemic monitoring and rapid pathogen detection (e.g., in large-scale epidemics such as Zika, Ebola, dengue, and the novel coronavirus), providing a highly timely response. In addition to viruses, nanopore sequencing can also be used for the rapid detection of other pathogens, such as bacteria and fungi.
[0100] Based on the common composition of proteins and nucleic acid molecules, nanopore sequencing platforms also have enormous application potential in the field of protein sequencing. For example, based on the research currently underway, by using protein unfolding enzymes as rate-control tools, characteristic protein signals have been successfully observed, and preliminary identification of protein types and modification states has been achieved, validating the feasibility of pore protein sequencing. In future developments, by further optimizing the rate-control system and developing suitable pore proteins and signal analysis algorithms, fingerprinting and even sequence identification of proteins at the single-molecule level may be achieved.
[0101] In addition to its applications in sequencing, the nanopore platform can also serve as a foundational detection platform, combined with sensing techniques to enable metabolomics analysis of various small and large molecules. By integrating genomics, proteomics, and metabolomics, the nanopore platform could ultimately develop into a universal measurement platform that meets the needs of comprehensive omics analysis, providing a powerful research tool for a deeper understanding of the laws of life and the mechanisms of disease.
[0102] The beneficial effects of the present application will be further explained in detail below with reference to specific embodiments.
[0103] Example 1 Predicted structure of AlphaFold2-Multimer of wild-type BCP20
[0104] Nine porin monomers (SEQ ID NO: 1) are non-covalently polymerized into nonamers to obtain porin BCP20. AlphaFold2-Multimer is used to predict the structure of BCP20. The prediction results are shown in Figures 1, 2, and 3. Figure 1 is a side view of the predicted structure of BCP20, Figure 2 is a top view of the predicted structure of BCP20, and Figure 3 is the side chain structure of the important amino acids in the sensor region of the predicted structure of BCP20, showing that the three amino acids in the amino acid side chain are G63, N68, and I69, and Figure 3 A shows that the distance between I69 is The distance between N68 is The distance between G63 is FIG3B shows an enlarged view of the side chain structure.
[0105] SEQ ID NO: 1:
[0106]
[0107] Example 2 Construction of expression vectors for porin monomers and their mutants
[0108] By the in-fusion method, the DNA sequence encoding the porin monomer (SEQ ID NO: 8) was inserted into the multiple cloning region of the vector pET24a after digestion with NdeI and XhoI. StrepII amino acids were added to the C-terminus of the amino acid sequence of the porin monomer (SEQ ID NO: 1) as a purification tag, wherein the screening tag was kanamycin. The constructed vector was named pET24a-BCP20. By the site-directed mutagenesis method, the Agilent site-directed mutagenesis kit was used, and the expression vector of the porin monomer was used as a template to construct the corresponding mutants. In this application, porin monomer mutants of G63A, N68Q and I69N were constructed, and the construction method of the mutants was consistent with that of the wild type.
[0109] Example 3 Cultivation and induction of porin monomer strains
[0110] The constructed porin monomer or its mutant expression plasmids were independently transformed into the E. coli expression strain E. coli BL21 (DE3), and the bacterial solution was evenly spread on a plate containing 50 μg / mL kanamycin and cultured at 37°C overnight. The next day, a single colony was picked and inoculated into 5 mL LB liquid culture medium containing 50 μg / mL kanamycin, and cultured at 37°C, 200 rpm, overnight. The bacterial solution obtained above was inoculated into 50 mL LB liquid culture medium containing 50 μg / mL kanamycin at a volume ratio of 1:100, and cultured at 37°C, 200 rpm for 4 hours. The expanded cultured bacterial solution was inoculated into 2 L LB liquid culture medium containing 50 μg / mL kanamycin at a volume ratio of 1:100, and cultured at 37°C, 200 rpm. Wait until OD 600 When the value reaches about 0.6-0.8, add IPTG to a final concentration of 0.5mM, culture at 16°C, 200rpm, and incubate for about 16-18 hours. Collect the bacterial solution by centrifugation at 8000rpm and freeze at -20°C until use.
[0111] Example 4 Extraction and purification of recombinant porin monomers
[0112] (1) Preparation of purification buffer
[0113] Buffer A: 20mM Tris-HCl, 250mM NaCl, 1% DDM, pH 8.0.
[0114] Buffer B: 20mM Tris-HCl, 250mM NaCl, 0.05% DDM, pH 8.0.
[0115] Buffer C: 20 mM Tris-HCl, 250 mM NaCl, 0.05% DDM, 5 mM desthiobiotin, pH 8.0.
[0116] (2) Purification step
[0117] Resuspend the cells thoroughly in 10 mL of Buffer A per 1 g of cells. Ultrasonicate the cells until the solution is clear. Rotate the solution overnight at 4°C. Centrifuge the next day at 18,000 rpm for 1 hour at 4°C. Remove the supernatant, filter through a 0.22 μm filter, and store at 4°C until ready for use.
[0118] A Strep-Tactin beads (IBA Lifesciences) column was equilibrated with Buffer A for 5 column volumes (CV) using an AKTA pure chromatography instrument and loaded with sample at 2 mL / min. After loading, the column was washed with Buffer B for 20 CV and eluted with Buffer C to collect the target protein.
[0119] The resulting protein was concentrated to 1 mL and passed through a Superdex 6increase 10 / 300GL (Cytiva) column equilibrated with buffer B to collect the target protein, which was then stored at -80°C. The purified target protein was subjected to SDS-PAGE electrophoresis, and the results are shown in Figure 4. These include electrophoretic bands for uncooked (porin nonamer BCP20) and 95°C boiled (porin monomers after denaturation). The results show that the target protein is in a polymeric state before boiling and in a monomeric state after boiling. The SDS-PAGE results of the mutant are consistent with those of the wild-type protein and are not presented in detail here.
[0120] Example 5 Library Construction
[0121] The sense and antisense strands of two partially complementary DNA chains (SEQ ID NO: 4) were annealed to form a linker. The linker was then ligated with the double-stranded target fragment pUC57 (SEQ ID NO: 5) using T4 DNA ligase at room temperature and purified to prepare a sequencing library. This library was then incubated with the helicase BCH105 (SEQ ID NO: 6) at 25°C for 1 hour (molar concentration ratio of 1:8) to form a sequencing library containing the BCH105 motor protein (as shown in Figure 5). During sequencing, this library was further able to complementarily bind to single-stranded DNA containing cholesterol (SEQ ID NO: 7, with cholesterol attached to the 5' end of the DNA), forming the structure shown in Figure 8.
[0122] Linker sequence sense chain: S1-(iSp18)4-S2.
[0123] The sequence of S1 is shown in SEQ ID NO: 2, the sequence of S2 is shown in SEQ ID NO: 3, and the structure of iSp18 is shown in FIG9 .
[0124] SEQ ID NO: 2:tttttttttttttttttttttttttttttttttttttt.
[0125] SEQ ID NO: 3: ggttgtttctgttggtgctgatattgct.
[0126] SEQ ID NO: 4 (linker sequence antisense strand):
[0127] gcaatatcagcaccaacagaaacaacctttgaggcgagcggtcaa.
[0128] SEQ ID NO: 5:
[0129]
[0130] SEQ ID NO: 6:
[0131]
[0132] Example 6 Construction of Nanopore Biosensor Using Porin BCP20 and Its Mutants
[0133] The current signal was collected using a patch clamp amplifier or other electrical signal amplifier. A single-channel nanopore detection system based on patch clamp and signal amplifier was constructed according to the method disclosed in the literature (Ji Z, Guo P. Channel from bacterial virus T7 DNA packaging motor for the differentiation of peptides composed of a mixture of acidic and basic amino acids. Biomaterials. 2019 May 21; 214: 119-222). Ag / AgCl electrodes were immersed in sequencing buffer and the electrodes were located in the cis and trans regions of the electrolytic cell, respectively. After the porin (i.e., the porin BCP20 purified in Example 4) was diluted a certain multiple using 1xPBS buffer, a single nanopore protein BCP20 was inserted into a phospholipid bilayer composed of diacylphosphatidylcholine (DPhPC, 1,2-diphytanoyl-sn-glycero-3-phosphocholine) under the action of an external electric field to form a nanopore biosensor. The dilution multiple was based on whether the porin was embedded in the membrane (i.e., embedded in the pore). Generally, a protein concentration of 0.1 mg / ml is used, diluted 100 times or 50 times or other multiples with PBS for trial. If a certain dilution concentration fails to embed the pore, it is necessary to reduce the dilution multiple and continue to try until the nanopore protein is successfully embedded in the membrane layer. Apply an external voltage to obtain the current amplitude value of a single pore protein. Figure 6 shows the biosensing current generated by the pore of the nanopore protein BCP20 when voltages of 0.02V, 0.04V, 0.10V, 0.14V and 0.18V are applied. Similarly, the pore protein BCP20 mutants (G63A, N68Q and I69N) also generated biosensing currents.
[0134] Example 7: Use of Porin BCP20 and Its Mutants for DNA Sequencing
[0135] The sequencing library containing the pUC57 sequence (SEQ ID NO: 5) prepared in Example 5 and cholesterol-linked single-stranded DNA (SEQ ID NO: 7, with cholesterol attached to the 5' end of the DNA) were mixed with sequencing buffer and added to the nanopore biosensor. Upon application of an external voltage of 0.14V or 0.18V, DNA was observed to be captured by the nanopore, generating a characteristic blockade current amplitude. The current amplitude varied as the DNA moved through the nanopore. Different DNA sequences produced different blockade current amplitudes. Cholesterol-linked single-stranded DNA can bind to the phospholipid bilayer, facilitating nanopore capture of the sequencing library and reducing the amount of sequencing library loaded. Figure 7 shows the current changes generated when the library DNA passed through the porin BCP20 under an applied voltage of 0.14V. Similarly, current changes were also generated when the library DNA passed through porin BCP20 mutants (G63A, N68Q, and I69N).
[0136] Single-stranded DNA sequence with cholesterol (SEQ ID NO: 7):
[0137] ttgaccgctcgcctc.
[0138] The cholesterol-containing single-stranded DNA can bind to the membrane and then complement the bottom chain of the library, thereby pulling the library to the vicinity of the nanopore, as shown in Figure 8, thereby increasing the capture efficiency during nanopore sequencing.
[0139] From the above description, it can be seen that the above-mentioned embodiments of the present invention achieve the following technical effects: the present invention has discovered a new porin monomer, which is polymerized into porin BCP20. The porin BCP20 and its mutants have good stability and can meet the requirements of single-molecule nanopore sequencing. The above-mentioned porin and its mutants can be used to form nanopore sensors and further nanopore sequencing devices to realize the detection of samples such as DNA, RNA, polypeptides, and proteins.
[0140] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A porin monomer, It is characterized in that The porin monomer comprises: (a) a protein consisting of the amino acid sequence shown in SEQ ID NO: 1; or (b) a protein mutant, wherein the amino acid sequence of the protein mutant undergoes substitution, deletion and / or addition of one or more amino acids at at least one of the following positions of the amino acid sequence shown in SEQ ID NO: 1: 68, 69, 63, 64, 65, 66, 67, 103, 107, 108, 109, 113, 116, 117, 123, 126, 153, 156, 167, 169, 201, 206, 209, 213, 216, 218, and the protein mutant has the function of forming a pore structure through polymerization; or (c) a porin monomer having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity to the protein described in (a) or (b), and having the function of forming a pore structure through polymerization.
2. The porin monomer according to claim 1, It is characterized in that In said (b), the type of amino acid substituted at each site is independently selected from the following: G63 mutation to G63A, G63S, or G63T; E64 mutation is E64G, E64A, E64S, E64T, E64N, or E64Q; K65 mutation to K65G, K65A, K65S, K65T, K65N, or K65Q; F66 mutation to F66G, F66A, F66S, F66T, F66N, or F66Q; A67 mutation to A67G, A67S, A67T, A67N, or A67Q; N68 mutation is N68G, N68A, N68S, N68T, or N68Q; I69 mutation is I69G, I69A, I69S, I69T, I69N or I69Q; D103 mutation is D103N, D103A, D103G, D103S, D103T, D103Q, D103R or D103K; K107 mutation is K107N, K107A, K107G, K107S, K107T, or K107Q; R109 mutation is R109W, R109F, R109Y, R109N, R109Q, R109S, R109T, R109A or R109G; R113 mutation is R113W, R113F, R113Y, R113N, R113Q, R113S, R113T, R113A or R113G; R116 mutation to R116N, R116A, R116G, R116S, R116T, or R116Q; E117 mutation is E117N, E117A, E117G, E117S, E117T, E117Q, E117R, or E117K; E123 mutation is E123N, E123A, E123G, E123S, E123T, E123Q, E123R or E123K; K126 mutations are K126N, K126Q, K126S, K126T, K126A, or K126G; D153 mutation is D153A, D153G, D153V, D153L, D153I, D153Y, D153F or D153W; R156 is mutated to R156A, R156G, R156N, R156Q, R156S, R156T, R156D or R156E; R167 mutation to R167N, R167Q, R167S, R167T, R167A, or R167G; D169 mutation is D169N, D169Q, D169S, D169T, D169A or D169G; D201 mutated to D201N, D201Q, D201S, D201T, D201A, or D201G; R206 is mutated to R206N, R206Q, R206S, R206T, R206A, R206G, R206D or R206E; D209 is mutated to D209N, D209Q, D209S, D209T, D209A, D209G, D209R or D209K; E213 mutation is E213N, E213Q, E213S, E213T, E213A or E213G; E216 mutation is E216N, E216Q, E216S, E216T, E216A or E216G; E218 mutation is E218N, E218Q, E218S, E218T, E218A or E218G.
3. A protein construct, It is characterized in that The protein construct is formed by covalently or non-covalently linking two or more porin monomers according to claim 1 or 2.
4. A porin, It is characterized in that The porin is formed by covalently or non-covalently linking 7 to 11 porin monomers as described in claim 1 or 2.
5. The porin according to claim 4, It is characterized in that The porin is composed of 9 porin monomers connected non-covalently.
6. The porin according to claim 4 or 5, It is characterized in that The pore diameter of the porin is 0.5-3 nm.
7. A kit, It is characterized in that The kit comprises the porin monomer of claim 1 or 2, or the protein construct of claim 3, or the porin of any one of claims 4 to 6.
8. The kit according to claim 7, It is characterized in that The kit further comprises a membrane layer, which comprises a lipid layer or an artificial polymer membrane.
9. The kit according to claim 7 or 8, It is characterized in that The kit also includes a sequencing buffer and / or single-stranded DNA linked to cholesterol.
10. An isolated DNA molecule, It is characterized in that The DNA molecule has: A nucleotide sequence encoding the porin monomer according to claim 1 or 2, or encoding the protein construct according to claim 3, or encoding the porin according to any one of claims 4 to 6.
11. The DNA molecule according to claim 10, It is characterized in that The DNA molecule has the nucleotide sequence shown in SEQ ID NO:
8.
12. The DNA molecule according to claim 11, It is characterized in that A DNA molecule that has 70% or more, preferably 80% or more, more preferably 90% or more, and even more preferably 95% or more identity with the nucleotide sequence shown in SEQ ID NO: 8 and encodes a protein having the same function.
13. A recombinant vector, It is characterized in that The recombinant vector comprises the DNA molecule according to any one of claims 10 to 12.
14. A host cell, It is characterized in that The host cell is transformed with the recombinant vector according to claim 13.
15. A nanopore sensor, It is characterized in that The nanopore sensor comprises: film layer; and a porin protein inserted in the middle of the membrane layer to form a pore, and when a voltage is applied across the membrane layer, the pore generates an electric current; Wherein, the porin comprises the porin according to any one of claims 4 to 6.
16. The nanopore sensor according to claim 15, It is characterized in that The membrane layer includes a lipid layer or an artificial polymer membrane; Preferably, the lipid layer comprises amphiphilic lipids; Preferably, the amphiphilic lipid comprises a phospholipid bilayer; Preferably, the lipid layer comprises a planar membrane layer or a liposome; Preferably, the liposomes comprise multilamellar liposomes or unilamellar liposomes; Preferably, the lipid layer comprises a phospholipid bilayer composed of diphytylphosphatidylcholine.
17. A nanopore sequencing device, It is characterized in that The nanopore sequencing device comprises the nanopore sensor according to claim 15 or 16.
18. The nanopore sequencing device according to claim 17, It is characterized in that The nanopore sequencing device comprises: an electrolytic cell containing a sequencing buffer; A nanopore sensor, wherein the nanopore sensor is located in the center of the electrolytic cell and divides the electrolytic cell and the sequencing buffer into a positive electrode electrolyte region and a negative electrode electrolyte region; A first electrode and a second electrode, wherein the first electrode and the second electrode are respectively arranged in the positive electrode electrolyte region and the negative electrode electrolyte region, and the first electrode and the second electrode are connected to a signal processing chip; Preferably, the first electrode and the second electrode comprise metal or composite electrode materials; Preferably, the first electrode and the second electrode are different, being silver and silver chloride respectively; or the first electrode and the second electrode are the same, and are each independently selected from gold, platinum, graphene or titanium nitride.
19. A sequencing method, It is characterized in that The sequencing method utilizes the pore protein described in any one of claims 4 to 6, or the nanopore sensor described in claim 15 or 16, or the nanopore sequencing device described in claim 17 or 18 to detect and analyze the electrical signal generated when the biological molecule to be tested passes through the pore of the pore protein, so as to determine the sequence of the biological molecule to be tested.
20. The sequencing method according to claim 19, It is characterized in that The biomolecule to be detected includes any one of the following modified or unmodified biomolecules: DNA, RNA or polypeptide.
21. The sequencing method according to claim 19, It is characterized in that The biomolecule to be detected is a target nucleic acid sequence, and the sequencing method comprises: (a) contacting the target nucleic acid sequence with a nucleic acid binding protein, wherein the nucleic acid binding protein controls the speed at which the target nucleic acid sequence moves through the pore of the porin; (b) when a voltage is applied across the pore, the target nucleic acid sequence moves through the pore, and an electrical signal passing through the pore is measured, wherein different types of nucleotides generate different electrical signals when passing through the pore, thereby determining the sequence information of the target nucleic acid based on the electrical signal; Preferably, the nucleic acid binding protein is selected from the group consisting of a nuclease, a polymerase, a topoisomerase, a ligase, a helicase and a single-stranded binding protein; Preferably, the electrical signal comprises an electric current.
22. Use of the porin monomer of claim 1 or 2, the porin of any one of claims 4 to 6, the kit of any one of claims 7 to 9, the DNA molecule of any one of claims 10 to 12, the recombinant vector of claim 13, the host cell of claim 14, the nanopore sensor of claim 15 or 16, or the nanopore sequencing device of claim 17 or 18 in small molecule detection, DNA sequencing, RNA sequencing or polypeptide sequencing.