Pore protein monomer, pore protein, mutant of pore protein and application of pore protein monomer and pore protein
Patent Information
- Application Number
- CN202280102619.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-07-22
AI Technical Summary
Existing nanopore sequencers have deficiencies in sequencing accuracy, throughput, and chip stability, and cannot meet the ultimate needs of molecular biology research. Moreover, the scarcity of natural porin types makes it difficult to meet the needs of single-molecule detection.
A new nanopore protein BCP35 was discovered through gene mining. It was used to form a pore structure and combined with lipid layer and electrical signal sequencing methods to construct highly stable nanopore sensors and sequencing devices to achieve the detection of nucleotides, amino acids, and sugars. Detection of small biological molecules such as vitamins and vitamins and sequencing of DNA, RNA and peptides.
It achieves high-accuracy and high-stability single-molecule nanopore sequencing, which can meet the needs of detection of small biological molecules and sequencing of nucleic acids and peptides, reduces sequencing time and cost, and provides more detailed genome and transcriptome modification maps. .
Smart Images

Figure CN120359234A_ABST
Abstract
Description
Porin monomer, porin and its mutant and its application Technical Field
[0001] The present invention relates to the field of single-molecule sequencing, and in particular to a porin monomer, a porin and a mutant thereof, and applications thereof. Background Art
[0002] Several commercial nanopore sequencers are currently available on the market, including the MinION, GridION, and PromethION from Oxford Nanopore Technologies in the UK, and the QNome-3841 from Qi Carbon. However, these still have significant shortcomings in sequencing accuracy, throughput, chip stability, and applicable scenarios, failing to meet the ultimate requirements of molecular biology research. Therefore, there is an urgent need to develop a single-molecule sequencer with high accuracy, high integration, and high stability. A nanopore-based single-molecule sequencer is a highly multidisciplinary and multi-technology detection system. The development of such an instrument requires deep cross-disciplinary collaboration and innovation across multiple disciplines, including physics, biology, chemistry, semiconductors, and computer science, building a high-precision single-molecule nanopore sequencing system from the underlying core modules.
[0003] Nanopore sequencing requires a sufficiently sharp sensing region within the pore protein to achieve high spatial resolution in both the horizontal and vertical directions. Currently, only a few natural proteins, such as the Mycobacterium smegmatis porin A (MspA) and the curli-specific transporter (CsgG), meet the industrial requirements for single-molecule detectors. Identifying more excellent porins suitable for single-molecule sequencing through gene mining remains an unresolved issue.
[0004] Summary of the Invention
[0005] The main purpose of the present invention is to provide a porin monomer, a porin and a mutant thereof and applications thereof, so as to provide a new porin and a nanopore sensor that can be used for nanopore sequencing.
[0006] To achieve the above objectives, according to a first aspect of the present invention, a porin monomer is provided, comprising: (a) a protein having an amino acid sequence as shown in SEQ ID NO: 1; or (b) a protein mutant, wherein the amino acid sequence of the protein mutant is substituted, deleted, and / or one or more amino acids are added at at least one of the following positions in SEQ ID NO: 1: 91, 98, 99, 125, 126, 127, 131, 134, 136, 140, 152, 155, 174, 177, 183, 185, 188, 219, 222, 223, 230, 232, 238, 239, and the protein mutant has the function of forming a pore structure upon polymerization; or (c) a protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity with the protein in (a) or (b), and has the function of forming a pore structure upon polymerization.
[0007] Furthermore, in b), the types of substituted amino acids are independently selected from the following: P91 is mutated into P91G, P91A or P91T; S98 is mutated into S98G, S98A, S98T, S98N or S98Q; N99 is mutated into N99G, N99A, N99S, N99T or N99Q; E125 is mutated into E125K, E125R, E125 G, E125A, E125S, E125T, E125N or E125Q; R126 is mutated to R126K, R126G, R126A, R126S, R126T, R126N or R126Q; K127 is mutated to K127R, K127G, K127A, K127S, K127T, K127N or K127Q; D131 Mutation: D131K, D131R, D131G, D131A, D131S, D131T, D131N or D131Q; K134 mutation: K134R, K134G, K134A, K134S, K134T, K134N or K134Q; R136 mutation: R136K, R136G, R136A, R136S, R136T, R136N or R136Q; R140 mutation: R140K, R140G, R140A, R140Q 40S, R140T, R140N or R140Q; K152 mutated to K152R, K152G, K152A, K152S, K152T, K152N or K152Q; D155 mutated to D155R, D155K, D155G, D155A, D155S, D155T, D155N or D155Q; T174 mutated to T174A, T174G, T174V, T174L, T174I, T174Y, T174F or T174Q 74W; E177 mutated to E177A, E177G, E177S, E177T, E177N or E177Q; H183 mutated to H183A, H183G, H183V, H183L, H183I, H183Y, H183F or H183W; K185 mutated to K185A, K185G, K185V, K185L, K185I, K185Y, K185F or K185W; R188 mutated to R188A, R188G, R188S, R188T, R188N or R188Q; K219 is mutated to K219A, K219G, K219V, K219L, K219I, K219Y, K219F or K219W; R222 is mutated to R222A, R222G, R222S, R222T, R222N or R222Q; S223 is mutated to S223A, S223G, S223V, S223L, S223I, S223Y, S223F or S223W;R230 is mutated to R230A, R230G, R230S, R230T, R230N or R230Q; K232 is mutated to K232A, K232G, K232S, K232T, K232N or K232Q; E238 is mutated to E238A, E238G, E238S, E238T, E238N or E238Q; R239 is mutated to R239A, R239G, R239V, R239L, R239I, R239Y, R239F or R239W.
[0008] Furthermore, the porin monomer includes a protein having any one of the amino acid sequences of SEQ ID NO: 2 to SEQ ID NO: 4.
[0009] In order to achieve the above object, according to a second aspect of the present invention, a protein construct is provided. The protein construct is composed of two or more porin monomers mentioned above, linked covalently or non-covalently.
[0010] In order to achieve the above object, according to the third aspect of the present invention, a porin is provided. The porin is composed of 7-11 porin monomers mentioned above, preferably 9, linked by covalent or non-covalent means.
[0011] Furthermore, the porin is composed of 9 porin monomers linked non-covalently, and the porin monomers include proteins having any one of the amino acid sequences of SEQ ID NO: 1 to SEQ ID NO: 4.
[0012] Furthermore, the pore diameter of the porin is 0.5 to 3 nm.
[0013] In order to achieve the above object, according to the fourth aspect of the present invention, a kit is provided, which comprises the above-mentioned porin monomer, or the above-mentioned protein construct, or the above-mentioned porin.
[0014] Furthermore, the kit further comprises a membrane layer, which comprises a lipid layer or an artificial polymer membrane.
[0015] Furthermore, the kit further comprises at least one of the following: a sequencing buffer, a nuclease, a polymerase, a topoisomerase, a ligase, a helicase, and a cholesterol-linked single-stranded DNA.
[0016] To achieve the above object, according to the fifth aspect of the present invention, an isolated DNA molecule is provided, which has: a nucleotide sequence encoding the above protein monomer; or a nucleotide sequence encoding the above protein construct, or a nucleotide sequence encoding the above porin.
[0017] Furthermore, a DNA molecule having 70% or more, preferably 80% or more, more preferably 90% or more, further preferably 99% or more, and most preferably 99% or more identity with the nucleotide sequence shown in SEQ ID NO: 5 and encoding a protein having the same function.
[0018] In order to achieve the above object, according to the sixth aspect of the present invention, a recombinant vector is provided, which comprises the above DNA molecule.
[0019] In order to achieve the above object, according to the seventh aspect of the present invention, a host cell is provided, wherein the host cell is transformed with the above recombinant vector.
[0020] In order to achieve the above-mentioned purpose, according to the eighth aspect of the present invention, a nanopore sensor is provided, which comprises: a membrane layer; and a pore protein inserted into the membrane layer and forming a pore, and when a voltage is applied across the membrane layer, the pore generates current; wherein the pore protein comprises the above-mentioned pore protein.
[0021] Furthermore, the membrane layer comprises a lipid layer or an artificial polymer membrane; preferably, the lipid layer comprises amphiphilic lipids; preferably, the amphiphilic lipids comprise a phospholipid bilayer; preferably, the lipid layer comprises a planar membrane layer or a liposome; preferably, the liposome comprises a multilayer liposome or a unilamellar liposome; preferably, the lipid layer comprises a phospholipid bilayer composed of diphytylphosphatidylcholine.
[0022] Furthermore, when a voltage is applied across the membrane layer, the biological molecules to be measured pass through the pores in the nanopore sensor and shift, and the pores generate a changing current; preferably, the biological molecules to be measured include DNA, RNA or polypeptides; preferably, the DNA and / or RNA include any one or more of the following modified bases: 5-methylcytosine, 6-methyladenine, 7-methylguanine, pseudouracil.
[0023] In order to achieve the above-mentioned objective, according to a ninth aspect of the present invention, a nanopore sequencing device is provided, which includes the above-mentioned nanopore sensor.
[0024] Furthermore, the nanopore sequencing device includes: an electrolytic cell, the electrolytic cell contains a sequencing buffer; a nanopore sensor, the nanopore sensor is located in the center of the electrolytic cell and divides the electrolytic cell and the sequencing buffer into a positive electrolyte region and a negative electrolyte region; a first electrode and a second electrode, the first electrode and the second electrode are respectively arranged in the positive electrolyte region and the negative electrolyte region, and the first electrode and the second electrode are connected to the signal processing chip; preferably, the first electrode and the second electrode include metal or composite electrode materials; preferably, the first electrode and the second electrode are different, namely silver and silver chloride, respectively; or the first electrode and the second electrode are the same, including gold, platinum, graphene or titanium nitride.
[0025] To achieve the above objectives, according to the tenth aspect of the present invention, a sequencing method is provided, which utilizes the above-mentioned porin, or the above-mentioned nanopore sensor, or the above-mentioned nanopore sequencing device to determine the sequence of the biomolecule to be tested by detecting and analyzing the electrical signal generated when the biomolecule to be tested passes through the pore of the porin.
[0026] Furthermore, the biomolecule to be detected includes modified or unmodified DNA, RNA or polypeptide; preferably, the electrical signal includes electric current.
[0027] Furthermore, the biological molecule to be detected is a target nucleic acid sequence, and the sequencing method includes: (a) contacting the nucleic acid sequence with the above-mentioned pore protein and nucleic acid binding protein, so that the nucleic acid binding protein controls the movement speed of the target nucleic acid sequence through the pore of the pore protein, wherein the nucleic acid binding protein is selected from any one or more of nucleases, polymerases, topoisomerases, ligases, helicases or single-stranded binding proteins; (b) when a voltage is applied across the pore, when the nucleic acid sequence moves through the pore, measuring the electrical signal passing through the pore, wherein different types of nucleotides generate different electrical signals when passing through the pore, thereby determining the sequence information of the nucleic acid based on the electrical signal.
[0028] In order to achieve the above-mentioned objectives, according to the eleventh aspect of the present invention, there is provided the above-mentioned porin monomer, or the above-mentioned porin, or the above-mentioned kit, or the above-mentioned DNA molecule, or the above-mentioned recombinant vector, or the above-mentioned host cell, or the above-mentioned nanopore sensor, or the above-mentioned nanopore sequencing device, or the above-mentioned sequencing method for use in biological small molecule detection, nucleic acid sequencing or polypeptide sequencing.
[0029] The present invention provides a new porin and its mutants that can be used for nanopore sequencing. The nanopore sensor composed of this protein or its mutants has good stability, can meet the needs of single-molecule nanosequencing, and can realize the detection of biological small molecules such as nucleotides, amino acids, sugars, vitamins, etc. It can also be used for sequencing modified or unmodified DNA, RNA or polypeptides. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0031] FIG1 shows a side view of the three-dimensional structure of BCP35 predicted according to Example 1 of the present invention.
[0032] FIG2 shows a top view of the three-dimensional structure of BCP35 predicted according to Example 1 of the present invention.
[0033] FIG3 shows a schematic diagram of the predicted structure of key amino acids in the BCP35 gating region according to Example 1 of the present invention and the distances between the key amino acids.
[0034] FIG4 shows an enlarged view of the structure of key amino acids in the gating region of BCP35 according to Example 1 of the present invention.
[0035] FIG5 shows a schematic diagram of key amino acids at the entrance of the predicted structure of BCP35 according to Example 1 of the present invention.
[0036] FIG6 shows a schematic diagram of key amino acids at the inner wall and outlet of the pore in the predicted structure of BCP35 according to Example 1 of the present invention.
[0037] FIG7 shows a schematic diagram of key amino acids in the transmembrane region of the predicted structure of BCP35 according to Example 1 of the present invention.
[0038] FIG8 shows an SDS-PAGE image obtained by purification of the BCP35 protein according to Example 4 of the present invention.
[0039] Figure 9 shows a schematic diagram of the sequencing library structure according to Example 5 of the present invention, wherein A is a schematic diagram of the structure of the sequencing library, and B is a schematic diagram of the structure of the sequencing library combined with single-stranded DNA containing cholesterol, a: positive chain; b: antisense chain; c: double-stranded target fragment to be tested; d: helicase BCH105; e: single-stranded DNA containing cholesterol.
[0040] FIG10 shows the specific structure of iSp18 according to Example 5 of the present invention.
[0041] FIG. 11 shows a graph of the pore opening current of BCP35 according to Example 6 of the present invention at different voltages in a phospholipid bilayer.
[0042] FIG12 shows a graph showing the current changes when the DNA to be tested passes through the nanopore BCP35 according to Example 7 of the present invention. DETAILED DESCRIPTION
[0043] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the embodiments.
[0044] As mentioned in the background art, there are very few types of pore proteins that can be used for single-molecule nanopore sequencing, and the selectivity is greatly limited. Therefore, in the present application, the inventors used gene mining methods based on computer-assisted structure prediction to dig out a new nanopore protein from deep-sea metagenomes (samples derived from 11,000 meters deep in the Mariana Trench), named BCP35, which is composed of nine monomers of the same type. Because it has a pore structure, it can be used as a detection protein and applied to the detection of small biological molecules such as nucleotides, amino acids, sugars, vitamins, or in nanopore-based DNA, RNA or peptide sequencing. Therefore, a series of protection schemes of the present application are proposed.
[0045] In a first exemplary embodiment of the present application, a porin monomer is provided, comprising: (a) a protein having an amino acid sequence as shown in SEQ ID NO: 1; or (b) a protein mutant, wherein the amino acid sequence of the protein mutant is substituted, deleted, and / or added with one or more amino acids at at least one of the following positions in SEQ ID NO: 1: 91, 98, 99, 125, 126, 127, 131, 134, 136, 140, 152, 155, 174, 177, 183, 185, 188, 219, 222, 223, 230, 232, 238, 239, and the protein mutant has the function of forming a pore structure upon polymerization; or (c) a protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity with the protein in (a) or (b), and has the function of forming a pore structure upon polymerization.
[0046] SEQ ID NO: 1:
[0047]
[0048] The porin monomers defined in (a) above can be polymerized to form the porin BCP35 having a pore structure. When applied to nanopore sequencing, they can allow the biomolecules to be tested to pass through the pore one by one, generating a current signal. Based on the sequence in (a), the protein can be mutated, for example, at other positions such as the mutation sites disclosed in (b), after substitution and / or deletion and / or addition of one or more amino acids, the pore structure and function of the porin can still be maintained. Mutating the porin monomers may affect the stability of the protein and aggregates, the inner diameter of the pore, and the amino acid residues on the inner wall of the pore, thereby affecting its physicochemical properties and the performance of the biomolecules to be tested. However, the conventional operation mode of mutation and the method of screening to obtain proteins with nanopore structure and functional activity are well known to those skilled in the art.
[0049] Identity as used herein refers to the "identity" between amino acid sequences or nucleic acid sequences, i.e., the total ratio of identical amino acid residues or nucleotides in an amino acid sequence or nucleic acid sequence. The identity of amino acid sequences or nucleic acid sequences can be determined using alignment programs such as BLAST (Basic Local Alignment Search Tool) and FASTA.
[0050] Proteins with 70%, 75%, 80%, 85%, 90%, 95%, 99% or more (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, or even 99.9% or more) identity and the same function have a high probability of being the same as the protein provided by the sequence in a).
[0051] As used herein, amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine (Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).
[0052] Generally speaking, according to the rules of substitution, replacement, etc., amino acids with similar properties will have similar effects when substituted with each other. For example, in the above protein, conservative amino acid substitutions may occur. "Conservative amino acid substitutions" include but are not limited to:
[0053] Hydrophobic amino acids (Ala, Cys, Gly, Pro, Met, Val, Ile, Leu) are replaced by other hydrophobic amino acids;
[0054] Substitution of bulky hydrophobic amino acids (Phe, Tyr, Trp) with other bulky hydrophobic amino acids;
[0055] Amino acids with positively charged side chains (Arg, His, Lys) are replaced by other amino acids with positively charged side chains;
[0056] Amino acids with polar and uncharged side chains (Ser, Thr, Asn, Gln) are replaced by other amino acids with polar and uncharged side chains.
[0057] Those skilled in the art may also perform conservative substitutions on amino acids according to amino acid substitution rules well known to those skilled in the art, such as the "blosum62 scoring matrix" in the prior art.
[0058] The "AlphaFold2-Multimer" used in this application is a publicly available artificial intelligence model that can predict the conformation of protein complexes. Its predictions of protein 3D structures are very close to those observed in real experiments using instruments such as cryo-electron microscopy. This allows for the acquisition of relatively realistic protein structures, thus guiding the study of protein structure and activity.
[0059] In a preferred embodiment, in b), the types of substituted amino acids are independently selected from the following: P91 is mutated into P91G, P91A or P91T; S98 is mutated into S98G, S98A, S98T, S98N or S98Q; N99 is mutated into N99G, N99A, N99S, N99T or N99Q; E125 is mutated into E125K, E125R, E125G, E125A, E125S, E125T, E125N or E125Q; R126 is mutated into R126K, R126G, R126A, R126S, R126T, R126N or R126Q; 26N or R126Q; K127 mutated to K127R, K127G, K127A, K127S, K127T, K127N or K127Q; D131 mutated to D131K, D131R, D131G, D131A, D131S, D131T, D131N or D131Q; K134 mutated to K134R, K134G, K134A, K134S, K134T, K134N or K134Q; R136 mutated to R136K, R136G, R136A, R136S, R136T, R136N or R136Q; R140 mutated to R140K, R140G, R140A, R140S, R140T, R140N or R140Q; K152 mutated to K152R, K152G, K152A, K152S, K152T, K152N or K152Q; D155 mutated to D155R, D155K, D155G, D155A, D155S, D155T, D155N or D155Q; T174 mutated to T174A, T174G, T174V, T174L, T174I, T174Y, T174F or T174W; E17 7 mutations to E177A, E177G, E177S, E177T, E177N or E177Q; H183 mutations to H183A, H183G, H183V, H183L, H183I, H183Y, H183F or H183W; K185 mutations to K185A, K185G, K185V, K185L, K185I, K185Y, K185F or K185W; R188 mutations to R188A, R188G, R188S, R188T, R188N or R188Q; K219 mutations to K219A, K219G, K219V, K219L, K219I, K219Y, K219F, or K219W; R222 is mutated to R222A, R222G, R222S, R222T, R222N, or R222Q; S223 is mutated to S223A, S223G, S223V, S223L, S223I, S223Y, S223F, or S223W;R230 is mutated to R230A, R230G, R230S, R230T, R230N or R230Q; K232 is mutated to K232A, K232G, K232S, K232T, K232N or K232Q; E238 is mutated to E238A, E238G, E238S, E238T, E238N or E238Q; R239 is mutated to R239A, R239G, R239V, R239L, R239I, R239Y, R239F or R239W.
[0060] Most of these amino acid sites are located at the porin's entrance, pore wall, and exit, and are crucial for library capture and the smooth passage of library DNA through the porin. Because nucleic acids are negatively charged, increasing or decreasing the amount of positive charge at the entrance can modulate the porin's library capture capacity.
[0061] Among the above-mentioned mutation sites, there are mainly three types of sites: mutation sites in the gate region (sensor region), mutation sites at the entrance of the porin, and mutation sites in the transmembrane region of the porin. Among them, the mutation sites in the sensor region mainly determine the opening of the porin, thereby directly determining the opening current. Because the nucleic acid to be tested is negatively charged, by adjusting the amino acids at the mutation sites at the entrance of the porin, including but not limited to mutating uncharged or negatively charged amino acids to positively charged amino acids, or mutating positively charged amino acids to other types of amino acids, the capture rate of the library can be adjusted. Mutations in the transmembrane region of the porin can enhance the insertion stability of the porin on lipid or polymer membranes. In addition, charged amino acids located on the inner wall or outlet of the pore structure of the porin can also affect the perforation of the sample to be tested.
[0062] Among the aforementioned mutation sites, P91, S98, and N99 are located in the sensor region. E125, R126, K127, D131, K134, R136, R140, K152, and D155 are located at the porin entrance. Mutating their charge properties can play a significant role in regulating library capture rate and sequencing noise. T174, H183, K185, K219, S223, and R239 are located on the outer wall of the porin transmembrane region. Mutating them to hydrophobic amino acids can enhance the stability of the porin pore. E177, R188, R222, and E238 are located on the inner wall of the pore structure and are charged amino acids on the inner wall of the porin barrel. Mutating their charge can facilitate the smooth passage of the DNA molecules being tested through the porin. R230 and K232 are located in the loop region (loop region) at the exit of the pore structure. Mutating them can promote the exit of the sequenced nucleic acid chain from the porin.
[0063] In a preferred embodiment, the porin monomer comprises a protein having any one of the amino acid sequences of SEQ ID NO: 2 to SEQ ID NO: 4.
[0064] SEQ ID NO: 2:
[0065]
[0066] SEQ ID NO: 3:
[0067]
[0068]
[0069] SEQ ID NO: 4:
[0070]
[0071] In a second typical embodiment of the present application, a protein construct is provided. The protein construct is composed of two or more porin monomers mentioned above, linked covalently or non-covalently.
[0072] In a third typical embodiment of the present application, a porin is provided. The porin is composed of 7-11 porin monomers mentioned above, preferably 9, linked by covalent or non-covalent means.
[0073] Porin monomers can spontaneously aggregate together through hydrogen bonds, ionic bonds, hydrophobic interactions, and other forces to form porins. Therefore, porin monomers expressed and purified exist as multimers, especially nonamers, under non-denaturing conditions, while denaturing the protein results in the existence of porin monomers.
[0074] In a preferred embodiment, the porin is composed of 9 porin monomers linked non-covalently, and the porin monomers include a protein having any one of the amino acid sequences of SEQ ID NO: 1 to SEQ ID NO: 4.
[0075] In a preferred embodiment, the pore diameter of the porin is 0.5 to 3 nm.
[0076] Within a certain range, the smaller the pore diameter of the nanopore protein, the higher its accuracy when used for sequencing. If the pore diameter of the nanopore protein is too large (more than one molecule may pass through the pore at a time), it is difficult to meet the needs of single-molecule sequencing. When the biological molecule to be tested passes through the overly large pore, the current signal generated may be missed or erroneous, resulting in low sequencing accuracy. In single-molecule sequencing, accurate sequencing results are obtained by sequencing the same molecule multiple times. Therefore, the higher the sequencing accuracy, the shorter the number of sequencing times and time required. The sequencing accuracy of pore proteins with pore diameters within the above range is relatively high. Using this nanopore protein with high sequencing accuracy for sequencing can greatly reduce sequencing time and reduce costs. This advantage is particularly evident in high-throughput sequencing.
[0077] In a fourth typical embodiment of the present application, a kit is provided, which includes the above-mentioned porin monomer, or the above-mentioned protein construct, or the above-mentioned porin.
[0078] To further improve operational convenience, in a preferred embodiment, the kit further comprises a membrane layer, which comprises a lipid layer or an artificial polymer membrane.
[0079] Preferably, the lipid layer comprises amphiphilic lipids; preferably, the amphiphilic lipids comprise a phospholipid bilayer; preferably, the lipid layer comprises a planar membrane layer or a liposome; preferably, the liposome comprises a multilamellar liposome or a unilamellar liposome; preferably, the lipid layer comprises a phospholipid bilayer composed of diphytylphosphatidylcholine.
[0080] Artificial polymer membranes include, but are not limited to, polysiloxanes, polyolefins, perfluoropolyethers, perfluoroalkyl polyethers, polystyrene, polyoxypropylene, polyvinyl acetate, polyoxybutylene, polyisoprene, polybutadiene, polyvinyl chloride, polyalkyl acrylates, polyalkyl methacrylates, polyacrylonitrile, polypropylene, PTHF, polymethacrylates, polyacrylates, polysulfones, polyethylene ethers, poly(propylene oxide) and copolymers thereof, alkyl-substituted C1-C6 alkyl acrylates and methacrylates, acrylamide, methacrylamide, (C1-C6 alkyl) acrylamide and methacrylamide, N,N-dialkyl-acrylamides, ethoxy acrylates and methacrylates, polyethylene glycol monomethacrylate and polyethylene glycol monomethyl ether methacrylate, hydroxy-substituted (C1-C6 alkyl) acrylamide and methacrylamide, hydroxy-substituted C1-C6 alkyl vinyl ether, sodium vinyl sulfonate, sodium styrene sulfonate, 2-acrylamide-2-methylpropanesulfonic acid, N-vinylpyrrole, N-vinyl- One or more of 2-pyrrolidone, 2-vinyloxazoline, 2-vinyl-4,4′-bisalkyloxazolinyl-5-one, 2,4-vinylpyridine, ethylenically unsaturated carboxylic acids having a total of 3 to 5 carbon atoms, amino(C1-C6 alkyl)-, mono(C1-C6 alkylamino)(C1-C6 alkyl)- and bis(C1-C6 alkylamino)(C1-C6 alkyl)-acrylates and methacrylates, allyl alcohol, 3-trimethylammonium 2-hydroxypropyl methacrylate chloride, dimethylaminoethyl methacrylate (DMAEMA), dimethylaminoethyl methacrylamide, glycerol methacrylate, N-(1,1-dimethyl-3-oxobutyl)acrylamide, cyclic imino ethers, vinyl ethers, cyclic ethers including epoxy derivatives, cyclic unsaturated ethers, N-substituted ethylenimines, β-lactones and β-lactams, ketene acetals, vinyl acetals or phosphoranes.
[0081] In a preferred embodiment, the kit further comprises at least one of the following: a sequencing buffer, a nuclease, a polymerase, a topoisomerase, a ligase, a helicase, and a cholesterol-linked single-stranded DNA.
[0082] Using the nanopore proteins, membrane layer, and sequencing buffer in the aforementioned kit, one or more nanopore proteins can be inserted into the membrane layer within the sequencing buffer to form a nanopore sensor. The nanopore assay buffer provides a neutral environment that maintains the stability of the nanopore proteins and the membrane layer, and the metal ions contained in the nanopore assay buffer provide good conductivity. A variety of membrane options are available, allowing nanopore proteins to be inserted into planar membrane layers, spherical liposomes, or lipid layers of varying compositions to form a nanopore sensor.
[0083] Cholesterol present on the single-stranded DNA molecules to be detected can bind to the lipid layer or artificial polymer membrane, facilitating nanopore capture of the sequencing library and reducing the amount of sample loaded. In actual use, the cholesterol in the above kit can be first bound to the molecules to be detected and then added to the space within the nanopore sensor for sequencing.
[0084] In a fifth typical embodiment of the present application, an isolated DNA molecule is provided, which has: a nucleotide sequence encoding the above-mentioned porin monomer; or a nucleotide sequence encoding the above-mentioned protein construct, or a nucleotide sequence encoding the above-mentioned porin.
[0085] In a preferred embodiment, the present invention relates to a DNA molecule that has more than 70%, preferably more than 80%, more preferably more than 90%, further preferably more than 99%, and most preferably more than 99% (for example, it can be 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, or even 99.9% or more) identity with the nucleotide sequence of SEQ ID NO: 5 and encodes a protein with the same function.
[0086] SEQ ID NO: 5:
[0087]
[0088] The above-mentioned DNA molecule can encode a nanopore protein monomer having the above-mentioned structure and function of this application. Based on the sequence of (a), the nucleotides are mutated and hybridized with the DNA molecule defined by (a) under strict conditions without frameshift mutation. If the mutation occurs in the nucleotide encoding the pore of the nanopore protein, it may result in the encoding of a nanopore protein with an altered pore, affecting the pore size of the nanopore protein and the properties of the amino acid residues on the inner wall of the pore. If the mutation occurs in the nucleotide encoding the non-pore part of the protein, it may affect the folding mode, three-dimensional structure and other properties of the encoded protein, thereby affecting the physicochemical properties and stability of the protein. (b) The nucleotide sequence that hybridizes with the DNA molecule defined under strict conditions includes a nucleotide sequence that is 80%, 85%, 90%, 95%, 98%, 99%, 99.9% or 100% complementary to the DNA molecule defined by (a).
[0089] As used herein, "isolated" means altered "by the hand of man" from its natural state, i.e., if it occurs in nature, it is altered and / or separated from its original environment. For example, a polynucleotide or polypeptide naturally present in a living organism is not "isolated," whereas the same polynucleotide or polypeptide separated from its coexisting components in its natural state is "isolated" (as the term is used herein).
[0090] In a sixth typical embodiment of the present application, a recombinant vector is provided, which comprises the above-mentioned DNA molecule.
[0091] The aforementioned DNA molecule, namely the nanopore protein expression gene, is inserted into a recombinant vector. The recombinant vector's ability to replicate itself is then used to replicate the nanopore protein expression gene in large quantities. "Recombinant" here refers to genetically engineered DNA produced by transplanting or splicing a gene from one species into the cells of a host organism of a different species. This DNA becomes part of the host's genetic structure and is replicated.
[0092] In a seventh typical embodiment of the present application, a host cell is provided, wherein the host cell is transformed with the above-mentioned recombinant vector.
[0093] The recombinant vector is transformed into a host cell, where it replicates, transcribes, and translates the nanopore protein expression gene on the recombinant vector, thereby producing large quantities of the nanopore protein. Common host cells include Escherichia coli, yeast, mammalian cells, insect cells, and other common host cells. The host cell folds the nanopore protein into the correct three-dimensional structure, resulting in a structurally and functionally normal nanopore protein.
[0094] In an eighth typical embodiment of the present application, a nanopore sensor is provided, which includes: a membrane layer; and a pore protein inserted into the membrane layer and forming a pore, and when a voltage is applied across the membrane layer, the pore generates current; wherein the pore protein includes the above-mentioned pore protein.
[0095] The nanopore sensor in this application specifically refers to a membrane layer with a nanopore protein inserted. This type of nanopore sensor, with a nanopore protein embedded in the membrane layer, can fix the orientation of the nanopore protein's pore. When an electric field is applied across the membrane layer, the pore diameter is perpendicular to the direction of the electric field. Ions on both sides of the membrane layer pass through the pore protein pore under the action of the electric field, generating an electric current.
[0096] In a preferred embodiment, the membrane layer comprises a lipid layer or an artificial polymer membrane; preferably, the lipid layer comprises amphiphilic lipids; preferably, the amphiphilic lipids comprise a phospholipid bilayer; preferably, the lipid layer comprises a planar membrane layer or a liposome; preferably, the liposome comprises a multilamellar liposome or a unilamellar liposome; preferably, the lipid layer comprises a phospholipid bilayer composed of diphytanoylphosphatidylcholine (DPhPC, 1,2-diphytanoyl-sn-glycero-3-phosphocholine).
[0097] There are many options for the selection of lipid layers. Nanopore proteins can be inserted into planar membrane layers or spherical liposomes, or into lipid layers formed of different components to form nanopore sensors.
[0098] In a preferred embodiment, when a voltage is applied across the membrane layer, the biomolecule to be detected passes through the pore in the nanopore sensor and shifts, and the pore generates a changing current; preferably, the biomolecule to be detected includes DNA, RNA or polypeptide; preferably, the DNA and / or RNA includes any one or more of the following modified bases: 5-methylcytosine (5mC), 6-methyladenine (m6A), 7-methylguanine (m7G), pseudouracil (pseudouridine, Ψ).
[0099] When an electric field is applied across the membrane, the biomolecules to be tested pass through the nanopore protein through the pores under the influence of the electric field. These biomolecules include biomacromolecules that carry genetic information, such as DNA, RNA, peptides, or proteins. These biomolecules can also be modified with groups. These groups include, but are not limited to, cholesterol, polyethylene glycols of varying degrees of polymerization, biotin, or fluorescent groups.
[0100] In a ninth typical embodiment of the present application, a nanopore sequencing device is provided, which includes the above-mentioned nanopore sensor.
[0101] In a preferred embodiment, the nanopore sequencing device includes: an electrolytic cell containing a sequencing buffer; a nanopore sensor, the nanopore sensor is located in the center of the electrolytic cell and divides the electrolytic cell and the sequencing buffer into a positive electrolyte region and a negative electrolyte region; a first electrode and a second electrode, the first electrode and the second electrode are respectively arranged in the positive electrolyte region and the negative electrolyte region, and the first electrode and the second electrode are connected to a signal processing chip; preferably, the first electrode and the second electrode include metal or composite electrode materials; preferably, the first electrode and the second electrode are different, namely silver and silver chloride, respectively; or the first electrode and the second electrode are the same, including gold, platinum, graphene or titanium nitride.
[0102] A nanopore sequencing device includes an electrolytic cell containing an electrolyte, a nanopore sensor, a first electrode, and a second electrode. The nanopore sensor is placed in the center of the electrolytic cell, which is then divided into a positive electrolyte region and a negative electrolyte region. Electrodes are positioned in each region, creating an electric field across the nanopore sensor. When the biomolecule to be tested passes through the nanopore protein on the membrane, a current amplitude is generated. This current amplitude is received and transmitted to a signal processing chip connected to the electrodes. Based on the current amplitude, the signal processing chip, i.e., the nanopore sequencing device containing the signal processing chip, can perform data analysis and determination of the sequence of the biomolecule to be tested.
[0103] In a tenth typical embodiment of the present application, a sequencing method is provided, which utilizes the above-mentioned porin, or the above-mentioned nanopore sensor, or the above-mentioned nanopore sequencing device to determine the sequence of the biological molecule to be tested by detecting and analyzing the electrical signal generated when the biological molecule to be tested passes through the pore of the porin.
[0104] In a preferred embodiment, the biomolecule to be detected includes modified or unmodified DNA, RNA or polypeptide; preferably, the electrical signal includes electric current.
[0105] In a preferred embodiment, the biomolecule to be detected is a target nucleic acid sequence, and the sequencing method comprises: (a) contacting the nucleic acid sequence with the above-mentioned pore protein and nucleic acid binding protein, so that the nucleic acid binding protein controls the movement speed of the target nucleic acid sequence through the pore of the pore protein, wherein the nucleic acid binding protein is selected from any one or more of a nuclease, a polymerase, a topoisomerase, a ligase, a helicase, or a single-stranded binding protein; (b) when a voltage is applied across the pore, when the nucleic acid sequence moves through the pore, measuring the electrical signal passing through the pore, wherein different types of nucleotides generate different electrical signals when passing through the pore, thereby determining the sequence information of the nucleic acid based on the electrical signal.
[0106] In the eleventh typical embodiment of the present application, a porin monomer, or the above-mentioned porin, or the above-mentioned kit, or the above-mentioned DNA molecule, or the above-mentioned recombinant vector, or the above-mentioned host cell, or the above-mentioned nanopore sensor, or the above-mentioned nanopore sequencing device, or the above-mentioned sequencing method is provided for use in biological small molecule detection, nucleic acid sequencing or polypeptide sequencing.
[0107] "Small molecules" in this application include but are not limited to small molecules such as nucleotides, amino acids, polysaccharides or vitamins.
[0108] Nanopore sequencing offers a major advantage over traditional sequencing: it avoids errors that can negatively impact accuracy, enabling extremely long read lengths. This can compensate for gaps that inevitably occur when assembling short fragments in traditional sequencing, allowing for the identification of long deletions, duplications, inversions, and translocations within chromosomes. This allows for full-length coverage of transcriptomes, typically several kilobases in length, providing a novel solution for scientific research on genome assembly, structural variation, and alternative splicing.
[0109] Since nanopore sequencing does not require PCR amplification, the original base modification information on the nucleic acid molecule to be tested can be retained, and the type, location and abundance of the modified base can be directly sequenced at one time. Therefore, the nanopore protein of the present application can also detect several nucleic acid molecules with DNA / RNA modified bases: including 5-methylcytosine (5mC), 6-methyladenine (m6A), 7-methylguanine (m7G), pseudouracil (pseudouridine, Ψ), etc. By performing specific model training and algorithm development on various modified bases, nanopore sequencing can complete the identification and positioning of more modified bases, thereby constructing a more complete genome / transcriptome modification map.
[0110] Furthermore, from a clinical application perspective, nanopore sequencing's long read lengths, high portability, fast sequencing speeds, and real-time readout make it highly suitable for critical epidemic monitoring and rapid pathogen detection (e.g., in large-scale epidemics such as Zika, Ebola, dengue, and the novel coronavirus), ensuring timely detection. In addition to viruses, nanopore sequencing can also be used for the rapid detection of other pathogens, such as bacteria and fungi.
[0111] Based on the common composition of proteins and nucleic acid molecules, nanopore sequencing platforms also have enormous application potential in the field of protein sequencing. For example, based on the research currently underway, by using protein unfolding enzymes as rate-control tools, characteristic protein signals have been successfully observed, and preliminary identification of protein types and modification states has been achieved, validating the feasibility of nanopore protein sequencing. In future developments, by further optimizing the rate-control system and developing suitable nanopore proteins and signal analysis algorithms, fingerprinting and even sequence identification of proteins at the single-molecule level may be achieved.
[0112] In addition to its applications in sequencing, the nanopore platform can also serve as a foundational detection platform, combined with sensing techniques to enable metabolomics analysis of various small and large molecules. By integrating genomics, proteomics, and metabolomics, the nanopore platform could ultimately develop into a universal measurement platform that meets the needs of comprehensive omics analysis, providing a powerful research tool for a deeper understanding of the laws of life and the mechanisms of disease.
[0113] The beneficial effects of the present application will be further explained in detail below in conjunction with specific examples. However, it will be understood by those skilled in the art that the following examples are merely illustrative of the present invention and should not be construed as limiting the scope of the present invention. Reagents or instruments used without manufacturer's indication are conventional products available on the market. The experimental methods used are conventional methods unless otherwise specified.
[0114] Example 1 Predicted structure of the Alphafold2-Multimer of wild-type BCP35
[0115] Porin BCP35 was obtained by non-covalently polymerizing nine porin monomers (SEQ ID NO: 1) into a nonamer. Alphafold2-Multimer was used to predict the structure of BCP35. The predicted results are shown in Figures 1, 2, 3, 4, 5, 6, and 7. Figure 1 shows a side view of the predicted BCP35 structure, and Figure 2 shows a top view of the predicted BCP35 structure.
[0116] Because the amino acid composition of the sensor region (gating region) plays a crucial role in generating current signals, the first mutants of BCP35 to be investigated were those targeting the sensor region. Figures 3 and 4 show the side chain structures of key amino acids in the predicted BCP35 sensor region, revealing the three amino acids in the side chains: P91, S98, and N99.
[0117] Because nucleic acids are negatively charged, positively charged amino acids at the porin entrance can increase the library capture efficiency of nanopore sequencing; conversely, negatively charged amino acids can reduce the library capture rate. Furthermore, the strength of the binding between the library and the porin also affects the sequencing speed and sequencing current signal. Figure 5 shows some important amino acids in the BCP35 entrance region: E125, R126, K127, D131, K134, R136, R140, K152, and D155.
[0118] At the same time, some amino acids on the inner wall and outlet of the barrel have an impact on the smooth passage of the analyte through and out of the nanopore. Figure 6 shows some amino acids on the inner wall and outlet of the barrel of BCP35 that may affect the smooth passage of the analyte through the pore, especially charged amino acids, which are E177, R188, R222, R230, K232, and E238.
[0119] Because the amino acids in the porin transmembrane region that face the membrane tend to be hydrophobic, charged and polar amino acids may affect pore insertion efficiency. Figure 7 shows several charged and polar amino acids in the BCP35 transmembrane region that face the membrane: T174, H183, K185, K219, S223, and R239.
[0120] Example 2 Construction of expression vectors for porin BCP35 and its mutants
[0121] Using an in-fusion method, the porin monomer-encoding DNA sequence (SEQ ID NO: 5) was digested with NdeI and XhoI, and then inserted into the multiple cloning region of the pET24a vector. A StrepII amino acid was added to the C-terminus of the porin monomer amino acid sequence (SEQ ID NO: 1) as a purification tag, with kanamycin as the screening tag. The constructed vector was named pET24a-BCP35. Using an Agilent site-directed mutagenesis kit and the porin BCP35 expression vector as a template, corresponding mutants were constructed. In this example, three mutants, P91A, S98N, and N99S, were constructed. The constructed mutant vectors were named pET24a-BCP35-P91A, pET24a-BCP35-S98N, and pET24a-BCP35-N99S, respectively.
[0122] Example 3 Cultivation and induction of porin monomer strains
[0123] The constructed porin monomer or mutant expression plasmids were independently transformed into the E. coli BL21(DE3) expression strain. The bacterial suspension was evenly spread on plates containing 50 μg / mL kanamycin and cultured overnight at 37°C. The next day, a single colony was picked and incubated in 5 mL of LB medium containing 50 μg / mL kanamycin at 37°C, 200 rpm, and incubated overnight. The resulting bacterial suspension was inoculated at a ratio of 1:100 into 50 mL of LB medium containing 50 μg / mL kanamycin and incubated at 37°C, 200 rpm, for 4 hours. The expanded culture was inoculated at a ratio of 1:100 into 2 L of LB medium containing 50 μg / mL kanamycin and incubated at 37°C, 200 rpm. When the OD600 value reached approximately 0.6-0.8, IPTG was added to a final concentration of 0.5 mM and the culture was incubated at 16°C, 200 rpm, for approximately 16-18 hours. The bacterial solution was collected by centrifugation at 8000 rpm and the bacteria were frozen at -20°C until use.
[0124] Example 4 Extraction and purification of recombinant porin monomers
[0125] (1) Buffer preparation:
[0126] Buffer A: 20mM Tris-HCl, 250mM NaCl, 1% DDM, pH 8.0.
[0127] Buffer B: 20mM Tris-HCl, 250mM NaCl, 0.05% DDM, pH 8.0.
[0128] Buffer C: 20 mM Tris-HCl, 250 mM NaCl, 0.05% DDM, 5 mM desthiobiotin, pH 8.0.
[0129] (2) Purification step:
[0130] Resuspend the cells thoroughly in 10 mL of Buffer A per 1 g of cells. Ultrasonicate the cells until the solution is clear. Rotate the solution overnight at 4°C. Centrifuge the next day at 18,000 rpm for 1 hour at 4°C. Remove the supernatant, filter through a 0.22 μm filter, and store at 4°C until ready for use.
[0131] On an AKTA pure chromatography instrument, a Strep-Tactin beads (IBA Lifesciences) column was equilibrated with Buffer A for 5 column volumes (CV) and loaded at 2 mL / min. After loading, the column was washed with Buffer B for 20 CV and eluted with Buffer C to collect the target protein.
[0132] The resulting protein was concentrated to 1 mL and passed through a Superdex 6increase 10 / 300GL (Cytiva) column equilibrated with buffer B to collect the target protein, which was then stored at -80°C. The purified target protein was subjected to SDS-PAGE electrophoresis, and the results of the wild-type protein monomer are shown in Figure 8. Lanes 2 and 3 and lanes 4 and 5 are the electrophoresis bands of the uncooked (porin nonamer BCP35) and the 95°C boiled (porin monomer after denaturation), respectively. The results show that the target protein is in a polymeric state when uncooked and in a monomeric state after boiling. The SDS-PAGE results of the mutant are consistent with those of the wild-type protein and are not presented in detail here.
[0133] Example 5 Library Construction
[0134] The sense and antisense strands of two partially complementary DNA chains (SEQ ID NO: 8) were ligated with the double-stranded target fragment pUC57 (SEQ ID NO: 9) at room temperature using T4 DNA ligase and purified to prepare a sequencing library. This library was then incubated with the helicase BCH105 (SEQ ID NO: 10) at 25°C for 1 hour (molar ratio of 1:8) to form a sequencing library containing the BCH105 motor protein, as shown in Figure 9A. During sequencing, this library was further able to bind to a single-stranded DNA containing cholesterol (SEQ ID NO: 11, with cholesterol attached to the 5' end of the DNA), forming the structure shown in Figure 9B.
[0135] Linker sequence sense chain: S1-(iSp18)4-S2.
[0136] The sequence of S1 is shown in SEQ ID NO: 6, the sequence of S2 is shown in SEQ ID NO: 7, and the structure of iSp18 is shown in FIG10 .
[0137] SEQ ID NO: 6: tttttttttttttttttttttttttttttttttttttt.
[0138] SEQ ID NO: 7: ggttgtttctgttggtgctgatattgct.
[0139] SEQ ID NO: 8 (linker sequence antisense strand):
[0140]
[0141] SEQ ID NO: 9:
[0142]
[0143]
[0144] SEQ ID NO: 10:
[0145]
[0146] Example 6 Construction of nanopore biosensors using porin BCP35 and its mutants
[0147] A patch clamp amplifier is used to collect current signals. Ag / AgCl electrodes are immersed in sequencing buffer and the electrodes are located in the cis and trans regions of the electrolytic cell, respectively. After diluting the porin (i.e., the porin BCP35 purified in Example 4) by a certain multiple using 1xPBS buffer, a single nanoporin BCP35 is inserted into a phospholipid bilayer composed of diacylphosphatidylcholine (DPhPC, 1,2-diphytanoyl-sn-glycero-3-phosphocholine) under the action of an external electric field force to form a nanopore biosensor. The dilution multiple is based on whether the porin is embedded in the membrane (i.e., embedded in the pore). Generally, a protein concentration of 0.1 mg / ml is used and diluted 100 times, 50 times, or other multiples with PBS for trial. If a certain dilution concentration fails to embed in the pore, it is necessary to reduce the dilution multiple and continue trying until the nanoporin is successfully embedded in the membrane layer. Apply an external voltage to obtain the current amplitude value of a single porin. Figure 11 shows the biosensing current generated by the nanoporin BCP35 pore when voltages of 0.02V, 0.04V, 0.10V, 0.14V, and 0.18V were applied. As can be seen, the BCP35 pore current is stable and has minimal noise at different voltages. Similarly, BCP35 mutants (P91A, S98N, and N99S) also generate stable biosensing currents.
[0148] Example 7 Use of Porin BCP35 and Its Mutants for DNA Sequencing
[0149] The sequencing library containing the pUC57 sequence prepared in Example 5 and cholesterol-linked single-stranded DNA (SEQ ID NO: 11, with cholesterol attached to the 5' end of the DNA) were mixed with sequencing buffer and added to the nanopore biosensor. After applying an external voltage of 0.14V or 0.18V, DNA was observed to be captured by the nanopore, generating a characteristic blockade current amplitude. The current amplitude varied as the DNA moved through the nanopore. Different DNA sequences produced different blockade current amplitudes. Cholesterol-linked single-stranded DNA can bind to the phospholipid bilayer, facilitating nanopore capture of the sequencing library and reducing the amount of sequencing library required. Figure 12 shows the current changes generated when the library DNA passed through the nanopore protein BCP35 under an applied voltage of 0.14V. It can be seen that the library DNA can pass through the wild-type BCP35 porin, generating a current that fluctuates with the passage of different nucleotides through the pore. Similarly, current changes were observed when library DNA passed through porin BCP35 mutants (P91A, S98N, and N99S), demonstrating that porin BCP35 and its mutants can be used for nanopore sequencing.
[0150] Single-stranded DNA sequence with cholesterol (SEQ ID NO: 11):
[0151] cholestero-ttgaccgctcgcctc.
[0152] Sequencing buffer: 0.47 M KCl, 25 mM HEPES, 1 mM EDTA, 5 mM ATP, 25 mM MgCl2, pH 7.6.
[0153] From the above description, it can be seen that the above-mentioned embodiments of the present invention achieve the following technical effects: the present invention has discovered a new nanopore protein monomer, which is polymerized into the porin BCP58. The porin and its mutants have good stability and can meet the requirements of single-molecule nanopore sequencing. The above-mentioned nanopore protein and its mutants can be used to form nanopore sensors and further nanopore sequencing devices, which can realize the detection of small molecules such as nucleotides, amino acids, sugars, vitamins, etc., and can also be used for sequencing samples such as DNA, RNA, and polypeptides.
[0154] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A porin monomer, characterized in that The porin monomer comprises: (a) a protein having the amino acid sequence shown in SEQ ID NO: 1; or (b) a protein mutant, wherein the amino acid sequence of the protein mutant is substituted, deleted and / or one or more amino acids are added at at least one of the following positions of SEQ ID NO: 1: 91, 98, 99, 125, 126, 127, 131, 134, 136, 140, 152, 155, 174, 177, 183, 185, 188, 219, 222, 223, 230, 232, 238, 239, and the protein mutant has the function of forming a pore structure through polymerization; or (c) has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity with the protein described in (a) or (b), and has the function of forming a pore structure through polymerization.
2. The porin monomer according to claim 1, characterized in that In b), the types of substituted amino acids are each independently selected from the following: P91 mutated to P91G, P91A, or P91T; S98 mutation to S98G, S98A, S98T, S98N, or S98Q; N99 mutation is N99G, N99A, N99S, N99T, or N99Q; E125 mutation is E125K, E125R, E125G, E125A, E125S, E125T, E125N, or E125Q; R126 mutation to R126K, R126G, R126A, R126S, R126T, R126N, or R126Q; K127 mutation is K127R, K127G, K127A, K127S, K127T, K127N, or K127Q; D131 mutation is D131K, D131R, D131G, D131A, D131S, D131T, D131N or D131Q; K134 mutation is K134R, K134G, K134A, K134S, K134T, K134N, or K134Q; R136 mutation to R136K, R136G, R136A, R136S, R136T, R136N, or R136Q; R140 mutation is R140K, R140G, R140A, R140S, R140T, R140N, or R140Q; K152 mutation is K152R, K152G, K152A, K152S, K152T, K152N, or K152Q; D155 is mutated to D155R, D155K, D155G, D155A, D155S, D155T, D155N, or D155Q; T174 mutation is T174A, T174G, T174V, T174L, T174I, T174Y, T174F, or T174W; E177 mutation is E177A, E177G, E177S, E177T, E177N, or E177Q; H183 mutated to H183A, H183G, H183V, H183L, H183I, H183Y, H183F, or H183W; K185 mutation is K185A, K185G, K185V, K185L, K185I, K185Y, K185F, or K185W; R188 mutation to R188A, R188G, R188S, R188T, R188N, or R188Q; K219 mutation is K219A, K219G, K219V, K219L, K219I, K219Y, K219F, or K219W; R222 mutated to R222A, R222G, R222S, R222T, R222N or R222Q; S223 mutation to S223A, S223G, S223V, S223L, S223I, S223Y, S223F, or S223W; R230 mutated to R230A, R230G, R230S, R230T, R230N, or R230Q; K232 mutation is K232A, K232G, K232S, K232T, K232N, or K232Q; E238 mutation to E238A, E238G, E238S, E238T, E238N, or E238Q; R239 is mutated to R239A, R239G, R239V, R239L, R239I, R239Y, R239F or R239W.
3. The porin monomer according to claim 1 or 2, characterized in that The porin monomer includes a protein having any one of the amino acid sequences of SEQ ID NO: 2 to SEQ ID NO:
4.
4. A protein construct, characterized in that The protein construct is formed by covalently or non-covalently linking two or more porin monomers according to any one of claims 1 to 3.
5. A porin, characterized in that The porin is composed of 7 to 11 porin monomers according to any one of claims 1 to 3, preferably 9 porin monomers, connected covalently or non-covalently.
6. The porin according to claim 5, characterized in that The porin is composed of 9 porin monomers connected non-covalently, and the porin monomers include proteins having any amino acid sequence in SEQ ID NO: 1-SEQ ID NO:
4.
7. The porin according to claim 5 or 6, characterized in that The pore diameter of the porin is 0.5 to 3 nm.
8. A kit, characterized in that: The kit comprises the porin monomer according to any one of claims 1 to 3, or the protein construct according to claim 4, or the porin according to any one of claims 5 to 7.
9. The kit according to claim 8, characterized in that The kit further comprises a membrane layer, which comprises a lipid layer or an artificial polymer membrane.
10. The kit according to claim 9, characterized in that The kit further comprises at least one of the following: a sequencing buffer, a nuclease, a polymerase, a topoisomerase, a ligase, a helicase, and a single-stranded DNA linked to cholesterol.
11. An isolated DNA molecule, characterized in that The DNA molecule has: A nucleotide sequence encoding a porin monomer according to any one of claims 1 to 3; or a nucleotide sequence encoding a protein construct according to claim 4, or a nucleotide sequence encoding a porin according to any one of claims 5 to 7.
12. The DNA molecule according to claim 11, characterized in that A DNA molecule that has 70% or more, preferably 80% or more, more preferably 90% or more, further preferably 99% or more, and most preferably 99% or more identity with the nucleotide sequence of SEQ ID NO: 5 and encodes a protein having the same function.
13. A recombinant vector, characterized in that: The recombinant vector comprises the DNA molecule according to claim 11 or 12.
14. A host cell, characterized in that The host cell is transformed with the recombinant vector according to claim 13.
15. A nanopore sensor, characterized in that: The nanopore sensor comprises: film layer; and a porin protein inserted into the membrane layer and forming a pore that generates an electric current when a voltage is applied across the membrane layer; Wherein, the porin comprises the porin according to any one of claims 5 to 7.
16. The nanopore sensor according to claim 15, characterized in that The membrane layer includes a lipid layer or an artificial polymer membrane; Preferably, the lipid layer comprises amphiphilic lipids; Preferably, the amphiphilic lipid comprises a phospholipid bilayer; Preferably, the lipid layer comprises a planar membrane layer or a liposome; Preferably, the liposomes comprise multilamellar liposomes or unilamellar liposomes; Preferably, the lipid layer comprises a phospholipid bilayer composed of diphytylphosphatidylcholine.
17. The nanopore sensor according to claim 15, characterized in that When a voltage is applied across the membrane layer, the biomolecule to be detected passes through the pore in the nanopore sensor and shifts, and the pore generates a changing current; Preferably, the biomolecule to be detected includes DNA, RNA or polypeptide; Preferably, the DNA and / or RNA comprises any one or more of the following modified bases: 5-methylcytosine, 6-methyladenine, 7-methylguanine, pseudouracil.
18. A nanopore sequencing device, characterized in that: The nanopore sequencing device comprises the nanopore sensor according to any one of claims 15 to 17.
19. The nanopore sequencing device according to claim 18, characterized in that The nanopore sequencing device comprises: an electrolytic cell containing a sequencing buffer; A nanopore sensor, wherein the nanopore sensor is located in the center of the electrolytic cell and divides the electrolytic cell and the sequencing buffer into a positive electrode electrolyte region and a negative electrode electrolyte region; A first electrode and a second electrode, wherein the first electrode and the second electrode are respectively arranged in the positive electrode electrolyte region and the negative electrode electrolyte region, and the first electrode and the second electrode are connected to a signal processing chip; Preferably, the first electrode and the second electrode comprise metal or composite electrode materials; Preferably, the first electrode and the second electrode are different, namely silver and silver chloride, respectively; or the first electrode and the second electrode are the same, including gold, platinum, graphene or titanium nitride.
20. A sequencing method, characterized in that: The sequencing method utilizes the pore protein described in any one of claims 5 to 7, or the nanopore sensor described in any one of claims 15 to 17, or the nanopore sequencing device described in claim 18 or 19 to determine the sequence of the biological molecule to be detected by detecting and analyzing the electrical signal generated when the biological molecule to be detected passes through the pore of the pore protein.
21. The sequencing method according to claim 20, characterized in that: The biomolecule to be detected includes modified or unmodified DNA, RNA or polypeptide; Preferably, the electrical signal comprises an electric current.
22. The sequencing method according to claim 20, characterized in that The biomolecule to be detected is a target nucleic acid sequence, and the sequencing method comprises: (a) contacting the nucleic acid sequence with a porin and a nucleic acid binding protein according to any one of claims 5 to 7, so that the nucleic acid binding protein controls the speed at which the target nucleic acid sequence moves through the pore of the porin, wherein the nucleic acid binding protein is selected from any one or more of a nuclease, a polymerase, a topoisomerase, a ligase, a helicase or a single-stranded binding protein; (b) When a voltage is applied across the pore, when the nucleic acid sequence moves through the pore, an electrical signal passing through the pore is measured, wherein different types of nucleotides generate different electrical signals when passing through the pore, thereby determining the sequence information of the nucleic acid based on the electrical signal.
23. Use of the porin monomer of any one of claims 1 to 3, or the porin of any one of claims 5 to 7, or the kit of any one of claims 8 to 10, or the DNA molecule of claim 11 or 12, or the recombinant vector of claim 13, or the host cell of claim 14, or the nanopore sensor of any one of claims 15 to 17, or the nanopore sequencing device of claim 18 or 19, or the sequencing method of any one of claims 20 to 22 in biological small molecule detection, nucleic acid sequencing or polypeptide sequencing.