Nanopore protein complex and use thereof
By designing nanoporin complexes containing multiple porin monomers and accessory proteins to form continuous channels, the problem of low sequencing accuracy in nanopore sequencers was solved, and higher precision multinucleotide sequencing was achieved.
Patent Information
- Application Number
- PCT/CN2024/119171
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-14
- Publication Date
- 2026-03-19
AI Technical Summary
Existing nanopore sequencers are insufficient in sequencing accuracy and the ability to read a large number of repetitive sequences in the genome, failing to meet the sequencing needs of high-precision applications such as disease diagnosis.
Design a nanoporin complex comprising multiple porin monomers and accessory proteins to form a continuous channel with two contraction zones, which is covalently or non-covalently linked. The N-terminus of the accessory protein is located within the nanoporous structure to enhance base recognition and current difference.
It improves the sequencing accuracy and resolution of polynucleotide homopolymers, enhances the ability to identify current differences among multiple nucleotides, and achieves higher-precision sequencing.
Smart Images

Figure CN2024119171_19032026_PF_FP_ABST
Abstract
Description
Nanopore protein complex and application thereof TECHNICAL FIELD
[0001] The present application relates to the technical field of nanopore sequencing, in particular to a nanopore protein complex and application thereof. BACKGROUND
[0002] The principle of nanopore sequencing technology is based on electrical signal changes. A transmembrane protein with a nanometer diameter is inserted into a biomimetic membrane, two electrodes are placed on both sides of the membrane, and a stable current through the nanopore is generated after power is applied between the two electrodes. When the measured object passes through the nanopore, it will hinder the flow of ions and cause characteristic changes in ion flow, that is, the size of the current changes (block current). Because different measured objects have different effects on the current, by detecting the current fluctuation signal of the nanopore in real time and converting the current signal with the help of machine learning analysis, real-time analysis of the composition of the measured object can be realized.
[0003] At present, the nanopore sequencers that have been released include MinION, GridION and PromethION of Oxford Nanopore Technologies in the United Kingdom and QNome-3841 nanopore sequencer of Ziqian Technology. These nanopore sequencers still have a large room for improvement in terms of sequencing accuracy, throughput, chip stability and applicable scenarios. Compared with the second-generation sequencing, the biggest deficiency of nanopore sequencing is that its sequencing accuracy needs to be improved, it cannot read a large number of repetitive sequences in the genome, and it cannot meet the sequencing needs of high-precision application scenarios such as disease diagnosis, so even with the advantage of long read length, it still limits its application in actual scientific research and industry.
[0004] It has been reported that adding auxiliary protein CsgF or truncating the N-terminal of CsgG, a single sensor nanopore protein used in commercial nanopore sequencers, can significantly improve signal step formation, especially for detecting samples containing homopolymer fragments, providing more detailed information, which provides theoretical support for creating nanopore sequencers with better performance, but the actual application effect needs to be verified. Related patent documents such as US20220056517A1 and CN113195736A.
[0005] Nanopore sequencers are high-integration single-molecule sequencers that apply multiple disciplines, and elements related to accuracy include nanopore proteins, adaptation of biomimetic membranes, performance of semiconductor materials, and optimization of algorithms. Among them, the pore protein that directly interacts with the measured object determines the limit of sequencing accuracy, so developing nanopore proteins with stable sequencing and high recognition accuracy is a top priority.
[0006] SUMMARY
[0007] The main purpose of the present application is to provide a nanopore protein complex and application thereof, so as to solve the problem of low accuracy of nanopore sequencing in the prior art.
[0008] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a nanopore protein complex is provided, which comprises: a nanopore protein and a plurality of auxiliary proteins, wherein the nanopore protein comprises a plurality of pore protein monomers and a nanopore channel structure formed by polymerization of the plurality of pore protein monomers, the N-terminus of each auxiliary protein is located within the nanopore channel structure, the plurality of auxiliary proteins are connected one-to-one with the plurality of pore protein monomers, and together with the nanopore channel structure form a continuous channel; wherein, along the direction of the to-be-measured substance passing through the continuous channel, the continuous channel comprises a first constriction region and a second constriction region connected in sequence, the first constriction region is formed by at least part of the nanopore protein, and the second constriction region is formed by part or all of the auxiliary proteins; the pore protein monomers are selected from any one or more of the following: (a) a protein having the amino acid sequence shown in SEQ ID NO: 1; or (b) a protein having substitution, deletion and / or addition of one or several amino acids at at least one of the following positions of SEQ ID NO: 1: 77, 81, 82, 176, 210, 214, 232, 66, 69, 70, 74, 109, 110, 113, 117, 118, 119, 120, 123, 127, 128, 168, 211, 221, 224, 227, 229, and having the function of being polymerized to form a nanopore channel structure; or (c) a protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity with the protein in (a) or (b), and having the function of being polymerized to form a nanopore channel structure.
[0009] Further, the auxiliary protein is selected from any one of the following: i) a polypeptide fragment having the amino acid sequence of any one of SEQ ID NOs: 3-8; ii) a polypeptide fragment having a signal peptide sequence added to the N-terminus of the amino acid sequence of any one of SEQ ID NOs: 3-8; iii) a protein having substitution and / or deletion and / or addition of one or several amino acids at at least one of the following positions of the amino acid sequence of any one of SEQ ID NOs: 3-8: L46, T49, N53, K62, D63, or P64; or iv) a polypeptide fragment having at least 50% identity with the amino acid sequence of any one of the polypeptide fragments in i), ii), or iii).
[0010] Further, in iv), the polypeptide fragment having at least 50% identity with the amino acid sequence of any one of the polypeptide fragments in i), ii), or iii) has the amino acid sequence shown in SEQ ID NO: 2.
[0011] Further, when the accessory protein is selected from iii), L46 is substituted with S, N, T, V or Q; T49 is substituted with N, S, V, L, A or Q; N53 is substituted with S, V, L, A, Q or T; K62 is substituted with N, S, V, L, A, I, Q, T, P, Y or W; D63 is substituted with N, S, V, L, A, I, Q, T, P, Y or W; and P64 is substituted with N, S, V, L, A, I, Q, T, Y or W.
[0012] Further, in (b), the types of the substituted amino acids are each independently selected from the following: S77 is mutated to S77G, S77A, or S77T; S81 is mutated to S81G, S81A, S81T, S81N, or S81Q; F82 is mutated to F82G, F82A, F82S, F82T, F82N, or F82Q; E176 is mutated to E176A, E176G, E176V, E176L, E176I, E176Y, E176F, or E176W; K210 is mutated to K210A, K210G, K210V, K210L, K210I, K210Y, K210F, or K210W; S214 is mutated to S214A, S214G, S214V, S214L, S214I, S214Y, S214F, or S214W; T232 is mutated to T232A, T232G, T232V, T232L, T232I, T232Y, T232F, or T232W; K66 is mutated to K66N, K66A, K66G, K66S, K66T, or K66Q; D69 is mutated to D69K, D69R, D69N, D69A, D69G, D69S, D69T, or D69Q; Q70 is mutated to Q70K, Q70R, Q70D, Q70E, Q70N, Q70A, Q70G, Q70S, Q70T, or Q70Q; Y74 is mutated to Y74N, Y74A, Y74G, Y74S, Y74T, or Y74Q; R109 is mutated to R109N, R109A, R109G, R109S, R109T, or R109Q; E110 is mutated to E110K, E110R, E110N, E110A, E110G, E110S, E110T, or E110Q; Q113 is mutated to Q113K, Q113R, Q113N, Q113A, Q113G, Q113S, or Q113T; T117 is mutated to T117K, T117R, T117N, T117A, T117G, T117S, or T117Q; E118 is mutated to E118K, E118R, E118N, E118A, E118G, E118S, E118T, or E118Q; R119 is mutated to R119N, R119A, R119G, R119S, R119T, or R119Q; K120 is mutated to K120N, K120A, K120G, K120S, K120T, or K120Q; R123 is mutated to R123N, R123A, R123G, R123S, R123T, or R123Q; K127 is mutated to K127N, K127A, K127G, K127S, K127T, or K127Q; K128 is mutated to K128N, K128A, K128G, K128S, K128T, or K128Q;R168 is mutated to R168N, R168Q, R168S, R168T, R168A or R168G; E211 is mutated to E211N, E211Q, E211S, E211T, E211A or E211G; E221 is mutated to E221N, E221Q, E221S, E221T, E221A or E221G; E224 is mutated to E224N, E224Q, E224S, E224T, E224A or E224G; E227 is mutated to E227N, E227Q, E227S, E227T, E227A or E227G; E229 is mutated to E229N, E229Q, E229S, E229T, E229A or E229G.
[0013] Further, in (b), the types of the substituted amino acids are each independently selected from the following: S77A, F82Q, S81N, F82A, K120N, K127S, E176I; preferably, the amino acid sequence of the pore protein monomer is selected from the following: an amino acid sequence that is substituted with F82Q based on SEQ ID NO: 1; or an amino acid sequence that is substituted with S77A+S81N+F82Q+K127S based on SEQ ID NO: 1.
[0014] Further, the pore protein monomer is selected from a protein having any one of the amino acid sequences in SEQ ID NO: 24-SEQ ID NO: 28.
[0015] Further, the plurality of auxiliary proteins are covalently connected to the plurality of pore protein monomers one by one through covalent bonds; or the plurality of auxiliary proteins are connected to the plurality of pore protein monomers one by one through non-covalent interactions.
[0016] Further, the plurality of auxiliary proteins and the plurality of pore protein monomers comprise a mutation to cysteine at any one of the following sites or an introduction of a non-natural amino acid at any one of the following sites to achieve covalent connection: the amino acid sites of the pore protein monomer include D181, R179, N159, E227, Q213, R168, F219, Y217 or L170 based on SEQ ID NO: 1; the amino acid sites of the auxiliary protein include S30, E31, L32, V33, V37, N38, S40, A55, Q58 or Q59 based on SEQ ID NO: 2.
[0017] Further, the covalent connection comprises any one or more of the following: spontaneous cross-linking, cross-linking agent connection, protein fusion connection, polypeptide molecule connection or chemical small molecule connection.
[0018] Further, the first constriction zone has a pore diameter of 13-21 angstroms, and the second constriction zone has a pore diameter of 15-25 angstroms.
[0019] According to a second aspect of the present application, there is provided a nanopore sensor comprising a membrane and any one of the nanopore protein complexes described above, wherein the nanopore protein in the nanopore protein complex is embedded in the membrane, and the nanopore channel structure of the nanopore protein and the plurality of auxiliary proteins together form a continuous channel across the membrane; and when a voltage is applied across the membrane, the continuous channel generates an electrical signal.
[0020] Further, the membrane comprises a layer of amphiphilic molecules; preferably, the membrane is a phospholipid bilayer composed of dipalmitoyl phosphatidylcholine.
[0021] Further, when a voltage is applied across the membrane, a biological molecule passes through the continuous channel in the nanopore sensor and is translocated, and the continuous channel generates a changed electrical current signal; preferably, the biological molecule comprises a polynucleotide, a polypeptide or a polysaccharide; preferably, the biological molecule is selected from polynucleotides containing homopolymers.
[0022] According to a third aspect of the present application, there is provided a kit comprising any one of the nanopore protein complexes described above or any one of the nanopore sensors described above, and the kit further comprises one or more optional components selected from the group consisting of a membrane, a sequencing buffer, a nuclease, a polymerase, a topoisomerase, a ligase, a helicase and a single-stranded DNA linked to a cholesterol.
[0023] According to a fourth aspect of the present application, there is provided a nanopore sequencing system comprising any one of the nanopore sensors described above, and the nanopore sequencing system further comprises: an electrically conductive solution, a positive electrode and a negative electrode for providing a voltage across the membrane, and a measuring device for measuring the electrical signal passing through the continuous channel; wherein the nanopore sensor is located in the electrically conductive solution and divides the electrically conductive solution into a first chamber and a second chamber, the negative electrode is in the first chamber, and the positive electrode is in the second chamber.
[0024] Further, the positive electrode and the negative electrode are connected to a signal processing chip to measure the electrical signal passing through the continuous channel.
[0025] Further, the positive electrode and the negative electrode comprise a metal electrode material or a composite electrode material.
[0026] According to a fifth aspect of the present application, there is provided a method for identifying or characterizing a biological molecule, the method comprising: processing the biological molecule using any one of the nanopore protein complexes described above, or any one of the nanopore sensors described above, or any one of the nanopore sequencing systems described above, and identifying or characterizing the biological molecule by measuring and analyzing the electrical signal generated when the biological molecule passes through the continuous channel of the nanopore protein complex.
[0027] Further, the method comprises: contacting any one of the nanopore sequencing systems with a biomolecule, applying a voltage across the membrane, such that the biomolecule enters the continuous channel and moves relative to the continuous channel; and measuring an electrical signal generated when the biomolecule moves relative to the continuous channel, thereby identifying or characterizing the biomolecule.
[0028] Further, the biomolecule comprises a polynucleotide, a polypeptide, or a polysaccharide; preferably, the biomolecule is selected from a polynucleotide containing a homopolymer; preferably, the electrical signal comprises an electric current.
[0029] Further, identifying the biomolecule comprises identifying whether the biomolecule is present; characterizing the biomolecule comprises determining a composition of the biomolecule; preferably, determining the composition of the biomolecule comprises: determining a nucleotide sequence of the polynucleotide, an amino acid sequence of the polypeptide, or a monosaccharide arrangement order of the polysaccharide.
[0030] Further, the biomolecule is a polynucleotide, nucleotides in the polynucleotide interact with the first constriction region and the second constriction region within the continuous channel, and wherein each of the first constriction region and the second constriction region is capable of distinguishing different nucleotides.
[0031] Further, a nucleic acid binding protein is used to control the movement of the polynucleotide relative to the continuous channel.
[0032] To achieve the above object, according to a sixth aspect of the present application, there is provided an application of any one of the nanopore protein complexes, or any one of the nanopore sensors, or any one of the kits, or any one of the nanopore sequencing systems in biomolecule detection, polynucleotide sequencing, polypeptide sequencing, or polysaccharide sequencing.
[0033] To achieve the above object, according to a seventh aspect of the present application, there is provided a preparation method of any one of the nanopore complexes, the preparation method comprising: respectively constructing an expression vector of a nanopore protein monomer and an expression vector of an auxiliary protein, co-expressing the nanopore protein monomer and the auxiliary protein in a competent cell, isolating and purifying the nanopore protein monomer and the auxiliary protein expressed by the competent cell, mixing the nanopore protein monomer and the auxiliary protein, and obtaining the nanopore complex; or constructing an expression vector of a nanopore protein monomer, expressing the nanopore protein monomer in a competent cell, isolating and purifying the nanopore protein monomer expressed by the competent cell, mixing the nanopore protein monomer and an auxiliary protein artificially synthesized in vitro, and obtaining the nanopore complex.
[0034] The technical scheme of the application provides a nanopore protein complex, which comprises nanopore proteins with a nanopore channel structure formed by polymerization of a plurality of nanopore monomers, and a plurality of auxiliary proteins connected one by one with the plurality of nanopore monomers, wherein the N terminal of each auxiliary protein is located in the nanopore channel structure and forms a continuous channel together with the nanopore channel structure. The nanopore protein complex with the continuous channel contains at least two receptors, and when high-throughput sequencing, especially nanopore sequencing based on electrical signal recognition, the nanopore protein complex can improve the current difference between a plurality of nucleotides, enhance the base recognition ability, and improve the sequencing accuracy of polynucleotides, so as to realize sequencing with higher accuracy resolution. BRIEF DESCRIPTION OF DRAWINGS
[0035] The drawings accompanying the specification of the present application form a part thereof and serve to provide further understanding of the present application, the illustrative embodiments of the present application and its description serve to explain the present application. The present application is not limited by the improper limitation of the drawings. In the drawings:
[0036] Fig. 1 shows a schematic diagram of the three-dimensional structure of BCP58-AP58 predicted in embodiment one of the present application;
[0037] Fig. 2 shows a schematic diagram of the three-dimensional structure of BCP58-AP58-N_30-64 predicted in embodiment one of the present application;
[0038] Fig. 3 shows a cross-sectional view of the nanopore constriction region in the three-dimensional structure of BCP58-AP58 predicted in embodiment one of the present application;
[0039] Fig. 4 shows a cross-sectional view of the nanopore constriction region in the three-dimensional structure of BCP58-AP58-N_30-64 predicted in embodiment one of the present application;
[0040] Fig. 5 shows a schematic diagram of the interaction site between BCP58 and AP58-N_30-64 predicted in embodiment one of the present application;
[0041] Fig. 6 shows a SDS-PAGE result diagram of the wild type BCP58 protein after purification in embodiment four of the present application;
[0042] Fig. 7A shows a schematic diagram of a sequencing library containing a helicase in embodiment six of the present application, and Fig. 7B shows a schematic diagram of the pairing structure of the sequencing library and the single-stranded DNA containing cholesterol (wherein a: upper strand; b: lower strand; c: double-stranded target fragment; d: helicase; e: cholesterol-labeled single-stranded DNA);
[0043] Fig. 8 shows a current signal diagram generated when the library DNA passes through the nanopore protein BCP58_1 according to embodiment six of the present application;
[0044] Figure 9 shows a current signal graph generated when the Chinese library DNA passes through the nanopore protein complex BCP58_1-AP58-N_30-55 according to Embodiment Six of the present application;
[0045] Figure 10 shows a current signal graph generated when the Chinese library DNA passes through the nanopore protein BCP58_2 according to Embodiment Six of the present application;
[0046] Figure 11 shows a current signal graph generated when the Chinese library DNA passes through the nanopore protein complex BCP58_2-AP58-N_30-64 according to Embodiment Six of the present application;
[0047] Figure 12 shows a current signal graph generated when the Chinese library homopolymer respectively passes through the nanopore protein BCP58_1 and BCP58_1-AP58-N_30-55 according to Embodiment Six of the present application;
[0048] Figure 13 shows a current signal graph generated when the Chinese library homopolymer respectively passes through the nanopore protein BCP58_2 and BCP58_2-AP58-N_30-64 according to Embodiment Six of the present application;
[0049] Figure 14 shows the structure of the spacer iSp18. DETAILED DESCRIPTION
[0050] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the embodiments.
[0051] As mentioned in the background section, the existing nanopore protein usually only contains one narrowest constriction as a sensor for detecting the analyte to be detected, but the ability to distinguish between polynucleotides is limited, which greatly limits the improvement of accuracy. If two or more sensors are introduced to further improve the current difference between multiple nucleotides, there will be an opportunity to enhance the base recognition ability. In the present application, the performance of the nanopore protein is modified by protein engineering or combination with polypeptides / enzymes to have at least two sensors, so as to meet the sequencing demand of higher precision resolution. On this basis, the applicant proposes a series of technical solutions in the present application.
[0052] In a first typical embodiment of the present application, a nanopore protein complex is provided, which includes: a nanopore protein and a plurality of auxiliary proteins, wherein the nanopore protein includes a plurality of pore protein monomers and a nanopore channel structure formed by polymerization of the plurality of pore protein monomers; the N-terminus of each auxiliary protein is located within the nanopore channel structure, the plurality of auxiliary proteins are connected one-to-one with the plurality of pore protein monomers, and together with the nanopore channel structure form a continuous channel; wherein, along the direction of the object to be detected passing through the continuous channel, the continuous channel includes a first constriction region and a second constriction region connected in sequence, the first constriction region is formed by at least part of the nanopore protein, and the second constriction region is formed by part or all of the auxiliary protein; the pore protein monomers are selected from any one or more of the following: (a) a protein having the amino acid sequence shown in SEQ ID NO: 1; or (b) a protein having substitution, deletion and / or addition of one or more amino acids at at least one of the following positions of SEQ ID NO: 1: 77, 81, 82, 176, 210, 214, 232, 66, 69, 70, 74, 109, 110, 113, 117, 118, 119, 120, 123, 127, 128, 168, 211, 221, 224, 227, 229, and having the function of forming a nanopore channel structure by polymerization; or (c) a protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity with the protein in (a) or (b), and having the function of forming the nanopore channel structure by polymerization.
[0053] In the above nanopore protein complex, the nanopore protein having a nanopore channel structure formed by polymerization of a plurality of pore protein monomers, and the auxiliary proteins connected one-to-one with the plurality of pore protein monomers, the N-terminus of each auxiliary protein is located within the nanopore channel structure, and together with the nanopore channel structure forms a continuous channel. The nanopore protein complex with such a continuous channel contains at least two sensors, which can meet the sequencing requirements of higher precision resolution (such as, can increase the current difference between a plurality of nucleotides, enhance the base recognition ability, and improve the sequencing accuracy of polynucleotide homopolymer).
[0054] The wild type of the above nanopore protein monomer is BCP58, and the gene sequence is derived from deep sea metagenome sequencing (derived from deep sea samples of Mariana Trench). The nanopore protein is polymerized by nine monomers of the same type, and the sequence of BCP58 is SEQ ID NO: 1 (which has been recorded in PCT / CN2022 / 143060). BCP58 has a nanopore channel structure and can be used as a detection protein for the detection of biological small molecules such as nucleotides, amino acids, sugars, and vitamins, or for nanopore-based DNA, RNA or polypeptide detection.
[0055] The amino acid sequence SEQ ID NO: 1 of the wild-type pore-forming monomer BCP58 is as follows:
[0056] It should be noted that the mutants of the pore-forming monomer in the above-mentioned nanopore complex include both the truncation mutants of BCP58 and the single-point or multi-point mutants of BCP58; similarly, the auxiliary protein can be an existing protein having the ability to form a continuous channel with the pore-forming monomer one by one or a mutant thereof, or a newly developed protein having such a function or a variant thereof. In a specific nanopore complex, the pore-forming monomer and the auxiliary protein can each be independently selected from the wild type or any one of the variants.
[0057] In a preferred embodiment, the pore-forming monomer is selected from the variants defined in (b), and the types of the substituted amino acids are each independently selected from the following: E176 is mutated to E176A, E176G, E176V, E176L, E176I, E176Y, E176F or E176W; K210 is mutated to K210A, K210G, K210V, K210L, K210I, K210Y, K210F or K210W; S214 is mutated to S214A, S214G, S214V, S214L, S214I, S214Y, S214F or S214W; and T232 is mutated to T232A, T232G, T232V, T232L, T232I, T232Y, T232F or T232W.
[0058] The above-mentioned four amino acid sites are located on the side facing the membrane of the transmembrane region of the pore-forming monomer, and mutation thereof can increase the performance of the nanopore complex in inserting the membrane.
[0059] In a preferred embodiment, the pore protein monomer is selected from the variants defined in (b) and the type of substituted amino acid is independently selected from each of the following: K66 is mutated to K66N, K66A, K66G, K66S, K66T, or K66Q; D69 is mutated to D69K, D69R, D69N, D69A, D69G, D69S, D69T, or D69Q; Q70 is mutated to Q70K, Q70R, Q70D, Q70E, Q70N, Q70A, Q70G, Q70S, Q70T, or Q70Q; Y74 is mutated to Y74N, Y74A, Y74G, Y74S, Y74T, or Y74Q; R109 is mutated to R109N, R109A, R109G, R109S, R109T, or R109Q; E110 is mutated to E110K, E110R, E110N, E110A, E110G, E110S, E110T, or E110Q; Q113 is mutated to Q113K, Q113R, Q113N, Q113A, Q113G, Q113S, or Q113T; T117 is mutated to T117K, T117R, T117N, T117A, T117G, T117S, or T117Q; E118 is mutated to E118K, E118R, E118N, E118A, E118G, E118S, E118T, or E118Q; R119 is mutated to R119N, R119A, R119G, R119S, R119T, or R119Q; K120 is mutated to K120N, K120A, K120G, K120S, K120T, or K120Q; R123 is mutated to R123N, R123A, R123G, R123S, R123T, or R123Q; K127 is mutated to K127N, K127A, K127G, K127S, K127T, or K127Q; K128 is mutated to K128N, K128A, K128G, K128S, K128T, or K128Q; R168 is mutated to R168N, R168Q, R168S, R168T, R168A, or R168G; E211 is mutated to E211N, E211Q, E211S, E211T, E211A, or E211G; E221 is mutated to E221N, E221Q, E221S, E221T, E221A, or E221G; E224 is mutated to E224N, E224Q, E224S, E224T, E224A, or E224G; E227 is mutated to E227N, E227Q, E227S, E227T, E227A, or E227G; and E229 is mutated to E229N, E229Q, E229S, E229T, E229A, or E229G.
[0060] Most of the above amino acid sites are located at the entrance, inner wall of the pore, and exit of the nanopore protein, which are important for the capture of the library and the smooth passage of the library DNA through the nanopore protein. Because nucleic acids are negatively charged, increasing or decreasing the amount of positive charge at the entrance can adjust the library capture ability of the nanopore protein.
[0061] Among the above mutation sites, there are mainly three types of sites: mutation sites in the sensor region, mutation sites at the entrance of the nanopore protein, and mutation sites in the transmembrane region of the nanopore protein. Among them, the mutation sites in the sensor region mainly determine the opening of the nanopore protein, thereby directly determining the opening current. Because the nucleic acid to be tested is negatively charged, by adjusting the amino acids of the mutation sites at the entrance of the pore protein, including but not limited to mutating uncharged or negatively charged amino acids to positively charged amino acids, or mutating positively charged amino acids to other types of amino acids, etc., such mutations can adjust the capture rate of the library. The mutation of the transmembrane region of the nanopore protein can enhance the stability of the pore protein inserted into the lipid or polymer membrane. In addition, the charged amino acids located at the inner wall of the pore structure of the nanopore protein or at the exit can also affect the pore of the sample to be tested.
[0062] Among the above mutation sites, S77, S81, and F82 are located in the sensor region. K66, D69, Q70, Y74, R109, E110, Q113, T117, E118, R119, K120, R123, K127, and K128 are located at the entrance of the nanopore protein, and the mutation of their charge properties can play an important role in adjusting the capture rate of the library and sequencing noise. E176, K210, S214, and T232 are located at the outer wall of the transmembrane region of the nanopore protein, and the mutation of hydrophobic amino acids can enhance the stability of the nanopore protein inserted into the pore. R168, E211, E227, and E229 are located at the inner wall of the nanopore structure, which are charged amino acids in the barrel wall of the nanopore protein, and the mutation of their charge can promote the smooth passage of the biological molecules such as DNA to be tested through the nanopore protein. E221 and E224 are located at the exit loop region of the pore structure, and the mutation of them can promote the nucleic acid chain after sequencing to leave the nanopore protein.
[0063] In a preferred embodiment, in (b), the type of substituted amino acid is each independently selected from the following: S77A, F82Q, S81N, F82A, K120N, K127S, E176I.
[0064] In a preferred embodiment, the pore protein monomer comprises a protein having an amino acid sequence of any one of SEQ ID NO: 24-SEQ ID NO: 28.
[0065] SEQ ID NO: 24:
[0066] SEQ ID NO: 25:
[0067] SEQ ID NO: 26:
[0068] SEQ ID NO: 27:
[0069] SEQ ID NO: 28:
[0070] The mutant of BCP58 described above can also be a protein having 70% or more (preferably 75% or more, 80% or more, 85% or more, 90% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more) identity to any of the amino acid sequences defined above and having the function of forming a nanopore structure by polymerization.
[0071] It should be noted that the identity in the present application refers to the "sequence identity" between two amino acid sequences, i.e., the percentage of identical amino acids between the sequences. Methods for assessing the degree of sequence identity between amino acids or nucleotides are known to those skilled in the art. For example, the degree of amino acid sequence identity is typically measured using sequence analysis software. For example, it can be determined by the BLAST program of the NCBI database. For determination of sequence identity, see, for example, Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987, and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991.
[0072] The protein having 70%, 75%, 80%, 85%, 90%, 95%, 99% or more (such as 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, even 99.9% or more) identity to the BCP58 nanopore protein shown in SEQ ID NO: 1 and having the same biological activity has a similar probability of having an active site, an active pocket, an active mechanism, a protein structure, etc. similar to the corresponding wild-type protein.
[0073] Amino acid residues can be represented according to the standard three-letter or one-letter amino acid code as is well known and conventional in the art. Herein, amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine (Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (lie; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).
[0074] Conservative amino acid substitutions or replacements are well known in the art, for example, conservative amino acid substitutions are preferably substitution of one amino acid residue by another amino acid residue within the same group (1)-(5) as follows: (1) smaller aliphatic nonpolar or weakly polar residues: Ala, Ser, Thr, Pro, and Gly; (2) polar negatively charged residues and their (uncharged) amides: Asp, Asn, Glu, and Gin; (3) polar positively charged residues: His, Arg, and Lys; (4) larger aliphatic nonpolar residues: Met, Leu, lie, Val, and Cys; and (5) aromatic residues: Phe, Tyr, and Trp. Particularly preferred conservative amino acid substitutions are as follows: Ala by Gly or Ser; Arg by Lys; Asn by Gin or His; Asp by Glu; Cys by Ser; Gin by Asn; Glu by Asp; Gly by Ala or Pro; His by Asn or Gin; lie by Leu or Val; Leu by lie or Val; Lys by Arg, Gin, or Glu; Met by Leu, Tyr, or lie; Phe by Met, Leu, or Tyr; Ser by Thr; Thr by Ser; Trp by Tyr.
[0075] Conservative amino acid substitutions can also be made according to well known rules for amino acid replacement as known to those skilled in the art, such as the“BLOSUM62 score matrix” and the like.
[0076] In other preferred embodiments, the amino acid sequence of the pore protein monomer is selected from the group consisting of: an amino acid sequence which has a substitution F82Q based on SEQ ID NO: 1; or an amino acid sequence which has substitutions S77A+S81N+F82Q+K127S based on SEQ ID NO: 1.
[0077] In the present application, the wild-type helper protein is denoted as AP58 (the full-length amino acid sequence is shown in SEQ ID NO: 2, wherein the 1st-29th is the N-terminal signal peptide, the 30th-64th is the core region forming the second sensor, and the remaining 65th-142nd is the non-core region), which is a helper protein of BCP58 mined in the corresponding species of BCP58, which can form a complex with BCP58, wherein the N-terminus of AP58 can be embedded into the nanopore channel of BCP58 to form an additional constriction zone, when the analyte passes through the nanopore complex formed by the two, respectively interacts with the two constriction zones, and the current signal change generated can provide more detailed information and higher resolution.
[0078] Correspondingly, various variants of BCP58 and / or various variants of AP58 can enhance the binding ability between the nanopore protein and the helper protein to different degrees, and then form a nanopore protein complex with two constriction zones, so as to improve the resolution of the analyte when applied to sequencing the analyte.
[0079] The amino acid sequence shown in SEQ ID NO: 2 is as follows:
[0080] MMILRVVKRAVPVLFVAIGQMCAGPVTFASELVYTPVNPSFGGNPLNGTYLLNNAQAQNKEKDPDGYSYESPSDLDRLASSLQSRLLGQLLADVGNGNTGSIVTDEFSMVVNDDGSGGLVIVITDLATGETTTIAVNGLIPD, wherein the 1st-29th (underlined part) is the N-terminal signal peptide, the 30th-64th is the core region forming the second sensor, and the remaining 65th-142nd is the non-core region.
[0081] In some preferred embodiments, the helper protein is selected from any one of: i) a polypeptide fragment having an amino acid sequence of any one of SEQ ID NOs: 3-8; ii) a polypeptide fragment having an amino acid sequence of any one of SEQ ID NOs: 3-8 with a signal peptide sequence added at the N-terminus thereof; iii) a protein in which at least one of the following positions of any one of SEQ ID NOs: 3-8 is substituted and / or deleted and / or added with one or more amino acids: L46, T49, N53, K62, D63, or P64; or iv) a polypeptide fragment having at least 50% identity to the amino acid sequence of any one of the polypeptide fragments of i), ii), or iii).
[0082] In the above i), the amino acid sequence of any one of SEQ ID NOs: 3-8 is the core region of the nanopore complex formed by one-to-one correspondence between the accessory protein and the pore protein monomer, and the specific sequence is as follows:
[0083] SEQ ID NO: 3: SELVYTPVNPSFGGNPLNGTYLLNNAQAQNKEKDP;
[0084] SEQ ID NO: 4: SELVYTPVNPSFGGNPLNGTYLLNNAQAQNKE;
[0085] SEQ ID NO: 5: SELVYTPVNPSFGGNPLNGTYLLNNAQAQ;
[0086] SEQ ID NO: 6: SELVYTPVNPSFGGNPLNGTYLLNNAQA;
[0087] SEQ ID NO: 7: SELVYTPVNPSFGGNPLNGTYLLNNAQ;
[0088] SEQ ID NO: 8: SELVYTPVNPSFGGNPLNGTYLLNNA.
[0089] The above 6 sequences can be artificially synthesized in vitro, or synthesized by genetic engineering means. When the latter is used for synthesis (such as expression in bacteria), in order to facilitate subsequent separation, a signal peptide sequence can be added to the N-terminus of the amino acid sequence of any one of SEQ ID NOs: 3-8 of the accessory protein. The specific signal peptide sequence is reasonably selected according to the cell used for expression. In this application, it can be the amino acid sequence of positions 1-29 of SEQ ID NO: 2.
[0090] In addition, when the accessory protein is selected from the protein in iii), that is, the protein in which at least one of the following positions of any one of the amino acid sequences of SEQ ID NOs: 3-8 is substituted and / or deleted and / or added with one or more amino acids: L46, T49, N53, K62, D63 or P64, as long as the mutation type is helpful to change the diameter of the nanopore channel structure, stability, affinity with the passing molecule, ability to generate electrical signal, etc. It is applicable to this application.
[0091] In some preferred embodiments of the present application, L46 includes but is not limited to mutations into S, N, T, V or Q, T49 includes but is not limited to mutations into N, S, V, L, A or Q; N53 includes but is not limited to mutations into S, V, L, A, Q, T; K62 includes but is not limited to mutations into N, S, V, L, A, I, Q, T, P, Y or W; D63 includes but is not limited to mutations into N, S, V, L, A, I, Q, T, P, Y or W; P64 includes but is not limited to mutations into N, S, V, L, A, I, Q, T, Y or W. These mutants can change the diameter, stability, affinity to pass-by molecules, ability to generate electrical signals, etc. of the protein pore channel structure, thus obtaining a nanopore protein more suitable for nanopore sequencing.
[0092] The above-mentioned helper protein can also be a polypeptide fragment having at least 50% identity (preferably at least 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 99%, such as at least 99.5%, 99.6%, 99.7%, 99.8%, or even more than 99.9%) to the amino acid sequence defined in any of i), ii) or iii) above. For example, it can be a full-length sequence as shown in SEQ ID NO: 2.
[0093] The above-mentioned helper protein in the form of various polypeptide fragments all have similar activity to wild-type AP58, the N-terminus can be embedded in the nanopore channel structure of the nanopore protein, and form a nanopore protein complex with at least two consecutive constriction zones, and further can generate different detection signals based on the differences of different biological molecules (such as nucleotides) interacting with different constriction zones, thus improving the detection accuracy, especially the detection accuracy of long homopolymer molecules.
[0094] The above-mentioned AP58 mutant can be a truncated body of different lengths including SEQ ID NO: 2, including but not limited to truncated bodies of any of the following N-terminal length intervals: 1-55, 1-56, 1-57, 1-58, 1-59, 1-60, 1-61, 1-62, 1-63 or 1-64, these mutants can change the pore-embedding ability of the nanopore complex and the adaptability to the membrane.
[0095] The constriction zone of the existing nanopore protein can distinguish different bases of the nucleic acid to be tested passing through the pore, and each nucleotide produces a current change when interacting with the constriction, which can be converted into corresponding sequence information. However, in the DNA homopolymer, the signal measured is not large enough in differentiation, and it is difficult to realize the single-base resolution of the current change, so that the length of the homopolymer cannot be accurately determined only according to the size of the signal measured. In the present application, a second constriction zone is introduced by using an auxiliary protein to bind a transmembrane nanopore protein, which can interact with nucleotides different from those interacting with the first constriction zone in space, and further produce more current change information. Therefore, compared with the case without auxiliary protein, the nanopore protein complex improved in the present application can more accurately determine the homopolymer sequence information, especially for longer homopolymer sequence fragments. That is, the above nanopore protein complex of the present application can provide improved sequencing signals for nucleotide homopolymers.
[0096] In the above nanopore protein complex, the connection between the pore protein monomer and the auxiliary protein can be covalent connection or non-covalent connection. In order to improve the stability of the nanopore protein complex as much as possible during the detection process and reduce the probability of dissociation under the action of electric field force, it is preferred that the two are connected by covalent connection. Therefore, any way that can achieve covalent connection of the two without affecting their activities is suitable for the present application.
[0097] In order to further improve the stability of the auxiliary protein attached to the nanopore protein, some amino acid residues in the interaction region of the two can be mutated or modified. In some specific embodiments of the present application, mutations are introduced in wild-type pore protein monomer BCP58 (and its mutants) and AP58 (and its mutants) to further stabilize the covalent connection. The mutations include but are not limited to: mutation of any one of the amino acid sites D181, R179, N159, E227, Q213, R168, F219, Y217, L170 of BCP58 to cysteine or introduction of a non-natural amino acid at one site. The amino acid sites of AP58 include but are not limited to mutation of any one of S30, E31, L32, V33, V37, N38, S40, A55, Q58, Q59 to cysteine or introduction of a non-natural amino acid at one site.
[0098] It should be noted that when a certain amino acid site of the pore protein monomer is mutated to Cys, a certain amino acid site of the auxiliary protein also needs to be mutated to Cys to form a disulfide bond covalent connection. However, there is no special one-to-one correspondence between the specific mutation sites, and they can be connected to each other at any site of the above two molecules.
[0099] The unnatural amino acids introduced by the above mutations include, but are not limited to, 4-azido-L-phenylalanine, 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselenyl)-L-alanine, O-2-propyn-l-yl-L-tyrosine, 4(dihydroxyboryl)-L-phenylalanine, 4-[(ethylsulfanyl)carbonyl]-L-phenylalanine, (2S)-2-amino-3-{4-[(propan-2-ylsulfanyl)carbonyl]phenyl}propanoic acid, (2S)-2-amino-3-{4-[(2-amino-3-sulfanylpropanoyl)amino]phenyl}propanoic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino-L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthalen-2-ylamino)propanoic acid, 6-(methylsulfanyl)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propanoic acid, (2R)-2-aminooctanoic acid 3-(2,2'-dipyridyl-5-yl)-D-alanine, 2-amino-3-(8-hydroxy-3-quinolyl)propanoic acid, 4-benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)sulfanyl]propanoic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propanoic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)oxy]carbonyl}amino)hexanoic acid, O-(2-nitrobenzyl)-L-tyrosine, 2-nitrophenylalanine, and the like.
[0100] The above covalent attachment includes, but is not limited to, a spontaneous cross-linking (e.g., cysteine spontaneously cross-linking into disulfide bond), a cross-linker attachment, a protein fusion attachment, a polypeptide molecular attachment, or a synthetic chemical small molecule attachment, and the like.
[0101] Suitable commercial cross-linking agents include, but are not limited to, chemical cross-linking agents including the following functional groups: maleimides, active esters, succinimides, azides, alkynes (e.g., dibenzocyclooctyne (DIBO or DBCO), difluorocycloalkynes, and linear alkynes), phosphines (e.g., those used in click chemistry, both traceless and non-traceless Staudinger ligation), haloacetyl groups (e.g., iodoacetamide), phosgene-type reagents, sulfonyl chloride reagents, isothiocyanates, acyl halides, hydrazines, disulfides, vinyl sulfones, aziridines, and photosensitive reagents (e.g., aryl azides, diaziridines).
[0102] In a second exemplary embodiment of the present application, a nanopore sensor is provided, which includes a membrane and any of the nanopore protein complexes described above, wherein the nanopore protein in the nanopore protein complex is embedded in the membrane, the nanopore channel structure of the nanopore protein and the plurality of auxiliary proteins together form a continuous channel across the membrane, and the continuous channel generates an electrical signal when a voltage is applied across the membrane.
[0103] The nanopore sensor formed by the nanopore protein complex containing the two constriction zones described above can improve the discrimination and sequencing accuracy of homopolymer nucleic acid molecules when applied to sequencing such molecules.
[0104] In the nanopore sensor described above, the membrane preferably includes a layer of amphiphilic molecules. More preferably, the membrane is a phospholipid bilayer composed of dipalmitoyl phosphatidylcholine.
[0105] In some embodiments, when a voltage is applied across the membrane, a biomolecule passes through the continuous channel in the nanopore sensor and is displaced, and the continuous channel generates a varying electrical current signal (the electrical current signal generated by the passage of a biomolecule through the continuous channel is different from the electrical current signal generated by the passage of no biomolecule). It should be noted that the biomolecule suitable for the present application includes, but is not limited to, polynucleotides, polypeptides, or polysaccharides; it is particularly suitable for nanopore sequencing of polynucleotides containing homopolymers, with relatively higher sequencing accuracy.
[0106] In a third exemplary embodiment of the present application, a kit is provided, which includes any of the nanopore protein complexes described above or any of the nanopore sensors described above, and in addition, the kit further includes one or more of the following optional components: a membrane, a sequencing buffer, a nuclease, a polymerase, a topoisomerase, a ligase, a helicase, and a cholesterol-linked single-stranded DNA.
[0107] In a fourth exemplary embodiment of the present application, a nanopore sequencing system is provided, which comprises the nanopore sensor as described above, and further comprises: a conductive solution, positive and negative electrodes for providing a voltage potential across the membrane, and a measuring device for measuring the electrical signal passing through the continuous channel; wherein the nanopore sensor is located in the conductive solution and divides the conductive solution into a first chamber and a second chamber, the negative electrode is in the first chamber, and the positive electrode is in the second chamber.
[0108] In some preferred embodiments, the positive and negative electrodes are connected to a signal processing chip for measuring the electrical signal passing through the continuous channel.
[0109] The specific materials of the positive and negative electrodes include, but are not limited to, metal electrode materials or composite electrode materials.
[0110] In a fifth exemplary embodiment of the present application, a method for identifying or characterizing a biomolecule is provided, which comprises: using the nanopore protein complex as described above, or the nanopore sensor as described above, or the nanopore sequencing system as described above, identifying or characterizing the biomolecule by measuring and analyzing the electrical signal generated when the biomolecule passes through the continuous channel of the nanopore protein complex.
[0111] In some preferred embodiments, the method comprises: contacting the nanopore sequencing system with a biomolecule (the specific biomolecule is selected from a polynucleotide, a polypeptide or a polysaccharide, and more preferably, the biomolecule is selected from a polynucleotide containing a homopolymer); applying a voltage across the membrane so that the biomolecule enters the continuous channel and moves relative to the continuous channel; and measuring the electrical signal generated when the biomolecule moves relative to the continuous channel, thereby identifying or characterizing the biomolecule.
[0112] In the present application, the specific types of electrical signals generated when the biomolecule passes through the first constriction region and the second constriction region include, but are not limited to, current signals. More specifically, it can be the strength of the current block signal, the duration of the current block signal, or the interval time of the occurrence of the current block signal event, etc.
[0113] In some specific embodiments, identifying the biomolecule includes identifying whether the biomolecule is present; and characterizing the biomolecule includes determining the composition of the biomolecule.
[0114] According to the specific type of the biomolecule, the determination of the composition of the biomolecule includes: determining the nucleotide sequence of the polynucleotide, the amino acid sequence of the polypeptide, or the monosaccharide arrangement order of the polysaccharide.
[0115] In some preferred embodiments, the biomolecule is a polynucleotide, the nucleotides in the polynucleotide interact with the first constriction region and the second constriction region in the continuous channel, and wherein each of the first constriction region and the second constriction region is capable of discriminating different nucleotides.
[0116] In particular, in some preferred embodiments, the polynucleotide moves through the continuous channel and translocates across the membrane. In some preferred embodiments, a nucleic acid binding protein (such as a helicase or a polymerase) is used to control the movement of the polynucleotide relative to the continuous channel pore.
[0117] In some preferred embodiments, the above method comprises determining the nucleotide sequence of the polynucleotide containing homopolymer. The "homopolymer" in the present application refers to a polymer polymerized from one monomer, such as a single repeated nucleic acid sequence: TTTTTTTTTTT (SEQ ID NO: 29), CCCCCC, GGGGGGG, AAAAAAAAAA (SEQ ID NO: 30), etc. The resolution of the method of the present application is relatively higher for polynucleotides containing homopolymers, and thus the sequencing accuracy is higher.
[0118] According to a sixth typical embodiment of the present application, there is provided a use of any one of the above nanopore protein complex, or any one of the above nanopore sensor, or any one of the above kit, or any one of the above nanopore sequencing system in biological small molecule detection, polynucleotide sequencing, polypeptide sequencing or polysaccharide sequencing.
[0119] In a seventh typical embodiment of the present application, there is provided a preparation method of the above nanopore complex, comprising: respectively constructing an expression vector of a nanopore protein monomer and an expression vector of an auxiliary protein, co-expressing the pore protein monomer and the auxiliary protein in a competent cell, isolating and purifying the pore protein monomer and the auxiliary protein expressed by the competent cell, mixing the pore protein monomer and the auxiliary protein to obtain the nanopore complex; or
[0120] constructing an expression vector of a nanopore protein monomer, expressing the pore protein monomer in a competent cell, isolating and purifying the pore protein monomer expressed by the competent cell, mixing the pore protein monomer and an auxiliary protein artificially synthesized in vitro to obtain the nanopore complex.
[0121] The above nanopore protein complex of the present application can be obtained by E. coli co-expression and dual-affinity purification, or wild-type BCP58 or its mutant and AP58 or its mutant are respectively expressed and purified in E. coli and then recombined to obtain, or only wild-type BCP58 or its mutant is expressed and recombined with a polypeptide artificially synthesized in vitro containing a partial fragment of AP58 to obtain.
[0122] It should be noted that the application of any one of the above nanopore protein complexes, or any one of the above nanopore sensors, or any one of the above kits, or any one of the above nanopore sequencing systems in biological small molecule detection, polynucleotide sequencing, polypeptide sequencing or polysaccharide sequencing is also within the scope of protection of the present application.
[0123] The beneficial effects of the present application will be further explained in detail below in conjunction with specific examples. The reagent materials used in the examples, which are not specifically listed with the manufacturers and item numbers, are all general reagent materials in the art, which can be sourced from any manufacturer in the art.
[0124] Example 1: Predicted structure of AlphaFold2-Multimer of wild-type BCP58-AP58
[0125] The structure of BCP5-AP58 was predicted using AlphaFold 2-Multimer. The prediction results are shown in Figures 1 and 2. Figure 1A is a side view of the predicted structure of BCP58-AP58, and Figure 1B is a top view of the predicted structure of BCP58-AP58. Figure 2A is a side view of the predicted structure of BCP58-AP58_N30-64 (AP58_N30-64 is SEQ ID NO: 3), and Figure 2B is a top view of the predicted structure of BCP58-AP58_N30-64.
[0126] The complex BCP58-AP58 has two constriction zones, which play a decisive role in the generation of current signals. Figures 3 and 4 show the side chain structures of important amino acids in the respective constriction zones of the predicted structures of BCP58-AP58 and BCP58-AP58-N_30-64 complexes, wherein the main constriction zone amino acid of AP58-N_30-64 is L46, and T49 and N53 also have an impact on the generation of current signals.
[0127] Figure 5 shows the important amino acids of BCP58 and AP58-N_30-64 that interact with each other, which can improve the stability of the complex through mutation or cross-linking. The BCP58 amino acid sites include N159, R179, D181, S209, Q213, E227, E229, T233 and N235; and the AP58-N_30-64 amino acid sites include S30, L32, V33, Y34, T35, V37, N38 and S40.
[0128] Example 2: Construction of expression vectors for pore protein BCP58 monomer and AP58 monomer and respective mutants
[0129] The DNA sequence (SEQ ID NO: 9) encoding the pore protein monomer was inserted into the multiple cloning region of the vector pET24a. StrepII amino acids were added at the C-terminus of the pore protein monomer amino acid sequence (SEQ ID NO: 1) as a purification tag, with kanamycin as a selection tag. The expression vector of the pore protein BCP58 monomer was used as a template to construct the corresponding mutants by the method of site-directed mutation using the Agilent site-directed mutation kit.
[0130] This example tested mutant BCP58_1 (F82Q), mutant BCP58_2 (S77A+S81N+F82Q+K127S), in which mutant 1 is a mutation in the gate region, and mutant 2 is a mutation in the gate region and the entry region. The method of constructing the expression vector of the mutant is consistent with the method of constructing the expression vector of the wild type.
[0131] The DNA sequence of the BCP58 monomer (WT) is shown as SEQ ID NO: 9:
[0132] The DNA sequence of BCP58_1 (F82Q) is shown as SEQ ID NO: 10:
[0133] The amino acid sequence of BCP58_1 (F82Q, del288-319) is shown as SEQ ID NO: 11:
[0134] The amino acid sequence of BCP58_1 (F82Q, del288-319) with the signal peptide cleaved and with a Strep II tag is shown as SEQ ID NO: 12:
[0135] The DNA sequence of BCP58_2 (S77A+S81N+F82Q+K127S) is shown as SEQ ID NO: 13:
[0136] The amino acid sequence of BCP58_2 (S77A+S81N+F82Q+K127S) is shown as SEQ ID NO: 14:
[0137] The amino acid sequence of BCP58_2 (S77A+S81N+F82Q+K127S) with the signal peptide cleaved and with a Strep II tag is shown as SEQ ID NO: 15:
[0138] Example Three: Culture and induction of pore protein monomer strains
[0139] The constructed expression plasmids of the pore protein monomer and its mutants were independently transformed into E. coli expression strain E. coli BL21 (DE3), and the bacterial liquid was uniformly smeared on a plate containing 50 μg / mL kanamycin and cultured at 37°C overnight. The next day, a single colony was inoculated in 5 mL LB liquid medium containing 50 μg / mL kanamycin, and cultured at 37°C, 200 rpm, overnight. The obtained bacterial liquid was inoculated in 50 mL LB liquid medium containing 50 μg / mL kanamycin at a volume ratio of 1:100, and cultured at 37°C, 200 rpm, for 4 h. The bacterial liquid was inoculated in 2 L LB liquid medium containing 50 μg / mL kanamycin at a volume ratio of 1:100, and cultured at 37°C, 200 rpm. When the OD value reached about 0.6-0.8, IPTG was added at a final concentration of 0.5 mM, and the culture was continued at 16°C, 200 rpm, for about 16-18 h. The bacterial liquid was collected by centrifugation at 8000 rpm, and the bacterial cells were stored at -20°C for use. 600
[0140] Example Four: Preparation of Recombinant Nanopore Protein Complex
[0141] 1. Extraction and purification of pore protein BCP58
[0142] (1) Buffer preparation
[0143] Buffer A: 20 mM Tris-HCl, 150 mM NaCl, 1% DDM, pH 8.0.
[0144] Buffer B: 20 mM Tris-HCl, 150 mM NaCl, 0.05% Tween20, 15% Glycerol, pH 8.0.
[0145] Buffer C: 20 mM Tris-HCl, 150 mM NaCl, 0.05% Tween20, 5 mM desthiobiotin, pH 8.0.
[0146] (2) Purification steps
[0147] The bacterial cells were resuspended at a ratio of 1 g of bacterial cells to 10 mL of Buffer A, and the cells were broken by ultrasonic treatment until the bacterial cell solution was clear. Then, the solution was placed on a rotator and rotated at 4°C overnight. The next day, the solution was centrifuged at 18000 rpm at 4°C for 1 h, and the supernatant was filtered through a 0.22 μm filter membrane and stored at 4°C.
[0148] Strep-Tactin beads (IBA Lifesciences) column was equilibrated with Buffer A for 5 column volumes (CV) using AKTA pure chromatograph, then loaded at 2 mL / min. After loading, the column was washed with Buffer B for 20 CV, and the target protein was eluted with Buffer C, and collected.
[0149] The obtained protein was concentrated to 1 mL, and then loaded onto a Superdex 6 increase 10 / 300 GL (Cytiva) equilibrated with buffer B. The target protein was collected and then stored at -80°C. The target protein obtained after purification was subjected to SDS-PAGE electrophoresis. Figure 6 shows the purification results of the wild-type pore protein BCP58, which shows that the protein is in a non-denatured state of a nonamer and in a denatured state of a monomer. The purification results of the mutant pore protein BCP58 are similar.
[0150] 2. In vitro recombination of pore protein complex BCP58 and N-terminal truncated polypeptide of AP58
[0151] The N-terminal truncated polypeptide (AP58-N_30-55, AP58-N_30-64) was synthesized by Genscript. The polypeptide powder was dissolved in Buffer B (concentration 1 mg / ml). The pore protein mutant and the polypeptide were mixed at a certain concentration ratio (e.g. 10:1, 50:1, 100:1, 200:1, etc.) at 4°C overnight to obtain the pore protein complex, which was stored at -80°C for later use.
[0152] Example Five: Construction of a nanopore biosensor using a recombinant nanopore protein complex
[0153] Current signal was collected by patch-clamp amplifier. The single-channel nanopore detection system based on patch-clamp and signal amplifier was built according to the method disclosed in the literature (Ji Z, Guo P. Channel from bacterial virus T7 DNA packaging motor for the differentiation of peptides composed of a mixture of acidic and basic amino acids. Biomaterials. 2019 May 21; 214: 119222). Ag / AgCl electrodes were immersed in the sequencing buffer and the electrodes were located in the cis and trans regions of the electrolytic cell, respectively. The double detector pore proteins (i.e. the pore protein complex obtained by recombinant in Example Four, here we will list the individual BCP58_1 and BCP58_2 and two complexes BCP58_1-AP58-N_30-55, BCP58_2-AP58-N_30-64) were diluted by a certain multiple using 1x PBS buffer. Under the action of an applied electric field, a single nanopore protein was inserted into a phospholipid bilayer composed of 1,2-diphytanoyl-sn-glycero-3-phosphocholine (DPhPC), forming a nanopore biosensor. Generally, 0.1 mg / mL protein concentration was used, diluted 100 times or 10 times or other multiples with PBS, and attempts were made. If a certain dilution concentration fails to embed the pore, the dilution multiple needs to be reduced for further attempts until the nanopore protein successfully embeds in the phospholipid bilayer. An applied voltage of 180 mV was applied to obtain the current amplitude value of a single pore protein (the highest value of the current graph shown below is the open pore current of the pore protein without DNA passing through).
[0154] Example Six: Use of nanopore sensor for DNA sequencing
[0155] Preparation of nucleic acid sequence: insert the artificially synthesized sequence SEQ ID NO: 20 into the multiple cloning site of PUC57 plasmid, use the primer combination of SEQ ID NO: 21 and SEQ ID NO: 22 to prepare the 3.5 kb sequence to be sequenced (SEQ ID NO: 20) by PCR amplification.
[0156] Sequence to be sequenced (SEQ ID NO: 20):
[0157] Primer F (SEQ ID NO: 21): gccatcagattgtgtttgttagt;
[0158] Primer R (SEQ ID NO: 22): gcttacggttcactactcacga.
[0159] Preparation of sequencing library: After the annealing of the sense strand of two partially region-complementary DNA strands (SEQ ID NO: 16-(iSp18)4-SEQ ID NO: 23, wherein iSp18 is a spacer) and the antisense strand (SEQ ID NO: 17), the linker was connected with the double-stranded target fragment PUC57 (SEQ ID NO: 20) to be detected using T4 DNA ligase at room temperature and purified to prepare the sequencing library. Then the sequencing library was incubated with the helicase BCH105 (SEQ ID NO: 18) at 25°C for 1 h (molar concentration ratio 1:8) to form a sequencing library containing the BCH105 motor protein (structure as shown in Figure 7A). During sequencing, the sequencing library can further complementarily pair and bind with the single-stranded DNA (SEQ ID NO: 19, with cholesterol connected at the 5' end of the DNA) containing cholesterol to form the structure shown in Figure 7B.
[0160] Linker sequence sense strand: SEQ ID NO: 16-(iSp18)4-SEQ ID NO: 23, wherein SEQ ID NO: 16: ttttttttttttttttttttttttttttttttttttttt; iSp18 is a spacer, which can be purchased commercially, and its structure is shown in Figure 14; SEQ ID NO: 23: ggttgtttctgttggtgctgatattgct.
[0161] Antisense strand of linker sequence (SEQ ID NO: 17): gcaatatcagcaccaacagaaacaacctttgaggcgagcggtcaa.
[0162] Amino acid sequence of helicase BCH105 (SEQ ID NO: 18):
[0163] Single-stranded DNA sequence with cholesterol (SEQ ID NO: 19): ttgaccgctcgcctc, with cholesterol modification at the 5' end.
[0164] The sequencing library and single-stranded DNA with cholesterol were mixed with sequencing buffer (0.47M KCl, 25mM HEPES, 1mM EDTA, 5mM ATP, 25mM MgCl2, pH7.6) and added to the nanopore biosensor obtained in Example 5; after applying an external voltage of 0.14V or 0.18V, it was observed that the DNA was captured by the nanopore, generating a characteristic current amplitude value of the block, as shown in Figures 8-13. As can be seen from the figures, as the DNA moves through the nanopore, the current amplitude value changes. Different DNA sequences produce different current amplitude values of the block. The single-stranded DNA with cholesterol can bind to the phospholipid bilayer, which helps the nanopore to capture the sequencing library and reduces the amount of sequencing library loaded.
[0165] Specifically, Figure 8 is the current change and local details generated when the library DNA passes through the nanopore protein BCP58_1 under the action of an external voltage of 0.18V. Figure 9 is the current change and local details generated when the library DNA passes through the nanopore protein complex BCP58_1-AP58-N_30-55 under the action of an external voltage of 0.18V. Figure 10 is the current change and local details generated when the library DNA passes through the nanopore protein BCP58_2 under the action of an external voltage of 0.18V. Figure 11 is the current change and local details generated when the library DNA passes through the nanopore protein complex BCP58_2-AP58-N_30-64 under the action of an external voltage of 0.18V. Figure 12 shows the current signals of the homopolymer passing through BCP58_1 and BCP58_1-AP58-N_30-55, respectively. Figure 13 shows the current signals of the homopolymer passing through BCP58_2 and BCP58_2-AP58-N_30-64, respectively.
[0166] It can be seen that under an external voltage of 0.18V, the open pore current of the nanopore protein BCP58_1 is 340pA, and the sequencing amplitude is about 100pA; the open pore current of the nanopore protein complex BCP58_1-AP58-N_30-55 is 220pA, and the sequencing amplitude is about 50pA; the open pore current of the mutant of the nanopore protein BCP58_2 is 230pA, and the sequencing amplitude is about 50pA; the open pore current of the nanopore protein complex BCP58_2-AP58-N_30-64 is 140pA, and the sequencing amplitude is about 20pA. The sequencing signal of the pore complex to the homopolymer region is different from the single nanopore protein BCP58, which exhibits more current change details and more step signals. Each step can be understood as a current signal generated by several bases passing through the pore at the same time, and the number of steps reflects the richness of the current signal generated by the base combination. The more bases that can be analyzed, the higher the accuracy, that is, the resolution is improved. It can be seen that the pore protein complex of the present application has the ability to improve the resolution of nanopore sequencing.
[0167] It should be noted here that the homopolymer shown in FIG. 12 and FIG. 13 refers to the repeated sequence at the beginning and end of the sequence to be sequenced, TTTTTTTTTTGGAATTTTTTTTTTGGAATTTTTTTTTT (SEQ ID NO: 31).
[0168] From the above description, it can be seen that the above-mentioned embodiments of the present application achieve the following technical effects: the present application provides a new pore protein complex that can be used for nanopore sequencing, and by adding the auxiliary protein AP58 or a variant on the basis of the wild-type BCP58 pore protein monomer variant, a nanopore protein complex with at least two sensors is formed.
[0169] The nanopore sensor composed of the protein or its mutant has at least two sensors (i.e. constriction regions in the nanopore channel structure), has a more optimal single-base recognition effect for nanopore sequencing, and thus can provide higher resolution for repeated sequences, especially homopolymers, meet the high-precision requirements of single-molecule nanopore sequencing, and realize the detection of biological small molecules such as nucleotides, amino acids, sugars, and vitamins, and can also be used for sequencing modified or unmodified DNA, RNA, or polypeptides.
[0170] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A nanopore protein complex, characterized in that, The nanopore protein complex comprises: a nanopore protein comprising a plurality of pore protein monomers and a nanopore channel structure formed by polymerization of the plurality of pore protein monomers; a plurality of accessory proteins, each of which has its N-terminus located within the nanopore channel structure, the plurality of accessory proteins being connected to the plurality of pore protein monomers one-to-one and collectively forming a continuous channel with the nanopore channel structure; wherein, in the direction of a target substance passing through the continuous channel, the continuous channel comprises a first constriction region and a second constriction region connected in sequence, the first constriction region being formed by at least part of the nanopore protein, and the second constriction region being formed by part or all of the accessory proteins; the pore protein monomers are selected from any one or more of the following: (a) a protein having the amino acid sequence shown in SEQ ID NO: 1; or (b) a protein having substitution, deletion and / or addition of one or several amino acids at at least one of the following positions of SEQ ID NO: 1: 77, 81, 82, 176, 210, 214, 232, 66, 69, 70, 74, 109, 110, 113, 117, 118, 119, 120, 123, 127, 128, 168, 211, 221, 224, 227, 229, and having the function of being polymerized to form the nanopore channel structure; or (c) a protein having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity to the protein described in (a) or (b), and having the function of being polymerized to form the nanopore channel structure.
2. The nanopore protein complex of claim 1, wherein, The accessory proteins are selected from any one of the following: i) a polypeptide fragment having the amino acid sequence of any one of SEQ ID NOs: 3-8; ii) a polypeptide fragment having a signal peptide sequence added to the N-terminus of the amino acid sequence of any one of SEQ ID NOs: 3-8; iii) a protein having substitution and / or deletion and / or addition of one or several amino acids at at least one of the following positions of the amino acid sequence of any one of SEQ ID NOs: 3-8: L46, T49, N53, K62, D63, or P64; or iv) a polypeptide fragment having at least 50% identity to the amino acid sequence of any one of the polypeptide fragments described in i), ii), or iii).
3. The nanopore protein complex of claim 2, wherein, In iv), the polypeptide fragment having at least 50% identity to the amino acid sequence of any one of the polypeptide fragments described in i), ii), or iii) has the amino acid sequence shown in SEQ ID NO:
2.
4. The nanopore protein complex of claim 2, wherein, When the accessory proteins are selected from iii), L46 is substituted with S, N, T, V, or Q; T49 is substituted with N, S, V, L, A, or Q; N53 is substituted with S, V, L, A, Q, or T; K62 is substituted with N, S, V, L, A, I, Q, T, P, Y, or W; D63 is substituted with N, S, V, L, A, I, Q, T, P, Y, or W; P64 is substituted with N, S, V, L, A, I, Q, T, Y, or W.
5. The nanopore protein complex according to any one of claims 1 to 4, wherein, In the (b), the types of the substituted amino acids are each independently selected from the following: S77 is mutated to S77G, S77A, or S77T; S81 is mutated to S81G, S81A, S81T, S81N, or S81Q; F82 is mutated to F82G, F82A, F82S, F82T, F82N, or F82Q; E176 is mutated to E176A, E176G, E176V, E176L, E176I, E176Y, E176F, or E176W; K210 is mutated to K210A, K210G, K210V, K210L, K210I, K210Y, K210F, or K210W; S214 is mutated to S214A, S214G, S214V, S214L, S214I, S214Y, S214F, or S214W; T232 is mutated to T232A, T232G, T232V, T232L, T232I, T232Y, T232F, or T232W; K66 is mutated to K66N, K66A, K66G, K66S, K66T, or K66Q; D69 is mutated to D69K, D69R, D69N, D69A, D69G, D69S, D69T, or D69Q; Q70 is mutated to Q70K, Q70R, Q70D, Q70E, Q70N, Q70A, Q70G, Q70S, Q70T, or Q70Q; Y74 is mutated to Y74N, Y74A, Y74G, Y74S, Y74T, or Y74Q; R109 is mutated to R109N, R109A, R109G, R109S, R109T, or R109Q; E110 is mutated to E110K, E110R, E110N, E110A, E110G, E110S, E110T, or E110Q; Q113 is mutated to Q113K, Q113R, Q113N, Q113A, Q113G, Q113S, or Q113T; T117 is mutated to T117K, T117R, T117N, T117A, T117G, T117S, or T117Q; E118 is mutated to E118K, E118R, E118N, E118A, E118G, E118S, E118T, or E118Q; R119 is mutated to R119N, R119A, R119G, R119S, R119T, or R119Q; K120 is mutated to K120N, K120A, K120G, K120S, K120T, or K120Q; R123 is mutated to R123N, R123A, R123G, R123S, R123T, or R123Q; K127 is mutated to K127N, K127A, K127G, K127S, K127T, or K127Q; K128 is mutated to K128N, K128A, K128G, K128S, K128T, or K128Q; R168 is mutated to R168N, R168Q, R168S, R168T, R168A or R168G; E211 is mutated to E211N, E211Q, E211S, E211T, E211A or E211G; E221 is mutated to E221N, E221Q, E221S, E221T, E221A or E221G; E224 is mutated to E224N, E224Q, E224S, E224T, E224A or E224G; E227 is mutated to E227N, E227Q, E227S, E227T, E227A or E227G; E229 is mutated to E229N, E229Q, E229S, E229T, E229A or E229G.
6. The nanopore protein complex of claim 5, wherein, In the (b), the types of the substituted amino acids are each independently selected from the following: S77A, F82Q, S81N, F82A, K120N, K127S, E176I; Preferably, the amino acid sequence of the pore protein monomer is selected from: an amino acid sequence substituted with F82Q on the basis of SEQ ID NO: 1; or an amino acid sequence substituted with S77A+S81N+F82Q+K127S on the basis of SEQ ID NO:
1.
7. The nanopore protein complex according to any one of claims 1 to 6, wherein, The pore protein monomer is selected from a protein having any one of the amino acid sequences in SEQ ID NO: 24-SEQ ID NO:
28.
8. The nanopore protein complex of claim 1, wherein, The plurality of the auxiliary proteins are covalently connected to the plurality of the pore protein monomers one-to-one through covalent bonds; or The plurality of the auxiliary proteins are connected to the plurality of the pore protein monomers one-to-one through non-covalent interactions.
9. The nanopore protein complex of claim 8, wherein, The plurality of the auxiliary proteins and the plurality of the pore protein monomers comprise a cysteine mutation at any one of the following sites or an unnatural amino acid introduced at any one of the following sites to achieve the covalent connection: The amino acid sites of the pore protein monomer include, based on SEQ ID NO: 1: D181, R179, N159, E227, Q213, R168, F219, Y217 or L170; The amino acid sites of the auxiliary protein include, based on SEQ ID NO: 2: S30, E31, L32, V33, V37, N38, S40, A55, Q58 or Q59.
10. The nanopore protein complex of claim 8 or 9, wherein, The covalent connection comprises any one or more of the following: spontaneous cross-linking, cross-linking agent connection, protein fusion connection, polypeptide molecule connection or chemical small molecule connection.
11. The nanopore protein complex according to any one of claims 1 to 10, wherein, The pore channel diameter of the first constriction region is 13-21 angstroms, and the pore channel diameter of the second constriction region is 15-25 angstroms.
12. A nanopore sensor, characterized in that, The nanopore sensor comprises a membrane and the nanopore protein complex of any one of claims 1-11, wherein the nanopore protein in the nanopore protein complex is embedded in the membrane, and the nanopore channel structure of the nanopore protein and the plurality of auxiliary proteins together form a continuous channel across the membrane; When a voltage is applied across the membrane, the continuous channel generates an electrical signal.
13. The nanopore sensor of claim 12, wherein, The membrane comprises a layer of amphiphilic molecules; Preferably, the membrane is a phospholipid bilayer consisting of diacylphosphatidylcholine.
14. The nanopore sensor of claim 12, wherein, When a voltage is applied across the membrane, a biomolecule translocates through a continuous channel in the nanopore sensor, the continuous channel generating an electrical signal that varies; Preferably, the biomolecule comprises a polynucleotide, a polypeptide, or a polysaccharide. Preferably, the biomolecule is selected from a polynucleotide containing a homopolymer.
15. A kit comprising, The kit comprises the nanopore protein complex of any one of claims 1 to 11 or the nanopore sensor of any one of claims 12 to 14, the kit further comprising one or more optional components: a membrane, a sequencing buffer, a nuclease, a polymerase, a topoisomerase, a ligase, a helicase, and a single-stranded DNA linked to a cholesterol.
16. A nanopore sequencing system, characterized in that, The nanopore sequencing system comprises the nanopore sensor of any one of claims 12 to 14, the nanopore sequencing system further comprising: an electrically conductive solution; positive and negative electrodes to provide a voltage across the membrane, and a measurement device to measure an electrical signal through the continuous channel; wherein the nanopore sensor is in the electrically conductive solution and partitions the electrically conductive solution into a first chamber and a second chamber, the negative electrode being in the first chamber and the positive electrode being in the second chamber.
17. The nanopore sequencing system of claim 16, wherein, The positive and negative electrodes are connected to a signal processing chip to measure an electrical signal through the continuous channel.
18. The nanopore sequencing system of claim 16, wherein, The positive and negative electrodes comprise a metal electrode material or a composite electrode material.
19. A method of identifying or characterizing a biomolecule, characterized by, The method comprises: the nanopore protein complex of any one of claims 1 to 11, or the nanopore sensor of any one of claims 12 to 14, or the nanopore sequencing system of any one of claims 16 to 18, identifying or characterizing the biomolecule by measuring and resolving an electrical signal generated by the biomolecule as it translocates through the continuous channel of the nanopore protein complex.
20. The method of claim 19, wherein, The method comprises: contacting the nanopore sequencing system of any one of claims 16 to 18 with the biomolecule, applying a voltage across the membrane to cause the biomolecule to enter the continuous channel and move relative to the continuous channel; and measuring an electrical signal generated by the biomolecule as it moves relative to the continuous channel, thereby identifying or characterizing the biomolecule.
21. The method of claim 19 or 20, wherein, The biomolecule comprises a polynucleotide, a polypeptide, or a polysaccharide; Preferably, the biomolecule is selected from a polynucleotide containing a homopolymer. Preferably, the electrical signal comprises an electrical current.
22. The method of claim 19, wherein, The identifying the biomolecule comprises identifying whether the biomolecule is present; the characterizing the biomolecule comprises determining a composition of the biomolecule; Preferably, determining the composition of the biomolecule comprises determining a nucleotide sequence of a polynucleotide, an amino acid sequence of a polypeptide, or an order of monosaccharides of a polysaccharide.
23. The method of claim 19, wherein, The biomolecule is a polynucleotide, nucleotides in the polynucleotide interact with the first constriction region and the second constriction region within the continuous channel, and wherein each of the first constriction region and the second constriction region is capable of distinguishing different nucleotides.
24. The method of claim 23, wherein, A nucleic acid binding protein is used to control movement of the polynucleotide relative to the continuous channel.
25. Use of the nanopore protein complex of any one of claims 1 to 11, or the nanopore sensor of any one of claims 12 to 14, or the kit of claim 15, or the nanopore sequencing system of any one of claims 16 to 18 in biological small molecule detection, polynucleotide sequencing, polypeptide sequencing or polysaccharide sequencing.
26. A method of preparing a nanopore complex according to any one of claims 1 to 11, characterised in that, The preparation method comprises: respectively constructing an expression vector of a nanopore protein monomer and an expression vector of an auxiliary protein, co-expressing the pore protein monomer and the auxiliary protein in a competent cell, isolating and purifying the pore protein monomer and the auxiliary protein expressed by the competent cell mixing the pore protein monomer and the auxiliary protein to obtain the nanopore complex; or constructing an expression vector of a nanopore protein monomer, expressing the pore protein monomer in a competent cell, isolating and purifying the pore protein monomer expressed by the competent cell, and mixing the pore protein monomer and an auxiliary protein artificially synthesized in vitro to obtain the nanopore complex.
Citation Information
Patent Citations
Pore
CN113195736A
Single molecule nanopore sequencing method
WO2023125605A1
De novo pores
WO2024033447A1
Porin monomer, porin, mutant thereof, and use thereof
WO2024138473A1