Nanopore protein complex, construction method therefor and use thereof

By constructing a nanopore protein complex, the number of sensing regions was increased, which solved the problem of low accuracy in nanopore sequencing, enabled high-resolution sequencing of homopolymer sequences, and improved sequencing accuracy.

WO2025255780A1PCT designated stage Publication Date: 2025-12-18SHENZHEN HUADA GENE INST
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/099017
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-12-18

AI Technical Summary

Technical Problem

In existing nanopore sequencing technologies, nanopore proteins have only a single contractile region, resulting in low sequencing accuracy. In particular, when processing homopolymer sequences, it is difficult to accurately distinguish repetitive bases, which affects sequencing accuracy and resolution.

Method used

A nanoporous protein complex was constructed by polymerizing nanoporous proteins with accessory proteins to form a continuous channel with multiple contraction zones. The accessory proteins were then embedded within the nanoporous protein cavity to increase the number of sensing regions, thus forming a nanoporous complex with two or more contraction zones.

Benefits of technology

It improves the accuracy of nanopore sequencing, especially the resolution of longer homopolymer sequences, outputs stable sequencing signals, and addresses the insufficient sequencing accuracy in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024099017_18122025_PF_FP_ABST
    Figure CN2024099017_18122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a nanopore protein complex, a construction method therefor, and a use thereof. The nanopore complex comprises: a nanopore protein and an accessory protein. The nanopore protein comprises a nanopore cavity, and the nanopore cavity is formed by polymerizing a plurality of pore protein monomers. The accessory protein is formed by polymerizing a plurality of accessory protein monomers. The N-terminus of the accessory protein is embedded in the nanopore cavity and forms a continuous channel together with the nanopore cavity. On the basis of the moving direction of an analyte passing through the continuous channel, the continuous channel comprises a first sensing region and a second sensing region which are communicated sequentially, wherein the first sensing region is formed by part of the nanopore protein, and the second sensing region is formed by part or all of the accessory protein. A nanopore protein monomer is selected from a protein having an amino acid sequence as shown in SEQ ID NO: 1 and a variant thereof, and an accessory protein monomer is selected from a protein having an amino acid sequence as shown in SEQ ID NO: 3 and a variant thereof.
Need to check novelty before this filing date? Find Prior Art

Description

Nanopore protein complex, construction method and application thereof TECHNICAL FIELD

[0001] The present application relates to the technical field of nanopore sequencing, in particular to a nanopore protein complex, a construction method and application thereof. BACKGROUND

[0002] As a newly emerging single-molecule sequencing technology, nanopore sequencing technology has brought a revolutionary change to the genetic sequencing industry with its unique advantages of high throughput, long read length, fast speed, in-situ detection and label-free operation, and has a wide range of applications in basic theoretical research and biomedical clinical practice in many fields such as molecular biology, medicine, epidemiology and ecology.

[0003] Nanopore sequencing technology is a sequencing technology based on electrical signals. This technology can be used to test the sequence of nucleotides, amino acids or glycans, as well as base, amino acid or glycan modifications (such as methylation and acylation, phosphorylation, hydroxylation, oxidation, reduction, glycosylation, decarboxylation, deamination, etc.). The core element of nanopore protein is inserted into the membrane to play the role of a signal sensor, which separates two electrolytic chambers containing electrolyte. When a voltage is applied between the electrolytic chambers, a stable current will be generated. When the analyte to be tested enters the nanopore, it will hinder the flow of ions, causing fluctuations in the current signal. By recording the continuous blocking electrical signals generated when the analyte passes through the nanopore protein one by one in real time, and with the help of machine learning analysis and decoding of the current signal, the sequence information of the analyte to be tested can be sequenced in real time.

[0004] The theoretical basis of nanopore sequencing is that nucleic acid molecules pass through the central channel of the nanopore protein embedded in the membrane under the control of the motor protein, and the current signal generated is amplified by the underlying circuit and converted into final sequence information by the base recognition algorithm. This route has been fully demonstrated and is feasible. The commercially available nanopore sequencers on the market are mainly a series of nanopore sequencers such as MinION, GridION and PromethION developed by Oxford Nanopore Technologies in the UK, and QNome-3841 nanopore sequencer developed by Zhi Carbon Technology. However, these nanopore sequencers still have great deficiencies in sequencing accuracy, throughput and chip stability, and cannot meet the ultimate needs of molecular biology research. Therefore, there is a need in the art for a single-molecule sequencer with high accuracy, high integration and high stability.

[0005] One of the main reasons for the low accuracy of nanopore sequencing is the single sensor and low resolution of the core element nanopore protein. The narrowest constriction region inside the nanopore protein is the most distinctive part of the analyte current signal change characteristics, which can be used as an internal sensor region of the current signal. This region needs to be sharp enough to have high spatial resolution in both lateral and longitudinal directions.

[0006] Among the studied pore proteins, there are mainly three types of proteins that can be potentially used for sequencing: transporter proteins that act as transport channels for various biological macromolecules and small molecules inside and outside cells; pore-forming toxin proteins that destroy the permeability of the cell membrane; and viral connectors that provide a genome transport channel for viruses to infect hosts. Currently, only a few natural proteins such as Mycobacterium smegmatis pore protein A (MspA) and curli-specific transport channel (CsgG) meet the sequencing requirements. The reasons for the difficulty of sequencing natural proteins include the test of protein stability in the in vitro expression and purification system of recombinant protein, the stability and symmetry of protein polymer, the shape and pore size of the protein lumen constriction region, etc.

[0007] Most of the pore proteins on the market have only one constriction region and can only contact a few parts of the analyte sequence. During sequencing, the sequences interact with each other, and the conductance map of the underlying sequence is very complex. Especially for homopolymer sequences, the current algorithm is difficult to accurately distinguish, which is one of the main reasons for the poor resolution of current nanopore sequencers, resulting in a relatively high error rate and slow popularization. Existing literature has shown that when there is more than one sensor in a single nanopore, it can help obtain additional sequence information, provide more opportunities to analyze homopolymer regions, and overcome the disadvantage of low accuracy of nanopore sequencing. Currently, there are only a few pore proteins with two sensors, but the sensor spacing is too short to provide sufficient spatial resolution. Especially in testing DNA homopolymers, the current nanopore signal often cannot accurately determine the number of repeated bases in the region, thereby reducing the accuracy of nanopore sequencing.

[0008] Therefore, it is urgent to develop nanopore proteins with more than two or more constriction regions and achieve high-precision sequencing.

[0009] SUMMARY

[0010] The main purpose of the present application is to provide a nanopore protein complex to solve the problem of low sequencing accuracy of nanopore proteins with only a single constriction region in the prior art when applied to nanopore sequencing.

[0011] In order to achieve the above object, according to one aspect of the present application, there is provided a nanopore protein complex, comprising: a nanopore protein, the nanopore protein being polymerized from a plurality of pore protein monomers, and the plurality of pore protein monomers being polymerized to form a hollow nanopore cavity; and an accessory protein, the accessory protein being polymerized from a plurality of accessory protein monomers, at least part of the accessory protein being located within the nanopore cavity, the accessory protein and the nanopore cavity jointly forming a continuous channel; wherein, according to a moving direction of an analyte passing through the continuous channel, the continuous channel comprises a first sensing region and a second sensing region connected in sequence, the first sensing region being formed by at least part of the nanopore protein, and the second sensing region being formed by at least part of the accessory protein; the pore protein monomer is selected from any one of the following proteins: 1) a protein having the amino acid sequence shown in SEQ ID NO: 1; 2) a protein having at least 50% identity with the amino acid sequence shown in SEQ ID NO: 1 and having a function of being polymerized to form the nanopore protein; or 3) a protein having one or more amino acids substituted, deleted or added on the basis of SEQ ID NO: 1 and having a function of being polymerized to form the nanopore protein; and the accessory protein monomer is selected from any one of the following proteins: i) a protein having the amino acid sequence shown in SEQ ID NO: 3; ii) a protein having at least 50% identity with the amino acid sequence shown in SEQ ID NO: 3 and having a function of being polymerized to form the accessory protein together with the pore protein monomers to form the nanopore protein; or iii) a protein having one or more amino acids substituted, deleted or added on the amino acid sequence shown in SEQ ID NO: 3 and having a function of being polymerized to form the accessory protein together with the pore protein monomers to form the nanopore protein.

[0012] Further, the nanopore protein and the accessory protein are derived from the same species; optionally, all or an N-terminal part of the accessory protein is located within the nanopore cavity of the nanopore protein; preferably, the accessory protein is attached within the nanopore cavity of the nanopore protein by covalent or non-covalent action; more preferably, the nanopore protein and the accessory protein exist in the form of a nonamer.

[0013] Further, the accessory protein monomer is a truncated mutant of the amino acid sequence shown in SEQ ID NO: 3, the truncated mutant being selected from any one of the following mutants: only retaining N-terminal 23-67 amino acids, only retaining N-terminal 23-57 amino acids, only retaining N-terminal 23-52 amino acids, only retaining N-terminal 23-54 amino acids, only retaining N-terminal 23-51 amino acids, only retaining N-terminal 23-50 amino acids, only retaining N-terminal 23-49 amino acids, only retaining N-terminal 23-48 amino acids, or only retaining N-terminal 23-47 amino acids.

[0014] Further, the length of the accessory protein monomer is 24-45 amino acids; preferably, the amino acid sequence of the accessory protein monomer is from the following residue position interval of SEQ ID NO: 3 or a mutant thereof: 23-48, 23-49, 23-50, 23-51, 23-54, or 23-57.

[0015] Further, the pore protein monomer is selected from proteins mutated at at least one amino acid position in any one or more of the following groups of SEQ ID NO: 1: 1) S71, N74, G75, and F76; preferably, S71 is mutated to G, A, or T; N74 is mutated to G, A, S, or T; G75 is mutated to A, S, T, or Q; F76 is mutated to A, S, T, N, or Q; 2) E162, R196, S200, and S216; preferably, E162 is mutated to A, G, V, L, I, Y, F, or W; R196 is mutated to A, G, V, L, I, Y, F, or W; S200 is mutated to A, G, V, L, I, Y, F, or W; S216 is mutated to A, G, V, L, I, Y, F, or W; 3) R103, E104, E112, R113, K114, R117, R119, D120, K122, D124, K154, R165, D199, R207, K209, K210, E213, or E215; preferably, R103 is mutated to A, G, S, T, N, or Q; E104 is mutated to K, R, A, G, S, T, N, or Q; E112 is mutated to K, R, A, G, S, T, N, or Q; R113 is mutated to A, G, S, T, N, or Q; K114 is mutated to A, G, S, T, N, or Q; R117 is mutated to A, G, S, T, N, or Q; R119 is mutated to A, G, S, T, N, or Q; D120 is mutated to K, R, A, G, S, T, N, or Q; K122 is mutated to A, G, S, T, N, or Q; D124 is mutated to K, R, A, G, S, T, N, or Q; K154 is mutated to A, G, S, T, N, or Q; R165 is mutated to A, G, S, T, N, or Q; D199 is mutated to A, G, S, T, N, or Q; R207 is mutated to A, G, S, T, N, Q, D, or E; K209 is mutated to A, G, S, T, N, Q, D, or E; K210 is mutated to A, G, S, T, N, Q, D, or E; E213 is mutated to A, G, S, T, N, or Q; E215 is mutated to A, G, S, T, N, or Q.

[0016] Further, the accessory protein monomer is selected from proteins mutated at at least one amino acid position in any one or more of the following groups of SEQ ID NO: 1: 1) S71, N74, G75, and F76; preferably, S71 is mutated to G, A, or T; N74 is mutated to G, A, S, or T; G75 is mutated to A, S, T, or Q; F76 is mutated to A, S, T, N, or Q; 2) E162, R196, S200, and S216; preferably, E162 is mutated to A, G, V, L, I, Y, F, or W; R196 is mutated to A, G, V, L, I, Y, F, or W; S200 is mutated to A, G, V, L, I, Y, F, or W; S216 is mutated to A, G, V, L, I, Y, F, or W; 3) R103, E104, E112, R113, K114, R117, R119, D120, K122, D124, K154, R165, D199, R207, K209, K210, E213, or E215; preferably, R103 is mutated to A, G, S, T, N, or Q; E104 is mutated to K, R, A, G, S, T, N, or Q; E112 is mutated to K, R, A, G, S, T, N, or Q; R113 is mutated to A, G, S, T, N, or Q; K114 is mutated to A, G, S, T, N, or Q; R117 is mutated to A, G, S, T, N, or Q; R119 is mutated to A, G, S, T, N, or Q; D120 is mutated to K, R, A, G, S, T, N, or Q; K122 is mutated to A, G, S, T, N, or Q; D124 is mutated to K, R, A, G, S, T, N, or Q; K154 is mutated to A, G, S, T, N, or Q; R165 is mutated to A, G, S, T, N, or Q; D199 is mutated to A, G, S, T, N, or Q; R207 is mutated to A, G, S, T, N, Q, D, or E; K209 is mutated to A, G, S, T, N, Q, D, or E; K210 is mutated to A, G, S, T, N, Q, D, or E; E213 is mutated to A, G, S, T, N, or Q; E215 is mutated to A, G, S, T, N, or Q.

[0017] Further, the first and second sensing zones have the same or different minimum pore diameters; preferably, the minimum pore diameter is 1-3 nm.

[0018] Further, the accessory protein monomer is selected from a protein in which at least one of the amino acids S23, S24, L25, T28, K30, N31, S33, F34, N47, A50, Q51, N52, or Q53 in the amino acid sequence shown in SEQ ID NO: 3 is mutated to cysteine or substituted by an unnatural amino acid; and / or the pore protein monomer is selected from a protein in which at least one of the amino acids N145, T148, K154, L156, L160, S161, R165, S195, Q197, D199, F203, Y205, K209, K210, L211, E213, E215, G217, S219, or N221 in SEQ ID NO: 1 is mutated to cysteine or substituted by an unnatural amino acid.

[0019] Further, the pore protein monomer and the accessory protein monomer have at least one of the following site combinations mutated, and the amino acid of the corresponding site in each site combination is mutated to cysteine or substituted by a non-natural amino acid: 1) R165 on SEQ ID NO: 1 and S23 on SEQ ID NO: 3; 2) S195 on SEQ ID NO: 1 and S24 on SEQ ID NO: 3; 3) S219 on SEQ ID NO: 1 and L25 on SEQ ID NO: 3; 4) N221 on SEQ ID NO: 1 and L25 on SEQ ID NO: 3; 5) N145C on SEQ ID NO: 1 and V26C on SEQ ID NO: 3; 6) D199 on SEQ ID NO: 1 and T28 on SEQ ID NO: 3; 7) D199 on SEQ ID NO: 1 and K30 on SEQ ID NO: 3; 8) E215 on SEQ ID NO: 1 and K30 on SEQ ID NO: 3; 9) E213 on SEQ ID NO: 1 and N31 on SEQ ID NO: 3; 10) K154 on SEQ ID NO: 1 and S33 on SEQ ID NO: 3; 11) E213 on SEQ ID NO: 1 and S33 on SEQ ID NO: 3; 12) E213 on SEQ ID NO: 1 and N47 on SEQ ID NO: 3; 13) L211 on SEQ ID NO: 1 and A50 on SEQ ID NO: 3; 14) L211 on SEQ ID NO: 1 and Q51 on SEQ ID NO: 3; 15) K209 on SEQ ID NO: 1 and N52 on SEQ ID NO: 3; 16) K210 on SEQ ID NO: 1 and Q53 on SEQ ID NO: 3.

[0020] Further, the amino acid of the corresponding site in each site combination is mutated to cysteine.

[0021] Further, the amino acid substitution is a non-natural amino acid by modification, which is direct modification or indirect modification; preferably, the direct modification includes modification by spontaneous reaction or oxidation reaction of the side chain group of the amino acid; preferably, the indirect modification includes chemical modification by attaching a chemical small molecule; preferably, the chemical small molecule includes a chemical crosslinking agent containing a functional group, and the functional group is resistant to dithiothreitol.

[0022] Further, the site mutated to cysteine in at least part of the accessory protein monomer forms a disulfide bond with the site mutated to cysteine in at least part of the pore protein monomer.

[0023] To achieve the above object, according to a second aspect of the present application, there is provided a nanopore sensor comprising a membrane and any one of the nanopore protein complexes described above, wherein the nanopore protein in the nanopore protein complex is positioned on the membrane, and the nanopore cavity of the nanopore protein and the accessory protein together form a continuous channel across the membrane.

[0024] Further, the membrane comprises a layer of amphiphilic molecules; preferably, the membrane is a phospholipid bilayer consisting of dipalmitoyl phosphatidylcholine, a diblock copolymer or a triblock copolymer.

[0025] According to a third aspect of the present application, there is provided a nanopore sequencing system comprising any one of the nanopore sensors described above, the nanopore sequencing system further comprising: a conductive solution, positive and negative electrodes for providing a voltage potential across the membrane, and a measuring device for measuring an electrical signal passing through the continuous channel; wherein the nanopore sensor is located in the conductive solution, and the conductive solution is divided into a first chamber and a second chamber.

[0026] Further, the nanopore sequencing system further comprises a test molecule selected from the group consisting of polynucleotides, polypeptides and polysaccharides, preferably, the test molecule is selected from the group consisting of polynucleotides containing homopolymers; preferably, the test molecule can be transiently located within the continuous channel, and one end of the test molecule is located in the first chamber and the other end is located in the second chamber.

[0027] According to a fourth aspect of the present application, there is provided a method of nanopore sequencing, the method comprising: contacting a nanopore sequencing system with a test molecule; applying an electric potential across the membrane, such that the test molecule enters the continuous channel; and measuring one or more times an electrical signal generated as the test molecule moves relative to the continuous channel, thereby obtaining a sequence of the test molecule.

[0028] Further, the electrical signal comprises a current block signal intensity, a current block duration or an interval time of current block events.

[0029] Further, the test molecule is a polynucleotide, the nucleotides in the polynucleotide interact with the first constriction region and the second constriction region within the continuous channel, and wherein each of the first constriction region and the second constriction region is capable of distinguishing different nucleotides, such that the total current passing through the continuous channel is affected by the interaction between each of the first constriction region and the second constriction region and the nucleotides positioned at each of the regions.

[0030] Further, a nucleic acid binding protein is used to control the movement of the polynucleotide relative to the continuous channel pore.

[0031] According to a fourth aspect of the present application, there is provided a method for preparing the nanopore protein complex as described above, comprising: separately constructing an expression vector for the pore protein monomer and an expression vector for the accessory protein monomer; obtaining the nanopore protein complex by co-expressing the expression vector for the pore protein monomer and the expression vector for the accessory protein monomer in a competent cell and then purifying by separation; or constructing the nanopore protein complex by in vitro recombination.

[0032] By using the technical solution of the present application, the accessory protein is embedded in the nanopore cavity of the nanopore protein, and a second constriction region is introduced on the basis of the one constriction region of the nanopore protein itself, forming a nanopore complex with a continuous channel of two constriction regions (or sensing regions). When such a complex is used for nanopore sequencing, the nucleotides at different spatial positions of the nucleic acid molecule to be measured will produce different interactions with the two constriction regions when passing through the first constriction region and the second constriction region in turn, and thus more electrical signal (such as current) change information is generated. Therefore, compared with the nanopore protein with only one constriction region, the nanopore protein complex of the present application can more accurately determine the information of homopolymer sequences, especially for longer homopolymer sequence fragments. BRIEF DESCRIPTION OF DRAWINGS

[0033] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application and are incorporated in and constitute a part of this application. The embodiments of the present application illustrated in the drawings and their descriptions are used to explain the present application and are not intended to limit the present application. In the drawings:

[0034] Figure 1 shows a schematic diagram of the predicted structure of the multimer formed by (signal peptide-containing) BCP34-AP34 in Example 1 of the present application, where A is a side view and B is a top view.

[0035] Figure 2 shows a schematic diagram of the predicted structure of the mature (signal peptide-free) BCP34-AP34 multimer formed in Example 1 of the present application, where A is a side view and B is a top view.

[0036] Figure 3 shows a schematic diagram of the nanopore constriction region in the predicted structure of the multimer formed by BCP34-AP34 in Example 1 of the present application.

[0037] Figure 4 shows a schematic diagram of the nanopore constriction region in the predicted structure of the multimer formed by BCP34-AP34 in Example 1 of the present application.

[0038] Figure 5 shows the purification results of the nanopore complex constructed by co-expression method in embodiment 3 of the application, wherein A is the result of Strep column purification, and B is the result of Ni column purification after Strep column purification.

[0039] Figure 6 shows the purification results of the TEV enzyme cleavage of the nanopore complex constructed by co-expression method in embodiment 3 of the application, wherein A is the result of the nanopore complex without denaturation, and B is the result of the nanopore complex after denaturation.

[0040] Figure 7 shows the purification results of the nanopore complex constructed by in vitro recombination method in embodiment 4 of the application, wherein A is the electrophoretogram of the nanopore complex before denaturation, and B is the electrophoretogram of the nanopore complex after denaturation.

[0041] Figure 8 shows the structural diagram of the interaction region of BCP34 and AP34 in embodiment 5 of the application.

[0042] Figure 9 shows the purification diagram of the cross-linking modified mutant of BCP34 in embodiment 5 of the application; A is the diagram before denaturation, and B is the diagram after denaturation.

[0043] Figure 10 shows the SDS-PAGE electrophoretogram of the nanopore complex constructed by cross-linking modification in embodiment 5 of the application; A shows the reaction efficiency of the nanopore complex formed by BCP34-2 and AP34-P2, and B shows the reaction efficiency of the nanopore complex formed by BCP34-3 and AP34-P3.

[0044] Figure 11 shows the SDS-PAGE electrophoretogram of the nanopore complex constructed by AP34 with different lengths in embodiment 5 of the application; A shows the reaction efficiency of the nanopore complex formed by AP34-P2, AP34-P5, AP34-P6 and AP34-P7 and BCP34-2 respectively; B shows the reaction efficiency of the nanopore complex formed by AP34-P4 and BCP34-2.

[0045] Figure 12 shows the schematic diagram of the helicase sequencing library and the pairing structure of the sequencing library and the single-stranded DNA containing cholesterol in embodiment 7 of the application, wherein A is the schematic diagram of the helicase sequencing library, and B is the pairing structure of the sequencing library and the single-stranded DNA containing cholesterol; a: upper strand; b: lower strand; c: double-stranded target fragment; d: helicase; e: cholesterol-labeled single-stranded DNA.

[0046] Figure 13 shows the sequencing current signal diagram of BCP34-1 nanopore protein in embodiment 7 of the application.

[0047] Figure 14 shows a BCP34-1-AP34-P1 nanopore complex sequencing current signal graph in Example 7 of the application.

[0048] Figure 15 shows a BCP34-1-AP34-P8 nanopore complex sequencing current signal graph in Example 7 of the application.

[0049] Figure 16 shows a BCP34-1-AP34-P9 nanopore complex sequencing current signal graph in Example 7 of the application.

[0050] Figure 17 shows a BCP34-2-AP34-P2 nanopore complex sequencing current signal graph in Example 7 of the application.

[0051] Figure 18 shows a BCP34-3-AP34-P3 nanopore complex sequencing current signal graph in Example 7 of the application.

[0052] Figure 19 shows a BCP34-2-AP34-P4 nanopore complex sequencing current signal graph in Example 7 of the application.

[0053] Figure 20 shows a BCP34-2-AP34-P5 nanopore complex sequencing current signal graph in Example 7 of the application.

[0054] Figure 21 shows a BCP34-2-AP34-P7 nanopore complex sequencing current signal graph in Example 7 of the application.

[0055] Figure 22 shows the characteristics of different nanopores on the current signal of DNA homopolymer region in Example 7 of the application; A is a comparison of the current signal of single constriction region pore BCP34-1 and double constriction region BCP34-1-AP34-1; B is a comparison of the current signal of different AP34 main signal regions; C is a comparison of the modified current signal of different cross-linking sites; D is a comparison of the current signal of AP34 polypeptides of different lengths. DETAILED DESCRIPTION

[0056] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the embodiments.

[0057] Explanation of terms:

[0058] Pore protein, also called nanopore protein in the present application, is a pore protein that can be used for nanopore sequencing.

[0059] Narrow region or sensing region: In the present application, it refers to the narrowest inner diameter region in the transmembrane channel formed by the nanopore protein or the nanopore protein lumen and the accessory protein. Among them, the transmembrane channel formed by the nanopore protein alone has a narrow region, and the transmembrane channel formed by the nanopore protein and the accessory protein together forms other narrow regions.

[0060] As mentioned in the background section, the existing nanopore protein for nanopore sequencing only has one narrow region, which is difficult to meet the needs of distinguishing the sequencing region containing homopolymer (oligonucleotide molecule containing multiple consecutive repeating nucleotides).

[0061] In order to improve this situation, the present application screens or constructs a new type of nanopore protein and its mutant with multiple sensing regions through pore protein database mining, structure prediction and analysis, protein assembly, protein modification, nanopore sequencing performance test and characterization and other technical means. The nanopore protein complex is formed by the pore protein monomer BCP34 and the accessory protein monomer AP34, and is polymerized to form a transmembrane protein with multiple narrow regions. Such nanopore protein complex can be co-expressed in vivo, or can be recombinantly expressed in vitro to form a stable complex with high efficiency. Further, a nanopore sensor is constructed by using the complex, a nanopore sequencing system is constructed, and a polyhomopolymer nucleic acid molecule is detected and verified, which proves that the improved nanopore protein complex containing two or more narrow regions provided by the present application helps to improve the accuracy of distinguishing homopolymer nucleic acid molecules. Thus, when applied to nanopore sequencing, it can output stable sequencing signals and has high resolution for homopolymer sequences, and has high potential in improving the accuracy of sequencing and improving the accuracy of nanopore sequencing.

[0062] On the basis of the above research results, the applicant proposes a series of technical solutions in the present application. In a typical embodiment, a nanopore complex is provided, which includes: a nanopore protein and an accessory protein, the nanopore protein is polymerized by a plurality of pore protein monomers, and the plurality of pore protein monomers are polymerized to form a hollow nanopore lumen; the accessory protein is polymerized by a plurality of accessory protein monomers, at least part of the accessory protein is embedded in the nanopore lumen, and the accessory protein and the nanopore lumen form a continuous channel together; wherein, according to the moving direction of the analyte passing through the continuous channel, the continuous channel includes a first sensing region and a second sensing region connected in sequence, the first sensing region is formed by a part of the nanopore protein, and the second sensing region is formed by part or all of the accessory protein.

[0063] The nanopore protein monomer is selected from any one of the following proteins: 1) a protein having the amino acid sequence shown in SEQ ID NO: 1; 2) a protein having at least 50% identity with the amino acid sequence shown in SEQ ID NO: 1 and having the function of polymerizing to form a nanopore protein; or 3) a protein having one or more amino acids substituted, deleted or added on the basis of SEQ ID NO: 1 and having the function of polymerizing to form a nanopore protein.

[0064] The accessory protein monomer is selected from any one of the following proteins: i) a protein having the amino acid sequence shown in SEQ ID NO: 3; ii) a protein having at least 50% identity with the amino acid sequence shown in SEQ ID NO: 3 and having the function of binding to the pore protein monomer and polymerizing to form the accessory protein together with the pore protein monomer polymerizing to form the nanopore protein described above; or iii) a protein having one or more amino acids substituted, deleted or added to the amino acid sequence shown in SEQ ID NO: 3 and having the function of binding to the pore protein monomer and polymerizing to form the accessory protein together with the pore protein monomer polymerizing to form the nanopore protein described above.

[0065] The constriction region of the nanopore can distinguish different nucleotides (i.e. having different bases) of the nucleic acid to be tested passing through the pore, and each nucleotide produces a current change when interacting with the constriction, which can be converted into corresponding sequence information. However, for DNA homopolymer sequencing, the signal measured by the existing single constriction region may not be sufficient to distinguish and resolve single base current changes, and thus the length of the homopolymer cannot be accurately determined only according to the size of the measured signal. The above-mentioned novel nanopore protein complex provided by the present application helps to improve the distinguishability of the sequencing signal of the nucleotide homopolymer.

[0066] The above-mentioned nanopore complex of the present application utilizes the embedding of the accessory protein into the nanopore cavity of the nanopore protein, and further introduces a second constriction region on the basis of the one constriction region possessed by the nanopore protein itself, forming a nanopore complex with a continuous channel having two constriction regions (or sensing regions). When such a complex is used for nanopore sequencing, the nucleotides at different spatial positions of the nucleic acid molecule sequentially passing through the first and second constriction regions interact differently with the two constriction regions, thereby producing more electrical signal (such as current) change information. Therefore, compared with the nanopore protein having only a single constriction region, the nanopore sequencing using the nanopore protein complex of the present application can more accurately determine the information of the homopolymer sequence, especially for longer homopolymer sequence fragments. That is, the nanopore protein complex can be applied to nanopore sequencing, output stable sequencing signals, and has high resolution for homopolymer sequences, has extremely high potential to improve the accuracy of sequencing and improve the disadvantage of insufficient accuracy of nanopore sequencing.

[0067] In the above nanopore protein complex, the nanopore protein and the accessory protein are preferably derived from the same species, which helps the two proteins to bind specifically.

[0068] It should be noted that the accessory protein in the present application is a type of protein without pore structure. However, its monomer can non-covalently bind to the monomer of the main pore protein, and also polymerize into a 9 main pore protein monomers + 9 accessory protein monomers complex during the self-assembly process of the main pore protein monomer.

[0069] It should be noted that in the above nanopore complex, the nanopore protein is based on the pore protein monomer BCP34 and its mutant provided by us in the past. The nanopore protein is a nonamer, and its monomer sequence is mined from the metagenomic database of the Mariana Trench 11000 meters deep. The wild type full-length protein sequence is SEQ ID NO: 1:

[0070] MKRFIAFAVSMTLVGCASFSPPKGQTSIRDRAQPLSATVTRNALTKLPPPLAP34IPAAVYNIKDQTGQYKPSPSNGFSTAVMQGATSVLVKALLDSRWFIPLEREGLQNLLTERKIIRARDSK KDLSNLAAASVIIEGSIIAYDSNVRTGGAGAKYLGIGLSEQYREDQVTVNLRAINVNNGRILQSVTSTKMIFSRQLDSGAFGYIRFKKLLEIESGYSYNEPAQLCVIDAIESALIQLIYEGVVGGTWKLKNPADIDSPIFTHYAGQNGPEGQAIM, wherein the first 15 amino acids (underlined part) are the signal peptide region of the pore protein monomer BCP34.

[0071] The mature pore protein monomer BCP34 without signal peptide region has an amino acid sequence of SEQ ID NO: 2:

[0072] It should be noted that homology in the present application refers to "sequence identity" between two amino acid sequences, i.e. the percentage of identical amino acids between the sequences. Methods for assessing the degree of sequence identity between amino acids or nucleotides are known to those skilled in the art. For example, the degree of amino acid sequence identity is typically measured using sequence analysis software. For example, it can be determined by the BLAST program of the NCBI database. For determination of sequence identity, see, for example: Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987 and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991.

[0073] The above-mentioned protein having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more (such as 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, even 99.9% or more) homology with the BCP34 nanopore protein monomer shown in SEQ ID NO: 1 and having the function of polymerizing to form a nanopore protein, has an active site, an active pocket, an active mechanism, a protein structure, etc. that are probably the same as those of the BCP34 nanopore protein monomer.

[0074] Similarly, the above-mentioned protein having 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99% or more (such as 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, even 99.9% or more) homology with the accessory protein monomer shown in SEQ ID NO: 3 and having the function of binding to the pore protein monomer and polymerizing to form an accessory protein together with the pore protein monomer polymerizing to form the above-mentioned nanopore protein, has an active site, an active pocket, an active mechanism, a protein structure, etc. that are probably the same as those of the AP34 monomer.

[0075] Amino acid residues can be represented according to the standard three letter or one letter amino acid code as is well known and conventional in the art. Herein, amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine (Cys; C), glutamic acid (Glu; E), glutamine (Gin; Q), glycine (Gly; G), histidine (His; H), isoleucine (lie; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).

[0076] Conservative amino acid substitutions or replacements are well known in the art, for example, conservative amino acid substitutions are preferably substitutions of one amino acid residue for another within one of the following groups (1) - (5): (1) smaller aliphatic nonpolar or weakly polar residues: Ala, Ser, Thr, Pro, and Gly; (2) residues with polar negative charge and their (uncharged) amides: Asp, Asn, Glu, and Gin; (3) residues with polar positive charge: His, Arg, and Lys; (4) larger aliphatic nonpolar residues: Met, Leu, lie, Val, and Cys; and (5) aromatic residues: Phe, Tyr, and Trp. Particularly preferred conservative amino acid substitutions are as follows: Ala for Gly or Ser; Arg for Lys; Asn for Gin or His; Asp for Glu; Cys for Ser; Gin for Asn; Glu for Asp; Gly for Ala or Pro; His for Asn or Gin; lie for Leu or Val; Leu for lie or Val; Lys for Arg, Gin, or Glu; Met for Leu, Tyr, or lie; Phe for Met, Leu, or Tyr; Ser for Thr; Thr for Ser; Trp for Tyr; Tyr for Trp or Phe; and Val for lie or Leu.

[0077] Conservative amino acid substitutions can also be made according to well known rules for amino acid substitution as known to those skilled in the art, such as the "blosum62 score matrix" and the like.

[0078] Mutants of BCP34 protein include, but are not limited to, variants with amino acid mutations at one or more of positions S71, N74, G75, and F76 of SEQ ID NO: 1; preferably, the mutation at position S71 includes mutation to G, A, or T; the mutation at position N74 includes mutation to G, A, S, or T; the mutation at position G75 includes mutation to A, S, T, N, or Q; the mutation at position F76 includes mutation to G, A, S, T, N, or Q; mutation of amino acid at one or more of positions E162, R196, S200, and S216; preferably, the mutation at position E162 includes mutation to A, G, V, L, I, Y, F, or W; the mutation at position R196 includes mutation to A, G, V, L, I, Y, F, or W; the mutation at position S200 includes mutation to A, G, V, L, I, Y, F, or W; the mutation at position S216 includes mutation to A, G, V, L, I, Y, F, or W; mutation of amino acid at one or more of positions R103, E104, E112, R113, K114, R117, R119, D120, K122, D124, K154, R165, D199, R207, K209, K210, E213, and E215; preferably, the mutation at position R103 includes mutation to A, G, S, T, N, or Q; the mutation at position E104 includes mutation to K, R, A, G, S, T, N, or Q; the mutation at position E112 includes mutation to K, R, A, G, S, T, N, or Q; the mutation at position R113 includes mutation to A, G, S, T, N, or Q; the mutation at position K114 includes mutation to A, G, S, T, N, or Q; the mutation at position R117 includes mutation to A, G, S, T, N, or Q; the mutation at position R119 includes mutation to A, G, S, T, N, or Q; the mutation at position D120 includes mutation to K, R, A, G, S, T, N, or Q; the mutation at position K122 includes mutation to A, G, S, T, N, or Q; the mutation at position D124 includes mutation to K, R, A, G, S, T, N, or Q; the mutation at position K154 includes mutation to A, G, S, T, N, or Q; the mutation at position R165 includes mutation to A, G, S, T, N, or Q; the mutation at position D199 includes mutation to A, G, S, T, N, or Q; the mutation at position R207 includes mutation to A, G, S, T, N, Q, D, or E; the mutation at position K209 includes mutation to A, G, S, T, N, Q, D, or E; the mutation at position K210 includes mutation to A, G, S, T, N, Q, D, or E; the mutation at position E213 includes mutation to A, G, S, T, N, or Q; the mutation at position E215 includes mutation to A, G, S, T, N, or Q (see PCT / CN2022 / 143298 for details).

[0079] The above-mentioned accessory protein monomer AP34 and its mutants are proteins from the same species as the pore protein monomer BCP34, and the sequence of the protein is mined from the Mariana Trench metagenomic database. It can form a central symmetrical nonamer under the interaction with BCP34, and its N-terminal (such as part of the N-terminal) or the whole is inserted into the nanopore cavity formed by the aggregation of the pore protein monomer BCP34. The full-length sequence of the wild-type sequence of the accessory protein monomer is SEQ ID NO: 3 as follows:

[0080] MSKVIFGFFAVLLMCFVVTASASSLVYTPKNPSFGGPAAYGTYLLNNANAQNQFKDPDAIRKDKTPIEEFNDRLQRSLLSRITSTISRSIIGIDGAVNPGSFETTDFLIDVTDLGGGQMSITTTDKVTGDQTSIVIETGL, wherein the first 22 amino acids (underlined part) are the signal peptide region of the accessory protein monomer AP34, and the protein fragment of the mature accessory protein monomer AP34 is free of the signal peptide region, and the sequence is shown as SEQ ID NO: 4:

[0081] The above-mentioned accessory protein monomer AP34 and its mutants can be truncated for changing the pore-embedding ability of the pore complex or / and the adaptability to the membrane. The truncation includes but is not limited to: retaining only the N-terminal 23-67 amino acids (i.e. lacking the signal peptide of the N-terminal 1-22 amino acids), retaining only the N-terminal 23-57 amino acids, retaining only the N-terminal 23-52 amino acids, retaining only the N-terminal 23-54 amino acids, retaining only the N-terminal 23-51 amino acids, retaining only the N-terminal 23-50 amino acids, retaining only the N-terminal 23-49 amino acids, retaining only the N-terminal 23-48 amino acids, and retaining only the N-terminal 23-47 amino acids. Preferably, the length of the mutant AP34 fragment can be 24-45 amino acids, and specifically can be 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44 or 45 amino acids. More preferably, its amino acid sequence is from the 23rd-48th, 23rd-49th, 23rd-50th, 23rd-51st, 23rd-54th or 23rd-57th corresponding residues of SEQ ID NO: 3 or any variant thereof.

[0082] In the nanopore complex described above, the constriction zone of the pore protein monomer BCP34 is composed of amino acids S71, N74, G75 and F76. Preferably, the mutation direction of the amino acid residues in the constriction zone is that the mutation at position S71 includes mutation to G, A or T; the mutation at position N74 includes mutation to G, A, S or T; the mutation at position G75 includes mutation to A, S, T or Q; and the mutation at position F76 includes mutation to A, S, T, N or Q. The main amino acid residue in the constriction zone of the accessory protein monomer AP34 is A39, and in addition, T42 and N46 can also have an impact on the sequencing signal. The mutation direction of this constriction zone includes but is not limited to one or more positions of A39, T42 and N46 of SEQ ID NO: 3 having an amino acid mutation. Preferably, the mutation at position A39 includes mutation to S, T, N, G, V, L, I or Q; the mutation at position T42 includes mutation to A, G or S; and the mutation at position N46 includes mutation to A, G or S. The above-mentioned mutations of the relevant amino acids in the constriction zone of BCP34 and the constriction zone of AP34 help to increase the signal-to-noise ratio of the current signal, thereby improving the resolution of base recognition, and ultimately improving the accuracy of the determination of the sequence of the analyte to be tested.

[0083] All or part (such as a part of the N-terminus) of the accessory protein monomer AP34 described above can be positioned within the nanopore lumen formed by the polymerization of the pore protein monomer BCP34. The central cavity or pore of the nonamer formed by the polymerization of the accessory protein monomer AP34 is aligned with the nanopore lumen formed by the polymerization of the nanopore protein monomer BCP34 to form a continuous channel, such that each interacting chemical group of the analyte translocating through the continuous channel first interacts with the constriction zone of the nanopore protein and then interacts with the constriction zone of the accessory protein.

[0084] The two constriction regions in the channel of the nanopore complex described above can have the same or different minimum pore diameters (i.e. the minimum pore diameters of the first sensing zone and the second sensing zone both belong to a parallel relationship, which can be that the minimum pore diameter of the first sensing zone is greater than that of the second sensing zone, or the minimum pore diameter of the first sensing zone is less than that of the second sensing zone). The size of the pore diameter includes but is not limited to the range of 1-3 nm. Compared with the current signal generated by a single constriction zone, the double constriction zone can provide a more distinguishable current signal between the observed currents, thereby having a higher resolution for the analyte to be tested in the form of a homopolymer.

[0085] In the nanopore protein complex described above, the accessory protein is attached to the nanopore lumen of the nanopore protein via covalent or non-covalent interactions; more preferably, the nanopore protein and the accessory protein exist in the form of a nonamer. Compared to the attachment mode of non-covalent interaction, the mode of attaching the accessory protein to the nanopore lumen of the nanopore protein via covalent interaction is relatively more stable, which helps to form a stable continuous channel with at least two constriction regions.

[0086] To further improve the stability of the attachment of the accessory protein to the nanopore protein, it is preferred in the present application to mutate or modify the partial amino acid residues of the interaction region of the two. The interaction and binding region of the accessory protein monomer AP34 and the pore protein monomer BCP34 generally includes residues 23 to 30 and / or 51 to 57 of AP34, and these residues can contain one or more modifications. The formation of the constriction region in the pore generally includes residues 30 to 48 of AP34, and these residues can contain one or more modifications. Residues 8 to 29 of AP34 form an a-helix. The AP34 constriction region is mainly in stable contact with the inner cavity of BCP34 at residues 23, 24, 25, 26, 31, 33, 34, 40, 43, and 45 of SEQ ID NO: 3. The accessory protein monomer AP34 is attached to the pore protein monomer BCP34 and forms a relatively stable complex through one or more covalent bonds and / or one or more non-covalent interactions. Preferably, the suitable non-covalent interaction includes salt bridges, electrostatic interactions, and π-π interactions. Covalent interaction refers to chemical modification of the mutant or modified monomer by attaching the molecule to one or more cysteines (cysteine linkage), attaching the molecule to one or more lysines, attaching the molecule to one or more unnatural amino acids, epitope enzyme modification, or modification of the terminus. Suitable methods for performing such modifications are well known in the art. The modification mode includes but is not limited to introducing cysteine, charged amino acid, unnatural reactive amino acid, or photoreactive amino acid at any one or more of these positions.

[0087] The above-mentioned positions include but are not limited to S23, S24, L25, T28, K30, N31, S33, F34, N47, A50, Q51, N52, and Q53 on AP34. These positions include but are not limited to N145, T148, K154, L156, L160, S161, R165, S195, Q197, D199, F203, Y205, K209, K210, L211, E213, E215, G217, S219, N221 on BCP34.

[0088] The site combinations of the interaction between the nanopore protein and the accessory protein include, but are not limited to, one or more of the following: R165 on BCP34 and S23 on AP34, S195 on BCP34 and S24 on AP34, S219 on BCP34 and L25 on AP34, N221 on BCP34 and L25 on AP34, N145C on BCP34 and V26C on AP34, D199 on BCP34 and T28 on AP34, D199 on BCP34 and K30 on AP34, E215 on BCP34 and K30 on AP34, E213 on BCP34 and N31 on AP34, K154 on BCP34 and S33 on AP34, E213 on BCP34 and S33 on AP34, E213 on BCP34 and N47 on AP34, L211 on BCP34 and A50 on AP34, L211 on BCP34 and Q51 on AP34, K209 on BCP34 and N52 on AP34, K210 on BCP34 and Q53 on AP34. Specific amino acid combinations are shown in Table 1.

[0089] Table 1

[0090] To further enhance the stability of the attachment of the accessory protein to the nanopore protein, the sites in any one or more of the above-mentioned site combinations can be mutated (e.g., all mutated to cysteine, thereby forming a disulfide bond between the combined sites within any one group) or modified. Specifically, the mutants of the modified nanopore protein monomer BCP34 or the accessory protein monomer AP34 can be directly chemically modified by the side chain groups of the amino acids through spontaneous reaction or oxidation reaction, or by attaching chemical groups. The chemical groups, such as amide groups, specifically include, but are not limited to, iodoacetamide or maleimide.

[0091] According to a second aspect of the present application, a nanopore sensor is provided, comprising a membrane and the above-mentioned nanopore protein complex, wherein the nanopore protein in the nanopore protein complex is positioned on the membrane, and the nanopore cavity of the nanopore protein and the accessory protein together form a continuous channel across the membrane.

[0092] The membrane on the above-mentioned nanopore sensor comprises a layer of amphiphilic molecules; preferably, the membrane is a phospholipid bilayer composed of diacylphosphatidylcholine, a two-block copolymer or a three-block copolymer.

[0093] The nanopore sensor formed by the nanopore protein complex containing the above-mentioned two constriction regions can improve the discrimination and sequencing accuracy of homopolymer nucleic acid molecules when applied to sequencing such molecules.

[0094] According to a third aspect of the present application, there is provided a nanopore sequencing system comprising the nanopore sensor, the conductive solution, positive and negative electrodes for providing a voltage potential across the membrane, and a measurement device for measuring the electrical signal passing through the continuous channel, wherein the nanopore sensor is located in the conductive solution and divides the conductive solution into the first chamber and the second chamber.

[0095] In some preferred embodiments, the nanopore sequencing system further comprises a molecule to be tested, wherein the molecule to be tested is selected from nucleic acid molecules, more preferably nucleic acid molecules containing homopolymer regions; preferably, the molecule to be tested is transiently located within the continuous channel, and one end of the molecule to be tested is located in the first chamber and the other end is located in the second chamber.

[0096] According to a fourth aspect of the present application, there is provided a method of nanopore sequencing, comprising: contacting the nanopore sequencing system with a nucleic acid molecule to be tested; applying an electric potential across the membrane so that the nucleic acid molecule to be tested enters the continuous channel; and measuring the electrical signal generated when the nucleic acid molecule to be tested moves relative to the continuous channel one or more times, thereby obtaining the sequence of the nucleic acid molecule to be tested.

[0097] In the present application, there are various types of specific electrical signals generated when the nucleic acid molecule to be tested passes through the first constriction region and the second constriction region, including but not limited to current block signal intensity, current block duration, or interval time of current block events.

[0098] The nucleic acid molecule to be tested is an oligonucleotide molecule, the nucleotides in the oligonucleotide molecule interact with the first constriction region and the second constriction region in the continuous channel, and wherein each of the first constriction region and the second constriction region is capable of distinguishing different nucleotides, so that the total current passing through the continuous channel is affected by the interaction between each of the first constriction region and the second constriction region and the nucleotides positioned in the corresponding region.

[0099] Specifically, the oligonucleotide molecule moves through the continuous channel and realizes transmembrane translocation. In some preferred embodiments, a nucleic acid binding protein is used to control the movement of the oligonucleotide molecule relative to the continuous channel pore.

[0100] In some preferred embodiments, the oligonucleotide molecule is a homopolymer-containing oligonucleotide molecule, and the method can determine the nucleotide sequence of the homopolymer-containing oligonucleotide molecule.

[0101] According to a fifth aspect of the present application, there is also provided a method for preparing a nanopore complex, comprising: separately constructing an expression vector of a nanopore protein monomer and an expression vector of an accessory protein; obtaining the nanopore complex by co-expressing the expression vector of the nanopore protein monomer and the expression vector of the accessory protein in a competent cell and through separation and purification; or constructing the nanopore complex by in vitro recombination.

[0102] In some preferred embodiments, the obtaining the nanopore complex by co-expressing the expression vector of the protein monomer and the expression vector of the accessory protein monomer in a competent cell and through separation and purification comprises: co-expressing the expression vector of the nanopore protein monomer and the expression vector of the accessory protein monomer in a competent cell and culturing the competent cell to a predetermined OD value, and collecting the co-expression cell, wherein the co-expression cell co-expresses the nanopore protein monomer and the accessory protein monomer, and the nanopore protein monomer and the accessory protein monomer spontaneously polymerize to form the nanopore complex; and sequentially performing cell disruption, cell membrane lysis and purification on the co-expression cell to obtain the nanopore complex; preferably, the predetermined OD value is an OD600 value of 0.6-0.8; preferably, the performing purification comprises sequentially performing strep tag purification and his tag purification; and preferably, the molar ratio of the nanopore protein monomer and the accessory protein monomer in the obtained nanopore complex is 1:1.

[0103] In some preferred embodiments, the constructing the nanopore complex by in vitro recombination comprises: separately expressing the expression vector of the nanopore protein monomer and the expression vector of the accessory protein monomer in a competent cell, and separately performing separation and purification on the cell after each expression to obtain the nanopore protein monomer and the accessory protein monomer; and mixing and incubating the nanopore protein monomer and the accessory protein monomer at a molar ratio of 1:30-50 to obtain the nanopore complex; preferably, the mixing and incubating the nanopore protein monomer and the accessory protein monomer is performed in a buffer solution; preferably, the buffer solution comprises 20mM Tris-HCl, 150mM NaCl, 0.05% Tween20 and 15% glycerol, pH 8.0; and preferably, the mixing and incubating is performed at 4°C.

[0104] To obtain the above-mentioned nanopore complex by in vitro expression, a protein purification tag can be added at the C-terminus of the pore protein monomer BCP34 and / or the accessory protein monomer AP34. Protein purification tags are well known in the art, including but not limited to a histidine tag (Poly His), a Strep tag II (non-biotinylated streptavidin), a Biotin Avitag (a short peptide that can be biotinylated), a calmodulin binding peptide (CBP), or an arginine tag (Poly Arg). The positions of the tag insertion include but are not limited to positions 48, 49, 50, 51, 54, 57, 62, or 67 of SEQ ID NO: 1 and mutants thereof and SEQ ID NO: 3 and mutants thereof.

[0105] The protein purification tag of AP34 can be inserted with a protease cleavage short peptide sequence, which helps to remove the extra residues that can interfere with the formation of the protein complex or the insertion of the protein complex into the pore. These protease cleavage sequences are also well known in the art, including but not limited to thrombin (recognition sequence SEQ ID NO: 30: LVPRG↓S), Factor Xa (recognition sequence SEQ ID NO: 31: IE / DG↓R), TEV protease (recognition sequence SEQ ID NO: 32: ENLYFQ↓G), HRV 3C protease (recognition sequence: SEQ ID NO: 33: LEVLFQ↓GP), wherein the arrow indicates the site of protease action.

[0106] The above-mentioned nanopore complex can also be obtained by co-expressing the pore protein monomer BCP34 and its mutants and the accessory protein monomer AP34 and its mutants. The method of co-expression refers to the simultaneous expression of the pore protein monomer and the accessory protein monomer in a suitable host cell and allowing the nanopore complex to form in vivo. Preferably, the host cell includes but is not limited to BL21 (DE3), BL21 Star (DE3) pLyss, Rossata (DE3), Lemo21 (DE3), etc. Specifically, at least one gene encoding the pore protein monomer and the gene encoding the accessory protein monomer in one vector, or the at least one accessory protein monomer in a second vector can be transformed together to express the proteins and prepare the nanopore complex in the transformed cells. The two vectors can be under the control of a single promoter or under the control of two independent promoters, the two independent promoters can be the same or different, the process can be carried out in vivo in a host cell or in a cell-free expression system. Preferably, the expression vector includes but is not limited to a T7 promoter vector such as PET.28a(+), PET.21a(+), PET.32a(+), etc.

[0107] When the nanopore complex is constructed in vivo, it can be purified by the corresponding protein purification tags of BCP34 and AP34 proteins. The purification steps are known methods in the art (such as ion exchange, gel filtration, hydrophobic interaction column chromatography, etc.), which can be used alone or in different combinations to purify the components of the nanopore complex. It should be noted that the protein purification tags of BCP34 and AP34 can be the same or different, and using two different tags can improve the purity of the nanopore complex to some extent.

[0108] In addition, the above nanopore complex can also be constructed by in vitro recombination. The pore protein monomer BP34 and its mutants can be encoded in a suitable vector and transformed into a suitable host cell for expression. Then it is purified by the corresponding protein purification tag, and the purification steps are known methods in the art (such as ion exchange, gel filtration, hydrophobic interaction column chromatography, etc.), which can be used alone or in different combinations to purify the components of the pore complex. The accessory protein monomer AP34 and its mutants can be obtained by the same in vitro expression and purification method as the pore protein monomer BP34, or by direct polypeptide solid-phase synthesis.

[0109] If obtained by in vivo expression, the protein sequence is SEQ ID NO: 3 or its mutants and truncates, and the signal peptide at the N-terminus in vivo can guide its secretion on the membrane protein. After purification by the purification tag, for AP34 with added protease cleavage site, the corresponding protease can be used to remove the C-terminal protein sequence unrelated to pore formation.

[0110] If obtained by polypeptide synthesis, the polypeptide sequence is the amino acid sequence directly without C-terminal signal peptide. The polypeptide synthesis method includes but is not limited to solid-phase synthesis, liquid-phase segmented synthesis, Staudinger ligation, native chemical ligation, photosensitive auxiliary group ligation, removable auxiliary group ligation, chemical region-selective ligation, carboxylic acid anhydride (NCA) method of amino acids, combinatorial chemistry, enzymatic method, genetic engineering method and fermentation method. Preferably, FMOC or BOC solid-phase polypeptide synthesis method is used.

[0111] When both BCP34 and AP34 are available, the nanopore complex can be formed by in vitro incubation. The ratio of incubated BCP34 and AP34 includes but is not limited to the range of 500:1 to 1:1. The temperature of incubation includes but is not limited to 4°C, 16°C, 20°C, 25°C and 37°C. The length of incubation includes but is not limited to 30 min, 1 h, 2 h, 3 h, 4 h, 5 h, 8 h, 16 h and 24 h. The salt concentration of the reaction solution for incubation includes but is not limited to 50 mM, 100 mM, 150 mM, 200 mM, 250 mM, 300 mM, 400 mM, 500 mM. The pH value of the reaction solution for incubation includes but is not limited to 7.0, 7.5, 8.0, 8.5. Oxidizing agents or chemical cross-linking agents that promote the stability of the nanopore complex can be added simultaneously. The incubation method can mix the pore protein monomer BCP34 and the accessory protein monomer AP34 together in the reaction solution, or the pore protein monomer can be inserted into the membrane, and then the accessory protein monomer is added, so that the nanopore complex can be formed in situ.

[0112] The nanopore protein complex with double constriction zones constructed by the above-mentioned method of the present application can be used to characterize different analytes, including but not limited to various biological or synthetic macromolecules and polymers of polynucleotides, polypeptides and polysaccharides. Preferably, for characterizing targets, polynucleotides include DNA and / or RNA and their modifications.

[0113] The present application provides a method for determining the presence or absence of one or more characteristics of a target analyte, the method comprising:

[0114] a. contacting the target analyte with the nanopore on the membrane on the nanopore sensor containing the above-mentioned nanopore protein complex BCP34-AP34 or its mutants with the first constriction zone and the second constriction zone, so that the target analyte moves relative to the nanopore;

[0115] b. acquiring one or more measurement values while the target analyte moves relative to the nanopore, thereby determining the presence or absence of one or more characteristics of the target analyte.

[0116] Further, the above-mentioned target analyte interacts with the nanopore protein complex so that the above-mentioned target analyte moves relative to the nanopore on the continuous channel.

[0117] It should be noted that the nanopore protein complex provided by the present application can be used to develop a single-molecule sequencer with higher accuracy, higher integration and higher stability. In addition, as a biosensor, the protein complex has extremely high potential to identify and analyze the composition and modification information of different organic and inorganic substances, and has broad application prospects in drug dynamics or drug screening, as well as in nucleic acid drug delivery and biological sensing. It can also be combined with genomics, proteomics, metabolomics, etc. to build a general measurement platform that meets the needs of full-omics analysis, helping us to better understand the laws of life and disease mechanisms.

[0118] The beneficial effects of the present application will be further illustrated below in conjunction with specific examples.

[0119] Example 1: Predicted structure of wild-type BCP34-AP34 AlphaFold2-Multimer

[0120] We used AlphaFold-Multimer-v3 to perform multimer structure prediction on the target protein sequence, and a series of predictions were performed using different model parameters provided by AlphaFold-Multimer-v3, and the multimer structure with the highest multimer confidence was selected as the final prediction result. The prediction results are shown in Figures 1 and 2. In Figure 1, A is the side view of the BCP34-AP34 multimer prediction structure, and B is the top view of the BCP34-AP34 multimer prediction structure. In Figure 2, A is the side view of the mature BCP34-AP34 (without signal peptide) multimer prediction structure, and B is the top view of the mature BCP34-AP34 (without signal peptide) multimer prediction structure. The structure shows that BCP34:AP34 is a nonamer formed by a 9:9 stoichiometric ratio through non-covalent interaction between the two, with C9 symmetry.

[0121] The complex has two constriction zones, which play a decisive role in the generation of current signals: the first constriction zone is formed by the aggregation of pore protein monomers BCP34, and the second constriction zone is formed by the aggregation of accessory protein monomers AP34, located below the first constriction zone, and there is a certain distance between the two constriction zones, as shown in Figure 3. Figure 4 shows the side chain structures of important amino acids in the respective constriction zones of the complex prediction structure, and the first constriction zone shows that the four amino acid side chains are S71, N74, G75, and F76 on BCP34, and the second constriction zone shows that the amino acid side chain is A39 on AP34, which is the narrowest part inside AP34, and two other narrow places, T42 and N46, also have a certain effect on the current signal.

[0122] Example 2: Construction of expression vectors for wild type BCP34 monomer and mutants and AP34 monomer and mutants

[0123] The DNA sequence of the BCP34 monomer (SEQ ID NO: 5) was inserted into the multiple cloning region of the vector pET24a by the In-fusion method after digestion with Ndel and Xhol. A StrepII amino acid was added as a purification tag at the C-terminus of the amino acid sequence of the wild type BCP34 (SEQ ID NO: 1), and the constructed vector was named pET24a-BCP34. The expression vector for the BCP34 monomer was used as a template to construct other mutants, such as BCP34-1 (SEQ ID NO: 6), by the method of site-directed mutagenesis using the Agilent site-directed mutagenesis kit.

[0124] The DNA sequence of the AP34-1 monomer (SEQ ID NO: 7) was inserted into the multiple cloning region of the vector pET21a by the In-fusion method after digestion with Ndel and Xhol, wherein the AP34-1 has a TEV cleavage recognition site added at position 57 of the amino acid sequence of the wild type AP34 for facilitating in vitro expression and purification of the protein. Six histidines were added as a purification tag at the C-terminus of the amino acid sequence of the AP34-1 (SEQ ID NO: 8), and the constructed vector was named pET21a-AP34-1. The expression vector for the AP34-1 monomer was used as a template to construct other mutants by the method of site-directed mutagenesis using the Agilent site-directed mutagenesis kit.

[0125] The amino acid sequence of BCP34 (SEQ ID NO: 1):

[0126] The DNA sequence of BCP34 (SEQ ID NO: 5)

[0127] The amino acid sequence of the BCP34-1 mutant (N74S, G75N and F76A SEQ ID NO: 6):

[0128] The amino acid sequence of AP34 (SEQ ID NO: 3):

[0129] The DNA sequence of AP34-1 (SEQ ID NO: 7):

[0130] Amino acid sequence of AP34-1 (AP34 with TEV cleavage recognition site added at position 57, SEQ ID NO: 8):

[0131] Example Three: Construction of multi-contracting zone complex pores by co-expression

[0132] To produce the nanopore complex, the two protein monomers can be co-expressed in a suitable gram-negative host, such as E. coli, and extracted and purified as a complex from the outer membrane.

[0133] In this example, we constructed plasmids pET24a-BCP34-1 and pET21a-AP34-1 following the method of Example Two, and co-transformed them into E. coli expression strain E. coli BL21(DE3) by heat shock method, then spread the bacterial solution evenly on plates containing 50 μg / mL kanamycin and 100 μg / mL ampicillin, and incubated at 37°C overnight. The next day, single colonies were picked and inoculated in 5 mL LB medium containing 50 μg / mL kanamycin and 100 μg / mL ampicillin, and incubated at 37°C, 200 rpm, overnight. The resulting bacterial solution was inoculated in 50 mL LB liquid medium containing 50 μg / mL kanamycin and 100 μg / mL ampicillin at a volume ratio of 1:100, and incubated at 37°C, 200 rpm, for 4 h. The expanded bacterial solution was inoculated in 2 L LB liquid medium containing 50 μg / mL kanamycin and 100 μg / mL ampicillin at a volume ratio of 1:100, and incubated at 37°C, 200 rpm. When the OD600 value reached about 0.6-0.8, IPTG was added to a final concentration of 0.5 mM, and incubated at 16°C, 200 rpm, for about 16-18 h. The bacterial solution was centrifuged at 8000 rpm to collect the bacterial cells, which were stored at -20°C for use.

[0134] To obtain the co-expressed protein complex in vitro, the purification and extraction steps were as follows:

[0135] (1) Buffer solution preparation

[0136] Buffer solution A: 20 mM Tris-HCl, 150 mM NaCl pH 8.0

[0137] Buffer solution B: 20 mM Tris-HCl, 150 mM NaCl 1% DDM, pH 8.0

[0138] Buffer solution C: 20 mM Tris-HCl, 150 mM NaCl, 0.05% Tween20, 15% gly pH 8.0

[0139] Buffer solution D: 20 mM Tris-HCl, 150 mM NaCl, 0.05% Tween20, 15% gly, 5 mM desthiobiotin, pH 8.0 (this buffer solution is prepared fresh)

[0140] Buffer solution E: 20 mM Tris-HCl, 150 mM NaCl, 0.05% Tween20, 15% gly, 300 mM imidazole, pH 8.0

[0141] (1) Cell disruption and cell membrane lysis

[0142] Resuspend the bacteria at a ratio of 1 g bacteria to 10 mL buffer solution A, and disrupt the cells by high pressure homogenizer until the bacteria solution is clear. Then use low temperature ultracentrifuge (Beckman, Optima TM XPN-90) to centrifuge at 40000 rpm, 4°C, for 1 h. Remove the supernatant and suspend the precipitate with buffer solution B, and then place it in a rotator at 4°C overnight. The next day, centrifuge at 18000 rpm, 4°C, for 1 h, take the supernatant, filter it through a 0.22 μm filter membrane, and then store it at 4°C.

[0143] (2) Purification steps

[0144] Use AKTA pure chromatograph to equilibrate the Strep-Tactin beads (IBA Lifesciences) chromatography column with buffer solution C for 5 column volumes (CV), and then load 2 mL / min. After loading is complete, use buffer solution C to flush for 20 CV, use buffer solution D to elute, and collect the target protein. Perform SDS-PAGE electrophoresis on the target protein obtained after purification, and Fig. 5, panel A, shows the protein characterization results of the wild-type complex after purification by the strep column. The denaturation condition is to heat the sample at 60°C for 15 minutes. The results show that the target protein is in a polymeric state without denaturation, and is in a monomeric state after denaturation. The monomer contains the mature BCP34-1 and AP34-1 proteins, respectively, proving that the complex is formed on the membrane and is purified by the strep tag of BCP34-1. At this time, there is excess BCP34-1 and part of the impurity protein.

[0145] Next, the eluted protein solution was loaded onto a 5 mL HisTrap column equilibrated in buffer solution C. The column was washed with >10 CVs 5% buffer E ion buffer A, and the target protein was eluted with 60 mL or more of a 5-100% gradient of buffer B, and collected. The target protein obtained after purification was subjected to SDS-PAGE electrophoresis, and Figure 5, B, shows the protein characterization results of the wild-type complex after his column purification. The results show that the target protein is in a polymeric state without denaturation, and in a monomeric state after denaturation, with the mature BCP34-1 and AP34-1 proteins, respectively, proving that the complex is formed on the membrane and further purified by the his tag of AP34-1, at which time the protein complex is in a 1:1 state.

[0146] Then, an appropriate amount of TEV enzyme was added to the protein elution solution to remove the C-terminal redundant protein sequence of AP34-1 (SEQ ID NO: 9 SD AIRKDKTPIEEFNDRLQRSLLSRITSTISRSIIGIDGAVNPGSFETTDFLIDVTDLGGGQMSITTTDKVTGDQTSIVIETGL), leaving the key sequence AP34-2 (SEQ ID NO: 10 MSKVIFGFFAVLLMCFVVTASASSLVYTPKNPSFGGPAAYGTYLLNNANAQNQFKDPENLYFQ), and the protein solution with the added enzyme was placed on a rotary instrument at 4°C overnight, and the target protein obtained was subjected to SDS-PAGE electrophoresis, and the protein running gel results are shown in Figure 6. Figure 6, A, shows the changes in the complex protein before and after the addition of TEV before and after denaturation, and Figure 6, B, shows the changes in the complex protein before and after the addition of TEV after denaturation, from which it can be concluded that the C-terminal redundant sequence of AP34-1 was successfully removed by the TEV enzyme, and the target protein was finally obtained. The obtained protein was concentrated to 1 mL, and was subjected to a Superdex 6 increase 10 / 300 GL (Cytiva) equilibrated with buffer solution C, and the target protein was collected and then stored at -80°C.

[0147] Example Four: Construction of a multi-contractile zone complex hole by in vitro recombination

[0148] The complex can also be constructed by in vitro recombination. In this embodiment, we express mutant BCP34-1 in a gram-negative host (e.g., E. coli) and purify the membrane, while the key fragment of AP34 is constructed by polypeptide synthesis. Since polypeptide synthesis is performed in vitro, the signal peptide at the N-terminus and the sequence at the C-terminus that may interfere with the insertion of the membrane of the complex pore are not required. Here, we use polypeptides synthesized by Genscript through solid-phase synthesis to synthesize the polypeptide sequence AP34-P1 (SEQ ID NO: 11: SSLVYTPKNPSFGGPAAYGTYLLNNANAQNQFKDP).

[0149] The expression and purification steps of BCP34-1 are the same as in Example 3, except that the second step of purifying histidine and the step of TEV enzyme cutting are removed, and the protein purified by Strep-Tactin beads (IBA Lifesciences) column is eluted and diluted to the appropriate concentration for use. At the same time, 2 mg of AP34-P1 lyophilized from Genscript is dissolved in 1 mL of buffer solution C in Example 3 to obtain a 2 mg / mL sample. The sample is vortexed until no peptide powder is visible. Since BCP34-1 contains a small amount of impurities, it is difficult to accurately measure the concentration. The intensity of the protein band on SDS-PAGE relative to the known marker protein can be used to obtain a rough estimate of the sample. BCP34-1 and AP34-P1 are mixed at a molar ratio of about 1:50 and incubated at 4°C at 700 rpm for 4 h and centrifuged at 13,000 rpm for 2 min. The obtained target protein is subjected to SDS-PAGE electrophoresis, as shown in Figure 7. Figure 7A shows the gel of the complex before denaturation, and Figure 7B shows the gel of the complex after denaturation. The results show that the target protein is in a polymer state before denaturation and in a monomer state after denaturation. Since there is no covalent interaction between BCP34-1 and AP34-P, and the size difference between BCP34-1 polymer and BCP34-AP34-P1 complex is not large, it cannot be directly shown from the gel. Therefore, the gel only indicates the state and stability of the protein.

[0150] Example Five: Further Stabilizing the Multi-Contraction Region Nanopore Complex by Covalent Cross-Linking

[0151] As can be seen from Example Four, although the complex pore is unstable, the two protein monomers that make up the complex pore have strong non-covalent interactions and can assemble into a polymer state before the protein is denatured. To further reduce the potential risks in actual nanopore testing applications, we will introduce cysteine into the interaction interface of the protein in this embodiment to further improve the stability of the complex through disulfide bonds.

[0152] According to the structure of Example One, we further analyzed, as shown in Figure 8, that there are many amino acids adjacent to the interaction region of BCP34 and AP34, and the side chain direction is close, the side chains of these amino acids are marked, and the potential interaction amino acid pairs have been listed in Table 1. In this example, we mainly present the results of BCP34-2 (SEQ ID NO: 12, R165C on BCP34-1) and AP34-P2 (SEQ ID NO: 13, S23C on AP34) and BCP34-3 (SEQ ID NO: 14, N145C on BCP34-1) and AP34-P3 (SEQ ID NO: 15, V26C on AP34), and the four sequences are as follows:

[0153] BPC34-2 (BCP34-1 based on increasing R165C mutation, SEQ ID NO: 12):

[0154] AP34-P2 (AP34 based on increasing S23C mutation, SEQ ID NO: 13):

[0155] BCP34-3 (BCP34 based on increasing N145C mutation, SEQ ID NO: 14):

[0156] AP34-P3 (AP34 based on increasing V26C mutation, SEQ ID NO: 15):

[0157] Among them, BCP34-2 and BCP34-3 are obtained by using the E. coli expression and membrane purification steps in Example Four, and the obtained target protein is subjected to SDS-PAGE electrophoresis, as shown in Figure 9, and the crude concentration of the two proteins is obtained by using the protein band on SDS-PAGE relative to the known marker protein intensity. Figure 9 A and B show the changes in BCP34-2 and BCP34-3 before and after denaturation, and the results show that the target protein BCP34-2 is in a polymeric state before denaturation, and is in a monomeric state after denaturation; BCP34-3 protein is in a polymeric state before denaturation, and is mostly in a monomeric state after denaturation.

[0158] AP34-P2 and AP34-P1 were synthesized by GenScript using the solid phase synthesis method of Example 4, and the dry powder was dissolved in solution. BCP34-2 and AP34-P2, BCP34-3 and AP34-P3 were mixed at a molar ratio of about 1:50, respectively, and incubated at 700 rpm for 4 h at 4°C, and centrifuged at 13,000 rpm for 2 min. The target protein obtained was characterized by SDS-PAGE electrophoresis, as shown in Figure 10. Figure 10A shows the reaction effect of BCP34-2 and AP34-P2, under the reaction conditions, about 90% of BCP34-2 and AP34-P2 form a stable complex BCP34-2-AP34-P2. Figure 10B shows the reaction effect of BCP34-3 and AP34-P3, under the reaction conditions, only about 20% of BCP34-3 and AP34-P3 form a stable complex BCP34-3-AP34-P3. And verified by DTT characterization, the complex is stabilized by disulfide bond, and the disulfide bond can be destroyed in the presence of reducing agent DTT.

[0159] Since the cross-linking efficiency of S23 on AP34 and R165C on BCP34 is high, we also tried to use the truncated AP34-P4, AP34-P5, AP34-P6, and AP34-P7 (sequences shown below) of AP34-P2, and react with protein BCP34-2 respectively, the reaction conditions remain the same, that is, the protein and the peptide are mixed at a molar ratio of about 1:50, and incubated at 700 rpm for 4 h at 4°C, and centrifuged at 13,000 rpm for 2 min. The target protein obtained was characterized by SDS-PAGE electrophoresis, as shown in Figure 11, which characterized the complexes BCP34-2-AP34-P4, BCP34-2-AP34-P5, BCP34-2-AP34-P6, BCP34-2-AP34-P7. Figure 11A characterizes the reaction efficiency of AP34-P2, AP34-P5, AP34-P6, AP34-P7 with BCP34-2, it can be seen that their reaction efficiency is high, among them, the reaction efficiency of AP34-P7 with BCP34-2 is the highest, close to complete reaction, and the complex formed can be broken in the presence of DTT. Figure 11B characterizes the reaction efficiency of AP34-P4 and BCP34-2 forming complex BCP34-2-AP34-P4, which verifies that the formation of the complex will not destroy the formation of the aggregate, and the formation of the complex will be broken by DTT.

[0160] AP34-P4 (SEQ ID NO: 16, truncated to retain 29 aa.): CSLVYTPKNPSFGGPAAYGTYLLNNANAQ.

[0161] AP34-P5 (SEQ ID NO: 17, 28 aa truncated): CSLVYTPKNPSFGGPAAYGTYLLNNANA.

[0162] AP34-P6 (SEQ ID NO: 18, 27 aa truncated): CSLVYTPKNPSFGGPAAYGTYLLNNAN.

[0163] AP34-P7 (SEQ ID NO: 19, 26 aa truncated): CSLVYTPKNPSFGGPAAYGTYLLNNA.

[0164] Example Six: Constructing nanopore biosensor using multi-shrinkage nanopore complex

[0165] Single nanopore current measurement is based on the amplifier of digital device, here using patch clamp amplifier to collect current signal. Ag / AgCl electrode is immersed in sequencing buffer (composition includes: 0.47M KCl, 25mM HEPES, 1mM EDTA, 5mM ATP, 25mM MgCl2, pH7.6), electrode is located in the cis and trans area of electrolytic cell respectively. Sequencing library and other reagents are added to the cis area. After diluting the nanopore protein or nanopore complex with 1x PBS buffer (usually the protein is diluted by 100 times or 10 times with 0.1 mg / ml protein concentration), under the action of external electric field force, a single nanopore or nanopore complex hole is inserted into the phospholipid bilayer composed of DPhPC (1,2-diphytanoyl-sn-glycero-3-phosphocholine), forming a nanopore biosensor. Apply external voltage to obtain the current amplitude value of single pore protein.

[0166] Example Seven: Using nanopore sensor in Example Six for DNA sequencing

[0167] a) DNA sequencing of nanopore BCP34-1 and BCP34-AP34-P1

[0168] Preparation of nucleic acid sequence: insert the artificially synthesized sequence SEQ ID NO: 20 into the multiple cloning site of PUC57 plasmid, and use the primer combination of SEQ ID NO: 21 and SEQ ID NO: 22 to prepare the 3.5 kb sequence to be sequenced SEQ ID NO: 20 by PCR amplification.

[0169] Preparation of sequencing library: construct the sequence to be sequenced SEQ ID NO: 20 into a sequencing library.

[0170] The sense strand (SEQ ID NO: 23 - (iSP18)4- SEQ ID NO: 29, wherein iSP18 is a spacer) and the antisense strand (SEQ ID NO: 24) of two partially region complementary DNA strands were annealed to form a linker, which was ligated with the double-stranded target fragment PUC57 (SEQ ID NO: 20) to be detected using T4 DNA ligase at room temperature and purified to prepare a sequencing library. Then the sequencing library was incubated with the helicase BCH105 (SEQ ID NO: 25) at 25°C for 1 h (molar ratio 1:8) to form a sequencing library with the structure shown in FIG. 12A containing the BCH105 motor protein. During sequencing, the sequencing library can further bind to the single-stranded DNA with cholesterol (SEQ ID NO: 26, cholesterol is connected to the 5' end of the DNA) in a complementary pairing manner to form the structure shown in FIG. 12B. The sequencing library was mixed with the sequencing buffer (sequencing buffer: 0.47 M KCl, 25 mM HEPES, 1 mM EDTA, 5 mM ATP, 25 mM MgCl2, pH 7.6) and added to a nanopore biosensor; after applying an external voltage of 0.14 V or 0.18 V, it was observed that the DNA was captured by the nanopore, producing a characteristic block current amplitude value. And as the DNA moves through the nanopore, the current amplitude value changes. Different DNA sequences produce different block current amplitude values. Single-stranded DNA with cholesterol can bind to the phospholipid bilayer, helping to capture the sequencing library by the nanopore and reducing the loading amount of the sequencing library.

[0171] SEQ ID NO: 20:

[0172] SEQ ID NO: 21: gccatcagattgtgtttgttagt.

[0173] SEQ ID NO: 22: gcttacggttcactactcacga.

[0174] SEQ ID NO: 23: ttttttttttttttttttttttttttttttttttttttt.

[0175] SEQ ID NO: 29: ggttgtttctgttggtgctgatattgct.

[0176] SEQ ID NO: 24: gcaatatcagcaccaacagaaacaacctttgaggcgagcggtcaa.

[0177] SEQ ID NO: 25:

[0178] SEQ ID NO: 26: cholesterol-ttgaccgctcgcctc.

[0179] Figure 13 is the current change and local details of the library DNA passing through the nanopore protein BCP34-1 under the action of an applied voltage of 0.18 V. Figure 14 is the current change and local details of the library DNA passing through the double-detector nanopore protein BCP34-1-AP34-P1 under the action of an applied voltage of 0.18 V. It can be seen that under an applied voltage of 0.18 V, the open pore current of the nanopore protein BCP34 is 230-250 pA, and the sequencing amplitude is about 40 pA; the open pore current of the double-detector nanopore protein is 130-150 pA, and the sequencing amplitude is about 20 pA. The current signal read in the homopolymer region of DNA is different from the mutant protein, and presents more stepped signals (this phenomenon can also be seen in different double-detector nanopore protein mutants, and their discrimination degrees are different. Comparison can be seen in Figure 22). It can be seen that under the same voltage, the open pore current of the double-detector pore protein BCP34-1-AP34-P1 is lower than that of the original single-detector protein BCP34-1, indicating that a second detector is formed outside the detector of the pore protein, and a double-detector nanopore protein is successfully constructed. The discrimination degree of the stepped homopolymer region indicates that the nanopore protein complex provides more information, so it is speculated that the double-detector nanopore protein has the potential to improve the resolution of nanopore sequencing.

[0180] b) Characterization of DNA sequencing by different AP34 main signal region nanopore complexes

[0181] According to the method of Example Four, we constructed two additional nanopore complexes BCP34-1-AP34-P8 and BCP34-1-AP34-P9. Among them, AP34-P8 is based on AP34 and A39S is modified in the main shrinkage region, and AP34-P9 is based on AP34 and A39N is modified in the main shrinkage region, and is obtained in vitro by polypeptide synthesis method. The sequences are as follows.

[0182] AP34-P8 (SEQ ID NO: 27, based on AP34 and A39S modified in the main shrinkage region):

[0183] AP34-P9 (SEQ ID NO: 28, based on AP34 and A39N modified in the main shrinkage region)

[0184] The complexes were sequenced by the test method of the present example. Figure 15 is the current change and local details of the library DNA when passing through the nanopore protein BCP34-1-AP34-P8 under the action of an applied voltage of 0.18 V. Figure 16 is the current change and local details of the library DNA when passing through the double-detector nanopore protein BCP34-1-AP34-P9 under the action of an applied voltage of 0.18 V. It can be seen that the modification of the main amino acids of the constriction region will cause the resolution ability of the characteristic sequence to change.

[0185] c) Sequencing characterization of DNA by nanopore complexes with different cross-linking site positions

[0186] The complexes BCP34-2-AP34-P2 and BCP34-3-AP34-P3 with different site cross-linking modifications constructed in Example Five were sequenced by the test method of the present example.

[0187] Figure 17 is the current change and local details of the library DNA when passing through the nanopore protein BCP34-2-AP34-P2 under the action of an applied voltage of 0.18 V.

[0188] Figure 18 is the current change and local details of the library DNA when passing through the double-detector nanopore protein BCP34-3-AP34-P3 under the action of an applied voltage of 0.18 V.

[0189] In comparison with the BCP34-1-AP34-P1 complex, the two complexes are easier to embed pores, the complexes are more stable, and there is no phenomenon of an increase in the opening current at a later stage. The output high-resolution sequencing current signal is also more. The opening current of the double-detector nanopore protein is 130-150 pA, and the sequencing amplitude is about 20 pA. The DNA homopolymer recognition current signal is different from that of the single pore protein, as shown in Figure 22 (wherein A is the current signal comparison of the single constriction zone pore BCP34-1 and the double constriction zone BCP34-1-AP34-1; B is the current signal comparison of the main signal region of different AP34s; C is the current signal comparison of different cross-linking site modifications; and D is the current signal comparison of different lengths of AP34 polypeptides).

[0190] d) Sequencing characterization of DNA by nanopore complexes with different lengths of truncated AP34 and BCP34-2

[0191] The different length of polypeptides constructed in Example 5 were used to construct nanopore complex BCP34-2-AP34-P4, BCP34-2-AP34-P5, BCP34-2-AP34-P6 and BCP34-2-AP34-P7 with BCP34-2. DNA sequencing was performed on the complex using the test method of the present example. Except for BCP34-2-AP34-P6, which did not obtain a characteristic double constriction zone sequencing current signal, the rest all obtained. This may be due to the poor folding of the polypeptide.

[0192] Figure 19 is the current change and local details generated when the library DNA passes through the nanopore complex BCP34-2-AP34-P4 under the action of an applied voltage of 0.18V.

[0193] Figure 20 is the current change and local details generated when the library DNA passes through the nanopore complex BCP34-2-AP34-P5 under the action of an applied voltage of 0.18V.

[0194] Figure 21 is the current change and local details generated when the library DNA passes through the nanopore complex BCP34-2-AP34-P7 under the action of an applied voltage of 0.18V.

[0195] Figure 22 is the characteristic of the DNA homopolymer region current signal of different nanopores in the present example 7; A is the comparison of the current signal of single constriction pore BCP34-1 and double constriction BCP34-1-AP34-1; B is the comparison of the current signal of different AP34 main signal regions; C is the comparison of the current signal of different cross-linking sites; D is the comparison of the current signal of different length of AP34 polypeptides.

[0196] The vertical coordinates in Figures 13 to 22 are all current (pA), and the horizontal coordinates are all time (min). The current signal in a certain time scale is shown in each figure (the specific value can be ignored because: 1. All mutations or truncations of the protein involved in the present example are mainly to change the sequencing signal and improve the stability of the complex, but not to change the sequencing speed, so we do not focus on the sequencing speed, i.e. we do not focus on the time (horizontal coordinate difference) used for sequencing a complete read; 2. The sequencing output is not generated at the same time, and the sequencing signal reads taken here are randomly selected and have the representative complete read sequencing graph of the sequencing output signal at this time, so the sequencing time (i.e. the horizontal coordinate) has no reference meaning, so it can be ignored).

[0197] As can be seen from the above figures, the length of different AP34 has an effect on the sequencing signal, and the signal-to-noise ratio and resolution of the homopolymer sequence obtained are different.

[0198] From the above description, it can be seen that the above-mentioned embodiments of the present application achieve the following technical effects:

[0199] a) The vast majority of pore proteins currently used in nanopore sequencers on the market have only one constriction region, and the resolving power for homopolymers is poor, resulting in a low accuracy rate of nanopore sequencing. By developing a new generation of nanopore proteins with two or more constriction regions, the present application further improves the current difference between multiple nucleotides, which has the potential to enhance the base recognition ability and effectively improve the accuracy of nanopore sequencing.

[0200] b) The present application develops a new nanopore complex with sequencing ability to break through the type limitation of pore proteins in the field of nanopore sequencing and to develop more single-molecule sequencers with high accuracy, high integration and high stability.

[0201] c) The present application is based on the new pore protein monomer BCP34 and its mutants, by excavating the accessory protein AP3434 of the same genus as the protein, and truncating and modifying the accessory protein monomer AP3434, the nanopore complex is constructed together with the pore protein monomer BCP34, creating a nanopore protein complex including but not limited to two constriction regions, and applying it in the field of nanopore sequencing.

[0202] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A nanopore protein complex, characterized in that, The nanopore protein complex comprises: a nanopore protein polymerized from a plurality of pore protein monomers, and the plurality of pore protein monomers polymerize to form a hollow nanopore cavity; an accessory protein polymerized from a plurality of accessory protein monomers, at least part of the accessory protein is located in the nanopore cavity, and the accessory protein and the nanopore cavity jointly form a continuous channel; wherein, according to the moving direction of the analyte passing through the continuous channel, the continuous channel comprises a first sensing region and a second sensing region connected in sequence, the first sensing region is formed by at least part of the nanopore protein, and the second sensing region is formed by at least part of the accessory protein; the pore protein monomers are selected from any one of the following proteins: 1) a protein having the amino acid sequence shown in SEQ ID NO: 1; 2) a protein having at least 50% identity with the amino acid sequence shown in SEQ ID NO: 1 and having the function of polymerizing to form a nanopore protein; or 3) a protein having one or more amino acids substituted, deleted or added on the basis of SEQ ID NO: 1, and having the function of polymerizing to form a nanopore protein; the accessory protein monomers are selected from any one of the following proteins: i) a protein having the amino acid sequence shown in SEQ ID NO: 3; ii) a protein having at least 50% identity with the amino acid sequence shown in SEQ ID NO: 3 and having the function of binding with the pore protein monomers and polymerizing with the pore protein monomers to form the accessory protein; or iii) a protein having one or more amino acids substituted, deleted or added on the amino acid sequence shown in SEQ ID NO: 3, and having the function of binding with the pore protein monomers and polymerizing with the pore protein monomers to form the accessory protein.

2. The nanopore protein complex of claim 1, wherein, The nanopore protein and the accessory protein are derived from the same species; Optionally, all of the accessory proteins or at least part of the N-terminal of the accessory proteins are located in the nanopore cavity of the nanopore protein; Preferably, the accessory protein is attached in the nanopore cavity of the nanopore protein by covalent or non-covalent action; More preferably, the nanopore protein and the accessory protein exist in the form of a nonamer.

3. The nanopore protein complex of claim 1, wherein, The accessory protein monomers are truncated mutants of the amino acid sequence shown in SEQ ID NO: 3, and the truncated mutants are selected from any one of the following mutants: only retaining N-terminal 23-67 amino acids, only retaining N-terminal 23-57 amino acids, only retaining N-terminal 23-52 amino acids, only retaining N-terminal 23-54 amino acids, only retaining N-terminal 23-51 amino acids, only retaining N-terminal 23-50 amino acids, only retaining N-terminal 23-49 amino acids, only retaining N-terminal 23-48 amino acids, or only retaining N-terminal 23-47 amino acids.

4. The nanopore protein complex of claim 1, wherein, The length of the accessory protein monomers is 24-45 amino acids. Preferably, the amino acid sequence of the accessory protein monomer is from the following residue position interval of SEQ ID NO: 3 or a mutant thereof: 23-48, 23-49, 23-50, 23-51, 23-54 or 23-57.

5. The nanopore protein complex of claim 1, wherein, The said pore protein monomer is selected from proteins mutated at at least one amino acid site in any one or more of the following groups of SEQ ID NO: 1: 1) S71, N74, G75 and F76; Preferably, S71 is mutated to G, A or T; N74 is mutated to G, A, S or T; G75 is mutated to A, S, T or Q; F76 is mutated to A, S, T, N or Q; 2) E162, R196, S200 and S216; Preferably, E162 is mutated to A, G, V, L, I, Y, F or W; R196 is mutated to A, G, V, L, I, Y, F or W; S200 is mutated to A, G, V, L, I, Y, F or W; S216 is mutated to A, G, V, L, I, Y, F or W; 3) R103, E104, E112, R113, K114, R117, R119, D120, K122, D124, K154, R165, D199, R207, K209, K210, E213 or E215; Preferably, R103 is mutated to A, G, S, T, N or Q; E104 is mutated to K, R, A, G, S, T, N or Q; E112 is mutated to K, R, A, G, S, T, N or Q; R113 is mutated to A, G, S, T, N or Q; K114 is mutated to A, G, S, T, N or Q; R117 is mutated to A, G, S, T, N or Q; R119 is mutated to A, G, S, T, N or Q; D120 is mutated to K, R, A, G, S, T, N or Q; K122 is mutated to A, G, S, T, N or Q; D124 is mutated to K, R, A, G, S, T, N or Q; K154 is mutated to A, G, S, T, N or Q; R165 is mutated to A, G, S, T, N or Q; D199 is mutated to A, G, S, T, N or Q; R207 is mutated to A, G, S, T, N, Q, D or E; K209 is mutated to A, G, S, T, N, Q, D or E; K210 is mutated to A, G, S, T, N, Q, D or E; E213 is mutated to A, G, S, T, N or Q; E215 is mutated to A, G, S, T, N or Q.

6. The nanopore protein complex of claim 1, wherein, The said accessory protein monomer is selected from substitution of at least one amino acid of the amino acid sequence set forth in SEQ ID NO: 3: A39, T42 or N46; Preferably, A39 is mutated to S, T, N, G, V, L, I or Q; T42 is mutated to A, G or S; N46 is mutated to A, G or S.

7. The nanopore protein complex of claim 1, wherein, The said first sensing zone and the said second sensing zone have the same or different minimum pore diameter; Preferably, the minimum pore diameter is 1-3 nm.

8. The nanopore protein complex of any one of claims 1-7, wherein, the accessory protein monomer is selected from the group consisting of at least one amino acid mutation to cysteine or substitution with an unnatural amino acid in the amino acid sequence set forth in SEQ ID NO: 3 at S23, S24, L25, T28, K30, N31, S33, F34, N47, A50, Q51, N52, or Q53; and / or the pore protein monomer is selected from the group consisting of at least one amino acid mutation to cysteine or substitution with an unnatural amino acid in the amino acid sequence set forth in SEQ ID NO: 1 at N145, T148, K154, L156, L160, S161, R165, S195, Q197, D199, F203, Y205, K209, K210, L211, E213, E215, G217, S219, or N221.

9. The nanopore protein complex of claim 8, wherein, the pore protein monomer and the accessory protein monomer have at least one site combination of mutations, and the amino acid of each of the corresponding sites in each of the site combinations is mutated to cysteine or substituted with an unnatural amino acid: 1) R165 in SEQ ID NO: 1 and S23 in SEQ ID NO: 3; 2) S195 in SEQ ID NO: 1 and S24 in SEQ ID NO: 3; 3) S219 in SEQ ID NO: 1 and L25 in SEQ ID NO: 3; 4) N221 in SEQ ID NO: 1 and L25 in SEQ ID NO: 3; 5) N145C in SEQ ID NO: 1 and V26C in SEQ ID NO: 3; 6) D199 in SEQ ID NO: 1 and T28 in SEQ ID NO: 3; 7) D199 in SEQ ID NO: 1 and K30 in SEQ ID NO: 3; 8) E215 in SEQ ID NO: 1 and K30 in SEQ ID NO: 3: 9) E213 in SEQ ID NO: 1 and N31 in SEQ ID NO: 3: 10) K154 in SEQ ID NO: 1 and S33 in SEQ ID NO: 3: 11) E213 in SEQ ID NO: 1 and S33 in SEQ ID NO: 3: 12) E213 in SEQ ID NO: 1 and N47 in SEQ ID NO: 3; 13) L211 in SEQ ID NO: 1 and A50 in SEQ ID NO: 3; 14) L211 in SEQ ID NO: 1 and Q51 in SEQ ID NO: 3; 15) K209 in SEQ ID NO: 1 and N52 in SEQ ID NO: 3; 16) K210 in SEQ ID NO: 1 and Q53 in SEQ ID NO:

3.

10. The nanopore protein complex of claim 9, wherein, the amino acid of each of the corresponding sites in any of the site combinations is mutated to cysteine.

11. The nanopore protein complex according to any one of claims 8-10, wherein, the amino acid is substituted by a non-natural amino acid through a modification, which is a direct modification or an indirect modification; preferably, the direct modification comprises modification through spontaneous reaction or oxidation reaction of a side chain group of the amino acid; preferably, the indirect modification comprises chemical modification through attachment of a chemical small molecule; preferably, the chemical small molecule comprises a chemical cross-linking agent containing a functional group, and the functional group is resistant to dithiothreitol. preferably, a disulfide bond is formed between the site mutated to cysteine in at least part of the accessory protein monomer and the site mutated to cysteine in at least part of the pore protein monomer.

12. A nanopore sensor, characterized in that, The nanopore sensor comprises a membrane and the nanopore protein complex of any one of claims 1-11, wherein the nanopore protein in the nanopore protein complex is positioned on the membrane, and the nanopore cavity of the nanopore protein and the accessory protein together form a continuous channel across the membrane.

13. The nanopore sensor of claim 12, wherein, The membrane comprises a layer of amphiphilic molecules; preferably, the membrane is a phospholipid bilayer composed of diacylphosphatidylcholine, a two-block copolymer or a three-block copolymer.

14. A nanopore sequencing system, characterized in that, The nanopore sequencing system comprises the nanopore sensor of claim 12 or 13, and further comprises: a conductive solution; positive and negative electrodes for providing a voltage potential across the membrane, and a measuring device for measuring an electrical signal passing through the continuous channel; wherein the nanopore sensor is located in the conductive solution, and the conductive solution is divided into a first chamber and a second chamber.

15. The nanopore sequencing system of claim 14, wherein, The nanopore sequencing system further comprises a test molecule selected from the group consisting of polynucleotides, polypeptides and polysaccharides, preferably, the test molecule is selected from polynucleotides containing homopolymers; preferably, the test molecule can be transiently located within the continuous channel, and one end of the test molecule is located in the first chamber and the other end is located in the second chamber. The method comprises:

16. A method of nanopore sequencing, characterized in that, contacting the nanopore sequencing system of claim 14 with the test molecule of claim 15; applying a potential across the membrane so that the test molecule enters the continuous channel; and measuring the electrical signal generated when the test molecule moves relative to the continuous channel one or more times, thereby obtaining the sequence of the test molecule. The electrical signal comprises current block signal strength, current block duration or interval time of current block events.

17. The method of claim 16, wherein, The test molecule is a polynucleotide, and the nucleotides in the polynucleotide interact with a first constriction region and a second constriction region within the continuous channel, and wherein each of the first constriction region and the second constriction region is capable of distinguishing different nucleotides, so that the total current passing through the continuous channel is affected by the interaction between each of the first constriction region and the second constriction region and the nucleotide positioned at each of the regions.

18. The method of claim 16, wherein, A nucleic acid binding protein is used to control the movement of the polynucleotide relative to the continuous channel pore.

19. The method of claim 18, wherein, The preparation method comprises:

20. A method of preparing a nanopore protein complex according to any one of claims 1 to 11, characterised in that, respectively constructing an expression vector of the pore protein monomer and an expression vector of the accessory protein monomer; ​ obtaining the nanopore protein complex by co-expressing an expression vector of the pore protein monomer and an expression vector of the accessory protein monomer in a competent cell, and separating and purifying; or constructing the nanopore protein complex by in vitro recombination.

Citation Information

Patent Citations

  • Nanopores with internal protein adaptors

    CN107735686A

  • Pore

    CN113195736A

  • Double-portal pore protein, pore protein mutant, nucleotide sequence and application of double-portal pore protein and pore protein mutant

    CN115974984A

  • Novel protein pores

    CN117106037A

  • pore

    US20220056517A1