Nanopore protein monomer and use thereof

By designing stable nanoporin monomers and constructs, the problems of limited nanoporin selection and unsuitable pore size were solved, thereby improving the stability and accuracy of DNA sequencing.

WO2025236238A1PCT designated stage Publication Date: 2025-11-20BEIJING POLYSEQ BIOTECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/093630
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

The selection of nanoporous proteins is limited, and different substrates require nanopores of different pore sizes, which affects the stability and accuracy of sequencing results.

Method used

We designed a nanoporous protein based on a naturally occurring non-shrinkable region. By artificially designing shrinkable regions and using amino acid sequence substitutions, deletions, and additions, combined with protein tags, we formed stable nanoporous protein monomers and constructs for DNA detection.

Benefits of technology

It improves the stability and pore current of nanoporous proteins, enhances DNA sequencing performance, and meets the detection needs of different substrates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024093630_20112025_PF_FP_ABST
    Figure CN2024093630_20112025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention are a nanopore protein monomer and a use thereof. The nanopore protein monomer is any one of the following polypeptides: (a1) a polypeptide having an amino acid sequence as set forth in SEQ ID NO: 4; (a2) a polypeptide derived from the amino acid sequence as set forth in SEQ ID NO: 4 by means of substitution and / or deletion and / or addition of one or more amino acids and having the same function; (a3) a polypeptide having 80% or above identity to the amino acid sequence defined in any one of (a1)-(a2) and having the same function; and (a4) a fusion polypeptide obtained after linking a tag to the terminus of the polypeptide defined in any one of (a1)-(a3). In the present invention, a nanopore protein which is stable and can be used for DNA detection is obtained on the basis of nanopores without a natural constriction region by a method of artificially designing a contraction region.
Need to check novelty before this filing date? Find Prior Art

Description

A nanopore protein monomer and application thereof TECHNICAL FIELD

[0001] The present application relates to the field of biotechnology, in particular to a nanopore protein monomer and application thereof. BACKGROUND

[0002] DNA, RNA, protein and polysaccharide, etc. biological macromolecules are the basic substances constituting living beings, and the sequence information and group post-modification determine their biological functions. The sequence identification technology of biological macromolecules is the core tool for analyzing the operation rules of life, and the single molecule sequencing technology emerges as the times require.

[0003] At present, single molecule sequencing technology can be mainly divided into two categories: one is optical zero mode waveguide sequencing, represented by Pacific Biosciences (Pacbio) in the United States; and the other is nanopore sequencing based on electricity, represented by Oxford Nanopore Technologies (ONT) in the United Kingdom. Nanopore sequencing technology is a new type of nucleic acid sequencing technology developed in recent years, which can be divided into solid-state pores and biological nanopores according to the types of pores. The pore protein of biological nanopore can allow the substrate to pass through. The following nanopore sequencing refers to biological nanopore sequencing technology.

[0004] Under the action of electric field force, charged nucleic acid substrate can pass through biological nanopore. When the nucleic acid passes through the nanopore, it can hinder the electric current of the nanopore, and different electric current signals are generated. By analyzing the electric current signal, the base information of the nucleic acid can be obtained. Compared with other sequencing methods, it has the advantages of low equipment cost, simple sample preparation, fast sequencing speed, etc. and has begun to be applied in various fields. The specific advantages are as follows: the library can be built simply without amplification; the signal reading speed is fast, usually reaching 200-300bp / s; the reading length is long, usually reaching thousands of bases; and the modification existing on the DNA can be directly detected; RNA can be directly detected, based on the characteristics of nanopore sequencing, RNA no longer needs to be reverse transcribed into DNA for sequence analysis, thereby the modification information on the RNA can be preserved. Because of the above advantages, nanopore sequencing technology has attracted widespread attention in recent years.

[0005] Nanopore channel protein is the core of nanopore sequencing. So far, there are relatively few nanopore proteins that can be used for nucleic acid detection, only a few natural proteins such as MspA (Mycobacterium smegmatis pore protein A), and CsgG, CsgG-CsgF (curli-specific transport channel) meet the requirements. The stability of natural nanopore protein, the diameter distribution of the pore, and the charge properties of the amino acids inside the channel have a great influence on the sequencing results.

[0006] SUMMARY

[0007] In order to solve the problems of few choices of nanopore proteins and different pore diameters required for different substrates, the patent proposes a nanopore based on a natural non-constriction region, and a nanopore protein capable of being used for DNA detection is obtained by artificially designing a constriction region.

[0008] The application provides a nanopore protein monomer, which is a polypeptide as described in any one of the following (a1) to (a4): (a1) a polypeptide having an amino acid sequence as shown in SEQ ID NO: 4; (a2) a polypeptide obtained by substituting, deleting and / or adding one or more amino acids in the amino acid sequence shown in SEQ ID NO: 4 and having the same function; (a3) a polypeptide having more than 80% identity with the amino acid sequence defined in any one of (a1) to (a2) and having the same function; and (a4) a fusion polypeptide obtained by connecting a tag to the terminal of the polypeptide defined in any one of (a1) to (a3).

[0009] The tag refers to a polypeptide or protein fused and expressed with the target protein by using DNA in vitro recombination technology, so as to facilitate the expression, detection, tracking and / or purification of the target protein. The protein tag can be a Flag tag, a His tag, an MBP tag, an HA tag, a myc tag, a GST tag and / or a SUMO tag, etc.

[0010] The nanopore protein monomer can be specifically R1-R10 protein monomers prepared in the following examples. For example, the amino acid sequence of the nanopore protein monomer is as shown in any one of SEQ ID NO: 2-11.

[0011] Alternatively, according to the nanopore protein monomer, the polypeptide of (a2) comprises at least one substitution of leucine at position 74 to proline, valine, threonine, serine, isoleucine, glycine, cysteine, aspartic acid, alanine, methionine, asparagine, glutamine, glutamic acid, arginine, lysine, histidine, tryptophan, tyrosine or phenylalanine.

[0012] The nanopore protein monomer can be specifically P6, P8-P16, P23-P25, P36-P46 protein monomers prepared in the following examples. For example, the amino acid sequence of the nanopore protein monomer is as shown in any one of SEQ ID NO: 13, 15-37.

[0013] Optionally, according to the above-mentioned nanopore protein monomer, the polypeptide of (a2) further comprises deletion of one or several amino acids of at least one of the following: deletion of one, two, three, four or five amino acids at positions 74-79 of SEQ ID NO: 4; and / or, deletion of one, two, three, four, five, six or seven amino acids at positions 67-73 of SEQ ID NO: 4.

[0014] The deletion of one or several amino acids can be deletion of the amino acid at position 74 of SEQ ID NO: 4; deletion of the amino acid at position 75 of SEQ ID NO: 4; deletion of the amino acids at positions 75-76 of SEQ ID NO: 4; deletion of the amino acids at positions 75-77 of SEQ ID NO: 4; deletion of the amino acids at positions 75-78 of SEQ ID NO: 4; deletion of the amino acids at positions 75-79 of SEQ ID NO: 4; deletion of the amino acid at position 73 of SEQ ID NO: 4; deletion of the amino acids at positions 72-73 of SEQ ID NO: 4; deletion of the amino acids at positions 71-73 of SEQ ID NO: 4; deletion of the amino acids at positions 70-73 of SEQ ID NO: 4; deletion of the amino acids at positions 69-73 of SEQ ID NO: 4; deletion of the amino acids at positions 68-73 of SEQ ID NO: 4; or deletion of the amino acids at positions 67-73 of SEQ ID NO: 4.

[0015] The above-mentioned nanopore protein monomer can be specifically the P7, P54-P64 protein monomers prepared in the following examples. For example, the amino acid sequence of the above-mentioned nanopore protein monomer is shown in any one of SEQ ID NO: 14, 38-48.

[0016] The present application also provides a construct comprising at least two linked nanopore protein monomers, at least one of which is the above-mentioned nanopore protein monomer. The construct retains the ability to form a pore.

[0017] The present application also provides a nanopore protein comprising at least two linked nanopore protein monomers, at least one of which is the above-mentioned nanopore protein monomer; or the nanopore protein comprises the above-mentioned construct.

[0018] Optionally, according to the above-mentioned construct or the above-mentioned nanopore protein, the nanopore protein monomers are the same or different; the linkage is covalent or non-covalent linkage; the nanopore protein comprises 10-12 nanopore protein monomers, for example, 10, 11 or 12 nanopore protein monomers.

[0019] The nanopore protein can be R1-R10, P6-P16, P23-P25, P36-P46 protein prepared in the following examples.

[0020] The application also provides the above-mentioned nanopore protein monomer, the above-mentioned construct or the above-mentioned nanopore protein related biological material, wherein the biological material is any one of the following: a1) a nucleic acid molecule encoding the above-mentioned nanopore protein monomer, the above-mentioned construct or the above-mentioned nanopore protein; a2) an expression cassette containing the nucleic acid molecule of a1); a3) a recombinant vector containing the nucleic acid molecule of a1) or containing the expression cassette of a2); a4) a recombinant cell containing the nucleic acid molecule of a1) or containing the expression cassette of a2) or containing the recombinant vector of a3).

[0021] In the above-mentioned biological material, the nucleic acid molecule can be DNA, such as cDNA, genomic DNA or recombinant DNA; the nucleic acid molecule can also be RNA, such as mRNA, siRNA, shRNA, sgRNA, miRNA or antisense RNA.

[0022] In the above-mentioned biological material, the expression cassette refers to DNA that can express genes in host cells, which not only includes promoters that initiate gene transcription, but also includes terminators that terminate gene transcription. Further, the expression cassette can also include enhancer sequences.

[0023] Alternatively, the nucleic acid molecule of a1) is any one of the following DNA molecules: 1) a DNA molecule encoding the sequence of R1-R10, P6-P16, P23-P25, P36-P46 encoding genes; 2) a DNA molecule with the nucleotide sequence of R1-R10, P6-P16, P23-P25, P36-P46 encoding genes; 3) a DNA molecule that hybridizes to the nucleotide sequence defined in 1) or 2) under stringent conditions, and encodes the above-mentioned nanopore protein monomer, the above-mentioned construct or the above-mentioned nanopore protein.

[0024] Alternatively, the recombinant vector of a3) is a vector with the DNA sequence of R1-R10, P6-P16, P23-P25, P36-P46 encoding genes, such as the R1-R10, P6-P16, P23-P25, P36-P46 expression vector prepared in the following examples.

[0025] The stringent conditions can be hybridization and membrane washing at 65°C in a solution of 0.1×SSPE (or 0.1×SSC), 0.1% SDS.

[0026] The use of the above-mentioned nanopore monomer, the above-mentioned construct, the above-mentioned nanopore or the above-mentioned related biomaterial in detecting the presence, absence or one or more characteristics of a target analyte or in preparing a product for detecting the presence, absence or one or more characteristics of a target analyte also falls within the protection scope of the present application.

[0027] The present application also provides a method for determining the presence, absence or one or more characteristics of a target analyte, which comprises: A. contacting the target analyte with the above-mentioned nanopore, so that the target analyte moves relative to the nanopore; B. obtaining one or more measurement values when the target analyte moves relative to the nanopore, thereby determining the presence, absence or one or more characteristics of the target analyte.

[0028] Optionally, the method further comprises the step of applying a potential difference when the target analyte is contacted with the above-mentioned nanopore.

[0029] Optionally, the measurement values are obtained by electrical measurement and / or optical measurement. For example, the electrical measurement includes but is not limited to current measurement, impedance measurement, tunnel measurement, wind tunnel measurement or field effect transistor (FET) measurement, etc. In a specific embodiment of the present application, the measurement values are current measurement values.

[0030] The present application also provides a kit for determining the presence, absence or one or more characteristics of a target analyte, which comprises the above-mentioned nanopore monomer, the above-mentioned construct, the above-mentioned nanopore or the above-mentioned related biomaterial, and a membrane.

[0031] The present application also provides a device for determining the presence, absence or one or more characteristics of a target analyte, which comprises the above-mentioned nanopore, and a membrane.

[0032] Preferably, the target analyte is one or more of a nucleotide, a nucleic acid, an amino acid, an oligopeptide, a polypeptide, a protein.

[0033] In the above-mentioned kit or device, the membrane and the nanopore can be packaged independently, or the nanopore can be embedded in the membrane.

[0034] The membrane can be any membrane existing in the prior art, and is preferably an amphiphilic layer, i.e. a layer formed by amphiphiles such as phospholipids having at least one hydrophilic part and at least one lipophilic or hydrophobic part, which can be synthetic or naturally occurring. For example, the membrane is a phospholipid monolayer membrane.

[0035] The kit or device can further comprise a rate controlling protein. The rate controlling protein can comprise one or more combinations of a nucleic acid binding protein, a helicase, an exonuclease, a telomerase, a topoisomerase, a transcriptase, a transposase, and / or a polymerase.

[0036] Optionally, the helicase is selected from the group consisting of a Hel308 family helicase and a modified Hel308 family helicase, a RecD helicase and a variant thereof, a TrwC helicase and a variant thereof, a Dda helicase and a variant thereof, a TraI Eco and a variant thereof, a XPD Mbu and a variant thereof, a Pif1-like helicase and a variant thereof.

[0037] Optionally, the target analyte is one or more of a nucleotide, a nucleic acid, an amino acid, an oligopeptide, a polypeptide, and a protein.

[0038] Optionally, the one or more characteristics are selected from the group consisting of (i) a length of the target analyte; (ii) an identity of the target analyte; (iii) a sequence of the target analyte; (iv) a secondary structure of the target analyte; and (v) whether the target analyte is modified. "Identity" refers to similarity between sequences. Identity can be assessed by eye or by computer software. Using computer software, identity between two or more sequences can be expressed as a percentage (%), which can be used to assess identity between related sequences.

[0039] Optionally, the nucleic acid can be naturally occurring or artificially synthesized. In particular, the nucleic acid can be natural DNA, RNA, or modified DNA or RNA, or artificially synthesized nucleic acid, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers with nucleoside side chains.

[0040] Optionally, the nucleic acid can be single-stranded, double-stranded, or at least a portion of which is double-stranded.

[0041] Optionally, the nucleic acid can be of any length. For example, the nucleic acid can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs, or 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs, or 100000 or more nucleotides or nucleotide pairs.

[0042] Optionally, one or more nucleotides in the nucleic acid can be modified, such as methylated, oxidized, damaged, abasic, protein-labeled, tagged, or have an intervening spacer.

[0043] The nanopore protein monomer, the construct or the protein can be artificially synthesized, or a gene encoding the same can be synthesized and then expressed biologically.

[0044] The application further provides a nanopore protein preparation method as described above, which comprises transforming a host cell with the recombinant vector as described above, and inducing the host cell to express the nanopore protein.

[0045] In the present application, "comprising" or "including" is used to describe the sequence of a protein or nucleic acid, which can be composed of the sequence, or can have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still has the activity described in the present application.

[0046] In the present application, the standard one-letter code for amino acids is used. These are as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y) and valine (V). The standard substitution notation is also used, i.e. L74V means that the L at position 74 of the sequence is replaced by V.

[0047] In the present application, the high-resolution three-dimensional structure of the NPB nanopore is analyzed by cryo-EM technology The structure analysis shows that the width of the constriction region of the NPB nanopore protein is Unlike other nanopore proteins, the NPB nanopore protein has almost no constriction region, has a large pore current on a phospholipid membrane and a large noise, and cannot be used for nucleic acid substrate detection. However, the wild-type NPB nanopore protein has strong stability, and the constriction region has no effect on the oligomerization of the pore protein, and is a good pore skeleton.

[0048] Currently, protein design is becoming mature, and its application is becoming widespread. The present application uses protein design to solve the problem of the absence of a constriction region in NPB, and attempts a more stable constriction region type (alpha helix type), which can effectively increase the types of nanopores that can be used for sequencing.

[0049] On the basis of the wild-type NPB nanopore, the present application designs 10 different sequences of constriction regions (R1-R10), and proves that the designed constriction regions can stabilize the overall pore current and improve the DNA sequencing performance through protein expression, purification, structure analysis and current detection. The stability and uniformity of the designed nanopore are improved through amino acid mutation, and the pore current and the amplitude of the DNA substrate pore signal are increased through amino acid deletion, so that the DNA sequencing requirements can be met. BRIEF DESCRIPTION OF DRAWINGS

[0050] Figure 1 is the NPB wild type protein purification and molecular sieve results, "+" is the protein sample heated at 100°C for 10 minutes; "-" is the protein sample not heated, placed at room temperature for 10 minutes.

[0051] Figure 2 is the NPB wild type protein structure analysis results.

[0052] Figure 3 is the design based on different shrinkage regions of NPB wild type.

[0053] Figure 4 is the R1-R10 protein purification and partial molecular sieve results, "+" is the protein sample heated at 100°C for 10 minutes; "-" is the protein sample not heated, placed at room temperature for 10 minutes.

[0054] Figure 5 is the NPB wild type and R1-R10 partial pore current and DNA translocation results.

[0055] Figure 6 is the R3 structure analysis results.

[0056] Figure 7 is the expression results of the modified proteins based on R3, "+" is the protein sample heated at 100°C for 10 minutes; "-" is the protein sample not heated, placed at room temperature for 10 minutes.

[0057] Figure 8 is the P8-P10 pore current and DNA translocation results.

[0058] Figure 9 is the P10 structure analysis results.

[0059] Figure 10 is the expression results of the deletion proteins based on P10.

[0060] Figure 11 is the P10, P54 and P61 translocation property comparison results.

[0061] Figure 12 is the P59, P62, P63 and P64 pore current and translocation signal results.

[0062] Figure 13 is the P61 DNA substrate translocation property identification results. DETAILED DESCRIPTION

[0063] The application will be further described in conjunction with the specific embodiments. The examples given are only to illustrate the application, and are not intended to limit the scope of the application. The examples provided below can serve as a guide for further improvement by those of ordinary skill in the art, and do not in any way constitute a limitation on the application. In the following examples, the experimental methods are conventional methods, unless otherwise specified, and are performed according to the techniques or conditions described in the literature in the art or according to the product instructions. In the following examples, the materials, reagents, etc. used, unless otherwise specified, can be obtained commercially. In the following examples, quantitative tests were set up in triplicate, and the results were averaged.

[0064] Example 1 Preparation of wild-type NPB nanopore expression vector and protein

[0065] 1. Construction of wild-type NPB nanopore protein vector

[0066] The wild-type NPB nanopore protein is derived from Thermodesulfobacteriota bacterium (ACCESSION: NPB10001, SEQ ID NO: 1), and the protein expression gene is obtained by artificial synthesis. The synthesis process is subjected to codon optimization suitable for E. coli expression. After synthesis, the gene is constructed into pBAD22 vector by seamless cloning to obtain the wild-type NPB nanopore protein vector. The protein C-terminal adds 6×his as an affinity purification tag. The wild-type NPB nanopore protein vector contains a wild-type NPB nanopore protein expression gene, which expresses a wild-type NPB nanopore protein with a 6×his tag. The wild-type NPB nanopore protein sequence is shown in SEQ ID NO: 1.

[0067] The vector construction steps are as follows:

[0068] Using the wild-type NPB nanopore protein expression gene (SEQ ID NO: 50) as the template, the target gene fragment was amplified by PCR using forward and reverse primers (primer F: CAGGAGGAATTAACCATGTTTCGCCTGCTGACCC (SEQ ID NO: 51); primer R: GAACTGCGGGTGGCTCCATTTGGTGCCGCTCGCCG (SEQ ID NO: 52)). After gel recovery, the linearized pBAD22 vector was ligated, and then transformed into DH5α competent cells for positive clone screening. Two clones were picked for sequencing, and after sequencing, the constructed plasmid was stored at -20°C for standby.

[0069] The PCR system is as follows (20 μL):

[0070] The PCR program is as follows:

[0071] Seamless ligation system (10 μL) as follows:

[0072] 2x Seamless ligation buffer 5 μL

[0073] Target fragment (50 ng / μL) 3 μL

[0074] Linearized vector (10 ng / μL) 2 μL

[0075] After 15 min of reaction at 50°C, 2 μL was transferred into DH5α cells, and the spots were sequenced.

[0076] 2. Preparation of wild-type NPB nanopore protein

[0077] After Ni column affinity chromatography and molecular sieve purification, the wild-type NPB nanopore protein had high purity (Figure 1, where A is the SDS-PAGE result), and the protein was uniform (Figure 1, where B is the molecular sieve result). On SDS-PAGE, heating produced monomeric protein (molecular weight about 26 kDa); without heating, the protein was mainly in an oligomeric state (i.e., in a pore state, molecular weight greater than the maximum band of marker 180 kDa), and there was also a small amount of monomer, indicating that the wild-type NPB protein had poor stability.

[0078] The expression and purification steps are as follows:

[0079] 1) After the wild-type NPB nanopore protein vector was sequenced correctly, it was transferred into BL21 (DE3) for expression. After obtaining 1 mL of seed liquid at 37°C and 200 rpm, it was transferred into 1 L of LB medium, and the culture was incubated at 37°C and 200 rpm until the OD 600 was 1.2. The temperature was then reduced to 26°C, and the final concentration of 0.4 g / L of arabinose was added for overnight induction. 2) The bacterial cells were collected at 4000 rpm, and 20 mL of lysis buffer was used to resuspend the cells per 1 L of bacteria. After high-pressure disruption, the membrane components were collected by centrifugation at 18000 rpm and 4°C for 1 hour. 3) The membrane components were resuspended with 10 mL of membrane solubilization buffer per 1 L of bacteria, and the membrane proteins were extracted by magnetic stirring at 4°C for 1 hour. The membrane protein components were collected by centrifugation at 18000 rpm and 4°C for 1 hour. 4) After the addition of imidazole to the supernatant to a final concentration of 30 mM, the Ni beads were incubated with the membrane solubilization buffer, and the affinity purification was performed after binding at 4°C for 1 hour. 5) The supernatant and Ni beads were introduced into the column, and the flow was naturally washed by gravity. The protein was eluted with 5 mL of elution buffer after washing with 10 mL of washing solution. 6) The purity of the target protein was detected by SDS-PAGE after molecular sieve purification, and the results are shown in Figure 1A, and the molecular sieve results are shown in Figure 1B.

[0080] Lysis buffer: 20 mM Tris-HCl, 150 mM NaCl, pH 8.0.

[0081] Solubilization buffer: 20 mM Tris-HCl, 150 mM NaCl, pH 8.0, 1% LDAO (lauryldimethylamine oxide).

[0082] Washing buffer: 20 mM Tris-HCl, 150 mM NaCl, pH 8.0, 0.5% LDAO, 50 mM imidazole.

[0083] Elution buffer: 20 mM Tris-HCl, 150 mM NaCl, pH 8.0, 0.1% LDAO, 200 mM imidazole.

[0084] Molecular sieve buffer: 20 mM Tris-HCl, 150 mM NaCl, pH 8.0, 0.06% LDAO.

[0085] Example 2: Wild-type NPB nanopore protein atomic level structure determination

[0086] After obtaining the wild-type NPB nanopore protein with high purity and homogeneity (Example 1), the atomic level structure thereof was resolved using cryo-EM technology, with a resolution of 3.2 A.

[0087] The steps for cryo-EM structure resolution are as follows:

[0088] 1) Prepare the sample after molecular sieve purification, and prepare several standby samples under the same frozen sample conditions, select 8 suitable samples to be placed in the Talos Arctica 200kV high-end electron microscope; 2) After the sample is placed, wait for the vacuum and temperature to stabilize, then open the lens and select the appropriate square hole under low magnification. This step is similar to the selection of frozen samples, and the purpose is to select the appropriate square hole for data collection in the following step; 3) According to the machine time space of the electron microscope and the number of photos that can be collected by each square hole, calculate the approximate number of square holes that need to be selected, and then select the square holes and take a map; 4) During the map taking process, the holes in the map that can be used for data collection can be selected offline, which can save a part of the time; 5) After the map taking is completed, the electron microscope can be adjusted for data collection. Mainly includes the co-axial adjustment of the electron microscope, the background subtraction and the setting of the basic parameters of data collection (underfocus amount -1.5 μm to -2.5 μm, electron dose 32 frames and pixel size ).

[0089] ​After data processing and structure analysis, the density map of the wild type NPB was obtained, as shown in FIG. 2. The overall map is a 11-mer structure. Through homology modeling, structure building and refinement, the atomic coordinates of the wild type NPB nanopore protein amino acid were obtained (FIG. 2, FIG. A is the top view and side view of the map, and FIG. B is the resolution information).

[0090] The NPB wild type constriction region has fewer amino acids (3, N58-C59-Q60) and a large pore diameter (FIG. 2, FIG. C is the model building result of the NPB wild type, top view and side view), and the overall stability is strong. Considering the current properties of the pore protein, which are largely determined by the constriction region, the NPB wild type can be used as the main framework for the design of new pore channels, and different types of constriction regions can be designed.

[0091] Example 3: Protein design based on the constriction region of NPB and preparation of the corresponding protein (NPB-R)

[0092] 1. Protein design

[0093] The protein design program involved in the present application includes RFdiffusion for generating protein skeleton structures, ProteinMPNN for generating one or more sequences from the skeleton, and AlphaFold2 for predicting structures from sequences. The predicted structure is analyzed to determine whether it meets the requirements, and the process is repeated until a reasonable structure of the amino acid sequence is generated. Finally, the feasibility of the designed sequence is verified by means of vector construction, protein expression, structure analysis, pore current detection and the like.

[0094] The structure of the nanopore protein (CsgG) currently capable of accurately sequencing nucleic acids has a constriction zone that is a loop structure with relatively large flexibility. This structure lacks rigidity and will swing to a large extent under the action of an electric field force and a DNA substrate, generating noise and affecting sequencing accuracy. According to the results of Example 2, the wild-type structure of NPB is used as the main framework to design a new constriction zone to improve the adverse effects of the flexible constriction zone. The main part of the designed constriction zone is mainly an alpha-helix, which can be divided into loop-alpha-helix, alpha-helix-loop-alpha-helix, etc. The conformation and pore diameter have diversity. FIG. 3 is a design based on different types of constriction zones of the wild-type NPB, in which A is a top view of a model of three nanopore proteins, NPB-WT is a model of the wild-type NPB, NPB-R3 and NPB-R4 are models of nanopore proteins with different constriction zone design types based on the wild-type structure of NPB as the main framework; B shows the structure characteristics of a monomer constriction zone, NPB-WT is a wild-type NPB monomer, and NPB-R1-R10 are R1-R10 nanopore protein monomers with different constriction zone design types based on the wild-type structure of NPB as the main framework. Table 1 is the theoretical diameter of the R1-R10 nanopore protein with different constriction zone design types based on the wild-type structure of NPB as the main framework. Table 2 is the pore amino acid sequence of the R1-R10 nanopore protein with different constriction zone design types based on the wild-type structure of NPB as the main framework.

[0095] Table 1: Theoretical diameter of the R1-R10 design pore

[0096] Table 2: Amino acid sequence of the R1-R10 design pore

[0097] 2. Vector and protein preparation

[0098] After the AlphaFold2 predicted structure is determined to be reasonable, the expression vector is constructed by the method described in Example 1. The wild-type NPB nanopore protein vector gene is used as a template, and the corresponding primers (f represents a forward primer, and r represents a reverse primer) are used to PCR amplify the target fragments (i.e., PCR fragments and PCR vectors). After gel recovery, the PCR fragments and the PCR vectors are ligated, and then transformed into DH5a competent cells for positive clone screening. Two clones are picked for sequencing, and after correct sequencing, the constructed plasmid is stored at -20°C for standby. The R1-R10 expression vectors are obtained according to the foregoing method.

[0099] The R1 expression vector contains the R1 coding gene, which expresses the R1 protein with a 6xhis tag. The sequence of the R1 protein is shown in SEQ ID NO: 2. The nucleotide sequence of the R1 coding gene is shown in SEQ ID NO: 64.

[0100] The R2 expression vector contains an R2-encoding gene that expresses an R2 protein with a 6xhis tag, the sequence of which is shown in SEQ ID NO: 3. The nucleotide sequence of the R2-encoding gene is shown in SEQ ID NO: 65.

[0101] The R3 expression vector contains an R3-encoding gene that expresses an R3 protein with a 6xhis tag, the sequence of which is shown in SEQ ID NO: 4. The nucleotide sequence of the R3-encoding gene is shown in SEQ ID NO: 66.

[0102] The R4 expression vector contains an R4-encoding gene that expresses an R4 protein with a 6xhis tag, the sequence of which is shown in SEQ ID NO: 5. The nucleotide sequence of the R4-encoding gene is shown in SEQ ID NO: 67.

[0103] The R5 expression vector contains an R5-encoding gene that expresses an R5 protein with a 6xhis tag, the sequence of which is shown in SEQ ID NO: 6. The nucleotide sequence of the R5-encoding gene is shown in SEQ ID NO: 68.

[0104] The R6 expression vector contains an R6-encoding gene that expresses an R6 protein with a 6xhis tag, the sequence of which is shown in SEQ ID NO: 7. The nucleotide sequence of the R6-encoding gene is shown in SEQ ID NO: 69.

[0105] The R7 expression vector contains an R7-encoding gene that expresses an R7 protein with a 6xhis tag, the sequence of which is shown in SEQ ID NO: 8. The nucleotide sequence of the R7-encoding gene is shown in SEQ ID NO: 70.

[0106] The R8 expression vector contains an R8-encoding gene that expresses an R8 protein with a 6xhis tag, the sequence of which is shown in SEQ ID NO: 9. The nucleotide sequence of the R8-encoding gene is shown in SEQ ID NO: 71.

[0107] The R9 expression vector contains an R9-encoding gene that expresses an R9 protein with a 6xhis tag, the sequence of which is shown in SEQ ID NO: 10. The nucleotide sequence of the R9-encoding gene is shown in SEQ ID NO: 72.

[0108] The R10 expression vector contains an R10-encoding gene that expresses an R10 protein with a 6xhis tag, the sequence of which is shown in SEQ ID NO: 11. The nucleotide sequence of the R10-encoding gene is shown in SEQ ID NO: 73.

[0109] The PCR system and PCR program were the same as in Example 1. The seamless ligation system was the same as in Example 1, except that the linearized vector was replaced with the PCR vector. Vector primers: V-f: gatatgctggcgaccgcgc (SEQ ID NO: 74); V-r: cgctttcgggccgcgat (SEQ ID NO: 75).

[0110] Fragment primers:

[0111] R1-f: CCAGGATGCGCGTGCTACGGAACTGGCAGCACTGGCTCTGGCGATCGCGGTTTAC gatatgctggcgaccgcgc (SEQ ID NO: 76);

[0112] R1-r: GCACGCGCATCCTGGCTACCAGTGTATACAGCAGTGAACAGTACAACAACAATAC G cgctttcgggccgcgat (SEQ ID NO: 77);

[0113] R2-f: TCAGGACGCACGTGCAACCGAACTGGCTGCAAAAGCTCTGGCTATCGCAGTTTAC gatatgctggcgaccgcgc (SEQ ID NO: 78);

[0114] R2-r: GCACGTGCGTCCTGAGAACCAGTATACACTGCGGTGAACAGAACAACAACAACAC G cgctttcgggccgcgat (SEQ ID NO: 79);

[0115] R3-f: ATGTTTCCGATGCAGAAGCTGCAGCTCTGGCTGCACGTCTGGCTGCTGCTGTCAT G GC Agatatgctggcgaccgcgc (SEQ ID NO: 80);

[0116] R3-r: CTGCATCGGAAACATCTGCCAGATCAGACATAACCATGCGCAGCGGCAGAACAACT ACACG cgctttcgggccgcgat (SEQ ID NO: 81);

[0117] R4-f: CTGATGAACATGAACGCGGTTACCGAACTGCGCCACACCATGGTTGTACTGGCGGCACTGCTGGCAGCAATG gatatgctggcgaccgcgc (SEQ ID NO: 82);

[0118] R4-r: GTTCATGTTCATCAGGATTTCCGCCAGCAGACGAACGATCTTCGCAAACAGATCGTTAGCGACTACAACGAA cgctttcgggccgcgat (SEQ ID NO: 83);

[0119] R5-f: CTGCAGGCCAGCAACGCGGTCGAGACCCTGGAAAAACAGGCTGCACGTCTGGCTGCCCTGCTGTCCCTGATG gatatgctggcgaccgcgc (SEQ ID NO: 84);

[0120] R5-r: GTTGCTGGCCTGCAGCAGTTTGATCATCATTTCAACGATTTTATACCAGGTGTCGTTGACAACAACAACGTA cgctttcgggccgcgat (SEQ ID NO: 85);

[0121] R6-f: TGCATATGTTGTGCGTGATCGTGACGCTGATAAAGCTTTCATGAAAATCCTGACGTGGCTGGCGCTGATG gatatgctggcgaccgcgc (SEQ ID NO: 86);

[0122] R6-r: CGCACAACATATGCAGCCAGTGCAGCAGCACGCAGAGCTTCTTCACGGTTCAGACCACCTACGATTACGTA cgctttcgggccgcgat (SEQ ID NO: 87);

[0123] R7-f: CTCAGGTAATGAACTCCGATAAGCTGGAACGCCTGGCGGCCGTACTGGCGGCGATGCTGGCGCTGATG gatatgctggcgaccgcgc (SEQ ID NO: 88);

[0124] R7-f: AGTTCATTACCTGAGCCGCCAGAGCGAAAGCGGTAACCAGCAGTTTCAGGTTGGCGATAACTACAACGT Acgctttcgggccgcgat (SEQ ID NO: 89);

[0125] R8-f: CTGGCGATGAAAGCGGCGGCTAAGAACGCGAACAACGGTCTGTACACCGCGTTCGCCCTGGTGGCGTACATG gatatgctggcgaccgcgc (SEQ ID NO: 90);

[0126] R8-r: CGCTTTCATCGCCAGTTTCAGGGCTTCCATGTAAACACGCATTACCTCCTGGTTGGACAGCAGAACCACGT Acgctttcgggccgcgat (SEQ ID NO: 91);

[0127] R9-f: CAAAACCCAGCTGCGCAAAGACAACCGTGCTGACGCCGCGCTGGTA gatatgctggcgaccgcgc (SEQ ID NO: 92);

[0128] R9-r: CGCAGCTGGGTTTTGATCAGTTCTTTGTTAATCACCACCACAACGT Acgctttcgggccgcgat (SEQ ID NO: 93);

[0129] R10-f: ATCGTGACGCAGATATGGCGGAACTGCTGCTGCGTGCTCTGGCTATCGCAGCATTC gatatgctggcgaccgcgc (SEQ ID NO: 94);

[0130] R10-r: TATCTGCGTCACGATCGCCGGTATAAACCGCACCCATCTTAACAACAACAACACG cgctttcgggccgcgat (SEQ ID NO: 95).

[0131] The sequencing correct vectors were used to express the proteins (R1-R10) according to the method described in Example 1. The expression of the proteins was different (Figure 4, where A is the SDS-PAGE result of 10 proteins), in which R1, R4, R5 and R7 were expressed in very low amounts; R2, R3, R6, R8 and R9 were expressed in relatively high amounts, and R10 was not oligomerized. Therefore, the oligomerization of the proteins was verified by molecular sieving (Figure 4, where B is the molecular sieving result of 4 proteins), and the results showed that the oligomerization of R3 and R9 was relatively uniform, and the oligomerization of R2 and R8 was not good. This result also reflects the diversity of the design results of the proteins.

[0132] 3. DNA detection ability identification

[0133] Further, the properties of the wild-type NPB and the designed 9 kinds of proteins (R1-R9) were identified. The recording process of the film-forming embedded hole and the hole current, DNA pore signal is as follows: the artificial phospholipid monolayer (dipalmitoyl phosphatidylcholine, DPhPC) is used, and then a single nanopore protein is embedded, and then the current change is recorded under a voltage of 180 mV.

[0134] 1. The embedding nanopore step is as follows: in the buffer (600 mM KCl, 75 mM K3[Fe(CN)6, 25 mM K4[Fe(CN)6]·3H2O, 100 mM Hepes, pH 8.0), the electrical signal measurement value is obtained from the nanopore embedded in the DPhPC phospholipid bilayer. After achieving single-hole insertion into the phospholipid bilayer, 2 mL of buffer (600 mM KCl, 75 mM K3[Fe(CN)6, 25 mM K4[Fe(CN)6]·3H2O, 100 mM Hepes, pH 8.0) is flowed through the system to remove residual excess nanopores to obtain a single nanopore signal acquisition system, and the current signal of the nanopore mutant protein on the phospholipid membrane is recorded.

[0135] 2. The preparation of the DNA sample to be tested and the process of DNA detection ability identification are as follows: after the single nanopore signal acquisition system is constructed, the DNA sample to be tested, ATP (final concentration 2 mM) and MgCl2 (final concentration 10 mM) of the assembled T4 Dda mutant protein are flowed into the single nanopore experimental system (total volume 100 μL), and the signal is measured under a constant voltage of +150 mV.

[0136] The DNA sample to be tested was assembled according to the method recorded in the patent (WO2014135838A1), and the T4 Dda mutant protein (T4 Dda mutant protein is a helicase, rate-limiting protein, specifically T4 Dda-E94C / C109A / C136A / A360C mutant protein, which is recorded in US20170283470A1) was used. The length of the DNA sample to be tested was 0.5 kb (https: / / doi.org / 10.1038 / s41587-020-0570-8; WO 2019002893A1, the sequence is shown as SEQ ID: 12), and it contained 5 repeated sequences, each containing 10 T (10 A in the complementary strand).

[0137] The detection results are shown in Figure 5, and the vertical coordinate is the current, with units of bits nA. The results show that the pore current results can correspond to the molecular sieve results. The current properties of R3 and R9 with better oligomerization state are better than those of other designed proteins, whether it is pore current or pore signal, and the current properties are significantly better than those of NPB wild type (Figure 5, upper WT_NPB, current is larger 0.7 nA, and noise is large), but there is a problem of uneven current size (Figure 5, R3 and R9 have two pore currents: pore current 1 and pore current 2), which is speculated to contain pore proteins with different diameters (oligomerization state); The signals of other proteins (R1, R4, R5, R6, R7, R2 and R8) are smaller than those of wild type NPB, but there are some problems such as large noise, serious spontaneous plugging, small DNA pore signal amplitude or large noise during pore process.

[0138] 4. Cryo-EM structure analysis of R3 and R9

[0139] In order to verify the rationality of protein design and provide a structural basis for subsequent modification based on R3 or R9, this embodiment further attempts to analyze the cryo-EM structure of R3 and R9. According to the method described in Example 2, after a large amount of screening, R3 can collect data, and R9 cannot obtain frozen data due to poor freezing properties.

[0140] The structure analysis results of R3 are shown in Figure 6, where A is the two-dimensional classification result, and B is the three-dimensional reconstruction result. From the two-dimensional classification, it can be seen that R3 has at least two types of pore channels with different diameters (Figure 6, the results in the A frame); from the three-dimensional reconstruction result, only one structure conforms to the original design (11-mer), but the density of the contraction zone is poor; the other one has the same oligomerization state (also 11-mer), but it is more loose and has a larger overall density diameter.

[0141] In summary, different contraction zones designed on the basis of NPB wild type can improve the stability of pore current, and R3 and R9 have greatly improved the detection performance of DNA substrate.

[0142] Example 4: NPB-R3 based sequencing with constriction zone modification and sequencing property identification

[0143] In view of the R3 cryo-EM structure in Example 3 and the unevenness of current detection, in this embodiment, the oligomerization of R3 is attempted to be improved by amino acid mutation and the like. According to the results in Example 3, the density of the constriction zone of R3 is poor, and it is speculated that there is a large degree of steric hindrance between individual molecules of R3. To change the spatial position between molecules, mainly rely on amino acid mutation and amino acid deletion and the like. Considering that the main problem is in the constriction zone, this embodiment mainly attempts to mutate or delete amino acids in this region.

[0144] In this embodiment, amino acid mutation or deletion is mainly concentrated in L74 and D64. For L74, it is mutated to other 19 kinds of amino acids and L74 is deleted (corresponding to P6-7, P8-11, P36-46, Table 3); D64 is mutated to P / A / V / T / S (corresponding to P12-P16, Table 3).

[0145] Table 3: Statistics table based on R3 corresponding mutation or deletion

[0146] 1. Preparation of expression vector

[0147] The method described in Example 1 was used to construct the expression vector, and the R3 expression vector gene was used as the template. The corresponding primers (f represents forward primer, r represents reverse primer) were used to PCR amplify the target fragments (i.e. PCR fragments and PCR vectors). After gel recovery, the PCR fragments and PCR vectors were ligated, and then transformed into DH5a competent cells for positive clone screening. Two clones were picked for sequencing, and the constructed plasmid was stored at -20℃ after sequencing. According to the above method, P6-P16, P23-P25, P36-P46 expression vectors were obtained respectively.

[0148] The P6 expression vector contains a P6-encoding gene that expresses a P6 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 13. The nucleotide sequence of the P6-encoding gene is set forth in SEQ ID NO: 96. The P7 expression vector contains a P7-encoding gene that expresses a P7 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 14. The nucleotide sequence of the P7-encoding gene is set forth in SEQ ID NO: 97. The P8 expression vector contains a P8-encoding gene that expresses a P8 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 15. The nucleotide sequence of the P8-encoding gene is set forth in SEQ ID NO: 98. The P9 expression vector contains a P9-encoding gene that expresses a P9 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 16. The nucleotide sequence of the P9-encoding gene is set forth in SEQ ID NO: 99. The P10 expression vector contains a P10-encoding gene that expresses a P10 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 17. The nucleotide sequence of the P10-encoding gene is set forth in SEQ ID NO: 100. The P11 expression vector contains a P11-encoding gene that expresses a P11 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 18. The nucleotide sequence of the P11-encoding gene is set forth in SEQ ID NO: 101. The P23 expression vector contains a P23-encoding gene that expresses a P23 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 19. The nucleotide sequence of the P23-encoding gene is set forth in SEQ ID NO: 102. The P24 expression vector contains a P24-encoding gene that expresses a P24 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 20. The nucleotide sequence of the P24-encoding gene is set forth in SEQ ID NO: 103. The P25 expression vector contains a P25-encoding gene that expresses a P25 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 21. The nucleotide sequence of the P25-encoding gene is set forth in SEQ ID NO: 104. The P36 expression vector contains a P36-encoding gene that expresses a P36 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 22. The nucleotide sequence of the P36-encoding gene is set forth in SEQ ID NO: 105. The P37 expression vector contains a P37-encoding gene that expresses a P37 protein with a 6xhis tag, the sequence of which is set forth in SEQ ID NO: 23. The nucleotide sequence of the P37-encoding gene is set forth in SEQ ID NO: 106.The P38 expression vector contains a P38 coding gene which expresses a P38 protein with a 6xhis tag, and the P38 protein sequence is shown as SEQ ID NO: 24. The nucleotide sequence of the P38 coding gene is shown as SEQ ID NO: 107. The P39 expression vector contains a P39 coding gene which expresses a P39 protein with a 6xhis tag, and the P39 protein sequence is shown as SEQ ID NO: 25. The nucleotide sequence of the P39 coding gene is shown as SEQ ID NO: 108. The P40 expression vector contains a P40 coding gene which expresses a P40 protein with a 6xhis tag, and the P40 protein sequence is shown as SEQ ID NO: 26. The nucleotide sequence of the P40 coding gene is shown as SEQ ID NO: 109. The P41 expression vector contains a P41 coding gene which expresses a P41 protein with a 6xhis tag, and the P41 protein sequence is shown as SEQ ID NO: 27. The nucleotide sequence of the P41 coding gene is shown as SEQ ID NO: 110. The P42 expression vector contains a P42 coding gene which expresses a P42 protein with a 6xhis tag, and the P42 protein sequence is shown as SEQ ID NO: 28. The nucleotide sequence of the P42 coding gene is shown as SEQ ID NO: 111. The P43 expression vector contains a P43 coding gene which expresses a P43 protein with a 6xhis tag, and the P43 protein sequence is shown as SEQ ID NO: 29. The nucleotide sequence of the P43 coding gene is shown as SEQ ID NO: 112. The P44 expression vector contains a P44 coding gene which expresses a P44 protein with a 6xhis tag, and the P44 protein sequence is shown as SEQ ID NO: 30. The nucleotide sequence of the P44 coding gene is shown as SEQ ID NO: 113. The P45 expression vector contains a P45 coding gene which expresses a P45 protein with a 6xhis tag, and the P45 protein sequence is shown as SEQ ID NO: 31. The nucleotide sequence of the P45 coding gene is shown as SEQ ID NO: 114. The P46 expression vector contains a P46 coding gene which expresses a P46 protein with a 6xhis tag, and the P46 protein sequence is shown as SEQ ID NO: 32. The nucleotide sequence of the P46 coding gene is shown as SEQ ID NO: 115. The P12 expression vector contains a P12 coding gene which expresses a P12 protein with a 6xhis tag, and the P12 protein sequence is shown as SEQ ID NO: 33. The nucleotide sequence of the P12 coding gene is shown as SEQ ID NO: 116. The P13 expression vector contains a P13 coding gene which expresses a P13 protein with a 6xhis tag, and the P13 protein sequence is shown as SEQ ID NO: 34. The nucleotide sequence of the P13 coding gene is shown as SEQ ID NO: 117.The P14 expression vector contains a P14 coding gene which expresses a P14 protein with a 6xhis tag, the P14 protein sequence is shown as SEQ ID NO: 35. The nucleotide sequence of the P14 coding gene is shown as SEQ ID NO: 118. The P15 expression vector contains a P15 coding gene which expresses a P15 protein with a 6xhis tag, the P15 protein sequence is shown as SEQ ID NO: 36. The nucleotide sequence of the P15 coding gene is shown as SEQ ID NO: 119. The P16 expression vector contains a P16 coding gene which expresses a P16 protein with a 6xhis tag, the P16 protein sequence is shown as SEQ ID NO: 37. The nucleotide sequence of the P16 coding gene is shown as SEQ ID NO: 120.

[0149] The PCR system and PCR procedure are the same as in Example 1. The seamless ligation system is the same as in Example 1, except that the linearized vector is replaced by the PCR vector.

[0150] Vector primers:

[0151] V-f: gatatgctggcgaccgcgc (SEQ ID NO: 121); V-r: cgctttcgggccgcgat (SEQ ID NO: 122).

[0152] Fragment primers: Primers used for P6 and P7:

[0153] P6-f: GCACGTccgGCTGCTGCTGTCATGGCA (SEQ ID NO: 123); P6-r: AGCAGCcggACGTGCAGCCAGAGCTG (SEQ ID NO: 124);

[0154] P7-f: GCTGCACGTGCTGCTGCTGTCATGGCA (SEQ ID NO: 125); P7-r: AGCAGCAGCACGTGCAGCCAGAGCTG (SEQ ID NO: 126);

[0155] The following P8-11, P36-46, forward primers are P8-f, P8-f: GCTGCTGCTGTCATGGCA (SEQ ID NO: 127); reverse is corresponding -r primer

[0156] P8-r: CATGACAGCAGCAGCcacACGTGCAGCCAGAGCTG (SEQ ID NO: 128);

[0157] P9-r: CATGACAGCAGCAGCggtACGTGCAGCCAGAGCTG (SEQ ID NO: 129);

[0158] P10-r: CATGACAGCAGCAGCgctACGTGCAGCCAGAGCTG (SEQ ID NO: 130);

[0159] P11-r: CATGACAGCAGCAGCaatACGTGCAGCCAGAGCTG (SEQ ID NO: 131);

[0160] P23-r: CATGACAGCAGCAGCaccACGTGCAGCCAGAGCTG (SEQ ID NO: 132);

[0161] P24-r: CATGACAGCAGCAGCacaACGTGCAGCCAGAGCTG (SEQ ID NO: 133);

[0162] P25-r: CATGACAGCAGCAGCatcACGTGCAGCCAGAGCTG (SEQ ID NO: 134);

[0163] P36-r: CATGACAGCAGCAGCagcACGTGCAGCCAGAGCTG (SEQ ID NO: 135);

[0164] P37-r: CATGACAGCAGCAGCcatACGTGCAGCCAGAGCTG (SEQ ID NO: 136);

[0165] P38-r: CATGACAGCAGCAGCgttACGTGCAGCCAGAGCTG (SEQ ID NO: 137);

[0166] P39-r: CATGACAGCAGCAGCctgACGTGCAGCCAGAGCTG (SEQ ID NO: 138);

[0167] P40-r: CATGACAGCAGCAGCttcACGTGCAGCCAGAGCTG (SEQ ID NO: 139);

[0168] P41-r: CATGACAGCAGCAGCacgACGTGCAGCCAGAGCTG (SEQ ID NO: 140);

[0169] P42-r: CATGACAGCAGCAGCtttACGTGCAGCCAGAGCTG (SEQ ID NO: 141);

[0170] P43-r: CATGACAGCAGCAGCgtgACGTGCAGCCAGAGCTG (SEQ ID NO: 142);

[0171] P44-r: CATGACAGCAGCAGCccaACGTGCAGCCAGAGCTG (SEQ ID NO: 143);

[0172] P45-r: CATGACAGCAGCAGCaaaACGTGCAGCCAGAGCTG (SEQ ID NO: 144);

[0173] P46-r: CATGACAGCAGCAGCataACGTGCAGCCAGAGCTG (SEQ ID NO: 145);

[0174] The following P12-16, reverse primer is P12-r, P12-r: GGAAACATCTGCCAGATCAGACA (SEQ ID NO: 146);

[0175] The forward is corresponding-f primer: P12-f: TGTTTCCcctGCAGAAGCTGCAGCTCTG (SEQ ID NO: 147);

[0176] P13-f: TGTTTCCGCAGCAGAAGCTGCAGCTCTG (SEQ ID NO: 148); P14-f: TGTTTCCgtgGCAGAAGCTGCAGCTCTG (SEQ ID NO: 149); P15-f: TGTTTCCaccGCAGAAGCTGCAGCTCTG (SEQ ID NO: 150); P16-f: TGTTTCCagcGCAGAAGCTGCAGCTCTG (SEQ ID NO: 151).

[0177] 2. Protein preparation: The correct vector was sequenced and the protein was prepared according to the method described in Example 1.

[0178] The protein expression is shown in Figure 7. The proteins are normally expressed, and the oligomerization state changes obviously with or without heating, and the mutant proteins form pores normally.

[0179] 3. Current property detection: The mutant proteins were detected for current property according to the method in Example 3.

[0180] The detection results are shown in Figure 8. The results show that after the L74 is mutated to V / T / S / I (P8-11), the pore current is uniform, only one current signal is detected, the noise and spontaneous blockage of the pore current are obviously improved, the DNA substrate can record a significant signal, and the substrate DNA capture rate is higher. Therefore, P8-P11 has basically been able to be used for sequencing of DNA substrate, wherein the pore current noise of P10 is relatively less. However, the substrate DNA trans-pore signal width of P8-P11 is small, the pore current is between 0.13 nA and 0.21 nA, and the signal width is less than 50 pA.

[0181] The other mutant proteins all have corresponding problems: for example, the three mutant proteins of P23-P25 can normally embed the pore, P23 and P25 show non-uniform pore current, contain more noisy pores, and the noise is frequently superimposed in the sequencing process, which is uncontrollable; the pore current of P24 is uniform as a whole, but the sequencing process is also unstable, and large noise may also occur. For example, P36-P39, the main problem of these mutant proteins is that the opening current and trans-pore signal amplitude are very small, which is similar to the case of P8-P11, but there are more spontaneous blockages of the pores. Overall, the mutation of L74 position shows large noise or frequent spontaneous blockage, except for P8-P11.

[0182] 4. Electron microscope observation: whether the oligomerization state of P10 is uniform is verified by cryo-EM, and the P10 cryo-data is collected by using the method described in Example 2.

[0183] The observation results are shown in Figure 9, wherein A is a two-dimensional classification result, and B is a three-dimensional reconstructed structure. The results show that both the two-dimensional classification and the three-dimensional classification show that P10 is a uniform 11-mer. Therefore, after the design protein is mutated and modified, a more uniform pore protein is obtained, and the sequencing performance is obviously improved.

[0184] Example 5: Identification of sequencing properties of deletion modification of constriction zone based on NPB-R

[0185] The trans-pore signal width represents the signal-to-noise ratio, which directly affects the accuracy of sequencing. Since the trans-pore signal width of the substrate DNA of P8-P11 is small (Figure 8, the pore current is between 0.13 nA and 0.21 nA, and the signal width is less than 50 pA), in this embodiment, the deletion of amino acids in the constriction zone of P10 (L74S) is attempted, the deletion length is 1-5 amino acids, the deletion position is before and after S74 (P10), the type of the deleted amino acid is shown in Table 4, and “-” in the table represents deleting one amino acid. The overall deleted protein corresponds to P54-P64.

[0186] Table 4: Type statistics table of P10 amino acid mutation

[0187] 1. Preparation of expression vector

[0188] The expression vector construction was performed using the method described in Example 1, taking the P10 expression vector gene as a template, using the corresponding primers (f represents forward primer, r represents reverse primer) to respectively PCR amplify the target fragments (i.e. PCR fragments and PCR vectors). After gel recovery, the PCR fragments and the PCR vectors were ligated, and then transferred into DH5a competent cells for positive clone screening. Two clones were picked for sequencing, and after correct sequencing, the constructed plasmid was stored at -20°C for standby. According to the foregoing method, P6-P16, P23-P25, P36-P46 expression vectors were obtained respectively.

[0189] The P54 expression vector contains a P54 coding gene which expresses a P54 protein with a 6xhis tag, and the sequence of the P54 protein is shown as SEQ ID NO: 38. The nucleotide sequence of the P54 coding gene is shown as SEQ ID NO: 164. The P55 expression vector contains a P55 coding gene which expresses a P55 protein with a 6xhis tag, and the sequence of the P55 protein is shown as SEQ ID NO: 39. The nucleotide sequence of the P55 coding gene is shown as SEQ ID NO: 165. The P56 expression vector contains a P56 coding gene which expresses a P56 protein with a 6xhis tag, and the sequence of the P56 protein is shown as SEQ ID NO: 40. The nucleotide sequence of the P55 coding gene is shown as SEQ ID NO: 166. The P57 expression vector contains a P57 coding gene which expresses a P57 protein with a 6xhis tag, and the sequence of the P57 protein is shown as SEQ ID NO: 41. The nucleotide sequence of the P57 coding gene is shown as SEQ ID NO: 167. The P58 expression vector contains a P58 coding gene which expresses a P58 protein with a 6xhis tag, and the sequence of the P58 protein is shown as SEQ ID NO: 42. The nucleotide sequence of the P58 coding gene is shown as SEQ ID NO: 168. The P59 expression vector contains a P59 coding gene which expresses a P59 protein with a 6xhis tag, and the sequence of the P59 protein is shown as SEQ ID NO: 43. The nucleotide sequence of the P59 coding gene is shown as SEQ ID NO: 169. The P60 expression vector contains a P60 coding gene which expresses a P60 protein with a 6xhis tag, and the sequence of the P60 protein is shown as SEQ ID NO: 44. The nucleotide sequence of the P60 coding gene is shown as SEQ ID NO: 170. The P61 expression vector contains a P61 coding gene which expresses a P61 protein with a 6xhis tag, and the sequence of the P61 protein is shown as SEQ ID NO: 45. The nucleotide sequence of the P61 coding gene is shown as SEQ ID NO: 171. The P62 expression vector contains a P62 coding gene which expresses a P62 protein with a 6xhis tag, and the sequence of the P62 protein is shown as SEQ ID NO: 46. The nucleotide sequence of the P62 coding gene is shown as SEQ ID NO: 172. The P63 expression vector contains a P63 coding gene which expresses a P63 protein with a 6xhis tag, and the sequence of the P63 protein is shown as SEQ ID NO: 47. The nucleotide sequence of the P63 coding gene is shown as SEQ ID NO: 173. The P64 expression vector contains a P64 coding gene which expresses a P64 protein with a 6xhis tag, and the sequence of the P64 protein is shown as SEQ ID NO: 48. The nucleotide sequence of the P64 coding gene is shown as SEQ ID NO: 174.

[0190] The PCR system and PCR procedure were the same as in Example 1. The seamless ligation system was the same as in Example 1, except that the linearized vector was replaced with the PCR vector.

[0191] The mutation primers used for the P54-P64 expression vectors were as follows:

[0192] The following P54-58 expression vectors, the forward primer was P54-f, P54-f: GCTGCTGTCATGGCAgatatgc (SEQ ID NO: 175); the reverse was the corresponding -r primer,

[0193] P54-r: TGCCATGACAGCAGCgctACGTGCAGCCAGAGC (SEQ ID NO: 176);

[0194] P55-r: atcTGCCATGACAGCgctACGTGCAGCCAGAGC (SEQ ID NO: 177);

[0195] P56-r: catatcTGCCATGACgctACGTGCAGCCAGAGC (SEQ ID NO: 178);

[0196] P57-r: cagcatatcTGCCATgctACGTGCAGCCAGAGC (SEQ ID NO: 179); P58-r: cgccagcatatcTGCgctACGTGCAGCCAGAGC (SEQ ID NO: 180);

[0197] The following P59-64 expression vectors, the forward primer was P59-f, P59-f: agcGCTGCTGCTGTCAT (SEQ ID NO: 181); the reverse was the corresponding -r primer,

[0198] P59-r: GACAGCAGCAGCgctTGCAGCCAGAGCTGCAGC (SEQ ID NO: 182);

[0199] P60-r: GACAGCAGCAGCgctAGCCAGAGCTGCAGCTTC (SEQ ID NO: 183);

[0200] P61-r: GACAGCAGCAGCgctCAGAGCTGCAGCTTCTGCAT (SEQ ID NO: 184);

[0201] P62-r: GACAGCAGCAGCgctAGCTGCAGCTTCTGCATCG (SEQ ID NO: 185);

[0202] P63-r: GACAGCAGCAGCgctTGCAGCTTCTGCATCGGAAAC (SEQ ID NO: 186);

[0203] P64-r: GACAGCAGCAGCgctAGCTTCTGCATCGGAAACATCTG (SEQ ID NO: 187).

[0204] 2. Protein preparation

[0205] The sequencing correct vectors were subjected to protein preparation according to the method described in Example 1. The protein expression is shown in Figure 10. The results show that, except for P56, the proteins with amino acid deletion can be normally expressed, and the oligomerization state changes obviously with or without heating, indicating that the mutant proteins form pores normally.

[0206] 3. Current property monitoring

[0207] The mutant proteins P10, P54-P55, P57-P64 were subjected to current property detection according to the method in Example 3. The results of pore property detection of P10, P54 and P61 are shown in Figure 11. Compared with P10 (P10 pore current is 0.12 nA, DNA pore signal amplitude is 25 pA, Figure 11, A), the pore current and DNA signal amplitude of P54 and P61 are obviously larger, the pore current of P54 is 0.22 nA, and the DNA pore signal amplitude is 65 pA (Figure 11, B); the pore current of P61 is 0.24 nA, and the DNA pore signal amplitude is 110 pA (Figure 11, C). Therefore, the DNA sequencing performance of P54 and P61 is greatly improved, and P61 is more obvious.

[0208] The detection results of other mutant parts are shown in Figure 12. Among them, P59 is the pore signal graph of P59 protein, P63 is the pore current signal graph of P63 protein, P62 is the pore current and pore signal graph of P62 protein, and P64 is the pore current and pore signal graph of P64 protein. P59 has a small pore signal for complementary DNA, and many signal fluctuations occur during the process of DNA re-entering and leaving the nanopore (Figure 12, P59, pore signal 2); P62 has two pore currents, and only under the larger current, the DNA substrate can pass through the pore, and the signal amplitude is small (Figure 12, P62 upper pore current, pore signal about 50 pA); P63 has a low substrate pore efficiency (Figure 12, P63); P64 has multiple pore currents, and none of them is stable, and can pass the DNA substrate, but the signal amplitude is small (Figure 12, P64).

[0209] Example 6: P61 DNA sequencing property identification

[0210] To further test the sequencing performance of P61, in this example, the sequencing properties of P61 were compared with MspA and ONT product R10.4.1 (MinION). The DNA sequencing properties of P61 and MspA were detected according to the method in Example 3.

[0211] The DNA sequencing properties of P61 and ONT product R10.4.1 (MinION) were detected according to the method in Example 3. The only difference was that the DNA sample to be detected was de Bruijn sequence (the sequence was shown as SEQ ID NO: 49).

[0212] The detection results were shown in FIG. 13, in which A was the sequencing signal diagram of P61 and MspA, B was the sequencing signal diagram of P61 and ONT R10.4.1, and C was the horizontal noise of P61 and ONT R10.4.1. The results showed that P61 had the advantages of high base resolution of sequencing signal, low noise, high signal-to-noise ratio, and good sequencing stability.

[0213] Compared with MspA, a transmembrane pore protein commonly used for DNA sequencing, the constriction zone (reading head) of P61 was more stable and more sensitive to DNA substrate; for the same DNA sequence, P61 showed more current signal (A in FIG. 13, the upper panel was the MspA current result, and the lower panel was the P61 current result), indicating that it had higher base recognition ability.

[0214] P61 and ONT R10.4.1 were detected under the same conditions for de Bruijn sequence (SEQ ID NO: 49), and the analysis results showed that the performance of P61 (signal width, step number, etc.) was comparable to that of ONT R10.4.1 (B in FIG. 13, the upper panel was the ONT R10.4.1 current result, and the lower panel was the P61 current result) and the signal noise was further reduced (C in FIG. 13, the noise level of P61 was lower), indicating that it had higher signal-to-noise ratio. At the same time, the capture efficiency and sequencing stability of P61 could meet the long-time sequencing standard through testing, which could meet the demand of high-throughput sequencing.

[0215] In summary, the overall properties of P61 (such as current stability, substrate capture efficiency, signal amplitude, etc.) could be directly used for DNA sequencing.

[0216] In this embodiment, the MspA nanopore is a mutant (D90N, D91N, D93N, D118R, D134R and D139K, the structural schematic diagram of which is shown in Figure 2 in the literature) introduced in the literature (https: / / doi.org / 10.1016 / j.ymeth.2016.03.026), and the embedded pore mode is the same as that in Example 3; R10.4.1 is a commercially available product of ONT.

[0217] The application has been described in detail above. For those skilled in the art, the application can be implemented in a wider range under equivalent parameters, concentrations and conditions without departing from the purpose and scope of the application and without unnecessary experiments. Although specific examples are given in the application, it should be understood that further improvements can be made to the application. In summary, according to the principle of the application, any changes, uses or improvements of the application are intended to be included, including changes made by conventional techniques known in the art outside the scope disclosed in the application. Some basic features can be applied within the scope of the following attached claims.

Claims

1. A nanopore protein monomer, characterized in that, The nanopore protein monomer is any one of the following: (a1) a polypeptide having an amino acid sequence as set forth in SEQ ID NO: 4; (a2) a polypeptide having an amino acid sequence as set forth in SEQ ID NO: 4 with one or several amino acid substitutions and / or deletions and / or additions and having the same function; (a3) a polypeptide having 80% or more identity with any one of the amino acid sequences defined in (a1)-(a2) and having the same function; (a4) a fusion polypeptide obtained by linking a tag to the end of the polypeptide defined in any one of (a1)-(a3).

2. The nanopore protein monomer of claim 1, wherein, (a2) the polypeptide comprises at least one substitution: leucine at position 74 is substituted with proline, valine, threonine, serine, isoleucine, glycine, cysteine, aspartic acid, alanine, methionine, asparagine, glutamine, glutamic acid, arginine, lysine, histidine, tryptophan, tyrosine or phenylalanine; aspartic acid at position 64 is substituted with proline, alanine, valine, threonine or serine; Preferably, the polypeptide of (a2) further comprises at least one deletion of one or several amino acids: one, two, three, four or five amino acids at positions 74-79 of SEQ ID NO: 4 are deleted; and / or one, two, three, four, five, six or seven amino acids at positions 67-73 of SEQ ID NO: 4 are deleted.

3. A construct, characterized in that, The construct comprises at least two linked nanopore protein monomers, at least one of which is the nanopore protein monomer of claim 1 or 2.

4. A nanopore protein, characterized in that, The nanopore protein comprises at least two linked nanopore protein monomers, at least one of which is the nanopore protein monomer of claim 1 or 2; or The nanopore protein comprises the construct of claim 3.

5. The construct of claim 3 or the nanopore protein of claim 4, wherein, The nanopore protein monomers are identical or different; the linkage is covalent or non-covalent; the nanopore protein comprises 10-12 nanopore protein monomers.

6. The nanopore protein monomer of claim 1 or 2, the construct of claim 3 or 5, or the biological material related to the nanopore protein of claim 4 or 5, characterized in that, The related biological material is any one of the following: a1) a nucleic acid molecule encoding the nanopore protein monomer of claim 1 or 2, the construct of claim 3 or 5 or the nanopore protein of claim 4 or 5; a2) an expression cassette comprising the nucleic acid molecule of a1); a3) a recombinant vector comprising the nucleic acid molecule of a1) or the expression cassette of a2); a4) a recombinant cell comprising the nucleic acid molecule of a1) or the expression cassette of a2) or the recombinant vector of a3).

7. Use of the nanopore protein monomer of claim 1 or 2, the construct of claim 3 or 5, the nanopore protein of claim 4 or 5 or the related biological material of claim 6 in detecting the presence, absence or one or more characteristics of a target analyte or in preparing a product for detecting the presence, absence or one or more characteristics of a target analyte; Preferably, the target analyte is one or more of a nucleotide, a nucleic acid, an amino acid, an oligopeptide, a polypeptide, a protein.

8. A method for determining the presence, absence, or one or more characteristics of a target analyte, comprising, The method comprises: A. the target analyte is contacted with the nanopore protein of claim 4 or 5, such that the target analyte moves relative to the nanopore protein; B. one or more measurements are taken as the target analyte moves relative to the nanopore protein, thereby determining the presence, absence or one or more characteristics of the target analyte; Preferably, the target analyte is one or more of a nucleotide, a nucleic acid, an amino acid, an oligopeptide, a polypeptide, a protein.

9. A kit or device for determining the presence, absence or one or more characteristics of a target analyte, characterised in that, The kit comprises the nanopore protein monomer of claim 1 or 2, the construct of claim 3 or 5, the nanopore protein of claim 4 or 5 or the related biological material of claim 6, and a membrane; The device comprises the nanopore protein of claim 4 or 5, and a membrane; Preferably, the target analyte is one or more of a nucleotide, a nucleic acid, an amino acid, an oligopeptide, a polypeptide, a protein.

10. A method of preparing a nanopore protein according to claim 4 or 5, wherein, comprises transforming a host cell with the recombinant vector of claim 6, and inducing the host cell to express the nanopore protein.

Citation Information

Patent Citations

  • Pore

    CN113195736A

  • Novel nanopore protein mutant and application thereof

    CN117384260A

  • Nanopore mutant with ultrahigh thermal stability and application thereof

    CN117417418A

  • PHT nanopore mutant protein and application thereof

    CN117886907A

  • Mutant of porin monomer, protein pore and application thereof

    WO2023019471A1