Nanopore protein monomer and application thereof
By designing stable nanoporin monomers and constructs, the problems of limited nanoporin selection and unsuitable pore size were solved, resulting in more efficient DNA sequencing performance.
Patent Information
- Application Number
- CN202410608665.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-11-18
AI Technical Summary
The limited selection of existing nanoporous proteins and the need for different pore sizes for different substrates affect the stability and accuracy of sequencing results.
A nanoporous protein based on a naturally occurring non-shrinkable region was designed. By artificially designing shrinkable regions and employing amino acid sequence mutations and deletions, combined with protein design, stable nanoporous protein monomers and constructs were formed for DNA detection.
It improves the stability and uniformity of nanoporous proteins, increases the amplitude of pore current and DNA substrate through-pore signal, meets the requirements of DNA sequencing, and enhances sequencing performance.
Smart Images

Figure CN120965826A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and in particular to a nanoporous protein monomer and its applications. Background Technology
[0002] Biological macromolecules such as DNA, RNA, proteins, and polysaccharides are the basic building blocks of living organisms, and their sequence information and post-functional modifications determine their biological functions. Sequence identification technology for biological macromolecules is a core tool for understanding the laws governing life, and single-molecule sequencing technology has emerged as a result.
[0003] Currently, single-molecule sequencing technologies can be mainly divided into two categories: one is optical zero-mode waveguide sequencing, represented by Pacific Biosciences (PacBio) in the United States; the other is electrical nanopore sequencing, represented by Oxford Nanopore Technologies (ONT) in the United Kingdom. Nanopore sequencing technology is a novel nucleic acid sequencing technology developed in recent years. Based on pore type, it can be divided into solid-state pores and biological nanopores. Biological nanopores are pore proteins that allow substrates to pass through. The following nanopore sequencing refers specifically to biological nanopore sequencing technology.
[0004] Under the influence of an electric field, charged nucleic acid substrates can pass through biological nanopores. When nucleic acids pass through the nanopore, they impede the current flowing through it, generating different current signals. By analyzing these signals, the base information of the nucleic acid can be obtained. Compared to other sequencing methods, it offers advantages such as low equipment cost, simple sample preparation, and fast sequencing speed, and has begun to be applied in various fields. Specific advantages include: easy library construction without amplification; fast signal readout speed, typically reaching 200-300 bp / s; long readout length, typically reaching thousands of bases; direct detection of modifications on DNA; and direct RNA sequencing. Due to the characteristics of nanopore sequencing, RNA no longer needs to be reverse transcribed into DNA for sequence analysis, thus preserving modification information on the RNA. Because of these advantages, nanopore sequencing technology has gained widespread attention in recent years.
[0005] Nanoporous proteins are the core of nanopore sequencing. To date, relatively few nanoporous proteins are suitable for nucleic acid detection; only a handful of natural proteins, such as MspA (Mycobacterium smegmatis porin A), CsgG, and CsgG-CsgF (curli-specific transport channels), meet the requirements. The stability, pore diameter distribution, and internal amino acid charge properties of natural nanoporous proteins significantly impact sequencing results. Summary of the Invention
[0006] To address the limitations of limited nanoporous protein selection and the need for nanopores with varying pore sizes for different substrates, this patent proposes a method for obtaining stable nanoporous proteins suitable for DNA detection by artificially designing shrinkage regions based on naturally occurring nanopores without shrinkage regions. This provides a foundation for nucleic acid sequencing and the detection of proteins or polysaccharides.
[0007] This invention provides a nanoporin monomer, wherein the nanoporin monomer is any of the following polypeptides: (a1) a polypeptide with an amino acid sequence as shown in SEQ ID NO: 4; (a2) a polypeptide having the same function by substituting and / or deleting and / or adding one or more amino acids of the amino acid sequence shown in SEQ ID NO: 4; (a3) a polypeptide having more than 80% identity with any of the amino acid sequences defined in (a1)-(a2) and having the same function; (a4) a fusion polypeptide obtained by attaching a tag to the end of any of the polypeptides defined in (a1)-(a3).
[0008] The tag refers to a polypeptide or protein that is fused with a target protein using in vitro DNA recombination technology for expression, to facilitate the expression, detection, tracing, and / or purification of the target protein. The protein tag may be a Flag tag, His tag, MBP tag, HA tag, myc tag, GST tag, and / or SUMO tag, etc.
[0009] The aforementioned nanoporous protein monomers may specifically be the R1-R10 protein monomers prepared in the following examples. For example, the amino acid sequences of the aforementioned nanoporous protein monomers are shown in any one of SEQ ID NO: 2-11.
[0010] Optionally, according to the nanoporous protein monomer described above, the polypeptide described in (a2) comprises at least one of the following substitutions: the leucine at position 74 is replaced by proline, valine, threonine, serine, isoleucine, glycine, cysteine, aspartic acid, alanine, methionine, asparagine, glutamine, glutamic acid, arginine, lysine, histidine, tryptophan, tyrosine, or phenylalanine.
[0011] The aforementioned nanoporous protein monomers may specifically be the P6, P8-P16, P23-P25, and P36-P46 protein monomers prepared in the following examples. For example, the amino acid sequences of the aforementioned nanoporous protein monomers are shown in any of SEQ ID NO: 13, 15-37.
[0012] Optionally, according to the nanoporin monomer described above, the polypeptide described in (a2) further comprises at least one or more amino acid deletions as follows: one, two, three, four or five amino acid deletions at positions 74-79 of SEQ ID NO: 4; and / or, one, two, three, four, five, six or seven amino acid deletions at positions 67-73 of SEQ ID NO: 4.
[0013] The deletion of one or more amino acids may be the deletion of amino acid position 74 of SEQ ID NO: 4; the deletion of amino acid position 75 of SEQ ID NO: 4; the deletion of amino acid positions 75-76 of SEQ ID NO: 4; the deletion of amino acid positions 75-77 of SEQ ID NO: 4; the deletion of amino acid positions 75-78 of SEQ ID NO: 4; the deletion of amino acid positions 75-79 of SEQ ID NO: 4; the deletion of amino acid position 73 of SEQ ID NO: 4; the deletion of amino acid positions 72-73 of SEQ ID NO: 4; the deletion of amino acid positions 71-73 of SEQ ID NO: 4; the deletion of amino acid positions 70-73 of SEQ ID NO: 4; the deletion of amino acid positions 69-73 of SEQ ID NO: 4; the deletion of amino acid positions 68-73 of SEQ ID NO: 4; or the deletion of amino acid positions 67-73 of SEQ ID NO: 4.
[0014] The aforementioned nanoporin monomers may specifically be the P7, P54-P64 protein monomers prepared in the following examples. For example, the amino acid sequences of the aforementioned nanoporin monomers are shown in any one of SEQ ID NO: 14, 38-48.
[0015] The present invention also provides a construct comprising at least two linked nanoporin monomers, at least one of which is the aforementioned nanoporin monomer. The construct retains the ability to form pores.
[0016] The present invention also provides a nanoporin comprising at least two linked nanoporin monomers, at least one of which is the nanoporin monomer described above; or the nanoporin comprises the construct described above.
[0017] Optionally, depending on the above-described construct or the above-described nanoporin, the nanoporin monomers may be the same or different; the connection may be covalent or non-covalent; the nanoporin may contain 10-12 nanoporin monomers, for example, 10, 11 or 12 nanoporin monomers.
[0018] The aforementioned nanoporous proteins can specifically be the R1-R10, P6-P16, P23-P25, and P36-P46 proteins prepared in the following examples.
[0019] The present invention also provides the above-mentioned nanoporin monomer, the above-mentioned construct, or related biomaterials of the above-mentioned nanoporin, wherein the related biomaterial is any one of the following: a1) a nucleic acid molecule encoding the above-mentioned nanoporin monomer, the above-mentioned construct, or the above-mentioned nanoporin; a2) an expression cassette containing the nucleic acid molecule of a1); a3) a recombinant vector containing the nucleic acid molecule of a1), or a recombinant vector containing the expression cassette of a2); a4) a recombinant cell containing the nucleic acid molecule of a1), or a recombinant cell containing the expression cassette of a2), or a recombinant cell containing the recombinant vector of a3).
[0020] In the above-mentioned biological materials, the nucleic acid molecule can be DNA, such as cDNA, genomic DNA or recombinant DNA; the nucleic acid molecule can also be RNA, such as mRNA, siRNA, shRNA, sgRNA, miRNA or antisense RNA.
[0021] In the aforementioned biological materials, the expression cassette refers to DNA capable of expressing genes in host cells. This DNA may include not only promoters that initiate gene transcription but also terminators that terminate gene transcription. Furthermore, the expression cassette may also include enhancer sequences.
[0022] Optionally, the nucleic acid molecule described in a1) is a DNA molecule that is any of the following: 1) a DNA molecule whose coding sequence is any of the R1-R10, P6-P16, P23-P25, or P36-P46 coding genes; 2) a DNA molecule whose nucleotide sequence is any of the R1-R10, P6-P16, P23-P25, or P36-P46 coding genes; 3) a DNA molecule that hybridizes to the nucleotide sequence defined in 1) or 2) under stringent conditions and encodes the above-mentioned nanoporin monomer, the above-mentioned construct, or the above-mentioned nanoporin.
[0023] Optionally, the recombinant vector described in a3) is a vector containing a DNA molecule having any of the encoding genes R1-R10, P6-P16, P23-P25, or P36-P46, such as the R1-R10, P6-P16, P23-P25, or P36-P46 expression vectors prepared in the following examples.
[0024] The stringent conditions can be hybridization and washing of the membrane at 65°C in a solution of 0.1×SSPE (or 0.1×SSC) and 0.1% SDS.
[0025] The application of the above-mentioned nanoporin monomers, the above-mentioned constructs, the above-mentioned nanoporins, or the above-mentioned related biomaterials in detecting the presence, absence, or one or more features of the target analyte, or in preparing products that detect the presence, absence, or one or more features of the target analyte, is also within the scope of protection of this invention.
[0026] The present invention also provides a method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising: A. contacting the target analyte with the aforementioned nanoporin, causing the target analyte to move relative to the nanoporin; B. acquiring one or more measurements while the target analyte is moving relative to the nanoporin, thereby determining the presence, absence, or one or more characteristics of the target analyte.
[0027] Optionally, the method further includes the step of applying a potential difference in contact between the target analyte and the nanoporous protein.
[0028] Optionally, the measured value is obtained through electrical and / or optical measurements. For example, the electrical measurements include, but are not limited to, current measurements, impedance measurements, tunnel measurements, wind tunnel measurements, or field-effect transistor (FET) measurements, etc. In one specific embodiment of the invention, the measured value is a current measurement.
[0029] The present invention also provides a kit for determining the presence, absence, or one or more characteristics of a target analyte, the kit comprising the above-described nanoporin monomer, the above-described construct, the above-described nanoporin or related biomaterial, and a membrane.
[0030] The present invention also provides an apparatus for determining the presence, absence, or one or more characteristics of a target analyte, the apparatus comprising the above-described nanoporous protein and a membrane.
[0031] Preferably, the target analyte is one or more of nucleotides, nucleic acids, amino acids, oligopeptides, polypeptides, and proteins.
[0032] In the above-mentioned kits or devices, the membrane and nanoporin can be packaged independently, or the nanoporin can be embedded in the membrane.
[0033] The membrane can be any membrane existing in the prior art, preferably an amphoteric molecular layer, that is, a layer formed by amphoteric molecules such as phospholipids having at least one hydrophilic portion and at least one lipophilic or hydrophobic portion, the amphoteric molecules being synthetic or naturally occurring. For example, the membrane is a phospholipid monolayer membrane.
[0034] The kits or devices described above may also include rate-controlling proteins. Rate-controlling proteins may include one or more combinations of nucleic acid-binding proteins, helicases, exonucleases, telomerases, topoisomerases, transcriptases, transloses, and / or polymerases.
[0035] Optionally, the helicase is selected from Hel308 family helicases and modified Hel308 family helicases, RecD helicase and its variants, TrwC helicase and its variants, Dda helicase and its variants, TraI Eco and its variants, XPD Mbu and its variants, Pif1-like helicase and its variants.
[0036] Optionally, the target analyte is one or more of nucleotides, nucleic acids, amino acids, oligopeptides, polypeptides, and proteins.
[0037] Optionally, the one or more features are selected from at least one of (i) the length of the target analyte; (ii) the identity of the target analyte; (iii) the sequence of the target analyte; (iv) the secondary structure of the target analyte; and (v) whether the target analyte is modified. "Identity" refers to the similarity between sequences. Identity can be evaluated visually or by computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences.
[0038] Optionally, the nucleic acid can be naturally occurring or artificially synthesized. Specifically, the nucleic acid can be natural DNA, RNA, or modified DNA or RNA, or it can be artificially synthesized nucleic acid, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threonine nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers with nucleoside side chains.
[0039] Optionally, the nucleic acid is single-stranded, double-stranded, or at least partially double-stranded.
[0040] Optionally, the nucleic acid can be of any length. For example, the length of the nucleic acid can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400 or at least 500 nucleotides or nucleotide pairs, or it can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs, or 100000 or more nucleotides or nucleotide pairs.
[0041] Optionally, one or more nucleotides in the nucleic acid may be modified, such as methylated, oxidized, damaged, debased, protein-labeled, tagged, or linked to a spacer in the middle of a polynucleotide sequence.
[0042] The aforementioned nanoporous protein monomers, constructs, or proteins can be synthesized artificially, or their encoding genes can be synthesized first and then expressed biologically.
[0043] The present invention also provides a method for preparing the above-mentioned nanoporous protein, comprising transforming host cells with the above-mentioned recombinant vector and inducing the host cells to express the nanoporous protein.
[0044] When the term "comprising" or "including" is used in this application to describe a protein or nucleic acid sequence, the protein or nucleic acid may be composed of the sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still have the activity described in this invention.
[0045] In this article, standard one-letter codes are used for amino acids. These are as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), gamma-glutamyl (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Standard substitution notation is also used, i.e., L74V means that the L at position 74 of the sequence is replaced by V.
[0046] In this invention, the high-resolution three-dimensional structure of NPB nanopores was determined using cryo-electron microscopy. Structural analysis revealed that the width of the NPB nanoporous protein contraction region was... Unlike other nanoporins, NPB nanoporins have almost no contraction regions, exhibiting large pore currents and significant noise on phospholipid membranes, making them unsuitable for nucleic acid substrate detection. However, wild-type NPB nanoporins demonstrate strong stability, and their contraction regions do not affect oligomerization, making them excellent pore frameworks.
[0047] Protein design is becoming increasingly sophisticated and its applications are becoming more widespread. This invention uses protein design to address the problem of the lack of a contraction region in NPB (non-blocking bioreactors), while also exploring a more stable contraction region type (α-helix). This approach can effectively increase the variety of nanopores that can be used for sequencing.
[0048] Based on NPB wild-type nanopores, this invention designs 10 different contraction regions (R1-R10). Through protein expression, purification, structural analysis, and current detection, it is demonstrated that the designed contraction regions can stabilize the overall pore current and improve DNA sequencing performance. The stability and uniformity of the designed nanopores are improved through amino acid mutation. The amplitude of pore current and DNA substrate through-pore signal is increased through amino acid deletion, so as to meet the requirements of DNA sequencing. Attached Figure Description
[0049] Figure 1 The results for NPB wild-type protein purification and molecular sieving are shown. "+" indicates protein samples heated at 100℃ for 10 minutes; "-" indicates protein samples left at room temperature for 10 minutes without heating.
[0050] Figure 2 The results of the NPB wild-type protein structure analysis.
[0051] Figure 3 This is a design based on different contraction region types of NPB wild-type.
[0052] Figure 4 The results are for protein purification (R1-R10) and some molecular sieves. "+" indicates protein samples heated at 100℃ for 10 minutes; "-" indicates protein samples left at room temperature for 10 minutes without heating.
[0053] Figure 5 The results show the pore currents and DNA perforation results for NPB wild-type and R1-R10 portions.
[0054] Figure 6 This is the analysis result of the R3 structure.
[0055] Figure 7 The results show the expression of modified proteins based on R3. "+" indicates protein samples heated at 100℃ for 10 minutes; "-" indicates protein samples left at room temperature for 10 minutes without heating.
[0056] Figure 8 The results show the current in the P8-P10 wells and the DNA perforation results.
[0057] Figure 9 This is the result of the structure analysis for P10.
[0058] Figure 10 The results are based on the expression of protein deleted from P10.
[0059] Figure 11 The results show the comparison of the via properties of P10, P54, and P61.
[0060] Figure 12 The results are for the current and via signals of holes P59, P62, P63, and P64.
[0061] Figure 13 Results of the identification of the permeability properties of P61 DNA substrate. Detailed Implementation
[0062] The present invention will be further described in detail below with reference to specific embodiments. The embodiments given are only for illustrating the present invention and are not intended to limit the scope of the present invention. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the present invention in any way. Unless otherwise specified, the experimental methods in the following embodiments are conventional methods, performed according to the techniques or conditions described in the literature in the art or according to the product instructions. Unless otherwise specified, the materials, reagents, etc. used in the following embodiments are all commercially available. The quantitative experiments in the following embodiments are all repeated three times, and the results are averaged.
[0063] Example 1: Preparation of wild-type NPB nanopore expression vector and protein
[0064] 1. Construction of wild-type NPB nanoporous protein carrier
[0065] Wild-type NPB nanoporous protein is derived from *Thermodesulfobacteriota bacterium* (ACCESSION: NPB10001, SEQ ID NO: 1). The expression gene for this protein was obtained through artificial synthesis, with codon optimization performed during the synthesis process for expression in *E. coli*. After synthesis, the gene was seamlessly cloned into the pBAD22 vector to obtain the wild-type NPB nanoporous protein vector. A 6×his tag was added to the C-terminus of the protein as an affinity purification tag. The wild-type NPB nanoporous protein vector contains the wild-type NPB nanoporous protein expression gene, which expresses wild-type NPB nanoporous protein with a 6×his tag. The wild-type NPB nanoporous protein sequence is shown in SEQ ID NO: 1.
[0066] The vector construction steps are as follows: Using the wild-type NPB nanoporous protein expression gene (SEQ ID NO: 50) as a template, the target gene fragment was amplified by PCR using forward and reverse primers (primer F: CAGGAGGAATTAACCATGTTTCGCCTGCTGACCC; primer R: GAACTGCGGGTGGCTCCATTTGGTGCCGCTCGCCG). After gel recovery, the fragment was ligated into the linearized pBAD22 vector and then transformed into DH5α competent cells for screening of positive clones. Two clones were selected for sequencing. After successful sequencing, the constructed plasmid was stored at -20℃ for later use.
[0067] The PCR system is as follows (20 μL):
[0068]
[0069] The PCR procedure is as follows:
[0070]
[0071] The seamless connection system (10μL) is as follows:
[0072] 2× Seamless Linkage Buffer 5μL
[0073] Target fragment (50 ng / μL) 3 μL
[0074] Linearization vector (10 ng / μL) 2 μL
[0075] After reacting at 50℃ for 15 min, 2 μL was transferred into DH5α cells and sequenced.
[0076] 2. Preparation of wild-type NPB nanoporous protein
[0077] Wild-type NPB nanoporous protein was purified to high purity using Ni column affinity chromatography and molecular sieves. Figure 1 (where A represents the SDS-PAGE result), the protein properties are uniform ( Figure 1 (where B represents the molecular sieve results). Furthermore, SDS-PAGE analysis shows that heating produces monomeric proteins (molecular weight approximately 26 kDa); without heating, the protein is primarily in an oligomeric state (i.e., a porous state, with a molecular weight greater than the marker's largest band of 180 kDa), and there are also instances of missing monomers, indicating that wild-type NPB protein has poor stability.
[0078] The expression and purification steps are as follows: 1) After confirming the correct sequencing of the wild-type NPB nanoporous protein vector, it was transformed into BL21(DE3) for expression. 1 mL of seed culture was obtained at 37℃ and 200 rpm, then transferred to 1 L of LB medium and incubated at 37℃ and 200 rpm until OD500 was reached. 600 1) Set the concentration to 1.2, cool to 26℃, add arabinose to a final concentration of 0.4 g / L and induce overnight; 2) Collect bacterial cells at 4000 rpm, resuspend each 1 L of cells in 20 mL of lysis buffer, autoclave and centrifuge at 18000 rpm at 4℃ for 1 hour, collect the precipitated membrane fraction; 3) Resuspend each 1 L of cells in 10 mL of membrane lysis buffer, magnetically stir at 4℃ for 1 hour to extract membrane proteins, centrifuge at 18000 rpm at 4℃ for 1 hour, collect the supernatant membrane protein fraction; 4) Add imidazole to the supernatant to a final concentration of 30 mM and incubate with Ni beads after equilibration with membrane lysis buffer, bind at 4℃ for 1 hour and then perform affinity purification; 5) Transfer the supernatant and Ni beads into a column and allow gravity flow-through; wash with 10 mL of washing buffer and elute with 5 mL of elution buffer; 6) Purify the target protein by molecular sieve purification and detect the purity by SDS-PAGE, the results are as follows. Figure 1 As shown in Figure A, the molecular sieve results are as follows: Figure 1 As shown in B.
[0079] Lysis buffer: 20mM Tris-HCl, 150mM NaCl, pH 8.0.
[0080] Dissolution buffer: 20 mM Tris-HCl, 150 mM NaCl, pH 8.0, 1% LDAO (dodecyl dimethylamine oxide).
[0081] Washing solution: 20mM Tris-HCl, 150mM NaCl, pH 8.0, 0.5% LDAO, 50mM imidazole.
[0082] Eluent: 20 mM Tris-HCl, 150 mM NaCl, pH 8.0, 0.1% LDAO, 200 mM imidazole.
[0083] Molecular sieve buffer: 20mM Tris-HCl, 150mM NaCl, pH 8.0, 0.06% LDAO.
[0084] Example 2: Atomic-level structure determination of wild-type NPB nanoporous protein
[0085] After obtaining wild-type NPB nanoporous protein with high purity and homogeneity (Example 1), its resolution was resolved using cryo-electron microscopy. The atomic-level structure.
[0086] The steps for cryo-electron microscopy structure analysis are as follows:
[0087] 1) Prepare the molecularly purified sample and several backup samples under the same frozen sample conditions. Select 8 suitable samples for loading into the Talos Arctica 200kV high-end electron microscope; 2) After loading the samples, wait for the vacuum and temperature to stabilize, then open the microscope tube and select suitable square wells at low magnification. This step is similar to the screening of frozen samples, and its purpose is to select suitable square wells for subsequent data collection; 3) Based on the available time of the electron microscope and the number of images that can be acquired from each square well, calculate the approximate number of square wells to be selected, and then take a map; 4) During the map taking process, wells that can be used for data collection can be selected offline, which can save some time; 5) After the map taking is completed, the electron microscope can be adjusted in preparation for data collection. This mainly includes the electron microscope's alignment, background subtraction, and basic data collection parameters (underfocus -1.5μm to -2.5μm, electron dose). 32 frames per second and pixel size ) settings.
[0088] Subsequent data processing and structural analysis yielded its density map, which is shown below. Figure 2As shown, the overall map is an 11-mer structure. Through homology modeling, the atomic coordinates of the amino acids in the wild-type NPB nanoporous protein were obtained through structure building and refinement. Figure 2 Figure A shows the top and side views of the map, and Figure B shows the resolution information.
[0089] NPB wild-type contraction zone has fewer amino acids (3, N58-C59-Q60) and a larger pore diameter. ( Figure 2 Figure C shows the model construction results of NPB wild-type (top and side views), demonstrating strong overall stability. Considering the current properties of pore proteins, which are largely determined by the contraction region, NPB wild-type can be used as the main framework for novel pore design, allowing for the design of contraction region types with different properties.
[0090] Example 3: Protein design and corresponding protein preparation based on NPB-based contraction regions (NPB-R)
[0091] 1. Protein design
[0092] The protein design procedure involved in this invention includes RFdiffusion for generating protein backbone structures; ProteinMPNN for generating one or more sequences from the backbone; and AlphaFold2 for predicting structures based on the sequences. The predicted structures are analyzed to determine if they meet the requirements, and this process is repeated until a reasonably structured amino acid sequence is generated. Finally, the feasibility of the designed sequences is verified through vector construction, protein expression, structure resolution, and pore current detection.
[0093] Analysis of the structure of nanoporous protein (CsgG), currently capable of accurate nucleic acid sequencing, reveals that its contraction region is a highly flexible loop structure. This structure lacks rigidity and exhibits significant oscillation under electric fields and the influence of DNA substrates, generating noise and affecting sequencing accuracy. Based on the results of Example 2, a novel contraction region was designed using the NPB wild-type structure as the main framework to mitigate the adverse effects of the flexible contraction region. The designed contraction region is predominantly α-helical, and can be categorized into main forms such as loop-α-helix and α-helix-loop-α-helix, exhibiting diversity in conformation and pore diameter. Figure 3This section presents designs based on different contraction zone types of NPB wild-type nanoporous proteins. A shows top views of three nanoporous protein models: NPB-WT is the wild-type NPB model, and NPB-R3 and NPB-R4 are models with different contraction zone designs using the wild-type NPB structure as the main framework. B shows the structural characteristics of the monomer contraction zone: NPB-WT is the wild-type NPB monomer, and NPB-R1-R10 are R1-R10 nanoporous protein monomers with different contraction zone designs using the wild-type NPB structure as the main framework. Table 1 shows the theoretical diameters of R1-R10 nanoporous proteins with different contraction zone designs using the wild-type NPB structure as the main framework. Table 2 shows the pore amino acid sequences of R1-R10 nanoporous proteins with different contraction zone designs using the wild-type NPB structure as the main framework.
[0094] Table 1: Theoretical Diameter of Design Channels for R1-R10
[0095]
[0096] Table 2: Amino acid sequences of the designed pores from R1 to R10
[0097]
[0098]
[0099] 2. Carrier and protein preparation
[0100] After AlphaFold2 predicted the structural rationality, the expression vector was constructed using the method described in Example 1. Using the wild-type NPB nanoporous protein vector gene as a template, the target fragment (i.e., the PCR fragment and the PCR vector) was amplified by PCR using corresponding primers (f represents the forward primer, r represents the reverse primer). After gel recovery, the PCR fragment and the PCR vector were ligated, and then transformed into DH5α competent cells for screening of positive clones. Two clones were selected for sequencing. After successful sequencing, the constructed plasmid was stored at -20℃ for later use. The R1-R10 expression vectors were obtained according to the aforementioned method. In the following protein-coding gene nucleotide sequences, lowercase letters indicate substitutions compared to the wild-type NPB nanoporous protein expression gene, "~" indicates deletions compared to the wild-type NPB nanoporous protein expression gene, and "+" represents a connector.
[0101] Where seq1 represents
[0102] ATGTTTCGCCTGCTGACCCTGCTGGCCGGCGTTCTGTTACTGGCGGTGAGCTGCGTGAGCAGCGGCGTGCAGACCCAAGTGGATACCACCGGCCCGACCGCGAGCCAAGTGCTGACCTATCGCGGCCCGAAAGCG;
[0103] seq2 represents
[0104] GATATGCTGGCGACCGCGCTGTTTCGCACCGGCCGCTTTATTGTGCTGGAACGCGGCGAAGGCCTGAAAGAAATTCAGAAAGAACTGGATCTGGCGCAGAGCGGCTATGTGCGCAAAGATCAAGCGCCGCAGACCGGTCAGATGGAAGGCGCGGATATTCTGGTGATTGGCGCGATTACCGCGTTTGAACCGAACGCGAGCGGCGTGGAAGGCGGTGGCGTGGTGATTCCGTATAAAATTCCGATTATTGGCGGCGCGCGCCTGAAAAAGAAAGAAGCGTATATTGCGGCGGATATTCGCCTGGTGGATGTGCGCACCGGTCGCATCATCAACGCGACCACCGTGGAAGGCAAAGCGAGCAGCTGGAAAGCGCAAGGCGTGCTGGGCGGCGTTATTGGCGACGTTGCCTTAGGCGGTGGCCTGGGCGTGTATCGCAACACCCCGATGGAAAAAGCGATTCGTCAGATGCTGTATGCGGCGGTGAACGCGATTGTGCAGATGGTGCCGCCGAACTATTATCGCTGGGGCGAAGATAACGTGCCGGCGAGCGGCACCAAA。
[0105] The R1 expression vector contains the R1 coding gene, which expresses the R1 protein with a 6×his tag. The amino acid sequence of the R1 protein is shown in SEQ ID NO: 2. The nucleotide sequence of the R1 coding gene is as follows:
[0106] 5’-seq1+cgtattgttgttgtactgttcactgctgtatacactggtagccaggatgcgcgtgctacggaactggcagcactggctctggcgatcgcggtttac+seq2-3’.
[0107] The R2 expression vector contains the R2 encoding gene and expresses the R2 protein with a 6×hi tag. The R2 protein sequence is shown in SEQ ID NO: 3. The nucleotide sequence of the R2 encoding gene is as follows:
[0108] 5'-seq1+cgtgttgttgttgttctgttcaccgcagtgtatactggttctcaggacgcacgtgcaaccgaactggctgcaaaagctctggctatcgcagtttac+seq2-3'.
[0109] The R3 expression vector contains the R3 encoding gene and expresses the R3 protein with a 6×hi tag. The R3 protein sequence is shown in SEQ ID NO: 4. The nucleotide sequence of the R3 encoding gene is as follows:
[0110] 5'-seq1+cgtgtagttgttctgccgctgcgcatggttatgtctgatctggcagatgtttccgatgcagaagctgcagctctggctgcacgtctggctgctgctgtcatggca+seq2-3'.
[0111] The R4 expression vector contains the R4 encoding gene and expresses the R4 protein with a 6×hi tag. The R4 protein sequence is shown in SEQ ID NO: 5. The nucleotide sequence of the R4 encoding gene is as follows:
[0112] 5'-seq1+ttcgttgtagtcgctaacgatctgtttgcgaagatcgttcgtctgctggcggaaatcctgatgaacatgaacgcggttaccgaactgcgccacaccatggttgtactggcggcactgctggcagcaatg+seq2-3'.
[0113] The R5 expression vector contains the R5 encoding gene and expresses the R5 protein with a 6×hi tag. The R5 protein sequence is shown in SEQ ID NO: 6. The nucleotide sequence of the R5 encoding gene is as follows:
[0114] 5'-seq1+tacgttgttgttgtcaacgacacctggtataaaatcgttgaaatgatgatcaaactgctgcaggccagcaacgcggtcgagaccctggaaaaacaggctgcacgtctggctgccctgctgtccctgatg+seq2-3'.
[0115] The R6 expression vector contains the R6 encoding gene and expresses the R6 protein with a 6×hi tag. The R6 protein sequence is shown in SEQ ID NO: 7. The nucleotide sequence of the R6 encoding gene is as follows:
[0116] 5'-seq1+tacgtaatcgtaggtggtctgaaccgtgaagaagctctgcgtgctgctgcactggctgcatatgttgtgcgtgatcgtgacgctgataaagctttcatgaaaatcctgacgtggctggcgctgatg+seq2-3'.
[0117] The R7 expression vector contains the R7 encoding gene and expresses the R7 protein with a 6×hi tag. The R7 protein sequence is shown in SEQ ID NO: 8. The nucleotide sequence of the R7 encoding gene is as follows:
[0118] 5'-seq1+tacgttgtagttatcgccaacctgaaactgctggttaccgctttcgctctggcggctcaggtaatgaactccgataagctggaacgcctggcggccgtactggcggcgatgctggcgctgatg+seq2-3'.
[0119] The R8 expression vector contains the R8 encoding gene and expresses the R8 protein with a 6×hi tag. The R8 protein sequence is shown in SEQ ID NO: 9. The nucleotide sequence of the R8 encoding gene is as follows:
[0120] 5'-seq1+tacgtggttctgctgtccaaccaggaggtaatgcgtgtttacatggaagccctgaaactggcgatgaaagcggcggctaagaacgcgaacaacggtctgtacaccgcgttcgccctggtggcgtacatg+seq2-3'.
[0121] The R9 expression vector contains the R9 encoding gene and expresses the R9 protein with a 6×hi tag. The R9 protein sequence is shown in SEQ ID NO: 10. The nucleotide sequence of the R9 encoding gene is as follows:
[0122] 5'-seq1+tacgttgtggtggtgattaacaaagaactgatcaaaacccagctgcgcaaagacaaccgtgctgacgccgcgctggta+seq2-3'.
[0123] The R10 expression vector contains the R10 encoding gene and expresses the R10 protein with a 6×hi tag. The R10 protein sequence is shown in SEQ ID NO: 11. The nucleotide sequence of the R10 encoding gene is as follows:
[0124] 5'-seq1+cgtgttgttgttgttaagatgggtgcggtttataccggcgatcgtgacgcagatatggcggaactgctgctgcgtgctctggctatcgcagcattc+seq2-3'.
[0125] The PCR system and PCR procedure were the same as in Example 1. The seamless ligation system was the same as in Example 1, except that the linearized vector was replaced with the PCR vector. Vector primers: Vf: gatatgctggcgaccgcgc; Vr: cgctttcgggccgcgat.
[0126] Fragment primers:
[0127] R1-f:
[0128] CCAGGATGCGCGTGCTACGGAACTGGCAGCACTGGCTCTGGCGATCGCGGTTTACgatatgctggcgaccgcgc;
[0129] R1-r:
[0130] GCACGCGCATCCTGGCTACCAGTGTATACAGCAGTGAACAGTACAACAACAATACGcgctttcgggccgcgat;
[0131] R2-f:
[0132] TCAGGACGCACGTGCAACCGAACTGGCTGCAAAAGCTCTGGCTATCGCAGTTTACgatatgctggcgaccgcgc;
[0133] R2-r:
[0134] GCACGTGCGTCCTGAGAACCAGTATACACTGCGGTGAACAGAACAACAACAACACGcgctttcgggccgcgat;
[0135] R3-f:
[0136] ATGTTTCCGATGCAGAAGCTGCAGCTCTGGCTGCACGTCTGGCTGCTGCTGTCATGGCAgatatgctggcgaccgcgc;
[0137] R3-r:
[0138] CTGCATCGGAAACATCTGCCAGATCAGACATAACCATGCGCAGCGGCAGAACAACTACACGcgctttcgggccgcgat;
[0139] R4-f:
[0140] CTGATGAACATGAACGCGGTTACCGAACTGCGCCACACCATGGTTGTACTGGCGGCACTGCTGGCAGCAATGgatatgctggcgaccgcgc;
[0141] R4-r:
[0142] GTTCATGTTCATCAGGATTTCCGCCAGCAGACGAACGATCTTCGCAAACAGATCGTTAGCGACTACAACGAAcgctttcgggccgcgat;
[0143] R5-f:
[0144] CTGCAGGCCAGCAACGCGGTCGAGACCCTGGAAAAACAGGCTGCACGTCTGGCTGCCCTGCTGTCCCTGATGgatatgctggcgaccgcgc;
[0145] R5-r:
[0146] GTTGCTGGCCTGCAGCAGTTTGATCATCATTTCAACGATTTTATACCAGGTGTCGTTGACAACAACAACGTAcgctttcgggccgcgat;
[0147] R6-f:
[0148] TGCATATGTTGTGCGTGATCGTGACGCTGATAAAGCTTTCATGAAAATCCTGACGTGGCTGGCGCTGATGgatatgctggcgaccgcgc;
[0149] R6-r:
[0150] CGCACAACATATGCAGCCAGTGCAGCAGCACGCAGAGCTTCTTCACGGTTCAGACCACCTACGATTACGTAcgctttcgggccgcgat;
[0151] R7-f:
[0152] CTCAGGTAATGAACTCCGATAAGCTGGAACGCCTGGCGGCCGTACTGGCGGCGATGCTGGCGCTGATGgatatgctggcgaccgcgc;
[0153] R7-r:
[0154] AGTTCATTACCTGAGCCGCCAGAGCGAAAGCGGTAACCAGCAGTTTCAGGTTGGCGATAACTACAACGTAcgctttcgggccgcgat;
[0155] R8-f:
[0156] CTGGCGATGAAAGCGGCGGCTAAGAACGCGAACAACGGTCTGTACACCGCGTTCGCCCTGGTGGCGTACATGgatatgctggcgaccgcgc;
[0157] R8-r:
[0158] CGCTTTCATCGCCAGTTTCAGGGCTTCCATGTAAACACGCATTACCTCCTGGTTGGACAGCAGAACCACGTAcgctttcgggccgcgat;
[0159] R9-f:CAAAACCCAGCTGCGCAAAGACAACCGTGCTGACGCCGCGCTGGTAgatatgctggcgaccgcgc;
[0160] R9-r: CGCAGCTGGGTTTTGATCAGTTCTTTGTTAATCACCACCACAACGTAcgctttcgggccgcgat;
[0161] R10-f:
[0162] ATCGTGACGCAGATATGGCGGAACTGCTGCTGCGTGCTCTGGCTATCGCAGCATTCgatatgctggcgaccgcgc;
[0163] R10-r:
[0164] TATCTGCGTCACGATCGCCGGTATAAACCGCACCCATCTTAACAACAACAACACGcgctttcgggccgcgat.
[0165] The sequencing vector was correctly used to prepare proteins (R1-R10) according to the method described in Example 1. Protein expression varied ( Figure 4 (A represents the SDS-PAGE results of 10 proteins). R1, R4, R5, and R7 showed extremely low expression levels; R2, R3, R6, R8, and R9 showed relatively high expression levels; and R10 did not oligomerize. Therefore, molecular sieve analysis was used to verify the oligomerization of this protein. Figure 4 (where B represents the results of four protein molecular sieves). The results show that the oligomerization states of R3 and R9 are relatively uniform, while the oligomerization states of R2 and R8 are poor. This result also reflects the diversity of protein design results.
[0166] 3. DNA testing capability assessment
[0167] Further characterization of wild-type NPB and nine designed (R1-R9) proteins was performed. The membrane formation and pore embedding, as well as the recording of pore current and DNA perforation signals, were performed as follows: an artificial phospholipid monolayer (diphyidophosphatidylcholine, DPhPC) was formed, and a single nanoporous protein was embedded in it. The current changes were then recorded at 180 mV.
[0168] 1. The nanopore embedding procedure is as follows: Electrochemical signals were measured from the nanopores embedded in the DPhPC phospholipid bilayer in a buffer solution (600 mM KCl, 75 mM K3[Fe(CN)6, 25 mM K4[Fe(CN)6]·3H2O, 100 mM Hepes, pH 8.0). After achieving single-pore insertion into the phospholipid bilayer, 2 mL of buffer solution (600 mM KCl, 75 mM K3[Fe(CN)6, 25 mM K4[Fe(CN)6]·3H2O, 100 mM Hepes, pH 8.0) was passed through the system to remove excess residual nanopores, obtaining a single nanopore signal acquisition system. The current signal of the nanopore mutant protein on the phospholipid membrane was recorded.
[0169] 2. The process of preparing the DNA sample to be tested and identifying the DNA detection capability is as follows: After the single nanopore signal acquisition system is constructed, the DNA sample to be tested with the T4Dda mutant protein assembled, ATP (final concentration 2mM) and MgCl2 (final concentration 10mM) are flowed into the single nanopore experimental system (total volume 100μL), and the signal is measured at a constant voltage of +150mV.
[0170] The DNA sample to be tested was assembled according to the method recorded in patent (WO2014135838A1) using T4Dda mutant protein (T4Dda mutant protein is a helicase, a rate-controlling protein, specifically T4Dda-E94C / C109A / C136A / A360C mutant protein, which is described in US20170283470A1). The DNA sample to be tested was 0.5kb in length (https: / / doi.org / 10.1038 / s41587-020-0570-8; WO 2019002893A1, sequence as shown in SEQ ID: 12), containing 5 repetitive sequences, each containing 10 Ts (10 As on the complementary strand).
[0171] Test results as follows Figure 5 As shown, the vertical axis represents current, in nA. The results show that the pore current results correspond to the molecular sieve results. The current properties of R3 and R9, which are in better oligomeric states, are better than other designed proteins in both pore current and permeation signal, and their current properties are significantly superior to those of the NPB wild-type protein. Figure 5 The above figure shows WT_NPB, which has a relatively large current of 0.7nA and high noise, but all of them exhibit uneven current distribution. Figure 5R3 and R9 exhibit two types of pore currents: pore current 1 and pore current 2, suggesting that they may contain pore proteins of different diameters (oligomeric state). Other proteins (R1, R4, R5, R6, R7, R2, and R8) have smaller signals than wild-type NPB, but they all have certain problems such as: high noise, severe spontaneous blocking, small DNA perforation signal amplitude, or significant noise during perforation.
[0172] 4. Cryo-electron microscopy structural analysis of R3 and R9
[0173] To verify the rationality of the protein design and to provide a structural basis for subsequent modifications based on R3 or R9, this embodiment further attempted to resolve the cryo-electron microscopy structures of R3 and R9. Following the method described in Example 2, the samples underwent extensive screening. Data could be collected for R3, but no cryo-electron microscopy data was obtained for R9 due to its poor cryo-resistance. The R3 structure resolution results are as follows... Figure 6 As shown, A represents the two-dimensional classification result, and B represents the three-dimensional reconstruction result. Data processing reveals that the two-dimensional classification shows that R3 has at least two different channel types with varying diameters. Figure 6 (The results shown in box A); From the 3D reconstruction results, only one structure conforms to the original design (11-mer), but the density of the shrinkage zone is poor; the other one, although the oligomer state has not changed (also 11-mer), is more loose and has a larger overall density diameter.
[0174] In summary, different contraction regions designed based on the NPB wild type can all improve the stability of pore current, with R3 and R9 showing significant improvements in the detection performance of DNA substrates.
[0175] Example 4: Modification of sequencing contraction regions and identification of sequencing properties based on NPB-R3
[0176] Given the inhomogeneity in the cryo-electron microscopy structure and current detection of R3 in Example 3, this example attempts to improve the oligomerization of R3 through amino acid mutation and other methods. Based on the results of Example 3, the density of the R3 contraction region is poor, suggesting significant steric hindrance between individual molecules in the contraction region. Changing the spatial positions between molecules mainly relies on amino acid mutation and deletion. Considering that this situation primarily exists in the contraction region, this example mainly focuses on attempting amino acid mutations or deletions in this region.
[0177] In this embodiment, amino acid mutations or deletions are mainly concentrated at the L74 and D64 positions. For the L74 position, it is mutated to one of the other 19 amino acids and the L74 position is deleted (corresponding numbers P6-7, P8-11, P36-46, Table 3); the D64 position is mutated to P / A / V / T / S (corresponding numbers P12-P16, Table 3).
[0178] Table 3: Statistics on mutations or deletions based on R3
[0179] serial number Corresponding mutation or deletion Corresponding protein SEQ ID NO: P6 L74P 13 P7 del-L74 14 P8-P11 L74V / T / S / I 15,16,17,18 P23-25 L74G / C / D 19,20,21 P36-39 L74A / M / N / Q 22,23,24,25 P40-43 L74E / R / K / H 26,27,28,29 P44-46 L74W / F / Y 30,31,32 P12-P16 D64P / A / V / T / S 33,34,35,36,37
[0180] 1. Preparation of expression vectors
[0181] Expression vectors were constructed using the method described in Example 1. Using the R3 expression vector gene as a template, the target fragment (i.e., the PCR fragment and the PCR vector) was amplified by PCR using corresponding primers (f represents the forward primer, r represents the reverse primer). After gel recovery, the PCR fragment and the PCR vector were ligated, and then transformed into DH5α competent cells for screening of positive clones. Two clones were selected for sequencing. After successful sequencing, the constructed plasmid was stored at -20℃ for later use. P6-P16, P23-P25, and P36-P46 expression vectors were obtained according to the aforementioned method. In the following protein-coding gene nucleotide sequences, lowercase letters indicate substitutions compared to the R3 coding gene, "~" indicates deletions compared to the R3 coding gene, and "+" represents a connector.
[0182] Where seq3 represents
[0183] ATGTTTCGCCTGCTGACCCTGCTGGCCGGCGTTCTGTTACTGGCGGTGAGCTGCGTGAGCAGCGGCGTGCAGACCCAAGTGGATACCACCGGCCCGACCGAGCCAAGTGCTGACCTATCGCGGCCCGAAAGCGCGTGTAGTTGTTCTGCCGCTGCGCATGGTTATGTCTGATCTGGCAGATGTTTCCGATGCAGAAGCTGCAGCTCTGGCTGCACGT;
[0184] seq4 represents
[0185] .
[0186] The P6 expression vector contains the P6 encoding gene, which expresses the P6 protein with a 6×his tag. The P6 protein sequence is shown in SEQ ID NO: 13.
[0187] The nucleotide sequence of the P6 encoding gene is as follows: 5'-seq3+ccg+seq4-3'.
[0188] The P7 expression vector contains the P7 encoding gene and expresses the P7 protein with a 6×his tag. The P7 protein sequence is shown in SEQ ID NO: 14.
[0189] The nucleotide sequence of the P7 encoding gene is as follows: 5'-seq3+~~~+seq4-3'.
[0190] The P8 expression vector contains the P8 encoding gene, which expresses the P8 protein with a 6×his tag. The P8 protein sequence is shown in SEQ ID NO: 15.
[0191] The nucleotide sequence of the P8 encoding gene is as follows: 5'-seq3+gtg+seq4-3'.
[0192] The P9 expression vector contains the P9 encoding gene and expresses the P9 protein with a 6×his tag. The P9 protein sequence is shown in SEQ ID NO: 16.
[0193] The nucleotide sequence of the P9 encoding gene is as follows: 5'-seq3+acc+seq4-3'.
[0194] The P10 expression vector contains the P10 encoding gene, which expresses the P10 protein with a 6×his tag. The P10 protein sequence is shown in SEQ ID NO: 17. The nucleotide sequence of the P10 encoding gene is as follows: 5'-seq3+agc+seq4-3'.
[0195] The P11 expression vector contains the P11 encoding gene, which expresses the P11 protein with a 6×hi tag. The P11 protein sequence is shown in SEQ ID NO: 18. The nucleotide sequence of the P11 encoding gene is as follows: 5'-seq3+att+seq4-3'.
[0196] The P23 expression vector contains the P23 encoding gene, which expresses the P23 protein with a 6×his tag. The P23 protein sequence is shown in SEQ ID NO: 19. The nucleotide sequence of the P23 encoding gene is as follows: 5'-seq3+gtt+seq4-3'.
[0197] The P24 expression vector contains the P24 encoding gene, which expresses the P24 protein with a 6×his tag. The P24 protein sequence is shown in SEQ ID NO: 20. The nucleotide sequence of the P24 encoding gene is as follows: 5'-seq3+tgt+seq4-3'.
[0198] The P25 expression vector contains the P25 encoding gene, which expresses the P25 protein with a 6×his tag. The P25 protein sequence is shown in SEQ ID NO: 21. The nucleotide sequence of the P25 encoding gene is as follows: 5'-seq3+gat+seq4-3'.
[0199] The P36 expression vector contains the P36 encoding gene, which expresses the P36 protein with a 6×his tag. The P36 protein sequence is shown in SEQ ID NO: 22. The nucleotide sequence of the P36 encoding gene is as follows: 5'-seq3+gct+seq4-3'.
[0200] The P37 expression vector contains the P37 encoding gene, which expresses the P37 protein with a 6×his tag. The P37 protein sequence is shown in SEQ ID NO: 23. The nucleotide sequence of the P37 encoding gene is as follows: 5'-seq3+atg+seq4-3'.
[0201] The P38 expression vector contains the P38 encoding gene, which expresses the P38 protein with a 6×his tag. The P38 protein sequence is shown in SEQ ID NO: 24. The nucleotide sequence of the P38 encoding gene is as follows: 5'-seq3+aac+seq4-3'.
[0202] The P39 expression vector contains the P39 encoding gene, which expresses the P39 protein with a 6×his tag. The P39 protein sequence is shown in SEQ ID NO: 25. The nucleotide sequence of the P39 encoding gene is as follows: 5'-seq3+cag+seq4-3'.
[0203] The P40 expression vector contains the P40 encoding gene, which expresses the P40 protein with a 6×his tag. The P40 protein sequence is shown in SEQ ID NO: 26. The nucleotide sequence of the P40 encoding gene is as follows: 5'-seq3+gaa+seq4-3'.
[0204] The P41 expression vector contains the P41 encoding gene, which expresses the P41 protein with a 6×hi tag. The P41 protein sequence is shown in SEQ ID NO: 27. The nucleotide sequence of the P41 encoding gene is as follows: 5'-seq3+cgt+seq4-3'.
[0205] The P42 expression vector contains the P42 encoding gene, which expresses the P42 protein with a 6×his tag. The P42 protein sequence is shown in SEQ ID NO: 28. The nucleotide sequence of the P42 encoding gene is as follows: 5'-seq3+aaa+seq4-3'.
[0206] The P43 expression vector contains the P43 encoding gene, which expresses the P43 protein with a 6×his tag. The P43 protein sequence is shown in SEQ ID NO: 29. The nucleotide sequence of the P43 encoding gene is as follows: 5'-seq3+cac+seq4-3'.
[0207] The P44 expression vector contains the P44 encoding gene, which expresses the P44 protein with a 6×his tag. The P44 protein sequence is shown in SEQ ID NO: 30. The nucleotide sequence of the P44 encoding gene is as follows: 5'-seq3+tgg+seq4-3'.
[0208] The P45 expression vector contains the P45 encoding gene, which expresses the P45 protein with a 6×his tag. The P45 protein sequence is shown in SEQ ID NO: 31. The nucleotide sequence of the P45 encoding gene is as follows: 5'-seq3+ttt+seq4-3'.
[0209] The P46 expression vector contains the P46 encoding gene, which expresses the P46 protein with a 6×his tag. The P46 protein sequence is shown in SEQ ID NO: 32. The nucleotide sequence of the P46 encoding gene is as follows: 5'-seq3+tat+seq4-3'.
[0210] Where seq5 represents
[0211] ATGTTTCGCCTGCTGACCCTGCTGGCCGGCGTTCTGTTACTGGCGGTGAGCTGCGTGAGCAGCGGCGTGCAGACCCAAGTGGATACCACCGGCCCGACCGAGCCAAGTGCTGACCTATCGCGGCCCGAAAGCGCGTGTAGTTGTTCTGCCGCTGCGCATGGTTATGTCTGATCTGGCAGATGTTTCC;
[0212] seq6 represents
[0213] GCAGAAGCTGCAGCTCTGGCTGCACGTCTGGCTGCTGCTGTCATGGCAGATATGCTGGCGACCGCGCTGTTTCGCACCGGCCGCTTTATTGTGCTGGAACGCGGCGAAGGCCTGAAAGAAATTCAGAAAGAACTGGATCTGGCGCAGAGCGGCTATGTGCGCAAAGATCAAGCGCCGCAGACCGGTCAGATGGAAGGCGCGGATATTCTGGTGATTGGCGCGATTACCGCGTTTGAA CCGAACGCGAGCGGCGTGGAAGGCGGTGGCGTGGTGATTCCGTATAAAATTCCGATTATTGGCGGCGCGCGCCTGAAAAAGAAAGAAGCGTATATTGCGGCGGATATTCGCCTGGTGGATGTGCGCACCGGTCGCATCATCAACGCGACCACCGTGGAAGGCAAAGCGAGCAGCTGGAAAGCGCAAGGCGTGCTGGGCGGCGTTATTGGCGACGTTGCCTTAGGCGGTGGCCTGGGCG
[0214] TGTATCGCAACACCCCGATGGAAAAAGCGATTCGTCAGATGCTGTATGCGGCGGTGAACGCGATTGTGCAGATGGTGCCGCCGAACTATTATCGCTGGGGCGAAGATAACGTGCCGGCGAGCGGCACCAAA.
[0215] The P12 expression vector contains the P12 encoding gene, which expresses the P12 protein with a 6×his tag. The P12 protein sequence is shown in SEQ ID NO: 33. The nucleotide sequence of the P12 encoding gene is as follows: 5'-seq5+cct+seq6-3'.
[0216] The P13 expression vector contains the P13 encoding gene, which expresses the P13 protein with a 6×his tag. The P13 protein sequence is shown in SEQ ID NO: 34. The nucleotide sequence of the P13 encoding gene is as follows: 5'-seq5+gca+seq6-3'.
[0217] The P14 expression vector contains the P14 encoding gene, which expresses the P14 protein with a 6×his tag. The P14 protein sequence is shown in SEQ ID NO: 35. The nucleotide sequence of the P14 encoding gene is as follows: 5'-seq5+gtg+seq6-3'.
[0218] The P15 expression vector contains the P15 encoding gene, which expresses the P15 protein with a 6×his tag. The P15 protein sequence is shown in SEQ ID NO: 36. The nucleotide sequence of the P15 encoding gene is as follows: 5'-seq5+acc+seq6-3'.
[0219] The P16 expression vector contains the P16 encoding gene, which expresses the P16 protein with a 6×his tag. The P16 protein sequence is shown in SEQ ID NO: 37. The nucleotide sequence of the P16 encoding gene is as follows: 5'-seq5+agc+seq6-3'.
[0220] The PCR system and PCR procedure were the same as in Example 1. The seamless ligation system was the same as in Example 1, except that the linearized vector was replaced with the PCR vector.
[0221] Vector primers: Vf: gatatgctggcgaccgcgc; Vr: cgctttcgggccgcgat.
[0222] Fragment primers: Primers used for P6 and P7:
[0223] P6-f: GCACGTccgGCTGCTGCTGTCATGGCA; P6-r: AGCAGCcggACGTGCAGCCAGAGCTG;
[0224] P7-f: GCTGCACGTGCTGCTGCTGTCATGGCA; P7-r: AGCAGCAGCACGTGCAGCCAGAGCTG;
[0225] For P8 - 11, P36 - 46 below, the forward primer is always P8-f, P8-f: GCTGCTGCTGTCATGGCA; the reverse primer is the corresponding -r primer
[0226] P8-r: CATGACAGCAGCAGCcacACGTGCAGCCAGAGCTG;
[0227] P9-r: CATGACAGCAGCAGCggtACGTGCAGCCAGAGCTG;
[0228] P10-r: CATGACAGCAGCAGCgctACGTGCAGCCAGAGCTG;
[0229] P11-r: CATGACAGCAGCAGCaatACGTGCAGCCAGAGCTG;
[0230] P23-r: CATGACAGCAGCAGCaccACGTGCAGCCAGAGCTG;
[0231] P24-r: CATGACAGCAGCAGCacaACGTGCAGCCAGAGCTG;
[0232] P25-r: CATGACAGCAGCAGCatcACGTGCAGCCAGAGCTG;
[0233] P36-r: CATGACAGCAGCAGCagcACGTGCAGCCAGAGCTG;
[0234] P37-r: CATGACAGCAGCAGCcatACGTGCAGCCAGAGCTG;
[0235] P38-r: CATGACAGCAGCAGCgttACGTGCAGCCAGAGCTG;
[0236] P39-r: CATGACAGCAGCAGCctgACGTGCAGCCAGAGCTG;
[0237] P40-r: CATGACAGCAGCAGCttcACGTGCAGCAGAGCTG;
[0238] P41-r: CATGACAGCAGCAGCacgACGTGCAGCAGAGCTG;
[0239] P42-r: CATGACAGCAGCAGCtttACGTGCAGCAGAGCTG;
[0240] P43-r: CATGACAGCAGCAGCgtgACGTGCAGCAGAGCTG;
[0241] P44-r: CATGACAGCAGCAGCccaACGTGCAGCAGAGCTG;
[0242] P45-r: CATGACAGCAGCAGCaaaACGTGCAGCAGAGCTG;
[0243] P46-r: CATGACAGCAGCAGCataACGTGCAGCAGAGCTG;
[0244] For pages 12-16 below, the reverse primer is P12-r, P12-r: GGAAACATCTGCCAGATCAGACA;
[0245] The forward direction corresponds to the -f primer: P12-f: TGTTTCCcctGCAGAAGCTGCAGCTCTG;
[0246] P13-f: TGTTTCCGCAGCAGAAGCTGCAGCTCTG; P14-f: TGTTTCCgtgGCAGAAGCTGCAGCTCTG; P15-f: TGTTTCCaccGCAGAAGCTGCAGCTCTG; P16-f: TGTTTCCagcGCAGAAGCTGCAGCTCTG.
[0247] 2. Protein preparation: Proteins were prepared according to the method described in Example 1 using the sequencing vector.
[0248] Protein expression status as follows Figure 7 As shown, all proteins were expressed normally, and the oligomeric state changed significantly with and without heating, while the mutant protein formed pores normally.
[0249] 3. Current property detection: The current properties of these mutant proteins were detected according to the method in Example 3.
[0250] Test results as follows Figure 8 As shown in the figure. The results indicate that after the L74 mutation to V / T / S / I (P8-11), the pore current became uniform, with only one current signal detected. The noise and spontaneous blockage of the pore current were significantly improved, and a clear signal could be recorded for the DNA substrate, with a high substrate DNA capture rate. Therefore, P8-P11 can be basically used for DNA substrate sequencing, with P10 exhibiting relatively less current noise. However, the substrate DNA signal width through the pores in P8-P11 is small, with pore currents between 0.13 nA and 0.21 nA and a signal width less than 50 pA.
[0251] Other mutant proteins also present corresponding problems: For example, the P23-P25 mutations, while allowing normal pore insertion, exhibit uneven pore currents in P23 and P25, containing many noisy pores. This noise frequently accumulates during sequencing and is uncontrollable. The P24 mutation shows generally uniform pore current, but the sequencing process is also unstable, with significant noise. Similarly, the P36-P39 mutations primarily suffer from very low pore opening current and through-pore signal amplitude, similar to the P8-P11 mutations, but with more spontaneous pore blocking. Overall, L74 mutations, except for P8-P11, manifest as either high noise or frequent spontaneous blocking.
[0252] 4. Electron microscopy observation: The uniformity of the oligomeric state of P10 was verified by cryo-electron microscopy. P10 cryo-electron data were collected using the method described in Example 2.
[0253] Observation results as follows Figure 9 As shown in the figure, A represents the two-dimensional classification result, and B represents the three-dimensional reconstructed structure. The results show that both two-dimensional and three-dimensional classifications indicate that P10 is a homogeneous 11-mer. Therefore, by mutating the designed protein, a more homogeneous pore protein was obtained, significantly improving sequencing performance.
[0254] Example 5: Identification of Sequencing Properties Based on NPB-R-Based Constriction Region Deletion Modification
[0255] The width of the well signal represents the signal-to-noise ratio, and this parameter directly affects the accuracy of sequencing. Given the relatively small width of the substrate DNA well signal in P8-P11 wells (…),… Figure 8 The pore current is between 0.13nA and 0.21nA, and the signal width is less than 50pA. In this embodiment, we mainly attempt to delete amino acids in the P10 (L74S) contraction region. The deletion length is 1-5 amino acids, and the deletion position is before and after S74 (P10). The amino acid types after deletion are shown in Table 4. In the table, "-" represents the deletion of one amino acid. The corresponding numbers of the deleted proteins are P54-P64.
[0256] Table 4: Statistical Table Based on P10 Amino Acid Mutation Types
[0257] serial number Corresponding mutation or deletion P10 <![CDATA[ 66 EAAAAAAR 74 ARRIVALS]]> P54 <![CDATA[ 66 EAAAR 74 S-AAVMAD]]> P55 <![CDATA[ 66 EAAAR 74 S--AVMAD]]> P56 <![CDATA[ 66 EAAAR 74 S---VMAD]]> P57 <![CDATA[ 66 EAAAAAAR 74 S----MAD]]> P58 <![CDATA[ 66 EAAALAAR 74 S-----AD]]> P59 <![CDATA[ 66 EAAAAAAA- 74 SAAAVMAD<!-- 16 --> ]]> P60 <![CDATA[ 66 EAAAAAA-- 74 ARRIVALS]]> P61 <![CDATA[ 66 EAAAAL--- 74 ARRIVALS]]> P62 <![CDATA[ 66 EAAA---- 74 SAAAVMAD]]> P63 <![CDATA[ 66 EAA----- 74 ARRIVALS]]> P64 <![CDATA[ 66 EA------ 74 SAAAVMAD]]>
[0258] 1. Preparation of expression vectors
[0259] Expression vectors were constructed using the method described in Example 1. Using the P10 expression vector gene as a template, the target fragment (i.e., the PCR fragment and the PCR vector) was amplified by PCR using corresponding primers (f represents the forward primer, r represents the reverse primer). After gel recovery, the PCR fragment and the PCR vector were ligated, and then transformed into DH5α competent cells for screening of positive clones. Two clones were selected for sequencing. After successful sequencing, the constructed plasmid was stored at -20℃ for later use. P6-P16, P23-P25, and P36-P46 expression vectors were obtained according to the aforementioned method.
[0260] In the following protein-coding gene nucleotide sequences, lowercase letters indicate substitutions compared to the P10-coding gene, "~" indicates deletions compared to the P10-coding gene, and "+" represents a connector. Here, seq7 represents...
[0261] ATGTTTCGCCTGCTGACCCTGCTGGCCGGCGTTCTGTTACTGGCGGTGAGCTGCGTGAGCAGCGGCGTGCAGACCCAAGTGGATACCACCGGCCCGACCGCGAGCCAAGTGCTGACCTATCGCGGCCCGAAAGCGCGTGTAGTTGTTCTGCCGCTGCGCATGGTTATGTCTGATCTGGCAGATGTTTCCGATGCAGAAGCT;
[0262] seq8 represents
[0263] .
[0264] The P54 expression vector contains the P54 encoding gene, which expresses the P54 protein with a 6×his tag. The P54 protein sequence is shown in SEQ ID NO: 38. The nucleotide sequence of the P54 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~GCTGCTGTCATG+seq8-3'.
[0265] The P55 expression vector contains the P55 encoding gene, which expresses the P55 protein with a 6×hi tag. The P55 protein sequence is shown in SEQ ID NO: 39. The nucleotide sequence of the P55 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~~~~~GCTGTCATG+seq8-3'.
[0266] The P56 expression vector contains the P56 encoding gene, which expresses the P56 protein with a 6×hi tag. The P56 protein sequence is shown in SEQ ID NO: 40. The nucleotide sequence of the P55 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~~~~~~~GTCATG+seq8-3'.
[0267] The P57 expression vector contains the P57 encoding gene, which expresses the P57 protein with a 6×hi tag. The P57 protein sequence is shown in SEQ ID NO: 41. The nucleotide sequence of the P57 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~~~~~~~~~~ATG+seq8-3'.
[0268] The P58 expression vector contains the P58 encoding gene, which expresses the P58 protein with a 6×hi tag. The P58 protein sequence is shown in SEQ ID NO: 42. The nucleotide sequence of the P58 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~~~~~~~~~~~~~~+seq8-3'.
[0269] The P59 expression vector contains the P59 encoding gene, which expresses the P59 protein with a 6×hi tag. The P59 protein sequence is shown in SEQ ID NO: 43. The nucleotide sequence of the P59 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCA~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0270] The P60 expression vector contains the P60 encoding gene, which expresses the P60 protein with a 6×his tag. The P60 protein sequence is shown in SEQ ID NO: 44. The nucleotide sequence of the P60 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCT~~~~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0271] The P61 expression vector contains the P61 encoding gene, which expresses the P61 protein with a 6×hi tag. The P61 protein sequence is shown in SEQ ID NO: 45. The nucleotide sequence of the P61 encoding gene is as follows: 5'-seq7+GCAGCTCTG~~~~~~~~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0272] The P62 expression vector contains the P62 encoding gene, which expresses the P62 protein with a 6×hi tag. The P62 protein sequence is shown in SEQ ID NO: 46. The nucleotide sequence of the P62 encoding gene is as follows: 5'-seq7+GCAGCT~~~~~~~~~~~~~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0273] The P63 expression vector contains the P63 encoding gene, which expresses the P63 protein with a 6×hi tag. The P63 protein sequence is shown in SEQ ID NO: 47. The nucleotide sequence of the P63 encoding gene is as follows: 5'-seq7+GCA~~~~~~~~~~~~~~~~~~~~~~~~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0274] The P64 expression vector contains the P64 encoding gene, which expresses the P64 protein with a 6×hi tag. The P64 protein sequence is shown in SEQ ID NO: 48. The nucleotide sequence of the P64 encoding gene is as follows: 5'-seq7+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~seq8-3'.
[0275] The PCR system and PCR procedure were the same as in Example 1. The seamless ligation system was the same as in Example 1, except that the linearized vector was replaced with the PCR vector.
[0276] The mutation primers used for the P54-P64 expression vector are as follows:
[0277] The following expression vectors (P54-58) use P54-f primers for both forward and reverse directions: P54-f: GCTGCTGTCATGGCAgatatgc; the reverse direction uses the corresponding -r primers.
[0278] P54-r: TGCCATGACAGCAGCgctACGTGCAGCCAGAGC;
[0279] P55-r: atcTGCCATGACAGCgctACGTGCAGCAGAGC;
[0280] P56-r: catatcTGCCATGACgctACGTGCAGCAGAGC;
[0281] P57-r: cagcatatcTGCCATgctACGTGCAGCCAGAGC; P58-r: cgccagcatatcTGCgctACGTGCAGCCAGAGC;
[0282] The following expression vectors (P59-64) use P59-f primers for both forward and reverse directions. P59-f: agcGCTGCTGCTGTCAT; the reverse direction uses the corresponding -r primers.
[0283] P59-r: GACAGCAGCAGCgctTGCAGCCAGAGCTGCAGC;
[0284] P60-r: GACAGCAGCAGCgctAGCCAGAGCTGCAGCTTC;
[0285] P61-r: GACAGCAGCAGCgctCAGAGCTGCAGCTTCTGCAT;
[0286] P62-r: GACAGCAGCAGCgctAGCTGCAGCTTCTGCATCG;
[0287] P63-r: GACAGCAGCAGCgctTGCAGCTCTCTGCATCGGAAAC;
[0288] P64-r: GACAGCAGCAGCgctAGCTTCTGCATCGGAAACATCTG.
[0289] 2. Protein preparation
[0290] The sequencing vector was correctly prepared for protein preparation according to the method described in Example 1. Protein expression was as follows: Figure 10 As shown in the figure. The results showed that, except for P56, all other proteins with deleted amino acids were expressed normally, and the oligomeric state changed significantly between heating and not heating, indicating that the mutant protein formed pores normally.
[0291] 3. Monitoring of Current Properties
[0292] The current properties of mutant proteins P10, P54-P55, and P57-P64 were detected according to the method in Example 3. The results of the via property detection for P10, P54, and P61 are as follows: Figure 11 As shown, compared to P10 (P10 pore current is 0.12 nA, DNA perforation signal amplitude is 25 pA), Figure 11 In the P54 and P61 channels, the current and DNA signal amplitude were significantly increased. The P54 channel current was 0.22 nA, and the DNA signal amplitude was 65 pA. Figure 11 (B in the text); the current in the P61 well was 0.24 nA, and the DNA perforation signal amplitude was 110 pA ( Figure 11 (C in the text). Therefore, the DNA sequencing performance of P54 and P61 was greatly improved, with P61 showing a more significant improvement.
[0293] Other mutation detection results are as follows Figure 12 As shown, P59 represents the pore current signal of the P59 protein, P63 represents the pore current signal of the P63 protein, P62 represents the pore current and pore current signal of the P62 protein, and P64 represents the pore current and pore current signal of the P64 protein. P59 shows a relatively small pore current signal for complementary DNA, and many signal fluctuations occur during DNA entry and exit from the nanopore. Figure 12 P59, through-hole signal 2); P62 has two types of pore currents, and only under a larger current can the DNA substrate pass through the pore, resulting in a smaller signal amplitude. Figure 12 The via current at the top of P62 is around 50pA; the via efficiency of the substrate at P63 is relatively low. Figure 12 P63); P64 exhibits multiple pore currents, all of which are unstable and can pass through DNA substrates, but the signal amplitude is relatively small. Figure 12 (P64).
[0294] Example 6: Identification of P61 DNA Sequencing Properties
[0295] To further examine the sequencing performance of P61, in this embodiment, the sequencing properties of P61 were compared with those of MspA and ONT product R10.4.1 (MinION). The DNA sequencing properties of P61 and MspA were tested according to the method described in Example 3.
[0296] The DNA sequencing properties of P61 and ONT product R10.4.1 (MinION) were detected according to the method in Example 3. The only difference was that the DNA sample to be tested was a de Bruijn sequence (as shown in SEQ ID: 49).
[0297] Test results as follows Figure 13 As shown in the figure, A represents the sequencing signal graphs of P61 and MspA, B represents the sequencing signal graphs of P61 and ONT R10.4.1, and C represents the horizontal noise of P61 and ONT R10.4.1. The results show that P61 sequencing signals have advantages such as high base resolution, low noise, high signal-to-noise ratio, and good sequencing stability.
[0298] Compared to MspA, a transmembrane pore protein commonly used for DNA sequencing, P61 has a more stable contraction region (reader head) and is more sensitive to DNA substrates; for the same DNA sequence, P61 exhibits a greater electrical signal. Figure 13 The figure above shows the MspA current result, and the figure below shows the P61 current result (A in the figure above shows the MspA current result, and P61 current result below shows the MspA current result), indicating that it has a higher ability to identify bases.
[0299] P61 and ONT R10.4.1 were used to detect the de Bruijn sequence (SEQ ID NO: 49) under the same conditions. Analysis showed that P61's performance (signal width, number of steps, etc.) was comparable to ONT R10.4.1. Figure 13 (The image above shows the current results from ONTR10.4.1, and the image below shows the current results from P61.) Furthermore, signal noise has been further reduced. Figure 13 The lower noise level of P61 (as seen in the C-type sequencing) indicates a higher signal-to-noise ratio. Furthermore, the capture efficiency and sequencing stability of P61 both meet the standards for long-duration sequencing, satisfying the demands of high-throughput sequencing.
[0300] In summary, the overall properties of P61 (such as current stability, substrate capture efficiency, signal amplitude, etc.) are already suitable for direct use in DNA sequencing.
[0301] In this embodiment, the MspA nanopores are the mutants (D90N, D91N, D93N, D118R, D134R, and D139K) described in the literature (https: / / doi.org / 10.1016 / j.ymeth.2016.03.026), and their structural schematic diagrams are shown in the literature. Figure 2 As shown), the hole embedding method is the same as in Example 3; R10.4.1 is the purchase of ONT commercial products.
[0302] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims.
Claims
1. A nanoporous protein monomer, characterized in that, The nanoporin monomer is any of the following polypeptides: (a1) A polypeptide with the amino acid sequence shown in SEQ ID NO: 4; (a2) A polypeptide having the same function by substituting and / or deleting and / or adding one or more amino acids of the amino acid sequence shown in SEQ ID NO: 4; (a3) is a polypeptide that has more than 80% identity with any of the amino acid sequences defined in (a1)-(a2) and has the same function; (a4) A fusion polypeptide obtained by attaching a tag to the end of any of the polypeptides defined in (a1)-(a3).
2. The nanoporous protein monomer according to claim 1, characterized in that, (a2) The polypeptide comprises at least one of the following substitutions: The leucine at position 74 is replaced by proline, valine, threonine, serine, isoleucine, glycine, cysteine, aspartic acid, alanine, methionine, asparagine, glutamine, glutamic acid, arginine, lysine, histidine, tryptophan, tyrosine, or phenylalanine. The 64th position of aspartic acid is replaced by proline, alanine, valine, threonine, or serine. Preferably, the polypeptide described in (a2) further comprises the deletion of at least one or more of the following amino acids: SEQ ID NO: 4 Delete one, two, three, four or five amino acids at positions 74-79; and / or SEQ ID NO:4 has one, two, three, four, five, six or seven amino acids deleted at positions 67-73.
3. A construct, characterized in that, The construct comprises at least two connected nanoporin monomers, at least one of which is the nanoporin monomer described in claim 1 or 2.
4. A nanoporous protein, characterized in that, The nanoporin comprises at least two linked nanoporin monomers, at least one of which is the nanoporin monomer described in claim 1 or 2; or The nanoporous protein comprises the construct according to claim 3.
5. The construct according to claim 3 or the nanoporous protein according to claim 4, characterized in that, The nanoporous protein monomers may be the same or different; The connection can be covalent or non-covalent; The nanoporin contains 10-12 nanoporin monomers.
6. The nanoporin monomer of claim 1 or 2, the construct of claim 3 or 5, or the related biomaterial of the nanoporin of claim 4 or 5, characterized in that, The relevant biomaterial is any one of the following: a1) Encoding the nanoporin monomer of claim 1 or 2, the construct of claim 3 or 5, or the nucleic acid molecule of the nanoporin of claim 4 or 5; a2) An expression cassette containing the nucleic acid molecule described in a1); a3) A recombinant vector containing the nucleic acid molecule described in a1), or a recombinant vector containing the expression cassette described in a2); a4) Recombinant cells containing the nucleic acid molecules described in a1), or recombinant cells containing the expression cassette described in a2), or recombinant cells containing the recombinant vector described in a3).
7. The use of the nanoporin monomer of claim 1 or 2, the construct of claim 3 or 5, the nanoporin of claim 4 or 5, or the related biomaterial of claim 6 in detecting the presence, absence, or one or more features of a target analyte or in preparing a product for detecting the presence, absence, or one or more features of a target analyte; Preferably, the target analyte is one or more of nucleotides, nucleic acids, amino acids, oligopeptides, polypeptides, and proteins.
8. A method for determining the presence, absence, or one or more characteristics of a target analyte, characterized in that, The method includes: A. The target analyte comes into contact with the nanoporous protein of claim 4 or 5, causing the target analyte to move relative to the nanoporous protein; B. Acquire one or more measurements as the target analyte moves relative to the nanoporin to determine the presence, absence, or one or more characteristics of the target analyte; Preferably, the target analyte is one or more of nucleotides, nucleic acids, amino acids, oligopeptides, polypeptides, and proteins.
9. A kit or device for determining the presence, absence, or one or more characteristics of a target analyte, characterized in that, The kit comprises the nanoporin monomer of claim 1 or 2, the construct of claim 3 or 5, the nanoporin of claim 4 or 5 or the related biomaterial of claim 6, and a membrane; The device comprises the nanoporous protein of claim 4 or 5, and a membrane; Preferably, the target analyte is one or more of nucleotides, nucleic acids, amino acids, oligopeptides, polypeptides, and proteins.
10. A method for preparing nanoporous protein according to claim 4 or 5, characterized in that, This includes transforming host cells with the recombinant vector as described in claim 6 to induce the host cells to express the nanoporous protein.
Citation Information
Patent Citations
Mutant csgg pores
US20170283470A1
Enzyme stalling method
WO2014135838A1
Novel protein pores
WO2019002893A1
Novel nanopore protein mutant and application thereof
CN117384260A
PHT nanopore mutant protein and application thereof
CN117886907A