Double-constriction-region nanopore protein compound and application thereof
By designing a nanopore protein complex with a dual contraction region, the problem of the scarcity of dual contraction regions in nanopore sequencing technology was solved, improving the accuracy and efficiency of nucleic acid detection, optimizing current signal detection, and enhancing sequencing performance.
Patent Information
- Application Number
- CN202410612481.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-11-18
AI Technical Summary
In existing nanopore sequencing technologies, the dual-contraction region nanopore protein complex is scarce, affecting the accuracy and efficiency of nucleic acid detection.
A dual-contraction nanoporin complex is provided, comprising a mutant NPB nanoporin and a binding agent. A stable contraction region is formed through amino acid sequence design. The binding agent is linked to the mutant NPB nanoporin to form a dual contraction region, thereby improving sequencing performance.
It improved the accuracy of oligopolynucleotide detection, enhanced the amount of sequencing information, and optimized current signal detection, stability, and sequencing performance.
Smart Images

Figure CN120965828A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and in particular to a dual-contraction zone nanoporous protein complex and its applications. Background Technology
[0002] Biological macromolecules such as DNA, RNA, proteins, and polysaccharides are the basic building blocks of living organisms, and their sequence information and post-functional modifications determine their biological functions. Sequence identification technology for biological macromolecules is a core tool for understanding the laws governing life, and single-molecule sequencing technology has emerged as a result.
[0003] Currently, single-molecule sequencing technologies can be mainly divided into two categories: one is optical zero-mode waveguide sequencing, represented by Pacific Biosciences (PacBio) in the United States; the other is electrical nanopore sequencing, represented by Oxford Nanopore Technologies (ONT) in the United Kingdom. Nanopore sequencing technology is a novel nucleic acid sequencing technology developed in recent years. Based on the pore type, it can be divided into solid pores and biological nanopores. Biological nanopores are pore proteins that allow substrates to pass through. The following nanopore sequencing refers specifically to biological nanopore sequencing technology.
[0004] Under the influence of an electric field, charged nucleic acid substrates can pass through biological nanopores. When nucleic acids pass through the nanopore, they impede the current flowing through it, generating different current signals. By analyzing these signals, the base information of the nucleic acid can be obtained. Compared to other sequencing methods, it offers advantages such as low equipment cost, simple sample preparation, and fast sequencing speed, and has begun to be applied in various fields. Specific advantages include: easy library construction without amplification; fast signal readout speed, typically reaching 200-300 bp / s; long readout length, typically reaching thousands of bases; direct detection of modifications on DNA; and direct RNA sequencing. Due to the characteristics of nanopore sequencing, RNA no longer needs to be reverse transcribed into DNA for sequence analysis, thus preserving modification information on the RNA. Because of these advantages, nanopore sequencing technology has gained widespread attention in recent years.
[0005] Nanopore proteins are the core of nanopore sequencing. To date, relatively few nanopore proteins can be used for nucleic acid detection. Only a few natural proteins, such as MspA (Mycobacterium smegmatis porin A), CsgG, and CsgG-CsgF (curli-specific transport channels), meet the requirements, and nanopore protein complexes with dual contraction zones are even rarer. Summary of the Invention
[0006] One of the objectives of this invention is to provide a dual-contraction zone nanoporous protein complex.
[0007] This invention provides a dual-contraction-region nanoporin complex comprising a mutant NPB nanoporin and a conjugate, wherein the conjugate links the mutant NPB nanoporin and forms a contraction region within the mutant NPB nanoporin, and the conjugate is a polypeptide with the amino acid sequence shown in SEQ ID NO: 63.
[0008] The mutant NPB nanoporous protein contains any of the following polypeptides: (a1) a polypeptide with the amino acid sequence shown in SEQ ID NO: 4; (a2) a polypeptide with the same function obtained by substituting and / or deleting and / or adding one or more amino acids of the amino acid sequence shown in SEQ ID NO: 4; (a3) a polypeptide with more than 80% identity to the amino acid sequence defined in (a1)-(a2) and with the same function; (a4) a fusion polypeptide obtained by attaching a tag to the end of any of the polypeptides defined in (a1)-(a3).
[0009] The aforementioned mutant NPB nanoporous protein can specifically be R1-R10 prepared in the following examples. For example, the aforementioned mutant NPB nanoporous protein contains a polypeptide with an amino acid sequence as shown in any of SEQ ID NO: 2-11.
[0010] Optionally, according to the above-described dual-contraction zone nanoporous protein complex, the polypeptide described in (a2) comprises at least one of the following substitutions: leucine at position 74 is replaced by proline, valine, threonine, serine, isoleucine, glycine, cysteine, aspartic acid, alanine, methionine, asparagine, glutamine, glutamic acid, arginine, lysine, histidine, tryptophan, tyrosine, or phenylalanine; and aspartic acid at position 64 is replaced by proline, alanine, valine, threonine, or serine.
[0011] The aforementioned mutant NPB nanoporous protein can specifically be the P6, P8-P16, P23-P25, and P36-P46 proteins prepared in the following examples. For example, the aforementioned mutant NPB nanoporous protein contains a polypeptide with an amino acid sequence as shown in any of SEQ ID NO: 13, 15-37.
[0012] Optionally, according to the above-described dual-contraction zone nanoporous protein complex, the polypeptide of (a2) further comprises at least one or more amino acid deletions as follows: one, two, three, four or five amino acid deletions at positions 74-79 of SEQ ID NO: 4; and / or one, two, three, four, five, six or seven amino acid deletions at positions 67-73 of SEQ ID NO: 4.
[0013] The deletion of one or more amino acids may be the deletion of amino acid position 74 of SEQ ID NO: 4; the deletion of amino acid position 75 of SEQ ID NO: 4; the deletion of amino acid positions 75-76 of SEQ ID NO: 4; the deletion of amino acid positions 75-77 of SEQ ID NO: 4; the deletion of amino acid positions 75-78 of SEQ ID NO: 4; the deletion of amino acid positions 75-79 of SEQ ID NO: 4; the deletion of amino acid position 73 of SEQ ID NO: 4; the deletion of amino acid positions 72-73 of SEQ ID NO: 4; the deletion of amino acid positions 71-73 of SEQ ID NO: 4; the deletion of amino acid positions 70-73 of SEQ ID NO: 4; the deletion of amino acid positions 69-73 of SEQ ID NO: 4; the deletion of amino acid positions 68-73 of SEQ ID NO: 4; or the deletion of amino acid positions 67-73 of SEQ ID NO: 4.
[0014] The aforementioned mutant NPB nanoporous protein can specifically be the P7, P54-P64 proteins prepared in the following examples. For example, the aforementioned mutant NPB nanoporous protein contains a polypeptide with an amino acid sequence as shown in any of SEQ ID NO: 14, 38-48.
[0015] Optionally, according to the above-described dual-contraction nanoporin complex, the conjugate is inserted into the cavity of the mutant NPB nanoporin.
[0016] Optionally, according to the above-described dual-contraction-region nanoporin complex, the mutant NPB nanoporin complex comprises at least two linked mutant NPB nanoporin monomers, for example, 10-12 mutant NPB nanoporin monomers, specifically 10, 11, or 12 mutant NPB nanoporin monomers. The mutant NPB nanoporin monomers may be the same or different. The linkage may be covalent or non-covalent.
[0017] Optionally, according to the above-described dual-contraction zone nanoporin complex, the ratio of the mutant NPB nanoporin monomer to the conjugate is 1:1.
[0018] Optionally, the mutant NPB nanoporin monomer is any of the following polypeptides: (b1) a polypeptide with an amino acid sequence as shown in SEQ ID NO: 4; (b2) a polypeptide with the same function obtained by substituting and / or deleting and / or adding one or more amino acids of the amino acid sequence shown in SEQ ID NO: 4; (b3) a polypeptide with more than 80% identity to the amino acid sequence defined in (b1)-(b2) and with the same function; (b4) a fusion polypeptide obtained by attaching a tag to the end of any of the polypeptides defined in (b1)-(b3).
[0019] In some embodiments, the mutant NPB nanoporous protein monomer described above comprises a polypeptide containing an amino acid sequence as shown in any of SEQ ID NO: 2-11.
[0020] Optionally, the polypeptide described in (b2) comprises at least one of the following substitutions: leucine at position 74 is replaced by proline, valine, threonine, serine, isoleucine, asparagine, tyrosine, aspartic acid, alanine, methionine, gamma-glutamyl, glutamic acid, arginine, lysine, histidine, tryptophan, or phenylalanine; and aspartic acid at position 64 is replaced by proline, alanine, valine, threonine, or serine.
[0021] In some embodiments, the mutant NPB nanoporin monomer described above contains a polypeptide with an amino acid sequence as shown in any of SEQ ID NO: 13, 15-37.
[0022] Optionally, the polypeptide described in (b2) further comprises at least one or more amino acid deletions as follows: one, two, three, four or five amino acid deletions at positions 74-79 of SEQ ID NO: 4; and / or one, two, three, four, five, six or seven amino acid deletions at positions 67-73 of SEQ ID NO: 4.
[0023] The deletion of one or more amino acids may be the deletion of amino acid position 74 of SEQ ID NO: 4; the deletion of amino acid position 75 of SEQ ID NO: 4; the deletion of amino acid positions 75-76 of SEQ ID NO: 4; the deletion of amino acid positions 75-77 of SEQ ID NO: 4; the deletion of amino acid positions 75-78 of SEQ ID NO: 4; the deletion of amino acid positions 75-79 of SEQ ID NO: 4; the deletion of amino acid position 73 of SEQ ID NO: 4; the deletion of amino acid positions 72-73 of SEQ ID NO: 4; the deletion of amino acid positions 71-73 of SEQ ID NO: 4; the deletion of amino acid positions 70-73 of SEQ ID NO: 4; the deletion of amino acid positions 69-73 of SEQ ID NO: 4; the deletion of amino acid positions 68-73 of SEQ ID NO: 4; or the deletion of amino acid positions 67-73 of SEQ ID NO: 4.
[0024] In some embodiments, the mutant NPB nanoporin monomer described above comprises a polypeptide containing an amino acid sequence such as that shown in SEQ ID NO: 14, 38-48.
[0025] Optionally, according to the above-described dual-contraction nanoporin complex, the connection is covalent or non-covalent, specifically hydrophobic and hydrogen bonding interactions. Specifically, the conjugate can be attached to the β-sheet of the mutant NPB nanoporin monomer, for example, the conjugate is attached to positions 143-169 and / or 198-223 of the mutant NPB nanoporin monomer (sequence as shown in SEQ ID NO:45).
[0026] In some embodiments, the above-described dual-contraction zone nanoporous protein complex structure is as follows: Figure 13 As shown on page 61D13 of A, it comprises: (1) a mutant NPB nanoporin, the mutant NPB nanoporin comprising a first opening, an intermediate segment, a second opening, and an inner cavity extending from the first opening through the intermediate segment to the second opening, wherein the inner cavity surface of the intermediate segment defines a contraction zone; and (2) a plurality of binding entities, each of the binding entities containing a mutant NPB nanoporin binding region, wherein the plurality of binding entities form another contraction zone within the intermediate segment of the mutant NPB nanoporin, and the two contraction zones are coaxially spaced apart within the intermediate segment of the mutant NPB nanoporin.
[0027] The aforementioned complexes, proteins, or protein monomers can be obtained by first synthesizing their encoding genes and then expressing them biologically, or they can be synthesized artificially through complete chemical processes.
[0028] In the aforementioned conjugates, proteins, or protein monomers, the tag may refer to a polypeptide or protein expressed in fusion with the target protein using in vitro DNA recombination technology, to facilitate the expression, detection, tracing, and / or purification of the target protein. The protein tag may be a Strep-TagII tag, Flag tag, His tag, MBP tag, HA tag, myc tag, GST tag, and / or SUMO tag, etc.
[0029] In the aforementioned conjugates, proteins, or protein monomers, identity refers to the identity of the amino acid sequences. The identity of amino acid sequences can be determined using homology search sites on the internet, such as the BLAST page on the NCBI homepage. For example, in Advanced BLAST 2.1, using blastp as the program, setting the Expect value to 10, setting all filters to OFF, using BLOSUM62 as the matrix, setting the Gap existence cost, Per residue gap cost, and Lambda ratio to 11, 1, and 0.85 (default values) respectively, and performing an identity search on a pair of amino acid sequences, the identity value (%) can then be obtained.
[0030] In this document, the 80% or more identity can be at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.
[0031] The aforementioned dual-contraction zone nanoporous protein complex can specifically be P61D13 prepared in the following examples.
[0032] The present invention also provides a method for generating the above-mentioned dual-contraction zone nanoporin complex, the method comprising: A1 or A2 as follows: A1 co-expressing one or more mutant NPB nanoporin monomers and the conjugate in a host cell, thereby allowing the formation of the dual-contraction zone nanoporin complex in the cell; A2 contacting one or more mutant NPB nanoporin monomers with the conjugate, thereby allowing the formation of the dual-contraction zone nanoporin complex in vitro.
[0033] Optionally, according to the method described above, the molar ratio of the mutant NPB nanoporin monomer and the conjugate in A2 is 1:5.
[0034] The present invention also provides related biomaterials of the above-mentioned dual-contraction-region nanoporin complex, wherein the related biomaterials are any one of the following: c1) a nucleic acid molecule encoding the above-mentioned dual-contraction-region nanoporin complex; c2) an expression cassette containing the nucleic acid molecule of c1); c3) a recombinant vector containing the nucleic acid molecule of c1), or a recombinant vector containing the expression cassette of c2); c4) a recombinant cell containing the nucleic acid molecule of c1), or a recombinant cell containing the expression cassette of c2), or a recombinant cell containing the recombinant vector of c3).
[0035] The present invention also provides a method for determining the presence, absence, or one or more characteristics of a target analyte, the method comprising: A. contacting the target analyte with the aforementioned dual-contraction zone nanoporin complex, causing the target analyte to move relative to the dual-contraction zone nanoporin complex; B. acquiring one or more measurements as the target analyte moves relative to the dual-contraction zone nanoporin complex, thereby determining the presence, absence, or one or more characteristics of the target analyte;
[0036] The present invention also provides a kit for determining the presence, absence, or one or more characteristics of a target analyte, the kit comprising the above-described dual-contraction zone nanoporous protein complex or the above-described related biomaterials, and a membrane.
[0037] The present invention also provides an apparatus for determining the presence, absence, or one or more characteristics of a target analyte, the apparatus comprising the above-described dual-contraction zone nanoporous protein complex and a membrane.
[0038] In the above-mentioned kits or devices, the membrane and the dual-contraction zone nanoporin complex can be packaged independently, or the dual-contraction zone nanoporin complex can be embedded in the membrane.
[0039] The membrane can be any membrane existing in the prior art, preferably a lipid bilayer. For example, the membrane is a lipid bilayer formed by the self-assembly of block copolymers / phospholipid molecules.
[0040] The kits or devices described above may also include rate-controlling proteins. Rate-controlling proteins may include one or more combinations of nucleic acid-binding proteins, helicases, exonucleases, telomerases, topoisomerases, transcriptases, transloses, and / or polymerases.
[0041] Optionally, the helicase is selected from Hel308 family helicases and modified Hel308 family helicases, RecD helicase and its variants, TrwC helicase and its variants, Dda helicase and its variants, TraI Eco and its variants, XPD Mbu and its variants, Pif1-like helicase and its variants.
[0042] The aforementioned conjugate and related biological materials are also within the scope of protection of this invention. The related biological materials are any one of the following: d1) a nucleic acid molecule encoding the aforementioned conjugate; d2) an expression cassette containing the nucleic acid molecule d1); d3) a recombinant vector containing the nucleic acid molecule d1) or a recombinant vector containing the expression cassette d2); d4) a recombinant cell containing the nucleic acid molecule d1) or a recombinant cell containing the expression cassette d2) or a recombinant cell containing the recombinant vector d3).
[0043] In the above-mentioned biological materials, the nucleic acid molecule can be DNA, such as cDNA, genomic DNA or recombinant DNA; the nucleic acid molecule can also be RNA, such as mRNA, siRNA, shRNA, sgRNA, miRNA or antisense RNA.
[0044] In the aforementioned biological materials, the expression cassette refers to DNA capable of expressing genes in host cells. This DNA may include not only promoters that initiate gene transcription but also terminators that terminate gene transcription. Furthermore, the expression cassette may also include enhancer sequences.
[0045] The application of the aforementioned conjugates or related biomaterials in the preparation of dual-contraction zone nanoporous protein complexes is also within the scope of protection of this invention.
[0046] The application of the aforementioned dual-contraction zone nanoporin complex, related biomaterials of the aforementioned dual-contraction zone nanoporin complex, the aforementioned conjugate, or related biomaterials of the aforementioned conjugate in detecting the presence, absence, or one or more features of the target analyte, or in preparing products for detecting the presence, absence, or one or more features of the target analyte, also falls within the scope of protection of this invention.
[0047] Optionally, the target analyte is one or more of nucleotides, nucleic acids, amino acids, oligopeptides, polypeptides, and proteins.
[0048] Optionally, the one or more features are selected from at least one of (i) the length of the target analyte; (ii) the identity of the target analyte; (iii) the sequence of the target analyte; (iv) the secondary structure of the target analyte; and (v) whether the target analyte is modified. "Identity" refers to the similarity between sequences. Identity can be evaluated visually or by computer software. Using computer software, the identity between two or more sequences can be expressed as a percentage (%), which can be used to evaluate the identity between related sequences.
[0049] Optionally, the nucleic acid can be naturally occurring or artificially synthesized. Specifically, the nucleic acid can be natural DNA, RNA, or modified DNA or RNA, or it can be artificially synthesized nucleic acid, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threonine nucleic acid (TNA), locked nucleic acid (LNA), or other synthetic polymers with nucleoside side chains.
[0050] Optionally, the nucleic acid is single-stranded, double-stranded, or at least partially double-stranded.
[0051] Optionally, the nucleic acid can be of any length. For example, the length of the nucleic acid can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400 or at least 500 nucleotides or nucleotide pairs, or it can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs, or 100000 or more nucleotides or nucleotide pairs.
[0052] Optionally, one or more nucleotides in the nucleic acid may be modified, such as methylated, oxidized, damaged, debased, protein-labeled, tagged, or linked to a spacer in the middle of a polynucleotide sequence.
[0053] When the term "comprising" or "including" is used in this application to describe a protein or nucleic acid sequence, the protein or nucleic acid may be composed of the sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still have the activity described in this invention.
[0054] In this article, standard one-letter codes are used for amino acids. These are as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), gamma-glutamyl (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Standard substitution notation is also used, i.e., L74V means that the L at position 74 of the sequence is replaced by V.
[0055] In this document, the contraction region (also referred to as the readout head or restriction region) refers to the pore defined by the inner surface of the pore or pore complex, which allows ions and target analytes (e.g., but not limited to polynucleotides or single nucleotides) to pass through the pore complex channel. In some embodiments, the contraction region is the narrowest pore in the pore or pore complex.
[0056] Dual readheads (dual restriction regions) significantly improve the detection accuracy of oligopolynucleotides. In this invention, a variety of binders capable of binding to mutant NPB nanoporin were designed using an artificially designed binder approach. Through in vitro assembly and evaluation methods such as current signal detection, binders that stably bind to mutant NPB nanoporin and effectively improve the sequencing information of oligopolynucleotides were obtained. This binder forms a dual-constriction nanoporin complex with the mutant NPB nanoporin, which exhibits superior sequencing performance compared to the mutant NPB nanoporin itself. Attached Figure Description
[0057] Figure 1 The results for NPB wild-type protein purification and molecular sieving are shown. "+" indicates protein samples heated at 100℃ for 10 minutes; "-" indicates protein samples left at room temperature for 10 minutes without heating.
[0058] Figure 2 The results of the NPB wild-type protein structure analysis.
[0059] Figure 3 This is a design based on different contraction region types of NPB wild-type.
[0060] Figure 4The results are for protein purification (R1-R10) and some molecular sieves. "+" indicates protein samples heated at 100℃ for 10 minutes; "-" indicates protein samples left at room temperature for 10 minutes without heating.
[0061] Figure 5 The results show the pore currents and DNA perforation results for NPB wild-type and R1-R10 portions.
[0062] Figure 6 This is the analysis result of the R3 structure.
[0063] Figure 7 The results show the expression of modified proteins based on R3. "+" indicates protein samples heated at 100℃ for 10 minutes; "-" indicates protein samples left at room temperature for 10 minutes without heating.
[0064] Figure 8 The results show the current in the P8-P11 wells and the DNA perforation results.
[0065] Figure 9 This is the result of the structure analysis for P10.
[0066] Figure 10 The results are based on the expression of protein deleted from P10.
[0067] Figure 11 The results show the comparison of the via properties of P10, P54, and P61.
[0068] Figure 12 The results are for the current and via signals of holes P59, P62, P63, and P64.
[0069] Figure 13 Results of the identification of the permeability properties of P61 DNA substrate.
[0070] Figure 14 The results show the changes in the current of the P61-binding peptide.
[0071] Figure 15 The results show a comparison of the sequencing properties of P61 and P61D13, with the dashed boxes indicating magnified views of the corresponding locations. Detailed Implementation
[0072] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0073] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available. All quantitative experiments in the following examples were performed in triplicate, and the results were averaged.
[0074] Example 1: Preparation of wild-type NPB nanopore expression vector and protein
[0075] 1. Construction of wild-type NPB nanoporous protein carrier
[0076] Wild-type NPB nanoporous protein is derived from *Thermodesulfobacteriota bacterium* (ACCESSION: NPB10001, SEQ ID NO: 1). The expression gene for this protein was obtained through artificial synthesis, with codon optimization performed during the synthesis process for expression in *E. coli*. After synthesis, the gene was seamlessly cloned into the pBAD22 vector to obtain the wild-type NPB nanoporous protein vector. A 6×his tag was added to the C-terminus of the protein as an affinity purification tag. The wild-type NPB nanoporous protein vector contains the wild-type NPB nanoporous protein expression gene, which expresses wild-type NPB nanoporous protein with a 6×his tag. The wild-type NPB nanoporous protein sequence is shown in SEQ ID NO: 1.
[0077] The carrier construction steps are as follows:
[0078] Using the wild-type NPB nanoporin expression gene (SEQ ID NO: 50) as a template, the target gene fragment was amplified by PCR using forward and reverse primers (primer F: CAGGAGGAATTAACCATGTTTCGCCTGCTGACCC; primer R: GAACTGCGGGTGGCTCCATTTGGTGCCGCTCGCCG). After gel recovery, the fragment was ligated into the linearized pBAD22 vector and then transformed into DH5α competent cells for screening of positive clones. Two clones were selected for sequencing. After successful sequencing, the constructed plasmid was stored at -20℃ for later use.
[0079] The PCR system is as follows (20 μL):
[0080]
[0081] The PCR procedure is as follows:
[0082]
[0083] The seamless connection system (10μL) is as follows:
[0084] 2× Seamless Linkage Buffer 5μL
[0085] Target fragment (50 ng / μL) 3 μL
[0086] Linearization vector (10 ng / μL) 2 μL
[0087] After reacting at 50℃ for 15 min, 2 μL was transferred into DH5α cells and sequenced.
[0088] 2. Preparation of wild-type NPB nanoporous protein
[0089] Wild-type NPB nanoporous protein was purified to high purity using Ni column affinity chromatography and molecular sieves. Figure 1 (where A represents the SDS-PAGE result), the protein properties are uniform ( Figure 1 (where B represents the molecular sieve results). Furthermore, SDS-PAGE analysis shows that heating produces monomeric proteins (molecular weight approximately 26 kDa); without heating, the protein is primarily in an oligomeric state (i.e., a porous state, with a molecular weight greater than the marker's largest band of 180 kDa), and there are also instances of missing monomers, indicating that wild-type NPB protein has poor stability.
[0090] The expression and purification steps are as follows: 1) After confirming the correct sequencing of the wild-type NPB nanoporous protein vector, it was transformed into BL21(DE3) for expression. 1 mL of seed culture was obtained at 37℃ and 200 rpm, then transferred to 1 L of LB medium and incubated at 37℃ and 200 rpm until OD500 was reached. 600 1) Set the concentration to 1.2, cool to 26℃, add arabinose to a final concentration of 0.4 g / L and induce overnight; 2) Collect bacterial cells at 4000 rpm, resuspend each 1 L of cells in 20 mL of lysis buffer, autoclave and centrifuge at 18000 rpm at 4℃ for 1 hour, collect the precipitated membrane fraction; 3) Resuspend each 1 L of cells in 10 mL of membrane lysis buffer, magnetically stir at 4℃ for 1 hour to extract membrane proteins, centrifuge at 18000 rpm at 4℃ for 1 hour, collect the supernatant membrane protein fraction; 4) Add imidazole to the supernatant to a final concentration of 30 mM and incubate with Ni beads after equilibration with membrane lysis buffer, bind at 4℃ for 1 hour and then perform affinity purification; 5) Transfer the supernatant and Ni beads into a column and allow gravity flow-through; wash with 10 mL of washing buffer and elute with 5 mL of elution buffer; 6) Purify the target protein by molecular sieve purification and detect the purity by SDS-PAGE, the results are as follows. Figure 1 As shown in Figure A, the molecular sieve results are as follows: Figure 1 As shown in B.
[0091] Lysis buffer: 20mM Tris-HCl, 150mM NaCl, pH 8.0.
[0092] Dissolution buffer: 20 mM Tris-HCl, 150 mM NaCl, pH 8.0, 1% LDAO (dodecyl dimethylamine oxide).
[0093] Washing solution: 20mM Tris-HCl, 150mM NaCl, pH 8.0, 0.5% LDAO, 50mM imidazole.
[0094] Eluent: 20 mM Tris-HCl, 150 mM NaCl, pH 8.0, 0.1% LDAO, 200 mM imidazole.
[0095] Molecular sieve buffer: 20mM Tris-HCl, 150mM NaCl, pH 8.0, 0.06% LDAO.
[0096] Example 2: Atomic-level structure determination of wild-type NPB nanoporous protein
[0097] After obtaining wild-type NPB nanoporous protein with high purity and homogeneity (Example 1), its resolution was resolved using cryo-electron microscopy. The atomic-level structure.
[0098] The steps for cryo-electron microscopy (cryo-EM) structural analysis are as follows: 1) Prepare samples purified by molecular sieves, and prepare several backup samples under the same freezing conditions. Select 8 suitable samples for loading into the Talos Arctica 200kV high-end electron microscope; 2) After loading the samples, wait for the vacuum and temperature to stabilize, then open the microscope tube and select suitable square apertures at low magnification. This step is similar to the screening of frozen samples, and its purpose is to select suitable square apertures for subsequent data collection; 3) Calculate the approximate number of square apertures to be selected based on the available time of the electron microscope and the number of images that can be acquired from each square aperture. After selection, take a map; 4) During map taking, you can select apertures that can be used for data collection offline, which can save some time; 5) After the map taking is completed, you can adjust the electron microscope to prepare for data collection. This mainly includes the electron microscope's alignment, background subtraction, and basic data collection parameters (underfocus -1.5μm to -2.5μm, electron dose). 32 frames per second and pixel size The settings were configured; subsequently, after data processing and structural analysis, its density map was obtained, as shown overall. Figure 2 As shown, the overall map is an 11-mer structure. Through homology modeling, the atomic coordinates of the amino acids in the wild-type NPB nanoporous protein were obtained through structure building and refinement. Figure 2 Figure A shows the top and side views of the map, and Figure B shows the resolution information.
[0099] NPB wild-type contraction zone has fewer amino acids (3, N58-C59-Q60) and a larger pore diameter. ( Figure 2 Figure C shows the model construction results of NPB wild-type (top and side views), demonstrating strong overall stability. Considering the current properties of pore proteins, which are largely determined by the contraction region, NPB wild-type can be used as the main framework for novel pore design, allowing for the design of contraction region types with different properties.
[0100] Example 3: Protein design and corresponding protein preparation based on NPB-based contraction regions (NPB-R)
[0101] 1. Protein design
[0102] The protein design procedure involved in this invention includes RFdiffusion for generating protein backbone structures; ProteinMPNN for generating one or more sequences from the backbone; and AlphaFold2 for predicting structures based on the sequences. The predicted structures are analyzed to determine if they meet the requirements, and this process is repeated until a reasonably structured amino acid sequence is generated. Finally, the feasibility of the designed sequences is verified through vector construction, protein expression, structure resolution, and pore current detection.
[0103] Analysis of the structure of nanoporous protein (CsgG), currently capable of accurate nucleic acid sequencing, reveals that its contraction region is a highly flexible loop structure. This structure lacks rigidity and exhibits significant oscillation under electric fields and the influence of DNA substrates, generating noise and affecting sequencing accuracy. Based on the results of Example 2, a novel contraction region was designed using the NPB wild-type structure as the main framework to mitigate the adverse effects of the flexible contraction region. The designed contraction region is predominantly α-helical, and can be categorized into main forms such as loop-α-helix and α-helix-loop-α-helix, exhibiting diversity in conformation and pore diameter. Figure 3 This section presents designs based on different contraction zone types of NPB wild-type nanoporous proteins. A shows top views of three nanoporous protein models: NPB-WT is the wild-type NPB model, and NPB-R3 and NPB-R4 are models with different contraction zone designs using the wild-type NPB structure as the main framework. B shows the structural characteristics of the monomer contraction zone: NPB-WT is the wild-type NPB monomer, and NPB-R1-R10 are R1-R10 nanoporous protein monomers with different contraction zone designs using the wild-type NPB structure as the main framework. Table 1 shows the theoretical diameters of R1-R10 nanoporous proteins with different contraction zone designs using the wild-type NPB structure as the main framework. Table 2 shows the pore amino acid sequences of R1-R10 nanoporous proteins with different contraction zone designs using the wild-type NPB structure as the main framework.
[0104] Table 1: Theoretical Diameter of Design Channels for R1-R10
[0105]
[0106] Table 2: Amino acid sequences of the designed pores from R1 to R10
[0107] serial number Amino acid modification section R1 RIVVVLFTAVYTGSQDARATELAALALAIAVY R2 RVVVVLFTAVYTGSQDARATELAAKALAIAVY R3 RVVVLPLRMVMSDLADVSDAEAAAALAARLAAAVMA R4 FVVVANDLFAKIVRLLAEILMNMNAVTELRHTMVVLAALLAAM R5 YVVVVNDTWYKIVEMMIKLLQASNAVETLEKQAARLAALLSLM R6 YVIVGGLNREEALRAAALAAYVVRDRDADKAFMKILTWLALM R7 YVVVIANLKLLVTAFALAAQVMNSDKLERLAAVLAAMLALM R8 YVVLLSNQEVMRVYMEALKLAMKAAAKNANNGLYTAFALVAYM R9 YVVVVINKELIKTQLRKDNRADAALV R10 RVVVVKMGAVYTGDRDADMAELLLRALAIAAF NPB RIAVASFKCKAANCQGIGEGIA
[0108] 2. Carrier and protein preparation
[0109] After AlphaFold2 predicted the structural rationality, expression vectors were constructed using the method described in Example 1. Using the wild-type NPB nanoporous protein vector gene as a template, the target fragment (i.e., the PCR fragment and the PCR vector) was amplified by PCR using corresponding primers (f represents the forward primer, r represents the reverse primer). After gel recovery, the PCR fragment and the PCR vector were ligated, and then transformed into DH5α competent cells for screening of positive clones. Two clones were selected for sequencing. After successful sequencing, the constructed plasmids were stored at -20℃ for later use. R1-R10 expression vectors were obtained according to the aforementioned method.
[0110] In the following protein-coding gene nucleotide sequences, lowercase letters indicate substitutions compared to the wild-type NPB nanoporous protein expression gene, "~" indicates deletions compared to the wild-type NPB nanoporous protein expression gene, and "+" represents a connector.
[0111] Where seq1 represents
[0112] ATGTTTCGCCTGCTGACCCTGCTGGCCGGCGTTCTGTTACTGGCGGTGAGCTGCGTGAGCAGCGGCGTGCAGACCCAAGTGGATACCACCGGCCCGACCGCGAGCCAAGTGCTGACCTATCGCGGCCCGAAAGCG;
[0113] seq2 represents
[0114] .
[0115] The R1 expression vector contains the R1 encoding gene and expresses the R1 protein with a 6×hi tag. The R1 protein sequence is shown in SEQ ID NO: 2. The nucleotide sequence of the R1 encoding gene is as follows:
[0116] 5'-seq1+cgtattgttgttgtactgttcactgctgtatacactggtagccaggatgcgcgtgctacggaactggcagcactggctctggcgatcgcggtttac+seq2-3'.
[0117] The R2 expression vector contains the R2 encoding gene and expresses the R2 protein with a 6×hi tag. The R2 protein sequence is shown in SEQ ID NO: 3. The nucleotide sequence of the R2 encoding gene is as follows:
[0118] 5'-seq1+cgtgttgttgttgttctgttcaccgcagtgtatactggttctcaggacgcacgtgcaaccgaactggctgcaaaagctctggctatcgcagtttac+seq2-3'.
[0119] The R3 expression vector contains the R3 encoding gene and expresses the R3 protein with a 6×hi tag. The R3 protein sequence is shown in SEQ ID NO: 4. The nucleotide sequence of the R3 encoding gene is as follows:
[0120] 5'-seq1+cgtgtagttgttctgccgctgcgcatggttatgtctgatctggcagatgtttccgatgcagaagctgcagctctggctgcacgtctggctgctgctgtcatggca+seq2-3'.
[0121] The R4 expression vector contains the R4 encoding gene and expresses the R4 protein with a 6×hi tag. The R4 protein sequence is shown in SEQ ID NO: 5. The nucleotide sequence of the R4 encoding gene is as follows:
[0122] 5'-seq1+ttcgttgtagtcgctaacgatctgtttgcgaagatcgttcgtctgctggcggaaatcctgatgaacatgaacgcggttaccgaactgcgccacaccatggttgtactggcggcactgctggcagcaatg+seq2-3'.
[0123] The R5 expression vector contains the R5 encoding gene and expresses the R5 protein with a 6×hi tag. The R5 protein sequence is shown in SEQ ID NO: 6. The nucleotide sequence of the R5 encoding gene is as follows:
[0124] 5'-seq1+tacgttgttgttgtcaacgacacctggtataaaatcgttgaaatgatgatcaaactgctgcaggccagcaacgcggtcgagaccctggaaaaacaggctgcacgtctggctgccctgctgtccctgatg+seq2-3'.
[0125] The R6 expression vector contains the R6 encoding gene and expresses the R6 protein with a 6×hi tag. The R6 protein sequence is shown in SEQ ID NO: 7. The nucleotide sequence of the R6 encoding gene is as follows:
[0126] 5'-seq1+tacgtaatcgtaggtggtctgaaccgtgaagaagctctgcgtgctgctgcactggctgcatatgttgtgcgtgatcgtgacgctgataaagctttcatgaaaatcctgacgtggctggcgctgatg+seq2-3'.
[0127] The R7 expression vector contains the R7 encoding gene and expresses the R7 protein with a 6×hi tag. The R7 protein sequence is shown in SEQ ID NO: 8. The nucleotide sequence of the R7 encoding gene is as follows:
[0128] 5'-seq1+tacgttgtagttatcgccaacctgaaactgctggttaccgctttcgctctggcggctcaggtaatgaactccgataagctggaacgcctggcggccgtactggcggcgatgctggcgctgatg+seq2-3'.
[0129] The R8 expression vector contains the R8 encoding gene and expresses the R8 protein with a 6×hi tag. The R8 protein sequence is shown in SEQ ID NO: 9. The nucleotide sequence of the R8 encoding gene is as follows:
[0130] 5'-seq1+tacgtggttctgctgtccaaccaggaggtaatgcgtgtttacatggaagccctgaaactggcgatgaaagcggcggctaagaacgcgaacaacggtctgtacaccgcgttcgccctggtggcgtacatg+seq2-3'.
[0131] The R9 expression vector contains the R9 encoding gene and expresses the R9 protein with a 6×hi tag. The R9 protein sequence is shown in SEQ ID NO: 10. The nucleotide sequence of the R9 encoding gene is as follows:
[0132] 5'-seq1+tacgttgtggtggtgattaacaaagaactgatcaaaacccagctgcgcaaagacaaccgtgctgacgccgcgctggta+seq2-3'.
[0133] The R10 expression vector contains the R10 encoding gene and expresses the R10 protein with a 6×hi tag. The R10 protein sequence is shown in SEQ ID NO: 11. The nucleotide sequence of the R10 encoding gene is as follows:
[0134] 5'-seq1+cgtgttgttgttgttaagatgggtgcggtttataccggcgatcgtgacgcagatatggcggaactgctgctgcgtgctctggctatcgcagcattc+seq2-3'.
[0135] The PCR system and PCR procedure were the same as in Example 1. The seamless ligation system was the same as in Example 1, except that the linearized vector was replaced with the PCR vector.
[0136] Vector primers: Vf: gatatgctggcgaccgcgc; Vr: cgctttcgggccgcgat.
[0137] Fragment primers:
[0138] R1-f: CCAGGATGCGCGTGCTACGGAACTGGCAGCACTGGCTCTGGCGATCGCGGTTTACgatatgctggcgaccgcgc;
[0139] R1-r: GCACGCGCATCCTGGCTACCAGTGTATACAGCAGTGAACAGTACAACAACAATACGcgctttcgggccgcgat;
[0140] R2-f: TCAGGACGCACGTGCAACCGAACTGGCTGCAAAAGCTCTGGCTATCGCAGTTTACgatatgctggcgaccgcgc;
[0141] R2-r: GCACGGCGTCCTGAGAACCAGTATACACTGCGGTGAACAGAACAACAACAACACGcgctttcgggccgcgat;
[0142] R3-f: ATGTTTCCGATGCAGAAGCTGCAGCTCTGGCTGCACGTCTGGCTGCTGCTGTCATGGCAgatatgctggcgaccgcgc;
[0143] R3-r: CTGCATCGGAAACATCTGCCAGATCAGACATAACCATGCGCAGCGGCAGAACAACTACACGcgctttcgggccgcgat;
[0144] R4-f:CTGATGAACATGAACGCGGTTACCGAACTGCGCCACACCATGGTTGTACTGGCGGCACTGCTGGCAGCAATGgatatgctggcgaccgcgc;
[0145] R4-r:GTTCATGTTCATCAGGATTTCCGCCAGCAGACGAACGATCTTCGCAAACAGATCGTTAGCGACTACAACGAAcgctttcgggccgcgat;
[0146] R5-f:CTGCAGGCCAGCAACGCGGTCGAGACCCTGGAAAAACAGGCTGCACGTCTGGCTGCCCTGCTGTCCCTGATGgatatgctggcgaccgcgc;
[0147] R5-r:GTTGCTGGCCTGCAGCAGTTTGATCATCATTTCAACGATTTTATACCAGGTGTCGTTGACAACAACAACGTAcgctttcgggccgcgat;
[0148] R6-f:TGCATATGTTGTGCGTGATCGTGACGCTGATAAAGCTTTCATGAAAATCCTGACGTGGCTGGCGCTGATGgatatgctggcgaccgcgc;
[0149] R6-r:CGCACAACATATGCAGCCAGTGCAGCAGCACGCAGAGCTTCTTCACGGTTCAGACCACCTACGATTACGTAcgctttcgggccgcgat;
[0150] R7-f:CTCAGGTAATGAACTCCGATAAGCTGGAACGCCTGGCGGCCGTACTGGCGGCGATGCTGGCGCTGATGgatatgctggcgaccgcgc;
[0151] R7-r:AGTTCATTACCTGAGCCGCCAGAGCGAAAGCGGTAACCAGCAGTTTCAGGTTGGCGATAACTACAACGTAcgctttcgggccgcgat;
[0152] R8-f: CTGGCGATGAAAGCGGCGGCTAAGAACGCGAACAACGGTCTGTACACCGCGTTCGCCCTGGTGGCGTACATGgatatgctggcgaccgcgc;
[0153] R8-r: CGCTTTCATCGCCAGTTTCAGGGCTTCCATGTAAACACGCATTACCTCCTGGTTGGACAGCAGAACCACGTAcgctttcgggccgcgat;
[0154] R9-f: CAAAACCCAGCTGCGCAAAGACAACCGTGCTGACGCCGCGCTGGTAgatatgctggcgaccgcgc;
[0155] R9-r: CGCAGCTGGGTTTTGATCAGTTCTTTGTTAATCACCACCACAACGTAcgctttcgggccgcgat;
[0156] R10-f:ATCGTGACGCAGATATGGCGGAACTGCTGCTGCGTGCTCTGGCTATCGCAGCATTCgatatgctggcgaccgcgc;
[0157] R10-r: TATCTGCGTCACGATCGCCGGTATAAACCGCACCCATCTTAACAACAACACGcgctttcgggccgcgat.
[0158] The sequencing vector was used to prepare proteins (R1-R10) according to the method described in Example 1.
[0159] Protein expression varies ( Figure 4 (A represents the SDS-PAGE results of 10 proteins). R1, R4, R5, and R7 showed extremely low expression levels; R2, R3, R6, R8, and R9 showed relatively high expression levels; and R10 did not oligomerize. Therefore, molecular sieve analysis was used to verify the oligomerization of this protein. Figure 4 (where B represents the results of four protein molecular sieves). The results show that the oligomerization states of R3 and R9 are relatively uniform, while the oligomerization states of R2 and R8 are poor. This result also reflects the diversity of protein design results.
[0160] 3. DNA testing capability assessment
[0161] Further characterization of wild-type NPB and nine designed (R1-R9) proteins was conducted.
[0162] The process for recording membrane pore formation, pore current, and DNA perforation signals is as follows: an artificial phospholipid monolayer membrane (diphyidophosphatidylcholine, DPhPC) is formed, and a single nanoporous protein is embedded in it. The current changes are then recorded at a voltage of 180mV.
[0163] 1) The steps for embedding nanopores are as follows:
[0164] Electron signal measurements were obtained from nanopores embedded in the DPhPC phospholipid bilayer in a buffer solution (600 mM KCl, 75 mM K3[Fe(CN)6, 25 mM K4[Fe(CN)6]·3H2O, 100 mM Hepes, pH 8.0). After achieving single-pore insertion into the phospholipid bilayer, 2 mL of buffer solution (600 mM KCl, 75 mM K3[Fe(CN)6, 25 mM K4[Fe(CN)6]·3H2O, 100 mM Hepes, pH 8.0) was passed through the system to remove residual excess nanopores, thus obtaining a single nanopore signal acquisition system. The current signal of the nanopore mutant protein on the phospholipid membrane was recorded.
[0165] 2) The process of preparing the DNA sample to be tested and identifying the DNA detection capability is as follows: After the single nanopore signal acquisition system is constructed, the DNA sample to be tested with the T4 Dda mutant protein assembled, ATP (final concentration 2mM) and MgCl2 (final concentration 10mM) are flowed into the single nanopore experimental system (total volume 100μL) and the signal is measured at a constant voltage of +150mV.
[0166] The DNA sample to be tested was assembled according to the method recorded in patent (WO2014135838A1) using T4Dda mutant protein (T4Dda mutant protein is a helicase, a rate-controlling protein, specifically T4Dda-E94C / C109A / C136A / A360C mutant protein, which is described in US20170283470A1). The DNA sample to be tested was 0.5kb in length (https: / / doi.org / 10.1038 / s41587-020-0570-8; WO 2019002893A1, sequence as shown in SEQ ID: 12), containing 5 repetitive sequences, each containing 10 Ts (10 As on the complementary strand).
[0167] Test results as follows Figure 5 As shown in the figure. The results indicate that the pore current results correspond to the molecular sieve results. The current properties of R3 and R9, which are in a better oligomeric state, are better than those of other designed proteins, both in terms of pore current and permeation signal, and the current properties are significantly better than those of NPB wild-type ( Figure 5The above figure shows WT_NPB, which has a relatively large current of 0.7nA and high noise, but all of them exhibit uneven current distribution. Figure 5 R3 and R9 exhibit two types of pore currents: pore current 1 and pore current 2, suggesting that they may contain pore proteins of different diameters (oligomeric state). Other proteins (R1, R4, R5, R6, R7, R2, and R8) have smaller signals than wild-type NPB, but they all have certain problems such as: high noise, severe spontaneous blocking, small DNA perforation signal amplitude, or significant noise during perforation.
[0168] 4. Cryo-electron microscopy structural analysis of R3 and R9
[0169] To verify the rationality of the protein design and to provide a structural basis for subsequent modifications based on R3 or R9, this embodiment further attempted to resolve the cryo-electron microscopy structures of R3 and R9. Following the method described in Example 2, the samples underwent extensive screening. Data could be collected for R3, but no cryo-data was obtained for R9 due to its poor freezing properties.
[0170] The analytical results of the R3 structure are as follows: Figure 6 As shown, A represents the two-dimensional classification result, and B represents the three-dimensional reconstruction result. Data processing reveals that the two-dimensional classification shows R3 has at least two different channel types with varying diameters. Figure 6 (The results shown in box A); From the 3D reconstruction results, only one structure conforms to the original design (11-mer), but the density of the shrinkage zone is poor; the other one, although the oligomer state has not changed (also 11-mer), is more loose and has a larger overall density diameter.
[0171] In summary, different contraction regions designed based on the NPB wild type can all improve the stability of pore current, with R3 and R9 showing significant improvements in the detection performance of DNA substrates.
[0172] Example 4: Modification of sequencing contraction regions and identification of sequencing properties based on NPB-R3
[0173] Given the inhomogeneity in the cryo-electron microscopy structure and current detection of R3 in Example 3, this example attempts to improve the oligomerization of R3 through amino acid mutation and other methods. Based on the results of Example 3, the density of the R3 contraction region is poor, suggesting significant steric hindrance between individual molecules in the contraction region. Changing the spatial positions between molecules mainly relies on amino acid mutation and deletion. Considering that this situation primarily exists in the contraction region, this example mainly focuses on attempting amino acid mutations or deletions in this region.
[0174] In this embodiment, amino acid mutations or deletions are mainly concentrated at the L74 and D64 positions. For the L74 position, it is mutated to one of the other 19 amino acids and the L74 position is deleted (corresponding numbers P6-7, P8-11, P36-46, Table 3); the D64 position is mutated to P / A / V / T / S (corresponding numbers P12-P16, Table 3).
[0175] Table 3: Statistics on mutations or deletions based on R3
[0176] serial number Corresponding mutation or deletion Corresponding protein SEQ ID NO: P6 L74P 13 P7 del-L74 14 P8-P11 L74V / T / S / I 15,16,17,18 P23-25 L74G / C / D 19,20,21 P36-39 L74A / M / N / Q 22,23,24,25 P40-43 L74E / R / K / H 26,27,28,29 P44-46 L74W / F / Y 30,31,32 P12-P16 D64P / A / V / T / S 33,34,35,36,37
[0177] 1. Preparation of expression vectors
[0178] Expression vectors were constructed using the method described in Example 1. Using the R3 expression vector gene as a template, the target fragment (i.e., the PCR fragment and the PCR vector) was amplified by PCR using corresponding primers (f represents the forward primer, r represents the reverse primer). After gel recovery, the PCR fragment and the PCR vector were ligated, and then transformed into DH5α competent cells for screening of positive clones. Two clones were selected for sequencing. After successful sequencing, the constructed plasmid was stored at -20℃ for later use. P6-P16, P23-P25, and P36-P46 expression vectors were obtained according to the aforementioned method.
[0179] In the following protein-coding gene nucleotide sequences, lowercase letters indicate substitutions compared to the R3-coding gene, "~" indicates deletions compared to the R3-coding gene, and "+" represents a connector.
[0180] Where seq3 represents
[0181] ATGTTTCGCCTGCTGACCCTGCTGGCCGGCGTTCTGTTACTGGCGGTGAGCTGCGTGAGCAGCGGCGTGCAGACCCAAGTGGATACCACCGGCCCGACCGAGCCAAGTGCTGACCTATCGCGGCCCGAAAGCGCGTGTAGTTGTTCTGCCGCTGCGCATGGTTATGTCTGATCTGGCAGATGTTTCCGATGCAGAAGCTGCAGCTCTGGCTGCACGT;
[0182] seq4 represents
[0183] .
[0184] The P6 expression vector contains the P6 encoding gene, which expresses the P6 protein with a 6×his tag. The P6 protein sequence is shown in SEQ ID NO: 13.
[0185] The nucleotide sequence of the P6 encoding gene is as follows: 5'-seq3+ccg+seq4-3'.
[0186] The P7 expression vector contains the P7 encoding gene and expresses the P7 protein with a 6×his tag. The P7 protein sequence is shown in SEQ ID NO: 14.
[0187] The nucleotide sequence of the P7 encoding gene is as follows: 5'-seq3+~~~+seq4-3'.
[0188] The P8 expression vector contains the P8 encoding gene, which expresses the P8 protein with a 6×his tag. The P8 protein sequence is shown in SEQ ID NO: 15.
[0189] The nucleotide sequence of the P8 encoding gene is as follows: 5'-seq3+gtg+seq4-3'.
[0190] The P9 expression vector contains the P9 encoding gene and expresses the P9 protein with a 6×his tag. The P9 protein sequence is shown in SEQ ID NO: 16.
[0191] The nucleotide sequence of the P9 encoding gene is as follows: 5'-seq3+acc+seq4-3'.
[0192] The P10 expression vector contains the P10 encoding gene, which expresses the P10 protein with a 6×his tag. The P10 protein sequence is shown in SEQ ID NO: 17. The nucleotide sequence of the P10 encoding gene is as follows: 5'-seq3+agc+seq4-3'.
[0193] The P11 expression vector contains the P11 encoding gene, which expresses the P11 protein with a 6×hi tag. The P11 protein sequence is shown in SEQ ID NO: 18. The nucleotide sequence of the P11 encoding gene is as follows: 5'-seq3+att+seq4-3'.
[0194] The P23 expression vector contains the P23 encoding gene, which expresses the P23 protein with a 6×his tag. The P23 protein sequence is shown in SEQ ID NO: 19. The nucleotide sequence of the P23 encoding gene is as follows: 5'-seq3+gtt+seq4-3'.
[0195] The P24 expression vector contains the P24 encoding gene, which expresses the P24 protein with a 6×his tag. The P24 protein sequence is shown in SEQ ID NO: 20. The nucleotide sequence of the P24 encoding gene is as follows: 5'-seq3+tgt+seq4-3'.
[0196] The P25 expression vector contains the P25 encoding gene, which expresses the P25 protein with a 6×his tag. The P25 protein sequence is shown in SEQ ID NO: 21. The nucleotide sequence of the P25 encoding gene is as follows: 5'-seq3+gat+seq4-3'.
[0197] The P36 expression vector contains the P36 encoding gene, which expresses the P36 protein with a 6×his tag. The P36 protein sequence is shown in SEQ ID NO: 22. The nucleotide sequence of the P36 encoding gene is as follows: 5'-seq3+gct+seq4-3'.
[0198] The P37 expression vector contains the P37 encoding gene, which expresses the P37 protein with a 6×his tag. The P37 protein sequence is shown in SEQ ID NO: 23. The nucleotide sequence of the P37 encoding gene is as follows: 5'-seq3+atg+seq4-3'.
[0199] The P38 expression vector contains the P38 encoding gene, which expresses the P38 protein with a 6×his tag. The P38 protein sequence is shown in SEQ ID NO: 24. The nucleotide sequence of the P38 encoding gene is as follows: 5'-seq3+aac+seq4-3'.
[0200] The P39 expression vector contains the P39 encoding gene, which expresses the P39 protein with a 6×his tag. The P39 protein sequence is shown in SEQ ID NO: 25. The nucleotide sequence of the P39 encoding gene is as follows: 5'-seq3+cag+seq4-3'.
[0201] The P40 expression vector contains the P40 encoding gene, which expresses the P40 protein with a 6×his tag. The P40 protein sequence is shown in SEQ ID NO: 26. The nucleotide sequence of the P40 encoding gene is as follows: 5'-seq3+gaa+seq4-3'.
[0202] The P41 expression vector contains the P41 encoding gene, which expresses the P41 protein with a 6×hi tag. The P41 protein sequence is shown in SEQ ID NO: 27. The nucleotide sequence of the P41 encoding gene is as follows: 5'-seq3+cgt+seq4-3'.
[0203] The P42 expression vector contains the P42 encoding gene, which expresses the P42 protein with a 6×his tag. The P42 protein sequence is shown in SEQ ID NO: 28. The nucleotide sequence of the P42 encoding gene is as follows: 5'-seq3+aaa+seq4-3'.
[0204] The P43 expression vector contains the P43 encoding gene, which expresses the P43 protein with a 6×his tag. The P43 protein sequence is shown in SEQ ID NO: 29. The nucleotide sequence of the P43 encoding gene is as follows: 5'-seq3+cac+seq4-3'.
[0205] The P44 expression vector contains the P44 encoding gene, which expresses the P44 protein with a 6×his tag. The P44 protein sequence is shown in SEQ ID NO: 30. The nucleotide sequence of the P44 encoding gene is as follows: 5'-seq3+tgg+seq4-3'.
[0206] The P45 expression vector contains the P45 encoding gene, which expresses the P45 protein with a 6×his tag. The P45 protein sequence is shown in SEQ ID NO: 31. The nucleotide sequence of the P45 encoding gene is as follows: 5'-seq3+ttt+seq4-3'.
[0207] The P46 expression vector contains the P46 encoding gene, which expresses the P46 protein with a 6×his tag. The P46 protein sequence is shown in SEQ ID NO: 32. The nucleotide sequence of the P46 encoding gene is as follows: 5'-seq3+tat+seq4-3'.
[0208] Where seq5 represents
[0209] ATGTTTCGCCTGCTGACCCTGCTGGCCGGCGTTCTGTTACTGGCGGTGAGCTGCGTGAGCAGCGGCGTGCAGACCCAAGTGGATACCACCGGCCCGACCGCGAGCCAAGTGCTGACCTATCGCGGCCCGAAAGCGCGTGTAGTTGTTCTGCCGCTGCGCATGGTTATGTCTGATCTGGCAGATGTTTCC;
[0210] seq6 represents
[0211] .
[0212] The P12 expression vector contains the P12 encoding gene, which expresses the P12 protein with a 6×his tag. The P12 protein sequence is shown in SEQ ID NO: 33. The nucleotide sequence of the P12 encoding gene is as follows: 5'-seq5+cct+seq6-3'.
[0213] The P13 expression vector contains the P13 encoding gene, which expresses the P13 protein with a 6×his tag. The P13 protein sequence is shown in SEQ ID NO: 34. The nucleotide sequence of the P13 encoding gene is as follows: 5'-seq5+gca+seq6-3'.
[0214] The P14 expression vector contains the P14 encoding gene, which expresses the P14 protein with a 6×his tag. The P14 protein sequence is shown in SEQ ID NO: 35. The nucleotide sequence of the P14 encoding gene is as follows: 5'-seq5+gtg+seq6-3'.
[0215] The P15 expression vector contains the P15 encoding gene, which expresses the P15 protein with a 6×his tag. The P15 protein sequence is shown in SEQ ID NO: 36. The nucleotide sequence of the P15 encoding gene is as follows: 5'-seq5+acc+seq6-3'.
[0216] The P16 expression vector contains the P16 encoding gene, which expresses the P16 protein with a 6×his tag. The P16 protein sequence is shown in SEQ ID NO: 37. The nucleotide sequence of the P16 encoding gene is as follows: 5'-seq5+agc+seq6-3'.
[0217] The PCR system and PCR procedure were the same as in Example 1. The seamless ligation system was the same as in Example 1, except that the linearized vector was replaced with the PCR vector.
[0218] Vector primers: Vf: gatatgctggcgaccgcgc; Vr: cgctttcgggccgcgat.
[0219] Fragment primers:
[0220] Primers used for P6 and P7:
[0221] P6-f: GCACGTccgGCTGCTGCTGTCATGGCA;
[0222] P6-r: AGCAGCcggACGTGCAGCAGAGCTG;
[0223] P7-f:GCTGCACGTGCTGCTGCTGTCATGGCA;
[0224] P7-r: AGCAGCAGCACGTGCAGCCAGAGCTG;
[0225] For pages 8-11 and 36-46 below, the forward primer is P8-f, and the reverse primer is the corresponding -r primer.
[0226] P8-f:GCTGCTGCTGTCATGGCA;
[0227] P8-r: CATGACAGCAGCAGCcacACGTGCAGCAGAGCTG;
[0228] P9-r: CATGACAGCAGCAGCggtACGTGCAGCAGAGCTG;
[0229] P10-r: CATGACAGCAGCAGCgctACGTGCAGCAGAGCTG;
[0230] P11-r:CATGACAGCAGCAGCaatACGTGCAGCCAGAGCTG;
[0231] P23-r:CATGACAGCAGCAGCaccACGTGCAGCCAGAGCTG;
[0232] P24-r:CATGACAGCAGCAGCacaACGTGCAGCCAGAGCTG;
[0233] P25-r:CATGACAGCAGCAGCatcACGTGCAGCCAGAGCTG;
[0234] P36-r:CATGACAGCAGCAGCagcACGTGCAGCCAGAGCTG;
[0235] P37-r:CATGACAGCAGCAGCcatACGTGCAGCCAGAGCTG;
[0236] P38-r:CATGACAGCAGCAGCgttACGTGCAGCCAGAGCTG;
[0237] P39-r:CATGACAGCAGCAGCctgACGTGCAGCCAGAGCTG;
[0238] P40-r:CATGACAGCAGCAGCttcACGTGCAGCCAGAGCTG;
[0239] P41-r:CATGACAGCAGCAGCacgACGTGCAGCCAGAGCTG;
[0240] P42-r:CATGACAGCAGCAGCtttACGTGCAGCCAGAGCTG;
[0241] P43-r:CATGACAGCAGCAGCgtgACGTGCAGCCAGAGCTG;
[0242] P44-r:CATGACAGCAGCAGCccaACGTGCAGCCAGAGCTG;
[0243] P45-r:CATGACAGCAGCAGCaaaACGTGCAGCCAGAGCTG;
[0244] P46-r: CATGACAGCAGCAGCataACGTGCAGCAGAGCTG;
[0245] For pages 12-16 below, the reverse primer is P12-r, and the forward primer is the corresponding -f primer.
[0246] P12-r: GGAAACATCTGCCAGATCAGACA;
[0247] P12-f: TGTTTCCcctGCAGAAGCTGCAGCTCTG;
[0248] P13-f: TGTTTCCGCAGCAGAAGCTGCAGCTCTG;
[0249] P14-f: TGTTTCCgtgGCAGAAGCTGCAGCTCTG;
[0250] P15-f: TGTTTCCaccGCAGAAGCTGCAGCTCTG;
[0251] P16-f: TGTTTCCagcGCAGAAGCTGCAGCTCTG.
[0252] 2. Protein preparation
[0253] The sequencing vector was prepared according to the method described in Example 1.
[0254] Protein expression status as follows Figure 7 As shown, all proteins were expressed normally, and there were significant changes in the oligomeric state between heating and not heating. The mutant protein formed pores normally.
[0255] 3. Current property detection
[0256] The current properties of these mutant proteins were detected according to the method in Example 3.
[0257] Test results as follows Figure 8 As shown in the figure. The results indicate that after the L74 mutation to V / T / S / I (P8-11), the pore current became uniform, with only one current signal detected. The noise and spontaneous blockage of the pore current were significantly improved, and a clear signal could be recorded for the DNA substrate, with a high substrate DNA capture rate. Therefore, P8-P11 can be basically used for DNA substrate sequencing, with P10 exhibiting relatively less current noise. However, the substrate DNA signal width through the pores in P8-P11 is small, with pore currents between 0.13 nA and 0.21 nA and a signal width less than 50 pA.
[0258] Other mutant proteins also present corresponding problems: For example, the P23-P25 mutations, while allowing normal pore insertion, exhibit uneven pore currents in P23 and P25, containing many noisy pores. This noise frequently accumulates during sequencing and is uncontrollable. The P24 mutation shows generally uniform pore current, but the sequencing process is also unstable, with significant noise. Similarly, the P36-P39 mutations primarily suffer from very low pore opening current and through-pore signal amplitude, similar to the P8-P11 mutations, but with more spontaneous pore blocking. Overall, L74 mutations, except for P8-P11, manifest as either high noise or frequent spontaneous blocking.
[0259] 4. Electron microscopy observation
[0260] The homogeneity of the oligomeric state of P10 was verified by cryo-electron microscopy. Cryo-electron data of P10 were collected using the method described in Example 2. The observation results are as follows: Figure 9 As shown in the figure, A represents the two-dimensional classification result, and B represents the three-dimensional reconstructed structure. The results show that both two-dimensional and three-dimensional classifications indicate that P10 is a homogeneous 11-mer. Therefore, by mutating the designed protein, a more homogeneous pore protein was obtained, significantly improving sequencing performance.
[0261] Example 5: Identification of Sequencing Properties Based on NPB-R-Based Constriction Region Deletion Modification
[0262] The width of the well signal represents the signal-to-noise ratio, and this parameter directly affects the accuracy of sequencing. Given the relatively small width of the substrate DNA well signal in P8-P11 wells (…),… Figure 8 The pore current is between 0.13nA and 0.21nA, and the signal width is less than 50pA. In this embodiment, we mainly attempt to delete amino acids in the P10 (L74S) contraction region. The deletion length is 1-5 amino acids, and the deletion position is before and after S74 (P10). The amino acid types after deletion are shown in Table 4. In the table, "-" represents the deletion of one amino acid. The corresponding numbers of the deleted proteins are P54-P64.
[0263] Table 4: Statistical Table Based on P10 Amino Acid Mutation Types
[0264] serial number Corresponding mutation or deletion P10 <![CDATA[ 66 EAAAAAAR 74 ARRIVALS]]> P54 <![CDATA[ 66 EAAAR 74 S-AAVMAD]]> P55 <![CDATA[ 66 EAAAR 74 S--AVMAD]]> P56 <![CDATA[ 66 EAAAR 74 S---VMAD]]> P57 <![CDATA[ 66 EAAAAAAR 74 S----MAD]]> P58 <![CDATA[ 66 EAAALAAR 74 S-----AD]]> P59 <![CDATA[ 66 EAAAAAAA- 74 SAAAVMAD]]> P60 <![CDATA[ 66 EAAAAAA-- 74 RECEIVABLES<!-- 17 --> ]]> P61 <![CDATA[ 66 EAAAAL--- 74 ARRIVALS]]> P62 <![CDATA[ 66 EAAA---- 74 SAAAVMAD]]> P63 <![CDATA[ 66 EAA----- 74 ARRIVALS]]> P64 <![CDATA[ 66 EA------ 74 SAAAVMAD]]>
[0265] 1. Preparation of expression vectors
[0266] Expression vectors were constructed using the method described in Example 1. Using the P10 expression vector gene as a template, the target fragment (i.e., the PCR fragment and the PCR vector) was amplified by PCR using corresponding primers (f represents the forward primer, r represents the reverse primer). After gel recovery, the PCR fragment and the PCR vector were ligated, and then transformed into DH5α competent cells for screening of positive clones. Two clones were selected for sequencing. After successful sequencing, the constructed plasmid was stored at -20℃ for later use. P6-P16, P23-P25, and P36-P46 expression vectors were obtained according to the aforementioned method.
[0267] In the following protein-coding gene nucleotide sequences, lowercase letters indicate substitutions compared to the P10-coding gene, "~" indicates deletions compared to the P10-coding gene, and "+" represents a connector. Here, seq7 represents...
[0268] ATGTTTCGCCTGCTGACCCTGCTGGCCGGCGTTCTGTTACTGGCGGTGAGCTGCGTGAGCAGCGGCGTGCAGACCCAAGTGGATACCACCGGCCCGACCGCGAGCCAAGTGCTGACCTATCGCGGCCCGAAAGCGCGTGTAGTTGTTCTGCCGCTGCGCATGGTTATGTCTGATCTGGCAGATGTTTCCGATGCAGAAGCT;
[0269] seq8 represents
[0270] .
[0271] The P54 expression vector contains the P54 encoding gene, which expresses the P54 protein with a 6×his tag. The P54 protein sequence is shown in SEQ ID NO: 38. The nucleotide sequence of the P54 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~GCTGCTGTCATG+seq8-3'.
[0272] The P55 expression vector contains the P55 encoding gene, which expresses the P55 protein with a 6×hi tag. The P55 protein sequence is shown in SEQ ID NO: 39. The nucleotide sequence of the P55 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~~~~~GCTGTCATG+seq8-3'.
[0273] The P56 expression vector contains the P56 encoding gene, which expresses the P56 protein with a 6×hi tag. The P56 protein sequence is shown in SEQ ID NO: 40. The nucleotide sequence of the P55 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~~~~~~~GTCATG+seq8-3'.
[0274] The P57 expression vector contains the P57 encoding gene, which expresses the P57 protein with a 6×hi tag. The P57 protein sequence is shown in SEQ ID NO: 41. The nucleotide sequence of the P57 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~~~~~~~~~~ATG+seq8-3'.
[0275] The P58 expression vector contains the P58 encoding gene, which expresses the P58 protein with a 6×hi tag. The P58 protein sequence is shown in SEQ ID NO: 42. The nucleotide sequence of the P58 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCACGTAGC~~~~~~~~~~~~~~~~+seq8-3'.
[0276] The P59 expression vector contains the P59 encoding gene, which expresses the P59 protein with a 6×hi tag. The P59 protein sequence is shown in SEQ ID NO: 43. The nucleotide sequence of the P59 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCTGCA~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0277] The P60 expression vector contains the P60 encoding gene, which expresses the P60 protein with a 6×his tag. The P60 protein sequence is shown in SEQ ID NO: 44. The nucleotide sequence of the P60 encoding gene is as follows: 5'-seq7+GCAGCTCTGGCT~~~~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0278] The P61 expression vector contains the P61 encoding gene, which expresses the P61 protein with a 6×hi tag. The P61 protein sequence is shown in SEQ ID NO: 45. The nucleotide sequence of the P61 encoding gene is as follows: 5'-seq7+GCAGCTCTG~~~~~~~~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0279] The P62 expression vector contains the P62 encoding gene, which expresses the P62 protein with a 6×hi tag. The P62 protein sequence is shown in SEQ ID NO: 46. The nucleotide sequence of the P62 encoding gene is as follows: 5'-seq7+GCAGCT~~~~~~~~~~~~~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0280] The P63 expression vector contains the P63 encoding gene, which expresses the P63 protein with a 6×hi tag. The P63 protein sequence is shown in SEQ ID NO: 47. The nucleotide sequence of the P63 encoding gene is as follows: 5'-seq7+GCA~~~~~~~~~~~~~~~~~~~~~~~~~~AGCGCTGCTGCTGTCATG+seq8-3'.
[0281] The P64 expression vector contains the P64 encoding gene, which expresses the P64 protein with a 6×hi tag. The P64 protein sequence is shown in SEQ ID NO: 48. The nucleotide sequence of the P64 encoding gene is as follows: 5'-seq7+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~seq8-3'.
[0282] The PCR system and PCR procedure were the same as in Example 1. The seamless ligation system was the same as in Example 1, except that the linearized vector was replaced with the PCR vector.
[0283] The mutation primers used for the P54-P64 expression vector are as follows:
[0284] For the expression vectors listed below (pages 54-58), the forward primer is P54-f, and the reverse primer is the corresponding -r primer.
[0285] P54-f:GCTGCTGTCATGGCAgatatgc;
[0286] P54-r: TGCCATGACAGCAGCgctACGTGCAGCCAGAGC;
[0287] P55-r: atcTGCCATGACAGCgctACGTGCAGCAGAGC;
[0288] P56-r: catatcTGCCATGACgctACGTGCAGCAGAGC;
[0289] P57-r: cagcatatcTGCCATgctACGTGCAGCAGAGC;
[0290] P58-r: cgccagcatatcTGCgctACGTGCAGCAGAGC;
[0291] For the expression vectors listed below (pages 59-64), the forward primer is P59-f, and the reverse primer is the corresponding -r primer.
[0292] P59-f: agcGCTGCTGCTGTCAT;
[0293] P59-r: GACAGCAGCAGCgctTGCAGCCAGAGCTGCAGC;
[0294] P60-r: GACAGCAGCAGCgctAGCCAGAGCTGCAGCTTC;
[0295] P61-r: GACAGCAGCAGCgctCAGAGCTGCAGCTTCTGCAT;
[0296] P62-r: GACAGCAGCAGCgctAGCTGCAGCTTCTGCATCG;
[0297] P63-r: GACAGCAGCAGCgctTGCAGCTCTCTGCATCGGAAAC;
[0298] P64-r: GACAGCAGCAGCgctAGCTTCTGCATCGGAAACATCT G.
[0299] 2. Protein preparation
[0300] The sequencing vector was correctly prepared according to the method described in Example 1. Protein expression was as follows. Figure 10 As shown in the figure. The results showed that, except for P56, all other proteins with deleted amino acids were expressed normally, and the oligomeric state changed significantly between heating and not heating, indicating that the mutant protein formed pores normally.
[0301] 3. Current property monitoring: The current properties of mutant proteins P10, P54-P55, and P57-P64 were detected according to the method in Example 3.
[0302] The test results for the via properties of P10, P54, and P61 are as follows: Figure 11 As shown, compared to P10 (P10 pore current is 0.12 nA, DNA perforation signal amplitude is 25 pA), Figure 11 In the P54 and P61 channels, the current and DNA signal amplitude were significantly increased. The P54 channel current was 0.22 nA, and the DNA signal amplitude was 65 pA. Figure 11(B in the text); the current in the P61 well was 0.24 nA, and the DNA perforation signal amplitude was 110 pA ( Figure 11 (C in the text). Therefore, the DNA sequencing performance of P54 and P61 was greatly improved, with P61 showing a more significant improvement.
[0303] Other mutation detection results are as follows Figure 12 As shown, P59 represents the pore current signal of the P59 protein, P63 represents the pore current signal of the P63 protein, P62 represents the pore current and pore current signal of the P62 protein, and P64 represents the pore current and pore current signal of the P64 protein. P59 shows a relatively small pore current signal for complementary DNA, and many signal fluctuations occur during DNA entry and exit from the nanopore. Figure 12 P59, through-hole signal 2); P62 has two types of pore currents, and only under a larger current can the DNA substrate pass through the pore, resulting in a smaller signal amplitude. Figure 12 The via current at the top of P62 is around 50pA; the via efficiency of the substrate at P63 is relatively low. Figure 12 P63); P64 exhibits multiple pore currents, all of which are unstable and can pass through DNA substrates, but the signal amplitude is relatively small. Figure 12 (P64).
[0304] Example 6: Identification of P61 DNA Sequencing Properties
[0305] To further examine the sequencing performance of P61, in this embodiment, the sequencing properties of P61 were compared with those of MspA and ONT product R10.4.1 (MinION). The DNA sequencing properties of P61 and MspA were tested according to the method described in Example 3.
[0306] The DNA sequencing properties of P61 and ONT product R10.4.1 (MinION) were detected according to the method in Example 3. The only difference was that the DNA sample to be tested was a de Bruijn sequence (as shown in SEQ ID: 49).
[0307] Test results as follows Figure 13 As shown in the figure, A represents the sequencing signal graphs of P61 and MspA, B represents the sequencing signal graphs of P61 and ONT R10.4.1, and C represents the horizontal noise of P61 and ONT R10.4.1. The results show that P61 sequencing signals have advantages such as high base resolution, low noise, high signal-to-noise ratio, and good sequencing stability.
[0308] Compared to MspA, a transmembrane pore protein commonly used for DNA sequencing, P61 has a more stable contraction region (reader head) and is more sensitive to DNA substrates; for the same DNA sequence, P61 exhibits a greater electrical signal. Figure 13The figure above shows the MspA current result, and the figure below shows the P61 current result (A in the figure above shows the MspA current result, and P61 current result below shows the MspA current result), indicating that it has a higher ability to identify bases.
[0309] P61 and ONT R10.4.1 were used to detect the de Bruijn sequence (SEQ ID NO: 49) under the same conditions. Analysis showed that P61's performance (signal width, number of steps, etc.) was comparable to ONT R10.4.1. Figure 13 (The image above shows the current results from ONTR10.4.1, and the image below shows the current results from P61.) Furthermore, signal noise has been further reduced. Figure 13 The lower noise level of P61 (as seen in the C-type sequencing) indicates a higher signal-to-noise ratio. Furthermore, the capture efficiency and sequencing stability of P61 both meet the standards for long-duration sequencing, satisfying the demands of high-throughput sequencing.
[0310] In summary, the overall properties of P61 (such as current stability, substrate capture efficiency, signal amplitude, etc.) are already suitable for direct use in DNA sequencing.
[0311] In this embodiment, the MspA nanopores are the mutants (D90N, D91N, D93N, D118R, D134R, and D139K) described in the literature (https: / / doi.org / 10.1016 / j.ymeth.2016.03.026), and their structural schematic diagrams are shown in the literature. Figure 2 As shown), the hole embedding method is the same as in Example 3; R10.4.1 is the purchase of ONT commercial products.
[0312] Example 7: Design and DNA sequencing identification of P61-based double contraction zone pore protein
[0313] 1. Design the contraction zone
[0314] The binder design procedure in RFdiffusion software was used, with the binder length set between 25 and 40 amino acid residues. After multiple rounds of optimization, the P61 binder length was selected as 34 amino acids (with a wider contact surface and including secondary structures such as α-helices); then ProteinMPNN was used to generate the corresponding peptide (Table 5).
[0315] Table 5: Statistical table of peptide sequences interacting with P61
[0316] serial number polypeptide amino acid types SEQ ID NO: Peptide 1 GEKKKVSESAEKGGNKENEKKLKEEKKAKDKTKK 51 Peptide 2 GPVVPVSLSAAAGGIAANQALLDALAAANDTTTE 52 Polypeptide 3 GPVVPVSLSGLSGGILANQAILTALLNALNTTTG 53 Peptide 4 GPVTPVSPSAAAGGRAENAAALAAAAAAADTTTA 54 Peptide 5 GPVVPVALSALAGGVAANLPALVAAAAAANATTE 55 Peptide 6 GPVVPVSLLGLLGGNLANLPLLVALLAALNRTTE 56 Polypeptide 7 GKVVPVSESAEEGGRKENQEELEKRKKEEDKTKE 57 Polypeptide 8 GEVVEESESAEEGGNKENQEKLVKEKAENDKTTK 58 Polypeptide 9 GEVVPVSESAADGGREENQAALEAAAAAADTRTA 59 Peptide 10 GEVKPVSKSAKDGGNKENQAKIDAEKAANDKTKE 60 Peptide 11 GEVVPVSLSGEKGGILVNQPILDLLKKLLDKTTK 61 Peptide 12 GKVVKVEESAEKGGKKENEKKLKEEKKKNDKTKE 62 Polypeptide 13 GPVVPVSLSALAGGVAANLPLLVALAAALDRRTA 63 Polypeptide 14 GPVVPVSVLAKEGGNAANQAIIDAAEKAADKTTK 64 Peptide 15 GEVVPVSLSGKEGGIEANQALLDALAAANDKTTK 65 Peptide 16 GKVVPKSLSAKDGGNKENQAVLDALAAANDKTTE 66
[0317] 2. Assembly of pore proteins in the double contraction zone
[0318] The aforementioned peptides 1-16 (95% purity) were synthesized and dissolved in solutions containing detergent (20 mM Tris-HCl, 150 mM NaCl, pH 8.0, 0.06% LDAO; peptides that could not be directly dissolved were first dissolved in 100% DMSO and then diluted in this solution). The P61 protein prepared in Example 5 and peptides 1-16 were incubated overnight at a molar ratio of 1:5 (P61:peptide). Excess peptides were then removed using molecular sieves to obtain the double-contraction pore protein complexes P61D1-P61D16, which are complexes formed by the binding of P61 protein with peptides 1-16, for example, P61D13 is a complex formed by the binding of P61 protein with peptide 13.
[0319] 3. Current property detection
[0320] The current properties of the double contraction zone pore protein complex P61D1-P61D16 were detected according to the method in Example 3. First, it was verified whether the peptide could stably bind to P61, and then DNA samples containing oligopolynucleotides (10 T, 5 replicates; complementary strand 10 A, 5 replicates) were tested. Figure 14 The image shows partial pore current detection results. The pore current of P61D13 decreases and stabilizes, indicating that D13 can stably bind to P61. For the binding of other peptides, the following situations exist: 1. The peptide cannot form a stable complex with P61, which is manifested by no observation of the effect of peptide binding on the pore current after embedding. The reason for non-binding may be that the peptide does not bind to P61, or the binding force is weak, causing the peptide to separate during the embedding process. Most peptides fall into this category; 2. After peptide binding, a decrease in pore current can be observed (indicating that the peptide has bound to P61). Figure 15 However, the current is unstable. Compared to P61D3, the via current exhibits more spikes (…). Figure 15 P61D2), spontaneous blockage is severe ( Figure 15 P61D9, P61D11, and P61D15), and the dissociation of peptides under an electric field (P61D9, P61D11, and P61D15). Figure 15 (The black box on page 61D14 shows the dissociation process).
[0321] Figure 15The results are shown in Figure A, where A represents the binding diagram of P61 and P61D13. The binding mechanism for other peptides is similar to that of P61D13. Figure B shows the permeation signal current maps; the upper map shows the permeation signal current map for P61, and the lower map shows the permeation signal current map for P61D13. As shown, peptide D13 binds (i.e., connects) to the β-sheet of the P61 monomer (indicated by the arrows in the figure, positions 143-169 and 198-223 of sequence SEQ ID NO:45). Comparing the permeation signal current maps of P61 and P61D13, P61D13 shows significantly more signal information (the part highlighted in the box in the figure, with the arrow indicating a magnified view). The results show that after P61 binds to peptide 13, the pore current decreases from 0.24 nA to 0.15 nA, and the resolution of oligopolynucleotides (10 nucleotides) by the dual-contraction nanopore is significantly improved. Therefore, P61D13 can meet the requirements for the detection of oligopolynucleotides.
[0322] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims.
Claims
1. A nanoporous protein complex with dual contraction zones, characterized in that, The nanoporin complex comprises a mutant NPB nanoporin and a conjugate, wherein the conjugate links the mutant NPB nanoporin and forms a contractile region within the mutant NPB nanoporin, and the conjugate is a polypeptide with the amino acid sequence shown in SEQ ID NO:
63. The mutant NPB nanoporous protein contains any of the following polypeptides: (a1) A polypeptide with the amino acid sequence shown in SEQ ID NO: 4; (a2) A polypeptide having the same function by substituting and / or deleting and / or adding one or more amino acids of the amino acid sequence shown in SEQ ID NO: 4; (a3) is a polypeptide that has more than 80% identity with any of the amino acid sequences defined in (a1)-(a2) and has the same function; (a4) A fusion polypeptide obtained by attaching a tag to the end of any of the polypeptides defined in (a1)-(a3).
2. The dual-contraction zone nanoporous protein complex according to claim 1, characterized in that, (a2) The polypeptide comprises at least one of the following substitutions: The leucine at position 74 is replaced by proline, valine, threonine, serine, isoleucine, glycine, cysteine, aspartic acid, alanine, methionine, asparagine, glutamine, glutamic acid, arginine, lysine, histidine, tryptophan, tyrosine, or phenylalanine. The 64th position of aspartic acid is replaced by proline, alanine, valine, threonine, or serine. Preferably, the polypeptide described in (a2) further comprises the deletion of at least one or more of the following amino acids: SEQ ID NO: 4 Delete one, two, three, four or five amino acids at positions 74-79; and / or SEQ ID NO: 4 may contain one, two, three, four, five, six, or seven amino acids deleted at positions 67-73. Preferably, the conjugate is inserted into the cavity of the mutant NPB nanoporous protein. Preferably, the nanoporin comprises at least two linked mutant nanoporin monomers.
3. A method for generating the dual-contraction zone nanoporous protein complex of claim 1 or 2, characterized in that, The method includes the following A1 or A2: A1 enables the co-expression of one or more mutant NPB nanoporin monomers and the conjugate in host cells, thereby allowing the formation of the dual-contraction zone nanoporin complex in the cells. A2 contacts one or more mutant NPB nanoporin monomers with the conjugate, thereby allowing the formation of the dual-contraction nanoporin complex in vitro.
4. The biomaterials related to the dual-contraction zone nanoporous protein complex according to claim 1 or 2, characterized in that, The relevant biomaterial is any one of the following: c1) A nucleic acid molecule encoding the dual-contraction zone nanoporous protein complex of claim 1 or 2; c2) An expression cassette containing the nucleic acid molecule described in c1); c 3) A recombinant vector containing the nucleic acid molecule described in c1), or a recombinant vector containing the expression cassette described in c 2); c 4) Recombinant cells containing the nucleic acid molecule described in c1), or recombinant cells containing the expression cassette described in c2), or recombinant cells containing the recombinant vector described in c3).
5. A method for determining the presence, absence, or one or more characteristics of a target analyte, characterized in that, The method includes: A. The target analyte comes into contact with the dual-contraction zone nanoporous protein complex according to claim 1 or 2, causing the target analyte to move relative to the dual-contraction zone nanoporous protein complex; B. Acquire one or more measurements as the target analyte moves relative to the dual-contraction zone nanoporous protein complex to determine the presence, absence, or one or more characteristics of the target analyte; Preferably, the target analyte is one or more of nucleotides, nucleic acids, amino acids, oligopeptides, polypeptides, and proteins.
6. A kit or device for determining the presence, absence, or one or more characteristics of a target analyte, characterized in that, The kit comprises the dual-contraction zone nanoporous protein complex of claim 1 or 2 or the related biomaterial of claim 4, and a membrane; The device comprises the dual-contraction zone nanoporous protein complex of claim 1 or 2, and a membrane; Preferably, the target analyte is one or more of nucleotides, nucleic acids, amino acids, oligopeptides, polypeptides, and proteins.
7. The combination as described in claim 1.
8. The biomaterial related to the complex of claim 7, characterized in that, The relevant biomaterial is any one of the following: d1) Encoding the nucleic acid molecule of the conjugate as described in claim 7; d2) An expression cassette containing the nucleic acid molecule described in d1); d3) A recombinant vector containing the nucleic acid molecule described in d1), or a recombinant vector containing the expression cassette described in d2); d4) Recombinant cells containing the nucleic acid molecules described in d1), or recombinant cells containing the expression cassette described in d2), or recombinant cells containing the recombinant vector described in d3).
9. The application of the complex of claim 7 or the related biomaterial of claim 8 in the preparation of nanoporous proteins.
10. The use of the dual-contraction zone nanoporous protein complex of claim 1 or 2, the related biomaterial of claim 4, the conjugate of claim 8, or the related biomaterial of claim 9 in detecting the presence, absence, or one or more features of a target analyte or in preparing a product for detecting the presence, absence, or one or more features of a target analyte; Preferably, the target analyte is one or more of nucleotides, nucleic acids, amino acids, oligopeptides, polypeptides, and proteins.
Citation Information
Patent Citations
Mutant csgg pores
US20170283470A1
Enzyme stalling method
WO2014135838A1
Novel protein pores
WO2019002893A1
Mutant pore
CN108779170A
Double-portal pore protein, pore protein mutant, nucleotide sequence and application of double-portal pore protein and pore protein mutant
CN115974984A