Nanopore-nucleic acid polymerase two-site coupling
By introducing dual-site coupling functional groups onto nanopores and nucleic acid polymerases, the problem of low coupling efficiency between nanopores and nucleic acid polymerases was solved, thereby improving the efficiency and accuracy of sequencing signal capture.
Patent Information
- Application Number
- PCT/CN2025/113621
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-08-08
- Publication Date
- 2026-02-12
AI Technical Summary
In existing technologies, the coupling efficiency between nanopores and nucleic acid polymerases is low, and the conformational freedom of the coupled complex is high, leading to a decrease in sequencing efficiency and accuracy.
By introducing two sets of specific coupling functional groups onto nanopores and nucleic acid polymerases, two-site coupling is formed, including click chemistry, antigen-antibody binding, ligand-receptor binding, and spytag-spycatcher binding, thereby controlling the relative distance and orientation between protein molecules.
This technology enables efficient coupling of nanopores and nucleic acid polymerases at low concentrations over a short time, reducing conformational changes and improving sequencing signal capture efficiency and sequencing accuracy.
Smart Images

Figure CN2025113621_12022026_PF_FP_ABST
Abstract
Description
Dual-site coupling of nanopore-nucleic acid polymerase
[0001] Cross-reference to Related Applications
[0002] This application claims priority to Chinese Patent Application No. 202411095705.6, filed on August 9, 2024, to the Chinese Patent Office, the contents of which are hereby incorporated by reference in their entirety. TECHNICAL FIELD
[0003] The present invention belongs to the field of protein chemistry / biochemistry. Specifically, the present invention introduces two sets of functional groups that can be coupled efficiently and spontaneously under mild conditions at selected sites on protein molecule A (nanopore) and protein molecule B (nucleic acid polymerase) through various biological or chemical means, thereby achieving dual coupling between protein molecules based on two independent sites. BACKGROUND
[0004] The concept of nanopore-based nucleic acid sequencing was proposed in 1995 and has gradually matured over the past 30 years. Oxford Nanopore Technologies, Inc. in the United Kingdom has already launched a commercial nanopore sequencer. In terms of technical route, Oxford Nanopore adopts the so-called "direct sequencing method", which involves introducing the nucleic acid chain molecule to be tested into the nanopore under the action of an electric field and passing it at a certain rate, detecting the current changes caused by different bases on the nucleic acid chain during this process to deduce the base sequence of the nucleic acid to be tested, and achieving the purpose of sequencing. In addition to the above-mentioned "direct sequencing method", some other teams around the world, including Sequencing Solutions, Roche Inc., Axbio Inc., and Geneus Technologies, Ltd., adopt a different technical route, which involves combining nanopore detection with the widely adopted "sequencing by synthesis" scheme to achieve the purpose of sequencing. Compared with "direct sequencing", the sequencing by synthesis scheme is more conducive to interpreting homopolymer sequences in complex genomes, thus being more conducive to accurate detection of human genome sequences and promoting the application of nanopore sequencing technology in the clinical market.
[0005] Coupling of nanopore and nucleic acid polymerase is an important prerequisite for the implementation of "sequencing by synthesis". In the field of protein chemistry, the simplest way to achieve the coupling between proteins is to directly fuse and express by genetic engineering. However, since the functional nanopore protein is usually expressed in the form of monomer, and the functional polymer is formed by subsequent stimulation of specific biochemical conditions. Such biochemical conditions are not friendly to polymerase, which usually leads to enzyme inactivation during this process. Therefore, the coupling of nanopore and nucleic acid polymerase is usually difficult to achieve by direct fusion, and the preferred way is to purify and prepare them separately and then couple them together by subsequent biochemical means. So far, the industry has adopted the "spytag-spycacther" system to couple the nanopore and the polymerase. In detail, a special polypeptide (spytag) is fused to the C-terminal of the nanopore protein by genetic engineering, and a small protein spycathcer is expressed with the polymerase. Since spytag and spycachter can spontaneously form a covalent bond under mild conditions, the above system can help the independently expressed and purified nanopore protein and the nucleic acid polymerase to achieve coupling.
[0006] Although the nanopore and polymerase coupling scheme based on Spytag-Spycatcher is proven to be effective and can better support the use of nanopores to achieve "sequencing by synthesis", the scheme itself still has some defects and needs to be further optimized. First, due to the speed limitation of the specific chemical reaction itself, the coupling of spytag-Spycatcher usually needs to be completed after several hours at uM concentration; second, even under the condition of overnight coupling at a higher concentration, there are still a small amount (~ 5%) of nanopores that are difficult to couple with polymerase. If these small amounts of nanopores are not effectively removed, they will often be preferentially embedded in the phospholipid membrane during the subsequent embedding process, resulting in a large number of invalid sequencing units, affecting sequencing efficiency and throughput; finally, the coupling scheme based on spytag-spycatcher has only one coupling site between the nanopore and the polymerase. Due to the inherent activity of protein, polypeptide connection part and other structures, the distance between the nanopore and the polymerase, the orientation, etc. have a high degree of freedom, and the entire complex may be in a conformation that is not conducive to the subsequent effective capture of nucleotide signals by the nanopore, resulting in signal leakage and an increase in sequencing error rate. SUMMARY
[0007] In one aspect, provided herein is a protein comprising two polypeptide fragments, comprising a first conjugation site and a second conjugation site between the two polypeptide fragments, wherein at least one of the conjugation sites is formed by any one of the following:
[0008] 1) fusion protein;
[0009] 2) click chemistry;
[0010] 3) antigen-antibody binding;
[0011] 4) ligand-receptor binding; and
[0012] 5) spytag-spycatcher binding.
[0013] In some embodiments, the two conjugation sites are independently formed by any one of the following:
[0014] 1) protein fusion;
[0015] 2) click chemistry;
[0016] 3) antigen-antibody binding;
[0017] 4) ligand-receptor binding; and
[0018] 5) spytag-spycatcher binding.
[0019] In some embodiments, the two polypeptide fragments are from a DNA polymerase and a nanopore protein, respectively.
[0020] In some embodiments, the nanopore protein is an alpha-hemolysin or a mutant thereof; preferably, the alpha-hemolysin comprises the amino acid sequence set forth in SEQ ID NO: 6, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, or 98% sequence identity to the amino acid sequence set forth in SEQ ID NO: 6.
[0021] In some embodiments, the nanopore protein is a NetB protein or a mutant thereof; preferably, the NetB protein or a mutant thereof comprises the amino acid sequence set forth in SEQ ID NO: 25 or 26, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, or 98% sequence identity to the amino acid sequence set forth in SEQ ID NO: 25 or 26.
[0022] In some embodiments, the DNA polymerase is a Phi29 nucleic acid polymerase or a mutant thereof; preferably, the Phi29 nucleic acid polymerase comprises an amino acid sequence as set forth in SEQ ID NO: 22, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 98% sequence identity to the amino acid sequence as set forth in SEQ ID NO: 22.
[0023] In some embodiments, the antigen-antibody binding is LM7-CL7 protein binding, and / or the ligand-receptor binding is biotin-avidin binding.
[0024] In some embodiments, the avidin is fused to a loop region of the DNA polymerase.
[0025] In some embodiments, the first coupling site is formed by a click chemistry reaction, and the second site is formed by a spytag-spycatcher binding; preferably, the Spycatcher protein is fused to the C-terminus of the DNA polymerase, and the spytag sequence is fused to the C-terminus of at least one monomer of the alpha-hemolysin and forms a click chemistry coupling site between N17 of the monomer and K143 of the DNA polymerase; more preferably, the DNA polymerase comprises a sequence as set forth in SEQ ID NO: 17, and the monomer comprises a sequence as set forth in SEQ ID NO: 7.
[0026] In some embodiments, the first coupling site is formed by a spytag-spycatcher binding, and the second site is formed by a LM7-CL7 protein pair binding; preferably, the Spycatcher protein is fused to the N-terminus of the DNA polymerase, and the I167-P171 sequence of the DNA polymerase is replaced by a GSG-CL7-SG sequence; the spytag sequence is fused to the C-terminus of at least one monomer of the alpha-hemolysin, and the 1A-T19 sequence of the monomer is replaced by a LM7-GSGT sequence; more preferably, the DNA polymerase comprises a sequence as set forth in SEQ ID NO: 18, and the monomer comprises a sequence as set forth in SEQ ID NO: 19.
[0027] In some embodiments, the first coupling site is formed by a click chemistry reaction, the second site is formed by biotin-avidin binding; preferably, the Spycatcher protein is fused to the C-terminus of the DNA polymerase, the at least one monomer of the alpha-hemolysin includes an E31C mutation for modification of biotin; and a click chemistry reaction binding is formed between Y106 of the DNA polymerase and N17 of the monomer; more preferably, the DNA polymerase includes the sequence set forth in SEQ ID NO: 20, and the monomer includes the sequence set forth in SEQ ID NO: 21.
[0028] In some embodiments, the first coupling site is formed by biotin-avidin binding, the second site is formed by LM7-CL7 protein pair binding; preferably, the avidin is fused to the N-terminus or C-terminus of the DNA polymerase, the CL7 protein is correspondingly fused to the C-terminus or N-terminus of the DNA polymerase; the LM7 protein is fused to the C-terminus of the at least one monomer of the alpha-hemolysin, and an N17C mutation is introduced in the monomer for modification of biotin; more preferably, the DNA polymerase fused with the avidin and the CL7 protein includes the sequence set forth in SEQ ID NO: 14 or 15; and the monomer includes the sequence set forth in SEQ ID NO: 16.
[0029] In some embodiments, the first coupling site is formed by spytag-spycatcher binding, the second site is formed by biotin-avidin binding; preferably, the Spycatcher protein is fused to the N-terminus or C-terminus of the DNA polymerase, the avidin is fused between Y148-H149, D160-Y161, D154-Q155, L51-Q52, or K20-K21 of the DNA polymerase; the spytag sequence is fused to the C-terminus of the at least one monomer of the alpha-hemolysin, and an N17C mutation is introduced in the monomer for modification of biotin; more preferably, the DNA polymerase fused with the Spycatcher protein and the avidin includes the sequence set forth in SEQ ID NO: 8, 10, 11, 12, or 13; and the monomer includes the sequence set forth in SEQ ID NO: 7.
[0030] In another aspect, provided herein are methods of making a protein comprising two polypeptide fragments, comprising forming a first coupling site and a second coupling site between the two polypeptide fragments, wherein at least one coupling site is formed by any of:
[0031] 1) protein fusion;
[0032] 2) click chemistry;
[0033] 3) antigen-antibody binding;
[0034] 4) ligand-receptor binding; and
[0035] 5) spytag-spycatcher binding.
[0036] In some embodiments, the two coupling sites are independently formed by any one of the following:
[0037] 1) protein fusion;
[0038] 2) click chemistry;
[0039] 3) antigen-antibody binding;
[0040] 4) ligand-receptor binding; and
[0041] 5) spytag-spycatcher binding.
[0042] In some embodiments, the two polypeptide fragments are from a DNA polymerase and a nanopore protein, respectively.
[0043] In another aspect, provided herein is use of the protein described above or the protein prepared by the method described above in nucleic acid sequencing.
[0044] In some embodiments, the sequencing is by a sequencing-by-synthesis approach.
[0045] In another aspect, provided herein is use of a DNA polymerase as a chaperone in expression of a protein of interest.
[0046] In some embodiments, the protein of interest is expressed in bacteria, and the DNA polymerase as a chaperone prevents the protein of interest from forming inclusion bodies.
[0047] [Corrected according to Rule 26 20.08.2025] In some embodiments, the bacteria is E. coli.
[0048] In some embodiments, the protein of interest is avidin.
[0049] In some embodiments, the avidin is fused to the C-terminus, N-terminus, or Loop region of the DNA polymerase.
[0050] In some embodiments, the DNA polymerase is a Phi29 nucleic acid polymerase or a mutant thereof; preferably, the Phi29 nucleic acid polymerase comprises an amino acid sequence as set forth in SEQ ID NO: 22, or an amino acid sequence having at least 80%, 85%, 90%, 95%, or 98% sequence identity to the amino acid sequence as set forth in SEQ ID NO: 22. BRIEF DESCRIPTION OF DRAWINGS
[0051] FIG. 1 is a schematic diagram of a template DNA for detecting polymerase activity.
[0052] FIG. 2 shows the results of purifying proteins P0065, P0072, and P0073 by gel filtration chromatography column.
[0053] FIG. 3 shows the results of detecting polymerase activity by a microplate reader.
[0054] FIG. 4 shows the results of a biotin function test of P0065.
[0055] FIG. 5 shows SDS-PAGE electrophoresis results of L0606, G0518, and a coupling sample.
[0056] FIG. 6 shows the results of purifying 1:6 heptamer nanopores of G0284 and G0037 by RESOURSE S chromatography.
[0057] FIG. 7 shows SDS-PAGE electrophoresis results of heptamer proteins.
[0058] FIG. 8 shows SDS-PAGE electrophoresis results of a comparison of coupling efficiency of single coupling and double coupling.
[0059] FIG. 9 is a schematic diagram of a sequencing principle model. (A) Model of a sequencing device; (B) Model of sequencing signal detection.
[0060] FIG. 10 shows the results of single coupling sequencing test data.
[0061] FIG. 11 shows the results of double coupling sequencing test data.
[0062] FIG. 12 is a schematic diagram of a nanopore and a polymerase forming a double coupling through biotin-avidin and SpyCatcher-spytag.
[0063] FIG. 13 shows the results of sequencing data based on a CL7-LM7 and biotin-avidin mediated double coupling scheme.
[0064] FIG. 14 shows the results of sequencing data based on a click chemistry reaction group N3-DBCO and SpyCatcher-Spytag mediated double coupling scheme.
[0065] Figure 15 shows sequencing data results based on a Spycatcher-Spytag and CL7-LM7 mediated double coupling protocol.
[0066] Figure 16 shows sequencing data results based on a click chemistry reaction group N3-DBCO and biotin-avidin mediated double coupling protocol. DETAILED DESCRIPTION
[0067] Unless otherwise indicated, all technical and scientific terms used herein have the meanings that are commonly understood by one of ordinary skill in the art.
[0068] The terms "comprising" or "including," or "having" when used in this document mean "including but not limited to," and are not intended to (and do not) exclude other moieties, integers or steps. Thus, particularly, when used in this specification the terms "comprising" or "including" or "having" have the same meaning.
[0069] The term "or" as used in this document means any singular, any two, any three, any more, or all of the listed items. The term "and / or" as used in this document means any singular, any two, any three, any more of the listed items, or any two, any three, any more of the listed items.
[0070] The term "polypeptide" is used interchangeably with "protein" and refers to a biomolecule composed of amino acids linked by peptide bonds. When referring to "polypeptide fragments," it is intended to indicate that they are part of a larger protein molecule, which can itself be a protein or even a fusion protein. These polypeptide fragments can be linked by one or more coupling sites. The term "coupling site" refers to a connection between two polypeptide fragments at a particular location, such as at a particular amino acid residue, which can be covalent or non-covalent. The presence of this connection can generally define the relative distance and / or orientation between the two polypeptide fragments. In the proteins provided herein, including DNA polymerases and nanopore proteins, two coupling sites can be formed between the DNA polymerase and the nanopore protein, thereby enabling a better definition of their positional relationship and orientation.
[0071] In the present application, the term "sequence identity", also known as "sequence identity", refers to the percentage of nucleotides / amino acid residues in a sequence that are identical to those in a reference sequence after sequence alignment, if necessary, introducing spaces in the sequence alignment to achieve the maximum percentage of sequence identity between the two sequences. Those skilled in the art can determine the percentage of sequence identity between two or more nucleic acid or amino acid sequences by various methods, for example, using computer software, such as Clustal Omega, T-coffee, Kalign and MAFFT, etc.
[0072] The present application relates to a coupling method between a nanopore protein and a nucleic acid polymerase. By a series of chemical and biological means, two sets of different coupling reaction functional groups are introduced at specific sites on the nanopore protein and the nucleic acid polymerase, respectively, to support efficient coupling of the corresponding protein molecules at sub-nM concentration in a short time. At the same time, since the coupling method is based on double-site anchoring, the conformational freedom between molecules is effectively reduced, so that the relative distance and orientation between the nanopore opening and the polymerase active center can be better controlled. Compared with the common single coupling technology, the double-site coupling technology based on the present application is more conducive to real-time monitoring of the chemical reactions catalyzed by the nanopore to the polymerase, and has application prospects in nanopore sequencers using the "sequencing by synthesis" scheme. Due to the universality of the present application, the above-mentioned double-site coupling technology can be further extended to the coupling between other protein molecules, and widely applied to other scenarios that require control of the relative distance and orientation between coupled molecules.
[0073] The present application proposes a method of introducing a second pair of specific coupling functional groups to improve coupling efficiency and effectively control the orientation and distance between two protein molecules on the basis of the conventional technology of coupling two protein molecules with a pair of matching functional groups, thereby better performing specific biochemical functions. In the embodiments of the present application, the introduction of the second coupling site enables the nanopore and the nucleic acid polymerase to complete efficient coupling at low concentration in a short time (including promoting the coupling rate of the first coupling site), and by adjusting the distance and orientation between the polymerase active center and the nanopore opening, the nanopore-polymerase complex based on double coupling is more suitable for single molecule sequencing systems, achieving higher data throughput and sequencing accuracy.
[0074] In some embodiments, the first coupling site between the nanopore and the polymerase can be achieved based on a matured spytag-spycatcher pair. By simple fusion expression, a spytag short peptide is introduced at the C-terminus of the nanopore alpha hemolysin, and a spycatcher domain is introduced at the C or N-terminus of the polymerase phi29 or its mutants or homologous protein mutants. Then the nanopore protein and the polymerase will form a spontaneous stable covalent bond coupling based on the interaction of spytag-spycatcher under mild conditions.
[0075] In some embodiments, the first coupling site can also be achieved by selecting a click chemistry scheme. By existing biochemical means, a Cys mutation is introduced at the characteristic site of the nanopore alpha bacterial toxin (alpha hemolysin). Then a DBCO (dibenzocyclooctyne) group commonly used in click chemistry is introduced on the Cys through the mediation of maleimide. At the same time, another click chemistry active group N3 (azido) group is introduced on the nucleic acid polymerase through unnatural amino acid-mediated means. Then under mild conditions, the paired click chemistry active groups can spontaneously form a stable covalent bond to achieve covalent coupling of the nanopore and the polymerase.
[0076] In some embodiments, the coupling between the nanopore and the polymerase can also be based on non-covalent means. High-affinity antigen-antibody interactions, such as the antigen-antibody pair of CL7 and LM7, can be used as an effective coupling method to support the formation of a relatively stable complex between two protein molecules. By fusion expression, a nanopore protein with LM7 at the C-terminus and a polymerase with CL7 (N-terminus or C-terminus) can be conveniently obtained. More importantly, this protein obtained by fusion expression exhibits normal biochemical functions consistent with separate expression for all independent domain entities, such as the nanopore, the polymerase, the CL7 and LM antigen-antibody pair, as described in detail in Example 2 below. After obtaining the nanopore with LM and the polymerase with CL, the high-affinity CL7-LM7 interaction can help the nanopore and the polymerase to be conveniently coupled non-covalently, as one of the optional embodiments of the present application.
[0077] In some embodiments, the coupling between the nanopore and the polymerase can also be achieved based on the tight binding between biotin and avidin. Biotin and tetrameric protein against biotin (streptavidin) have extremely high affinity, which is considered as a nearly "semi-covalent" interaction. In order to facilitate the 1:1 coupling between the nanopore protein and the polymerase, in some embodiments, the selected monomeric protein against biotin is Monomeric Rhizavidin (hereinafter abbreviated as MR). Although the affinity between the monomeric protein against biotin and biotin is reduced compared with the complete tetramer of streptavidin, it is still sufficient to support efficient coupling at sub-nM concentration and stable maintenance of the complex.
[0078] In some embodiments, the monomeric protein against biotin is directly expressed in a soluble manner after being fused with the nucleic acid polymerase and exhibits normal function of binding biotin. In a conventional E. coli-based recombinant protein expression system, the monomeric protein against biotin is expressed in the form of insoluble inclusion bodies. If a functional protein is to be obtained, a cumbersome step of dissolving the inclusion bodies by a denaturant, purification and then renaturation is required. The present application proposes a way of directly expressing the MR protein against biotin by fusion with the nucleic acid polymerase, using the polymerase molecule which is easy to fold and has high expression as a molecular chaperone to assist the folding of the MR protein against biotin, so as to directly express the soluble protein against biotin with normal biochemical function in the form of a fusion protein. In some embodiments, the protein against biotin can be fused at the N-terminus of the nucleic acid polymerase; in other embodiments, the protein against biotin can be fused at the C-terminus of the nucleic acid polymerase. Since the N-terminus and the C-terminus of the monomeric protein against biotin MR are close in spatial structure, in some embodiments, the MR can also be directly fused in the loop of the nucleic acid polymerase. Here, the loop refers to the secondary structure connecting other secondary structures such as alpha helix and beta sheet, which usually has great flexibility. The present application proves that for the monomeric protein against biotin MR, the nucleic acid polymerase is a good molecular chaperone for assisting folding. In all the above embodiments, the monomeric protein against biotin MR which is difficult to be expressed in a soluble manner can achieve correct folding after being fused with the nucleic acid polymerase, and the expressed fusion protein of the polymerase and the protein against biotin is not only soluble, but also the two respective biological functions of the polymerase and the protein against biotin are well maintained.
[0079] In some embodiments, the nanopore and the nucleic acid polymerase can form a stable, oriented complex through a double-coupling approach based on click chemistry and spytag-spycather covalent linkage. In a preferred embodiment, the C-terminus of the nanopore generates a spytag sequence through fusion expression and introduces a Cys mutation at a specific position near the nanopore opening. Then, through the mediation of maleimide, an azido group DBCO (dibenzocyclooctyne) commonly used in click chemistry is introduced on Cys. At the same time, the nucleic acid polymerase generates a spycather through fusion expression at its C-terminus, and introduces a click chemistry functional group N3 near the active center of the polymerase through unnatural amino acid-mediated means. In such a preferred embodiment, the active center pocket of the nucleic acid polymerase, i.e., the channel for the entry and exit of nucleotide molecules, can be at a preferred position relative to the nanopore opening, facilitating efficient capture of nucleotide signals during sequencing and ensuring sequencing accuracy. At the same time, the template DNA and the DNA molecules of the synthesis reaction are in a conformation that is far from the nanopore opening and is not easily interacted with the nanopore to produce interference signals. Therefore, in this preferred scheme, due to the greatly reduced interference signal density produced by the DNA template / product and the nanopore, the negative impact on the correct sequencing signal is also greatly reduced, ensuring better signal recognition and higher sequencing accuracy.
[0080] In some embodiments, the nanopore and the nucleic acid polymerase can also form a stable, oriented complex through a double-coupling approach based on spytag-spycather covalent linkage and LM7-CL7 high-affinity binding. In this preferred embodiment, the C-terminus of the nanopore generates a specific polypeptide sequence spytag through fusion expression, and introduces LM7 at the N-terminus of the nanopore opening through fusion expression. At the same time, the nucleic acid polymerase introduces spycatcher at its N-terminus through genetic recombination, and introduces CL7 in the loop region near its active pocket. In such an embodiment, a stable complex can also be formed efficiently and quickly based on the synergistic effect of the two pairs of coupling groups, and the relative distance and orientation between the nanopore and the polymerase active center can be effectively controlled to achieve more powerful signal capture and interference from the DNA template.
[0081] In some embodiments, the nanopore and the nucleic acid polymerase can also form a stable, relatively fixed complex through a double coupling mode of biotin-avidin high affinity binding plus LM7-CL7 high affinity binding. In this preferred embodiment, the C-terminus of the nanopore is fused to express the LM7 sequence, and a Cys mutation is introduced at a specific position near the nanopore opening, so as to introduce biotin at this position through maleimide mediation. At the same time, the nucleic acid polymerase is introduced with CL7 at its N-terminus and avidin at its C-terminus through genetic recombination. In such an embodiment, a stable complex can also be formed efficiently and quickly based on the synergistic effect of the two pairs of coupling groups, and the relative distance and orientation between the nanopore and the active center of the polymerase can be effectively controlled to achieve more powerful signal capture and interference from the DNA template.
[0082] In summary, to overcome the defects of the prior art based on single coupling, the present application uses two different coupling pairs of functional groups between the nanopore and the polymerase to achieve double-site coupling between proteins. Compared with the existing single-site coupling mode, the introduction of the second efficient coupling functional group pair not only increases the coupling efficiency at low concentration in a short time, but also limits the conformational freedom of the polymerase relative to the nanopore opening to some extent, which significantly improves the nucleotide signal capture and effectively reduces the serious interference signal caused by the interaction between the template or product DNA molecule and the nanopore, thereby significantly improving the signal quality of the sequencing as a whole.
[0083] In some embodiments, the method of the present application includes specifically introducing two pairs of spontaneously coupling functional groups on the nanopore and the polymerase, respectively, to form a stable complex between the two protein molecules depending on the two coupling sites.
[0084] In some embodiments, the relative distance and orientation between the two protein molecules in the coupling complex are controlled by adjusting the position of the coupling functional group introduced on the protein molecule, selecting the structure of the functional group, and optimizing the length of the connecting molecule.
[0085] In some embodiments, the nanopore-polymerase coupling complex of the present application can be applied to nucleic acid sequencing.
[0086] In some embodiments, the two coupling sites can be based on covalent coupling, such as spytag-spycatcher, click chemistry, etc., or non-covalent coupling, such as biotin-avidin, antigen-antibody interaction, etc.
[0087] In some embodiments, the monomeric avidin and polymerase are directly obtained in soluble and functional avidin activity by means of fusion expression, without the need for inclusion body expression and renaturation process to obtain the functional monomeric avidin.
[0088] In some embodiments, the avidin can be selected to be fused at the N-terminus, C- terminus or in the middle Loop region of the polymerase.
[0089] The present application is further illustrated by the following specific examples.
[0090] EXAMPLE
[0091] Example 1. Avidin expression by fusion at different positions of the polymerase
[0092] This example relates to fusion site selection, definition and characterization of the enzyme activity after expression and purification, and verification of avidin activity.
[0093] In this example, we fused the streptavidin monomeric Rhizavidin to the C-terminal of the optimized phi29 DNA polymerase (SEQ ID NO:22) at P0065, Y148-H149 (P0073) and V509-G511 (P0072) respectively, the fusion protein junction was linked by (GS)n (n varied from 1 to 8, see specific sequence), the fusion protein sequences were SEQ ID NO:1, SEQ ID NO:2 and SEQ ID NO:3. To test the polymerase activity, we modified the 5' end of a hairpin DNA (SEQ ID NO:23) (CCTACCGTATCGTTTTACCAGTCAGTCAGTCAGTCAGTCAGTCAGTCAGTCAGTCAGCTAAGCACCTGTCTTCGCTTTTGCGAAGACAGGTGCTTAG) with CY3 fluorescent group, and the 3' end of another oligonucleotide DNA (SEQ ID NO:24) (GTAAAACGATACGGTAGG) complementary to the hairpin DNA with BHQ2 fluorescent quenching group, when the polymerase extended along the 3' end of the hairpin DNA, the newly synthesized DNA strand would replace the complementary oligonucleotide strand, making the quenching group release away from the fluorescent group, thus generating fluorescence (Figure 1). At the same time, to test the streptavidin function, we modified the Biotin group on an alpha hemolysin monomer protein (G255-Biotin), after mixing the streptavidin fusion polymerase with G255-Biotin, if the two proteins formed a coupled complex, the molecular weight would be larger, and the retention volume of gel filtration chromatography column would be smaller, thus to determine whether the streptavidin function was normal. The experiment showed that the above three fusion proteins could be soluble expressed in E. coli to obtain correctly folded proteins (Figure 2), and the polymerase activity and streptavidin function were normal (Figure 3 and Figure 4).
[0094] Figure 1 is a schematic diagram of the template DNA for polymerase activity detection, the polymerase extends along the 3' end of the hairpin DNA, the newly synthesized DNA strand replaces the complementary oligonucleotide strand, making the quenching group release away from the fluorescent group, thus generating fluorescence.
[0095] Figure 2 shows the results of purifying P0065, P0072 and P0073 by gel filtration chromatography column. The 13.5-15.0 mL fraction is the protein of interest, the peak shape of the three fusion proteins is symmetrical, without tailing and delay, indicating that the protein is folded regularly.
[0096] Figure 3 shows the results of the enzyme detector detecting the activity of the polymerase. The reaction system is 50 mM Tris-HCl pH 7.5, 200 mM KCl, 0.2 uM dNTP, 10 mM MgCl2, 25 nM template DNA, 50 nM polymerase, total volume 100 uL. The detection method is that the detection temperature is 25°C, the automatic injector is used to add MgCl2 at the end, the detection starts after shaking for 3 s, the excitation wavelength is 540 nm, and the detection emission wavelength is 570 nm. B1-mother phi29 DNA polymerase, B2-P0065, B3-P0072, B4-P0073. The test shows that the activity of the three fusion polymerases is not affected compared with the parent. Figure 4 shows the results of the P0065 avidin function test. A: G255-Biotin gel filtration chromatography purification, take the main peak fraction for standby; B: P0065 gel filtration chromatography purification, take the main peak fraction for standby; C: mix the purified P0065 with G255-Biotin according to a molar concentration of 1:4, and immediately inject, most of the P0065 and G255-Biotin form a complex, and the peak position moves to about 13 mL.
[0097] Example 2. Description of the implementation scheme based on antigen-antibody coupling
[0098] This example fuses LM7 on the nanopore and CL7 on the polymerase, and provides a coupling verification process and data.
[0099] In this example, a phi29 DNA polymerase homolog protein is used, and CL7 protein (number L0606, sequence see SEQ ID NO: 4) is fused at the N-terminus of the protein. On the other hand, we fuse LM7 protein (number G0518, sequence see SEQ ID NO: 5) at the C-terminus of an alpha hemolysin nanopore monomer protein. CL7 and LM7 are a pair of antigen-antibody proteins with very high specificity and affinity. To test whether the LM7-CL7 fusion expression can couple the polymerase and the nanopore protein, we use heparin as a medium for pull-down experiment. Since the polymerase has a very strong affinity with the heparin medium, and the alpha hemolysin nanopore monomer protein hardly binds to the heparin medium, we mix L0606 with G0518 for heparin affinity purification. Those free G0518 will directly flow through the heparin chromatography column, while the G0518 coupled with L0606 will be bound to the medium, and then eluted together with L0606. The experiment shows that the fusion expression scheme of this example can effectively couple the polymerase and the nanopore monomer.
[0100] Figure 5 shows SDS-PAGE electrophoresis diagram of L0606, G0518 and conjugated sample. Lane 1. G0518 only; Lane 2. L0606 only; Lane 3. L0606 mixed with G0518 at 1:2; 4. mixed sample was flowed through heparin medium column, indicating that G0518 does not bind to heparin medium; 5. heparin medium column elution sample (L0606 mixed with G0518 at 1:2), it can be seen that the sample contains L0606 and G0518, indicating that the two proteins are successfully conjugated to form a pull-down effect; M. Protein Maker.
[0101] Example 3. Preparation of a nanopore protein with a spytag at the C-terminus and biotin on one monomer
[0102] This example relates to the whole process steps of alpha hemolysin sequence, spytag sequence, linker, expression and purification, biotin modification, preparation of 7-mer, purification of 1:6 seven-mer, and shows a gel map.
[0103] In this example, the improved alpha hemolysin monomer protein (number G0037, sequence see SEQ ID NO: 6) is used as 6 parts of the seven-mer nanopore, and the 1 part monomer used for coupling with the polymerase is the spytag sequence fused at the C-terminus of G0037, with the introduction of N17C mutation for modification of GMBS-Biotin (number G0284, sequence see SEQ ID NO: 7).
[0104] Expression and purification of G0037 monomer protein:
[0105] pET26b was used as an expression vector when constructing the G0037 protein monomer expression vector, TEV and His-Tag tags were added at the C-terminus, the vector was transformed into E. coli BL21 strain, a single colony was picked and inoculated into 10 mL of LB medium containing 50 ug / mL kanamycin, 37°C, 250 rpm for 16 h, then the culture was transferred to 400 mL of TB autoinduction medium containing 50 ug / mL kanamycin, 25°C, 250 rpm for 16 h.
[0106] The culture was transferred to centrifuge bottles and centrifuged at 4000 rpm, 4°C for 15 min to pellet the cells, then the cell pellet was resuspended with 40 mL Binding buffer 1 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol). The resuspension was then subjected to sonication treatment in an ice-water bath, followed by centrifugation at 15000 rpm, 4°C for 15 min to pellet the residual cell debris, the supernatant was loaded into a 1 mL pre-equilibrated Ni-Agarose gel gravity column. After loading was completed, the column was washed with 10 mL of Washing buffer 1 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol, 30 mM imidazole), finally the target protein was eluted with 3-10 mL of Elution buffer-1 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol, 300 mM imidazole).
[0107] After the protein was eluted, the OD280 concentration was determined by spectrophotometer, then TEV protease was added at a concentration of 100:1, incubated at room temperature for 16 h to remove the C-terminal His-Tag label. After the enzyme digestion, the protein was dialyzed into 1 L of Binding Buffer-1, dialyzed at room temperature for 3 times, 1 h each time. Finally, the dialyzed sample was loaded again into a 1 mL pre-equilibrated Ni-Agarose gel gravity column, and the flow-through sample was collected to obtain G0037 monomer. The sample was tested by 12% SDS-PAGE gel electrophoresis and stored for later use.
[0108] G0284 monomer protein expression and purification:
[0109] The vector construction and expression of G0284 monomer were the same as G0037.
[0110] After expression the culture was transferred to centrifuge bottles and spun at 4000 rpm, 4°C for 15 min to pellet the cells, then the cell pellet was resuspended in 40 mL Binding buffer 2 (50 mM PB (pH 8.0), 200 mM NaCl, 10% glycerol). The resuspension was then sonicated in an ice water bath, then spun at 15000 rpm, 4°C for 15 min to pellet the remaining cell debris, the supernatant was loaded onto a 1 mL pre-equilibrated Ni-sepharose gravity column. After loading the column was washed with 10 mL of Washing buffer 2 (50 mM PB (pH 8.0), 200 mM NaCl, 10% glycerol, 30 mM imidazole), then the protein of interest was eluted with 3-10 mL of Elution buffer-2 (50 mM PB (pH 8.0), 200 mM NaCl, 10% glycerol, 300 mM imidazole).
[0111] After protein elution, the sample was dialyzed into 1 L of Binding buffer 2, 3 times for 1 h at room temperature. Then 400 uM of GMBS-Biotin was added to the dialyzed sample to a final concentration, and incubated at room temperature for 3 h, finally 1 mM of DTT was added to stop the reaction. The sample was checked by 12% SDS-PAGE before storage.
[0112] 1 :6 heptamer preparation and purification:
[0113] Protein concentration was measured by OD280. G0037 and G0284 were mixed at a ratio of 5:1, DPHPC phospholipid was added to a final concentration of 0.75 mg / mL, and the sample was incubated at 37°C for 16 h. The sample was then tested for heptamer formation by 12% SDS-PAGE. Then, β-OG (n-octyl-β-D-glucopyranoside) was added to a final concentration of 2.5%, and after the phospholipid was completely dissolved, Tween 20 was added to a final concentration of 0.2%. The sample was mixed well, centrifuged at 3500 rpm for 3 min, and the precipitate was removed. The supernatant was then dialyzed against 1 L of Buffer A (20 mM acetate-pH 5.0, 0.1% Tween-20) for 3 times, each for 1 h at room temperature. The dialyzed sample was loaded onto a cation exchange column (RESOURSE S1ML) and eluted with a gradient of 10-30% Buffer B (20 mM acetate-pH 5.0, 2 M NaCl, 0.1% Tween-20). The fractions with 35S conductivity were collected, and the 1:6 heptamer nanopore was obtained. Figure 6 shows the results of RESOURSE S chromatography purification of the 1:6 heptamer nanopore of G0284 and G0037. Peak 1 is 0:7, peak 2 is the target 1:6 component, peak 3 is the 2:5 component, and peak 4 is the 3:4 and above components. Figure 7 shows the SDS-PAGE electrophoresis results of the heptamer protein. M, Marker; Lane 1, G0037 monomer; Lane 2, G0284 monomer; Lane 3, G0037 and G0284 after incubation at 37°C for 16 h in DPHPC phospholipid; Lane 4, 1:6 heptamer of G0284 and G0037 after RESOURSE S chromatography purification. The heptamer nanopore is very stable and can tolerate SDS, so the electrophoresis process maintains the heptamer state, and the molecular weight is large. Lane 4 shows that the purified 1:6 heptamer has a purity of greater than 95% and can maintain the heptamer state.
[0114] Example 4. Coupling efficiency of SpyCatcher-spyTag single coupling, comparison of double coupling efficiency
[0115] In this example, we fused SpyCatcher protein at the C-terminal of optimized phi29 DNA polymerase, and fused Monomeric Rhizavidin at the position of Y148-H149, which is mutant number P0165, sequence see SEQ ID NO:8. Because SpyCatcher-spytag is covalent reaction, it can keep connection in denatured protein electrophoresis, while the coupling of avidin and biotin is non-covalent, which will dissociate in denatured electrophoresis. To test whether the coupling reaction of avidin and biotin can promote SpyCatcher-spytag coupling, we used the 1:6 heptamer described in Example 3 to mix with different proportions of P0165 and then ran SDS-PAGE electrophoresis for testing. In the experiment, P0165 in control group B was pre-added with biotin to block avidin, while experimental group C was not added with biotin for blocking. The results showed that the coupling reaction of avidin and biotin can effectively promote SpyCatcher-spytag coupling to form covalent connection.
[0116] Figure 8 shows the SDS-PAGE electrophoresis results of the comparison of coupling efficiency of single coupling and double coupling. After mixing, ice bath for 10 min, immediately add loading buffer to terminate the reaction, then load electrophoresis. A Pore only; B control group; C experimental group; D P0165 only. Comparing the control group and the experimental group, it can be seen that when the ratio is less than 1:2, the same ratio, the experimental group Pore has less remaining amount, and the proportion of covalent coupling is higher, indicating that the second coupling site of avidin and biotin can effectively shorten the distance between two proteins, and the coupling speed of SpyCatcher-spytag is accelerated.
[0117] Example 5. Double coupling scheme sequencing VS single coupling scheme sequencing process, experimental conditions and data
[0118] In this embodiment, the single-coupling scheme polymerase is an optimized phi29 DNA polymerase homologous protein, which is fused with SpyCatcher protein at its N-terminal (No. L0538, sequence see SEQ ID NO: 9), and the single-coupling scheme nanopore is the G0284:G0037 (1:6) heptamer of unmodified GMBS-Biotin; the double-coupling scheme polymerase is selected from, but not limited to, the following cases: 1. fused with SpyCatcher protein at its N-terminal and fused with avidin at D160-Y161 (No. L0479, sequence see SEQ ID NO: 10), 2. fused with SpyCatcher protein at its N-terminal and fused with avidin at D154-Q155 (No. L0545, sequence see SEQ ID NO: 11), 3. fused with SpyCatcher protein at its N-terminal and fused with avidin at L51-Q52 (No. L0842, sequence see SEQ ID NO: 12), 4. fused with SpyCatcher protein at its N-terminal and fused with avidin at K20-K21 (No. L0843, sequence see SEQ ID NO: 13). The double-coupling scheme nanopore is the G0284:G0037 (1:6) heptamer of modified GMBS-Biotin. The sequencing process and sequencing signal detection are shown in FIG. 9, in which the polymerase uses the DNA template to be tested and four kinds of nucleotides labeled with different labels as raw materials to synthesize a new DNA chain, in which process the labels enter the nanopore in turn, which affects the current passing through the nanopore, and the change of the current is read by the detector to become DNA sequence information.
[0119] FIG. 9 is a model diagram of the sequencing principle. A model of the sequencing device, an artificial phospholipid membrane divides the salt solution into two parts, and the salt ions cannot pass freely. Two salt solution pools are connected to the electrodes, respectively. When the nanopore is embedded in the phospholipid membrane, the salt ions can pass through the phospholipid membrane through the nanopore channel. At this time, the voltage is applied to the electrode to form an electric current. When the state of the nanopore channel changes, the current can be monitored in real time by the detector. B. Model of sequencing signal detection, 1. DNA polymerase; 2. DNA to be sequenced; 3. Nucleotides with labels; 4. Connection bridge between DNA polymerase and nanopore; 5. Nanopore protein; 6. Artificial phospholipid membrane; 7. Change of current size through the nanopore channel corresponding to the model diagram of four kinds of bases.
[0120] Preparation of sequencing complex:
[0121] In 100 uL of Buffer C (50 mM Tris-HCl pH 8.0, 300 mM KAc, 0.1% Tween-20, 1 mM DTT), the DNA template to be tested, polymerase and nanopore were added in a molar ratio of 2:2:1 (more than 10 uL of total volume was added), mixed and incubated at 37°C for 30 min, and then centrifuged at 15000 rpm, 4°C for 15 min to precipitate the coupled polymerase-coupled nanopore protein that did not bind DNA. The supernatant was the desired sequencing complex.
[0122] Sequencing test:
[0123] After the device shown in FIG. 9A was ready, the prepared sequencing complex sample was diluted with Buffer-Seq (50 mM Tris-HCl pH 8.0, 300 mM KAc) to the appropriate concentration and then added to the device. After the sequencing complex was embedded in the artificial phospholipid membrane, the excess DPN complex was washed away with Buffer-Seq. Finally, the prepared sequencing reagent (4 labeled nucleotides with a final concentration of 5 uM and a final concentration of 0.4 mM MnCl2 dissolved in Buffer-Seq) was added, and the test recording began (FIG. 9B). The test used an alternating current excitation voltage of 180 mV and 500 Hz, and the detector measured the current value through the nanopore every 200 us.
[0124] Sequencing effect comparison:
[0125] The test results show that the single coupling scheme has more serious interference signals (mainly from DNA or polymerase interference), and the label capture rate is low, and the signal boundary is blurred, which increases the error rate of signal recognition. Compared with the single coupling scheme, the interference signal of the double coupling scheme test is significantly reduced, the capture rate is higher, the signal is clearer and distinguishable, and the sequencing accuracy is improved (FIG. 10 and FIG. 11).
[0126] FIG. 10 shows the single coupling sequencing test data results. The open pore current signal band is severely weakened by the interference signal, the base label signal is blurred, the signal event is easily misidentified as multiple events, and the C base label signal is close to the interference signal and is easily missed or misidentified.
[0127] FIG. 11 shows the double coupling sequencing test data results using case 1 polymerase. Compared with single coupling, the interference signal is significantly reduced, and the base label signal is basically not affected, the signal is clear and easy to identify and distinguish. The test data using other case polymerases (2-4) is similar.
[0128] Figure 12 shows the three-dimensional structure of the possible double-coupling molecule formed by the nanopore and polymerase through the avidin-biotin and SpyCatcher-spytag.
[0129] Example 6. Double-coupling scheme based on CL7-LM7 and biotin-avidin
[0130] In this example, the polymerase used for double-coupling reaction is an optimized phi29 homologous enzyme, and the coupling design of the polymerase is selected from, but not limited to, the following cases: 1. The avidin is fused at the N-terminus, and the CL7 protein (number L0690, sequence see SEQ ID NO: 14) is fused at the C-terminus; 2. The CL7 protein is fused at the N-terminus, and the avidin is fused at the C-terminus (number L0691, sequence see SEQ ID NO: 15). On the other hand, we fused the LM7 protein at the C-terminus of an alpha hemolysin nanopore monomer protein, and introduced the N17C mutation for modification of GMBS-biotin (number G0556, sequence see SEQ ID NO: 16), and then prepared a 1:6 heptamer with G0037, the preparation method is the same as described in Example 3.
[0131] The preparation and testing of sequencing complexes are the same as described in Example 5. The test results of the two polymerases show that the sequencing results based on this scheme have similar effects to the double-coupling scheme of Example 5, as shown in Figure 13.
[0132] Example 7. Double-coupling scheme based on click chemistry reaction groups N3-DBCO and SpyCatcher-spytag
[0133] In this example, the polymerase used for double-coupling reaction is an optimized phi29 polymerase, which is fused with SpyCatcher at the C-terminus, and the codon of the amino acid at position K143 is mutated to the stop codon TAG. By co-expression with a plasmid carrying a special transfer RNA and aminoacyl-tRNA synthetase, azidomethyl-L-phenylalanine (pAMF) carrying an N3 group can be inserted into the amino acid sequence of the polymerase (number P0318, sequence see SEQ ID NO: 17). On the other hand, we modified GMBS-DBCO at the N17C site of G0284 as described above, and prepared a 1:6 heptamer with G0037. The preparation and testing of sequencing complexes are the same as described in Example 5. The test results show that the sequencing results based on this scheme have similar effects to the double-coupling scheme of Example 5, as shown in Figure 14.
[0134] Example 8. Double-coupling scheme based on SpyCatcher-spytag and CL7-LM7
[0135] In this example, the polymerase used for the dual coupling reaction is an optimized phi29 polymerase, and the coupling design is to fuse SpyCatcher at its N-terminus and to fuse GSG-CL7-SG sequence (SEQ ID NO: 18) between I167-P171 sequence in the loop region by replacement (SEQ ID NO: 18). On the other hand, we fused spytag sequence at the C-terminus of an alpha hemolysin nanopore monomer protein and replaced 1A-T19 sequence at the N-terminus with LM7-GSGT sequence (SEQ ID NO: 19) (SEQ ID NO: 19). Then we made 1:6 heptamer with G0037 as described in Example 3. The sequencing complex preparation and testing were as described in Example 5. The testing results showed that the sequencing results based on this scheme had similar effects as the dual coupling scheme of Example 5, as shown in Figure 15.
[0136] Example 9. Dual coupling scheme based on click chemistry reaction group N3-DBCO and biotin-avidin
[0137] In this example, the polymerase used for the dual coupling reaction is an optimized phi29 polymerase, and the coupling design is to fuse SpyCatcher at its N-terminus and to fuse GSG-CL7-SG sequence (SEQ ID NO: 18) between I167-P171 sequence in the loop region by replacement (SEQ ID NO: 18). On the other hand, we fused spytag sequence at the C-terminus of an alpha hemolysin nanopore monomer protein and replaced 1A-T19 sequence at the N-terminus with LM7-GSGT sequence (SEQ ID NO: 19) (SEQ ID NO: 19). Then we made 1:6 heptamer with G0037 as described in Example 3. The sequencing complex preparation and testing were as described in Example 5. The testing results showed that the sequencing results based on this scheme had similar effects as the dual coupling scheme of Example 5, as shown in Figure 15.
[0138] The following are the partial amino acid sequences referred to herein (for sequences containing a histidine tag, the sequence listed also encompasses the corresponding sequence without the histidine tag, unless the histidine tag is required to function when the sequence is used; for sequences containing a methionine (M) at the N-terminus, the sequence listed also encompasses the corresponding sequence without the methionine):
[0139] SEQ ID NO: 22 DNA polymerase
[0140] SEQ ID NO: 1 Fusion protein 1
[0141] SEQ ID NO: 2 Fusion protein 2
[0142] SEQ ID NO: 3 Fusion protein 3
[0143] SEQ ID NO: 4 Number L0606
[0144] SEQ ID NO: 5 Number G0518
[0145] SEQ ID NO: 6 Number G0037
[0146] SEQ ID NO: 7 Number G0284
[0147] SEQ ID NO: 8 Number P0165
[0148] SEQ ID NO: 9 Number L0538
[0149] SEQ ID NO: 10 Number L0479
[0150] SEQ ID NO: 11 Number L0545
[0151] SEQ ID NO: 12 Number L0842
[0152] SEQ ID NO: 13 Number L0843
[0153] SEQ ID NO: 14 Numbered L0690
[0154] SEQ ID NO: 15 Numbered L0691
[0155] SEQ ID NO: 16 Numbered G0556
[0156] SEQ ID NO: 17 Numbered P0318
[0157] SEQ ID NO: 18 Numbered L0604
[0158] SEQ ID NO: 19 Numbered G0434
[0159] SEQ ID NO: 20 Numbered P0188
[0160] SEQ ID NO: 21 Numbered G0320
[0161] SEQ ID NO: 25
[0162] SEQ ID NO: 26
Claims
1. A protein comprising two polypeptide fragments, between which are included a first coupling site and a second coupling site, wherein at least one coupling site is formed by any one of the following: 1) protein fusion; 2) click chemistry; 3) antigen-antibody binding; 4) ligand-receptor binding; and 5) spytag-spycatcher binding.
2. The protein of claim 1, wherein the two coupling sites are independently formed by any one of the following: 1) protein fusion; 2) click chemistry; 3) antigen-antibody binding; 4) ligand-receptor binding; and 5) spytag-spycatcher binding.
3. The protein of claim 1 or 2, wherein the two polypeptide fragments are from a DNA polymerase and a nanopore protein, respectively.
4. The protein of any one of claims 1-3, wherein the nanopore protein is an alpha-hemolysin or a mutant thereof; preferably, the alpha-hemolysin comprises the amino acid sequence set forth in SEQ ID NO: 6, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, or 98% sequence identity to the amino acid sequence set forth in SEQ ID NO:
6.
5. The protein of any one of claims 1-4, wherein the nanopore protein is a NetB protein or a mutant thereof; preferably, the NetB protein or a mutant thereof comprises the amino acid sequence set forth in SEQ ID NO: 25 or 26, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, or 98% sequence identity to the amino acid sequence set forth in SEQ ID NO: 25 or 26.
6. The protein of any one of claims 1-5, wherein the DNA polymerase is a Phi29 nucleic acid polymerase or a mutant thereof; preferably, the Phi29 nucleic acid polymerase comprises the amino acid sequence set forth in SEQ ID NO: 22, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, or 98% sequence identity to the amino acid sequence set forth in SEQ ID NO:
22.
7. The protein of any one of claims 1-6, wherein the antigen-antibody binding is a LM7-CL7 protein binding, and / or the ligand-receptor binding is a biotin-avidin binding.
8. The protein of any one of claims 1-7, wherein the avidin is fused to a loop region of the DNA polymerase.
9. The protein of any one of claims 1-8, wherein the first coupling site is accomplished by click chemistry reaction, and the second site is formed by spytag-spycatcher binding; preferably, the Spycatcher protein is fused to the C-terminus of the DNA polymerase, and the spytag sequence is fused to the C-terminus of at least one monomer of the alpha-hemolysin, and forms a click chemistry coupling site between the N17 of the monomer and the K143 of the DNA polymerase. More preferably, the DNA polymerase comprises the sequence set forth in SEQ ID NO: 17, and the monomer comprises the sequence set forth in SEQ ID NO:
7.
10. The protein of any one of claims 1-9, wherein the first coupling site is formed by a spytag-spycatcher binding, and the second site is formed by a LM7-CL7 protein pair binding. Preferably, the SpyCatcher protein is fused at the N-terminus of the DNA polymerase and replaces the I167-P171 sequence of the DNA polymerase with a GSG-CL7-SG sequence; the spytag sequence is fused at the C-terminus of at least one monomer of the alpha-hemolysin and replaces the 1A-T19 sequence of the monomer with a LM7-GSGT sequence. More preferably, the DNA polymerase comprises the sequence set forth in SEQ ID NO: 18, and the monomer comprises the sequence set forth in SEQ ID NO:
19.
11. The protein of any one of claims 1-10, wherein the first coupling site is completed by a click chemistry reaction, and the second site is formed by a biotin-avidin binding. Preferably, the avidin is fused at the C-terminus of the DNA polymerase, at least one monomer of the alpha-hemolysin comprises an E31C mutation for biotin modification, and a click chemistry reaction binding is formed between Y106 of the DNA polymerase and N17 of the monomer. More preferably, the DNA polymerase comprises the sequence set forth in SEQ ID NO: 20, and the monomer comprises the sequence set forth in SEQ ID NO:
21.
12. The protein of any one of claims 1-11, wherein the first coupling site is formed by a biotin-avidin binding, and the second site is formed by a LM7-CL7 protein pair binding. Preferably, the avidin is fused at the N-terminus or C-terminus of the DNA polymerase, and the CL7 protein is correspondingly fused at the C-terminus or N-terminus of the DNA polymerase; the LM7 protein is fused at the C-terminus of at least one monomer of the alpha-hemolysin and an N17C mutation is introduced in the monomer for biotin modification. More preferably, the DNA polymerase fused with the avidin and the CL7 protein comprises the sequence set forth in SEQ ID NO: 14 or 15; and the monomer comprises the sequence set forth in SEQ ID NO:
16.
13. The protein of any one of claims 1-7, wherein the first coupling site is formed by a spytag-spycatcher binding, and the second site is formed by a biotin-avidin binding. Preferably, the Spycatcher protein is fused at the N- or C-terminus of the DNA polymerase, the avidin is fused between Y148-H149, between D160-Y161, between D154-Q155, between L51-Q52 or between K20-K21 of the DNA polymerase; the spytag sequence is fused at the C-terminus of at least one monomer of the alpha-hemolysin and an N17C mutation is introduced in the monomer for biotin modification; More preferably, the DNA polymerase fused with the Spycatcher protein and the avidin comprises the sequence set forth in SEQ ID NO: 8, 10, 11, 12 or 13; the monomer comprises the sequence set forth in SEQ ID NO:
7.
14. A method of making a protein comprising two polypeptide fragments, comprising forming a first coupling site and a second coupling site between the two polypeptide fragments, wherein at least one coupling site is formed by any one of: 1) protein fusion; 2) click chemistry; 3) antigen-antibody binding; 4) ligand-receptor binding; and 5) spytag-spycatcher binding.
15. The method of claim 14, wherein the two coupling sites are independently formed by any one of: 1) protein fusion; 2) click chemistry; 3) antigen-antibody binding; 4) ligand-receptor binding; and 5) spytag-spycatcher binding.
16. The method of claim 14 or 15, wherein the two polypeptide fragments are from a DNA polymerase and a nanopore protein, respectively.
17. Use of the protein of any one of claims 1-13 or the protein made by the method of any one of claims 14-16 in nucleic acid sequencing.
18. The use of claim 17, wherein the sequencing is in a sequencing-by-synthesis mode.
19. Use of a DNA polymerase as a chaperone in expression of a protein of interest.
20. The use of claim 19, wherein the protein of interest is expressed in bacteria, and the DNA polymerase as a chaperone prevents the protein of interest from forming inclusion bodies.
21. [Amended according to Rule 26 20.08.2025] The use of claim 19 or 20, wherein the bacteria is E. coli.
22. The use of any one of claims 19-21, wherein the protein of interest is an avidin.
23. The use of any one of claims 19-22, wherein the avidin is fused at the C-terminus, N-terminus or Loop region of the DNA polymerase.
24. The use of any one of claims 19-23, wherein the DNA polymerase is a Phi29 nucleic acid polymerase or a mutant thereof; preferably, the Phi29 nucleic acid polymerase comprises the amino acid sequence set forth in SEQ ID NO: 22, or comprises an amino acid sequence having at least 80%, 85%, 90%, 95%, or 98% sequence identity to the amino acid sequence set forth in SEQ ID NO: 22.
Citation Information
Patent Citations
Hybridization linkers
CN102439043A
Site-specific bio-conjugation methods and compositions useful for nanopore systems
CN109069662A
Method for controlling speed of polypeptide passing through nanopore and application thereof
CN112147185A
Nanopore preparation and detection method and detection device thereof
CN115398008A
Nanopore sequencing method and kit
CN118207308A