Mutant of helicase Dda, preparation method of mutant and application of mutant in sequencing
Patent Information
- Application Number
- CN202280102752.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-08-08
AI Technical Summary
In the existing nanopore sequencing technology, the protein yield, stability, sequencing speed uniformity, sequencing throughput and other indicators of helicase Dda are insufficient, which cannot meet the growing demand for sequencing, affecting the stability of sequencing accuracy and speed. .
By carrying out site-specific mutations on the amino acid sequence of Dda helicase, such as replacing positively charged, negatively charged or large side chain amino acids with neutral amino acids, the properties of its protein and interaction with nanopores are optimized, thereby improving its performance in Performance in nanopore sequencing.
It significantly improves the protein yield, sequencing speed and sequencing speed uniformity of helicase Dda, improves the capability and efficiency of nanopore sequencing, and facilitates the further development of nanopore sequencing technology.
Smart Images

Figure CN120457201A_ABST
Abstract
Description
A mutant of helicase Dda, its preparation method and its application in sequencing Technical Field
[0001] The present invention belongs to the technical fields of gene sequencing, molecular detection and clinical detection, and particularly relates to a mutant of helicase Dda, a preparation method thereof and application thereof in sequencing. Background Art
[0002] Nanopore detection technology, a novel platform, is an emerging single-molecule sequencing technology with the advantages of low cost, high throughput, and label-free operation. Its emergence has revolutionized the genetic sequencing industry. The lightweight and portable equipment used in this technology can meet diverse sequencing scenarios. Furthermore, due to its non-amplified direct sequencing nature, it has no length limit on the sequenceable DNA, allowing for real-time base calling and direct sequencing of modifications such as RNA and methylation, as well as other single molecules. This technology has the potential to overcome the limitations of second-generation sequencing and serve as a supplement to address diverse sequencing needs. Nanopore sequencing technology has broad applications in many fields, including molecular biology, medicine, epidemiology, and ecology.
[0003] The basic principle of nanopore sequencing is to use a nanopore that can provide an ion current channel. The nanopore is inserted into a membrane (protein or solid) as a signal sensor to separate two electrolyte chambers filled with electrolyte. When voltage is applied between the two electrolyte chambers, a stable perforation current is generated. Driven by electrophoresis, the molecule to be tested passes through the nanopore, hindering the flow of ions and causing fluctuations in the current signal. Different bases have different effects on the current. By detecting the current fluctuation signal of the nanopore in real time, and using machine learning to analyze and decode the current signal, the sequence information of the molecule to be tested can be read in real time for gene sequencing.
[0004] During the sequencing process, it is impossible to accurately obtain polynucleotide sequence information due to the extremely high speed at which nucleic acid molecules pass through the nanopore channel. Therefore, effectively reducing and controlling the perforation motion of nucleic acid molecules is a key technical issue in realizing nanopore sequencing. At present, the most common and effective method is to use the idea of helicase unwinding to control the perforation motion of nucleic acid molecules, improve detection accuracy, and maintain sequencing speed and sequencing uniformity. In this process, the properties of the helicase itself, such as high protein quality, good DNA binding ability, and appropriate DNA unwinding ability, are very critical. In addition, as a rate-controlling protein, the stability of its interaction interface with the nanopore has a positive effect on improving the uniformity, stability, and signal-to-noise ratio of the sequencing signal.
[0005] The helicases currently used in commercial nanopore sequencers are either Dda helicases derived from bacteriophage T4 or Dda helicase mutants modified with their pin and tower domains. These helicases, including protein yield, stability, sequencing speed, sequencing speed uniformity, sequencing throughput, and pore compatibility, leave much to be desired and cannot meet the rapidly evolving sequencing needs. There is an urgent need for high-performance helicases to improve and enhance nanopore sequencing technology, enabling it to leverage its advantages and contribute to the advancement of science and technology.
[0006] Summary of the Invention
[0007] In order to improve the protein yield, stability, sequencing speed, sequencing speed uniformity, sequencing flux, and pore compatibility of the helicase Dda in currently commercialized nanopore sequencers in nanopore sequencing applications, the present invention provides a mutant of the helicase Dda, a preparation method thereof, and its application in sequencing. The mutant of the helicase Dda can be used to control the movement of polynucleotides and can be applied to nanopore sequencing. Compared with the wild-type Dda helicase or the modified Dda helicase, the mutant has significant improvements in one or more indicators such as protein yield, sequencing speed, and sequencing speed uniformity, thereby greatly improving its sequencing capability and contributing to the further promotion of nanopore sequencing.
[0008] In order to solve the above technical problems, the first aspect of the present invention provides a mutant of helicase Dda, wherein the mutant has one or more mutations selected from the following in the amino acid sequence shown in SEQ ID NO: 4:
[0009] (1) replacing the positively charged amino acids in the amino acid sequence with neutral amino acids;
[0010] (2) replacing the negatively charged amino acids in the amino acid sequence with neutral amino acids;
[0011] (3) truncating the N-terminal amino acids in the amino acid sequence or replacing them with short side chain neutral amino acids;
[0012] (4) Substituting amino acids with larger side chain steric hindrance in the amino acid sequence with amino acids with shorter side chain.
[0013] In some technical solutions, the mutated site is selected from one or more of the following:
[0014] (1), the mutated sites include K11, K19, K22, K24, K25, K38, K43, K67, K68, K72, K76, K86, K101, K108, K123, K126, K145, K166, K177, K194, K199, K227, K243, K247, K254, K255, K261, K277, K280, K284, K310, K339, K341, one or more of K351, K358, K364, K368, K371, K381, K386, K388, K397, R110, R122, R148, R191, R207, R216, R236, R253, R298, R312, R321, R337, R405, R431, R433, H26, H27, H64, H82, H165, H204, H322, H396, and H414;
[0015] Preferably, the mutated sites include K11 deletion, K11A, K11L, K11I, K19A, K177A, K194A, K199A, K227A, K254A, K386A, K397A, R191A, R216A, R405A, H204A, H322A, K11S, K19S, K177S, K194S , K199S, K227S, K194I, K199L, K254L, R191S, R207A, R207S, R216S, H204S, R216I, R236A, R253A, K11G, K177G, K194G, K199G, K227G, R191G, R207G, and R405S;
[0016] (2), the mutated sites include D4, D5, D105, D143, D167, D185, D189, D198, D202, D212, D217, D230, D231, D246, D260, D262, D282, D324, D332, D333, D346, D376, D379, D404, D417, one or more of D435, E8, E23, E47, E54, E77, E99, E102, E116, E154, E172, E175, E234, E258, E267, E273, E288, E301, E303, E317, E328, E334, E338, E347, E348, E361, and E419;
[0017] Preferably, the mutated sites include D4 deletion, D5 deletion, E8 deletion, D4A, D5A, D185A, D189A, D198A, D202A, D212A, D217A, D346A, E8A, E54A, E172A, E175A, E234A, E258A, E288A, D5S, D185S, D189S, D198S, D202S, D212S, E8S, E8I, E54S, E172S, E175S, D217S, D5L, D185L, D189L, D198L, D202L, D212L, E8L, D5I, D4S, D5S, D4G, One or more of D5G, D185G, D189G, D198G, D202G, D212G, E8G, E54G, E192G, E175G, and E54L;
[0018] (3), wherein the mutated sites include one or more of M1, T2, F3, L6, T7, G9 and Q10;
[0019] Preferably, the mutated sites include one or more of M1 deletion, F3 deletion, L6 deletion, Q10 deletion, M1-K11 deletion, M1-F3 deletion, M1-T7 deletion, M1A, T2A, Q10A, Q10S, Q10L, T7I and Q10I;
[0020] (4), wherein the mutated sites include one or more of M219, M238, M271, W195, W323, W366, F3, F14, F209, F218, F223, F233, F240, F257, F263, Y169, Y197, Y279, Y304, Y318, Q10, Q272, N12, N14, N88, N180, N192, N221, N235, N242, and N249;
[0021] Preferably, the mutated sites include one or more of F3A, F209A, F218A, F223A, F233A, F240A, F257A, F263A, Y197A, Y279A, N12A, N14A, N192A, N221A, N235A, N242A, N249A, F3S, F257S, Y197S, N192S, N221S, F3G, Y197G, H204G, Q10G, N12G, N192G and N221G.
[0022] In some technical embodiments, the mutation comprises one or more of the following amino acid substitutions: D5, E8, K11, K177, D189, R191, D202 and R216.
[0023] In some preferred technical solutions, the mutation is selected from one or more of the following groups:
[0024] (i) D5T, D5H, D5Y, D5Q or D5N;
[0025] (ii) E8T, E8H, E8Y, E8Q or E8N;
[0026] (iii) K11A, K11T, K11H, K11Y, K11Q or K11N;
[0027] (iv) K177E, K177A, K177Q, K177S, K177L, K177T, K177H, K177Y or K177N;
[0028] (v) D189A, D189T, D189H, D189Y, D189Q or D189N;
[0029] (vi) R191A, R191S, R191L, R191T, R191H, R191Y, R191Q or R191N;
[0030] (vii)K199Q;
[0031] (viii) D202T, D202Q, D202Y, D202H or D202N;
[0032] (ix) R216L, R216G, R216T, R216H, R216Y, R216Q or R216N.
[0033] In some technical embodiments, the mutation is K11A, R191A, R191S, D189A, K177H, K177Y, K177E, K177A, K177Q or K177S.
[0034] In order to solve the above technical problems, the second aspect of the present invention is a nucleic acid, which encodes the mutant as described in the first aspect of the present invention.
[0035] In some preferred technical solutions, the sequence of the nucleic acid is a sequence shown in SEQ ID NO: 5 in which one or more nucleotides are replaced, deleted or inserted.
[0036] In some technical solutions, the mutation is to replace CGT at positions 571-573 with GCC; or,
[0037] Replace CGT with AGT at positions 571-573; or,
[0038] Replace GAT with GCC at positions 565-567; or,
[0039] Replace AAA with CAT at positions 529-531; or,
[0040] Replace AAA with TAT at positions 529-531; or,
[0041] Replace AAA with GAA in positions 529-531; or,
[0042] Replace AAA with GCC at positions 529-531; or,
[0043] Replace AAA with CAG at positions 529-531; or,
[0044] Replace AAA with AGT at positions 529-531.
[0045] In order to solve the above technical problems, the third aspect of the present invention provides a recombinant expression vector, which comprises the nucleic acid as described in the second aspect of the present invention.
[0046] In some preferred technical solutions, the backbone plasmid of the recombinant expression vector is PET.28a(+), PET.21a(+) or PET.32a(+).
[0047] In order to solve the above technical problems, the fourth aspect of the present invention provides a transformant, which comprises the nucleic acid as described in the second aspect of the present invention or the recombinant expression vector as described in the third aspect of the present invention.
[0048] In some preferred technical solutions, the host cell of the transformant is Escherichia coli.
[0049] In some more preferred technical solutions, the host cell of the transformant is BL21(DE3), BL21Star(DE3)pLyss, Rossata(DE3) or Lemo21(DE3).
[0050] In order to solve the above technical problems, the fifth aspect of the present invention provides a method for preparing the mutant as described in the first aspect of the present invention, which comprises culturing the transformant as described in the fourth aspect of the present invention in a culture medium and fermenting it to produce the mutant.
[0051] In order to solve the above technical problems, the sixth aspect of the present invention provides a helicase-linker complex, which includes the mutant as described in the first aspect of the present invention and a sequencing linker; the sequencing linker includes two DNA chains with complementary partial regions, for example, as shown in SEQ ID NO: 1 and 2, respectively.
[0052] In order to solve the above technical problems, the seventh aspect of the present invention provides a kit, which includes the mutant as described in the first aspect of the present invention.
[0053] In some preferred technical solutions, the kit further comprises one or more of the following: single-stranded DNA containing a cholesterol anchor molecule at the 5' end, a nanopore, a nanopore protein, an electrical signal detector, a membrane and / or a sequencing buffer.
[0054] In some technical solutions, the anchoring molecule is a hydrophobic molecule, preferably selected from any one or more of the following: lipids, fatty acids, sterols, carbon nanotubes, polypeptides, proteins and / or amino acids, such as cholesterol, palmitate or tocopherol; the membrane is an amphiphilic membrane (e.g., a phospholipid bilayer), a polymer membrane (e.g., a diblock copolymer, a triblock copolymer), or any combination thereof; the nanopore is a transmembrane protein pore or a solid-state pore; the transmembrane protein pore is selected from hemolysin, MspA, MspB, MspC, MspD, FraC, ClyA, PA63, CsgG, CsgD, XcpQ, SP1, phi29 connector protein, InvG, GspD, or any combination thereof; and the sequencing buffer is a dihydrogen phosphate-hydrogen phosphate buffer system, a carbonic acid-sodium bicarbonate buffer system, a Tris-HCl buffer system, a HEPES buffer system, a MOPS buffer system, or any combination thereof.
[0055] In order to solve the above technical problems, the eighth aspect of the present invention provides a DNA unwinding method, which comprises unwinding double-stranded DNA using the mutant described in the first aspect of the present invention or the helicase-linker complex described in the sixth aspect of the present invention.
[0056] In order to solve the above technical problems, the ninth aspect of the present invention provides a sequencing method, which includes the following steps: using the DNA unwinding method as described in the eighth aspect of the present invention to sequence the double-stranded DNA while unwinding it.
[0057] In order to solve the above technical problems, the tenth aspect of the present invention provides a mutant as described in the first aspect of the present invention, the nucleic acid as described in the second aspect of the present invention, the recombinant expression vector as described in the third aspect of the present invention, the transformant as described in the fourth aspect of the present invention, the helicase-linker complex as described in the sixth aspect of the present invention, or the kit as described in the seventh aspect of the present invention for use in unwinding or sequencing or in preparing reagents for unwinding or sequencing.
[0058] In some preferred technical solutions, the sequencing is nanopore sequencing.
[0059] The beneficial effects brought about by the technical solution of the present invention are:
[0060] The present invention optimizes sequencing performance by modifying key amino acid sites in a helicase derived from bacteriophage T4 or a Dda helicase mutant modified with its pin and tower domains, based on the helicase used in currently commercialized nanopore sequencers. The present invention primarily provides a Dda helicase mutant, a method for preparing the mutant, and its application in sequencing. These mutants can be used to control the movement of polynucleotides and are applicable to nanopore sequencing. These mutants significantly improve protein yield, sequencing speed, sequencing speed uniformity, and sequencing throughput compared to wild-type or modified Dda helicases, significantly enhancing their sequencing capabilities and contributing to the further advancement of nanopore sequencing. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 shows the purification results of Dda-1 using molecular sieve Superdex 200. (A) Elution profile of Dda-1 using molecular sieve Superdex 200; (B) Elution profile of Dda-1 using molecular sieve gel.
[0062] Figure 2 is a histogram of the Dda-1 sequencing speed distribution.
[0063] Figure 3 shows the purification results of DDA-20 using molecular sieve Superdex 200. (A) Elution profile of DDA-20 using molecular sieve Superdex 200; (B) Elution profile of DDA-20 using molecular sieve gel.
[0064] FIG4 is a histogram of the sequencing speed distribution of Dda-20.
[0065] Figure 5 shows the purification results of Dda-32 using molecular sieve Superdex 200. (A) Elution profile of Dda-32 using molecular sieve Superdex 200; (B) Elution profile of Dda-32 using molecular sieve gel.
[0066] FIG6 is a histogram of the sequencing speed distribution of Dda-32.
[0067] Figure 7 shows the purification results of Dda-33 using molecular sieve Superdex 200. (A) Elution profile of Dda-33 using molecular sieve Superdex 200; (B) Elution profile of Dda-33 using molecular sieve gel.
[0068] FIG8 is a histogram of the sequencing speed distribution of Dda-33.
[0069] Figure 9 shows the purification results of Dda-45 using molecular sieve Superdex 200. (A) Elution profile of Dda-45 using molecular sieve Superdex 200; (B) Elution profile of Dda-45 using molecular sieve gel.
[0070] FIG10 is a histogram of the sequencing speed distribution of Dda-45.
[0071] Figure 11 shows the purification results of Dda-56 using molecular sieve Superdex 200. (A) Elution profile of Dda-56 using molecular sieve Superdex 200; (B) Elution profile of Dda-56 using molecular sieve gel.
[0072] FIG12 is a histogram of the sequencing speed distribution of Dda-56.
[0073] Figure 13 shows the purification results of Dda-58 using molecular sieve Superdex 200. (A) Elution profile of Dda-58 using molecular sieve Superdex 200; (B) Elution profile of Dda-58 using molecular sieve gel.
[0074] FIG14 is a histogram of the sequencing speed distribution of Dda-58.
[0075] Figure 15 shows the purification results of Dda-62 using molecular sieve Superdex 200. (A) Elution profile of Dda-62 using molecular sieve Superdex 200; (B) Elution profile of Dda-62 using molecular sieve gel.
[0076] FIG16 is a histogram of the sequencing speed distribution of Dda-62.
[0077] Figure 17 shows the purification results of Dda-66 using molecular sieve Superdex 200. (A) Elution profile of Dda-66 using molecular sieve Superdex 200; (B) Elution profile of Dda-66 using molecular sieve gel.
[0078] FIG18 is a histogram of the sequencing speed distribution of Dda-66.
[0079] Figure 19 shows the purification results of DDA-70 using molecular sieve Superdex 200. (A) Elution profile of DDA-70 using molecular sieve Superdex 200; (B) Elution profile of DDA-70 using molecular sieve gel.
[0080] Figure 20 is a histogram of the sequencing speed distribution of Dda-70.
[0081] Figure 21 shows the purification results of Dda-77 using molecular sieve Superdex 200. (A) Molecular sieve Superdex 200 elution profile of Dda-77; (B) Molecular sieve gel elution profile of Dda-77.
[0082] FIG22 is a histogram of the sequencing speed distribution of Dda-77.
[0083] FIG23 is a schematic diagram of a nanopore sequencing patch clamp amplifier. DETAILED DESCRIPTION
[0084] Example 1 Cloning, expression and purification of optimized helicase Dda mutant
[0085] 1. Cloning and Expression of the Optimized Helicase Dda Mutant
[0086] The DNA sequences of several optimized helicase Dda mutants were respectively connected to the PET.28a(+) plasmid, using double enzyme cleavage sites Nde1 and Xho1, so that the N-terminus of the expressed optimized helicase Dda mutant protein had a 6*His tag and a thrombin enzyme cleavage site.
[0087] The cloned PET.28a(+)-Dda-XX mutant plasmid was transformed into E. coli expression bacteria BL21(DE3) or its derivatives. A single colony was picked and inoculated into 20 mL of LB medium containing kanamycin resistance and cultured at 37°C with shaking overnight. Then, the colony was transferred into 2 L of LB medium containing kanamycin resistance and cultured at 37°C with shaking until the OD 600 =0.6-0.8, cooled to 16°C, added with IPTG at a final concentration of 500 μM to induce expression overnight, and obtained Dda-XX bacteria.
[0088] 2. Purification of the Optimized Helicase Dda Mutant
[0089] Buffer A: 20mM Tris-HCl pH 7.5, 250mM NaCl, 20mM imidazole
[0090] Buffer B: 20mM Tris-HCl pH 7.5, 250mM NaCl, 300mM imidazole Buffer C: 20mM Tris-HCl pH 7.5, 80mM NaCl
[0091] Buffer D: 20mM Tris-HCl pH 7.5, 1000mM NaCl
[0092] Buffer E: 20mM Tris-HCl pH 7.5, 200mM NaCl
[0093] Bacteria expressing the optimized helicase Dda mutant, Dda-XX, were harvested and resuspended in Buffer A. The cells were disrupted using a cell disruptor, and the supernatant was collected by centrifugation. The supernatant was mixed with Ni-NTA medium equilibrated in Buffer A and allowed to bind for 1 hour. The medium was collected and washed extensively with Buffer A until all contaminants were removed. Next, Buffer B was added to the medium to elute the protein. The eluted protein was passed through a desalting column (Cytiva, Sephadex G-25) equilibrated in Buffer C, and the buffer was changed from Buffer B to Buffer C. The protein solution from the desalting column was then added to ssDNA cellulose medium equilibrated in Buffer C. An appropriate amount of coagulation protease was added. This enzyme specifically recognizes the thrombin cleavage site amino acid sequence LVPRGS in the vector sequence PET28(a)+, thereby cleaving the affinity His tag carried by the protein. This procedure was performed at 4°C and incubated overnight on a rotary shaker.
[0094] The next day, the ssDNA cellulose filler was collected. At this time, the target protein was specifically adsorbed to the ssDNA filler. The ssDNA cellulose filler was washed 3-4 times with Buffer C to remove impurities that were not adsorbed to the ssDNA cellulose filler. Then, it was eluted with buffer D to destroy the specific adsorption of the target protein to the ssDNA filler and elute the target protein into the solution. The protein purified from ssDNA cellulose was concentrated using a 30K ultrafiltration concentrator tube (Merck Millipore) in a centrifuge pre-cooled at 4°C. The parameters were set to 3000g speed, each centrifugation time was 10 minutes, and the final protein volume was concentrated to 2mL. Finally, it was passed through a molecular sieve Superdex 200 (Cytiva) using Buffer E as the molecular sieve buffer. The target protein peak was collected, concentrated, and frozen.
[0095] In this example, 11 Dda mutants were expressed and purified: Dda-1 (SEQ ID NO: 4), and the purification results are shown in Figure 1 . Dda-20 (K11A), and the purification results are shown in Figure 3 . Dda-32 (R191A), and the purification results are shown in Figure 5 . Dda-33 (R191S), and the purification results are shown in Figure 7 . Dda-45 (D189A), and the purification results are shown in Figure 9 . Dda-56 (K177H), and the purification results are shown in Figure 11 . Dda-58 (K177Y), and the purification results are shown in Figure 13 . Dda-62 (K177E), and the purification results are shown in Figure 15 . Dda-66 (K177A), and the purification results are shown in Figure 17 . Dda-70 (K177Q), and the purification results are shown in Figure 19 . The purification results of Dda-77 (K177S) are shown in Figure 21 .
[0096] Example 2 Construction and purification of the optimized helicase Dda mutant and DNA complex
[0097] First, two partially complementary DNA strands (top strand, SEQ ID NO:1 and bottom strand, SEQ ID NO:2) were annealed to form a linker. This annealed linker was then reacted with an optimized helicase Dda mutant in a 20mM HEPES (pH 7.2) and 50mM NaCl buffer system at 28°C in a metal bath for 1 hour to allow thorough mixing of the optimized helicase Dda mutant and DNA linker. A suitable amount of the oxidative crosslinker TMAD was then added, and the reaction was continued in a 28°C metal bath for 1.5 hours to stably crosslink the pin and tower domains of the helicase Dda mutant, thereby securing the linker in the resulting "arch" cavity and facilitating subsequent sequencing. Finally, an appropriate amount of unwinding buffer (50mM HEPES (pH 8.0), 1000mM KCl, 4mM ATP, and 20mM MgCl2) was added, and the reaction was continued in a 30°C metal bath for 1.5 hours to remove any non-target helicase Dda mutant-DNA complexes. Finally, magnetic beads (VAHTSTM DNA Clean Beads#N411-03) were used to purify and remove free proteins, cross-linking agents, and non-specifically bound non-target complexes in the reaction system.
[0098] Example 3 CsgG nanopore testing of the optimized helicase Dda mutant
[0099] 1. Incubate the helicase-containing sequencing library with single-stranded DNA (ssDNA-chol, SEQ ID NO: 3) containing cholesterol at its 5' end at room temperature for 10 minutes. The ssDNA-chol sequence is complementary to a portion of the bottom strand of the adaptor. Cholesterol binding to the phospholipid membrane reduces library loading and improves capture efficiency.
[0100] 2. Use a patch clamp amplifier (as shown in Figure 23) or other electrical signal amplifier to collect current signals. A Teflon membrane with a micrometer-sized pore (50-200 μm in diameter) in the center divides the electrolytic cell into two chambers: the cis chamber and the trans chamber. A pair of Ag / AgCl electrodes is placed in each chamber. A bilayer of phospholipid membranes is formed at the micropores of both chambers, and a nanopore protein (CsgG-Eco-(Y51A / F56Q / R97W / R192D-StrepII(C))) is added. Electrical measurements are obtained after a single nanopore protein is inserted into the phospholipid membrane. The sequencing library incubated in step 1 is added, and 180 mV is applied. The sequencing library is captured by the nanopore, and the nucleic acid passes through the nanopore under the control of the helicase. The buffer used in this experiment was: 470 mM KCl, 25 mM HEPES, 10 mM MgCl2, 30 mM ATP, pH = 8.10, and the sequencing temperature was 30°C.
[0101] Example 4 Characterization test of helicase Dda-1
[0102] The protein sequence of Dda-1 is SEQ ID NO:4 (the gene sequence is SEQ ID NO:5). This Dda mutant, serving as the basic modified control group, was cloned, expressed, and purified according to the methods described in Example 1. The protein exhibited moderate homogeneity and high purity, with a protein yield of approximately 0.3 mg / L. A DNA-helicase complex was constructed according to the complex construction method described in Example 2, and finally sequenced using the on-machine sequencing method described in Example 3. The sequencing speed was evaluated, and statistical and fitting analyses were performed. The resulting speed fitting plot is shown in Figure 2.
[0103] Example 5 Optimizing helicase to significantly increase sequencing speed
[0104] Dda-20, Dda-32, Dda-33, Dda-45, Dda-56, Dda-58, Dda-62, Dda-66, Dda-70, and Dda-77 were all mutated based on the sequence of Dda-1. The mutation information is shown in Table 1. These proteins were cloned, expressed, and purified according to the method in Example 1. The DNA and helicase complexes were constructed according to the method for constructing the complex in Example 2. Finally, the sequencing was performed according to the on-machine sequencing method of Example 3. The sequencing speed was evaluated, and statistical and fitting analysis were performed to obtain speed fitting graphs, which correspond to Figures 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, and 22, respectively. Based on the speed fitting, the average speed value was obtained, as shown in Table 1. It can be seen that these optimized helicases Dda have significantly improved sequencing speed. Compared with Dda-1, the speed of Dda-20 increased by more than 27%, the speed of Dda-32 increased by more than 27%, the speed of Dda-33 increased by more than 24%, the speed of Dda-45 increased by more than 83%, the speed of Dda-56 increased by more than 31%, the speed of Dda-58 increased by more than 31%, the speed of Dda-62 increased by more than 28%, the speed of Dda-66 increased by more than 41%, the speed of Dda-70 increased by more than 60%, and the speed of Dda-77 increased by more than 79%. The improvement of sequencing speed indicators is expected to increase the sequencing speed and overall sequencing throughput, and further promote the development of nanopore sequencing technology.
[0105] Dda-32 and Dda-33 underwent two different mutations at the R191 site. Both changes improved the properties of the helicase Dda, with the trend being to replace amino acids with long, positive side chains at this key site with either short, neutral or polar side chains. Since these two mutations did not produce significant differences in sequencing results, it is speculated that the optimization of this site primarily altered the length of its side chain, thereby improving overall sequencing data.
[0106] Dda-56, Dda-58, Dda-62, Dda-66, Dda-70, and Dda-77 underwent multiple mutations at the K177 site, all of which improved the properties of the helicase Dda. The trend of change was to replace amino acids with long positive side chains at this key site with those with shorter positive side chains, long polar amino acids, long negative amino acids, short neutral side chains, and short polar amino acids. The performance improvement was comparable for long polar amino acids and long negative amino acids, suggesting that replacing positive amino acids at this site with polar or negative amino acids is more beneficial for improving sequencing performance. The optimization of sequencing performance in these three directions—positive amino acids with short side chains, neutral amino acids with short side chains, and polar amino acids with short side chains—increased in a step-by-step manner. This suggests that shortening the amino acid side chain at this site is beneficial for improving sequencing performance, and that polar amino acids are superior to nonpolar amino acids.
[0107] Example 6 Optimizing Helicase to Improve Sequencing Speed Distribution Uniformity
[0108] When fitting the sequencing speed, the inventors also calculated the coefficient of variation CV of the sequencing speed as shown in Table 1. It can be seen that compared with the commonly used modified helicase Dda mutant Dda-1 on the market, under the same test conditions, the speed dispersion coefficients of the several optimized Dda helicase mutants mentioned in the embodiments of this patent: Dda-32, Dda-33, Dda-45, Dda-58, Dda-62, Dda-66 and Dda-76 are smaller than those of Dda-1, that is, the speed distribution is more concentrated and the uniformity is better. This performance improvement, to a certain extent, reflects the improvement in the quality of the overall sequencing output signal, the sequencing is more stable, and it also helps to decode the current signal and improve the sequencing accuracy, which is very important for the development of nanopore sequencing technology.
[0109] In particular, the speed of Dda-45 (D189A) was significantly improved compared to Dda-1, while maintaining a very uniform distribution of sequencing speeds, with a speed dispersion coefficient CV value of only 0.04, and the overall speed distribution was also relatively concentrated and showed a normal distribution.
[0110] Example 7 Optimizing Helicase to Increase Protein Yield
[0111] Wild-type Dda is difficult to express and produces very low yields. While Dda-1 was successfully expressed and purified, 1 L of expression bacteria only produced approximately 0.3 mg of the target protein. In this example, the optimized helicase of the present invention also exhibited superior protein yields, as summarized in Table 1.
[0112] Dda-20, Dda-62, Dda-66, and Dda-70 showed a 70-95% increase in protein yield compared to Dda-1. Dda-32 doubled the protein yield of Dda-1, while Dda-77 achieved nearly three times the protein yield of Dda-1. The most significant improvement was Dda-45, which achieved approximately 4.63 times the protein yield of Dda-1.
[0113] These proteins have good quality, good uniformity, high purity and excellent sequencing performance, which is very beneficial for nanopore sequencing applications.
[0114] Table 1 Summary of Dda mutant protein manifestations
[0115]
[0116] The sequence involved in the present invention is as follows:
[0117] SEQ ID NO: 1: 5'-XXXXXXXXXXXXXXXXXXXXXXXXXXXXTTTTTTTTTTTYYYYGGTTGTTTCTGTTGGTGCTGATATTGCT-3' (X=SPC3, Y=iSP18)
[0118] SEQ ID NO:2: 5'-GCAATATCAGCACCAACAGAAACAACCTTGAGGCGAGCGGTCAA-3'
[0119] SEQ ID NO:3: 5'-cholesterol-TTGACCGCTCGCCTC-3'
[0120] Amino acid sequence of Dda-1 (SEQ ID NO:4): MTFDDLTEGQKNAFNIVMKAIKEKKHHVTINGPAGTGKTTLTKFIIEALISTGETGIILAAPTHAAKKILSKLSGKEASTIHSILKINPVTYECNVLFEQKEVPDLAKARVLICDEVSMYDRKLFKILLSTIPPWATIIGIGDNKQIRPVDPGENTAYISPFFTHKDFYQCELTEVKRSNAPIIDVATDVRNGKWIYDKVVDGHGVRGFTGDTALRDFMVNYFSIVKSLDDLFENRVMAFTNKSVDKLNSIIRKKIFETDKDFIVGEIIVMQEPLFKTYKIDGKPVSEIIFNNGQLVRIIEAEYTSTFVKARGVPGEYLIRHWDLTVETYGDDEYYREKIKIISSDEELYKFNLFLGKTCETYKNWNKGGKAPWSDFWDAKSQFSKVKALPASTFHKAQGMSVDRAFIYTPCIHYADVELAQQLLYVGVTRGRYDVFYV
[0121] Nucleic acid sequence of Dda-1 (SEQ ID NO: 5): ATGACGTTTGATGATCTGACGGAAGGTCAGAAAAATGCCTTTAATATCGTTATGAAAGCCATCAAAGAAAAAAAACATCATGTGACGATCAATGGTCCAGCGGGAACCGGCAAAACGACACTGACCAAATTTATCATCGAAGCGCTGATTAGTACGGGCGAAACGGGTATCATCCTGGCCGCTCCGACACATGCCGCCAAAAAAATTCTGTCTAAACTGTCAGGCAAAGAAGCCTCTACAATTCATAGCATCCTGAAAATCAATCCAGTGACCTATGAATGCAATGTGCTGTTTGAACAGAAAGAAGTTCCGGATCTGGCCAAAGCCCGCGTGCTGATTTGCGATGAAGTGAGCATGTATGATCGCAAACTGTTTAAAATCTTACTGAGTACCATTCCTCCTTGGGCTACCATCATCGGTATTGGCGATAATAAACAGATTCGTCCGGTTGATCCAGGCGAAAATACGGCCTATATTAGTCCGTTTTTTACACATAAAGATTTTTATCAGTGCGAACTGACGGAAGTTAAACGTAGTAATGCACCTATTATCGATGTTGCTACAGATGTTCGTAATGGCAAATGGATCTATGATAAAGTTGTGGATGGGCATGGGCTGCGCGGCTTTACAGGTGATACAGCCCTGCGCGATTTTATGGTTAATTATTTTAGCATTGTGAAATCTCTGGATGATCTGTTTGAAAATCGCGTTATGGCCTTTACCAATAAATCAGTGGATAAACTGAATAGTATTATCCGCAAAAAAATCTTTGAAACCGATAAA GATTTTATTGTGGGCGAAATCATTGTTATGCAGGAACCACTGTTTAAAACATATAAAATCGATGGTAAACCGGTGTCAGAAATTATTCTTTAATAATGGTCAGCTGGTTCGCATCATCGAAGCGGAATATACGAGTACGTTTGTTAAAGCACGCGGCGTTCCAGGCGAATATCTGATTCG CCATTGGGATCTGACCGTGGAAACGTATGGTGATGATGAATATTATCGTGAAAAAATCAAAATCATTAGCTCAGATGAAGAACTGTATAAATTTAATCTGTTTCTGGGCAAAACCTGCGAAACCTATAAAAATTGGAATAAAGGTGGCAAAGCTCCTTGGAGCGATTTTTGGGATGCCA AATCTCAGTTTTCTAAAGTTAAAGCCCTGCCGGCCTCTACCTTTCATAAAGCCCAGGGCATGAGCGTTGATCGTGCCTTTATCTATACCCCTTGCATTCATTATGCAGATGTGGAATTAGCGCAGCAGCTGCTGTATGTGGGTGTGACACGCGGTCGCTATGATGTGTTTTATGTGTAA
[0122] The gene sequence of Dda-20 is based on Dda-1, with AAA at positions 31-33 replaced by GCA;
[0123] The gene sequence of Dda-32 is based on Dda-1, with CGT at positions 571-573 replaced by GCC;
[0124] The gene sequence of Dda-33 is based on Dda-1, with CGT at positions 571-573 replaced by AGT;
[0125] The gene sequence of Dda-45 is based on Dda-1, with GAT at positions 565-567 replaced by GCC;
[0126] The gene sequence of Dda-56 is based on Dda-1, with AAA at positions 529-531 replaced by CAT;
[0127] The gene sequence of Dda-58 is based on Dda-1, with AAA at positions 529-531 replaced by TAT;
[0128] The gene sequence of Dda-62 is based on Dda-1, with AAA at positions 529-531 replaced by GAA;
[0129] The gene sequence of Dda-66 is based on Dda-1, with AAA at positions 529-531 replaced by GCC;
[0130] The gene sequence of Dda-70 is based on Dda-1, with AAA at positions 529-531 replaced by CAG;
[0131] The gene sequence of Dda-77 is based on Dda-1, with AAA at positions 529-531 replaced by AGT.
[0132] Although the above describes specific embodiments of the present invention, it should be understood by those skilled in the art that these are merely illustrative and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present invention. Therefore, the scope of protection of the present invention is defined by the appended claims.
Claims
1. A mutant of helicase Dda, characterized in that The mutant is one or more mutations selected from the following on the amino acid sequence shown in SEQ ID NO: 4: (1) replacing the positively charged amino acids in the amino acid sequence with neutral amino acids; (2) replacing the negatively charged amino acids in the amino acid sequence with neutral amino acids; (3) truncating the amino acids in the N-terminal region of the amino acid sequence or replacing them with short side chain neutral amino acids; (4) Substituting amino acids with larger side chain steric hindrance in the amino acid sequence with amino acids with shorter side chain.
2. The mutant according to claim 1, characterized in that The mutation site is selected from one or more of the following: (1), the mutated sites include K11, K19, K22, K24, K25, K38, K43, K67, K68, K72, K76, K86, K101, K108, K123, K126, K145, K166, K177, K194, K199, K227, K243, K247, K254, K255, K261, K277, K280, K284, K310, K339, K341, one or more of K351, K358, K364, K368, K371, K381, K386, K388, K397, R110, R122, R148, R191, R207, R216, R236, R253, R298, R312, R321, R337, R405, R431, R433, H26, H27, H64, H82, H165, H204, H322, H396, and H414; Preferably, the mutated sites include K11 deletion, K11A, K11L, K11I, K19A, K177A, K194A, K199A, K227A, K254A, K386A, K397A, R191A, R216A, R405A, H204A, H322A, K11S, K19S, K177S, K194S , K199S, K227S, K194I, K199L, K254L, R191S, R207A, R207S, R216S, H204S, R216I, R236A, R253A, K11G, K177G, K194G, K199G, K227G, R191G, R207G, and R405S; (2), the mutated sites include D4, D5, D105, D143, D167, D185, D189, D198, D202, D212, D217, D230, D231, D246, D260, D262, D282, one or more of D324, D332, D333, D346, D376, D379, D404, D417, D435, E8, E23, E47, E54, E77, E99, E102, E116, E154, E172, E175, E234, E258, E267, E273, E288, E301, E303, E317, E328, E334, E338, E347, E348, E361, and E419; Preferably, the mutated sites include D4 deletion, D5 deletion, E8 deletion, D4A, D5A, D185A, D189A, D198A, D202A, D212A, D217A, D346A, E8A, E54A, E172A, E175A, E234A, E258A, E288A, D5S, D185S, D189S, D198S, D202S, D212S, one or more of E8S, E8I, E54S, E172S, E175S, D217S, D5L, D185L, D189L, D198L, D202L, D212L, E8L, D5I, D4S, D5S, D4G, D5G, D185G, D189G, D198G, D202G, D212G, E8G, E54G, E192G, E175G, and E54L; (3), wherein the mutated sites include one or more of M1, T2, F3, L6, T7, G9 and Q10; Preferably, the mutated sites include one or more of M1 deletion, F3 deletion, L6 deletion, Q10 deletion, M1-K11 deletion, M1-F3 deletion, M1-T7 deletion, M1A, T2A, Q10A, Q10S, Q10L, T7I and Q10I; (4), wherein the mutated sites include one or more of M219, M238, M271, W195, W323, W366, F3, F14, F209, F218, F223, F233, F240, F257, F263, Y169, Y197, Y279, Y304, Y318, Q10, Q272, N12, N14, N88, N180, N192, N221, N235, N242, and N249; Preferably, the mutated sites include one or more of F3A, F209A, F218A, F223A, F233A, F240A, F257A, F263A, Y197A, Y279A, N12A, N14A, N192A, N221A, N235A, N242A, N249A, F3S, F257S, Y197S, N192S, N221S, F3G, Y197G, H204G, Q10G, N12G, N192G and N221G.
3. The mutant according to claim 1 or 2, characterized in that The mutation comprises one or more of the following amino acid substitutions: D5, E8, K11, K177, D189, R191, D202, and R216; Preferably, the mutation is selected from one or more of the following groups: (i) D5T, D5H, D5Y, D5Q or D5N; (ii) E8T, E8H, E8Y, E8Q or E8N; (iii) K11A, K11T, K11H, K11Y, K11Q or K11N; (iv) K177E, K177A, K177Q, K177S, K177L, K177T, K177H, K177Y or K177N; (v) D189A, D189T, D189H, D189Y, D189Q or D189N; (vi) R191A, R191S, R191L, R191T, R191H, R191Y, R191Q or R191N; (vii) K199Q; (viii) D202T, D202Q, D202Y, D202H or D202N; (ix) R216L, R216G, R216T, R216H, R216Y, R216Q or R216N.
4. The mutant according to any one of claims 1 to 3, characterized in that The mutation is K11A, R191A, R191S, D189A, K177H, K177Y, K177E, K177A, K177Q or K177S.
5. A nucleic acid, characterized in that The nucleic acid encodes the mutant according to any one of claims 1 to 4; Preferably, the sequence of the nucleic acid is a sequence as shown in SEQ ID NO:5 in which one or more nucleotides are replaced, deleted or inserted.
6. The nucleic acid according to claim 5, characterized in that The mutation is to replace CGT at position 571-573 with GCC; or, Replace CGT with AGT at positions 571-573; or, Replace GAT with GCC at positions 565-567; or, Replace AAA with CAT at positions 529-531; or, Replace AAA with TAT at positions 529-531; or, Replace AAA with GAA at positions 529-531; or, Replace AAA with GCC at positions 529-531; or, Replace AAA with CAG at positions 529-531; or, Replace AAA with AGT at positions 529-531.
7. A recombinant expression vector, characterized in that: The recombinant expression vector comprises the nucleic acid according to claim 5 or 6; Preferably, the backbone plasmid of the recombinant expression vector is PET.28a(+), PET.21a(+) or PET.32a(+).
8. A transformant, characterized in that: The transformant comprises the nucleic acid according to claim 5 or 6 or the recombinant expression vector according to claim 7; Preferably, the host cell of the transformant is Escherichia coli; More preferred is BL21(DE3), BL21 Star(DE3)pLyss, Rossata(DE3) or Lemo21(DE3).
9. A method for preparing the mutant according to any one of claims 1 to 4, characterized in that: The transformant according to claim 8 is cultured in a culture medium to ferment and produce the mutant.
10. A helicase-adapter complex, characterized in that The helicase-adapter complex comprises the mutant according to any one of claims 1 to 4 and a sequencing adapter, wherein the sequencing adapter comprises two DNA chains with complementary partial regions.
11. A kit, characterized in that: The kit comprises the mutant according to any one of claims 1 to 4; preferably further comprises one or more of the following: single-stranded DNA containing a cholesterol anchor molecule at the 5' end, a nanopore, a nanopore protein, an electrical signal detector, a membrane and / or a sequencing buffer.
12. The kit according to claim 11, characterized in that The anchoring molecule is a hydrophobic molecule, preferably selected from any one or more of the following: lipids, fatty acids, sterols, carbon nanotubes, polypeptides, proteins and / or amino acids, such as cholesterol, palmitate or tocopherol; the membrane is an amphiphilic membrane, a high molecular polymer membrane, such as a diblock copolymer di-block, a triblock copolymer tri-block, or any combination thereof; the nanopore is a transmembrane protein pore or a solid pore; the transmembrane protein pore is selected from hemolysin, MspA, MspB, MspC, MspD, FraC, ClyA, PA63, CsgG, CsgD, XcpQ, SP1, phi29 connector protein, InvG, GspD or any combination thereof; the sequencing buffer is a dihydrogen phosphate-hydrogen phosphate buffer system, a carbonate-sodium bicarbonate buffer system, a Tris-HCl buffer system, a HEPES buffer system, a MOPS buffer system or any combination thereof.
13. A DNA unwinding method, characterized in that: The method comprises using the mutant according to any one of claims 1 to 4 or the helicase-adapter complex according to claim 10 to unwind double-stranded DNA.
14. A sequencing method, characterized in that: It includes the following steps: The double-stranded DNA is sequenced while being unwound using the DNA unwinding method as described in claim 13.
15. Use of the mutant according to any one of claims 1 to 4, the nucleic acid according to claim 5 or 6, the recombinant expression vector according to claim 7, the transformant according to claim 8, the helicase-linker complex according to claim 10, or the kit according to claim 11 or 12 in unwinding or sequencing or in preparing unwinding or sequencing reagents; Preferably, the sequencing is nanopore sequencing.