Helicase as well as preparation method and application thereof in high-throughput sequencing

CN120344656APending Publication Date: 2025-07-18BGI HANGZHOU CYCLONESEQ TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102499.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The unwinding activity of helicases in existing nanopore sequencers decreases in high-salt environments, resulting in reduced sequencing speed and efficiency. There is a lack of better-performing helicases to improve sequencing performance.

Method used

A new helicase was developed. Its gene is derived from the deep-sea metagenome. It has high salt tolerance and stability. It improves protein uniformity and ATP hydrolysis activity through the mutation of amino acids and unnatural amino acids, and is suitable for high salt environments. DNA unwinding and nanopore sequencing.

Benefits of technology

This new helicase exhibits superior unwinding activity in high-salt environments, can stably control the perforation movement of nucleic acid molecules, improves the uniformity and efficiency of sequencing, and is suitable for high-throughput nanopore sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000020_0000
    Figure 00000020_0000
  • Figure 00000020_0001
    Figure 00000020_0001
  • Figure 00000020_0002
    Figure 00000020_0002
Patent Text Reader

Abstract

The invention provides helicase as well as a preparation method and application thereof in high-throughput sequencing. The amino acid sequence of the helicase is as shown in SEQ ID NO: 1 or has at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% of identity with the amino acid sequence as shown in SEQ ID NO: 1. The helicase still has high unwinding activity in a high-salt environment, can be used for control and characterization of nucleic acid, and is applied to nanopore sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

A helicase and its preparation method and application in high-throughput sequencing Technical Field

[0001] The present invention belongs to the field of biotechnology or sequencing, and particularly relates to a helicase and an application thereof. Background Art

[0002] Nanopore sequencing, an emerging single-molecule sequencing technology, has revolutionized the genetic sequencing industry with its unique advantages, including high throughput, long read lengths, rapid speed, in situ detection, and label-free operation. Devices based on this technology are lightweight and portable, adapting to diverse sequencing scenarios. Furthermore, due to its non-amplified direct sequencing nature, there is no length limit on the sequenceable DNA. Furthermore, real-time base calling enables direct sequencing of modified molecules such as RNA and methylation, as well as other single molecules. Nanopore sequencing technology has broad applications in many fields, including molecular biology, medicine, epidemiology, and ecology, such as genome mapping, epidemic monitoring, rare species detection, and rapid and cost-effective protein sequencing.

[0003] The principle of nanopore sequencing technology is based on changes in electrical signals. A nanopore inserted into a membrane (protein or solid) serves as a signal sensor to separate two electrolyte chambers. When voltage is applied to the two electrolyte chambers, a stable current is generated through the nanopore. When the nucleic acid molecule to be tested enters the nanopore, the flow of ions is hindered, resulting in fluctuations in the current signal. Because different nucleotides have different effects on the current, the sequence information of the nucleic acid molecule to be tested is identified by detecting the current fluctuation signal of the nanopore in real time and analyzing and decoding the current signal with the help of machine learning.

[0004] During the sequencing process, due to the extremely high speed at which nucleic acid molecules pass through the nanopore channel, it is impossible to accurately obtain the sequence information of the nucleic acid molecules. Therefore, effectively reducing and controlling the perforation motion of nucleic acid molecules is a key technical issue in achieving nanopore sequencing. Currently, the most common and effective method is to use helicase unwinding to control the perforation motion of nucleic acid molecules, improve detection accuracy, and maintain sequencing speed and sequencing uniformity. In addition to having highly conserved regions (mainly II and VI regions) that bind to ATP, exert ATP hydrolase activity, and provide power for unwinding, helicase also has two key domains: Pin (pin structure) and Tower (tower structure). These two structures interact to form an "arch" structure that allows single-stranded DNA to pass through it. In addition, Pin also plays a role in unwinding double-stranded DNA into single-stranded DNA and assisting in promoting the directional movement of nucleic acid single strands.

[0005] The helicase currently used in commercial nanopore sequencers is the DDA helicase derived from bacteriophage T4. This helicase has inherent limitations, such as its stability, salt tolerance, and unwinding speed. In particular, high salt levels inhibit the unwinding activity of the DDA helicase, reducing its unwinding speed and its inability to fully utilize its unwinding capacity. This, in turn, impairs the sequencing speed and efficiency of nanopore sequencing applications. Therefore, there is a market demand for newer, higher-performing helicases.

[0006] Summary of the Invention

[0007] To address the problem of the lack of better helicase performance in the prior art, the present invention provides a novel helicase with high salt tolerance and stability, a preparation method thereof, and its application in high-throughput sequencing. Specifically, the present invention provides:

[0008] a) A helicase, whose amino acid sequence is as shown in SEQ ID NO: 1 or has at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identity with the amino acid sequence shown in SEQ ID NO: 1. The helicase can be successfully expressed in an Escherichia coli recombinant protein expression system, the protein itself has high homogeneity and purity, and has good ATP hydrolysis activity and dsDNA unwinding activity. Because its gene is derived from the deep-sea metagenome, the proteins derived from this genome have extremely high thermal stability and salt tolerance. Therefore, the helicase exhibits better unwinding activity in a high-salt environment than in a low-salt environment, can bind well to single-stranded DNA, and unwind double-stranded DNA. The helicase can be used for the control and characterization of nucleic acids, and is applied to single-molecule nanopore sequencing to output a stable sequencing current signal.

[0009] At least one cysteine ​​on the surface of the three-dimensional structure of the helicase is mutated, and the mutation is that the cysteine ​​is replaced by alanine, glutamine, glycine, histidine, isoleucine, leucine, valine, serine, threonine or methionine, so as to improve the uniformity of the protein, thereby improving indicators such as sequencing uniformity.

[0010] Preferably, the mutated sites include at least one of C21, C50, C56, C91, C156, C279, C367 and C379.

[0011] The helicase has at least one amino acid mutation in the pin domain and / or the tower domain, wherein the amino acid mutation is a substitution of the original amino acid with cysteine ​​or a non-natural amino acid.

[0012] Preferably, the mutation site of the pin domain is at least one of 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102 and 103; and / or the mutation site of the tower domain is at least one of 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367 and 368.

[0013] More preferably, the mutation site of the pin domain is at least one of G85, I86, S87, P88, T89, V90, D91, K92, K93, E94, L95, E96, F97, E98, H99, V97, N98, I99, P100, S101, L102 and W103; and / or the mutation site of the tower domain is at least one of L337, Y338, At least one of E339, V340, A341, N342, Y343, Y344, D345, Y346, Q347, Q348, 1347, A348, D349, Y350, Y351, E352, H353, 1354, A355, W356, N357, M358, K359, T360, P361, Q362, A363, K364, A365, K366, A367, and Y368.

[0014] Such unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (PAZF), 4-azido-L-phenylalanine (PAZF-HCl), 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselenoyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4-(dihydroxyboryl)-L-phenylalanine, 4-[(ethylsulfanyl)carbonyl]-L-phenylalanine, (2S)-2 -amino-3-{4-[(propan-2-ylsulfanyl)carbonyl]phenyl}propanoic acid, (2S)-2-amino-3-{4-[(2-amino-3-sulfanylpropionyl)amino]phenyl}propanoic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine, 3-hydroxy-L-tyrosine, 3 -amino-L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthyl-2-ylamino)propionic acid, 6-(methylsulfanyl)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propionic acid, (2R)-2-aminooctanoate 3-(2,2′-bipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxy- At least one of (2R)-2-amino-3-[(2-nitrobenzyl)sulfanyl]propionic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propionic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)oxy]carbonyl}amino)hexanoic acid, O-(2-nitrobenzyl)-L-tyrosine and 2-nitrophenylalanine.

[0015] The helicase has at least one mutation in the DNA binding region and / or near the ATP catalytic active center, wherein the mutation includes substitution of the original amino acid with an amino acid with a larger side chain, such as increasing the number of carbon atoms, increasing the length and / or increasing the molecular volume, thereby increasing (i) electrostatic interaction; (ii) hydrogen bonding and / or (iii) cation-pi (cation-π) interaction between at least one amino acid and one or more nucleotides in ssDNA.

[0016] Preferably, the mutation site in the DNA binding region is at least one of 63, 73, 79, 80, 81, 82, 83, 84, 85, 86, 87, 95 and 96; and / or the mutation site near the ATP catalytic active center is at least one of 154, 155, 156, 158, 159, 160, 161, 174, 175, 177, 178, 179 and 181.

[0017] More preferably, the mutation site in the DNA binding region is at least one of H63, E73, T79, V80, H81, S82, A83, L84, G85, I86, S87, L95 and E96; and / or the mutation site near the ATP catalytic active center is at least one of D154, D155, P156, I158, S159, P160, V161, P174, M175, N177, T178, G179 and L181.

[0018] The helicase undergoes at least one mutation in the DNA binding region and the nanopore binding region, wherein the mutation includes replacement of the original amino acid by an amino acid with a non-positive surface charge and / or an amino acid with a side chain length shorter than the original amino acid, thereby reducing the repulsion between the motor protein and the pore, etc.

[0019] Preferably, the mutation occurs in at least one of positions 1, 2, 3, 5, 7, 8, 9, 10, 36, 42, 45, 67, 74, 97, 207, 208, 209, 220, 221, 222, 374, 408 and 411 of the DNA binding region and the nanopore binding region.

[0020] More preferably, the mutation occurs in at least one of M1, N2, S3, N5, D7, Q8, Q9, K10, K36, K42, K45, D67, K74, F97, R207, K208, D209, K220, K221, D222, H374, D408 and K411 in the DNA binding region and the nanopore binding region.

[0021] b) An isolated nucleic acid encoding a helicase as described in a).

[0022] Preferably, the nucleotide sequence of the helicase is as shown in SEQ ID NO:2 or has at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identity with the nucleotide sequence shown in SEQ ID NO:2.

[0023] c) A recombinant expression vector comprising a promoter and the nucleic acid as described in b).

[0024] Preferably, the promoter is T7; and / or the backbone plasmid of the recombinant expression vector is PET.28a(+), PET.21a(+) or PET.32a(+).

[0025] d) A transformant comprising a host cell and the nucleic acid as described in b).

[0026] Preferably, the host cell is Escherichia coli, more preferably BL21(DE3), BL21Star(DE3)pLyss, Rossata(DE3) or Lemo21(DE3).

[0027] e) A method for preparing the helicase described in a), comprising culturing the transformant described in d) in a culture medium to ferment and produce the helicase.

[0028] f) A helicase-sequencing adapter complex, comprising the helicase as described in a) and a sequencing adapter.

[0029] g) A kit comprising the helicase described in a) and / or the helicase-sequencing adapter complex described in f); preferably further comprising a single-stranded DNA containing an anchor molecule at the 5' end, a nanopore, a nanopore protein, an electrical signal detector, a membrane and / or a buffer.

[0030] Preferably, the anchoring molecule is a hydrophobic molecule, preferably selected from any one or more of the following: lipids, fatty acids, sterols, carbon nanotubes, polypeptides, proteins and / or amino acids, such as cholesterol, palmitate or tocopherol.

[0031] The nanopore is a transmembrane protein pore or a solid-state pore; preferably, the transmembrane protein pore is selected from hemolysin, MspA, MspB, MspC, MspD, FraC, ClyA, PA63, CsgG, CsgD, XcpQ, SP1, phi29 connector protein (phi29connector), InvG, GspD or any combination thereof.

[0032] The membrane is an amphiphilic membrane (such as a phospholipid bilayer), a high molecular polymer membrane (such as a diblock copolymer, a triblock copolymer) or any combination thereof.

[0033] The buffer is a dihydrogen phosphate-hydrogen phosphate buffer system, a carbonic acid-sodium bicarbonate buffer system, a Tris-HCl buffer system, a HEPES buffer system, a MOPS buffer system or any combination thereof.

[0034] h) Use of the helicase described in a), the helicase-sequencing adapter complex described in f), or the kit described in g) in high-throughput sequencing.

[0035] Preferably, the high-throughput sequencing is nanopore sequencing.

[0036] i) A DNA unwinding method, comprising unwinding double-stranded DNA using the helicase described in a), the helicase-sequencing adapter complex described in f), or the kit described in g).

[0037] j) A sequencing method comprising the following steps:

[0038] The DNA is unwound using the DNA unwinding method described in i), and the obtained single-stranded DNA is sequenced simultaneously.

[0039] The beneficial effects brought about by the technical solution of the present invention are:

[0040] The present invention provides a new helicase, named BCH666, whose gene is derived from the deep-sea metagenome. The protein itself has good salt tolerance and stability, as well as DNA unwinding activity. It can have high unwinding activity in high-salt environments and can be used for the control and characterization of nucleic acids and applied to nanopore sequencing. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] FIG1 shows the results of the molecular sieve Superdex 200 purification of BCH666 protein; FIG1A shows the elution profile of the molecular sieve Superdex 200 purification of BCH666; FIG1B shows the SDS-PAGE electrophoresis results obtained after molecular sieve elution of BCH666.

[0042] FIG2 is a diagram of the structure of the BCH666 protein predicted using Alphafold 2 software.

[0043] FIG3 shows the ATPase activity detection results of BCH666 protein.

[0044] FIG4 shows the dsDNA melting activity test results of BCH666 protein (low salt reaction buffer 1).

[0045] FIG5 shows the dsDNA melting activity test results of BCH666 protein (high salt reaction buffer 2).

[0046] FIG6 shows the detection result of the restriction sequence blocking the depolymerization activity of BCH666 protein (low salt reaction buffer).

[0047] FIG7 shows the detection result of the restriction sequence blocking the BCH666 protein depolymerization activity (high salt reaction buffer).

[0048] FIG8 is a schematic diagram of the connector structure (a: upper chain; b: lower chain).

[0049] Figure 9 is a schematic diagram of the structure of a sequencing library containing helicase (a: upper chain; b: lower chain; c: double-stranded target fragment; d: helicase; e: cholesterol-labeled double-stranded DNA).

[0050] FIG10 is a schematic diagram of an electrical signal amplifier.

[0051] FIG11 is a diagram of the current signal obtained when the BCH666 protein was used for sequencing. DETAILED DESCRIPTION

[0052] Example 1 Cloning, expression and purification of BCH666 protein

[0053] 1. Cloning and Expression of BCH666 Protein

[0054] The amino acid sequence of the BCH666 protein is shown in SEQ ID NO: 1, and its full-length DNA sequence is shown in SEQ ID NO: 2. The full-length DNA sequence was synthesized (Liuhe BGI) and ligated into the PET.28a(+) plasmid using the double restriction enzyme sites Nde1 and Xho1. The resulting plasmid was labeled PET.28a(+)-BCH666. The BCH666 protein expressed from this plasmid has a thrombin restriction site and a 6×His tag at its N-terminus.

[0055] Transform the PET.28a(+)-BCH666 plasmid into Escherichia coli BL21(DE3) or its derivatives. Pick a single colony and inoculate it into 20 mL of LB medium containing kanamycin resistance. Cultivate with shaking at 37°C overnight. Then, inoculate the colony into 2 L of LB medium containing kanamycin resistance and incubate with shaking at 37°C until the OD600 reaches 0.6-0.8. Cool the culture to 16°C and induce expression overnight with IPTG at a final concentration of 500 μM to obtain BCH666 cells.

[0056] 2. Purification of BCH666 Protein

[0057] The buffer solution used is as follows:

[0058] (1) Buffer A: 20mM Tris-HCl pH 7.5, 250mM NaCl, 20mM imidazole;

[0059] (2) Buffer B: 20 ​​mM Tris-HCl pH 7.5, 250 mM NaCl, 300 mM imidazole;

[0060] (3) Buffer C: 20mM Tris-HCl pH 7.5, 80mM NaCl;

[0061] (4) Buffer D: 20mM Tris-HCl pH 7.5, 1000mM NaCl;

[0062] (5) Buffer E: 20mM Tris-HCl pH 7.5, 200mM NaCl.

[0063] BCH666 cells were harvested, resuspended in buffer A, disrupted using a cell disruptor, and centrifuged to collect the supernatant. The supernatant was mixed with Ni-NTA filler pre-equilibrated with buffer A and allowed to bind for 1 hour. The filler was collected and washed extensively with buffer A until all contaminants were removed. Buffer B was then added to the filler to elute the protein. The eluted protein was passed through a HiTrap desalting column (Sephadex G-25, Cat. No. 29048684, Cytiva) equilibrated in buffer C. The protein solution from the desalting column was then added to ssDNA cellulose filler equilibrated in buffer C. Thrombin was added to an appropriate amount of thrombin and digested at 4°C overnight. Thrombin specifically recognizes the thrombin cleavage site amino acid sequence LVPRGS in the vector sequence PET28(a)+, thereby removing the affinity His tag from the protein. The ssDNA cellulose filler was collected, and the target protein was now specifically adsorbed to the filler. Wash the ssDNA cellulose packing material 3-4 times with Buffer C to remove any unadsorbed proteins. Then, elute with Buffer D to disrupt the specific adsorption of the target protein to the ssDNA packing material and elute the target protein into the solution. The protein purified from the ssDNA cellulose was concentrated using a 30K ultrafiltration concentrator (Merck Millipore) in a 4°C pre-cooled centrifuge at 3000g for 10 minutes each. Repeat this process until the final protein volume is reduced to 2 mL. The protein was then filtered through a Superdex 200 molecular sieve using Buffer E. The target protein peak was collected, concentrated, and frozen.

[0064] Figure 1 shows the results of the purification of BCH666 protein using Superdex 200 molecular sieves. As shown in Figure 1, a relatively large amount of BCH666 protein with good purity was obtained after purification. The protein exhibited uniform peak shape and high purity.

[0065] 3. Use AlphaFold 2 to predict the structure of BCH666 protein

[0066] The structure of BCH666 protein was predicted using AlphaFold 2 software, and the results are shown in Figure 2. The root mean square error (RMSD) between the predicted value and the true value of the protein skeleton structure is The protein consists of a helix, a sheet, and a loop. The structure of the BCH666 protein is similar to that of conventional 5'-3' helicases, with a long and flexible pin domain.

[0067] Example 2 ATPase activity detection of BCH666 protein

[0068] SEQ ID NO: 3-8 were synthesized (Liuhe BGI).

[0069] 1. Preparation of double-stranded DNA (ovDNA-1) and single-stranded DNA (ssDNA)

[0070] Dissolve SEQ ID NO:3 and SEQ ID NO:4 in TE buffer (pH = 8) to a final concentration of 100 μM. Anneal SEQ ID NO:3 and SEQ ID NO:4 to ovDNA-1 with 20 T residues at the 5' end, at a concentration of 10 μM. The annealing process was incubation at 95°C for 5 minutes, followed by a cooling rate of 0.1°C / s to 25°C, and continued incubation for 30 minutes. The annealing recipe is shown in Table 1. Dilute 100 μM of SEQ ID NO:4 to 10 μM in TE buffer (pH = 8) to prepare ssDNA.

[0071] Table 1 ovDNA-1 annealing formula

[0072] Solution volume 100 μM SEQ ID NO: 3 5 μL 100 μM SEQ ID NO: 4 5 μL TE buffer (pH = 8) 40 μL

[0073] 2. Prepare high salt reaction buffer

[0074] High salt reaction buffer (2×): 20 mM HEPES (pH 8.0), 4 mM ATP, 4 mM MgCl2, 1.0 M KCl.

[0075] 3. Dilute protein

[0076] Dilute BCH666 protein to 10 μM in 1× PBS.

[0077] 4. ATP hydrolysis reaction

[0078] Add the corresponding reagents according to the reaction system in Table 2, incubate at 30°C for 30 minutes for ATP hydrolysis test, and inactivate at 80°C for 5 minutes. ①② are experimental groups, ③④⑤⑥ are corresponding control groups, and each group has 3 replicates.

[0079] Table 2 ATP hydrolysis reaction system

[0080] No. Reaction buffer (2×) DNA BCH666 Protein H2O ①10μL1μL (ovDNA-1)1μL8μL ②10μL1μL (ssDNA)1μL8μL ③10μL——1μL9μL ④10μL1Μl (ovDNA-1)——9μL ⑤10μL1μL (ssDNA)——9μL

[0081] ⑥10μL————10μL

[0082] 5. Detection of Remaining ATP in the Reaction

[0083] The ATP concentration remaining in the reaction was determined using an ATP detection kit (Beyotime, S0026B) according to the manufacturer's instructions.

[0084] 6. Experimental Results

[0085] The results are shown in FIG3 . Under high salt conditions, the BCH666 protein has the activity of hydrolyzing ATP.

[0086] Example 3 Detection of dsDNA Melting Activity of BCH666 Protein

[0087] 1. Preparation of double-stranded DNA (ovDNA-2)

[0088] Dissolve SEQ ID NO:5, SEQ ID NO:6, and SEQ ID NO:7 in TE buffer (pH 8) to a final concentration of 100 μM. Anneal SEQ ID NO:5 and SEQ ID NO:6 to ovDNA-2 with 20 T residues at the 5' end at a concentration of 10 μM. The annealing process was incubation at 95°C for 5 minutes, followed by a cooling rate of 0.1°C / s to 25°C, and incubation for 30 minutes. The annealing recipe is shown in Table 3.

[0089] Table 3 ovDNA-2 annealing formula

[0090] Solution volume 100 μM SEQ ID NO: 5 5 μL 100 μM SEQ ID NO: 6 5 μL TE buffer (pH = 8) 40 μL

[0091] 2. Prepare reaction buffer

[0092] Low salt reaction buffer: 100 mM HEPES (pH = 8.0), 1 mg / mL BSA, 10 mM MgCl2, 150 mM KCl;

[0093] High salt reaction buffer: 100 mM HEPES (pH=8.0), 1 mg / mL BSA, 10 mM MgCl2, 500 mM KCl.

[0094] 3. Prepare the reaction solution

[0095] Experimental reaction solution: Add 3 μL of 10 μM ovDNA-2, 6 μL of 100 μM SEQ ID NO:7 (20x competitor DNA), and 6 μL of 100 mM ATP to 582 μL of low salt reaction buffer. Add 3 μL of 10 μM ovDNA-2, 6 μL of 100 μM SEQ ID NO:7 (20x competitor DNA), and 6 μL of 100 mM ATP to 582 μL of high salt reaction buffer. SEQ ID NO:7 acts as excess capture DNA, preferentially annealing with complementary DNA to prevent re-annealing of the initial substrate and loss of fluorescence.

[0096] Positive control: Add 1 μL of 10 μM SEQ ID NO: 6, 2 μL of 100 μM SEQ ID NO: 7 (20x competitor DNA), and 2 μL of 100 mM ATP to 195 μL of low-salt reaction buffer. Add 1 μL of 10 μM SEQ ID NO: 6, 2 μL of 100 μM SEQ ID NO: 7 (20x competitor DNA), and 2 μL of 100 mM ATP to 195 μL of high-salt reaction buffer.

[0097] 4. Dilute protein

[0098] Dilute BCH666 protein to 4.8 μM with 1× PBS.

[0099] 5. Prepare the Melting Reaction

[0100] Add the corresponding reagents according to Table 4. ①② are experimental groups, ③④ are negative control groups, and ⑤⑥ are positive control groups. Use a microplate reader to detect the kinetic changes of fluorescence intensity within 30 minutes at 30°C. Each group has 3 replicates.

[0101] Table 4 Melting reaction formula

[0102] No. Category Solution 1 Solution 2 ① Experimental group 58.5μL experimental reaction solution (low salt) 1.5μL protein ② Experimental group 58.5μL experimental reaction solution (high salt) 1.5μL protein ③ Negative control group 58.5μL experimental reaction solution (low salt) 1.5μL PBS ④ Negative control group 58.5μL experimental reaction solution (high salt) 1.5μL PBS ⑤ Positive control group 58.5μL positive control solution (low salt) 1.5μL PBS ⑥ Positive control group 58.5μL positive reaction solution (high salt) 1.5μL PBS

[0103] 6. Data Analysis

[0104] The percentages of the fluorescence values ​​of the experimental group and the negative control group relative to the fluorescence value of the positive control group were calculated.

[0105] 7. Experimental Results

[0106] Within the error range and the allowable instrument fluctuation, the experimental results were plotted by calculating the ratio of the fluorescence value of the experimental group to the fluorescence value of the positive control group, as well as the ratio of the fluorescence value of the negative control group to the fluorescence value of the positive control group (due to the sensitivity of the instrument, the negative control group had fluorescence absorption readings), and the results are shown in Figures 4 and 5.

[0107] From the experimental results in Figures 4 and 5 , it can be seen that the negative control group in each experiment remained unchanged during the measurement process, while the fluorescence value of the experimental group gradually increased with the increase of reaction time, indicating that the BCH666 protein has the activity of unwinding double-stranded DNA, and its unwinding direction is 5'-3'.

[0108] Furthermore, Figure 4 shows the results under low salt conditions, and Figure 5 shows the results under high salt conditions. Comparison of the two shows that the activity of BCH666 protein in unwinding dsDNA increases with increasing salt concentration.

[0109] Example 4 Detection of the melting activity of BCH666 by blocking the restriction sequence

[0110] 1. Preparation of double-stranded DNA containing restriction sequences (ovDNA-3)

[0111] Sequence ID NO: 5 and Sequence ID NO: 8 were annealed to ovDNA-3 (containing a restriction endonuclease sequence) with 20 T residues at the 5' end at a concentration of 10 μM. The annealing process consisted of incubation at 95°C for 5 minutes, followed by a cooling rate of 0.1°C / s to 25°C for 30 minutes. The annealing recipe is shown in Table 5.

[0112] Table 5 ovDNA-3 (containing restriction sequence) annealing formula

[0113] Solution volume 100 μM SEQ ID NO: 5 5 μL 100 μM SEQ ID NO: 8 5 μL TE buffer (pH=8) 40 μL

[0114] 2. Prepare reaction buffer

[0115] The low-salt reaction buffer was 100 mM HEPES (pH = 8.0), 1 mg / mL BSA, 10 mM MgCl2, and 150 mM KCl;

[0116] The high salt reaction buffer consisted of 100 mM HEPES (pH=8.0), 1 mg / mL BSA, 10 mM MgCl2, and 500 mM KCl.

[0117] 3. Prepare the reaction solution

[0118] Experimental reaction solution: Add 3 μL of 10 μM ovDNA-3 (containing a restriction endonuclease sequence), 6 μL of 100 μM SEQ ID NO: 7 (20x competitor DNA), and 6 μL of 100 mM ATP to 582 μL of low-salt reaction buffer. Add 3 μL of 10 μM ovDNA-3, 6 μL of 100 μM SEQ ID NO: 78 (20x competitor DNA), and 6 μL of 100 mM ATP to 585 μL of high-salt reaction buffer.

[0119] Positive control: Add 1 μL of 10 μM SEQ ID NO: 8, 2 μL of 100 μM SEQ ID NO: 7 (20x competitor DNA), and 2 μL of 100 mM ATP to 195 μL of low-salt reaction buffer. Add 1 μL of 10 μM SEQ ID NO: 8, 2 μL of 100 μM SEQ ID NO: 7 (20x competitor DNA), and 2 μL of 100 mM ATP to 195 μL of high-salt reaction buffer.

[0120] 4. Dilute protein

[0121] Dilute BCH666 protein to 4.8 μM with 1× PBS.

[0122] 5. Prepare the Melting Reaction

[0123] Add the corresponding reagents according to Table 6. ①② are experimental groups, ③④ are negative control groups, and ⑤⑥ are positive control groups. Use a microplate reader to detect the kinetic changes of fluorescence intensity within 30 minutes at 30°C. Each group has 3 replicates.

[0124] Table 6 Melting reaction formula

[0125] No. Category Solution 1 Solution 2 ① Experimental group 58.5 μL experimental reaction solution (low salt) 1.5 μL protein ② Experimental group 58.5 μL experimental reaction solution (high salt) 1.5 μL protein ③ Negative control group 58.5 μL experimental reaction solution (low salt) 1.5 μL 1× PBS ④ Negative control group 58.5 μL experimental reaction solution (high salt) 1.5 μL 1× PBS ⑤ Positive control group 58.5 μL positive control solution (low salt) 1.5 μL 1× PBS

[0126] ⑥ Positive control group 58.5 μL positive control solution (high salt) 1.5 μL 1× PBS

[0127] 6. Data Analysis

[0128] The percentages of the fluorescence values ​​of the experimental group and the negative control group relative to the fluorescence value of the positive control group were calculated.

[0129] 7. Experimental Results

[0130] The experimental results were plotted using the same method as in Example 3, as shown in Figures 6 and 7. Under low-salt conditions, the restriction sequence almost completely blocked BCH666 protein from unwinding dsDNA. As shown in Figure 7, under high-salt conditions, the restriction sequence failed to block BCH666 protein from unwinding dsDNA.

[0131] Example 5 Nanopore Sequencing Application of BCH666 Protein

[0132] 1. Two partially complementary DNA strands (upper strand, SEQ ID NO: 9 and lower strand, SEQ ID NO: 10) were annealed to form a linker (as shown in Figure 8), which was then ligated to the double-stranded target fragment using a Rapid T4 DNA Ligase Kit (NEB, E6057AVIAL) and purified to obtain a sequencing library.

[0133] The ligation steps are as follows: Take the rapid T4 DNA ligase out of the -20°C freezer, gently tap the tube wall to mix, centrifuge briefly, and place on ice. Thaw the rapid ligation reaction buffer, pipette to mix, centrifuge briefly, and then place on ice. Prepare the reaction mixture (120μL rapid ligation reaction buffer, 60μL T4 DNA ligase, 30μL 10μM adapter). Then, add 390μL of the purified end-repaired, "A"-added, and purified product of the double-stranded target fragment to the ligation reaction mixture. Use a flared pipette tip to gently pipette and mix 6 times, centrifuge briefly to collect the reaction solution at the bottom of the tube, and then place it in a metal bath preheated at 25°C for the ligation reaction, with a timer for 30 minutes. After the reaction is completed, centrifuge the reaction tube briefly to collect the reaction solution at the bottom of the tube.

[0134] Purification steps are as follows: 30 minutes in advance, remove Ampure XP magnetic beads (Beckman Coulter, A63882) from a 4°C refrigerator, shake thoroughly, and bring to room temperature. Shake thoroughly before use. Pipette 240 μL of magnetic beads into a DNA low-binding tube (Eppendorf, 0030108051) containing the sample ligation product. Mix thoroughly by gently tapping the tube or pipetting gently with a flared pipette tip at least six times, ensuring that all liquid and beads in the pipette tip are transferred to the tube. Incubate on a rotary mixer at room temperature for 5 minutes. Briefly centrifuge the DNA low-binding tube (Eppendorf, 0030108051) and place on a magnetic rack. Let stand for 2–5 minutes until the liquid is clear. Carefully remove the supernatant with a pipette and discard. Place the DNA low-adsorption tube (Eppendorf, 0030108051) on the magnetic stand and add 900 μL of wash buffer [20 mM Tris (pH = 7.5), 2500 mM NaCl]. Remove the DNA low-adsorption tube (Eppendorf, 0030108051) from the magnetic stand and gently tap the tube to mix the magnetic beads. After mixing, return the tube to the magnetic stand and let it stand for 2-5 minutes until all the magnetic beads are attached to the wall. Carefully aspirate and discard the supernatant. Remove the centrifuge tube from the magnetic stand and centrifuge briefly. After separation on the magnetic stand, use a small-range pipette to remove the remaining liquid at the bottom of the tube. Remove the DNA low-adsorption tube (Eppendorf, 0030108051) from the magnetic stand and add 68 μL of elution buffer [20 mM Tris (pH = 7.5), 50 mM NaCl] to elute the DNA. Gently tap the tube to mix. Centrifuge briefly for 3 seconds and collect the liquid at the bottom of the tube. Incubate at room temperature for 10 minutes. After brief centrifugation, place the DNA low-binding tube (Eppendorf, 0030108051) on a magnetic stand and let it stand for 2-5 minutes until the liquid clears. Transfer 66 μL of the supernatant to a new 1.5 mL DNA low-binding tube (Eppendorf, 0030108051). The remaining sample can be used for concentration determination. The Qubit-dsDNA HS Assay Kit (Thermofisher, Q32854) is recommended for concentration determination.

[0135] 2. The BCH666 protein was incubated with the sequencing library at 25°C for 1 h (molar concentration ratio 1:8) to form a sequencing library containing the helicase BCH666.

[0136] 3. The sequencing library containing the helicase BCH666 was incubated with single-stranded DNA (ssDNA-chol, SEQ ID NO: 11) containing cholesterol at its 5' end at room temperature for 10 minutes. The ssDNA-chol sequence is complementary to a portion of the lower strand of the adapter. Cholesterol binding to the phospholipid membrane reduces the amount of sequencing library loaded and improves library capture efficiency (a schematic diagram of the adapter is shown in Figure 9, where the star represents cholesterol and the triangle represents the helicase BCH666).

[0137] 4. Use a patch clamp amplifier or other electrical signal amplifier to collect the current signal (as shown in Figure 10). According to the method disclosed in the literature (Ji Z, Guo P. Channel from bacterial virus T7 DNA packaging motor for the differentiation of peptides composed of a mixture of acidic and basic amino acids. Biomaterials. 2019 May 21; 214: 119-222), a single-channel nanopore detection system based on patch clamp and signal amplifier was constructed. A Teflon membrane with a micrometer-sized pore (50-200 μm in diameter) divides the electrolytic cell into two chambers: the cis chamber and the trans chamber. A pair of Ag / AgCl electrodes is placed in each chamber. A bilayer of phospholipid membranes is formed at the micropores of each chamber, and the nanopore protein CsgG-Eco-(Y51A / F56Q / R97W / R192D-StrepII(C)) is added. Electrical measurements are obtained after a single nanopore protein inserts into the membrane. The sequencing library obtained in step 3 is added, and 180 mV is applied. The library is captured by the nanopore, and the nucleic acids pass through the nanopore under the control of the helicase BCH666. The buffer used in this experiment is: 0.47 M KCl, 25 mM HEPES, 1 mM EDTA, 5 mM ATP, 25 mM MgCl2, pH 7.6, and the sequencing temperature is 28°C.

[0138] 5. The sequencing electrical signal is shown in Figure 11. As the helicase BCH666 guides a single DNA strand into the nanopore, the current is partially blocked, decreasing. Because different nucleotides have different sizes, the amount of current blocked also varies, resulting in the visible fluctuations in the current signal. This example demonstrates the potential of the helicase BCH666 for nanopore sequencing.

[0139] The sequences used in the present invention are as follows:

[0140] Amino acid sequence of BCH666 protein (SEQ ID NO: 1)

[0141]

[0142]

[0143] DNA sequence corresponding to BCH666 protein (SEQ ID NO: 2)

[0144]

[0145] SEQ ID NO: 3:

[0146] 5’-GCGTCGAAAAGCAGTACTTAGGCATT-3’

[0147] SEQ ID NO: 4:

[0148] 5’-TTTTTTTTTTTTTTTTTTTTTAATGCCTAAGTACTGCTTTTCGACGC-3’

[0149] SEQ ID NO: 5:

[0150] 5’-BHQ-1-GCGTCGAAAAGCAGTACTTAGGCATT-3’

[0151] SEQ ID NO: 6:

[0152] 5’-TTTTTTTTTTTTTTTTTTTTTAATGCCTAAGTACTGCTTTTCGACGC-FAM-3’

[0153] SEQ ID NO: 7:

[0154] 5’-AATGCCTAAGTACTGCTTTTCGACGCT-3’

[0155] SEQ ID NO: 8:

[0156] 5’-TTTTTTTTTTTTTTTTTTTTTYYYY-AATGCCTAAGTACTGCTTTTCGACGC-FAM-3’ (Y = iSp18)

[0157] SEQ ID NO: 9:

[0158] 5’-TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTYYYY-GGTTGTTTCTGTTGGTGCTGATATTGCT-3’ (Y = iSp18)

[0159] SEQ ID NO: 10:

[0160] 5'-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3'

[0161] SEQ ID NO: 11:

[0162] 5'-cholesterol-TTGACCGCTCGCCTC-3'

[0163] Wherein, iSp18 is represented by the following formula I:

[0164]

[0165] Although the above describes specific embodiments of the present invention, it should be understood by those skilled in the art that these are merely illustrative and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present invention. Therefore, the scope of protection of the present invention is defined by the appended claims.

Claims

1. A helicase, characterized in that The amino acid sequence of the helicase is as shown in SEQ ID NO:1 or has at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identity with the amino acid sequence shown in SEQ ID NO:

1.

2. The helicase according to claim 1, characterized in that At least one cysteine ​​on the three-dimensional structure surface of the helicase is mutated, and the mutation is that the cysteine ​​is replaced by alanine, glutamine, glycine, histidine, isoleucine, leucine, valine, serine, threonine or methionine; Preferably, the mutated sites include at least one of C21, C50, C56, C91, C156, C279, C367 and C379.

3. The helicase according to claim 1, characterized in that The helicase has at least one amino acid mutation in the pin domain and / or the tower domain, wherein the amino acid mutation is a substitution of the original amino acid with cysteine ​​or a non-natural amino acid; Preferably, the mutation site of the pin domain is at least one of 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102 and 103; and / or the mutation site of the tower domain is at least one of 337, 338, 339, 340, 341, 342, 343, 344, 345, 346, 347, 348, 347, 348, 349, 350, 351, 352, 353, 354, 355, 356, 357, 358, 359, 360, 361, 362, 363, 364, 365, 366, 367 and 368; More preferably, the mutation site of the pin domain is at least one of G85, I86, S87, P88, T89, V90, D91, K92, K93, E94, L95, E96, F97, E98, H99, V97, N98, I99, P100, S101, L102 and W103; and / or the mutation site of the tower domain is at least one of L337, Y338, E339, V340, A34 1. At least one of N342, Y343, Y344, D345, Y346, Q347, Q348, I347, A348, D349, Y350, Y351, E352, H353, I354, A355, W356, N357, M358, K359, T360, P361, Q362, A363, K364, A365, K366, A367 and Y368.

4. The helicase according to claim 3, characterized in that The unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (PAZF), 4-azido-L-phenylalanine (PAZF-Hcl), 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselenoyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4-(dihydroxyboryl)-L-phenylalanine, 4-[(ethylsulfanyl)carbonyl]-L-phenylalanine, (2S)-2-amino-3-{4-[ [(propan-2-ylsulfanyl)carbonyl]phenyl}propionic acid, (2S)-2-amino-3-{4-[(2-amino-3-sulfanylpropionyl)amino]phenyl}propionic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino-L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthylamino)propionic acid, 6-(methylsulfanyl)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propionic acid, (2R)-2-aminooctanoate 3-(2,2′-bipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxy-3-quinolyl)propionic acid, 4 -benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)sulfanyl]propanoic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propanoic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)oxy]carbonyl}amino)hexanoic acid, O-(2-nitrobenzyl)-L-tyrosine and 2-nitrophenylalanine.

5. The helicase according to claim 1, characterized in that The helicase has at least one mutation in the DNA binding region and / or near the ATP catalytic active center, wherein the mutation includes replacement of the original amino acid with an amino acid with a larger side chain; Preferably, the mutation site of the DNA binding region is at least one of 63, 73, 79, 80, 81, 82, 83, 84, 85, 86, 87, 95 and 96; and / or the mutation site near the ATP catalytic active center is at least one of 154, 155, 156, 158, 159, 160, 161, 174, 175, 177, 178, 179 and 181; More preferably, the mutation site in the DNA binding region is at least one of H63, E73, T79, V80, H81, S82, A83, L84, G85, I86, S87, L95 and E96; and / or the mutation site near the ATP catalytic active center is at least one of D154, D155, P156, I158, S159, P160, V161, P174, M175, N177, T178, G179 and L181.

6. The helicase according to claim 1, characterized in that The helicase has at least one mutation in the DNA binding region and the nanopore binding region, wherein the mutation includes substitution of the original amino acid with an amino acid having a non-positive surface charge and / or an amino acid having a side chain shorter than the original amino acid; Preferably, the mutation occurs in at least one of positions 1, 2, 3, 5, 7, 8, 9, 10, 36, 42, 45, 67, 74, 97, 207, 208, 209, 220, 221, 222, 374, 408 and 411 of the DNA binding region and the nanopore binding region; More preferably, the mutation occurs in at least one of M1, N2, S3, N5, D7, Q8, Q9, K10, K36, K42, K45, D67, K74, F97, R207, K208, D209, K220, K221, D222, H374, D408 and K411 in the DNA binding region and the nanopore binding region.

7. An isolated nucleic acid, characterized in that The isolated nucleic acid encodes the helicase according to any one of claims 1 to 6; Preferably, the nucleotide sequence of the helicase is as shown in SEQ ID NO:2 or has at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identity with the nucleotide sequence shown in SEQ ID NO:

2.

8. A recombinant expression vector, characterized in that: The recombinant expression vector comprises a promoter and the nucleic acid according to claim 7; Preferably, the promoter is T7; and / or the backbone plasmid of the recombinant expression vector is PET.28a(+), PET.21a(+), or PET.32a(+).

9. A transformant, characterized in that: The transformant comprises a host cell and the nucleic acid according to claim 7 or the recombinant expression vector according to claim 8; Preferably, the host cell is Escherichia coli, more preferably BL21(DE3), BL21 Star(DE3)pLyss, Rossata(DE3) or Lemo21(DE3).

10. A method for preparing the helicase according to any one of claims 1 to 6, characterized in that: The transformant according to claim 9 is cultured in a culture medium to ferment and produce the helicase.

11. A helicase-sequencing adapter complex, characterized in that The method comprises the helicase according to any one of claims 1 to 6, and a sequencing adapter.

12. A kit, characterized in that: The kit comprises the helicase according to any one of claims 1 to 6 and / or the helicase-sequencing adapter complex according to claim 11; preferably also comprises a single-stranded DNA containing an anchor molecule at the 5' end, a nanopore, a nanopore protein, an electrical signal detector, a membrane and / or a buffer; Preferably, the anchoring molecule is a hydrophobic molecule, preferably selected from any one or more of the following: lipids, fatty acids, sterols, carbon nanotubes, polypeptides, proteins and / or amino acids, such as cholesterol, palmitate or tocopherol; The nanopore is a transmembrane protein pore or a solid-state pore; preferably, the transmembrane protein pore is selected from hemolysin, MspA, MspB, MspC, MspD, FraC, ClyA, PA63, CsgG, CsgD, XcpQ, SP1, phi29 connector protein, InvG, GspD or any combination thereof; The membrane is an amphiphilic membrane (such as a phospholipid bilayer), a high molecular polymer membrane (such as a di-block copolymer, a tri-block copolymer) or any combination thereof; The buffer is a dihydrogen phosphate-hydrogen phosphate buffer system, a carbonic acid-sodium bicarbonate buffer system, a Tris-HCl buffer system, a HEPES buffer system, a MOPS buffer system or any combination thereof.

13. Use of the helicase according to any one of claims 1 to 6, the helicase-sequencing adapter complex according to claim 11, or the kit according to claim 12 in high-throughput sequencing; Preferably, the high-throughput sequencing is nanopore sequencing.

14. A DNA denaturing method, characterized in that: The method comprises using the helicase according to any one of claims 1 to 6, the helicase-sequencing adapter complex according to claim 11 or the kit according to claim 12 to unwind the double-stranded DNA.

15. A sequencing method, characterized in that: It includes the following steps: The DNA is sequenced while being unwound using the DNA unwinding method as described in claim 14.