Helicase and application thereof

CN120303401APending Publication Date: 2025-07-11BGI HANGZHOU CYCLONESEQ TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102038.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The helicase salt tolerance of existing nanopore sequencers is poor, resulting in a decrease in unwinding speed in high-salt environments, which impairs sequencing speed and efficiency.

Method used

A helicase with two tower domains and one pin domain was designed. Its stability and unwinding activity in high-salt environment were improved through amino acid mutation and introduction of unnatural amino acids, and the towers were connected through chemical cross-linking. domain and pin domain to enhance DNA binding ability.

Benefits of technology

It achieves the maintenance of high unwinding activity in high-salt environments, improves the stability and persistence of nucleic acid perforation movement, and improves the efficiency and signal-to-noise ratio of nanopore sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303401A_ABST
    Figure CN120303401A_ABST
Patent Text Reader

Abstract

The invention discloses helicase and application thereof. Wherein the helicase is provided with two tower structural domains and a pin structural domain, and the two tower structural domains are positioned on the same side of a three-dimensional structure of the helicase. According to the technical scheme, the brand-new helicase BCH3X with the special spiral characteristic structural domain is provided, and the helicase BCH3X has good salt tolerance and stability, can have high unwinding activity under high salt, can be used for control and characterization of nucleic acid and is applied to nanopore sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

Helicase and its applications Technical Field

[0001] The present invention relates to the field of biotechnology, and in particular to a helicase and applications thereof. Background Art

[0002] Nanopore sequencing, an emerging single-molecule sequencing technology, has revolutionized the genome sequencing industry with its unique advantages, including high throughput, long read lengths, rapid speed, in situ detection, and label-free operation. This technology eliminates the need for imaging equipment, enabling portable systems to meet diverse sequencing scenarios. Furthermore, due to its non-amplified direct sequencing nature, there is no length limit for sequenceable DNA, enabling real-time base calling and direct sequencing of RNA, methylated and other modified molecules, and other single molecules. Nanopore sequencing has broad applications in molecular biology, medicine, epidemiology, and ecology, including genome mapping, epidemic monitoring, rare species detection, identification of hidden intermediates, monitoring the dynamics of non-covalent interactions, characterizing epigenetic and post-translational modifications, and rapid and cost-effective protein sequencing.

[0003] Nanopore sequencing technology is based on electrical signals. A nanopore (protein or solid) inserted into a membrane as a signal sensor separates two electrolyte chambers. When voltage is applied between the two electrolyte chambers, a stable perforation current is generated. When the molecule to be tested enters the nanopore, the flow of ions is hindered, resulting in current signal fluctuations. Different bases have different effects on the current. By detecting the current fluctuation signal of the nanopore in real time and using machine learning to analyze and decode the current signal, real-time sequencing of the molecule to be tested can be achieved.

[0004] During this sequencing process, due to the extremely high speed at which nucleic acid molecules pass through the nanopore channel, it is impossible to accurately obtain polynucleotide sequence information. Therefore, effectively reducing and controlling the perforation motion of nucleic acid molecules is a key technical issue in achieving nanopore sequencing. Currently, the most common and effective method is to use the idea of ​​helicase unwinding to control the perforation motion of nucleic acid molecules and improve detection accuracy. In addition, in order to better maintain sequencing speed and sequencing uniformity, the helicase needs to have good salt tolerance and thermal stability in high-salt electrolyte solutions.

[0005] The helicase used in current commercial nanopore sequencers is the Dda helicase derived from the bacterial phage T4. Its production, stability, and salt tolerance are all poor. In particular, high salt will inhibit the unwinding activity of the Dda helicase, causing its unwinding speed to decrease and unable to fully exert its unwinding ability, thereby weakening its sequencing speed in nanopore sequencing applications and reducing sequencing efficiency.

[0006] Summary of the Invention

[0007] The present invention aims to provide a helicase and an application thereof, so as to solve the technical problem of poor salt tolerance of helicase in the prior art.

[0008] To achieve the above object, according to one aspect of the present invention, a helicase is provided. The helicase has two tower domains and one pin domain, wherein the two tower domains are located on the same side of the three-dimensional structure of the helicase.

[0009] Furthermore, the helicase includes at least one of the following: A) BCH326, BCH326 is a protein having the amino acid sequence shown in SEQ ID NO: 1; B) BCH338, BCH338 is a protein having the amino acid sequence shown in SEQ ID NO: 3; C) a protein in which at least one cysteine ​​on the surface of the protein defined in A) or B) is mutated to alanine, glutamine, glycine, histidine, isoleucine, leucine, valine, serine, threonine or methionine; D) a protein in which the amino acid at at least one site on the tower domain and / or pin domain of the amino acid sequence of any of the proteins defined in A), B) and C) is mutated to cysteine ​​or at least one non-natural amino acid is introduced, and the protein has the ability to unwind DNA; and E) a protein having more than 70% homology to the amino acid sequence of any of the proteins defined in A), B), C) and D) and having the same function.

[0010] Furthermore, C) includes: proteins in which the C at position 319 of BCH326 is substituted with A, S, T, V, I, L or G; and proteins in which the C at position 326 or 459 of BCH338 is substituted with A, S, T, V, I, L or G.

[0011] Further, in D), the amino acid mutation to cysteine ​​or the introduction of at least one unnatural amino acid at at least one site on the tower domain and / or the pin domain includes at least one of the following: S389, R340, K341, S342, N343, K343, S344, I345, V346, I347, D348, K349, D350, G351, K352, A353, K354, E355, F356, L357, R358, K359, F360, L361, N362, F363, A364, K365, I366, Y367, N368, F369, T370, The amino acid at least one of the following positions is mutated into cysteine ​​or at least one unnatural amino acid is introduced into the amino acid residues of at least one of the following positions: 70, N371, K372, G373, G374, H378, G379, R380, R381, I382, T383, K384, K385, S386, K387, K388, E389, L390 and W391; D87, I88, G89, T90, I91, H92, S93, Y94, F95, D96, I97, K98, P99, D100, I101, D102, D103, N104, G105, N106, R107, V108, The amino acid at least one of F109, K110, P111 or S112 is mutated to cysteine ​​or at least one unnatural amino acid is introduced; S405, K406, F407, L408, V409, P410, L411, G412, D413, G414, S415, K416, E417, D418, L419, F420, P421, L422, Y423, K424, E425, A426, V427, F428, D429, I430, A431, K432, T433, M434, N435, N436, Q437, R438, The amino acid at least one of K439, I440, S441, K442, N443, S444, K445, K446, N447, F448, or W449 is mutated to cysteine, or at least one unnatural amino acid is introduced; the amino acid at least one of E93, I94, R95, P96, D97, I98, N99, E100, F101, G102, E103, R104, I105, F106, V107, P108, K109, L110, R111, D112, M113, or M114 on the pin domain of BCH338 is mutated to cysteine, or at least one unnatural amino acid is introduced;Preferably, in E), the protein has an amino acid sequence homology of 70%, 80%, 90%, 95% or 99% or more and the same function as the protein defined in any one of A), B), C) and D).

[0012] Further, the unnatural amino acid is selected from 4-azido-L-phenylalanine, 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselenoyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4(dihydroxyboryl)-L-phenylalanine, 4-[(ethylsulfanyl)carbonyl]-L-phenylalanine, (2S)-2-amino-3-{4-[(propan-2-ylsulfanyl)carbonyl]phenyl}propanoic acid, (2S)-2-amino-3-{4-[(2-amino-3-sulfanylpropionyl)amino]phenyl}propanoic acid, O-methyl-L-tyrosine, 4- amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino-L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthyl-2-ylamino)propionic acid, 6-(methylsulfanyl)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propionic acid at least one of (2R)-2-aminooctanoate 3-(2,2′-bipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxy-3-quinolinyl)propionic acid, 4-benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)sulfanyl]propionic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propionic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)oxy]carbonyl}amino)hexanoic acid, and O-(2-nitrobenzyl)-L-tyrosine or 2-nitrophenylalanine; preferably , BCH326 introduces at least one unnatural amino acid including at least one of the following: D100 introduces 4-azido-L-phenylalanine, I101 introduces 4-azido-L-phenylalanine, D102 introduces 4-azido-L-phenylalanine, D103 introduces 4-azido-L-phenylalanine, N104 introduces 4-azido-L-phenylalanine, G105 introduces 4-azido-L-phenylalanine, N106 introduces 4-azido-L-phenylalanine, R107 introduces 4-azido-L-phenylalanine, D103 introduces 4-acetyl-L-phenylalanine, G105 introduces 4-acetyl-L-phenylalanine and N106 introduces 4-acetyl-L-phenylalanine;Preferably, the at least one unnatural amino acid introduced into BCH338 includes at least one of the following: A431 introduces 4-azido-L-phenylalanine, K432 introduces 4-azido-L-phenylalanine, T433 introduces 4-azido-L-phenylalanine, M434 introduces 4-azido-L-phenylalanine, N435 introduces 4-azido-L-phenylalanine, S441 introduces 4-azido-L-phenylalanine, K442 introduces 4-azido-L-phenylalanine, N443 introduces 4-azido-L-phenylalanine, and S444 introduces 4-azido-L-phenylalanine.

[0013] Furthermore, there is at least one amino acid mutation at the amino acid site of the DNA binding region of the helicase and / or the amino acid site near the ATP catalytic active center, and the mutation includes mutating the original amino acid to an amino acid with a larger side chain; preferably, mutating the original amino acid to an amino acid with a larger side chain includes at least one of the following: asparagine is replaced by glutamine, histidine, arginine or lysine; proline is replaced by arginine, lysine, phenylalanine or leucine; histidine is replaced by arginine, lysine, glutamine, asparagine phenylalanine, tyrosine or tryptophan; proline is replaced by arginine, lysine, glutamine, asparagine or histidine; phenylalanine is replaced by arginine, lysine, histidine, tyrosine or tryptophan; isoleucine is replaced by phenylalanine, tryptophan, histidine, lysine or arginine; tyrosine is replaced by arginine, lysine, or tryptophan; the amino acid site of the DNA binding region of BCH326 includes The amino acid sites near the ATP catalytic active center include: L157, V160, L294, G296, N299, L303, A304, I328, F329, T330, N331, G332, G333 and E334, and the amino acid sites near the ATP catalytic active center include: K211, E212, E213, N214, Y215, K216, A217, P218, L219, K220, D221, I222, N22 3 and N224; the amino acid sites in the DNA binding region of BCH338 include: H89, S90, Y91, F92, E93, I94, R95 and P96; the amino acid sites near the ATP catalytic active center include: Y152, Q153, L154, P155, P156, V157, F193, L194, I195, K196, E197, Y198, E199, E200 and N201.

[0014] Furthermore, the amino acid on the surface of the helicase that interacts with the nanopore binding region has a mutation at at least one site, and the mutation includes mutating the original amino acid to an amino acid with a shorter side chain; preferably, the mutation of the original amino acid to an amino acid with a shorter side chain includes: asparagine is replaced by isoleucine, valine, isoleucine, alanine, serine or glycine; lysine is replaced by isoleucine, valine, isoleucine, alanine, serine or glycine; lysine is replaced by isoleucine, valine, isoleucine, alanine, serine or glycine; arginine is replaced by isoleucine, valine, isoleucine, alanine, serine or glycine; preferably, the amino acids on the surface of BCH326 that interact with the nanopore binding region include : M1, E2, S3, K4, I5, N6, L7, T8, E9, D10, Q11, L12, K13, I14, I15, K16, I189, I190, R191, T192, Q193, N194, K195, N196, and S197; amino acids on the BCH338 surface that interact with the nanopore binding region include: M1, G2, E3, I4, K5, L6, N7, E8, E9, Q10, Q11, K12, K177, I177, L178, R179, T180, K181, N182, L213, I214, D215, H216, F217, H218, V219, Y220, G221, D248, L249, T250, D251, S252, T253, E254 and S255.

[0015] According to another aspect of the present invention, an isolated DNA molecule is provided. The DNA molecule comprises (a) a nucleotide sequence encoding the helicase of any one of claims 2 to 3; or (b) a nucleotide sequence that hybridizes under stringent conditions with the DNA molecule defined in (a); or (c) a nucleotide sequence as set forth in SEQ ID NO: 2 or SEQ ID NO: 4; or (d) a nucleotide sequence that has 70% or greater homology with the nucleotide sequence defined in any one of (a) to (c) and encodes a protein having the same function as the helicase.

[0016] Furthermore, the DNA molecule has a nucleotide sequence that has 75% or more, preferably 85% or more, more preferably 95% or more, and even more preferably 99% or more homology with any one of the nucleotide sequences defined in (a) to (c) and encodes a protein having the same function.

[0017] According to another aspect of the present invention, a recombinant vector is provided, which comprises any one of the above-mentioned DNA molecules.

[0018] Furthermore, the recombinant vector is selected from a plasmid, a virus or a carrier expression vector; further, the recombinant vector includes a regulatory element for controlling the expression of the DNA molecule; further, the regulatory element includes a promoter operably linked to the DNA molecule; preferably, the promoter includes T7, trc, lac, ara or λL; more preferably, the recombinant vector is selected from plasmid PET.28a(+), PET.21a(+) or PET.32a(+).

[0019] According to another aspect of the present invention, a host cell is provided, which comprises any one of the aforementioned DNA molecules of the present invention, or any one of the aforementioned recombinant vectors of the present invention.

[0020] Further, the host cell includes Escherichia coli; preferably, the host cell includes BL21(DE3), BL21Star(DE3)pLysS, Rossata(DE3) or Lemo21(DE3).

[0021] According to another aspect of the present invention, the application of the above-mentioned helicase in nucleic acid control or characterization is provided; further, nucleic acid control includes controlling the speed of nucleic acid passing through the nanopore, controlling the stability of nucleic acid perforation, or controlling the continuity of nucleic acid perforation; further, the application includes the application of nanosensors and single-molecule nanopore sequencing applications.

[0022] According to another aspect of the present invention, a nanopore sequencing kit is provided, wherein the kit comprises a helicase, which is any of the above-mentioned helicases of the present invention.

[0023] According to another aspect of the present invention, a nanopore sequencing method is provided, comprising sequencing a nucleic acid molecule to be sequenced under the control of a helicase, wherein the helicase is any of the above-mentioned helicases of the present invention.

[0024] By applying the technical solution of the present invention, a new type of helicase BCH3X with a special helical characteristic domain is provided. It has good salt tolerance and stability, can have high unwinding activity under high salt conditions, can be used for the control and characterization of nucleic acids, and can be applied to nanopore sequencing. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0026] FIG1 shows the results of the molecular sieve Superdex 200 purification of BCH326, wherein (A) the molecular sieve Superdex 200 elution profile of BCH326; (B) the molecular sieve gel elution profile of BCH326.

[0027] FIG2 shows the results of the molecular sieve Superdex 200 purification of BCH338, wherein (A) the molecular sieve Superdex 200 elution profile of BCH338; (B) the molecular sieve gel elution profile of BCH338.

[0028] FIG3 shows the Alphafold2 predicted structure of BCH326.

[0029] FIG4 shows the Alphafold2 predicted structure of BCH338.

[0030] FIG5 shows the ATPase activity detection of BCH326 protein.

[0031] FIG6 shows the ATPase activity detection of BCH338 protein.

[0032] FIG7 shows the dsDNA melting activity detection of BCH326 protein (low salt reaction buffer).

[0033] FIG8 shows the dsDNA melting activity detection of BCH326 protein (high salt reaction buffer).

[0034] FIG9 shows the dsDNA melting activity detection of BCH338 protein (low salt reaction buffer).

[0035] FIG10 shows the dsDNA melting activity detection of BCH338 protein (high salt reaction buffer).

[0036] FIG11 shows the detection of the restriction sequence blocking the melting activity of BCH326 protein (low salt reaction buffer).

[0037] FIG12 shows the detection of the restriction sequence blocking the melting activity of BCH326 protein (high salt reaction buffer).

[0038] FIG13 shows the detection of the restriction sequence blocking the melting activity of BCH338 protein (low salt reaction buffer).

[0039] FIG14 shows the detection of the restriction sequence blocking the melting activity of BCH338 protein (high salt reaction buffer).

[0040] FIG15 shows a schematic diagram of the linker (a: upper chain; b: lower chain).

[0041] Figure 16 shows a schematic diagram of a sequencing library containing helicase (a: upper strand; b: lower strand; c: double-stranded target fragment; d: helicase; e: cholesterol-labeled double-stranded DNA).

[0042] FIG17 shows a schematic diagram of a patch clamp amplifier.

[0043] FIG18 shows a BCH326 sequencing current signal diagram.

[0044] FIG19 shows a BCH338 sequencing current signal diagram.

[0045] FIG20 shows the crystal structure of Dda. DETAILED DESCRIPTION

[0046] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0047] According to a typical embodiment of the present application, a helicase is provided that has two tower domains and a pin domain, with the two tower domains located on the same side of the helicase's three-dimensional structure. During nanopore sequencing, the tower and pin domains of a helicase with this structure can be cross-linked to form a DNA binding region, facilitating rate-controlled sequencing, increasing sequencing continuity and stability, and preventing fluctuations in sequencing signals caused by DNA slippage or fluctuations during the sequencing process. Furthermore, the helicase has high salt tolerance, enabling superior unwinding capabilities at high salt concentrations, thereby improving sequencing efficiency.

[0048] According to a typical embodiment of the present application, a helicase is provided, the gene of which is derived from a deep-sea metagenome and has high salt tolerance. The helicase comprises at least one of the following: A) BCH326, BCH326 is a protein having the amino acid sequence shown in SEQ ID NO: 1; B) BCH338, BCH338 is a protein having the amino acid sequence shown in SEQ ID NO: 3; C) a protein in which at least one cysteine ​​on the surface of the protein defined in A) or B) is mutated to alanine, glutamine, glycine, histidine, isoleucine, leucine, valine, serine, threonine, or methionine; D) a protein in which at least one amino acid in the tower domain and / or pin domain of the amino acid sequence of any of the proteins defined in A), B), and C) is mutated to cysteine ​​or at least one unnatural amino acid is introduced, and the protein has DNA unwinding ability; and E) a protein having an amino acid sequence with greater than 70% homology to the protein defined in A), B), C), and D) and having the same function.

[0049] The helicase defined in A) or B) above is derived from the deep-sea metagenome and has high salt tolerance.

[0050] The helicase defined in C) above can improve protein uniformity by mutating at least one cysteine ​​on the surface of the helicase to alanine, glutamine, glycine, histidine, isoleucine, leucine, valine, serine, threonine or methionine, thereby improving indicators such as sequencing uniformity.

[0051] According to a typical embodiment of the present invention, the helicase defined in C) above includes: a protein in which the C at position 319 of the BCH326 is replaced by A, S, T, V, I, L or G; and a protein in which the C at position 326 or 459 of the BCH338 is replaced by A, S, T, V, I, L or G.

[0052] According to a typical embodiment of the present invention, based on the sequence defined in A) or B), the protein is mutated to stably link the tower domain and the pin domain, thereby immobilizing the DNA in the region formed by the tower domain and the pin domain during sequencing, thereby improving the stability and continuity of sequencing. For example, the amino acid mutation to cysteine ​​or the introduction of at least one unnatural amino acid at at least one site in the tower domain and / or the pin domain of any of the amino acid sequences in A), B), and C) includes at least one of the following:

[0053] On the tower domain of BCH326, S389, R340, K341, S342, N343, K343, S344, I345, V346, I347, D348, K349, D350, G351, K352, A353, K354, E355, F356, L357, R358, K359, F360, L361, N362, F363, A364, K3 at least one of the amino acid residues in the residues 65, I366, Y367, N368, F369, T370, N371, K372, G373, G374, H378, G379, R380, R381, I382, T383, K384, K385, S386, K387, K388, E389, L390, and W391 is mutated to cysteine ​​or at least one unnatural amino acid is introduced;

[0054] The amino acid of at least one of D87, I88, G89, T90, I91, H92, S93, Y94, F95, D96, I97, K98, P99, D100, I101, D102, D103, N104, G105, N106, R107, V108, F109, K110, P111, or S11 on the pin domain of BCH326 is mutated to cysteine ​​or at least one unnatural amino acid is introduced;

[0055] On the tower domain of BCH338, S405, K406, F407, L408, V409, P410, L411, G412, D413, G414, S415, K416, E417, D418, L419, F420, P421, L422, Y423, K424, E425, A426, at least one of V427, F428, D429, I430, A431, K432, T433, M434, N435, N436, Q437, R438, K439, I440, S441, K442, N443, S444, K445, K446, N447, F448, or W449 is mutated to cysteine ​​or at least one unnatural amino acid is introduced;

[0056] The amino acid of at least one of E93, I94, R95, P96, D97, I98, N99, E100, F101, G102, E103, R104, I105, F106, V107, P108, K109, L110, R111, D112, M113, and M114 on the pin domain of BCH338 is mutated to cysteine ​​or at least one unnatural amino acid is introduced;

[0057] The above mutations enable chemical linkage between the tower domain and the pin domain, including covalent or non-covalent linkage.

[0058] In addition, in F), the protein preferably has 70%, 80%, 90%, 95%, or 99% or greater homology and the same function as the amino acid sequence of any of the proteins defined in A), B), and C). The term "homology" as used herein has a meaning generally known in the art, and those skilled in the art are familiar with the rules and standards for determining homology between different sequences. Sequences defined by varying degrees of homology in the present invention must also possess helicase function. Those skilled in the art can obtain such variant sequences based on the teachings of this disclosure.

[0059] According to a typical embodiment of the present invention, the non-natural amino acids mentioned are not limited to 4-azido-L-phenylalanine, 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselenoyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4-(dihydroxyboryl)-L-phenylalanine, 4-[(ethylsulfanyl)carbonyl]-L-phenylalanine, (2S)-2-amino-3-{4- [(Propan-2-ylsulfanyl)carbonyl]phenyl}propionic acid, (2S)-2-amino-3-{4-[(2-amino-3-sulfanylpropionyl)amino]phenyl}propionic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino- L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthylamino)propionic acid, 6-(methylsulfanyl)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propionic acid, (2R)-2-aminooctanoate 3-(2,2′-bipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxy- 3-quinolinyl)propionic acid, 4-benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)sulfanyl]propionic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propionic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)oxy]carbonyl}amino)hexanoic acid, O-(2-nitrobenzyl)-L-tyrosine or 2-nitrophenylalanine, etc.

[0060] In one embodiment of the present invention, cysteine ​​is mutated at the above-mentioned position or a non-natural amino acid is introduced, and one or more ends of one or more connectors are preferably used to covalently connect the tower domain and the pin domain of the helicase. If one end is covalently connected, the one or more connectors can instantaneously connect the above-mentioned two or more cysteines and / or non-natural amino acids. If two or all ends are covalently connected, the one or more connectors permanently connect the two or more cysteines and / or non-natural amino acids. Among them, the connector is capable of producing a medium containing covalent or non-covalent effects. This can be a commercial cross-linking agent, or it can be a small protein, a polypeptide, a synthetic small molecule, etc. It is used to connect the tower domain and the pin domain, so that after the DNA binds to the helicase, the two domains are connected, and the DNA can not leave the area during the perforation movement, thereby improving the stability and sustainability of nanopore sequencing.

[0061] One of the covalent attachment methods is to cross-link the tower domain or the pin domain of the helicase during sequencing using a cross-linking agent to improve the continuity and stability of sequencing. Suitable chemical cross-linking agents are well known in the art. Suitable chemical cross-linking agents include, but are not limited to, those comprising the following functional groups: maleimide, active ester, succinimide, azide, alkyne (e.g., dibenzocyclooctyne (DIBO or DBCO), difluorocycloalkyne and linear alkyne), phosphine (e.g., those used in traceless and non-traceless Staudinger binding), haloacetyl (e.g., iodoacetamide), phosgene-type reagents, sulfonyl chloride reagents, isothiocyanates, acyl halides, hydrazine, disulfide, vinyl sulfone, aziridine and photosensitizer (e.g., aryl azide, diaziridine).

[0062] In addition, the present invention can also mutate the amino acid sites in the DNA binding region of this type of helicase and the amino acid sites near the ATP catalytic active center. The mutation direction includes but is not limited to mutation to larger side chain amino acids, thereby increasing the (i) electrostatic interaction between at least one amino acid and one or more nucleotides in ssDNA; (ii) hydrogen bond and / or (iii) cation-pi (cation-π) interaction; substitution to increase positively charged amino acids to reduce the repulsion between the motor protein and the pore, etc.

[0063] Among them, the mutation of the original amino acid to an amino acid with a larger side chain includes at least one of the following: asparagine (N) is replaced by glutamine (Q), histidine (H), arginine (R) or lysine (K); proline (P) is replaced by arginine (R), lysine (K), phenylalanine (F) or leucine (I); histidine (H) is replaced by arginine (R), lysine (K), glutamine (Q), asparagine (N) phenylalanine (F), tyrosine (Y) or tryptophan (W); proline (P) is replaced by arginine (R), lysine (K), glutamine (Q), asparagine (N) phenylalanine (F), tyrosine (Y) or tryptophan (W); Acid (P) is replaced by (i) arginine (R), lysine (K), glutamine (Q), asparagine (N) or histidine (H); phenylalanine (F) is replaced by arginine (R), lysine (K), histidine (H), tyrosine (Y) or tryptophan (W); isoleucine (I) is replaced by phenylalanine (F), tryptophan (W), histidine (H), lysine (K) or arginine (R); tyrosine (Y) is replaced by arginine (R), lysine (K), or tryptophan (W), etc.

[0064] These sites include but are not limited to the DNA binding region: BCH326: L157, V160, L294, G296, N299, L303, A304, I328, F329, T330, N331, G332, G333, E334; BCH338: H89, S90, Y91, F92, E93, I94, R95, P96; amino acid sites near the ATP catalytic active center: BCH326: K 198, E199, E200, N201.

[0065] According to one embodiment of the present invention, the long side chain amino acids at the binding amino acid sites of the surface and nanopore binding region of this type of helicase can be mutated into short side chain amino acids to reduce the repulsion between the helicase and the nanopore during sequencing.

[0066] Common mutation directions include asparagine (N) replaced by isoleucine (I), valine (V), isoleucine (L), alanine (A), serine (S) or glycine (G); lysine (K) replaced by isoleucine (I), valine (V), isoleucine (L), alanine (A), serine (S) or glycine (G); lysine (K) replaced by isoleucine (I), valine (V), isoleucine (L), alanine (A), serine (S) or glycine (G); arginine (R) replaced by isoleucine (I), valine (V), isoleucine (L), alanine (A), serine (S) or glycine (G), etc. Preferably, these amino acid positions include but are not limited to BCH326: M1, E2, S3, K4, I5, N6, L7, T8, E9, D10, Q11, L12, K13, I14, I15, K16, I189, I190, R191, T192, Q193, N194, K195, N196, S197; BCH338: M1, G2, E3, I4, K5, L6, N7, E8, E9, Q10, Q11, K12, K177, One or more of I177, L178, R179, T180, K181, N182, L213, I214, D215, H216, F217, H218, V219, Y220, G221, D248, L249, T250, D251, S252, T253, E254, S255.

[0067] According to a typical embodiment of the present application, an isolated DNA molecule is provided, which has (a) a nucleotide sequence encoding any of the above-mentioned helicases; or (b) a nucleotide sequence that hybridizes with the DNA molecule defined in (a) under stringent conditions; or (c) a nucleotide sequence shown in SEQ ID NO: 2 or SEQ ID NO: 4; or (d) a nucleotide sequence having more than 70% (preferably more than 80%, more preferably more than 85%, further preferably more than 90%, most preferably more than 95%, for example, it can be 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 98.5%, 99%, 99.5%, 99.6%, 99.7%, 99.8% or more, or even more than 99.9%) homology with any of the nucleotide sequences defined in (a) to (c) and encoding a nucleotide sequence having the same function as the above-mentioned helicase.

[0068] It should be noted that the “homology” in this application refers to the identity between any two nucleotide sequences or amino acid sequences from the first amino acid to the last amino acid encoded by the corresponding genes when compared.

[0069] As used herein, "isolated" means altered "by the hand of man" from its natural state, i.e., if it occurs in nature, it is altered and / or separated from its original environment. For example, a polynucleotide or polypeptide naturally present in a living organism is not "isolated," whereas the same polynucleotide or polypeptide separated from its coexisting components in its natural state is "isolated" (as the term is used herein).

[0070] The DNA molecules of the present invention hybridize with the helicase encoding gene of the present invention under "stringent conditions", which refers to the conditions under which the presence of the helicase encoding gene of the present invention can be identified by nucleic acid hybridization. In the present invention, if two DNA molecules can form an antiparallel double-stranded nucleic acid structure, it can be said that the two DNA molecules can specifically hybridize with each other. If two DNA molecules show complete complementarity, one of the DNA molecules is said to be the "complement" of the other DNA molecule. In the present invention, if two DNA molecules can hybridize with each other with sufficient stability so that they anneal and bind to each other under conventional "high stringency" conditions, the two DNA molecules are said to have "complementarity". Deviation from complete complementarity is permissible as long as such deviation does not completely prevent the two molecules from forming a double-stranded structure. In order for a DNA molecule to be able to serve as a primer or probe, it is only necessary to ensure that it has sufficient complementarity in sequence so that a stable double-stranded structure can be formed under the specific solvent and salt concentration used.

[0071] In the present invention, a substantially homologous sequence is a DNA molecule that can specifically hybridize with the complementary strand of another matching DNA molecule under highly stringent conditions. Suitable stringent conditions that promote DNA hybridization, for example, treatment with 6.0× sodium chloride / sodium citrate (SSC) at approximately 45°C, followed by washing with 2.0×SSC at 50°C, are well known to those skilled in the art. For example, the salt concentration in the washing step can be selected from about 2.0×SSC at 50°C for low stringency conditions to about 0.2×SSC at 50°C for high stringency conditions. In addition, the temperature in the washing step can be increased from room temperature of about 22°C for low stringency conditions to about 65°C for high stringency conditions. Both the temperature and the salt concentration can be varied, or one of them can be kept constant while the other is varied. Preferably, the stringent conditions in the present invention may be specific hybridization with the nucleotide sequence encoding the helicase of the present application in a 6×SSC, 0.5% SDS solution at 65°C, followed by washing the membrane once with 2×SSC, 0.1% SDS and once with 1×SSC, 0.1% SDS.

[0072] According to a typical embodiment of the present application, a recombinant vector is provided, comprising the aforementioned DNA molecule, i.e., the helicase expression gene. The recombinant vector is selected from a plasmid, a virus, or a carrier expression vector; the recombinant vector includes a regulatory element for controlling the expression of the aforementioned DNA molecule; the regulatory element includes a promoter operably linked to the DNA molecule; preferably, the promoter includes T7, trc, lac, ara, or λL; more preferably, the recombinant vector is selected from plasmids PET.28a(+), PET.21a(+), or PET.32a(+).

[0073] A helicase-expressing gene is inserted into a recombinant vector, and the recombinant vector's ability to replicate itself is exploited to replicate the helicase gene in large quantities. "Recombinant" here refers to genetically engineered DNA produced by transplanting or splicing a gene from one species into the cells of a host organism of a different species. This DNA becomes part of the host's genetic structure and is replicated.

[0074] According to a typical embodiment of the present application, a host cell is provided, wherein the host cell is transformed with the above-mentioned DNA molecule or recombinant vector.

[0075] The recombinant vector is transformed into a host cell, and the host cell replicates, transcribes, and translates the helicase expression gene on the recombinant vector to produce a large amount of helicase. The host cell includes Escherichia coli, which can be BL21(DE3), BL21Star(DE3)pLysS, Rossata(DE3), Lemo21(DE3), etc. The helicase of the present invention can be successfully expressed in the E. coli recombinant protein expression system, and the protein is uniform and high in purity.

[0076] According to a typical embodiment of the present application, this type of helicase exhibits superior unwinding activity in a high-salt environment than in a low-salt environment, can bind well to single-stranded DNA, and unwind double-stranded DNA. This type of helicase has strong unwinding activity, and the limit sequence blocking Spacer-18 (Sp18) cannot completely block its unwinding activity. This helicase can be used for the control and characterization of nucleic acids and applied to single-molecule nanopore sequencing. Nucleic acid control includes controlling the speed at which nucleic acids pass through the nanopore, controlling the stability of nucleic acid perforation, or controlling the persistence of nucleic acid perforation; further, the application of this helicase includes the application of nanosensors and single-molecule nanopore sequencing applications. Among them, sequencing stability refers to that the DNA to be tested enters the nanopore at a constant speed, the sequencing signal is stable and complete, the signal-to-noise ratio is high, and the sequencing quality will not decrease significantly as the sequencing time increases. Sequencing persistence refers to the continuous output of the sequencing signal until the sequencing of the molecule to be tested is completed, the library is continuously captured, and there is no sudden interruption resulting in low coverage and accuracy of sequencing.

[0077] The beneficial effects of the present application will be further explained in detail below with reference to specific embodiments.

[0078] Example 1

[0079] Cloning, expression, and purification of BCH326 and BCH338

[0080] 1. Cloning and Expression of BCH326 and BCH338

[0081] The full-length BCH326 and BCH338 DNA sequences were ligated into the PET.28a(+) plasmid, respectively, using double restriction enzyme sites Nde1 and Xho1, so that the expressed BCH326 and BCH338 proteins had a 6*His tag and a thrombin restriction site at the N-terminus.

[0082] Transform the cloned PET.28a(+)-BCH326 and PET.28a(+)-BCH328 plasmids into E. coli expression bacteria BL21(DE3) or its derivatives. Pick a single colony and inoculate it into 20 mL of LB medium containing kanamycin resistance and culture it at 37°C with shaking overnight. Then transfer it into 2 L of LB medium containing kanamycin resistance and culture it at 37°C with shaking until the OD 600 =0.6-0.8, cooled to 16°C, and IPTG was added to a final concentration of 500 μM to induce expression overnight.

[0083] 2. Purification of BCH326 and BCH338

[0084] Buffer A: 20mM Tris-HCl pH 7.5, 250mM NaCl, 20mM imidazole

[0085] Buffer B: 20 ​​mM Tris-HCl pH 7.5, 250 mM NaCl, 300 mM imidazole

[0086] Buffer C: 20mM Tris-HCl pH 7.5, 80mM NaCl

[0087] Buffer D: 20mM Tris-HCl pH 7.5, 1000mM NaCl

[0088] Buffer E: 20mM Tris-HCl pH 7.5, 200mM NaCl

[0089] The expressed BCH326 and BCH338 cells were harvested, resuspended in Buffer A, disrupted using a cell disruptor, and centrifuged to collect the supernatant. The supernatant was mixed with Ni-NTA medium pre-equilibrated with Buffer A and allowed to bind for 1 hour. The medium was collected and washed extensively with Buffer A until all contaminants were removed. Next, Buffer B was added to the medium to elute the protein. The eluted protein was passed through a desalting column (Cytiva, Sephadex G-25) equilibrated with Buffer C, replacing the buffer from Buffer B to Buffer C. The protein solution from the desalting column was then added to ssDNA cellulose medium equilibrated with Buffer C. An appropriate amount of coagulation protease was added. This enzyme specifically recognizes the thrombin cleavage site amino acid sequence LVPRG↓S in the vector sequence PET28(a)+, thereby cleaving the affinity His tag from the protein. This was performed at 4°C and incubated overnight on a rotary shaker. The next day, the ssDNA cellulose medium was collected, at which point the target protein was specifically adsorbed to the ssDNA medium. The ssDNA cellulose packing material was washed 3-4 times with Buffer C to remove any contaminating proteins that were not adsorbed to the packing material. Elution was then performed with Buffer D to disrupt the specific adsorption of the target protein to the packing material and elute the target protein into the solution. The protein purified from the ssDNA cellulose was concentrated using a 30K ultrafiltration concentrator (Merck Millipore) in a pre-cooled centrifuge at 4°C. The centrifugation parameters were set at 3000g for 10 minutes each, and the final volume was concentrated to 2 mL. Finally, the protein was filtered through a Superdex 200 molecular sieve (Cytiva) using Buffer E as the molecular sieve buffer. The target protein peak was collected, concentrated, and frozen. As shown in Figure 1, the purification yielded a relatively high amount of pure BGH326 protein with a uniform peak shape. An average yield of 0.28 mg of the target protein was obtained per liter of expression bacteria, comparable to the yield of helicase Dda (0.23 mg per liter of expression bacteria).

[0090] As shown in Figure 2, the purified protein, BGH338, was purified in large quantities with a uniform peak shape. On average, 0.42 mg of the target protein was purified per liter of expression bacteria, which is higher than the yield of helicase Dda (0.23 mg per liter of expression bacteria).

[0091] 3. Amino acid sequence of BCH326 (SEQ ID NO. 1)

[0092]

[0093] 4. DNA sequence of BCH326 (SEQ ID NO. 2)

[0094]

[0095] 5. Amino acid sequence of BCH338 (SEQ ID NO. 3)

[0096]

[0097] 6. DNA sequence of BCH338 (SEQ ID NO. 4)

[0098]

[0099]

[0100] 7. AlphaFold2 structure prediction of BCH326 and BCH338

[0101] With the help of AlphaFold2, BCH326 and BCH338 were predicted and their structure diagrams were obtained. The root mean square error (RMSD) between the predicted value and the true value of the protein skeleton structure reached and As shown in Figures 3 and 4, different secondary structures are indicated by different shapes: helix, sheet, and loop. It can be seen that, compared with the Dda helicase (shown in Figure 20, PDB number 3UPU), these two helicases each have two tower domains, both located on the same side of the protein, while the Dda helicase has only one tower domain.

[0102] Example 2

[0103] ATPase activity detection of BCH326 and BCH338 proteins

[0104] 1. Preparation of double-stranded DNA (ovDNA) and single-stranded DNA (ssDNA): Sequence ID No. 5 and Sequence ID No. 6 were annealed to form ovDNA with 20 T residues overhanging the 5' end. The annealing process was incubation at 95°C for 5 minutes, followed by a cooling rate of 0.1°C / s to 25°C, and continued incubation for 30 minutes. The annealing recipe is shown in Table 1. Sequence ID No. 6 (100 μM) was diluted to 10 μM in TE buffer (pH 8) to prepare ssDNA.

[0105] Table 1. ovDNA annealing recipe

[0106] Solution volume: 100 μM SEQ ID NO. 5 5 μL 100 μM SEQ ID NO. 6 5 μL TE buffer (pH = 8) 40 μL

[0107] 2. Prepare high salt reaction buffer (2×): 20 mM HEPES (pH 8.0), 4 mM ATP, 4 mM MgCl2, 1.0 M KCl.

[0108] 3. Protein dilution: Dilute BCH326 and BCH338 proteins to 10 μM using 1× PBS.

[0109] 4. Perform ATP hydrolysis reaction: Add the corresponding reagents according to the reaction system in Table 2, incubate at 30°C for 30 minutes, and inactivate at 80°C for 5 minutes. ①② are experimental groups, and ③④⑤⑥ are corresponding control groups, with three replicates per group.

[0110] Table 2. ATP hydrolysis reaction system

[0111] No. Reaction buffer (2×) DNA Protein H2O ①10μL 1μL (ovDNA) 1μL 8μL ②10μL 1μL (ssDNA) 1μL 8μL ③10μL——1μL9μL ④10μL 1μL (ovDNA)——9μL ⑤10μL 1μL (ssDNA)——9μL ⑥10μL————10μL

[0112] 5. Detection of the remaining ATP in the reaction: According to the manufacturer's instructions, the ATP detection kit (Biyuntian, S0026B) was used to determine the remaining ATP concentration in the reaction.

[0113] 6. Experimental results: As shown in Figures 5 and 6, under high salt conditions, both BCH326 and BCH338 have the activity of hydrolyzing ATP.

[0114] Example 3

[0115] Detection of dsDNA melting activity of BCH326 and BCH338 proteins

[0116] 1. Preparation of double-stranded DNA (ovDNA): Sequence ID No. 7 and Sequence ID No. 8 were annealed to form ovDNA with 20 T residues overhanging the 5' end. The annealing process was as follows: incubate at 95°C for 5 minutes, cool to 25°C at a rate of 0.1°C / s, and incubate for 30 minutes. The annealing recipe is shown in Table 3.

[0117] Table 3. ovDNA annealing recipe

[0118] Solution volume: 100 μM SEQ ID NO. 7 5 μL 100 μM SEQ ID NO. 8 5 μL TE buffer (pH = 8) 40 μL

[0119] 2. Prepare reaction buffer: low salt reaction buffer 1 is 100 mM HEPES (pH = 8.0), 1 mg / mL BSA, 10 mM MgCl2, 150 mM KCl; high salt reaction buffer 2 is 100 mM HEPES (pH = 8.0), 1 mg / mL BSA, 10 mM MgCl2, 500 mM KCl.

[0120] 3. Prepare the reaction solution: Add 3 μL of 10 μM annealed ovDNA, 6 μL of 100 μM SEQ ID NO. 9 (as competitor DNA to capture unwound single DNA strands), and 6 μL of 100 mM ATP to 585 μL of low-salt reaction buffer or high-salt reaction buffer as the experimental reaction solution. Add 1 μL of 10 μM SEQ ID NO. 8, 2 μL of 100 μM SEQ ID NO. 9, and 2 μL of 100 mM ATP to 195 μL of low-salt reaction buffer or high-salt reaction buffer as the positive control solution.

[0121] 4. Protein dilution: Dilute BCH326 and BCH338 proteins to 4.8 μM using 1× PBS.

[0122] 5. Prepare the melt reaction: Divide into experimental group ①, negative group ②, and positive group ③. Add the corresponding reagents according to Table 4. Use a microplate reader to detect the kinetic changes of fluorescence intensity within 30 minutes at 30°C. Repeat three times for each group.

[0123] 6. Data analysis: Calculate the percentage of the fluorescence value of the experimental group and the negative control group relative to the fluorescence value of the positive control group.

[0124] Table 4. Melting reaction recipe

[0125] No. Solution 1 Solution 2 ① 58.5 μL experimental reaction solution 1.5 μL protein ② 58.5 μL experimental reaction solution 1.5 μL reaction buffer ③ 58.5 μL positive control solution 1.5 μL reaction buffer

[0126] 7. Experimental results:

[0127] Within the tolerance range and instrument fluctuations, the experimental results were plotted by calculating the ratio of the measured fluorescence value to the fluorescence value of the positive control (due to instrument sensitivity, the negative control group had fluorescence absorption readings). The experimental results show that the negative control group in each experiment remained unchanged during the measurement process, while the fluorescence value of the experimental group gradually increased with the increase in reaction time, indicating that the activity of the cell to unwind double-stranded DNA is 5'-3', and the unwinding direction is 5'-3'.

[0128] As shown in Figures 7 and 8 , BCH326 exhibits dsDNA unwinding activity under both low-salt (final KCl concentration of 150 mM) and high-salt (final KCl concentration of 500 mM) conditions, and this activity increases with increasing salt concentration. As shown in Figures 9 and 10 , BCH338 exhibits dsDNA unwinding activity under both low-salt and high-salt conditions, and this activity increases with increasing salt concentration.

[0129] Example 4

[0130] Limiting sequence blocking BCH326 and BCH338 melting activity detection

[0131] 1. Preparation of double-stranded DNA (ovDNA) containing a restriction sequence: Sequence ID No. 7 and Sequence ID No. 10 were annealed to form ovDNA (containing a restriction sequence) with 20 T residues overhanging the 5' end. The annealing process was incubation at 95°C for 5 minutes, cooling at a rate of 0.1°C / s to 25°C, and incubation for 30 minutes. The annealing recipe is shown in Table 5.

[0132] Table 5. Annealing formula of ovDNA (containing restriction sequences)

[0133] Solution volume: 100 μM SEQ ID NO. 7 5 μL 100 μM SEQ ID NO. 10 5 μL TE buffer (pH = 8) 40 μL

[0134] 2. Prepare reaction buffer: low salt reaction buffer is 100 mM HEPES (pH = 8.0), 1 mg / mL BSA, 10 mM MgCl2, 150 mM KCl; high salt reaction buffer is 100 mM HEPES (pH = 8.0), 1 mg / mL BSA, 10 mM MgCl2, 500 mM KCl.

[0135] 3. Prepare the reaction solution: Add 3 μL of 10 μM annealed ovDNA (containing the restriction endonuclease sequence), 6 μL of 100 μM SEQ ID NO. 9 (20x competitor DNA), and 6 μL of 100 mM ATP to 585 μL of low-salt reaction buffer or high-salt reaction buffer for the experimental reaction. Add 1 μL of 10 μM SEQ ID NO. 11, 2 μL of 100 μM SEQ ID NO. 9 (20x competitor DNA), and 2 μL of 100 mM ATP to 195 μL of low-salt reaction buffer or high-salt reaction buffer for the positive control.

[0136] 4. Protein dilution: Dilute BCH326 and BCH338 proteins to 4.8 μM using 1× PBS.

[0137] 5. Prepare the melt reaction: Divide into experimental group ①, negative group ②, and positive group ③. Add the corresponding reagents according to Table 6. Use a microplate reader to detect the kinetic changes of fluorescence intensity within 30 minutes at 30°C. Repeat three times for each group.

[0138] 6. Data analysis: Calculate the percentage of the fluorescence value of the experimental group and the negative control group relative to the fluorescence value of the positive control group.

[0139] Table 6. Melting reaction recipe

[0140] No. Solution 1 Solution 2 ① 58.5 μL experimental reaction solution 1.5 μL protein ② 58.5 μL experimental reaction solution 1.5 μL 1× PBS ③ 58.5 μL positive control solution 1.5 μL 1× PBS

[0141] 7. Experimental results:

[0142] As shown in Figures 11 and 12, under both low-salt (final KCl concentration of 150 mM) and high-salt (final KCl concentration of 500 mM) conditions, the restriction sequence weakened BCH326's dsDNA unwinding activity but did not completely block its unwinding activity. Furthermore, BCH326 maintained a sustained unwinding activity trend under high-salt conditions. As shown in Figure 13, under low-salt conditions, the restriction sequence almost completely blocked BCH338's dsDNA unwinding activity. As shown in Figure 14, under high-salt conditions, the restriction sequence weakened BCH338's dsDNA unwinding activity but did not completely block its dsDNA unwinding activity.

[0143] Example 5

[0144] Nanopore sequencing applications of BCH326 and BCH338 proteins

[0145] 1. Two partially complementary DNA strands (top strand, SEQ ID NO. 11 and bottom strand, SEQ ID NO. 12) were annealed to form a linker. The linker was then ligated to the double-stranded target fragment using T4 DNA ligase at room temperature and purified to prepare a sequencing library. Figure 15 shows a schematic diagram of the linker (a: top strand; b: bottom strand).

[0146] 2. Incubate the BCH326 or BCH338 protein with the sequencing library at 25°C for 1 hour (molar concentration ratio 1:8) to form a sequencing library containing helicase. Figure 16 shows a schematic diagram of the sequencing library containing helicase (a: top strand; b: bottom strand; c: double-stranded target fragment; d: helicase; e: cholesterol-labeled double-stranded DNA).

[0147] 3. The helicase-containing sequencing library was incubated with single-stranded DNA (ssDNA-chol, SEQ ID NO. 13) containing cholesterol at its 5' end at room temperature for 10 minutes. The ssDNA-chol sequence is complementary to a portion of the bottom strand of the adaptor. Cholesterol binding to the phospholipid membrane reduces library loading and improves capture efficiency.

[0148] 4. Use a patch clamp amplifier or other electrical signal amplifier (as shown in Figure 17) to collect current signals. A Teflon membrane with a micrometer-sized pore (50-200 μm in diameter) in the center divides the electrolytic cell into two chambers: the cis chamber and the trans chamber. A pair of Ag / AgCl electrodes is placed in each chamber. A bilayer of phospholipid membranes is formed at the micropores of each chamber, and the nanopore protein CsgG-Eco-(Y51A / F56Q / R97W / R192D-StrepII(C)) is added. Electrical measurements are obtained after a single nanopore protein is inserted into the phospholipid membrane. The reaction product from step 3 is added, and 180 mV is applied. The sequencing library is captured by the nanopore, and the nucleic acids pass through the nanopore under the control of the helicase. The buffer used in this experiment is: 0.47 M KCl, 25 mM HEPES, 1 mM EDTA, 5 mM ATP, 25 mM MgCl2, pH 7.6, and the sequencing temperature is 28°C.

[0149] 5. Sequencing experiments were performed using the BCH326, and the sequencing signals are shown in Figure 18. Sequencing experiments were performed using the BCH338, and the sequencing signals are shown in Figure 19. The results show that as the helicase guides the DNA single strand into the nanopore, the current is partially blocked, decreasing. Due to the varying sizes of different nucleotides, the magnitude of the blocked current also varies, resulting in the visible fluctuations in the current signal. Both Figures 18 and 19 show intact adapter and response signals, with a high signal-to-noise ratio, indicating good sequencing signal stability.

[0150] SEQ ID NO.5: 5'-GCGTCGAAAAGCAGTACTTAGGCATT-3'

[0151] SEQ ID NO.6: 5'-TTTTTTTTTTTTTTTTTTTTTAATGCCTAAGTACTGCTTTTCGACGC-3'

[0152] SEQ ID NO.7: 5'-BHQ-1-GCGTCGAAAAGCAGTACTTAGGCATT-3'

[0153] SEQ ID NO.8: 5'-TTTTTTTTTTTTTTTTTTTTTAATGCCTAAGTACTGCTTTTCGACGC-FAM-3'

[0154] SEQ ID NO.9: 5'-AATGCCTAAGTACTGCTTTTTCGACGCT-3'

[0155] SEQ ID NO.10: 5'-TTTTTTTTTTTTTTTTTTTTTNNNNAATGCCTAAGTACTGCTTTTCGACGC-FAM-3' (N=iSP18)

[0156] SEQ ID NO.11: 5'-TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTNNNNGGTTGTTTCTGTTGGTGCTGATATTGCT-3' (N=iSP18)

[0157] SEQ ID NO.12: 5'-GCAATATCAGCACCAACAGAAACAACCTTTGAGGCGAGCGGTCAA-3'

[0158] SEQ ID NO. 13: 5'-cholesterol-TTGACCGCTCGCCTC-3'.

[0159] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A helicase, characterized in that The helicase has two tower domains and one pin domain, and the two tower domains are located on the same side of the three-dimensional structure of the helicase.

2. The helicase according to claim 1, characterized in that The helicase comprises at least one of the following: A) BCH326, wherein BCH326 is a protein having an amino acid sequence as shown in SEQ ID NO: 1; B) BCH338, wherein BCH338 is a protein having an amino acid sequence as shown in SEQ ID NO: 3; C) a protein in which at least one cysteine ​​on the surface of the protein defined in A) or B) is mutated to alanine, glutamine, glycine, histidine, isoleucine, leucine, valine, serine, threonine or methionine; D) a protein having DNA unwinding ability, wherein the amino acid at at least one site on the tower domain and / or the pin domain of the amino acid sequence of the protein defined in any one of A), B) and C) is mutated to cysteine ​​or at least one unnatural amino acid is introduced; and E) A protein having an amino acid sequence homology of more than 70% with any of the proteins defined in A), B), C) and D) and having the same function.

3. The helicase according to claim 2, characterized in that Said C) comprises: A protein in which the amino acid C at position 319 of BCH326 is substituted with A, S, T, V, I, L or G; and A protein in which the amino acid C at position 326 or position 459 of BCH338 is substituted with A, S, T, V, I, L or G.

4. The helicase according to claim 2, characterized in that In said D), the amino acid mutation at at least one site on said tower domain and / or said pin domain to cysteine ​​or the introduction of at least one non-natural amino acid comprises at least one of the following: The tower domain of BCH326 has S389, R340, K341, S342, N343, K343, S344, I345, V346, I347, D348, K349, D350, G351, K352, A353, K354, E355, F356, L357, R358, K359, F360, L361, N362, F363, A364, K the amino acid in at least one of 365, I366, Y367, N368, F369, T370, N371, K372, G373, G374, H378, G379, R380, R381, I382, T383, K384, K385, S386, K387, K388, E389, L390 and W391 is mutated to cysteine ​​or at least one unnatural amino acid is introduced; The amino acid of at least one of D87, I88, G89, T90, I91, H92, S93, Y94, F95, D96, I97, K98, P99, D100, I101, D102, D103, N104, G105, N106, R107, V108, F109, K110, P111 or S112 on the pin domain of BCH326 is mutated to cysteine ​​or at least one unnatural amino acid is introduced; The tower domain of BCH338 has S405, K406, F407, L408, V409, P410, L411, G412, D413, G414, S415, K416, E417, D418, L419, F420, P421, L422, Y423, K424, E425, A426, V427, F428, D the amino acid in at least one of 429, I430, A431, K432, T433, M434, N435, N436, Q437, R438, K439, I440, S441, K442, N443, S444, K445, K446, N447, F448, or W449 is mutated to cysteine ​​or at least one unnatural amino acid is introduced; The amino acid of at least one of E93, I94, R95, P96, D97, I98, N99, E100, F101, G102, E103, R104, I105, F106, V107, P108, K109, L110, R111, D112, M113, and M114 on the pin domain of BCH338 is mutated to cysteine ​​or at least one unnatural amino acid is introduced; Preferably, in E), the protein has 70%, 80%, 90%, 95% or 99% or more homology and the same function as the amino acid sequence of the protein defined in any one of A), B), C) and D).

5. The helicase according to any one of claims 2 to 4, characterized in that The unnatural amino acid is selected from 4-azido-L-phenylalanine, 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselenoyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4-[(ethylsulfanyl)carbonyl]-L-phenylalanine, (2S)-2-amino-3-{4-[(propan-2-ylsulfanyl)carbonyl]phenyl 4-[(2-amino-3-sulfanylpropionyl)amino]phenyl}propionic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino-L-tyrosine, 3-iodo-L- Tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphth-2-ylamino)propionic acid, 6-(methylsulfanyl)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propionic acid, (2R)-2-aminooctanoate 3-(2,2′-bipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxy-3-quinolyl)propionic acid , 4-benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)sulfanyl]propionic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propionic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)oxy]carbonyl}amino)hexanoic acid and O-(2-nitrobenzyl)-L-tyrosine or 2-nitrophenylalanine; Preferably, the BCH326 introduces at least one unnatural amino acid including at least one of the following: D100 introduces 4-azido-L-phenylalanine, I101 introduces 4-azido-L-phenylalanine, D102 introduces 4-azido-L-phenylalanine, D103 introduces 4-azido-L-phenylalanine, N104 introduces 4-azido-L-phenylalanine, G105 introduces 4-azido-L-phenylalanine, N106 introduces 4-azido-L-phenylalanine, R107 introduces 4-azido-L-phenylalanine, D103 introduces 4-acetyl-L-phenylalanine, G105 introduces 4-acetyl-L-phenylalanine and N106 introduces 4-acetyl-L-phenylalanine; Preferably, the BCH338 introduces at least one unnatural amino acid including at least one of the following: A431 introduces 4-azido-L-phenylalanine, K432 introduces 4-azido-L-phenylalanine, T433 introduces 4-azido-L-phenylalanine, M434 introduces 4-azido-L-phenylalanine, N435 introduces 4-azido-L-phenylalanine, S441 introduces 4-azido-L-phenylalanine, K442 introduces 4-azido-L-phenylalanine, N443 introduces 4-azido-L-phenylalanine and S444 introduces 4-azido-L-phenylalanine.

6. The helicase according to any one of claims 1 to 4, characterized in that There is at least one amino acid mutation at the amino acid site of the DNA binding region of the helicase and / or the amino acid site near the ATP catalytic active center, wherein the mutation includes mutating the original amino acid to an amino acid with a larger side chain; Preferably, the mutation of the original amino acid to an amino acid with a larger side chain comprises at least one of the following: asparagine is replaced by glutamine, histidine, arginine or lysine; proline is replaced by arginine, lysine, phenylalanine or leucine; histidine is replaced by arginine, lysine, glutamine, asparagine phenylalanine, tyrosine or tryptophan; proline is replaced by arginine, lysine, glutamine, asparagine or histidine; phenylalanine is replaced by arginine, lysine, histidine, tyrosine or tryptophan; isoleucine is replaced by phenylalanine, tryptophan, histidine, lysine or arginine; tyrosine is replaced by arginine, lysine or tryptophan; The amino acid sites in the DNA binding region of BCH326 include: L157, V160, L294, G296, N299, L303, A304, I328, F329, T330, N331, G332, G333, and E334, and the amino acid sites near the ATP catalytic active center include: K211, E212, E213, N214, Y215, K216, A217, P218, L219, K220, D221, I222, N223, and N224; The amino acid sites in the DNA binding region of BCH338 include: H89, S90, Y91, F92, E93, I94, R95 and P96; the amino acid sites near the ATP catalytic active center include: Y152, Q153, L154, P155, P156, V157, F193, L194, I195, K196, E197, Y198, E199, E200 and N201.

7. The helicase according to any one of claims 1 to 4, characterized in that The amino acid on the surface of the helicase that interacts with the nanopore binding region has at least one mutation, wherein the mutation includes mutating the original amino acid to an amino acid with a shorter side chain; Preferably, the mutation of the original amino acid to an amino acid with a shorter side chain includes: asparagine is replaced by isoleucine, valine, isoleucine, alanine, serine or glycine; lysine is replaced by isoleucine, valine, isoleucine, alanine, serine or glycine; lysine is replaced by isoleucine, valine, isoleucine, alanine, serine or glycine; arginine is replaced by isoleucine, valine, isoleucine, alanine, serine or glycine; Preferably, the amino acids on the surface of BCH326 that interact with the nanopore binding region include: M1, E2, S3, K4, I5, N6, L7, T8, E9, D10, Q11, L12, K13, I14, I15, K16, I189, I190, R191, T192, Q193, N194, K195, N196 and S197; The amino acids on the surface of BCH338 that interact with the nanopore binding region include: M1, G2, E3, I4, K5, L6, N7, E8, E9, Q10, Q11, K12, K177, I177, L178, R179, T180, K181, N182, L213, I214, D215, H216, F217, H218, V219, Y220, G221, D248, L249, T250, D251, S252, T253, E254, and S255.

8. An isolated DNA molecule, characterized in that The DNA molecule has (a) a nucleotide sequence encoding the helicase according to any one of claims 2 to 3; or (b) a nucleotide sequence that hybridizes under stringent conditions to the DNA molecule defined in (a); or (c) having the nucleotide sequence shown in SEQ ID NO: 2 or SEQ ID NO: 4; or (d) A nucleotide sequence having 70% or more homology with any one of the nucleotide sequences defined in (a) to (c) and encoding a protein having the same function as the helicase.

9. The DNA molecule according to claim 8, characterized in that The DNA molecule has a nucleotide sequence that has 75% or more, preferably 85% or more, more preferably 95% or more, and further preferably 99% or more homology with any one of the nucleotide sequences defined in (a) to (c) and encodes a protein having the same function.

10. A recombinant vector, characterized in that: The recombinant vector comprises the DNA molecule according to claim 8 or 9.

11. The recombinant vector according to claim 10, characterized in that The recombinant vector is selected from a plasmid, a virus or a carrier expression vector; Furthermore, the recombinant vector includes a regulatory element for controlling the expression of the DNA molecule; Furthermore, the regulatory element includes a promoter operably linked to the DNA molecule; Preferably, the promoter comprises T7, trc, lac, ara or λL; More preferably, the recombinant vector is selected from plasmid PET.28a(+), PET.21a(+) or PET.32a(+).

12. A host cell, characterized in that The host cell contains the DNA molecule according to claim 8 or 9, or the recombinant vector according to claim 10 or 11.

13. The host cell according to claim 12, characterized in that The host cell includes Escherichia coli; Preferably, the host cell comprises BL21(DE3), BL21 Star(DE3)pLysS, Rossata(DE3) or Lemo21(DE3).

14. Use of the helicase according to any one of claims 1 to 7 in nucleic acid control or characterization; Further, the nucleic acid control includes controlling the speed of the nucleic acid passing through the nanopore, controlling the stability of the nucleic acid perforation, or controlling the continuity of the nucleic acid perforation; Furthermore, the application includes application in nanosensors and / or application in single-molecule nanopore sequencing.

15. A nanopore sequencing kit, comprising a helicase, characterized in that: The helicase is the helicase according to any one of claims 1 to 7.

16. A method for nanopore sequencing, comprising sequencing a nucleic acid molecule to be sequenced under the control of a helicase, characterized in that: The helicase is the helicase according to any one of claims 1 to 7.