Ski2-like helicase and use thereof

By performing amino acid replacement and deletion of specific sites on Ski2-like helicase, a stable pore is formed, and the decoil function from the 3' end to the 5' end is achieved, solving the problem of localization of helicase direction in nanopore sequencing technology, and improving the accuracy and application range of sequencing.

WO2025138248A1PCT designated stage expired Publication Date: 2025-07-03SHENZHEN HUADA GENE INST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/143610
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Among the existing nanopore sequencing technologies, there is a lack of helicase suitable for dislodging from the 3' end to the 5' end, which limits the application scenarios and accuracy of the sequencing technology.

Method used

A Ski2-like helicase is developed to form stable pores by replacing and deleting cysteine ​​or non-natural amino acids at specific sites, controlling the speed of nucleic acids passing through nanopores, and having the function of unrotating from the 3' end to the 5' end.

Benefits of technology

The application scenarios of nanopore sequencing have been broadened, and the accuracy and efficiency of nucleic acid sequence detection have been improved, especially the detection of nucleic acid sequences with base modification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023143610_03072025_PF_FP_ABST
    Figure CN2023143610_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a Ski2-like helicase and a use thereof. The Ski2-like helicase comprises: 1) a protein having an amino acid sequence as shown in any one of SEQ ID NO: 1 to 4; or 2) a protein in which at least one site in the 2A domain and / or the Ratchet domain of the amino acid sequence in 1) is replaced by cysteine or an unnatural amino acid, and / or at least one cysteine in any one or more domains of the 2A domain, WH domain, and Ratchet domain is deleted or replaced by a natural amino acid other than cysteine, and which has helicase activity for unwinding nucleic acids in the 3' to 5' direction. The described helicase can form stable channels, has the function of unwinding nucleic acids in the 3' to 5' direction, and can broaden the application scenarios for nanopore sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

Ski2-like helicase and its application Technical Field

[0001] The present invention relates to the field of enzyme technology, and in particular to a Ski2-like helicase and applications thereof. Background Art

[0002] Nanopore gene sequencing technology is a new generation of sequencing technology that has emerged in recent years. Compared with mature second-generation sequencing technology, nanopore sequencing has the advantages of providing relatively low-cost genotyping, high test mobility, long read length, rapid sample processing and real-time monitoring. It has important application prospects in the rapid identification of viral pathogens, environmental monitoring, food safety monitoring, human genome sequencing, plant genome sequencing, antibiotic resistance monitoring, etc.

[0003] Nanopore sequencing technology relies on a nanoscale protein pore, or "nanopore," which acts as a biosensor embedded in a biological or artificial membrane. A motor protein then pulls nucleic acids through the nanopore. Powering nanopore nucleic acid sequencing is a pair of membrane-separated electrolytic chambers. When a voltage is applied to the membrane, one side becomes positively charged and the other negatively charged. Driven by the motor protein, nucleic acid molecules experience the positive charge, forcing them into the pore and occupying most of the space within it. Only a relatively small number of small ions are allowed to pass through the membrane, resulting in a change in the current signal received by the electrodes on either side of the membrane. Because the nanopore only allows the passage of a single nucleic acid polymer, and the individual ATCG bases have different charge properties, different bases interfere with the current flow differently when passing through the protein nanopore. By monitoring and decoding these current signals in real time, the base sequence can be determined, enabling sequencing.

[0004] However, if left untreated, a linear nucleic acid chain can rapidly pass through the nanopore, resulting in significant errors in current signal analysis. Therefore, slowing the passage of nucleic acids through the nanopore is crucial for the development of nanopore sequencing technology. In addition to controlling the rate at which nucleic acid chains pass through the nanopore, motor proteins also possess unwinding activity, allowing double-stranded DNA or RNA-DNA hybrids to unwind into single-stranded molecules that pass through the nanopore. Furthermore, currently available helicases for nanopore sequencing mostly sequence from the 5' to the 3' end of nucleic acids, while few helicases are suitable for sequencing from the 3' to the 5' end of nucleic acids. This limits the application of nanopore sequencing technology. Therefore, the development of newer helicases for nanopore sequencing is particularly urgent.

[0005] Summary of the Invention

[0006] The main purpose of the present invention is to provide a Ski2-like helicase and its application to solve the problem of immaturity of nanopore sequencing technology from the 3' end to the 5' end in the prior art.

[0007] In order to achieve the above object, according to the first aspect of the present invention, a Ski2-like helicase is provided, wherein the Ski2-like helicase comprises: 1) a Ski2-like helicase having SEQ ID NOs: A protein having an amino acid sequence as shown in any one of 1 to 4, and comprising a 1A domain, a 2A domain, a WH domain, a Ratchet domain and a HLH domain; or 2) a protein having at least one of the following variations in the amino acid sequence in 1), and having helicase activity that unwinds the nucleic acid from the 3' end to the 5' end: i) at least one site is substituted with cysteine ​​or a non-natural amino acid, and at least one site is selected from at least one of the following domains: a 2A domain or a Ratchet domain; ii) at least one cysteine ​​site is deleted or substituted with a natural amino acid, and at least one cysteine ​​site is selected from at least one of the following domains: a 2A domain, a WH domain or a Ratchet domain; or 3) an amino acid sequence having more than 50%, preferably more than 80%, more preferably more than 90%, and further preferably more than 95% homology to the amino acid sequence defined in any one of 1) and 2), and having helicase activity that unwinds the nucleic acid from the 3' end to the 5' end.

[0008] Further, at least one site in i) of 2) is independently selected from at least one of the following: sites in the 2A domain of the amino acid sequence as shown in SEQ ID NO: 1: E279, G280, I281, S282, L283, V284, D285, G286, E287, P288, S289 or S290; sites in the 2A domain of the amino acid sequence as shown in SEQ ID NO: 2: S275, A276, D277, E278, G279, N280, I281, A282, L283, V284, D285, G286, E287, P288, S289, S290 or I291; sites in the 2A domain of the amino acid sequence as shown in SEQ ID NO: The sites in the 2A domain of the amino acid sequence shown in SEQ ID NO: 3 are: K269, G270, I271, D272, E273, D274, L275, G276, D277, D278 or K279; the sites in the 2A domain of the amino acid sequence shown in SEQ ID NO: 4 are: A279, P280, G281, E282, D283, V284, M285, M286, G287, E288, V289, D290, V291, E292, E293, A294, S295 or S296; preferably, the sites in the 2A domain of the amino acid sequence shown in SEQ ID NO: 5 are: the sites in the Ratchet domain of the amino acid sequence set forth in SEQ ID NO:1: E586, I587, P588, E589, K590, E591, M592, E593, K594, L595, Y596, or G597; the sites in the Ratchet domain of the amino acid sequence set forth in SEQ ID NO:2: T590, P591, E592, K593, E594, M595, E596, K597, R598, Y599, G600, V601, G602, or P603; the sites in the Ratchet domain of the amino acid sequence set forth in SEQ ID NO:3: E568, E569, V570, P571, E572, K573, R574, I575, T576, D577, R578, or Y579; the sites in the Ratchet domain of the amino acid sequence set forth in SEQ ID NO:4: Sites in the Ratchet domain of the amino acid sequence shown in NO:4: E599, E600, K601, D602, D603, D604, T605, L606, A607, E608, S609 or F610.

[0009] Further, at least one cysteine ​​site in ii) of 2) is independently selected from at least one of the following: site C298 or C329 in the 2A domain of the amino acid sequence shown in SEQ ID NO: 1; site C298 or C329 in the 2A domain of the amino acid sequence shown in SEQ ID NO: 2; site C304, C335 or C379 in the 2A domain of the amino acid sequence shown in SEQ ID NO: 4; preferably, site C441 or C458 in the WH domain of the amino acid sequence shown in SEQ ID NO: 4; preferably, site C529 in the Ratchet domain of the amino acid sequence shown in SEQ ID NO: 1; site C532 in the Ratchet domain of the amino acid sequence shown in SEQ ID NO: 2.

[0010] Furthermore, the unnatural amino acids include: 4-azido-L-phenylalanine, 4-azido-L-phenylalanine, 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselenoyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4-(dihydroxyboryl)-L-phenylalanine, 4-[(ethylsulfanyl)carbonyl]-L-phenylalanine, (2S)-2-amino-3-{4-[( propan-2-ylsulfanyl)carbonyl]phenyl}propionic acid, (2S)-2-amino-3-{4-[(2-amino-3-sulfanylpropionyl)amino]phenyl}propionic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino-L -tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthylamino)propionic acid, 6-(methylsulfanyl)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propionic acid, (2R)-2-aminooctanoate 3-(2,2′-bipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxy-

[00145] The following examples include 4-[(2-nitrobenzyl)-3-[(2-nitrobenzyl)-1-[(2-nitrobenzyl)-1-[(2-nitrobenzyl)-2-thiazolyl)]-1-[(2-nitrobenzyl)-3-[(2-nitrobenzyl)-1-oxy]propionic acid, (2S)-2-amino-3-[(2-nitrobenzyl)-1-oxy]propionic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)-1-oxy]carbonyl}amino)hexanoic acid, O-(2-nitrobenzyl)-L-tyrosine, or 2-nitrophenylalanine.

[0011] Furthermore, there is a connection relationship between the site in the 2A domain of the amino acid sequence in i) of 2) that is substituted with cysteine ​​or a non-natural amino acid and the site in the Ratchet domain that is substituted with cysteine ​​or a non-natural amino acid, preferably, the connection includes non-covalent connection, covalent connection, or covalent and non-covalent connection; preferably, the covalent connection is performed using a cross-linker, and the cross-linker includes: maleimide, succinimide, active ester, azide, alkyne, phosphine, haloacetyl, phosgene-type reagent, sulfonyl chloride reagent, isothiocyanate, acyl halide, hydrazine, disulfide, vinyl sulfone, aziridine or a photosensitizer.

[0012] In order to achieve the above object, according to the second aspect of the present invention, a DNA molecule encoding the above-mentioned Ski2-like helicase is provided.

[0013] Furthermore, the DNA molecule is selected from: A) a polynucleotide comprising a nucleotide sequence as shown in any one of SEQ ID NOs: 5 to 8; or B) a polynucleotide comprising a nucleotide sequence as shown in any one of SEQ ID NOs: 5 to 8 having a homology of more than 80%, more preferably more than 90%, and even more preferably more than 95% to a polynucleotide.

[0014] In order to achieve the above object, according to the third aspect of the present invention, a recombinant vector is provided, which contains the above DNA molecule.

[0015] In order to achieve the above object, according to a fourth aspect of the present invention, a host cell is provided, wherein the host cell contains the above recombinant vector.

[0016] Furthermore, the host cell includes a prokaryotic cell or a eukaryotic cell.

[0017] In order to achieve the above-mentioned purpose, according to the fifth aspect of the present invention, a kit is provided, comprising a Ski2-like helicase, wherein the Ski2-like helicase is the above-mentioned Ski2-like helicase.

[0018] Furthermore, the kit also includes a sequencing chip with nanopore protein and a sequencing buffer; preferably, the sequencing buffer includes: 0.1-1.0M KCl, 10-100mM HEPES, 1-50mM ATP and 10-100mM MgCl2; preferably, the pH value of the sequencing buffer is 6-10.

[0019] Furthermore, the kit also includes a sequencing adapter and a single-stranded DNA with cholesterol modification at the 5' end; wherein the single-stranded DNA and the sequencing adapter have a complementary region; preferably, the complementary region is 10-50 nt long.

[0020] To achieve the above-mentioned object, according to the sixth aspect of the present invention, a sequencing method is provided, which comprises: using the above-mentioned Ski2-like helicase or the Ski2-like helicase in the above-mentioned kit to control the nucleic acid fragment to be tested to pass through the nanopore, and sequencing from the 3' end to the 5' end; sequencing comprises: analyzing the electrical signal generated when the nucleic acid fragment passes through the nanopore, thereby determining the sequence of the nucleic acid fragment.

[0021] Furthermore, the method further comprises: preparing the nucleic acid sample to be tested into a nucleic acid fragment to be tested with a sequencing adapter; and combining the nucleic acid fragment to be tested with the sequencing adapter with a Ski2-like helicase.

[0022] Furthermore, nanopore sequencing is performed at 20-40° C.; preferably, the molar concentration ratio of Ski2-like helicase to the nucleic acid fragment to be detected is 100:1-1:1.

[0023] To achieve the above-mentioned object, according to the seventh aspect of the present invention, a sequencing library complex is provided, which comprises: a sequencing library to be tested and a Ski2-like helicase; wherein the Ski2-like helicase is the above-mentioned Ski2-like helicase or the Ski2-like helicase in the above-mentioned kit; the sequencing library to be tested comprises a nucleic acid fragment to be tested with a sequencing adapter, and the Ski2-like helicase is bound to the sequencing adapter.

[0024] To achieve the above-mentioned object, according to an eighth aspect of the present invention, a method for preparing a sequencing library complex is provided, the preparation method comprising: preparing a nucleic acid sample to be tested into a nucleic acid fragment to be tested with a sequencing adapter, incubating the nucleic acid fragment to be tested with the sequencing adapter with a Ski2-like helicase to obtain a sequencing library complex, and binding the Ski2-like helicase of the sequencing library complex to the sequencing adapter; or, incubating the sequencing adapter with the Ski2-like helicase to form an adapter complex, preparing the nucleic acid sample to be tested into a nucleic acid fragment to be tested with the adapter complex, and obtaining a sequencing library complex, and binding the Ski2-like helicase of the sequencing library complex to the sequencing adapter; wherein the Ski2-like helicase is the above-mentioned Ski2-like helicase or the Ski2-like helicase in the above-mentioned kit.

[0025] By applying the technical solution of the present invention, the interaction between the binding domains of the helicase and its mutants of the present application can form a stable pore, which can allow nucleic acids to pass through the pore stably during nanopore sequencing and control the speed of nucleic acids passing through the nanopore. In addition, the helicase of the present application has the function of unwinding the nucleic acid from the 3' end to the 5' end. The use of the helicase of the present application and its related mutants can broaden the application scenarios of nanopore sequencing. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0027] FIG1 is a schematic diagram showing the corresponding relationship of the amino acid sequences of Ski2-like helicases 1 to 4 of the present invention;

[0028] Among them, in Figure 1 , there are motif Q, motif Ⅰ, motif Ⅰa, motif Ⅰb, motif Ⅱ, motif Ⅲ, motif Ⅳa, motif Ⅳ, motif Ⅴ, and motif VI of the Ski2-like helicase, and the numbers above the sequence indicate the specific amino acid positions;

[0029] FIG2 shows the predicted tertiary structure of the unmutated parent of Ski2-like helicase 1 of the present invention;

[0030] FIG3 shows the predicted tertiary structure of the unmutated parent of Ski2-like helicase 2 of the present invention;

[0031] FIG4 shows the predicted tertiary structure of the unmutated parent of Ski2-like helicase 3 of the present invention;

[0032] FIG5 shows the predicted tertiary structure of the unmutated parent of Ski2-like helicase 4 of the present invention;

[0033] Figure 6 shows the purification diagram (A) and SDS-PAGE gel electrophoresis results (B) of Ski2-like helicase 1 in Example 1 of the present invention;

[0034] FIG7 shows a purification diagram (A) and SDS-PAGE gel electrophoresis results (B) of Ski2-like helicase 2 in Example 1 of the present invention;

[0035] FIG8 shows a purification diagram (A) and SDS-PAGE gel electrophoresis results (B) of Ski2-like helicase 3 in Example 1 of the present invention;

[0036] Figure 9 shows the purification diagram (A) and SDS-PAGE gel electrophoresis results (B) of Ski2-like helicase 4 in Example 1 of the present invention;

[0037] FIG10 shows the binding of Ski2-like helicase 1 and ssDNA at different mixing ratios in Example 2 of the present invention;

[0038] FIG11 shows the binding of Ski2-like helicase 2 and ssDNA at different mixing ratios in Example 2 of the present invention;

[0039] FIG12 shows the dsDNA depolymerization activity detection results of Ski2-like helicase 1 in Example 3 of the present invention;

[0040] FIG13 shows the dsDNA unwinding activity detection results of Ski2-like helicase 2 in Example 3 of the present invention;

[0041] FIG14 shows the dsDNA depolymerization activity detection results of Ski2-like helicase 3 in Example 3 of the present invention;

[0042] FIG15 shows the dsDNA depolymerization activity detection results of Ski2-like helicase 4 in Example 3 of the present invention;

[0043] FIG16 is a schematic diagram showing the sequencing electrical signals of Ski2-like helicase 1 in Example 4 of the present invention;

[0044] FIG17 is a schematic diagram showing the sequencing electrical signals of Ski2-like helicase 2 in Example 4 of the present invention;

[0045] FIG18 is a schematic diagram showing the sequencing electrical signals of Ski2-like helicase 3 in Example 4 of the present invention;

[0046] FIG19 is a schematic diagram showing the sequencing electrical signal of Ski2-like helicase 4 in Example 4 of the present invention. DETAILED DESCRIPTION

[0047] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the embodiments.

[0048] As mentioned in the background art, during the nanopore sequencing process, the nucleic acid molecules are driven by the motor protein and are subjected to a positively charged pulling force, entering the pore to complete the current signal monitoring of different bases, thereby achieving sequencing. If the motor protein can control and slow down the speed at which the nucleic acid passes through the nanopore, the accuracy of the current signal analysis can be improved. In addition to controlling the rate at which the nucleic acid chain passes through the nanopore, the motor protein also has unwinding activity, allowing double-stranded DNA or RNA-DNA hybrid double strands to unwind into single-stranded molecules through the nanopore. The sequencing direction of existing helicases is mostly from the 5' end to the 3' end of the nucleic acid, and there are very few helicases suitable for nanopore sequencing from the 3' end to the 5' end of the nucleic acid, which makes the application of nanopore sequencing technology relatively limited. Therefore, the present application attempts to develop a new type of helicase to broaden the application scenarios of nanopore sequencing (for example, nucleic acid sequences with base modifications at the 5' end, etc.) and / or more accurately detect current changes in nucleic acid sequences.

[0049] In a first typical embodiment of the present invention, a Ski2-like helicase is provided, wherein the Ski2-like helicase comprises: 1) a protein having SEQ ID NOs: A protein having an amino acid sequence as shown in any one of 1 to 4, and comprising a 1A domain, a 2A domain, a WH domain, a Ratchet domain and a HLH domain; or 2) a protein having at least one of the following variations in the amino acid sequence in 1), and having helicase activity that unwinds the nucleic acid from the 3' end to the 5' end: i) at least one site is substituted with cysteine ​​or a non-natural amino acid, and at least one site is selected from at least one of the following domains: a 2A domain or a Ratchet domain; ii) at least one cysteine ​​site is deleted or substituted with a natural amino acid, and at least one cysteine ​​site is selected from at least one of the following domains: a 2A domain, a WH domain or a Ratchet domain; or 3) an amino acid sequence having more than 50%, preferably more than 80%, more preferably more than 90%, and further preferably more than 95% homology to the amino acid sequence defined in any one of 1) and 2), and having helicase activity that unwinds the nucleic acid from the 3' end to the 5' end.

[0050] SEQ ID NO: 1: Wild-type amino acid sequence of Ski2-like helicase 1 before mutation

[0051] SEQ ID NO: 2: Wild-type amino acid sequence of Ski2-like helicase 2 before mutation

[0052] SEQ ID NO: 3: Wild-type amino acid sequence of Ski2-like helicase 3 before mutation

[0053] SEQ ID NO: 4: Wild-type amino acid sequence of Ski2-like helicase 4 before mutation

[0054] The helicases of the present application have the same structural domains, namely, the 1A domain, the 2A domain, the WH domain, the Ratchet domain, and the HLH domain. The predicted tertiary structures of the unmutated parent Ski2-like helicases 1 to 4 are shown in Figures 2 to 5 . The specific amino acid residue positions of each domain of the unmutated parent Ski2-like helicases 1 to 4 are shown in the table below:

[0055] The 1A and 2A domains of the helicase contain 10 conserved motifs that participate in the interaction between ATP and nucleic acids and are closely related to helicase activity. The conserved motifs in the unmutated parent of Ski2-like helicases 1 to 4 were identified through sequence alignment as shown in Figure 1. The WH domain and Ratchet domain of the helicase can connect with the 1A and 2A domains, respectively, to form a ring structure to accommodate the passage of single-stranded DNA. Both the HLH domain and the Ratchet domain can interact with nucleic acids.

[0056] The Ski2-like helicase of the present application is a variant protein in which amino acid substitutions or deletions are made to the wild type protein in 1) (such as the amino acid sequence shown in SEQ ID NO: 1-4), including: replacing any amino acid site in the 2A domain or the Ratchet domain with cysteine ​​or a non-natural amino acid; or replacing any amino acid site in the 2A domain and the Ratchet domain with cysteine ​​or a non-natural amino acid. The above mutations can enable the amino acids with sulfhydryl groups on the side chains in the 2A domain and the Ratchet domain to be covalently linked, thereby strengthening the binding ability of the 2A domain and the Ratchet domain, stabilizing the pore formed therefrom, and allowing the nucleic acid to pass through the pore stably during nanopore sequencing, thereby controlling the speed at which the nucleic acid passes through the nanopore.

[0057] In addition, on the basis of the above mutations, or separately, the protein in 1) also includes the following mutations: the cysteine ​​in at least one of the 2A domain, the WH domain or the Ratchet domain is deleted or replaced with a natural amino acid other than cysteine. It is possible to prevent the cysteine ​​in the above domains from producing unnecessary cross-linking with the cysteine ​​or non-natural amino acid after the replacement of the 2A domain or the Ratchet domain, thereby affecting the cross-linking between the 2A domain and the Ratchet domain. Since the helicase of the present application also has the function of unwinding from the 3' end to the 5' end of the nucleic acid (i.e., it can bind to the DNA single strand in the 3' end to the 5' end direction to unwind), the application scenarios of nanopore sequencing can be broadened by using the helicase of the present application and its related mutants.

[0058] As used herein, amino acid residues are abbreviated as follows: alanine (Ala; A), asparagine (Asn; N), aspartic acid (Asp; D), arginine (Arg; R), cysteine ​​(Cys; C), glutamic acid (Glu; E), glutamine (Gln; Q), glycine (Gly; G), histidine (His; H), isoleucine (Ile; I), leucine (Leu; L), lysine (Lys; K), methionine (Met; M), phenylalanine (Phe; F), proline (Pro; P), serine (Ser; S), threonine (Thr; T), tryptophan (Trp; W), tyrosine (Tyr; Y), and valine (Val; V).

[0059] Substitution and replacement rules generally state that amino acids with similar properties will have similar effects after substitution. For example, conservative amino acid substitutions may occur in the homologous proteins mentioned above. "Conservative amino acid substitutions" include but are not limited to:

[0060] Hydrophobic amino acids (Ala, Cys, Gly, Pro, Met, Val, Ile, Leu) are replaced by other hydrophobic amino acids;

[0061] Substitution of bulky hydrophobic amino acids (Phe, Tyr, Trp) with other bulky hydrophobic amino acids;

[0062] Amino acids with positively charged side chains (Arg, His, Lys) are replaced by other amino acids with positively charged side chains;

[0063] Amino acids with polar and uncharged side chains (Ser, Thr, Asn, Gln) are replaced by other amino acids with polar and uncharged side chains.

[0064] Those skilled in the art may also perform conservative substitutions on amino acids according to amino acid substitution rules well known to those skilled in the art, such as the "blosum62 scoring matrix" in the prior art.

[0065] In order to further strengthen the interaction between the 2A domain and the Ratchet domain, allow the nucleic acid to pass through the pore stably, and control the speed of the nucleic acid passing through the nanopore, in a preferred embodiment, at least one site in i) of 2) is independently selected from at least one of the following: sites E279, G280, I281, S282, L283, V284, D285, G286, E287, P288, S289, or S290 in the 2A domain of the amino acid sequence of SEQ ID NO: 1; sites S275, A276, D277, E278, G279, N280, I281, A282, L283, V284, D285, G286, E287, P288, S289, S290, or I291 in the 2A domain of the amino acid sequence of SEQ ID NO: 2; The following are the sites in the 2A domain of the amino acid sequence shown in SEQ ID NO:3: K269, G270, I271, D272, E273, D274, L275, G276, D277, D278 or K279; the following are the sites in the 2A domain of the amino acid sequence shown in SEQ ID NO:4: A279, P280, G281, E282, D283, V284, M285, M286, G287, E288, V289, D290, V291, E292, E293, A294, S295 or S296.

[0066] In order to further strengthen the interaction between the 2A domain and the Ratchet domain, allow the nucleic acid to pass through the pore stably, and control the speed of the nucleic acid passing through the nanopore, in a preferred embodiment, the sites in the Ratchet domain of the amino acid sequence as shown in SEQ ID NO: 1 are: E586, I587, P588, E589, K590, E591, M592, E593, K594, L595, Y596 or G597; the sites in the Ratchet domain of the amino acid sequence as shown in SEQ ID NO: 2 are: T590, P591, E592, K593, E594, M595, E596, K597, R598, Y599, G600, V601, G602 or P603; The sites in the Ratchet domain of the amino acid sequence shown in NO:3 are: E568, E569, V570, P571, E572, K573, R574, I575, T576, D577, R578 or Y579; the sites in the Ratchet domain of the amino acid sequence shown in SEQ ID NO:4 are: E599, E600, K601, D602, D603, D604, T605, L606, A607, E608, S609 or F610.

[0067] In order to prevent intramolecular or intermolecular cross-linking of cysteines in the 2A domain, WH domain and Ratchet domain, in a preferred embodiment, at least one site in ii) of 2) is independently selected from at least one of the following: site C298 or C329 in the 2A domain of the amino acid sequence shown in SEQ ID NO: 1; site C298 or C329 in the 2A domain of the amino acid sequence shown in SEQ ID NO: 2; site C304, C335 or C379 in the 2A domain of the amino acid sequence shown in SEQ ID NO: 4; preferably, site C441 or C458 in the WH domain of the amino acid sequence shown in SEQ ID NO: 4; preferably, site C529 in the Ratchet domain of the amino acid sequence shown in SEQ ID NO: 1; site C532 in the Ratchet domain of the amino acid sequence shown in SEQ ID NO: 2.

[0068] Any unnatural amino acid that can be linked to each other or to cysteine ​​is suitable for the present application. In a preferred embodiment, the unnatural amino acid includes: 4-azido-L-phenylalanine (PAZF), 4-azido-L-phenylalanine (PAZF-Hcl), 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselenoyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4-dihydroxyboryl)-L-phenylalanine, alanine, 4-[(ethylsulfanyl)carbonyl]-L-phenylalanine, (2S)-2-amino-3-{4-[(propan-2-ylsulfanyl)carbonyl]phenyl}propanoic acid, (2S)-2-amino-3-{4-[(2-amino-3-sulfanylpropionyl)amino]phenyl}propanoic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-Nitro L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino-L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphth-2-ylamino)propionic acid, 6-(methylsulfanyl)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propionic acid, (2R)-2-aminooctanoate 3-(2,2′-bipyridin-5-yl)-D-alanine , 2-amino-3-(8-hydroxy-3-quinolyl)propionic acid, 4-benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)sulfanyl]propionic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propionic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)oxy]carbonyl}amino)hexanoic acid, O-(2-nitrobenzyl)-L-tyrosine or 2-nitrophenylalanine.

[0069] The binding between the 2A domain and the Ratchet domain is further strengthened by the connection between the mutated amino acids. In a preferred embodiment, a connection exists between the site in the 2A domain of the amino acid sequence in 2) i) substituted with cysteine ​​or a non-natural amino acid and the site in the Ratchet domain substituted with cysteine ​​or a non-natural amino acid. Due to the different groups carried by the mutated amino acids, in a preferred embodiment, the connection includes non-covalent bonding, covalent bonding, or a combination of covalent and non-covalent bonding.

[0070] Covalent linkage includes direct linkage between amino acid groups, or indirect linkage through a cross-linking agent, a polypeptide molecule or a synthetic small molecule compound. In a preferred embodiment, covalent linkage is performed using a cross-linking agent, and the cross-linking agent includes: maleimide, succinimide, active ester, azide, alkyne (dibenzocyclooctyne (DIBO or DBCO), difluorocycloalkyne and linear alkyne), phosphine, haloacetyl (such as iodoacetamide), phosgene-type reagents, sulfonyl chloride reagents, isothiocyanate, acyl halide, hydrazine, disulfide, vinyl sulfone, aziridine or photosensitive reagents (such as aryl azide and diaziridine).

[0071] In a second typical embodiment of the present invention, a DNA molecule is provided, which encodes any one of the above-mentioned Ski2-like helicases.

[0072] In a preferred embodiment, the above-mentioned DNA molecule is selected from: A) a polynucleotide comprising a nucleotide sequence as shown in any one of SEQ ID NOs: 5 to 8; or B) a polynucleotide comprising a nucleotide sequence as shown in any one of SEQ ID NOs: 5 to 8 having a homology of more than 80%, more preferably more than 90%, and even more preferably more than 95% to a polynucleotide.

[0073] SEQ ID NO: 5: Nucleic acid sequence of Ski2-like helicase 1

[0074] SEQ ID NO:6: Nucleic acid sequence of Ski2-like helicase 2

[0075] SEQ ID NO: 7: Nucleic acid sequence of Ski2-like helicase 3

[0076] SEQ ID NO:8: Nucleic acid sequence of Ski2-like helicase 4

[0077] Due to the principle of codon degeneracy, the nucleotide sequence obtained by translation of the above amino acid sequence is not limited to the nucleotide sequence shown in the above SEQ ID NO: 5-8. Any nucleotide sequence that can encode the above Ski2-like helicase falls within the protection scope of the DNA molecule of this application.

[0078] In a third exemplary embodiment of the present invention, a recombinant vector is provided, comprising the aforementioned DNA molecule. In a preferred embodiment, the recombinant plasmid comprises a vector with a T7 promoter, such as pET-28a(+), pET-21a(+), or pET-32a(+).

[0079] The DNA encodes the Ski2-like helicase and can be linked to a recombinant plasmid to form circular DNA. Both the DNA and the recombinant plasmid can be transcribed and translated under the action of RNA polymerase, ribosomes, tRNA, etc., to produce the helicase capable of unwinding nucleic acids from the 3' end to the 5' end.

[0080] In a fourth typical embodiment of the present invention, a host cell is provided, wherein the host cell contains the above-mentioned recombinant vector. In a preferred embodiment, the host cell includes a prokaryotic cell or a eukaryotic cell. Specifically, the prokaryotic cell can be Escherichia coli, and the eukaryotic cell can be yeast. In order to improve the expression accuracy of the helicase and obtain a helicase with high protein homogeneity and purity, in a preferred embodiment, the Escherichia coli includes BL21 (DE3), BL21 Star (DE3) pLyss, Rossata (DE3) or Lemo21 (DE3).

[0081] The host cells described above can replicate recombinant plasmids and transcribe and translate the DNA molecules carried by the recombinant plasmids, thereby obtaining a large amount of Ski2-like helicase. Ski2-like helicase can be obtained by fragmenting the host cells for protein purification, using crude enzyme catalysis after fragmentation, or other methods using existing technologies. The host cells are non-plant-derived.

[0082] In a fifth exemplary embodiment of the present invention, a kit is provided, comprising the aforementioned Ski2-like helicase. This kit can be used to efficiently perform nanopore sequencing on samples requiring 3'-5' end sequencing, such as samples with 5'-end modifications, and has promising application prospects.

[0083] To further ensure the sequencing efficiency of Ski2-like helicase-involved sequencing, in a preferred embodiment, the kit also includes a sequencing chip with a nanopore protein and a sequencing buffer; preferably, the sequencing buffer includes 0.1-1.0M KCl, 10-100mM HEPES, 1-50mM ATP, and 10-100mM MgCl2; preferably, the pH of the sequencing buffer is 6-10. HEPES (4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid) stabilizes the ion concentration and pH of the reaction system, and ATP provides energy for nucleic acid unwinding. Specifically, the types of the above-mentioned nanopore proteins include but are not limited to: α-hemolysin, Aeromonas hydrophila toxin (Aerolysin), Mycobacterium smegmatis porin A (MspA), curli-specific transport channel (Curli production assembly / transport component CsgG), and Fragaceatoxin C (FraC).

[0084] To further improve the capture rate of sequencing libraries, in a preferred embodiment, the kit also includes sequencing adapters and single-stranded DNA with a cholesterol-modified 5' end. The single-stranded DNA and sequencing adapter have a complementary region, preferably 10-50 nt in length. The cholesterol-modified single-stranded DNA can bind to the phospholipid membrane on the sequencing chip, thereby immobilizing the nucleic acids to be sequenced onto the chip, reducing the library sample load and improving the capture rate of nucleic acids in the library.

[0085] In a sixth exemplary embodiment of the present invention, a sequencing method is provided, comprising: drawing a nucleic acid fragment to be sequenced from a sequencing library to a nanopore of a nanopore chip; using the aforementioned Ski2-like helicase or the Ski2-like helicase in the aforementioned kit to control the passage of the nucleic acid fragment through the nanopore, and performing sequencing from the 3' end to the 5' end; the sequencing comprises: analyzing the electrical signal generated when the nucleic acid fragment passes through the nanopore, thereby determining the sequence of the nucleic acid fragment. Using the Ski2-like helicase or a kit containing the Ski2-like helicase enables efficient and convenient nanopore sequencing, and has promising application prospects.

[0086] In a preferred embodiment, the method further comprises: preparing the nucleic acid sample into nucleic acid fragments with sequencing adapters; and performing PCR amplification on the nucleic acid fragments with sequencing adapters to obtain a sequencing library.

[0087] To further efficiently complete the sequencing reaction, in a preferred embodiment, the sequencing temperature is 20-40°C; preferably, the molar concentration ratio of the Ski2-like helicase to the nucleic acid fragment to be tested is 100:1-1:1. The Ski2-like helicase of the present application is capable of performing nanopore sequencing reactions at the aforementioned temperature and concentration.

[0088] In a seventh exemplary embodiment of the present invention, a sequencing library complex is provided, comprising: a sequencing library to be tested and a Ski2-like helicase; wherein the Ski2-like helicase is the aforementioned Ski2-like helicase or the Ski2-like helicase in the aforementioned kit; the sequencing library to be tested comprises a nucleic acid fragment to be tested with a sequencing adapter, and the Ski2-like helicase is bound to the sequencing adapter. Specifically, the Ski2-like helicase is bound to the 3' end of the sequencing adapter.

[0089] In an eighth exemplary embodiment of the present invention, a method for preparing a sequencing library complex is provided, the method comprising: preparing a nucleic acid sample to be tested into a nucleic acid fragment to be tested with a sequencing adapter, incubating the nucleic acid fragment to be tested with a sequencing adapter with a Ski2-like helicase to obtain a sequencing library complex, wherein the Ski2-like helicase in the sequencing library complex binds to the sequencing adapter; or incubating the sequencing adapter with a Ski2-like helicase to form an adapter complex, preparing the nucleic acid sample to be tested into a nucleic acid fragment to be tested with the adapter complex, thereby obtaining a sequencing library complex, wherein the Ski2-like helicase in the sequencing library complex binds to the sequencing adapter; wherein the Ski2-like helicase is the aforementioned Ski2-like helicase or the Ski2-like helicase in the aforementioned kit. Specifically, the Ski2-like helicase binds to the 3' end of the sequencing adapter.

[0090] The present application is further described in detail below with reference to specific embodiments. These embodiments should not be construed as limiting the scope of protection claimed in this application.

[0091] Example 1: Cloning, expression and purification of Ski2-like helicase

[0092] 1. Cloning and expression of Ski2-like helicase

[0093] The Ski2-like helicase sequence was connected to the PET.28a(+) plasmid using the double enzyme cleavage sites Nde1 and Xho1. The Ski2-like helicase expressed by the plasmid had a 6×His tag and a thrombin cleavage site at the N-terminus.

[0094] Among them, Ski2-like helicase 1 is a mutant having mutation sites of V284C, P588C, C298A and C529A based on SEQ ID NO: 1; Ski2-like helicase 2 is a mutant having mutation sites of N280C, P591C, C298A and C532A based on SEQ ID NO: 2; Ski2-like helicase 3 is a mutant having mutation sites of D274C and P571C based on SEQ ID NO: 3; Ski2-like helicase 4 is a mutant having mutation sites of M286C, D602C, C304A, C379A, C441A and C458A based on SEQ ID NO: 4.

[0095] Transform the recombinant plasmid containing the Ski2-like helicase into Escherichia coli BL21(DE3) or its derivatives. Pick a single colony and inoculate it into 20 mL of LB medium containing kanamycin resistance. Cultivate with shaking at 37°C overnight. Then, inoculate the colony into 2 L of LB medium containing kanamycin resistance and incubate with shaking at 37°C until the OD600 reaches 0.6-0.8. Cool the culture to 16°C, add IPTG at a final concentration of 500 μM, induce expression overnight, and harvest the cells.

[0096] 2. Purification of Ski2-like helicase

[0097] The buffer solution used is as follows:

[0098] 1) Buffer A: 20mM Tris-HCl pH 8.0, 500mM NaCl, 20mM imidazole

[0099] 2) Buffer B: 20 ​​mM Tris-HCl pH 8.0, 500 mM NaCl, 300 mM imidazole

[0100] 3) Buffer C: 20mM Tris-HCl pH 8.0

[0101] 4) Buffer D: 20mM Tris-HCl pH 8.0, 100mM NaCl

[0102] 5) Buffer E: 20mM Tris-HCl pH 8.0, 1000mM NaCl

[0103] 6) Buffer F: 20mM Tris-HCl pH 8.0, 200mM NaCl

[0104] After harvesting cells expressing the Ski2-like helicase using a centrifuge at 5000 rpm at 4°C, resuspend the cells in Buffer A, disrupt them using a cell disrupter, and collect the supernatant after high-speed centrifugation. Mix the supernatant with Ni-NTA medium equilibrated in Buffer A and incubate at 4°C for 1 hour. Rinse the medium with 5-10 column volumes of Buffer A until all contaminants are eluted. Next, add Buffer B to the medium and elute the target protein in batches. The eluted target protein is diluted with Buffer C to a salt concentration of 50 mM. Swell the ssDNA cellulose medium with TE buffer for 3-4 hours and equilibrate with Buffer D. Mix the diluted protein with the ssDNA cellulose medium, add an appropriate amount of protease, and incubate overnight at 4°C. After centrifugation, remove the supernatant and flow-through. Remove nonspecifically bound contaminants with Buffer D three times, and then elute the target protein with Buffer E. The eluted target protein was then concentrated to 1 ml and further purified using Superdex 200 increase 10 / 300GL (Cytiva) molecular sieves, using Buffer F. The target protein's elution peak, identified by the molecular sieve elution curve and SDS-PAGE analysis, was collected and concentrated (as shown in Figures 6-9). The resulting target protein was frozen at -80°C.

[0105] Figures 6 and 7 show the Superdex 200 increase 10 / 300GL purification results (A) and SDS-PAGE gel electrophoresis (B), respectively, of Ski2-like helicase 1 and Ski2-like helicase 2. The different channels in the SDS-PAGE gel electrophoresis (B) represent samples at different elution volumes. The resulting Ski2-like helicase 1 and Ski2-like helicase 2 variants demonstrate high protein purity, with uniform protein peaks.

[0106] Figure 8 shows the purification results of Ski2-like helicase 3 using Superdex 200 increase 10 / 300GL (A) and an SDS-PAGE gel electrophoresis image (B). The different channels in the SDS-PAGE gel electrophoresis image (B) represent samples at different elution volumes. The resulting Ski2-like helicase 3 variant exhibits a broad protein peak but exhibits high protein purity.

[0107] Figure 9 shows the Superdex 200 increase 10 / 300 molecular sieve purification results (A) and SDS-PAGE gel electrophoresis (B). The different channels in the SDS-PAGE gel electrophoresis (B) represent samples at different elution volumes. The molecular sieve elution curve of the Ski2-like helicase 4 variant protein shows multiple absorption peaks, but the primary protein peak is relatively clear, indicating average protein purity with a few minor bands.

[0108] Example 2: DNA binding ability test of Ski2-like helicase 1 and Ski2-like helicase 2

[0109] Purified Ski2-like helicase 1 and Ski2-like helicase 2 were mixed with 0.5 μM ssDNA (SEQ ID NO: 9) in Buffer D (20 mM Tris-HCl pH 8.0, 100 mM NaCl) at different ratios, with the protein concentrations in each tube being 0.25 μM, 0.5 μM, 1 μM, 2.5 μM, and 5 μM, respectively, for a total volume of 20 μl. After the components were prepared in a centrifuge tube, they were incubated at 30°C for 30 minutes and then immediately run on an 8% Native PAGE gel for electrophoresis. The substrate ssDNA sequence used in the experiment (SEQ ID NO: 9) was: 5'-CAAGCTCTCAACTGCAGTCTAGACTCGAGC-3'.

[0110] Figures 10 and 11 show the binding of Ski2-like helicase 1 and Ski2-like helicase 2 to ssDNA at different mixing ratios. As can be seen, as the mixing ratio of Ski2-like helicase 1 and Ski2-like helicase 2 to DNA increases, the binding of the proteins to ssDNA becomes more complete. As shown in Figure 10, when the concentration ratio of Ski2-like helicase 1 to ssDNA is greater than or equal to 5:1, all ssDNA is completely bound. As shown in Figure 11, when the concentration ratio of Ski2-like helicase 2 to ssDNA is greater than or equal to 2:1, all ssDNA is bound to the proteins. These results indicate that both Ski2-like helicase 1 and Ski2-like helicase 2 have strong DNA binding abilities.

[0111] Example 3: dsDNA melting activity detection

[0112] 1. Preparation of double-stranded DNA (overhang DNA, ovDNA)

[0113] SEQ ID NO:10: 5'-CCAACAAAGACAACCACGACTATAACG-BHQ-1-3' and SEQ ID NO:11: 5'-FAM-CGTTATAGTCGTGGTTGTCTTTGTTGGTTTTTTTTTTTTTTTTTTTT-3' were synthesized, and SEQ ID NO:10 and SEQ ID NO:11 were annealed to ovDNA with 20 Ts overhanging the 3' end. The annealing process was incubation at 95°C for 5 minutes, cooling at a rate of 0.1°C / s to 25°C, and incubation for 30 minutes. The annealing formula is shown in Table 1.

[0114] Table 1 ovDNA annealing formula

[0115] 2. Prepare reaction buffer

[0116] Reaction buffer: 100 mM HEPES (pH=8.0), 1 mg / mL BSA, 10 mM MgCl2, 500 mM KCl.

[0117] 3. Prepare the reaction solution

[0118] Experimental reaction solution: 3 μL of 10 μM ovDNA, 6 μL of 100 μM SEQ ID NO:12: 5'-CGTTATAGTCGTGGTTGTCTTTGTTGG-3' (20x competitor DNA), and 6 μL of 100 mM ATP were added to 585 μL of reaction buffer. The competitor DNA binds to SEQ ID NO:10 modified with the quencher group BHQ-1, allowing the free FAM-modified SEQ ID NO:11 to recover fluorescence after unwinding and be detected.

[0119] Positive control solution: 1 μL 10 μM SEQ ID NO: 11, 2 μL 100 μM SEQ ID NO: 12 (20-fold competitor DNA), and 2 μL 100 mM ATP were added to 195 μL reaction buffer.

[0120] 4. Dilute protein

[0121] The protein was diluted to 4.8 μM with 1× PBS.

[0122] 5. Prepare the Melting Reaction

[0123] The corresponding reagents were added according to Table 2. ① was the experimental group, ② was the negative control group, and ③ was the positive control group. The kinetic changes of fluorescence intensity within 30 min of the reaction were detected using a microplate reader at 30°C. Each group was repeated three times.

[0124] Table 2 Melting reaction formula

[0125] 6. Data Analysis

[0126] The fluorescence value of the negative control group was first subtracted from the fluorescence value of the experimental group and the positive control group, and then the percentage of the fluorescence value of the experimental group relative to the positive control group was calculated.

[0127] 7. Experimental Results

[0128] As can be seen from Figures 12, 13, 14, and 15, the fluorescence value of the experimental group gradually increased with the increase in reaction time, indicating that the proteins of Ski2-like helicase 1, Ski2-like helicase 2, Ski2-like helicase 3, and Ski2-like helicase 4 all have the activity of unwinding double-stranded DNA. Furthermore, since ovDNA contains single-stranded nucleic acid in the 3' direction, the fluorescence signal can be detected using the Ski2-like helicase of the present application, indicating that the Ski2-like helicase of the present application has unwinding activity in the 3'-5' direction.

[0129] Example 4: Nanopore sequencing

[0130] 1. Two partially complementary DNA strands (the top strand consists of 5'-SEQ ID NO: 13-(iSP18)8-SEQ ID NO: 14-(iSPC3) 30 -3', wherein SEQ ID NO: 13: CGTTATAGTCGTGGTTGTCTTTGTTGG, SEQ ID NO: 14: TTTTTTTTTTTTTTTTTTTT, iSP18 is a spacer modified linker composed of 6 consecutive ethylene glycols, and iSPC3 is a spacer modified linker composed of 3 CH2; the bottom strand, SEQ ID NO: 15: 5'-AACTGGCGAGCGGAGTTTCCAACAAAGACAACCACGACTATAACGT-3') is annealed to form a linker, which is then ligated with the double-stranded target fragment to be tested using T4 DNA ligase and purified to obtain a sequencing library.

[0131] 2. Incubate the Ski2-like helicase and sequencing library (i.e., the nucleic acid fragment to be tested) at 25°C for 1 hour (molar concentration ratio 8:1) to form a sequencing library containing helicase.

[0132] 3. The helicase-containing sequencing library was incubated with single-stranded DNA containing cholesterol at its 5' end (ssDNA-chol, SEQ ID NO: 16: 5'-cholesterol-TTGACCGCTCGCCTC-3') at room temperature for 10 minutes. The ssDNA-chol sequence is complementary to a portion of the bottom strand of the adaptor. Cholesterol binding to the phospholipid membrane reduces library loading and improves capture efficiency.

[0133] 4. Use a patch clamp amplifier or other electrical signal amplifier to collect the current signal. A Teflon membrane with a micrometer-sized pore (50-200μm in diameter) in the center divides the electrolytic cell into two chambers: the cis chamber and the trans chamber. A pair of Ag / AgCl electrodes is placed in each chamber. A bimolecular phospholipid membrane is formed at the micropores of the two chambers, and the nanopore protein is added. Electrical measurements are obtained after a single nanopore protein is inserted into the phospholipid membrane. The reaction product from step 3 is added, and 180mV is applied. The sequencing library is captured by the nanopore, and the nucleic acid passes through the nanopore under the control of the helicase. The buffer used in this experiment is: 0.3M KCl, 25mM HEPES, 30mM ATP, 25mM MgCl2, pH 8, and the sequencing temperature is 30°C.

[0134] 5. The sequencing electrical signals of Ski2-like helicases 1-4 are shown in Figures 16, 17, 18, and 19, respectively. As can be seen, as the helicase guides the DNA single strand into the nanopore, some of the current is blocked, causing the current to decrease. Because different nucleotides have different sizes, the magnitude of the blocked current also varies, resulting in a fluctuating current signal. This example demonstrates that all Ski2-like helicases 1-4 can be used for nanopore sequencing.

[0135] From the above description, it can be seen that the above-mentioned embodiments of the present invention achieve the following technical effects: the interaction between the binding domains of the helicase and its mutants of the present application can form a stable pore, which can allow nucleic acids to pass through the pore stably during nanopore sequencing and control the speed of nucleic acids passing through the nanopore. In addition, the helicase of the present application has the function of unwinding the nucleic acid from the 3' end to the 5' end. The use of the helicase of the present application and its related mutants can broaden the application scenarios of nanopore sequencing.

[0136] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A Ski2-like helicase, characterized in that, The Ski2-like helicase includes: 1) A protein having an amino acid sequence shown in any one of SEQ ID NOs: 1 to 4, and including a 1A domain, a 2A domain, a WH domain, a Ratchet domain, and an HLH domain; or 2) A protein having at least one of the following mutations in the amino acid sequence in 1) and having helicase activity for unwinding in the direction from the 3'-end to the 5'-end of nucleic acid: i) At least one site is substituted with cysteine or a non-natural amino acid, and the at least one site is selected from at least one of the following domains: the 2A domain or the Ratchet domain; ii) At least one cysteine site is deleted or substituted with a natural amino acid, and the at least one cysteine site is selected from at least one of the following domains: the 2A domain, the WH domain, or the Ratchet domain; or 3) A protein having an amino acid sequence with a homology of more than 50%, preferably more than 80%, more preferably more than 90%, and further preferably more than 95% to the amino acid sequence defined in any one of 1) and 2), and having helicase activity for unwinding in the direction from the 3'-end to the 5'-end of nucleic acid.

2. The Ski2-like helicase according to claim 1, characterized in that, In the i) of 2), the at least one site is independently selected from at least one of the following: Sites in the 2A domain of the amino acid sequence shown in SEQ ID NO: 1: E279, G280, I281, S282, L283, V284, D285, G286, E287, P288, S289, or S290; Sites in the 2A domain of the amino acid sequence shown in SEQ ID NO: 2: S275, A276, D277, E278, G279, N280, I281, A282, L283, V284, D285, G286, E287, P288, S289, S290, or I291; Sites in the 2A domain of the amino acid sequence shown in SEQ ID NO: 3: K269, G270, I271, D272, E273, D274, L275, G276, D277, D278, or K279; Sites in the 2A domain of the amino acid sequence shown in SEQ ID NO: 4: A279, P280, G281, E282, D283, V284, M285, M286, G287, E288, V289, D290, V291, E292, E293, A294, S295, or S296; Preferably, sites in the Ratchet domain of the amino acid sequence shown in SEQ ID NO: 1: E586, I587, P588, E589, K590, E591, M592, E593, K594, L595, Y596, or G597; Sites in the Ratchet domain of the amino acid sequence as shown in SEQ ID NO: 2: T590, P591, E592, K593, E594, M595, E596, K597, R598, Y599, G600, V601, G602 or P603; Sites in the Ratchet domain of the amino acid sequence as shown in SEQ ID NO: 3: E568, E569, V570, P571, E572, K573, R574, I575, T576, D577, R578 or Y579; Sites in the Ratchet domain of the amino acid sequence as shown in SEQ ID NO: 4: E599, E600, K601, D602, D603, D604, T605, L606, A607, E608, S609 or F610.

3. The Ski2-like helicase according to claim 1, characterized in that, Each of the at least one cysteine site in ii) of 2) is independently selected from at least one of the following: Sites in the 2A domain of the amino acid sequence as shown in SEQ ID NO: 1: C298 or C329; Sites in the 2A domain of the amino acid sequence as shown in SEQ ID NO: 2: C298 or C329; Sites in the 2A domain of the amino acid sequence as shown in SEQ ID NO: 4: C304, C335 or C379; Preferably, sites in the WH domain of the amino acid sequence as shown in SEQ ID NO: 4: C441 or C458; Preferably, sites in the Ratchet domain of the amino acid sequence as shown in SEQ ID NO: 1: C529; Sites in the Ratchet domain of the amino acid sequence as shown in SEQ ID NO: 2: C532.

4. The Ski2-like helicase according to claim 1, wherein The unnatural amino acids include: 4-azido-L-phenylalanine, 4-azido-L-phenylalanine, 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetylacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselanyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4(dihydroxyboryl)-L-phenylalanine, 4-[(ethylthio)carbonyl]-L-phenylalanine, (2S)-2-amino-3-{4-[(propan-2-ylthio)carbonyl]phenyl}propanoic acid, (2S)-2-amino-3-{4-[(2-amino-3-sulfanylpropanoyl)amino]phenyl}propanoic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino-L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthalen-2-ylamino)propanoic acid, 6-(methylthio)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propanoic acid, (2R)-2-aminooctanoate 3-(2,2′-dipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxy-3-quinolyl)propanoic acid, 4-benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)thio]propanoic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propanoic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)oxy]carbonyl}amino)hexanoic acid, O-(2-nitrobenzyl)-L-tyrosine or 2-nitrophenylalanine.

5. The Ski2-like helicase according to claim 1, wherein There is a connection relationship between the site substituted with cysteine or an unnatural amino acid in the 2A domain of the amino acid sequence in i) of 2) and the site substituted with cysteine or an unnatural amino acid in the Ratchet domain. Preferably, the connection includes non-covalent connection, covalent connection or a combination of covalent and non-covalent binding connection; Preferably, the covalent connection is carried out using a cross-linking agent, and the cross-linking agent includes: maleimide, succinimide, active ester, azide, alkyne, phosphine, haloacetyl, phosgene-type reagent, sulfonyl chloride reagent, isothiocyanate, acyl halide, hydrazine, disulfide, vinyl sulfone, aziridine or photosensitive reagent.

6. A DNA molecule, characterized in that, The DNA molecule encodes the Ski2-like helicase according to any one of claims 1 to 5.

7. The DNA molecule according to claim 6, wherein The DNA molecule is selected from: A) A polynucleotide comprising a nucleotide sequence as shown in any one of SEQ ID NOs: 5 to 8; or B) A polynucleotide comprising a polynucleotide having a homology of more than 80%, more preferably more than 90%, and further preferably more than 95% with any one of the nucleotide sequences shown in SEQ ID NOs: 5 to 8.

8. A recombinant vector, characterized in that, The recombinant vector contains the DNA molecule according to claim 6 or 7.

9. A host cell, characterized in that, The host cell contains the recombinant vector according to claim 8.

10. The host cell according to claim 9, characterized in that, The host cell includes a prokaryotic cell or a eukaryotic cell.

11. A kit, comprising a Ski2-like helicase, characterized in that, The Ski2-like helicase is the Ski2-like helicase according to any one of claims 1 to 5.

12. The kit according to claim 11, wherein The kit further includes a sequencing chip with nanopore protein and a sequencing buffer; Preferably, the sequencing buffer includes: 0.1 - 1.0 M KCl, 10 - 100 mM HEPES, 1 - 50 mM ATP, and 10 - 100 mM MgCl2; Preferably, the pH value of the sequencing buffer is 6 - 10.

13. The kit according to claim 11, wherein The kit further includes a sequencing adapter and a single-stranded DNA modified with cholesterol at the 5' end; Wherein, there is a complementary region between the single-stranded DNA and the sequencing adapter; Preferably, the length of the complementary region is 10 - 50 nt.

14. A sequencing method, characterized in that, The method includes: Using the Ski2-like helicase according to any one of claims 1 to 5 or the Ski2-like helicase in the kit according to any one of claims 11 to 13 to control the nucleic acid fragment to be tested to pass through the nanopore for sequencing from the 3' end to the 5' end; The sequencing includes: analyzing the electrical signal generated when the nucleic acid fragment passes through the nanopore channel to determine the sequence of the nucleic acid fragment.

15. The method according to claim 14, wherein The method further includes: Preparing the nucleic acid sample to be tested into a nucleic acid fragment to be tested with a sequencing adapter, and binding the nucleic acid fragment to be tested with the sequencing adapter to the Ski2-like helicase; or, Binding the sequencing adapter to the Ski2-like helicase to form an adapter complex, Preparing the nucleic acid sample to be tested into a nucleic acid fragment to be tested with the adapter complex.

16. The method according to claim 14 or 15, characterized in that, Performing the nanopore sequencing at 20 - 40 °C; Preferably, the molar concentration ratio of the Ski2-like helicase to the nucleic acid fragment to be tested is 100:1 - 1:

1.

17. A sequencing library complex, characterized in that, The sequencing library complex includes: a sequencing library to be tested and a Ski2-like helicase; Wherein, the Ski2-like helicase is the Ski2-like helicase according to any one of claims 1 to 5 or the Ski2-like helicase in the kit according to any one of claims 11 to 13; The sequencing library to be tested includes a nucleic acid fragment to be tested with a sequencing adapter, and the Ski2-like helicase binds to the sequencing adapter.

18. A method for preparing a sequencing library complex, characterized in that, The preparation method includes: Preparing the nucleic acid sample to be tested into a nucleic acid fragment to be tested with a sequencing adapter, Incubate the nucleic acid fragment to be sequenced with the sequencing adapter with Ski2-like helicase to obtain the sequencing library complex, wherein the Ski2-like helicase of the sequencing library complex binds to the sequencing adapter; or, Incubate the sequencing adapter with the Ski2-like helicase to form an adapter complex, Prepare the nucleic acid sample to be sequenced into a nucleic acid fragment to be sequenced with the adapter complex to obtain the sequencing library complex, wherein the Ski2-like helicase of the sequencing library complex binds to the sequencing adapter; Wherein, the Ski2-like helicase is the Ski2-like helicase according to any one of claims 1 to 5 or the Ski2-like helicase in the kit according to any one of claims 11 to 13.

Citation Information

Patent Citations

  • Modified CfM HL4 helicase and application thereof

    CN116334030A

  • Bio-engineered hyper-functional "super" helicases

    US20170335297A1

  • Modified helicases

    US20230227799A1

  • Modified prp43 helicase and use thereof

    WO2022213253A1