NS3 helicase mutants for characterizing biomolecules and uses thereof

By developing the NS3 helicase mutant as a motor protein, the problem of unwell binding force of the biomolecule to be tested and the motor protein in nanopore sequencing is solved, and the stable current signal and efficient movement of the biomolecule to be tested is achieved, which improves the reliability of sequencing.

CN120230736APending Publication Date: 2025-07-01BEIJING QITAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411524782.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-10-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In nanopore sequencing technology, the binding force of the biomolecules to be tested and the motor protein is too weak or too strong, resulting in unstable current signal and it is difficult to achieve continuous and stable current signal.

Method used

A new motor protein, namely the NS3 helicase mutant, is developed. Through specific mutations in the amino acid sequence, it has moderate binding force with the biomolecule to be tested, thereby achieving stable binding of one or two NS3 helicase mutants, controlling the movement of the biomolecule to be tested, and obtaining a stable current signal.

Benefits of technology

Through the use of NS3 helicase mutants, the stable movement and continuous and stable current signals of the biomolecules to be tested are achieved, and the analysis reliability of nanopore sequencing is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120230736A_ABST
    Figure CN120230736A_ABST
Patent Text Reader

Abstract

The invention relates to a nanopore sequencing technology, in particular to NS3 helicase mutants for characterizing biomolecules and application of the NS3 helicase mutants. According to the present invention, the NS3 helicase is modified so as to provide the moderate binding force with the biomolecule to be detected, such that the one or two NS3 helicase mutants are stably bound to the biomolecule to be detected, such that the movement of the biomolecule to be detected can be continuously and stably controlled so as to obtain the stable current signal;
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to nanopore sequencing technology, and in particular to NS3 helicase mutants and uses thereof for characterizing biomolecules. Background Art

[0002] Nanopore sequencing technology is an emerging third-generation sequencing technology. Under the condition of applying a certain voltage, the biological molecules to be tested (such as nucleotides, polypeptides, polysaccharides, lipids) pass through the nanopore one by one and make a short stop under the action of motor proteins (also called molecular motors or rate-controlling enzymes), generating current signals corresponding to different bases, amino acids, etc., which are recorded and identified and then analyzed and restored to the sequence information of the biological molecules to be tested. Nanopore sequencing technology is different from previous sequencing technologies in that it has the following characteristics: ultra-long read length, no need for amplification, short time required, low cost, etc. It has broad prospects in the fields of genome assembly, structural variation detection, epigenetic modification detection, pathogenic or environmental microbial detection, etc.

[0003] In the process of realizing the present invention, the inventors found that there are at least the following problems in the prior art: One of the challenges faced by nanopore sequencing technology is that the binding force between the biomolecule to be detected and the motor protein is too weak or too strong, which is not conducive to obtaining a stable current signal. If the binding force is too weak, the motor protein is easy to separate from the biomolecule to be detected, resulting in an interruption of the current signal. If the binding force is too strong, multiple motor proteins are stably bound to the biomolecule to be detected, resulting in the blockage of the nanopore and the interruption of the current signal. Therefore, developing a motor protein with moderate binding force to the biomolecule to be detected is a technical problem that needs to be solved in this field. Summary of the invention

[0004] In order to solve the above technical problems, the inventors, after long-term exploration and continuous attempts, provide a new motor protein, which is an NS3 helicase mutant, and has a moderate binding force with the biological molecule to be tested, so that one or two NS3 helicase mutants are stably bound to the biological molecule to be tested, and there is no situation where the NS3 helicase mutant cannot stably bind to the biological molecule to be tested or multiple NS3 helicase mutants are bound to the biological molecule to be tested, and it can continuously and stably control the movement of the biological molecule to be tested, and obtain a continuous, stable and more analytically friendly current signal.

[0005] In a first aspect of the present disclosure, an NS3 helicase mutant can be provided, the NS3 helicase mutant comprising an amino acid sequence as shown in SEQ ID NO: 1, or an amino acid sequence having at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity with the amino acid sequence as shown in SEQ ID NO: 1.

[0006] In a second aspect of the present disclosure, a construct can be provided, which comprises the NS3 helicase mutant described in the first aspect of the present disclosure.

[0007] In a third aspect of the present disclosure, a nucleic acid can be provided, which encodes the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure.

[0008] In a fourth aspect of the present disclosure, an expression vector can be provided, which comprises the nucleic acid described in the third aspect of the present disclosure.

[0009] In a fifth aspect of the present disclosure, a host cell can be provided, which comprises the nucleic acid described in the third aspect of the present disclosure or comprises the expression vector described in the fourth aspect of the present disclosure.

[0010] In a sixth aspect of the present disclosure, a method for preparing the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure can be provided, comprising: culturing the host cell described in the fifth aspect of the present disclosure, inducing expression, and then purifying the obtained expression product.

[0011] In a seventh aspect of the present disclosure, a method for controlling the movement of a biomolecule to be measured can be provided, the method comprising: contacting the biomolecule to be measured with the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure, thereby controlling the movement of the biomolecule to be measured.

[0012] In an eighth aspect of the present disclosure, a method for characterizing a biomolecule to be measured can be provided, the method comprising:

[0013] (a) contacting the biomolecule to be measured with the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure, such that the NS3 helicase mutant or the construct controls the movement of the biomolecule to be measured through a pore; and (b) performing one or more measurements while the biomolecule to be measured moves relative to the pore, wherein the measurements indicate one or more characteristics of the biomolecule to be measured, thereby characterizing the biomolecule to be measured.

[0014] In a ninth aspect of the present disclosure, there can be provided the use of the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure for characterizing a biomolecule to be tested or controlling the movement of the biomolecule to be tested through a pore.

[0015] In a tenth aspect of the present disclosure, there can be provided an analysis device for characterizing a biomolecule to be tested, comprising: (a) one or more pores; and (b) one or more NS3 helicase mutants described in the first aspect of the present disclosure or constructs described in the second aspect of the present disclosure.

[0016] In an eleventh aspect of the present disclosure, there can be provided a method for forming an analysis device for characterizing a biomolecule to be tested, the method comprising: forming a complex of (a) a pore and (b) an NS3 helicase mutant described in the first aspect of the present disclosure or a construct described in the second aspect of the present disclosure, thereby forming an analysis device for characterizing the biomolecule to be tested.

[0017] In a twelfth aspect of the present disclosure, there can be provided a complex comprising a linker and an NS3 helicase mutant described in the first aspect of the present disclosure or a construct described in the second aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present disclosure, the drawings required for use in the description of the specific embodiments will be briefly introduced below.

[0019] Figure 1 The purification results of wild-type NS3 helicase and its mutants are shown, where lane 1 corresponds to wild-type NS3 helicase (SEQ ID NO: 1), lanes 2 - 57 correspond to NS3 helicase mutants 1 - 56 respectively, and lanes 59 - 60 correspond to NS3 helicase mutants 57 - 58 respectively.

[0020] Figure 2 The electrophoresis pattern verifying the preparation result of the Y-shaped linker is shown, where lane 1 corresponds to the protein molecular weight standard, lane 2 corresponds to the Y-shaped linker, lane 3 corresponds to Y1, lane 4 corresponds to YB, and lane 5 corresponds to Y2.

[0021] Figure 3 The binding effect of wild-type NS3 helicase and its mutants with the sequencing linker is shown, where lanes 1 - 37 correspond to the EMSA results of wild-type NS3 helicase (SEQ ID NO: 1), NS3 helicase mutants 1 - 35, and the free linker respectively.

[0022] Figure 4Shows the binding effects of wild-type NS3 helicase and its mutants with sequencing adapters, where lanes 1-20 correspond to the EMSA results of NS3 helicase mutants 36-51, free adapter, and NS3 helicase mutants 10, 11, and 18 respectively.

[0023] Figure 5 Shows the binding effects of wild-type NS3 helicase and its mutants with sequencing adapters, where lanes 1-9 correspond to the EMSA results of NS3 helicase mutants 11, 52-58, and free adapter respectively.

[0024] Figure 6 Shows the flow chart for preparing the construct for RNA sequencing and the structure of the construct for RNA sequencing.

[0025] Figure 7 Shows the raw current signal map obtained by nanopore sequencing of RNA using NS3 helicase mutant 11.

[0026] Figure 8 Shows the current distribution map obtained by nanopore sequencing of RNA using NS3 helicase mutant 11.

[0027] Figure 9 Shows the raw current signal map obtained by nanopore sequencing of RNA using NS3 helicase mutant 57.

[0028] Figure 10 Shows the current distribution map obtained by nanopore sequencing of RNA using NS3 helicase mutant 57.

[0029] Figure 11 Shows the raw current signal map obtained by nanopore sequencing of DNA using NS3 helicase mutant 11.

[0030] Figure 12 Shows the current distribution map obtained by nanopore sequencing of DNA using NS3 helicase mutant 11.

[0031] Figure 13 Shows the raw current signal map obtained by nanopore sequencing of DNA using NS3 helicase mutant 57.

[0032] Figure 14 Shows the current distribution map obtained by nanopore sequencing of DNA using NS3 helicase mutant 57.

[0033] Figure 15 Shows the raw current signal map obtained by nanopore sequencing of RNA using NS3 helicase mutant 41.

[0034] Figure 16Shows the current distribution map obtained by nanopore sequencing of RNA using the NS3 helicase mutant 41.

[0035] Figure 17 Shows the original current signal map obtained by nanopore sequencing of RNA using the wild-type NS3 helicase.

[0036] Figure 18 Shows the current distribution map obtained by nanopore sequencing of RNA using the wild-type NS3 helicase. Detailed implementation manners

[0037] In order to enable those of ordinary skill in the art to better understand the technical solutions of this application, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. It should be understood that the specific implementation manners described herein are only intended to explain this application, rather than limiting this application. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the implementation manners is only to provide a better understanding of this application by showing examples of this application.

[0038] In order to more clearly explain the implementation manners of the present invention, some scientific terms and proprietary nouns are used herein. Unless clearly defined herein, all these terms and nouns should be understood to have the meanings commonly understood by those skilled in the art. For the sake of clarity, the following definitions are given for some of the terms used herein.

[0039] In this article, the terms "protein", "protein", "polypeptide", "peptide" are considered to have the same meaning and can be used interchangeably to refer to a molecule containing amino acid residues linked by peptide bonds and containing more than five amino acid residues, usually containing 20 or more amino acids, optionally containing 50 or more amino acids, or containing 100 or more amino acids. Proteins can be optionally modified (e.g., glycosylation, phosphorylation, acylation, farnesylation, isoprenylation, sulfonation, etc.) to increase their functionality or activity. A protein that exhibits activity under certain conditions and in the presence of a specific substrate can be called an "enzyme". It should be understood that due to the degeneracy of the genetic code, multiple nucleotide sequences encoding a given protein can be generated.

[0040] The "nucleic acid" described herein is a general term for deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), which are macromolecular compounds polymerized by many nucleotide monomers. In this article, the term "nucleic acid" is considered to have the same meaning as the term "polynucleotide"; therefore, the term "nucleic acid" and the term "polynucleotide" can be used interchangeably.

[0041] Nucleotide monomers are composed of a pentose sugar, a phosphate group, and a nitrogenous base. If the pentose sugar is ribose, the resulting polymer is RNA; if the pentose sugar is deoxyribose, the resulting polymer is DNA. The nitrogenous bases in nucleotides can include, but are not limited to: adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C). The nucleotides can be naturally occurring or synthetic. Thus, the "nucleotides" described herein include, but are not limited to: adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), and deoxycytidine monophosphate (dCMP). Preferably, the nucleotides are selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, or dCMP.

[0042] The term "expression" includes any step involved in protein production, including but not limited to transcription, post-transcriptional modification, translation, post-translational modification, and secretion.

[0043] An "expression vector" contains a polynucleotide encoding a protein, which is operably linked to appropriate control sequences (such as a promoter, as well as transcriptional and translational termination signals) for expression and / or translation in vitro. The expression vector can be any vector (e.g., a plasmid or a virus), which can conveniently undergo recombinant DNA procedures and can cause the expression of the polynucleotide. The choice of vector will usually depend on the compatibility of the vector with the cell into which the vector is to be introduced. The vector can be a linear or closed circular plasmid. The vector can be an autonomously replicating vector, i.e., a vector that exists as an extrachromosomal entity, the replication of which is independent of chromosomal replication, such as a plasmid, an extrachromosomal element, a minichromosome, or an artificial chromosome. Alternatively, the vector can be a vector that integrates into the genome when introduced into the host cell and replicates with the chromosome into which it has integrated. The integrative cloning vector can integrate at a random or predetermined target locus in the chromosome of the host cell. The vector system can be a single vector or plasmid or two or more vectors or plasmids, which together contain the total DNA to be introduced into the genome of the host cell, or a transposon.

[0044] As used herein, the term "control sequence" refers to components that participate in the regulation of the expression of a coding sequence, either in a particular organism or in vitro. Examples of control sequences are transcriptional initiation sequences, termination sequences, promoters, leader sequences, signal peptides, propeptides, prepropeptides or enhancer sequences; Shine-Delgarno sequences, repressor or activator sequences; efficient RNA processing signals, such as splicing and polyadenylation signals; sequences that stabilize cytoplasmic mRNA; sequences that enhance translation efficiency (e.g., ribosome binding sites); sequences that enhance protein stability; and, when required, sequences that enhance protein secretion.

[0045] As defined herein, a "host cell" is an organism suitable for genetic manipulation and capable of being used in the production of a target product, such as the NS3 helicase mutant described in the first aspect of the present disclosure. A host cell can be a host cell found in nature, or a host cell derived from a parental host cell after genetic manipulation or classical mutagenesis. Advantageously, the host cell is a recombinant host cell. The host cell can be a prokaryotic, archaeal or eukaryotic host cell. Prokaryotic host cells can be, but are not limited to, bacterial host cells. Eukaryotic host cells can be, but are not limited to, yeast, fungal, amoebal, algal, plant, animal, or insect host cells.

[0046] When used in reference to a nucleic acid or a protein (or an enzyme), the term "recombinant" means that the nucleic acid or protein (or enzyme) has been modified in sequence by artificial intervention compared to its native form. When referring to a cell (e.g., a host cell), the term "recombinant" means that the genome of the cell has been modified in sequence by artificial intervention compared to its native form.

[0047] When used in reference to an NS3 helicase mutant, the term "substitution" means that a natural amino acid residue present in the corresponding wild-type NS3 helicase is replaced by another amino acid residue.

[0048] The similarity between two protein sequences or nucleic acid sequences can be represented by their identity. To determine the percentage of sequence identity between two amino acid sequences or two nucleic acid sequences, the sequences are aligned to achieve the best match, and sequence identity is the percentage of identical matches in the aligned regions between the two sequences. The percentage of sequence identity between two amino acid sequences or between two polynucleotide sequences can be determined using well-known algorithms, such as the Needleman and Wunsch algorithm (Needleman, S.B. and Wunsch, C.D. (1970) J. Mol. Biol. 48, 443-453) for aligning two sequences. For example, it can be performed using the NEEDLE program from the EMBOSS package. Those skilled in the art will understand that slightly different results may be obtained when using different algorithms or different parameters of a specific algorithm, but the percentage of identity between the two sequences will not change significantly.

[0049] As used in this application, "at least one", "at least one", "one or more", "one or more", "one or more" or "one or more" includes: one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve or more.

[0050] As used herein, "and / or" includes any of the listed items and any number of any combination of items. The expression "A and / or B" includes three cases: (1) A; (2) B; and (3) A and B. The expression "A, B and / or C" includes seven cases: (1) A; (2) B; (3) C; (4) A and B; (5) A and C; (6) B and C; and (7) A, B and C. The meaning of similar expressions can be inferred accordingly.

[0051] As used in this application, "comprising", "containing" or "including" is an open-ended description, indicating the presence of the specified component or step described, as well as other specified components or steps that do not have a substantial impact. In particular, when the above terms are used to describe the sequence of a protein or nucleic acid, it means that the protein or nucleic acid can either consist of the sequence, or can have additional amino acid residues or nucleotides at one or both ends of the protein or nucleic acid, but the protein or nucleic acid still has the activity described in this application (such as moderate binding affinity to the biomolecule to be tested, etc.).

[0052] In a first aspect of the present disclosure, an NS3 helicase mutant can be provided, where the NS3 helicase mutant includes the amino acid sequence shown in SEQ ID NO: 1, or an amino acid sequence having at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity with the amino acid sequence shown in SEQ ID NO: 1.

[0053] As used herein, the term "NS3 helicase" should be understood in its broad sense and is intended to cover wild-type NS3 helicase as well as NS3 helicase mutants.

[0054] As used herein, an "NS3 helicase mutant" refers to a substitution, insertion, deletion, truncation, transversion, etc. at one or more positions relative to a reference sequence (i.e., the wild-type NS3 helicase mentioned). Mutants can be generated by, for example, site-saturation mutagenesis, scanning mutagenesis, insertion mutagenesis, random mutagenesis, site-directed mutagenesis, and directed evolution, as well as various other recombinant methods known to those skilled in the art.

[0055] For research and reports on NS3 helicase, see, for example, Joseph L Kim et.al, "Hepatitis C virus NS3 RNA helicase domain with a bound oligonucleotide: the crystal structure provides insights into the mode of unwinding", Structure, 1998, 6, 89 - 100; Ting Zhou et.al, "NS3 from Hepatitis C Virus Strain JFH-1 Is an Unusually Robust Helicase That Is Primed To Bind and Unwind Viral RNA", Journal of Virology, 2018, 92, e01253 - 17; and Todd C. Appleby et.al, "Visualizing ATP-Dependent RNA Translocation by the NS3 Helicase from HCV", Journal of Molecular Biology, 2011, 405, 1139 - 1153.

[0056] In some embodiments, the NS3 helicase is derived from Pestivirus sp. The source and sequence information of wild-type homologous proteins similar to the NS3 helicase from Pestivirus sp. are shown in the following table.

[0057]

[0058]

[0059] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes a first amino acid substitution at one or more positions selected from the following, optionally two or more positions: 70, 71, 72, 73, 74, 191, 192, 193, 194, 195, 196, 197, 198, 212, 213, 214, 215, 216, 217, 218, 219, 368, 369, 370, 371, 372, 373, 374, 397, 398, 399, 400, 401, 403, 404, 405, 406, and 407, where the positions are defined with reference to SEQ ID NO:1.

[0060] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes a first amino acid substitution at one or more positions selected from the following: 70, 71, 72, 73, 74, 191, 192, 193, 194, 195, 196, 197, 198, 212, 213, 214, 215, 216, 217, 218, and 219, and includes a first amino acid substitution at one or more positions selected from the following: 368, 369, 370, 371, 372, 373, 374, 397, 398, 399, 400, 401, 403, 404, 405, 406, and 407, where the positions are defined with reference to SEQ ID NO:1.

[0061] In some embodiments, the first amino acid substitution includes: the original amino acid at the position is replaced with any different amino acid. In some embodiments, the first amino acid substitution includes: the original amino acid at the position is replaced with cysteine or a non-natural amino acid.

[0062] The unnatural amino acids described in the present application include, but are not limited to: 4-azido-L-phenylalanine (Faz), 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselanyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4(dihydroxyboryl)-L-phenylalanine, 4-[(ethylthio)carbonyl]-L-phenylalanine, (2S)-2-amino-3-{4-[(propan-2-ylthio)carbonyl]phenyl}propanoic acid, (2S)-2-amino-3-{4-[(2-amino-3-sulfanylpropanoyl)amino]phenyl}propanoic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine, 3-hydroxy-L-tyrosine, 3-amino-L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthalen-2-ylamino)propanoic acid, 6-(methylthio)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propanoic acid, (2R)-2-aminooctanoate 3-(2,2'-bipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxyquinolin-3-yl)propanoic acid, 4-benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)thio]propanoic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propanoic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-({[(2-nitrobenzyl)oxy]carbonyl}amino)hexanoic acid, O-(2-nitrobenzyl)-L-tyrosine, 2-nitrophenylalanine, 4-[(E)-phenyldiazenyl]-L-phenylalanine, 4-[3-(trifluoromethyl)-3H-diazirin-3-yl]-D-phenylalanine, 2-amino-3-[[5-(dimethylamino)-1-naphthyl]sulfonylamino]propanoic acid, (2S)-2-amino 4-(7-hydroxy-2-oxo-2H-chromen-4-yl)butanoic acid, (2S)-3-[(6-acetylnaphthalen-2-yl)amino]-2-aminopropanoic acid, 4(carboxymethyl)phenylalanine, 3-nitro-L-tyrosine, O-thio-L-tyrosine, (2R)-6-acetamido-2-aminohexanoate, 1-methylhistidine, 2-aminononanoic acid, 2-aminodecanoic acid, L-homocysteine, 5-sulfanyl-norvaline, 6-sulfanyl-L-norleucine, 5-(methylthio)-L-norvaline, N6-{[(2R,(3R)-3-Methyl-3,4-dihydro-2H-pyrrol-2-yl]carbonyl}-L-lysine, N6-[(benzyloxy)carbonyl]lysine, (2S)-2-amino-6-[(cyclopentylcarbonyl)amino]hexanoic acid, N6-[(cyclopentyloxy)carbonyl]-L-lysine, (2S)-2-amino-6-{[(2R)-tetrahydrofuran-2-ylcarbonyl]amino}hexanoic acid, (2S)-2-amino-8-[(2R,3S)-3-ethynyltetrahydrofuran-2-yl]-8-oxooctanoic acid, N6-(tert-butoxycarbonyl)-L-lysine, (2S)-2-hydroxy-6-({[(2-methyl-2-propanyl)oxy]carbonyl}amino)hexanoic acid, N6-[(allyloxy)carbonyl]lysine, (2S)-2-amino-6-({[(2-azidobenzyl)oxy]carbonyl}amino)hexanoic acid, N6L-prolyl-L-lysine, (2S)-2-amino-6-{[(prop-2-yn-1-yloxy)carbonyl]amino}hexanoic acid or N6-[(2-azidoethoxy)carbonyl]-L-lysine.,

[0063] In some embodiments, a cysteine or other natural or unnatural amino acid generated by a first amino acid substitution at one position is linked to a cysteine or other natural or unnatural amino acid generated by a first amino acid substitution at another position.

[0064] In some embodiments, any natural amino acid generated by a first amino acid substitution at one position is linked to any natural amino acid generated by a first amino acid substitution at another position. In some embodiments, a cysteine or unnatural amino acid generated by a first amino acid substitution at one position is linked to a cysteine or unnatural amino acid generated by a first amino acid substitution at another position. In some embodiments, a cysteine generated by a first amino acid substitution at one position is linked to a cysteine generated by a first amino acid substitution at another position. In some embodiments, an unnatural amino acid generated by a first amino acid substitution at one position is linked to an unnatural amino acid generated by a first amino acid substitution at another position. In some embodiments, an unnatural amino acid generated by a first amino acid substitution at one position is linked to a cysteine generated by a first amino acid substitution at another position.

[0065] In some embodiments, a cysteine or other natural or unnatural amino acid generated by a first amino acid substitution at one position is linked to a cysteine originally present at another position.

[0066] The connection described above can be any type of connection, including temporary or permanent connections, such as covalent bonds, hydrogen bonds, electrostatic interactions, π-π interactions, or hydrophobic interactions, etc. In some embodiments, the connection can be permanent, such as a covalent bond. In some embodiments, a linker can be used for the connection. The linker can be an amino acid sequence and / or a chemical crosslinker. In some embodiments, the connection can be formed through disulfide bond formation, through isopeptide bond formation, and / or through the use of a linker.

[0067] Suitable amino acid linkers are known in the art and include, but are not limited to, flexible linkers and rigid linkers. Flexible linkers have 2 - 20, such as 4, 6, 8, 10, or 16 serine (S) and / or glycine (G), including (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, (SG)8, (SG)10, (SG)15, or (SG)20. Rigid linkers have 2 - 30, such as 4, 6, 8, 16, or 24 proline (P), such as (P)12.

[0068] Suitable chemical crosslinkers are known in the art and include, but are not limited to, those having the following functional groups: maleimide, BM(PEG)3 (a bis - maleimide - activated PEG compound), active ester, succinimide, azide, alkyne (such as dibenzocyclooctynol (DIBO or DBCO), difluorocycloalkyne, and linear alkyne), phosphine (such as those used in traceless and non - traceless Staudinger ligation), haloacetyl (such as iodoacetamide), phosgene - type reagent, sulfonyl chloride reagent, isothiocyanate, acyl halide, hydrazine, disulfide, vinyl sulfone, aziridine, and photoreactive reagents (such as aryl azide, diazirine).

[0069] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes a first amino acid substitution at one or more positions selected from: 212, 214, 215, 216, 217, and 218, and includes a first amino acid substitution at one or more positions selected from: 371, 372, 373, 374, 399, 400, 403, 404, 405, and 406, where the positions are defined with reference to SEQ ID NO:1.

[0070] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes a first amino acid substitution at one or more positions selected from: 214, 215, 216, 217, and 218, and includes a first amino acid substitution at one or more positions selected from: 371, 372, 373, 399, 403, 404, 405, and 406, where the positions are defined with reference to SEQ ID NO:1.

[0071] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes a first amino acid substitution at one or more positions selected from: 215 and 216, and includes a first amino acid substitution at one or more positions selected from: 372 and 373, where the positions are defined with reference to SEQ ID NO:1.

[0072] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes a first amino acid substitution at the following positions: 215 and 372, or includes a first amino acid substitution at the following positions: 216 and 372, or includes a first amino acid substitution at the following positions: 215 and 373, or includes a first amino acid substitution at the following positions: 214 and 404, or includes a first amino acid substitution at the following positions: 215 and 399, or includes a first amino acid substitution at the following positions: 216 and 399, or includes a first amino acid substitution at the following positions: 216 and 403, or includes a first amino acid substitution at the following positions: 216 and 371, or includes a first amino acid substitution at the following positions: 217 and 399, or includes a first amino acid substitution at the following positions: 217 and 403, or includes a first amino acid substitution at the following positions: 217 and 372, or includes a first amino acid substitution at the following positions: 217 and 373, or includes a first amino acid substitution at the following positions: 218 and 399, or includes a first amino acid substitution at the following positions: 218 and 405, or includes a first amino acid substitution at the following positions: 218 and 406, where the positions are defined with reference to SEQ ID NO:1.

[0073] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes a first amino acid substitution at the following positions: 215 and 372, or includes a first amino acid substitution at the following positions: 216 and 372, or includes a first amino acid substitution at the following positions: 215 and 373, where the positions are defined with reference to SEQ ID NO:1.

[0074] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant comprises a second amino acid substitution at one or more positions selected from the following, optionally two or more positions: 25, 47, 68, 70, 71, 95, 112, 117, 130, 131, 143, 191, 232, 251, 321, 325, 327, 330, 331, 333, 375, 376, and 377, where the positions are defined with reference to SEQ ID NO:1.

[0075] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant comprises a second amino acid substitution at one or more positions selected from the following, optionally two or more positions: 71, 112, 143, 191, 232, 251, and 327, where the positions are defined with reference to SEQ ID NO:1.

[0076] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant comprises a second amino acid substitution at one position selected from the following: 71, 112, 143, 191, 232, 251, and 327, or comprises a second amino acid substitution at two positions selected from the following: 71 and 143, or comprises a second amino acid substitution at two positions selected from the following: 191 and 232, where the positions are defined with reference to SEQ ID NO:1.

[0077] In some embodiments, the second amino acid substitutions include: K25H, K25R, V47H, V47Y, V47F, F68H, F68R, G70K, G70R, T71E, T71H, T71N, T71K, T71R, N95E, N95D, H112F, H112K, H112R, H112W, H112Y, T117S, T117V, T117I, E130G, E130L, E130A, Q131N, Q131S, Q131T, T143V, T143I, T143S, N191H, N191R, N191W, N191Y, N191F, A232H, A232W, A232R, A232F, V251H, V251W, V251R, V251F, Y321F, Y321W, Y321H, E325L, E325Q, E325N, E325V, D327G, D327L, D327N, D327I, K330G, K330R, K330H, Y331R, Y331G, Y331F, Y331W, Y331H, E333L, E333Q, E333N, E333V, E333D, T375G, T375A, T375V, T375Y, Q376R, Q376H, Q376N, Q376Y, W377F, W377Y and W377H, where the positions are defined with reference to SEQ ID NO:1.

[0078] In some embodiments, the second amino acid substitutions include: T71E, T71H, T71N, H112W, T143V, N191H, N191F, N191R, N191W, N191Y, A232H, V251H, D327N, D327G and D327L, where the positions are defined with reference to SEQ ID NO:1.

[0079] In some embodiments, compared with SEQ ID NO:1, the NS3 helicase mutant includes the following second amino acid substitutions: T71E; T71H; T71N; H112W; T143V; T71H / T143V; N191H / A232H; N191H; N191F; N191R; N191W; N191Y; A232H; V251H; D327N; D327G; D327L, where the positions are defined with reference to SEQ ID NO:1.

[0080] In the present application, "T71H / T143V" means that both the amino acid substitutions T71H and T143V are included, and so on.

[0081] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes a third amino acid substitution at one or more positions selected from the following, optionally two or more positions: 108, 111, 187, 193, 246, 247, 248, 284, 338, 364, 374, and 388, where the positions are defined with reference to SEQ ID NO:1.

[0082] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes a third amino acid substitution at one or more positions selected from the following, optionally two or more positions: 187, 246, 247, 284, 338, and 388, where the positions are defined with reference to SEQ ID NO:1.

[0083] In some embodiments, the third amino acid substitution includes: C108S, C111S, C111V, C111L, C111T, C187S, C187V, C187L, C187T, C193S, C193V, C193L, C193T, C246S, C247N, C247S, C247V, C247L, C247T, C248N, C248S, C248V, C248L, C248T, C284S, C284V, C284L, C284T, C338S, C338V, C338L, C338T, C364S, C364V, C364L, C364T, C374S, C374V, C374L, C374T, C388S, C388V, C388L, and C388T, where the positions are defined with reference to SEQ ID NO:1.

[0084] In some embodiments, the third amino acid substitution includes: C187S, C246S, C247N, C284T, C338S, and C388S, where the positions are defined with reference to SEQ ID NO:1.

[0085] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant includes the following third amino acid substitutions: C187S; C246S / C247N; C284T; C338S; C388S, where the positions are defined with reference to SEQ ID NO:1.

[0086] In some embodiments, the NS3 helicase mutant comprises at least one of the first amino acid substitutions and at least one of the second amino acid substitutions.

[0087] In some embodiments, the NS3 helicase mutant comprises at least one of the first amino acid substitutions and at least one of the third amino acid substitutions.

[0088] In some embodiments, the NS3 helicase mutant comprises at least one of said second amino acid substitutions and at least one of said third amino acid substitutions.

[0089] In some embodiments, the NS3 helicase mutant comprises at least one of said first amino acid substitutions, at least one of said second amino acid substitutions, and at least one of said third amino acid substitutions.

[0090] In some embodiments, compared to SEQ ID NO:1, the NS3 helicase mutant comprises the following amino acid substitutions: K214C / G404C; S215C / D399C; S215C / L372C; S215C / R373C; P216C / D399C; P216C / A403C; P216C / A371C; P216C / L372C; S217C / D399C; S217C / A403C; S217C / L372C; S217C / R373C; V218C / D399C; V218C / V405C; V218C / H406C; T71E / S215C / L372C; T71H / S215C / L372C; H112W / S215C / L372C; T143V / S215C / L372C; S215C / R373C / T71H / T143V; S215C / R373C / N191H / A232H; S215C / R373C / D327N; S215C / R373C / N191R; S215C / R373C / N191W; S215C / R373C / N191Y; P216C / L372C / A232H; P216C / L372C / V251H; P216C / L372C / D327G; P216C / L372C / D327L; P216C / L372C / N191F; C187S; C246S / C247N; C284T; S215C / L372C / N191F / A232H / C388S; S215C / L372C / N191F / A232H / C187S / C284T, where the positions are defined with reference to SEQ ID NO:1.

[0091] In some embodiments, the NS3 helicase mutant comprises an amino acid sequence as shown in any one of SEQ ID NOs: 2-7.

[0092] In a second aspect of the present disclosure, a construct can be provided that comprises the NS3 helicase mutant described in the first aspect of the present disclosure.

[0093] In some embodiments, the construct further comprises a sequence that is more conducive to controlling the translocation of a biomolecule to be measured through a nanopore during nanopore sequencing.

[0094] In some embodiments, the construct comprises one or more NS3 helicase mutants described in the first aspect of the present disclosure.

[0095] The construct according to the present disclosure can be modified to facilitate identification or purification, for example, by adding histidine residues (His-tag), aspartic acid residues (asp-tag), streptavidin tag, Flag tag, SUMO tag, GST tag or MBP tag or Strep TagII tag, or by adding a signal sequence to promote their secretion from cells in which the polypeptide is not naturally contained in the signal sequence. The replacement method for introducing a genetic tag is to chemically link the tag to a natural or artificial site on the construct.

[0096] In the third aspect of the present disclosure, a nucleic acid encoding the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure can be provided.

[0097] In some embodiments, the nucleic acid is contained in a vector selected from plasmids, viruses, and bacteriophages.

[0098] In the fourth aspect of the present disclosure, an expression vector comprising the nucleic acid described in the third aspect of the present disclosure can be provided.

[0099] In some embodiments, the expression vector includes plasmids, viruses, and bacteriophages. In some embodiments, the expression vector further comprises regulatory elements for controlling the expression of the nucleic acid. In some embodiments, the regulatory element is a promoter operably linked to the nucleic acid. In some embodiments, the promoter includes T7, trc, lac, ara, and λL.

[0100] There are various methods known to those skilled in the art for inserting nucleic acids into expression vectors. See, for example, Sambrook and Russell, Molecular Cloning: A Laboratory Manual, 3rd Edition, CSHL Press, Cold Spring Harbor, NY, 2001.

[0101] In the fifth aspect of the present disclosure, a host cell comprising the nucleic acid described in the third aspect of the present disclosure or comprising the expression vector described in the fourth aspect of the present disclosure can be provided.

[0102] In some embodiments, the host cell is Escherichia coli. In some embodiments, the host cell is selected from BL21(DE3), JM109(DE3), B834(DE3), TUNER, C41(DE3), Rosetta2(DE3), Origami, Origami B, and the like.

[0103] In a sixth aspect of the present disclosure, a method for preparing the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure can be provided, including: culturing the host cell described in the fifth aspect of the present disclosure, inducing expression, and then purifying the obtained expression product.

[0104] Genetic engineering techniques, such as overexpression of enzymes in host cells, genetic modification of host cells, or hybridization techniques, are methods known in the art, such as those described in Sambrook and Russell, Molecular Cloning: A Laboratory Manual, 3rd Edition, CSHL Press, Cold Spring Harbor, NY, 2001 or F. Ausubel et al., "Current protocols in molecular biology", Green Publishing and Wiley Interscience, New York, 1987.

[0105] In a seventh aspect of the present disclosure, a method for controlling the movement of a biomolecule to be measured can be provided, the method including: contacting the biomolecule to be measured with the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure, thereby controlling the movement of the biomolecule to be measured.

[0106] In some embodiments, the biomolecule to be measured includes polynucleotides, polypeptides, polysaccharides, or lipids. The biomolecule to be measured can pass through the pores on the molecular membrane in sequence under the action of an electric potential. When the biomolecule to be measured passes through the pores, it will cause a change in the electrical signal, and the characterization information of the biomolecule to be measured can be further obtained according to the electrical signal. For example, the size information, sequence information, identity information, modification information, etc. of the biomolecule to be measured can be obtained according to the current change information. In some embodiments, the biomolecule to be measured is DNA or RNA. In some embodiments, the biomolecule to be measured includes fully double-stranded polynucleotides, partially double-stranded polynucleotides, or single-stranded polynucleotides.

[0107] The binding of the biomolecule to be detected to the NS3 helicase can be direct or indirect. When the biomolecule to be detected is a polynucleotide, the biomolecule to be detected can directly bind to the NS3 helicase; when the biomolecule to be detected is, for example, a polypeptide, a polysaccharide or a lipid, the biomolecule to be detected can indirectly bind to the NS3 helicase. For example, the biomolecule to be detected can be linked to a polynucleotide (which can directly bind to the NS3 helicase) and move as the NS3 helicase controls the movement of the polynucleotide.

[0108] In some embodiments, controlling the movement of the biomolecule to be detected is controlling the movement of the biomolecule to be detected through a pore. In some embodiments, the pore can be a nanopore. In some embodiments, the pore can be a transmembrane pore. The pore can be natural or artificial, including but not limited to biological pores, solid pores, or pores that are a hybrid of biological and solid.

[0109] In some embodiments, the method can include using one or more NS3 helicase mutants to jointly control the movement of the biomolecule to be detected. In some embodiments, when more than one NS3 helicase mutant is used, the NS3 helicase mutants can be the same or different.

[0110] In an eighth aspect of the present disclosure, a method for characterizing a biomolecule to be detected can be provided. The method includes: (a) contacting the biomolecule to be detected with the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure, such that the NS3 helicase mutant or the construct controls the movement of the biomolecule to be detected through a pore; and (b) performing one or more measurements when the biomolecule to be detected moves relative to the pore, wherein the measurements indicate one or more characteristics of the biomolecule to be detected, thereby characterizing the biomolecule to be detected.

[0111] In some embodiments, steps (a) and (b) can be repeated one or more times.

[0112] In some embodiments, any number of the NS3 helicase mutants described in the first aspect of the present disclosure can be used in the method. In some embodiments, one or more of the NS3 helicase mutants described in the first aspect of the present disclosure are used. In some embodiments, one or more of the NS3 helicase mutants described in the first aspect of the present disclosure are used. In some embodiments, when more than one NS3 helicase mutant is used, the NS3 helicase mutants can be the same or different. In some embodiments, more than one NS3 helicase mutants can be linked to each other or only arranged by binding to the biomolecule to be detected respectively to perform the function of controlling the movement of the biomolecule to be detected.

[0113] In some embodiments, the method further comprises the following step: applying a potential difference across a pore in contact with the NS3 helicase mutant or construct and the biomolecule to be tested.

[0114] In some embodiments, the pore is a structure that allows hydrated ions to flow from one side of the membrane to the other layer of the membrane under the drive of the applied potential. In some embodiments, the pore is a nanopore. In some embodiments, the pore is a transmembrane pore. The pore provides a channel for the movement of the biomolecule to be tested. In some embodiments, the pore is selected from biological pores, solid-state pores, or pores hybridized by biological and solid-state materials.

[0115] In some embodiments, the biological pores include but are not limited to those derived from Mycobacterium smegmatis porin A, Mycobacterium smegmatis porin B, Mycobacterium smegmatis porin C, Mycobacterium smegmatis porin D, CSGG, hemolysin, cytolysin, interleukin, outer membrane porin F, outer membrane porin G, outer membrane phospholipase A, WZA, or Neisseria autotransporter lipoprotein, etc. The solid-state pores are derived from graphene nanopores, MoS2 nanopores, BN nanopores, or PA63 nanopores.

[0116] The membrane can be any membrane existing in the prior art, preferably an amphiphilic molecular layer, that is, a layer formed by amphiphilic molecules such as phospholipids having at least one hydrophilic part and at least one lipophilic or hydrophobic part. The amphiphilic molecules can be synthetic or naturally occurring. In some embodiments, the membrane is a lipid bilayer membrane. The biomolecule to be tested can be connected to the membrane using any known method. If the membrane is an amphiphilic molecular layer, such as a lipid bilayer, the biomolecule to be tested is preferably connected to the membrane through a polypeptide present in the membrane or through a hydrophobic anchor present in the membrane. Among them, the hydrophobic anchor is preferably a lipid, fatty acid, sterol, carbon nanotube, or amino acid.

[0117] In some embodiments, when a force (such as voltage) is applied across the pore, the rate of the biomolecule to be tested passing through the pore is controlled by the mutant or construct, so as to obtain a recognizable and stable current level for determining the characteristics of the biomolecule to be tested.

[0118] In some embodiments, the biomolecule to be tested is single-stranded, double-stranded, or at least partially double-stranded.

[0119] In some embodiments, the biomolecule to be tested can be modified by means of tags, spacers, methylation, oxidation, or damage.

[0120] In some embodiments, the length of the biomolecule to be measured can be 10 - 100,000 bases or amino acids. In some embodiments, the length of the biomolecule to be measured can be at least 10, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 1,000, at least 2,000, at least 5,000, at least 10,000, at least 50,000 or at least 100,000 bases or amino acids.

[0121] In some embodiments, one or more of the features are selected from the source, length, identity, sequence, secondary structure of the biomolecule to be measured, or whether the biomolecule to be measured is modified. In some embodiments, one or more of the features are performed by electrical measurement and / or optical measurement. In some embodiments, electrical measurements include, but are not limited to, current measurement, impedance measurement, tunneling measurement, wind tunnel measurement, or field effect transistor (FET) measurement, etc.

[0122] The electrical signals described in the present disclosure are selected from the measured values of current, voltage, tunneling, resistance, potential, conductivity, or lateral electrical measurement. In some embodiments, the electrical signal is the current passing through the pore.

[0123] In some embodiments, the characterization further includes applying an improved Viterbi algorithm.

[0124] In a ninth aspect of the present disclosure, there can be provided the use of the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure for characterizing a biomolecule to be measured or controlling the movement of the biomolecule to be measured through a pore.

[0125] In a tenth aspect of the present disclosure, there can be provided an analytical device for characterizing a biomolecule to be measured, comprising: (a) one or more pores; and (b) one or more of the NS3 helicase mutants described in the first aspect of the present disclosure or the constructs described in the second aspect of the present disclosure.

[0126] In some embodiments, the analytical device is a sensor or a kit.

[0127] In some embodiments, the analytical device is a kit. In some embodiments, the kit further includes a chip comprising a lipid bilayer. In some embodiments, the pore spans the lipid bilayer. In some embodiments, the kit comprises one or more lipid bilayers, each lipid bilayer comprising one or more of the pores. In some embodiments, the kit further includes reagents for performing the characterization of the biomolecule to be measured. In some embodiments, the reagents include buffers, enzyme tools required for PCR amplification, etc.

[0128] In the eleventh aspect of the present disclosure, a method for forming an analytical device for characterizing a biomolecule to be measured can be provided. The method includes: forming a complex of (a) a pore and (b) the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure, thereby forming an analytical device for characterizing the biomolecule to be measured.

[0129] In the twelfth aspect of the present disclosure, a complex can be provided, which includes a linker and the NS3 helicase mutant described in the first aspect of the present disclosure or the construct described in the second aspect of the present disclosure.

[0130] The various embodiments and preferred options described above for the NS3 helicase mutant of the present disclosure can be combined with each other (as long as they are not inherently contradictory to each other), and are equally applicable to the constructs, nucleic acids, expression vectors, host cells, methods, uses, analytical devices, and complexes of the present disclosure (as long as they are not inherently contradictory to each other), and vice versa. All the various embodiments formed by such combinations are regarded as part of the disclosure of this application.

[0131] The technical solutions of the present disclosure will be more clearly and explicitly illustrated below by way of examples. It should be understood that these examples are only for illustrative purposes and are by no means intended to limit the protection scope of the present disclosure. The protection scope of the present disclosure is only defined by the claims.

[0132] Examples

[0133] Example 1: Expression and purification of NS3 helicase or its mutant

[0134] The nucleic acid sequences corresponding to each wild-type NS3 helicase and its mutant were ligated into the vector pET29a by restriction enzyme digestion and ligation or homologous recombination. After correct sequencing verification, they were transformed into the expression competent host cell BL21(DE3). Single colonies were picked from the plate and inoculated into 10 ml of liquid LB medium with kanamycin resistance. After overnight culture at 37 °C, they were transferred to a large bottle of medium for scale-up culture the next day. When the OD 600 reached about 0.4 - 0.8, isopropyl-β-D-thiogalactoside (IPTG) with a final concentration of 0.5 mM was added, and induction expression was carried out overnight at 16 °C for about 16 hours. The cells collected by low-temperature centrifugation were resuspended in Ni-A buffer and then homogenized under high pressure for fragmentation. The supernatant was collected by centrifugation at 16000 g and 4 °C for subsequent protein chromatography purification, including nickel ion affinity chromatography and ion exchange chromatography. The purified target protein was detected by SDS-PAGE gel electrophoresis (without excision of the His tag), and the results were as Figure 1As shown, lane 1 corresponds to wild-type NS3 helicase (SEQ ID NO:1), lanes 2-57 correspond to NS3 helicase mutants 1-56 respectively, lanes 59-60 correspond to NS3 helicase mutants 57-58 respectively, and to the left of lane 1 is a protein molecular weight standard, where the bands from top to bottom represent 250Kd, 150Kd, 100Kd, 70Kd, 50Kd, 40Kd, 35Kd, 25Kd, and 20Kd respectively. Figure 1 Show: The wild-type NS3 helicase and NS3 helicase mutants 1-58 were successfully purified. The amino acid substitutions of NS3 helicase mutants 1-58 relative to SEQ ID NO:1 are shown in Table 1 below:

[0135] Table 1: Amino acid substitutions of NS3 helicase mutants 1-58 relative to SEQ ID NO:1

[0136]

[0137]

[0138] Example 2: Preparation of sequencing adapters

[0139] Anneal the three synthetic single strands Y1, Y2, and YB in a ratio of 1:1.1:1.1 (slowly cooling from 95°C to 25°C, with a cooling rate not exceeding 0.1°C / s). The final annealing system includes 60 mM HEPES pH 7.0; 200 mM NaCl, and the final concentration of Y1 is 4-8 μM, finally forming a sequencing adapter: Y-shaped adaptor. Detect the prepared Y-shaped adaptor using SDS-PAGE, and the results are as Figure 2 shown. Figure 2 Show: The Y-shaped adaptor was successfully obtained.

[0140]

[0141] Example 3: Binding effect of NS3 helicase or its mutants with sequencing adapters

[0142] Use gel mobility shift assay to test the affinity of NS3 helicase or its mutants with sequencing adapters. The specific method is as follows: Mix the Y-shaped adaptor (500 nM) with 7.5-fold molar amount of NS3 helicase or its mutants in a buffer (100 mM NaAc (pH 7); 1.5 mM BM(PEG)3) and incubate at room temperature for 30 minutes, then electrophorese for 40 min under the condition of 1*TBE buffer concentration, and the results are as Figure 3 , Figure 4 and Figure 5 shown.

[0143] In Figure 3Among them, lanes 1-37 respectively correspond to the EMSA results of wild-type NS3 helicase (SEQ ID NO: 1), NS3 helicase mutants 1-35, and free linker. Figure 3 It shows that: after the wild-type NS3 helicase binds to the sequencing linker, there is no single-enzyme binding band; mutants 3, 4, 10, 11, 12, 13, 17, 18, 20, 21, 26, 27, 28, 31, 32 have single-enzyme binding bands after binding to the sequencing linker; mutants 3, 10, 11, 17, 18, 20, 21, 27, 31 have obvious single-enzyme binding bands after binding to the sequencing linker; mutants 10, 11, 18 have particularly obvious single-enzyme binding bands after binding to the sequencing linker, and the single-enzyme binding ratio is relatively high.

[0144] In Figure 4 Among them, lanes 1-20 respectively correspond to the EMSA results of NS3 helicase mutants 36-51, free linker, and NS3 helicase mutants 10, 11, 18. Figure 4 It shows that: mutants 36, 40, 41, 42, 43, 47 have obvious single-enzyme binding bands and a little double-enzyme binding bands after binding to the sequencing linker; 39, 51 have obvious single-enzyme binding bands and trace double-enzyme binding bands after binding to the sequencing linker; 37, 44, 45, 46, 48, 49, 50, 10, 11, 18 have obvious single-enzyme binding bands after binding to the sequencing linker, and the double-enzyme binding bands are invisible or almost invisible.

[0145] In Figure 5 Among them, lanes 1-9 respectively correspond to the EMSA results of NS3 helicase mutants 11, 52-58, and free linker. Figure 5 It shows that: mutants 52, 54 have obvious single-enzyme binding bands and trace double-enzyme binding bands after binding to the sequencing linker; mutants 11, 53, 57, 58 have obvious single-enzyme binding bands after binding to the sequencing linker, and the double-enzyme binding bands are invisible or almost invisible; in addition, the single-enzyme binding ratio of mutants 57, 58 exceeds 50%, showing a significant improvement compared with mutant 11.

[0146] Example 4: Using NS3 helicase for nanopore sequencing of RNA

[0147] Anneal Oligo A1 (5’-(phos)GGCTTCTTCTTGCTCTTAGGTAGTAGGTTC-3’) and Oligo B (3’-TTTTTTTTTT CCGAAGAAGAACGAGAATCC AGTCCAGCACCGACC-5’), then ligate the 5’ end of Oligo A1 in the annealed product to a 1.2 kb fixed-length RNA sequence with poly-A at the 3’ end, and then perform reverse transcription on the fixed-length RNA sequence to form an RNA / DNA double strand, obtaining a nucleic acid construct.

[0148] Mix the Y-shaped adaptor (500 nM) with 7.5-fold molar amount of NS3 helicase mutant 11 in a buffer (100 mM NaAc (pH 7); 1.5 mM BM(PEG)3) and incubate at room temperature for 30 minutes to form an enzyme complex.

[0149] Ligate the nucleic acid construct to the 5’ end of Y1 in the Y-shaped adaptor in the enzyme complex, and finally form a construct for RNA sequencing as Figure 6 shown.

[0150] The library construction process for RNA nanopore sequencing follows that described in Wang, Y., Zhao, Y., Bollas, A. et al. Nanopore sequencing technology, bioinformatics and applications. Nat Biotechnol 39, 1348–1365 (2021). Add the construct for RNA sequencing to the Qitian carbon nanopore sequencer Qnome3841 and perform sequencing using the sequencing kit provided with the instrument.

[0151] The sequencing results are as Figure 7 and Figure 8 shown. Figure 7 and Figure 8 respectively show: an example current trace (y-axis coordinate = current (pA), x-axis coordinate = number of sampling points (5000 sampling points / s)) and current distribution map (y-axis coordinate = number of current values (count), x-axis coordinate = current (pA)) when NS3 helicase mutant 11 controls the RNA / DNA double strand and the RNA strand translocates through the nanopore in the 3’-5’ direction. The results show that: when using NS3 helicase mutant 11 as a motor protein for nanopore sequencing of RNA, a relatively stable sequencing current signal was obtained. In addition, when sequencing RNA, the peak helicase rate of NS3 helicase mutant 11 reached 207 bp / s. Additionally, the number of sequencing reads obtained during 1 hour of sequencing was 85 (reads).

[0152] Example 5: Nanopore sequencing of RNA using NS3 helicase

[0153] Example 5 is different from Example 4 in that mutant 57 is used instead of mutant 11.

[0154] The sequencing results are as Figure 9 and Figure 10 shown. Figure 9 and Figure 10 respectively show: an example current trace (y-axis coordinate = current (pA), x-axis coordinate = number of sampling points (5000 sampling points / s)) and a current distribution map (y-axis coordinate = number of current values (count), x-axis coordinate = current (pA)) when mutant 57 controls the RNA / DNA double strand and the RNA strand translocates through the nanopore in the 3'-5' direction. The results show that when using NS3 helicase mutant 57 as the motor protein for nanopore sequencing of RNA, a relatively stable sequencing current signal is obtained, and the signal quality is less noisy than that of NS3 helicase mutant 11. In addition, when sequencing RNA, the peak unwinding rate of NS3 helicase mutant 57 reaches 168 bp / s, which is lower than the peak unwinding rate of NS3 helicase mutant 11. In addition, the number of sequencing reads obtained during 1 hour of sequencing is 159 (pieces).

[0155] Example 6: Nanopore sequencing of DNA using NS3 helicase

[0156] Mix the Y-shaped adaptor (500 nM) with 6-fold molar amount of NS3 helicase mutant 11 in buffer (100 mM NaAc (pH 7); 1.5 mM BM(PEG)3) and incubate at room temperature for 30 minutes, then ligate another 1.2 kb of double-stranded DNA (single-stranded is also possible) to one 5'-end of the Y-shaped adaptor to obtain a DNA sequencing construct. Add the DNA sequencing construct to the Qitian nanopore sequencer Qnome3841 and perform sequencing using the sequencing kit provided with the instrument.

[0157] The sequencing results are as Figure 11 and Figure 12 shown. Figure 11 and Figure 12Respectively shown are: an example current trace (y-axis coordinate = current (pA), x-axis coordinate = number of sampling points (5000 sampling points / s)) when the NS3 helicase mutant 11 controls the translocation of a DNA duplex through the nanopore in the 3'-5' direction, and a current distribution diagram (y-axis coordinate = number of current values (count), x-axis coordinate = current (pA)). The results show that when using the NS3 helicase mutant 11 as the motor protein for nanopore sequencing of DNA, a relatively stable sequencing current signal was obtained. In addition, when sequencing DNA, the peak helicase rate of the NS3 helicase mutant 11 reached 85 bp / s, which is lower than the peak helicase rate when sequencing RNA.

[0158] In addition, the current distribution when the NS3 helicase mutant controls the translocation of DNA through the nanopore is quite different from that when it controls the translocation of RNA through the nanopore. When DNA passes through the pore, an obvious bimodal current value is presented.

[0159] Example 7: Nanopore sequencing of DNA using NS3 helicase

[0160] The difference between Example 7 and Example 6 is that mutant 57 is used instead of mutant 11.

[0161] The sequencing results are as Figure 13 and Figure 14 shown. Figure 13 and Figure 14 Respectively shown are: an example current trace (y-axis coordinate = current (pA), x-axis coordinate = number of sampling points (5000 sampling points / s)) when the NS3 helicase mutant 57 controls the translocation of a DNA duplex through the nanopore in the 3'-5' direction, and a current distribution diagram (y-axis coordinate = number of current values (count), x-axis coordinate = current (pA)). The results show that when using the NS3 helicase mutant 57 as the motor protein for nanopore sequencing of DNA, a relatively stable sequencing current signal was obtained. In addition, when sequencing DNA, the peak helicase rate of the NS3 helicase mutant 57 reached 80 bp / s, which is lower than the peak helicase rate of the NS3 helicase mutant 11 and also lower than the peak helicase rate when sequencing RNA.

[0162] Example 8: Nanopore sequencing of RNA using NS3 helicase

[0163] The difference between Example 8 and Example 4 is that mutant 41 is used instead of mutant 11.

[0164] The sequencing results are as Figure 15 and Figure 16 shown. Figure 15 and Figure 16Respectively shown are: an example current trace (y-axis coordinate = current (pA), x-axis coordinate = number of sampling points (5000 sampling points / s)) and a current distribution diagram (y-axis coordinate = number of current values (count), x-axis coordinate = current (pA)) when mutant 41 controls the RNA / DNA double strand and the RNA strand translocates through the nanopore in the 3'-5' direction. The results show that when using NS3 helicase mutant 41 as the motor protein for nanopore sequencing of RNA, a relatively stable sequencing current signal is obtained. In addition, when sequencing RNA, the peak unwinding rate of NS3 helicase mutant 41 reaches 150 bp / s, which is lower than the peak unwinding rate of NS3 helicase mutant 11. Additionally, the number of sequencing reads obtained during 1 hour of sequencing is 109 (pieces).

[0165] Comparative Example 1: Using NS3 helicase for nanopore sequencing of RNA

[0166] The difference between Comparative Example 1 and Example 4 is that wild-type NS3 helicase is used instead of NS3 helicase mutant 11.

[0167] The sequencing results are as Figure 17 and Figure 18 shown. Figure 17 and Figure 18 Respectively shown are: an example current trace (y-axis coordinate = current (pA), x-axis coordinate = number of sampling points (5000 sampling points / s)) and a current distribution diagram (y-axis coordinate = number of current values (count), x-axis coordinate = current (pA)) when wild-type NS3 helicase controls the RNA / DNA double strand and the RNA strand translocates through the nanopore in the 3'-5' direction. The results show that when using wild-type NS3 helicase as the motor protein for nanopore sequencing of RNA, the obtained sequencing current signal is unstable, and the speed is 240 bp / s.

[0168] In addition, it is statistically obtained that the number of sequencing reads obtained during 1 hour of sequencing is 5 (pieces).

[0169] Furthermore, in Example 4, 5, 8 and Comparative Example 1, the comparison of the number of sequencing reads obtained during 1 hour of sequencing shows that the number of sequencing reads of NS3 helicase mutants 11, 41, 57 is significantly higher than that of wild-type NS3 helicase.

Claims

1. An NS3 helicase mutant, comprising the amino acid sequence as shown in SEQ ID NO:1, or an amino acid sequence having at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% sequence identity to the amino acid sequence as shown in SEQ ID NO:

1.

2. The NS3 helicase mutant of claim 1, wherein the NS3 helicase mutant comprises a first amino acid substitution at one or more positions, optionally two or more positions, selected from the group consisting of: 70, 71, 72, 73, 74, 191, 192, 193, 194, 195, 196, 197, 198, 212, 213, 214, 215, 216, 217, 218, 219, 368, 369, 370, 371, 372, 373, 374, 397, 398, 399, 400, 401, 403, 404, 405, 406 and 407, said positions being defined with reference to SEQ ID NO:

1.

3. The NS3 helicase mutant of claim 1, wherein the NS3 helicase mutant comprises a first amino acid substitution at one or more positions selected from the group consisting of: 70、71、72、73、74、191、192、193、194、195、196、197、198、212、 213, 214, 215, 216, 217, 218 and 219, and comprising a first amino acid substitution at one or more positions selected from the group consisting of: 368, 369, 370, 371, 372, 373, 374, 397, 398, 399, 400, 401, 403, 404, 405, 406 and 407, said positions being with reference to SEQ ID NO:1 defined.

4. The NS3 helicase mutant of any one of the preceding claims, wherein the first amino acid substitution comprises: The native amino acid at that position is substituted with cysteine ​​or an unnatural amino acid.

5. An NS3 helicase mutant according to any of the preceding claims, wherein compared to SEQ ID NO:1, the NS3 helicase mutant comprises a first amino acid substitution at one or more positions selected from the group consisting of: 212, 214, 215, 216, 217 and 218, and comprises a first amino acid substitution at one or more positions selected from the group consisting of: 371, 372, 373, 374, 399, 400, 403, 404, 405 and 406, wherein the positions are defined with reference to SEQ ID NO:

1.

6. An NS3 helicase mutant according to any of the preceding claims, wherein compared to SEQ ID NO:1, the NS3 helicase mutant comprises a first amino acid substitution at one or more positions selected from the group consisting of: 214, 215, 216, 217 and 218, and comprises a first amino acid substitution at one or more positions selected from the group consisting of: 371, 372, 373, 399, 403, 404, 405 and 406, wherein the positions are defined with reference to SEQ ID NO:

1.

7. An NS3 helicase mutant according to any of the preceding claims, wherein compared to SEQ ID NO:1, the NS3 helicase mutant comprises a first amino acid substitution at one or more positions selected from the group consisting of: 215 and 216, and comprises a first amino acid substitution at one or more positions selected from the group consisting of: 372 and 373, wherein the positions are defined with reference to SEQ ID NO:

1.

8. An NS3 helicase mutant according to any of the preceding claims, wherein compared to SEQ ID NO:1, the NS3 helicase mutant comprises a first amino acid substitution at the following positions: 215 and 372, or comprises a first amino acid substitution at the following positions: 216 and 372, or comprises a first amino acid substitution at the following positions: 215 and 373, wherein the positions are defined with reference to SEQ ID NO:

1.

9. The NS3 helicase mutant of any of the preceding claims, wherein the NS3 helicase mutant comprises a second amino acid substitution at one or more positions, optionally two or more positions, selected from the group consisting of: 25, 47, 68, 70, 71, 95, 112, 117, 130, 131, 143, 191, 232, 251, 321, 325, 327, 330, 331, 333, 375, 376 and 377, said positions being defined with reference to SEQ ID NO:

1.

10. The NS3 helicase mutant of any of the preceding claims, wherein the NS3 helicase mutant comprises a second amino acid substitution at one or more positions, optionally two or more positions, selected from the group consisting of: 71, 112, 143, 191, 232, 251 and 327, compared to SEQ ID NO:1, wherein the positions are defined with reference to SEQ ID NO:

1.

11. The NS3 helicase mutant of any one of the preceding claims, wherein the second amino acid substitution comprises: K25H, K25R, V47H, V47Y, V47F, F68H, F68R, G70K, G70R, T71E, T71H, T71N, T71K, T71R, N95E, N95D, H112F, H112K, H112R, H112W, H112Y, T117S, T117V, T117I, E130G, E130L, E130A, Q131N, Q131S, Q131T, T143V, T143I, T143S, N191H, N191R, N191W, N191Y, N191F, A232H, A232W, A232R, A232F, V251H, V251W, V251R, V251F, Y321F, Y321W, Y321H, E325L, E325Q, E325N, E325V, D327G, D327L, D327N, D327I, K330G, K330R, K330H, Y331R, Y331G, Y331F, Y331W, Y331H, E333L, E333Q, E333N, E333V, E333D, T375G, T375A, T375V, T375Y, Q376R, Q376H, Q376N, Q376Y, W377F, W377Y and W377H, said positions being defined with reference to SEQ ID NO:

1.

12. The NS3 helicase mutant of any one of the preceding claims, wherein the second amino acid substitution comprises: T71E, T71H, T71N, H112W, T143V, N191H, N191F, N191R, N191W, N191Y, A232H, V251H, D327N, D327G and D327L, said positions being defined with reference to SEQ ID NO:

1.

13. An NS3 helicase mutant according to any of the preceding claims, wherein compared to SEQ ID NO:1, the NS3 helicase mutant comprises a third amino acid substitution at one or more positions, optionally two or more positions, selected from the group consisting of: 108, 111, 187, 193, 246, 247, 248, 284, 338, 364, 374 and 388, wherein the positions are defined with reference to SEQ ID NO:

1.

14. An NS3 helicase mutant according to any of the preceding claims, wherein compared to SEQ ID NO:1, the NS3 helicase mutant comprises a third amino acid substitution at one or more positions, optionally two or more positions, selected from the group consisting of: 187, 246, 247, 284, 338 and 388, wherein the positions are defined with reference to SEQ ID NO:

1.

15. The NS3 helicase mutant of any one of the preceding claims, wherein the third amino acid substitution comprises: C108S, C111S, C111V, C111L, C111T, C187S, C187V, C187L, C187T, C193S, C193V, C1 93L, C193T, C246S, C247N, C247S, C247V, C247L, C247T, C248N, C248S, C248V, C248L , C248T, C284S, C284V, C284L, C284T, C338S, C338V, C338L, C338T, C364S, C364V, C364L, C364T, C374S, C374V, C374L, C374T, C388S, C388V, C388L and C388T, said positions being defined with reference to SEQ ID NO:

1.

16. The NS3 helicase mutant of any one of the preceding claims, wherein the third amino acid substitution comprises: C187S, C246S, C247N, C284T, C338S and C388S, said positions being defined with reference to SEQ ID NO:

1.

17. The NS3 helicase mutant of any one of the preceding claims, comprising at least one of the first amino acid substitutions and at least one of the second amino acid substitutions.

18. The NS3 helicase mutant of any one of the preceding claims, comprising at least one of the first amino acid substitutions and at least one of the third amino acid substitutions.

19. The NS3 helicase mutant of any one of the preceding claims, comprising at least one of the second amino acid substitutions and at least one of the third amino acid substitutions.

20. The NS3 helicase mutant of any of the preceding claims, comprising at least one of the first amino acid substitutions, at least one of the second amino acid substitutions, and at least one of the third amino acid substitutions.

21. The NS3 helicase mutant according to any one of the preceding claims, comprising The amino acid sequence shown in any one of NOs:2-7.

22. A construct comprising the NS3 helicase mutant according to any one of claims 1-21.

23. A nucleic acid encoding the NS3 helicase mutant of any one of claims 1-21 or the construct of claim 22.

24. The nucleic acid of claim 23, wherein the nucleic acid is contained in a vector selected from the group consisting of a plasmid, a virus and a bacteriophage.

25. An expression vector comprising the nucleic acid according to claim 23 or 24.

26. The expression vector of claim 25, wherein the expression vector comprises a plasmid, a virus and a bacteriophage.

27. The expression vector according to claim 25 or 26, wherein the expression vector further comprises a regulatory element for controlling the expression of the nucleic acid.

28. The expression vector of claim 27, wherein the regulatory element is a promoter operably linked to the nucleic acid.

29. The expression vector of claim 28, wherein the promoter comprises T7, trc, lac, ara and λL.

30. A host cell comprising the nucleic acid according to claim 23 or 24 or comprising the expression vector according to any one of claims 25-29. The host cell according to claim 30 , which is Escherichia coli.

32. A method for preparing the NS3 helicase mutant according to any one of claims 1 to 21 or the construct according to claim 22, comprising: The method comprises culturing the host cell according to claim 30 or 31, inducing expression, and then purifying the obtained expression product.

33. A method for controlling the movement of a biomolecule to be detected, the method comprising: The biomolecule to be detected is contacted with the NS3 helicase mutant according to any one of claims 1 to 21 or the construct according to claim 22, thereby controlling the movement of the biomolecule to be detected.

34. The method of claim 33, wherein the biomolecule to be detected comprises a polynucleotide, a polypeptide, a polysaccharide or a lipid.

35. The method according to claim 33 or 34, wherein the biomolecule to be detected comprises DNA or RNA.

36. The method according to any one of claims 33 to 35, wherein the biomolecule to be detected comprises a completely double-stranded polynucleotide, a partially double-stranded polynucleotide, or a single-stranded polynucleotide.

37. A method for characterizing a biomolecule to be detected, the method comprising: (a) contacting the test biomolecule with the NS3 helicase mutant of any one of claims 1-21 or the construct of claim 22, such that the NS3 helicase mutant or construct controls the movement of the test biomolecule through the pore; and (b) performing one or more measurements as the test biomolecule moves relative to the pore, wherein the measurements indicate one or more characteristics of the test biomolecule, thereby characterizing the test biomolecule.

38. Use of the NS3 helicase mutant of any one of claims 1-21 or the construct of claim 22 for characterizing a test biomolecule or controlling the movement of a test biomolecule through a pore.

39. An analytical device for characterizing a biomolecule to be detected, comprising: (a) one or more holes; and (b) one or more The NS3 helicase mutant according to any one of claims 1-21 or the construct according to claim 22.

40. The analytical device according to claim 39, which is a sensor or a reagent kit.

41. A method of forming an analytical device for characterizing a biomolecule to be detected, the method comprising: (a) the pore and (b) the NS3 helicase mutant according to any one of claims 1 to 21 or the construct according to claim 22 form a complex, thereby forming an analytical device for characterizing the biomolecule to be detected.

42. A complex comprising a linker and the NS3 helicase mutant of any one of claims 1-21 or the construct of claim 22.