Novel modified protein pores and enzymes
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- OXFORD NANOPORE TECH LTD
- Filing Date
- 2023-04-14
- Publication Date
- 2026-04-22
AI Technical Summary
Existing methods for detecting and characterizing analytes, such as polynucleotides, through nanopore sensing face challenges in controlling the movement of analytes and accurately identifying their components due to variations in residence time and signal noise.
The use of modified Dda helicases with specific mutations and novel protein pores or pore complexes, including modified CsgG pores and CsgF peptides, to control the movement of analytes and improve the accuracy of characterization by reducing residence time variations and enhancing signal clarity.
The modified Dda helicases and pore complexes achieve improved accuracy in analyte characterization, with error rates reduced to less than 1%, while maintaining or improving the rate of analyte passage through the pores.
Abstract
Description
[Technical Field]
[0001] The present invention relates to modified Dda helicases that can be used to control the movement of analytes, such as polynucleotides. The modified Dda helicases are used in the detection and characterization of analytes. The present invention also relates to novel protein pores or pore complexes and their use in the detection and characterization of analytes. [Background technology]
[0002] Two of the key components of analyte, particularly polymer, characterization using nanopore sensing are (1) controlling the movement of the polymer through the pore and (2) discriminating between the constituent building blocks as the polymer passes through the pore. During nanopore sensing, the narrowest portion of the pore typically corresponds to the most discriminating portion of the nanopore in terms of the change in measured signal as a function of the analyte moving relative to the nanopore. CsgG was identified as an ungated, nonselective protein secretion channel from Escherichia coli (Goyal et al., 2014) and has been used as a nanopore for analyte detection and characterization. Mutations to the wild-type CsgG pore that improve the properties of the pore in this context have also been disclosed (WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893, which are incorporated by reference in their entireties. WO2015 / 055981, WO2015 / 166276, and WO2016 / 055777, which are incorporated by reference in their entireties, also describe polynucleotide-binding proteins, specifically Dda helicases, that can be used to control the movement of analytes through transmembrane protein pores such as the CsgG pores described herein. Summary of the Invention
[0003] The inventors have surprisingly identified specific Dda mutants with improved ability to control analyte translocation through the pore. When using a pore to sequence a polynucleotide, the system jointly estimates the number and identity of bases / nucleotides passing through the pore. Better control over translocation rate variability can reduce one source of statistical noise and simplify the estimation task. Successive short dwells of a polynucleotide within the pore can cause a failure to call the underlying nucleotide / base, resulting in a deletion error. An abnormally long dwell can lead to an insertion error. Ensuring that each nucleotide / base spends a sufficient time interval within the pore helps resolve statistical uncertainty in nucleotide / base identity from noisy signal levels. Further information can be extracted from the dependence of dwell time on nucleotide / base identity, for example, through interactions with motor enzymes. Reducing the overall variability of dwell times can help extract more accurate information through this channel. During regions where signal levels provide limited information about translocation (e.g., long homopolymer regions), multi-nucleotide / base dwell times can be used to infer the number of bases passing through the pore. Reducing dwell time variability can make these inferences more accurate.
[0004] In some embodiments, the variants of the present invention exhibit improved accuracy when used in methods for controlling analyte translocation through a transmembrane pore and in methods for characterizing analytes using a transmembrane pore. In the context of analyte characterization (particularly polynucleotides), accuracy is understood to mean raw read accuracy, i.e., a single passage of a single molecule through a transmembrane pore. Accuracy is a useful measure for tracking platform improvements in sequencing devices. Accuracy may also refer to consensus accuracy, or the accuracy of detecting specific things, such as mutations in polynucleotide analytes. Additionally or alternatively, accuracy is understood to mean the percentage of bases above a certain confidence level, where the confidence level has been pre-calibrated. In some embodiments, the variants of the present invention exhibit improved accuracy with minimal or no change in speed. In some embodiments, accuracy is improved to provide less than 10% error, less than 5% error, less than 4% error, less than 3% error, less than 2% error, less than 1% error, or less than 0.1% error. The mutants identified by the inventors typically contain a combination of mutations, i.e., one or more modifications in the portion of the mutant that interacts with the transmembrane pore. Precision can also be affected by the rate at which the polymer translocates through the pore under enzyme control, which can be altered by changing the concentration of ATP provided to the enzyme. Surprisingly, the inventors have found that enzymes can exhibit rate changes during successive polymer translocations within the same sequencing run under the same conditions, potentially resulting in a decrease in precision.
[0005] Accuracy can be affected by many factors, including the shape and composition of the nanopore, the enzyme, and the interaction between the enzyme and the nanopore. Accuracy is also affected by the rate at which the polymer translocates through the pore under enzyme control, and the translocation rate can be increased or decreased by changing the concentration of ATP provided to the enzyme. The inventors have surprisingly found that rate variations occur during successive polymer translocations within the same sequencing run under the same sequencing conditions, which can result in decreased sequencing accuracy. The variation in sequencing rate for many polymers can be measured to obtain a normalized rate distribution, and the inventors have surprisingly found that some modified enzymes can result in a lower normalized rate distribution and, therefore, increased sequencing accuracy.
[0006] The present invention provides the following: a modified DNA-dependent ATPase (Dda) helicase, wherein the helicase comprises a modification or substitution at one or more of the positions corresponding to amino acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993; - a construct comprising a helicase of the invention and an additional polynucleotide binding moiety, wherein the helicase is bound to the polynucleotide binding moiety and the construct is capable of controlling the movement of an analyte; a polynucleotide comprising a sequence encoding a helicase of the invention or a construct of the invention, a vector comprising a polynucleotide of the invention operably linked to a promoter, a host cell comprising a vector of the invention, a method for producing a helicase of the invention or a construct of the invention, the method comprising expressing a polynucleotide of the invention, transfecting a cell with a vector of the invention or culturing a host cell of the invention, a method for controlling the movement of an analyte, the method comprising contacting the analyte with a helicase of the invention or a construct of the invention, thereby controlling the movement of the analyte, - a method for characterizing a target analyte, comprising: (a) contacting a target analyte with a transmembrane pore and a helicase of the invention or a construct of the invention such that the helicase or construct controls translocation of the target analyte through the pore; (b) obtaining one or more measurements as the polynucleotide moves relative to the pore, the measurements being indicative of one or more characteristics of the target analyte, thereby characterizing the target analyte; - a method for forming a sensor for characterizing a target analyte, the method comprising forming a complex between (a) a pore and (b) a helicase of the invention or a construct of the invention, thereby forming a sensor for characterizing the target analyte; a sensor for characterizing a target analyte, the sensor comprising a complex between (a) a pore and (b) a helicase of the invention or a construct of the invention; - the use of a helicase of the invention or a construct of the invention for controlling the translocation of a target analyte through a pore; - a kit for characterizing a target analyte, comprising: (a) a pore and helicase of the invention, or a construct of the invention, or (b) a kit comprising a helicase of the invention or a construct of the invention and one or more loading moieties; - a device for characterizing a target analyte in a sample, the device comprising: (a) a plurality of pores; and (b) a plurality of helicases of the invention or a plurality of constructs of the invention; - a method for producing a helicase of the invention, comprising: (a) providing a helicase; (b) modifying a helicase to produce a helicase of the invention; - a method for producing a construct of the invention, comprising binding a helicase of the invention to an additional polynucleotide binding moiety, thereby producing the construct; a series of two or more helicases bound to a polynucleotide, wherein at least one of the two or more helicases is a helicase of the invention; and A method for improving the translocation of a target analyte through a transmembrane pore when the translocation is controlled by a DNA-dependent ATPase (Dda) helicase, wherein the DNA-dependent ATPase (Dda) helicase has been modified to include substitutions at one or more of the positions corresponding to amino acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993 and / or the position corresponding to amino acid position 40 in Dda1993, thereby improving the translocation of the target analyte through the transmembrane pore.
[0007] The inventors have also surprisingly identified new transmembrane pore mutations that improve or alter the rate at which an analyte passes through / relative to the transmembrane pore, preferably where the translocation of the analyte is under the control of a polynucleotide binding protein. In one embodiment of the invention, the transmembrane pore mutation increases the rate at which the analyte passes through / relative to the transmembrane pore. In another embodiment, the transmembrane pore mutation decreases the rate at which the analyte passes through / relative to the transmembrane pore. The rate at which the analyte passes through the pore / relative to the pore may be increased by 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 100%, 150%, 200%, or 300% or more relative to the rate at which the analyte moves through a pore that does not contain a mutation of the invention. The rate at which the analyte passes through the pore / relative to the pore may be decreased by 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or 90% relative to the rate at which the analyte moves through a pore that does not contain a mutation of the invention. The inventors have surprisingly found that these changes in velocity, such as increases or decreases in velocity caused by pore modification, have little or no effect on accuracy readings. This is particularly advantageous in methods of characterizing analytes in which the analyte is contacted with a pore and a polynucleotide-binding protein, such as a helicase of the present invention, such that the polynucleotide-binding protein controls the movement of the target analyte through / relative to the pore. In one embodiment, the mutant pore interacts with the polynucleotide-binding protein in a manner different from other transmembrane pores that do not contain the mutation. The pore mutant may alter the distribution of rates at which DNA translocates through the pore, resulting in a tighter distribution of rates compared to other transmembrane pores that do not contain the mutation, leading to reduced sequencing errors. In a preferred embodiment of the present invention, the modified DNA-dependent ATPase (Dda) helicase of the present invention is used to control the movement of an analyte, such as a polynucleotide, through a transmembrane pore of the present invention.
[0008] The present invention provides an isolated CsgG pore or a homologue or variant thereof, or an isolated pore complex comprising a CsgG pore or a homologue or variant thereof and a modified CsgF peptide or a homologue or variant thereof, wherein the CsgG pore comprises at least one monomer comprising a modification at one or more of positions W97, Q100, E101, N102, and T104 in SEQ ID NO: 117; In one embodiment, the CsgF peptide comprises a CsgG-binding region and a region that forms a constriction within the pore. In one embodiment, the CsgF peptide is a truncated CsgF peptide lacking the C-terminal head domain of CsgF. In another embodiment, the CsgF peptide is a truncated CsgF peptide lacking a portion of the C-terminal head and neck domains of CsgF. In another embodiment, the CsgF peptide is a truncated CsgF peptide lacking the C-terminal head and neck domains of CsgF. The CsgG / CsgF pore is also referred to herein as a pore complex and an isolated pore complex. The isolated pore complex comprises a CsgG pore or a homologue or variant thereof, and a modified CsgF peptide or a homologue or variant thereof, particularly a truncated CsgF fragment or a homologue or variant thereof. In one embodiment, the modified CsgF peptide, homolog, or variant is located in the lumen of the CsgG pore or homolog or variant thereof. In another embodiment, the isolated pore complex has two or more channel constrictions, one located or provided by the CsgG pore formed by its constriction loop, and another additional channel constriction or leader head introduced by the modified CsgF peptide, or homolog, or variant thereof. In one embodiment, the CsgG pore or CsgG-like pore is a mutant CsgG pore rather than a wild-type pore, and in certain embodiments, mutations are present, for example, in the channel constriction loop. In other embodiments, mutations are alternatively or additionally present in the upper part of the pore, in the region where the pore interacts with the polynucleotide binding protein. Mutations may affect how the polynucleotide binding protein interacts with the pore and / or how the pore interacts with the polynucleotide binding protein. In another embodiment, the isolated pore complex comprising a modified CsgF peptide or a homologue or variant thereof has a CsgF channel constriction with a diameter in the range of 0.5 nm to 2.0 nm.In one embodiment, the pore complex comprises: (i) a CsgG pore comprising a first opening, a midsection comprising a beta-barrel, a second opening, and a lumen extending from the first opening through the midsection to the second opening, wherein the luminal surface of the midsection defines a CsgG constriction; and (ii) a plurality of modified CsgF peptides, each having a CsgF constriction region and a CsgF binding region (also referred to herein as a CsgG binding domain or region of CsgF), wherein the modified CsgF peptides form a CsgF constriction within the beta-barrel of the CsgG pore, and the CsgG constriction and the CsgF constriction are coaxially spaced apart within the beta-barrel of the CsgG pore. The luminal surface of the CsgG pore may comprise one or more loop regions of a CsgG monomer that define the CsgG constriction. The CsgF constriction region and the CsgF binding region typically correspond to the N-terminal portion of the CsgF mature peptide. In one embodiment, the pore complex excludes CsgA, CsgB, and CsgE.
[0009] One embodiment relates to a pore comprising a CsgG pore and a modified CsgF peptide, wherein the modified CsgF peptide is bound to CsgG and forms a constriction in the pore, and the pore is mutated to alter the interaction of the pore and a polynucleotide-binding enzyme, and / or the pore is mutated to improve the rate at which an analyte passes through the pore. In one embodiment of the invention, the rate at which an anite passes through the pore is increased. In another embodiment of the invention, the rate at which the analyte passes through the pore is decreased.
[0010] Another embodiment relates to an isolated pore complex in which a modified CsgF peptide and a CsgG pore or a monomer of the pore, or a homologue or variant thereof, are covalently coupled, and even more particularly, the coupling is via a cysteine residue or via a non-modified reactive or photoreactive amino acid in the CsgG monomer, or a homologue thereof, at a position corresponding to 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 117 or SEQ ID NO: 3.
[0011] The present invention also provides an isolated transmembrane pore or pore complex, or a membrane composition comprising an isolated pore or pore complex of the present invention and a membrane component, particularly the transmembrane pore or pore complex or membrane composition consisting of an isolated pore or pore complex of the present invention and a membrane or insulating layer component.
[0012] The present invention also provides the following: a membrane comprising a pore or pore complex according to the invention, an array comprising a plurality of membranes of the invention; - A system comprising: (a) a membrane of the invention or an array of the invention; (b) means for applying a potential across the membrane(s); and (c) means for detecting an electrical or optical signal across the membrane(s).
[0013] The present invention also provides a method for producing a transmembrane pore complex of the present invention, comprising co-expressing a CsgG pore, or a homologue or variant thereof, and a modified CsgF peptide, or a homologue or variant thereof, in a suitable host cell, thereby allowing transmembrane pore complex formation in vivo.
[0014] The present invention also provides a method for producing an isolated pore complex of the present invention, comprising contacting a CsgG monomer, or a homologue or variant thereof, with a modified CsgF peptide, or a homologue or variant thereof, thereby allowing in vitro reconstitution of the isolated pore complex. The modified CsgF peptide may contain an enzyme cleavage site at a suitable position in its amino acid sequence and may be cleaved before or after pore formation.
[0015] In certain embodiments, the modified CsgF peptide or homolog or variant thereof comprises SEQ ID NO: 12 or SEQ ID NO: 14, or a homolog or variant thereof. In certain embodiments, the modified CsgF peptide of the method comprises SEQ ID NO: 15 or SEQ ID NO: 16, or a homolog or variant thereof.
[0016] The present invention also provides a method for determining the presence, absence, or one or more characteristics of a target analyte, comprising: (i) contacting a target analyte with an isolated pore or isolated pore complex of the invention, or a transmembrane pore complex of the invention, such that the target analyte migrates into the pore channel; and (ii) obtaining one or more measurements of the analyte as it moves through the pore channel, thereby determining the presence, absence, or one or more characteristics of the analyte. In one embodiment, the analyte is a polynucleotide. In particular, the method alternatively using a polynucleotide as the analyte includes determining one or more characteristics selected from (i) the length of the analyte or polynucleotide, (ii) the identity of the analyte or polynucleotide, (iii) the sequence of the analyte or polynucleotide, (iv) the secondary structure of the analyte or polynucleotide, and (v) whether the analyte or polynucleotide is modified.
[0017] In another embodiment, the analyte is a protein, (poly)peptide, or peptide. In a further embodiment, the analyte is a polymer, oligosaccharide, polysaccharide, or small organic or inorganic compound, such as, but not limited to, pharmacologically active compounds, toxic compounds, and pollutants.
[0018] The present invention also provides a method for characterizing a polynucleotide or (poly)peptide using an isolated pore or isolated pore complex of the invention, or a transmembrane pore complex of the invention, in particular the CsgG pore or homologue or variant thereof comprising 6 to 10 CsgG monomers that form the CsgG pore channel.
[0019] The present invention also provides the use of an isolated pore or isolated pore complex of the present invention, or a transmembrane pore complex of the present invention, for determining the presence, absence, or one or more characteristics of a target analyte. Furthermore, the present invention also relates to a kit for characterizing a target analyte, comprising (a) the isolated pore or pore complex, and (b) a membrane component.
[0020] The present invention also provides the following: a method for modifying the rate at which a target analyte passes through a pore, the method comprising contacting the target analyte with an isolated pore or isolated pore complex of the invention, or with a transmembrane pore complex of the invention, such that the target analyte moves relative to or towards the pore complex; - a kit for characterizing a target analyte, comprising (a) an isolated pore or isolated pore complex of the present invention, and one or both of (b) a membrane component and (c) a polynucleotide binding protein; - a method for characterizing a target analyte, comprising: (a) contacting a target analyte with an isolated pore or isolated pore complex of the invention and a DNA-dependent ATPase (Dda) helicase of the invention, or a helicase construct of the invention, such that the helicase or construct controls the movement of the target analyte through the pore or pore complex; (b) obtaining one or more measurements as the polynucleotide translocates relative to the pore or pore complex, the measurements being indicative of one or more characteristics of the target analyte, thereby characterizing the target analyte; - a kit for characterizing a target analyte, comprising: (a) a DNA-dependent ATPase (Dda) helicase of the invention, or a helicase construct of the invention; and (b) an isolated CsgG pore or isolated pore complex of the invention; a device comprising a pore or pore complex of the invention inserted into an in vitro membrane, - A device manufactured by a method comprising: (i) obtaining an isolated pore or isolated pore complex of the present invention; and (ii) contacting the isolated pore or isolated pore complex with an in vitro membrane such that the pore is inserted into the in vitro membrane. DETAILED DESCRIPTION OF THE INVENTION
[0021] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. All publications, patents, and patent applications mentioned herein are hereby incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, the present specification is intended to supersede and / or supersede any such conflicting material.
[0022] While the present invention will be described with respect to particular embodiments and with reference to certain drawings, the present invention is not limited thereto, but rather only by the claims. Any reference signs in the claims should not be construed as limiting the scope thereof. Of course, it should be understood that not necessarily all aspects or advantages can be achieved in accordance with any particular embodiment of the invention. Thus, for example, one skilled in the art will recognize that the invention can be embodied or carried out in a manner that achieves or optimizes one advantage or group of advantages taught herein, without necessarily achieving other aspects or advantages that may be taught or suggested herein.
[0023] The present invention, both as to its organization and method of operation, together with its features and advantages, may best be understood by reference to the following detailed description when read in connection with the accompanying drawings. Aspects and advantages of the present invention will be apparent and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment, although they may. Similarly, in the description of exemplary embodiments of the present invention, it should be understood that various features of the invention may be grouped together in a single embodiment or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects may comprise fewer than all features of a single foregoing disclosed embodiment.
[0024] Additionally, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly indicates otherwise. Thus, for example, reference to a "polynucleotide" includes two or more polynucleotides, reference to a "polynucleotide-binding protein" includes two or more such proteins, reference to a "helicase" includes two or more helicases, reference to a "monomer" refers to two or more monomers, reference to a "pore" includes two or more pores, etc.
[0025] In all discussions herein, the standard single-letter codes for amino acids are used, as follows: alanine (A), arginine (R), asparagine (N), aspartic acid (D), cysteine (C), glutamic acid (E), glutamine (Q), glycine (G), histidine (H), isoleucine (I), leucine (L), lysine (K), methionine (M), phenylalanine (F), proline (P), serine (S), threonine (T), tryptophan (W), tyrosine (Y), and valine (V). Standard substitution notation is also used, i.e., Q42R means that Q at position 42 is substituted with R.
[0026] In paragraphs herein where different amino acids at a particular position are separated by a / symbol, the / symbol means "or." For example, Q87R / K means Q87R or Q87K. When different positions are separated by a / symbol, the / symbol means "and," e.g., Y51 / N55 means Y51 and N55.
[0027] definition The following terms or definitions are provided solely to aid in the understanding of the present invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one of ordinary skill in the art of the present invention. Skilled artisans should refer to the following references for definitions and technical terms, especially those in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 thed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016). The definitions provided herein should not be construed to be narrower than understood by one of ordinary skill in the art.
[0028] As used herein, "about" when referring to a measurable value, e.g., amount, temporal duration, etc., is meant to encompass variations of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% from the specified value, where such variations are appropriate for performing the disclosed methods.
[0029] As used herein, "nucleotide sequence," "DNA sequence," or "nucleic acid molecule(s)" refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. The term refers only to the primary structure of the molecule. Thus, the term includes double- and single-stranded DNA and RNA. As used herein, the term "nucleic acid" refers to a single- or double-stranded covalently linked nucleotide sequence in which the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. A polynucleotide can be composed of deoxyribonucleotide or ribonucleotide bases. Nucleic acids can be synthetically produced in vitro or isolated from natural sources. Nucleic acids can further include modified DNA or RNA, e.g., DNA or RNA that is methylated, or RNA that has undergone post-translational modifications, e.g., 5'-capping with 7-methylguanosine, 3'-processing such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNAs), such as hexitol nucleic acids (HNAs), cyclohexene nucleic acids (CeNAs), threose nucleic acids (TNAs), glycerol nucleic acids (GNAs), locked nucleic acids (LNAs), and peptide nucleic acids (PNAs). The size of a nucleic acid, also referred to herein as a "polynucleotide," is typically expressed as the number of base pairs (bp) for double-stranded polynucleotides or the number of nucleotides (nt) for single-stranded polynucleotides. 1000 bp or nt is equivalent to a kilobase (kb). Polynucleotides less than about 40 nucleotides in length are typically referred to as "oligonucleotides" and may include primers for use in manipulating DNA, such as via polymerase chain reaction (PCR).
[0030] As used herein, "gene" includes both the promoter region and coding sequence of a gene. It refers to both the genomic sequence (including possible introns) and the cDNA derived spliced messenger that is operably linked to the promoter sequence.
[0031] A "coding sequence" is a nucleotide sequence that is transcribed into mRNA and / or translated into a polypeptide when placed under the control of appropriate regulatory sequences. The boundaries of the coding sequence are determined by a translation start codon at the 5'-terminus and a translation stop codon at the 3'-terminus. A coding sequence can comprise, but is not limited to, mRNA, cDNA, recombinant nucleotide sequence, or genomic DNA, and can also contain introns under certain circumstances.
[0032] The term "amino acid" in the context of this disclosure is used in its broadest sense and is intended to include organic compounds containing an amine (NH) and a carboxyl (COOH) functional group, along with a side chain (e.g., an R group) specific to each amino acid. In some embodiments, amino acid refers to a naturally occurring L α-amino acid or residue. Commonly used one- and three-letter abbreviations for naturally occurring amino acids are used herein: A = Ala, C = Cys, D = Asp, E = Glu, F = Phe, G = Gly, H = His, I = Ile, K = Lys, L = Leu, M = Met, N = Asn, P = Pro, Q = Gln, R = Arg, S = Ser, T = Thr, V = Val, W = Trp, and Y = Tyr (Lehninger, A.L., (1975) Biochemistry, 2nd ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes chemically modified amino acids, such as D-amino acids, retro-inverso amino acids, and amino acid analogs, naturally occurring amino acids not normally incorporated into proteins, such as norleucine, and chemically synthesized compounds having properties known in the art to be characteristic of amino acids, such as β-amino acids. For example, analogs or mimetics of phenylalanine or proline that allow the same conformational restriction of peptide compounds as natural Phe or Pro are included within the definition of amino acid. Such analogs and mimetics are referred to herein as "functional equivalents" of the respective amino acids. Other examples of amino acids are listed by Roberts and Vellaccio, "The Peptides: Analysis, Synthesis, Biology," Gross and Meiehofer, eds., Vol. 5, p. 341, Academic Press, Inc., NY 1983, which is incorporated herein by reference.
[0033] The terms "protein," "polypeptide," and "peptide" are further used interchangeably herein to refer to a polymer of amino acid residues, and to variants and synthetic analogs thereof. Thus, these terms apply to amino acid polymers in which one or more amino acid residues are non-naturally occurring synthetic amino acids, such as chemical analogs of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. Polypeptides may also undergo maturation or post-translational modification processes, including, but not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, and the like. "Recombinant polypeptide" refers to a polypeptide made using recombinant techniques, e.g., via expression of a recombinant or synthetic polynucleotide. When a chimeric polypeptide or a biologically active portion thereof is recombinantly produced, it is also preferably substantially free of culture medium; e.g., culture medium represents less than about 20%, more preferably less than about 10%, and most preferably less than about 5% of the volume of a protein preparation. "Isolated" refers to material that is substantially or essentially free from components that normally accompany it in its native state. For example, as used herein, an "isolated polypeptide" refers to a polypeptide purified from a protein complex or CsgF peptide that has been removed from molecules that flank it in its naturally occurring state, e.g., molecules present in a production host that flank the polypeptide. An isolated CsgF peptide (optionally a truncated CsgF peptide) can be produced by amino acid chemical synthesis or by recombinant production. An isolated complex can be produced by in vitro reconstitution of components of the complex, e.g., the CsgG pore and CsgF peptide(s), following purification, or can be produced by recombinant co-expression.
[0034] "Ortholog" and "paralog" encompass evolutionary concepts used to describe the ancestral relationships of genes. Paralogs are genes within the same species that arose by duplication of an ancestral gene, while orthologs are genes from different organisms that arose by speciation and are also derived from a common ancestral gene.
[0035] Protein "homologues" include peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions compared to the unmodified or wild-type protein in question and that have similar biological and functional activity as the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which sequences are identical amino acid-for-amino acid over a window of comparison. Thus, "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a window of comparison, determining the number of positions at which an identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occurs in both sequences to obtain the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (i.e., the window size), and multiplying the result by 100 to obtain the percentage of sequence identity.
[0036] The term "CsgG pore" defines a pore comprising multiple CsgG monomers. Each CsgG monomer can be a wild-type homolog of E. coli CsgG, such as a wild-type monomer from E. coli (SEQ ID NO: 3), e.g., a monomer having any one of the amino acid sequences set forth in SEQ ID NOs: 68-88, or a variant of any of them (e.g., a variant of any one of SEQ ID NOs: 3, 117, and 68-88). Variant CsgG monomers may also be referred to as modified CsgG monomers or mutant CsgG monomers. Modifications or mutations in the variants include, but are not limited to, any one or more of the modifications disclosed herein, or combinations of such modifications.
[0037] For all aspects and embodiments of the invention, a CsgG homologue refers to a polypeptide having at least 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% complete sequence identity to the wild-type E. coli CsgG set forth in SEQ ID NO:117 or SEQ ID NO:3. A CsgG homologue also refers to a polypeptide that contains the PFAM domain PF03783, which is characteristic of CsgG-like proteins. A list of currently known CsgG homologues and CsgG architectures can be found at http: / / pfam.xfam.org / / family / PF03783. Similarly, a CsgG homologue polynucleotide can include a polynucleotide having at least 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity to the wild-type E. coli CsgG set forth in SEQ ID NO:1. Examples of homologues of CsgG shown in SEQ ID NO: 3 have the sequences shown in SEQ ID NOs: 68-88.
[0038] The term "modified CsgF peptide" or "CsgF peptide" defines a CsgF peptide that has been truncated from its C-terminus (e.g., an N-terminal fragment) and / or modified to include a cleavage site. The CsgF peptide may be a fragment of wild-type E. coli CsgF (SEQ ID NO: 5 or SEQ ID NO: 6) or a wild-type homologue of E. coli CsgF, such as a peptide comprising any one of the amino acid sequences set forth in SEQ ID NOs: 17-36, or a variant of any of these (e.g., modified to include a cleavage site).
[0039] For all aspects and embodiments of the present invention, a CsgF homolog refers to a polypeptide having at least 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity to the wild-type E. coli CsgF set forth in SEQ ID NO:6. In some embodiments, a CsgF homolog also refers to a polypeptide comprising the PFAM domain PF10614, characteristic of CsgF-like proteins. A list of currently known CsgF homologs and CsgF architectures can be found at http: / / pfam.xfam.org / / family / PF10614. Similarly, a CsgF homologous polynucleotide can include a polynucleotide having at least 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity to the wild-type E. coli CsgF set forth in SEQ ID NO:4. Examples of truncated regions of the CsgF homologue shown in SEQ ID NO: 6 have the sequences shown in SEQ ID NOs: 17-36.
[0040] The term "N-terminal portion of the CsgF mature peptide" refers to a peptide having an amino acid sequence corresponding to the first 60, 50, or 40 amino acid residues starting from the N-terminus of the CsgF mature peptide (without the signal sequence). The CsgF mature peptide can be wild-type or mutant (e.g., with one or more mutations).
[0041] Sequence identity may be relative to a fragment or portion of a full-length polynucleotide or polypeptide. Thus, a sequence may have only 50% overall sequence identity with a full-length reference sequence, but the sequence of a particular region, domain, or subunit may share 80%, 90%, or even 99% sequence identity with the reference sequence. Homology to the nucleic acid sequence of SEQ ID NO: 1 for a CsgG homolog or SEQ ID NO: 4 for a CsgF homolog, respectively, is not limited solely to sequence identity. Despite apparently low sequence identity, many nucleic acid sequences can exhibit significant biological homology to each other. Homologous nucleic acid sequences are considered to hybridize to each other under low stringency conditions (M.R. Green, J. Sambrook, 2012, Molecular Cloning: A Laboratory Manual, Fourth Edition, Books 1-3, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).
[0042] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is the gene most frequently observed in a population and is therefore an arbitrarily designed "normal" or "wild-type" form of the gene. In contrast, the terms "modified," "mutant," or "variant" refer to a gene or gene product that exhibits modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications, and / or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. Note that naturally occurring mutants can be isolated and are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) can be substituted with arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express the mutant monomers. Alternatively, they can be introduced by expressing mutant monomers in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e., non-naturally occurring) analogs of those amino acids. They can also be produced by naked ligation when mutant monomers are produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acids can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acids they replace. Alternatively, conservative substitutions can introduce another amino acid that is aromatic or aliphatic in place of an existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 major amino acids defined in Table 1 below.If the amino acids have similar polarities, this can also be determined by reference to the hydrophobicity scale for amino acid side chains in Table 2. Table 1 - Chemical properties of amino acids [Table 1] Table 2 - Hydrophobicity scale [Table 2]
[0043] Mutant or modified proteins, monomers, or peptides can also be chemically modified in any manner and at any site. Mutant or modified monomers or peptides are preferably chemically modified by conjugation of a molecule to one or more cysteines (cysteine ligation), one or more lysines, one or more unnatural amino acids, enzymatic modification of an epitope, or terminal modification. Suitable methods for performing such modifications are well known in the art. Mutants of modified proteins, monomers, or peptides can be chemically modified by conjugation of any molecule. For example, mutants of modified proteins, monomers, or peptides can be chemically modified by conjugation of a dye or fluorophore. In some embodiments, mutant or modified monomers or peptides are chemically modified with a molecular adaptor that facilitates interaction between a pore containing the monomer or peptide and a target nucleotide or target polynucleotide sequence. The molecular adaptor is preferably a cyclic molecule, a cyclodextrin, a species capable of hybridization, a DNA binder or interchelator, a peptide or peptide analog, a synthetic polymer, an aromatic planar molecule, a small positively charged molecule, or a small molecule capable of hydrogen bonding.
[0044] The presence of the adaptor improves the host-guest chemistry of the pore and the nucleotide or polynucleotide sequence, thereby improving the sequencing capability of the pore formed from the mutant monomer. The principles of host-guest chemistry are well known in the art. The adaptor affects the physical or chemical properties of the pore, improving its interaction with the nucleotide or polynucleotide sequence. The adaptor may alter the charge of the barrel or channel of the pore or may specifically interact with or bind to the nucleotide or polynucleotide sequence, thereby promoting its interaction with the pore. Thus, the modified CsgF peptides provided in the present disclosure can be coupled to enzymes or proteins, providing better accessibility of the protein or enzyme to the pore, which may facilitate specific applications of pore complexes containing the modified CsgF peptide.
[0045] In this context, a protein may also be a fusion protein, which refers specifically to a genetic fusion produced, for example, by recombinant DNA technology. A protein may also be conjugated, or "conjugated to," as used herein, which refers specifically to chemical and / or enzymatic conjugation resulting in a stable covalent bond.
[0046] When several polypeptides or protein monomers bind or interact with each other, the proteins may form a protein complex. "Bind" refers to any interaction, whether direct or indirect. A direct interaction implies contact between binding partners, for example, via a covalent bond or coupling. An indirect interaction refers to any interaction in which interacting partners interact in a complex of three or more compounds. The interaction may be fully indirect with the aid of one or more bridging molecules, or it may be partially indirect, where there is still direct contact between the partners that is stabilized by the additional interaction of one or more compounds. A "complex," as referred to in this disclosure, is defined as a group of two or more related proteins that may have different functions. The association between different polypeptides of a protein complex may be through non-covalent interactions, such as hydrophobic or ionic forces, or may be through covalent bonds or couplings, such as disulfide bridges or peptide bonds. Covalent "bonding" or "coupling" are used interchangeably herein and may also refer to "cysteine coupling" or "reactive or photoreactive amino acid coupling," respectively, which refer to bioconjugation between cysteines or (photo)reactive amino acids, which are chemical covalent bonds to form stable complexes. Examples of photoreactive amino acids include azidohomoalanine, homopropargylglycine, homoallelic glycine, p-acetyl-Phe, p-azido-Phe, p-propargyloxy-Phe, and p-benzoyl-Phe (Wang et al. 2012, in Protein Engineering, DOI: 10.5772 / 28719; Chin et al. 2002, Proc. Nat. Acad. Sci. USA 99(17); 11020-24).
[0047] A "biological pore" is a transmembrane protein structure that defines a channel or hole that allows the translocation of molecules and ions from one side of a membrane to the other. Translocation of ionic species through the pore can be driven by a potential difference applied to either side of the pore. A "nanopore" is a biological pore in which the minimum diameter of the channel through which molecules or ions pass is on the order of nanometers (10-9 nanometers). In some embodiments, the biological pore can be a transmembrane protein pore. The transmembrane protein structure of a biological pore can be monomeric or oligomeric in nature. Typically, the pore contains multiple polypeptide subunits arranged around a central axis, thereby forming a protein-lined channel that extends substantially perpendicular to the membrane in which the nanopore resides. The number of polypeptide subunits is not limited. Typically, the number of subunits is 5-30, and preferably 6-10. Alternatively, the number of subunits is not defined, as in the case of perfringolysin or related large membrane pores. The portion of the protein subunits within the nanopore that forms the protein-lined channel typically contains secondary structural motifs that may include one or more transmembrane β-barrel and / or α-helical segments.
[0048] The terms "pore," "pore complex," or "complex pore," as used interchangeably herein, refer to an oligomeric pore, e.g., at least a CsgG monomer (e.g., comprising one or more CsgG monomers, such as two or more CsgG monomers, three or more CsgG monomers, etc.) or a CsgG pore (composed of CsgG monomers) and a CsgF peptide (e.g., a modified or truncated CsgF peptide) associate in a complex to form a pore or nanopore. The pore complex of the present disclosure has the properties of a biological pore, i.e., it has a typical transmembrane protein structure. When the pore complex is provided in an environment with a membrane component, a membrane, a cell, or an insulating layer, the pore complex inserts into the membrane or insulating layer to form a "transmembrane pore complex."
[0049] The pores, pore complexes, transmembrane pores, or transmembrane pore complexes of the present disclosure are suitable for characterizing analytes. In some embodiments, the pores, pore complexes, transmembrane pores, or transmembrane pore complexes described herein can be used to sequence polynucleotide sequences, for example, because they can distinguish between different nucleotides with a high degree of sensitivity. A pore or pore complex can be isolated, substantially isolated, purified, or substantially purified. A pore or pore complex is "isolated" or purified if it is completely free of any other components, such as lipids or other pores, or other proteins with which it normally associates in its native state, such as CsgE, CsgA, or CsgB, or if it is sufficiently concentrated from the membrane compartment. A pore or pore complex is substantially isolated if it is mixed with a carrier or diluent that does not interfere with its intended use. For example, a pore or pore complex is substantially isolated or substantially purified if it exists in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of other components, such as triblock copolymers, lipids, or other pores. Alternatively, the pore complex of the present disclosure, when present in a membrane, may be a transmembrane pore or transmembrane pore complex. The present disclosure provides isolated pores and isolated pore complexes, including homo-oligomeric pores derived from CsgG containing the same mutant monomer, which may also contain mutant forms of the CsgG monomer as its homolog. Alternatively, isolated pores or isolated pore complexes are provided, including hetero-oligomeric CsgG pores, which may be CsgG pores composed of mutant and wild-type CsgG monomers, or different forms of CsgG variants, mutants, or homologs. An isolated pore complex typically comprises at least 7, at least 8, at least 9, or at least 10 CsgG monomers and one or more (modified) CsgF peptides, for example, 2, 3, 4, 5, 6, 7, 8, 9, or 10 CsgF peptides. The pore complex may comprise any ratio of CsG monomers to CsgF peptides. In one embodiment, the ratio of CsG monomers to CsgF peptides is 1:1.
[0050] As used interchangeably herein, the terms "constriction," "opening," "constriction region," "channel constriction," or "constriction site" refer to an opening defined by the luminal surface of a pore or pore complex that acts to allow the passage of ions and target molecules (e.g., but not limited to, polynucleotides or individual nucleotides) but not other non-target molecules through the pore or pore complex channel. In some embodiments, the constriction(s) are the narrowest opening(s) within the pore or pore complex. In this embodiment, the constriction(s) may serve to limit the passage of molecules through the pore. The size of the constriction is typically a critical factor in determining the suitability of a nanopore for nucleic acid sequencing applications. If the constriction is too small, the molecules to be sequenced will not be able to pass. However, to achieve the greatest effect on ion flow through the channel, the constriction should not be too large. For example, the constriction should not be wider than the solvent-accessible lateral diameter of the target analyte. Ideally, any constriction should be as close as possible to the lateral diameter of the passing analyte. For nucleic acid and nucleic acid base sequencing, preferred constriction diameters are in the nanometer range (10-meter range). Preferably, the diameter should be in the range of 0.5-2.0 nm, and typically, the diameter is in the range of 0.7-1.2 nm. The constriction in wild-type E. coli CsgG has a diameter of approximately 9 Å (0.9 nm). The CsgF constriction formed in a pore complex comprising a CsgG-like pore and a modified CsgF peptide, or a homolog or variant thereof, has a diameter in the range of 0.5-2 nm, or 0.7-1.2 nm, and is therefore suitable for nucleic acid sequencing.
[0051] When two or more constrictions are present and spaced apart, each constriction can simultaneously interact with or "read" a separate nucleotide in the nucleic acid strand. In this situation, the reduction in ion flow through the channel results from the combined flow restriction of all constrictions containing nucleotides. Thus, in some cases, double constrictions can lead to a composite current signal. In certain situations, the current readout of one constriction, or "leading head," may not be individually determinable when two such leading heads are present. The constriction of wild-type E. coli CsgG (SEQ ID NO: 3) is composed of two cyclic rings formed by the juxtaposition of tyrosine residues at position 51 (Tyr51) in adjacent protein monomers, and phenylalanine and asparagine residues at positions 56 and 55 (Phe56 and Asn55), respectively. The wild-type pore structure of CsgG has most often been reengineered via recombinant genetic techniques to expand, alter, or remove one of the two cyclic rings that make up the CsgG constriction (referred to herein as the "CsgG channel constriction"), leaving a well-defined, single leading head. The constriction motif of the CsgG oligomeric pore is located at amino acid residues 38-63 of the wild-type monomeric E. coli CsgG polypeptide, as shown in SEQ ID NO: 3. Considering this region, mutations at any of amino acid residues 50-53, 54-56, and 58-59, as well as the key positioning of the side chains of Tyr51, Asn55, and Phe56 within the channel of the wild-type CsgG structure, have been shown to be advantageous for modifying or altering the characteristics of the leading head. The present disclosure, which relates to a pore complex comprising a CsgG pore and a modified CsgF peptide, or a homologue or variant thereof, surprisingly adds another constriction (herein referred to as a "CsgF channel constriction") to the CsgG-containing pore complex, forming a suitable additional second leader head within the pore via complexation with the modified CsgF peptide, where the additional CsgF channel constriction or leader head is positioned adjacent to the constriction loop of the CsgG pore or mutated GcsG pore.The additional CsgF channel constriction or reader head is positioned approximately 10 nm or less, e.g., 5 nm or less, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 nm, from the constriction loop of the CsgG pore or mutated GcsG pore. The pore complexes or transmembrane pore complexes of the present disclosure include pore complexes with two reader heads, i.e., channel constrictions positioned to provide suitable separate reader heads without interfering with the accuracy of the other constriction channel reader head. Thus, the pore complex may comprise the wild-type CsgG pore, or a homologue thereof, together with the CsgG mutant pores WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2019 / 002893, WO2017 / 149318, WO2018 / 211241, WO2019 / 002893 (all of which are incorporated by reference in their entirety herein, each of which lists mutations to the wild-type CsgG pore that improve the properties of the pore), and a modified CsgF peptide, or a homologue or variant thereof, wherein the CsgF peptide has a separate constriction channel that forms the leader head.
[0052] Pores and Pore Complexes The present invention provides an isolated CsgG pore or a homologue or variant thereof, or an isolated pore complex comprising a CsgG pore or a homologue or variant thereof and a modified CsgF peptide or a homologue or variant thereof, wherein the CsgG pore comprises at least one mutant CsgG monomer comprising a modification at one or more of positions W97, Q100, E101, N102, and T104 in SEQ ID NO: 117. The CsgG pore may be the pore of SEQ ID NO: 3 or 117, or a homologue or variant thereof. The at least one mutant monomer preferably comprises a variant of SEQ ID NO: 117 comprising a modification at one or more of positions W97, Q100, E101, N102, and T104.
[0053] The present invention provides an isolated CsgG pore, or a homologue or variant thereof, wherein the CsgG pore comprises at least one mutant CsgG monomer comprising a modification at one or more of positions W97, Q100, E101, N102 and T104 in SEQ ID NO: 117. The CsgG pore may be the pore of SEQ ID NO: 3 or 117, or a homologue or variant thereof. The at least one mutant monomer preferably comprises a variant of SEQ ID NO: 117 comprising a modification at one or more of positions W97, Q100, E101, N102 and T104.
[0054] The present invention provides an isolated pore complex comprising a CsgG pore, or a homologue or variant thereof, and a modified CsgF peptide, or a homologue or variant thereof, wherein the CsgG pore comprises at least one mutant CsgG monomer comprising a modification at one or more of positions W97, Q100, E101, N102, and T104 in SEQ ID NO: 117. The CsgG pore may be the pore of SEQ ID NO: 3 or 117, or a homologue or variant thereof. The at least one mutant monomer preferably comprises a variant of SEQ ID NO: 117 comprising a modification at one or more of positions W97, Q100, E101, N102, and T104.
[0055] At least one mutant monomer or variant may include any number and combination of modifications at one or more of positions (a) W97, (b) Q100, (c) E101, (d) N102, and (e) T104 in SEQ ID NO: 117. For example, at least one mutant monomer may include any number and combination of modifications at one or more of positions (a) W97, (b) Q100, (c) E101, (d) N102, and (e) T104 in SEQ ID NO: 117. For example, at least one mutant monomer may include any of the following: (a);(b);(c);(d);(e);(a) and (b);(a) and (c);(a) and (d);(a) and (e);(b) and (c);(b) and (d);(b) and (e);(c) and (d);(c) and (e);(d) and (e);(a),(b) and (c);(a),(b) and (d);(a),(b) and (e);(a),(c) ) and (e); (a), (d) and (e); (b), (c) and (d); (b), (c) and (e); (b), (d) and (e); (c), (d) and (e); (a), (b), (c) and (d); (a), (b), (c) and (e); (a), (b), (c) and (e); (a), (b), (c), (d) and (e); (b), (c), (d) and (e); and (a), (b), (c), (d) and (e). At least one mutant monomer or variant preferably contains a modification at one or more positions (a) W97, (b) Q100, (c) E101, and (d) N102 in SEQ ID NO: 117, including all combinations of (a)-(d) listed above.
[0056] The modification at one or more of positions W97, Q100, E101, N102, and T104 in SEQ ID NO: 117 can be any of the modifications discussed in more detail below. The modification can be a deletion, such as a deletion of E101. The deletion of E101 increases the translocation rate through the pore (Example 1). The modification is preferably a substitution.
[0057] The W at position 97 is preferably substituted with R, H, K, A, V, I, L, M, F, Y, S, T, Q, D, E, N, C, P, or G. The W at position 97 is more preferably substituted with D, G, N, R, or S. These substitutions increase the rate at which the analyte passes through the pore / relative to the pore (Example 1). The W at position 97 is more preferably substituted with D or R. These substitutions increase the rate at which the analyte passes through the pore / relative to the pore, and increase the normalized rate distribution (Example 1). The W at position 97 is more preferably substituted with G, N, or S. These substitutions increase the rate at which the analyte passes through the pore / relative to the pore, and decrease the normalized rate distribution (Example 1).
[0058] Q at position 100 is preferably substituted with R, H, K, W, A, V, I, L, M, F, Y, T, N, or S. Q at position 100 is more preferably substituted with A, K, or S. These substitutions increase the rate at which an analyte passes through the pore / relative to the pore (Example 1). Q at position 100 is more preferably substituted with K (Q100K). This substitution increases the rate at which an analyte passes through the pore / relative to the pore, and increases the normalized rate distribution (Example 1). Q at position 100 is more preferably substituted with A or S (Q100A or Q100S). These substitutions increase the rate at which an analyte passes through the pore / relative to the pore, and decreases the normalized rate distribution (Example 1). Q at position 100 is most preferably substituted with A (Q100A). This substitution also has the effect shown in Example 4 when used with the modified Dda helicase of the present invention.
[0059] E at position 101 is preferably substituted with A, V, I, L, M, F, Y, or W. E at position 101 is more preferably substituted with A (E101A). This substitution decreases the rate at which the analyte passes through / relative to the pore, decreasing the normalized rate distribution (Example 1).
[0060] E at position 101 is preferably substituted with S, T, N, Q, C, G, or P. E at position 101 is more preferably substituted with G or S. These substitutions increase the rate at which an analyte passes through the pore / relative to the pore (Example 1). E at position 101 is more preferably substituted with G (E101G). This substitution increases the rate at which an analyte passes through the pore / relative to the pore, and decreases the normalized rate distribution (Example 1). E at position 101 is more preferably substituted with S (E101S). This substitution increases the rate at which an analyte passes through the pore / relative to the pore, and increases the normalized rate distribution (Example 1).
[0061] The N at position 102 is preferably substituted with D, E, R, H, K, S, T, Q, V, I, L, M, F, Y, W, or A. The N at position 102 is more preferably substituted with A, D, R, S, or W. These substitutions increase the rate at which the analyte passes through the pore / relative to the pore (Example 1). The N at position 102 is preferably substituted with A, R, or S. These substitutions increase the rate at which the analyte passes through the pore / relative to the pore, and decrease the normalized rate distribution (Example 1). The N at position 102 is preferably substituted with D or W (N102D or N102W). These substitutions increase the rate at which the analyte passes through the pore / relative to the pore, and increase the normalized rate distribution (Example 1). The N at position 102 is most preferably substituted with A or S (N102A or N102S). These substitutions also have the effect shown in Example 4 when used with the modified Dda helicases of the invention.
[0062] The T at position 104 is preferably substituted with R, H, or K.
[0063] As described in more detail below, the pore preferably comprises 6 to 10 monomers. Any number of these, for example 6, 7, 8, 9, or 10, may be mutant monomers comprising a modification at one or more of positions W97, Q100, E101, N102, and T104 in SEQ ID NO: 117, or may comprise a variant of SEQ ID NO: 117 comprising a modification at one or more of positions W97, Q100, E101, N102, and T104. All 6 to 10 monomers may comprise a modification at one or more of positions W97, Q100, E101, N102, and T104 in SEQ ID NO: 117, or may comprise a variant of SEQ ID NO: 117 comprising a modification at one or more of positions W97, Q100, E101, N102, and T104.
[0064] A mutant CsgG monomer is a monomer whose sequence differs from that of a wild-type CsgG monomer and which retains the ability to form a pore. Mutant monomers may also be referred to herein as variants. Methods for confirming the ability of mutant monomers to form a pore are well known in the art and are discussed in more detail below. At least one mutant monomer or variant may have any of the percentages of homology / sequence identity to SEQ ID NO: 117 or SEQ ID NO: 3 listed below. At least one mutant monomer may contain any of the additional modifications, mutations, or substitutions described below, including the types of modifications and substitutions described for the Dda helicases of the invention. At least one mutant monomer may contain any of the additional modifications, mutations, or substitutions described in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, and WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893 (all of which are incorporated by reference in their entirety).
[0065] The present invention surprisingly relates to a CsgG pore complexed with an optionally extracellularly positioned CsgF peptide, which introduces an additional channel constriction or leader head into the pore complex. Furthermore, the present disclosure provides location information for the constriction effected by the CsgF peptide within the pore complex, where the peptide is inserted into the lumen of the CsgG pore and the constriction site is in the N-terminal portion of the CsgF protein. Furthermore, modified or truncated CsgF peptides of the present disclosure have been shown to be sufficient for pore complex formation, providing means and methods for biosensing applications. The present disclosure includes wild-type and mutant CsgG pores (e.g., as disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, and International Patent Application No. PCT / GB2018 / 051191), or homologs or variants thereof, optionally in combination with modified or truncated CsgF peptides and mutants or homologs thereof, all of which together improve the ability of the CsgG pore or CsgG-like pore complex to interact with analytes such as polynucleotides. The additional constriction introduced into the CsgG-like nanopore channel by complexation with the (modified or truncated) CsgF peptide increases the contact surface with passing analytes and can act as a second reader head for analyte detection and characterization. Pores containing mutant CsgG monomers combined with novel mutant or modified forms of CsgF can improve the characterization of analytes such as polynucleotides, providing a more discriminatory and direct relationship between the current observed as the polynucleotide translocates through the pore. In particular, by having two stacked reader heads spaced a defined distance apart, the CsgG:CsgF pore complex can facilitate the characterization of polynucleotides containing at least one homopolymeric stretch, e.g., several consecutive copies of the same nucleotide, that would otherwise exceed the interaction length of a single CsgG leader head. Additionally, by having two stacked constrictions at a defined distance, small molecule analytes, including organic or inorganic drugs and pollutants, passing through the CsgG:CsgF complex pore pass sequentially through two independent reader heads.The chemistry of either reader head can be independently modified, each imparting unique interaction properties with analytes and thus providing additional discrimination during analyte detection.
[0066] The present invention relates to an isolated pore complex comprising a CsgG pore, or a homolog or mutant thereof, or a CsgG-like pore and a modified CsgF peptide, or a homolog or mutant thereof. Indeed, the present disclosure relates to a modified CsgG biological pore comprising a modified CsgF peptide, which may be a truncated, mutant, and / or variant thereof. In one embodiment, the interaction region between the modified CsgF peptide or a homolog or mutant thereof is located in the lumen of the CsgG pore or a homolog or mutant thereof. In another embodiment, the pore complex has two or more constriction sites or leader heads provided by at least one constriction of the CsgG pore and at least one introduced by a CsgF peptide that forms a complex with the CsgG pore. N-terminal CsgF positions including amino acid residues 39-64 of SEQ ID NO:5, or more specifically, the position ranging from amino acid residues 49-64 of SEQ ID NO:5, have been shown to enable detectable amounts of stable CsgG:CsgF complexes. In one embodiment, the CsgF constriction produced by a modified CsgF peptide (such as those described herein) is adjacent to or opposite a first constriction in the CsgG pore of the pore complex. For CsgG or CsgG-like protein pores, the constriction site has been determined to be formed by the loop region of the beta strand.
[0067] In one embodiment, a modified CsgF peptide refers to a truncated CsgF protein or fragment, including an N-terminal CsgF peptide fragment, where the modification specifically includes the constriction region and is defined by the constraint that it binds to a CsgG monomer, or a homolog or variant thereof. The modified CsgF peptide may additionally contain mutations or homologous sequences that may promote specific properties of the pore complex. In certain embodiments, the modified CsgF peptide comprises a truncation of the CsgF protein compared to the wild-type preprotein (SEQ ID NO: 5) or mature protein (SEQ ID NO: 6) sequence, or a homolog thereof. These modified peptides are intended to function as pore complex components that introduce an additional constriction site or leader head into the CsgG-like pore formed by CsgG and the modified or truncated CsgF peptide. Examples of truncated modified peptides are described below.
[0068] Examples of modified CsgF peptide homologs are disclosed in WO 2019 / 002893 (incorporated herein by reference in its entirety), which identifies CsgF-like proteins or CsgF peptides containing homologous or similar constriction regions in different bacterial strains, which may be useful for use in similar pore complexes. The structural features and CsgG-binding elements in CsgF peptides derived from various CsgF homologs are conserved so that the CsgF peptides can be used in combination with different wild-type or mutant CsgG pores. This includes complexes of the CsgG pore with non-cognate CsgFs, meaning that the CsgG pore and the parent CsgF homolog from which the CsgF is derived do not need to originate from the same operon, bacterial species, or strain.
[0069] In alternative embodiments, the CsgG pore in the pore complex is not a wild-type pore, but also contains mutations or modifications to enhance pore properties. The isolated pore complexes of the present disclosure formed by a CsgG pore or its homolog and a modified CsgF peptide or its homolog may be formed from the wild-type form of the CsgG pore, or may be further modified within the CsgG pore, such as by directed mutagenesis of specific amino acid residues, to further enhance the desired properties of the CsgG pore for use within the pore complex. For example, in embodiments of the present invention, mutations are contemplated to alter the number, size, shape, placement, or orientation of constrictions within the channel. Pore complexes containing modified mutant CsgG pores can be prepared by known genetic engineering techniques that result in the insertion, substitution, and / or deletion of specific targeted amino acid residues in the polypeptide sequence. In the case of oligomeric CsgG pores, mutations can be made in each monomeric polypeptide subunit, any one of the monomers, or all of the monomers. Suitably, in one embodiment of the present invention, the described mutations are made to all monomer polypeptides within the oligomeric protein structure. A mutant CsgG monomer is a monomer whose sequence differs from that of a wild-type CsgG monomer but retains the ability to form a pore. Methods for confirming the ability of mutant monomers to form a pore are well known in the art. The present disclosure includes wild-type and mutant CsgG pores (e.g., as disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, and International Patent Application No. PCT / GB2018 / 051191), or homologs thereof, in combination with modified or truncated CsgF peptides and their mutants or homologs, all of which together improve the ability of the CsgG-like pore complex to interact with an analyte, such as a polynucleotide. A mutant CsgG pore may comprise one or more mutant monomers. The CsgG pore may be a homopolymer, comprising identical monomers, or a heteropolymer, comprising two or more different monomers, which may have one or more of the mutations described below in any combination.
[0070] In certain embodiments, nanopore complexes comprising modified CsgF peptides differ from wild-type CsgF protein set forth in SEQ ID NO: 6 because the modified CsgF peptides comprise only an N-terminal fragment or truncation of the wild-type CsgF protein. However, the modified CsgF peptides may be additionally or alternatively mutated CsgF peptides, in that mutations, such as amino acid substitutions, are made to allow for a better second constriction site within the pore formed by the complex comprising the CsgF pore and the modified CsgF peptide. Thus, the mutant monomers may have improved polynucleotide reading properties when the complex is used in nucleotide sequencing, i.e., they may exhibit improved polynucleotide capture and nucleotide discrimination in addition to the improved characteristics of complexes comprising two reader heads. In particular, pores constructed from the mutant peptides capture nucleotides and polynucleotides more easily than the wild-type. Additionally, pores constructed from the mutant peptides may exhibit an increased current range, making it easier to distinguish between different nucleotides, and reduced state fluctuations, increasing the signal-to-noise ratio. In addition, as a polynucleotide translocates through a pore constructed from the mutant, the number of nucleotides contributing to the current may decrease. This facilitates identifying a direct relationship between the current observed as the polynucleotide translocates through the pore and the polynucleotide sequence. In addition, pores constructed from mutant peptides may exhibit increased throughput and are more likely to interact with analytes, such as polynucleotides, making it easier to characterize analytes using the pore. Pores constructed from mutant peptides may be more easily inserted into membranes or may provide an easier way to retain additional proteins in close proximity to the pore complex.
[0071] In an alternative embodiment, the CsgF constriction site provided in the pore complex of the invention has a diameter in the range of 0.5 nm to 2.0 nm, thereby providing a pore complex suitable for nucleic acid sequencing as described above.
[0072] The pore can be stabilized by covalent attachment of the CsgF peptide to the CsgG pore. The covalent attachment can be, for example, a disulfide bond or click chemistry. The CsgF peptide and the CsgG pore can be covalently attached, for example, via residues at positions corresponding to one or more of the following pairs of positions in SEQ ID NO:6 and SEQ ID NO:3 or SEQ ID NO:117, respectively: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144.
[0073] In the pore, the interaction between the CsgF peptide and the CsgG pore may be stabilized by hydrophobic or electrostatic interactions, for example, at positions corresponding to one or more of the following pairs of positions in SEQ ID NO: 6 and SEQ ID NO: 3 or SEQ ID NO: 117, respectively: 1 and 153, 4 and 133, 5 and 136, 8 and 187, 8 and 203, 9 and 203, 11 and 142, 11 and 201, 12 and 149, 12 and 203, 26 and 191, and 29 and 144.
[0074] Residues of CsgF and / or CsgG at one or more of the positions listed above may be modified to enhance the interaction between CsgG and CsgF in the pore.
[0075] In one embodiment, the pore of the present invention can be isolated, substantially isolated, purified, or substantially purified.The pore of the present invention is isolated or purified when it does not contain any other components such as lipids or other pores.The pore is substantially isolated when it is mixed with a carrier or diluent that does not interfere with its intended use.For example, the pore is substantially isolated or substantially purified when it exists in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of other components, such as triblock copolymers, lipids, or other pores.Alternatively, the pore of the present invention can be present in a membrane.Suitable membranes are discussed below.
[0076] The pores of the present invention may exist as individual or single pores, alternatively, the pores of the present invention may exist in a homogeneous or heterogeneous population of two or more pores.
[0077] CsgF peptide The isolated pore complexes of the present invention include modified CsgF monomers (peptides) or truncated CsgF proteins, or modified or truncated peptides of CsgF homologs or variants. These novel modified CsgF peptides can be used in pore complexes to integrate second or additional leader heads. The modification or truncation preferably results in a fragment, more preferably an N-terminal fragment, of wild-type CsgF or a mutant or homolog CsgF protein. The modified CsgF peptide of SEQ ID NO: 5, or a homolog or variant thereof, can be any of those disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893 (all of which are incorporated by reference in their entireties).
[0078] CsgF peptides forming part of the present invention are truncated CsgF peptides lacking the C-terminal head, or lacking a portion of the C-terminal head and neck domain of CsgF (e.g., a truncated CsgF peptide may include only a portion of the neck domain of CsgF), or lacking the C-terminal head and neck domain of CsgF. The CsgF peptide may lack a portion of the CsgF neck domain; for example, the CsgF peptide may include a portion of the neck domain from amino acid residue 36 at the N-terminus of the neck domain (see SEQ ID NO: 6) (e.g., residues 36-40, 36-41, 36-42, 36-43, 36-45, 36-46, residues 36-50, or 36-60 of SEQ ID NO: 6). The CsgF peptide preferably includes the CsgG-binding region and the region that forms the constriction within the pore. The CsgG-binding region typically comprises residues 1-8 and / or 29-32 of the CsgF protein (SEQ ID NO: 6 or a homologue from another species) and may contain one or more modifications. The region that forms the constriction within the pore typically comprises residues 9-28 of the CsgF protein (SEQ ID NO: 6 or a homologue from another species) and may contain one or more modifications. Residues 9-17 contain the conserved motif N9PXFGGXXX17 and form a turn region. Residues 9-28 form an alpha-helix. X17 (N17 in SEQ ID NO: 6) forms the apex of the constriction region, which corresponds to the narrowest part of the CsgF constriction within the pore. The CsgF constriction region also makes stabilizing contacts with the CsgG beta-barrel, primarily at residues 9, 11, 12, 18, 21, and 22 of SEQ ID NO: 6.
[0079] The CsgF peptide typically has a length of 28 to 50 amino acids, for example, 29 to 49, 30 to 45, or 32 to 40 amino acids. Preferably, the CsgF peptide contains 29 to 35 amino acids, or 29 to 45 amino acids. The CsgF peptide contains all or part of the FCP corresponding to residues 1 to 35 of SEQ ID NO: 6. When the CsgF peptide is shorter than the FCP, truncation is preferably at the C-terminus.
[0080] A CsgF fragment of SEQ ID NO: 6 or a homologue or variant thereof can have a length of 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, or 55 amino acids.
[0081] The CsgF peptide can comprise the amino acid sequence of SEQ ID NO: 6 from residue 1 to any one of residues 25-60, e.g., 27-50, e.g., 28-45 of SEQ ID NO: 6, or the corresponding residues from a homolog of SEQ ID NO: 6, or any variant thereof. More specifically, the CsgF peptide can comprise SEQ ID NO: 39 (residues 1-29 of SEQ ID NO: 6), or a homolog or variant thereof.
[0082] Examples of such CsgF peptides include, consist essentially of, or consist of SEQ ID NO:15 (residues 1-34 of SEQ ID NO:6), SEQ ID NO:54 (residues 1-30 of SEQ ID NO:6), SEQ ID NO:40 (residues 1-45 of SEQ ID NO:6), or SEQ ID NO:55 (residues 1-35 of SEQ ID NO:6), and any homologs or variants thereof. Other examples of CsgF peptides include, consist essentially of, or consist of SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO:10, SEQ ID NO:11, SEQ ID NO:12, SEQ ID NO:13, SEQ ID NO:14, SEQ ID NO:16.
[0083] In the CsgF peptide, for example, one or more residues in SEQ ID NO: 15, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 54, or SEQ ID NO: 55 can be modified.
[0084] For example, the CsgF peptide can include modifications at positions corresponding to one or more of the following positions in SEQ ID NO: 6: G1, T4, F5, R8, N9, N11, F12, A26, and Q29.
[0085] The CsgF peptide can be modified to introduce, for example, a cysteine, a hydrophobic amino acid, a charged amino acid, a non-denatured reactive amino acid, or a photoreactive amino acid at positions corresponding to one or more of the following positions in SEQ ID NO: 6: G1, T4, F5, R8, N9, N11, F12, A26, and Q29.
[0086] For example, a CsgF peptide may contain modifications at positions corresponding to one or more of the following positions in SEQ ID NO: 6: N15, N17, A20, N24, and A28. A CsgF peptide may contain a modification at position corresponding to D34 to stabilize the CsgG-CsgF complex. In certain embodiments, the CsgF peptide contains one or more of the following substitutions: N15S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, N17S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, A20S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, N24S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, A28S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, and D34F / Y / W / R / K / N / Q / C. The CsgF peptide can contain, for example, one or more of the following substitutions: G1C, T4C, N17S, and D34Y or D34N.
[0087] CsgF peptides can be produced by cleaving a longer protein, such as full-length CsgF, using an enzyme. Cleavage at a specific site can be directed by modifying a longer protein, such as full-length CsgF, to include an enzyme cleavage site at an appropriate position. Examples of CsgF amino acid sequences modified to include such enzyme cleavage sites are shown in SEQ ID NOS: 56-67. After cleavage, all or part of the added enzyme cleavage site can be present in the CsgF peptide that associates with CsgG to form a pore. Thus, the CsgF peptide can further include all or part of the enzyme cleavage site at its C-terminus.
[0088] Some examples of suitable CsgF peptides are shown in Table 3 below. Table 3: CsgF peptides [Table 3]
[0089] In a particular embodiment, the CsgF fragment comprises the amino acid sequence SEQ ID NO: 39, or a variant or homologue thereof. In particular, SEQ ID NO: 39 comprises the first 29 amino acids of the mature CsgF peptide (SEQ ID NO: 6). In another embodiment, the modified CsgF peptide of the invention is a truncated peptide comprising SEQ ID NO: 40. In particular, SEQ ID NO: 40 comprises the first 45 amino acids of the mature CsgF peptide (SEQ ID NO: 6). In particular, the CsgF constriction site and the binding site for CsgG are located within the N-terminal CsgF peptide region, and amino acids 39-64 of SEQ ID NO: 5 (present in SEQ ID NO: 39 and SEQ ID NO: 40), or in particular amino acids 49-64 of SEQ ID NO: 5 (present in SEQ ID NO: 40 but absent in SEQ ID NO: 39; the latter fragment encoded by SEQ ID NO: 39 exhibits a weaker interaction with CsgG (see Examples)), are further characterized as conferring greater stability to the complex. Thus, the present disclosure provides for modification of the CsgF protein by truncating the protein to the N-terminal fragment or to the peptide or peptides comprising the constriction site region to enable complex formation with the CsgG pore, or a homolog or variant thereof, in vivo. Further limitations are provided in one embodiment relating to modified CsgF peptides comprising SEQ ID NO: 37 or SEQ ID NO: 38. Finally, identification of CsgF homologous peptides aligned specifically within the constriction site (FCP peptide) also provides modified CsgF peptide homologs that can form part of the isolated complex.
[0090] A further embodiment relates to a modified or truncated CsgF peptide comprising SEQ ID NO: 15, which contains a region of the CsgF protein including several residues from the CsgG binding site and / or constriction region sufficient for in vitro reconstitution of a complex pore comprising CsgG or a homolog thereof and the modified CsgF peptide to yield an isolated pore complex comprising a CsgF channel constriction. Another embodiment describes the modified CsgF peptide comprising SEQ ID NO: 16, which contains an N-terminal fragment of the CsgF protein, and two additional amino acids (KD) that increase the solubility and stability of the (synthetic) peptide, also allowing for in vitro reconstitution of the complex pore. Further embodiments are provided in which the modified CsgF peptide comprises SEQ ID NO: 15, SEQ ID NO: 16, or a homolog or variant thereof, and the modified CsgF peptide is further mutated but still retains a minimum of 35% amino acid identity, for example, 40%, 50%, 60%, 70%, 80%, 85%, 90% amino acid identity to SEQ ID NO: 15 or SEQ ID NO: 16, respectively, within the region of the modified CsgF peptide corresponding to SEQ ID NO: 15 or 16. Further embodiments are provided in which the modified CsgF peptide comprises SEQ ID NO: 15, SEQ ID NO: 16, or a homolog or variant thereof, and the modified CsgF peptide is further mutated but still retains a minimum of 40%, 45%, 50%, 60%, 70%, 80%, 85%, or 90% amino acid identity to SEQ ID NO: 15 or SEQ ID NO: 16, respectively, within the region of the modified CsgF peptide corresponding to SEQ ID NO: 15 or 16. These mutated regions are intended to alter and / or improve the characteristics of the CsgF constriction site, as discussed above, thereby, for example, allowing for more accurate target analysis. Another embodiment discloses modified CsgF peptides, in which one or more positions within the region comprising SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:54, or SEQ ID NO:55 are modified, and the mutation(s) retain a minimum of 35% amino acid identity, or 40%, 50%, 60%, 70%, 80%, 85%, 90%, or 95% amino acid identity to SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:54, or SEQ ID NO:55 in a peptide fragment corresponding to the region comprising SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:54, or SEQ ID NO:55.
[0091] An additional embodiment relates to an isolated pore complex, in which the CsgG pore via at least one monomer and the modified CsgF peptide are coupled via a covalent bond. The covalent bond or bond can, in one example, be via a cysteine linkage, where the sulfhydryl side group of the cysteine is covalently bonded to another amino acid residue or moiety. In a second possibility, the covalent bond is obtained via interactions between non-denatured (photo)reactive amino acids. (Photo)reactive amino acids refer to artificial analogs of natural amino acids that can be used to crosslink protein complexes and can be incorporated into proteins and peptides in vivo or in vitro. Commonly used photoreactive amino acid analogs are photoreactive diazirine analogs of leucine and methionine, as well as parabenzoyl-phenyl-alanine, azidohomoalanine, homopropargylglycine, homoallelicglycine, p-acetyl-Phe, p-azido-Phe, p-propargyloxy-Phe, and p-benzoyl-Phe (Wang et al. 2012, Chin et al. 2002). Upon exposure to UV light, they become activated and covalently bind to interacting proteins within a few angstroms of the photoreactive amino acid analog. However, the location in the CsgG monomer at which this covalent binding can occur depends on exposure to the modified CsgF peptide. Some amino acids are at positions for providing a covalent bond, i.e., positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, or 209 of SEQ ID NO: 3 or SEQ ID NO: 117, or homologs thereof.
[0092] Another aspect of the present invention relates to a construct comprising the modified CsgF peptide, wherein the peptide is covalently linked. A "construct" includes two or more covalently linked monomers derived from modified CsgF and / or CsgG, or homologs thereof. In other words, a construct may contain more than one monomer. In another aspect, the present invention also provides a pore complex comprising at least one construct of the present invention. The pore complex contains sufficient constructs, and optionally monomers, to form a pore. For example, an octameric pore may contain (a) four constructs, each containing two monomers; (b) two constructs, each containing four monomers; (c) one construct containing two monomers and six monomers that do not form part of the construct; or (d) one construct having one or two CsgF monomers and six to seven CsgG monomers in one construct; or even (e) a construct having CsgF and CsgG monomers in addition to another construct containing only CsgG monomers. For example, a nonameric pore offers the same and additional possibilities. Other combinations of constructs and monomers can be envisioned by those skilled in the art. One or more constructs of the present invention can be used to form a pore complex for characterizing polynucleotides, such as sequencing. The construct can include at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten monomers. The construct preferably includes two monomers. The two or more monomers may be the same or different and may be CsgF, CsgG, a CsgG / CsgF fusion monomer, or a homolog thereof, or any combination thereof.
[0093] Another embodiment relates to a polynucleotide or nucleic acid molecule encoding the pore or pore complex of the invention, or a homologue or variant thereof, or a polynucleotide encoding the construct described above.
[0094] Certain embodiments relate to an isolated pore complex of the invention or an isolated transmembrane pore complex comprising an isolated pore complex and a membrane component. The isolated transmembrane pore complex is directly applicable for use in molecular sensing, such as nucleic acid sequencing. Alternatively, a membrane composition is provided comprising a modified CsgG / CsgF biological pore as described herein, according to the isolated pore complex of the invention, and a membrane, membrane component, or insulating layer. One embodiment relates to an isolated transmembrane pore complex consisting of an isolated pore complex of the invention and a membrane component.
[0095] The CsgG:CsgF complex is very stable, but when CsgF is truncated, the stability of the CsgG:CsgF complex decreases compared to a complex containing full-length CsgF. Therefore, to make the complex more stable, for example, a disulfide bond can be created between CsgG and CsgF after introducing a cysteine residue at the position identified herein. The pore complex can be created by any of the methods described above, and disulfide bond formation can be induced by using an oxidizing agent (e.g., copper-orthophenanthroline). Instead of cysteine interactions, other interactions (e.g., hydrophobic interactions, charge-charge interactions / electrostatic interactions) can also be used at those positions.
[0096] In another embodiment, unnatural amino acids can also be incorporated at these positions. In this embodiment, the covalent bond is created via click chemistry. For example, unnatural amino acids bearing azides or alkynes, or bearing dibenzocyclooctyne (DBCO) and / or bicyclo[6.1.0]nonyne (BCN) groups can be introduced at one or more of these positions.
[0097] Such stabilizing mutations can be combined with any other modifications to CsgG and / or CsgF, such as those disclosed herein.
[0098] The CsgG pore may comprise at least one, for example 2, 3, 4, 5, 6, 7, 8, 9, or 10, CsgG monomers that have been modified to facilitate binding to a CsgF peptide. For example, cysteine residues may be introduced at one or more of the positions corresponding to positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3 or SEQ ID NO: 117 to facilitate covalent binding to CsgG. Alternatively, or in addition to covalent binding via cysteine residues, the pore may be stabilized by hydrophobic or electrostatic interactions. To facilitate such interactions, non-denatured reactive or photoreactive amino acids at positions corresponding to one or more of positions 132, 133, 136, 138, 140, 142, 144, 145, 147, 149, 151, 153, 155, 183, 185, 187, 189, 191, 201, 203, 205, 207, and 209 of SEQ ID NO: 3 or SEQ ID NO: 117.
[0099] The CsgF peptide can be modified to facilitate binding to the CsgG pore. For example, cysteine residues can be introduced into one or more of the positions corresponding to positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO:6 to facilitate covalent binding to CsgG. Alternatively or in addition to covalent binding via cysteine residues, the pore can be stabilized by hydrophobic or electrostatic interactions. To facilitate such interactions, non-denatured reactive or photoreactive amino acids can be introduced into positions corresponding to one or more of positions 1, 4, 5, 8, 9, 11, 12, 26, or 29 of SEQ ID NO:6.
[0100] Preferred exemplary CsgF peptides include those containing the following mutations relative to SEQ ID NO:6: N15X1 / N17X2 / A20X3 / N24X4 / A28X5 / D34X6, where X1 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X2 is N / S / A / T / Q / G / L / V / I / F / Y / W / R / K / D / C, X3 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, X4 is N / S / T / Q / A / G / L / V / I / F / Y / W / R / K / D / C, X5 is A / S / T / Q / N / G / L / V / I / F / Y / W / R / K / D / C, and X6 is D / F / Y / W / R / K / N / Q / C. The mutations at positions N15, N17, A20, N24, and A28 are constriction mutations, and the mutation at position 34 affects the interaction pf CsgF with the bottom of the CsgG pore to stabilize the interaction.
[0101] CsgG pore The CsgG pore may be a homo-oligomeric pore comprising the same mutant monomers of the invention. The CsgG pore may be a hetero-oligomeric pore derived from CsgG, for example comprising at least one mutant monomer disclosed herein.
[0102] The CsgG pore may contain any number of mutant monomers. The pore typically contains at least 7, at least 8, at least 9, or at least 10 identical mutant monomers, for example 7, 8, 9, or 10 mutant monomers. The CsgG pore preferably contains 8 or 9 identical mutant monomers.
[0103] In a preferred embodiment, all of the monomers in the hetero-oligomeric CsgG pore (e.g., 10, 9, 8, or 7 of the monomers) are mutant monomers disclosed herein, at least one of which is different from the other monomers, and they may all be different from each other.
[0104] The mutant monomers within the CsgG pore are preferably all about the same length, or are the same length. The barrels of the mutant monomers of the invention within the pore are preferably about the same length, or are the same length. Length may be measured in number of amino acids and / or in length units.
[0105] The mutant monomer may be a variant of SEQ ID NO:3 or SEQ ID NO:117 that includes a modification at one or more of positions W97, Q100, E101, N102, and T104. Over the entire length of the amino acid sequence of SEQ ID NO:3 or SEQ ID NO:117, the variant is preferably at least 40% homologous to that sequence based on amino acid identity. More preferably, the variant may be at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of SEQ ID NO:3 or SEQ ID NO:117 over the entire sequence based on amino acid identity. Over the entire length of the amino acid sequence of SEQ ID NO:3 or SEQ ID NO:117, the variant is preferably at least 40% identical to that sequence. More preferably, variants may be at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% identical to SEQ ID NO: 3 or SEQ ID NO: 117 over the entire sequence. There may be at least 80%, for example at least 85%, 90%, or 95% amino acid identity ("hard homology") over a stretch of 100 or more, for example 125, 150, 175, or 200 or more contiguous amino acids.
[0106] CsgG monomers are highly conserved (as can be readily seen from Figures 45-47 of WO2017 / 149317). Furthermore, from knowledge of mutations relative to SEQ ID NO:3 or SEQ ID NO:117, it is possible to determine equivalent positions for mutations in CsgG monomers other than those of SEQ ID NO:3 or SEQ ID NO:117.
[0107] Thus, as described in the claims and elsewhere in the specification, reference to a mutant CsgG monomer comprising a variant of the sequence set forth in SEQ ID NO: 3 or SEQ ID NO: 117 and specific amino acid mutations thereof also encompasses mutant CsgG monomers comprising variants of the sequence set forth in SEQ ID NOs: 68-88 and their corresponding amino acid mutations. Similarly, as described in the claims and elsewhere in the specification, reference to a construct, pore, or method involving the use of a pore associated with a mutant CsgG monomer comprising a variant of the sequence set forth in SEQ ID NO: 3 or SEQ ID NO: 117 and specific amino acid mutations thereof also encompasses constructs, pores, or methods associated with mutant CsgG monomers comprising variants of the sequence according to the SEQ ID NOs disclosed above and their corresponding amino acid mutations. It will be further understood that the present invention extends to other variant CsgG monomers not expressly identified herein that exhibit highly conserved regions.
[0108] Standard methods in the art can be used to determine homology. For example, the UWGCG Package provides the BESTFIT program, which can be used to calculate homology, for example, using its default settings (Devereux et al. (1984) Nucleic Acids Research 12, p387-395). For example, the PILEUP and BLAST algorithms can be used to calculate homology or align sequences (identify equivalent residues or corresponding sequences, typically using their default settings), as described in Altschul SF (1993) J Mol Evol 36:290-300, Altschul, SF et al. (1990) J Mol Biol 215:403-10. Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ).
[0109] SEQ ID NO:3 is the wild-type CsgG monomer from Escherichia coli strain K-12 substrain MC4100. Variants of SEQ ID NO:3 or SEQ ID NO:117 may include any of the substitutions present in other CsgG homologs. Preferred CsgG homologs are set forth in SEQ ID NOs:68-88. Variants may include a combination of one or more of the substitutions present in SEQ ID NOs:68-88 compared to SEQ ID NO:3 or SEQ ID NO:117. For example, a mutation may be made at any one or more of the positions in SEQ ID NO:3 or SEQ ID NO:117 that differ between SEQ ID NO:3 or SEQ ID NO:117 and any one of SEQ ID NOs:68-88. Such a mutation may be a substitution of an amino acid in SEQ ID NO:3 or SEQ ID NO:117 with an amino acid from the corresponding position in any one of SEQ ID NOs:68-88. Alternatively, the mutation at any one of these positions may be a substitution with any amino acid, or may be a deletion or insertion mutation, such as a deletion or insertion of 1 to 10 amino acids, such as 2 to 8 or 3 to 6 amino acids. Other than the mutations disclosed herein, amino acids that are conserved between SEQ ID NO:3 or SEQ ID NO:117 and all of SEQ ID NOs:66-88 are preferably present in variants of the invention. However, conservative mutations may be made at any one or more of these positions that are conserved between SEQ ID NO:3 or SEQ ID NO:117 and all of SEQ ID NOs:66-88.
[0110] The present invention provides pore-forming CsgG mutant monomers that contain any one or more of the amino acids described herein substituted at a specific position in SEQ ID NO:3 or SEQ ID NO:117 at a position in the structure of the CsgG monomer that corresponds to the specific position in SEQ ID NO:3 or SEQ ID NO:117. Corresponding positions can be determined by standard techniques in the art. For example, the PILEUP and BLAST algorithms described above can be used to align the sequence of the CsgG monomer with SEQ ID NO:3 or SEQ ID NO:117 and thus identify corresponding residues.
[0111] Pore-forming mutant monomers typically retain the ability to form the same 3D structure as a wild-type CsgG monomer, such as the same 3D structure as a CsgG monomer having the sequence of SEQ ID NO: 3 or SEQ ID NO: 117. The 3D structure of CsgG is known in the art and is disclosed, for example, in Goyal et al (2014) Nature 516(7530):250-3. In addition to the mutations described herein, any number of mutations can be made in the wild-type CsgG sequence, provided that the CsgG mutant monomer retains the improved properties conferred by the mutations of the invention.
[0112] Typically, the CsgG monomer retains the ability to form a structure comprising three alpha helices and five beta sheets. Mutations can be made at least in the region of CsgG N-terminal to the first alpha helix (starting at S63 in SEQ ID NO:3), in the second alpha helix (G85 to A99 in SEQ ID NO:3 or SEQ ID NO:117), in the loop between the second alpha helix and the first beta sheet (Q100 to N120 in SEQ ID NO:3 or SEQ ID NO:117), in the fourth and fifth beta sheets (S173 to R192 and R198 to T107 in SEQ ID NO:3 or SEQ ID NO:117, respectively), and in the loop between the fourth and fifth beta sheets (F193 to Q197 in SEQ ID NO:3 or SEQ ID NO:117) without affecting the ability of the CsgG monomer to form a transmembrane pore through which the polypeptide can translocate. It is therefore contemplated that further mutations may be made in any of these regions in any CsgG monomer without affecting the ability of the monomer to form a pore through which a polynucleotide can be translocated. It is also contemplated that mutations may be made in other regions, for example, in any of the alpha helices (S63 to R76, G85 to A99, or V211 to L236 of SEQ ID NO: 3 or SEQ ID NO: 117) or in any of the beta sheets (I121 to N133, K135 to R142, I146 to R162, S173 to R192, or R198 to T107 of SEQ ID NO: 3 or SEQ ID NO: 117), without affecting the ability of the monomer to form a pore through which a polynucleotide can be translocated. It is also contemplated that deletion of one or more amino acids may be made in any of the loop regions connecting the alpha helices and beta sheets, and / or in the N- and / or C-terminal regions of the CsgG monomer, without affecting the ability of the monomer to form a pore through which a polynucleotide can be translocated.
[0113] In addition to those discussed above, amino acid substitutions may be made to the amino acid sequence of SEQ ID NO:3 or SEQ ID NO:117, for example, up to 1, 2, 3, 4, 5, 10, 20, or 30 substitutions. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side chain volume. The introduced amino acids may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acids they replace. Alternatively, conservative substitutions may introduce another amino acid that is aromatic or aliphatic in place of an existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 major amino acids defined in Table 1 above. If the amino acids have similar polarity, this can also be determined by referring to the hydrophobicity scale for amino acid side chains in Table 2.
[0114] One or more amino acid residues of the amino acid sequence of SEQ ID NO: 3 or SEQ ID NO: 117 may additionally be deleted from the polypeptides described above. Up to 1, 2, 3, 4, 5, 10, 20, 30 or more residues may be deleted.
[0115] Variants may include fragments of SEQ ID NO:3 or SEQ ID NO:117. Such fragments retain pore-forming activity. Fragments may be at least 50, at least 100, at least 150, at least 200, or at least 250 amino acids in length. Such fragments may be used to produce pores. Fragments preferably include the membrane-spanning domains of SEQ ID NO:3 or SEQ ID NO:117, i.e., K135 to Q153 and S183 to S208.
[0116] Alternatively or additionally, one or more amino acids may be added to the above-described polypeptides. The extension may be provided at the amino or carboxy terminus of the amino acid sequence of SEQ ID NO: 3 or SEQ ID NO: 117, or a polypeptide variant or fragment thereof. The extension may be very short, for example, 1 to 10 amino acids in length. Alternatively, the extension may be longer, for example, up to 50 or 100 amino acids. A carrier protein may be fused to the amino acid sequence according to the invention. Other fusion proteins are discussed in more detail below.
[0117] The CsgG pore described herein includes a wild-type CsgG pore or a homologue or mutant / variant thereof. A variant is a polypeptide having an amino acid sequence different from that of SEQ ID NO: 3 or SEQ ID NO: 117 and retaining its ability to form a pore. A variant typically contains the region of SEQ ID NO: 3 or SEQ ID NO: 117 involved in pore formation. The pore-forming ability of CsgG, which contains a β-barrel, is provided by the β-sheet in each subunit. A variant of SEQ ID NO: 3 or SEQ ID NO: 117 typically contains the regions of SEQ ID NO: 3 or SEQ ID NO: 117 that form β-sheets, i.e., K134 to Q154 and S183 to S208. One or more modifications can be made to the regions of SEQ ID NO: 3 or SEQ ID NO: 117 that form β-sheets, as long as the resulting variant retains its ability to form a pore. A variant of SEQ ID NO: 3 or SEQ ID NO: 117 preferably contains one or more modifications, such as substitutions, additions, or deletions, within its α-helical and / or loop regions.
[0118] The mutant CsgG monomer may be a mutant CsgG monomer whose sequence differs from that of a wild-type CsgG monomer and which retains the ability to form a pore. Mutant monomers may also be referred to herein as variants. Methods for determining the ability of mutant monomers to form pores are well known in the art and are discussed in more detail below.
[0119] At least one monomer, or any or all of the 6 to 10 monomers in a CsgG pore or pore complex of the invention may have any of the specific modifications or substitutions disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893 (all of which are incorporated by reference in their entirety).
[0120] Preferred additional modifications in at least one monomer / variant of SEQ ID NO: 117 in a pore or pore complex of the invention include, but are not limited to, one or more of the following, such as two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, or all of the following: (a) substitution at position Y51, e.g., Y51A, Y51I, Y51L, or Y51T; (b) a substitution at position N55, e.g., N55V, N55Q, N55S, or N55A; (c) a substitution at position F56, e.g., F56Q, F56A, or F56N; (d) a substitution at position L90, e.g., L90N, L90R, or L90K; (e) a substitution at position N91, e.g., N91N, N91R, or N91K; (f) a substitution at position K94, e.g., K94Q, K94R, K94F, K94Y, K94W, K94L, K94S, or K94N; (g) a substitution at position R192, e.g., R192D, R192Q, R192F, R192S, or R192T; (h) substitution at position Q153, e.g., Q153C, and (i) Substitution at position C215, for example, C215T.
[0121] At least one monomer / variant of SEQ ID NO: 117 may further comprise a deletion of one or more positions, for example a deletion of V105 to I107, a deletion of F193 to L199, or a deletion of F195 to L199.
[0122] Any number of the monomers in the pore or pore complex, for example 6, 7, 8, 9, or 10, can be mutant monomers / variants of SEQ ID NO: 117 that further comprise one or more of these additional modifications in addition to a modification at one or more of positions W97, Q100, E101, N102, and T104 in SEQ ID NO: 117. All 6 to 10 of the monomers in the pore or pore complex can be mutant monomers / variants of SEQ ID NO: 117 that further comprise one or more of these additional substitutions in addition to a modification at one or more of positions W97, Q100, E101, N102, and T104 in SEQ ID NO: 117.
[0123] Dual-pore complex The pore or pore complex of the present invention may be a dual pore complex comprising a first pore or complex and a second pore or complex. The dual pore complex may comprise a pore-pore, a pore-pore complex, or a pore complex-pore complex. Either of the pores or complexes may be a pore or complex of the present invention. In one embodiment, both the first pore complex and the second pore complex are CsgG / CsgF pore complexes of the present invention. In another embodiment, both the first pore and the second pore are CsgG pores of the present invention. The first and second pores or complexes may be the same or different. In addition to any of the mutations disclosed herein, in the dual pore, at least one CsgG monomer may comprise one or more of the additional mutations described in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, and WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893 (all of which are incorporated by reference in their entirety). At least one CsgG monomer preferably comprises any of the additional substitutions disclosed above.
[0124] Methods for producing modified proteins Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express mutant monomers. Alternatively, they can be introduced by expressing mutant monomers in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (i.e., non-naturally occurring) analogs of those amino acids. They can also be produced by naked ligation when mutant monomers are produced using partial peptide synthesis.
[0125] Monomers derived from CsgG may be modified to aid in their identification or purification, for example by the addition of a streptavidin tag or a signal sequence to facilitate their secretion from cells that do not naturally contain such a sequence. Other suitable tags are discussed in more detail below. The monomers may be labeled with a revealing label. The revealing label may be any suitable label that allows the monomer to be detected. Suitable labels are described below.
[0126] Monomers derived from CsgG can also be produced using D-amino acids, for example, they can contain a mixture of L- and D-amino acids, which is conventional in the art for producing such proteins or peptides.
[0127] Monomers derived from CsgG contain one or more specific modifications to facilitate nucleotide discrimination. Monomers derived from CsgG may also contain other non-specific modifications as long as they do not interfere with pore formation. Many non-specific side chain modifications are known in the art and can be made to the side chains of monomers derived from CsgG. Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde followed by reduction with NaBH4, amidination with methylacetimidate, or acylation with acetic anhydride.
[0128] Monomers derived from CsgG can be produced using standard methods known in the art. Monomers derived from CsgG can be made synthetically or by recombinant means. For example, monomers can be synthesized by in vitro translation and transcription (IVTT). Suitable methods for producing pores and monomers are discussed in International Application Nos. 2010 / 004273, 2010 / 004265, or 2010 / 086603 (incorporated herein by reference in their entireties). Methods for inserting pores into membranes are known.
[0129] Two or more CsgG monomers in the pore may be covalently linked to each other. For example, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 monomers may be covalently linked. The covalently linked monomers may be the same or different.
[0130] The monomers may be genetically fused, optionally via a linker, or chemically fused, for example via a chemical crosslinker. Methods for covalently linking monomers are disclosed in WO2017 / 149316, WO2017 / 149317, and WO2017 / 149318, which are incorporated herein by reference in their entireties.
[0131] In some embodiments, the mutant monomer is chemically modified. The mutant monomer can be chemically modified in any manner and at any site. The mutant monomer is preferably chemically modified by conjugation of a molecule to one or more cysteines (cysteine linkage), one or more lysines, one or more unnatural amino acids, enzymatic modification of an epitope, or terminal modification. Suitable methods for performing such modifications are well known in the art. The mutant monomer can be chemically modified by conjugation of any molecule. For example, the mutant monomer can be chemically modified by conjugation of a dye or fluorophore.
[0132] In some embodiments, the mutant monomer is chemically modified with a molecular adaptor that facilitates interaction between a pore containing the monomer and a target nucleotide or target polynucleotide sequence. The presence of the adaptor improves the host-guest chemistry of the pore and the nucleotide or polynucleotide sequence, thereby improving the sequencing capability of the pore formed from the mutant monomer. The principles of host-guest chemistry are well known in the art. The adaptor affects the physical or chemical properties of the pore, improving its interaction with the nucleotide or polynucleotide sequence. The adaptor may alter the charge of the pore barrel or channel, or may specifically interact with or bind to the nucleotide or polynucleotide sequence, thereby facilitating its interaction with the pore.
[0133] The molecular adaptor is preferably a cyclic molecule, a cyclodextrin, a species capable of hybridization, a DNA binder or interchelator, a peptide or peptide analog, a synthetic polymer, an aromatic planar molecule, a small positively charged molecule, or a small molecule capable of hydrogen bonding.
[0134] The adaptor may be cyclic. The cyclic adaptor preferably has the same symmetry as the pore. As CsgG typically has 8 or 9 subunits around a central axis, the adaptor preferably has 8-fold or 9-fold symmetry. This is discussed in more detail below.
[0135] Adapters typically interact with nucleotide or polynucleotide sequences through host-guest chemistry. Adapters are typically capable of interacting with nucleotide or polynucleotide sequences. Adapters include one or more chemical groups capable of interacting with nucleotide or polynucleotide sequences. The one or more chemical groups preferably interact with nucleotide or polynucleotide sequences through non-covalent interactions, such as hydrophobic interactions, hydrogen bonds, van der Waals forces, π-cation interactions, and / or electrostatic forces. The one or more chemical groups capable of interacting with nucleotide or polynucleotide sequences are preferably positively charged. More preferably, the one or more chemical groups capable of interacting with nucleotide or polynucleotide sequences include amino groups. The amino groups may be attached to primary, secondary, or tertiary carbon atoms. Even more preferably, the adapters include a ring of amino groups, such as a ring of 6, 7, or 8 amino groups. Most preferably, the adapters include a ring of 8 amino groups. The protonated amino group ring may interact with a negatively charged phosphate group in a nucleotide or polynucleotide sequence.
[0136] Correct positioning of the adaptor within the pore can be facilitated by host-guest chemistry between the adaptor and the pore containing the mutant monomer. The adaptor preferably contains one or more chemical groups capable of interacting with one or more amino acids in the pore. More preferably, the adaptor contains one or more chemical groups capable of interacting with one or more amino acids in the pore through non-covalent interactions such as hydrophobic interactions, hydrogen bonds, van der Waals forces, π-cation interactions, and / or electrostatic forces. The chemical groups capable of interacting with one or more amino acids in the pore are typically hydroxyl or amine. The hydroxyl group can be attached to a primary, secondary, or tertiary carbon atom. The hydroxyl group can form hydrogen bonds with uncharged amino acids in the pore. Any adaptor that facilitates interaction between the pore and a nucleotide or polynucleotide sequence can be used.
[0137] Suitable adaptors include, but are not limited to, cyclodextrins, cyclic peptides, and cucurbiturils. The adaptor is preferably a cyclodextrin or a derivative thereof. The cyclodextrin or derivative thereof may be any of those disclosed in Eliseev, AV, and Schneider, HJ. (1994) J. Am. Chem. Soc. 116, 6081-6088. The adaptor is more preferably heptakis-6-amino-β-cyclodextrin (am7-βCD), 6-monodeoxy-6-monoamino-β-cyclodextrin (am1-βCD), or heptakis-(6-deoxy-6-guanidino)-cyclodextrin (gu7-βCD). The guanidino group of gu7-βCD has a much higher pKa than the primary amine of am7-βCD and is therefore more positively charged. This gu7-βCD adapter can be used to increase the residence time of the nucleotide within the pore, increasing the precision of the measured residual current and increasing the rate of base detection at high temperatures or low data acquisition rates.
[0138] When succinimidyl 3-(2-pyridyldithio)propionate (SPDP) crosslinker is used as discussed in more detail below, the adapter is preferably heptakis(6-deoxy-6-amino)-6-N-mono(2-pyridyl)dithiopropanoyl-β-cyclodextrin (am6amPDP1-βCD).
[0139] More preferred adaptors include γ-cyclodextrin, which contains nine sugar units (and thus has nine-fold symmetry), which may contain a linker molecule or may be modified to include all or many of the modified sugar units used in the β-cyclodextrin example discussed above.
[0140] The molecular adaptor can be covalently linked to the mutant monomer. The adaptor can be covalently linked to the pore using any method known in the art. The adaptor is typically linked via a chemical bond. When the molecular adaptor is linked via a cysteine bond, one or more cysteines are preferably introduced into the mutant (e.g., into the barrel) by substitution. The mutant monomer can be chemically modified by linking a molecular adaptor to one or more cysteines in the mutant monomer. The one or more cysteines can be naturally occurring, i.e., at positions 1 and / or 215 in SEQ ID NO: 3 or SEQ ID NO: 117. Alternatively, the mutant monomer can be chemically modified by linking a molecule to one or more cysteines introduced at other positions. The cysteine at position 215 can be removed, for example, by substitution, to ensure that the molecular adaptor does not bind to that position, rather than to the cysteine at position 1 or a cysteine introduced at another position.
[0141] The reactivity of cysteine residues can be enhanced by modifying neighboring residues. For example, the basic group of an adjacent arginine, histidine, or lysine residue can change the pKa of the cysteine thiol group to that of the more reactive S-group. The reactivity of cysteine residues can be protected by thiol-protecting groups such as dTNB. These can be reacted with one or more cysteine residues of the mutant monomer before the linker is attached.
[0142] The molecule may be attached directly to the mutant monomer, or preferably, the molecule is attached to the mutant monomer using a linker, such as a chemical cross-linker or a peptide linker.
[0143] Suitable chemical cross-linkers are well known in the art. Preferred cross-linkers include 2,5-dioxopyrrolidin-1-yl 3-(pyridin-2-yldisulfanyl)propanoate, 2,5-dioxopyrrolidin-1-yl 4-(pyridin-2-yldisulfanyl)butanoate, and 2,5-dioxopyrrolidin-1-yl 8-(pyridin-2-yldisulfanyl)octananoate. The most preferred cross-linker is succinimidyl 3-(2-pyridyldithio)propionate (SPDP). Typically, the molecule is covalently bound to the bifunctional cross-linker before the molecule / cross-linker complex is covalently bound to the mutant monomer; however, it is also possible to covalently bind the bifunctional cross-linker to the monomer before the bifunctional cross-linker / monomer complex is bound to the molecule.
[0144] The linker is preferably dithiothreitol (DTT) resistant. Suitable linkers include, but are not limited to, iodoacetamide-based and maleimide-based linkers.
[0145] In other embodiments, the monomers can be attached to polynucleotide binding proteins, which form modular sequencing systems that can be used in the sequencing methods of the invention. Polynucleotide binding proteins are discussed below.
[0146] The polynucleotide-binding protein is preferably covalently linked to the mutant monomer. The protein can be covalently linked to the monomer using any method known in the art. The monomer and protein can be chemically fused or genetically fused. If the entire construct is expressed from a single polynucleotide sequence, the monomer and protein are genetically fused. Genetic fusion of a monomer to a polynucleotide-binding protein is discussed in WO2010 / 004265 (incorporated herein by reference in its entirety).
[0147] When the polynucleotide-binding protein is bound via a cysteine bond, one or more cysteines are preferably introduced into the mutant by substitution. The one or more cysteines are preferably introduced into loop regions with low conservation among homologs, indicating that mutations or insertions can be tolerated. They are therefore suitable for binding polynucleotide-binding proteins. In such embodiments, the naturally occurring cysteine at position 251 can be removed. The reactivity of the cysteine residue can be enhanced by the above-mentioned modifications.
[0148] The polynucleotide-binding protein can be attached to the mutant monomer directly or via one or more linkers. The molecule can be attached to the mutant monomer using a hybridization linker as described in WO 2010 / 086602 (incorporated herein by reference in its entirety). Alternatively, a peptide linker can be used. The peptide linker is an amino acid sequence. The length, flexibility, and hydrophilicity of the peptide linker are typically designed so that it does not interfere with the function of the monomer and molecule. Preferred flexible peptide linkers are stretches of 2 to 20, e.g., 4, 6, 8, 10, or 16, serine and / or glycine amino acids. More preferred flexible linkers include (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, and (SG)8, where S is serine and G is glycine. Preferred rigid linkers are stretches of 2 to 30, e.g., 4, 6, 8, 16, or 24, proline amino acids. More preferred rigid linkers include (P)12, where P is proline.
[0149] chemical modification The mutant CsgG monomer or CsgF peptide may be chemically modified with molecular adaptors and polynucleotide binding proteins.
[0150] The molecule (wherein the monomer or peptide is chemically modified) can be directly attached to the monomer or peptide or can be attached via a linker, as disclosed in WO2010 / 004273, WO2010 / 004265, or WO2010 / 086603 (which are incorporated by reference in their entirety).
[0151] Any of the proteins described herein, such as the CsgG monomer and / or CsgF peptide, can be modified to aid in their identification or purification, for example, by the addition of histidine residues (his tag), aspartic acid residues (asp tag), streptavidin tag, Flag tag, SUMO tag, GST tag, or MBP tag, or by the addition of a signal sequence to facilitate secretion of the polypeptide from cells that do not naturally contain such sequences. An alternative to introducing a genetic tag is to chemically react the tag onto a native or engineered position on the protein. One example of this would be reacting a gel shift reagent to a cysteine engineered onto the outside of the protein. This has been shown as a method for isolating hemolysin hetero-oligomers (Chem Biol. 1997 Jul;4(7):497-505).
[0152] Any of the proteins described herein, such as the CsgG monomer and / or CsgF peptide, can be labeled with a revealing label. The revealing label can be any suitable label that allows the protein to be detected. Suitable labels include, but are not limited to, fluorescent molecules, radioisotopes, such as I and S, enzymes, antibodies, antigens, polynucleotides, and ligands such as biotin.
[0153] Any of the proteins described herein, such as the CsgG monomer and / or CsgF peptide, can be made synthetically or by recombinant means. For example, the protein can be synthesized by in vitro translation and transcription (IVTT). The amino acid sequence of the protein can be modified to include non-naturally occurring amino acids or to increase the stability of the protein. When the protein is produced by synthetic means, such amino acids may be introduced during production. The protein can also be altered following either synthetic or recombinant production.
[0154] Proteins can also be produced using D-amino acids, for example, proteins can contain a mixture of L- and D-amino acids, which is conventional in the art for producing such proteins or peptides.
[0155] Proteins may also contain other non-specific modifications as long as they do not interfere with the function of the protein. Many non-specific side chain modifications are known in the art and can be made to the side chains of proteins (several). Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde followed by reduction with NaBH4, amidination with methylacetimidate, or acylation with acetic anhydride.
[0156] Any of the proteins described herein, such as the CsgG monomer and / or CsgF peptide, can be produced using standard methods known in the art. Polynucleotide sequences encoding the proteins can be obtained and replicated using standard methods in the art. Polynucleotide sequences encoding the proteins can be expressed in bacterial host cells using standard methods in the art. Proteins can be produced intracellularly by expressing the polypeptide in situ from a recombinant expression vector. The expression vector optionally carries an inducible promoter to control expression of the polypeptide. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.
[0157] Proteins can be produced on a large scale from protein-producing organisms following purification by any protein liquid chromatography system, or following recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA systems, Bio-Cad systems, Bio-Rad BioLogic systems, and Gilson HPLC systems.
[0158] Method for producing pores The present invention provides methods for producing a CsgG:modified CsgF pore complex that retains two or more constriction sites in vivo and in vitro. One embodiment provides a method for producing a transmembrane pore complex comprising a CsgG pore, or a homolog or mutant form thereof, and a modified CsgF peptide, or a homolog or mutant thereof, via co-expression. The method comprises expressing a CsgG monomer (expressed as the preprotein provided in SEQ ID NO: 2, or a homolog or mutant thereof) and expressing a modified or truncated CsgF monomer, both in a suitable host cell, to allow complex pore formation in vivo. The complex comprises a modified CsgF peptide in complex with the CsgG pore to provide an additional leader head for the pore. The resulting pore complex produced by the method using the modified CsgF peptide allows the passage of analytes, particularly polynucleotide chains, providing a structure sufficient for use of the pore complex in characterizing target analytes, such as nucleic acid sequencing, and includes two or more reader heads for improved reading of the polynucleotide sequence when used in an appropriate setting for the application. Methods for making the pores and complexes of the present invention, as well as methods for tagging them, are disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, and WO2017 / 149318, WO2018 / 211241, and WO2019 / 002893 (all of which are incorporated herein by reference in their entireties).
[0159] Methods for characterizing analytes The present invention provides methods for determining the presence, absence, or one or more characteristics of a target analyte. The methods involve contacting the target analyte with an isolated pore complex, such as a pore of the present invention, or a transmembrane pore, such that the target analyte migrates relative to the pore channel, e.g., into or through the pore channel, and obtaining one or more measurements as the analyte migrates relative to the pore, thereby determining the presence, absence, or one or more characteristics of the analyte. The target analyte may also be referred to as a template analyte or analyte of interest. The isolated pore complex typically contains at least seven, at least eight, at least nine, or at least ten monomers, e.g., seven, eight, nine, or ten CsgG monomers. The isolated pore complex preferably contains eight or nine identical CsgG monomers. One or more of the CsgG monomers, e.g., two, three, four, five, six, seven, eight, nine, or ten, are preferably chemically modified, or the CsgF peptide is chemically modified. The isolated pore complex monomers, such as CsgG monomers or homologs or variants thereof, and modified CsgF monomers or homologs or variants thereof, can be derived from any organism. The analyte can pass through the CsgG constriction followed by the CsgF constriction. In an alternative embodiment, the analyte can pass through the CsgF constriction followed by the CsgG constriction, depending on the orientation of the CsgG / CsgF complex within the membrane.
[0160] The method is for determining the presence, absence, or one or more characteristics of a target analyte. The method may be for determining the presence, absence, or one or more characteristics of at least one analyte. The method may relate to determining the presence, absence, or one or more characteristics of two or more analytes. The method may include determining the presence, absence, or one or more characteristics of any number of analytes, for example, 2, 5, 10, 15, 20, 30, 40, 50, 100 or more analytes. Any number of characteristics of one or more analytes may be determined, for example, 1, 2, 3, 4, 5, 10 or more characteristics.
[0161] The binding of molecules within the channel of the pore complex or near any opening of the channel affects the open channel ion flow through the pore, which is the essence of the "molecular sensing" of the pore channel. In a manner similar to nucleic acid sequencing applications, the fluctuation of the open channel ion flow can be measured using a suitable measurement technique by changing the current (e.g., WO2000 / 28312 and D. Stoddart et al., Proc. Natl. Acad. Sci., 2010, 106, 7702-7 or WO2009 / 077734; all of which are incorporated herein by reference in their entirety). The degree of reduction in ion flow, measured by the reduction in current, is related to the size of the obstacle within or near the pore. Thus, the binding of a molecule of interest, also referred to as an "analyte," within or near the pore provides a detectable and measurable event, thereby forming the basis of a "biological sensor." Molecules suitable for nanopore sensing include nucleic acids, proteins, peptides, polysaccharides, and small molecules (here referring to organic or inorganic compounds of low molecular weight (e.g., <900 Da or <500 Da)), such as pharmaceuticals, toxins, cytokines, and pollutants. Detecting the presence of biomolecules finds applications in personalized drug development, medicine, diagnostics, life science research, environmental monitoring, and the security and / or defense industries.
[0162] In another aspect, an isolated pore complex or transmembrane pore complex containing a wild-type or modified E. coli CsgG nanopore or a homolog or mutant thereof and a modified CsgF peptide, which provides a channel constriction to the pore within the complex, can function as a molecular or biological sensor. In some embodiments, the CsgG nanopore can be derived or isolated from a bacterial protein (e.g., E. coli, Salmonella typhi). In some embodiments, the CsgG nanopore can be recombinantly produced. Procedures for analyte detection are described in Howaka et al. Nature Biotechnology (2012) Jun 7;30(6):506-7. The analyte molecule to be detected can bind either to the surface of the channel or within the lumen of the channel itself. The location of binding can be determined by the size of the molecule being sensed.
[0163] The target analyte is preferably a metal ion, an inorganic salt, a polymer, an amino acid, a peptide, a polypeptide, a protein, a nucleotide, an oligonucleotide, a polynucleotide, a polysaccharide, a dye, a bleaching agent, a pharmaceutical, a diagnostic agent, a recreational drug, an explosive, a toxic compound, or an environmental pollutant. The method may involve determining the presence, absence, or one or more characteristics of two or more analytes of the same type, such as two or more proteins, two or more nucleotides, or two or more pharmaceuticals. Alternatively, the method may involve determining the presence, absence, or one or more characteristics of two or more different types of analytes, such as one or more proteins, one or more nucleotides, and one or more pharmaceuticals.
[0164] The target analyte may be secreted from the cell. Alternatively, the target analyte may be an analyte that is present inside the cell, and therefore must be extracted from the cell before the method can be performed.
[0165] Wild-type pores can act as sensors, but are often modified via recombinant or chemical methods to increase the strength of binding, the location of binding, or the specificity of binding of the molecule to be sensed. Typical modifications include the addition of a specific binding moiety complementary to the structure of the molecule to be sensed. When the analyte molecule comprises a nucleic acid, this binding moiety may comprise a cyclodextrin or an oligonucleotide. For small molecules, this may be the antigen-binding portion of an antibody or non-antibody molecule containing a known complementary binding region, for example, a single-chain variable fragment (scFv) region or antigen recognition domain from a T-cell receptor (TCR), or for proteins, it may be a known ligand of the target protein. In this way, wild-type or modified E. coli CsgG nanopores or homologs thereof can be endowed with the ability to act as molecular sensors for detecting the presence in a sample of suitable antigens (including epitopes), which may include receptors, cell surface antigens including markers of solid tumors or blood cancer cells (e.g., lymphoma or leukemia), viral antigens, bacterial antigens, protozoan antigens, allergens, allergy-related molecules, albumin (e.g., human, rodent, or bovine), fluorescent molecules (including fluorescein), blood group antigens, small molecules, drugs, enzymes, catalytic sites of enzymes or enzyme substrates, and transition-state analogs of enzyme substrates. As noted above, modifications can be achieved using known genetic engineering and recombinant DNA techniques. The location of any adaptations will depend on the properties of the molecule being sensed, such as its size, three-dimensional structure, and its biochemical properties. Selection of an adapted structure can utilize computational structural design. The determination and optimization of protein-protein or protein-small molecule interactions can be investigated using technologies such as BIAcore®, which uses surface plasmon resonance to detect molecular interactions (BIAcore, Inc., Piscataway, NJ; see also www.biacore.com).
[0166] In one embodiment, the analyte is an amino acid, peptide, polypeptide, or protein. The amino acid, peptide, polypeptide, or protein may be naturally occurring or non-naturally occurring. The polypeptide or protein may include synthetic or modified amino acids therein. Several different types of modifications to amino acids are known in the art. Suitable amino acids and their modifications are described above. It should be understood that the target analyte may be modified by any method available in the art.
[0167] In another embodiment, the analyte is a polynucleotide, such as a nucleic acid, which is defined as a macromolecule containing two or more nucleotides. Nucleic acids are particularly suitable for nanopore sequencing. Naturally occurring nucleic acid bases in DNA and RNA can be distinguished by their physical size. As a nucleic acid molecule or individual base passes through the nanopore channel, the size difference between the bases directly correlates to a reduction in ion flow through the channel. The fluctuations in ion flow can be recorded. Suitable electrical measurement techniques for recording the fluctuations in ion flow are discussed above. With suitable calibration, the characteristic reduction in ion flow can be used to identify specific nucleotides and related bases passing through the channel in real time. In typical nanopore nucleic acid sequencing, the open channel ion flow decreases as individual nucleotides of a nucleic acid sequence of interest pass sequentially through the nanopore channel due to partial blockage of the channel by the nucleotide. It is this reduction in ion flow that is measured using the suitable recording techniques described above. The reduction in ion flow can be calibrated to the reduction in ion flow measured for known nucleotides passing through the channel, which provides a means for determining which nucleotides pass through the channel. Thus, when performed sequentially, it provides a method for determining the nucleotide sequence of the nucleic acid passing through the nanopore. To accurately determine individual nucleotides, it is typically necessary that the reduction in ion flow through the channel be directly correlated to the size of the individual nucleotides passing through the constriction (or "leading head"). It will be understood that sequencing can be performed on intact nucleic acid polymers that are "threaded" through the pore, for example, through the action of the associated polymerase. Alternatively, the sequence can be determined by the passage of nucleotide triphosphate groups that are successively removed from the target nucleic acid adjacent to the pore (see, for example, WO2014 / 187924, the entire contents of which are incorporated herein by reference).
[0168] A polynucleotide or nucleic acid can contain any combination of any nucleotides. Nucleotides can be naturally occurring or artificial. One or more nucleotides in a polynucleotide can be oxidized or methylated. One or more nucleotides in a polynucleotide can be damaged. For example, a polynucleotide can contain pyrimidine dimers. Such dimers are typically associated with UV damage and are a major cause of skin melanoma. One or more nucleotides in a polynucleotide can be modified, for example, with a label or tag, suitable examples of which are known to those skilled in the art. A polynucleotide can contain one or more spacers. A nucleotide typically contains a nucleobase, a sugar, and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, and more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. Polynucleotides preferably contain the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC). Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically contain monophosphate, diphosphate, or triphosphate. Nucleotides may contain four or more phosphates, for example, four or five phosphates. The phosphates may be attached to the 5' or 3' side of the nucleotide. Nucleotides in a polynucleotide may be linked to each other in any manner. Nucleotides are typically linked by their sugar and phosphate groups, similar to nucleic acids. Nucleotides may be connected through their nucleobases, similar to pyrimidine dimers. Polynucleotides may be single-stranded or double-stranded. At least a portion of the polynucleotide is preferably double-stranded.The polynucleotide is most preferably ribonucleic acid (RNA) or deoxyribonucleic acid (DNA). In particular, the method alternatively using a polynucleotide as an analyte includes determining one or more characteristics selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether the polynucleotide is modified.
[0169] The polynucleotide can be of any length (i). For example, the polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs in length. The polynucleotide can be 1,000 nucleotides or nucleotide pairs or more, 5,000 nucleotides or nucleotide pairs or more, or 100,000 nucleotides or nucleotide pairs or more in length. Any number of polynucleotides can be investigated. For example, the method can involve characterizing 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100, or more polynucleotides. When two or more polynucleotides are characterized, they can be different polynucleotides or two instances of the same polynucleotide. The polynucleotide can be naturally occurring or artificial. For example, the method can be used to verify the sequence of a manufactured oligonucleotide. The method is typically performed in vitro.
[0170] The nucleotide may have any identity (ii), including, but not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotide is preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP. A nucleotide may be abasic (i.e., lacking a nucleobase). A nucleotide may also lack a nucleobase and a sugar (i.e., a C3 spacer). The sequence (iii) of a nucleotide is determined by the sequential identities of the following nucleotides, linked together in the 5' to 3' direction of the strand, throughout the polynucleotide strand:
[0171] CsgG pores and pores containing CsgF peptides are particularly useful for analyzing homopolymers. For example, pores can be used to determine the sequence of polynucleotides that contain two or more identical, e.g., at least 3, 4, 5, 6, 7, 8, 9, or 10 consecutive nucleotides. For example, pores can be used to sequence polynucleotides that contain poly-A, poly-T, poly-G, and / or poly-C regions.
[0172] The CsgG pore constriction consists of residues 51, 55, and 56 of SEQ ID NO:3 or SEQ ID NO:117. The reader head of CsgG and its constriction variants are generally sharp. When DNA passes through the constriction, the interaction of approximately five bases of DNA with the pore's reader head at any given time dominates the current signal. These sharper reader heads are very good at reading mixed-sequence regions of DNA (where A, T, G, and C are mixed), but when there are homopolymeric regions within the DNA (e.g., poly-T, poly-G, poly-A, poly-C), the signal flattens and lacks information. Because five bases dominate the signal for CsgG and its constriction variants, it is difficult to distinguish photopolymers longer than five without using additional dwell time information. However, when DNA passes through a second reader head, more DNA bases interact with the combined reader head, increasing the length of homopolymers that can be distinguished.
[0173] Modified Dda helicase The present invention provides modified Dda helicases. One or more specific modifications are discussed in more detail below. Modifications according to the present invention include one or more substitutions as discussed below.
[0174] one or more of positions 55, 114, 156, 177, 210, 221, 350, and 358 of Dda1993 The present invention provides modified DNA-dependent ATPase (Dda) helicases in which one or more of the positions corresponding to amino acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993 have been modified or substituted. Dda helicases and the positions corresponding to positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993 are discussed in more detail below. Positions 55, 114, 156, and 177 are in the 1A domain of Dda1993. Positions 210 and 221 are in the 2A domain of Dda1993. Positions 350 and 358 are in the tower domain of Dda1993. The modified Dda helicases of the present invention may be any of the following: (a);(b);(c);(d);(e);(f);(g);(h);(a) and (b);(a) and (c);(a) and (d);(a) and (e);(a) and (f);(a) and (g);(a) and (h);(b) and (c);(b) and (d);(b) and (e);(b and (f);(b and (g);(b and (h);(c and (d);(c ) and (e);(c) and (f);(c) and (g);(c) and (h);(d) and (e);(d) and (f);(d) and (g);(d) and (h);(e) and (f);(e) and (g);(e) and (h);(f) and (g);(f) and (h);(g) and (h);(a), (b), and (c);(a), (b), and (d);(a), (b), and (e);(a), (b), and ( f);(a), (b), and (g);(a), (b), and (h);(a), (c), and (d);(a), (c), and (e);(a), (c), and (f);(a), (c), and (g);(a), (c), and (h);(a), (d), and (e);(a), (d), and (f);(a), (d), and (g);(a), (d), and (h);(a), (e), and (f);(a), ( e), and (g);(a), (e), and (h);(a), (f), and (g);(a), (f), and (h);(a), (g), and (h);(b), (c), and (d);(b), (c), and (e);(b), (c), and (f);(b), (c), and (g);(b), (c), and (h);(b), (d), and (e);(b), (d), and (f);(b), (d), and (g);(b), (d), and (h);(b), (e), and (f);(b), (e), and (g);(b), (e), and (h);(b), (f), and (g);(b), (f), and (h);(b), (g), and (h);(c), (d), and (e);(c), (d), and (f);(c), (d), and (g);(c), (d), and (h);(c), (e), and (f);(c), (e), and (g);(c), (f), and (g);(c), (f), and (h);(c), (g), and (h);(d), (e), and and (f);(d), (e), and (g);(d), (e), and (h);(d), (f), and (g);(d), (f), and (h);(d), (g), and (h);(e), (f), and (g);(e), (f), and (h);(e), (g), and (h);(f), (g), and (h);(a), (b), (c), and (d);(a), (b), (c), and (e);(a), (b), (c), and (f);(a), (b), (c), and (g);(a), (b), (c), and (h);(a), (b), (d), and (e);(a), (b), ( d), and (f);(a), (b), (d), and (g);(a), (b), (d), and (h);(a), (b), (e), and (f);(a), (b), (e), and (g);(a), (b), (e), and (h);(a), (b), (f), and (g);(a), (b), (f), and (h);(a), (b), (g), and (h);(a), (c), (d), and (e);(a), (c), (d), and (f);(a), (c), (d), and (g);(a), (c), (d), and (h);(a), (c), (e), and (f);(a), (c), (e), and (g);(a), (c), (e), and (h);(a), (c), (f), and (g);(a), (c), (f), and (h);(a), (c), (g), and (h);(a), (d), (e), and (f);(a), (d), (e), and (g);(a), (d), (e), and (h);(a), (d), (f), and (g);(a), (d), (f), and (h);(a), (d), (g), and (h);(a), (e), (f), and (g);(a), (e), (f), and (g);(a), (e), (f), and (h);(a), (e), (f), and (h);(a), (f), (g), and (h);(b), (c), (d), and (e);(b), (c), (d), and (f);(b), (c), (d), and (g);(b), (c), (d), and (h);(b), (c), (e), and (f);(b), (c), (e), and (g);(b), (c), (e), and (h);(b), (c), (f), and (g);(b), (c), (f), and (h);(b), (c), (g), and (h);(b), (d), (e), and (f);(b), (d), (e), and (g);(b), (d), (e), and (h) ;(b),(d),(f), and(g);(b),(d),(f), and(h);(b),(d),(g), and(h);(b),(e),(f), and(g);(b),(e),(f), and(h);(b),(e),(g), and(h);(b),(f),(g), and(h);(c),(d),(e), and(f);(c),(d),(e), and(g);(c),(d),(f), and(g);(c),(d),(f), and(h);(c),(d),(f), and(h);(c),(d),(f), and (g);(c), (e), (f), and (h);(c), (e), (g), and (h);(c), (f), (g), and (h);(d), (e), (f), and (g);(d), (e), (f), and (h);(d), (e), (g), and (h);(d), (f), (g), and (h);(e), (f), (g), and (h);(a), (b), (c), (d), and (e);(a), (b), (c), (d), and (f);(a), (b), (c), (d), and (g);(a), (b), (c), (d), and (h);(a), (b), (c), (e ), and (f);(a), (b), (c), (e), and (g);(a), (b), (c), (e), and (h);(a), (b), (c), (f), and (g);(a), (b), (c), (f), and (h);(a), (b), (c), (g), and (h);(a), (b), (d), (e), and (f);(a), (b), (d), (e), and (g);(a), (b), (d), (e), and (h);(a), (b), (d), (f), and (g);(a), (b), (d), (f), and (h);(a), (b), (d), (f), and (g);(a), (b), (d), (f), and (h);(a), (b), (d), (f), and (h);(a), (b), (e), (f), and (g);(a), (b), (e), (f), and (h);(a), (b), (e), (g), and (h);(a), (b), (f), (g), and (h);(a), (c), (d), (e), and (f);(a), (c), (d), (e), and (g);(a), (c), (d), (e), and (h);(a), (c), (d), (f), and (g);(a), (c), (d), (f), and (h);(a), (c), (d), (f), and (h);(a), (c), (d), (f), and (h);(a), (c), (e), (f), and (g);(a), (c), (e), (f), and (h);(a), (c), (e), (g), and (h);(a), (c), (f), (g), and (h);(a), (d), (e), (f), and (g);(a), (d), (e), (f), and (h);(a), (d), (e), (g), and (h);(a), (d), (f), (g), and (h);(a), (e), (f), (g), and (h);(b), (c), (d), (e), and (f);(b), (c), (d), (e), and (g);(b), (c), (d), (f), and (g);(b), (c), (d), (f), and (h);(b), (c), (d), (g), and (h);(b), (c), (e), (f), and (g);(b), (c), (e), (f), and (h);(b), (c), (e), (g), and (h);(b), (c), (f), (g), and (h);(b), (d), (e), (f), and (g);(b), (d), (e), (f), and (h);(b), (d), (e), (g), and (h);(b), (d), (f), (g), and (h);(b), (e), (f), (g), and (h);(c), (d), (e), (f), and (g);(c), (d), (e), (f), and (h);(c), (d), (e), (g), and (h);(c), (d), (f), (g), and (h);(c), (e), (f), (g), and (h);(d), (e), (f), (g), and (h);(a), (b), (c), (d), (e), and (f);(a), (b), (c), (d), (e), and (g);(a), (b), (c), (d), (e), and (h);(a), (b), (c), (d), (f), and (g);(a), (b), (c), (d), (f), and (h);(a), (b), (c), (d), (g), and (h);(a), (b), (c), (e), (f), and (g);(a), (b), (c), (e), (f), and (h);(a), (b), (c), (e), (g), and (h);(a), (b), (c), (f), (g), and (h);(a), (b), (d), (e), (f), and (g);(a), (b), (d), (e), (f), and (h);(a), (b), (d), (e), (f), and (h);(a), (b), (d), (e), (f), and (h);(a), (b), (d), (f) , (g), and (h);(a), (b), (e), (f), (g), and (h);(a), (c), (d), (e), (f), and (g);(a), (c), (d), (e), (f), and (h);(a), (c), (d), (e), (g), and (h);(a), (c), (d), (e), (g), and (h);(a), (c), (e), (f), (g), and (h);(a), (d), (e), (f), (g), and (h);(b), (c), (d), (e), (f), and (g);(b), (c), (d), (e), (f), and (h);(b), ( c), (d), (e), (g), and (h);(b), (c), (d), (f), (g), and (h);(b), (c), (e), (f), (g), and (h);(b), (d), (e), (f), (g), and (h);(c), (d), (e), (f), (g), and (h);(a), (b), (c), (d), (e), (f), and (g);(a), (b), (c), (d), (e), (f), and (h);(a), (b), (c), (d), (e), (g), and (h);(a), (b), (c), (d), (e), (g), and (h);(a), (b), (c), (e), (f), (g), and (h); (a), (b), (d), (e), (f), (g), and (h); (a), (c), (d), (e), (f), (g), and (h); (b), (c), (d), (e), (f), (g), and (h); or modifications or substitutions at any number and combination of positions corresponding to amino acid positions (a) 55, (b) 114, (c) 156, (d) 177, (e) 210, (f) 221, (g) 350, and (h) 358, including (a), (b), (c), (d), (e), (f), (g), and (h);
[0175] The present invention provides modified DNA-dependent ATPase (Dda) helicases in which one or more of the positions corresponding to amino acid positions 114, 177, 350, and 358 in Dda1993 have been modified or substituted. Dda helicases and the positions corresponding to positions 114, 177, 350, and 358 in Dda1993 are discussed in more detail below. Positions 114 and 177 are in the 1A domain of Dda1993. Positions Y350 and K358 are in the tower domain of Dda1993. The modified Dda helicases of the present invention may comprise modifications or substitutions at any number and combination of positions corresponding to amino acid positions (a) 114, (b) 177, (c) 350, and (d) 358 in Dda1993, including (a); (b); (c); (d); (a) and (b); (a) and (c); (a) and (d); (b) and (c); (b) and (d); (c) and (d); (a), (b), and (c); (a), (b), and (d); (b), (c), and (d); or (a), (b), (c), and (d).
[0176] The position corresponding to amino acid position 55 in Dda1993 is preferably substituted with D, E, K, N, or S. The position corresponding to amino acid position 114 in Dda1993 is preferably substituted with K (T55K). These substitutions increase speed and precision when used to characterize polynucleotide analytes (Example 5). These substitutions also decrease the normalized rate distribution when used to characterize polynucleotide analytes (Example 5).
[0177] The position corresponding to amino acid position 114 in Dda1993 is preferably substituted with A, V, I, L, M, F, Y, W, G, P, S, T, N, or Q. The position corresponding to amino acid position 114 in Dda1993 is preferably substituted with A, G, I, L, M, P, S, T, or V. The position corresponding to amino acid position 114 in Dda1993 is preferably substituted with G, L, S, or T. These substitutions decrease the speed when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 114 in Dda1993 is preferably substituted with A, I, M, P, or V. These substitutions increase the speed when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 114 in Dda1993 is preferably substituted with G (C11G). This substitution decreases the speed and increases the precision when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 114 in Dda1993 is preferably substituted with I or P. These substitutions increase the speed and decrease the precision when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 114 in Dda1993 is preferably substituted with G, I, or P. These substitutions decrease the normalized speed distribution when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 114 in Dda1993 is most preferably substituted with I (C114I).
[0178] The position corresponding to amino acid position 156 in Dda1993 is preferably substituted with A, E, F, G, I, L, M, P, S, V, Y, D, K, or N. The position corresponding to amino acid position 156 in Dda1993 is preferably substituted with F (T156F). This substitution increases speed and increases precision when used to characterize polynucleotide analytes (Example 5). This substitution also decreases the normalized rate distribution when used to characterize polynucleotide analytes (Example 5).
[0179] The position corresponding to amino acid position 177 in Dda1993 is preferably substituted with D, E, F, G, H, I, L, M, N, Q, R, S, T, V, W, or Y. The position corresponding to amino acid position 177 in Dda1993 is preferably substituted with F, G, S, V, W, or Y. These substitutions decrease the speed when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 177 in Dda1993 is preferably substituted with D, E, G, H, I, L, M, N, Q, R, or T. These substitutions increase the speed when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 177 in Dda1993 is preferably substituted with F, H, I, L, M, N, or W. These substitutions decrease the precision and normalized speed distribution when used to characterize polynucleotide analytes (Example 2). They have different effects on velocity (Example 2). The position corresponding to amino acid position 177 in Dda1993 is preferably substituted with N (K177N). This substitution reduces precision and increases the normalized velocity distribution when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 177 in Dda1993 is most preferably substituted with M (K177M).
[0180] The position corresponding to amino acid position 210 in Dda1993 is preferably substituted with D, E, K, S, N, R, H, or Y. The position corresponding to amino acid position 210 in Dda1993 is preferably substituted with R (T210R), H (T210H), or K (T210K). The position corresponding to amino acid position 210 in Dda1993 is preferably substituted with K (T210K). This substitution increases speed and precision when used to characterize polynucleotide analytes (Example 5). This substitution also decreases the normalized rate distribution when used to characterize polynucleotide analytes (Example 5).
[0181] The position corresponding to amino acid position 221 in Dda1993 is preferably substituted with D, K, E, Q, R, A, H, L, T, or Y. The position corresponding to amino acid position 221 in Dda1993 is preferably substituted with D (N221D) or E (N221E). The position corresponding to amino acid position 221 in Dda1993 is preferably substituted with E (N221E). This substitution increases speed and increases precision when used to characterize polynucleotide analytes (Example 5). This substitution also decreases the normalized rate distribution when used to characterize polynucleotide analytes (Example 5).
[0182] The position corresponding to amino acid position 350 in Dda1993 is preferably substituted with D, E, A, V, I, L, M, F, W, R, H, K, L, S, T, N, or Q. The position corresponding to amino acid position 350 in Dda1993 is preferably substituted with I, F, W, or S. The position corresponding to amino acid position 350 in Dda1993 is preferably substituted with I or S (Y350I or Y350S). The position corresponding to amino acid position 350 in Dda1993 is preferably substituted with I (Y350I). This substitution increases the rate and decreases the precision and normalized rate distribution when used to characterize polynucleotide analytes (Example 3). The position corresponding to amino acid position 350 in Dda1993 is preferably substituted with I or S (Y350I or Y350S). These substitutions have the effects shown in Example 4 when used with the pore complexes of the invention.
[0183] The position corresponding to amino acid position 350 in Dda1993 is preferably substituted with A, D, E, G, K, L, N, Q, R, T, V, H, or M. The position corresponding to amino acid position 350 in Dda1993 is preferably substituted with D (Y350D) or E (Y350E). The position corresponding to amino acid position 350 in Dda1993 is preferably substituted with E (Y350E). This substitution increases speed and increases precision when used to characterize polynucleotide analytes (Example 5). This substitution also decreases the normalized rate distribution when used to characterize polynucleotide analytes (Example 5).
[0184] The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with D, E, A, V, I, L, M, F, Y, W, R, H, L, S, T, N, or Q.
[0185] The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with E, I, L, or M. These substitutions decrease the rate when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with I or M. These substitutions decrease the rate and increase the precision when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with M (K358M). This substitution decreases the rate and increases the precision and normalized velocity distribution when used to characterize polynucleotide analytes (Example 2). The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with I (K358I). This substitution decreases the rate and normalized velocity distribution and increases the precision when used to characterize polynucleotide analytes (Example 2). Example 2 uses a CsgG pore that does not contain a CsgF peptide.
[0186] The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with A, E, F, I, M, or S. These substitutions increase accuracy when used to characterize polynucleotide analytes (Example 3). The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with A, E, F, I, or M. These substitutions decrease speed when used to characterize polynucleotide analytes (Example 3). The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with S (K358S). This substitution increases speed when used to characterize polynucleotide analytes (Example 3). The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with A, E, I, M, or S. These substitutions decrease the normalized speed distribution when used to characterize polynucleotide analytes (Example 3). The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with (K358F). These substitutions increase the normalized velocity distribution when used to characterize polynucleotide analytes (Example 3), which uses a CsgG pore complex containing a CsgF peptide.
[0187] The position corresponding to amino acid position 358 in Dda1993 is preferably substituted with I, L, or Q. These substitutions decrease the rate and increase the precision and normalized rate distribution when used to characterize polynucleotide analytes (Example 4), which uses a pore complex of the invention.
[0188] The position corresponding to amino acid position 358 in Dda1993 is most preferably substituted with I (K358I).
[0189] The modified helicases of the present invention may further comprise any of the modifications, mutations, or substitutions discussed below.
[0190] A Dda helicase modified according to the present invention may be any of SEQ ID NOs: 118-133. SEQ ID NO: 118 is Dda1993. The modified helicase preferably comprises a variant of any of SEQ ID NOs: 118-133. The variant may have any percentage of sequence homology / identity to any of SEQ ID NOs: 118-113, as listed below. Table 4 below summarizes preferred Dda helicases that may be modified according to the present invention. [Table 4]
[0191] Table 5 shows the amino acids in SEQ ID NOs: 119-133 that correspond to positions 40, 55, 114, 156, 177, 210, 221, 350, and 358 in SEQ ID NO: 118. [Table 5]
[0192] The helicase of the present invention preferably comprises a variant of SEQ ID NO: 118 that includes one or more of (a) to (h) below. (a) T55D, T55E, T55K, T55N, or T55S, or T55K. (b) C114A, C114V, C114I, C114L, C114M, C114F, C114Y, C114W, C114G, C114P, C114S, C114T, C114N, or C114Q, C114A, C114G, C114I, C114L, C114M, C114P, C114S, C114T, or C114V, C114G, C114L, C114S, or C114T, C114A, C114I, C114M, C114P, or C114V C11G, C114I or C114P C114G, C114I, or C114P, or C114I, (c) T156A, T156E, T156F, T156G, T156I, T156L, T156M, T156P, T156S, T156V, T156Y, T156D, T156K, or T156N, or T156F, (d) K177D, K177E, K177F, K177G, K177H, K177I, K177L, K177M, K177N, K177Q, K177R, K177S, K177T, K177V, K177W, or K177Y, K177F, K177G, K177S, K177V, K177W, or K177Y, K177D, K177E, K177G, K177H, K177I, K177L, K177M, K177N, K177Q, K177R, or K177T, K177F, K177H, K177I, K177L, K177M, K177N, or K177W, K177N, or K177M, (e) T210D, T210E, T210K, T210S, T210N, T210R, T210H, or T210Y, T210R, T210H, or T210Y, or T210K, (f) N221D, N221K, N221E, N221Q, N221R, N221A, N221H, N221L, N221T, or N221Y, N221D or N221E, or N221E, (g) Y350D, Y350E, Y350A, Y350V, Y350I, Y350L, Y350M, Y350F, Y350W, Y350R, Y350H, Y350K, Y350L, Y350S, Y350T, Y350N, or Y350Q, Y350I or Y350S, Y350I, Y350S, Y350I, Y350F, Y350W, or Y350S, Y350A, Y350D, Y350E, Y350G, Y350K, Y350L, Y350N, Y350Q, Y350R, Y350T, Y350V, Y350H, or Y350M, Y350D or Y350E, or Y350E, and (h) K358D, K358E, K358A, K358V, K358I, K358L, K358M, K358F, K358Y, K358W, K358R, K358H, K358L, K358S, K358T, K358N, or K358Q, K358E, K358I, K358L, or K358M, K358I or K358M, K358M, K358I, K358A, K358E, K358F, K358I, K358M, or K358S, K358A, K358E, K358F, K358I, or K358M, K358S, K358A, K358E, K358I, K358M, or K358S, K358F, K358I, K358L, or K358Q, or K358I. Variants may include any combination and permutation of (a)-(h) as listed above.
[0193] The helicase of the present invention preferably comprises a variant of SEQ ID NO: 118 comprising one or more of (a), (b), (c), and (d) as follows: (a) C114A, C114V, C114I, C114L, C114M, C114F, C114Y, C114W, C114G, C114P, C114S, C114T, C114N, or C114Q, C114A, C114G, C114I, C114L, C114M, C114P, C114S, C114T, or C114V, C114G, C114L, C114S, or C114T, C114A, C114I, C114M, C114P, or C114V C11G, C114I or C114P C114G, C114I, or C114P, or C114I, (b) K177D, K177E, K177F, K177G, K177H, K177I, K177L, K177M, K177N, K177Q, K177R, K177S, K177T, K177V, K177W, or K177Y, K177F, K177G, K177S, K177V, K177W, or K177Y, K177D, K177E, K177G, K177H, K177I, K177L, K177M, K177N, K177Q, K177R, or K177T, K177F, K177H, K177I, K177L, K177M, K177N, or K177W, K177N, or K177M, (c) Y350D, Y350E, Y350A, Y350V, Y350I, Y350L, Y350M, Y350F, Y350W, Y350R, Y350H, Y350K, Y350L, Y350S, Y350T, Y350N, or Y350Q, Y350I or Y350S, Y350I, or Y350S, and (d) K358D, K358E, K358A, K358V, K358I, K358L, K358M, K358F, K358Y, K358W, K358R, K358H, K358L, K358S, K358T, K358N, or K358Q, K358E, K358I, K358L, or K358M, K358I or K358M, K358M, K358I, K358A, K358E, K358F, K358I, K358M, or K358S, K358A, K358E, K358F, K358I, or K358M, K358S, K358A, K358E, K358I, K358M, or K358S, K358F, K358I, K358L, or K358Q, or K358I.
[0194] Variants may include (a); (b); (c); (d); (a) and (b); (a) and (c); (a) and (d); (b) and (c); (b) and (d); (c) and (d); (a), (b), and (c); (a), (b), and (d); (a), (c), and (d); (b), (c), and (d); or (a), (b), (c), and (d).
[0195] Preferred variants of SEQ ID NO: 118 include: C114I; K177M; Y350I; K358I; C114I and K177M; C114I and Y350I; C114I and K358I; K177M and Y350I; K177M and K358I; Y350I and K358I; C114I, K177M, and Y350I; C114I, K177M, and K358I; C114I, Y350I, and K358I; K177M, Y350I, and K358I; or C114I, K177M, Y350I, and K358I.
[0196] The helicase preferably comprises a variant of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, wherein one or more of the positions corresponding to amino acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993 have been modified or substituted (including specific substitutions) as defined above. The various combinations and permutations of one or more of positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993 are defined above with reference to (a) to (h).
[0197] The helicase preferably comprises a variant of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, wherein one or more of the positions corresponding to amino acid positions 114, 177, 350, and 358 in Dda1993 have been modified or substituted (including specific substitutions) as defined above. The various combinations and permutations of one or more of positions 114, 177, 350, and 358 in Dda1993 are defined above with reference to (a) to (d).
[0198] The helicase preferably comprises a variant of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, in which one or more of the positions corresponding to amino acid positions 114, 177, and 358 in Dda1993 have been modified or substituted (including specific substitutions) as defined above.
[0199] Position 40 in Dda1993 Any of the modified helicases of the invention may further comprise a modification or substitution at a position corresponding to amino acid position 40 in Dda1993. Position 40 or a corresponding position may be substituted as A, V, I, L, M, F, Y, or W. The position corresponding to position T40 in Dda1993 is set forth in Table 5 above. Helicases of the invention preferably include variants of SEQ ID NO: 118 that, in addition to the modifications / substitutions listed above, further comprise a substitution at T40, e.g., T40A, T40V, T40I, T40L, T40M, T40F, T40Y, or T40W. The substitution is preferably T40Y.
[0200] The present invention provides modified DNA-dependent ATPase (Dda) helicases, which contain a modification or substitution at a position corresponding to amino acid position 40 in Dda1993. Position T40 is in the tower domain of Dda1993. Helicases of the present invention preferably contain variants of SEQ ID NO: 118, which contain a substitution at T40, e.g., T40A, T40V, T40I, T40L, T40M, T40F, T40Y, or T40W. The substitution is preferably T40Y. Modified Dda helicases of the present invention may further contain a modification or substitution at one or more positions corresponding to amino acid positions (a) 55, (b) 114, (c) 156, (d) 177, (e) 210, (f) 221, (g) 350, and (h) 358, including any combination and permutation of (a) through (h) listed above. The modified Dda helicases of the present invention may further comprise modifications or substitutions at one or more of the positions corresponding to amino acid positions (a) 114, (b) 177, (c) 350, and (d) 358 in Dda1993, including (a); (b); (c); (d); (a) and (b); (a) and (c); (a) and (d); (b) and (c); (b) and (d); (c) and (d); (a), (b), and (c); (a), (b), and (d); (b), (c), and (d); or (a), (b), (c), and (d). The helicase preferably comprises a variant of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, wherein the position corresponding to amino acid position 40 in Dda1993 has been modified or substituted (including specific substitutions) as defined above.
[0201] The modified helicases of the present invention may further comprise any of the modifications, substitutions, combinations of modifications, or combinations of substitutions discussed below.
[0202] Other helicases of the present invention The present invention also provides modified DNA-dependent ATPase (Dda) helicases having any of the modifications, substitutions, combinations of modifications, or combinations of substitutions discussed below, alone or in combination. In other words, these helicases of the present invention do not necessarily have substitutions at any of amino acid positions 40, 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993, or at positions corresponding to any of amino acid positions 40, 114, 350, 177, and K358 in Dda1993. Such modified helicases of the present invention are preferably variants of any of SEQ ID NOS: 118-133. Variants can have any percentage of sequence homology / identity to any of SEQ ID NOS: 118-133, as listed below.
[0203] The use of Dda helicase in the characterization of analytes is described in WO2015 / 055981, WO2015 / 166276, and WO2016 / 055777 (all incorporated by reference).
[0204] The modified helicases of the present invention provide more consistent movement of a target analyte relative to, for example, through, a transmembrane pore, leading to improved accuracy. The helicases preferably provide more consistent movement of a target analyte, such as a polynucleotide, from one k-mer to another k-mer or from k-mer to k-mer as it moves relative to, for example, through, a pore. The helicases allow a target analyte, such as a target polynucleotide, to move more smoothly relative to, for example, through a transmembrane pore. The helicases preferably provide more regular or less irregular movement of a target analyte, such as a target polynucleotide, relative to, for example, through a transmembrane pore.
[0205] The modification(s) typically increase accuracy by at least 0.1%, at least 0.5%, at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% compared to a helicase without the modification.
[0206] The ability of a helicase to control polynucleotide movement can be determined as described in the Examples.
[0207] The modified helicase has the ability to control the movement of a polynucleotide. The ability of a helicase to control the movement of a polynucleotide can be assayed using any method known in the art. For example, the helicase can be contacted with a polynucleotide, and the position of the polynucleotide can be determined using standard methods. The ability of a modified helicase to control the movement of a polynucleotide is typically assayed in a nanopore system such as those described below, particularly as described in the Examples.
[0208] The modified helicases of the present invention can be isolated, substantially isolated, purified, or substantially purified. A helicase is isolated or purified when it is completely free of any other components, such as lipids, polynucleotides, pore monomers, or other proteins. A helicase is substantially isolated when it is mixed with a carrier or diluent that does not interfere with its intended use. For example, a helicase is substantially isolated or substantially purified when it is present in a form that contains less than 10%, less than 5%, less than 2%, or less than 1% of other components, such as lipids, polynucleotides, pore monomers, or other proteins.
[0209] Dda helicase Any Dda helicase may be modified according to the present invention. Preferred Dda helicases are discussed below and described in WO2015 / 055981, WO2015 / 166276, and WO2016 / 055777 (all incorporated by reference).
[0210] Dda helicases typically contain five domains: 1A (RecA-like motor) domain, 2A (RecA-like motor) domain, tower domain, pin domain, and hook domain (Xiaoping He et al., 2012, Structure; 20: 1189-1200). Domains can be identified by protein modeling, X-ray diffraction measurements of proteins in the crystalline state (Rupp B (2009). Biomolecular Crystallography: Principles, Practice and Application to Structural Biology. New York: Garland Science.), nuclear magnetic resonance (NMR) spectroscopy of proteins in solution (Mark Rance; Cavanagh, John; Wayne J. Fairbrother; Arthur W. Hunt III; Skelton, N. Cholas J. (2007). Protein NMR spectroscopy: principles and practice (2nd ed.). Boston: Academic Press.), or cryo-electron microscopy of proteins in the frozen hydrated state (van Heel M, Gowen B, Matadeen R, Orlova EV, Finn R, Pape T, Cohen D, Stark H, Schmidt R, Schatz M, Patwardhan A (2000). "Single-particle electron cryo-microscopy: towards atomic Protein structural information determined by the above-mentioned methods is publicly available from the Protein Bank (PDB) database.
[0211] In addition to modifications or substitutions at one or more positions corresponding to amino acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993, modifications or substitutions at one or more positions corresponding to amino acid positions 114, 177, 350, and 358 in Dda1993, modifications or substitutions at one or more positions corresponding to amino acid positions 114, 177, and 350 in Dda1993, and / or modifications or substitutions at a position corresponding to position 40 in Dda1993, modified helicases of the invention preferably include any of the following additional modifications, substitutions, combinations of modifications, or combinations of substitutions: As explained above, the present invention also provides modified helicases that have any of the modifications, substitutions, combinations of modifications, or combinations of substitutions listed below, either alone (i.e., not necessarily having a substitution at any of positions 40, 55, 114, 156, 177, 210, 221, 350, and 358 of Dda1993, or at any of positions 40, 114, 177, 350, and 358 of Dda1993).
[0212] Modifications of the present invention The helicase of the present invention may be a helicase in which at least one amino acid that interacts with the transmembrane pore has been substituted. Any number of amino acids may be substituted, for example, 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, or 6 or more. 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acids may be substituted. The amino acid that interacts with the transmembrane pore may be identified using protein modeling, as discussed above.
[0213] Base and / or sugar interactions The helicase of the present invention is preferably a helicase in which at least one amino acid that interacts with the sugar and / or base of one or more nucleotides in single-stranded DNA (ssDNA) is substituted with an amino acid containing a larger side chain (R group). Any number of amino acids can be substituted, for example, one or more, two or more, three or more, four or more, five or more, or six or more amino acids. Each amino acid can interact with the base, the sugar, or a base and a sugar. The amino acids that interact with the sugar and / or base of one or more nucleotides in single-stranded DNA can be identified using protein modeling, as discussed above.
[0214] Helicases of the present invention preferably comprise variants of SEQ ID NO: 118, wherein at least one amino acid that interacts with the sugar and / or base of one or more nucleotides in ssDNA is at least one of H82, N88, P89, F98, D121, V150, P152, F240, F276, S287, H396, and Y415. These numbers correspond to the relevant positions in SEQ ID NO: 118 and may need to be changed in the case of variants in which one or more amino acids are inserted or deleted compared to SEQ ID NO: 118. Those skilled in the art will be able to determine the corresponding positions in variants, as discussed above. Helicases of the invention preferably comprise a variant of SEQ ID NO: 118, wherein at least one amino acid that interacts with the sugar and / or base of one or more nucleotides in the ssDNA is F98 and one or more of H82, N88, P89, D121, V150, P152, F240, F276, S287, H396, and Y415, e.g., F98 / H82, F98 / N88, F98 / P89, F98 / D121, F98 / V150, F98 / P152, F98 / F240, F98 / F276, F98 / S287, or F98 / H396.
[0215] The helicase of the present invention is preferably a variant of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, wherein at least one amino acid that interacts with the sugar and / or base of one or more nucleotides in the ssDNA is at least one of the amino acids corresponding to H82, N88, P89, F98, D121, V150, P152, F240, F276, S287, H396, and Y415 in SEQ ID NO: 118. The helicase of the present invention preferably comprises a variant of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, wherein at least one amino acid that interacts with the sugar and / or base of one or more nucleotides in ssDNA is the amino acid corresponding to F98 in SEQ ID NO: 118, and and Y415, e.g., amino acids corresponding to F98 / H82, F98 / N88, F98 / P89, F98 / D121, F98 / V150, F98 / P152, F98 / F240, F98 / F276, F98 / S287, or F98 / H396. Table 6 shows the amino acids in SEQ ID NOs: 119 to 133 that correspond to H82, N88, P89, F98, D121, V150, P152, F240, F276, S287, H396, and Y415 in SEQ ID NO: 118. [Table 6]
[0216] The at least one amino acid that interacts with the sugar and / or base of one or more nucleotides in the ssDNA is preferably at least one amino acid that intercalates between nucleotides in the ssDNA. The amino acid that intercalates between nucleotides in the ssDNA can be modeled as discussed above. The at least one amino acid that intercalates between nucleotides in the ssDNA is preferably at least one of P89, F98, and V150 in SEQ ID NO: 118, such as P89, F98, V150, P89 / F98, P89 / V150, F98 / V150, or P89 / F98 / V150.
[0217] At least one amino acid intercalating between nucleotides in the ssDNA in SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133 is preferably at least one of the amino acids corresponding to P89, F98, and V150 in SEQ ID NO: 118, e.g., P89, F98, V150, P89 / F98, P89 / V150, F98 / V150, or P89 / F98 / V150. The corresponding amino acids are shown in Table 6 above.
[0218] Larger R groups The larger side chain (R group) preferably (a) contains an increased number of carbon atoms, (b) has an increased length, (c) has an increased molecular volume, and / or (d) has an increased van der Waals volume. The larger side chain (R group) is preferably (a); (b); (c); (d); (a) and (b); (a) and (c); (a) and (d); (b) and (c); (b) and (d); (c) and (d); (a), (b), and (c); (a), (b), and (d); (a), (c), and (d); or (a), (b), (c), and (d). Each of (a) through (d) can be measured using standard methods in the art.
[0219] A larger side chain (R group) preferably increases (i) electrostatic interactions, (ii) hydrogen bonding, and / or (iii) cation-pi (cation-π) interactions between at least one amino acid and one or more nucleotides in the ssDNA, e.g., (i); (ii); (iii); (i) and (ii); (i) and (iii); (ii) and (iii); and (i), (ii), and (iii). Those skilled in the art can determine whether an R group increases any of these interactions. For example, in (i), positively charged amino acids such as arginine (R), histidine (H), and lysine (K) have R groups that increase electrostatic interactions. For example, in (ii), amino acids such as asparagine (N), serine (S), glutamine (Q), threonine (T), and histidine (H) have R groups that increase hydrogen bonding. For example, in (iii), aromatic amino acids such as phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H) have R groups that increase cation-pi (cation-π) interactions. Specific substitutions below are labeled (i) through (iii) to reflect these changes. Other possible substitutions are labeled (iv). These (iv) substitutions typically increase the length of the side chain (R group).
[0220] Amino acids containing larger side chains (R) can be unnatural amino acids, which can be any of those discussed below.
[0221] The amino acid containing the larger side chain (R group) is preferably not alanine (A), cysteine (C), glycine (G), selenocysteine (U), methionine (M), aspartic acid (D), or glutamic acid (E).
[0222] Histidine (H) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q) or asparagine (N), or (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W). Histidine (H) is more preferably substituted with (a) N, Q, or W, or (b) Y, F, Q, or K.
[0223] Asparagine (N) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q) or histidine (H), or (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W). Asparagine (N) is more preferably substituted with R, H, W, or Y.
[0224] Proline (P) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), threonine (T), or histidine (H), (iii) tyrosine (Y), phenylalanine (F), or tryptophan (W), or (iv) leucine (L), valine (V), or isoleucine (I). Proline (P) is more preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), threonine (T), or histidine (H), (iii) phenylalanine (F) or tryptophan (W), or (iv) leucine (L), valine (V), or isoleucine (I). Proline (P) is more preferably substituted with (a) F; (b) L, V, I, T, or F; or (c) W, F, Y, H, I, L, or V.
[0225] Valine (V) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), or histidine (H), (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W), or (iv) isoleucine (I) or leucine (L). Valine (V) is more preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), or histidine (H), (iii) tyrosine (Y) or tryptophan (W), or (iv) isoleucine (I) or leucine (L). Valine (V) is more preferably substituted with I or H, or I, L, N, W, or H.
[0226] Phenylalanine (F) is preferably substituted with (i) arginine (R) or lysine (K), (ii) histidine (H), or (iii) tyrosine (Y) or tryptophan (W). Phenylalanine (F) is more preferably substituted with (a) W, (b) W, Y, or H, (c) W, R, or K, or (d) K, H, W, or R.
[0227] Glutamine (Q) is preferably substituted with (i) arginine (R) or lysine (K), or (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W).
[0228] Alanine (A) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), or histidine (H), (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W), or (iv) isoleucine (I) or leucine (L).
[0229] Serine (S) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), or histidine (H), (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W), or (iv) isoleucine (I) or leucine (L). Serine (S) is preferably substituted with K, R, W, or F.
[0230] Lysine (K) is preferably substituted with (i) arginine (R), or (iii) tyrosine (Y) or tryptophan (W).
[0231] Arginine (R) is preferably replaced with (iii) tyrosine (Y) or tryptophan (W).
[0232] Methionine (M) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), or histidine (H), or (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W).
[0233] Leucine (L) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q) or asparagine (N), or (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W).
[0234] Aspartic acid (D) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), or histidine (H), or (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W). Aspartic acid (D) is more preferably substituted with H, Y, or K.
[0235] Glutamic acid (E) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), or histidine (H), or (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W).
[0236] Isoleucine (I) is preferably substituted with (i) arginine (R) or lysine (K), (ii) glutamine (Q), asparagine (N), or histidine (H), (iii) phenylalanine (F), tyrosine (Y), or tryptophan (W), or (iv) leucine (L).
[0237] Tyrosine (Y) is preferably substituted with (i) arginine (R) or lysine (K), or (iii) tryptophan (W). Tyrosine (Y) is more preferably substituted with W or R.
[0238] The helicase more preferably comprises a variant of SEQ ID NO: 118, including (a) P89F, (b) F98W, (c) V150I, (d) V150H, (e) P89F and F98W, (f) P89F and V150I, (g) P89F and V150H, (h) F98W and V150I, (i) F98W and V150H, (j) P89F, F98W, and V150I, or (k) P89F, F98W, and V150H.
[0239] The helicase more preferably comprises a variant of SEQ ID NO: 118, including H82N, H82Q, H82W, N88R, N88H, N88W, N88Y, P89L, P89V, P89I, P89E, P89T, P89F, D121H, D121Y, D121K, V150I, V150L, V150N, V150W, V150H, P152W, P152F, P152Y, P152H, P152I, P152L, P152V, F 240W, F240Y, F240H, F276W, F276R, F276K, F276H, S287K, S287R, S287W, S287F, H396Y, H396F, H396Q, H396K, Y415W , Y415R, F98W / H82N, F98W / H82Q, F98W / H82W, F98W / N88R, F98W / N88H, F98W / N88W, F98W / N88Y, F98W / P89L, F98W / P89 V, F98W / P89I, F98W / P89T, F98W / P89F, F98W / D121H, F98W / D121Y, F98W / D121K, F98W / V150I, F98W / V150L, F98W / V1 50N, F98W / V150W, F98W / V150H, F98W / P152W, F98W / P152F, F98W / P152Y, F98W / P152H, F98W / P152I, F98W / P152L, F98 Includes W / P152V, F98W / F240W, F98W / F240Y, F98W / F240H, F98W / F276W, F98W / F276R, F98W / F276K, F98W / F276H, F98W / S287K, F98W / S287R, F98W / S287W, F98W / S287F, F98W / H396Y, F98W / H396F, F98W / H396Q, F98W / Y415W, or F98W / Y415R.
[0240] Phosphate Interactions The helicase of the present invention is preferably a helicase in which at least one amino acid that interacts with one or more phosphate groups in one or more nucleotides in ssDNA has been substituted. Any number of amino acids can be substituted, for example, one or more, two or more, three or more, four or more, five or more, or six or more amino acids. Each nucleotide in ssDNA contains three phosphate groups. Each substituted amino acid can interact with any number of phosphate groups at once, for example, one, two, or three phosphate groups at once. Amino acids that interact with one or more phosphate groups can be identified using protein modeling, as discussed above.
[0241] The substitutions preferably increase (i) electrostatic interactions, (ii) hydrogen bonding, and / or (iii) cation-pi interactions between at least one amino acid and one or more phosphate groups in the ssDNA. Preferred substitutions that increase (i), (ii), and (iii) are discussed below using the labels (i), (ii), and (iii).
[0242] The substitution preferably increases the net positive charge at the position. The net charge at any position can be measured using methods known in the art. For example, the isoelectric point can be used to define the net charge of an amino acid. The net charge is typically measured at about 7.5. The substitution preferably involves replacing a negatively charged amino acid with a positively charged, uncharged, nonpolar, or aromatic amino acid. A negatively charged amino acid is an amino acid that has a net negative charge. Negatively charged amino acids include, but are not limited to, aspartic acid (D) and glutamic acid (E). A positively charged amino acid is an amino acid that has a net positive charge. Positively charged amino acids may be naturally occurring or non-naturally occurring. Positively charged amino acids may be synthetic or modified. For example, modified amino acids with a net positive charge can be specifically designed for use in the present invention. Many different types of modifications to amino acids are well known in the art. Preferred naturally occurring positively charged amino acids include, but are not limited to, histidine (H), lysine (K), and arginine (R).
[0243] An uncharged amino acid, apolar amino acid, or an aromatic amino acid may be naturally occurring or non-naturally occurring. It may be synthetic or modified. An uncharged amino acid has no net charge. Suitable uncharged amino acids include, but are not limited to, cysteine (C), serine (S), threonine (T), methionine (M), asparagine (N), and glutamine (Q). A nonpolar amino acid has a nonpolar side chain. Suitable nonpolar amino acids include, but are not limited to, glycine (G), alanine (A), proline (P), isoleucine (I), leucine (L), and valine (V). An aromatic amino acid has an aromatic side chain. Suitable aromatic amino acids include, but are not limited to, histidine (H), phenylalanine (F), tryptophan (W), and tyrosine (Y).
[0244] The helicase preferably comprises a variant of SEQ ID NO: 118, wherein at least one amino acid that interacts with one or more phosphates in one or more nucleotides in the ssDNA is at least one of H64, T80, S83, N242, K243, N293, T394, and K397. These numbers correspond to the relevant positions in SEQ ID NO: 89 and may need to be changed in the case of a variant in which one or more amino acids are inserted or deleted compared to SEQ ID NO: 118. Those skilled in the art will be able to determine the corresponding positions in the variant, as discussed above.
[0245] The helicase preferably comprises a variant of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, wherein at least one amino acid that interacts with one or more phosphates in one or more nucleotides in the ssDNA is at least one of the amino acids corresponding to H64, T80, S83, N242, K243, N293, T394, and K397 in SEQ ID NO: 118. Table 7 shows the amino acids in SEQ ID NOs: 119 to 133 that correspond to H64, T80, S83, N242, K243, N293, T394, and K397 in SEQ ID NO: 118. [Table 7]
[0246] Histidine (H) is preferably substituted with (i) arginine (R) or lysine (K), (ii) asparagine (N), serine (S), glutamine (Q), or threonine (T), (iii) phenylalanine (F), tryptophan (W), or tyrosine (Y). Histidine (H) is preferably substituted with (a) N, Q, K, or F, or (b) N, Q, or W.
[0247] Threonine (T) is preferably substituted with (i) arginine (R), histidine (H), or lysine (K), (ii) asparagine (N), serine (S), glutamine (Q), or histidine (H), or (iii) phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H). Threonine (T) is more preferably substituted with (a) K, Q, or N, or (b) K, H, or N.
[0248] Serine (S) is preferably substituted with (i) arginine (R), histidine (H), or lysine (K), (ii) asparagine (N), glutamine (Q), threonine (T), or histidine (H), or (iii) phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H). Serine (S) is more preferably substituted with H, N, K, T, R, or Q.
[0249] Asparagine (N) is preferably substituted with (i) arginine (R), histidine (H), or lysine (K), (ii) serine (S), glutamine (Q), threonine (T), or histidine (H), or (iii) phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H). Asparagine (N) is more preferably substituted with (a) H or Q, or (b) Q, K, or H.
[0250] Lysine (K) is preferably substituted with (i) arginine (R) or histidine (H), (ii) asparagine (N), serine (S), glutamine (Q), threonine (T), or histidine (H), or (iii) phenylalanine (F), tryptophan (W), tyrosine (Y), or histidine (H). Lysine (K) is more preferably substituted with (a) Q or H, or (b) R, H, or Y.
[0251] The helicase more preferably comprises a variant of SEQ ID NO: 118, including one or more, for example all of: (a) H64N, H64Q, H64K, or H64F; (b) T80K, T80Q, or T80N; (c) S83H, S83N, S83K, S83T, S83R, or S83Q; (d) N242H or N242Q; (e) K243Q or K243H; (f) N293Q, N293K, or N293H; (g) T394K, T394H, or T394N; or (h) K397R, K397H, or K397Y.
[0252] combination The helicase is preferably a variant of SEQ ID NO: 118 comprising the following substitutions: -F98 / H64, for example F98W / H64N, F98W / H64Q, F98W / H64K, or F98W / H64F, -F98 / T80, for example, F98W / T80K, F98W / T80Q, F98W / T80N, F98 / H82, for example F98W / H82N, F98W / H82Q or F98W / H82W, F98 / S83, for example F98W / S83H, F98W / S83N, F98W / S83K, F98W / S83T, F98W / S83R or F98W / S83Q, F98 / N242, for example F98W / N242H, F98W / N242Q, F98W / K243Q or F98W / K243H, F98 / N293, for example F98W / N293Q, F98W / N293K, F98W / N293H, F98W / T394K, F98W / T394H, F98W / T394N, F98W / H396Y, F98W / H396F, F98W / H396Q or F98W / H396K, or F98 / K397, for example, F98W / K397R, F98W / K397H, or F98W / K397Y.
[0253] Preferred combinations in SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133 include amino acid combinations corresponding to the combinations in SEQ ID NO: 118 listed above.
[0254] Pore Interactions The helicases of the present invention are further helicases in which the portion of the helicase that interacts with the transmembrane pore comprises one or more modifications, preferably one or more substitutions. The portion of the helicase that interacts with the transmembrane pore is typically the portion of the helicase that interacts with the transmembrane pore, for example, when the helicase is used to control the movement of a polynucleotide through the pore, as discussed in more detail below. This portion typically includes amino acids that interact with or contact the pore, for example, when the helicase is used to control the movement of a polynucleotide through the pore, as discussed in more detail below. This portion typically includes amino acids that interact with or contact the pore when the helicase is bound or attached to an analyte, such as a polynucleotide, that is moving through the pore under an applied potential.
[0255] In SEQ ID NO: 118, the transmembrane pore interacting moieties are typically located at positions 1, 2, 3, 4, 5, 6, 51, 176, 177, 178, 179, 180, 181, 185, 189, 191, 193, 194, 195, 197, 198, 199, 200, 201, 202, 203, 204, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267, 268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 13, 216, 219, 220, 221, 223, 224, 226, 227, 228, 229, 247, 254, 255, 256, 257, 258, 259, 260, 261, 298, 300, 304, 308, 318, 319, 321, 337, 347, 350, 351, 405, 415, 422, 434, 437, 438. These numbers correspond to the relevant positions in SEQ ID NO: 118 and may need to be changed in the case of variants in which one or more amino acids are inserted or deleted compared to SEQ ID NO: 118. Those skilled in the art will be able to determine the corresponding positions in variants as discussed above. The moiety that interacts with the transmembrane pore preferably comprises: (a) positions 1, 2, 4, 51, 177, 178, 179, 180, 185, 193, 195, 197, 198, 199, 200, 202, 203, 204, 207, 208, 209, 210, 211, 212, 216, 221, 223, 224, 226, 227, 228, 229, 254, 255, 256, 257, 258, 260, 304, 318, 321, 347, 350, 351, 405, 415, 422, 434, 437, and 438 in SEQ ID NO: 118, or (b) comprising the amino acids at positions 1, 2, 178, 179, 180, 185, 195, 197, 198, 199, 200, 202, 203, 207, 209, 210, 212, 216, 221, 223, 226, 227, 255, 258, 260, 304, 350, and 438 in SEQ ID NO: 118.
[0256] The portion that interacts with the transmembrane pore preferably comprises one or more, such as two, three, four, or five, of the amino acids at positions K194, W195, K198, K199, and E258 in SEQ ID NO: 118. Variants of SEQ ID NO: 118 preferably comprise a modification at one or more of: (a) K194, (b) W195, (c) D198, (d) K199, and (d) E258. Variants of SEQ ID NO: 118 preferably comprise a substitution at one or more of: (a) K194, such as K194L, (b) W195, such as W195A, (c) D198, such as D198V, (d) K199, such as K199L, and (e) E258, such as E258L. Variants are {a}, {b}, {c}, {d}, {e}, {a, b}, {a, c}, {a, d}, {a, e}, {b, c}, {b, d}, {b, e}, {c, d}, {c, e}, {d, e}, {a, b, c}, {a, b, d}, {a, b, e}, {a, c, d} , {a, c, e}, {a, d, e}, {b, c, d}, {b, c, e}, {b, d, e}, {c, d, e}, {a, b, c, d}, {a, b, c, e}, {a, b, d, e}, {a, c, d, e}, {b, c, d, e}, or {a, b, c, d, e}. The modifications or substitutions listed in this paragraph are preferred when the modified polynucleotide binding protein interacts with a pore derived from MspA, particularly any of the modified pores discussed below.
[0257] The portion of the polynucleotide binding protein that interacts with the transmembrane pore preferably comprises the amino acid at position 194 or 199 of SEQ ID NO: 118. Variants preferably comprise K194A, K194V, K194F, K194D, K194S, K194W, or K194L, and / or K199A, K199V, K199F, K199D, K199S, K199W, or K199L. In SEQ ID NOs: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, the portion that interacts with the transmembrane pore typically comprises amino acids at positions corresponding to those in SEQ ID NO: 118 listed above. The amino acids in SEQ ID NOs: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, and 133 that correspond to these positions in SEQ ID NO: 118 can be identified using the alignment in Table 8 below. [Table 8]
[0258] Preferred combinations Helicases of the present invention preferably comprise a variant of SEQ ID NO: 118, comprising a substitution with F98, e.g., F98R, F98K, F98Q, F98N, F98H, F98Y, F98F, or F98W, and with K194, e.g., K194A, K194V, K194F, K194D, K194S, K194W, or K194L, and / or K199, e.g., K199A, K199V, K199F, K199D, K199S, K199W, or K199L. Helicases of the invention preferably include variants of SEQ ID NOs: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, which include a substitution at the position corresponding to F98 in SEQ ID NO: 118, and a substitution at position(s) corresponding to K194 and / or K199 in SEQ ID NO: 118. These corresponding positions may be replaced with any of the amino acids listed above for F98, K194, and K119 in SEQ ID NO: 118.
[0259] The helicase is preferably a variant of SEQ ID NO: 118 comprising the following substitutions: F98 / K194 / H64, such as F98W / K194L / H64N, F98W / K194L / H64Q, F98W / K194L / H64K, or F98W / K194L / H64F; F98 / K194 / T80, for example F98W / K194L / T80K, F98W / K194L / T80Q or F98W / K194L / T80N, F98 / K194 / H82, for example F98W / K194L / H82N, F98W / K194L / H82Q or F98W / K194L / H82W, F98 / S83 / K194, for example F98W / S83H / K194L, F98W / S83T / K194L, F98W / S83R / K194L, F98W / S83Q / K194L, F98W / S83N / K194L, F98W / S83K / K194L, F98W / N88R / K194L, F98W / N88H / K194L, F98W / N88W / K194L or F98W / N88Y / K194L, -F98 / S83 / K194 / F276, for example, F98W / S83H / K194L / F276K, F98 / P89 / K194, for example F98W / P89L / K194L, F98W / P89V / K194L, F98W / P89I / K194L or F98W / P89T / K194L, F98 / D121 / K194, for example F98W / D121H / K194L, F98W / D121Y / K194L or F98W / D121K / K194L, F98 / V150 / K194, for example F98W / V150I / K194L, F98W / V150L / K194L, F98W / V150N / K194L, F98W / V150W / K194L, or F98W / V150H / K194L, F98 / P152 / K194, for example F98W / P152W / K194L, F98W / P152F / K194L, F98W / P152Y / K194L, F98W / P152H / K194L, F98W / P152I / K194L, F98W / P152L / K194L or F98W / P152V / K194L, F98 / F240 / K194, for example F98W / F240W / K194L, F98W / F240Y / K194L or F98W / F240H / K194L, F98 / N242 / K194, for example F98W / N242H / K194L or F98W / N242Q / K194L, F98 / K194 / F276, for example F98W / K194L / F276K, F98W / K194L / F276H, F98W / K194L / F276W, or F98W / K194L / F276R, F98 / K194 / S287, for example F98W / K194L / S287K, F98W / K194L / S287R, F98W / K194L / S287W or F98W / K194L / S287F, F98 / N293 / K194, for example F98W / N293Q / K194L, F98W / N293K / K194L or F98W / N293H / K194L, F98 / T394 / K194, for example F98W / T394K / K194L, F98W / T394H / K194L or F98W / T394N / K194L, F98 / H396 / K194, for example F98W / H396Y / K194L, F98W / H396F / K194L, F98W / H396Q / K194L or F98W / H396K / K194L, F98 / K397 / K194, for example F98W / K397R / K194L, F98W / K397H / K194L or F98W / K397Y / K194L, or F98 / Y415 / K194, for example, F98W / Y415W / K194L or F98W / Y415R / K194L.
[0260] In any of the above combinations, K194 may be replaced with any of W195, D198, K199, and E258.
[0261] The modified helicase preferably comprises a modification or substitution at position(s) corresponding to amino acid positions 98 and / or 194 in Dda1993. This is preferably in addition to a modification or substitution at one or more of positions corresponding to amino acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993, a modification or substitution at one or more positions corresponding to amino acid positions 114, 177, 350, and 358 in Dda1993, and / or a modification or substitution at a position corresponding to position 40 in Dda1993. Position 98 or a corresponding position may be substituted with R, H, K, S, T, N, Q, A, V, I, L, M, Y, or W. Position 98 or a corresponding position is preferably substituted with R, K, Q, N, H, Y, or W. Position 194 or a corresponding position may be substituted with A, V, I, L, M, F, Y, W, D, E, S, T, N, or Q. Position 194 or a corresponding position is preferably substituted with A, V, F, D, S, W, or L. Helicases of the present invention preferably comprise a variant of SEQ ID NO: 118 comprising a substitution with F98, e.g., F98R, F98K, F98Q, F98N, F98H, F98Y, or F98W, and / or a substitution with K194, e.g., K194A, K194V, K194F, K194D, K194S, K194W, or K194L. Helicases of the present invention preferably comprise a variant of SEQ ID NO: 118 comprising F98W and K194L. Helicases of the invention preferably include variants of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, which include a substitution at a position corresponding to F98 in SEQ ID NO: 118 and / or a substitution at a position corresponding to K194 in SEQ ID NO: 118.
[0262] In any of the above combinations, K194 may be replaced with any of W195, D198, K199, and E258.
[0263] Modifications in the tower domain and / or pin domain and / or 1A domain The modified helicase preferably comprises a modification or substitution at a position corresponding to amino acid position 360 in Dda1993. This may be in addition to modifications or substitutions at one or more positions corresponding to amino acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993, modifications or substitutions at one or more positions corresponding to amino acid positions 114, 177, 350, and 358 in Dda1993, and / or modifications or substitutions at a position corresponding to position 40 in Dda1993. A360 is in the tower domain of Dda1993, as is Y350 and K358. Position 360 or a corresponding position may be substituted with C, G, P, A, V, I, L, M, F, Y, or W. Position 360 or a corresponding position is preferably substituted with C or Y. Helicases of the invention preferably include variants of SEQ ID NO: 118 that include a substitution at A360, e.g., A360C or A360Y. Helicases of the invention preferably include variants of SEQ ID NO: 118 that include K358I and A360C. Helicases of the invention preferably include variants of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, which include a substitution at the position corresponding to A360 in SEQ ID NO: 118.
[0264] The modified helicase preferably includes a modification or substitution at one or more of amino acid positions 94, 98, and 109 in Dda1993, for example, position(s) 94, 98, 109, 94 and 98, 94 and 109, 98 and 109, and positions corresponding to 94, 98, and 109. This may be in addition to modifications or substitutions at one or more positions corresponding to amibo acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993, modifications or substitutions at one or more positions corresponding to amino acid positions 114, 177, 350, and 358 in Dda1993, and / or modifications or substitutions at a position corresponding to position 40 in Dda1993. All of these positions are in the pin domain. Position 94 or a corresponding position may be substituted with C, G, P, A, V, I, L, M, F, Y, or W. Position 94 or a corresponding position is preferably substituted with C or Y. Position 98 or a corresponding position may be substituted with R, H, K, S, T, N, Q, A, V, I, L, M Y, or W. Position 98 or a corresponding position is preferably substituted with R, K, Q, N, H, Y, or W. Position 109 or a corresponding position may be substituted with A, V, I, L, M, F, Y, or W. Position 109 or a corresponding position is preferably substituted with A or V. Helicases of the present invention preferably include variants of SEQ ID NO: 118 that include substitutions in one or more of E94, F98, and C109 (including all combinations listed above). Preferred variants include the following substitutions: E94 and F98, for example E94C or E94Y and F98R, F98K, F98Q, F98N, F98H, F98Y, F98F or F98W, E94 and C109, for example E94C or E94Y and C109A or C109V, F98 and C109, for example F98R, F98K, F98Q, F98N, F98H, F98Y, F98F or F98W, and C109A or C109V, or E94, F98 and C109, for example E94C or E94Y and F98R, F98K, F98Q, F98N, F98H, F98Y, F98F or F98W and C109A or C109V.
[0265] More preferred variants include: E94C and F98W, E94C and C109A, F98W and C109A, or E94C, F98W, and C109A. Helicases of the invention preferably include variants of SEQ ID NOs: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, which include substitutions at position(s) corresponding to one or more of E94, F98, and C109 in SEQ ID NO: 118. Table 9 includes information regarding E94, C109, C136, and A360 (see the modified helicases disclosed above and below). [Table 9]
[0266] The helicase of the present invention is preferably a helicase in which at least one cysteine residue (i.e., one or more cysteine residues) and / or at least one unnatural amino acid (i.e., one or more unnatural amino acids) have been introduced into (i) the tower domain, and / or (ii) the pin domain, and / or (iii) the 1A (RecA-like motor) domain, such that the helicase has the ability to control polynucleotide movement. These types of modifications are disclosed in WO2015 / 055981, which is incorporated herein by reference in its entirety. At least one cysteine residue and / or at least one unnatural amino acid can be introduced into the tower domain, the pin domain, the 1A domain, the tower domain and the pin domain, the tower domain and the 1A domain, or the tower domain, the pin domain, and the 1A domain.
[0267] The helicase of the present invention is preferably a helicase in which at least one cysteine residue and / or at least one unnatural amino acid is introduced into each of (i) the tower domain and (ii) the pin domain and / or the 1A (RecA-like motor) domain, i.e., the tower domain and the pin domain, the tower domain and the 1A domain, or the tower domain, the pin domain, and the 1A domain.
[0268] Any number of cysteine residues and / or unnatural amino acids can be introduced into each domain. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more cysteine residues can be introduced and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more unnatural amino acids can be introduced. Only one or more cysteine residues can be introduced. Only one or more unnatural amino acids can be introduced. A combination of one or more cysteine residues and one or more unnatural amino acids can be introduced.
[0269] The at least one cysteine residue and / or the at least one unnatural amino acid is preferably introduced by substitution. Methods for doing this are known in the art.
[0270] These modifications do not prevent the helicase from binding to the polynucleotide. Instead, they reduce the ability of the polynucleotide to unbind or disassociate from the helicase. In other words, one or more modifications increase the processivity of the helicase by preventing dissociation from the polynucleotide strand. The thermal stability of the enzyme is also typically increased by one or more modifications, improving structural stability, which is beneficial for strand sequencing.
[0271] An unnatural amino acid is an amino acid that is not naturally found in helicases. The unnatural amino acid is preferably not histidine, alanine, isoleucine, arginine, leucine, asparagine, lysine, aspartic acid, methionine, cysteine, phenylalanine, glutamic acid, threonine, glutamine, tryptophan, glycine, valine, proline, serine, or tyrosine. The unnatural amino acid is more preferably not one of the 20 amino acids in the previous sentence or selenocysteine.
[0272] Preferred unnatural amino acids for use in the present invention include 4-azido-L-phenylalanine (Faz), 4-acetyl-L-phenylalanine, 3-acetyl-L-phenylalanine, 4-acetoacetyl-L-phenylalanine, O-allyl-L-tyrosine, 3-(phenylselanyl)-L-alanine, O-2-propyn-1-yl-L-tyrosine, 4-(dihydroxyboryl)-L-phenylalanine, 4-[(ethylsulfanyl)carbonyl]-L-phenylalanine, (2S)-2-amino-3-4-[ (Propan-2-ylsulfanyl)carbonyl]phenyl;Propanoic acid, (2S)-2-amino-3-4-[(2-amino-3-sulfanylpropanoyl)amino]phenyl;Propanoic acid, O-methyl-L-tyrosine, 4-amino-L-phenylalanine, 4-cyano-L-phenylalanine, 3-cyano-L-phenylalanine, 4-fluoro-L-phenylalanine, 4-iodo-L-phenylalanine, 4-bromo-L-phenylalanine, O-(trifluoromethyl)tyrosine, 4-nitro-L-phenylalanine Leucine, 3-hydroxy-L-tyrosine, 3-amino-L-tyrosine, 3-iodo-L-tyrosine, 4-isopropyl-L-phenylalanine, 3-(2-naphthyl)-L-alanine, 4-phenyl-L-phenylalanine, (2S)-2-amino-3-(naphthalen-2-ylamino)propanoic acid, 6-(methylsulfanyl)norleucine, 6-oxo-L-lysine, D-tyrosine, (2R)-2-hydroxy-3-(4-hydroxyphenyl)propanoic acid, (2R)-2-ammoniooctanoate 3- (2,2'-bipyridin-5-yl)-D-alanine, 2-amino-3-(8-hydroxy-3-quinolyl)propanoic acid, 4-benzoyl-L-phenylalanine, S-(2-nitrobenzyl)cysteine, (2R)-2-amino-3-[(2-nitrobenzyl)sulfanyl]propanoic acid, (2S)-2-amino-3-[(2-nitrobenzyl)oxy]propanoic acid, O-(4,5-dimethoxy-2-nitrobenzyl)-L-serine, (2S)-2-amino-6-([(2-nitrobenzyl)oxy]carbonyl;(2S)-2-amino-4-(7-hydroxy-2-oxo-2H-chromen-4-yl)butanoic acid, ... -3-[(6-acetylnaphthalen-2-yl)amino]-2-aminopropanoic acid, 4-(carboxymethyl)phenylalanine, 3-nitro-L-tyrosine, O-sulfo-L-tyrosine, (2R)-6-acetamido-2-ammoniohexanoate, 1-methylhistidine, 2-aminononanoic acid, 2-aminodecanoic acid, L-homocysteine, 5-sulfanylnorvaline, 6-sulfanyl-L-norleucine, 5-(methylsulfanyl)-L-norvaline, N; 6 -[(2R,3R)-3-methyl-3,4-dihydro-2H-pyrrol-2-yl]carbonyl;-L-lysine, N 6 -[(benzyloxy)carbonyl]lysine, (2S)-2-amino-6-[(cyclopentylcarbonyl)amino]hexanoic acid, N 6 -[(Cyclopentyloxy)carbonyl]-L-lysine, (2S)-2-amino-6-[(2R)-tetrahydrofuran-2-ylcarbonyl]amino; Hexanoic acid, (2S)-2-amino-8-[(2R,3S)-3-ethynyltetrahydrofuran-2-yl]-8-oxooctanoic acid, N 6 -(tert-Butoxycarbonyl-L-lysine, (2S)-2-hydroxy-6-([(2-methyl-2-propanyl)oxy]carbonyl;amino)hexanoic acid, N 6 -[(allyloxy)carbonyl]lysine, (2S)-2-amino-6-([(2-azidobenzyl)oxy]carbonyl;amino)hexanoic acid, N 6 -L-prolyl-L-lysine, (2S)-2-amino-6-[(prop-2-yn-1-yloxy)carbonyl]amino; hexanoic acid, and N 6The most preferred unnatural amino acid is 4-azido-L-phenylalanine (Faz), including, but not limited to, -[(2-azidoethoxy)carbonyl]-L-lysine. Table 10 below (separated into two parts) identifies the residues that make up each domain in each Dda homologue (SEQ ID NOS: 118-133). [Table 10-1] [Table 10-2]
[0273] Helicases of the invention preferably include variants of SEQ ID NO: 118 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues D260-P274 and N292-A389), and / or (ii) the pin domain (residues K86-E102), and / or (iii) the 1A domain (residues M1-L85 and V103-K177). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues N292-A389 of the tower domain.
[0274] Helicases of the invention preferably include variants of SEQ ID NO: 119 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues G295-N309 and F316-Y421), and / or (ii) the pin domain (residues Y85-L112), and / or (iii) the 1A domain (residues M1-I84 and R113-Y211). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues F316-Y421 of the tower domain.
[0275] Helicases of the invention preferably include variants of SEQ ID NO: 120 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues V328-P342 and N360-Y448), and / or (ii) the pin domain (residues K148-N165), and / or (iii) the 1A domain (residues M1-L147 and S166-V240). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues N360-Y448 of the tower domain.
[0276] Helicases of the invention preferably include variants of SEQ ID NO: 121 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues A261-T275 and T285-Y370), and / or (ii) the pin domain (residues G91-E107), and / or (iii) the 1A domain (residues M1-L90 and E108-H173). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues T285-Y370 of the tower domain.
[0277] Helicases of the invention preferably include variants of SEQ ID NO: 122 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues G294-1307 and T314-Y407), and / or (ii) the pin domain (residues G116-T135), and / or (iii) the 1A domain (residues M1-L115 and N136-V205). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues T314-Y407 of the tower domain.
[0278] Helicases of the invention preferably include variants of SEQ ID NO: 123 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues V288-E301 and N307-N393), and / or (ii) the pin domain (residues G97-P113), and / or (iii) the 1A domain (residues M1-L96 and F114-V194). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues N307-N393 of the tower domain.
[0279] Helicases of the invention preferably include variants of SEQ ID NO: 124 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues S250-P264 and E278-S371), and / or (ii) the pin domain (residues K78-E95), and / or (iii) the 1A domain (residues M1-L77 and V96-V166). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues E278-S371 of the tower domain.
[0280] Helicases of the invention preferably include variants of SEQ ID NO: 125 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues K255-P269 and T284-S380), and / or (ii) the pin domain (residues K82-K98), and / or (iii) the 1A domain (residues M1-M81 and L99-M171). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues T284-S380 of the tower domain.
[0281] Helicases of the invention preferably include variants of SEQ ID NO: 126 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues D242-P256 and T271-S366), and / or (ii) the pin domain (residues K69-K85), and / or (iii) the 1A domain (residues M1-M68 and M86-M158). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues T271-S366 of the tower domain.
[0282] Helicases of the invention preferably include variants of SEQ ID NO: 127 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues T263-P277 and N295-P392), and / or (ii) the pin domain (residues K88-K107), and / or (iii) the 1A domain (residues M1-L87 and A108-M181). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues N295-P392 of the tower domain.
[0283] Helicases of the invention preferably include variants of SEQ ID NO: 128 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues D263-P277 and N295-A391), and / or (ii) the pin domain (residues K88-K107), and / or (iii) the 1A domain (residues M1-L87 and A108-M181). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues N295-A391 of the tower domain.
[0284] Helicases of the invention preferably include variants of SEQ ID NO: 129 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues A258-P272 and N290-P386), and / or (ii) the pin domain (residues K86-G102), and / or (iii) the 1A domain (residues M1-L85 and T103-K176). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues N290-P386 of the tower domain.
[0285] Helicases of the invention preferably include variants of SEQ ID NO: 130 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues L266-P280 and N298-A392), and / or (ii) the pin domain (residues K92-D108), and / or (iii) the 1A domain (residues M1-L91 and V109-M183). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues N298-A392 of the tower domain.
[0286] Helicases of the invention preferably include variants of SEQ ID NO: 131 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues D262-P276 and N294-A392), and / or (ii) the pin domain (residues K88-E104), and / or (iii) the 1A domain (residues M1-L87 and M105-M179). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues N294-A392 of the tower domain.
[0287] Helicases of the invention preferably include variants of SEQ ID NO: 132 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues D261-P275 and N293-A389), and / or (ii) the pin domain (residues K87-E103), and / or (iii) the 1A domain (residues M1-L86 and V104-K178). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues N293-A389 of the tower domain.
[0288] Helicases of the invention preferably include variants of SEQ ID NO: 133 in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into (i) the tower domain (residues E261-P275 and T293-A390), and / or (ii) the pin domain (residues K87-E103), and / or (iii) the 1A domain (residues M1-L86 and V104-M178). The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into residues T293-A390 of the tower domain.
[0289] The helicase of the present invention preferably comprises a variant of any one of SEQ ID NOs: 118-133, in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into each of (i) the tower domain, and (ii) the pin domain and / or the 1A domain. The helicase of the present invention more preferably comprises a variant of any one of SEQ ID NOs: 118-133, in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into each of (i) the tower domain, (ii) the pin domain, and (iii) the 1A domain. Any number and combination of cysteine residues and unnatural amino acids can be introduced, as discussed above.
[0290] The helicase of the present invention preferably comprises (i) E94C and / or A360C, (ii) E93C and / or K358C, (iii) E93C and / or A360C, (iv) E93C and / or E361C, (v) E93C and / or K364C, (vi) E94C and / or L354C, (vii) E94C and / or K358C, (viii) E93C and / or L354C, (ix) E94C and / or E361C, (x) E94C and / or K364C, (xi) L97C and / or L354C, (xii) L97C and / or K358C, (xiii) L97C and / or A360C, (xiv) L97C and / or E361C, (xv) L97C and / or K364C, (xvi) K123C and / or L354C, (xvii) K123C and / or K358C, (xviii) K123C and / or A360C, (xix) K123C and / or E361C, (xx) K123C and / or K364C, (xxi) N155C and / or L354C, (xxii) N155C and / or K358C , (xxiii) N155C and / or A360C, (xxiv) N155C and / or E361C, (xxv) N155C and / or K364C, (xxvi) any of (i) to (xxv) and G357C, (xxvii) any of (i) to (xxv) and Q100C, (xxviii) any of (i) to (xxv) and I127C, (xxix) any of (i) to (xxv) and Q100C and I127C, (xxx) E94C and / or F377C, (xxxi) N95C, (xxxii) T91C, (xxxi ii) Y92L, E94Y, Y350N, A360C, and Y363N, (xxxiv) E94Y and A360C, (xxxv) A360C, (xxxvi) Y92L, E94C, Y350N, A360Y, and Y363N, (xxxvii) Y92L, E94C, and A360Y, (xxxviii) E94C and / or A360C and F276A, (xxxix) E94C and / or L356C, (xl) E93C and / or E356C, (xli) E93C and / or G357C, (xlii) E93C and / or A360C,(xliii) N95C and / or W378C, (xliv) T91C and / or S382C, (xlv) T91C and / or W378C, (xlvi) E93C and / or N353C, (xlvii) E93C and / or S382C, (xlviii) E93C and / or K381C, (xlix) E93C and / or D379C, (l) E93C and / or S375C, (li) E93C and / or W378C, (lii) E93C and / or W374C, (liii) E (liv) E94C and / or N353C, (liv) E94C and / or S382C, (lv) E94C and / or K381C, (lvi) E94C and / or D379C, (lvii) E94C and / or S375C, (lviii) E94C and / or W378C, (lix) E94C and / or W374C, (lx) E94C and A360Y, (lxi) E94C, G357C, and A360C, or (lxii) T2C, E94C, and A360C. In any one of (i) to (lxii), and / or are preferably and.
[0291] Helicases of the present invention preferably comprise a variant of any one of SEQ ID NOs: 119 to 133, which comprises a cysteine residue at a position corresponding to that in SEQ ID NO: 118 as defined in any one of (i) to (lxii). The positions in any one of SEQ ID NOs: 119 to 133 that correspond to those in SEQ ID NO: 118 can be identified using the alignment of SEQ ID NOs: 118 to 133 below. Helicases of the present invention preferably comprise a variant of SEQ ID NO: 92, which comprises (a) D99C and / or L341C, (b) Q98C and / or L341C, or (d) Q98C and / or A340C. Helicases of the present invention preferably comprise a variant of SEQ ID NO: 96, which comprises D90C and / or A349C. Helicases of the present invention preferably comprise a variant of SEQ ID NO: 102, which comprises D96C and / or A362C.
[0292] The helicase of the present invention preferably comprises a variant of any one of SEQ ID NOs: 118-133 as defined in any one of (i) to (lxii), in which Faz is introduced in place of cysteine at one or more of the specific positions. Faz can be introduced in place of cysteine at each of the specific positions. The helicase of the present invention preferably comprises a variant of SEQ ID NO: 118, including: (i) E94Faz and / or A360C, (ii) E94C and / or A360Faz, (iii) E94Faz and / or A360Faz, (iv) Y92L, E94Y, Y350N, A360Faz, and Y363N, (v) A360Faz, (vi) E94Y and A360Faz, (vii) Y92L, E94Faz, Y350N, A360Y, and Y363N, (viii) Y92L, E94Faz, and A360Y, (ix) E94Faz and A360Y, and (x) E94C, G357Faz, and A360C.
[0293] Helicases of the present invention preferably further comprise one or more single amino acid deletions from the pin domain. Any number of single amino acid deletions, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more, can be made. More preferably, the helicase comprises a variant of SEQ ID NO: 118 comprising a deletion of E93, a deletion of E95, or a deletion of E93 and E95. More preferably, the helicase comprises a variant of SEQ ID NO: 118 comprising: (a) a deletion of E94C, a deletion of N95, and A360C; (b) a deletion of E93, a deletion of E94, a deletion of N95, and A360C; (c) a deletion of E93, a deletion of E94C, a deletion of N95, and A360C; or (d) a deletion of E93C, a deletion of N95, and A360C. The helicase of the present invention preferably comprises a variant of any one of SEQ ID NOs: 119 to 133, which variant comprises a deletion of a position corresponding to E93 in SEQ ID NO: 118, a deletion of a position corresponding to E95 in SEQ ID NO: 118, or a deletion of positions corresponding to E93 and E95 in SEQ ID NO: 118.
[0294] The helicase of the present invention preferably further comprises one or more single amino acid deletions from the hook domain. Any number of single amino acid deletions can be made, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. The helicase more preferably comprises a variant of SEQ ID NO: 118 comprising any number of deletions from positions T278 to S287. The helicase more preferably comprises a variant of SEQ ID NO: 118 comprising: (a) a deletion of E94C, Y279 to K284, and A360C; (b) a deletion of E94C, T278, Y279, V286, and S287, and A360C; (c) a deletion of E94C, I281, and K284, and a replacement with a single G, and A360C; (d) a deletion of E94C, K280, and P2845, and a replacement with a single G, and A360C; or (e) a deletion of Y279 to K284, E94C, F276A, and A230C. The helicase of the present invention preferably comprises a variant of any one of SEQ ID NOs: 119 to 133, comprising deletions of any number of positions corresponding to 278 to 287 in SEQ ID NO: 118.
[0295] The helicase of the present invention preferably further comprises one or more single amino acid deletions from the pin domain and one or more single amino acid deletions from the hook domain.
[0296] The helicases of the present invention are preferably helicases further comprising at least one cysteine residue and / or at least one unnatural amino acid introduced into the hook domain and / or the 2A (RecA-like) domain. Any number and combination of cysteine residues and unnatural amino acids can be introduced as discussed above for the tower, pin, and 1A domains.
[0297] Helicases of the present invention preferably include variants of SEQ ID NO: 118, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues L275 to F291) and / or the 2A (RecA-like) domain (residues R178 to T259 and L390 to V439).
[0298] Helicases of the present invention preferably include variants of SEQ ID NO: 119, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues A310 to L315) and / or the 2A (RecA-like) domain (residues R212 to E294 and G422 to S678).
[0299] Helicases of the present invention preferably include variants of SEQ ID NO: 120, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues V343 to L359) and / or the 2A (RecA-like) domain (residues R241 to N327 and A449 to G496).
[0300] Helicases of the present invention preferably include variants of SEQ ID NO: 121, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues W276 to L284) and / or the 2A (RecA-like) domain (residues R174 to D260 and A371 to V421).
[0301] Helicases of the present invention preferably include variants of SEQ ID NO: 122, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues R308 to Y313) and / or the 2A (RecA-like) domain (residues R206 to K293 and I408 to L500).
[0302] Helicases of the present invention preferably include variants of SEQ ID NO: 123, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues M302 to W306) and / or the 2A (RecA-like) domain (residues R195 to D287 and V394 to Q450).
[0303] Helicases of the present invention preferably include variants of SEQ ID NO: 124, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues V265 to I277) and / or the 2A (RecA-like) domain (residues R167 to T249 and L372 to N421).
[0304] Helicases of the present invention preferably include variants of SEQ ID NO: 125, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues V270 to F283) and / or the 2A (RecA-like) domain (residues R172 to T254 and L381 to K434).
[0305] Helicases of the present invention preferably include variants of SEQ ID NO: 126, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues V257 to F270) and / or the 2A (RecA-like) domain (residues R159 to T241 and L367 to K420).
[0306] Helicases of the present invention preferably include variants of SEQ ID NO: 127, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues L278 to Y294) and / or the 2A (RecA-like) domain (residues R182 to T262 and L393 to V443).
[0307] Helicases of the present invention preferably include variants of SEQ ID NO: 128, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues L278 to Y294) and / or the 2A (RecA-like) domain (residues R182 to T262 and L392 to V442).
[0308] Helicases of the present invention preferably include variants of SEQ ID NO: 129 in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues L273 to F289) and / or the 2A (RecA-like) domain (residues R177 to N257 and L387 to V438).
[0309] Helicases of the present invention preferably include variants of SEQ ID NO: 130, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues L281 to F297) and / or the 2A (RecA-like) domain (residues R184 to T265 and L393 to I442).
[0310] Helicases of the present invention preferably include variants of SEQ ID NO: 131, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues H277 to F293) and / or the 2A (RecA-like) domain (residues R180 to T261 and L393 to V442).
[0311] Helicases of the present invention preferably include variants of SEQ ID NO: 132, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues L276 to F292) and / or the 2A (RecA-like) domain (residues R179 to T260 and L390 to I439).
[0312] Helicases of the invention preferably include variants of SEQ ID NO: 133, in which at least one cysteine residue and / or at least one unnatural amino acid has been further introduced into the hook domain (residues L276 to F292) and / or the 2A (RecA-like) domain (residues R179 to T260 and L391 to V441).
[0313] The helicase of the present invention preferably comprises a variant of SEQ ID NO: 118 comprising one or more of (i) I181C, (ii) Y279C, (iii) I281C, and (iv) E288C. The helicase may comprise any combination of (i) to (iv), for example, (i), (ii), (iii), (iv), (i) and (ii), (i) and (iii), (i) and (iv), (ii) and (iii), (ii) and (iv), (iii) and (iv), or (i), (ii), (iii), and (iv). The helicase more preferably comprises a variant of SEQ ID NO: 118 comprising (a) E94C, I281C, and A360C, or (b) E94C, I281C, G357C, and A360C. The helicase of the present invention preferably comprises a variant of any one of SEQ ID NOs: 119-133, which comprises a cysteine residue at one or more of the positions corresponding to those in SEQ ID NO: 118 as defined in (i)-(iv), (a) and (b). The helicase may comprise any of these variants in which Faz is introduced at one or more of the specified positions (or each specified position) in place of cysteine.
[0314] The helicase of the present invention is further modified to reduce its surface negative charge. Surface residues can be identified in the same manner as for the Dda domain disclosed above. The surface negative charge is typically a surface negatively charged amino acid, such as aspartic acid (D) or glutamic acid (E).
[0315] The helicase is preferably modified to neutralize one or more surface negative charges by substituting one or more negatively charged amino acids with one or more positively charged amino acids, uncharged amino acids, nonpolar amino acids, and / or aromatic amino acids, or preferably by introducing one or more positively charged amino acids adjacent to one or more negatively charged amino acids. Suitable positively charged amino acids include, but are not limited to, histidine (H), lysine (K), and arginine (R). Uncharged amino acids have no net charge. Suitable uncharged amino acids include, but are not limited to, cysteine (C), serine (S), threonine (T), methionine (M), asparagine (N), and glutamine (Q). Nonpolar amino acids have nonpolar side chains. Suitable nonpolar amino acids include, but are not limited to, glycine (G), alanine (A), proline (P), isoleucine (I), leucine (L), and valine (V). Aromatic amino acids have an aromatic side chain. Suitable aromatic amino acids include, but are not limited to, histidine (H), phenylalanine (F), tryptophan (W), and tyrosine (Y).
[0316] Preferred substitutions include, but are not limited to, substitution of E for R, substitution of E for K, substitution of E for N, substitution of D for K, and substitution of D for R.
[0317] The helicase of the present invention preferably comprises a variant of SEQ ID NO: 118, wherein the one or more negatively charged amino acids correspond to one or more of D5, E8, E23, E47, D167, E172, D202, D212, and E273. Any number of these amino acids may be neutralized, for example, 1, 2, 3, 4, 5, 6, 7, or 8 amino acids. Any combination may be neutralized. The helicase of the present invention preferably comprises a variant of any one of SEQ ID NOs: 119 to 133, wherein the one or more negatively charged amino acids correspond to one or more of D5, E8, E23, E47, D167, E172, D202, D212, and E273 in SEQ ID NO: 118. The amino acids in SEQ ID NOs: 119 to 133 corresponding to D5, E8, E23, E47, D167, E172, D202, D212, and E273 in SEQ ID NO: 118 can be determined using the alignment in WO 2015 / 055981. The helicase of the present invention preferably comprises a variant of SEQ ID NO: 118 comprising (a) E94C, E273G, and A360C, or (b) E94C, E273G, N292G, and A360C.
[0318] The helicase of the present invention is preferably further modified by removal of one or more naturally occurring cysteine residues. Any number of naturally occurring cysteine residues may be removed. The one or more cysteine residues are preferably removed by substitution. The one or more cysteine residues are preferably substituted with alanine (A), serine (S), or valine (V). The helicase of the present invention preferably comprises a variant of SEQ ID NO: 118, wherein the one or more naturally occurring cysteine residues are one or more of C109, C114, C136, C171, and C412. Any number and combination of these cysteine residues may be removed. For example, variants of SEQ ID NO: 118 include: C109; C114; C136; C171; C412; C109 and C114; C109 and C136; C109 and C171; C109 and C412; C114 and C136; C114 and C171; C114 and C412; C136 and C171; C136 and C412; C171 and C412; C109, C114, and C136; C109, C114, and C171; C109, C114, and C412; C109, C136, and C171; C109, C136, and C412; C109, C171, and C412; C114, C136, and C171; C114, C136, and C412; C114, C171, and C412; C136, C171, and C412; C109, C114, C136, and C171; C109, C114, C136, and C412; C109, C114, C171, and C412; C109, C136, C171, and C412; C114, C136, C171, and C412; or C109, C114, C136, C171, and C412.
[0319] The modified helicase preferably includes a modification or substitution at position(s) corresponding to amino acid position(s) 109 and / or 136 in Dda1993, which removes one or two cysteine residues. This may be in addition to a modification or substitution at one or more positions corresponding to amino acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in Dda1993, a modification or substitution at one or more positions corresponding to amino acid positions 114, 177, 350, and 358 in Dda1993, and / or a modification or substitution at a position corresponding to position 40 in Dda1993. Position 109 or a corresponding position may be substituted with A, V, I, L, M, F, Y, or W. Position 109 or a corresponding position is preferably substituted with A or V. Position 136 or a corresponding position may be substituted with A, V, I, L, M, F, Y, or W. Position 136, or a corresponding position, is preferably substituted with A or V. Helicases of the invention preferably comprise variants of SEQ ID NO: 118 comprising substitutions with C109, e.g., C109A, C109V, C109I, C109L, C109M, C109F, C109Y, or C109W, and / or C136, e.g., C136A, C136V, C136I, C136L, C136M, C136F, C136Y, or C136W. Helicases of the invention preferably comprise variants of SEQ ID NO: 118 comprising C109A and / or C136A. Helicases of the present invention preferably include variants of SEQ ID NO: 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, or 133, which contain substitutions at position(s) corresponding to C109 and / or C136 in SEQ ID NO: 118. Helicases of the present invention are preferably helicases in which at least one cysteine residue (i.e., one or more cysteine residues) and / or at least one unnatural amino acid (i.e., one or more unnatural amino acids) have been introduced only into the tower domain. Suitable modifications are discussed above.
[0320] Helicases of the present invention preferably include variants of SEQ ID NO: 118 comprising the following mutations: E93C and K364C; E94C and K364C; E94C and A360C; L97C and E361C; L97C and E361C and C412A; K123C and E361C; K123C, E361C, and C412A; N155C and K358C; N155C, K358C, and C412A; N155C and L354C; N155C, L354C, and C412A; deltaE93, E94C, deltaN95, and A360C; E94C, deltaN95, and A360C; E94C, Q100C, I127C, and A360C; L354C; G357C; E94C, G357C, and A360C; E94C, Y279C, and A360C; E94C, I281C, and A360C; E94C, Y279Faz, and A360C; Y279C and G357C; I281C and G357C; E94C, Y279C, G357C, and A360C; E94C, I281C, G357C, and A360C; E8R, E47K, E94C, D202K, and A360C; D5K, E23N, E94C, D167K, E1 72R, D212R, and A360C; D5K, E8R, E23N, E47K, E94C, D167K, E172R, D202K, D212R, and A360C; E94C, C114A, C171A, A360C, and C412D; E94C, C114A, C171A, A360C, and C412S; E94C, C109A, C136A, and A360C; E94C, C109A, C114A, C136A, C171A, A360C, and C412S; E94C, C109V, C114V, C171A, A360 C, and C412S; C109A, C114A, C136A, G153C, C171A, E361C, and C412A; C109A, C114A, C136A, G153C, C171A, E361C, and C412D; C109A, C114A, C136A, G153C, C171A, E361C, and C412S; C109A, C114A, C136A, G153C, C171A, K358C, and C412A; C109A, C114A, C136A, G153C, C171A, K358C, and C412D;C109A, C114A, C136A, G153C, C171A, K358C, and C412S; C109A, C114A, C136A, N155C, C171A, K358C, and C412A; C109A, C114A, C136A, N155C, C171A, K358C, and C412D; C109A, C114A, C136A, N155C, C171A, K358C, and C412S; C109A, C114A, C136A, N155C, C171A, L354C, and C412A ;C109A, C114A, C136A, N155C, C171A, L354C, and C412D;C109A, C114A, C136A, N155C, C171A, L354C, and C412S;C109A, C114A, K123C, C136A, C171A, E361C, and C412A;C109A, C114A, K123C, C136A, C171A, E361C, and C412D;C109A, C114A, K123C, C136A, C171A, E361C, and C412 S; C109A, C114A, K123C, C136A, C171A, K358C, and C412A; C109A, C114A, K123C, C136A, C171A, K358C, and C412D; C109A, C114A, K123C, C136A, C171A, K358C, and C412S; C109A, C114A, C136A, G153C, C171A, E361C, and C412A; E94C, C109A, C114A, C136A, C171A, A360C, and C412 D; E94C, C109A, C114V, C136A, C171A, A360C, and C412D; E94C, C109V, C114A, C136A, C171A, A360C, and C412D; L97C, C109A, C114A, C136A, C171A, E361C, and C412A; L97C, C109A, C114A, C136A, C171A, E361C, and C412D; or L97C, C109A, C114A, C136A, C171A, E361C, and C412S. ;
[0321] Modifications in the hook domain and / or 2A domain In one embodiment, the helicase of the present invention is a helicase in which at least one cysteine residue and / or at least one unnatural amino acid has been introduced into the hook domain and / or the 2A (RecA-like motor) domain, and the helicase is capable of controlling the movement of a polynucleotide. The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced into the hook domain and the 2A (RecA-like motor) domain.
[0322] Any number of cysteine residues and / or unnatural amino acids can be introduced into each domain. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more cysteine residues can be introduced and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more unnatural amino acids can be introduced. Only one or more cysteine residues can be introduced. Only one or more unnatural amino acids can be introduced. A combination of one or more cysteine residues and one or more unnatural amino acids can be introduced.
[0323] The at least one cysteine residue and / or at least one unnatural amino acid is preferably introduced by substitution. Methods for doing this are known in the art. Suitable modifications of the hook domain and / or 2A (RecA-like motor) domain are discussed above.
[0324] The helicase of the present invention is preferably a variant of SEQ ID NO: 118 comprising (a) Y279C, I181C, E288C, Y279C, and I181C, (b) Y279C and E288C, (c) I181C and E288C, or (d) Y279C, I181C, and E288C. The helicase of the present invention preferably comprises a variant of any one of SEQ ID NOs: 199 to 133 comprising a mutation at one or more of the positions corresponding to those in SEQ ID NO: 118 as defined in (a) to (d).
[0325] surface modification In one embodiment, the helicase is modified to reduce its surface negative charge, and the helicase has the ability to control polynucleotide movement. Suitable modifications are discussed above. Any number of surface negative charges can be neutralized.
[0326] Helicases of the invention preferably comprise variants of SEQ ID NO: 118 comprising the following mutations: E273G; E8R, E47K, and D202K; D5K, E23N, D167K, E172R, and D212R; or D5K, E8R, E23N, E47K, D167K, E172R, D202K, and D212R.
[0327] Other modified helicases In one embodiment, the helicase of the invention comprises a variant of SEQ ID NO: 118, including: A360K; Y92L and / or A360Y; Y92L, Y350N, and Y363N; Y92L and / or Y363N; or Y92L.
[0328] Other modifications In addition to the specific mutations disclosed above, variants of SEQ ID NO: 118 may include one or more of the following mutations: K38A, T91F, T91N, T91Q, T91W, V96E, V96F, V96L, V96Q, V96R, V96W, V96Y, P274G, V286F, V286W, V286Y, F291G, N292F, N292G, N292P, N292Y, G294Y, G294F, K364A, and W378A.
[0329] In addition to the specific mutations disclosed above, variants of SEQ ID NO: 118 may include: K38A, E94C, and A360C; H64K; E94C and A360C; H64N; E94C and A360C; H64Q; E94C and A360C; H64S; E94C and A360C; H64W, E94C, and A360C; T80K, E94C, and A360C; T80K, S83K, E94C, N242K, N293K, and A360C; T80K, S83K, E94C, N242K, N293K, A360C, and T394K; T80K, S83K, E94C, N293K, and A360C; T80K, S83K, E 94C, A360C, and T394K; T80K, S83K, E94C, A360C, and T394N; T80K, E94C, N242K, and A360C; T80K, E94C, N242K, N293K, and A360C; T80K, E94C, N293K, and A360C; T80N, E94C, and A360C; H82A, E94C, and A360C; H82A, P89A, E94C, F98A, and A360C; H82F, E94C, and A360C; H82Q, E94C, A360C; H82R, E94C, and A36 0C;H82W, E94C, and A360C;H82W, P89W, E94C, F98W, and A360C;H82Y, E94C, and A360C;S83K, E94C, and A360C;S83K, T80K, E94C, A360C, and T394K;S83N, E94C, and A360C;S83T, E94C, and A360C;N88H, E94C, and A360C;N88Q, E94C, and A360C;P89A, E94C, and A360C;P89A, F98W, E94C, and A360C;P89A, E94C, F 98Y, and A360C; P89A, E94C, F98A, and A360C; P89F, E94C, and A360C; P89S, E94C, and A360C; P89T, E94C, and A360C; P89W, E94C, F98W, and A360C; P89Y, E94C, and A360C; T91F, E94C, and A360C; T91N, E94C, and A360C; T91Q, E94C, and A360C; T91W, E94C, and A360C; E94C, V96E, and A360C; E94C, V96F, and A360C;E94C, V96L, and A360C; E94C, V96Q, and A360C; E94C, V96R, and A360C; E94C, V96W, and A360C; E94C, V96Y, and A360C; E94C, F98A, and A360C; E94C, F98L, and A360C; E94C, F98V, and A360C; E94C, F98Y, and A360C; E94C; F98W and A360C; E94C, V150A, and A360C; E94C, V150F, and A360C; E94C, V150I, and A360C; E94C, V150K, and A360C; E94C, V150L, and A360C; E94C, V150S, and A360C; E94C, V150T, and A360C; E94C, V150W, and A360C; E94C, V150Y, and A360C; E94C, F240Y, and A360C; E94C, F240W, and A360C; E94C, N242K, and A360C; E94C, N242K, N293K, and A360C; E94C, P274G, and A360C; E94C, L275G, and A360C; E94C, F276A, and A360C; E94C, F276I , and A360C;E94C, F276M, and A360C;E94C, F276V, and A360C;E94C, F276W, and A360C;E94C, F276Y, and A360C;E94C, V286F, and A360C;E94C, V286W, and A360C;E94C, V286Y, and A360C;E94C, S287F, and A360C;E94C, S287W, and A360C;E94C, S287Y, and A360C;E94C, F291G, and A360C;E94C, N292F, and A360C;E94C, N292G, and A360C; E94C, N292P, and A360C; E94C, N292Y, and A360C; E94C, N293F, and A360C; E94C, N293K, and A360C; E94C, N293Q, and A360C; E94C, N293Y, and A360C; E94C, G294F, and A360C; E94C, G294Y, and A360C; E94C, A36C, and K364A; E94C, A360C, W378A; E94C, A360C, and T394K; E94C, A360C, and H396Q; E94C, A360C, and H396S;E94C, A360C, and H396W; E94C, A360C, and Y415F; E94C, A360C, and Y415K; E94C, A360C, and Y415M; or E94C, A360C, and Y415W.
[0330] The helicase of the present invention preferably comprises a variant of SEQ ID NO: 118 comprising: (a) E94C / A360C / W378A, or (b) E94C / A360C / C109A / C136A / W378A, or (d) E94C / A360C / C109A / C136A / W378A, then (ΔM1)G1G2 (i.e., deletion of M1, then addition of G1 and G2).
[0331] A preferred variant of any one of SEQ ID NOs: 118-133 has (in addition to the modifications of the invention) an N-terminal methionine (M) replaced with one glycine residue (G). In the examples, this is shown as (ΔM1)G1. This may also be referred to as M1G. Any of the variants discussed above may further comprise M1G.
[0332] Most preferred helicases of the present invention comprise variants of SEQ ID NO: 118, including (a) E94C / F98W / A360C / C109A / C136A / K194L, (b) M1G / E94C / F98W / A360C / C109A / C136A / K194L, (c) E94C / F98W / A360C / C109A / C136A / K199L, or (d) M1G / E94C / F98W / A360C / C109A / C136A / K199L.
[0333] Other Preferred Helicases Other preferred helicases of the present invention include variants of SEQ ID NO: 118 comprising the following substitutions: T40 / E94 / F98, for example, T40Y / E94C / F98W, T40 / E94 / F98 / C114, for example, T40Y / E94C / F98W / C114I, T40 / E94 / F98 / K177, for example, T40Y / E94C / F98W / K177M, T40 / E94 / F98 / Y350, for example, T40Y / E94C / F98W / Y350, T40Y / E94C / F98W / Y350I, or T40Y / E94C / F98W / Y350E, T40 / E94 / F98 / K358, for example, T40Y / E94C / F98W / K358I, E94 / F98 / C114, for example, E94C / F98W / C114I, E94 / F98 / K177, for example, E94C / F98W / K177M, E94 / F98 / Y350, for example, E94C / F98W / Y350I or E94C / F98W / Y350E, E94 / F98 / K358, for example, E94C / F98W / K358I, T40 / E94 / F98 / C114 / K177, for example, T40Y / E94C / F98W / C114I / K177M, T40 / E94 / F98 / C114 / Y350, for example, T40Y / E94C / F98W / C114I / Y350I, or T40Y / E94C / F98W / C114I / Y350E, T40 / E94 / F98 / C114 / K358, for example, T40Y / E94C / F98W / C114I / K358I, T40 / E94 / F98 / K177 / Y350, for example, T40Y / E94C / F98W / K177M / Y350I, or T40Y / E94C / F98W / K177M / Y350E, T40 / E94 / F98 / K177 / K358, for example, T40Y / E94C / F98W / K177M / K358I, T40 / E94 / F98 / Y350 / K358, for example, T40Y / E94C / F98W / Y350I / K358I, or T40Y / E94C / F98W / Y350E / K358I, E94 / F98 / C114 / K177, for example, E94C / F98W / C114I / K177M, E94 / F98 / C114 / Y350, for example, E94C / F98W / C114I / Y350I, or E94C / F98W / C114I / Y350E, E94 / F98 / C114 / K358, for example, E94C / F98W / C114I / K358I, E94 / F98 / K177 / Y350, for example, E94C / F98W / K177M / Y350I or E94C / F98W / K177M / Y350E, E94 / F98 / K177 / K358, for example, E94C / F98W / K177M / K358I, E94 / F98 / Y350 / K358, for example, E94C / F98W / Y350I / K358I, or E94C / F98W / Y350E / K358I, T40 / E94 / F98 / C114 / K177 / Y350, for example, T40Y / E94C / F98W / C114I / K177M / Y350I, or T40Y / E94C / F98W / C114I / K177M / Y350E, T40 / E94 / F98 / C114 / K177 / K358, for example, T40Y / E94C / F98W / C114I / K177M / K358I, T40 / E94 / F98 / C114 / Y350 / K358, for example, T40Y / E94C / F98W / C114I / Y350I / K358I, or T40Y / E94C / F98W / C114I / Y350E / K358I, T40 / E94 / F98 / K177 / Y350 / K358, for example, T40Y / E94C / F98W / K177M / Y350I / K358I, or T40Y / E94C / F98W / K177M / Y350E / K358I, E94 / F98 / C114 / K177 / Y350, for example, E94C / F98W / C114I / K177M / Y350I, or E94C / F98W / C114I / K177M / Y350E, E94 / F98 / C114 / K177 / K358, for example, E94C / F98W / C114I / K177M / K358I, E94 / F98 / C114 / Y350 / K358, for example, E94C / F98W / C114I / Y350I / K358I, or E94C / F98W / C114I / Y350E / K358I, E94 / F98 / K177 / Y350 / K358, for example, E94C / F98W / K177M / Y350I / K358I, or E94C / F98W / K177M / Y350E / K358I, T40 / E94 / F98 / C114 / K177 / Y350 / K358, for example, T40Y / E94C / F98W / C114I / K177M / Y350I / K358I, or T40Y / E94C / F98W / C114I / K177M / Y350E / K358I, E94 / F98 / C114 / K177 / Y350 / K358, for example, E94C / F98W / C114I / K177M / Y350I / K358I, or E94C / F98W / C114I / K177M / Y350E / K358I, T40 / F98 / K194, for example, T40Y / E94C / F98W / K194L, T40 / F98 / C114 / K194, for example, T40Y / E94C / F98W / C114I K194L, T40 / F98 / K177 / K194, for example, T40Y / E94C / F98W / K177M K194L, T40 / F98 / K194 / Y350, for example, T40Y / E94C / F98W / K194L / Y350I, or T40Y / E94C / F98W / K194L / Y350E, T40 / F98 / K194 / K358, for example, T40Y / E94C / F98W / K194L / K358I, F98 / C114 / K194, for example, E94C / F98W / C114I / K194L, F98 / K177 / K194, for example, E94C / F98W / K177M / K194L, F98 / K194 / Y350, for example, E94C / F98W / K194L / Y350I, or E94C / F98W / K194L / Y350E, F98 / K194 / K358, for example, E94C / F98W / K194L / K358I, T40 / F98 / C114 / K177 / K194, for example, T40Y / E94C / F98W / C114I / K177M / K194L, T40 / F98 / C114 / K194 / Y350, for example, T40Y / E94C / F98W / C114I / K194L / Y350I, or T40Y / E94C / F98W / C114I / K194L / Y350E, T40 / F98 / C114 / K194 / K358, for example, T40Y / E94C / F98W / C114I / K194L / K358I, T40 / F98 / K177 / K194 / Y350, for example, T40Y / E94C / F98W / K177M / K194L / Y350I, or T40Y / E94C / F98W / K177M / K194L / Y350E, T40 / F98 / K177 / K194 / K358, for example, T40Y / E94C / F98W / K177M / K194L / K358I, T40 / F98 / K194 / Y350 / K358, for example, T40Y / E94C / F98W / K194L / Y350I / K358I, or T40Y / E94C / F98W / K194L / Y350E / K358I, F98 / C114 / K177 / K194, for example, E94C / F98W / C114I / K177M / K194L, F98 / C114 / K194 / Y350, for example, E94C / F98W / C114I / K194L / Y350I, or E94C / F98W / C114I / K194L / Y350E, F98 / C114 / K194 / K358, for example, E94C / F98W / C114I / K194L / K358I, F98 / K177 / K194 / Y350, for example, E94C / F98W / K177M / K194L / Y350I, or E94C / F98W / K177M / K194L / Y350E, F98 / K177 / K194 / K358, for example, E94C / F98W / K177M / K194L / K358I, F98 / K194 / Y350 / K358, for example, E94C / F98W / K194L / Y350I / K358I, or E94C / F98W / K194L / Y350E / K358I, T40 / F98 / C114 / K177 / K194 / Y350, for example, T40Y / E94C / F98W / C114I / K177M / K194L / Y350I, or T40Y / E94C / F98W / C114I / K177M / K194L / Y350E, T40 / F98 / C114 / K177 / K194 / K358, for example, T40Y / E94C / F98W / C114I / K177M / K194L / K358I, T40 / F98 / C114 / K194 / Y350 / K358, for example, T40Y / E94C / F98W / C114I / K194L / Y350I / K358I, or T40Y / E94C / F98W / C114I / K194L / Y350E / K358I, T40 / F98 / K177 / K194 / Y350 / K358, for example, T40Y / E94C / F98W / K177M / K194L / Y350I / K358I, or T40Y / E94C / F98W / K177M / K194L / Y350E / K358I, F98 / C114 / K177 / K194 / Y350, for example, E94C / F98W / C114I / K177M / K194L / Y350I, or E94C / F98W / C114I / K177M / K194L / Y350E, F98 / C114 / K177 / K194 / K358, for example, E94C / F98W / C114I / K177M / K194L / K358I, F98 / C114 / K194 / Y350 / K358, for example, E94C / F98W / C114I / K194L / Y350I / K358I, or E94C / F98W / C114I / K194L / Y350E / K358I, F98 / K177 / K194 / Y350 / K358, for example, E94C / F98W / K177M / K194L / Y350I / K358I, or E94C / F98W / K177M / K194L / Y350E / K358I, T40 / F98 / C114 / K177 / K194 / Y350 / K358, for example, T40Y / E94C / F98W / C114I / K177M / K194L / Y350I / K358I, or T40Y / E94C / F98W / C114I / K177M / K194L / Y350E / K358I, F98 / C114 / K177 / K194 / Y350 / K358, for example, E94C / F98W / C114I / K177M / K194L / Y350I / K358I, or E94C / F98W / C114I / K177M / K194L / Y350E / K358I, T40 / A360, for example, T40Y / A360C, T40 / C114 / A360, for example, T40Y / C114I / A360C, T40 / K177 / A360, for example, T40Y / K177M / A360C, T40 / Y350 / A360, for example, T40Y / Y350I / A360C, or T40Y / Y350E / A360C, T40 / K358 / A360, for example, T40Y / K358I / A360C, C114 / A360, e.g., C114I / A360C, K177 / A360, for example, K177M / A360C, Y350 / A360, for example, Y350I / A360C, or Y350E / A360C, K358 / A360, for example, K358I / A360C, T40 / C114 / K177 / A360, for example, T40Y / C114I / K177M / A360C, T40 / C114 / Y350 / A360, for example, T40Y / C114I / Y350I / A360C, or T40Y / C114I / Y350E / A360C; T40 / C114 / K358 / A360, for example, T40Y / C114I / K358I / A360C, T40 / K177 / Y350 / A360, for example, T40Y / K177M / Y350I / A360C, or T40Y / K177M / Y350E / A360C, T40 / K177 / K358 / A360, for example, T40Y / K177M / K358I / A360C, T40 / Y350 / K358 / A360, for example, T40Y / Y350I / K358I / A360C, or T40Y / Y350E / K358I / A360C, C114 / K177 / A360, for example, C114I / K177M / A360C, C114 / Y350 / A360, for example, C114I / Y350I / A360C, or C114I / Y350E / A360C; C114 / K358 / A360, for example, C114I / K358I / A360C, K177 / Y350 / A360, for example, K177M / Y350I / A360C, or K177M / Y350E / A360C, K177 / K358 / A360, for example, K177M / K358I / A360C, Y350 / K358 / A360, for example, Y350I / K358I / A360C, or Y350E / K358I / A360C, T40 / C114 / K177 / Y350 / A360, for example, T40Y / C114I / K177M / Y350I / A360C, or T40Y / C114I / K177M / Y350E / A360C, T40 / C114 / K177 / K358 / A360, for example, T40Y / C114I / K177M / K358I / A360C, T40 / C114 / Y350 / K358 / A360, for example, T40Y / C114I / Y350I / K358I / A360C, or T40Y / C114I / Y350E / K358I / A360C, T40 / K177 / Y350 / K358 / A360, for example, T40Y / K177M / Y350I / K358I / A360C, or T40Y / K177M / Y350E / K358I / A360C, C114 / K177 / Y350 / A360, for example, C114I / K177M / Y350I / A360C, or C114I / K177M / Y350E / A360C; C114 / K177 / K358 / A360, for example, C114I / K177M / K358I / A360C, C114 / Y350 / K358 / A360, for example, C114I / Y350I / K358I / A360C, or C114I / Y350E / K358I / A360C; K177 / Y350 / K358 / A360, for example, K177M / Y350I / K358I / A360C, or K177M / Y350E / K358I / A360C, T40 / C114 / K177 / Y350 / K358 / A360, for example, T40Y / C114I / K177M / Y350I / K358I / A360C, or K177M / Y350E / K358I / A360C, C114 / K177 / Y350 / K358 / A360, for example, C114I / K177M / Y350I / K358I / A360C, or C114I / K177M / Y350E / K358I / A360C; T40 / E94 / F98 / C109, for example, T40Y / E94C / F98W / C109A, T40 / E94 / F98 / C109 / C114, for example, T40Y / E94C / F98W / C109A / C114I, T40 / E94 / F98 / C109 / K177, for example, T40Y / E94C / F98W / C109A / K177M, T40 / E94 / F98 / C109 / Y350, for example, T40Y / E94C / F98W / C109A / Y350I, or T40Y / E94C / F98W / C109A / Y350E, T40 / E94 / F98 / C109 / K358, for example, T40Y / E94C / F98W / C109A / K358I, E94 / F98 / C109 / C114, for example, E94C / F98W / C109A / C114I, E94 / F98 / C109 / K177, for example, E94C / F98W / C109A / K177M, E94 / F98 / C109 / Y350, for example, E94C / F98W / C109A / Y350I, or E94C / F98W / C109A / Y350E, E94 / F98 / C109 / K358, for example, E94C / F98W / C109A / K358I, T40 / E94 / F98 / C109 / C114 / K177, for example, T40Y / E94C / F98W / C109A / C114I / K177M, T40 / E94 / F98 / C109 / C114 / Y350, for example, T40Y / E94C / F98W / C109A / C114I / Y350I, or T40Y / E94C / F98W / C109A / C114I / Y350E, T40 / E94 / F98 / C109 / C114 / K358, for example, T40Y / E94C / F98W / C109A / C114I / K358I, T40 / E94 / F98 / C109 / K177 / Y350, for example, T40Y / E94C / F98W / C109A / K177M / Y350I oe T40Y / E94C / F98W / C109A / K177M / Y350E, T40 / E94 / F98 / C109 / K177 / K358, for example, T40Y / E94C / F98W / C109A / K177M / K358I, T40 / E94 / F98 / C109 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / Y350I / K358I, or T40Y / E94C / F98W / C109A / Y350E / K358I, E94 / F98 / C109 / C114 / K177, for example, E94C / F98W / C109A / C114I / K177M, E94 / F98 / C109 / C114 / Y350, for example, E94C / F98W / C109A / C114I / Y350I, or E94C / F98W / C109A / C114I / Y350E, E94 / F98 / C109 / C114 / K358, for example, E94C / F98W / C109A / C114I / K358I, E94 / F98 / C109 / K177 / Y350, for example, E94C / F98W / C109A / K177M / Y350I, or E94C / F98W / C109A / K177M / Y350E, E94 / F98 / C109 / K177 / K358, for example, E94C / F98W / C109A / K177M / K358I, E94 / F98 / C109 / Y350 / K358, for example, E94C / F98W / C109A / Y350I / K358I, or E94C / F98W / C109A / Y350E / K358I, T40 / E94 / F98 / C109 / C114 / K177 / Y350, for example, T40Y / E94C / F98W / C109A / C114I / K177M / Y350I, or T40Y / E94C / F98W / C109A / C114I / K177M / Y350E, T40 / E94 / F98 / C109 / C114 / K177 / K358, for example, T40Y / E94C / F98W / C109A / C114I / K177M / K358I, T40 / E94 / F98 / C109 / C114 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / C114I / Y350I / K358I, or T40Y / E94C / F98W / C109A / C114I / Y350E / K358I, T40 / E94 / F98 / C109 / K177 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / K177M / Y350I / K358I, or T40Y / E94C / F98W / C109A / K177M / Y350E / K358I, E94 / F98 / C109 / C114 / K177 / Y350, for example, E94C / F98W / C109A / C114I / K177M / Y350I, or E94C / F98W / C109A / C114I / K177M / Y350E, E94 / F98 / C109 / C114 / K177 / K358, for example, E94C / F98W / C109A / C114I / K177M / K358I, E94 / F98 / C109 / C114 / Y350 / K358, for example, E94C / F98W / C109A / C114I / Y350I / K358I, or E94C / F98W / C109A / C114I / Y350E / K358I, E94 / F98 / C109 / K177 / Y350 / K358, for example, E94C / F98W / C109A / K177M / Y350I / K358I, or E94C / F98W / C109A / K177M / Y350E / K358I, T40 / E94 / F98 / C109 / C114 / K177 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / C114I / K177M / Y350I / K358I, or T40Y / E94C / F98W / C109A / C114I / K177M / Y350E / K358I, E94 / F98 / C109 / C114 / K177 / Y350 / K358, for example, E94C / F98W / C109A / C114I / K177M / Y350I / K358I, or E94C / F98W / C109A / C114I / K177M / Y350E / K358I, T40 / C109 / C136, for example, T40Y / C109A / C136A, T40 / C109 / C114 / C136, for example, T40Y / C109A / C114I / C136A, T40 / C109 / C136 / K177, for example, T40Y / C109A / C136A / K177M, T40 / C109 / C136 / Y350, for example, T40Y / C109A / C136A / Y350I or T40Y / C109A / C136A / Y350E; T40 / C109 / C136 / K358, for example, T40Y / C109A / C136A / K358I, C109 / C114 / C136, for example, C109A / C136A / C114I, C109 / C136 / K177, for example, C109A / C136A / K177M, C109 / C136 / Y350, for example, C109A / C136A / Y350I or C109A / C136A / Y350E; C109 / C136 / K358, for example, C109A / C136A / K358I, T40 / C109 / C114 / C136 / K177, for example, T40Y / C109A / C114I / C136A / K177M, T40 / C109 / C114 / C136 / Y350, for example, T40Y / C109A / C114I / C136A / Y350, or T40Y / C109A / C114I / C136A / Y350I; T40 / C109 / C114 / C136 / K358, for example, T40Y / C109A / C114I / C136A / K358I, T40 / C109 / C136 / K177 / Y350, for example, T40Y / C109A / C136A / K177M / Y350E, or T40Y / C109A / C136A / K177M / Y350I; T40 / C109 / C136 / K177 / K358, for example, T40Y / C109A / C136A / K177M / K358I, T40 / C109 / C136 / Y350 / K358, for example, T40Y / C109A / C136A / Y350I / K358I, or T40Y / C109A / C136A / Y350E / K358I; C109 / C114 / C136 / K177, for example, C109A / C114I / C136A / K177M, C109 / C114 / C136 / Y350, for example, C109A / C114I / C136A / Y350I, or C109A / C114I / C136A / Y350E, C109 / C114 / C136 / K358, for example, C109A / C114I / C136A / K358I, C109 / C136 / K177 / Y350, for example, C109A / C136A / K177M / Y350I or C109A / C136A / K177M / Y350E; C109 / C136 / K177 / K358, for example, C109A / C136A / K177M / K358I, C109 / C136 / Y350 / K358, for example, C109A / C136A / Y350I / K358I, or C109A / C136A / Y350E / K358I; T40 / C109 / C114 / C136 / K177 / Y350, for example, T40Y / C109A / C114I / C136A / K177M / Y350I, or T40Y / C109A / C114I / C136A / K177M / Y350E, T40 / C109 / C114 / C136 / K177 / K358, for example, T40Y / C109A / C114I / C136A / K177M / K358I, T40 / C109 / C114 / C136 / Y350 / K358, for example, T40Y / C109A / C114I / C136AY350I / K358I, or T40Y / C109A / C114I / C136AY350E / K358I, T40 / C109 / C136 / K177 / Y350 / K358, for example, T40Y / C109A / C136A / K177M / Y350I / K358I, or T40Y / C109A / C136A / K177M / Y350E / K358I; C109 / C114 / C136 / K177 / Y350, for example, C109A / C114I / C136A / K177M / Y350I or C109A / C114I / C136A / K177M / Y350E; C109 / C114 / C136 / K177 / K358, for example, C109A / C114I / C136A / K177M / K358I, C109 / C114 / C136 / Y350 / K358, for example, C109A / C114I / C136AY350I / K358I, or C109A / C114I / C136AY350E / K358I, C109 / C136 / K177 / Y350 / K358, for example, C109A / C136A / K177M / Y350I / K358I, or C109A / C136A / K177M / Y350E / K358I; T40 / C109 / C114 / C136 / K177 / Y350 / K358, for example, T40Y / C109A / C114I / C136A / K177M / Y350I / K358I, or T40Y / C109A / C114I / C136A / K177M / Y350E / K358I, C109 / C114 / C136 / K177 / Y350 / K358, for example, C109A / C114I / C136A / K177M / Y350I / K358I C109A / C114I / C136A / K177M / Y350E / K358I, T40 / E94 / F98 / C109 / K194, for example, T40Y / E94C / F98W / C109A / K194L, T40 / E94 / F98 / C109 / C114 / K194, for example, T40Y / E94C / F98W / C109A / C114I / K194L, T40 / E94 / F98 / C109 / K177 / K194, for example, T40Y / E94C / F98W / C109A / K177M / K194L, T40 / E94 / F98 / C109 / K194 / Y350, for example, T40Y / E94C / F98W / C109A / K194L / Y350I, or T40Y / E94C / F98W / C109A / K194L / Y350E, T40 / E94 / F98 / C109 / K194 / K358, for example, T40Y / E94C / F98W / C109A / K194L / K358I, E94 / F98 / C109 / C114 / K194, for example, E94C / F98W / C109A / C114I / K194L, E94 / F98 / C109 / K177 / K194, for example, E94C / F98W / C109A / K177M / K194L, E94 / F98 / C109 / K194 / Y350, for example, E94C / F98W / C109A / K194L / Y350I, or E94C / F98W / C109A / K194L / Y350E, E94 / F98 / C109 / K194 / K358, for example, E94C / F98W / C109A / K194L / K358I, T40 / E94 / F98 / C109 / C114 / K177 / K194, for example, T40Y / E94C / F98W / C109A / C114I / K194L / K177M, T40 / E94 / F98 / C109 / C114 / K194 / Y350, for example, T40Y / E94C / F98W / C109A / C114I / K194L / Y350I, or T40Y / E94C / F98W / C109A / C114I / K194L / Y350E, T40 / E94 / F98 / C109 / C114 / K194 / K358, for example, T40Y / E94C / F98W / C109A / C114I / K194L / K358I, T40 / E94 / F98 / C109 / K177 / K194 / Y350, for example, T40Y / E94C / F98W / C109A / K177M / K194L / Y350I, or T40Y / E94C / F98W / C109A / K177M / K194L / Y350E, T40 / E94 / F98 / C109 / K177 / K194 / K358, for example, T40Y / E94C / F98W / C109A / K177M / K194L / K358I, T40 / E94 / F98 / C109 / K194 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / K194L / Y350I / K358I, or T40Y / E94C / F98W / C109A / K194L / Y350E / K358I, E94 / F98 / C109 / C114 / K177 / K194, for example, E94C / F98W / C109A / C114I / K194L / K177M, E94 / F98 / C109 / C114 / K194 / Y350, for example, E94C / F98W / C109A / C114I / K194L / Y350I, or E94C / F98W / C109A / C114I / K194L / Y350E, E94 / F98 / C109 / C114 / K194 / K358, for example, E94C / F98W / C109A / C114I / K194L / K358I, E94 / F98 / C109 / K177 / K194 / Y350, for example, E94C / F98W / C109A / K177M / K194L / Y350I o E94C / F98W / C109A / K177M / K194L / Y350E, E94 / F98 / C109 / K177 / K194 / K358, for example, E94C / F98W / C109A / K177M / K194L / K358I, E94 / F98 / C109 / K194 / Y350 / K358, for example, E94C / F98W / C109A / K194L / Y350I / K358I, or E94C / F98W / C109A / K194L / Y350E / K358I, T40 / E94 / F98 / C109 / C114 / K177 / K194 / Y350, for example, T40Y / E94C / F98W / C109A / C114I / K177M / K194L / Y350I, or T40Y / E94C / F98W / C109A / C114I / K177M / K194L / Y350E, T40 / E94 / F98 / C109 / C114 / K177 / K194 / K358, for example, T40Y / E94C / F98W / C109A / C114I / K177M / K194L / K358I, T40 / E94 / F98 / C109 / C114 / K194 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / C114I / K194L / Y350I / K358I, or T40Y / E94C / F98W / C109A / C114I / K194L / Y350E / K358I, T40 / E94 / F98 / C109 / K177 / K194 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / K177M / K194L / Y350I / K358I, or T40Y / E94C / F98W / C109A / K177M / K194L / Y350E / K358I, E94 / F98 / C109 / C114 / K177 / K194 / Y350, for example, E94C / F98W / C109A / C114I / K177M / K194L / Y350I, or E94C / F98W / C109A / C114I / K177M / K194L / Y350E, E94 / F98 / C109 / C114 / K177 / K194 / K358, for example, E94C / F98W / C109A / C114I / K177M / K194L / K358I, E94 / F98 / C109 / C114 / K194 / Y350 / K358, for example, E94C / F98W / C109A / C114I / K194L / Y350I / K358I, or E94C / F98W / C109A / C114I / K194L / Y350E / K358I, E94 / F98 / C109 / K177 / K194 / Y350 / K358, for example, E94C / F98W / C109A / K177M / K194L / Y350I / K358I, or E94C / F98W / C109A / K177M / K194L / Y350E / K358I, T40 / E94 / F98 / C109 / C114 / K177 / K194 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / C114I / K177M / K194L / Y350I / K358I, or T40Y / E94C / F98W / C109A / C114I / K177M / K194L / Y350E / K358I, E94 / F98 / C109 / C114 / K177 / K194 / Y350 / K358, for example, E94C / F98W / C109A / C114I / K177M / K194L / Y350I / K358I, or E94C / F98W / C109A / C114I / K177M / K194L / Y350E / K358I, T40 / E94 / F98 / C109 / C136, for example, T40Y / E94C / F98W / C109A / C136A, T40 / E94 / F98 / C109 / C114 / C136, for example, T40Y / E94C / F98W / C109A / C114I / C136A, T40 / E94 / F98 / C109 / C136 / K177, for example, T40Y / E94C / F98W / C109A / C136A / K177M, T40 / E94 / F98 / C109 / C136 / Y350, for example, T40Y / E94C / F98W / C109A / C136A / Y350I, or T40Y / E94C / F98W / C109A / C136A / Y350E, T40 / E94 / F98 / C109 / C136 / K358, for example, T40Y / E94C / F98W / C109A / C136A / K358I, E94 / F98 / C109 / C114 / C136, for example, E94C / F98W / C109A / C114I / C136A, E94 / F98 / C109 / C136 / K177, for example, E94C / F98W / C109A / C136A / K177M, E94 / F98 / C109 / C136 / Y350, for example, E94C / F98W / C109A / C136A / Y350I, or T40Y / E94C / F98W / C109A / C136A / Y350E, E94 / F98 / C109 / C136 / K358, for example, E94C / F98W / C109A / C136A / K358I, T40 / E94 / F98 / C109 / C114 / C136 / K177, for example, T40Y / E94C / F98W / C109A / C114I / C136A / K177M, T40 / E94 / F98 / C109 / C114 / C136 / Y350, for example, T40Y / E94C / F98W / C109A / C114I / C136A / Y350I, or T40Y / E94C / F98W / C109A / C114I / C136A / Y350E, T40 / E94 / F98 / C109 / C114 / C136 / K358, for example, T40Y / E94C / F98W / C109A / C114I / C136A / K358I, T40 / E94 / F98 / C109 / C136 / K177 / Y350, for example, T40Y / E94C / F98W / C109A / C136AK177M / Y350I, or T40Y / E94C / F98W / C109A / C136AK177M / Y350E, T40 / E94 / F98 / C109 / C136 / K177 / K358, for example, T40Y / E94C / F98W / C109A / C136AK177M / K358I, T40 / E94 / F98 / C109 / C136 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / C136AY350I / K358I, or T40Y / E94C / F98W / C109A / C136AY350E / K358I, E94 / F98 / C109 / C114 / C136 / K177, for example, E94C / F98W / C109A / C114I / C136A / K177M, E94 / F98 / C109 / C114 / C136 / Y350, for example, E94C / F98W / C109A / C114I / C136A / Y350I, or E94C / F98W / C109A / C114I / C136A / Y350E, E94 / F98 / C109 / C114 / C136 / K358, for example, E94C / F98W / C109A / C114I / C136A / K358I, E94 / F98 / C109 / C136 / K177 / Y350, for example, E94C / F98W / C109A / C136AK177M / Y350I, or E94C / F98W / C109A / C136AK177M / Y350E, E94 / F98 / C109 / C136 / K177 / K358, for example, E94C / F98W / C109A / C136AK177M / K358I, E94 / F98 / C109 / C136 / Y350 / K358, for example, E94C / F98W / C109A / C136AY350I / K358I, or E94C / F98W / C109A / C136AY350E / K358I, T40 / E94 / F98 / C109 / C114 / C136 / K177 / Y350, for example, T40Y / E94C / F98W / C109A / C114I / C136A / K177M / Y350I, or T40Y / E94C / F98W / C109A / C114I / C136A / K177M / Y350E, T40 / E94 / F98 / C109 / C114 / C136 / K177 / K358, for example, T40Y / E94C / F98W / C109A / C114I / C136A / K177M / K358I, T40 / E94 / F98 / C109 / C114 / C136 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / C114I / C136A / Y350I / K358I, or T40Y / E94C / F98W / C109A / C114I / C136A / Y350E / K358I, T40 / E94 / F98 / C109 / C136 / K177 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / C136A / K177M / Y350I / K358I, or T40Y / E94C / F98W / C109A / C136A / K177M / Y350E / K358I, E94 / F98 / C109 / C114 / C136 / K177 / Y350, for example, E94C / F98W / C109A / C114I / C136A / K177M / Y350I, or E94C / F98W / C109A / C114I / C136A / K177M / Y350E, E94 / F98 / C109 / C114 / C136 / K177 / K358, for example, E94C / F98W / C109A / C114I / C136A / K177M / K358I, E94 / F98 / C109 / C114 / C136 / Y350 / K358, for example, E94C / F98W / C109A / C114I / C136A / Y350I / K358I, or E94C / F98W / C109A / C114I / C136A / Y350E / K358I, E94 / F98 / C109 / C136 / K177 / Y350 / K358, for example, E94C / F98W / C109A / C136A / K177M / Y350I / K358I, or E94C / F98W / C109A / C136A / K177M / Y350E / K358I, T40 / E94 / F98 / C109 / C114 / C136 / K177 / Y350 / K358, for example, T40Y / E94C / F98W / C109A / C114I / C136A / K177M / Y350I / K358I, or T40Y / E94C / F98W / C109A / C114I / C136A / K177M / Y350E / K358I, E94 / F98 / C109 / C114 / C136 / K177 / Y350 / K358, for example, E94C / F98W / C109A / C114I / C136A / K177M / Y350I / K358I, or E94C / F98W / C109A / C114I / C136A / K177M / Y350E / K358I, T40 / E94 / F98 / C109 / A360, for example, T40Y / E94C / F98W / C109A / A360C, T40 / E94 / F98 / C109 / C114 / A360, for example, T40Y / E94C / F98W / C109A / C114I / A360C, T40 / E94 / F98 / C109 / K177 / A360, for example, T40Y / E94C / F98W / C109A / K177M / A360C, T40 / E94 / F98 / C109 / Y350 / A360, for example, T40Y / E94C / F98W / C109A / Y350I / A360C, or T40Y / E94C / F98W / C109A / Y350E / A360C, T40 / E94 / F98 / C109 / K358 / A360, for example, T40Y / E94C / F98W / C109A / K358I / A360C, E94 / F98 / C109 / C114 / A360, for example, E94C / F98W / C109A / C114I / A360C, E94 / F98 / C109 / K177 / A360, for example, E94C / F98W / C109A / K177M / A360C, E94 / F98 / C109 / Y350 / A360, for example, E94C / F98W / C109A / Y350I / A360C, or E94C / F98W / C109A / Y350E / A360C, E94 / F98 / C109 / K358 / A360, for example, E94C / F98W / C109A / K358I / A360C, T40 / E94 / F98 / C109 / C114 / K177 / A360, for example, T40Y / E94C / F98W / C109A / C114I / K177M / A360C, T40 / E94 / F98 / C109 / C114 / Y350 / A360, for example, T40Y / E94C / F98W / C109A / C114I / Y350I / A360C, or T40Y / E94C / F98W / C109A / C114I / Y350E / A360C, T40 / E94 / F98 / C109 / C114 / K358 / A360, for example, T40Y / E94C / F98W / C109A / C114I / K358I / A360C, T40 / E94 / F98 / C109 / K177 / Y350 / A360, for example, T40Y / E94C / F98W / C109A / K177M / Y350I / A360C, or T40Y / E94C / F98W / C109A / K177M / Y350E / A360C, / A360C T40 / E94 / F98 / C109 / Y350 / K358 / A360, for example, T40Y / E94C / F98W / C109A / Y350I / K358I / A360C, or T40Y / E94C / F98W / C109A / Y350E / K358I / A360C, E94 / F98 / C109 / C114 / K177 / A360, for example, E94C / F98W / C109A / C114I / K177M / A360C, E94 / F98 / C109 / C114 / Y350 / A360, for example, E94C / F98W / C109A / C114I / Y350I / A360C, or E94C / F98W / C109A / C114I / Y350E / A360C, E94 / F98 / C109 / C114 / K358 / A360, for example, E94C / F98W / C109A / C114I / K358I / A360C, E94 / F98 / C109 / K177 / Y350 / A360, for example, E94C / F98W / C109A / K177M / Y350I / A360C, or E94C / F98W / C109A / K177M / Y350E / A360C, E94 / F98 / C109 / Y350 / K358 / A360, for example, E94C / F98W / C109A / Y350I / K358I / A360C, or E94C / F98W / C109A / Y350E / K358I / A360C, T40 / E94 / F98 / C109 / C114 / K177 / Y350 / A360, for example, T40Y / E94C / F98W / C109A / C114I / K177M / Y350I / A360C, or T40Y / E94C / F98W / C109A / C114I / K177M / Y350E / A360C, T40 / E94 / F98 / C109 / C114 / K177 / K358 / A360, for example, T40Y / E94C / F98W / C109A / C114I / K177M / K358I / A360C, T40 / E94 / F98 / C109 / C114 / Y350 / K358 / A360, for example, T40Y / E94C / F98W / C109A / C114I / Y350I / K358I / A360C, or T40Y / E94C / F98W / C109A / C114I / Y350E / K358I / A360C, T40 / E94 / F98 / C109 / K177 / Y350 / K358 / A360, for example, T40Y / E94C / F98W / C109A / K177M / Y350I / K358I / A360C, or T40Y / E94C / F98W / C109A / K177M / Y350E / K358I / A360C, E94 / F98 / C109 / C114 / K177 / Y350 / A360, for example, E94C / F98W / C109A / C114I / K177M / Y350I / A360C, or E94C / F98W / C109A / C114I / K177M / Y350E / A360C, E94 / F98 / C109 / C114 / K177 / K358 / A360, for example, E94C / F98W / C109A / C114I / K177M / K358I / A360C, E94 / F98 / C109 / C114 / Y350 / K358 / A360, for example, E94C / F98W / C109A / C114I / Y350I / K358I / A360C, or E94C / F98W / C109A / C114I / Y350E / K358I / A360C, E94 / F98 / C109 / K177 / Y350 / K358 / A360, for example, E94C / F98W / C109A / K177M / Y350I / K358I / A360C, or E94C / F98W / C109A / K177M / Y350E / K358I / A360C, T40 / E94 / F98 / C109 / C114 / K177 / Y350 / K358 / A360, for example, T40Y / E94C / F98W / C109A / C114I / K177M / Y350I / K358I / A360C, or T40Y / E94C / F98W / C109A / C114I / K177M / Y350E / K358I / A360C, or E94 / F98 / C109 / C114 / K177 / Y350 / K358 / A360, for example, E94C / F98W / C109A / C114I / K177M / Y350I / K358I / A360C, or E94C / F98W / C109A / C114I / K177M / Y350E / K358I / A360C.
[0334] Any of these variants of SEQ ID NO: 118 may further comprise modifications or substitutions at any number and combination of positions (a) 55, (b) 156, (c) 210, and (d) 221, including (a); (b); (c); (d); (a) and (b); (a) and (c); (a) and (d); (b) and (c); (b) and (d); (c) and (d); (a), (b) and (c); (a), (b) and (d); (a), (c) and (d); (b), (c) and (d); or (a), (b), (c) and (d). Any of these variants of SEQ ID NO: 118 may further comprise any number and combination of (a) T55K, (b) T156F, (c) T210K, and (d) N221E, including (a);(b);(c);(d);(a) and (b);(a) and (c);(a) and (d);(b) and (c);(b) and (d);(c) and (d);(a), (b) and (c);(a), (b) and (d);(a), (c) and (d);(b), (c) and (d); or (a), (b), (c) and (d).
[0335] Additional Modifications / Substitutions The present invention also provides modified DNA-dependent ATPase (Dda) helicases in which one or more of the positions corresponding to the following amino acid positions in Dda1993 have been modified or substituted: 86, 90, 92, 97, 101, 102, 273, 293, 300, 301, 303, 305, 308, 310, 312 317, 323, 328, 332, 334, 335, 336, 337, 339, 351, 354, 359, 361, 364, 366, 368, 371, 374, 376, 377, 379, and 388. The present invention also provides modified DNA-dependent ATPase (Dda) helicases in which one or more of the positions corresponding to the following amino acid positions in Dda1993 have been modified or substituted. 351, 354, and 361. Any number and combination of these modifications / substitutions may be made, including 351, 354, 361, 351 and 354, 351 and 361, 354 and 361, or 351, 354, and 361. These positions may be modified or substituted alone or in combination with any of the inventive modifications or substitutions described above.
[0336] Helicases of the invention preferably comprise variants of SEQ ID NO: 118 comprising one or more of the following substitutions: K86A, V90T, Y92A, Y92D, Y92F, Y92G, Y92H, Y92N, Y92Q, Y92S, Y92T, Y92V, Y92W, L97H, K101A, E102G, E273A, N293H, I200A, I300F, E301I, E303I, Y305L, F308I, F308L, K310A, K310I, K310L, R312I, R312L, R312M, E317I, E317Y, W323H, E328I, D332A, D332L, E334A, E334I , E334Y, Y335A, Y335I, Y335L, Y336L, R337I, R337L, R337M, K339A, K339I, K339L, K351I, K351Q, L354Q, L354A, T359L, E361T, E361I, E361Q, K364A, K364R, W366L, K368A, K368I, K368L, K371I, K371L, K371M, W374A, W374L, W374L, D376S, F377A, D379A, D379L, or K388R. Helicases of the invention preferably comprise variants of SEQ ID NO: 118 comprising one or more of the following substitutions: (a) K351I or K351Q, (b) L354A or L354Q, and (c) E361I or E361Q. Variants can include (a), (b), (c), (a) and (b), (a) and (c), (b) and (c), or (a), (b), and (c). These substitutions can be made alone or in combination with any of the modifications or substitutions of the invention described above.
[0337] variant A helicase variant is an enzyme that has an amino acid sequence different from that of the wild-type helicase and has polynucleotide binding activity. In particular, a variant of any one of SEQ ID NOs: 118-133 is an enzyme that has an amino acid sequence different from that of any one of SEQ ID NOs: 118-133 and has polynucleotide binding activity. Polynucleotide binding activity can be measured using methods known in the art. Suitable methods include, but are not limited to, fluorescence anisotropy, tryptophan fluorescence, and electrophoretic mobility shift assay (EMSA). For example, the ability of a variant to bind to a single-stranded polynucleotide can be determined as described in the Examples.
[0338] The variant has helicase activity, which can be measured in various ways. For example, the ability of the variant to translocate along a polynucleotide can be measured using electrophysiology, a fluorescence assay, or ATP hydrolysis.
[0339] The variant may contain modifications that facilitate handling of the polynucleotide encoding the helicase and / or that facilitate its activity at high salt concentrations and / or at room temperature.
[0340] A variant is preferably at least 20% homologous to the amino acid sequence of any one of SEQ ID NOs: 118-133 based on amino acid similarity or identity over the entire length of that sequence. More preferably, the variant polypeptide can be at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% homologous to the amino acid sequence of any one of SEQ ID NOs: 118-133 based on amino acid identity over the entire sequence. More preferably, a variant polypeptide can be at least 30%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, more preferably at least 95%, 97%, or 99% identical to the amino acid sequence of any one of SEQ ID NOs: 118-133 over the entire sequence. There can also be at least 70%, e.g., at least 80%, at least 85%, at least 90%, or at least 95% amino acid identity ("hard homology") over a stretch of 100 or more, e.g., 150, 200, 300, 400, or 500 or more contiguous amino acids. Homology is determined as described below. Notably, in addition to the specific modifications discussed above, a variant of any one of SEQ ID NOs: 118-133 can include one or more substitutions, one or more deletions, and / or one or more additions, as discussed below.
[0341] A preferred variant of any one of SEQ ID NOs: 118-133 has an unnatural amino acid, such as Faz, at the amino (N)-terminus and / or carboxy (C)-terminus. A preferred variant of any one of SEQ ID NOs: 118-133 has a cysteine residue at the amino (N)-terminus and / or carboxy (C)-terminus. A preferred variant of any one of SEQ ID NOs: 118-133 has a cysteine residue at the amino (N)-terminus and an unnatural amino acid, such as Faz, at the carboxy (C)-terminus, or vice versa.
[0342] Preferred variants of SEQ ID NO: 118 contain one or more, for example all of the following modifications: E54G, D151E, I196N, and G357A.
[0343] No connection In one preferred embodiment, none of the introduced cysteines and / or unnatural amino acids in the modified helicase of the invention are connected to each other.
[0344] Connecting two more of the introduced cysteines and / or unnatural amino acids In another preferred embodiment, two more of the introduced cysteines and / or unnatural amino acids in the modified helicase of the invention are connected to each other, which typically reduces the ability of the helicase of the invention to unbind from a polynucleotide.
[0345] Any number and combination of two or more of the introduced cysteines and / or unnatural amino acids can be connected to each other. For example, 3, 4, 5, 6, 7, 8 or more cysteines and / or unnatural amino acids can be connected to each other. One or more cysteines can be connected to one or more cysteines. One or more cysteines can be connected to one or more unnatural amino acids, such as Faz. One or more unnatural amino acids, such as Faz, can be connected to one or more unnatural amino acids, such as Faz.
[0346] The two or more cysteines and / or unnatural amino acids can be connected in any manner. The connection can be temporary, for example, by a non-covalent bond. Even a temporary connection reduces unbinding of the polynucleotide from the helicase.
[0347] The two or more cysteines and / or unnatural amino acids are preferably connected by an affinity molecule. Suitable affinity molecules are known in the art. The affinity molecule is preferably (a) a complementary polynucleotide (WO2010 / 086602, the entire contents of which are incorporated herein by reference), (b) an antibody or a fragment thereof and a complementary epitope (Biochemistry 6th Ed, W.H. Freeman and co (2007) pp953-954), (c) a peptide zipper (O'Shea et al., Science 254(5031):539-544), (d) capable of interacting through β-sheet expansion (Remaut and Waksman Trends Biochem. Sci. (2006) 31 436-444), (e) capable of forming hydrogen bonds, pi-stacking, or salt bridges, or (f) a rotaxane (Xiang Ma and He Tian Chem. Soc. Rev., 2010, 39, 70-80), (g) aptamers and complementary proteins (James, W. in Encyclopedia of Analytical Chemistry, RAMeyers (Ed.) pp. 4848-4871 John Wiley & Sons Ltd, Chichester, 2000), or (h) half-chelators (Hammerstein et al. J. Biol. Chem. 2011 April 22;286(16):14324-14334). For (e), a hydrogen bond occurs between a proton attached to an electronegative atom and another electronegative atom. Pi-stacking requires two aromatic rings that can be stacked together when the planes of the rings are parallel. Salt bridges are between groups that can delocalize their electrons on several atoms, for example, between aspartic acid and arginine.
[0348] The two or more moieties can be transiently connected by a hexa-his tag or Ni-NTA.
[0349] The two or more cysteines and / or unnatural amino acids are preferably permanently connected. In the context of the present invention, a connection is permanent if it is not disrupted during use of the helicase or if it cannot be disrupted without intervention on the part of the user, such as opening a -SS- bond using reduction.
[0350] The two or more cysteines and / or unnatural amino acids are preferably covalently linked. The two or more cysteines and / or unnatural amino acids can be covalently linked using any method known in the art.
[0351] Two or more cysteines and / or non-natural amino acids can be covalently linked via their naturally occurring amino acids, such as cysteine, threonine, serine, aspartic acid, asparagine, glutamic acid, and glutamine. Naturally occurring amino acids can be modified to facilitate linkage. For example, naturally occurring amino acids can be modified by acylation, phosphorylation, glycosylation, or farnesylation. Other suitable modifications are known in the art. Modifications to naturally occurring amino acids can be post-translational modifications. Two or more cysteines and / or non-natural amino acids can be linked via amino acids introduced into their sequences. Such amino acids are preferably introduced by substitution. The introduced amino acids can be cysteines or non-natural amino acids that facilitate linkage. Suitable unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz), any one of the amino acids numbered 1-71 contained in Figure 1 of Liu CC and Schultz PG, Annu. Rev. Biochem., 2010, 79, 413-444, or any one of the amino acids listed below. The introduced amino acid can be modified as discussed above.
[0352] In a preferred embodiment, two or more cysteines and / or unnatural amino acids are connected using a linker. Linker molecules are discussed in more detail below. One suitable connection method is cysteine linkage, which is discussed in more detail below. Two or more cysteines and / or unnatural amino acids are preferably connected using one or more, for example, two or three, linkers. The one or more linkers can be designed to reduce the size of the opening or to close the opening, as discussed above. When one or more linkers are used to close the opening, as discussed above, at least a portion of the one or more linkers is preferably oriented such that it is not parallel to the polynucleotide when bound by the helicase. More preferably, all of the linkers are oriented in this manner. When one or more linkers are used to close the opening, as discussed above, at least a portion of the one or more linkers preferably traverses the opening in an orientation that is not parallel to the polynucleotide when bound by the helicase. More preferably, all of the linkers traverse the opening in this manner. In these embodiments, at least a portion of one or more linkers can be perpendicular to the polynucleotide, which effectively closes the opening so that the polynucleotide cannot unbind from the helicase through the opening.
[0353] Each linker can have two or more functional ends, for example, two, three, or four functional ends. Suitable configurations of ends within a linker are well known in the art.
[0354] One or more ends of the one or more linkers are preferably covalently attached to the helicase. If one end is covalently attached, the one or more linkers may temporarily connect two or more cysteines and / or unnatural amino acids, as discussed above. If both or all ends are covalently attached, the one or more linkers permanently connect two or more cysteines and / or unnatural amino acids.
[0355] The one or more linkers are preferably amino acid sequences and / or chemical cross-linkers.
[0356] Suitable amino acid linkers, such as peptide linkers, are known in the art. The length, flexibility, and hydrophilicity of the amino acid or peptide linker are typically designed so that it reduces the size of the opening but does not interfere with the function of the helicase. Preferred flexible peptide linkers are stretches of 2 to 20, e.g., 4, 6, 8, 10, or 16, serine and / or glycine amino acids. More preferred flexible linkers are (SG)1, (SG)2, (SG)3, (SG)4, (SG)5, (SG)8, (SG) 10 , (SG) 15 , or (SG) 20 wherein S is serine and G is glycine. A preferred rigid linker is a stretch of 2 to 30, e.g., 4, 6, 8, 16, or 24, proline amino acids. A more preferred rigid linker is (P) 12 wherein P is proline. The amino acid sequence of the linker preferably comprises a polynucleotide-binding moiety. Such moieties and the advantages associated with their use are discussed below.
[0357] Suitable chemical cross-linking agents are well known in the art and include, but are not limited to, those containing the following functional groups: maleimide, active ester, succinimide, azide, alkyne (such as dibenzocyclooctynol (DIBO or DBCO), difluorocycloalkyne, and linear alkyne), phosphine (such as those used in traceless and non-traceless Staudinger ligation), haloacetyl (such as iodoacetamide), phosgene-type reagents, sulfonyl chloride reagents, isothiocyanate, acyl halides, hydrazine, disulfide, vinyl sulfone, aziridine, and photoreactive reagents (such as aryl azide, diaziridine).
[0358] The reaction between the amino acid and the functional group can be spontaneous, such as cysteine / maleimide, or can require an external reagent, such as Cu(I) to link an azide and a linear alkyne.
[0359] Linkers can include any molecule that spans the required distance. Linkers can vary in length from one carbon (phosgene-type linkers) to many angstroms. Examples of linear molecules include, but are not limited to, polyethylene glycol (PEG), polypeptides, polysaccharides, deoxyribonucleic acid (DNA), peptide nucleic acid (PNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), saturated and unsaturated hydrocarbons, and polyamides. These linkers can be inert or reactive; in particular, they can be chemically cleavable at defined positions or can themselves be modified with fluorophores or ligands. Linkers are preferably resistant to dithiothreitol (DTT).
[0360] Preferred crosslinkers include 2,5-dioxopyrrolidin-1-yl 3-(pyridin-2-yldisulfanyl)propanoate, 2,5-dioxopyrrolidin-1-yl 4-(pyridin-2-yldisulfanyl)butanoate, and 2,5-dioxopyrrolidin-1-yl 8-(pyridin-2-yldisulfanyl)octananoate, di-maleimide PEG 1k, di-maleimide PEG 3.4k, di-maleimide PEG 5k, di-maleimide PEG 10k, bis(maleimido)ethane (BMOE), bis-maleimidohexane (BMH), 1,4-bis-maleimidobutane (BMB), 1,4 bis-maleimidyl-2,3-dihydroxybutane (BMDB), BM[PEO]2 (1,8-bis-maleimidodiethylene glycol), BM[PEO]3 (1,11-bis-maleimidotriethylene glycol), tris[2-maleimidoethyl]amine (TMEA), DTME dithiobismaleimidoethane, bis-maleimidoPEG3, bis-maleimidoPEG11, DBCO-maleimide, DBCO-PEG4-maleimide, DBCO-PEG4-NH2, DBCO-PEG4-NHS, DBCO-NHS, DBCO-PEG-DBCO 2.8kDa, DBCO-PEG-DBCO Examples of crosslinkers include 4.0 kDa, DBCO-15 atoms-DBCO, DBCO-26 atoms-DBCO, DBCO-35 atoms-DBCO, DBCO-PEG4-SS-PEG3-biotin, DBCO-SS-PEG3-biotin, DBCO-SS-PEG11-biotin, (succinimidyl 3-(2-pyridyldithio)propionate (SPDP)), and maleimide-PEG(2 kDa)-maleimide (alpha, omega-bis-maleimide poly(ethylene glycol)). The most preferred crosslinker is maleimide-propyl-SRDFWRS-(1,2-diaminoethane)-propyl-maleimide.
[0361] One or more of the linkers may be cleavable, which is discussed in more detail below.
[0362] Two or more cysteines and / or unnatural amino acids can be connected using two different linkers that are specific for each other. One of the linkers is attached to one moiety and the other is attached to another moiety. The linkers should react to form a modified helicase of the present invention. Two or more cysteines and / or unnatural amino acids can be connected using hybridization linkers described in WO2010 / 086602 (incorporated herein by reference in its entirety). In particular, two or more cysteines and / or unnatural amino acids can be connected using two or more linkers, each of which comprises a hybridizable region and a group capable of forming a covalent bond. The hybridizable region in the linker hybridizes and connects the two or more cysteines and / or unnatural amino acids. The connected cysteines and / or unnatural amino acids are then coupled via the formation of a covalent bond between the groups. Any of the specific linkers disclosed in WO2010 / 086602 (incorporated herein by reference in its entirety) can be used in accordance with the present invention.
[0363] Two or more cysteines and / or unnatural amino acids can be modified and then linked using a chemical cross-linker that is specific for the two modifications. Any of the cross-linkers discussed above can be used.
[0364] The linker may be labeled. Suitable labels include fluorescent molecules (such as Cy3 or AlexaFluor® 555), radioisotopes, e.g. 125 I, 35 Examples of suitable labels include, but are not limited to, ligands such as S, enzymes, antibodies, antigens, polynucleotides, and biotin. Such labels allow for the quantification of the amount of linker. The label can also be a cleavable purification tag such as biotin, or a specific sequence that is not present in the protein itself but appears in an identification method, such as a peptide released by trypsin digestion.
[0365] A preferred method of connecting two or more cysteines is via cysteine linkage, which can be mediated by bifunctional chemical cross-linkers or by amino acid linkers with terminally presented cysteine residues.
[0366] The length, reactivity, specificity, rigidity, and solubility of any bifunctional linker can be designed to ensure that the size of the opening is sufficiently reduced and that helicase function is maintained. Suitable linkers include bismaleimide crosslinkers, such as 1,4-bis(maleimido)butane (BMB) or bis(maleimido)hexane. One disadvantage of bifunctional linkers is that, if conjugation at a specific site is preferred, the helicase must not contain additional surface-accessible cysteine residues, because conjugation of the bifunctional linker to a surface-accessible cysteine residue can be difficult to control and may affect substrate binding or activity. If the helicase contains several accessible cysteine residues, modification of the helicase to remove them may be required while ensuring that the modification does not affect the folding or activity of the helicase. This is discussed in WO 2010 / 086603, which is incorporated herein by reference in its entirety. The reactivity of cysteine residues can be enhanced by modification of adjacent residues, for example, on the peptide linker. For example, the basic group of an adjacent arginine, histidine, or lysine residue may shift the pKa of the cysteine thiol group toward the more reactive S - The pKa of the cysteine residue can be changed. The reactivity of cysteine residues can be protected by thiol protecting groups such as 5,5'-dithiobis-(2-nitrobenzoic acid) (dTNB). These can be reacted with one or more cysteine residues of the helicase before the linker is attached. Selective deprotection of surface-accessible cysteines can be possible using a reducing reagent (e.g., immobilized tris(2-carboxyethyl)phosphine, TCEP) immobilized on the beads. Cysteine linkages are discussed in more detail below.
[0367] Another preferred method of attachment is via Faz ligation, which can be mediated by a bifunctional chemical linker or by a polypeptide linker with a terminally displayed Faz residue.
[0368] Other modifications The helicases of the present invention can also be modified to increase the attraction between (i) the tower domain and (ii) the pin domain and / or the 1A domain. Any known chemical modification can be performed in accordance with the present invention. These types of modifications are disclosed in WO2015 / 055981, which is incorporated herein by reference in its entirety.
[0369] In particular, the present invention provides helicases of the present invention in which at least one charged amino acid has been introduced into (i) the tower domain, and / or (ii) the pin domain, and / or (iii) the 1A (RecA-like motor) domain, and the helicase has the ability to control polynucleotide movement. The ability of a helicase to control polynucleotide movement can be measured as discussed above. The present invention preferably provides helicases of the present invention in which at least one charged amino acid has been introduced into (i) the tower domain, and (ii) the pin domain and / or the 1A domain.
[0370] The at least one charged amino acid may be negatively or positively charged. Preferably, the at least one charged amino acid is oppositely charged relative to any amino acid(s) that interact with it in the helicase. For example, at least one positively charged amino acid may be introduced into the tower domain at a position where it interacts with a negatively charged amino acid in the pin domain. The at least one charged amino acid is typically introduced into an uncharged position in a wild-type (i.e., unmodified) helicase. The at least one charged amino acid may be used to replace at least one oppositely charged amino acid in the helicase. For example, a positively charged amino acid may be used to replace a negatively charged amino acid.
[0371] Suitable charged amino acids are discussed above. At least one charged amino acid may be natural, such as arginine (R), histidine (H), lysine (K), aspartic acid (D), or glutamic acid (D). Alternatively, at least one charged amino acid may be artificial or non-natural. Any number of charged amino acids may be introduced into each domain. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more charged amino acids may be introduced into each domain.
[0372] The helicase preferably comprises a variant of SEQ ID NO: 118 comprising a positively charged amino acid at one or more of the following positions: (i) 93, (ii) 354, (iii) 360, (iv) 361, (v) 94, (vi) 97, (vii) 155, (viii) 357, (ix) 100, and (x) 127. The helicase preferably comprises a variant of SEQ ID NO: 118 comprising a negatively charged amino acid at one or more of the following positions: (i) 354, (ii) 358, (iii) 360, (iv) 364, (v) 97, (vi) 123, (vii) 155, (viii); 357, (ix) 100, and (x) 127. The helicase preferably comprises a variant of any one of SEQ ID NOs: 119-133, which comprises a positively charged amino acid or a negatively charged amino acid at a position corresponding to that in SEQ ID NO: 118 as defined in any of (i)-(x). The positions in any one of SEQ ID NOs: 119-133 that correspond to those in SEQ ID NO: 118 can be identified using the alignment of SEQ ID NOs: 118-133 below.
[0373] Helicases preferably include variants of SEQ ID NO: 118 that have been modified by the introduction of at least one charged amino acid to contain oppositely charged amino acids at the following positions: (i) 93 and 354, (ii) 93 and 358, (iii) 93 and 360, (iv) 93 and 361, (v) 93 and 364, (vi) 94 and 354, (vii) 94 and 358, (viii) 94 and 360, (ix) 94 and 361, (x) 94 and 364, (xi) 97 and 354, (xii) 97 and 358, (xiii) 97 and 360, (xiv) (xxiv) 155 and 361, (xxv) 155 and 364. The helicase of the present invention preferably comprises a variant of any one of SEQ ID NOs: 119 to 133, which comprises an oppositely charged amino acid at a position corresponding to that in SEQ ID NO: 118 as defined in any of (i) to (xxv).
[0374] The present invention also provides helicases in which (i) at least one charged amino acid has been introduced into the tower domain, and (ii) at least one oppositely charged amino acid has been introduced into the pin domain and / or the 1A (RecA-like motor) domain, wherein the helicase has the ability to control the movement of a polynucleotide. At least one charged amino acid can be negatively charged, and at least one oppositely charged amino acid can be positively charged, or vice versa. Suitable charged amino acids are discussed above. Any number of charged amino acids and any number of oppositely charged amino acids can be introduced. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more charged amino acids can be introduced, and / or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more oppositely charged amino acids can be introduced.
[0375] Charged amino acids are typically introduced at uncharged positions in the wild-type helicase. One or both of the charged amino acids may be used to replace charged amino acids in the helicase. For example, a positively charged amino acid may be used to replace a negatively charged amino acid. Charged amino acids may be introduced at any of the positions within (i) the tower domain and (ii) the pin domain and / or the 1A domain discussed above. Oppositely charged amino acids are typically introduced so that they interact in the resulting helicase. Helicases preferably include variants of SEQ ID NO: 118 in which oppositely charged amino acids are introduced at the following positions: (i) 97 and 354, (ii) 97 and 360, (iii) 155 and 354, or (iv) 155 and 360. Helicases of the present invention preferably include variants of any one of SEQ ID NOs: 119-133, which include oppositely charged amino acids at positions corresponding to those in SEQ ID NO: 118 as defined in any of (i)-(iv).
[0376] construct The present invention also provides a construct comprising a modified helicase of the present invention and an additional polynucleotide binding moiety, wherein the helicase is bound to the polynucleotide binding moiety, and the construct has the ability to control the movement of a polynucleotide. The construct may be artificial or non-naturally occurring.
[0377] The constructs of the present invention are useful tools for controlling polynucleotide translocation during strand sequencing. The constructs of the present invention are less likely to disassociate from the polynucleotide to be sequenced than the modified helicases of the present invention. The constructs can provide a greater read length of the polynucleotide because they control the translocation of the polynucleotide through the nanopore.
[0378] Targeting constructs can also be designed to bind to specific polynucleotide sequences. As discussed in more detail below, the polynucleotide binding moiety can bind to a specific polynucleotide sequence, thereby targeting the helicase portion of the construct to the specific sequence.
[0379] The construct has the ability to control the movement of the polynucleotide, which can be determined as discussed above.
[0380] The constructs of the present invention can be isolated, substantially isolated, purified, or substantially purified. A construct is iso...
Claims
1. (a) Modified DNA-dependent ATPase (Dda) helicase, wherein the helicase comprises modification or substitution at one or more of the positions corresponding to amino acid positions 55, 114, 156, 177, 210, 221, 350, and 358 in SEQ ID NO: 118, and (b) Isolated CsgG pores, or isolated pore complexes comprising CsgG pores and modified CsgF peptides, wherein the CsgG pores contain at least one monomer modified at one or more positions W97, Q100, E101, N102, and T104 in SEQ ID NO: 117, A kit for characterizing target analytes, including those containing the specified substance.
2. (a) The amino acid corresponding to position 55 of the helicase is substituted with D, E, K, N, or S; (b) The amino acid corresponding to position 114 of the helicase is substituted with A, V, I, L, M, F, Y, W, G, P, S, T, N, or Q; (c) The amino acid corresponding to position 156 of the helicase is substituted with A, E, F, G, I, L, M, P, S, V, Y, D, K, or N; (d) The amino acid corresponding to position 177 of the helicase is substituted with D, E, F, G, H, I, L, M, N, Q, R, S, T, V, W, or Y; (e) The amino acid corresponding to position 210 of the helicase The kit according to claim 1, wherein (f) the amino acid at position 221 of the helicase is substituted with D, E, K, S, N, R, H, or Y, (g) the amino acid corresponding to position 350 of the helicase is substituted with A, D, E, G, K, L, N, Q, R, T, V, H, or M, or with D, E, A, V, I, L, M, F, W, R, H, K, L, S, T, N, or Q, and / or (h) the amino acid corresponding to position 358 of the helicase is substituted with D, E, A, V, I, L, M, F, Y, W, R, H, L, S, T, N, or Q.
3. The kit according to claim 1, wherein the helicase further comprises modification or substitution at the position corresponding to amino acid position 40 in SEQ ID NO:
118.
4. The kit according to claim 3, wherein the amino acid corresponding to position 40 is substituted with A, V, I, L, M, F, Y, or W.
5. The helicase is part of a construct comprising the helicase and an additional polynucleotide binding portion, wherein the helicase is bound to the polynucleotide binding portion, and the construct has the ability to control the movement of the analyte. The kit according to claim 1, wherein the structure may include two or more helicases as described in claim 1.
6. The kit according to claim 1, wherein (i) W at position 97 of the monomer is replaced with R, H, K, A, V, I, L, M, F, Y, S, T, Q, D, E, N, C, P, or G; (ii) Q at position 100 of the monomer is replaced with R, H, K, W, A, V, I, L, M, F, Y, T, N, or S; (iii) E at position 101 of the monomer is replaced with A, V, I, L, M, F, Y, or W, or S, T, N, Q, C, G, or P; (iv) N at position 102 of the monomer is replaced with D, E, R, H, K, S, T, Q, V, I, L, M, F, Y, W, or A; or (v) T at position 104 of the monomer is replaced with R, H, or K.
7. The kit according to claim 1, wherein the modified CsgF peptide is inserted into the lumen of the CsgG pore.
8. The kit according to claim 1, wherein the pore complex has two or more channel constrictions, including CsgG channel constriction and CsgF channel constriction.
9. The kit according to claim 8, wherein the CsgF channel constriction has a diameter in the range of 0.5 nm to 2.0 nm.
10. The kit according to claim 1, wherein the isolated pore or isolated pore complex comprises Q100A and / or N102A or N102S, and the Dda helicase comprises any of Y350I, Y350S, K351I, K351Q, L354A, L354Q, K358I, K358L, K358Q, E361I, and E361Q.
11. The kit according to claim 1, wherein the Dda helicase comprises Y350I and / or K358I.
12. A method for characterizing a target analyte, (a) The target analyte is brought into contact with the kit according to any one of claims 1 to 11 such that the helicase or construct controls the movement of the target analyte through the pore or pore complex, (b) A method comprising obtaining one or more measurements as a target analyte moves toward the pores or pore complex, wherein the measurements represent one or more features of the target analyte, thereby characterizing the target analyte.
13. The method according to claim 12, wherein the target analyte is a polynucleotide.
14. The method according to claim 13, comprising determining one or more features selected from (i) the length of the polynucleotide, (ii) the identity of the polynucleotide, (iii) the sequence of the polynucleotide, (iv) the secondary structure of the polynucleotide, and (v) whether or not the polynucleotide is modified.
15. The method according to claim 12, wherein the CsgG pore contains 6 to 10 monomers.