Methods and kits
Patent Information
- Application Number
- PCT/EP2026/059141
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-28
- Filing Date
- 2026-03-30
- Publication Date
- 2026-10-01
Smart Images

Figure IMGF000013_0001_TABLE 
Figure IMGF000028_0001_TABLE 
Figure IMGF000068_0001_TABLE
Abstract
Description
[0001] METHODS AND KITS
[0002] TECHNICAL FIELD
[0003] The invention relates to polynucleotide-binding proteins based on fragments of DNA Topoisomerase V from Methanopyrus kandleri (Topo V). The invention also relates to methods using the proteins to couple polynucleotides to membranes and kits comprising the proteins. The methods and kits are particularly suited to determining the presence, absence or one or more characteristics of polynucleotide analytes, especially using nanopores.
[0004] INTRODUCTION
[0005] Biological pores (and other nanopores) have great potential as direct, electrical biosensors for polymers and a variety of small molecules. In particular, recent focus has been given to nanopores as technology for sequencing polymeric analytes, such as polynucleotides. When a potential is applied across a nanopore, there is a change in the current flow as the monomeric nucleotide components of the polynucleotide reside transiently in the barrel for a certain period of time. Nanopore detection of the nucleotide gives a current change of known signature and duration. In the strand sequencing method, a single polynucleotide strand is passed through the pore and the identities of the nucleotides are derived. Strand sequencing can involve the use of a molecular brake to control the movement of the polynucleotide through the pore.
[0006] When a potential is applied across a nanopore, there is a drop in the current flow when polynucleotide resides transiently in the barrel for a certain period of time. Nanopore detection of the polynucleotide gives a current blockade of known signature and duration. The concentration of the polynucleotide can then be determined by the number of blockade events per unit time to a single pore. For nanopore applications, efficient capture of polynucleotides from solution is required. For instance, the number of interactions between the polynucleotide and the nanopore needs to be maximal. Therefore, it is preferred to have the polynucleotide at as high a concentration as is possible. This becomes a particular problem as the concentration of polynucleotides in some samples can be limiting, especially when the sample contains other molecules or impurities. Methods for increasing the concentration of the polynucleotide in the proximity of the nanopore by coupling the polynucleotide to the membrane containing the nanopore are known in the art (e.g., WO 2012 / 164270 incorporated by reference in its entirety). Known methods typically involve modifying the polynucleotides with sequencing adaptors which are capable of coupling to the membrane. Improved methods of coupling are needed in the art.
[0007] Polynucleotides are typically prepared for nanopore sensing by removing proteins bound to the polynucleotide and removing other components from the sample. This purification allows the polynucleotide to be captured effectively and move effectively and efficiently through the nanopore. Polynucleotide purification and enrichment steps are known in the art. Thereis a need to identify novel ways to allow nanopore sensing of polynucleotides without the need for modification with sequencing adaptors, purification or enrichment.
[0008] DNA Topoisomerase V from Methanopyrus kandleri (Topo V) is a protein containing a topoisomerase domain (residues 1-268) and a DNA-binding domain from residues 269-984. Fragments of Topo V are known in the art (e.g., Rajan, R., Prasad, R., Taneja, B., Wilson, S. H., & Mondragon, A. (2013). Identification of one of the apurinic / apyrimidinic lyase active sites of topoisomerase V by structural and functional studies. Nucleic Acids Research, 41 1), 657-666; and Rakhi Rajan, Amy Osterman, Alfonso Mondragon, Methanopyrus kandleri topoisomerase V contains three distinct AP lyase active sites in addition to the topoisomerase active site, Nucleic Acids Research, Volume 44, Issue 7, 20 April 2016, Pages 3464-3474). However, these are all formed from the amino-terminal (N-terminal) of the protein and contain at least the topoisomerase domain.
[0009] SUMMARY OF THE INVENTION
[0010] The inventors have developed novel polynucleotide-binding proteins based on fragments of DNA Topoisomerase V from Methanopyrus kandleri (Topo V). The inventors have surprisingly demonstrated increased polynucleotide delivery by coupling the polynucleotide to a membrane in which the relevant detector is present using a polynucleotide-binding protein of the invention. This lowers by several orders of magnitude the amount of polynucleotide required in order to be detected.
[0011] The polynucleotide-binding proteins of the invention are capable of binding any polynucleotide with any sequence. In other words, the polynucleotide-binding proteins of the invention do not require specific sequences for binding. This allows the polynucleotide-binding proteins to be used to couple any polynucleotide to a membrane. The polynucleotide-binding proteins of the invention can bind directly to and couple the polynucleotide itself without having to modify or functionalise the polynucleotide, for instance using a sequencing adaptor. This avoids the need for a modification or functionalisation step in any method using the polynucleotide-binding proteins of the invention. The direct binding of the polynucleotide-binding proteins of the invention to polynucleotides means they may also be used to couple a polynucleotide to a membrane without purification or enrichment of the polynucleotide. Overall, using the polynucleotide-binding protein results in an efficient method of coupling the polynucleotide to the membrane.
[0012] The invention provides a polynucleotide-binding protein comprising (a) a fragment of DNA Topoisomerase V from Methanopyrus kandleri (Topo V) or (b) a variant thereof, wherein the protein lacks the amino-terminal (N-terminal) topoisomerase domain and one or more of tandem helix-hairpin-helix ((HhH)?) domains 1-9 from Topo V.The invention also provides:
[0013] A method for coupling a polynucleotide to a membrane, comprising coupling the polynucleotide to the membrane using one or more first coupling adaptors each comprising a polynucleotide-binding protein of the invention;
[0014] One or more coupling adaptors for coupling a polynucleotide to a membrane, each comprising a polynucleotide-binding protein of the invention and a membrane protein or a hydrophobic anchor;
[0015] A kit for coupling a polynucleotide to a membrane, comprising (1) one or more first coupling adaptors each comprising a polynucleotide-binding protein of the invention and a first oligonucleotide, and (2) one or more second coupling adaptors each comprising a membrane protein or a hydrophobic anchor and a second oligonucleotide which hybridizes to the first oligonucleotide;
[0016] A method for determining the presence, absence or one or more characteristics of a polynucleotide analyte, comprising (a) coupling the polynucleotide analyte to a membrane using a polynucleotide-binding protein of the invention, a method of the invention, one or more coupling adaptors of the invention or a kit of the invention and (b) allowing the coupled polynucleotide analyte to interact with a detector present in the membrane and thereby determining the presence, absence or one or more characteristics of the polynucleotide analyte;
[0017] Use of a polynucleotide-binding protein of the invention, a method of the invention, one or more coupling adaptors of the invention or a kit of the invention for coupling a polynucleotide or a plurality of polynucleotides to a membrane;
[0018] A membrane comprising a polynucleotide or a plurality of polynucleotides coupled to it using a polynucleotide-binding protein of the invention, a method of the invention, one or more coupling adaptors of the invention or a kit of the invention;
[0019] An array comprising a plurality of membranes of the invention;
[0020] A system comprising (a) a membrane of the invention or an array of the invention, (b) means for applying a potential across the membrane(s) and (c) means for detecting electrical or optical signals across the membrane(s);
[0021] An apparatus comprising a polynucleotide or a plurality of polynucleotides coupled to an in vitro membrane using a polynucleotide-binding protein of the invention, a method of the invention, one or more coupling adaptors of the invention or a kit of the invention;An apparatus produced by a method comprising coupling a polynucleotide or a plurality of polynucleotides to an in vitro membrane using a polynucleotide-binding protein of the invention, a method of the invention, one or more coupling adaptors of the invention or a kit of the invention;
[0022] A polynucleotide encoding a polynucleotide-binding protein of the invention;
[0023] An expression vector comprising a polynucleotide of the invention;
[0024] A host cell comprising a polynucleotide of the invention or an expression vector of the invention;
[0025] A kit for characterising a polynucleotide analyte comprising (a) a polynucleotide- binding protein of the invention, one or more coupling adaptors of the invention or a kit of the invention and (b) the components of a membrane;
[0026] A kit for characterising a polynucleotide analyte comprising (a) a membrane comprising a detector and (b) a polynucleotide-binding protein of the invention, one or more coupling adaptors of the invention or a kit of the invention; and
[0027] An apparatus for characterising a polynucleotide analyte in a sample, comprising (a) a plurality of membranes each comprising a detector and (b) a plurality of polynucleotide-binding proteins of the invention, a plurality of one or more coupling adaptors of the invention or a plurality of kits of the invention.
[0028] DESCRIPTION OF THE FIGURES
[0029] It is to be understood that Figures are for the purpose of illustrating particular embodiments of the invention only and are not intended to be limiting.
[0030] Figure 1: The number of reads obtained using a MinlON flow cell from ONT over a 12h run depends on the number of HhH motifs present in Topo V DNA-binding domain. One of the data points (3 motifs) is repeated because motifs may be removed either from the C- or the N-terminus, resulting in two distinct proteins with the same number of HhH motifs.
[0031] Figure 2: Generic scheme of a DNA-binding protein functionalised as a generic tether. The DNA-binding protein is functionalised with a hydrophobic group via a linker (e.g., polyethylene glycol-based) via one of several available bioconjugation chemistries. Details on the nature of the conjugation chemistry and hydrophobic groups are found in the main text.
[0032] DESCRIPTION OF THE SEQUENCE LISTING SEQ ID NOs: 1-17 are shown in the Table below.DETAILED DESCRIPTION
[0033] It is to be understood that different applications of the disclosed products and methods may be tailored to the specific needs in the art. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments of the invention only and is not intended to be limiting.
[0034] All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.
[0035] The invention, both as to organization and method of operation, together with features and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying Figures. The aspects and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout this specification to "embodiment(s)" means that a particular feature, structure, or characteristic described in connection with the embodiment(s) is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment", "in another embodiment" or "in a preferred embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may do so. Similarly, it should be appreciated that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.
[0036] Definitions
[0037] Where an indefinite or definite article is used when referring to a singular noun e.g., "a” or "an", "the", this includes a plural of that noun unless something else is specifically stated. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. The term "comprising" is interchangeable with "consisting of" or "consisting essentially of".Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein.
[0038] The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.
[0039] "About" as used herein when referring to a measurable value such as an amount, percentage, a temporal duration, and the like, is meant to encompass variations of ± 20 % or ± 10 %, more preferably ± 5 %, even more preferably ± 1 %, and still more preferably ± 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods. Any embodiment containing the term "about" includes the same feature without the term. For example, about 40% includes 40%.
[0040] "Polynucleotide" as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RIMA. The term "polynucleotide" as used herein, is a single or double stranded covalently linked sequence of nucleotides in which the 3' and 5' ends on each nucleotide are joined by phosphodiester bonds. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Polynucleotides may be manufactured synthetically in vitro or isolated from natural sources. Polynucleotides may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post-translational modification, for example 5'-capping with 7-methylguanosine, 3'-processing such as cleavage and polyadenylation, and splicing. Polynucleotides may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA), peptide nucleic acid (PNA) and (poly)ADPribose modified DNA. Sizes of polynucleotides are typically expressed as the number of base pairs (bp) or nucleotide pairs for double stranded polynucleotides, or in the case of single stranded polynucleotides as the number of nucleotides (nt). One thousand bp or nt equal a kilobase (kb). Polynucleotides of less than around 40 nucleotides in length are typically called "oligonucleotides" and may compriseprimers for use in manipulation of DNA such as via polymerase chain reaction (PCR). The term "polynucleotide" is interchangeable with "polynucleotide sequence", "nucleotide sequence", "DNA sequence", "nucleic acid", or "nucleic acid molecule(s)".
[0041] The term "amino acid" in the context of the present disclosure is used in its broadest sense and is meant to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., a R group) specific to each amino acid. The amino acids typically refer to naturally occurring L a-amino acids or residues. The commonly used one and three letter abbreviations for naturally occurring amino acids are used herein: A=Ala; C=Cys; D=Asp; E=Glu; F=Phe; G=Gly; H = His; I = Ile; K=Lys; L=Leu; M = Met;
[0042] N=Asn; P=Pro; Q=Gln; R=Arg; S=Ser; T=Thr; V=Val; W=Trp; and Y=Tyr (Lehninger, A. L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes D-amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurring amino acids that are not usually incorporated into proteins such as norleucine, and chemically synthesised compounds having properties known in the art to be characteristic of an amino acid, such as p-amino acids. For example, analogues or mimetics of phenylalanine or proline, which allow the same conformational restriction of the peptide compounds as do natural Phe or Pro, are included within the definition of amino acid. Such analogues and mimetics are referred to herein as "functional equivalents" of the respective amino acid. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.
[0043] The terms "polypeptide", and "peptide" are interchangeably used herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. Polypeptides can also undergo maturation or post-translational modification processes that may include, but are not limited to glycosylation, proteolytic cleavage, lipidization, signal peptide cleavage, propeptide cleavage, phosphorylation, and such like. A peptide can be made using recombinant techniques, e.g. , through the expression of a recombinant or synthetic polynucleotide. A recombinantly produced peptide it typically substantially free of culture medium, e.g., culture medium represents less than about 20 %, more preferably less than about 10 %, and most preferably less than about 5 % of the volume of the protein preparation.
[0044] In all discussion below, the term "residue" in connection with a polypeptide or protein is interchangeable with the term "position" (especially in relation to a specific sequence) or "amino acid".The term "protein" is used to describe a folded polypeptide having a secondary or tertiary structure. The protein may be composed of a single polypeptide or may comprise multiple polypepties that are assembled to form a multimer. The multimer may be a homooligomer, or a heterooligmer. The protein may be a naturally occurring or wild type protein, or a modified, or non-naturally, occurring protein. The protein may, for example, differ from a wild type protein by the addition, substitution, or deletion of one or more amino acids.
[0045] The term "polypeptide" when being used in the context of what is being coupled to the membrane in accordance with the invention is interchangeable with "protein". Similarly, the term "polypeptide analyte" is interchangeable with "protein analyte".
[0046] A "variant" of a polypeptide or protein encompasses peptides, oligopeptides, polypeptides, proteins, and enzymes having structural similarity with the unmodified or wild-type polypeptide or protein. Structural variants are defined below with reference to root mean square deviation (RMSD). Standard methods in the art can be used to determine structural similarity, including RMSD. Suitable methods include, but are not limited to, AlphaFold, PSIPRED, TM-Align, US-Align or FATCAT. In all instances herein, RMSD is preferably measured over all alpha carbon atoms (Co atoms). RMSD may be measured over all heavy atoms.
[0047] A "variant" of a polypeptide or protein encompasses peptides, oligopeptides, polypeptides, proteins, and enzymes having amino acid substitutions, deletions and / or insertions relative to the unmodified or wild-type polypeptide or protein in question and having similar biological and functional activity as the unmodified polypeptide or protein from which they are derived. The term "amino acid identity" as used herein refers to the extent that sequences are identical on an amino acid-by-amino acid basis over a window of comparison. Thus, a "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched residues, dividing the number of matched residues by the total number of residues in the window of comparison ( / .e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.
[0048] Standard methods in the art may be used to determine identity and / or homology. For example, the UWGCG Package provides the BESTFIT program which can be used to calculate homology, for example used on its default settings (Devereux et al (1984) Nucleic Acids Research 12, p387-395). The PILEUP and BLAST algorithms can be used to calculate homology or line up sequences (such as identifying equivalent residues or corresponding sequences (typically on their default settings)), for example as described in Altschul S. F. (1993) J Mol Evol 36:290-300; Altschul, S.F et al (1990) J Mol Biol 215:403-10. Softwarefor performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). In all instances herein, identity or homology is typically measured over the entire length of the reference sequence.
[0049] Corresponding positions or regions may be determined by standard techniques in the art. For example, the PILEUP and BLAST algorithms mentioned below can be used to align the sequence of the invention, such as a fragment, with the reference sequence, e.g., SEQ ID NO: 1, and identify corresponding positions or regions.
[0050] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designed the "normal" or "wild-type" form of the gene. In contrast, the term "modified", "mutant" or "variant" refers to a gene or gene product that displays modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications and / or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For instance, methionine (M) may be substituted with arginine (R) by replacing the codon for methionine (ATG) with a codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For instance, non-naturally occurring amino acids may be introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express the mutant monomer. Alternatively, they may be introduced by expressing the mutant monomer in E. coli that are auxotrophic for specific amino acids in the presence of synthetic ( / .e., non-naturally occurring) analogues of those specific amino acids. They may also be produced by naked ligation if the mutant monomer is produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side-chain volume. The amino acids introduced may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acids they replace. Alternatively, the conservative substitution may introduce another amino acid that is aromatic or aliphatic in the place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well-known in the art.
[0051] A variant protein can also be chemically modified in any way and at any site. A variant protein is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art. The variant protein may be chemically modified by the attachment of any molecule. For instance, the variant protein may be chemically modified by attachment of a dye or a fluorophore.
[0052] In the context of the invention, the terms "couple", "coupling" or the like are interchangeable with "tether", "tethering" or the like.
[0053] In the context of the invention, the term "one or more" includes the singular. For instance, one or more coupling adaptors each comprising a polynucleotide-binding protein of the invention includes a coupling adaptor comprising a polynucleotide-binding protein of the invention.
[0054] Polynucleotide-binding proteins of the invention
[0055] The invention provides a polynucleotide-binding protein. The protein is capable of binding to a polynucleotide. Polynucleotides are described below. The protein may bind any of the polynucleotides described below. The ability of the protein to bind a polynucleotide may be measured using routine methods known in the art, including those described in the Examples.
[0056] The polynucleotide-binding protein may be any length. The polynucleotide-binding protein may be from about 70 to about 400 amino acids in length. The polynucleotide-binding protein may be from about 79 to about 350 amino acids, from about 80 to about 300 amino acids, from about 100 to about 250 amino acids, from about 120 to about 200 amino acids or from about 140 to about 180 amino acids in length. The polynucleotide-binding protein is preferably less than about 400 amino acids in length. The polynucleotide-binding protein may be less than about 350 amino acids, less than about 300 amino acids, less than about 250 amino acids or less than about 200 amino acids in length.
[0057] The polynucleotide-binding protein comprises (a) a fragment of DNA Topoisomerase V from Methanopyrus kandleri (Topo V) or (b) a variant thereof. The variant in (b) is a variant of a fragment in (a). Variants are defined in more detail below.
[0058] The protein or fragment preferably does not comprise the sequences of Topo-78, Topo-86, Topo-91, Topo-97, Topo-103 and full length Topo V. Full length Topo V is shown in SEQ ID NO: 1. Topo-78, Topo-86, Topo-91, Topo-97 and Topo-103 are described in Rajan et al. (2013) supra and Rakhi Rajan et al. (2016) supra. Methanopyrus kandleri topoisomerase V contains three distinct AP lyase active sites in addition to the topoisomerase active site, Nucleic Acids Research, Volume 44, Issue 7, 20 April 2016, Pages 3464-3474. The residues in SEQ ID NO: 1 that relate to Topo-78, Topo-86, Topo-91, Topo-97 and Topo-103 are shown in Table 1.
[0059] Table 1 - Known Topo V fragmentsFragment Residues in SEQ ID NO: 1
[0060] Topo-78 1-685
[0061] Topo-86 1-751
[0062] Topo-91 1-802
[0063] Topo-97 1-854
[0064] Topo-103 1-907
[0065]
[0066] The protein or fragment preferably does not comprise the sequences shown in residues 1-685 of SEQ ID NO: 1, residues 1-751 of SEQ ID NO: 1, residues 1-802 of SEQ ID NO: 1, residues 1-854 of SEQ ID NO: 1, residues 1-907 of SEQ ID NO: 1 and SEQ ID NO: 1.
[0067] First embodiment
[0068] In one embodiment, the protein or the fragment lacks the amino-terminal (N-terminal) topoisomerase domain and one or more of tandem helix-hairpin-helix ((HhH)?) domains 1-9 from Topo V.
[0069] The N-terminal topoisomerase domain corresponds to residues 1-268 of SEQ ID NO: 1.
[0070] The protein or the fragment preferably lacks the amino-terminal (N-terminal) topoisomerase domain and repair site I. The N-terminal topoisomerase domain and repair site I correspond to residues 1-291 of SEQ ID NO: 1. The N-terminal topoisomerase domain and repair site I correspond to SEQ ID NO: 2.
[0071] Wild-type Topo V contains twelve tandem helix-hairpin-helix ((HhH)?) domains. These are shown in Table 3. Helix-hairpin-helix (HhH) motifs are a protein structures that bind to DNA without sequence specificity. Their structures and function are known in the art. HhH motifs can be identified or predicted using structural similarity searches. HhH motifs in a row form a common, larger structure containing 5 alpha-helices that is termed a tandem helix-hairpin-helix ((HhH)?) domains. These motifs and domains are known in the art (Aidan J. Doherty, Louise C. Serpell, Christopher P. Ponting, The Helix-Hairpin-Helix DNA-Binding Motif: A Structural Basis for Non-Sequence-Specific Recognition of DNA, Nucleic Acids Research, Volume 24, Issue 13, 1 July 1996, Pages 2488-2497; and Xuguang Shao, Nick V. Grishin, Common fold in helix-hairpin-helix proteins, Nucleic Acids Research, Volume 28, Issue 14, 15 July 2000, Pages 2643-2650).The protein or fragment may lack any number and combination of (HhH)? domains 1-9 from Topo V. The protein or fragment may lack 1, 2, 3, 4, 5, 6, 7, 8 or 9 of (HhH)? domains 1-9 from Topo V. The protein or fragment preferably lacks all of (HhH)? domains 1-9 from Topo V.
[0072] The protein or fragment preferably lacks (HhH)? domain 1 or (HhH)? domains 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8 or 1-9 from Topo V.
[0073] The protein or fragment preferably lacks (HhH)? domain 1 from Topo V. (HhH)? domain 1 corresponds to residues 292-350 of SEQ ID NO: 1. (HhH)? domain 1 corresponds to SEQ ID NO: 3.
[0074] The protein or fragment preferably lacks (HhH)? domains 1-2 from Topo V. (HhH)? domains 1-2 correspond to residues 292-417 of SEQ ID NO: 1. (HhH)? domains 1-2 correspond to SEQ ID NOs: 3-4.
[0075] The protein or fragment preferably lacks (HhH)? domains 1-3 from Topo V. (HhH)? domains 1-3 correspond to residues 292-465 of SEQ ID NO: 1. (HhH)? domains 1-3 correspond to SEQ ID NOs: 3-5.
[0076] The protein or fragment preferably lacks (HhH)? domains 1-4 from Topo V. (HhH)? domains 1-4 correspond to residues 292-518 of SEQ ID NO: 1. (HhH)? domains 1-4 correspond to SEQ ID NOs: 3-6.
[0077] The protein or fragment preferably lacks (HhH)? domains 1-5 from Topo V. (HhH)? domains 1-5 correspond to residues 292-567 of SEQ ID NO: 1. (HhH)? domains 1-5 correspond to SEQ ID NOs: 3-7.
[0078] The protein or fragment preferably lacks (HhH)? domains 1-6 from Topo V. (HhH)? domains 1-6 correspond to residues 292-618 of SEQ ID NO: 1. (HhH)? domains 1-6 correspond to SEQ ID NOs: 3-8.
[0079] The protein or fragment preferably lacks (HhH)? domains 1-7 from Topo V. (HhH)? domains 1-7 correspond to residues 292-685 of SEQ ID NO: 1. (HhH)? domains 1-7 correspond to SEQ ID NOs: 3-9.
[0080] The protein or fragment preferably lacks (HhH)? domains 1-8 from Topo V. (HhH)? domains 1-8 correspond to residues 292-751 of SEQ ID NO: 1. (HhH)? domains 1-8 correspond to SEQ ID NOs: 3-10.
[0081] The protein or fragment preferably lacks (HhH)? domains 1-2 from Topo V. (HhH)? domains 1-9 correspond to residues 292-802 of SEQ ID NO: 1. (HhH)? domains 1-9 correspond to SEQ ID NOs: 3-11.The protein or fragment preferably further lacks a portion of (HhH)? domain 10 from Topo V. In the context of the invention the term "portion" is used to refer to a region of Topo V that is removed from the full-length protein, and the term "part" is used to refer to the region that remains when a portion is removed. The protein or fragment preferably lacks the N-terminal topoisomerase domain, repair site I, (HhH)? domains 1-9 and a portion of (HhH)? domain 10 from Topo V. (HhH)? domain 10 from Topo V corresponds to residues 803-852 of SEQ ID NO: 1. (HhH)? domain 10 from Topo V corresponds to SEQ ID NO: 12.
[0082] The protein or fragment may lack any amount of (HhH)? domain 10 or SEQ ID NO: 12. The portion may correspond to any amount of (HhH)? domain 10 or SEQ ID NO: 12. The portion preferably comprises a HhH motif. These can be identified as described above. The portion may be at least about 5% of (HhH)? domain 10 or SEQ ID NO: 12. The portion may be at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 98% or at least about 99% of the (HhH)? domain 10 or SEQ ID NO: 12. %s are usually calculated on the basis on the number of amino acids in the portion compared with the number of amino acids in the (HhH)? domain 10 or SEQ ID NO: 12. The portion may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, or at least 53 amino acids of (HhH)? domain 10 or SEQ ID NO: 12. The portion may be any part of the (HhH)? domain 10 or SEQ ID NO: 12. The portion may be at the N-terminus of the (HhH)? domain 10 or SEQ ID NO: 12. The portion may be at the C-terminus of the (HhH)? domain 10 or SEQ ID NO: 12. The portion may be consecutive amino acids within the sequence of the (HhH)? domain 10 or SEQ ID NO: 12. The portion is preferably 26 amino acids at the N-terminus of (HhH)? domain 10 or SEQ ID NO: 12.
[0083] The protein or fragment preferably comprises (HhH)? domains 2-12, 3-12, 4-12, 5-12, 6-12, 7-12, 8-12, 9-12, 10-12 or 11-12 from Topo V.
[0084] The protein or fragment preferably comprises (HhH)? domains 2-12 from Topo V. (HhH)? domains 2-12 correspond to residues 351-984 of SEQ ID NO: 1. (HhH)? domains 2-12 correspond to SEQ ID NOs: 4-14.
[0085] The protein or fragment preferably comprises (HhH)? domains 3-12 from Topo V. (HhH)? domains 3-12 correspond to residues 418-984 of SEQ ID NO: 1. (HhH)? domains 3-12 correspond to SEQ ID NOs: 5-14.The protein or fragment preferably comprises (HhH)? domains 4-12 from Topo V. (HhH)? domains 4-12 correspond to residues 466-984 of SEQ ID NO: 1. (HhH)? domains 4-12 correspond to SEQ ID NOs: 6-14.
[0086] The protein or fragment preferably comprises (HhH)? domains 5-12 from Topo V. (HhH)? domains 5-12 correspond to residues 519-984 of SEQ ID NO: 1. (HhH)? domains 5-12 correspond to SEQ ID NOs: 7-14.
[0087] The protein or fragment preferably comprises (HhH)? domains 6-12 from Topo V. (HhH)? domains 6-12 correspond to residues 568-984 of SEQ ID NO: 1. (HhH)? domains 6-12 correspond to SEQ ID NOs: 8-14.
[0088] The protein or fragment preferably comprises (HhH)? domains 7-12 from Topo V. (HhH)? domains 7-12 correspond to residues 619-984 of SEQ ID NO: 1. (HhH)? domains 7-12 correspond to SEQ ID NOs: 9-14.
[0089] The protein or fragment preferably comprises (HhH)? domains 8-12 from Topo V. (HhH)? domains 8-12 correspond to residues 686-984 of SEQ ID NO: 1. (HhH)? domains 8-12 correspond to SEQ ID NOs: 10-14.
[0090] The protein or fragment preferably comprises (HhH)? domains 9-12 from Topo V. (HhH)? domains 9-12 correspond to residues 752-984 of SEQ ID NO: 1. (HhH)? domains 9-12 correspond to SEQ ID NOs: 11-14.
[0091] The protein or fragment preferably comprises (HhH)? domains 10-12 from Topo V. (HhH)? domains 10-12 correspond to residues 803-984 of SEQ ID NO: 1. (HhH)? domains 10-12 correspond to SEQ ID NOs: 12-14.
[0092] The protein or fragment preferably comprises (HhH)? domains 11-12 from Topo V. (HhH)? domains 11-12 correspond to residues 853-984 of SEQ ID NO: 1. (HhH)? domains 11-12 correspond to SEQ ID NOs: 13-14.
[0093] The protein or fragment preferably comprises (i) a part of (HhH)? domain 10 from Topo V and (ii) (HhH)? domains 11-12 from Topo V. (HhH)? domain 10 from Topo V corresponds to residues 803-852 of SEQ ID NO: 1. (HhH)? domain 10 from Topo V corresponds to SEQ ID NO: 12. (HhH)? domains 11-12 correspond to residues 853-984 of SEQ ID NO: 1. (HhH)? domains 11-12 correspond to SEQ ID NOs: 13-14. The part is complementary to the portion described above, i.e., the part is what remains when the portion is removed from (HhH)? domain 10 or SEQ ID NO: 12.
[0094] The protein or fragment may comprise any amount of (HhH)? domain 10 or SEQ ID NO: 12. The part may correspond to any amount of (HhH)? domain 10 or SEQ ID NO: 12. The part preferably comprises a HhH motif. These can be identified as described above. The part maybe at least about 5% of (HhH)? domain 10 or SEQ ID NO: 12. The part may be at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 98% or at least about 99% of the (HhH)? domain 10 or SEQ ID NO: 12. %s are usually calculated on the basis on the number of amino acids in the part compared with the number of amino acids in the (HhH)? domain 10 or SEQ ID NO: 12. The part may comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, or at least 53 amino acids of (HhH)? domain 10 or SEQ ID NO: 12. The portion may be any part of the (HhH)? domain 10 or SEQ ID NO: 12. The part may be at the N-terminus of the (HhH)? domain 10 or SEQ ID NO: 12. The part may be at the C-terminus of the (HhH)? domain 10 or SEQ ID NO: 12. The part may be consecutive amino acids within the sequence of the (HhH)? domain 10 or SEQ ID NO: 12. The part is preferably 28 amino acids at the C-terminus of (HhH)2 domain 10 or SEQ ID NO: 12.
[0095] Second embodiment
[0096] In another embodiment, the protein or fragment comprises (HhH)2 domains 11-12 from Topo V or a variant thereof. As explained above, (HhH)2 domains 11-12 correspond to residues 853-984 of SEQ ID NO: 1. (HhH)2 domains 11-12 correspond to SEQ ID NOs: 13-14. Any of the (HhH)2 domains 11-12 embodiments described above for the first embodiment equally apply to the second embodiment.
[0097] The protein or fragment preferably further comprises a part of (HhH)2 domain 10 from Topo V. The part of (HhH)2 domain 10 from Topo V may be any of those described above with reference to the first embodiment.
[0098] The protein or fragment preferably lacks the N-terminal topoisomerase domain and one or more, or all, of (HhH)2 domains 1-9 from Topo V. The protein or fragment may lack any number and combination of (HhH)2 domains 1-9 from Topo V. The protein or fragment may lack 1, 2, 3, 4, 5, 6, 7, 8 or 9 of (HhH)2 domains 1-9 from Topo V. The protein or fragment preferably lacks all of (HhH)2 domains 1-9 from Topo V. The protein or fragment preferably lacks (HhH)2 domain 1 or (HhH)2 domains 1-2, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8 or 1-9 from Topo V. Any of the embodiments relating to (HhH)2 domains 1-9 from Topo V described above with reference to the first embodiment equally apply to the second embodiment.Preferred fragments
[0099] The protein or fragment preferably comprises or consists of the sequence shown in SEQ ID NO: 15. This sequence is formed from (i) a part (the C-terminal 28 amino acids) of (HhH)? domain 10 from Topo V and (ii) (HhH)? domains 11-12 from Topo V. The protein preferably comprises a variant of SEQ ID NO: 15. The variant may be any of those described below.
[0100] Variants
[0101] The variant preferably comprises (i) a sequence having at least about 40% identity and / or (ii) homology to the sequence of the fragment and / or (b) has a root mean square deviation (RMSD) of less than about 4.0 Angstroms (A) when compared with the fragment. The fragment may be any of those described above.
[0102] In (i), the variant preferably comprises a sequence having at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98% or at least about 99% homology and / or identity to the sequence of the fragment. The variant preferably comprises a sequence having at least about 80% homology and / or identity to the sequence of the fragment. The variant preferably comprises a sequence having at least about 90% homology and / or identity to the sequence of the fragment. Homology and / or identity is typically measured over the entire length of the reference sequence. The fragment may be any of those described above.
[0103] Sequence homology and / or identity can also relate to a region of the fragment. The region preferably comprises at least one (HhH)? domain. The (HhH)? domain may be any of the (HhH)? domains described above. A variant may have less than 40% overall sequence homology and / or identity with the fragment, but the sequence of a particular region may share at least about 80%, at least about 90%, or as much as about 99% sequence homology and / or identity with the corresponding region of the fragment. There may be at least about 80%, at least 85%, at least about 90%, at least about 95%, at least about 97%, at least about 98% or at least about 99% homology and / or identity over a stretch of about 10 or more, for example about 12, about 15, about 17, about 18, about 20, about 25 or more, contiguous amino acids in the fragment ("hard homology"). The fragment may be any of those described above.
[0104] In (ii), variant preferably has a root mean square deviation (RMSD) of less than about 3.5 A, less than about 3.0 A, less than about 2.5 A, less than about 2.0 A, less than about 1.5 A, less than about 1.0 A or less than about 0.5 A, when compared with the fragment. The fragment may be any of those described above.The variant may comprise one or more amino acid substitutions. Amino acid substitutions may be made to the amino acid sequence of the fragment, for example up to 1, 2, 3, 4, 5, 10, 20 or 30 substitutions. The one or more substitutions are preferably one or more conservative substitutions. Conservative substitutions are defined above.
[0105] The variant may be modified to introduce one or more cysteines, one or more hydrophobic amino acids, one or more charged amino acids, one or more non-native amino acids, or one or more polar amino acids. Any number and combination of such introductions may be made. The introduction is preferably by substitution or addition. Preferred substitutions are described in more detail below.
[0106] One or more amino acid residues of the sequence of the fragment may additionally be deleted from the polypeptides described above. Up to 1, 2, 3, 4, 5, 10, 20 or 30 or more residues may be deleted.
[0107] One or more amino acids may be alternatively or additionally added to the fragments described above. An extension may be provided at the N-terminus or C-terminus of the sequence of the fragment. The extension may be quite short, for example from about 1 to about 10 amino acids in length. Alternatively, the extension may be longer, for example up to about 50 or about 100 amino acids. The variant may comprises an amino acid extension of from about 10 to about amino acids, such as from about 50 to about 200 amino acids or about 100 to about 180 amino acids. The variant shown in SEQ ID NO: 17 comprises an extension of 160 amino acids at the C-terminus.
[0108] The variant preferably comprises or consists of the sequence shown in SEQ ID NO: 17.
[0109] The variant preferably comprises one or more cysteine residues introduced into the fragment by substitution. The variant may comprise any number of one or more cysteine residues introduced by substitution, such as 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 cysteine residues. The one or more cysteine residues are preferably introduced into the last 20 residues at the C-terminus of the fragment. The one or more cysteine residues may be introduced at any one or more positions in the last 20 residues at the C-terminus of the fragment. The variant preferably comprises a cysteine at the position corresponding to position 973 in SEQ ID NO: 1. The variant shown in SEQ ID NO: 16 differs from the fragment shown in SEQ ID NO: 15 by a cysteine at the position corresponding to position 973 in SEQ ID NO: 1. Corresponding positions can be determined as described above.
[0110] A variant of the fragment is a polypeptide that has an amino acid sequence which varies from that of the fragment and which retains its ability to bind a polynucleotide. A variant typically contains the (HhH)? domains that are responsible for polynucleotide binding.
[0111] Polynucleotides of the inventionThe invention also provides a polynucleotide encoding a polynucleotide-binding protein of the invention. The polynucleotide-binding protein may be any of those described above, including all the fragments and variants. The invention also provides an expression vector comprising a polynucleotide of the invention. The invention also provides a host cell comprising a polynucleotide of the invention or an expression vectors of the invention. Suitable vectors and host cells are known in the art.
[0112] Coupling methods of the invention
[0113] The invention provides a method for coupling a polynucleotide to a membrane. This method may also be called the coupling method of the invention.
[0114] Membranes are defined in more detail below. The method comprises coupling the polynucleotide to the membrane using one or more first coupling adaptors each comprising a polynucleotide-binding protein of the invention. The polynucleotide-binding protein of the invention may be any of those described above, including all the fragments and variants.
[0115] Coupling the polynucleotide to the membrane containing a detector lowers by several orders of magnitude the amount of polynucleotide required. The method is of course advantageous for detecting polynucleotides that are present at low concentrations. The method preferably allows the presence, absence or characteristics of the polynucleotide to be determined when the polynucleotide (also called the polynucleotide analyte below) is present at a concentration of from about O.OOlpM to about InM, such as less than O.OlpM, less than O.lpM, less than lpM, less than lOpM or less than lOOpM.
[0116] Coupling one end of the polynucleotide to the membrane, even temporarily, also means the end is prevented from interfering with any detector in the membrane or any detector-based process. Other advantages of the method of the invention are described above.
[0117] The coupling may be permanent or stable. In other words, the coupling may be such that the polynucleotide remains coupled to the membrane during the method. The coupling may be transient. In other words, the coupling may be such that the polynucleotide decouples from the membrane during the method. For certain applications, especially the detection or characterisation methods described below, the transient nature of the coupling is preferred. If the coupling is permanent or stable, then some polynucleotide data may be lost as the detector cannot continue to the end of the polynucleotide. If the coupling is transient, then when the coupled end of the polynucleotide randomly becomes free of the membrane, then the polynucleotide can be processed to completion. Chemical groups that form permanent / stable or transient links with the membrane are described in more detail below. The polynucleotide may be transiently coupled to the membrane using cholesterol or a fatty acyl chain. Any fatty acyl chain having a length of from about 6 to 30 about carbon atoms, such as hexadecanoic acid, may be used.Polynucleotide
[0118] The polynucleotide being coupled to the membrane can be any polynucleotide. The polynucleotide may be a polynucleotide that is secreted from cells. Alternatively, the polynucleotide can be a polynucleotide that is present inside cells such that the polynucleotide must be extracted from the cells before the invention can be carried out. The polynucleotide may be present in blood or an extract thereof.
[0119] The polynucleotide can be naturally-occurring or non-naturally-occurring. The polynucleotide can include within it synthetic or modified nucleotides. A number of different types of modification to nucleotides are known in the art. For the purposes of the invention, it is to be understood that the polynucleotide can be modified by any method available in the art.
[0120] The polynucleotide can be provided as an impure mixture of one or more polynucleotides and one or more impurities. Impurities may comprise truncated forms of the target polynucleotide which are distinct from the "polynucleotide analyte" or "target polynucleotide analyte" for characterisation. For example, the polynucleotide may be a full-length polynucleotide and impurities may comprise fractions of the polynucleotide. Impurities may also comprise polynucleotides other than the target polynucleotide, e.g., which may be copurified from a cell culture or obtained from a sample.
[0121] The polynucleotide or nucleic acid may comprise any combination of any nucleotides. The nucleotides can be naturally occurring or artificial. One or more nucleotides in the polynucleotide can be oxidized or methylated. One or more nucleotides in the polynucleotide may be damaged. For instance, the polynucleotide may comprise a pyrimidine dimer. Such dimers are typically associated with damage by ultraviolet light and are the primary cause of skin melanomas. One or more nucleotides in the polynucleotide may be modified, for instance with a label or a tag, for which suitable examples are known by a skilled person. The polynucleotide may comprise one or more spacers. A nucleotide typically contains a nucleobase, a sugar and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines and more specifically adenine (A), guanine (G), thymine (T), uracil (U) and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably a deoxyribose. The polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dll) and / or thymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC). The nucleotide is typically a ribonucleotide or deoxyribonucleotide. The nucleotide typically contains a monophosphate, diphosphate, or triphosphate. The nucleotide may comprise more than three phosphates, such as 4 or 5 phosphates.Phosphates may be attached on the 5' or 3' side of a nucleotide. The nucleotides in the polynucleotide may be attached to each other in any manner. The nucleotides are typically attached by their sugar and phosphate groups as in nucleic acids. The nucleotides may be connected via their nucleobases as in pyrimidine dimers. The polynucleotide may be single stranded or double stranded. At least a portion of the polynucleotide is preferably double stranded. The polynucleotide is most preferably ribonucleic nucleic acid (RIMA) or deoxyribonucleic acid (DNA).
[0122] The polynucleotide can be any length. For example, the polynucleotide can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400 or at least 500 nucleotides or nucleotide pairs in length. The polynucleotide can be 1000 or more nucleotides or nucleotide pairs, 5000 or more nucleotides or nucleotide pairs in length or 100000 or more nucleotides or nucleotide pairs in length. Any number of polynucleotides can be investigated. For instance, the method may concern characterising 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50, 100 or more polynucleotides. If two or more polynucleotides are characterised, they may be different polynucleotides or two instances of the same polynucleotide. The polynucleotide can be naturally occurring or artificial. For instance, the method may be used to verify the sequence of a manufactured oligonucleotide. The method is typically carried out in vitro.
[0123] Nucleotides can have any identity (ii), and include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP) and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP and dUMP. A nucleotide may be abasic ( / .e., lack a nucleobase). A nucleotide may also lack a nucleobase and a sugar ( / .e., is a C3 spacer). The sequence of the nucleotides (iii) is determined by the consecutive identity of following nucleotides attached to each other throughout the polynucleotide strain, in the 5' to 3' direction of the strand.
[0124] Any number of polynucleotides may be coupled to the membrane in accordance with the invention. From about 2 to about 1 x 1013, from about 5 to about 1 x 1012, from about 10 to about 1 x 1011, from about 20 to about 1 x 1010, from about 50 to 1 x 1010, from about 100 to 1 x 109, from about 100 to 1 x 108, from about 500 to 1 x 107, from about 1,000 to 1 x 108, from about 5,000 to 1 x 107or from about 10,000 to 1 x 105polynucleotides may be coupled to the membrane. At least about 10, at least about 50, at least about 100, at least about 500, at least about 1,000, at least about 5,000, at least about 10,000, at least about50,000, at least about 1 x 105, at least about 1 x 106, at least about 1 x 107, at least about 1 x 108, at least about 1 x 109, at least about 1 x IO10, at least about 1 x 1011, at least about 1 x 1012or at least about 1 x 1013or more polynucleotides can be coupled to the membrane. In the context of the invention, multiple instances of a polynucleotide are called a plurality of polynucleotides. The plurality may comprise any of the numbers of polynucleotides in this paragraph. The polynucleotides may be the same. The polynucleotides may be a homologous plurality of polynucleotides. The polynucleotides may be different. The polynucleotides may be a heterologous plurality of different polynucleotides.
[0125] The invention provides a coupling a plurality of polynucleotides to a membrane, comprising coupling the plurality of polynucleotides to the membrane using one or more first coupling adaptors each comprising a polynucleotide-binding protein of the invention. The polynucleotide-binding protein of the invention may be any of those described above, including all the fragments and variants.
[0126] The polynucleotide may be eukaryotic, prokaryotic, viral, an artificially designed polynucleotide or an engineered polynucleotide. The polynucleotide may be a bacterial protein, fungal protein, virus protein or parasite-derived protein. The plurality of polynucleotides may comprise any of these polynucleotides.
[0127] The polynucleotide or the plurality of polynucleotides is typically present in or derived from any suitable sample. The invention is typically carried out on a sample that is known to contain or suspected to contain the polynucleotide or the plurality of polynucleotides.
[0128] The sample may be a biological sample. The invention may be carried out in vitro on a sample obtained from or extracted from any organism or microorganism. The organism or microorganism is typically archaeal, prokaryotic or eukaryotic and typically belongs to one of the five kingdoms: plantae, animalia, fungi, monera and protista.
[0129] The sample is preferably a fluid sample. The sample typically comprises a body fluid of the patient. The sample may be urine, lymph, saliva, mucus or amniotic fluid but is preferably blood, plasma or serum. Typically, the sample is human in origin, but alternatively it may be from another mammal animal such as from commercially farmed animals such as horses, cattle, sheep or pigs or may alternatively be pets such as cats or dogs. Alternatively a sample of plant origin is typically obtained from a commercial crop, such as a cereal, legume, fruit or vegetable, for example wheat, barley, oats, canola, maize, soya, rice, bananas, apples, tomatoes, potatoes, grapes, tobacco, beans, lentils, sugar cane, cocoa or cotton.
[0130] The polynucleotide or the plurality of polynucleotides may be derived from one or more cells. The cells may be any type of cells. The cells may be a prokaryotic cells. The cells maybe bacterial or archaeal. The cells are typically eukaryotic cells. The cells may be a protozoan, algal, fungal, plant or animal cells. The animal cells may be derived from the ectoderm, endoderm, or mesoderm. The cells may be stem cells, such as embryonic stem cells, induced pluripotent stem cells or mesenchymal stem cells, bone cells, such as osteoclasts, osteoblasts or osteocytes, tendon cells, such as tenoblasts or tenocytes, chondrocytes, synovial cells, vascular cells, blood cells, such as red blood cells, immune cells, platelet, neutrophils or basophils, muscle cells, such as skeletal muscle cells, cardiac muscle cells or smooth muscle cells, reproductive cells, such as sperm, oocytes, duct cells or epididymal cells, secretory cells, adipocytes, liver lipocytes, epithelial cells, odontoblasts, cementoblasts, hormone-secreting cells, barrier cells, exocrine secretory epithelial cells, nerve cells, astrocytes, oligodendrocytes, or neurons.
[0131] The immune cells may be neutrophil granulocyte and precursors, such as myeloblasts, promyelocytes, myelocytes, or metamyelocytes, eosinophil granulocyte and precursors, basophil granulocyte and precursors, mast cells, leukocytes, lymphocytes, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, B cells, macrophages, dendritic cells, plasma cells, neutrophils, or monocytes.
[0132] The cells may be wild-type or naturally occurring. The cells may be genetically modified or genetically engineered. For instance, the immune cells may be genetically engineered to express a recombinant chimeric antigen receptor (CAR) or T cell receptor (TCR). The cells may be genetically modified or genetically engineered using transduction or transfection or any other common techniques known to those skilled in the art.
[0133] The cells may be healthy cells or obtained from a healthy donor or source. The cells may be diseased or damaged, associated with a disease or damage or obtained from a diseased or damaged donor or source.
[0134] The sample may be a non-biological sample. The non-biological sample is preferably a fluid sample. Examples of a non-biological sample include surgical fluids, water such as drinking water, sea water or river water, and reagents for laboratory tests.
[0135] The sample is typically processed prior to being assayed, for example by centrifugation or by passage through a membrane that filters out unwanted molecules or cells, such as red blood cells. The sample may be measured immediately upon being taken. The sample may also be typically stored prior to assay, preferably below -70°C. The polynucleotide or the plurality of polynucleotides is typically extracted from the sample before it is used in the coupling method of the invention. Polynucleotide extraction kits are commercially available from, for instance, Thermo Fischer Scientific® and QIAGEN®.
[0136] Before step (a), the polynucleotide analyte is preferably not purified or is preferably not separated from other components in a sample.Membrane
[0137] Any suitable membrane may be used in the coupling method of the invention. The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties. The amphiphilic molecules may be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form a monolayer are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer sub-units that are polymerized together to create a single polymer chain. Block copolymers typically have properties that are contributed by each monomer sub-unit. However, a block copolymer may have unique properties that polymers formed from the individual sub-units do not possess. Block copolymers can be engineered such that one of the monomer subunits is hydrophobic ( / .e., lipophilic), whilst the other sub-unit(s) are hydrophilic whilst in aqueous media. In this case, the block copolymer may possess amphiphilic properties and may form a structure that mimics a biological membrane. The block copolymer may be a diblock (consisting of two monomer sub-units) but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphipiles. The copolymer may be a triblock, tetrablock or pentablock copolymer. The membrane may be a triblock copolymer membrane.
[0138] Archaebacterial bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipid forms a monolayer membrane. These lipids are generally found in extremophiles that survive in harsh biological environments, thermophiles, halophiles and acidophiles. Their stability is believed to derive from the fused nature of the final bilayer. It is straightforward to construct block copolymer materials that mimic these biological entities by creating a triblock polymer that has the general motif hydrophilic-hydrophobic-hydrophilic. This material may form monomeric membranes that behave similarly to lipid bilayers and encompass a range of phase behaviours from vesicles through to laminar membranes. Membranes formed from these triblock copolymers hold several advantages over biological lipid membranes. Because the triblock copolymer is synthesised, the exact construction can be carefully controlled to provide the correct chain lengths and properties required to form membranes and to interact with pores and other proteins.
[0139] Block copolymers may also be constructed from sub-units that are not classed as lipid submaterials; for example, a hydrophobic polymer may be made from siloxane or other non-hydrocarbon-based monomers. The hydrophilic sub-section of block copolymer can also possess low protein binding properties, which allows the creation of a membrane that is highly resistant when exposed to raw biological samples. This head group unit may also be derived from non-classical lipid head-groups.Triblock copolymer membranes also have increased mechanical and environmental stability compared with biological lipid membranes, for example a much higher operational temperature or pH range. The synthetic nature of the block copolymers provides a platform to customise polymer-based membranes for a wide range of applications.
[0140] The membrane may be one of the membranes disclosed in WO 2014 / 064443 or WO 2014 / 064444 (both of which are incorporated herein by reference in their entireties).
[0141] The amphiphilic molecules may be chemically modified or functionalised to facilitate coupling of the polynucleotide. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.
[0142] Amphiphilic membranes are typically naturally mobile, essentially acting as two-dimensional fluids with lipid diffusion rates of approximately IO-8cm s’1. This means that the pore and coupled polynucleotide can typically move within an amphiphilic membrane.
[0143] The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, a planar lipid bilayer, a supported bilayer, or a liposome. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, WO 2009 / 077734, and WO 2006 / 100484 (incorporated herein by reference in their entireties).
[0144] Methods for forming lipid bilayers are known in the art. Lipid bilayers are commonly formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566).
[0145] A lipid bilayer may be formed as described in WO 2009 / 077734 (incorporated herein by reference in its entirety). In this method, the lipid bilayer is formed from dried lipids. A lipid bilayer may be formed across an opening as described in W02009 / 077734.
[0146] The membrane may comprise a solid-state layer. Solid state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as SisN4, AI2O3, and SiO, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber, and glasses. The solid-state layer may be formed from graphene. Suitable graphene layers are disclosed in WO 2009 / 035647 (incorporated herein by reference in its entirety). If the membrane comprises a solid-state layer, the pore is typically present in an amphiphilic membrane or layer contained within the solid-state layer, for instance within ahole, well, gap, channel, trench or slit within the solid-state layer. The skilled person can prepare suitable solid state / amphiphilic hybrid systems. Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857 (incorporated herein by reference in their entireties). Any of the amphiphilic membranes or layers described above may be used.
[0147] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated, naturally occurring lipid bilayer comprising a pore, or (iii) a cell having a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. The layer may comprise other transmembrane and / or intramembrane proteins as well as other molecules in addition to the pore. Suitable apparatus and conditions are described below. The coupling method of the invention and / or the detection or characterisation method of the invention is typically carried out in vitro.
[0148] The polynucleotide or the plurality of polynucleotides may be coupled to the membrane via the adaptors using any known method including those described in WO 2012 / 164270 (incorporated by reference in its entirety). If the membrane is an amphiphilic layer, such as a lipid bilayer or any of the copolymer membranes described above, the polynucleotide plurality of polynucleotides is preferably coupled to the membrane via a protein present in the membrane (also known as a membrane protein) or one or more hydrophobic anchors present in the membrane. These may be form part of the one or more second coupling adaptors as described in more detail below. Any number of one or more hydrophobic anchors may be used, such as 1, 2, 3, 4, 5 or more hydrophobic anchors. One or two hydrophobic anchors are preferably used. The one or more hydrophobic anchors are preferably selected from a lipid, fatty acid, sterol, carbon nanotube, polypeptide, protein or amino acid. The one or more hydrophobic anchors may be cholesterol, palmitate or tocopherol. The one or more hydrophobic anchors are preferably tocopherol. The one or more hydrophobic anchors are preferably octotocopherol. The one or more hydrophobic anchors are preferably one or more cholesterols. In preferred embodiments, the polypeptide is not coupled to the membrane via the detector.
[0149] The hydrophobic modification may for example comprise a phosphorothioate such as a charge-neutralized alkyl-phosphorothioate (PPT) as described in Jones et al, J. Am. Chem. Soc. 2021, 143, 22, 8305 (incorporated by reference in its entirety). Suitable alkyl groups include for example C1-C10 alkyl groups such as C2-C6 alkyl groups, e.g., methyl, ethyl, propyl, butyl, pentyl and hexyl groups. Incorporation of the charge-neutralized alkyl-phosphorothioate into an adaptor allows the adaptor to couple to a hydrophobic region such the membrane.
[0150] The components of the membrane, such as the copolymer monomers, amphiphilic molecules or lipids, may be chemically-modified or functionalised to facilitate coupling of thepolynucleotide or the plurality of polynucleotides to the membrane via the adaptors.
[0151] Examples of suitable chemical modifications and suitable ways of functionalising the components of the membrane are described in more detail below. Any proportion of the membrane components may be functionalized, for example at least 0.01%, at least 0.1%, at least 1%, at least 10%, at least 25%, at least 50% or 100%.
[0152] Coupling of molecules to synthetic lipid bilayers has been carried out previously with various different tethering strategies. These are summarised in Table 2 below. Any of these may be used to couple the polynucleotide or the plurality of polynucleotides to the membrane via the adaptors
[0153] Table 2 - Coupling chemistries
[0154] Attachment Type of coupling Reference
[0155] group
[0156] Thiol Stable Yoshina-Ishii, C. and S. G. Boxer (2003).
[0157] "Arrays of mobile tethered vesicles on supported lipid bilayers." J Am Chem Soc 125(13): 3696-7.
[0158] Biotin Stable Nikolov, V., R. Lipowsky, et al. (2007).
[0159] "Behavior of giant vesicles with anchored DNA molecules." Biophvs J 92(12): 4356-68
[0160] Cholesterol Transient Pfeiffer, I. and F. Hook (2004). "Bivalent cholesterol-based coupling of oligonucletides to lipid membrane assemblies." J Am Chem Soc 126(33): 10224-5
[0161] Surfactant Transient van Lengerich, B., R. J. Rawle, et al.
[0162] "Covalent attachment of lipid vesicles to a fluid-supported bilayer allows observation of DNA-mediated vesicle interactions." Langmuir 26(11): 8666-72
[0163] Palmitate Transient van Lengerich, B., R. J. Rawle, et al.
[0164] "Covalent attachment of lipid vesicles to a fluid-supported bilayer allows observation of DNA-mediated vesicle interactions." Langmuir 26(11): 8666-72
[0165]
[0166] Any of the adaptors described below may be functionalised using a modified phosphoramidite in the synthesis reaction, which is easily compatible for the direct addition of suitable coupling moieties, such as cholesterol, tocopherol or palmitate, as well as for reactive groups, such as thiol, cholesterol, lipid and biotin groups. The one or more adaptors below preferably comprise a tocopherol, such as octotocopherol. These different attachment chemistries give a suite of options for attachment to adaptors. Each different modification group tethers the adaptors in a slightly different way and coupling is not always permanent so giving different dwell times for the polynucleotide or the plurality of polynucleotides to the membrane. The advantages of transient coupling are described above.
[0167] Coupling to a membrane or functionalised membrane can also be achieved by a number of other means provided that a complementary reactive group or a tether can be added to the relevant adaptor. The addition of reactive groups to either end of polynucleotide adaptors has been reported previously. A thiol group can be added to the 5' of ssDNA or dsDNA using T4 polynucleotide kinase and ATPyS (Grant, G. P. and P. Z. Qin (2007). "A facile method for attaching nitroxide spin labels at the 5’ terminus of nucleic acids." Nucleic Acids Res 35(10): e77). An azide group could be added to the 5'-phosphate of ssDNA or dsDNA using T4 polynucleotide kinase and y-[2-Azidoethyl]-ATP or y-[6-Azidohexyl]-ATP. Using thiol or Click chemistry a tether, containing either a thiol, iodoacetamide OPSS or maleimide group (reactive to thiols) or a DIBO (dibenzocyclooxtyne) or alkyne group (reactive to azides), can be covalently attached to the analyte. A more diverse selection of chemical groups, such as biotin, thiols and fluorophores, can be added using terminal transferase to incorporate modified oligonucleotides to the 3' of ssDNA (Kumar, A., P. Tchen, et al. (1988).
[0168] "Nonradioactive labeling of synthetic oligonucleotide probes with terminal deoxynucleotidyl transferase." Anal Biochem 169(2): 376-82). Streptavidin / biotin coupling may be used. It may also be possible that tethers could be directly added to polynucleotide adaptors using terminal transferase with suitably modified nucleotides (e.g. cholesterol or palmitate).
[0169] The reactive group may be ligated to a single strand or double stranded polynucleotide adaptor. Ligation of short pieces of ssDNA have been reported using T4 RNA ligase I (Troutt, A. B., M. G. McHeyzer-Williams, et al. (1992). "Ligation-anchored PCR: a simple amplification technique with single-sided specificity." Proc Natl Acad Sci U S A 89(20):
[0170] 9823-5).
[0171] Adenylated nucleic acids (AppDNA) are intermediates in ligation reactions, where an adenosine-monophostate is attached to the 5'-phosphate of the nucleic acid. Various kits are available for generation of this intermediate, such as the 5' DNA Adenylation Kit from NEB. By substituting ATP in the reaction for a modified nucleotide triphosphate, then addition of reactive groups (such as thiols, amines, biotin, azides, etc) to the 5' of polynucleotide adaptors should be possible. It may also be possible that tethers could bedirectly added to polynucleotide adaptors using a 5' DNA adenylation kit with suitably modified nucleotides (e.g. cholesterol or palmitate).
[0172] A common technique for the amplification of sections of genomic DNA is using polymerase chain reaction (PCR). Here, using two synthetic oligonucleotide primers, a number of copies of the same section of DNA can be generated, where for each copy the 5' of each strand in the duplex will be a synthetic polynucleotide. By using an antisense primer single or multiple nucleotides can be added to 3' end of single or double stranded DNA by employing a polymerase. Examples of polymerases which could be used include, but are not limited to, Terminal Transferase, Klenow and E. coli Poly(A) polymerase). By substituting ATP in the reaction for a modified nucleotide triphosphate then reactive groups, such as a cholesterol, thiol, amine, azide, biotin or lipid, can be incorporated into the polynucleotide adaptors. Therefore, each copy of the amplified polynucleotide adaptors will contain a reactive group for coupling.
[0173] Coupling polynucleotide adaptors to a membrane can also be achieved by anchoring a binding group, such as a polynucleotide binding protein or a chemical group, to the membrane and allowing the binding group to interact with the polynucleotide adaptors or by functionalizing the membrane. The binding group may be coupled to the membrane by any of the methods described herein. In particular, the binding group may be coupled to the membrane using one or more linkers, such as maleimide functionalised linkers.
[0174] The binding group can be any group that interacts with single or double stranded nucleic acids, specific nucleotide sequences within the polynucleotide adaptors or patterns of modified nucleotides within the polynucleotide adaptors, or any other ligand that is present on the polynucleotide.
[0175] Suitable binding proteins include E. coli single stranded binding protein, P5 single stranded binding protein, T4 gp32 single stranded binding protein, the TOPO V dsDNA binding region, human histone proteins, E. coli HU DNA binding protein and other archaeal, prokaryotic or eukaryotic single- or double-stranded nucleic acid binding proteins, including those listed below.
[0176] The specific nucleotide sequences in the polynucleotide adaptors could be sequences recognised by transcription factors, ribosomes, endonucleases, topoisomerases or replication initiation factors. The patterns of modified nucleotides could be patterns of methylation or damage.
[0177] The chemical group can be any group which intercalates with or interacts with a polynucleotide adaptor. The group may intercalate or interact with the polynucleotide adaptor via electrostatic, hydrogen bonding or Van der Waals interactions. Such groups include a lysine monomer, poly-lysine (which will interact with ssDNA or dsDNA), ethidiumbromide (which will intercalate with dsDNA), universal bases or universal nucleotides (which can hybridise with any polynucleotide analyte) and osmium complexes (which can react to methylated bases). A polynucleotide adaptor may therefore be coupled to the membrane using one or more universal nucleotides attached to the membrane. Each universal nucleotide residue may be attached to the membrane using one or more linkers. A universal nucleotide is preferably one which will hybridise or bind to some degree to nucleotides comprising the nucleosides adenosine (A), thymine (T), uracil (U), guanine (G) and cytosine (C). Universal nucleotides may hybridise or bind more strongly to some nucleotides than to others. For instance, a universal nucleotide (I) comprising the nucleoside, 2'-deoxyinosine, will show a preferential order of pairing of I-C>I-A>I-G approximately=I-T. The polymerase will replace a nucleotide species with a universal nucleotide if the universal nucleotide takes the place of the nucleotide species in the population. For instance, the polymerase will replace dGMP with a universal nucleotide, if it is contacted with a population of free dAMP, dTMP, dCMP and the universal nucleotide.
[0178] The universal nucleotide preferably comprises one of the following nucleobases: hypoxanthine, 4-nitroindole, 5-nitroindole, 6-nitroindole, formylindole, 3-nitropyrrole, nitroimidazole, 4-nitropyrazole, 4-nitrobenzimidazole, 5-nitroindazole, 4-aminobenzimidazole or phenyl (C6-aromatic ring). The universal nucleotide more preferably comprises one of the following nucleosides: 2'-deoxyinosine, inosine, 7-deaza-2'-deoxyinosine, 7-deaza-inosine, 2-aza-deoxyinosine, 2-aza-inosine, 2-O'-methylinosine, 4-nitroindole 2'-deoxyribonucleoside, 4-nitroindole ribonucleoside, 5-nitroindole 2'-deoxyribonucleoside, 5-nitroindole ribonucleoside, 6-nitroindole 2'-deoxyribonucleoside, 6-nitroindole ribonucleoside, 3-nitropyrrole 2'-deoxyribonucleoside, 3-nitropyrrole ribonucleoside, an acyclic sugar analogue of hypoxanthine, nitroimidazole 2'-deoxyribonucleoside, nitroimidazole ribonucleoside, 4-nitropyrazole 2'-deoxyribonucleoside, 4-nitropyrazole ribonucleoside, 4-nitrobenzimidazole 2'-deoxyribonucleoside, 4-nitrobenzimidazole ribonucleoside, 5-nitroindazole 2'-deoxyribonucleoside, 5-nitroindazole ribonucleoside, 4-aminobenzimidazole 2'-deoxyribonucleoside, 4-aminobenzimidazole ribonucleoside, phenyl C-ribonucleoside, phenyl C-2'-deoxyribosyl nucleoside, 2'-deoxynebularine, 2'-deoxyisoguanosine, K-2'-deoxyribose, P-2'-deoxyribose and pyrrolidine. The universal nucleotide more preferably comprises 2'-deoxyinosine. The universal nucleotide is more preferably IMP or dIMP. The universal nucleotide is most preferably dPMP (2'-Deoxy-P-nucleoside monophosphate) or dKMP (N6-methoxy-2, 6-diaminopurine monophosphate).
[0179] Where the binding group is a protein, it may be able to anchor directly into the membrane without further functionalisation, for example if it already has an external hydrophobic region which is compatible with the membrane. Examples of such proteins include transmembrane proteins. Alternatively the protein may be expressed with a geneticallyfused hydrophobic region which is compatible with the membrane. Such hydrophobic protein regions are known in the art.
[0180] Examples of methods of coupling adaptors to membranes are disclosed in WO 2012 / 164270 and WO 2015 / 150786 (incorporated herein by reference in their entireties).
[0181] Any of the adaptors may be coupled to the membrane using a tethering complex comprising one or more hydrophilic components connected by a hydrophobic linker as described in WO 2021 / 111139 (incorporated herein by reference in its entirety). The tethering complex represents a hydrophobic anchor as described herein.
[0182] First coupling adaptor(s)
[0183] The method comprises coupling the polynucleotide or the plurality of polynucleotides to the membrane using one or more first coupling adaptors each comprising a polynucleotide-binding protein of the invention. The polynucleotide-binding protein of the invention may be any of those described above, including all the fragments and variants.
[0184] Any number of one or more first coupling adaptors may be used. The method may comprise using 1, 2, 3, 4, 5 or more first coupling adaptors. The method preferably comprises using 1 or 2 first coupling adaptors. Any of these may be referred to as "types of" first coupling adaptors to distinguish them from multiple instances of the one or more first coupling adaptors in the populations described below. The one or more first coupling adaptors or one or more types of first coupling adaptors may differ based on one or more of their (a) length, (b), structure, (c) number of polynucleotide-binding proteins of the invention and (d) identities of the polynucleotide-binding proteins of the invention. The one or more first coupling adaptors or one or more types of first coupling adaptors may differ in terms of (a); (b); (c); (d); (a) and (b); (a) and (c); (a) and (d); (b) and (c); (b) and (d); (c) and (d); (a), (b) and (c); (a), (b) and (d); (a), (c) and (d); (b), (c) and (d); or (a), (b), (c) and (d). The embodiments in this paragraph apply to any of the first coupling adaptors described below.
[0185] The method preferably comprises using one or more first coupling adaptors comprising one or more polynucleotide-binding proteins of the invention. The one or more polynucleotide-binding proteins of the invention may be any of those described above, including all the fragments and variants. The method may comprise using 1, 2, 3, 4, 5 or more polynucleotide-binding proteins of the invention. The one or more polynucleotide-binding proteins of the invention may be the same or different. The method preferably comprises using 1 polynucleotide-binding protein of the invention. The one or more first coupling adaptors or one or more types of first coupling adaptors may comprise the same polynucleotide-binding protein of the invention.The one or more first coupling adaptors may be one or more polypeptide adaptors. The one or more polypeptide adaptors may be modified in any of the way known in the art. One or more of the amino acids / derivatives / analogs in the one or more first coupling adaptors may be modified. Any one or more post-translational modifications may be present in the one or more first coupling adaptors. The one or more first coupling adaptors may be labelled with a molecular label. The one or more first coupling adaptors may contain one or more crosslinked sections, e.g. , C-C bridges. The one or more first coupling adaptors may comprise sulphide-containing amino acids and thus have the potential to form disulphide bonds.
[0186] The one or more first coupling adaptors can be any suitable length. The one or more first coupling adaptors preferably comprise one or more polypeptide adaptors having a length of from about 2 to about 500 amino acids. The one or more polypeptide adaptors preferably have a length of from about 5 to about 450 amino acids, from about 10 to about 400 amino acids, from about 20 to about 300 amino acids, from about 50 to about 200 amino acids or from about 60 to about 100 amino acids. The one or more polypeptide adaptors may have a length of at least about 10 amino acids, at least about 20 amino acids, at least about 30 amino acids, at least about 40 amino acids, at least about 50 amino acids, at least about 60 amino acids, at least about 70 amino acids, at least about 80 amino acids, at least about 90 amino acids, at least about 100 amino acids, at least about 150 amino acids, at least about 200 amino acids, at least about 300 amino acids, at least about 400 amino acids or at least about 500 amino acids.
[0187] The one or more first coupling adaptors may comprise a polynucleotide and a polypeptide. The one or more first coupling adaptors may be one or more polynucleotide-polypeptide conjugates. The one or more conjugates preferably comprise a polynucleotide conjugated to a polypeptide. The polypeptide section typically comprises or consists of the polynucleotide binding-protein of the invention.
[0188] The polypeptide can be conjugated to the polynucleotide at any suitable position. For example, the polypeptide can be conjugated to the polynucleotide at the N-terminus or the C-terminus of the polypeptide. The polypeptide can be conjugated to the polynucleotide via a side chain group of a residue (e.g., an amino acid residue) in the polypeptide. The polypeptide may have a naturally occurring reactive functional group which can be used to facilitate conjugation to the polynucleotide. For example, a cysteine residue can be used to form a disulphide bond to the polynucleotide or to a modified group thereon.
[0189] The polypeptide may be modified in order to facilitate its conjugation to the polynucleotide. For example, the polypeptide may be modified by attaching a moiety comprising a reactive functional group for attaching to the polynucleotide. For example, the polypeptide can be extended at the N-terminus or the C-terminus by one or more residues (e.g., amino acid residues) comprising one or more reactive functional groups for reacting with acorresponding reactive functional group on the polynucleotide. For example, the polypeptide can be extended at the N-terminus and / or the C-terminus by one or more cysteine residues. Such residues can be used for attachment to the polynucleotide portion of the conjugate, e.g., by maleimide chemistry (e.g., by reaction of cysteine with an azido-maleimide compound such as azido-[Pol]-maleimide wherein [Pol] is typically a short chain polymer such as PEG, e.g., PEG2, PEG3, or PEG4; followed by coupling to appropriately functionalised polynucleotide e.g., polynucleotide carrying a BCN group for reaction with the azide). Such chemistry is described in Example 2. For avoidance of doubt, when the polypeptide comprises an appropriate naturally occurring residue at the N- and / or C-terminus (e.g., a naturally occurring cysteine residue at the N- and / or C-terminus) then such residue(s) can be used for attachment to the polynucleotide.
[0190] A residue in the polypeptide may be modified to facilitate attachment of the polypeptide to the polynucleotide. A residue (e.g., an amino acid residue) in the polypeptide may be chemically modified for attachment to the polynucleotide. A residue (e.g., an amino acid residue) in the polypeptide may be enzymatically modified for attachment to the polynucleotide.
[0191] The conjugation chemistry between the polynucleotide and the polypeptide in the conjugate is not particularly limited. Any suitable combination of reactive functional groups can be used. Many suitable reactive groups and their chemical targets are known in the art. Some exemplary reactive groups and their corresponding targets include aryl azides which may react with amine, carbodiimides which may react with amines and carboxyl groups, hydrazides which may react with carbohydrates, hydroxmethyl phosphines which may react with amines, imidoesters which may react with amines, isocyanates which may react with hydroxyl groups, carbonyls which may react with hydrazines, maleimides which may react with sulfhydryl groups, NHS-esters which may react with amines, PFP-esters which may react with amines, psoralens which may react with thymine, pyridyl disulfides which may react with sulfhydryl groups, vinyl sulfones which may react with sulfhydryl amines and hydroxyl groups, vinylsulfonamides, and the like. Other suitable chemistry for conjugating the polypeptide to the polynucleotide includes click chemistry. Many suitable click chemistry reagents are known in the art. Suitable examples of click chemistry include, but are not limited to, the following:
[0192] copper(I)-catalyzed azide-alkyne cycloadditions (azide alkyne Huisgen cycloadditions);
[0193] strain-promoted azide-alkyne cycloadditions; including alkene and azide [3 + 2] cycloadditions; alkene and tetrazine inverse-demand Diels-Alder reactions; and alkene and tetrazole photoclick reactions;copper-free variant of the 1,3 dipolar cycloaddition reaction, where an azide reacts with an alkyne under strain, for example in a cyclooctane ring such as in bicycle[6.1.0]nonyne (BCN);
[0194] the reaction of an oxygen nucleophile on one linker with an epoxide or aziridine reactive moiety on the other; and
[0195] the Staudinger ligation, where the alkyne moiety can be replaced by an aryl phosphine, resulting in a specific reaction with the azide to give an amide bond.
[0196] Any reactive group may be used to form the conjugate. Some suitable reactive groups include [1, 4-Bis[3-(2-pyridyldithio)propionamido]butane; 1,1 1-bis-maleimidotriethyleneglycol; 3,3'-dithiodipropionic acid di(N-hydroxysuccinimide ester); ethylene glycol-bis(succinic acid N-hydroxysuccinimide ester); 4,4'-diisothiocyanatostilbene-2,2'-disulfonic acid disodium salt; Bis[2-(4-azidosalicylamido)ethyl] disulphide; 3-(2-pyridyldithio)propionic acid N-hydroxysuccinimide ester; 4-maleimidobutyric acid N-hydroxysuccinimide ester; lodoacetic acid N-hydroxysuccinimide ester; S-acetylthioglycolic acid N-hydroxysuccinimide ester; azide-PEG-maleimide; and alkyne-PEG-maleimide. The reactive group may be any of those disclosed in WO 2010 / 086602, particularly in Table 3 of that application.
[0197] The reactive functional group may be comprised in the polynucleotide and the target functional group may be comprised in the polypeptide prior to the conjugation step. The reactive functional group may be comprised in the polypeptide and the target functional group may be comprised in the polynucleotide prior to the conjugation step. The reactive functional group may be attached directly to the polypeptide. The reactive functional group may be attached to the polypeptide via a spacer. Any suitable spacer can be used. Suitable spacers include for example alkyl diamines such as ethyl diamine, etc.
[0198] Any polynucleotide, such as a nucleic acid, in the one or more first coupling adaptors is a macromolecule comprising two or more nucleotides. The polynucleotide or nucleic acid may comprise any combination of any nucleotides. The nucleotides can be naturally occurring or artificial. One or more nucleotides in the polynucleotide can be oxidized or methylated. One or more nucleotides in the polynucleotide may be modified, for instance with a label or a tag, for which suitable examples are known by a skilled person. The polynucleotide may comprise one or more spacers. A nucleotide typically contains a nucleobase, a sugar and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines and more specifically adenine (A), guanine (G), thymine (T), uracil (U) and cytosine (C). The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably a deoxyribose. The polynucleotidepreferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dll) and / or thymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC). The nucleotide is typically a ribonucleotide or deoxyribonucleotide. The nucleotide typically contains a monophosphate, diphosphate, or triphosphate. The nucleotide may comprise more than three phosphates, such as 4 or 5 phosphates. Phosphates may be attached on the 5' or 3' side of a nucleotide. The nucleotides in the polynucleotide may be attached to each other in any manner. The nucleotides are typically attached by their sugar and phosphate groups as in nucleic acids. The nucleotides may be connected via their nucleobases as in pyrimidine dimers. The polynucleotide may be single stranded or double stranded. At least a portion of the polynucleotide is preferably double stranded. The polynucleotide is most preferably ribonucleic nucleic acid (RIMA) or deoxyribonucleic acid (DNA).
[0199] The polynucleotide in the one or more first coupling adaptors can be any length. For example, the polynucleotide can have a length of from about 5 to about 500 nucleotides or nucleotide pairs, from about 10 to about 200 nucleotides or nucleotide pairs, from about 30 to about 150 nucleotides or nucleotide pairs, from about 50 to about 100 nucleotides or nucleotide pairs or from about 60 to about 90 nucleotides or nucleotide pairs. The polynucleotide can be at least about 10, at least about 20, at least about 25, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 400 or at least about 500 nucleotides or nucleotide pairs in length.
[0200] Suitable nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP) and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP and dUMP. A nucleotide may be abasic (i.e., lack a nucleobase). A nucleotide may also lack a nucleobase and a sugar (i.e., is a C3 spacer).
[0201] The one or more first coupling adaptors typically couple the polynucleotide or the plurality of polynucleotides directly to the membrane. Preferably, the one or more first coupling adaptors each further comprise a membrane protein or a hydrophobic anchor. The membrane protein or hydrophobic anchor may be any of those described above.
[0202] Leader sequenceThe one or more first coupling adaptors may further comprise a leader. Any suitable leader may be used. The leader may be a charged polymer, e.g. , a negatively charged polymer. The leader may comprise a polymer such as PEG or a polysaccharide. The leader may be a polypeptide. The leader may be a polynucleotide. The leader may be a polypeptidepolynucleotide conjugate. The leader may be any of the polypeptides, polynucleotides, or polypeptide-polynucleotide conjugates described above.
[0203] The leader may be from about 10 to about 150 monomer units (e.g., ethylene glycol or saccharide units, amino acids, nucleotides or a combination thereof) in length. The leader may be from about 12 to about 120, from about 15 to about 100, from about 18 to about 80 or about 20 to about 60 monomer units in length.
[0204] Second coupling adaptor(s)
[0205] The method preferably further comprises using one or more second coupling adaptors each comprising a membrane protein or a hydrophobic anchor to couple the one or more first coupling adaptors to the membrane. The membrane protein or hydrophobic anchor may be any of those described above.
[0206] Any number of one or more second coupling adaptors may be used. The method may comprise using 1, 2, 3, 4, 5 or more second coupling adaptors. The method preferably comprises using 1 or 2 second coupling adaptors. Any of these may be referred to as "types of" second coupling adaptors to distinguish them from multiple instances of the one or more second coupling adaptors in the populations described below. The one or more second coupling adaptors or one or more types of second coupling adaptors may differ based on one or more of their (a) length, (b), structure, (c) number of coupling domains and (d) identity of coupling domains. Coupling domains are described in more detail below. The one or more second coupling adaptors or one or more types of second coupling adaptors may differ in terms of (a); (b); (c); (d); (a) and (b); (a) and (c); (a) and (d); (b) and (c); (b) and (d); (c) and (d); (a), (b) and (c); (a), (b) and (d); (a), (c) and (d); (b), (c) and (d); or (a), (b), (c) and (d). The embodiments in this paragraph apply to any of the second coupling adaptors described below.
[0207] The one or more second coupling adaptors may be any of the types of adaptors described above for the one or more first coupling adaptors. The one or more second coupling adaptors may be one or more polypeptide adaptors, one or more polynucleotide-polypeptide conjugate adaptors or one or more polynucleotide adaptors. The polypeptide may be any of the types and / or lengths described above with reference to the one or more first coupling adaptors. The polynucleotide-polypeptide conjugate may be any of the types and / or lengths described above with reference to the one or more first coupling adaptors. The one or more polynucleotide adaptors may be any of the types and / or lengths described above withreference to the one or more first coupling adaptors. The polynucleotide in a second coupling adaptor is preferably at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 11 nucleotides, at least about 12 nucleotides, at least about 13 nucleotides, at least about 14 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 27 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, or at least about 50 nucleotides in length.
[0208] The one or more second coupling adaptors typically bind to the one or more first coupling adaptors. The one or more second coupling adaptors typically specifically bind to the one or more first coupling adaptors. Typically, the polynucleotide-binding protein(s) in the one or more first coupling adaptors bind to the polynucleotide or the plurality of polynucleotides, the one or more second coupling adaptors bind or specifically bind the one or more first coupling adaptors and thereby couple the polynucleotide or the plurality of polynucleotides to the membrane.
[0209] The term "specifically binds to" means binding that is measurably different from a nonspecific or non-selective interaction (e.g., with a non-target molecule). Specific binding can be measured, for example, by measuring binding of one or more second coupling adaptors to the one or more first coupling adaptors (or vice versa and comparing it to binding to a non-target molecule. Specific binding can also be determined by competition with a control molecule that mimics the epitope recognized on the target molecule. Such methods are routine in the art. All instances herein of the term "specifically binds to" is interchangeable with "specifically interacts with," "specific for," "selectively binds to" "selectively interacts with" and "selective for".
[0210] The one or more second coupling adaptors specifically bind to the one or more first coupling adaptors with preferential or high affinity, but do not bind or bind with only low affinity to other or different molecules, such as other or different polypeptides, other or different polypeptides or proteins or other or different polynucleotides. Preferably, the one or more second coupling adaptors bind to the one or more first coupling adaptors with an affinity that is at least about 10 times, such as at least about 50 times, at least about 100 times, at least about 200 times, at least about 300 times, at least about 400 times, at least about 500 times, at least about 1000 times or at least about 10,000 times, greater than its / their affinity for other molecules.
[0211] The term "Kassoc" or "Kon", as used herein, is intended to refer to the association rate of a particular binding molecule-target, whereas the term "Kdis" or "Koff," as used herein, is intended to refer to the dissociation rate of a particular binding molecule-target interaction. The term "Kd", as used herein, is intended to refer to the dissociation constant, which isobtained from the ratio of Kon to Koff (i.e. Kon / Koff) and is expressed as a molar concentration (M).
[0212] The one or more second coupling adaptors have affinity for the one or more first coupling adaptors. The one or more second coupling adaptors preferably haves high affinity for the one or more first coupling adaptors. The one or more second coupling adaptors have high affinity if it / they bind(s) to the one or more first coupling adaptors x IO-8M or less, about 1 x IO-8M or less, or about 5 x IO-9M or less. The one or more second coupling adaptors preferably have a Kd for the one or more first coupling adaptors of about 100 nM or less, such as about 50 nM or less, about 20 nM or less, about 10 nM or less, about 1 nM or less, about 100 pM or less, about 50 pM or less, about 20 pM or less or about 10 pM or less. A molecule or group binds with low affinity if it binds with a Kd of about 1 x IO-6M or more, about 1 x IO-5M or more, about 1 x 10-4M or more, about 1 x IO-3M or more, or about 1 x IO2M or more.
[0213] Affinity can be measured using known binding assays, such as those that make use of fluorescence and radioisotopes. Competitive binding assays are also known in the art. The strength of binding between peptides or proteins and polynucleotides can be measured using nanopore force spectroscopy as described in Hornblower et al., Nature Methods. 4: 315-317. (2007) or Isothermal Titration Calorimetry (ITC), which is a label-free quantification technique used in studies of a wide variety of biomolecular interactions. ITC works by directly measuring the heat that is either released or absorbed during a biomolecular binding event. Kd values for antibodies and other binding domains can be determined using surface plasmon resonance, such as a Biacore® system, or solution equilibrium titration (SET) (see Friguet, et al., (1985) J. Immunol. Methods, 77(2):305-319, and Hanel et al., (2005) Anal. Biochem., 339(1) : 182-184).
[0214] The one or more second coupling adaptors may specifically bind to at least a part of the one or more first coupling adaptors and vice versa. The one or more second coupling adaptors may specifically bind to at least about 5% of the one or more first coupling adaptors, such as at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90% or about 100% of the one or more first coupling adaptors and vice versa. %s are usually calculated on the basis on the number of nucleotides or amino acids to which the one or more second coupling adaptors bind. The one or more second coupling adaptors may specifically bind to at least about 3 nucleotides or amino acids in the one or more first coupling adaptors and vice versa. The one or more second coupling adaptors may specifically bind to at least about 4 nucleotides or amino acids, at least about 5 nucleotides or amino acids, at least about 10 nucleotides or amino acids, at least about 15 nucleotides or amino acids, at least about 20 nucleotides or amino acids, at least about 30 nucleotides or amino acids, at least about 40 nucleotides or amino acids, at least about 50 nucleotidesor amino acids, at least about 60 nucleotides or amino acids, at least about 70 nucleotides or amino acids, at least about 80 nucleotides or amino acids, at least about 90 nucleotides or amino acids, at least about 100 nucleotides or amino acids, at least about 110 nucleotides or amino acids, at least about 120 nucleotides or amino acids, at least about 130 nucleotides or amino acids, at least about 140 nucleotides or amino acids or at least about 150 nucleotides or amino acids in the one or more first coupling adaptors and vice versa.
[0215] The one or more second coupling adaptors may specifically bind to the one or more first coupling adaptors and vice versa using coupling domains. The one or more second coupling adaptors may comprise one or more coupling domains which bind or specifically bind to one or more coupling domains in the one or more first coupling adaptors and vice versa. The one or more coupling domains may be peptides, polypeptides, proteins, nucleotides, oligonucleotide, polynucleotides, polynucleotide-polypeptide conjugates, monosaccharides, oligosaccharides, polysaccharides, one or more polyethyleneglycols (PEGs), saturated and unsaturated hydrocarbons or polyamides. The skilled person is capable of designing coupling domains which bind or specifically bind to each other.
[0216] When one or more second coupling adaptors are used, the one or more first coupling adaptors may further comprise a first oligonucleotide and the one or more second coupling adaptors may further comprises a second oligonucleotide which hybridizes to the first oligonucleotide. The first and / or second oligonucleotide may be any length. The first and / or second oligonucleotide is preferably at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 11 nucleotides, at least about 12 nucleotides, at least about 13 nucleotides, at least about 14 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 27 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, or at least about 50 nucleotides in length.
[0217] The first oligonucleotide preferably specifically hybridises to the second oligonucleotide. Oligonucleotides "specifically hybridise" when they hybridise with preferential or high affinity each other but do not substantially hybridise, does not hybridise, or hybridises with only low affinity to other polynucleotide sequences, especially other polynucleotide sequences used in the invention. Conditions that permit the hybridisation are well-known in the art (for example, Sambrook et al., 2001, Molecular Cloning: a laboratory manual, 3rd edition, Cold Spring Harbour Laboratory Press; and Current Protocols in Molecular Biology, Chapter 2, Ausubel et al., Eds., Greene Publishing and Wiley-lnterscience, New York (1995)).
[0218] Hybridisation can be carried out under low stringency conditions, for example in the presence of a buffered solution of 30 to 35% formamide, 1 M NaCI and 1 % SDS (sodium dodecyl sulfate) at 37 °C followed by a 20 wash in from IX (0.1650 M Na + ) to 2X (0.33 MNa + ) SSC (standard sodium citrate) at 50 °C. Hybridisation can be carried out under moderate stringency conditions, for example in the presence of a buffer solution of 40 to 45% formamide, 1 M NaCI, and 1 % SDS at 37 °C, followed by a wash in from 0.5X (0.0825 M Na + ) to IX (0.1650 M Na + ) SSC at 55 °C. Hybridisation can be carried out under high stringency conditions, for example in the presence of a buffered solution of 50% formamide, 1 M NaCI, 1% SDS at 37 °C, followed by a wash in 0.1X (0.0165 M Na + ) SSC at 60 °C. The oligonucleotides "specifically hybridise" if they hybridise with a melting temperature (Tm) that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C or at least 10 °C, greater than its Tm for other polynucleotide sequences. More preferably, the oligonucleotides hybridise with a Tm that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C, greater than its Tm for other polynucleotide sequences. Preferably, the oligonucleotides hybridise with a Tm that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C, greater than its Tm for a polynucleotide which differs from the oligonucleotides by one or more nucleotides, such as by 1, 2, 3, 4 or 5 or more nucleotides. The strands or parts thereof typically hybridise with a Tm of at least 90 °C, such as at least 92 °C or at least 95 °C. Tm can be measured experimentally using known techniques, including the use of DNA microarrays, or can be calculated using publicly available Tm calculators, such as those available over the internet.
[0219] Linkers
[0220] Any of the coupling adaptors may comprise one or more linkers such as 2, 3, 4 or more linkers. The one or more first coupling adaptors and / or the one or more second coupling adaptors may comprise one or more linkers.
[0221] Preferred linkers include, but are not limited to, polymers, such as polynucleotides, polyethylene glycols (PEGs), polysaccharides and polypeptides. These linkers may be linear, branched, or circular. For instance, the linker may be a circular polynucleotide. The adaptor(s) may hybridise to a complementary sequence on a circular polynucleotide linker. The one or more linkers may comprise a component that can be cut or broken down, such as a restriction site or a photolabile group. The one or more linkers may be functionalised with maleimide groups to attach to cysteine residues in proteins. Suitable linkers are described in WO 2010 / 086602 (incorporated herein by reference in its entirety).
[0222] Adaptor synthesis
[0223] The adaptors of the invention are typically synthetic or semi-synthetic. For example, DNA or RNA may be purely synthetic, synthesised by conventional DNA synthesis methods such asphosphoramidite based chemistries. Synthetic polynucleotides subunits may be joined together by known means, such as ligation or chemical linkage, to produce longer strands. Internal self-forming structures (e.g., hairpins, quadruplexes) can be designed into the substrate, e.g., by ligating appropriate sequences. Synthetic polynucleotides can be copied and scaled up for production by means known in the art, including PCR, incorporation into bacterial factories, and the like. Synthetic polypeptides may be produced using any of the methods described herein. The production of polynucleotide-polypeptide conjugates is described above.
[0224] Spacers
[0225] Any of the coupling adaptors may comprise one or more spacers. The one or more first coupling adaptors and / or the one or more second coupling adaptors may comprise one or more spacers.
[0226] Any of the adaptors may comprise from about one to about 10 spacers, e.g., from about 1 to about 5 spacers, e.g., about 1, 2, 3, 4 or 5 spacers.
[0227] One or more spacers are typically included in the adaptors to provide a distinctive signal when they pass through or across a nanopore. One or more spacers may be used to define or separate one or more regions of an adaptor, e.g., to separate an adaptor from the target polypeptide.
[0228] A spacer may comprise a linear molecule, such as a polymer, e.g., a polypeptide or a polyethylene glycol (PEG). Typically, such a spacer has a different structure from the target polynucleotide. For instance, if the target polynucleotide is DNA, the or each spacer typically does not comprise DNA. In particular, if the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the or each spacer preferably comprises peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) or a synthetic polymer with nucleotide side chains. A spacer may comprise one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dTs), one or more inverted dideoxy-thymidines (ddTs), one or more dideoxy-cytidines (ddCs), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-Methyl RNA bases, one or more Isodeoxycytidines (Iso-dCs), one or more Iso-deoxyguanosines (Iso-dGs), one or more C3 (OC3H6OPO3) groups, one or more photo-cleavable (PC) [OCsHe- OjNHCl-h-CeHsNO?-CH(CH3)OPC>3] groups, one or more hexandiol groups, one or more spacer 9 (iSp9) [(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSplS) [(OCH2CH2)6OPO3] groups; or one or more thiol connections. A spacer may comprise any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNATechnologies®). For example, C3, iSp9 and iSpl8 spacers are all available from IDT®. A spacer may comprise any number of the above groups as spacer units.
[0229] A spacer may comprise one or more chemical groups, e.g., one or more pendant chemical groups. The one or more chemical groups may be attached to one or more nucleobases in an adaptor. The one or more chemical groups may be attached to the backbone of an adaptor. Any number of appropriate chemical groups may be present, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenols (DNPs), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctyne groups.
[0230] A spacer may comprise one or more abasic nucleotides ( / .e., nucleotides lacking a nucleobase), such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more abasic nucleotides. The nucleobase can be replaced by -H (idSp) or -OH in the abasic nucleotide. Abasic spacers can be inserted into target polynucleotides by removing the nucleobases from one or more adjacent nucleotides. For instance, polynucleotides may be modified to include 3-methyladenine, 7-methylguanine, l,N6-ethenoadenine inosine or hypoxanthine and the nucleobases may be removed from these nucleotides using Human Alkyladenine DNA Glycosylase (hAAG). Alternatively, polynucleotides may be modified to include uracil and the nucleobases removed with Uracil-DNA Glycosylase (UDG). The one or more spacers preferably do not comprise any abasic nucleotides.
[0231] Suitable spacers can be designed or selected depending on the nature of the adaptor, the polynucleotide binding protein, and the conditions under which the method is to be carried out.
[0232] Tags
[0233] Any of the adaptors used in the invention may comprise a tag. The tag may comprise or be an oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino). The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) can have about 10-30 nucleotides in length or about 10-20 nucleotides in length. The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) for use in the tag can have at least one end (e.g., 3'- or 5'-end) modified for conjugation to other modifications or to a solid substrate surface including, e.g., a bead. The end modifiers may add a reactive functional group which can be used for conjugation. Examples of functional groups that can be added include, but are not limited to amino, carboxyl, thiol, maleimide, aminooxy, and any combinations thereof. The functional groups can be combined with different length of spacers (e.g. , C3, C9, C12, Spacer 9 and 18) to add physical distance of the functional group from the end of the oligonucleotide sequence.The tag may comprise or be a morpholino oligonucleotide. The morpholino oligonucleotide can have about 10-30 nucleotides in length or about 10-20 nucleotides in length. The morpholino oligonucleotides can be modified or unmodified. For example, the morpholino oligonucleotide can be modified on the 3' and / or 5' ends of the oligonucleotides. Examples of modifications on the 3' and / or 5' end of the morpholino oligonucleotides include, but are not limited to 3' affinity tag and functional groups for chemical linkage (including, e.g., 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyl dithio, and any combinations thereof); 5' end modifications (including, e.g., 5'-primary ammine, and / or 5'-dabcyl), modifications for click chemistry (including, e.g., 3'-azide, 3'-alkyne, 5'-azide, 5'-alkyne), and any combinations thereof.
[0234] The tag may further comprise a polymeric linker, e.g. , to facilitate coupling to a detector e.g., a nanopore. An exemplary polymeric linker includes, but is not limited to, polyethylene glycol (PEG). The polymeric linker may have a molecular weight of about 500 Da to about 10 kDa (inclusive), or about 1 kDa to about 5 kDa (inclusive). The polymeric linker (e.g., PEG) can be functionalized with different functional groups including, e.g., but not limited to maleimide, NHS ester, dibenzocyclooctyne (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combinations thereof. The tag may further comprise a 1 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag may further comprise a 2 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag may further comprise a 3 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag may further comprise a 5 kDa PEG with a 5'-maleimide group and a 3'-DBCO group.
[0235] The tag may be a peptide, polypeptide or protein. The tag may be a CsgA polypeptide. The CsgA polypeptide may be any of those described in WO 2024 / 100270 (incorporated herein in its entirety).
[0236] Other examples of a tag include, but are not limited to His tags, biotin or streptavidin, antibodies that bind to analytes, aptamers that bind to analytes, analyte binding domains such as DNA binding domains (including, e.g., peptide zippers such as leucine zippers, single-stranded DNA binding proteins (SSB)), and any combinations thereof.
[0237] The tag may be attached to the external surface of a nanopore, e.g., on the cis side of a membrane, using any methods known in the art. For example, one or more tags can be attached to the nanopore via one or more cysteines (cysteine linkage), one or more primary amines such as lysines, one or more non-natural amino acids, one or more histidines (His tags), one or more biotin or streptavidin, one or more antibody-based tags, one or more enzyme modification of an epitope (including, e.g., acetyl transferase), and any combinations thereof. Suitable methods for carrying out such modifications are well-known in the art. Suitable non-natural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz) and any one of the amino acids numbered 1-71 in Figure 1 of Liu C. C. and Schultz P. G., Annu. Rev. Biochem., 2010, 79, 413-444.
[0238] Where one or more tags are attached to a nanopore via cysteine linkage(s), the one or more cysteines can be introduced to one or more monomers that form the nanopore by substitution. The nanopore may be chemically modified by attachment of (i) Maleimides including diabromomaleimides such as: 4-phenylazomaleinanil, l.N-(2-Hydroxyethyl)maleimide, N-Cyclohexylmaleimide, 1.3-Maleimidopropionic Acid, 1.1-4-Ami nophenyl- lH-pyrrole, 2,5, dione, 1.1-4-Hydroxy phenyl- lH-pyrrole, 2,5, dione, N-Ethylmaleimide, N-Methoxycarbonylmaleimide, N-tert-Butylmaleimide, N-(2-Aminoethyl)maleimide , 3-Maleimido-PROXYL , N-(4-Chlorophenyl)maleimide, l-[4-(dimethylamino)-3,5-dinitrophenyl]-lH-pyrrole-2,5-dione, N-[4-(2-Benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(l-naphthyl)-maleimide, N-(2,4-xylyl)maleimide, N-(2,4-difluorophenyl)maleimide , N-(3-chloro-para-tolyl)-maleimide, l-(2-amino-ethyl)-pyrrole-2, 5-dione hydrochloride, l-cyclopentyl-3-methyl-2,5-dihydro-lH-pyrrole-2,5-dione, l-(3-a mi nopropyl)-2,5-di hydro- lH-pyrrole-2,5-dione hydrochloride, 3- methyl- l-[2-oxo-2-(piperazin-l-yl)ethyl]-2,5-di hydro- lH-pyrrole-2, 5-dione hydrochloride, l-benzyl-2,5-dihydro-lH-pyrrole-2, 5-dione, 3-methyl-l-(3,3,3-trifluropropyl)-2,5-dihydro-lH-pyrrole-2,5-dione, l-[4-(methylamino)cyclohexyl]-2,5-dihydro-lH-pyrrole-2, 5-dione triflu roacetic acid, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2, l-benzyl-3-methyl-2,5-dihydro-lH-pyrrole-2,5-dione, l-(2-fluorophenyl)-3-methyl-2,5-dihydro 1H-pyrrole-2, 5-dione, N-(4-phenoxyphenyl)maleimide , N-(4-nitrophenyl)maleimide (ii) lodocetamides such as :3-(2-Iodoacetamido)-proxyl, N-(cyclopropylmethyl)-2-iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2-trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, N-(4-(aminosulfonyl)phenyl)-2-iodoacetamide, N-(l,3-benzothiazol-2-yl)-2-iodoacetamide, N-(2,6-diethylphenyl)-2-iodoacetamide, N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide, (iii) Bromoacetamides: such as N-(4-(acetylamino)phenyl)-2-bromoacetamide , N-(2-acetylphenyl)-2-bromoacetamide , 2-bromo-n-(2-cyanophenyl)acetamide, 2-bromo-N-(3-(trifluoromethyl)phenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide , 2-bromo-N-(4-fluorophenyl)-3-methylbutanamide, N-Benzyl-2-bromo-N-phenylpropionamide, N-(2-bromo-butyryl)-4-chloro-benzenesulfonamide, 2-Bromo-N-methyl-N-phenylacetamide, 2-bromo-N-phenethyl-acetamide,2-adamantan-l-yl-2-bromo-N-cyclohexyl-acetamide, 2-bromo-N-(2-methylphenyl)butanamide, Monobromoacetanilide, (iv) Disulphides such as: aldrithiol-2 , aldrith iol-4 , isopropyl disulfide, l-(Isobutyldisulfanyl)-2-methylpropane, Dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-Pyridyldithio)propionic acid, 3-(2-Pyridyldithio)propionic acid hydrazide, 3-(2-Pyridyldithio)propionic acid N-succinimidyl ester, am6amPDPl-pCD and (v) Thiols such as: 4-Phenylthiazole-2-thiol, Purpald, 5, 6, 7, 8-tetrahydro-quinazoline-2-thiol.The tag may be attached directly to a nanopore or via one or more linkers. The tag may be attached to the nanopore using the hybridization linkers described in WO 2010 / 086602 (incorporated herein by reference in its entirety). Alternatively, peptide linkers may be used. Peptide linkers are amino acid sequences. The length, flexibility and hydrophilicity of the peptide linker are typically designed such that it does not to disturb the functions of the monomer and pore. Preferred flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10 or 16, serine and / or glycine amino acids. More preferred flexible linkers include (SG)i, (SG)2, (SG)3, (SG)4, (SG)5and (SG)s wherein S is serine and G is glycine. Preferred rigid linkers are stretches of 2 to 30, such as 4, 6, 8, 16 or 24, proline amino acids. More preferred rigid linkers include (P)i2 wherein P is proline.
[0239] Suitable pore tags are also described in WO 2018 / 100370, which describes non-hairpin methods for characterising double-stranded polynucleotides and is herein incorporated by reference in its entirety.
[0240] Biotin enrichment
[0241] Any of the adaptors may comprise biotin. The biotin may be used to isolate the adaptors. Suitable methods for biotin-based enrichment are known in the art. For instance, a surface, such as a bead, comprising avidin / or streptavidin may be used to barcoded constructs comprising biotin.
[0242] Method steps
[0243] The coupling method may comprise (a) allowing the one or more first coupling adaptors to bind to the polynucleotide or the plurality of polynucleotides. The one or more first coupling adaptors may be any of those described above. The one or more first coupling adaptors typically bind to the polynucleotide or the plurality of polynucleotides via the polynucleotide-binding protein of the invention.
[0244] The method preferably does not comprise modifying or functionalising the polynucleotide to facilitate binding of the one or more first coupling adaptors, for instance using a sequencing adaptor. The method preferably comprises (a) allowing the one or more first coupling adaptors to bind to the polynucleotide or the plurality of polynucleotides without modifying or functionalising the polynucleotide to facilitate binding of the one or more first coupling adaptors, for instance using a sequencing adaptor.
[0245] The method may also comprise (b) coupling the one or more first coupling adaptors to the membrane. The one or more first coupling adaptors may be coupled to the membrane directly. The one or more first coupling adaptors may comprise a membrane protein or a hydrophobic anchor. The method may also comprise (b) allowing the one or more first coupling adaptors to couple to the membrane.The method may also comprise (b) coupling the one or more first coupling adaptors to the membrane using the one or more second coupling adaptors. The method may also comprise (b) allowing the one or more first coupling adaptors to couple to the membrane using the one or more second coupling adaptors. The one or more second coupling adaptors may be any of those described above. The one or more second coupling adaptors are functionalised to couple to the membrane because they comprise a membrane protein or a hydrophobic anchor. Step (b) typically comprises contacting the polynucleotide-first coupling adaptor conjugate or the plurality of polynucleotide-first coupling adaptor conjugates from step (a) with the one or more second coupling adaptors under conditions which allow the one or more second coupling adaptors to bind, specifically bind or hybridise to the one or more first coupling adaptors and to couple the polynucleotide or plurality of polynucleotides to the membrane.
[0246] Steps (a) and (b) are carried out simultaneously or sequentially. These steps may be carried out in either order, including (a) then (b) or (b) then (a). Steps (a) and (b) may be carried out at the same time.
[0247] The method may comprise using a population of one or more first coupling adaptors and / or a population of one or more second coupling adaptors. The one or more first coupling adaptors in the population are typically all the same. The one or more second coupling adaptors in the population are typically the same. The one or more first coupling adaptors in the population may be any of the adaptors defined above. The one or more second coupling adaptors in the population may be any of the adaptors defined above.
[0248] The population(s) may contain any number of one or more first coupling adaptors and / or one or more second coupling adaptors. For instance, the method may use a (1) a population of at least about 2, at least about 5, at least about 10, at least about 50, at least about 100, at least about 500, at least about 1,000, at least about 5,000, at least about 10,000, at least about 50,000, at least about 1 x 105, at least about 1 x 106, at least about 1 x 107, at least about 1 x 108, at least about 1 x 109, at least about 1 x 1010, at least about 1 x 1011, at least about 1 x 1012or at least about 1 x 1013one or more first coupling adaptors and / or (2) a population of at least about 2, at least about 5, at least about 10, at least about 50, at least about 100, at least about 500, at least about 1,000, at least about 5,000, at least about 10,000, at least about 50,000, at least about 1 x 105, at least about 1 x 106, at least about 1 x 107, at least about 1 x 108, at least about 1 x 109, at least about 1 x 1010, at least about 1 x 1011, at least about 1 x 1012or at least about 1 x 1013one or more second coupling adaptors. The number of adaptors in the populations are typically the same as or greater than the number of polypeptides in the plurality. The number of polypeptides in the plurality may not be known.The coupling method may comprise (a) allowing a population of one or more first coupling adaptors to bind to the plurality of polynucleotides. The method may also comprise (b) coupling the population of one or more first coupling adaptors to the membrane. The method may also comprise (b) allowing the population of one or more first coupling adaptors to couple to the membrane.
[0249] The coupling method may comprise (a) allowing a population of one or more first coupling adaptors to bind to the plurality of polynucleotides. The method may also comprise (b) coupling the population of one or more first coupling adaptors to the membrane using a population of one or more second coupling adaptors. The method may also comprise (b) allowing the population of one or more first coupling adaptors to couple to the membrane using a population of one or more second coupling adaptors. The one or more second coupling adaptors may be any of the second coupling adaptors described above and they may be coupled to the membrane in any of the ways described above.
[0250] Adaptors
[0251] The invention also provides one or more coupling adaptors for coupling a polynucleotide or a plurality of polynucleotides to a membrane, each comprising a polynucleotide-binding protein according of the invention and a membrane protein or a hydrophobic anchor. The invention also provides one or more coupling adaptors for coupling a polynucleotide or a plurality of polynucleotides to a membrane, each comprising one or more polynucleotide-binding proteins of the invention and a membrane protein or a hydrophobic anchor. The one or more coupling adaptors may be any of the first coupling adaptors described above with reference to the coupling method of the invention. The polynucleotide-binding protein or the one or more polynucleotide-binding proteins may be any of those described above. The membrane protein or hydrophobic anchor may be any of those described above.
[0252] The invention also provides a population of coupling adaptors for coupling a plurality of polynucleotides to a membrane, each comprising a polynucleotide-binding protein according of the invention and a membrane protein or a hydrophobic anchor. The invention also provides a population of coupling adaptors for coupling a plurality of polynucleotides to a membrane, each comprising one or more polynucleotide-binding proteins of the invention and a membrane protein or a hydrophobic anchor. The population may comprise any of the numbers of first coupling adaptors described above with reference to the coupling method of the invention. The coupling adaptors may be any of the first coupling adaptors described above with reference to the coupling method of the invention. The polynucleotide-binding protein or the one or more polynucleotide-binding proteins may be any of those described above. The membrane protein or hydrophobic anchor may be any of those described above.
[0253] KitsThe invention provides a kit for coupling a polynucleotide or a plurality of polynucleotides to a membrane. The kit comprises (1) one or more first coupling adaptors each comprising a polynucleotide-binding protein of the invention and a first oligonucleotide. The kit also comprises (2) one or more second coupling adaptors each comprising a membrane protein or a hydrophobic anchor and a second oligonucleotide which hybridizes to the first oligonucleotide. The one or more first coupling adaptors may be any of those described above with reference to the coupling method of the invention. The one or more second coupling adaptors may be any of those described above with reference to the coupling method of the invention. The membrane protein or hydrophobic anchor may be any of those described above. The first and / or second oligonucleotides may be any of those described above.
[0254] The kit may comprise any numbers of (the types of) one or more first coupling adaptors and / or one or more second coupling adaptors described above with reference to the method of the invention.
[0255] The kit may comprise (1) a population of first coupling adaptors each comprising a polynucleotide-binding protein of the invention and a first oligonucleotide and (2) a population of second coupling adaptors each comprising a membrane protein or a hydrophobic anchor and a second oligonucleotide which hybridizes to the first oligonucleotide. The population(s) may contain any number of one or more first coupling adaptors and / or one or more second coupling adaptors. For instance, the kit may comprise (1) a population of at least about 2, at least about 5, at least about 10, at least about 20, at least about 50, at least about 100, at least about 500, at least about 1000, at least about 5000, at least about 10000, at least about 50000, at least about 100000 one or more first coupling adaptors and / or (2) a population of at least about 2, at least about 5, at least about 10, at least about 20, at least about 50, at least about 100, at least about 500, at least about 1000, at least about 5000, at least about 10000, at least about 50000, at least about 100000 one or more second coupling adaptors.
[0256] The number of adaptors in the populations are typically the same as or greater than the number of polynucleotides in the plurality. The number of polynucleotides in the plurality may not be known.
[0257] The invention provides a kit for coupling a polynucleotide or a plurality of polynucleotides to a membrane. The kit comprises (1) one or more polynucleotides encoding one or more first coupling adaptors each comprising a polynucleotide-binding protein of the invention. The kit also comprises (2) one or more polynucleotide encoding one or more second coupling adaptors each comprising a membrane protein or a hydrophobic anchor and a second oligonucleotide which hybridizes to the first oligonucleotide. The one or more first coupling adaptors may be any of those described above with reference to the coupling method of theinvention. The one or more second coupling adaptors may be any of those described above with reference to the coupling method of the invention. The membrane protein or hydrophobic anchor may be any of those described above. The first and / or second oligonucleotides may be any of those described above.
[0258] The one or more polynucleotides in the kit may be any of those described above. The kit may comprise any number of one or more polynucleotides, such as 1, 2, 3, 4, 5 or more polynucleotides. The kit may encode any of the numbers of (the types of) one or more first coupling adaptors and / or one or more second coupling adaptors described above with reference to the method of the invention.
[0259] The invention also provides one or more expression vectors comprising a kit of the invention. The polynucleotides in the kit defined in (1) and (2) above may be in the same or different expression vectors. The invention also provides a host cell comprising a kit of the invention or one or more expression vectors of the invention. Suitable vectors and host cells are known in the art.
[0260] The invention also provides a kit for characterising a polynucleotide analyte or a plurality of polynucleotide analytes comprising (a) a polynucleotide-binding protein of the invention, one or more coupling adaptors of the invention or a kit of the invention and (b) the components of a membrane. The components in (a) may be any of those described above. The membrane in (b) may be any of those described above.
[0261] The invention also provides a kit for characterising a polynucleotide analyte or a plurality of polynucleotide analytes comprising (a) a membrane comprising a detector and (b) a polynucleotide-binding protein of the invention, one or more coupling adaptors of the invention or a kit of the invention. The components in (a) may be any of those described above and below. The components in (b) may be any of those described above.
[0262] The kit may additionally comprise one or more other reagents or instruments which enable any of the embodiments mentioned above or below to be carried out. Such reagents or instruments include one or more of the following: suitable buffer(s) (aqueous solutions), means to obtain a sample from a subject (such as a vessel or an instrument comprising a needle), means to amplify and / or express polynucleotides, a membrane as defined below or voltage or patch clamp apparatus. Reagents may be present in the kit in a dry state such that a fluid sample is used to resuspend the reagents. The kit may also, optionally, comprise instructions to enable the kit to be used in the methods described herein or details regarding for which organism the method may be used. The kit may comprise a magnet or an electromagnet. The kit may, optionally, comprise nucleotides.
[0263] Characterisation methodsThe invention provides a method for determining the presence, absence or one or more characteristics of a polynucleotide analyte. This method may also be known as the detection or characterisation method of the invention.
[0264] The polynucleotide analyte may be any of the polynucleotides defined above with reference to the coupling method of the invention.
[0265] The invention may be for determining the presence, absence or one or more characteristics of a plurality of polynucleotide analytes. There may be any number of polynucleotide analytes in the plurality. For instance, there may be at least about 2, at least about 5, at least about 10, at least about 20, at least about 50, at least about 100, at least about 500, at least about 1000, at least about 5000, at least about 10000, at least about 50000, at least about 100000 or more polynucleotides analytes in the plurality.
[0266] The polynucleotide analyte or the plurality of polynucleotide analytes may be called the target polynucleotide analyte or the plurality of target polynucleotide analytes.
[0267] The method comprises (a) coupling the polynucleotide analyte or the plurality of polynucleotide analytes to a membrane using a polynucleotide-binding protein of the invention, a kit of the invention or a coupling method of the invention. The polynucleotide-binding protein of the invention may be any of those described above, including all the fragments and variants. The kit may be any of the kits of the invention described above. The method may be any of the coupling methods of the invention described above.
[0268] The method also comprises (b) allowing the coupled polynucleotide analyte or the plurality of coupled polynucleotide analytes to interact with a detector present in the membrane and thereby determining the presence, absence or one or more characteristics of the polynucleotide analyte or the plurality of polynucleotide analytes.
[0269] The invention also provides a method for detecting analyte polynucleotides, comprising (a) providing a membrane in which is present a nanopore that provides a channel through the membrane; (b) contacting the membrane, in an ionic solution, with analyte polynucleotides, wherein following contact with the membrane the analytes polynucleotides are coupled or tethered to the membrane using a polynucleotide-binding protein of the invention, a kit of the invention or a coupling method of the invention; (c) applying a potential difference across the membrane and detecting a first analyte polynucleotide using the nanopore, from among the analytes polynucleotides coupled or tethered to the membrane; and (d) detecting a second analyte polynucleotide, from among the analytes polynucleotides coupled or tethered to the membrane, using the same nanopore, wherein the second polynucleotide is not the first polynucleotide. The polynucleotide-binding protein of the invention may be any of those described above, including all the fragments and variants.The kit may be any of the kits of the invention described above. The method may be any of the coupling methods of the invention described above.
[0270] Any method of characterisation may be used. The method preferably uses next generation sequencing (NGS).
[0271] The coupled polynucleotide analyte or the plurality of coupled polynucleotide analytes are allowed to interact with a detector present in the membrane. The coupled polynucleotide analyte or the plurality of coupled polynucleotide analytes are preferably moved with respect to the detector in the membrane. Any suitable measurements can be taken using the detector as the polynucleotide analyte or the plurality of polynucleotide analytes move with respect to the detector. The detector preferably detects the polynucleotide analyte or the plurality of polynucleotide analytes via electrical or optical means. The detector may be selected from (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube; and (v) a nanopore. Preferably, the detector comprises a nanopore.
[0272] The coupled polynucleotide analyte or the plurality of coupled polynucleotide analytes may be characterised in the detection or characterisation method of the invention in any suitable manner. The coupled polynucleotide analyte or the plurality of coupled polynucleotide analytes are preferably characterised by detecting an ionic current or optical signal as they move with respect to a nanopore. This is described in more detail herein. The method is amenable to these and other methods of characterising analytes.
[0273] Nanopore characterisation
[0274] The coupled polynucleotide or the coupled polynucleotides are preferably characterised using a nanopore. In step (b), the method preferably comprises (i) allowing the coupled polynucleotide analyte or the plurality of coupled polynucleotide analytes to interact with, such as move through, the nanopore and (ii) measuring the current passing through the pore during the interaction and thereby determining the presence, absence or one or more characteristics of the polynucleotide analyte or the plurality of polynucleotide analytes. The coupled polynucleotide analyte or the plurality of coupled polynucleotide analytes typically become(s) uncoupled from the membrane as they interact with, such as move through, the nanopore.
[0275] The one or more characteristics are preferably selected from (i) the length(s) of the polynucleotide analyte or the plurality of polynucleotide analytes, (ii) the identity / identities of the polynucleotide analyte or the plurality of polynucleotide analytes, (iii) the sequence(s) of the polynucleotide analyte or the plurality of polynucleotide analytes, (iv) the secondary structure(s) of the polynucleotide analyte or the plurality of polynucleotide analytes and (v) whether or not the polynucleotide analyte or the plurality of polynucleotide analytes is / aremodified. The polynucleotide analyte or the plurality of polynucleotide analytes may be modified in any of the ways described above. The one or more characteristics of the polynucleotide analyte or the plurality of polynucleotide analytes are preferably measured by electrical measurement and / or optical measurement. The electrical measurement is preferably a current measurement, an impedance measurement, a tunnelling measurement, or a field effect transistor (FET) measurement.
[0276] The method more preferably comprises (i) contacting the coupled polynucleotide analyte or the plurality of coupled polynucleotide analytes with a nanopore such that the polynucleotide analyte or the plurality of polynucleotide analytes move(s) through the pore and (ii) measuring the current moving through the pore as the polynucleotide analyte or the plurality of polynucleotide analytes move(s) through the pore wherein the current is indicative of one or more characteristics of the polynucleotide analyte or the plurality of polynucleotide analytes and thereby characterising the polynucleotide analyte or the plurality of polynucleotide analytes. The one or more characteristics may be any of those defined above.
[0277] Any suitable nanopore can be used. The nanopore is preferably a transmembrane pore. A transmembrane pore is a structure that crosses the membrane to some degree. It permits hydrated ions driven by an applied potential to flow across or within the membrane. The transmembrane pore typically crosses the entire membrane so that hydrated ions may flow from one side of the membrane to the other side of the membrane. However, the transmembrane pore does not have to cross the membrane. It may be closed at one end. For instance, the pore may be a well, gap, channel, trench or slit in the membrane along which or into which hydrated ions may flow.
[0278] The nanopore or transmembrane pore typically has a first opening and a second opening. The first opening is typically the cis opening and the second opening is typically the trans opening. However, the first opening may be the trans opening and the second opening may be the cis opening. Any movement control protein used in the detection or characterisation method of the invention is typically provided at the first opening of the nanopore and thus controls the movement of the target polynucleotide in the direction from the second opening of the nanopore towards the first opening of the nanopore.
[0279] Any transmembrane pore may be used in the detection or characterisation method of the invention. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores and solid-state pores. The pore may be a DNA origami pore (Langecker et al., Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983.The nanopore may be a transmembrane protein pore which is a monomer or an oligomer. The pore is preferably made up of several repeating subunits, such as at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, or at least about 16 subunits. The pore is preferably a hexameric, heptameric, octameric or nonameric pore. The pore may be a homo-oligomer or a hetero-oligomer.
[0280] The transmembrane protein pore may comprise a barrel or channel through which the ions may flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane -barrel or channel or a transmembrane a-helix bundle or channel.
[0281] Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polynucleotide (as described herein). These amino acids are preferably located near a constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate the interaction between the pore and nucleotides, polynucleotides, or nucleic acids.
[0282] The transmembrane protein pore may be from or derived from Wza, Iota toxin, Anthrax protective antigen, Vibrio cholerae cytolysin, Cytotoxin K (CytK), CELIII, CsgG, CsgF, CsgG-CsgF, Aerolysin, alpha hemolysin, MspA, MspB, MspC, PorARr, PorBRr, PorARc, PilQ, necrotic enteritis B-like toxin (NetB), FraC, portal proteins including G20c, P23_45, T4, SPP1, P22 and Phi29, gamma hemolysin, Monalysin, Lysenin, ClyA, an actinoporin, Clostridium perfringens beta toxin, parasporin-2, epsilon toxin, lectin from the parasitic mushroom Laetiporus sulphureus (LSL), volvatoxin, Cry toxins, CytlAa, Cyt2Aa, Complement component 9 (C9), Perfringolysin O, Pleurotolysin, Listeriolysin, Perforin-2, Gasdermin-A3, L-, P- and M-ring protein, Type II secretion system protein D, GspD, InvG, VirB7, SpoIIIAG, Cag8, Cag3, Cag or other proteins in the Type IV secretion system apparatus protein CagY, WzzB, Pentraxin, Afp2, Major vault protein, Thioredoxin-dependent peroxidase reductase, Arf-GAP, Respiratory syncytial virus ribonucleoprotein, Chikungunya virus nonstructural protein 1, PRC, YaxA, XaxA, HfaB, NfpAB, leukocidin or PrgH.
[0283] Suitable transmembrane protein pores for use in the invention include those described in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO 2018 / 211241, WO 2019 / 002893, WO 2023 / 118404, WO 2023 / 198911, WO 2024 / 033421, WO 2024 / 033422, WO 2024 / 033443, and WO 2024 / 089270 (all incorporated by reference herein in their entirety).
[0284] The transmembrane protein pore may also be any of the CsgG pores described in WO 2023 / 060420, WO 2023 / 60418, WO 2023 / 60422, WO 2023 / 060421, WO 2023 / 019470,CN114957412, WO 2023 / 019471, W02023 / 060419 and WO 2023 / 050031 (all incorporated herein by reference in their entireties) or a variant thereof.
[0285] The transmembrane protein pore may be any of the pores described in WO 2023 / 123370, WO 2024 / 138470, WO 2024 / 138472, WO 2024 / 138424, WO 2024 / 138425, WO 2024 / 138512 and WO 2024 / 138565 (all incorporated herein by reference in their entireties) or a variant thereof.
[0286] The transmembrane pore may be formed from a chimeric pore monomer comprising two or more regions, wherein at least two of the two or more regions are from at least two different pores. The chimeric pore monomer may comprise any number of regions, such as three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more or ten or more regions, from different pores. The chimeric pore monomer may comprise two or three regions. The regions are preferably selected from a cap region, a constriction region, and a transmembrane region. The regions may be a cap region and a constriction region. The regions may be a cap region, a constriction region, and a transmembrane region. The at least two different pores are typically at least two different pores that appear in nature. The at least two different pores are typically at least two different wild-type or naturally occurring pores. The at least two different pores are preferably different before any artificial or synthetic modifications, such as additions, deletions and / or substitutions, are made to them. The at least two different pores are preferably homologues, for example structural homologues. A structural homologue refers to a protein or molecule that shares a similar three-dimensional structure with another protein or molecule. This can be determined using standard methods in the art (e.g., AlphaFold or PSIPRED). Structural homologues typically have similar sequences. Structural homologues are normally identified in similar species. The at least two different pores may be selected from any of the pores listed above. The at least two different pores may be two different PorARc pores or three different PorARc pores. The at least two different pores may be two different CsgG pores or three different CsgG pores. The chimeric pore monomer may be any of those described in PCT / EP2023 / 080135 (incorporated by reference herein in its entirety).
[0287] The transmembrane pore may be formed from a pore monomer comprising (a) a CsgG monomer and (b) a fusion polypeptide comprising a first portion comprising a CsgF peptide and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the pore monomer. The pore monomer may be derived from a protein transmembrane pore complex comprising (a) a CsgG transmembrane pore comprising a lumen and (b) a fusion polypeptide comprising a first portion comprising a CsgF protein and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the transmembrane pore. The auxiliary protein can be designed de novo using computer-based structural analysis tools to confer certain desirable features to the CsgGmonomer (e.g., modulation of pore width, lengthening of pore lumen, formation of one or more additional constrictions, etc.). The de novo designed auxiliary protein may form one or more additional constrictions in the lumen of a CsgG pore formed from the monomer, and improve discrimination of polymer units as an analyte moves through the pore. The pore monomer may be any of the pore monomers described in WO 2024 / 033447 (incorporated by reference herein in its entirety).
[0288] Movement control protein
[0289] The movement of the polynucleotide analyte or the plurality of polynucleotide analytes with respect to, such as through, the detector, nanopore or transmembrane is preferably controlled using a movement control protein.
[0290] As those skilled in the art will appreciate, any suitable movement control protein can be used in the invention. The movement control protein may be any protein that is capable of binding to a polynucleotide and controlling its movement with respect to a detector, e.g., a nanopore.
[0291] In more detail, movement control proteins such as helicases can typically control the movement of polynucleotides in at least two active modes of operation (when is provided with all the necessary components to facilitate movement e.g., ATP and Mg2+) and one inactive mode of operation (when not provided with the necessary components to facilitate movement; or when the movement control protein is modified in order to prevent the active mode).
[0292] When provided with all the necessary components to facilitate movement, a movement control protein may move along a polynucleotide in either a 5'-3' direction or a 3'-5' direction. Many movement control proteins process polynucleotides in a 5'-3' direction. Movement control proteins which control the movement of polynucleotides in this manner are typically suitable for use in the detection or characterisation method of the invention.
[0293] However, when a movement control protein is not provided with the necessary components to facilitate movement or is modified in order to prevent it from actively controlling the movement of the polynucleotide with respect to the nanopore, it can still passively control the movement of the polynucleotide with respect to the nanopore. For example, the movement control protein can bind to the polynucleotide and act as a brake slowing the movement of the polynucleotide when it is pulled into the pore by an applied field (e.g., by the first force in the detection or characterisation method of the invention). In the "inactive" mode it typically does not matter whether the polynucleotide is captured either 3' or 5' down ( / .e., moves through the nanopore in a 5'-3' direction or in a 3'-5' direction), as the applied force provides the impetus to move the polynucleotide through the nanopore.
[0294] However, in such embodiments, the movement control protein may still control themovement of the polynucleotide with respect to the nanopore e.g., by acting as a brake. When in the inactive mode the movement control of a polynucleotide by a movement control protein can be described in a number of ways including ratcheting, sliding, and braking. Typically the detection or characterisation method of the invention do not comprise the use of a movement control protein operating in the passive mode. However, when a movement control protein the movement control protein is used, it may be a movement control protein operating in the passive mode.
[0295] Some methods of the invention may comprise use of a movement control protein as a pausing moiety to impede the movement of the polynucleotide strand through the nanopore. The movement control protein may be a protein which binds to polynucleotides but which does not have polynucleotide processing capacity, / .e., it is not a movement control protein.
[0296] A polynucleotide-handling enzyme is a polypeptide that is capable of interacting with a polynucleotide. The enzyme may modify the polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. The enzyme may modify the polynucleotide by orienting it or moving it to a specific position. A movement control protein as used herein may be, or may be derived from a polynucleotide handling enzyme. A movement control protein may be, or may be derived from a polynucleotide-handling enzyme.
[0297] The movement control protein may be derived from a member of any of the Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30 and 3.1.31.
[0298] Typically, the movement control protein is a helicase, a polymerase, an exonuclease, a topoisomerase, or a variant thereof.
[0299] The movement control protein may be modified to prevent the movement control protein disengaging from the polynucleotide. Thus, the target polynucleotide preferably does not disengage from the movement control protein.
[0300] As used herein, the term "disengaging" refers to the dissociation of the movement control protein from the target polynucleotide. Thus, a movement control protein may be modified to prevent it from dissociating from the target polynucleotide, e.g., into the reaction medium. It is important to distinguish potential "disengagement" of a movement control protein from "unbinding" of a movement control protein from a target polynucleotide. As used herein, "unbinding" refers to the transient release of the target polynucleotide the active site of the movement control protein (described in more detail herein) but does not imply disengagement. Thus, for example, a movement control protein may be modified to prevent the movement control protein from disengaging from a polynucleotide, but withoutpreventing the movement control protein from unbinding from the polynucleotide. When unbound, the movement control protein remains engaged with the target polynucleotide. For example, the movement control protein may remain engaged with the target polynucleotide ( / '.e., it may be prevented from disengaging from the target polynucleotide) because it is topologically closed around the target polynucleotide. The movement control site may remain free to bind or unbind the target polynucleotide such that the movement control protein may bind or unbind to the target polynucleotide, whilst the movement control protein remains engaged with the target polynucleotide. When the movement control protein is unbound from the target polynucleotide it may be able to move on (e.g., along) the target polynucleotide under an applied force and may be capable of re-binding to the target polynucleotide. When engaged on the target polynucleotide but unbound from the target polynucleotide, the movement control protein is not capable of dissociating from the target polynucleotide.
[0301] The movement control protein can be adapted to prevent disengagement in any suitable way. For example, the movement control protein can be loaded on the polynucleotide and then modified in order to prevent it from disengaging from the polynucleotide. Alternatively, the movement control protein can be modified to prevent it from disengaging from the polynucleotide before it is loaded onto the polynucleotide. Modification of a movement control protein and / or a movement control protein in order to prevent it from disengaging from a polynucleotide can be achieved using methods known in the art, such as those described in WO 2014 / 013260, which is hereby incorporated by reference in its entirety, and with particular reference to passages describing the modification of movement control proteins such as helicases in order to prevent them from disengaging with polynucleotide strands. For example, a movement control protein can be modified by treating with tetramethylazodicarboxamide (TMAD). Various other closing moieties are described in WO 2021 / 255476 (incorporated herein by reference in its entirety).
[0302] For example, a movement control protein and / or a movement control protein may have a polynucleotide-unbinding opening, e.g., a cavity, cleft or void through which a polynucleotide strand may pass when the movement control protein disengages from the strand. The polynucleotide-unbinding opening may be the opening through which a polynucleotide may pass when the movement control protein disengages from the polynucleotide. The polynucleotide-unbinding opening for a given movement control protein can be determined by reference to its structure, e.g., by reference to its X-ray crystal structure. The X-ray crystal structure may be obtained in the presence and / or the absence of a polynucleotide substrate. The location of a polynucleotide-unbinding opening in a given movement control protein may be deduced or confirmed by molecular modelling using standard packages known in the art. The polynucleotide-unbinding opening may betransiently produced by movement of one or more parts e.g., one or more domains of the movement control protein.
[0303] The movement control protein may be modified by closing the polynucleotide-unbinding opening. The polynucleotide-unbinding opening may be closed with a closing moiety.
[0304] Closing the polynucleotide-unbinding opening may therefore prevent the movement control protein from disengaging from the polynucleotide. For example, the movement control protein may be modified by covalently closing the polynucleotide-unbinding opening.
[0305] However, as explained above closing the polynucleotide-unbinding opening does not necessarily prevent the target polynucleotide from unbinding from the movement control site of the movement control protein. A preferred protein for addressing in this way is a helicase.
[0306] The movement control protein may be modified with a closing moiety for (i) topologically closing the movement control site of the movement control protein around the target polynucleotide and (ii) promoting unbinding of the target polynucleotide from the movement control site of the movement control protein and / or retarding re-binding of the target polynucleotide to the movement control site of the movement control protein. The movement control protein may be modified in any suitable manner to facilitate attachment of such a closing moiety.
[0307] A closing moiety may comprise a bifunctional cross-linking moiety. The closing moiety may comprise a bifunctional cross-linker. The bifunctional crosslinker may attach at two points on the movement control protein and close the polynucleotide-unbinding opening of the movement control protein thereby preventing disengagement of the polynucleotide from the movement control protein whilst allowing unbinding of the polynucleotide from the polynucleotide-binding site of the movement control protein.
[0308] The closing moiety may attach at any suitable positions on the movement control protein. For example, the closing moiety may crosslink two amino acid residues of the movement control protein. Typically, at least one amino acid crosslinked by the closing moiety is a cysteine or a non-natural amino acid. The cysteine or non-natural amino acid may be introduced into the movement control protein by substitution or modification of a naturally occurring amino acid residue of the movement control protein. Methods for introducing non-natural amino acids are well known in the art and include for example native chemical ligation with synthetic polypeptide strands comprising such non-natural amino acids.
[0309] Methods for introducing cysteines into a movement control protein are likewise within the capability of one of skill in the art, for example using techniques disclosed in references such as Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).The closing moiety may have a length of from about 1 A to about 100 A. The length of the closing moiety may be calculated according to static bond lengths or more preferably using molecular dynamics simulations. The length may for example be from about 2 A to about 80 A, such as from about 5 A to about 50 A, e.g., from about 8 to about 30 A such as from about 10 to about 25 A or about 20 A, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 A.
[0310] Movement control proteins suitable for being closed using a closing moiety as described above are described in more detail herein. The movement control protein is preferably a helicase.
[0311] The movement control protein may be or may be derived from an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I from E. coli, exonuclease III enzyme from E. coli, Reel from T. thermophilus and bacteriophage lambda exonuclease, TatD exonuclease and variants thereof.
[0312] The movement control protein may be a polymerase. The polymerase may be PyroPhage® 3173 DNA Polymerase (which is commercially available from Lucigen® Corporation), SD Polymerase (commercially available from Bioron®), Klenow from NEB or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase or a variant thereof. Modified versions of Phi29 polymerase that may be used in the invention are disclosed in US Patent No. 5,576,204.
[0313] The movement control protein may be a topoisomerase. In one embodiment, the topoisomerase is a member of any of the Moiety Classification (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase may be a reverse transcriptase, which are enzymes capable of catalysing the formation of cDNA from a RNA template. They are commercially available from, for instance, New England Biolabs® and Invitrogen®.
[0314] The movement control protein is preferably a helicase. Any suitable helicase can be used in accordance with the detection or characterisation method of the invention. The helicase is preferably a member of superfamily 1 or superfamily 2. The helicase is more preferably a member of one of the following families: Pifl-like, Upfl-like, UvrD / Rep, Ski-like, Rad3 / XPD, NS3 / NPH-II, DEAD, DEAH / RHA, RecG-like, REcQ-like, TIR-like, Swi / Snf-like and Rig-I-like. The first three of those families are in superfamily 1 and the second ten families are in superfamily 2. The helicase is more preferably a member of one of the following subfamilies: RecD, Upfl , PcrA, Rep, UvrD, Hel308, Mtr4, XPD, NS3, Mssll6, Prp43, RecG, RecQ, T1R, RapA and Hef. The first five of those subfamilies are in superfamily 1 and the second eleven subfamilies are in superfamily 2. Members of the Upfl, Mtr4, NS3, Mssll6, Prp43 and Hef subfamilies are RNA helicases. Members of the remaining subfamilies are DNA helicases.The movement control protein is preferably a NS3 helicase or a modified NS3 helicase. The NS3 helicase or modified NS3 helicase may be any of those described in PCT / EP2024 / 063118 (incorporated herein by reference in its entirety).
[0315] For example, the or each enzyme used in accordance with the present disclosure may be independently selected from a Hel308 helicase, a RecD helicase, a Tral helicase, a TrwC helicase, an XPD helicase, and a Dda helicase, or a variant thereof. Monomeric helicases may comprise several domains attached together. For instance, Tral helicases and Tral subgroup helicases may contain two RecD helicase domains, a relaxase domain and a C-terminal domain. The domains typically form a monomeric helicase that is capable of functioning without forming oligomers. Particular examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pifl and Tral. These helicases typically work on single stranded DNA. Examples of helicases that can move along both strands of a double stranded DNA include FtfK and hexameric enzyme complexes, or multisubunit complexes such as RecBCD. The movement control protein may be a Dda (DNA-dependent ATPase) helicase.
[0316] Hel308 helicases are described in publications such as WO 2013 / 057495, the entire contents of which are incorporated by reference. RecD helicases are described in publications such as WO 2013 / 098562, the entire contents of which are incorporated by reference. XPD helicases are described in publications such as WO 2013 / 098561, the entire contents of which are incorporated by reference. Dda helicases are described in publications such as WO 2015 / 055981 and WO 2016 / 055777, the entire contents of each of which are incorporated by reference.
[0317] The helicase may be Trwc Cba or a variant thereof, Hel308 Mbu or a variant thereof or Dda or a variant thereof. Variants may differ from the native sequences in any of the ways described herein.
[0318] General methods
[0319] As mentioned above, the detection or characterisation method of the invention may be operated using any suitable detector, and as such any suitable apparatus for detecting polynucleotides can be used.
[0320] The detection or characterisation method of the invention may be carried out using any apparatus that is suitable for nanopore sensing. For example, the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections. The barrier may have an aperture in which a membrane containing a transmembrane pore is formed. Transmembrane pores are described herein.The methods may be carried out using the apparatus described in WO 2008 / 102120, WO 2010 / 122293, or WO 00 / 28312 (incorporated herein by reference in their entireties). In brief, the binding of the polynucleotide analyte or the plurality of polynucleotide analytes in the channel of a pore will have an effect on the open-channel ion flow through the pore, which is the essence of "molecular sensing" of pore channels. Variation in the open-channel ion flow can be measured using suitable measurement techniques by the change in electrical current. The degree of reduction in ion flow, as measured by the reduction in electrical current, is related to the size of the obstruction within, or in the vicinity of, the pore. Binding of the polynucleotide analyte or the plurality of polynucleotide analytes in or near the pore therefore provides a detectable and measurable event, thereby forming the basis of a "biological sensor".
[0321] When used to characterize the polynucleotide analyte or the plurality of polynucleotide analytes, the presence, absence or one or more characteristics of the polynucleotide analyte or the plurality of polynucleotide analytes are determined. The methods may be for determining the presence, absence or one or more characteristics of the polynucleotide analyte or the plurality of polynucleotide analytes. The methods may concern determining the presence, absence or one or more characteristics of two or more the plurality of polynucleotide analytes. The methods may comprise determining the presence, absence or one or more characteristics of any number of polynucleotide analytes, such as 2, 5, 10, 15, 20, 30, 40, 50, 100 or more polynucleotide analytes. Any number of characteristics of the one or more polynucleotide analytes may be determined, such as 1, 2, 3, 4, 5, 10 or more characteristics. Characteristics amenable to being detected in the methods provide herein include the identity or sequence of the polynucleotide analyte or the plurality of polynucleotide analytes, the length of the polynucleotide analyte or the plurality of polynucleotide analytes, whether or not the polynucleotide analyte or the plurality of polynucleotide analytes is / are modified, etc. In some embodiments, the detection or characterisation method of the invention is a method of sequencing the polynucleotide analyte or the plurality of polynucleotide analytes. In some embodiments the sequences of the polynucleotide analyte or the plurality of polynucleotide analytes may be determined in real-time by aligning real-time signal or amino calling to known references. Exemplary methods of determining a polynucleotide sequence are described in WO 2016 / 059427 (incorporated herein by reference in its entirety).
[0322] The method may involve measuring the ion current flow through the pore, typically by measurement of a current. Alternatively, the ion flow through the pore may be measured optically, such as disclosed by Heron etal: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Therefore, the apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore. Thecharacterisation methods may be carried out using a patch clamp or a voltage clamp. The characterisation methods preferably involve the use of a voltage clamp.
[0323] The method may involve measuring an optical signal as described in Chen et al, Nature Communications (2018)9:1733, the entire contents of which are hereby incorporated by reference. For example, a nanopore such as an optically engineered nanopore structure (e.g., a plasmonic nanoslit) may be used to locally enable single-molecule surface enhanced Raman spectroscopy (SERS) to allow the characterisation of the polynucleotide through direct Raman spectroscopic detection.
[0324] The method may be carried out on a silicon-based array of wells where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.
[0325] The method may involve the measuring of a current flowing through the pore. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically from +2 V to -2 V, typically -400 mV to +400mV. The voltage used is preferably in a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20mV and 0 mV and an upper limit independently selected from + 10 mV, + 20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV. The voltage used is more preferably in the range 100 mV to 240mV and most preferably in the range of 120 mV to 220 mV. It is possible to increase discrimination between different nucleotides by a pore by using an increased applied potential.
[0326] The methods are typically carried out in the presence of any charge carriers, such as metal salts, for example alkali metal salts, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3-methyl imidazolium chloride. In the exemplary apparatus described above, the salt is present in the aqueous solution in the chamber. Potassium glutamate (KCsHsNC ), potassium chloride (KCI), sodium chloride (NaCI) or caesium chloride (CsCI) is typically used. KCI is preferred. The salt may be an alkaline earth metal salt such as calcium chloride (CaCI2). The salt concentration may be at saturation. The salt concentration may be 3M or lower and is typically from 0.1 to 2.5 M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M. The salt concentration is preferably from 150 mM to 1 M. The method is preferably carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M. High salt concentrations provide a high signal to noise ratio and allow for currents indicative of binding / no binding to be identified against the background of normal current fluctuations.The methods are typically carried out in the presence of a buffer. In the exemplary apparatus described above, the buffer is present in the aqueous solution in the chamber. Any suitable buffer may be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCI buffer. The methods are typically carried out at a pH of from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5. The pH used is preferably about 7.5.
[0327] The methods may be carried out at from 0 °C to 100 °C, from 15 °C to 95 °C, from 16 °C to 90 °C, from 17 °C to 85 °C, from 18 °C to 80 °C, 19 °C to 70 °C, or from 20 °C to 60 °C. The methods are typically carried out at room temperature. The methods are optionally carried out at a temperature that supports enzyme function, such as about 37 °C.
[0328] Any of the proteins described herein, such as the protein pores, may be made synthetically or by recombinant means. For example, the pore may be synthesised by in vitro translation and transcription (IVTT). The amino acid sequence of the pore may be modified to include non-naturally occurring amino acids or to increase the stability of the protein. When a protein is produced by synthetic means, such amino acids may be introduced during production. The pore may also be altered following either synthetic or recombinant production.
[0329] Any of the proteins described herein, such as the protein pores, can be produced using standard methods known in the art. Polynucleotide sequences encoding a pore or construct may be derived and replicated using standard methods in the art. Polynucleotide sequences encoding a pore or construct may be expressed in a bacterial host cell using standard techniques in the art. The pore may be produced in a cell by in situ expression of the polypeptide from a recombinant expression vector. The expression vector optionally carries an inducible promoter to control the expression of the polypeptide. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.
[0330] The protein may be produced in large scale following purification by any protein liquid chromatography system from protein producing organisms or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA systems, the Bio-Cad system, the Bio-Rad BioLogic system, and the Gilson HPLC system.
[0331] Methods for introducing or substituting non-naturally occurring amino acids in proteins are also well known in the art and described in WO 2019 / 002893 (incorporated by reference herein in its entirety). The proteins may be modified to assist their identification or purification, for example by the addition of a streptavidin tag or by the addition of a signal sequence to promote their secretion from a cell where the monomer does not naturally contain such a sequence. The proteins may also be produced using D-amino acids or amixture of L-amino acids and D-amino acids. This is conventional in the art for producing such proteins or peptides.
[0332] Any protein of the invention may be chemically modified. The protein can be chemically modified in any way and at any site. The protein may be chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art. The protein may be chemically modified by the attachment of any molecule, such as a dye or a fluorophore.
[0333] The protein may be chemically modified with a molecular adaptor that facilitates the interaction between a pore comprising the monomer and a target nucleotide or target polynucleotide sequence. Suitable adaptors, including a cyclic molecule, a cyclodextrin, a species that is capable of hybridization, a DNA binder or interchelator, a peptide or peptide analogue, a synthetic polymer, an aromatic planar molecule, a small positively charged molecule or a small molecule capable of hydrogen-bonding, are described in WO 2019 / 002893 (incorporated by reference herein in its entirety). The molecular adaptor may be attached using any of the methods and linkers described above.
[0334] Any of the proteins may be modified to assist their identification or purification, for example by the addition of histidine residues (a his tag), aspartic acid residues (an asp tag), a streptavidin tag, a flag tag, a SUMO tag, a GST tag or a MBP tag, or by the addition of a signal sequence to promote their secretion from a cell where the polypeptide does not naturally contain such a sequence. An alternative to introducing a genetic tag is to chemically react a tag onto a native or engineered position on the protein. An example of this would be to react a gel-shift reagent to a cysteine engineered on the outside of the protein. This has been demonstrated as a method for separating hemolysin heterooligomers (Chem Biol. 1997 Jul;4(7):497-505).
[0335] Any of the proteins may be labelled with a revealing label. The revealing label may be any suitable label which allows the protein to be detected. Suitable labels include, but are not limited to, fluorescent molecules, radioisotopes, e.g., 1251, 35S, enzymes, antibodies, antigens, polynucleotides, and ligands such as biotin.
[0336] The protein may also contain other non-specific modifications as long as they do not interfere with the function of the protein. A number of non-specific side chain modifications are known in the art and may be made to the side chains of the protein(s). Such modifications include, for example, reductive alkylation of amino acids by reaction with an aldehyde followed by reduction with NaBH4, amidation with methylacetimidate or acylation with acetic anhydride.Membrane
[0337] The invention also provides a membrane comprising a polynucleotide or a plurality of polynucleotides coupled to it using a polynucleotide-binding protein of the invention, a kit of the invention or a coupling method of the invention. The membrane may be any of those described above. The polynucleotide or the plurality of polynucleotides may be any of those described above. The polynucleotide-binding protein of the invention may be any of those described above, including all the fragments and variants. The kit or the coupling method may be any of those described above, especially with reference to the preferred kits or preferred coupling methods. The polynucleotide or the plurality of polynucleotides may be coupled to the membrane in any of the ways described above. The membrane preferably further comprises a detector, a nanopore or a transmembrane pore.
[0338] The invention also provides an array comprising a plurality of membranes of the invention. Any of the embodiments described above with respect to the membranes of the invention, the polynucleotide-binding proteins of the invention, kits of the invention and coupling methods of the invention equally apply the array of the invention. The array may be set up to perform any of the methods described herein.
[0339] In a preferred embodiment, each membrane in the array comprises one detector, nanopore or transmembrane pore. Due to the manner in which the array is formed, for example, the array may comprise one or more membranes that do not comprise a detector, nanopore or transmembrane pore, and / or one or more membranes that comprise two or more detectors, nanopores or transmembrane pores. The array may comprise from about 2 to about 1000, such as from about 10 to about 800, from about 20 to about 600 or from about 30 to about 500 membranes.
[0340] Systems
[0341] The invention provides a system comprising (a) a membrane of the invention or an array of the invention, (b) means for applying a potential across the membrane(s) and (c) means for detecting electrical or optical signals across the membrane(s). The membranes may be any as described herein.
[0342] In one embodiment, the system further comprises a first chamber and a second chamber, wherein the first and second chambers are separated by the membrane(s). When used to characterise a polynucleotide analyte or the plurality of polynucleotide analytes, the system may further comprise a polynucleotide analyte, wherein the polynucleotide analyte is transiently located within the continuous channel and wherein one end of the polynucleotide analyte is located in the first chamber and one end of the polynucleotide analyte is located in the second chamber.In one embodiment, the system further comprises an electrically conductive solution in contact with the oligomeric pore construct(s), electrodes providing a voltage potential across the membrane(s), and a measurement system for measuring the current through the oligomeric pore construct (s). The voltage applied across the membranes and construct is preferably from +5 V to -5 V, such as -600 mV to +600mV or -400 mV to +400 mV. The voltage used is preferably in the range 100 mV to 240 mV and more preferably in the range of 120 mV to 220 mV. It is possible to increase discrimination between different amino acids or nucleotides by the oligomeric pore construct by using an increased applied potential. Any suitable electrically conductive solution may be used. For example, the solution may comprise charge carriers, such as metal salts, for example alkali metal salt, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3-methyl imidazolium chloride. In an exemplary system, salt is present in the aqueous solution in the chamber. Potassium chloride (KCI), sodium chloride (NaCI), caesium chloride (CsCI) or a mixture of potassium ferrocyanide and potassium ferricyanide is typically used. KCI, NaCI and a mixture of potassium ferrocyanide and potassium ferricyanide are preferred. The charge carriers may be asymmetric across the membrane. For instance, the type and / or concentration of the charge carriers may be different on each side of the membrane, e.g., in each chamber.
[0343] The salt concentration may be at saturation. The salt concentration may be 3 M or lower and is typically from 0.1 to 2.5 M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M. The salt concentration is preferably from 150 mM to 1 M. The method is preferably carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M. High salt concentrations provide a high signal to noise ratio and allow for currents indicative of the presence of an amino acid or nucleotide to be identified against the background of normal current fluctuations.
[0344] A buffer may be present in the electrically conductive solution. Typically, the buffer is phosphate buffer. Other suitable buffers are HEPES and Tris-HCI buffer. The pH of the electrically conductive solution may be from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5. The pH used is preferably about 7.5.
[0345] The system may be comprised in an apparatus. The apparatus may be any conventional apparatus for analyte analysis, such as an array or a chip. The apparatus is preferably set up to carry out the disclosed method. For example, the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections. The barrier typically has an aperture in which the membrane(s) containing the oligomericpore construct (s) are formed. Alternatively, the barrier forms the membrane in which the oligomeric pore construct is present.
[0346] The apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore.
[0347] The apparatus may be any of those described in WO 2008 / 102120, WO 2009 / 077734, WO 2010 / 122293, WO 2011 / 067559, or WO 00 / 28312 (all incorporated herein by reference in their entirety).
[0348] Apparatus
[0349] The invention also provides an apparatus comprising a polynucleotide or a plurality of polynucleotides coupled to an in vitro membrane using a polynucleotide-binding protein of the invention, a kit of the invention or a coupling method of the invention.
[0350] The invention also provides an apparatus produced by a method comprising coupling a polynucleotide or a plurality of polynucleotides to an in vitro membrane using a polynucleotide-binding protein of the invention, a kit of the invention or a coupling method of the invention.
[0351] The invention also provides an apparatus for characterising a polynucleotide analyte or a plurality of polynucleotide analytes in a sample, comprising (a) a plurality of membranes each comprising a detector and (b) a plurality of polynucleotide-binding proteins of the invention, a plurality of one or more coupling adaptors of the invention or a plurality of kits of the invention.
[0352] The membrane preferably comprises a detector, a nanopore or a transmembrane pore.
[0353] Any of the specific embodiments described above are equally applicable to the apparatuses of the invention.
[0354] SEQUENCE LISTING
[0355] Table 3 - Sequences of the invention
[0356] SEQ Description Sequence
[0357] ID NO:
[0358] 1 Full length wildMALVYDAEFVGSEREFEEERETFLKGVKAYDGVLATRYLMERSSS type Topo V AKNDEELLELHQNFILLTGSYACSIDPTEDRYQNVIVRGVNFDERV QRLSTGGSPARYAIVYRRGWRAIAKALDIDEEDVPAIEVRAVKRN PLQPALYRILVRYGRVDLMPVTVDEVPPEMAGEFERLIERYDVPIDE KEERILEILRENPWTPHDEIARRLGLSVSEVEGEKDPESSGIYSLW SRVVVNIEYDERTAKRHVKRRDRLLEELYEHLEELSERYLRHPLTR RWIVEHKRDIMRRYLEQRIVECALKLQDRYGIREDVALCLARAFD GSISMIATTPYRTLKDVCPDLTLEEAKSVNRTLATLIDEHGLSPDA
[0359]
[0360] ADELIEHFESIAGILATDLEEIERMYEEGRLSEEAYRAAVEIQLAELTKKEGVGRKTAERLLRAFGNPERVKQLAREFEIEKLASVEGVGERVL RSLVPGYASLISIRGIDRERAERLLKKYGGYSKVREAGVEELREDG LTDAQIRELKGLKTLESIVGDLEKADELKRKYGSASAVRRLPVEEL RELGFSDDEIAEIKGIPKKLREAFDLETAAELYERYGSLKEIGRRLS YDDLLELGATPKAAAEIKGPEFKFLLNIEGVGPKLAERILEAVDYDL ERLASLNPEELAEKVEGLGEELAERVVYAARERVESRRKSGRQER SEEEWKEWLERKVGEGRARRLIEYFGSAGEVGKLVENAEVSKLLE VPGIGDEAVARLVPGYKTLRDAGLTPAEAERVLKRYGSVSKVQEG ATPDELRELGLGDAKIARILGLRSLVNKRLDVDTAYELKRRYGSVS AVRKAPVKELRELGLSDRKIARIKGIPETMLQVRGMSVEKAERLLE RFDTWTKVKEAPVSELVRVPGVGLSLVKEIKAQVDPAWKALLDVK GVSPELADRLVEELGSPYRVLTAKKSDLMRVERVGPKLAERIRAA GKRYVEERRSRRERIRRKLRG
[0361] Topoisomerase MALVYDAEFVGSEREFEEERETFLKGVKAYDGVLATRYLMERSSS domain and AKNDEELLELHQNFILLTGSYACSIDPTEDRYQNVIVRGVNFDERV repair site I QRLSTGGSPARYAIVYRRGWRAIAKALDIDEEDVPAIEVRAVKRN PLQPALYRILVRYGRVDLMPVTVDEVPPEMAGEFERLIERYDVPIDE
[0362] Residues 1-291 KEERILEILRENPWTPHDEIARRLGLSVSEVEGEKDPESSGIYSLW of SEQ ID NO: 1 SRVVVNIEYDERTAKRHVKRRDRLLEELYEHLEELSERYLRHPLTR RWIVEHKRDIMRRYLE
[0363] (HhH)? domain 1 QRIVECALKLQDRYGIREDVALCLARAFDGSISMIATTPYRTLKDV CPDLTLEEAKSVN
[0364] Residues 292- 350 of SEQ ID
[0365] NO: 1
[0366] (HhH)? domain 2 RTLATLIDEHGLSPDAADELIEHFESIAGILATDLEEIERMYEEGRL SEEAYRAAVEIQLAELTKKE
[0367] Residues 351- 417 of SEQ ID
[0368] NO: 1
[0369] (HhH)2 domain 3 GVGRKTAERLLRAFGNPERVKQLAREFEIEKLASVEGVGERVLRSL VP
[0370] Residues 418- 465 of SEQ ID
[0371] NO: 1
[0372] (HhH)2 domain 4 GYASLISIRGIDRERAERLLKKYGGYSKVREAGVEELREDGLTDAQ IRELKGL
[0373] Residues 466- 518 of SEQ ID
[0374] NO: 1
[0375] (HhH)2 domain 5 KTLESIVGDLEKADELKRKYGSASAVRRLPVEELRELGFSDDEIAEI KG
[0376] Residues 519- 567 of SEQ ID
[0377] NO: 1
[0378] (HhH)2 domain 6 IPKKLREAFDLETAAELYERYGSLKEIGRRLSYDDLLELGATPKAAA EIKG
[0379] Residues 568- 618 of SEQ ID
[0380] NO: 1
[0381]
[0382] (HhH)? domain 7, PEFKFLLNIEGVGPKLAERILEAVDYDLERLASLNPEELAEKVEGLG region divide and EELAERVVYAARERVESRRK
[0383] repair site II
[0384] Residues 619-685 of SEQ ID
[0385] NO: 1
[0386] (HhH)? domain 8 SGRQERSEEEWKEWLERKVGEGRARRLIEYFGSAGEVGKLVENA EVSKLLEVPGIGDEAVARLVPG
[0387] Residues 686-751 of SEQ ID
[0388] NO: 1
[0389] (HhH)2 domain 9 YKTLRDAGLTPAEAERVLKRYGSVSKVQEGATPDELRELGLGDAK IARILG
[0390] Residues 752-802 of SEQ ID
[0391] NO: 1
[0392] (HhH)2 domain LRSLVNKRLDVDTAYELKRRYGSVSAVRKAPVKELRELGLSDRKIA 10 RIKG
[0393] Residues 803-852 of SEQ ID
[0394] NO: 1
[0395] (HhH)2 domain IPETMLQVRGMSVEKAERLLERFDTWTKVKEAPVSELVRVPGVGL 11 SLVKEIKAQVD
[0396] Residues 853-908 of SEQ ID
[0397] NO: 1
[0398] (HhH)2 domain PAWKALLDVKGVSPELADRLVEELGSPYRVLTAKKSDLMRVERVG 12 PKLAERIRAAGKRYVEERRSRRERIRRKLRG
[0399] Residues 909-984 of SEQ ID
[0400] NO: 1
[0401] Preferred SVSAVRKAPVKELRELGLSDRKIARIKGIPETMLQVRGMSVEKAE fragment RLLERFDTWTKVKEAPVSELVRVPGVGLSLVKEIKAQVDPAWKAL LDVKGVSPELADRLVEELGSPYRVLTAKKSDLMRVERVGPKLAERI RAAGKRYVEERRSRRERIRRKLRG
[0402] Preferred SVSAVRKAPVKELRELGLSDRKIARIKGIPETMLQVRGMSVEKAE fragment RLLERFDTWTKVKEAPVSELVRVPGVGLSLVKEIKAQVDPAWKAL LDVKGVSPELADRLVEELGSPYRVLTAKKSDLMRVERVGPKLAERI SEQ ID NO: 15 RAAGKRYVEERRCRRERIRRKLRG
[0403] with cysteine
[0404] substitution
[0405] (underlined)
[0406] Fragment used in MSDSEVNQEAKPEVKPEVKPETHINLKVSDGSSEIFFKIKKTTPLR the Example RLMEAFAKRQGKEMDSLRFLYDGIRIQADQTPEDLDMEDNDIIEA HREQIGGSTSGSGENLYFQGSSGGSAWSHPQFEKGGGSGGGSG GSAWSHPOFEKGGSSVSAVRKAPVKELRELGLSDRKIARIKGIPE
[0407]
[0408] TM LO VRG M S V E K A E RLLE RFDTWTK V K E A PVS E LV RV PG VG LS LVThe underlined KEIKAOVDPAWKALLDVKGVSPELADRLVEELGSPYRVLTAKKSD sequence is SEQ LMRVERVGPKLAERIRAAGKRYVEERRCRRERIRRKLRG
[0409] ID NO: 16
[0410] The remaining
[0411] sequence is a
[0412] SUMO tag plus a
[0413] TEV protease
[0414] recognition site,
[0415] linkers, and
[0416] Strep-tag
[0417]
[0418] The following Examples illustrate the invention. It is to be understood that although particular embodiments, specific configurations as well as materials and / or molecules, have been described herein for methods according to the invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of this invention. The following examples are provided to better illustrate particular embodiments, and they should not be considered limiting the application. The application is limited only by the claims.
[0419] EXAMPLES
[0420] Example 1
[0421] This Example describes possible engineering routes of Topoisomerase V's DNA-binding domain for the development of a protein-based, generic DNA tether.
[0422] Topoisomerase V (Topo V) is a protein found in the archaeon Methanopyrus kandleri and has no known homologues. This protein has a topoisomerase domain (residues 1-268), and a DNA-binding domain from residue 269-984. This region can be further subdivided into two regions: 269-683, and 684-984. These two regions are separated by a loop, and both consist of tandem helix-turn-helix motifs (HhH2). The DNA-binding motifs are the regions of interest. Since these regions are made from tandem DNA-binding motifs, the interaction between this protein and DNA can be modulated by the number of helix-hairpin-helix motifs remaining in the protein. Two approaches can be pursued here: the domain may be truncated by sequentially removing HhH motifs, or by sequentially removing HhH2 motifs. These motifs may be sequentially removed either starting from the N- or the C-terminus of the truncated protein. The different domains exhibit different DNA-binding activity (see Figure 1). The data presented in Figure 1 was obtained sequencing a standard analyte as indicated in Example 3.
[0423] This domain of Topoisomerase V can further be modified by the replacement of a surface-exposed residue by a cysteine or any other amenable residue to allow for the conjugation of a hydrophobic group-carrying linker at different locations on the protein's surface. This can include cysteines artificially introduced in the protein to allow for specific modification at itsC- or N-terminus, or single point mutation of amino acids whose replacement is unlikely to negatively affect the protein's function, such as serines. The protein may also be modified to introduce a non-natural amino acid (e.g., 4-azidophenylalanine) that would allow downstream protein modification. Other potential protein modifications include the removal of amino acids with exposed free amines in order to allow targeted labelling of lysine residues. In addition to protein engineering that would allow for further protein processing (see Example 2), the protein may also be modified in order to better modulate its DNA-binding activity.
[0424] Example 2
[0425] This Example describes the functionalisation of a polynucleotide-binding protein with a hydrophobic moiety to generate a generic DNA-binding tether (Figure 2).
[0426] The protein is engineered to have a single, surface-exposed cysteine and a strep affinity tag. The protein sample is incubated with TCEP for 30 minutes at room temperature to ensure any disulfide bridge is reduced to a thiol group. Excess TCEP is removed by buffer exchanging against 50 mM Tris-HCI, 400 mM NaCI, 5% glycerol. The protein is then incubated with a linker which consists of a hydrophobic linker linked to a maleimide group by a polyethylene glycol (PEG) polymer of varying number of subunits - we have observed differences between using ~ 16 or ~ 114 PEG subunits -, where the conjugation with the protein will take place. This linker is added in a 15x excess to the protein. The conjugation reaction takes place for 18 hours at 37°C. After this incubation, excess linker is removed by affinity purification with MagStrep® Strep-Tactin®XT beads. The wash buffer used in this purification is 50 mM Tris-HCI, 400 mM NaCI, 5% glycerol, and the elution buffer has the same composition, but it is further supplemented with 50 mM biotin. After purification, the protein is buffer exchanged against 50 mM Tris-HCI, 400 mM NaCI, 5% glycerol with the aim of removing excess biotin.
[0427] Although the protocol above is the preferred method of functionalising Topo V, there are other available strategies. In one of them, Topo V is modified with a PEG that instead of carrying a maleimide on one end, is instead modified with N-hydroxysuccinimide (NHS). This allows the modification of amino acids displaying free amines, especially lysines. We have successfully modified Topo V using this chemistry, which resulted in a similar or lower number of strands as those obtained with Topo V modified via thiol-maleimide chemistry. In the former case, Topo V has potentially several hydrophobic group-carrying linkers on its surface, compared to a single one obtained by thiol-maleimide conjugation. The conjugation of several hydrophobic groups to the same protein may be beneficial in specific conditions. Other alternative conjugation chemistries that may be used include unnatural amino acids, such as 4-phenylazidophenylalanine, which could then be conjugated to a hydrophobic group which is itself modified with a dibenzocyclooctyne (DBCO) group.The nature of the hydrophobic group may as well be varied, and different groups used under similar conditions were shown to result in considerably different outcomes. The hydrophobic groups tested are cholesterol, palmitate, stearate, pyrene, and tocopherol, but this is not an exhaustive list of the hydrophobic groups that may be used. The experiments' outcome differed depending on the hydrophobic group used, indicating that the choice of hydrophobic anchor ultimately depends on the desired application of the generic tether whose activity may, apart from the ways mentioned before, be modulated by the nature of the hydrophobic group attached to it.
[0428] Example 3
[0429] This Example describes the use of a protein-based, generic tether in an Oxford Nanopore Technologies flow cell.
[0430] The flow cell is primed with Flow Cell Flush buffer. The functionalised protein is diluted to an appropriate amount in Flow Cell Flush buffer, which is then flowed into the flow cell and incubated for 5 minutes. Unbound protein is removed by flushing the flow cell with 0.5 ml of Flow Cell Flush buffer.
Claims
1. CLAIMS1. A polynucleotide-binding protein comprising (a) a fragment of DNA Topoisomerase V from Methanopyrus kandleri (Topo V) or (b) a variant thereof, wherein the protein lacks the amino-terminal (N-terminal) topoisomerase domain and one or more of tandem helix-hairpin-helix ((HhH)?) domains 1-9 from Topo V.
2. A polynucleotide-binding protein according to claim 1, wherein the protein comprises (a) (HhH)2domains 2-12, 3-12, 4-12, 5-12, 6-12, 7-12, 8-12, 9-12, 10-12 or 11-12 from Topo V or (b) a variant thereof.
3. A polynucleotide-binding protein according to claim 1 or 2, wherein the protein lacks all of (HhH)? domains 1-9 from Topo V.
4. A polynucleotide-binding protein according to claim 3, wherein the protein further lacks a portion of (HhH)? domain 10 from Topo V.
5. A polynucleotide-binding protein comprising (a) a fragment of DNA topoisomerase V from Methanopyrus kandleri (Topo V) or (b) a variant thereof, wherein the protein comprises (HhH)? domains 11-12 from Topo V or a variant thereof.
6. A polynucleotide-binding protein according to claim 5, wherein the protein further comprises a part of (HhH)? domain 10 from Topo V.
7. A polynucleotide-binding protein according to claim 5 or 6, wherein the protein lacks the N-terminal topoisomerase domain and one or more, or all, of (HhH)? domains 1-9 from Topo V.
8. A polynucleotide-binding protein according to any one of the preceding claims, wherein the fragment comprises or consists of the sequence shown in SEQ ID NO: 15.
9. A polynucleotide-binding protein according to any one of the preceding claims, wherein the protein is less than about 400 amino acids in length.
10. A polynucleotide-binding protein according to any one of the preceding claims, wherein the protein does not comprise the sequences of Topo-78, Topo-86, Topo-91, Topo-97, Topo-103 and full length Topo V.
11. A polynucleotide-binding protein according to any one of the preceding claims, wherein the variant (i) comprises a sequence having at least about 40% identity and / or homology to the sequence of the fragment and / or (ii) has a root mean square deviation (RMSD) of less than about 4.0 Angstroms (A) when compared with the fragment.
12. A polynucleotide-binding protein according to any one of the preceding claims, wherein the variant comprises one or more cysteine residues introduced into the fragment by substitution.
13. A polynucleotide-binding protein according to claim 12, wherein the one or more cysteine residues are introduced into the last 20 residues at the carboxyl-terminus (C-terminus) of the fragment.
14. A method for coupling a polynucleotide to a membrane, comprising coupling the polynucleotide to the membrane using one or more first coupling adaptors each comprising a polynucleotide-binding protein according to any one of the preceding claims.
15. A method according to claim 14, wherein the one or more first coupling adaptors each further comprises a membrane protein or a hydrophobic anchor.
16. A method according to claim 14 or 15, wherein the one or more first coupling adaptors each further comprise a linker.
17. A method according to claim 14, wherein the method further comprises using one or more second coupling adaptors each comprising a membrane protein or a hydrophobic anchor to couple the one or more first coupling adaptors to the membrane.
18. A method according to claim 17, wherein the one or more first coupling adaptors each further comprise a first oligonucleotide and the one or more second coupling adaptors each further comprise a second oligonucleotide which hybridizes to the first oligonucleotide.
19. A method according to any one of claims 14-18, wherein the method comprises simultaneously or sequentially (a) allowing the polynucleotide-binding protein in the one or more first coupling adaptors to bind to the polynucleotide and (b) coupling the one or more first coupling adaptors to the membrane, optionally using the one or more second coupling adaptors.
20. A method according to any one of claims 14-19, wherein the membrane is an amphiphilic layer or a solid state layer.
21. A method according to any one of claims 14-20, wherein the polynucleotide is coupled transiently or permanently to the membrane.
22. One or more coupling adaptors for coupling a polynucleotide to a membrane, each comprising a polynucleotide-binding protein according to any one of claims 1-13 and a membrane protein or a hydrophobic anchor.
23. A kit for coupling a polynucleotide to a membrane, comprising (1) one or more first coupling adaptors each comprising a polynucleotide-binding protein according to any one of claims 1-13 and a first oligonucleotide, and (2) one or more second coupling adaptors each comprising a membrane protein or a hydrophobic anchor and a second oligonucleotide which hybridizes to the first oligonucleotide.
24. A method for determining the presence, absence or one or more characteristics of a polynucleotide analyte, comprising (a) coupling the polynucleotide analyte to a membrane using a polynucleotide-binding protein according to any one of claims 1-13, a method according to any one of claims 14-21, one or more coupling adaptors according to claim 22 or a kit according to claim 23 and (b) allowing the coupled polynucleotide analyte to interact with a detector present in the membrane and thereby determining the presence, absence or one or more characteristics of the polynucleotide analyte.
25. A method according to claim 24, wherein before step (a) the polynucleotide analyte is not purified or is not separated from other components in a sample.
26. A method according to claim 24 or 25, wherein the detector comprises a transmembrane pore.
27. A method according to any one of claims 24-26, wherein step (b) comprises (i) allowing the polynucleotide analyte to interact with the transmembrane pore and (ii) measuring the current passing through the pore during the interaction and thereby determining the presence, absence or one or more characteristics of the polynucleotide analyte.
28. Use of a polynucleotide-binding protein according to any one of claims 1-13, a method according to any one of claims 14-21, one or more coupling adaptors according to claim 22 or a kit according to claim 23 for coupling a polynucleotide or a plurality of polynucleotides to a membrane.
29. A membrane comprising a polynucleotide or a plurality of polynucleotides coupled to it using a polynucleotide-binding protein according to any one of claims 1-13, a method according to any one of claims 14-21, one or more coupling adaptors according to claim 22 or a kit according to claim 23.
30. A membrane according to claim 29, wherein the membrane further comprises a detector.
31. An array comprising a plurality of membranes according to claim 29 or 30.
32. A system comprising (a) a membrane according to claim 29 or 30 or an array according to claim 31, (b) means for applying a potential across the membrane(s) and (c) means for detecting electrical or optical signals across the membrane(s).
33. An apparatus comprising a polynucleotide or a plurality of polynucleotides coupled to an in vitro membrane using a polynucleotide-binding protein according to any one of claims 1-13, a method according to any one of claims 14-21, one or more coupling adaptors according to claim 22 or a kit according to claim 23.
34. An apparatus produced by a method comprising coupling a polynucleotide or a plurality of polynucleotides to an in vitro membrane using a polynucleotide-binding protein according to any one of claims 1-13, a method according to any one of claims 14-21, one or more coupling adaptors according to claim 22 or a kit according to claim 23.
35. An apparatus according to claim 33 or 34, wherein the membrane comprises a detector.
36. A polynucleotide encoding a polynucleotide-binding protein according to any one of claims 1-13.
37. An expression vector comprising a polynucleotide according to claim 36.
38. A host cell comprising a polynucleotide according to claim 36 or an expression vector according to claim 37.
39. A kit for characterising a polynucleotide analyte comprising (a) a polynucleotide-binding protein according to any one of claims 1-13, one or more coupling adaptors according to claim 22 or a kit according to claim 23 and (b) the components of a membrane.
40. A kit for characterising a polynucleotide analyte comprising (a) a membrane comprising a detector and (b) a polynucleotide-binding protein according to any one of claims 1-13, one or more coupling adaptors according to claim 22 or a kit according to claim 23.
41. An apparatus for characterising a polynucleotide analyte in a sample, comprising (a) a plurality of membranes each comprising a detector and (b) a plurality of polynucleotide-binding proteins according to any one of claims 1-13, a plurality of one or more coupling adaptors according to claim 22 or a plurality of kits according to claim 23.