Method and Adapter

JP2025517736A5Pending Publication Date: 2026-03-11OXFORD NANOPORE TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-05-16
Publication Date
2026-03-11

AI Technical Summary

Technical Problem

Current methods for sequencing telomeres are inefficient due to their repetitive and non-linear nature, which makes it difficult to accurately characterize them without the need for restriction digestion and PCR.

Method used

A method involving the ligation of a polynucleotide telomere adapter to the 5' end of the non-overhanging strand at the telomere end, allowing for specific characterization of the telomere in the 5' to 3' direction without the need for restriction digestion or PCR.

Benefits of technology

This method enables specific and effective characterization of telomeres with high resolution of sequence and length, allowing for characterization beyond the telomere into the chromosome, and identifying modifications such as methylation and the presence of telomere-binding proteins.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a method for characterizing (e.g., sequencing) at least a part of a telomere, and an adapter for use in such a method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for characterizing (e.g., sequencing) at least a part of a telomere, and an adapter for use in such a method.

Background Art

[0002] Telomeres, which are regions of repetitive DNA sequences at the ends of eukaryotic chromosomes, are difficult to characterize (e.g., sequence). This is due to various reasons including, but not limited to, telomeres being highly repetitive and non-linear, themselves looping back to form displacement loops (D-loops), and telomere-binding proteins binding along the D-loops. Methods for sequencing telomeres are known, but these methods typically involve restriction digestion and / or PCR and Southern blot analysis (Lai, T.P., Zhang, N., Noh, J. et al. Nat Commun 8, 1356 (2017), Bendix L, Horn PB, Jensen UB, Rubelj I, Kolvraa S. Aging Cell. 2010 Jun;9(3):383-97, and Bennett HW, Liu N, Hu Y, King MC. FEBS Lett. 2016 Dec;590(23):4159-4170). The sequences of telomeres have to be reconstructed from shorter sequences in these methods, and these methods focus on characterizing short telomeres. The use of PCR is problematic because polymerases cannot effectively copy long stretches of repetitive sequences such as those found in telomeres.

[0003] Biological pores (and other nanopores) have great potential as direct electrical biosensors for polymers and various small molecules. In particular, recently, nanopores have attracted attention as potential DNA sequencing technologies. When a potential is applied across a nanopore, if an analyte such as a nucleotide temporarily resides in the barrel for a certain period of time, there is a change in the current flow. Nanopore detection of nucleotides results in current changes with known signatures and durations. In the strand sequencing method, a single polynucleotide strand passes through the pore and the identity of the nucleotide is deduced. Strand sequencing may involve the use of a molecular brake to control the movement of the polynucleotide through the pore.

[0004] Nanopores have been used to characterize telomeres by extending the 3' overhang of telomeres with a polyA tail (Sholes SL, Karimian K, Gershman A, Kelly TJ, Timp W, Greider CW. Genome Res. 2022 Apr;32(4):616 - 628. doi: 10.1101 / gr.275868.121.), but this method is not specific to telomeres and many other 3' overhangs in genomic DNA are also extended. This method also used polymerase to fill the overhangs before additional tagging. An improved method for characterizing telomeres is needed. Summary of the Invention

[0005] The inventors have identified certain methods for characterization, such as sequencing at least a portion of a telomere. This method involves ligating a polynucleotide telomere adapter to the 5' end of the non-overhanging strand at the end of the telomere. As will be described in more detail below, the 3' end of the adapter specifically hybridizes to the first portion of the overhanging strand, so that only the ends of the telomere are adapted. The telomere adapter can then be used to characterize at least a portion of the ligated non-overhanging strand of the telomere in the 5' to 3' direction from the end of the telomere. Since the telomere adapter is not ligated to any other part of the chromosome, this method specifically characterizes at least a portion of the telomere. In essence, the method is an enrichment strategy for specifically characterizing at least a portion of the telomere. One or more parts or all of the telomere can be specifically characterized, and the method can also involve characterizing one or more parts or all of the subtelomere, chromatin (i.e., genomic DNA), the opposite subtelomere, and the opposite telomere. The method can be used in combination with nanopore sequencing, but it is not essential. The method of the present invention has several advantages including specifically and effectively characterizing at least a portion of the telomere, high resolution of sequence and length, the ability to characterize beyond the telomere and further into the chromosome (including the possibility of characterizing the entire chromosome), the absence of the need for restriction digestion and fragment size control, the absence of the need for PCR, the ability to identify methylated nucleotides in the telomere or other modifications, the ability to identify the presence or absence of telomere-binding proteins and characterize such proteins, but is not limited to these.

[0006] In a preferred embodiment, the method also involves creating a double-strand break from the telomere end to the opposite end of at least a portion of the telomere and characterizing the unligated overhanging strand of at least a portion of the telomere in the opposite direction, i.e., towards the end of the telomere. This allows both strands of at least a portion of the telomere to be characterized, increasing the resolution particularly with respect to sequencing of repetitive telomere sequences and identification of modifications such as methylation.

[0007] The present invention relates to a method for characterizing at least a part of a telomere, comprising: (a) ligating a polynucleotide telomere adapter to the 5'-end of the non-overhang strand at the end of the telomere, wherein the 3'-end of the adapter specifically hybridizes to the first part of the overhang strand and the 5'-end of the adapter does not hybridize to the opposite part of the overhang strand; (b) using the telomere adapter to characterize at least a part of the ligated non-overhang strand of the telomere in the 5' to 3' direction from the end of the telomere. A method is provided that includes these steps.

[0008] The present invention also relates to a method for characterizing at least a part of a telomere, comprising: (a) ligating a polynucleotide telomere adapter to the 5'-end of the non-overhang strand at the end of the telomere, wherein the 3'-end of the adapter specifically hybridizes to the first part of the overhang strand and the 5'-end of the adapter does not hybridize to the opposite part of the overhang strand; (b) using a polymerase-induced effector protein to create a double-strand break from the telomere end to the opposite end of at least a part of the telomere and attaching a sequencing adapter to the opposite end; (c) using the telomere adapter to characterize at least a part of the ligated non-overhang strand of the telomere in the 5' to 3' direction from the end of the telomere, and using the sequencing adapter to characterize at least a part of the unligated overhang strand of the telomere in the 5' to 3' direction up to the end of the telomere. A method is provided that includes these steps.

[0009] The present invention also provides the following. - A polynucleotide telomere adapter, wherein the 3' end of the adapter specifically hybridizes to the first part of the overhang strand at the end of the telomere, and the 5' end of the adapter does not hybridize to the opposite part of the overhang strand at the end of the telomere, the polynucleotide telomere adapter, - A population of six telomere adapters, each having a 3' end that specifically hybridizes to one of six possible sequences of the first part of the overhang strand at the end of the telomere, and a 5' end that does not hybridize to the opposite part of the overhang strand at the end of the telomere, the population of six telomere adapters, - A kit for characterizing at least a part of a telomere, comprising: (a) one or more polynucleotide telomere adapters of the present invention or a population of six telomere adapters of the present invention; and (b) one or more sprint polynucleotides or one or more polynucleotide extensions, the kit, and - A system comprising: (a) one or more polynucleotide telomere adapters of the present invention or a population of six telomere adapters of the present invention; and (b) a nanopore.

[0010] It should be understood that the figures are for the sole purpose of illustrating certain embodiments of the present invention and are not intended to be limiting.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

[0012] Description of the Sequence Listing SEQ ID NOs: 1-6 show the sequences of the telomere adapters used in Example 1 (Table 3 and FIG. 1).

[0013] SEQ ID NO: 7 shows the preferred 5' end of the telomere adapter of the present invention. This is present in all of SEQ ID NOs: 1-6 and 14-19.

[0014] SEQ ID NO: 8 shows the preferred sequence of the sprint polynucleotide, which is the reverse complement of SEQ ID NO: 7.

[0015] SEQ ID NOs: 9-10 show the sprint polynucleotides used in the examples (Table 5).

[0016] SEQ ID NO: 11 shows the polynucleotide extension used in the examples (Table 7).

[0017] SEQ ID NOs: 12-13 show the sequences of the biotinylated telomere adapters used in the examples (Table 9).

[0018] SEQ ID NOs: 14-19 show the sequences of the telomere adapters used in Example 10 (Table 14).

[0019] SEQ ID NO: 20 shows the sprint polynucleotide used in Example 10 (Table 15).

[0020] Sequence number 21 shows the sequence of the overhang strand above the exemplary chromosome ends in FIGS. 1 - 6.

[0021] Sequence number 22 shows the sequence (in the 5' to 3' direction) of the underhang (non - overhang) strand of the exemplary chromosome ends in FIGS. 1 - 6.

[0022] Sequence number 23 shows the sequence (in the 5' to 3' direction) formed by the attachment of the T1 telomere adapter to the non - overhang strand in FIG. 1.

[0023] Sequence number 24 shows the sequence (in the 5' to 3' direction) formed by the attachment of the telomere adapter in FIG. 3 to the sequencing adapter. This sequence includes sequence number 23.

[0024] Sequence numbers 25 - 30 show the sequences (in the 5' to 3' direction) of the extended telomere adapters in FIG. 4.

[0025] Sequence number 31 shows the sequence (in the 5' to 3' direction) formed by the attachment of the extended T1 telomere adapter to the non - overhang strand in FIG. 4.

[0026] Sequence number 32 shows the sequence of the upper strand in FIG. 6.

[0027] Sequence number 33 shows the sequence (in the 5' to 3' direction) of the lower strand in FIG. 6. DETAILED DESCRIPTION OF THE INVENTION

[0028] It should be understood that different applications of the disclosed products and methods can be tailored to specific needs in the art. It should also be understood that the terms used herein are for the purpose of describing only particular embodiments of the invention and are not intended to be limiting.

[0029] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. All publications, patents, and patent applications mentioned herein are incorporated by reference herein as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent that the publications and patents or patent applications incorporated by reference conflict with the disclosure contained herein, this specification is intended to supersede and / or take precedence over any such conflicting material.

[0030] Definitions When an indefinite or definite article, e.g., "a" or "an", "the", is used to refer to a singular noun, unless otherwise specified, this includes the plural form of that noun. When the term "comprising" is used in this specification and the claims, this does not exclude other elements or steps. Further, the terms first, second, third, etc. in this specification and the claims are used to distinguish similar elements and are not necessarily used to describe the order in which they occur or the chronological order. The terms so used are interchangeable under appropriate circumstances, and it should be understood that the embodiments of the invention described herein are operable in an order other than that described or illustrated herein. The following terms or definitions are provided only to assist in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as would be understood by one of ordinary skill in the art of the invention. Practitioners are referred to the definitions and terms of the art, in particular, Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 thCited are ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016). The definitions provided herein should not be construed as being narrower in scope than would be understood by one of ordinary skill in the art.

[0031] As used herein when referring to measurable values such as amounts, temporal durations, etc., "about" means encompassing a variation of ±20% or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1% from the specified value, such variations being appropriate to practice the disclosed methods.

[0032] As used herein, the terms "nucleotide sequence," "DNA sequence," or "nucleic acid molecule(s)" refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double-stranded and single-stranded DNA and RNA. As used herein, the term "nucleic acid" is a single-stranded or double-stranded covalently linked nucleotide sequence in which the 3' and 5' ends of each nucleotide are linked by phosphodiester bonds. Polynucleotides can be composed of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids can be synthetically produced in vitro or isolated from natural sources. Nucleic acids can further include modified DNA or RNA, such as methylated DNA or RNA, or post-translational modifications, such as 5'-capping with 7-methylguanosine, 3'-processing such as cleavage and polyadenylation, and RNA that has been subjected to splicing. Nucleic acids can also include synthetic nucleic acids (XNAs) such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA), and peptide nucleic acid (PNA). The size of a nucleic acid, also referred to herein as a "polynucleotide," is typically represented by the number of base pairs (bp) of a double-stranded polynucleotide or the number of nucleotides (nt) in the case of a single-stranded polynucleotide. 1000 bp or nt corresponds to a kilobase (kb). Polynucleotides less than about 40 nucleotides in length are typically referred to as "oligonucleotides" and can include primers for use in the manipulation of DNA, such as via polymerase chain reaction (PCR).

[0033] The term "amino acid" as used in the context of the present disclosure is used in its broadest sense and includes, along with the side chain (e.g., R group) specific to each amino acid, an amine (NH 2It means including an organic compound containing an amino (NH₂) and a carboxyl (COOH) functional group. Amino acids typically refer to naturally occurring Lα-amino acids or residues. One-letter and three-letter abbreviations commonly used for naturally occurring amino acids are used herein: A = Ala, C = Cys, D = Asp, E = Glu, F = Phe, G = Gly, H = His, I = Ile, K = Lys, L = Leu, M = Met, N = Asn, P = Pro, Q = Gln, R = Arg, S = Ser, T = Thr, V = Val, W = Trp, and Y = Tyr (Lehninger, A.L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes chemically modified amino acids such as D-amino acids, retro-inverso amino acids, and amino acid analogs, naturally occurring amino acids not normally incorporated into proteins such as norleucine, and chemically synthesized compounds having properties known in the art that exhibit the characteristics of amino acids. For example, analogs or mimetics of phenylalanine or proline that allow for the same conformational restrictions in a peptide compound as natural Phe or Pro are included within the definition of an amino acid. Such analogs and mimetics are referred to herein as "functional equivalents" of the respective amino acids. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.

[0034] The terms "polypeptide" and "peptide" are used interchangeably herein to refer to a polymer of amino acid residues, as well as variants and synthetic analogs thereof. Thus, these terms apply to amino acid polymers in which one or more amino acid residues are synthetic non-naturally occurring amino acids, such as chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. A polypeptide can also undergo maturation or post-translational modification processes including, but not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, and the like. A peptide can be made using recombinant techniques, for example, by expression of a recombinant or synthetic polynucleotide. A recombinantly produced peptide typically contains substantially no culture medium, for example, the culture medium corresponds to less than about 20%, more preferably less than about 10%, and most preferably less than about 5% of the volume of the protein preparation.

[0035] The term "protein" is used to describe a folded polypeptide having a secondary or tertiary structure. A protein can be composed of a single polypeptide or can contain multiple polypeptides that aggregate to form a multimer. The multimer can be a homo-oligomer or a hetero-oligomer. A protein can be a naturally occurring protein or wild-type protein, or it can be a modified protein or a non-naturally occurring protein. A protein can differ from a wild-type protein, for example, by the addition, substitution, or deletion of one or more amino acids.

[0036] A "variant" of a protein includes peptides, oligopeptides, polypeptides, proteins, and enzymes that have amino acid substitutions, deletions, and / or insertions compared to the unmodified or wild-type protein in question and have biological and functional activities similar to those of the unmodified protein from which they are derived. As used herein, the term "amino acid identity" refers to the degree to which sequences are identical amino acid by amino acid over a comparison window. Thus, the "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over a comparison window, determining the number of positions at which identical amino acid residues (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, Ile, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gln, Cys, and Met) occur in both sequences to calculate the number of matching positions, dividing the number of matching positions by the total number of positions within the comparison window (i.e., window size), and multiplying the result by 100 to calculate the percentage of sequence identity.

[0037] For all aspects and embodiments of the present invention, a "variant" has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% complete sequence identity to the amino acid sequence of the corresponding wild-type protein. Sequence identity can also be with respect to a fragment or portion of a full-length polynucleotide or polypeptide. Thus, a sequence can have only 50% overall sequence identity to a full-length reference sequence, but the sequences of specific regions, domains, or subunits can share 80%, 90%, or 99% sequence identity with the reference sequence.

[0038] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is the most frequently observed in a population and is thus the arbitrarily designed "normal" or "wild-type" form of that gene. In contrast, the terms "modified", "variant", or "mutant" refer to a gene or gene product that exhibits a modification of the sequence (e.g., substitution, cleavage, or insertion), post-translational modification, and / or functional characteristics (e.g., altered features) compared to the wild-type gene or gene product. It should be noted that naturally occurring variants can be isolated and are identified by the fact that they have altered features compared to the wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, methionine (M) can be substituted with arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the variant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, non-naturally occurring amino acids can be introduced by including synthetic aminoacyl-tRNA in the IVTT system used to express the variant monomer. Alternatively, they can be introduced by expressing the variant monomer in E. coli that is auxotrophic for a particular amino acid in the presence of a synthetic (i.e., non-naturally occurring) analogue of that amino acid. They may also be produced by naked ligation if the variant monomer is produced using solid-phase peptide synthesis. Conservative substitutions replace an amino acid with another amino acid of similar chemical structure, similar chemical properties, or similar side-chain volume. The amino acids introduced can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge as the amino acids they replace. Alternatively, conservative substitutions can introduce another amino acid that is aromatic or aliphatic in place of an existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 major amino acids defined in Table 1 below.When the amino acids have similar polarities, this can also be determined with reference to the hydrophobicity scale of the amino acid side chains in Table 2.

Table 1

Table 2

[0039] Variants or modified proteins, monomers, or peptides can also be chemically modified in any manner and at any site. Variants or modified monomers or peptides are preferably chemically modified by attachment of a molecule to one or more cysteines (cystine bond), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzymatic modification of an epitope, or modification of the termini. Suitable methods for performing such modifications are well known in the art. Variants of modified proteins, monomers, or peptides can be chemically modified by attachment of any molecule. For example, variants of modified proteins, monomers, or peptides can be chemically modified by attachment of a dye or fluorophore.

[0040] The method of the present invention The present invention provides a method for characterizing at least a portion of a telomere. A telomere is a region of repetitive DNA sequences at the ends of eukaryotic chromosomes. Each chromosome has telomeres at each end. Telomeres protect the terminal regions of chromosomal DNA from progressive degradation and ensure the integrity of linear chromosomes by preventing DNA repair systems from misinterpreting the ends of DNA strands as double-strand breaks. Chromosomes are structures formed from long DNA molecules that contain all or part of the genetic material within eukaryotic cells. A section of DNA called subtelomere typically separates the telomere from the chromatin (i.e., genomic DNA) in the chromosome. The structure of a chromosome is typically telomere - subtelomere - chromatin - subtelomere - telomere.

[0041] The term "part" in at least a part of a telomere is interchangeable with "portion". This part can be any amount of the telomere. At least a part of the telomere is preferably at least about 5% of the telomere, such as at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The method is preferably for characterizing all of the telomere (or the entire telomere).

[0042] This method is preferably for characterizing all of the telomere (or the entire telomere) and additional parts of the chromosome. The additional parts of the chromosome can be any additional amount of the chromosome. The additional part can be any percentage of the remaining part of the chromosome as discussed above in relation to at least a part of the telomere.

[0043] This method is preferably for characterizing (i) all of the telomere, (ii) all of the telomere and at least a part or all of the subtelomere, (iii) all of the telomere, all of the subtelomere, and at least a part or all of the chromatin, (iv) all of the telomere, all of the subtelomere, all of the chromatin, and at least a part or all of the opposite subtelomere, or (v) all of the telomere, all of the subtelomere, all of the chromatin, all of the opposite subtelomere, and at least a part or all of the opposite telomere. Chromatin is interchangeable with genomic DNA.

[0044] This method is preferably for (i), i.e., for characterizing all telomeres. This method is preferably for characterizing (ii). This method is preferably for characterizing all telomeres and all subtelomeres. This method is preferably for characterizing (iii). This method is preferably for characterizing all telomeres, all subtelomeres, and all chromatin. This method is preferably for characterizing (iv). This method is preferably for characterizing all telomeres, all subtelomeres, all chromatin, and all opposite subtelomeres. This method is preferably for characterizing (v). This method is preferably for characterizing all telomeres, all subtelomeres, all chromatin, all opposite subtelomeres, and all opposite telomeres.

[0045] In any of (ii) to (v), at least a part of the subtelomere / chromatin / opposite subtelomere / opposite telomere can be any of the percentages considered above in relation to at least a part of the telomere.

[0046] This method is preferably for characterizing all of the chromosome (or the entire chromosome). The method of the present invention is preferably repeated at the other end of the chromosome, and this method includes characterizing both strands of the entire chromosome. The method of the present invention preferably includes performing the method of the present invention at both ends of the chromosome and includes characterizing both strands of the entire chromosome. Any of the methods of the present invention can be used at each end of the chromosome. The methods used at each end may be the same or different. These methods are preferably the same.

[0047] The method preferably includes, in step (b), characterizing the ligated non-overhang strand of (i) a telomere, (ii) at least a part or all of a subtelomere, (iii) at least a part or all of chromatin, (iv) at least a part or all of the opposite subtelomere, or (v) at least a part or all of the opposite telomere, in the 5' to 3' direction from the end of the telomere. Any of the embodiments discussed above for (i) to (v) are equally applicable to step (b) of the method.

[0048] At least a part of the telomere, including any one of (i) to (v) discussed above, can be of any length. For example, at least a part of the telomere can be at least 10, at least 50, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides or nucleotide pairs in length. At least a part of the telomere can be 1000 or more nucleotides or nucleotide pairs in length, or 5000 or more nucleotides or nucleotide pairs in length, or 100000 or more nucleotides or nucleotide pairs in length, or 500,000 or more nucleotides or nucleotide pairs in length, or 1,000,000 or more nucleotides or nucleotide pairs in length, or 10,000,000 or more nucleotides or nucleotide pairs in length, or 100,000,000 or more nucleotides or nucleotide pairs in length, or 200,000,000 or more nucleotides or nucleotide pairs in length, or the full length of the chromosome.

[0049] This method is preferably for characterizing at least a part of one or both telomeres on each of two or more different chromosomes. Thereby, it becomes possible to compare the characteristics of telomeres on two or more different chromosomes. The two or more different chromosomes can be from the same cell, tissue, organism or taxonomic rank (such as genus or species), or from different cells, tissues, organisms or taxonomic ranks (such as genus or species). The cell(s), tissue(s), organism(s), or taxonomic rank(s) are typically eukaryotic. This method can be for characterizing at least a part of one or both telomeres on each of any number of two or more different chromosomes, for example, about three or more, about four or more, about five or more, about ten or more, about fifteen or more, about twenty or more, about twenty-five or more, about thirty or more, or about forty or more different chromosomes. Preferred numbers of different chromosomes include, but are not limited to, 2, 4, 6, 7, 8, 9, 10, 11, 12, 14, 16, 17, 18, 20, 22, 24, 26, 28, 30, 31, 32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62, 63, 64, 66, 68, 70, 72, 74, 76, 78, 80 and 82. The number of different chromosomes can be even larger in polyploid cells, tissues, organisms, or taxonomic ranks. This method is preferably for characterizing at least a part of one or both telomeres on each of 46 different chromosomes. In these embodiments, the method comprises: (a) ligating a polynucleotide telomere adapter to the 5' end of the non-overhang strand at the end of each telomere, wherein the 3' end of the adapter specifically hybridizes to the first part of the overhang strand and the 5' end of the adapter does not hybridize to the opposite part of the overhang strand; and (b) using each telomere adapter to characterize at least a part of the ligated non-overhang strand of each telomere in the 5' to 3' direction from the end of each telomere.

[0050] The method preferably includes (a) sequencing at least a part of the telomere, (b) measuring the length of at least a part of the telomere, (c) telomere-to-telomere assembly of chromosomes or genomes, (d) identifying telomere or chromosome fusions, (e) identifying one or more modifications in at least a part of the telomere, (f) identifying at least a part of the telomere as a variant, or (g) linking at least a part of the telomere to a specific cell or tissue type. The method can include any number and combination of (a)-(g).

[0051] The method preferably includes (a). The method of sequencing the ligated non-overhanging strand will be considered in more detail below.

[0052] The method preferably includes (b). The lack of digestion in the method of the present invention means that the length of at least a part of the telomere or all of the telomere can be easily measured in one step.

[0053] The method preferably includes (c). As described above, according to the method, it becomes possible to characterize the entire chromosome from telomere to telomere. Also, as described above, the method can be applied to two or more different chromosomes, which means that the entire genome can be characterized.

[0054] The method preferably includes (d). By characterizing at least a part of the telomere, it is possible to identify telomere fusions using the method.

[0055] This method preferably includes (e). One or more modifications preferably include (i) methylation of one or more nucleotides, (ii) oxidation of one or more nucleotides, and (iii) one or more of the damages to one or more nucleotides. This method can identify (i), (ii), (iii), (i) and (ii), (i) and (iii), (ii) and (iii), or (i), (ii) and (iii). For example, at least a part of the telomere may contain one or more pyrimidine dimers. Such dimers are typically associated with damage by ultraviolet light and are the main cause of skin melanoma. Nanopore sequencing can identify methylated, oxidized, and damaged nucleotides.

[0056] This method preferably includes (f). This method can identify at least a part of the telomere as a variant including one or more of (i) telomere deletion, (ii) telomere addition, and / or (ii) telomere substitution. This method can identify at least a part of the telomere as a variant including (i), (ii), (iii), (i) and (ii), (i) and (iii), (ii) and (iii), or (i), (ii) and (iii). (i) can include deletions of any number of nucleotides from the telomere. (i) can include a single deletion or multiple deletions. The deletion(s) can be on one or both strands. (ii) can include additions of any number of nucleotides to the telomere. (ii) can include a single addition or multiple additions. The addition(s) can be on one or both strands. (iii) can include substitutions of any number of nucleotides within the telomere. The substitution(s) can be on one or both strands.

[0057] This method preferably includes (g). When this method is performed on a specific cell, tissue, organism, or taxonomic rank (such as a genus or species), at least a part of the telomere can be linked to the cell, tissue, organism, or taxonomic rank (such as a genus or species). The method of the present invention can be performed on different chromosomes from different cells, tissues, or taxonomic ranks as described above, and thus each at least a part of the telomere can be linked to different cells, tissues, or taxonomic ranks. The cell(s), tissue(s), organism(s), or taxonomic rank(s) are typically eukaryotic.

[0058] Telomere adapter The method of the present invention includes, in step (a), ligating a polynucleotide telomere adapter to the 5' end of the non-overhanging strand of the telomere end. Methods for ligating polynucleotides are known in the art. The telomere adapter is formed from at least one polynucleotide as discussed below. The polynucleotide telomere adapter is compatible with the telomere adapter.

[0059] Telomeres typically have a 3' overhang (see Figure 1). This means that the strand in the 5' to 3' direction typically overhangs at the end of the telomere. The strand in the 3' to 5' direction at the end of the telomere is typically in a non-overhanging state. "Non-overhanging" is interchangeable with "underhanging". Both strands at the end of the telomere are typically DNA.

[0060] The telomere adapter can be any type of polynucleotide. A polynucleotide such as a nucleic acid is a macromolecule containing two or more nucleotides. The polynucleotide can be single-stranded or double-stranded. A double-stranded polynucleotide is made up of two single-stranded polynucleotides hybridized together. The polynucleotide telomere adapter can be a single-stranded polynucleotide or a double-stranded polynucleotide.

[0061] A polynucleotide can contain any combination of any nucleotides. The nucleotides can be either naturally occurring or artificial. Nucleotides typically contain a nucleobase, a sugar, and at least one phosphate group. The nucleobase and the sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, more specifically, adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C).

[0062] The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. The polynucleotide preferably contains the following nucleosides: deoxyadenosine (dA), deoxythymidine (dU) and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC).

[0063] Nucleotides are typically ribonucleotides or deoxyribonucleotides. Nucleotides typically contain mono-, di-, or tri-phosphate. Nucleotides can contain more than three phosphates, for example, four or five phosphates. The phosphate can be attached to the 5’ or 3’ side of the nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. Nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP.

[0064] The nucleotide can be abasic (i.e., lacking a nucleobase). The nucleotide may also lack a nucleobase and a sugar (i.e., be a C3 spacer).

[0065] The nucleotides in the polynucleotide can be attached to each other in any manner. The nucleotides are typically attached by their sugars and phosphate groups, as in the case of nucleic acids. The nucleotides can be connected via nucleobases, as in the case of pyrimidine dimers.

[0066] The polynucleotide can be a nucleic acid such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The polynucleotide can include a single strand of RNA hybridized to a single strand of DNA. The polynucleotide can be any synthetic nucleic acid known in the art, for example, peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), bridged nucleotide (BNA), locked nucleic acid (LNA), or other synthetic polymers having nucleotide side chains. The PNA backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone is composed of repeating glycol units linked by phosphodiester bonds. The TNA backbone is composed of repeating threose sugars linked together by phosphodiester bonds. LNA is formed from ribonucleotides having an extra bridge connecting the 2'-oxygen and 4'-carbon in the ribose moiety, as discussed above.

[0067] The polynucleotide is preferably DNA, RNA, or a DNA or RNA hybrid, most preferably DNA. The DNA / RNA hybrid can contain DNA and RNA on the same strand. Preferably, the DNA / RNA hybrid contains a single strand of DNA hybridized to an RNA strand.

[0068] The backbone of the polynucleotide can be modified to reduce the possibility of strand breakage. For example, DNA is known to be more stable than RNA under many conditions. The backbone of the polynucleotide strand can be modified to avoid damage caused by aggressive chemical substances such as free radicals. DNA or RNA containing unnatural or modified bases can be generated by amplifying natural DNA or RNA polynucleotides in the presence of modified NTPs using appropriate polymerases.

[0069] The telomere adapter can be of any length. For example, the telomere adapter can be at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, or at least 50 nucleotides or nucleotide pairs in length.

[0070] The telomere adapter is typically a single-stranded oligonucleotide. An oligonucleotide is typically a short nucleotide polymer having 50 or fewer nucleotides, for example, 45 or fewer, 42 or fewer, 41 or fewer, 40 or fewer, 35 or fewer, or 30 or fewer nucleotides. The telomere adapter is preferably about 15 to about 50 nucleotides in length, for example, about 20 to about 45 nucleotides in length. For example, the oligonucleotide can be about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 nucleotides, about 35 nucleotides, about 40 nucleotides, about 41 nucleotides, about 42 nucleotides, or 45 nucleotides in length. The polynucleotide telomere adapter is most preferably about 30 nucleotides or about 41 nucleotides in length. Different regions of the telomere adapter are considered in more detail below.

[0071] Telomere adapters are typically synthetic or semi-synthetic. For example, DNA or RNA can be a pure synthetic compound synthesized by conventional DNA synthesis methods such as phosphoramidite-based chemical reactions. Synthetic polynucleotide subunits can be ligated together by known means such as ligation or chemical bonding to generate longer strands. Internal self-forming structures (e.g., hairpins, quadruplexes) can be designed into the substrate, for example, by ligating appropriate sequences. Synthetic polynucleotides can be replicated and scaled up for production by means known in the art, including PCR, incorporation into bacterial factories, and the like.

[0072] The polynucleotide telomere adapter has a 5' to 3' directionality. The 3' end of the adapter specifically hybridizes to the first part of the overhang strand. The 3' end can be any part or portion of the polynucleotide telomere adapter, for example, at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 30%, at least about 40%, or at least about 50%. The 3' end of the adapter can be of any length, for example, at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 10 nucleotides, or at least about 20 nucleotides in length.

[0073] The 3' end of the adapter is preferably about 18 nucleotides or less, for example, about 16 nucleotides or less, about 15 nucleotides or less, about 14 nucleotides or less, about 10 nucleotides or less, about 8 nucleotides or less, or about 7 nucleotides or less. The 3' end is preferably about 4 to about 18 nucleotides in length, for example, about 5 to about 16, about 6 to about 15, or about 7 to about 10 nucleotides in length. For example, the 3' end can be about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, or about 18 nucleotides in length. The 3' end of the polynucleotide telomere adapter is most preferably about 7 nucleotides in length or about 18 nucleotides in length.

[0074] The 3'-end of the telomere adapter specifically hybridizes to the first portion of the overhang strand. The first portion of the overhang strand may be any of the lengths discussed above in relation to the 3'-end. The first portion of the overhang strand is preferably the same length as the 3'-end of the telomere adapter.

[0075] The 3' end "specifically hybridizes" to the first portion of the overhang strand with a preferential or high affinity, but does not hybridize, hybridizes with very low affinity, or substantially does not hybridize to other polynucleotide sequences, particularly other sequences in telomeres or chromosomes. Conditions that enable hybridization are well known in the art (e.g., Sambrook et al., 2001, Molecular Cloning: a laboratory manual, 3rd edition, Cold Spring Harbour Laboratory Press, and Current Protocols in Molecular Biology, Chapter 2, Ausubel et al., Eds., Greene Publishing and Wiley-Interscience, New York (1995)). Hybridization can be carried out under low stringency conditions, for example, in the presence of a buffer solution of 30 - 35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate) at 37°C, followed by washing 20 times in 1× (0.1650 M Na+) - 2× (0.33 M Na+) SSC (standard sodium citrate) at 50°C. Hybridization can be carried out under moderately stringent conditions, for example, in the presence of a buffer solution of 40 - 45% formamide, 1 M NaCl, and 1% SDS at 37°C, followed by washing once in 0.5× (0.0825 M Na+) - 1× (0.1650 M Na+) SSC at 55°C. Hybridization can be carried out under high stringency conditions, for example, in the presence of a buffer solution of 50% formamide, 1 M NaCl, 1% SDS at 37°C, followed by washing once in 0.1× (0.0165 M Na+) SSC at 60°C.The 3' end "specifically hybridizes" when it hybridizes to the first portion of the overhang strand at a melting temperature (Tm) that is at least 2°C, for example, at least 3°C, at least 4°C, at least 5°C, at least 6°C, at least 7°C, at least 8°C, at least 9°C, or at least 10°C higher than the Tm for other polynucleotide sequences. More preferably, the 3' end hybridizes to the first portion of the overhang strand at a Tm that is at least 2°C, for example, at least 3°C, at least 4°C, at least 5°C, at least 6°C, at least 7°C, at least 8°C, at least 9°C, at least 10°C, at least 20°C, at least 30°C, or at least 40°C higher than the Tm for other polynucleotide sequences. Preferably, the 3' end hybridizes to the first portion of the overhang strand at a Tm that is at least 2°C, for example, at least 3°C, at least 4°C, at least 5°C, at least 6°C, at least 7°C, at least 8°C, at least 9°C, at least 10°C, at least 20°C, at least 30°C, or at least 40°C higher than the Tm for a polynucleotide that differs from the first portion of the overhang strand by one or more nucleotides, for example, 1, 2, 3, 4, or 5 or more nucleotides. The 3' end typically hybridizes to the first portion of the overhang strand at a Tm of at least 90°C, for example, at least 92°C or at least 95°C. The Tm can be measured experimentally using known techniques, including the use of DNA microarrays, or can be calculated using publicly available Tm calculators, such as those available on the Internet.

[0076] The 3' end of the adapter typically comprises, or consists of, a sequence that is at least about 80% identical to the reverse complement of the sequence of the first portion of the overhang strand. The 3' end of the adapter preferably comprises, or consists of, a sequence that is at least about 85% identical or at least about 85.7% identical to the reverse complement of the sequence of the first portion of the overhang strand. The 3' end of the adapter more preferably comprises, or consists of, a sequence that is at least about 90%, at least about 95%, at least about 98%, or at least about 99% identical to the reverse complement of the sequence of the first portion of the overhang strand. The 3' end of the adapter most preferably comprises, or consists of, a sequence that is the reverse complement of the sequence of the first portion of the overhang DNA strand at the end of the telomere. Complementary means that the 3' end of the adapter comprises, or consists of, a sequence that is 100% identical to the reverse complement of the sequence of the first portion of the overhang strand. Complementarity is typically determined using standard Watson-Crick base pairing. In the most preferred embodiment, the 3' end is single-stranded DNA and comprises, or consists of, a sequence that is the reverse complement of the sequence of the first portion of the overhang DNA strand at the end of the telomere.

[0077] Human telomeres typically contain multiple repeats of the sequence TTAGGG. The 3’ end of the telomere adapter preferably specifically hybridizes to one of six possible sequences of the first part of the overhang strand or contains at least one sequence that is the reverse complement thereof. The six possible sequences in the 5’ to 3’ direction are typically TTAGGG, TAGGGT, AGGGTT, GGGTTA, GGTTAG, and GTTAGG. The 3’ end of the telomere adapter preferably contains CCCTAA, ACCCTA, AACCCT, TAACCC, CTAACC, or CCTAAC. The 3’ end of the telomere adapter preferably contains two or more consecutive repeats of CCCTAA, ACCCTA, AACCCT, TAACCC, CTAACC, or CCTAAC. The 3’ end of the telomere adapter may contain any number of consecutive repeats of CCCTAA, ACCCTA, AACCCT, TAACCC, CTAACC, or CCTAAC, for example, three or more, four or more, or five or more. The 3’ end preferably contains or consists of three consecutive repeats of CCCTAA, ACCCTA, AACCCT, TAACCC, CTAACC, or CCTAAC.

[0078] The method can include the use of two or more polynucleotide telomere adapters. Any number, for example, about 2 or more, about 3 or more, about 4 or more, about 5 or more, about 6 or more, or about 10 or more telomere adapters may be used. The method preferably includes, prior to step (a), contacting the telomeres with a population of six telomere adapters, each having a 3′ end that specifically hybridizes to or includes the reverse complement of one of the six possible sequences of the first portion of the overhang strand and a 5′ end that does not hybridize to the opposite portion of the overhang strand. The population includes sequences that specifically hybridize to all of the six possible sequences or include sequences that are the reverse complements thereof. The six possible sequences in the 5′ to 3′ direction are typically TTAGGG, TAGGGT, AGGGTT, GGGTTA, GGTTAG, and GTTAGG. The population of six telomere adapters preferably each specifically hybridizes to one of the six possible sequences of the first portion of the overhang strand (i.e., one of CCCTAA, ACCCTA, AACCCT, TAACCC, CTAACC, and CCTAAC) or has a 3′ end that is the reverse complement thereof. The population typically includes adapters that include sequences that specifically hybridize to all six possible sequences or all six possible reverse complement sequences. The 3′ end of each adapter can have two or more, for example, three or more, four or more, or five or more consecutive repeats of one of the six possible reverse complements (i.e., CCCTAA, ACCCTA, AACCCT, TAACCC, CTAACC, or CCTAAC). The 3′ end of each adapter preferably has three consecutive repeats of CCCTAA, ACCCTA, AACCCT, TAACCC, CTAACC, or CCTAAC. All of the adapters in the population can have a plurality of consecutive repeats of one of the six possible reverse complements, preferably having the same number of repeats.

[0079] The population preferably comprises six telomere adapters, the 3' end of one adapter comprises or consists of CCCTAA, the 3' end of one adapter comprises or consists of ACCCTA, the 3' end of one adapter comprises or consists of AACCCT, the 3' end of one adapter comprises or consists of TAACCC, the 3' end of one adapter comprises or consists of CTAACC, the 3' end of one adapter comprises or consists of CCTAAC. The population preferably comprises six telomere adapters, the 3' end of one adapter comprises or consists of two or more, e.g., three consecutive repeats of CCCTAA, the 3' end of one adapter comprises or consists of two or more, e.g., three consecutive repeats of ACCCTA, the 3' end of one adapter comprises or consists of two or more, e.g., three consecutive repeats of AACCCT, the 3' end of one adapter comprises or consists of two or more, e.g., three consecutive repeats of TAACCC, the 3' end of one adapter comprises or consists of two or more, e.g., three consecutive repeats of, e.g., CTAACC, the 3' end of one adapter comprises or consists of two or more, e.g., three consecutive repeats of CCTAAC. The population preferably comprises six telomere adapters, the 3' end of one adapter comprises or consists of ACCCTAA, the 3' end of one adapter comprises or consists of AACCCTA, the 3' end of one adapter comprises or consists of TAACCCT, the 3' end of one adapter comprises or consists of CTAACCC, the 3' end of one adapter comprises or consists of CTAACC, the 3' end of one adapter comprises or consists of CCCTAAC. All of the specific sequences are given in the 5' to 3' direction.

[0080] The 5' end(s) of the polynucleotide telomere adapter(s) do(es) not hybridize to the opposite portion of the overhang strand. This means that the 5' end(s) of the adapter(s) do(es) not form a duplex with the opposite portion of the overhang strand and can be freely characterized or modified as will be discussed in more detail below. The lack of hybridization can be measured as discussed above. The 5' end can be any portion of the polynucleotide telomere adapter, for example, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 85%. The 5' end of the adapter can be of any length, for example, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, or at least about 30 nucleotides in length.

[0081] The 5' end of the adapter is preferably 60 nucleotides or less, for example, about 50 nucleotides or less, about 40 nucleotides or less, about 30 nucleotides or less, or about 25 nucleotides or less. The 5' end is preferably about 15 to about 35 nucleotides in length, for example, about 18 to about 30 or about 20 to about 25 nucleotides in length. For example, the 5' end can be about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, or about 28 nucleotides in length. The 5' end of the polynucleotide telomere adapter is most preferably about 23 nucleotides in length.

[0082] One skilled in the art can design a 5' end that sufficiently mismatches with the opposite part of the overhang strand so as not to hybridize. The 5' end of the adapter typically includes, or consists of, a sequence that is less than about 20% identical to the reverse complementary sequence of the sequence of the opposite part of the overhang strand. The 5' end of the adapter preferably includes, or consists of, a sequence that is less than about 15%, less than about 10%, less than about 5%, less than about 2%, or less than about 1% identical to the reverse complementary sequence of the sequence of the opposite part of the overhang strand. The 5' end of the adapter most preferably includes, or consists of, a sequence that is 0% identical (i.e., non-complementary) to the reverse complementary sequence of the sequence of the opposite part of the overhang strand. Complementarity or its absence is typically determined using standard Watson-Crick base pairing.

[0083] For human telomeres, the 5' end does not contain the CCCTAA repeat. A preferred 5' end includes, or consists of, AGCAATACGTAACTGAACGAAGT (SEQ ID NO: 7). This is in the 5' to 3' direction. In embodiments where two or more polynucleotide telomere adapters are used, the adapters can have the same 5' end or different 5' ends. The nucleotides at the 5' end of the telomere adapter are preferably modified, for example, with a phosphate group.

[0084] The telomere adapter preferably includes, or consists of, a sequence shown in any one of SEQ ID NOs: 1-6. These sequences are shown in Table 3 and FIG. 1 below.

[0085] The method includes contacting the telomere with a population of six telomere adapters that include, or consist of, the following sequences prior to step (a): AGCAATACGTAACTGAACGAAGTACCCTAA (SEQ ID NO: 1), AGCAATACGTAACTGAACGAAGTAACCCTA (SEQ ID NO: 2), AGCAATACGTAACTGAACGAAGTTAACCCT (SEQ ID NO: 3), AGCAATACGTAACTGAACGAAGTCTAACCC (SEQ ID NO: 4), AGCAATACGTAACTGAACGAAGTCCTAACC (SEQ ID NO: 5), and AGCAATACGTAACTGAACGAAGTCCCTAAC (SEQ ID NO: 6).

[0086] These are shown in Table 3 and FIG. 1 of Example 1.

[0087] The telomere adapter preferably comprises, or consists of, a sequence shown in any one of SEQ ID NOs: 14 to 19. These sequences are shown below and in Table 14.

[0088] This method includes contacting the telomere with a population of six telomere adapters comprising, or consisting of, the following sequences before step (a): AGCAATACGTAACTGAACGAAGTCCCTAACCCTAACCCTAA (SEQ ID NO: 14), AGCAATACGTAACTGAACGAAGTTAACCCTAACCCTAACCC (SEQ ID NO: 15), AGCAATACGTAACTGAACGAAGTCTAACCCTAACCCTAACC (SEQ ID NO: 16), AGCAATACGTAACTGAACGAAGTCCTAACCCTAACCCTAAC (SEQ ID NO: 17), AGCAATACGTAACTGAACGAAGTAACCCTAACCCTAACCCT (SEQ ID NO: 18), and AGCAATACGTAACTGAACGAAGTACCCTAACCCTAACCCTA (SEQ ID NO: 19).

[0089] These are shown in Table 14.

[0090] Any of the telomere adapters used in the present invention may contain click chemistry groups for promoting covalent bonding, as discussed below.

[0091] Sprint polynucleotide Step (a) preferably further comprises hybridizing a sprint polynucleotide to the 5' end of the telomere adapter. This is possible because the 5' end of the telomere adapter is not hybridized to the overhang strand of the telomere end. The sprint polynucleotide may be any of the polynucleotides discussed above in connection with the telomere adapter. The sprint polynucleotide may be any of the lengths discussed above in connection with the 5' of the telomere adapter. The sprint polynucleotide may be the same length as the 5' end or a different length. All or part of the sprint polynucleotide may hybridize to all or part of the 5' end.

[0092] The sprint polynucleotide typically hybridizes specifically to the 5' end of the telomere adapter. Specific hybridization has been discussed above. The sprint polynucleotide typically comprises or consists of a sequence that is at least 80% identical to a partial or reverse complementary sequence of a portion of the sequence of the 5' end of the telomere adapter. The sprint polynucleotide preferably comprises or consists of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% identical to a partial or reverse complementary sequence of a portion of the 5' end of the telomere adapter. The sprint polynucleotide preferably comprises or consists of a sequence that is complementary to a partial or reverse complementary sequence of a portion of the 5' end of the telomere adapter. Complementary means that the sprint polynucleotide comprises or consists of a sequence that is 100% identical to a partial or reverse complementary sequence of a portion of the 5' end of the telomere adapter. Complementarity is typically determined using standard Watson-Crick base pairing. The partial or portion of the 5' end can be of any length, for example, at least about 5, at least about 10, at least about 15, or at least about 20 nucleotides in length.

[0093] The splint polynucleotide preferably comprises, or consists of, a sequence that specifically hybridizes to a part, a portion, or all of AGCAATACGTAACTGAACGAAGT (SEQ ID NO: 7). A preferred splint polynucleotide comprises, or consists of, ACTTCGTTCAGTTACGTATTGCT (SEQ ID NO: 8). This is in the 5' to 3' direction. SEQ ID NO: 8 is the reverse complement of SEQ ID NO: 7.

[0094] The splint polynucleotide is preferably compatible with a sequencing adapter. The splint polynucleotide is typically adapted to form a 3' overhang when hybridized to the 5' end of a telomere adapter. This example is shown in Figure 2. The overhang can be of any length, for example, at least about 1 nucleotide, at least about 2 nucleotides, at least about 3 nucleotides, at least about 4 nucleotides, at least about 5 nucleotides, or at least about 6 nucleotides in length. The 3' overhang preferably hybridizes specifically to a part of the sequencing adapter, for example, an overhang on the sequencing adapter. Specific hybridization is discussed above. The 3' overhang is preferably complementary to a part of the sequencing adapter, for example, an overhang on the sequencing adapter. Suitable sequencing adapters are discussed in more detail below.

[0095] The splint polynucleotide preferably comprises, or consists of, ACTTCGTTCAGTTACGTATTGCTAGCAAT (SEQ ID NO: 9) or ACTTCGTTCAGTTACGTATTGCTA (SEQ ID NO: 10). These are shown in Table 5 and specifically hybridize to the 5' end having the sequence shown in SEQ ID NO: 7. The overhang on the sequencing adapter preferably contains ATTGCT.

[0096] The sprint polynucleotide preferably comprises or consists of ACTTCGTTCGGTGACTTGAGGACAGCAAT (SEQ ID NO: 20), which is shown in Table 15. The overhang on the sequencing adapter preferably comprises ATTGCT.

[0097] Sequencing adapter Step (a) preferably further comprises attaching the sequencing adapter to the telomere adapter and, if present, to the sprint polynucleotide. Step (b) preferably comprises characterizing at least a portion of the ligated non-overhang strand of the telomere in the 5' to 3' direction using the sequencing adapter. Step (a) preferably further comprises attaching the sequencing adapter to the telomere adapter and, if present, to the sprint polynucleotide. Step (b) preferably comprises characterizing at least a portion of the ligated non-overhang strand of the telomere in the 5' to 3' direction using the sequencing adapter. Step (a) preferably further comprises attaching the sequencing adapter to the telomere adapter and, if present, to the sprint polynucleotide. Step (b) preferably comprises characterizing at least a portion of the ligated non-overhang strand of the telomere in the 5' to 3' direction using the sequencing adapter. In all cases, the attachment is preferably a covalent bond. The sequencing adapter can be ligated or annealed to the telomere adapter and, if present, to the sprint polynucleotide. The sequencing adapter is preferably attached to the telomere adapter by ligation and, if present, to the sprint polynucleotide. The sequencing adapter may be attached to the telomere adapter using click chemistry. Suitable click chemistry groups are considered in more detail below.

[0098] Step (a) more preferably further comprises specifically hybridizing the 3' overhang formed by the sprint polynucleotide hybridized to the telomere adapter with the overhang on the sequencing adapter, and attaching, preferably covalently, the sequencing adapter to the telomere adapter and the sprint polynucleotide. This example is shown in FIG. 3. The sequencing adapter can be ligated or annealed to the telomere adapter and the sprint polynucleotide. The sequencing adapter is preferably attached to the telomere adapter and the sprint polynucleotide by ligation.

[0099] The sequencing adapter typically comprises a polynucleotide strand capable of binding to the ends of the target polynucleotide. The target polynucleotide is typically intended for characterization by the methods disclosed herein and comprises a telomere adapter, a sprint polynucleotide, and / or a polynucleotide extension.

[0100] The sequencing adapter can be added to both ends of the target polynucleotide. Alternatively, different adapters can be added to those two ends of the target polynucleotide. The adapter can be added to only one end of the target polynucleotide. Methods for adding an adapter to a polynucleotide are known in the art. The adapter can be attached to the polynucleotide, for example, by ligation, by click chemistry, by tagmentation, by topoisomerase conversion, or by any other suitable method.

[0101] The adapter may be a composite or an artifact. Typically, the adapter comprises the polymers described herein. The adapter preferably comprises a polynucleotide. The adapter may comprise a single-stranded polynucleotide chain. The adapter may comprise a double-stranded polynucleotide. The sequencing adapter may comprise any of the polynucleotides discussed above in connection with the telomere adapter and may include DNA, RNA, modified DNA (such as methylated DNA), RNA, PNA, LNA, BNA, and / or PEG. Usually, the adapter comprises single-stranded and / or double-stranded DNA or RNA.

[0102] The sequencing adapter may be a Y adapter. The Y adapter is typically double-stranded and comprises (a) a region where two strands hybridize together at one end and (b) a region where the two strands are not complementary at the other end. The non-complementary portions of those strands form an overhang. The hybridized stem of the adapter typically attaches to the 5' end of the first strand of the double-stranded polynucleotide and the 3' end of the second strand of the double-stranded polynucleotide, or to the 3' end of the first strand of the double-stranded polynucleotide and the 5' end of the second strand of the double-stranded polynucleotide. The presence of the non-complementary region in the Y adapter gives the adapter a Y shape because, unlike the double-stranded portion, the two strands typically do not hybridize to each other. The hybridized stem ends of the Y adapter may also include short overhangs that specifically hybridize to and attach to the telomere adapter, the splint polynucleotide, or the polynucleotide extension.

[0103] Some of the methods of the present invention use polynucleotide-binding proteins to control the movement of telomeric strands relative to nanopores. The polynucleotide-binding proteins can bind to the overhangs of adapters such as Y-adapters. The polynucleotide-binding proteins can bind to double-stranded regions. The polynucleotide-binding proteins can bind to single-stranded and / or double-stranded regions of the adapter. A first polynucleotide-binding protein can bind to the single-stranded region of such an adapter, and a second polynucleotide-binding protein can bind to the double-stranded region of the adapter.

[0104] The sequencing adapter preferably includes a membrane anchor or a pore anchor. The anchor is complementary to the overhang to which the polynucleotide-binding protein binds and can thus attach to the polynucleotide hybridized thereto.

[0105] One of the non-complementary strands of a sequencing adapter such as a Y-adapter can include a leader sequence that can enter the nanopore when contacting the nanopore.

[0106] The leader sequence includes a polymer such as a polynucleotide, for example, DNA or RNA, a modified polynucleotide (such as abasic DNA), PNA, LNA, polyethylene glycol (PEG), or a polypeptide. The leader sequence preferably includes a single strand of DNA, for example, a poly dT section. The leader sequence can be of any length, but is typically 10 to 150 nucleotides in length, for example, 20 to 120, 30 to 100, 40 to 80, or 50 to 70 nucleotides in length.

[0107] The array determination adapter can be a hairpin loop adapter. A hairpin loop adapter is an adapter that includes a single polynucleotide strand, where the ends of the polynucleotide strand can hybridize to each other or are hybridized to each other, and the central section of the polynucleotide forms a loop. Suitable hairpin loop adapters can be designed using methods known in the art. Typically, the 3' end of the hairpin loop adapter binds to the 5' end of the first strand of the double-stranded polynucleotide, and the 5' end of the hairpin loop adapter binds to the 3' end of the second strand of the double-stranded polynucleotide, or the 5' end of the hairpin loop adapter binds to the 3' end of the first strand of the double-stranded polynucleotide, and the 3' end of the hairpin loop adapter binds to the 5' end of the second strand of the double-stranded polynucleotide. As described in more detail below, the sequencing adapter can be attached to the target polynucleotide to characterize the target polynucleotide.

[0108] One of ordinary skill in the art will also understand that when the adapter includes a polynucleotide strand, the sequence of the adapter is typically not critical and can be controlled or selected according to other experimental conditions, such as polynucleotide-binding proteins and any polynucleotide to be characterized. Exemplary sequences are provided only as examples in the examples. For example, the adapter can include one or more of SEQ ID NOs: 21-26 or 28-33 in WO2021 / 255476 (which is incorporated herein by reference in its entirety), or a polynucleotide sequence having at least 20%, for example, at least 30% (e.g., at least 40%, e.g., at least 50%, e.g., at least 60%, e.g., at least 70%, e.g., at least 80%, e.g., at least 90%, at least 95%) sequence similarity or identity with one or more of SEQ ID NOs: 21-26 or 28-33 in WO2021 / 255476 (which is incorporated herein by reference in its entirety). The sequence of the adapter can typically be changed without adversely affecting the effectiveness of the methods of the invention.

[0109] The array determination adapter may include a loading site for loading a polynucleotide-binding protein. The loading site can be, for example, a single-stranded region that can be targeted by a polynucleotide-binding protein. The loading site is a region of the array determination adapter to which an exogenous polynucleotide strand containing a polynucleotide-binding protein can bind in order to translocate the polynucleotide-binding protein onto the polynucleotide to be evaluated by the method of the present invention.

[0110] The polynucleotide-binding protein, if present, can be provided on the array determination adapter. WO2015 / 110813 and WO2020 / 234612 describe the loading of polynucleotide-binding proteins onto target polynucleotides such as adapters, and the entire contents of which are incorporated herein by reference.

[0111] Spacer The polynucleotides such as telomere adapters, array determination adapters, sprint polynucleotides, or polynucleotide extensions used in the present invention may include one or more spacers, for example, about 1 to about 10 spacers, for example, about 1 to about 5 spacers, for example, about 1, 2, 3, 4, or 5 spacers. The spacer can include any suitable number of spacer units. The spacer typically provides an energy barrier that hinders the movement of the polynucleotide-binding protein. For example, the spacer can hinder the movement of the polynucleotide-binding protein by reducing the pulling force of the protein, for example, using a deoxyribose spacer. The spacer can physically block the movement of the polynucleotide-binding protein by introducing, for example, a bulky chemical group to physically block the movement of the protein.

[0112] One or more spacers are included in the polynucleotide or sequencing adapter to provide a unique signal when they pass through the nanopore. One or more spacers can be used to define or separate one or more regions of the polynucleotide, for example, the adapter can be separated from the target polynucleotide.

[0113] The spacer can include a linear molecule such as a polymer that is, for example, a polypeptide or polyethylene glycol (PEG). Typically, such a spacer has a structure different from that of the target polynucleotide. For example, if the target polynucleotide is DNA, the spacer or each spacer typically does not contain DNA. Specifically, when the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the spacer or each spacer preferably includes a peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), or a synthetic polymer having a nucleotide side chain. The spacer can be one or more nitroindoles, one or more inosines, one or more acridines, one or more 2-aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridine, one or more inverted thymidine (inverted dT), one or more inverted dideoxy-thymidine (ddT), one or more dideoxy-cytidine (ddC), one or more 5-methylcytidine, one or more 5-hydroxymethylcytidine, one or more 2'-O-methyl RNA bases, one or more iso-deoxycytidine (Iso-dC), one or more iso-deoxyguanosine (Iso-dG), one or more C3(OC 3 H 6 OPO 3 ) group, one or more photocleavable (PC)[OC 3 H 6 -C(O)NHCH 2 -C 6 H 3 NO 2 -CH(CH 3 )OPO 3 group, one or more hexanediol groups, one or more spacer 9 (iSp9)[(OCH 2 CH 2 )3 OPO 3 group, or one or more spacers 18 (iSp18) [(OCH 2 CH 2 ) 6 OPO 3 group, or may contain one or more thiol bonds. The spacer may contain any combination of these groups. Many of these groups are commercially available from IDT (registered trademark) (Integrated DNA Technologies (registered trademark)). For example, C3, iSp9, and iSp18 spacers are all available from IDT (registered trademark). The spacer may contain any number of the above groups as spacer units.

[0114] The spacer may contain one or more chemical groups, for example, one or more pendant chemical groups. One or more chemical groups may be attached to one or more nucleobases in the sequencing adapter. One or more chemical groups may be attached to the backbone of the sequencing adapter. Any number of suitable chemical groups may be present, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxigenin and / or anti-digoxigenin, and dibenzylcyclooctyne groups.

[0115] The spacer can include one or more abasic nucleotides (i.e., nucleotides lacking a nucleobase), for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more abasic nucleotides. The nucleobase can be replaced by -H (idSp) or -OH in the abasic nucleotide. The abasic spacer can be inserted into the target polynucleotide by removing the nucleobase from one or more adjacent nucleotides. For example, the polynucleotide can be modified to include 3-methyladenine, 7-methylguanine, 1,N6-ethenoadenine inosine, or hypoxanthine, and the nucleobase can be removed from these nucleotides using human alkyladenine DNA glycosylase (hAAG). Alternatively, the polynucleotide can be modified to include uracil, and the nucleobase can be removed by uracil DNA glycosylase (UDG). One or more spacers do not contain any abasic nucleotides.

[0116] Suitable spacers can be designed or selected according to the nature of the polynucleotide or sequencing adapter, the polynucleotide-binding protein, and the conditions under which the method is to be performed.

[0117] Tag Polynucleotides such as the telomere adapter, sequencing adapter, sprint polynucleotide, or polynucleotide extension used in the present invention can include a tag or tether. For example, the polynucleotide can bind to a tag on the nanopore via its adapter and can be released, for example, at some point during the characterization of the polynucleotide by the nanopore. Strong non-covalent bonds (e.g., biotin / avidin) are still reversible and can be useful in some embodiments of the methods described herein.

[0118] A pair of a pore tag and an array determination adapter may be configured such that the binding strength or affinity of a binding site on a polynucleotide to a tag on a nanopore (e.g., a binding site provided by an anchor or leader sequence of the adapter or by a capture sequence within a double-stranded stem of the adapter) is sufficient to maintain the coupling between the nanopore and the polynucleotide until an additional force is applied to release the bound polynucleotide from the nanopore.

[0119] The tag or tether is preferably uncharged. This can ensure that the tag or tether is not drawn into the nanopore under the influence of a potential difference.

[0120] One or more molecules that attract or bind to a polynucleotide or an adapter may be linked to a detector (e.g., a pore). Any molecule that hybridizes to the adapter and / or the target polynucleotide may be used. The molecule attached to the pore may be selected from PNA tags, PEG linkers, short oligonucleotides, positively charged amino acids, and aptamers. Pores having such molecules attached thereto are known in the art. For example, pores with short oligonucleotides attached are disclosed in Howarka et al (2001) Nature Biotech. 19: 636-639 and WO2010 / 086620, and pores containing PEG attached within the lumen of the pore are disclosed in Howarka et al (2000) J. Am. Chem. Soc. 122(11):2411-2416.

[0121] The capture of a target polynucleotide in the methods described herein may be enhanced using an oligonucleotide that is a short oligonucleotide attached to a detector (e.g., a nanopore) and that contains a sequence complementary to the sequence of a leader sequence or another single-stranded sequence of an adapter.

[0122] The tag or tether may include or be an oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino). The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) can have a length of about 10 - 30 nucleotides or about 10 - 20 nucleotides. The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) for use as a tag or tether can have at least one end (e.g., 3´- or 5´-end) modified for attachment to other modification sites or to the surface of a solid substrate, such as a bead. The end modifier may add a reactive functional group that can be used for attachment. Examples of functional groups that can be added include, but are not limited to, amino, carboxyl, thiol, maleimide, aminooxy, and any combination thereof. The functional group can be combined with spacers of different lengths (e.g., C3, C9, C12, spacer 9 and 18) to add a physical distance of the functional group from the end of the oligonucleotide sequence.

[0123] The tag or tether may include or be a morpholino oligonucleotide. The morpholino oligonucleotide can have a length of about 10 - 30 nucleotides or about 10 - 20 nucleotides. The morpholino oligonucleotide may or may not be modified. For example, the morpholino oligonucleotide may be modified at the 3´ and / or 5´ ends of the oligonucleotide. Examples of 3´ and / or 5´ end modifications of the morpholino oligonucleotide include 3´ affinity tags and functional groups for chemical attachment (e.g., including 3´-biotin, 3´-primary amine, 3´-disulfide amide, 3´-pyridyldithio, and any combination thereof); 5´ end modifications (e.g., including 5´-primary amine, and / or 5´-dabsyl); modifications for click chemistry (e.g., including 3´-azide, 3´-alkyne, 5´-azide, 5´-alkyne), and any combination thereof, but are not limited thereto.

[0124] The tag or tether may further include a polymer linker, for example, to facilitate binding to a detector, such as a nanopore. Exemplary polymer linkers include, but are not limited to, polyethylene glycol (PEG). The polymer linker may have a molecular weight of about 500 Da to about 10 kDa (including both ends), or about 1 kDa to about 5 kDa (including both ends). The polymer linker (e.g., PEG) can be functionalized with different functional groups including, but not limited to, maleimide, NHS ester, dibenzocyclooctyne (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combination thereof. The tag or tether may also include 1 kDa PEG having a 5'-maleimide group and a 3'-DBCO group. The tag or tether may further include 2 kDa PEG having a 5'-maleimide group and a 3'-DBCO group. The tag or tether may further include 3 kDa PEG having a 5'-maleimide group and a 3'-DBCO group. The tag or tether may further include 3 kDa PEG having a 5'-maleimide group and a 5'-DBCO group.

[0125] Other examples of tags or tethers include, but are not limited to, His tags, biotin or streptavidin, antibodies that bind to the analyte, aptamers that bind to the analyte, analyte binding domains such as DNA binding domains (e.g., peptide zippers such as leucine zippers, single-stranded DNA binding protein (SSB)), and any combination thereof.

[0126] Any method known in the art can be used to attach a tag or tether to the outer surface of the nanopore, e.g., on the cis side of the membrane. For example, one or more tags or tethers can be attached to the nanopore via one or more cysteines (cysteine bonding), primary amines such as one or more lysines, one or more non-natural amino acids, one or more histidines (His tags), one or more biotins or streptavidins, one or more antibody-based tags, one or more enzymatic modifications of epitopes (e.g., including acetyltransferase), and any combination thereof. Suitable methods for performing such modifications are well known in the art. Suitable non-natural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz), and any one of the amino acids numbered 1 to 71 in FIG. 1 of Liu C.C. and Schultz P.G., Annu. Rev. Biochem., 2010, 79, 413-444, which is incorporated herein by reference in its entirety.

[0127] When one or more tags or tethers are attached to the nanopore via cysteine binding(s), one or more cysteines can be introduced by substitution into one or more monomers forming the nanopore. The nanopore can be chemically modified by the following attachments: (i) 4-phenylazomaleinanil, 1.N-(2-hydroxyethyl)maleimide, N-cyclohexylmaleimide, 1.3-maleimidopropionic acid, 1.1-4-aminophenyl-1H-pyrrole,2,5,dione, 1.1-4-hydroxyphenyl-1H-pyrrole,2,5,dione, N-ethylmaleimide, N-methoxycarbonylmaleimide, N-tert-butylmaleimide, N-(2-aminoethyl)maleimide, 3-maleimide-proxyl, N-(4-chlorophenyl)maleimide, 1-[4-(dimethylamino)-3,5-dinitrophenyl]-1H-pyrrole-2,5-dione, N-[4-(2-benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(1-naphthyl)-maleimide, N-(2,4-xylyl)maleimide, N-(2,4-difluorophenyl)maleimide, N-(3-chloro-para-tolyl)-maleimide, 1-(2-amino-ethyl)-pyrrole-2,5-dione hydrochloride, 1-cyclopentyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(3-aminopropyl)-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 3-methyl-1-[2-oxo-2-(piperazin-1-yl)ethyl]-2,5-dihydro-1H-pyrrole-2,5-dione hydrochloride, 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-dione, 3-methyl-1-(3,3,3-trifluoropropyl)-2,5-dihydro-1H-pyrrole-2,5-dione, 1-[4-(methylamino)cyclohexyl]-2,5-dihydro-1H-pyrrole-2,5-dione trifluoroacetic acid, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2, 1-benzyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-dione, 1-(2-fluorophenyl)-3-methyl-2,5-dihydro 1H-pyrrole-2,Maleimides containing dibromomaleimide such as 5-dion, N-(4-phenoxyphenyl)maleimide, N-(4-nitrophenyl)maleimide, etc., (ii) iodoacetamides such as 3-(2-iodoacetamido)-proxyl, N-(cyclopropylmethyl)-2-iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2-trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, N-(4-(aminosulfonyl)phenyl)-2-iodoacetamide, N-(1,3-benzothiazol-2-yl)-2-iodoacetamide, N-(2,6-(diethylphenyl)-2-iodoacetamide, N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide, etc., (iii) bromoacetamides such as N-(4-(acetylamino)phenyl)-2-bromoacetamide, N-(2-acetylphenyl)-2-bromoacetamide, 2-bromo-n-(2-cyanophenyl)acetamide, 2-bromo-N-(3-(trifluoromethyl)phenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide, 2-bromo-N-(4-fluorophenyl)-3-methylbutanamide, N-benzyl-2-bromo-N-phenylpropionamide, N-(2-bromo-butyryl)-4-chloro-benzenesulfonamide, 2-bromo-N-methyl-N-phenylacetamide, 2-bromo-N-phenethyl-acetamide, 2-adamantan-1-yl-2-bromo-N-cyclohexyl-acetamide, 2-bromo-N-(2-methylphenyl)butanamide, monobromoacetanilide, etc., (iv) disulfides such as aldritol-2, aldritol-4, isopropyldisulfide, 1-(isobutyldisulfanyl)-2-methylpropane, dibenzyldisulfide, 4-aminophenyldisulfide, 3-(2-pyridyldithio)propionic acid, 3-(2-pyridyldithio)propionic acid hydrazide, 3-(2-pyridyldithio)propionic acid N-succinimidyl ester, am6amPDP1-βCD, etc., and (v) thiols such as 4-phenylthiazole-2-thiol, Purpald, 5,6,7,8-tetrahydro-quinazoline-2-thiol, etc.,

[0128] The tag or tether may be attached directly to the nanopore or via one or more linkers. The tag or tether may be attached to the nanopore using a hybridization linker as described in WO2010 / 086602, which is hereby incorporated by reference in its entirety. Alternatively, a peptide linker may be used. The peptide linker is an amino acid sequence. The length, flexibility, and hydrophilicity of the peptide linker are typically designed so as not to interfere with the function of the monomer and the pore. Preferred flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10, or 16, serine, and / or glycine amino acids. A more preferred flexible linker is (SG) 1 , (SG) 2 , (SG) 3 , (SG) 4 , (SG) 5 , and (SG) 8 , where S is serine and G is glycine. Preferred rigid linkers are stretches of 2 to 30, such as 4, 6, 8, 16, or 24, proline amino acids. A more preferred rigid linker is (P) 12 , where P is proline.

[0129] Suitable pore tags are also described in WO2018 / 100370, which describes a non-hairpin method for characterizing double-stranded polynucleotides and is hereby incorporated by reference in its entirety.

[0130] Anchor Polynucleotides such as the telomere adapters, sequencing adapters, sprint polynucleotides, or polynucleotide extensions used in the present invention may include a membrane anchor. The anchor typically assists in the characterization of the target polynucleotide according to the methods disclosed herein. For example, the membrane anchor may facilitate the localization of the selected polynucleotide around the nanopore.

[0131] The anchor can be a polypeptide anchor and / or a hydrophobic anchor that can be inserted into the membrane. The hydrophobic anchor is preferably a lipid, fatty acid, sterol, carbon nanotube, polypeptide, protein, or amino acid, such as cholesterol, palmitic acid, or tocopherol. The anchor can include a thiol, biotin, or surfactant.

[0132] The anchor can be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or a fusion protein), Ni-NTA (for binding to poly-histidine or a protein with a poly-histidine tag), or a peptide (such as an antigen).

[0133] The anchor preferably can include one linker, or two, three, four or more linkers. Preferred linkers include, but are not limited to, polymers such as polynucleotides, polyethylene glycol (PEG), polysaccharides, and polypeptides. These linkers can be linear, branched, or cyclic. For example, the linker can be a cyclic polynucleotide. The adapter can hybridize to a complementary sequence on the cyclic polynucleotide linker. One or more anchors or one or more linkers can include a component that can be cleaved or degraded, such as a restriction site or a photocleavable group. The linker is functionalized with a maleimide group to attach to a cysteine residue of a protein. Suitable linkers are described in WO2010 / 086602, which is hereby incorporated by reference in its entirety.

[0134] The anchor is preferably cholesterol or a fatty acyl chain. For example, any fatty acyl chain having 6 to 30 carbon atoms in length, such as hexadecanoic acid, can be used.

[0135] Examples of suitable anchors and methods for attaching an anchor to an adapter are disclosed in WO2012 / 164270 and WO2015 / 150786, which are hereby incorporated by reference in their entireties.

[0136] The anchor may consist of, or may include, a hydrophobic modification to a polynucleotide or a sequencing adapter. The hydrophobic modification may include a modified phosphate group contained within the polynucleotide or the polynucleotide anchor. The hydrophobic modification may include, for example, phosphorothioates such as charge-neutralized alkyl phosphorothioates (PPTs), which are described in Jones et al, J. Am. Chem. Soc. 2021, 143, 22, 8305, the entire content of which is incorporated herein by reference. Suitable alkyl groups include, for example, C 2 ~C 6 alkyl groups such as C 1 ~C 10 alkyl groups, such as methyl, ethyl, propyl, butyl, pentyl, and hexyl groups. Incorporation of charge-neutralized alkyl-phosphorothioates into a polynucleotide allows the polynucleotide to engage in hydrophobic regions such as lipid bilayers.

[0137] Polynucleotide extension As an alternative to the use of the sprint polynucleotide, step (a) preferably further comprises ligating or covalently attaching a polynucleotide extension to the telomere adapter. The polynucleotide extension is typically ligated or covalently attached to the 5' end of the telomere adapter. The polynucleotide extension may be any of the polynucleotides discussed above in connection with the telomere adapter. The polynucleotide extension may be of any length, including any of the lengths discussed above in connection with the 5' of the telomere adapter. The polynucleotide extension is preferably covalently attached to the telomere adapter using click chemistry. Suitable click chemistries are known in the art and are discussed herein.

[0138] A preferred polynucleotide extension comprises or consists of GCTTGGGTGTTTAACC (SEQ ID NO: 11). This is used in the click adapters of Table 7. The polynucleotide extension may contain or further contain one or more, such as two or more, three or more, or four or more universal nucleotides, such as Int 5-nitroindole (i5NiTInd). One or more universal nucleotides are typically at the 3' end of the polynucleotide extension. The polynucleotide extension can be ligated or covalently attached to the telomere adapter using one or more universal nucleotides.

[0139] The method of the present invention preferably comprises, in step (a), ligating a polynucleotide telomere adapter to the 5' end of the non-overhanging strand at the end of the telomere, wherein the 3' end of the adapter specifically hybridizes to the first portion of the overhanging strand and the 5' end of the adapter does not hybridize to the opposite portion of the overhanging strand, and the telomere adapter contains a polynucleotide extension at the 5' end. Any of the embodiments discussed above in relation to the telomere adapter and the polynucleotide extension are equally applicable to this embodiment of the "extended" telomere adapter.

[0140] The extended telomere adapter preferably comprises or consists of the sequence shown in SEQ ID NO: 11 covalently bound to any one of SEQ ID NOs: 1 to 6. The extended telomere adapter preferably comprises or consists of the sequence shown in SEQ ID NO: 11 covalently bound to any one of SEQ ID NOs: 14 to 19. The two sequences are preferably covalently bound by one or more, for example, two or more, three or more, or four or more universal nucleotides, for example, Int 5-nitroindole (i5NiTInd). The two sequences are more preferably covalently bound by four i5NiTInds. The extended telomere adapter preferably comprises or consists of the sequence shown in any one of SEQ ID NOs: 25 to 30. The extended telomere adapter preferably contains a click chemistry group such as DBCOTEG at the 5' end. Examples of these extended telomere adapters are shown in Table 7.

[0141] This method includes contacting the telomere with a population of six extended telomere adapters comprising or consisting of the following sequences before step (a): SEQ ID NO: 11 covalently bound to AGCAATACGTAACTGAACGAAGTACCCTAA (SEQ ID NO: 1), SEQ ID NO: 11 covalently bound to AGCAATACGTAACTGAACGAAGTAACCCTA (SEQ ID NO: 2), SEQ ID NO: 11 covalently bound to AGCAATACGTAACTGAACGAAGTTAACCCT (SEQ ID NO: 3), SEQ ID NO: 11 covalently bound to AGCAATACGTAACTGAACGAAGTCTAACCC (SEQ ID NO: 4), SEQ ID NO: 11 covalently bound to AGCAATACGTAACTGAACGAAGTCCTAACC (SEQ ID NO: 5), and SEQ ID NO: 11 covalently bound to AGCAATACGTAACTGAACGAAGTCCCTAAC (SEQ ID NO: 6).

[0142] This method includes, prior to step (a), contacting the telomere with a population of six extended telomere adapters comprising or consisting of the following sequences: GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTACCCTAA (SEQ ID NO: 25), GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTAACCCTA (SEQ ID NO: 26), GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTTAACCCT (SEQ ID NO: 27), GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTCTAACCC (SEQ ID NO: 28), GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTCCTAACC (SEQ ID NO: 29), and GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTCCCTAAC (SEQ ID NO: 30) (wherein N is Int 5-nitroindole (i5NiTInd)).

[0143] This method includes, prior to step (a), contacting the telomere with a population of six extended telomere adapters comprising or consisting of the following sequences: SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTCCCTAACCCTAACCCTAA (SEQ ID NO: 14), SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTTAACCCTAACCCTAACCC (SEQ ID NO: 15), SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTCTAACCCTAACCCTAACC (SEQ ID NO: 16), SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTCCTAACCCTAACCCTAAC (SEQ ID NO: 17), SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTAACCCTAACCCTAACCCT (SEQ ID NO: 18), and SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTACCCTAACCCTAACCCTA (SEQ ID NO: 19).

[0144] The two sequences within the extension adapter are preferably covalently linked by one or more, such as two or more, three or more, or four or more universal nucleotides, such as Int 5-nitroindole (i5 NiTInd). The two sequences within the extension adapter are more preferably covalently linked by four i5NiTInds. The extension telomere adapter preferably contains a click chemistry group such as DBCOTEG at the 5' end. Examples of these extension adapters are shown in Table 7 and Figure 4.

[0145] The polynucleotide extension preferably comprises a sequencing adapter, and step (b) preferably comprises characterizing at least a partially ligated non-overhang strand of the telomere in the 5' to 3' direction from the end of the telomere using the sequencing adapter. Step (a) preferably further comprises attaching (e.g., covalently) a sequencing adapter to the polynucleotide extension, and step (b) preferably comprises characterizing at least a partially ligated non-overhang strand of the telomere in the 5' to 3' direction using the sequencing adapter. More preferably, step (a) further comprises specifically hybridizing a 5' overhang on the sequencing adapter to the polynucleotide extension and attaching (e.g., covalently) the sequencing adapter to the polynucleotide extension. In any of these embodiments of the sequencing adapter, the polynucleotide extension may be ligated or covalently attached to the telomere adapter, or the method may use the extended telomere adapter as defined above. The sequencing adapter may be ligated or annealed to the polynucleotide extension. The sequencing adapter is preferably attached to the polynucleotide extension by ligation. The polynucleotide extension is preferably covalently attached to the sequencing adapter using click chemistry. The polynucleotide extension preferably contains a click chemistry group such as DBCOTEG at the 5' end and is attached to the sequencing adapter using click chemistry. Other possible click chemistry groups are considered in more detail below.

[0146] Biotin enrichment Telomere adapter(s) comprising an extended telomere adapter preferably contain biotin. Biotin is preferably at the 5' end of the telomere adapter(s) or extended telomere adapter(s). This example is shown in Figure 6. In such an embodiment, step (a) preferably further comprises enriching at least a portion of the ligated non-overhanging strands of the telomere using biotin. Suitable methods for biotin-based enrichment are known in the art. For example, the telomere adapter(s) preferably contain biotin, and the surface such as beads contains avidin / streptavidin, and this surface can be used to enrich at least a portion of the ligated non-overhanging strands of the telomere.

[0147] Linker or spacer One or more linkers or spacers are preferably present between the 3' end of the telomere adapter and the 5' end of the telomere adapter. One or more linkers or spacers are preferably flexible. The telomere adapter can contain any number of one or more spacers, for example, about 1 to about 10 spacers, for example, about 1 to about 5 spacers, for example, about 1, 2, 3, 4, or 5 spacers. The spacer can contain any suitable number of spacer units. The spacer typically provides an energy barrier that hinders the movement of polynucleotide-binding proteins. For example, the spacer can hinder the movement of polynucleotide-binding proteins by reducing the pulling force of the protein, for example, using a deoxyribose spacer. The spacer can physically block the movement of polynucleotide-binding proteins by, for example, introducing a bulky chemical group to physically hinder the movement of the protein. Suitable spacers and their use are discussed above with reference to sequencing adapters, and these are equally applicable to telomere adapters.

[0148] Bidirectional characterization The method of the present invention preferably further comprises characterizing at least a partially unligated overhang strand of a telomere in the 5' to 3' direction up to the end of the telomere. This example is shown in FIG. 7. Any method can be used to characterize the unligated overhang strand, including any of the methods discussed below. The same or different methods can be used to characterize the ligated non-overhang strand and the unligated overhang stand. It is preferred to use the same method to characterize the ligated non-overhang strand and the unligated overhang stand.

[0149] The method of the present invention preferably uses a polymer-induced effector protein such as Cas9 protein to create a double-strand break from the telomere end to at least a part of the opposite end of the telomere, attach a sequencing adapter to the opposite end, and use the sequencing adapter to characterize at least a partially unligated overhang strand of the telomere in the 5' to 3' direction up to the end of the telomere. The sequencing adapter can be ligated or annealed to the opposite end. The sequencing adapter can be any of those discussed above. Attaching a Y-sequencing adapter to the double-strand break is known in the art.

[0150] The present invention also provides a method for characterizing at least a part of a telomere, comprising: (a) ligating a polynucleotide telomere adapter to the 5'-end of the non-overhang strand at the end of the telomere, wherein the 3'-end of the adapter is complementary to the first part of the overhang strand and the 5'-end of the adapter is not complementary to the opposite part of the overhang strand; (b) using a polymerase-induced effector protein such as Cas9 protein to create a double-strand break from the telomere end to the end opposite to at least a part of the telomere, and attaching a sequencing adapter to the opposite end; and (c) using the telomere adapter to characterize the ligated non-overhang strand of at least a part of the telomere in the 5' to 3' direction from the telomere end, and using the sequencing adapter to characterize the unligated overhang strand of at least a part of the telomere in the 5' to 3' direction up to the telomere end.

[0151] Polymerase-induced effector protein The polymerase-induced effector protein can be any protein that binds to the end opposite to at least a part of the telomere via a guide polymer. The polymerase-induced effector protein can bind to or attach to, as non-limiting examples, guide oligonucleotides such as aptamers, or guide polypeptides such as antibodies that bind to a part of at least a part of the telomere.

[0152] The polymer-induced effector protein is preferably a polynucleotide-induced effector protein. In this embodiment, the guide polymer is preferably a guide polynucleotide. The polynucleotide-induced effector protein can be any protein that binds to or attaches to the guide polynucleotide and binds to the target polynucleotide sequence at the opposite end of at least a part of the telomere, preferably at the opposite end of at least a part of the telomere to which the guide polynucleotide binds. The polynucleotide-induced effector protein preferably comprises a target polynucleotide sequence recognition domain and at least one nuclease domain. The recognition domain binds the guide polynucleotide (e.g., RNA) and the target polynucleotide (e.g., DNA). The polynucleotide-induced effector protein can contain one nuclease domain that cleaves one or both strands of the double-stranded polynucleotide, or can contain two nuclease domains, where the first nuclease domain is arranged for cleavage of one strand of the target polynucleotide sequence and the second nuclease domain is arranged for cleavage of the complementary strand of the target polynucleotide sequence. The nuclease domain can be active or inactive. For example, the nuclease domain, or one or both of the two nuclease domains, can be inactivated by mutation.

[0153] The guide polynucleotide can be a guide RNA, a guide DNA, or a guide containing both DNA and RNA. The guide polynucleotide is preferably a guide RNA. Thus, the polynucleotide-induced effector protein is preferably an RNA-induced effector protein.

[0154] An RNA-guided effector protein can be any protein that binds to or attaches to a guide RNA. An RNA-guided effector protein typically binds to a region of the guide RNA that is not the region of the guide RNA that binds to the target polynucleotide sequence. For example, when the guide RNA includes a crRNA and a tracrRNA, the RNA-guided effector protein typically binds to the tracrRNA, and the crRNA typically binds to at least a portion of the opposite end of the telomere, also known as the target polynucleotide sequence. The RNA-guided effector protein preferably also binds to the target polynucleotide sequence. The region of the guide RNA that binds to the target polynucleotide sequence can also bind to the RNA-guided effector protein. The RNA-guided effector protein typically binds to the double-stranded region of the target polynucleotide sequence. The region of the target polynucleotide sequence to which the RNA-guided effector protein binds is typically located near the sequence to which the guide RNA hybridizes. The guide RNA and the RNA-guided effector protein typically form a complex, which then binds to the target polynucleotide sequence at a site determined by the sequence of the guide RNA.

[0155] The RNA-guided effector protein can bind upstream or downstream of the sequence to which the guide RNA binds. For example, the RNA-guided effector protein can bind to a protospacer adjacent motif (PAM) in the DNA located adjacent to the sequence to which the guide RNA binds. The PAM is a short (less than 10, typically 2-6 base pairs) sequence such as 5'-NGG-3' (where N is any base), 5'-NGA-3', 5'-YG-3' (where Y is a pyrimidine), 5'-TTN-3', or 5'-YTN-3'. Different RNA-guided effector proteins bind to different PAMs. Specifically, the RNA-guided effector protein can bind to a target polynucleotide sequence that does not contain a PAM when the target is RNA or a DNA / RNA hybrid.

[0156] RNA-guided effector proteins are typically nucleases such as RNA-guided endonucleases. RNA-guided effector proteins are typically Cas proteins. RNA-guided effector proteins can be Cas, Csn2, Cpf1, Csf1, Cmr5, Csm2, Csy1, Cse1, or C2c2. Cas proteins can be Cas3, Cas4, Cas8a, Cas8b, Cas8c, Cas9, Cas10, or Cas10d. Preferably, the Cas protein is Cas9. Cas, Csn2, Cpf1, Csf1, Cmr5, Csm2, Csy1, or Cse1 are preferably used when the target polynucleotide sequence contains a double-stranded DNA region. C2c2 is preferably used when the target polynucleotide sequence contains a double-stranded RNA region.

[0157] DNA-guided effector proteins, such as proteins from the RecA family, can be used to target DNA. Examples of proteins from the RecA family that can be used are RecA, RadA, and Rad51. The nuclease activity of an RNA-guided endonuclease can be inactivated. One or more catalytic nuclease sites of an RNA-guided endonuclease can be inactivated. For example, if an RNA-guided endonuclease contains two catalytic nuclease sites, one or both of the catalytic sites can be inactivated. Typically, one of the catalytic sites will cleave one strand of the polynucleotide to which it specifically binds, and the other catalytic site will cleave the opposite strand of the polynucleotide. Thus, an RNA-guided endonuclease may cleave both strands, one strand, or neither strand of the double-stranded region of the target polynucleotide. An RNA-guided endonuclease preferably cleaves both strands of at least a portion of a telomere.

[0158] The polynucleotide-induced effector protein is preferably Cas9. Cas9 has a two-lobe multi-domain protein structure including a target recognition and a nuclease lobe. The recognition lobe binds guide RNA and DNA. The nuclease lobe contains HNH and RuvC nuclease domains that are arranged for cleavage of the complementary and non-complementary strands of the target DNA. The structure of Cas9 is detailed in Nishimasu, H., et al., (2014) Crystal Structure of Cas9 in Complex with Guide RNA and Target DNA. Cell 156, 935-949. The relevant PDB reference for Cas9 is 5F9R (Crystal structure of catalytically active Streptococcus pyogenes CRISPR-Cas9 complexed with a single-guide RNA and double-stranded DNA primed for target DNA cleavage).

[0159] Cas9 can be a "specificity-enhanced" Cas9 that shows reduced off-target binding compared to wild-type Cas9. An example of such a "specificity-enhanced" Cas9 is S. pyogenes Cas9 D10A / H840A / K848A / K1003A / R1060A. ONLP12296 is the amino acid sequence of S. pyogenes Cas9 D10A / H840A / K848A / K1003A / R1060A with a C-terminal Twin-Strep tag together with a TEV-cleavable linker.

[0160] The catalytic site of an RNA-guided endonuclease can be inactivated by mutations. The mutations can be substitution, insertion, or deletion mutations. For example, one or more, such as 2, 3, 4, 5, or 6 amino acids can be substituted or inserted into or deleted from the catalytic site. The mutation is preferably a substitution insertion in the case of a single amino acid at the catalytic site, more preferably a substitution. A person skilled in the art will be able to easily identify the catalytic sites of RNA-guided endonucleases and the mutations that inactivate them. For example, when the RNA-guided endonuclease is Cas9, one catalytic site can be inactivated by a mutation at D10 and the other by a mutation at H640.

[0161] An inactivated ("dead") polynucleotide-guided effector protein that does not cleave the target polynucleotide sequence thus does not exhibit directional bias. An active ("alive") polynucleotide-guided effector protein that cleaves the target polynucleotide sequence can remain bound to only one of the two ends of the cleavage site and thus can exhibit some directional bias.

[0162] The polymer-guided effector protein can be specifically modified for use in the methods of the invention. The protein can include an anchor that can bind to a membrane, such as cholesterol. The polymer-guided effector protein preferably has a binding moiety that can bind to a surface to which it is attached. The surface is preferably the surface of beads.

[0163] Guide polymer The guide polymer can bind to the opposite end of at least a part of the telomere and mediate the binding of the polymer-induced effector protein to the opposite end of at least a part of the telomere. The guide polymer can have any structure that enables it to bind to the opposite end of at least a part of the telomere, which is also known as the target polynucleotide sequence. It can be any of the polymers discussed above in relation to the telomere adapter. The guide polymer is preferably an oligonucleotide, polynucleotide, polypeptide, protein, oligosaccharide, or polysaccharide. The guide polypeptide or protein is preferably a zinc finger binding protein, transcription activator-like effector (TALE), transcription factor, restriction enzyme, DNA binding protein or enzyme, antibody, or antibody fragment. Suitable antibody fragments are known in the art and include, but are not limited to, Fab, Fab’, (Fab’)2, FV, scFv, diabody, triabody, tetrabody, Bis-scFv, minibody, Fab2, and Fab3. The guide oligonucleotide is preferably an aptamer.

[0164] The guide polymer can bind to the polymer-induced effector protein. In this context, the guide polymer can bind to the polymer-induced effector protein, be bound by the polymer-induced effector protein, or both. The guide polymer can attach to the polymer-induced effector protein. For example, suitable methods for attaching the guide polymer to the polymer-induced effector protein using streptavidin / biotin are known in the art. The polymer-induced effector is preferably covalently bound to the guide polymer. For example, covalent bonds using click chemistry are discussed in WO2018 / 060740, which is incorporated herein by reference in its entirety.

[0165] The guide polymer is preferably a guide polynucleotide. In this case, the polymer-induced effector protein is preferably a polynucleotide-induced effector protein. The guide polynucleotide preferably contains a sequence that can bind to at least a part of the opposite end of the telomere or can specifically hybridize to a target polynucleotide sequence. The guide polynucleotide preferably contains a sequence at least a part of the opposite end of the telomere and a nucleotide sequence that binds to the polynucleotide-induced effector protein. The guide polynucleotide preferably contains a nucleotide sequence that specifically hybridizes to a sequence at least a part of the opposite end of the telomere and a nucleotide sequence that binds to the polynucleotide-induced effector protein. The guide polynucleotide can have any structure that allows it to specifically bind / hybridize to at least a part of the opposite end of the telomere and bind to the polynucleotide-induced effector protein.

[0166] The guide polynucleotide typically specifically hybridizes to a sequence of about 20 nucleotides in the target polynucleotide sequence. The sequence to which the guide polynucleotide binds can be from about 10 to about 40, such as from about 15 to about 30, preferably from about 18 to about 25, such as about 19, 20, 21, 22, 23, or 24 nucleotides. The guide polynucleotide is typically complementary to one strand of at least a part of the opposite end of the telomere. The guide polynucleotide preferably contains a nucleotide sequence of from about 10 to about 40, such as from about 15 to about 30, preferably from about 18 to about 25, such as about 19, 20, 21, 22, 23, or 24 nucleotides that is complementary to the sequence of the target polynucleotide sequence or a sequence in the target polynucleotide sequence. The degree of complementarity is preferably exact.

[0167] The guide polynucleotide is preferably a guide RNA. The guide RNA can be complementary to a region in the target polynucleotide sequence that is 5' to the PAM. This is preferred when the target polynucleotide contains DNA, particularly when the RNA effector protein is Cas9 or Cpf1. The guide RNA can be complementary to a region in the target polynucleotide sequence adjacent to guanine. This is preferred when the target polynucleotide contains RNA, particularly when the RNA effector protein is C2c2.

[0168] The guide RNA can have any structure that enables it to bind to the target polynucleotide sequence and the RNA-guided effector protein. The guide RNA can include a crRNA and a tracrRNA that bind to a sequence in the target polynucleotide sequence. The tracrRNA typically binds to the RNA-guided effector protein. Typical structures of guide RNAs are known in the art. For example, the crRNA is typically single-stranded RNA, and the tracrRNA typically has a double-stranded region where the single strand binds to the 3' end of the crRNA and a portion that forms a hairpin loop at the 3' end of the strand that does not bind to the crRNA. The crRNA and tracrRNA can be transcribed in vitro as a single-piece sgRNA. The guide RNA is preferably an sgRNA.

[0169] The guide RNA can contain other components such as additional RNA bases or DNA bases or other nucleobases. The RNA and DNA bases in the guide RNA can be natural bases or modified bases. Guide DNA can be used instead of guide RNA, and DNA-guided effector proteins can be used instead of RNA-guided effector proteins. The use of guide DNA and DNA-guided effector proteins can be preferred when the target polynucleotide is RNA.

[0170] Characterization Step (b) involves characterizing at least a partially ligated non-overhanging strand of the telomere in the 5' to 3' direction from the end of the telomere using a telomere adapter. The method preferably also involves characterizing at least a non-ligated overhanging strand of the telomere in the 5' to 3' direction up to the end of the telomere using a sequencing adapter. Any characterization method may be used. The method preferably uses next-generation sequencing (NGS).

[0171] The ligated non-overhanging strand and / or the non-ligated overhanging strand preferably move relative to a detector such as a nanopore. The detector can be selected from (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor, (iii) an AFM tip, (iv) a nanotube, optionally a carbon nanotube, and (v) a nanopore. Preferably, the detector is a nanopore.

[0172] The ligated non-overhanging strand and / or the non-ligated overhanging strand can be characterized in any suitable manner in the method of the present invention. The ligated non-overhanging strand and / or the non-ligated overhanging strand are preferably characterized by detecting an ion current or an optical signal that moves relative to the nanopore. This is described in more detail herein. The method is suitable for these and other methods of characterizing polynucleotides.

[0173] In another non-limiting example, the ligated non-overhang strand and / or the unligated overhang strand are characterized by detecting by-products of polynucleotide processing reactions such as sequencing by synthesis reactions. Thus, the method can include detecting the products of the sequential addition of (poly)nucleotides by an enzyme such as polymerase to a nucleic acid strand. The product can be a change in one or more properties of the enzyme, such as the three-dimensional structure of the enzyme. Thus, such methods involve subjecting a polynucleotide(s) to an enzyme such as polymerase or reverse transcriptase under conditions such that the template-dependent incorporation of nucleotide bases into the elongating oligonucleotide strand causes a conformational change in the enzyme in response to the template encountered sequentially, the incorporation of strand nucleic acid bases and / or the incorporation of template-specific native or analogous bases (i.e., incorporation events), the detection of conformational changes in the enzyme in response to such incorporation events, and thereby the detection of the sequence of the template strand. In such methods, the polynucleotide strand can be moved according to the methods of the present invention. Such methods can include detecting and / or measuring incorporation events using methods known to those of skill in the art, such as the methods described in US2017 / 0044605.

[0174] In another embodiment, a phosphate-labeled species can be released when a nucleotide is added to a synthetic nucleic acid strand complementary to a template strand, and the by-product can be labeled such that the phosphate-labeled species is detected using a detector described herein. Polynucleotides characterized in this way can be moved according to the methods of this specification. Suitable labels can be optical labels that are detected using nanopores or zero-mode waveguides, or by Raman spectroscopy or other detectors. Suitable labels can be non-optical labels that are detected using nanopores or other detectors.

[0175] In another approach, the nucleoside phosphates (nucleotides) are not labeled, and a native by-product species is detected when a nucleotide is added to a synthetic nucleic acid strand complementary to a template strand. Suitable detectors can be ion-sensitive field effect transistors, or other detectors.

[0176] These and other detection methods are suitable for use in the methods described herein. When the polynucleotide moves relative to the detector, any suitable measurement can be made using the detector.

[0177] Characterization of Nanopores Ligated non-overhang strands and / or unligated overhang strands are preferably characterized using nanopores.

[0178] This method preferably comprises: (i) contacting a ligated non-overhang strand and / or an unligated overhang strand with a nanopore such that the ligated non-overhang strand and / or the unligated overhang strand moves relative to the nanopore; and (ii) taking one or more measurements as the ligated non-overhang strand and / or the unligated overhang strand moves relative to the nanopore, the one or more measurements indicative of one or more characteristics of the ligated non-overhang strand and / or the unligated overhang strand, thereby characterizing the ligated non-overhang strand and / or the unligated overhang strand. The one or more characteristics are preferably selected from: (i) the length of the ligated non-overhang strand and / or the unligated overhang strand; (ii) the identity of the ligated non-overhang strand and / or the unligated overhang strand; (iii) the sequence of the ligated non-overhang strand and / or the unligated overhang strand; (iv) the secondary structure of the ligated non-overhang strand and / or the unligated overhang strand; and (v) whether the ligated non-overhang strand and / or the unligated overhang strand is modified. The ligated non-overhang strand and / or the unligated overhang strand may be modified by methylation, oxidation, damage, one or more proteins, or one or more labels, tags, or spacers. The one or more characteristics of the ligated non-overhang strand and / or the unligated overhang strand are preferably measured by electrical and / or optical measurements. The electrical measurement is preferably a current measurement, an impedance measurement, a tunneling measurement, or a field effect transistor (FET) measurement.

[0179] This method preferably comprises: (i) contacting a ligated non-overhang strand and / or an unligated overhang strand with a nanopore such that the ligated non-overhang strand and / or the unligated overhang strand moves through the nanopore; and (ii) measuring a current passing through the nanopore as the ligated non-overhang strand and / or the unligated overhang strand moves through the nanopore, the current being indicative of one or more characteristics of the ligated non-overhang strand and / or the unligated overhang strand, thereby characterizing the ligated non-overhang strand and / or the unligated overhang strand. The one or more characteristics can be any of those described above.

[0180] Movement of the ligated non-overhang strand and / or the unligated overhang strand relative to or through the nanopore is preferably controlled using a polynucleotide-binding protein. The use of such proteins in nanopore sequencing is known. Examples of suitable proteins are considered in more detail below.

[0181] The present invention also provides a method for characterizing at least a part of a telomere, comprising: (a) ligating a polynucleotide telomere adapter to the 5' end of the non-overhang strand at the end of the telomere, wherein the 3' end of the adapter specifically hybridizes to the first part of the overhang strand and the 5' end of the adapter does not hybridize to the opposite part of the overhang strand; (b) using a polymerase-induced effector protein such as Cas9 protein to create a double-strand break from the telomere end to the end opposite to at least a part of the telomere, and attaching a sequencing adapter to the opposite end; (c) contacting the telomere adapter with a nanopore so that the ligated non-overhang strand moves relative to the nanopore, and taking one or more measurements when the ligated non-overhang strand moves relative to the nanopore, the measurements indicating one or more characteristics of the ligated non-overhang strand, thereby characterizing the ligated non-overhang strand; (d) contacting the sequencing adapter with the nanopore so that the unligated overhang strand moves relative to the nanopore, and taking one or more measurements when the unligated overhang strand moves relative to the nanopore, the measurements indicating one or more characteristics of the unligated overhang strand, thereby characterizing the unligated overhang strand. The ligated non-overhang strand and the unligated overhang strand preferably move through the nanopore. The method preferably includes measuring a current that moves through the nanopore, the current indicating one or more characteristics of the ligated non-overhang strand and the unligated overhang strand. The movement of the ligated non-overhang strand and the unligated overhang strand relative to / through the nanopore is preferably controlled using a polynucleotide-binding protein. The telomere adapter and the sequencing adapter can be any of those discussed above.

[0182] Any suitable nanopores can be used. The nanopores are preferably transmembrane pores. A transmembrane pore is a structure that traverses the membrane to some extent. It enables hydrated ions driven by an applied potential to flow across or within the membrane. Transmembrane pores typically traverse the entire membrane, thereby allowing hydrated ions to flow from one side of the membrane to the opposite side. However, transmembrane pores do not necessarily have to traverse the membrane. They may be closed at one end. For example, the pore can be a well, gap, channel, trench, or slit within the membrane, along which or within which hydrated ions can flow.

[0183] Nanopores typically have a first opening and a second opening. The first opening is typically the cis opening, and the second opening is typically the trans opening. However, the first opening can be the trans opening, and the second opening can be the cis opening. The polynucleotide-binding protein used in the method of the present invention is typically provided at the first opening of the nanopore and thus controls the movement of the target polynucleotide in the direction from the second opening of the nanopore to the first opening of the nanopore.

[0184] In the method of the present invention, any transmembrane pores may be used. The pores can be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid-state pores. The pores can be DNA origami pores (Langecker et al., Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983.

[0185] The nanopores are preferably transmembrane protein pores. A transmembrane protein pore is a polypeptide or an aggregate of polypeptides that allows hydrated ions, such as polynucleotides, to flow from one side of the membrane to the opposite side of the membrane. In the method of the present invention, the transmembrane protein pore can form a pore that allows hydrated ions driven by an applied potential to flow from one side of the membrane to the opposite side. The transmembrane protein pore preferably allows a polynucleotide to flow from one side of the membrane, such as a triblock copolymer membrane, to the opposite side. The transmembrane protein pore allows a polynucleotide to move through the pore.

[0186] The nanopores can be transmembrane protein pores that are monomers or oligomers. The pores are preferably composed of several repeating subunits, such as at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, or at least about 16 subunits. The pores are preferably hexamer, heptamer, octamer, or nanomer pores. The pores can be homo-oligomers or hetero-oligomers.

[0187] The transmembrane protein pore can include a barrel or channel through which ions can flow. The subunits of the pore typically surround a central axis and contribute to a transmembrane β-barrel or channel or a transmembrane α-helix bundle or channel.

[0188] Typically, the barrel or channel of a transmembrane protein pore contains amino acids that facilitate interaction with an analyte, such as a target polynucleotide (as described herein). These amino acids are preferably located near the constriction of the barrel or channel. Transmembrane protein pores typically contain one or more positively charged amino acids, such as arginine, lysine, or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate interaction between the pore and a nucleotide, polynucleotide, or nucleic acid.

[0189] The nanopore can be a transmembrane protein pore derived from a β-barrel pore or an α-helix bundle pore. A β-barrel pore includes a barrel or channel formed from β-strands. Suitable β-barrel pores include β-toxins such as α-hemolysin, anthrax toxin, and leukocidin, and bacterial outer membrane proteins / polysins such as Mycobacterium smegmatis porin (Msp), e.g., MspA, MspB, MspC, or MspD, CsgG, outer membrane polysin F (OmpF), outer membrane polysin G (OmpG), outer membrane phospholipase A, and Neisseria autotransporter lipoprotein (NalP), as well as other pores such as lysenin, but are not limited thereto. An α-helix bundle pore includes a barrel or channel formed from α-helices. Suitable α-helix bundle pores include inner membrane proteins and α outer membrane proteins, e.g., WZA and ClyA toxins, but are not limited thereto.

[0190] The nanopore can be a transmembrane pore derived from or based on Msp, α-hemolysin (α-HL), lysenin, CsgG, ClyA, Sp1, or hemolytic protein fragaceatoxin C (FraC).

[0191] The nanopore can be a transmembrane protein pore derived from CsgG, for example, CsgG from E. coli strain K-12 sub-strain MC4100. Such pores are oligomers and typically contain 7, 8, 9, or 10 monomers derived from CsgG. The pore can be a homo-oligomeric pore derived from CsgG containing the same monomers. Alternatively, the pore can be a hetero-oligomeric pore derived from CsgG containing at least one monomer different from the others. Examples of suitable pores derived from CsgG are disclosed in WO2016 / 034591, WO2017 / 149316, WO2017 / 149317, WO2017 / 149318, and WO2019 / 002893 (all of which are hereby incorporated by reference in their entirety).

[0192] The nanopore can be a transmembrane pore derived from lysenin. Examples of suitable pores derived from lysenin are disclosed in WO2013 / 153359 (which is hereby incorporated by reference in its entirety).

[0193] The nanopore can be a transmembrane pore derived from or based on α-hemolysin (α-HL). The wild-type α-hemolysin pore is formed from 7 identical monomers or subunits (i.e., it is a heptamer). The α-hemolysin pore can be α-hemolysin-NN or a variant thereof. The variant preferably contains N residues at positions E111 and K147.

[0194] The nanopore can be a transmembrane protein pore derived from Msp, for example, MspA. Examples of suitable pores derived from MspA are disclosed in WO2012 / 107778 (which is hereby incorporated by reference in its entirety).

[0195] The nanopore can be a transmembrane pore derived from or based on ClyA.

[0196] Membrane The detector or nanopore is typically present in a membrane. Any suitable membrane can be used.

[0197] The membrane is preferably an amphiphilic layer. The amphiphilic layer is a layer formed from amphiphilic molecules such as phospholipids that have both hydrophilic and lipophilic properties. The amphiphilic molecules can be synthetic or naturally occurring. Amphiphilic substances that do not occur naturally and amphiphilic substances that form monolayers are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). A block copolymer is a polymeric material in which two or more monomer subunits are polymerized together to make a single polymer chain. Block copolymers typically have properties contributed by each monomer subunit. However, block copolymers can have unique properties not possessed by the polymers formed from the individual subunits. Block copolymers can be manipulated in an aqueous medium such that one of the monomer subunits is hydrophobic (i.e., lipophilic) and the other subunit(s) is hydrophilic. In this case, the block copolymer can have amphiphilic properties and can form a structure that mimics a biological membrane. Block copolymers can be diblocks (consisting of two monomer subunits), but can also be constructed from more than two monomer subunits to form more complex arrangements that behave as amphiphilic substances. The copolymer can be a triblock, tetrablock, or pentablock copolymer. The membrane can be a triblock copolymer membrane.

[0198] Archaeal bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipids form a monolayer membrane. These lipids are generally found in extremeophilic bacteria, thermophiles, halophiles, and acidophiles that survive in harsh biological environments. Their stability is thought to derive from the fused nature of the resulting bilayer. It is straightforward to construct block copolymers that mimic these biological entities by creating triblock polymers with the general motif hydrophilic-hydrophobic-hydrophilic. This material forms monomeric membranes that behave like lipid bilayers and can exhibit a wide range of phase behavior from vesicles to lamellar membranes. Membranes formed from these triblock copolymers have several advantages over biological lipid membranes. Since the triblock copolymers are synthetic, their exact structure can be carefully controlled to provide the correct chain lengths and properties necessary to form membranes and interact with pores and other proteins.

[0199] Block copolymers may also be constructed from subunits that are not classified as lipid submaterials. For example, the hydrophobic polymer can be made from siloxanes or other non-hydrocarbon-based monomers. The hydrophilic subsections of block copolymers can also have low protein-binding properties, which enables the fabrication of membranes that are highly resistant when exposed to raw biological samples. This head group unit may also be derived from unclassified lipid head groups.

[0200] Triblock copolymer membranes also have increased mechanical and environmental stability compared to biological lipid membranes, e.g., much higher operating temperatures or pH ranges. The synthetic nature of the block copolymers provides a platform for customizing polymer-based membranes for a wide range of applications.

[0201] The membrane may also be one of the membranes disclosed in International Application No. 2014 / 064443 or 2014 / 064444 (both of which are hereby incorporated by reference in their entirety).

[0202] Amphiphilic molecules can be chemically modified or functionalized to facilitate the coupling of polynucleotides. The amphiphilic layer can be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.

[0203] Amphiphilic membranes are typically naturally mobile and act essentially as two-dimensional fluids with a lipid diffusion rate of approximately 10 -8 cm s -1 . This means that pores and coupled polynucleotides can typically move within the amphiphilic membrane.

[0204] The membrane can be a lipid bilayer. The lipid bilayer is a model of the cell membrane and functions as an excellent basis for a wide range of experimental studies. For example, the lipid bilayer can be used for in vitro investigation of membrane proteins by single-channel recording. Alternatively, the lipid bilayer can be used as a biosensor for detecting the presence of a wide range of substances. The lipid bilayer can be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers, or liposomes. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO2008 / 102121, WO2009 / 077734, and WO2006 / 100484 (which are hereby incorporated by reference in their entirety).

[0205] Methods for forming lipid bilayers are known in the art. Lipid bilayers are typically formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561 - 3566).

[0206] The lipid bilayer can be formed as described in WO2009 / 077734, which is hereby incorporated by reference in its entirety. In this method, the lipid bilayer is formed from dry lipids. The lipid bilayer can be formed across an opening as described in WO2009 / 077734.

[0207] The membrane can include a solid state layer. The solid state layer can be formed from both organic materials and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as Si 3 N 4 、 A1 2 O 3 、 and SiO, organic and inorganic polymers such as polyamides, plastics such as Teflon®, or elastomers such as two-component addition-curing silicone rubber, and glass. The solid state layer may be formed from graphene. Suitable graphene layers are disclosed in WO2009 / 035647, which is hereby incorporated by reference in its entirety. When the membrane includes a solid state layer, the pores are typically present in an amphiphilic membrane or layer contained within the solid state layer, such as in holes, wells, gaps, channels, trenches, or slits within the solid state layer. One of ordinary skill in the art can prepare a suitable solid state / amphiphilic hybrid system. Suitable systems are disclosed in WO2009 / 020682 and WO2012 / 005857, which are hereby incorporated by reference in their entirety. Any of the amphiphilic membranes or layers discussed above can be used.

[0208] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer containing pores, (ii) an isolated naturally occurring lipid bilayer containing pores, or (iii) a cell with pores inserted therein. The method is typically carried out using an artificial amphiphilic layer such as an artificial triblock copolymer layer. The layer can include other transmembrane and / or intramembrane proteins, as well as other molecules in addition to the pores. Suitable devices and conditions are discussed below. The methods of the invention are typically carried out in vitro.

[0209] Polynucleotide-binding protein As will be understood by those skilled in the art, any suitable polynucleotide-binding protein can be used in the methods and products of the present invention. A polynucleotide-binding protein can be any protein that can bind to a polynucleotide and control its movement relative to a detector, such as a nanopore.

[0210] More specifically, polynucleotide-binding proteins, such as helicases, typically have at least two active modes of operation (when all components necessary to promote movement, such as ATP and Mg 2+ are provided) and one inactive mode of operation (when the components necessary to facilitate movement are not provided, or when the polynucleotide-binding protein is modified to prevent the active mode), in which they can control the movement of DNA.

[0211] When all components necessary to promote movement are provided, the polynucleotide-binding protein can move along a polynucleotide, such as DNA, in either the 5'-3' or 3'-5' direction. Many polynucleotide-binding proteins process polynucleotides, such as DNA, in the 5'-3' direction. Polynucleotide-binding proteins that thus control the movement of polynucleotides are typically suitable for use in the methods of the present invention.

[0212] However, if the polynucleotide-binding protein does not have the components necessary to facilitate movement, or is modified to prevent the active control of the movement of the polynucleotide relative to the nanopore, it can still passively control the movement of the polynucleotide relative to the nanopore. For example, a polynucleotide-binding protein can bind to a polynucleotide and act as a brake to slow the movement of the polynucleotide when the polynucleotide is drawn into the pore by the applied field (e.g., by the first force in the methods of the present invention). In the "inactive" mode, since the applied force provides the driving force to move the polynucleotide through the nanopore, it is usually not a problem whether the DNA is captured at the 3' or 5' end (i.e., whether the nanopore moves in the 5'-3' or 3'-5' direction). However, in such embodiments, the polynucleotide-binding protein can still control the movement of the polynucleotide relative to the nanopore, for example, by acting as a brake. In the case of the inactive mode, the control of the movement of the polynucleotide by the polynucleotide-binding protein can be described in several ways including ratcheting, sliding, and braking. Typically, the methods of the present invention do not include the use of polynucleotide-binding proteins that operate in the passive mode. However, if a polynucleotide-binding protein is used, it can be a polynucleotide-binding protein that operates in the passive mode.

[0213] Some methods of the present invention can include the use of a polynucleotide-binding protein as a stop portion that prevents the movement of a polynucleotide strand through a nanopore. The polynucleotide-binding protein can be a protein that binds to a polynucleotide but does not have polynucleotide processing capabilities, i.e., it is not a polynucleotide-processing protein.

[0214] A polynucleotide handling enzyme is a polypeptide that can interact with a polynucleotide. The enzyme may modify the polynucleotide by cleaving the polynucleotide to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. The enzyme may modify the polynucleotide by orienting it or moving it to a specific location. As used herein, a polynucleotide-binding protein may be or be derived from a polynucleotide handling enzyme. A polynucleotide-binding protein may be or be derived from a polynucleotide handling enzyme.

[0215] A polynucleotide-binding protein may be derived from any member of enzyme classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31.

[0216] Typically, a polynucleotide-binding protein is a helicase, polymerase, exonuclease, topoisomerase, or a variant thereof.

[0217] A polynucleotide-binding protein may be modified to prevent dissociation of the polynucleotide-binding protein from the polynucleotide. Thus, the target polynucleotide preferably does not dissociate from the polynucleotide-binding protein.

[0218] As used herein, the term "dissociation" refers to the dissociation of a polynucleotide-binding protein from a target polynucleotide. Thus, a polynucleotide-binding protein can be modified to prevent it from dissociating from the target polynucleotide, e.g., into the reaction medium. It is important to distinguish the potential "departure" of a polynucleotide-binding protein from the "unbinding" of the polynucleotide-binding protein from the target polynucleotide. As used herein, "unbinding" refers to the temporary release of the active site of the polynucleotide-binding protein of the target polynucleotide (described in more detail herein), but does not mean dissociation. Thus, for example, a polynucleotide-binding protein can be modified such that it prevents the polynucleotide-binding protein from dissociating from the polynucleotide, but does not prevent the polynucleotide-binding protein from unbinding from the polynucleotide. When not bound, the polynucleotide-binding protein remains associated with the target polynucleotide. For example, a polynucleotide-binding protein can maintain its binding to the target polynucleotide (i.e., prevent dissociation from the target polynucleotide) because it is topologically closed around the target polynucleotide. The polynucleotide-binding site can remain freely bindable or unbindable to the target polynucleotide such that the polynucleotide-binding protein can bind or unbind to the target polynucleotide while remaining associated with the target polynucleotide. When the polynucleotide-binding protein unbinds from the target polynucleotide, it can move (e.g., along) on the target polynucleotide under the applied force and can rebind to the target polynucleotide. When associated with the target polynucleotide but unbound from it, the polynucleotide-binding protein cannot dissociate from the target polynucleotide.

[0219] The polynucleotide-binding protein can be adapted to prevent dissociation in any suitable manner. For example, the polynucleotide-binding protein can be loaded onto the polynucleotide and then modified to prevent its dissociation from the polynucleotide. Alternatively, the polynucleotide-binding protein can be modified to prevent its dissociation from the polynucleotide before it is loaded onto the polynucleotide. The modification of the polynucleotide-binding protein and / or the polynucleotide-binding protein to prevent the dissociation of the polynucleotide-binding protein from the polynucleotide can be achieved using methods known in the art, such as those discussed in WO2014 / 013260, which is hereby incorporated by reference in its entirety, and with particular reference to the section describing the modification of polynucleotide-binding proteins, such as helicases, to prevent their dissociation from the polynucleotide strand. For example, the polynucleotide-binding protein can be modified by treatment with tetramethylazodicarboxamide (TMAD). Various other blocking moieties are described in WO2021 / 255476, which is hereby incorporated by reference in its entirety.

[0220] For example, a polynucleotide-binding protein and / or a polynucleotide-binding protein may have a polynucleotide dissociation opening, such as a cavity, groove, or void, through which the strand can pass when the polynucleotide-binding protein dissociates from the strand. The polynucleotide dissociation opening can be an opening through which a nucleotide can pass when the polynucleotide-binding protein dissociates from the nucleotide. The polynucleotide dissociation opening of a given polynucleotide-binding protein can be determined by reference to its structure, for example, by reference to its X-ray crystal structure. The X-ray crystal structure can be obtained in the presence and / or absence of a polynucleotide substrate. The position of the polynucleotide non-binding opening in a given polynucleotide-binding protein can be estimated or confirmed by molecular modeling using standard packages known in the art. The polynucleotide dissociation opening can be transiently generated by the movement of one or more portions of the polynucleotide-binding protein, such as one or more domains.

[0221] A polynucleotide-binding protein can be modified by closing the polynucleotide dissociation opening. The polynucleotide dissociation opening can be closed with a closing moiety. Thus, by closing the polynucleotide dissociation opening, dissociation of the polynucleotide-binding protein from the polynucleotide can be prevented. For example, a polynucleotide-binding protein can be modified by covalently closing the polynucleotide dissociation opening. However, as explained above, closing the polynucleotide dissociation opening does not necessarily prevent the target polynucleotide from dissociating from the polynucleotide-binding site of the polynucleotide-binding protein. In some embodiments, a protein preferred for dealing with this way is a helicase.

[0222] The polynucleotide-binding protein can be modified at the closed portion to (i) topologically close the polynucleotide-binding site of the polynucleotide-binding protein around the target polynucleotide, (ii) facilitate the dissociation of the target polynucleotide from the polynucleotide-binding site of the polynucleotide-binding protein, and / or delay the reassociation of the target polynucleotide to the polynucleotide-binding site of the polynucleotide-binding protein. The polynucleotide-binding protein can be modified in any suitable manner to facilitate the attachment of such a closed portion.

[0223] The closed portion can include a bifunctional crosslinking portion. The closed portion can include a bifunctional crosslinking agent. The bifunctional crosslinking agent attaches at two points on the polynucleotide-binding protein and can close the polynucleotide dissociation opening of the polynucleotide-binding protein, thereby preventing the polynucleotide from dissociating from the polynucleotide-binding protein while allowing the dissociation of the polynucleotide from the polynucleotide-binding site of the polynucleotide-binding protein.

[0224] The closing moiety can be attached at any suitable position on the polynucleotide-binding protein. For example, the closing moiety can crosslink two amino acid residues of the polynucleotide-binding protein. Typically, at least one amino acid crosslinked by the closing moiety is cysteine or a non-natural amino acid. Cysteine or a non-natural amino acid can be introduced into the polynucleotide-binding protein by substitution or modification of a naturally occurring amino acid residue of the polynucleotide-binding protein. Methods for introducing non-natural amino acids are well known in the art and include, for example, native chemical ligation with a synthetic polypeptide chain containing such non-natural amino acids. Methods for introducing cysteine into a polynucleotide-binding protein are also within the capabilities of one of ordinary skill in the art and use techniques described in references such as Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Press, Plainsview, New York (2012), and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).

[0225] The closing moiety can have a length of from about 1 Å to about 100 Å. The length of the closing moiety can be calculated according to the static bond length or, more preferably, using molecular dynamics simulations. The length can be, for example, from about 2 Å to about 80 Å, such as from about 5 Å to about 50 Å, such as from about 8 to about 30 Å, such as from about 10 to about 25 Å or about 20 Å, such as about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 Å.

[0226] Polynucleotide-binding proteins suitable for being closed using a closing moiety as described above are discussed in more detail herein. The polynucleotide-binding protein is preferably a helicase as described herein, such as the Dda helicase.

[0227] The polynucleotide-binding protein may be an exonuclease or derived therefrom. Suitable enzymes include, but are not limited to, exonuclease I from E. coli, exonuclease III enzyme from E. coli, RecJ from T. thermophilus, bacteriophage lambda exonuclease, TatD exonuclease, and variants thereof.

[0228] The polynucleotide-binding protein can be a polymerase. The polymerase can be, for example, PyroPhage® 3173 DNA polymerase (commercially available from Lucigen® Corporation), SD polymerase (commercially available from Bioron®), Klenow from NEB, or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase or a variant thereof. A modified version of the Phi29 polymerase that can be used in the present invention is disclosed in U.S. Patent No. 5,576,204.

[0229] The polynucleotide-binding protein can be a topoisomerase. In one embodiment, the topoisomerase is a member of either of the subclassification (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase can be a reverse transcriptase, which is an enzyme capable of catalyzing the formation of cDNA from an RNA template. They are commercially available, for example, from New England Biolabs® and Invitrogen®.

[0230] The polynucleotide-binding protein is preferably a helicase. Any suitable helicase can be used in the method of the present invention. For example, each or the said enzyme used in accordance with the present disclosure can independently be selected from Hel308 helicase, RecD helicase, TraI helicase, TrwC helicase, XPD helicase, and Dda helicase, or variants thereof. The monomeric helicase can include several domains attached together. For example, TraI helicase and TraI subgroup helicases can include two RecD helicase domains, a relaxase domain, and a C-terminal domain. These domains typically form monomeric helicases that can function without forming oligomers. Specific examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pif1, and TraI. These helicases typically act on single-stranded DNA. Examples of helicases that can move along both strands of double-stranded DNA include FtfK and hexameric enzyme complexes, or multi-subunit complexes such as RecBCD. The polynucleotide-binding protein is preferably a Dda (DNA-dependent ATPase) helicase.

[0231] Hel308 helicase is described in publications such as WO2013 / 057495, the entire contents of which are incorporated by reference. RecD helicase is described in publications such as WO2013 / 098562, the entire contents of which are incorporated by reference. XPD helicase is described in publications such as WO2013 / 098561, the entire contents of which are incorporated by reference. Dda helicase is described in publications such as WO2015 / 055981 and WO2016 / 055777, the entire contents of each of which are incorporated by reference.

[0232] The helicase can be Trwc Cba or a variant thereof, Hel308Mbu or a variant thereof, or Dda or a variant thereof. The variant can differ from the native sequence in any of the ways discussed herein. Exemplary variants of Dda include E94C / A360C. Further exemplary variants of Dda include E94C / A360C, followed by (ΔM1)G1G2 (i.e., deletion of M1, followed by addition of G1 and G2).

[0233] Disclaimer The 3' end of the telomere adapter preferably does not contain one or more universal nucleotides. A universal nucleotide is a nucleotide that hybridizes or binds to some extent to all nucleotides in a template polynucleotide. A universal nucleotide preferably hybridizes or binds to some extent to nucleotides including the nucleosides adenosine (A), thymine (T), uracil (U), guanine (G), and cytosine (C). A universal nucleotide may hybridize more strongly to some nucleotides than to other nucleotides. For example, a universal nucleotide containing the nucleoside, 2'-deoxyinosine (I), will show a preferential pairing order of I-C>I-A>I-G approximately = I-T. If a universal nucleotide substitutes for a nucleotide species within a population, the polymerase will replace the nucleotide species with the universal nucleotide. For example, the polymerase will replace dGMP with the universal nucleotide when in contact with a population of free dAMP, dTMP, dCMP, and the universal nucleotide.

[0234] Universal nucleotides preferably include one of the following nucleobases: hypoxanthine, 4-nitroindole, 5-nitroindole, 6-nitroindole, formylindole, 3-nitropyrrole, nitroimidazole, 4-nitropyrazole, 4-nitrobenzimidazole, 5-nitroindazole, 4-aminobenzimidazole, or phenyl (C6 aromatic ring). More preferably, universal nucleotides include one of the following nucleosides: 2'-deoxyinosine, inosine, 7-deaza-2'-deoxyinosine, 7-deaza-inosine, 2-aza-deoxyinosine, 2-aza-inosine, 2-O'-methylinosine, 4-nitroindole 2'-deoxyribonucleoside, 4-nitroindole ribonucleoside, 5-nitroindole 2'-deoxyribonucleoside, 5-nitroindole ribonucleoside, 6-nitroindole 2'-deoxyribonucleoside, 6-nitroindole ribonucleoside, 3-nitropyrrole 2'-deoxyribonucleoside, 3-nitropyrrole ribonucleoside, acyclic sugar analog of hypoxanthine, nitroimidazole 2'-deoxyribonucleoside, nitroimidazole ribonucleoside, 4-nitropyrazole 2'-deoxyribonucleoside, 4-nitropyrazole ribonucleoside, 4-nitrobenzimidazole 2'-deoxyribonucleoside, 4-nitrobenzimidazole ribonucleoside, 5-nitroindazole 2'-deoxyribonucleoside, 5-nitroindazole ribonucleoside, 4-aminobenzimidazole 2'-deoxyribonucleoside, 4-aminobenzimidazole ribonucleoside, phenyl C-ribonucleoside, phenyl C-2'-deoxyribosyl nucleoside, 2'-deoxynebularine, 2'-deoxyisoguanosine, K-2'-deoxyribose, P-2'-deoxyribose, and pyrrolidine. More preferably, universal nucleotides include 2'-deoxyinosine. More preferably, universal nucleotides are IMP or dIMP. Most preferably, universal nucleotides are dPMP (2'-deoxy-P-nucleoside monophosphate) or dKMP (N6-methoxy-2,6-diaminopurine monophosphate).

[0235] The method of the present invention preferably does not include restriction digestion. The method of the present invention preferably does not include amplifying at least a part of telomeres or polymerase chain reaction (PCR). The method of the present invention preferably does not include (i) restriction digestion and / or (ii) amplifying at least a part of telomeres or polymerase chain reaction (PCR). The method of the present invention preferably does not include (i) restriction digestion, (ii) amplifying at least a part of telomeres, and (iii) or polymerase chain reaction (PCR). These embodiments provide the advantages discussed above.

[0236] Telomere source The method of the present invention can be carried out on any suitable sample. The sample is typically one known or suspected to contain at least a part of telomeres.

[0237] The sample may be a biological sample. The present invention can be carried out in vitro on a sample obtained or extracted from any organism or microorganism. Telomeres are typically present only in eukaryotes. At least a part of telomeres can be derived from any eukaryote, including any of those listed below.

[0238] The sample is preferably a fluid sample. The sample usually contains body fluids. The body fluids may be obtained from humans or animals. The human or animal may have a disease, be suspected of having a disease, or be at risk of a disease. The sample may be urine, lymph fluid, saliva, mucus, semen, or amniotic fluid, but is preferably whole blood, plasma, or serum. Typically, the sample is of human origin, but alternatively, it may be of another mammalian origin, such as that of commercially raised animals such as horses, cows, sheep or pigs, or alternatively, it may be a pet such as a cat or a dog.

[0239] Alternatively, plant-derived samples are typically obtained from commercially available crops such as grains, legumes, fruits, or vegetables, for example, wheat, barley, oats, rapeseed, corn, soybeans, rice, bananas, apples, tomatoes, potatoes, grapes, tobacco, beans, lentils, sugarcane, cocoa, cotton, tea, or coffee.

[0240] The sample may be a non-biological sample. The non-biological sample is preferably a liquid sample. Examples of non-biological samples include surgical fluids, water, such as drinking water, seawater, or river water, and reagents for clinical tests.

[0241] Typically, the sample can be processed, for example, by centrifugation or by passage through a membrane that filters and removes unwanted molecules or cells, such as red blood cells, before being assayed. The sample may be measured immediately after collection. The sample can also typically be stored, preferably at less than -70 °C, before the assay.

[0242] Typically, the sample contains genomic DNA. The sample may contain T cell DNA.

[0243] At least a portion of the telomere can be derived from common organisms such as plants or animals. At least a portion of the telomere is often obtained from humans or animals, for example, from urine, lymph fluid, saliva, mucus, semen, or amniotic fluid, or from whole blood, plasma, or serum. At least a portion of the telomere can be obtained from plants, such as grains, legumes, fruits, or vegetables.

[0244] General method As described above, the method of the present invention can be operated using any suitable detector and thus any suitable device for detecting polynucleotides can be used.

[0245] The method of the present invention can be carried out using any device suitable for nanopore sensing. For example, the device can comprise a chamber containing an aqueous solution and a barrier separating the chamber into two sections. The barrier can have an opening through which a membrane containing transmembrane pores is formed. The transmembrane pores are described herein.

[0246] The method can be carried out using the devices described in WO2008 / 102120, WO2010 / 122293, or WO00 / 28312 (which are hereby incorporated by reference in their entirety). Briefly, the binding of molecules (e.g., target polynucleotides) within the pore channel affects the open-channel ion current through the pore, which is the essence of "molecular sensing" of the pore channel. Variations in the open-channel ion current can be measured using suitable measurement techniques by changes in the current. The degree of decrease in the ion current measured by the decrease in the current is related to the size of the obstacle within or near the pore. Thus, the binding of the target molecule (e.g., target polynucleotide) within or near the pore provides a detectable and measurable event, thereby forming the basis of a "biological sensor". By detecting the presence of biomolecules, applications are found in personalized drug development, medicine, diagnostics, life science research, environmental monitoring, and in the security and / or defense industries.

[0247] When used to characterize a polynucleotide, the presence, absence, or one or more characteristics of the target polynucleotide are determined. The method can be for determining the presence, absence, or one or more characteristics of at least one target polynucleotide. The method can relate to determining the presence, absence, or one or more characteristics of two or more target polynucleotides. The method can include determining the presence, absence, or one or more characteristics of any number of target polynucleotides, such as 2, 5, 10, 15, 20, 30, 40, 50, 100 or more target polynucleotides. Any number of characteristics of one or more target polynucleotides, such as 1, 2, 3, 4, 5, 10 or more characteristics, can be determined. Characteristics that can be detected by the methods provided herein include polynucleotide identity or sequence, polynucleotide length, whether the polynucleotide is modified, and the like. In some embodiments, the method of the invention is a method for sequencing at least a portion of a telomere. In some embodiments, the sequence of at least a portion of the telomere can be determined in real time by aligning a real-time signal or base calling to a known reference. Exemplary methods for determining polynucleotide sequence are described in WO2016 / 059427, which is incorporated herein by reference in its entirety.

[0248] When used to characterize a polynucleotide, the method can typically include measuring the flow of ionic current through a pore, by measuring an electric current. Alternatively, the flow of ions through the pore can be measured optically, as disclosed, for example, by Heron et al: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Thus, the device can also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore. The characterization method can be carried out using a patch clamp or a voltage clamp. The characterization method preferably involves the use of a voltage clamp.

[0249] This method can include measuring an optical signal, as described in Chen et al, Nature Communications (2018) 9:1733, the entire content of which is incorporated herein by reference. For example, nanopores such as optically designed nanopore structures (such as plasmonic nanoslits) can be used to locally enable single-molecule surface-enhanced Raman spectroscopy (SERS) and characterize polynucleotides by direct Raman spectroscopic detection.

[0250] This method can be implemented in a silicon-based well array where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.

[0251] This method can include measuring the current flowing through the pores. The method is typically carried out using a voltage applied across the membrane and pores. The voltage used is typically in the range of +2V to -2V, typically -400mV to +400mV. The voltage used is preferably in a range having a lower limit independently selected from -400mV, -300mV, -200mV, -150mV, -100mV, -50mV, -20mV, and 0mV and an upper limit independently selected from +10mV, +20mV, +50mV, +100mV, +150mV, +200mV, +300mV, and +400mV. The voltage used is more preferably in the range of 100mV to 240mV, most preferably in the range of 120mV to 220mV. By using an increased applied potential, it is possible to increase the identification between different nucleotides for each pore.

[0252] This method is typically carried out in the presence of a metal salt, such as an alkali metal salt, a halogen salt, such as a chloride salt, such as an alkali metal chloride salt, or any charge carrier. Examples of charge carriers include ionic liquids or organic salts, such as tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary apparatus discussed above, the salt is present in the aqueous solution within the chamber. Potassium chloride (KCl), sodium chloride (NaCl), or cesium chloride (CsCl) is typically used. KCl is preferred. The salt can be an alkaline earth metal salt such as calcium chloride (CaCl2). The salt concentration can be saturated. The salt concentration can be 3 M or less, typically 0.1 - 2.5 M, 0.3 - 1.9 M, 0.5 - 1.8 M, 0.7 - 1.7 M, 0.9 - 1.6 M, or 1 M - 1.4 M. The salt concentration is preferably 150 mM - 1 M. The characterization method is preferably carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M, or at least 3.0 M. High salt concentrations provide a high signal-to-noise ratio and enable the identification of currents that show binding / no binding to the background of normal current fluctuations.

[0253] The method is usually carried out in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is present in the aqueous solution within the chamber. Any suitable buffer can be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCl buffer. The method is typically carried out at a pH of 4.0 - 12.0, 4.5 - 10.0, 5.0 - 9.0, 5.5 - 8.8, 6.0 - 8.7, or 7.0 - 8.8, or 7.5 - 8.5. The pH used is preferably about 7.5.

[0254] The method can be carried out at 0 °C to 100 °C, 15 °C to 95 °C, 16 °C to 90 °C, 17 °C to 85 °C, 18 °C to 80 °C, 19 °C to 70 °C, or 20 °C to 60 °C. Optionally, the method is carried out at a temperature that supports enzymatic function, for example, at about 37 °C.

[0255] Any of the proteins described herein, such as protein pores, can be made synthetically or by recombinant means. For example, the pores can be synthesized by in vitro translation and transcription (IVTT). The amino acid sequence of the pores can be modified to include non-natural amino acids or to increase the stability of the protein. Such amino acids may be introduced during production if the protein is produced by synthetic means. The pores can also be altered after either synthetic or recombinant production.

[0256] Any of the proteins described herein, such as protein pores, can be generated using standard methods known in the art. The polynucleotide sequence encoding the pores or structures can be derived and replicated using standard methods in the art. The polynucleotide sequence encoding the pores or structures can be expressed in bacterial host cells using standard methods in the art. The pores can be generated in cells by in situ expression of polypeptides from recombinant expression vectors. Optionally, the expression vectors carry inducible promoters to control the expression of the polypeptides. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.

[0257] The pores may be produced on a large scale after purification by any protein liquid chromatography system from a protein-producing organism, or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA system, Bio-Cad system, Bio-Rad BioLogic system, and Gilson HPLC system.

[0258] Polynucleotide telomere adapter The present invention also provides a polynucleotide telomere adapter, wherein the 3' end of the adapter specifically hybridizes to the first part of the overhang strand at the end of the telomere, and the 5' end of the adapter is not complementary to the opposite part of the overhang strand at the end of the telomere. The telomere adapter may be any of those defined above with reference to the method of the present invention, including the extended telomere adapter. The telomere adapter may further comprise a sprint polynucleotide and / or a sequencing adapter.

[0259] The present invention also provides a population of six telomere adapters, each having a 3' end that specifically hybridizes to one of six possible sequences of the first part of the overhang strand at the end of the telomere, and a 5' end that is not complementary to the opposite part of the overhang strand at the end of the telomere. The population may be any of those defined above with respect to the method of the present invention, including the extended telomere adapter.

[0260] The population preferably comprises or consists of six telomere adapters having the following sequences: AGCAATACGTAACTGAACGAAGTACCCTAA (SEQ ID NO: 1), AGCAATACGTAACTGAACGAAGTAACCCTA (SEQ ID NO: 2), AGCAATACGTAACTGAACGAAGTTAACCCT (SEQ ID NO: 3), AGCAATACGTAACTGAACGAAGTCTAACCC (SEQ ID NO: 4), AGCAATACGTAACTGAACGAAGTCCTAACC (SEQ ID NO: 5), and AGCAATACGTAACTGAACGAAGTCCCTAAC (SEQ ID NO: 6).

[0261] These are shown in Table 3 and FIG. 1 of Example 1.

[0262] The population preferably comprises or consists of six telomere adapters comprising the following sequences: AGCAATACGTAACTGAACGAAGTCCCTAACCCTAACCCTAA (SEQ ID NO: 14), AGCAATACGTAACTGAACGAAGTTAACCCTAACCCTAACCC (SEQ ID NO: 15), AGCAATACGTAACTGAACGAAGTCTAACCCTAACCCTAACC (SEQ ID NO: 16), AGCAATACGTAACTGAACGAAGTCCTAACCCTAACCCTAAC (SEQ ID NO: 17), AGCAATACGTAACTGAACGAAGTAACCCTAACCCTAACCCT (SEQ ID NO: 18), and AGCAATACGTAACTGAACGAAGTACCCTAACCCTAACCCTA (SEQ ID NO: 19).

[0263] These are shown in Table 14.

[0264] The population preferably comprises or consists of six extended telomere adapters comprising the following sequences: SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTACCCTAA (SEQ ID NO: 1), SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTAACCCTA (SEQ ID NO: 2), SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTTAACCCT (SEQ ID NO: 3), SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTCTAACCC (SEQ ID NO: 4), SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTCCTAACC (SEQ ID NO: 5), and SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTCCCTAAC (SEQ ID NO: 6).

[0265] The population preferably comprises or consists of six extended telomere adapters comprising the following sequences: GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTACCCTAA (SEQ ID NO: 25), GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTAACCCTA (SEQ ID NO: 26), GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTTAACCCT (SEQ ID NO: 27), GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTCTAACCC (SEQ ID NO: 28), GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTCCTAACC (SEQ ID NO: 29), and GCTTGGGTGTTTAACCNNNNAGCAATACGTAACTGAACGAAGTCCCTAAC (SEQ ID NO: 30) (wherein N is Int 5-nitroindole (i5NiTInd)).

[0266] The population preferably comprises or consists of six extended telomere adapters comprising the following sequences: SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTCCCTAACCCTAACCCTAA (SEQ ID NO: 14), SEQ ID NO: 11 covalently linked to AGCAATACGTAACTGAACGAAGTTAACCCTAACCCTAACCC (SEQ ID NO: 15), SEQ ID NO:11 covalently attached to AGCAATACGTAACTGAACGAAGTCTAACCCTAACCCTAACC (SEQ ID NO:16), SEQ ID NO:11 covalently attached to AGCAATACGTAACTGAACGAAGTCCTAACCCTAACCCTAAC (SEQ ID NO:17), SEQ ID NO:11 covalently attached to AGCAATACGTAACTGAACGAAGTAACCCTAACCCTAACCCT (SEQ ID NO:18), and SEQ ID NO:11 covalently attached to AGCAATACGTAACTGAACGAAGTACCCTAACCCTAACCCTA (SEQ ID NO:19).

[0267] The two sequences in the extension adapter are preferably covalently linked by one or more, such as two or more, three or more, or four or more universal nucleotides, such as Int 5-nitroindole (i5NiTInd). The two sequences in the extension adapter are more preferably covalently linked by four i5NiTInds. The extension telomere adapter preferably contains a click chemistry group such as DBCOTEG at the 5' end. Examples of these extension adapters are shown in Table 7 and FIG. 4.

[0268] The telomere adapters in the population may each further comprise a primer polynucleotide and / or a sequencing adapter.

[0269] Kits and Systems The present invention also provides a kit for characterizing at least a part of a telomere, the kit comprising: (a) one or more polynucleotide telomere adapters of the present invention or a population of six telomere adapters of the present invention; and (b) one or more splint polynucleotides or one or more polynucleotide extensions. The one or more telomere adapters or the population of six telomere adapters can be any of those defined above with reference to the method of the present invention. The one or more splint polynucleotides or the one or more polynucleotide extensions can be any of those defined above with reference to the method of the present invention. The one or more telomere adapters and the one or more polynucleotide extensions preferably contain click chemistry groups.

[0270] The kit of the present invention preferably further comprises one or more sequencing adapters. The one or more sequencing adapters can be any of those discussed above with reference to the method of the present invention.

[0271] The kit preferably further comprises a polymer-induced effector protein and one or more guide polymers. The polymer-induced effector protein and the one or more guide polymers can be any of those discussed above with reference to the method of the present invention. The polymer-induced effector protein is preferably a Cas9 protein. The one or more guide polymers are preferably one or more guide RNAs.

[0272] The present invention also provides a system for carrying out the method of the present invention. The system is for characterizing at least a part of a telomere. The system comprises: (a) one or more polynucleotide telomere adapters of the present invention or a population of six telomere adapters of the present invention; and (b) a nanopore. Any of the above-described embodiments related to the method of the present invention are equally applicable to the system of the present invention.

[0273] The nanopores are preferably present within the membrane. Suitable membranes have been described above. The system can include any of the membranes disclosed above, such as an amphiphilic layer, a triblock copolymer membrane, or a solid state layer. The membrane is typically part of an array of membranes, each of which preferably contains nanopores. The array can be any of those described in WO2018 / 060740, which is hereby incorporated by reference in its entirety.

[0274] The system is preferably adapted to apply a voltage across the membrane and take one or more electrical measurements. Suitable adaptations are discussed in WO2018 / 060740, which is hereby incorporated by reference in its entirety.

[0275] The system can further include a polynucleotide-binding protein. The kit can further include microparticles. Any of the above-described embodiments related to the method of the present invention are equally applicable to the system of the present invention.

[0276] The system can further include one or more primer polynucleotides, one or more polynucleotide extensions, and / or one or more sequencing adapters. Any of the above-described embodiments related to the method of the present invention are equally applicable to the system of the present invention.

[0277] The system or kit may additionally include one or more other reagents or instruments that enable the implementation of any of the above-mentioned embodiments. Such reagents or instruments may include, hereinafter, suitable buffer(s) (aqueous solution), means for obtaining a sample from a subject (such as an instrument including a container or a needle), means for amplifying and / or expressing a polynucleotide, a membrane as defined above, or one or more of a voltage or patch clamp device. The reagent may be present in the system or kit in a dry state such that a fluid sample is used to resuspend the reagent. The system or kit may also optionally include instructions for enabling the use of the system or kit in the methods described herein, or details regarding the organism for which the method may be used. The system or kit may include a magnet or an electromagnet. The system or kit may optionally include nucleotides.

[0278] The following examples illustrate the present invention. It should be understood that while specific embodiments, specific configurations, and materials and / or molecules are being considered herein for the methods according to the present invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of the present invention. The following examples are provided to better illustrate specific embodiments and should not be considered as limiting the present application. The present application is limited only by the claims.

[0279] Example 1 This example demonstrates the assembly of the telomere adapter described herein and its use in combination with a sequencing adapter for nanopore sequencing. The sequencing adapter is attached via a ligation reaction.

[0280] Assembly and Ligation of Telomere Adapter to Chromosome Ends (As shown in Figure 1) The telomere adapter was assembled using the following sequences in Table 3. [Table 3]

[0281] The underlined bases hybridize with the G-rich strand. 1) T1 to T6 at 1 μM were mixed in equal amounts to prepare "Telo-Mix". 2) The ligation reaction products were summarized in Table 4 below. [Table 4] 3) Incubation was carried out at 35 °C for 16 hours, and then T4 was heat-inactivated at 65 °C for 10 minutes.

[0282] The protocol includes an optional step of digesting adapter-ligation DNA: 1) 4 μL of EcoRI was added to 200 μL of the ligation reaction product. 2) Incubation was carried out at 37 °C for 15 minutes.

[0283] Annealing of the "sprint" oligo Annealing of the sprint oligo was carried out as follows (also shown in Figure 2). 1) The above assembly reaction product was washed with 0.4X SPRI beads. 2) The DNA was eluted in 190 μL of DI water (37 °C, 15 minutes, mixing occasionally). 3) 10 μL of NaCl (1 M) solution was added to obtain a final NaCl concentration of 50 mM. 4) 2 μL of 10 μM "sprint" oligo was added to obtain a final "sprint" concentration of 100 nM. [Table 5]

[0284] The underlined bases can hybridize with the nanopore sequencing adapter. 5) Incubation was carried out at 50 °C for 1 hour. 6) The reaction product was washed with 0.4X SPRI beads. 7) Elution in 30 μL of water. (37 °C, 15 minutes, mixing occasionally).

[0285] Ligation of Nano-Pore Array Determination Adapter 1) The sequencing adapter ligation reaction was assembled as follows using the Adapter Mix II extension kit from Oxford Nanopore Technologies (Figure 3): [Table 6] 2) Incubation was carried out at room temperature for 20 minutes. 3) The reaction was washed with 20 μL of AXP beads and LFB. 4) Elution in 15 μL of elution buffer (37 °C, 15 minutes).

[0286] Nanopore sequencing was performed using the GridION device from Oxford Nanopore Technologies equipped with a FLO-MIN106 flow cell.

[0287] Example 2 This example demonstrates the assembly and use of the telomere adapter described herein in combination with a sequencing adapter for nanopore sequencing. The sequencing adapter is attached via a 'click' chemical reaction.

[0288] The telomere adapter was assembled using the following sequences (Figure 4). 1) Click all telomere adapters with 5’ DBCOTEG (Table 7): [Table 7]

[0289] The underlined bases hybridize with the G-rich strand. 2) 1 μM of T1 - T6 were mixed in equal amounts to prepare 'Telo-Mix'. 3) The ligation reaction was assembled as follows (Table 8): [Table 8] 4) Incubation was carried out at 35 °C for 16 hours, and then T4 was heat-inactivated at 65 °C for 10 minutes. 5) 1 μL of NEB Exonuclease I was added to the reaction. 6) Incubation was carried out at 37 °C for 15 minutes, followed by 80 °C for 15 minutes.

[0290] The protocol may include an optional step of digesting the "Telo Mix" ligation DNA. 1) 4 μL of EcoRI was added to 200 μL of the ligation reaction. 2) Incubation was carried out at 37 °C for 15 minutes.

[0291] Addition of "Click" nanopore sequencing adapter 1) The above reaction was washed with 0.4X SPRI beads. 2) Elution in 12 μL of elution buffer. (37 °C, 15 minutes) 3) 1 μL of the sequencing adapter RAP-T from the Oxford Nanopore Technologies kit SQK-PCS111 was added. 4) Incubation was carried out at room temperature for 10 minutes.

[0292] Nanopore sequencing was performed using an Oxford Nanopore Technologies GridION device fitted with a FLO-MIN106 flow cell.

[0293] Example 3 This example demonstrates the assembly of the biotinylated telomere adapter described herein and its use in combination with a sequencing adapter for nanopore sequencing. The sequencing adapter is attached via a ligation reaction.

[0294] The biotinylated telomere adapter can anneal at the chromosome ends and enable sequencing of both strands.

[0295] As shown in Fig. 6, the biotinylated telomere adapter was assembled using the following array. 1) Biotinylated telomere adapter (Table 9): [Table 9]

[0296] The underlined bases hybridize with the G-rich strand. 1) 1 μM of T1 and T2 were mixed in equal amounts to prepare "Telo-Mix". 2) The annealing reaction mixture was assembled as follows (Table 10): [Table 10] 3) Incubation was carried out at 35 °C for 16 hours. 4) The streptavidin Dynabeads® biotinylation pull-down protocol was followed. The telomere strand was eluted in 48 μL of water by heating the beads at 50 °C.

[0297] At this point, the telomere reads of interest were concentrated in the supernatant, the telomere adapter was attached to the beads, and separated from the free-floating strand with the telomere overhang. The protocol proceeds using the supernatant according to the ONT "Genomic DNA By Ligation (SQK-LSK110)" protocol. 5) The telomere overhang was filled using the NEBNext® FFPE DNA Repair Mix and the NEBNext® Ultra™ II End Repair / dA-Tailing Module (Table 11). [Table 11]

[0298] The reaction mixture was incubated using a thermal cycler at 20 °C for 5 minutes and at 65 °C for 5 minutes. 6) The above reactants were washed with 1X SPRI beads and 70% ethanol. 7) DNA was eluted in 61 μL of DI water (room temperature, 10 minutes, mixed occasionally). 8) The sequencing adapter ligation reactants were assembled as follows (Table 12) using sequencing adapters from Oxford Nanopore Technologies' LSK109-112 kit:

Table 12

[0299] The ends were dA-tailed and the ONT sequencing adapter was attached by ligation. Sequencing was performed using an Oxford Nanopore Technologies GridION device fitted with a FLO-MIN106 flow cell.

[0300] Example 4 This example demonstrates the analysis of data obtained from the experiments described in Example 1.

[0301] The data output in Fast5 file format was base-called using Bonito basecaller* (Oxford Nanopore Technologies) with a telomere-trained model.

[0302] *https: / / github.com / nanoporetech / bonito.

[0303] Potential telomere reads were identified in base space using a noise-canceling repeat finder with a minimum motif repeat length of 500 bp. Reads were classified using a custom script and the order of the reads was determined in advance by the nature of the chemistry.

[0304] Reads were mapped to CHM13 v1.1 and an alternative subtelomere assembly with a minimap aligner was added to enable anchoring of reads from subtelomeric regions to the P and Q arms.

[0305] The mapped reads need to have telomeric and subtelomeric alignments. At least 20% of the reads are aligned to the criteria considered in future steps.

[0306] When a read has multiple mapping alignments, a single alignment is selected by choosing the longest read coverage and identity.

[0307] Methylation was called using the Remora tool (Oxford Nanopore Technologies). https: / / github.com / nanoporetech / remora.

[0308] The results are shown in Figures 8 - 13.

[0309] Example 5 Table 13 below shows the total number of telomeric reads and the length of the subtelomere important for uniquely mapping reads obtained by several different methods including the restriction digestion method and the "Teloseq" method described herein. MobI + AluI has the shortest subtelomere length that makes it difficult to uniquely map reads, which is evident from the percentage mapped. [Table 13]

[0310] Example 6 This example demonstrates the use of a customized base caller to identify and map telomeric sequences enriched using the method of the present invention.

[0311] The Bonito-based calling model was trained to be specifically optimized for telomere repeats and reads filtered using a noise-canceling repeat finder (Figure 14a). Isolated telomere reads were aligned using minimap2 against the CHM13 v2.0 reference and alternative subtelomere assemblies. Alignments to subtelomere regions were used to anchor reads to specific chromosomal arms. Multiple telomere enrichment methods were compared using the human cell line HG002 (Figure 14b). The "Teloseq" method of the present invention was compared to other approaches including restriction digestion with AluI and MboI that do not cleave telomeres, Cas9-based sequencing using guide RNAs targeting subtelomere regions, and whole-genome sequencing. Figure 14b demonstrates that the "Teloseq" method of the present invention enabled the identification of the largest number of telomere sequences.

[0312] Example 7 Using sequencing data obtained using an Oxford Nanopore Technologies R10 flow cell, telomere sequencing was performed using the techniques described above.

[0313] The Bonito-based calling model was trained to be specifically optimized for telomere repeats using a stringent filtering stage used to isolate telomere motifs. Isolated telomere reads were aligned using minimap2 to a custom reference including CHM13 v2.0, HG002, and an alternative subtelomere assembly that was virtually digested and retained the P and Q arms (Figure 15a).

[0314] Using the human cell line HG002, 23,236 reads containing telomere motifs with at least 50 bp of continuous repeats were captured, and 18,519 telomere reads were uniquely mapped to chromosomal arms (Figure 15b). The mapped telomere reads demonstrated a median telomere length of 3,826 bp (Figure 15c).

[0315] Example 8 This example shows the fading of telomere reads that can provide valuable telomere variant information.

[0316] HG002 telomere reads were uniquely mapped to the faded HG002 assembly of 88 autosomes and haplotyped chromosomal arms, and successful alignment to 86 arms was achieved. In chromosome 18.q, telomere SNP variants were observed only in the maternal copy and highlighted in dark gray (Figure 16a). When applying a mapping quality threshold of 30 or higher, each haplotyped chromosomal arm has a coverage of approximately 200-fold (Figure 16b). Poor-quality mappings were observed for certain chromosomes such as 1.p and 21.p. The light boxes indicate chromosomes of the paternal haplotype that were confirmed not to uniquely map the telomere alignment to the reference (chr13, chr22).

[0317] Example 9 This example demonstrates the mapping of haplotype telomere reads to provide high resolution of telomere length measurement.

[0318] HG002 telomere reads were uniquely mapped to the faded HG002 assembly to evaluate telomere length (Figure 17). Without fading, the average P-arm telomere length is 3,768 bp and the average Q-arm telomere length is 4,062 bp. On average, the delta between the maternal and paternal telomere lengths is 91. However, when the haplotype is divided by chromosomal arm, there is a significant difference between paternal haploid arms (Δ523). The average paternal P-arm is 3,690 bp and the average paternal Q-arm is 4,213 bp. There was no significant difference in the telomere length of the maternal haploid arms (Δ73). The average maternal P-arm is 3,850 bp while the average maternal Q-arm is 3,923 bp.

[0319] Example 10 Example 1 was repeated using the telomere adapters shown in Table 14 and the sprint polynucleotides shown in Table 15. [Table 14] [Table 15]

Claims

1. 1. A method for characterizing at least a portion of a telomere, comprising: (a) ligating a polynucleotide telomere adaptor to the 5' end of a non-overhanging strand at the end of the telomere, wherein the 3' end of the adaptor specifically hybridizes to a first portion of the overhanging strand and the 5' end of the adaptor does not hybridize to an opposite portion of the overhanging strand; (b) using the telomere adapter to characterize the ligated non-overhanging strand of the at least a portion of the telomere in a 5' to 3' direction from the end of the telomere; characterizing the at least a portion of the telomere comprises (a) sequencing the at least a portion of the telomere; The method.

2. 2. The method of claim 1, wherein characterizing the at least a portion of the telomere further comprises: (a) measuring the length of the at least a portion of the telomere; (b) intertelomeric assembly of a chromosome or genome; (c) identifying telomeres or chromosomal fusions; (d) identifying one or more modifications in the at least a portion of the telomere; (e) identifying the at least a portion of the telomere as a variant; or (f) linking the at least a portion of the telomere to a particular cell or tissue type.

3. 2. The method of claim 1, wherein the method is for characterizing (i) all of the telomeres, (ii) all of the telomeres and at least some or all of the subtelomeres, (iii) all of the telomeres, all of the subtelomeres, and at least some or all of the chromatin, (iv) all of the telomeres, all of the subtelomeres, all of the chromatin, and at least some or all of the opposite subtelomere, or (v) all of the telomeres, all of the subtelomeres, all of the chromatin, all of the opposite subtelomere, and at least some or all of the opposite telomere.

4. 10. The method of claim 1, wherein the method is for characterizing at least a portion of one or both telomeres on each of two or more different chromosomes.

5. 2. The method of claim 1, wherein step (a) further comprises hybridizing a splint polynucleotide to the 5' end of the telomere adapter, and optionally, the splint polynucleotide is compatible with a sequencing adapter.

6. 2. The method of claim 1, wherein step (a) further comprises attaching a sequencing adapter to the telomere adapter and, if present, to the splint polynucleotide, and wherein step (b) comprises using the sequencing adapter to characterize the ligated non-overhanging strand of the at least a portion of the telomere in a 5' to 3' direction.

7. wherein step (a) further comprises ligating or covalently attaching a polynucleotide extension to said telomeric adapter, or step (a) uses a telomeric adapter further comprising a polynucleotide extension at its 5' end; Optionally, (i) the polynucleotide extension comprises a sequencing adaptor, or step (a) further comprises covalently attaching a sequencing adaptor to the polynucleotide extension, and step (b) comprises using the sequencing adaptor to characterize the ligated non-overhanging strand of the at least a portion of the telomere in a 5' to 3' direction from the end of the telomere; and / or ii) the polynucleotide extension is covalently attached to the telomere adaptor or the sequencing adaptor using click chemistry; The method of claim 1.

8. 2. The method of claim 1, wherein the telomere adapter comprises biotin, and optionally, step (a) further comprises enriching the ligated non-overhanging strand of the at least a portion of the telomere using the biotin. (i) the 3' end of the telomere adapter is at least 5 or at least about 7 nucleotides in length; and / or (ii) one or more linkers or spacers are present between the 3' end of the telomere adaptor and the 5' end of the telomere adaptor, and optionally the linkers are flexible linkers or spacers; and / or (iii) the method comprises, prior to step (a), contacting the telomere with a population of six telomere adapters, each having a 3' end that specifically hybridizes to one of six possible sequences of the first portion of the overhanging strand and a 5' end that does not hybridize to the opposite portion of the overhanging strand; and / or (iv) the ligated non-overhanging strands are characterized using a nanopore; and / or (v) the method does not involve (i) restriction digestion, and / or (ii) amplifying said at least a portion of a telomere or polymerase chain reaction (PCR); and / or (vi) the method further comprises characterizing the unligated overhanging strand of the at least some of the telomeres in a 5' to 3' direction to the end of the telomere, and optionally the method comprises using a polymer-induced effector protein to create a double-stranded break from the telomere end to the opposite end of the at least some of the telomeres and attaching a sequencing adaptor to the opposite end, and using the sequencing adaptor to characterize the unligated overhanging strand of the at least some of the telomeres in a 5' to 3' direction to the end of the telomere; and / or (vii) the method is repeated at the other end of the chromosome, wherein the method comprises characterizing both strands of the entire chromosome. The method of claim 1.

10. 1. A method for characterizing at least a portion of a telomere, the method comprising: (a) ligating a polynucleotide telomere adaptor to a 5' end of a non-overhanging strand at an end of the telomere, wherein a 3' end of the adaptor specifically hybridizes to a first portion of the overhanging strand and a 5' end of the adaptor does not hybridize to an opposite portion of the overhanging strand; (b) using a polymer-derived effector protein to create a double-stranded break from the telomere end to an opposite end of the at least portion of the telomere and attaching a sequencing adaptor to the opposite end; and (c) using the telomere adaptor to characterize the ligated non-overhanging strand of at least a portion of the telomere in a 5' to 3' direction from the end of the telomere and using the sequencing adaptor to characterize the unligated overhanging strand of the at least a portion of the telomere in a 5' to 3' direction to the end of the telomere.

11. A polynucleotide telomere adaptor, wherein a 3' end of the adaptor is configured to specifically hybridize to a first portion of an overhanging strand at an end of a telomere and a 5' end of the adaptor is configured not to hybridize to an opposite portion of the overhanging strand at said end of the telomere; and one or more linkers or spacers are present between the 3' end of the adaptor and the 5' end of the adaptor, and optionally the telomere adaptor comprises biotin, and / or the 3' end of the telomere adaptor is at least 5 nucleotides or at least about 7 nucleotides in length. The polynucleotide telomere adaptor.

12. a population of six telomere adapters, each having a 3' end configured to specifically hybridize to one of six possible sequences of a first portion of an overhanging strand at the end of a telomere and a 5' end configured not to hybridize to an opposite portion of the overhanging strand at the end of the telomere; one or more linkers or spacers are present between the 3' end of the adaptor and the 5' end of the adaptor, and optionally each telomeric adaptor comprises biotin, and / or the 3' end of each telomeric adaptor is at least 5 nucleotides or at least about 7 nucleotides in length. A group of the six telomere adaptors.

13. 13. A kit for characterizing at least a portion of a telomere, comprising: (a) one or more polynucleotide telomere adaptors according to claim 11 or a population of six telomere adaptors according to claim 12; and (b) one or more splint polynucleotides or one or more polynucleotide extensions, optionally (i) the telomeric adaptor(s) and the one or more polynucleotide extensions comprise click chemistry groups; and / or (ii) the kit further comprises one or more sequencing adaptors; and / or (iii) the kit further comprises a polymer-derivatized effector protein and one or more guide polymers; The kit.

14. 13. A system comprising: (a) one or more polynucleotide telomere adaptors according to claim 11 or a population of six telomere adaptors according to claim 12; and (b) a nanopore.