Method to measure allele-specific telomere length
Telogator2 leverages long-read sequencing to accurately measure allele-specific telomere lengths and TVR sequences, addressing the limitations of existing methods by providing high-resolution characterization of individual telomeres and TVRs, enhancing the understanding of aging and disease-related telomere dynamics.
Patent Information
- Application Number
- PCT/US2025/029136
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-17
- Filing Date
- 2025-05-13
- Publication Date
- 2025-11-20
AI Technical Summary
Current methods for measuring telomere lengths and characterizing telomere variant regions (TVRs) are limited by low throughput and resolution, failing to accurately account for individual telomere lengths and TVR sequences, which affects the precision of telomere length measurements and hinders understanding of their role in aging and disease.
A method called Telogator2, which utilizes long-read sequencing data to identify and characterize telomere variant repeats (TVRs), allowing for high-resolution measurement of allele-specific telomere lengths (ATL) by clustering reads based on TVR sequences and aligning them to subtelomere boundaries, thereby determining the length of each of the 92 telomeric alleles.
Enables accurate, high-throughput measurement of individual telomere lengths at nucleotide resolution, facilitating the identification of chromosome-end fusions and assessing genomic instability, tumorigenesis, and providing deeper insights into aging and disease-related telomere dynamics.
Smart Images

Figure IMGF000027_0001 
Figure IMGF000030_0001 
Figure IMGF000033_0001
Abstract
Description
METHOD TO MEASURE ALLELE-SPECIFIC TELOMERE LENGTHFIELD
[0001] The present disclosure relates to methods for nucleic acid analysis. In particular, the disclosure relates to methods of measuring individual telomere lengths and identifying their physically linked telomere variant sequences (TVRs). The disclosure falls within the fields of molecular biology, computational genomics, and medicine.BACKGROUND
[0002] Telomeres are regions of repetitive DNA at the ends of linear chromosomes which are part of a biological system for protecting chromosome ends from degradation. Human telomeric DNA is comprised of (TTAGGG)n repeats, referred to as ’canonical repeats,’ which shorten with cell division in normal somatic cells. Telomeres can shorten to a point where they no longer protect chromosome ends, initiating cellular senescence, a state where cells do not further divide but retain biological function [1 , 2],
[0003] The genomic regions adjacent to telomeres include telomere variant repeat (TVR) and subtelomere regions. In humans, subtelomeres are informally defined as the most distal ~ 500kb of each chromosome arm and are repeat-rich and highly variable across populations [3, 4], TVR regions are found between telomeres and subtelomeres (Figure. 1) and are composed of canonical repeats interspersed with blocks of ’variant repeats.’ The most frequent variant repeats are canonical TTAGGG sequences modified by a single base substitution (e.g. TCAGGG) or insertion / deletion (e.g. TTAAGGG). Early analysis of TVR regions identified the most common variant repeats [5], and subsequent studies reported a large diversity of repeat patterns both within individuals and across distantly related samples [6-8]. TVR regions are described as not having telomeric function due to the reduced binding affinity of variant repeats to shelterin genes TRF1 and TRF2 [9-11]. From targeted studies of telomeres 12q and Xp / Yp, it was found that the distributions of variant repeats within human TVR regions are highly variable across chromosome arms, but also subject to Mendelian inheritance. Thus, it is thought that TVR regions exhibit relatively high de novo mutation rates and are specific to individual chromosomes [12, 13].
[0004] While the size and sequence composition of TVR regions has been studied in detail at a limited number of chromosome arms, TVR regions have not been broadly characterized genome-wide due to the lack of high-throughput methods capable of analyzing telomere sequences of individual alleles. The variable size and composition of TVR regions potentially impacts analyses of telomeres themselves: e.g., if a TVR region does not have telomeric function but is considered to be part of a telomere, from the perspective of atelomere length (TL) measurement method, then the TLs reported by the method will be an overestimate proportional to the size of the TVR region
[0014] ,SUMMARY
[0005] One aspect of this disclosure is a method that enables high-resolution characterization of individual telomeres and their adjacent TVR regions and that leverages TVRs to determine telomere lengths.
[0006] The present disclosure provides, in one embodiment, a method for identifying the hypervariable telomere variant repeat (TVR) sequences that are linked to each telomeric allele and leveraging those TVR sequence to determine allele specific telomere length (ATL). The methods disclosed herein include, for example, identifying TVR sequences, clustering telomeric reads using the TVR sequences. The methods disclosed here may further include aligning the clustered reads to the subtelomere-telomere boundary, thereby providing an accurate, high-throughput approach for estimating the ATL. The estimated ATLs may comprise all 92 telomeric alleles.
[0007] Accordingly, in one embodiment provided herein is a method for determining allelespecific telomere length (ATL), the method comprising obtaining a long read sequencing data set. The method may further comprise extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set. The method may further comprise identifying TVR regions by locating boundaries between telomere repeat regions and subtelomere regions. The method may further comprise processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR). The method may further comprise grouping the cTVR sequences into allele-specific read clusters. The method may further comprise determining ATL from the read clusters.
[0008] The disclosure further provides in one embodiment a method for determining allele-specific telomere length (ATL) from allele-specific telomere variant repeat (TVR) sequences, the method comprising obtaining a long read sequencing data set. The method may further comprise extracting from the long read sequencing data set a subset of sequencing reads identifying telomere region boundaries. The method may further comprise identifying individual alleles based on the presence of TVRs, comprising: (i) querying sequencing extracted reads of (b) for matches to a known set of telomere repeats, (ii) identifying TVR regions in the extracted sequencing reads, and (iii) clustering the extracted sequencing reads based on their TVR sequences and subtelomere sequences. The method may further comprise assigning each telomere allele from (c) to a chromosome arm by aligning the subtelomere sequences to a collection of subtelomere reference sequences.The method may further comprise producing a report containing the lengths and sequences of each telomere alleles, thereby determining allele-specific telomere length (ATL) and allele-specific telomere variant repeat (TVR) sequences from a biological sample.
[0009] In some embodiments, obtaining a long read sequencing data set comprises any one or more of (a) collecting a biological sample and running one or more long read sequencing reactions, and / or (b) collecting a long read sequencing data set from a database.
[0010] In some embodiments, the database comprises assembled genome data or raw sequence read data.
[0011] In various embodiments, the data is in Fasta, Fastq, BAM or CRAM formats.
[0012] In one embodiment, the biological sample is from a human.
[0013] In some embodiments, the subset of sequencing reads extracted in (b) comprises at least approximately 5, 10, 15, 20, or 25 or more copies of canonical telomere repeats.
[0014] In some embodiments, the at least one reference genome containing subtelomere sequences comprises a telomere-to-telomere reference genome. In various embodiments, the telomere-to-telomere reference genome is selected from the group consisting of T2T- CHM13, T2T-Yao, T2T-CN1 , T2T-HG002, T2T-ksa001 , and T2T-i002c.
[0015] In some embodiments, operations (b) - (e) are performed using a suitably programmed computer.
[0016] The present disclosure also provides in one embodiment a method for measuring all 92 allele-specific telomere lengths at nucleotide resolution, the method comprising: (a) obtaining a long read sequencing data set; (b) extracting from the long read sequencing data set a subset of sequencing reads identifying telomere region boundaries; (c) identifying individual alleles based on the presence of TVRs, comprising: (i) querying sequencing extracted reads of (b) for matches to a known set of telomere repeats, (ii) identifying TVR regions in the extracted sequencing reads, and (iii) clustering the extracted sequencing reads based on their TVR sequences and subtelomere sequences; (d) assigning each telomere allele from (c) to a chromosome arm by aligning the subtelomere sequences to a collection of subtelomere reference sequences; (e) producing a report containing the lengths and sequences of all anchored telomere alleles, thereby determining allele-specific telomere length (ATL) and allele-specific telomere variant repeat (TVR) sequences of a DNA molecule from a biological sample, thereby measuring all 92 allele-specific telomere lengths at nucleotide resolution.
[0017] Further provided is a method of identifying individual alleles based on the composition of telomere variant repeat sequences in a sample, the method comprising: (a) obtaining a long read sequencing data set; (b) extracting from the long read sequencing data set a subset of sequencing reads identifying telomere region boundaries; (c) identifying individual alleles based on the presence of TVRs, comprising: (i) querying sequencing extracted reads of (b) for matches to a known set of telomere repeats, (ii) identifying TVR regions in the extracted sequencing reads, and (iii) clustering the extracted sequencing reads based on their TVR sequences and subtelomere sequences; (d) assigning each telomere allele from (c) to a chromosome arm by aligning the subtelomere sequences to a collection of subtelomere reference sequences; (e) producing a report containing the lengths and sequences of all anchored telomere alleles, thereby determining allele-specific telomere length (ATL) and allele-specific telomere variant repeat (TVR) sequences of a DNA molecule from a biological sample, thereby identifying individual alleles based on the composition of telomere variant repeat sequences in a sample.
[0018] The disclosure also provides in one embodiment a method for assigning TVR and telomere regions to their respective allele chromosome arm, the method comprising: (a) obtaining a long read sequencing data set; (b) extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set; c) identifying TVR regions by locating boundaries between telomere repeat regions and subtelomere regions; d) processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR); (e) grouping the cTVR sequences into allele-specific read clusters; (f) determining ATL from the read clusters; and (g) identifying the respective allele chromosome arm from the read clusters that extend into the subtelomere regions.
[0019] Also provided herein is a method for determining the degree of relationship between individuals, the method comprising: (a) obtaining a long read sequencing data set from at least two individuals; (b) extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set; c) identifying TVR regions by locating boundaries between telomere repeat regions and subtelomere regions; d) processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR); (e) grouping the cTVR sequences into allelespecific read clusters; (f) determining ATL from the read clusters; and (g) determining the degree of relationship between individuals.
[0020] The disclosure further provides in one embodiment a method for assessing the risk of genomic instability and tumorigenesis by determining the amount of telomere and TVR erosion, the method comprising: (a) obtaining a long read sequencing data set; (b)extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set; c) identifying TVR regions by locating boundaries between telomere repeat regions and subtelomere regions; d) processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR); (e) grouping the cTVR sequences into allele-specific read clusters; (f) determining ATL from the read clusters; and (g) determining the amount of telomere and TVR erosion.
[0021] Also provided herein is a method for identifying chromosome-end fusions, the method comprising: (a) obtaining a long read sequencing data set; (b) extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set; c) identifying TVR regions by locating boundaries between telomere repeat regions and subtelomere regions; d) processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR); (e) grouping the cTVR sequences into allele-specific read clusters; (f) determining ATL from the read clusters; and (g) identifying chromosome-end fusions.
[0022] In some embodiments, the chromosome-end fusions are markers for genomic instability and / or tumorigenesis.
[0023] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the invention will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 illustrates the structure of example telomeres, which are the roughly IQ- 20 kb repetitive regions capping the end of eukaryotic chromosomes, and the telomeric region that is used to determine allele-specific telomere length (ATL) by the methods disclosed herein. The schematic diagram depicts the structure of: i.) example telomeric regions, which comprise (TTAGGG)ncanonical repeats; ii.) an example subtelomeric region, which is immediately adjacent to the telomeric repeats; this region is informally defined as the most distal ~ 500kb of each chromosome arm and is repeat-rich and highly variable across populations; iii.) an example region between telomeres and subtelomeres defined as a telomere variant repeat (TVR) region, which is composed of canonical repeats interspersed with blocks of variant repeats; and iv.) example subtelomere / telomere and TVR / telomere boundaries. In the presently disclosed methods, allele specific telomere length may be determined by mapping reads to the TVR / telomere boundaries boundary point of the reference genome sequence.
[0025] Figure 2 is an overview of “Telogator2” - a method for measuring allele-specific telomere length (ATL) and identifying their physically linked TVR sequences from long read sequencing data, consistent with some embodiments. Figure 2A is a workflow diagram outlining 8 operations for identifying individual alleles based on their composition of telomere variant repeats (TVRs) in accordance with the presently disclosed methods. Figure 2B is a pictorial diagram for each operation of the bioinformatics analysis pipeline shown in Figure 2A, from the initial input of raw data in various formats to the output of the processed data, in this case, a violin plot of the allele-specific telomere lengths (ATLs) for each chromosome arm.
[0026] Figure 3 provides plots showing the unique TVR patterns for each allele and how these are processed bioinformatically. Figure 3A is a plot of all telomere alleles in the functionally haploid CHM13hTERT cell line (CHM13) and their corresponding TVR regions, the repeats in both regions are color coded as indicated in Table 1 . The telomere lengths (TLs) shown represent the 90th-percentile TL at each chromosome arm. Figure 3B shows an example conversion from a nucleotide string i.) to a repeat symbol representation ii.). Figures 3C & D show an example of TVR region clustering and boundary detection. Figure 3C left panel is a dendrogram of a hierarchical clustering of reads based on TVR patterns. Figure 3C right panel shows a clustering of reads from different telomeres of different lengths. Figure 3D shows estimation of TVR region boundary and allele-specific telomere length (ATL).
[0027] Figure 4 is a plot showing the distribution of average TVR region lengths by chromosome arm for the 45 Human Pangenome Reference Consortium (HPRC) samples.
[0028] Figure 5 provides data showing the utility of TVRs in studying heritability and population genomics. Figure 5A is a visualization of TVR and telomere regions at 18q for HG005, HG006 and HG007 Genome in a Bottle Consortium (GIAB) standard reference samples. The figure shows the utility of TVRs in inferring allelic parent-of-origin, wherein allele (1 ) is paternally inherited, and allele (2) is maternally inherited, making TVRs suitable for use in estimating heritability. Figure 5B is a phylogenetic tree resulting from the clustering of TVRs from the same 12p allele across an ethnically diverse population of samples from the Human Pangenome Reference Consortium (HPRC), illustrating the utility of TVRs in studying population genetics (i.e., comparative / population genomics).
[0029] Figure 6 is a violin plot showing the distribution of allele-specific telomere lengths (ATLs) at each chromosome arm of the CHM13 haploid cell line.
[0030] Figure 7 is a scatterplot showing the correlation between ATLs reported by Telogator2 to T2T telomere reference annotations.
[0031] Figure 8 is a visualization of TVR and telomere regions in the HG005 GIAB standard reference sample that have variation with respect to the corresponding allele in the parent.
[0032] Figure 9 shows a cluster of reads from a telomere from the 2p chromosome arm of the HG007 GIAB standard reference sample, with two reads exhibiting abnormally long telomeres.
[0033] Figure 10 is a scatterplot showing the correlation between average ATL computed by Telogator2 from long reads vs. tel_content computed by TelomereHunter from short reads.
[0034] Figure 11 is a visualization of a telomere from HG005, 3p. Reads numbered (2), (6), and (5) exhibit a common systematic artifact found in PacBio Hi Fi reads in telomere regions where canonical CCTAA repeats are miscalled as CCCCAA, CCCCCA, and similar variations.
[0035] Figure 12 is an image of a scoring matrix used during multiple sequence alignment of TVR regions.
[0036] Figure 13 is a plot showing the distribution of all ATLs sorted by length for the HG00609 human genome of a Han Chinese male individual.
[0037] Figure 14 shows a comparison of a sample’s chromosome-specific TL (red / dark shade) to an aged-matched average of healthy controls (gray / light shade).
[0038] Figure 15 shows allele telomeres, sorted by length, from two bone marrow samples. One with a normal TL distribution (left) and one with eroded telomeres (right).
[0039] Figure 16 shows an example of comparing chromosome-specific TL on MRC5 cell line at different passages.DETAILED DESCRIPTION
[0040] Methods for measuring TLs are of great medical interest due to the extensive relationships established between telomeres and aging, disease, and behavioral health. Typically, studies involving TL measurements use results from methods that estimate the average length of a sample’s telomeres across all chromosome arms. These average values have proven useful in establishing correlations but are of limited use in characterizing the underlying mechanisms relating telomeres to disease [15, 16]. The length of a sample’s shortest telomere is often more informative than their average TL [17, 18], and the lengths of telomeres on specific chromosome arms have been observed to correlate with disease risk [19-21], Since conventional methods for measuring TLs of individual chromosome arms aregenerally low throughput, low resolution, or labor intensive [14, 22], recent approaches using nanochannel arrays or long read sequencing have been developed to address these limitations [23-25]. Long read sequencing, in particular, has garnered significant interest for telomere analysis, as recent improvements in cost and throughput make it increasingly accessible to clinical laboratories.
[0041] Current long read sequencing platforms can generate high-quality reads ~ 20kb in length that are capable of spanning telomeres and their adjacent regions. Several methods have recently been developed for analyzing telomeres with this data: TLD
[0026] demonstrated the viability of aligning long reads to telomere boundaries for estimating telomere lengths, and EdgeCase
[0024] further showed that hierarchically clustering reads based on their sequence can separate individual telomeres. A method known as “Telogator”
[0027] builds upon these ideas and leverages the telomere-to-telomere (T2T) reference genome to report telomere lengths at each chromosome arm. ChArmTelo
[0028] is a recently described method for analyzing telomeres at individual chromosomes using 10X linked-reads, however the software is not publicly available and the protocols for sequencing 10X linked-reads is no longer supported. A new method for selectively enriching telomere regions using ’telobaits’ prior to long read sequencing has shown potential in reducing the cost of large studies via sample multiplexing
[0025] . However, the 5-8kb DNA fragments generated by this approach are limited in their ability to span entire telomeres or to extend far enough into subtelomeres such that reads can be uniquely mapped to specific chromosome arms. Additionally, the associated bioinformatics workflow only reports average TL alongside a simplified representation of TVR sequence composition that is limited in its ability to represent the variety of TVR patterns found in human telomeres.
[0042] To address the aforementioned unmet needs and facilitate high-resolution characterization of individual telomeres and their adjacent TVR regions, a method described and referred to as “Telogator2” - a method for measuring allele-specific telomere length (ATL) and identifying their physically linked TVR sequences from long read sequencing data - is provided herein. Telogator2 significantly expands the functionality of previously released Telogator, going beyond chromosome arm specificity to advantageously report TVR and telomere sequences for individual alleles. Advantageously, Telogator2 can identify distinct telomere alleles in the presence of sequencing errors and alignments where reads may be mapped (or multi-mapped) to chromosome arms different from where they originated [58, incorporated by reference herein in its entirety].
[0043] The present disclosure is based in part on the discovery that advantageously allows for the identification and characterization of telomere variant repeats (TVRs) in telomere sequences which enables the accurate, high-throughput measurement of individualtelomere lengths. It has particularly been determined by the present disclosure that by finding reads with telomere sequences that can be aligned to that arm’s subtelomere, the TL at each chromosome arm can be measured from the subtelomere boundary.Advantageously, the disclosed methods may identify telomere alleles even in the presence of sequencing errors and multi-mapped reads resulting from duplication telomeric reads. Duplicated genomic regions represent a challenge when analyzing high-throughput sequencing data as the reads align equally well at more than one such genomic location, i.e., “multi-mapped reads.” When reads cannot be unambiguously aligned to a reference, genes or genetic elements cannot be accurately quantified. The ability to accurately map individual DNA molecules to discrete subtelomere regions provides a unique method of estimating telomere lengths, and ultimately connecting telomere length measurements with individual subtelomere haplotypes
[0023] .
[0044] Thus, some embodiments of the present disclosure provide methods for determining allele-specific telomere length (ATL) (e.g., from specific telomere variant repeat (TVR) sequences). The methods provided herein may also advantageously enable measuring all 92 allele-specific telomere lengths at nucleotide resolution, and identifying individual alleles based on the composition of telomere variant repeat sequences in a sample. Methods for assigning TVR and telomere regions to their respective allele chromosome arm are also provided herein. In other embodiments, the methods provided herein allow determination of the degree of relationship between individuals. The methods further allow one to assess the risk of genomic instability and tumorigenesis by determining the amount of telomere and TVR erosion.
[0045] As provided herein, the disclosure may be used to demonstrate that the TVR regions that are located between telomeres and subtelomeres are in most cases unique to each telomere allele (92 total). The methods provided herein thus may increase the resolution at which telomere length can be estimated. Indeed, telomere length measurement was conventionally restricted to the average telomere length. The unique properties of the TVRs (i.e. being unique) enables leveraging long-read sequencing and bioinformatics method operations as described herein to accurately measure the length of the 92 allele telomeres.
[0046] In some embodiments, the present methods leverage the TVR regions that are unique to each allele telomere to cluster telomeric reads. Within each cluster, one can: (1) derive the estimated length of the allele telomere from the canonical repeat section of the reads belonging to the cluster, and (2) use the section of the reads that extends in the subtelomeres to determine the chromosome arm provenance of the cluster. From theprovenance and the length obtained in the aforementioned operations, one can define the telomere length of the 92 allele telomeres.
[0047] Measuring allele telomere lengths instead of average telomere length may be desirable in some instances. For example, telomere shortening is related to an important physiological mechanism: senescence. There is growing evidence that senescence is triggered by the shortest telomere(s). Therefore, allele telomere lengths should provide deeper insight into the state of telomeres in a sample than the average telomere does. This information should be of interest to aging research. Allele telomeres length is also relevant to conditions and diseases such as Telomere Biology Disorders (TBD) that affects patients with short telomeres as well as cancers since abnormally short telomeres have been associated with genomic instability and tumorigenesis.
[0048] In some embodiments, the methods described herein include extracting from a long-read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set. In one embodiment, the sequencing reads and boundaries are identified using telomere canonical repeats.
[0049] In some embodiments, the methods include a operation whereby TVR regions are “processed” to, for example, identify and remove sequencing errors and to identify consensus TVR (cTVR) sequences. In this way, “processing” includes multiple sequence alignment, and filtering to remove isolated incidences of variant repeats that are attributable to sequencing artifacts.
[0050] The methods provided herein use, in some embodiments, assembled genome data including sequencing data (as a typical input) or previously assembled genome data.
[0051] In still another embodiment, methods applicable to population analyses are provided. For example, a method is provided herein for determining the degree of relationship between individuals. After identifying and processing TVR regions and cTVR regions, the cTVR regions are grouped into allele-specific read clusters (where ATL may be determined). The degree of relationship between individuals is accomplished, according to one embodiment, by comparing the distance between TVR sequences using pairwise alignment.
[0052] The present disclosure also provides methods for identifying telomere erosion and chromosome-end fusions. After identifying and processing TVR regions and cTVR regions, the cTVR regions are grouped into allele-specific read clusters (where ATL may be determined). Determining the presence or amount of telomere and TVR erosion by measuring the distance of the proximal TVR boundary to the end of the telomere or chromosome-end fusion by reporting the presence of 2 juxtaposed TVRs from different allelein the same read. The 2 TVRs can be (or not) separated by a stretch of canonical repeat. In this way, the methods described herein allows studying of aging and senescence, diagnosis of disorders, and risk of genomic instability and tumorigenesis.
[0053] Unless otherwise indicated, the practice of some specific herein may involve conventional techniques and apparatus commonly used in molecular biology, microbiology, protein purification, protein engineering, protein and DNA sequencing, and recombinant DNA fields that are within the skill of the art. Such techniques and apparatus are known to those of skill in the art and are described in numerous texts and reference works (See e.g., Sambrook et al., “Molecular Cloning: A Laboratory Manual,” Third Edition (Cold Spring Harbor),
[2001] ); and Ausubel et al., “Current Protocols in Molecular Biology”
[1987] ).
[0054] Numeric ranges are inclusive of the numbers defining the range. It is intended that every maximum numerical limitation given throughout this specification includes every lower numerical limitation, as if such lower numerical limitations were expressly written herein. Every minimum numerical limitation given throughout this specification will include every higher numerical limitation, as if such higher numerical limitations were expressly written herein. Every numerical range given throughout this specification will include every narrower numerical range that falls within such broader numerical range, as if such narrower numerical ranges were all expressly written herein.
[0055] Unless defined otherwise herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. Various scientific dictionaries that include the terms included herein are well known and available to those in the art. Although any methods and materials similar or equivalent to those described herein find use in the practice or testing of the embodiments disclosed herein, some methods and materials are described.
[0056] To assist in understanding the present invention, certain terms are first defined. The terms defined immediately below are more fully described by reference to the specification as a whole. Additional definitions are provided throughout the application. It is to be understood that this disclosure is not limited to the particular methodology, protocols, and reagents described, as these may vary, depending upon the context they are used by those of skill in the art.
[0057] As used herein the term “chromosome” refers to the heredity-bearing gene carrier of a living cell, which is derived from chromatin strands comprising DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is employed herein.
[0058] The term “genome,” as used herein, generally refers to genomic information from a subject, which may be, for example, at least a portion or an entirety of a subject's hereditary information. A genome can be encoded either in DNA or in RNA. A genome can include the sequence of all chromosomes together in an organism. For example, the human genome ordinarily has a total of 46 chromosomes. The sequence of all these together may constitute a human genome.
[0059] An “allele” refers to any of two or more alternative forms of a gene that occupy the same locus on a chromosome. If two alleles within a diploid individual are identical by descent (that is, both alleles are direct descendants of a single allele in an ancestor), such alleles are called autozygous. If the alleles are not identical by descent, they are called allozygous. If two copies of same allele is present in an individual, the individual is homozygous for that gene. If different alleles are present in an individual, the individual is heterozygous for that gene.
[0060] As used herein, “telomeric region” means the DNA segment at the ends of a chromosome with repeat telomeric sequence. In the case of vertebrates, it can be the (TTAGGG)n repeat sequence at the ends of chromosomes. The term “telomeric sequence,” as used herein, refers to a repeating sequence found in a chromosomal telomere. The sequence is recognized to vary from one species to the next. In humans and other mammals (as well as most other vertebrates), the sequence is known to be made up of repeats of the sequence TTAGGG. A telomeric sequence may include sequence differences from the typical repeat unit which are though to result from errors introduced during DNA replication and / or telomere length extension, whether mediated by telomerase or the ALT pathway.
[0061] As used herein, “sub-telomeric region” or “subtelomere” region means the segment of DNA immediately adjacent to telomere at the centromeric end of telomeres. This signifies chromosomal DNA located adjacent to (preferably in the range of from 100 to 500 bp but can be up to 20 kb, although is preferably up to 1 kb or even 2 kb, from the tandem telomeric repeats of the telomeric DNA. Subtelomeric region often contains degenerate telomeric repeats. In the case of humans, repeats of TGAGGG and TCAGGG can be present in subtelomeric region.
[0062] The expression “telomere variant repeat sequence” or TVR as used herein, refers to non-canonical repeat sequences located between telomeric and subtelomeric regions of the chromosome. “Allele specific variant repeat sequence” or “Allele specific TVR” refers to TVR sequences for each of the 92 telomeric alleles in a cell.
[0063] As used herein, the expression “polynucleotide length” refers to the absolute number of nucleotides in a sequence or in a region of a reference genome. The expression “allele-specific telomere length” or “ATL” refers to the length of a particular telomere allele in a cell of a sample. Healthy diploid human cells contain 46 chromosomes (23 pairs), which have 2 arms each, for a total of 92 telomeres, each with its own length. Accordingly, the term ATL as used herein refers to the length or absolute number of nucleotides for each of those 92 telomeres.
[0064] The term “sequencing,” as used herein, generally refers to methods and technologies for determining the sequence of nucleotide bases in one or more polynucleotides. The polynucleotides can be, for example, nucleic acid molecules such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including variants or derivatives thereof (e.g., single stranded DNA). Sequencing can be performed by various systems currently available, such as, without limitation, sequencing system by Illumina®, Pacific Biosciences (PacBio®), Oxford Nanopore®, Life Technologies (Ion Torrent®), Roche®, Genapsys®, and MGI Tech®. Sequencing may be performed without using nucleic acid amplification. Alternatively, or in addition, sequencing may be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g. digital PCR, quantitative PCR, or real time PCR), or isothermal amplification. Such systems may provide a plurality of raw genetic data corresponding to the genetic information of a subject (e.g., human), as generated by the systems from a sample provided by the subject. In some examples, such systems provide sequencing reads (also “reads” herein). A read may include a string of nucleic acid bases corresponding to a sequence of a nucleic acid molecule that has been sequenced.
[0065] The term “genomic read” or “read” is used in reference to a sequence read of any segment in the entire genome of an individual or from a portion of a nucleic acid sample. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in ATCG) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region, e.g., that can be aligned and mapped to a chromosome or genomic region or gene.
[0066] The term “short read,” as used herein, generally refers to a read length of a DNA or RNA polynucleotide of about 100 to about 600 bp.
[0067] The term “long read,” as used herein generally refers to a read length of a DNA or RNA polynucleotide of greater than about 1 kbp.
[0068] The term “Next Generation Sequencing (NGS)” as used herein refers to sequencing methods that allow for massively parallel sequencing of clonally amplified molecules and of single nucleic acid molecules. Non-limiting examples of NGS include sequencing-by-synthesis using reversible dye terminators, and sequencing-by-ligation.
[0069] As used herein, a “short-read sequencing method” shall be understood to mean sequencing methods which are capable of producing single reads of up to about 1000 bases, such as from about 35 bases to about 1000 bases. However, typically short-read sequencing method produce reads of 500 bases or less. Example short-read sequencing methods are described herein and include next generation sequencing methods such as sequencing-by-hybridization, sequencing-by-synthesis, sequencing-by-ligation platform, combinatorial probe anchor synthesis, and ion semiconductor sequencing. However, chain termination (Sanger sequencing) may also be used to produce short read sequences. In one particular example, a sequencing-by-synthesis method using the Illumina platform is used to produce short read sequences. It also follows that “short-read sequence data” is data comprising and / or relating to sequences of about 35 bases to about 1000 bases in length, and typically 500 bases or less.
[0070] Conversely, a long-read sequencing method shall be understood to mean a method capable of producing sequence reads in excess of 1000 bases. Example long-read sequencing methods are described herein and include nanopore sequencing and single molecule real time (SMRT) sequencing. Of course, it will be appreciated that read length achieved using a long-read sequencing method is also dependent on preparation of the nucleic acid molecule library e.g., cDNA library, not just the sequencing platform. For example, if a library is produced with an average fragment length of about 500 bases, then the average length of reads obtained from such a library will not exceed that length. However, library preparation aside, it will be appreciated that long-read sequencing methods are capable of producing long-read sequences e.g., over 1 Mb. It also follows that “long- read sequence data” is data comprising and / or relating to sequences in excess of 1000 bases in length, such as for example, between 1000 bases and 500 kb.
[0071] In some embodiments, the disclosed methods utilize long-read sequences that potentially can span long repeat sequences having complete telomeric and subtelomeric sequences. In some embodiments, single molecule sequencing or synthetic long-read sequencing are used to obtain long reads.
[0072] An example of a sequencing technology that can be used in the methods of the provided invention includes the single molecule, real-time (SMRT) technology of Pacific Biosciences (Menlo Park, Calif.). In SMRT, each of the four DNA bases is attached to one of four different fluorescent dyes. These dyes are phospholinked. A single DNA polymerase is immobilized with a single molecule of template single stranded DNA at the bottom of a zero-mode waveguide (ZMW). A ZMW is a confinement structure which enables observation of incorporation of a single nucleotide by DNA polymerase against the background of fluorescent nucleotides that rapidly diffuse in and out of the ZMW (in microseconds). It takes several milliseconds to incorporate a nucleotide into a growing strand. During this time, the fluorescent label is excited and produces a fluorescent signal, and the fluorescent tag is cleaved off. Detection of the corresponding fluorescence of the dye indicates which base was incorporated. The process is repeated.
[0073] Another example of a sequencing technique that can be used in the methods of the disclosure is nanopore sequencing (Soni, G. V., and Meller, A., Clin Chem 53: 1996- 2001 (2007)). A nanopore is a small hole, of the order of 1 nanometer in diameter. Immersion of a nanopore in a conducting fluid and application of a potential across it results in a slight electrical current due to conduction of ions through the nanopore. The amount of current which flows is sensitive to the size of the nanopore. As a DNA molecule passes through a nanopore, each nucleotide on the DNA molecule obstructs the nanopore to a different degree.
[0074] As used herein, the terms “alignment” and “aligning” refer to the process of comparing a read to a reference sequence and thereby determining whether the reference sequence contains the read sequence. An alignment process attempts to determine if a read can be mapped to a reference sequence but it does not always result in a read aligned to the reference sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. In some cases, alignment simply indicates whether or not a read is a member of a particular reference sequence (i.e., whether the read is present or absent in the reference sequence). For example, the alignment of a read to the reference sequence for human chromosome 13 will indicate whether the read is present in the reference sequence for chromosome 13. A tool that provides this information may be called a set membership tester. In some cases, an alignment additionally indicates a location in the reference sequence where the read maps to. For example, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13.
[0075] As used herein, the expression “reference genome” or “reference sequence” refers to any genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length human genome. Such sequences may be referred to as genomic reference sequences. In another example, the reference sequence is limited to a specific human chromosome such as chromosome 13. In some embodiments, a reference Y chromosome is the Y chromosome sequence from human genome version hg19. Such sequences may be referred to as chromosome reference sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In some embodiments, the reference genome is T2T-CHM13, T2T-Yao, T2T-CN1 , T2T-HG002, T2T-ksa001 , and / or T2T-i002c.
[0076] In some embodiments, a reference sequence for alignment may have a sequence length from about 1 to about 100 times the length of a read. In such embodiments, the alignment and sequencing are considered a targeted alignment or sequencing, instead of a whole genome alignment or sequencing. In these embodiments, the reference sequence typically include a gene and / or a repeat sequence of interest. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual.
[0077] Aligned reads are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known reference sequence such as a reference genome. An aligned read and its determined location on the reference sequence constitute a sequence tag. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period for implementing the methods disclosed herein. One example of an algorithm from aligning sequences is the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysis pipeline. Alternatively, a Bloom filter or similar set membership tester may be employed to align reads to reference genomes. See U.S. Patent No. 9845552, which is incorporated herein by reference in its entirety. Thematching of a sequence read in aligning can be a 100% sequence match or less than 100% (i.e., a non-perfect match).
[0078] The term “mapping” used herein refers to assigning a read sequence to a larger sequence, e.g., a reference genome, by alignment.
[0079] As used herein, the term “k-mer” refers to substrings of length k in a given string (can be DNA, RNA, protein, or any string sequence), where k is a positive integer. As the number and depth of high-throughput sequencing experiments grow, efficient methods to map, store, and search DNA sequences have become critical for their analysis. Selecting a subset of k-mers in a string in a local manner is a common task in bioinformatics tools for speeding up computation. However, with a large size of k-mers of length k, the number of k- mers is non-trivial and storing them is costly. At the same time, due to the overlap of consecutive k-mers, the information they contain is highly redundant. Therefore, selecting a minimal subset of k-mers in a sequence can for many applications lead to a dramatic increase in computing efficiency while only losing a small amount of information. Arguably the most well-known and common method for selecting subsets of k-mers is the minimizer technique, which generates a small representation of a long sequence. This small representation is designed to preserve some of the structure and information from the original string. Minimizers sample a sequence by comparing the forward and reverse k-mers in a window simultaneously according to a predefined selection scheme and select k-mers, with a minimal value in a window of a specified size, or the “lowest-ordered” k-mers. By storing only these minimal k-mers, the storage cost is significantly reduced while maintaining a similar amount of information. Further, by using minimizers, it is possible to design algorithms that perform operations between sequences, such as sequence search, indexing, and sequence alignment, with reduced storage, memory, or computational costs and with provably little or no cost in accuracy.
[0080] The term “window (w)” as used herein refers to one or more genomic sections chosen for analysis, and sometimes used as a reference for comparison (e.g., used for normalization and / or other mathematical or statistical manipulation). For example, it can refer to a sequence of length (w+k-1) consisting of exactly w k-mers. A “static” window generally includes a predetermined set of genomic sections that do not change during manipulations and / or analysis. A “sliding” or “moving” window includes repeatedly moving or sliding to an adjacent test genomic section.
[0081] As used herein, the expression “sequencing masking” refers to a procedure that identifies and masks sequences that provide low information in a search, such as non- informative repetitive and low-complexity sequences, by substituting an N for the actualnucleotide (i.e. G, A, T, or C). Unmasked, such sequences cause problems for search and clustering algorithms based on matching words or patterns. For example, as mentioned elsewhere, k-mers are used to seed sequence alignments, and they speed up genome sequence analysis. However, sequences that are mostly or all composed of a single letter such as AAAAAA or TTTTCTTT often have high frequencies. K-mers of such sequences with a high frequency may lead to too many false-positive seed matches. The large number of false hits may lead to high memory-usage and runtime. Masking highly repetitive k-mers may help deal with this issue. Low information sequences, although not necessarily informative in comparative analysis, are a part of the actual sequence, and thus are masked in the edited sequence instead of removed so that the low information sequence can be obtained in the database if necessary. This masks the low information sequences for search purposes but preserves the spacing of the DNA molecule. The actual sequences corresponding to the masked sequences are stored for informational purposes.
[0082] In certain embodiments, sequence masks are employed to reduce data noise. In some embodiments, both the sequence of interest and its normalizing sequences are masked. In some embodiments, different masks may be employed when different chromosomes or segments of interest are considered. For example, one mask (or group of masks) may be employed when chromosome 13 is the chromosome of interest and a different mask (or group of masks) may be employed with chromosome 21 is the chromosome of interest. In certain embodiments, the masks are defined at the resolution of bins. Therefore, in one example, the mask resolution is 100 kb. In some embodiments, a distinct mask may be applied to chromosome Y. The masked exclusion regions for chromosome Y may be provided at a finer resolution (1 kb) than for other chromosomes of interest, as described in U.S. Patent Publication No. 2014 / 0371078. The masks are provided in the form of files identifying excluded genomic regions.
[0083] As used herein, the expression “sequence anchoring” refers to the use of fast local homology search tool by alignment algorithms to identify “anchor points,” i.e. high-scoring local alignments of the input sequences, for genome alignment. To align longer sequences, most programs for genomic alignment rely on some sort of anchoring. In a first operation, they may use a fast local alignment method to identify high-scoring local homologies (as opposed to global), so called anchor points. Next, chains of such local alignments may be calculated and, finally, sequence segments between the selected anchor points may be aligned with a slower but more sensitive alignment method. For multiple sequence sets, either pairwise or multiple local alignments can be used as anchor points. A pioneering tool to find anchor points for genomic alignment is MU Mmer (Delcher et al., 1999); the current version of the program is considered the state-of-the-art in alignment anchoring (Kurtz et al.,2004). MLIMmer uses maximal unique matches as pairwise anchor points. The genome aligner MGA, by contrast, uses maximal exact matches involving all input sequences (Ho"hl et al., 2002). Finding anchor points is typically a significant operation in whole genome sequence alignment. Here, a trade-off between speed, sensitivity and precision may be made. A sufficient number of anchor points is typically helpful to reduce the run time of the subsequent, more sensitive alignment routine. Wrongly chosen anchor points, on the other hand, can substantially deteriorate the quality of the final output alignment. They may not only lead to misalignments of nonhomologous parts of the sequences but may also prevent biologically relevant, true homologies from being aligned. Also, if the number of anchor points is too large, finding optimal chains of anchor points can become computationally expensive
[0053] .
[0084] As used herein the term “sample” refers to any biological sample that comprises nucleic acid molecules, typically comprising DNA and / or RNA. Samples may be tissues, cells or extracts thereof, or may be purified samples of nucleic acid molecules.
[0085] As used herein, “nucleic acids” include polymeric molecules such as deoxyribonucleic acid (DNA), ribonucleic acid (RNA), peptide nucleic acid (PNA), or any sequence of what are commonly referred to as bases joined by a chemical backbone where the bases have the ability to form base pairs or hybridize with a complementary chemical structure. Suitable non-nucleotidic backbones include, for example, polyamide and polymorpholino backbones. The term “nucleic acids” includes oligonucleotide, nucleotide, or polynucleotide sequences, and fragments or portions thereof. The nucleic acid can be provided in any suitable form, e.g., isolated from natural sources, recombinantly produced, or artificially synthesized, can be single- or double-stranded.
[0086] Telomeres and telomere length
[0087] As described herein, a telomere is a region of repetitive nucleotide sequences associated with specialized proteins at the ends of linear chromosomes. Telomeres are a genetic feature most commonly found in eukaryotes, and are known to protect the terminal regions of chromosomal DNA from progressive degradation and ensure the integrity of linear chromosomes by preventing DNA repair systems from mistaking the very ends of the DNA strand for a double strand break. Telomeres are also critical for maintaining genomic integrity and may be factors for age-related diseases (e.g., Bone marrow failure, pulmonary fibrosis, immunodeficiency and other related Telomere Syndromes). Laboratory studies and clinical observations have shown that telomere dysfunction or shortening is commonly acquired due process of cellular aging and tumor development, wherein observational studies have found shortened telomeres in many types of experimental cancers (e.g.,squamous cell carcinomas of the skin and upper aerodigestive tract, myelodysplasia, and acute myeloid leukemia). Despite having short telomeres, cancer cells typically have activated telomerase that will maintain telomeres. In some cases, long telomeres inherited in families predispose to cancer.
[0088] Most vertebrate telomeric DNA consists of long (TTAGGG)n repeats of variable length, often around 3-20kb. In one embodiment of the present disclosure, a “subtelomere” refers to a segment of DNA between telomeric caps and coding regions of DNA. In vertebrates, each chromosome has two subtelomeres immediately adjacent to the long (TTAGGG)n repeats at each chromosome end. In some embodiments, subtelomeres are considered to be the most distal (farthest from the centromere) region of unique DNA on a chromosome, and they are unusually dynamic and variable mosaics of highly repeated blocks of sequence.
[0089] Telomere length varies greatly between species, from approximately 300 base pairs in yeast to many kilobases in humans, and usually is composed of arrays of guaninerich, six- to eight- base-pair-long repeats. Eukaryotic telomeres normally terminate with 3’ single-stranded-DNA overhang, which is essential for telomere maintenance and capping.
[0090] Provided herein are methods for determining allele-specific telomere length (ATL) and allele-specific telomere variant repeat (TVR) sequences of a DNA molecule from a biological sample, the method comprising: (a) obtaining a long read sequencing data set; (b) extracting from the long read sequencing data set a subset of sequencing reads containing a minimum number of canonical telomere repeats and identifying telomere region boundaries; (c) identifying individual alleles based on the presence of TVRs, comprising: (i) querying extracted reads for matches to a known set of telomere repeats, (ii) identifying TVR regions in the extracted reads, and (iii) clustering the extracted reads based on their TVR sequences and subtelomere sequences; (d) identifying anchored telomeres by aligning the subtelomere sequences of each extracted read with at least one reference genome containing subtelomere sequences; (e) producing a report containing the lengths and sequences of all anchored telomere alleles, thereby determining allele-specific telomere length (ATL) and allele-specific telomere variant repeat (TVR) sequences of a DNA molecule from a biological sample.
[0091] Samples and Sample Processing / Telomere source
[0092] The described methods can detect and quantify individual telomere length in a subject from a telomere possessing species, or in a DNA-containing sample taken from a telomere possessing species. Such species include any organism that has linearchromosomes and therefore require telomeres to buffer against the shortening of chromosomal ends that naturally result from the process of lagging strand DNA replication.
[0093] A telomere possessing species can be a human or non-human subject. Nonhuman subjects encompass any suitable organism, including non-human mammals, vertebrate and invertebrate animals, plants, microorganisms, and the like. The described methods are directed to subjects from which quantification of short and / or critically short telomeres can provide clinical and research diagnostic information, such as for human and non-human veterinary subjects. However, the described methods are equally applicable for use in quantifying short and / or critically short telomeres in invertebrate and non -animal organisms.
[0094] In particular embodiments, the described methods are used to quantify individual telomeres in any suitable sample from a telomere possessing species. Such samples include, but are not limited to, samples isolated from a subject and from which DNA is isolated for use in the described methods. In other embodiments, such samples include tissue(s) or cells of a telomere possessing species, which are then cultured in vitro, and from which DNA is isolated. Although, telomere lengths may be determined for all eukaryotes, in a preferred embodiment, telomere lengths are determined for vertebrates, including without limitation, amphibians, birds, and mammals, for example rodents, ungulates, and primates, particularly humans. Preferred are organisms in which longevity is a desirable trait or where longevity and susceptibility to disease are correlated.
[0095] Any cellular or tissue sample with telomere-containing chromosomes could potentially be processed for use in the described methods. The telomere containing samples may be obtained from any tissue of any organism, including tissues of blood, brain, bone marrow, lymph, liver spleen, breast, and other tissues, including those obtained from biopsy samples. The samples may also comprise bodily fluids, such as saliva, urine, feces, cerebrospinal fluid, semen, etc. Typically, the tissue or cells are non-stem cells, i.e., somatic cells since the telomeres of stem cells generally do not decrease over time due to continued expression of telomerase activity. However, in some embodiments, telomeres may be measured for stems cells in order to assess inherited telomere characteristics of an organism.
[0096] In some embodiments, the sample is newly obtained from the subject. In other embodiments, the sample is a stored sample, preferably below -70°C. In some embodiments in which multiple samples are used to compare individual telomere lengths over time, the samples can be a combination of sample(s) stored by standard methods in the art and newly isolated samples. In other embodiments, the samples have all been stored.
[0097] DNA for use in the described methods can be isolated from the telomere containing samples by any standard method for preparing chromosomal DNA from cellular material. For instance, the sample may be treated using detergents, sonication, electroporation, denaturants, etc., to disrupt the cells. The target nucleic acids may be purified as needed. The samples may also be extracted using kits or automated procedures or instruments. The automated instruments may include, but not limited to, QIAcube, QIAsymphony, EZ1 advanced, Biorobot Universal, QIAextractor, or QI Agility.
[0098] Purified nucleic acid molecules can be sequenced with long-read sequencing technologies at, for example, 2x, 3x, 4x, 7x, 10x, 20x, 38x, or 60x coverage. The accuracy of a read can improve with increasing coverage. Sequence reads can be output in FASTA, FASTQ, SAM, or other file types used for storing nucleic acid sequence and mapping data.
[0099] Sequencing methods
[0100] As indicated above, the prepared samples (e.g., sequencing libraries) may be sequenced as part of the procedure for identifying TVRs and measuring ATLs. Any number of sequencing technologies can be utilized.
[0101] Some sequencing technologies are available commercially, such as the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, Calif.) and the sequencing-by-synthesis platforms from 454 Life Sciences (Bradford, Conn.), Illumina / Solexa (Hayward, Calif.) and Helicos Biosciences (Cambridge, Mass.), and the sequencing-by-ligation platform from Applied Biosystems (Foster City, Calif.), as described below. In addition to the single molecule sequencing performed using sequencing-by- synthesis of Helicos Biosciences, other single molecule sequencing technologies include, but are not limited to, the SMRT™ technology of Pacific Biosciences, the ION TORRENT™ technology, and nanopore sequencing developed for example, by Oxford Nanopore Technologies. Illustrative sequencing technologies are described in greater detail below.
[0102] In another illustrative, but non-limiting, embodiment, the methods described herein comprise obtaining sequence information for the nucleic acids in the test sample, e.g., cfDNA in a maternal test sample, cfDNA or cellular DNA in a subject being screened for a cancer, and the like, using the single molecule, real-time (SMRT™) sequencing technology of Pacific Biosciences. In SMRT sequencing, the continuous incorporation of dye-labeled nucleotides is imaged during DNA synthesis. Single DNA polymerase molecules are attached to the bottom surface of individual zero-mode wavelength detectors (ZMW detectors) that obtain sequence information while phospholinked nucleotides are being incorporated into the growing primer strand. A ZMW detector comprises a confinement structure that enables observation of incorporation of a single nucleotide by DNApolymerase against a background of fluorescent nucleotides that rapidly diffuse in an out of the ZMW (e.g., in microseconds). It typically takes several milliseconds to incorporate a nucleotide into a growing strand. During this time, the fluorescent label is excited and produces a fluorescent signal, and the fluorescent tag is cleaved off. Measurement of the corresponding fluorescence of the dye indicates which base was incorporated. The process is repeated to provide a sequence.
[0103] In another illustrative, but non-limiting embodiment, the methods described herein comprise obtaining sequence information for the nucleic acids in the test sample from a subject being screened for a cancer, and the like, using nanopore sequencing (e.g. as described in Soni G V and Meller A. Clin Chem 53: 1996-2001
[2007] ). Nanopore sequencing DNA analysis techniques are developed by a number of companies, including, for example, Oxford Nanopore Technologies (Oxford, United Kingdom), Sequenom, NABsys, and the like. Nanopore sequencing is a single-molecule sequencing technology whereby a single molecule of DNA is sequenced directly as it passes through a nanopore. A nanopore is a small hole, typically of the order of 1 nanometer in diameter. Immersion of a nanopore in a conducting fluid and application of a potential (voltage) across it results in a slight electrical current due to conduction of ions through the nanopore. The amount of current that flows is sensitive to the size and shape of the nanopore. As a DNA molecule passes through a nanopore, each nucleotide on the DNA molecule obstructs the nanopore to a different degree, changing the magnitude of the current through the nanopore in different degrees. Thus, this change in the current as the DNA molecule passes through the nanopore provides a read of the DNA sequence.
[0104] After sequencing of DNA fragments, sequence reads of predetermined length, e.g., 100 bp, are mapped or aligned to a known reference genome. The mapped or aligned reads and their corresponding locations on the reference sequence are also referred to as tags. In one embodiment, the reference genome sequence is the NCBI36 / hg18 sequence, which is publicly available. Alternatively, the reference genome sequence is the GRCh37 / hg19, which is publicly available. Other sources of public sequence information include GenBank, dbEST, dbSTS, EMBL (the European Molecular Biology Laboratory), and the DDBJ (the DNA Databank of Japan).
[0105] Methods provided herein can be used with genetic data, such as deoxyribonucleic acid (DNA) data. Such genetic data can be provided by a genetic sequencing device, such as a Pacific Biosciences, or Oxford Nanopore sequencing device. Such devices may provide a plurality of raw genetic data corresponding to the genetic information of a subject, as generated by the device from a sample provided by the subject.
[0106] In some embodiments, the disclosure provides a workflow, pipeline, computational tools or algorithms for whole genome sequencing data analysis. In some embodiments, the pipeline is optimized for executing methods such as compression, processing, and analysis of raw data generated from various sequencers, such as, e.g., HiSeq, MiSeq, Genome Analyzer, SOLID, 454, Ion Torrent, and Pacific Biosciences. One of skill in the art will understand that the disclosed methods can be used not only for genomic data, but can be utilized for the compression, processing, and analysis of any dataset.
[0107] An example workflow is depicted in Figure 2A, illustrating the execution of several methods by the pipeline described herein. As shown in Figure 2A, a subset of reads containing telomeric reads are extracted from long reads obtained from a sample. This is accomplished by mapping all raw reads in the appropriate data format to the CHM13 draft genome assembly (v1 .0) obtained from the telomere-to-telomere consortium. In some embodiments, the raw reads are in BAM format. In other embodiments, the raw reads are in Fastq format.
[0108] Advantageously, extracting a subset of reads may reduce the complexity of analyzing the sequencing data. In one embodiment, a subset of reads containing a minimum number of canonical telomere repeats are then extracted, with a default minimum of 10 copies, to construct a reduced set of reads.
[0109] Reads can be aligned with any known gene alignment tools in the art such as pbmm2, NGMLR, minimap2, etc. In some embodiments, the PacBio Complete Long Reads (CLR) are aligned to the Telomere-to-Telomere CHM13 (T2T-CHM13) build of the human reference genome, using the pbmm2 alignment software, with the -preset SUBREAD parameter. The T2T-CHM13 genome includes gapless telomere-to-telomere assemblies for all 22 human autosomes and chromosome X. In an alternative embodiment, Nanopore long reads are aligned to the T2T-CHM13 reference genome using minimap2 with the parameters -ax map-ont -N 5 -Y -L.
[0110] In one embodiment, a reduced set of reads is realigned to a copy of the T2T reference genome or other reference genome provided herein, which contains subtelomeres and the alternate subtelomere assemblies. In some embodiments, the telomere sequences in the reference are masked and appended as additional reference contigs. Reads which did not contain any expected combination of sequences can be discarded at this operation.
[0111] The disclosure is further described in the following examples, which do not limit the scope of the disclosure described in the claimsEXAMPLES
[0112] Mammalian telomeres are regions of repetitive DNA at the ends of linear chromosomes, which protect chromosome ends from degradation. Progressive telomere shortening occurs in all dividing normal cells owing to incomplete lagging strand DNA synthesis (i.e., the “end replication problem”), oxidative damage, exonucleolytic processing events and other factors, and this shortening eventually results in replicative senescence or apoptosis. Telomere shortening or telomere uncapping results in telomere dysfunction, a key contributor to organismal ageing and multiple age-related diseases such as infertility, arthritis, diabetes, cancer, cardiovascular and neurodegenerative diseases. Additionally, germline mutations affecting genes involved in telomere maintenance lead to a heterogeneous group of diseases termed telomere biology disorders (TBD). Telomere lengths (TLs) have thus been extensively studied in the context of aging and disease.
[0113] The most common techniques to measure TL mostly provide information only about average TL (i.e., the mean of 92 telomeres, assuming a normal human chromosomal complement) or relative telomere length for a pool of telomeres. However, it is well established that it is the shortest telomeres, rather than the average TL, that constitute telomere dysfunction and limit cellular survival in the absence of telomerase. Therefore, characterizing the entire distribution of TL within a sample may provide particularly valuable information about the relative increases in the number of short telomeres. Additionally, human telomeres vary in both chromosomal-specific and allelic telomere lengths, yet most current methods lack resolution at the level of individual telomeres or they are laborious, technically demanding, and have low throughput. Thus, there is need for a highly sensitive, high-throughput method that can quantify allelic telomere lengths and identify their physically linked DNA sequence.
[0114] The present disclosure addresses these limitations by providing a method for measuring the lengths of all 92 telomere alleles from long read sequencing data, whose workflow is presented in Figure 2A. Long read sequencing allows the capture of the full telomere sequence as well as the sequences of subtelomeres and the unique variant repeats proximal to telomere regions termed telomere variant repeats (TVRs) in one long read. The structure of telomeric regions and the positions of these three regions is depicted in Figure 1 . In one embodiment of the present disclosure, individual telomeres are identified using TVRs which are unique across alleles and allele specific telomere length is calculated. Accordingly, the method further provides the TVR sequences for each allele, which were previously uncharacterized regions of the human genome due to the lack of high-throughput methods capable of analyzing telomere sequences of individual alleles. The method can also positively detect telomere loss, characterize novel subtelomeric variants, haplotypes, and previously uncharacterized recombined subtelomeres. This high-throughput and high-resolution characterization of telomeres could be foundational to future studies investigating the roles of specific telomeres in aging and disease. This includes but is not limited to population genomics studies, levels of telomere and TVRs erosion associated to genomics instability and study of chromosome fusion events that are precursor to chromothripsis and tumorigenesis.
[0115] Materials and Methods
[0116] Telomere repeat set:
[0117] As provided herein Telogator2 uses a curated set of telomere repeats for identifying TVR regions (Table 1).
[0118] Table 1 : Telomere repeats and their corresponding character mappings used by Telogator2.1 Also includes TAGGG and TAGGGG.2 Also includes TGGG and TTGGG.3 The repeat symbols are drawn from the amino acid alphabet so that external methods for sequence alignment can be used in subsequent operations.
[0119] These sequences were derived from the analysis of PacBio HiFi long reads from multiple publicly available human samples, starting with the Han Chinese trio (acquired from the Sequence Read Archive under project accession number PRJNA200694), which were chosen for their high coverage. Reads were first queried for the canonical TTAGGG repeat and common variants such as TCAGGG (C-type), TGAGGG (G-type) and TTGGGG (J- type). Collectively, these four repeats comprised 90% of all read sequences beyond the subtelomere boundary. The remaining regions that did not match any of these repeats weremanually examined, and additional variant repeats were added to the set as they were found.
[0120] In the process of identifying variant repeats from this data, an increased frequency of sequencing errors in telomere regions was observed, compared to reads aligning elsewhere in the genome. Specifically, an increased frequency of insertion / deletion sequencing errors was observed in homopolymer sequences (Figure 11), a commonly known sequencing artifact in PacBio long reads
[0047] , The J-type TTGGGG repeats are particularly prone to this error. To increase variant repeat matches in the presence of these errors additional patterns were added to the set representing T-type and J-type repeats modified by insertions and deletions.
[0121] Identifying telomere regions:
[0122] Telomere regions of each read were queried for exact matches to sequences from the telomere repeat set. Overlapping matches were prevented by searching for the longest repeats before searching for shorter repeats, while disallowing matches to a repeat if a larger repeat (that contains the smaller one) had already been found at that position. Partially overlapping matches were trimmed such that each position in the read was matched to a single telomere repeat from the variant repeat set. From these matches, each read was converted from a nucleotide string of [A,C,G,T] characters to a new string with symbols indicating which telomere variant repeat was found at that position (Figure 3B). The symbols corresponding to each repeat are shown in Table 1 . This approach is similar to the ’telomere variant repeat codes’ used in visualizations by Baird et al.
[0013] and provides a representation of telomere sequences that simplifies downstream analysis via reducing repetitive DNA elements to contiguous blocks of identical symbols.
[0123] TVR region clustering:
[0124] TVR regions of all reads were then compared pairwise. Similarity scores were computed via pairwise alignment using Biopython
[0048] with a match score of +5 and mismatch / gap penalties of -4. These pairwise similarity scores were then converted to pairwise distances in the same manner as is done in progressive multiple sequence alignments
[0049] . Hierarchical clustering was then applied to the pairwise distances, and reads were assigned a cluster by cutting the resultant dendrogram (example shown in Figure. 3C (left panel)). Like Grigorev et al.
[0024] , the Ward variance minimization algorithm was found to best separate the telomere alleles.
[0125] Cluster refinement:
[0126] After forming initial clusters of reads, each cluster was processed individually in two operations: First, the TVR regions of reads were compared pairwise in the same manner as described above, but only to other reads in the same cluster, and using more aggressive clustering thresholds. This was done to separate distinct alleles that may have been clustered together in the first operation, due to their TVR regions sharing a common prefix or otherwise being similar enough to be grouped together. Second, the subtelomere portions of reads were compared to each other, again only to other reads in the same cluster. This was done to separate alleles where the TVR regions were highly similar, but the dissimilarity of the subtelomeres made it clear that the supporting reads originate from a different allele. Overall, it was rare for clusters of reads to have highly similar telomeres and divergent subtelomeres, and this operation was found to only affect one or two clusters per sample, on average.
[0127] Consensus TVR sequences and ATL distributions:
[0128] For each refined cluster, a consensus TVR sequence was obtained via multiple sequence alignment of its constituent reads using Muscle
[0050] . Muscle is run with a custom scoring matrix (Figure 12) specifying the same match / mismatch penalties as were used for the pairwise alignments, with the exception that the matching score of canonical repeats is reduced to 0. This scoring matrix was designed to produce multiple alignments that prioritize matching variant repeats and reduce the reward for aligning canonical repeats in between the blocks of variant repeats. This produced consensus sequences that more closely resemble the variant repeat patterns in the individual reads.
[0129] The TVR boundary for each allele was then determined from its consensus sequence. The boundary is defined to be the position of the most distal variant repeat. In practice, the TVR boundary was chosen to be the position at which 95% of cumulative variant repeats have been observed (example shown in Figure. 3D), due to sporadic variant repeats further into the telomere that are likely attributable to sequencing errors.
[0130] ATL distributions were reported for each cluster based on how far each supporting read extended beyond the TVR boundary. A representative ATL for each cluster can be chosen via a user-specified method, either the mean, median, max, or a percentile of an ATL distribution (default: mean).
[0131] Subtelomere anchoring:
[0132] To assign each cluster a chromosome arm, the subtelomere portion of all supporting reads for all clusters were extracted and aligned to a collection of curated subtelomere reference sequences. By default, minimap2
[0051] was used for read alignment, but winnowmap
[0052] or pbmm2 (https: / / github.com / PacificBiosciences / pbmm2) can be usedinstead via input options. The curated subtelomere reference contains the most distal 500kb of each chromosome arm of contigs from three recently released T2T references [29-31] and is extensible to include the subtelomeres of additional references as they become available.
[0133] In most cases, all subtelomere sections of reads from the same cluster will align to the same subtelomere reference sequence, making chromosome arm assignment straightforward. However, in approximately 5-10% of clusters, subtelomere reads will be mapped to multiple subtelomeres. In these cases, chromosome arm assignment will be ambiguous, and all candidate arms will be included in the final output. By default, Telogator2 requires reads have at least 1 kb of subtelomere sequence in order to be aligned. Long reads from whole genome sequencing will nearly always have sufficient subtelomere sequence, but this may not be the case for reads sequenced using telomere-capture strategies.
[0134] After chromosome arms were assigned to each cluster, the TVR consensuses and ATL distributions were written to an output report, along with plots showing TVR and telomere regions colorized (grouped) based on their composition of variant repeats.
[0135] Sequencing data:
[0136] Whole genome PacBio HiFi long reads for the CHM13 human haploid cell line were downloaded from the Sequence Read Archive (SRA) under project accession PRJNA530776. Whole genome HiFi reads for the Han Chinese trio (HG005, HG006, HG007) were downloaded from SRA under project accession PRJNA200694. Whole genome HiFi reads for 46 human samples (sequenced as part of the Human Pangenome Project) were downloaded from SRA under project accession PRJNA701308.Corresponding short read data for these 46 samples was downloaded from SRA at http colon forward slash forward slash ftp dot sra dot ebi dot ac dot uk forward slash vol1 forward slash run forward slash ERR398 forward slash in the form of BAM files (reads aligned to GRChg38). Telomere reads from HG002 were included in the Telogator2 repository. A full list of samples and their coverage is provided in Table 2.
[0137] Table 2: HG002 PB vs. ONT comparison
[0138] Telogator2
[0139] In one embodiment, Telogator2 begins by extracting a subset of reads containing a minimum number of canonical telomere repeats, with a default minimum of 10 copies (canonical repeats). In one embodiment, telomere region boundaries are estimated based on the density of telomere repeats in a 100bp sliding window (described previously
[0027] ) and reads that terminate in telomere sequence on one end and non-telomere sequence on the other are selected for further analysis. In another embodiment, these reads are further processed to identify individual alleles based on their composition of telomere variant repeats. In some embodiments, this is done by, i) querying reads for matches to a curated set of telomere repeats, ii) identifying TVR regions, and iii) clustering reads based on their TVR and subtelomere sequences (Figure. 2B). After clustering, in some embodiments, the sub-telomere portions of each read are aligned to a customized reference genome containing subtelomere sequences from three different telomere-to-telomere reference assemblies: T2T-CHM13
[0029] , as well as two diploid assemblies derived from Han Chinese individuals [30, 31]. In one embodiment, reads that can be aligned to a specific subtelomere are referred to as “anchored” telomeres. In yet another embodiment, an output report is produced containing the lengths and sequence compositions of all anchored telomere alleles. In some embodiments, plots showing each allele are also produced, with TVR and telomere regions colorized based on their composition of variant repeats.
[0140] Validation of telomere alleles in CHM13
[0141] As an initial validation, HiFi reads of the CHM13 haploid human cell line were analyzed and the telomere alleles identified by Telogator2 were compared to that of the T2T reference genome.
[0142] From this data, Telogator2 identified all 46 telomeres, (Figure. 3A). 45 / 46 telomeres were anchored at the expected reference coordinates at their respective arm’s subtelomere-telomere boundary. The one exception was the telomere of 18p, where telomere reads were instead mapped about 100kb inside the subtelomere. Interestingly, the telomere of 18p was also absent from the initial de novo assemblies used to construct the T2T reference, and needed to be resolved manually
[0029] .
[0143] Examining TVR regions in a larger sample set
[0144] To examine TVR regions in a larger set of samples, HiFi long reads from 46 samples, which were sequenced as part of the Human Pangenome Reference Project
[0032] , which selected healthy individuals from a variety of demographics, were downloaded. Each sample had 30-40x coverage in subtelomere regions, with the exception of HG02572 which was discarded as an outlier due to its lower coverage, leaving 45 samples in total.
[0145] Processing these samples with Telogator2, identified an average of 91 (std=2) unique alleles per sample, which is very close to the expected total number of alleles. This suggests that, assuming sufficient coverage and read lengths, and outside of situations where parents are closely related, each of an individual’s 92 telomeres can be distinguished based on their unique TVR and subtelomere sequences.
[0146] On average 3 alleles per sample (std=2) were reported by Telogator2 as ’blank,’ indicating that not enough variant repeats were present to identify a TVR region between subtelomere and telomere. Via visual inspection, it was found that many blank TVR regions do contain a small number of variant repeats, but not enough to be reliably detected by Telogator2 in the presence of noise. The major exception is 22p, where 18 / 45 samples were found to have a telomere allele with no observable variant repeats.
[0147] Thus, the vast majority of telomeres from the samples analyzed have unique TVR regions that can be clearly distinguished. The non-unique TVRs fall into three categories; False negatives: TVR regions that have unique patterns, but are comprised of very few variant repeats such that they labeled by Telogator2 as 'blank'. These alleles are instead clustered using their subtelomere sequences; No variant repeats: Rare alleles where the canonical telomere repeats are directly adjacent to subtelomeres with no TVR region in between. Most frequently found in 22p; Homozygous TVR regions: TVR regions are inherited, thus it is possible that two copies of a particular allele could be inherited that are highly similar or identical if they have not acquired random mutations. E.g., if two parents are closely related, the telomeres of their offspring at some chromosome arms may be homozygous and will be reported by Telogator2 as a single allele.
[0148] The default parameters in Telogator2 were tuned such that, in the presence of noise, reads with the same TVR regions are still clustered together in a vast majority of cases. Numerous output plots are produced at each clustering operation, so that individual thresholds can be adjusted if necessary.
[0149] TVR region length and repeat composition.
[0150] It was observed that the lengths of TVR regions are generally similar across chromosome arms when averaged across many samples (Figure 4). The exceptions to this were the TVR regions in the short arms of acrocentric chromosomes (13p, 14p, 15p, 21 p and 22p) which on average are shorter than TVR regions from the other arms.Approximately 70% of TVR regions comprised of canonical repeats. The remaining 30% are variant repeats, with the most frequent occurrences being G-type and J-type (Table 3).
[0151] Table 3: TVR repeat composition as determined by Telogator2
[0152] While previous work has described TVR regions as being between 0-2kb [7, 33, 34], in the present study it was observed that they can be as long as ~ 8kb in some alleles (Figure. 4). Theoretically, these larger TVR regions could potentially affect the accuracy of experimental methods for measuring the lengths of individual telomeres. For example, when using TRF
[0035] or STELA
[0036] , it is common practice to subtract a constant corresponding to the distance between the telomere-subtelomere boundary and a reference point further into the subtelomere (an enzyme restriction site in the case of TRF, and subtelomere-specific primer sets for STELA). While this constant is typically in the range of 2-4kb
[0014] , for alleles with very large TVR regions this may be insufficient and could result in overestimated ATLs.
[0153] TVR region inheritance
[0154] To validate the sizes and sequence compositions of TVR regions reported by Telogator2, long reads from the Han Chinese trio (HG005, HG006 and HG007)
[0037] were processed. TVR regions are known to be inherited, thus the variant repeat patterns of a child should also be found in one of their parents with minimal variation. In total, 91 unique alleles were identified in child (HG005), 45 of which could be traced to the father (HG006), 45 traced to the mother (HG007), and a single allele containing very few variant repeats anchored to 22p where inheritance could not be determined. The 92nd allele was possibly missed either due to insufficient coverage from that particular allele, or it may have been highly similar to another allele such that the reads were clustered together.
[0155] 87 of the 90 inherited TVR regions appeared to be virtually identical between parent and child (Figure 5A). In a previous study of telomeres in trios, it was observed that the combined TVR + telomere regions in child alleles were closer in sequence similarity to their father than their mother
[0024] , However, when restricted solely to TVR regions no significant differences were found in the distances between HG005 / HG006 (average sequence similarity: 98.3%, std=1 .8%) and HG005 / HG007 (average sequence similarity: 98.0%, std=2.0%).
[0156] A majority of the inherited TVR regions that were observed in the Han trio were extremely similar in parent and child (average sequence similarity 98.2%, std=1 .9%). However, three alleles had notable differences which could possibly be attributable to de novo variation (Figure 8): HG005 1q: Deletion of the most distal block of CCCTAG repeats with respect to maternal allele; HG005 21 q: Deletion of the most distal block of CCCCAA repeats with respect to paternal allele; and HG005 22q: Deletion within most distal block of canonical repeats within the TVR region, with respect to maternal allele.
[0157] ATL estimation from the uniqueness of TVR regions
[0158] For each set of reads from the same allele, a TL distribution can be estimated by simply reporting the length of the canonical repeats beyond the TVR boundary in each read. From these distributions a ’representative’ length for the allele can be chosen, e.g. by taking the mean or a percentile.
[0159] While the usage of whole genome sequencing for estimating average TL has been widely adopted and used in a large number of studies, the TL distributions for individual alleles are potentially more sensitive and could be affected by multiple factors, including how the sample was prepared for sequencing. For example, fragmenting DNA to achieve a targeted size may introduce reads with incomplete telomere regions that have fewer canonical repeats than the allele from which it originated. To address this, several telomere capture protocols have recently been developed for preparing DNA for long read sequencing while ensuring that each molecule contains entire telomeres. These include an approach using ’telobaits’
[0025] and PacBio sequencing, as well as the ‘Telo-seq’ protocol for ONT sequencing
[0054] , However, the publicly available data from these preparations show limitations: Most notably, the average read lengths are comparatively short (5 - 8 kb), and a majority of reads do not contain enough subtelomere sequence to be uniquely assigned a chromosome arm.
[0160] So while the ATLs presented in this work are all derived from whole genome sequencing, and thus should be interpreted as estimates and not precise measurements of absolute lengths, Telogator2 is readily applicable to data from telomere capture preparations, and as those protocols improve, we anticipate that Telogator2 will be a valuable method for analyzing the data.
[0161] For example, using the CHM13 telomere alleles identified earlier, the ATL estimates reported by Telogator2 were compared to the telomere sizes derived from T2T reference genome annotations. Since the same whole genome data that was used to construct the reference was being utilized, it was expected that the ATL distributions found at each arm would support the amount of telomere sequence that was included in the finalassembly. The ATLs reported by Telogator2 (Figure 6) correlate well with the reference annotations (R = 0.87), with an average difference of 288bp (std = 224bp) across all alleles (Figure 7).
[0162] Variance in ATL across sequencing technologies
[0163] To assess the variability of ATLs across different sequencing technologies, Telogator2 was run on long reads from the same sample sequenced on both the PacBio Revio platform and the ONT Prometh ION 2 Solo platform. The HG002 sample was chosen for this experiment, the DNA of which was acquired from the Coriell Institute for Medical Research.
[0164] The PacBio reads were combined from multiple runs for a total coverage of 75x, the ONT reads were from a single run with a coverage of 20x. In total 90 unique alleles were identified from the PacBio reads while 84 unique alleles were identified from the ONT reads. The difference was attributed to coverage depth. All 84 alleles found from the ONT data could also be found in the results from PacBio data. For these 84 alleles pairwise comparisons of TL distributions at each allele using multiple strategies were used to select a representative TL value from each distribution (Table 4).
[0165] Table 4: Correlation of ATLs computed from PacBio and ONT datasets for HG002, using different methods for choosing representative ATL from ATL distributions.
[0166] Multiple strategies were tested to identify the most consistent method for characterizing TL distributions from long reads across multiple runs of the same sample (which are expected to have the same ATLs) and is thus a suitable default for Telogator2. It was found that, among the strategies for selecting a representative ATL, the 75th percentile results in the highest correlation between the two datasets. The average difference in ATL when comparing 75th percentiles is 588 bp (std = 618 bp).
[0167] Finally, the ATLs reported by Telogator2 on the 45 pangenome samples were compared to the average TL estimates produced by TelomereHunter
[0038] when applied to short reads from the same samples. For each sample, ATLs derived from the long readswere averaged together and compared against TelomereHunter’s “tel-content” output, yielding a correlation of R = 0.89 (Figure 12).
[0168] Distribution of allele-specific telomere length in patients with Telomere Biology Disorders (TBD)
[0169] Only a few experimental procedures report TL at the chromosome arm level, most of which are low throughput or labor-intensive. High throughput approaches based on optical mapping using fluorescent probes to label repetitive sequences have recently been proposed [23, 40], though the measurements exhibit high variance. Recently proposed methods such as the Telomere length Combing Assay (TCA)
[0041] or single telomere absolute-length rapid assay (STAR)
[0042] show potential for more accurate measurement of individual telomeres, but these methods have not yet seen broader adoption to confirm the consistency of their results
[0043] .
[0170] Allele-level characterization of telomeres enables the establishment of baselines for characterizing individuals with abnormal ATLs. For example, in the context of telomere biology disorders it is largely unknown whether the condition is driven by low average TL or the shortening of a specific telomere or group of telomeres, though it has generally been found that the shortest individual telomeres are the most relevant [17, 33]. It has been observed that as cultured cells approach senescence, the distribution of TL skews towards shorter values
[0044] , and it has been suggested that cellular senescence can be triggered by the shortening of a subset of telomeres beyond a critical length
[0045] . A more comprehensive characterization of ATL (beyond the limited number of arms that are usually studied in detail) using long reads and Telogator2 could corroborate or clarify these findings.
[0171] Comparing ATL between unrelated individuals
[0172] As described herein, the present disclosure provides methods for ATL measurement which can be used to compare the distribution of telomere length between individuals or one individual and a group of individuals used as reference. In some embodiments, applications include identification of patients with Telomere Biology Disorders (TBD) and an analysis of population heterogeneity. An example comparison of a sample’s chromosome-specific TL (red / dark shade) to an aged-matched average of healthy controls (gray / light shade) is shown in Figure 14.
[0173] The methods described herein may also enable visualization of sorted allele TLs to identify erosion subtypes (e.g., cancer, HCC, MM). Figure 15 shows allele telomeres, sorted by length, from two bone marrow samples. One with a normal TL distribution (left) and one with eroded telomeres (right).
[0174] Comparing ATL from the same individual at specific time points
[0175] When comparing sequencing data from the same individual, the TVRs from the same allele telomere are identical. Using the techniques described herein, the TVRs can be precisely aligned to measure changes in ATL between two time points. An example of comparing chromosome-specific TL on MRC5 cell line at different passages is shown in Figure 16.
[0176] Following the same principle as describe above, TVRs can be used to compute the precise difference in ATLs between two tissues from the same individual that have been biopsied at similar time.
[0177] REFERENCES[1] Harley, C.B., Futcher, A.B., Greider, C.W. (1990) Nature. 345(6274) :458-460.[2] Shay, J.W., Wright, W.E. (2005) Carcinogenesis. 26(5):867-874.[3] Riethman et al. (2004) Genome Res. 14(1 ):18-28.[4] Mewborn, S., Lese Martin, C., Ledbetter, D. (2005) Cytogenet Genome Res. 108(1 - 3):22-25.[5] Allshire, R.C., Dempster, M., Hastie, N.D. (1989) Nucleic Acids Res. 17(12):4611 - 4627.[6] Coleman, J., Baird, D.M., Royle, N.J. (1999) Hum Mol Genet. 8(9):1637-1646.[7] Lee et al. (2018) Nucleic Acids Res. 2018;46(10):4903-4918.[8] Bluhm et al. (2019) Nucleic Acids Res. 47(4) :1896-1907.[9] Nishikawa et al. (2001 ) Structure. 9(12):1237-1251.
[0010] Hanaoka, S., Nagadoi, A., Nishimura, Y. (2006) Comparison of DNA-Binding Activities Between hTRF2 and hTRFI with hTRF2 Mutants. In: Webb, G.A. (eds) Modern Magnetic Resonance. 743-751 Springer, Dordrecht.
[0011] Conomos et al. (2012) J Cell Biol. 199(6):893-906.
[0012] Baird, D.M., Jeffreys, A., Royle, N. (1995) EMBO J. 14(21 ):5433-5443.
[0013] Baird, D.M., Coleman, J., Rosser, Z.H., Royle, N.J. (2000) Am J Hum Genet. 66(1 ):235-250.
[0014] Aubert, G., Hills, M., Lansdorp, P.M. (2012) Mutat Res. 730(1 -2):59-67.
[0015] Vera, E., Blasco, M.A. (2012) Aging (Albany NY). 4(6):379-392.
[0016] Lansdorp, P.M. (2022) Blood. 139(6):813-821 .
[0017] Hemann, M.T., Strong, M.A., Hao, L.-Y., Greider, C.W. (2001 ) Cell. 107(1 ):67-77.
[0018] Xu, Z., Due, K.D., Holcman, D., Teixeira, M.T. (2013) Genetics. 194(4):847-857.
[0019] Zheng, Y.-L., Loffredo, C.A., Shields, P.G., Selim, S.M. (2009) Carcinogenesis. 30(8):1380-1386.
[0020] Xing et al. (2009) Cancer Prev Res (Phila). 2(5):459-465.
[0021] Barkovskaya et al. (2019) Bulletin of Siberian Medicine. 18(1 ): 164-174.
[0022] Montpetit et al. (2014) Nurs Res. 63(4):289-299.
[0023] Young et al. (2017) Nucleic Acids Res. 45(9):e73.
[0024] Grigorev et al. (2021 ) Genome Res. 31 (7) :1269-1279.
[0025] Tham et al. (2023) Nat Commun. 14(1 ) :281 .
[0026] Reed et al. (2021 ) iScience. 24(2):102082.
[0027] Stephens, Z., Ferrer, A., Boardman, L., Iyer, R. K., & Kocher, J. A. (2022). Bioinformatics. 38(7):1788-1793.
[0028] Guo, M., Songyang, Z., Xiong, Y. (2023) Small Methods. 7(1 1 ):e2300385.
[0029] Nurk et al. (2022) Science. 376(6588) :44-53.
[0030] He et al. (2023) Genomics, proteomics & bioinformatics, S1672-0229(23)00100-6.
[0031] Yang et al. (2023) Cell Res. 33(10):745-761 .
[0032] Wang et al. (2022) Nature. 604(7906) :437-446.
[0033] Capper et al. (2007) Genes Dev. 21 (19):2495-2508.
[0034] Dejardin, J., Kingston, R.E. (2009) Cell. 136(1 ):175-186.
[0035] Moyzis et al. (1988) Proc Natl Acad Sci U S A. 85(18):6622-6626.
[0036] Baird, D.M., Rowson, J., Wynford-Thomas, D., Kipling, D. (2003) Nat Genet. 33(2):203-207.
[0037] Wang et al. (2019) Sci Data. 6(1 ) :91.
[0038] Feuerbach et al. (2019) BMC Bioinformatics. 20(1 ):272.
[0039] Dubocanin et al. (2022) Single-molecule architecture and heterogeneity of human telomeric dna and chromatin. BioRxiv, pp. 2022-05.
[0040] Uppuluri, L., Varapula, D., Young, E., Riethman, H., Xiao, M. (2021 ) Environ Toxicol Pharmacol. 82:103562.
[0041] Kahl et al. (2020) Front Cell Dev Biol. 8:493.
[0042] Luo et al. (2020) Sci Adv. 6(34):eabb7944.
[0043] Dweck, A., Maitra, R. (2021 ) Mol Biol Rep. 48(7):5621 -5627.
[0044] Martens et al. (2000) Exp Cell Res. 256(1 ):291 -299.
[0045] Kaul et al. (201 1 ) EMBO Rep. 13(1 ):52-59.
[0046] Tan, K.-T., Slevin, M.K., Meyerson, M., Li, H. (2022) Genome Biol. 23(1 ):180.
[0047] Goodwin, S., McPherson, J.D., McCombie, W.R. (2016) Nat Rev Genet. 17(6):333- 351.
[0048] Cock et al. (2009) Bioinformatics. 25(1 1 ):1422-1423.
[0049] Feng, D. F., & Doolittle, R. F. (1987) J Mol Evol. 1987;25(4):351 -360.
[0050] Edgar, R.C. (2004) Nucleic Acids Res. 32(5):1792-1797.
[0051] Li, H. (2018) Bioinformatics. 34(18):3094-3100.
[0052] Jain et al. (2020) Bioinformatics. 36(Suppl_1 ):i1 11 -i118.
[0053] Leimeister, C. A., Dencker, T., & Morgenstern, B. (2019). Bioinformatics. 35(2):211 - 218
[0054] Schmidt et al. (2023) High resolution long-read telomere sequencing reveals dynamic mechanisms in aging and cancer. bioRxiv. 2023.11.28.569082.
[0055] Graakjaer, J., Londono-Vallejo, J. A., Christensen, K., & Kolvraa, S. (2006). Ann N Y Acad Sci. 1067:311 -316.
[0056] Norris et al. (2021 ). Hum Genet. 140(6):945-955.
[0057] Ferrer, A., Stephens, Z. D., & Kocher, J. A. (2023). Curr Hematol Malig Rep. 18(6):284-291.
[0058] Stephens, Z., Kocher, JP. Characterization of telomere variant repeats using long reads enables allele-specific telomere length estimation. BMC Bioinformatics 25, 194 (2024). https: / / doi.Org / 10.1 186 / S12859-024-05807-5
Claims
What is claimed is:1 . A method for determining allele-specific telomere length (ATL), the method comprising:(a) obtaining a long read sequencing data set;(b) extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set; c) identifying telomere variant repeat (TVR) regions by locating boundaries between telomere repeat regions and subtelomere regions; d) processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR);(e) grouping the cTVR sequences into allele-specific read clusters; and(f) determining ATL from the read clusters.
2. A method for determining allele-specific telomere length (ATL) from allelespecific telomere variant repeat (TVR) sequences, the method comprising:(a) obtaining a long read sequencing data set;(b) extracting from the long read sequencing data set a subset of sequencing reads identifying telomere region boundaries;(c) identifying individual alleles based on the presence of TVRs, comprising:(i) querying sequencing extracted reads of (b) for matches to a known set of telomere repeats,(ii) identifying TVR regions in the extracted sequencing reads, and(iii) clustering the extracted sequencing reads based on their TVR sequences and subtelomere sequences;(d) assigning each telomere allele from (c) to a chromosome arm by aligning the subtelomere sequences to a collection of subtelomere reference sequences;(e) producing a report containing the lengths and sequences of each telomere alleles, thereby determining allele-specific telomere length (ATL) and allele-specific telomere variant repeat (TVR) sequences from a biological sample.
3. The method of any one of claims 1 or 2, wherein obtaining a long read sequencing data set comprises any one or more of (a) collecting a biological sample andrunning one or more long read sequencing reactions, and / or (b) collecting a long read sequencing data set from a database.
4. The method of claim 3, wherein the database comprises assembled genome data or raw sequence read data.
5. The method of claim 4, wherein the data is in Fasta and / or Fastq formats.
6. The method of claim 3, wherein the biological sample is from a human.
7. The method of claim 2, wherein the subset of sequencing reads extracted in(b) comprises at least approximately 5 copies of canonical telomere repeats.
8. The method of claim 2, wherein the at least one reference genome containing subtelomere sequences comprises a telomere-to-telomere reference genome.
9. The method of claim 8, wherein the telomere-to-telomere reference genome is selected from the group consisting of T2T-CHM13, T2T-Yao, T2T-CN1 , T2T-HG002, T2T- ksa001 , and T2T-i002c.
10. The method of any one of the preceding claims, wherein at least operations (b) - (e) are performed using a suitably programmed computer.
11. A method for measuring all 92 allele-specific telomere lengths at nucleotide resolution, the method comprising:(a) obtaining a long read sequencing data set;(b) extracting from the long read sequencing data set a subset of sequencing reads identifying telomere region boundaries;(c) identifying individual alleles based on the presence of TVRs, comprising:(i) querying sequencing extracted reads of (b) for matches to a known set of telomere repeats,(ii) identifying TVR regions in the extracted sequencing reads, and(iii) clustering the extracted sequencing reads based on their TVR sequences and subtelomere sequences;(d) assigning each telomere allele from (c) to a chromosome arm by aligning the subtelomere sequences to a collection of subtelomere reference sequences;(e) producing a report containing the lengths and sequences of all anchored telomere alleles, thereby determining allele-specific telomere length (ATL) and allele-specific telomere variant repeat (TVR) sequences of a DNA molecule from a biological sample, thereby measuring all 92 allele-specific telomere lengths at nucleotide resolution.
12. A method of identifying individual alleles based on the composition of telomere variant repeat sequences in a sample, the method comprising:(a) obtaining a long read sequencing data set;(b) extracting from the long read sequencing data set a subset of sequencing reads identifying telomere region boundaries;(c) identifying individual alleles based on the presence of TVRs, comprising:(i) querying sequencing extracted reads of (b) for matches to a known set of telomere repeats,(ii) identifying TVR regions in the extracted sequencing reads, and(iii) clustering the extracted sequencing reads based on their TVR sequences and subtelomere sequences;(d) assigning each telomere allele from (c) to a chromosome arm by aligning the subtelomere sequences to a collection of subtelomere reference sequences;(e) producing a report containing the lengths and sequences of all anchored telomere alleles, thereby determining allele-specific telomere length (ATL) and allele-specific telomere variant repeat (TVR) sequences of a DNA molecule from a biological sample, thereby identifying individual alleles based on the composition of telomere variant repeat sequences in a sample.
13. A method for assigning TVR and telomere regions to their respective allele chromosome arm, the method comprising:(a) obtaining a long read sequencing data set;(b) extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set;c) identifying TVR regions by locating boundaries between telomere repeat regions and subtelomere regions; d) processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR);(e) grouping the cTVR sequences into allele-specific read clusters;(f) determining ATL from the read clusters; and(g) identifying the respective allele chromosome arm from the read clusters that extend into the subtelomere regions.
14. A method for determining the degree of relationship between individuals, the method comprising:(a) obtaining a long read sequencing data set from at least two individuals;(b) extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set; c) identifying TVR regions by locating boundaries between telomere repeat regions and subtelomere regions; d) processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR);(e) grouping the cTVR sequences into allele-specific read clusters;(f) determining ATL from the read clusters; and(g) determining the degree of relationship between individuals.
15. A method for assess the risk of genomic instability and tumorigenesis by determining the amount of telomere and TVR erosion, the method comprising:(a) obtaining a long read sequencing data set;(b) extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set; c) identifying TVR regions by locating boundaries between telomere repeat regions and subtelomere regions; d) processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR);(e) grouping the cTVR sequences into allele-specific read clusters;(f) determining ATL from the read clusters; and(g) determining the amount of telomere and TVR erosion.
16. A method for identifying chromosome-end fusions, the method comprising:(a) obtaining a long read sequencing data set;(b) extracting from the long read sequencing data set a subset of sequencing reads spanning telomere regions present in the long read sequencing data set; c) identifying TVR regions by locating boundaries between telomere repeat regions and subtelomere regions; d) processing the TVR regions to eliminate sequencing errors and identifying consensus TVR sequences (cTVR);(e) grouping the cTVR sequences into allele-specific read clusters;(f) determining ATL from the read clusters; and(g) identifying chromosome-end fusions.
17. The method of claim 16 wherein the chromosome-end fusions are markers for genomic instability and / or tumorigenesis.
Citation Information
Patent Citations
A T2T assembly method for plant and animal genomes based on high-throughput sequencing
CN115810395B
Method for determination of telomere length
US20040265815A1
Detection of telomere fusion events
WO2023118606A1
Methods of measuring absolute length of individual telomeres
WO2024098029A2
Cited By
Lucid ganoderma binuclear genome assembly method based on haplotype analysis
CN121565256A