Genetic risk factors for atypical frontotemporal lobar degeneration with ubiquitin inclusions
Genetic analysis of the chrl5ql4 locus using PCR and sequencing identifies aFTLD-U risk alleles, addressing the lack of ante-mortem diagnostic biomarkers for aFTLD-U, enabling early and accurate diagnosis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-03-26
AI Technical Summary
Current diagnostic methods for atypical frontotemporal lobar degeneration with ubiquitin inclusions (aFTLD-U) are limited to post-mortem neuropathologic examination, lacking reliable ante-mortem biomarkers for accurate diagnosis and differentiation from other neurodegenerative disorders.
Detection of specific genetic markers, including haplotypes and tandem repeat expansions at the chrl5ql4 locus, through genome and transcriptome analysis using methods such as PCR and sequencing, to identify aFTLD-U risk alleles and diagnose the condition ante-mortem.
Enables precise ante-mortem identification of aFTLD-U, facilitating early clinical intervention and therapeutic strategies by providing reliable genetic biomarkers for this form of FTLD.
Smart Images

Figure IMGF000028_0001 
Figure IMGF000036_0001 
Figure 00000044_0000
Abstract
Description
[0001] RoRad / REPEAT / 848
[0002] GENETIC RISK FACTORS FOR ATYPICAL FRONTOTEMPORAL LOBAR DEGENERATION WITH UBIQUITIN
[0003] INCLUSIONS
[0004] This invention was made with US Government support from the National Institutes of Health (NIH), more specifically from NINDS, with grant# UG3 NS103870. The US Government has certain rights in the invention.
[0005] FIELD OF THE INVENTION
[0006] The invention relates to genetic risk factors that are associated with frontotemporal lobar degeneration (FTLD), in particular with atypical FTLD with ubiquitin inclusions (aFTLD-U). These genetic risk factors are located on the chrl5ql4 locus and are detectable beyond the brain, thus opening avenues for more precise ante-mortem identification of this type of FTLD patients.
[0007] BACKGROUND TO THE INVENTION
[0008] Whereas frontotemporal lobar degeneration (FTLD or FTD) is accounting for about 5-15% of all dementias, it is a common cause of early-onset dementia (under age of 65) and presents with early social-emotional-behavioural and / or language changes that can be accompanied by a motor disorder depending on the brain region that is primarily affected.
[0009] The most common underlying diagnoses are linked to neuronal and glial inclusions containing tau (FTLD- tau) or TDP-43 (FTLD-TDP), although 5-10% of FTD patients may have FET pathology characterized by inclusions of multiple proteins from the FUS-Ewing sarcoma (EWS)-TAF15 family (Grossman et al. 2023, Nat Rev Disease Primers 9:40) and also transportin 1 (TNPO1), along with several other DNA / RNA- binding proteins that use TNPO1 as their nuclear import receptor.
[0010] The FTLD-FET neuropathological subtype can be further divided into atypical FTLD with ubiquitinated inclusions (aFTLD-U), neuronal intermediate filament inclusion body disease (NIFID), and basophilic inclusion body disease (BIBD) based on differences in the morphology, subcellular localization, and anatomic distribution of FET+ inclusions and immunoreactivity for other aggregating proteins such as a- internexin (in the case of NIFID). aFTLD-U is most common among these subtypes and stands out for its very characteristic clinical presentation of severe and progressive early-onset behavioural variant FTD (bvFTD), often with psychiatric symptoms, without language or motor problems. Based on this clinical presentation and the distinct feature of extensive caudate atrophy on magnetic resonance imaging (MRI), aFTLD-U can be suspected during life; yet, a definitive diagnosis can only be obtained using immunohistochemical analysis at autopsy.
[0011] The gold standard for diagnosis of the different FTLD types is post-mortem neuropathologic examination. Nevertheless, a number of biomarkers are available, especially for FTLD-tau and FTLD-TDP, allowing RoRad / REPEAT / 848 reasonably accurate ante-mortem diagnosis of the disease. Correct stratification of the patients with different FTLD forms (and distinguishing them from other neurodegenerative or psychiatric disorders with sometimes overlapping features) is important when testing candidate therapeutic agents. Clinical trials have and are being designed in the field of FTLD and defining clinically homogenous groups is of utmost importance in order to increase the success of being able to halt disease progression or of curing the disease (Panza et al. 2020, Nat Rev Neurol 16:213-228).
[0012] Genetic screening for FTD is usually involving detection of mutations in the C9orf72, MAPT and PGRN genes. In order to distinguish from other neurodegenerative or psychiatric disorders, serum and cerebrospinal fluid (CSF) biomarkers can be assessed, and furthermore imaging techniques ( M Rl, PET, ...) can be used in helping to diagnose FTD (e.g. Gifford et al. 2023, Biomarkers Neuropsychiatry 8:100065). Recently, the plasma GFAP / NfL ratio has been suggested as biomarker to distinguish sporadic FTLD-tau and FTLD-TDP during life (Cousins et al. 2022, JAMA Neurol 79:1155-1164). Even more recently, analysis of the contents of plasma extracellular vesicles (EVs) revealed that a high ratio of 3-repeat (3R) to 4- repeat (4R) tau isoforms in plasma EVs is indicative for FTLD-tau; and high levels of TDP-43 in plasma EVs being indicative for FTLD-TDP-43 (Chatterjee et al. 2024, Nature Medicine 30:1771-1783). No biomarkers, including genetic biomarkers, appear to have been described for aFTLD-U leaving postmortem autopsy as only option to conclusively diagnose this form of FTLD .
[0013] Recurrent deletions of chromosome 15ql3.3 have been associated with intellectual disability, schizophrenia, autism and epilepsy, Prader-Willi / Angelman syndromes and developmental delay (Antonacci et al. 2014, Nat Genet 46:1293-1302; Paparella et al. 2023, Int J Mol Sci 24:15818).
[0014] SUMMARY OF THE INVENTION
[0015] This disclosure relates to methods of analyzing the genome of a subject, the method comprising detecting in the DNA of a genomic DNA sample obtained from the subject one or more of:
[0016] - the haplotype of the single nucleotide variant rs549846383;
[0017] - the haplotype of the single nucleotide variant rsl48687709;
[0018] - the presence of a tandem repeat expansion relative to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1) or equivalent in an alternative human genome build version; and
[0019] - the presence of a repeat length polymorphism relative to the wild-type genomic DNA region of chrl5:34480576-34480608 according to the human genome build GRCh38 (SEQ ID NO:2) or equivalent in an alternative human genome build version.
[0020] This disclosure further relates to methods of or for analyzing or assessing the transcriptome of a subject, such methods comprising detecting, assessing or analyzing the presence of a tandem repeat expansion in the transcriptome of a transcriptomic sample obtained from the subject, wherein the tandem repeat RoRad / REPEAT / 848 expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wildtype genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1). In one embodiment, this method is included as an option to the genome analysis methods described hereinabove.
[0021] Such methods can further comprise a step of diagnosing the subject to have atypical frontotemporal lobar dementia with ubiquitinated inclusions (aFTLD-U), to be a carrier of aFTLD-U risk alleles, or to be at risk or susceptible of developing aFTLD-U when the genomic DNA of the subject comprises one or more of:
[0022] - the minor allele of the single nucleotide variant rsl48687709;
[0023] - the minor allele of the single nucleotide variant rs549846383;
[0024] - a tandem repeat expansion relative to the wild-type genomic DNA region of chrl5:34419425- 34419451 according to the human genome build GRCh38 (SEQ ID NO:1) or equivalent in an alternative human genome build version; and
[0025] - a repeat length polymorphism relative to the wild-type genomic DNA region of chrl5:34480576- 34480608 according to the human genome build GRCh38 (SEQ ID NO:2) or equivalent in an alternative human genome build version; and / or when the transcriptome of the subject comprises a tandem repeat expansion, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1).
[0026] In one embodiment to these methods, the tandem repeat expansion relative to the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1), or as transcribed into the transcriptome, comprises:
[0027] - a CT dimer-rich expansion;
[0028] - a CnT-rich sequence, optionally combined with or flanked by CT dimers;
[0029] - a CCCCT pentamer repeat, optionally combined with or flanked by CT dimers; or
[0030] -a CCCTCT hexamer repeat combined with or flanked by CT dimers.
[0031] In a particular embodiment to these methods, the tandem repeat expansion comprises a CT dimer-rich expansion.
[0032] In another particular embodiment to these methods, the tandem repeat expansion comprises at least 150 CT dimers wherein the CT dimers are contiguously or non-contiguously present in the tandem repeat expansion and are counted after deletion of CCCTCT, CCCCT, CCCT, CTTT and CCTT motifs from the tandem repeat expansion; or, alternatively, the total length of the tandem repeat expansion is at least 400 bp and wherein the expanded repeat comprises a CT dimer content of at least 80%.
[0033] In any of these methods, the detection of a single nucleotide variant haplotype, of the presence of a tandem repeat expansion and / or of the presence of the repeat length polymorphism comprises a DNA genotyping step. In one embodiment, the DNA genotyping step comprises a PCR amplification step RoRad / REPEAT / 848 and / or a DNA sequencing step, or comprises a PCR amplification step and capillary electrophoresis, or comprises a repeat-primer based PCR amplification step.
[0034] This disclosure further relates to diagnostic kits comprising an oligonucleotide, wherein the oligonucleotide enables the detection of :
[0035] - the haplotype of the single nucleotide variant rs549846383 in the DNA of a genomic DNA sample obtained from a subject;
[0036] - the haplotype of the single nucleotide variant rsl48687709 in the DNA of a genomic DNA sample obtained from a subject;
[0037] - the presence of a tandem repeat expansion in the DNA of a genomic DNA sample obtained from a subject relative to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1) or equivalent in an alternative human genome build version;
[0038] - the presence of a repeat length polymorphism in the DNA of a genomic DNA sample obtained from a subject relative to the wild-type genomic DNA region of chrl5:34480576-34480608 according to the human genome build GRCh38 (SEQ ID NO:2) or equivalent in an alternative human genome build version; or
[0039] - the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1); wherein the oligonucleotide enables the detection by means of a DNA genotyping method.
[0040] This disclosure further relates to use of a diagnostic kit in any one of the above methods wherein the detecting comprises the use of an oligonucleotide in a DNA genotyping method.
[0041] BRIEF DESCRIPTION OF THE DRAWINGS
[0042] FIGURE 1. Manhattan plot indicating that the association analysis for aFTLD-U shows a highly significant locus at chrl5ql4, with many variants supporting the signal. Each dot is representing a single nucleotide variant and its level of association with aFTLD-U. Dots in odd-numbered chromosomes are indicated in black, and dots in even-numbered chromosomes are indicated in grey.
[0043] FIGURE 2. (A) Schematic visualization of the chrl5ql4 associated locus, showing the GOLGA8A and GOLGA8B genes, the identified haplotypes A and B and the rs549846383 and rsl48687709 variants that tag those haplotypes as identified by GWAS. The segmental duplications with the highest identity are indicated by the triangle lines, leading to low mappability for short-read sequencing. The location of the RoRad / REPEAT / 848 tandem repeat is indicated, with the sequence composition observed in expansions. The GOLGA8A-B deletion, with genomic coordinates according to the HPRC, is indicated by a bar.
[0044] (B) information about the rs549846383 single nucleotide variant as retrieved from www.ncbi.nlm. nih.gov / snp / ?term=rs549846383 (top) and as retrieved from the NCBI Genome Data Viewer (human genome build GRC38) (bottom; double-stranded sequence shown, top strand sequence defined by SEQ ID NO:26).
[0045] (C) information about the rsl48687709 single nucleotide as retrieved from www.ncbi.nlm. nih.gov / snp / ?term=rsl48687709 (top) and as retrieved from the NCBI Genome Data Viewer (human genome build GRC38) (bottom; double-stranded sequence shown, top strand sequence defined by SEQ ID NO:28).
[0046] FIGURE 3. GOLGA8A repeat characteristics: length consensus with the length in nucleotides of the repeat consensus sequence of the longest allele for each individual (relative or compared to the wild-type repeat at chrl5:34419425-34419451).
[0047] FIGURE 4. Wild-type repeat sequences
[0048] (A) Information about the wild-type repeat at chrl5:34419425-34419451 as retrieved from as retrieved from the NCBI Genome Data Viewer (human genome build GRC38). The top-strand sequences of chrl5:34419425-34419451 is provided hereinafter as being defined by SEQ ID NO:1 (full top strand sequence defined by SEQ ID NO:29 and complementary sequence shown).
[0049] (B) Information about the wild-type repeat at chrl5:34480576-34480608 as retrieved from as retrieved from the NCBI Genome Data Viewer (human genome build GRC38). The top-strand sequences of chrl5:34480576-34480608 is provided hereinafter as being defined by SEQ ID NO:2 (full top strand sequence defined by SEQ ID NQ:30 and complementary sequence shown).
[0050] FIGURE 5. Repeat sequence composition plot showing a heatmap of the frequency of 12-mer motifs, which enables representing dimers, tetramers, and hexamer motifs. In both panels, every row is an individual with an expanded allele (>100bp). The "Haplotype" columns indicate the chrl5ql4 haplotype. Panel A: haplotype A in black, no associated haplotype indicated in white. Panel B: haplotype B, no associated haplotype indicated in white. In both panels, individuals with an aFTLD-U diagnosis are displayed in black in the "Phenotype" column, non-aFTLD-U controls are indicated in white.
[0051] FIGURE 6. Electropherograms of repeat-primed PCR (RP-PCR) products of 3 aFTLD-U patients (A: aFTLD- Ul; B: aFTLD-U2; C: aFTLD-U3) and one non-aFTLD-U subject (D: non-aFTLD-U) wherein the PCR primers were adjusted to detect the presence of CT-repeats at the left or right end of the region of repeat expansion ("CT-left", "CT-right", respectively), the presence of CCCTCT-motifs at the left or right end of the region of repeat expansion ("CCCTCT -left", "CCCTCT -right", respectively), the presence of CCCT- or CCCCT-motifs at the left end of the region of repeat expansion ("CCCT-left", "CCCCT-left", respectively). RoRad / REPEAT / 848
[0052] FIGURE 7. Electropherograms of repeat-primed PCR (RP-PCR) products of 22 different samples, of a blank sample, and of a run without sample and without reagents. The PCR primers were adjusted to detect the presence of CT-repeats at the left (Al to A4) or right (Bl to B4) end of the region of repeat expansion ("CT-left", "CT-right", respectively), the presence of CCCTCT-motifs at the left (Cl to C4) or right (DI to D4) end of the region of repeat expansion ("CCCTCT -left", "CCCTCT -right", respectively). Al, Bl, Cl, DI: samples 1 to 6; A2,B2,C2,D2: samples 7 to 12; A3,B3,C3,D3: samples 13 to 18; A4,B4,C4,D4: samples 18 to 22, blank, and run without sample and without reagents.
[0053] FIGURE 8. (A) Schematic representation of the chrl5:34419425-34419451 in DNA isolated from blood obtained from an aFTLD-U patient and in which the CT dimer repeat is expanded, as determined by nanopore sequencing. (B) Electropherograms of repeat-primed PCR (RP-PCR) products starting from the DNA of the same blood sample as in (A). The PCR primers were adjusted to detect the presence of CT- repeats at the left (top panel) or right (bottom panel) end of the region of repeat.
[0054] DETAILED DESCRIPTION
[0055] The current disclosure is based on work aiming at deciphering the molecular underpinning of frontotemporal lobar degeneration (FTLD or FTD) with FET pathology, FTLD-FET, characterized by inclusions of multiple proteins from the FUS-Ewing sarcoma (EWS)-TAF15 family (Grossman et al. 2023, Nat Rev Disease Primers 9:40) and more in particular to its most common variant atypical FTLD with ubiquitinated inclusions (aFTLD-U). Thereto, an international consortium to gather samples and clinicopathological information from the largest cohort of FTLD-FET patients was established. Using a common variant genome-wide association study based on short-read whole genome sequencing data from 59 aFTLD-U patients and 3150 controls, a major association locus at chrl5ql4 (a notoriously repetitive and complex genomic region; Antonacci et al. 2014, Nat Genet 46:1293-1302) and more specifically 2 haplotypes were identified that could not be thoroughly investigated with short reads alone. Subsequently, in-house and public long-read genome sequencing (ONT) data from more than 1500 individuals were leveraged, which led to the identification of a tandem repeat expansion intronic in the GOLGA8A gene and an intergenic repeat length polymorphism on the associated haplotypes. A repeat-primed PCR (RP-PCR) method was subsequently designed enabling analysis of the tandem repeat expansion, in particular the tandem repeat expansion composition. The identified tandem repeat expansion was also detected beyond the brain tissue such as in DNA from blood or lymphoblast cell line cultures of aFTLD-U patients, and was also detected in transcriptomic samples of aFTLD-U patients (i.e. in cDNA synthesized from transcribed RNA). RoRad / REPEAT / 848
[0056] Therefore, this disclosure in essence relates to methods of or for analyzing or assessing the genome of a subject, such methods comprising detecting, assessing or analyzing in a genomic DNA sample obtained from the subject or in the DNA of a genomic DNA sample obtained from the subject one or more of:
[0057] - the haplotype of the single nucleotide variant rs549846383;
[0058] - the haplotype of the single nucleotide variant rsl48687709;
[0059] - the presence of a tandem repeat expansion relative or compared to the wild-type genomic
[0060] DNA region of chrl5:34419425-34419451 (SEQ ID NO:1 / relative or compared to SEQ ID NO:1); and
[0061] - the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608 (SEQ ID NO:2 / relative or compared to SEQ ID NO:2).
[0062] Alternatively, this disclosure relates to methods of or for analyzing or assessing the transcriptome of a subject, such methods comprising detecting, assessing or analyzing the presence of a tandem repeat expansion in the transcriptome of a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1).
[0063] Alternatively, this disclosure relates to methods of or for analyzing or assessing the genome and / or transcriptome of a subject, such methods comprising detecting, assessing or analyzing one or more of:
[0064] - the haplotype of the single nucleotide variant rs549846383 in the DNA of a genomic DNA sample obtained from the subject;
[0065] - the haplotype of the single nucleotide variant rsl48687709 in the DNA of a genomic DNA sample obtained from the subject;
[0066] - the presence of a tandem repeat expansion in the DNA of a genomic DNA sample obtained from the subject relative or compared to the wild-type genomic DNA region of chrl5:34419425- 34419451 (SEQ ID NO:1 / relative or compared to SEQ ID NO:1);
[0067] - the presence of a repeat length polymorphism in the DNA of a genomic DNA sample obtained from the subject relative or compared to the wild-type genomic DNA region of chrl5:34480576- 34480608 (SEQ ID NO:2 / relative or compared to SEQ ID NO:2); and
[0068] - the presence of a tandem repeat expansion in the RNA or transcriptome of a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1).
[0069] Example 3 herein provides the nucleotide sequences as defined by SEQ ID Nos: 1 and 2.
[0070] A genomic DNA sample is a biological sample that contains genomic DNA. A transcriptomic sample in general is a biological sample that contains RNA or that at least contains transcribed RNA. The origin of RoRad / REPEAT / 848 a genomic DNA sample and a transcriptomic sample may be different or the same, e.g. any biological sample such as a cell, tissue sample, biofluid sample. The same biological sample may optionally be used to analyze or assess genomic DNA and the transcriptome. With "transcriptome corresponding to or covering a genomic DNA region" wherein the genomic DNA region is a genomic DNA region of interest, is meant the RNA / transcriptome (or the fraction of the whole RNA / transcriptome) that is transcribed from a genomic DNA region comprising the DNA region of interest. In particular, the wild-type genomic DNA region of interest is the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1). The phrasing "tandem repeat expansion in a transcriptome relative or compared to a wild-type transcriptome corresponding to or covering a wild-type genomic DNA region (of interest)" can throughout this document be exchanged for "tandem repeat expansion in the RNA or transcriptome of a biological sample (comprising the RNA or transcriptome)(sample obtained from a subject) relative or compared to a wild-type RNA or transcriptome (or RNA or transcriptome profile) mapped or aligned to a wild-type genomic DNA region (of interest)"; or can be exchanged for ""tandem repeat expansion in the RNA or transcriptome of a biological sample (comprising the RNA or transcriptome)(sample obtained from a subject) relative or compared to a wild-type genomic DNA region (of interest) after mapping or aligning the RNA or transcriptome to the wild-type genomic DNA region (of interest)".
[0071] In particular, the reference to the human genome positions hereinabove and hereinafter are relative to or according to the human genome build GRCh38 (Genome Reference Consortium Human Build 38) unless mentioned otherwise, or to equivalent positions in an alternative human genome build version. A defined sequence as present in human genome build GRCh38 can routinely be traced back in an alternative version or build of the human genome by means of an appropriate bioinformatics tool (e.g. LiftOver tool, Assembly Converter).
[0072] Information and sequence contexts (GRCh38 sequence contexts) of the haplotypes of the single nucleotide variants rs549846383 and rsl48687709 are depicted in Figures 2B and 2C, respectively. In particular, the minor allelic variant of rs549846383 is lacking 1 thymine (T) nucleobase compared to the major allelic variant and is therefore referred to herein as single nucleotide variant although the information in Figure 2B defines rs549846383 as "indel" (insertion / deletion) or "delin" (deletion / insertion) variant. In particular, the minor allelic variant of rsl48687709 is comprising a cytosine (C) instead of a thymine (T) nucleobase in the major allelic variant.
[0073] The presence of a tandem repeat expansion relative to the wild-type genomic DNA region of chrl5:34419425-34419451 as defined by SEQ ID NO:1 is thus in particular according to the human genome build GRCh38. The equivalent of SEQ ID NO:1 can easily be assigned in an alternative human genome build version - Figures 2B and 2C, for instance, provide the start positions of rs549846383 and rsl48687709, respectively, in human genome builds GRCh38 and GRCh37. Likewise, the presence of a repeat length polymorphism relative to the wild-type genomic DNA region of chrl5:34480576-34480608 RoRad / REPEAT / 848 as defined by SEQ ID NO:2 is thus in particular according to the human genome build GRCh38. The equivalent of SEQ ID NO:2 can easily be assigned in an alternative human genome build version.
[0074] When not specifically mentioned, throughout this document, "wild-type genomic DNA region of chrl5:34419425-34419451" is the genomic DNA region according to the human genome build GRCh38 and is defined by SEQ ID NO:1. When not specifically mentioned, throughout this document, "wild-type genomic DNA region of chrl5:34480576-34480608" is the genomic DNA region according to the human genome build GRCh38 and is defined by SEQ ID NO:2.
[0075] The wild-type genomic DNA region of chrl5:34419425-34419451 is intronic to the GOLGA8A gene and is, according to GRCh38, defined by SEQ ID NO:1 (see Example 3 and Figure 4A, comprised in SEQ ID NO:29), or alternatively by the sequence complementary to SEQ ID NO:1 (as is likewise depicted in Figure 4A) or by the reverse complement sequence of SEQ ID NO:1. In case of transcription of the GOLG8A gene region comprising an expanded repeat in this region into (m)RNA, the direction of the transcription is determining the sequence of the (m)RNA, and any thymine (T) is appearing as a uridine (U).
[0076] The wild-type genomic DNA region of chrl5:34480576-34480608, according to GRCh38, is defined by SEQ ID NO:2 (see Example 3 and Figure 4B, comprised in SEQ ID NQ:30), or alternatively by the sequence complementary to SEQ ID NO:2 (as is likewise depicted in Figure 4B) or by the reverse complement sequence of SEQ ID NO:2. In case of transcription into (m)RNA, the direction of the transcription is determining the sequence of the (m)RNA, and any thymine (T) is appearing as a uridine (U).
[0077] Those skilled in the art will readily recognize that nucleic acid molecules may be double-stranded molecules and that reference to a particular site on one strand refers as well to the corresponding site on a complementary strand. Nucleic acid molecules such as RNA, mRNA and denatured DNA are singlestranded. In defining a variant position, allele, or nucleotide sequence, reference to an adenine, a thymine (uridine in case of (m)RNA), a cytosine, or a guanine at a particular site on one strand of a nucleic acid molecule also defines the thymine (uridine in case of (m)RNA), adenine, guanine, or cytosine, respectively, at the corresponding site on a complementary strand of the nucleic acid molecule. Thus, reference may be made to either strand in order to refer to a particular variant position, allele, or nucleotide sequence. Probes and primers may be designed and / or routinely adapted to hybridize to either strand and nucleic acid fingerprinting or genotyping methods referred to herein may generally target either strand, again routinely adapted to the nucleic acid region of interest to be detected, assessed or analyzed.
[0078] "Tandem repeat expansion" refers to expansion in length of a wild-type tandem repeat present in a genome due to multiplication of the wild-type repeat unit or of part of the wild-type repeat unit. A number of diseases are known as "repeat expansion" diseases including myotonic dystrophy types 1 and 2, dentatorubral-pallidoluysian atrophy, progressive myoclonus epilepsy type I, Fragile X syndrome, mental retardation, Friedreich ataxia, Huntington disease, Huntington disease-like 2, spinobulbar RoRad / REPEAT / 848 muscular atrophy and spinocerebellar ataxia (Usdin 2008, Genome Res 18:1011-1019). The generalized term "tandem repeat expansion / repeat length polymorphism relative or compared to a wild-type sequence" is referring to the presence of an expanded repeat / polymorphism in the tandem repeat expansion of a test sample (genomic or transcriptomic) whereas the expanded repeat / polymorphism is absent in the wild-type sequence (genomic or transcriptomic). As evident from the above, such expansion can be detected in a DNA sample (such as a genomic DNA sample), and can be detected in a transcriptomic sample (transcribed RNA or cDNA synthesized from transcribed RNA).
[0079] The above methods thus alternatively relate to methods of or for analyzing the genome of a subject, comprising detecting or analyzing in a genomic DNA sample obtained from the subject or in the DNA of a genomic DNA sample obtained from the subject:
[0080] - the haplotype of the single nucleotide variant rs549846383; or
[0081] - the haplotype of the single nucleotide variant rsl48687709; or
[0082] - the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0083] - the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; or
[0084] - the haplotype of the single nucleotide variant rs549846383 and the haplotype of the single nucleotide variant rsl48687709; or
[0085] - the haplotype of the single nucleotide variant rs549846383 and the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425- 34419451; or
[0086] - the haplotype of the single nucleotide variant rs549846383 and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576- 34480608; or
[0087] - the haplotype of the single nucleotide variant rsl48687709 and the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425- 34419451; or
[0088] - the haplotype of the single nucleotide variant rsl48687709 and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576- 34480608; or
[0089] - the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451 and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; or RoRad / REPEAT / 848
[0090] - the haplotype of the single nucleotide variant rs549846383, the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425- 34419451, and the presence of a repeat length polymorphism relative or compared to the wildtype genomic DNA region of chrl5:34480576-34480608; or
[0091] - the haplotype of the single nucleotide variant rsl48687709, the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425- 34419451, and the presence of a repeat length polymorphism relative or compared to the wildtype genomic DNA region of chrl5:34480576-34480608; or
[0092] - the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, and the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0093] - the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; or
[0094] - the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, the presence of a tandem repeat expansion relative or compared to the wildtype genomic DNA region of chrl5:34419425-34419451, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576- 34480608.
[0095] Alternatively, the above methods alternatively relate to methods of or for analyzing or assessing the genome and / or transcriptome of a subject, such methods comprising detecting, assessing or analyzing one or more of:
[0096] - the haplotype of the single nucleotide variant rs549846383 in the DNA of a genomic DNA sample obtained from the subject; or
[0097] - the haplotype of the single nucleotide variant rsl48687709 in the DNA of a genomic DNA sample obtained from the subject; or
[0098] - the presence of a tandem repeat expansion in the DNA of a genomic DNA sample obtained from the subject relative or compared to the wild-type genomic DNA region of chrl5:34419425- 34419451; or
[0099] - the presence of a repeat length polymorphism in the DNA of a genomic DNA sample obtained from the subject relative or compared to the wild-type genomic DNA region of chrl5:34480576- 34480608; or
[0100] - the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or RoRad / REPEAT / 848 the haplotype of the single nucleotide variant rs549846383 and the haplotype of the single nucleotide variant rsl48687709 in the DNA of a genomic DNA sample obtained from the subject; or
[0101] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383 and the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0102] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383 and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; or
[0103] - the haplotype of the single nucleotide variant rs549846383 in the DNA of a genomic DNA sample obtained from the subject; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0104] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rsl48687709 and the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0105] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rsl48687709 and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; or
[0106] - the haplotype of the single nucleotide variant rsl48687709 in the DNA of a genomic DNA sample obtained from the subject; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0107] - in the DNA of a genomic DNA sample obtained from the subject: the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425- 34419451 and the presence of a repeat length polymorphism relative or compared to the wildtype genomic DNA region of chrl5:34480576-34480608; or
[0108] - the presence of a tandem repeat expansion in the DNA of a genomic DNA sample obtained from the subject relative or compared to the wild-type genomic DNA region of chrl5:34419425- 34419451; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or RoRad / REPEAT / 848
[0109] - the presence of a repeat length polymorphism in the DNA of a genomic DNA sample obtained from the subject relative or compared to the wild-type genomic DNA region of chrl5:34480576- 34480608; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0110] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, and the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0111] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; or
[0112] - the haplotype of the single nucleotide variant rs549846383 and the haplotype of the single nucleotide variant rsl48687709 in the DNA of a genomic DNA sample obtained from the subject; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0113] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451, and the presence of repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; or
[0114] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0115] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, and the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the RoRad / REPEAT / 848 tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0116] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rsl48687709, the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451, and the presence of repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; or
[0117] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rsl48687709, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0118] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rsl48687709, and the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0119] - in the DNA of a genomic DNA sample obtained from the subject: the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425- 3441945, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0120] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; or
[0121] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, and the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451; and the presence of a tandem repeat expansion in a RoRad / REPEAT / 848 transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0122] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0123] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451, the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0124] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rsl48687709, the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451; or
[0125] - in the DNA of a genomic DNA sample obtained from the subject: the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451, and the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608; and the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451. RoRad / REPEAT / 848
[0126] The terms "allele" and "allelic variant" are used interchangeably herein. An allele is any one of two or of a number of alternative forms or sequences of the same gene occupying a given locus or position on a chromosome. A single allele for each locus is inherited separately from each parent, resulting in two alleles for each gene (germline alleles). An individual having two copies of the same allele of a particular gene is homozygous at that locus, whereas an individual having two different alleles of a particular gene is heterozygous. Somatic allelic variations readily occur or accumulate post conception.
[0127] The term "complement" as used herein means the complementary sequence to a nucleic acid sequence according to standard Watson / Crick base pairing rules. A complement sequence can also be a sequence of RNA complementary to the DNA sequence or its complement sequence, and can also be a cDNA. The term "substantially complementary" as used herein means that two sequences hybridize under stringent hybridization conditions. The skilled artisan will understand that substantially complementary sequences need not hybridize along their entire length. In particular, substantially complementary sequences comprise a contiguous sequence of bases that do not hybridize to a target or marker sequence, positioned 3 ' or 5' to a contiguous sequence of bases that hybridize under stringent hybridization conditions to a target or marker sequence.
[0128] The term "genomic DNA" refers to the entire DNA or to part of the entire DNA originating from the nucleus of a cell (as opposed to mitochondrial DNA). Genomic DNA may be intact or fragmented. In some embodiments, genomic DNA may include the nucleic acid sequence from all or a portion of a single gene or from multiple genes, the nucleic acid sequence from one or more chromosomes, or the nucleic acid sequence from all chromosomes of a cell. The term "total genomic nucleic acid" is used herein to refer to the full DNA contained in the genome of a cell. As is well known, genomic nucleic acid includes gene coding regions, introns, 5' and 3' untranslated regions, 5' and 3' flanking DNA and structural segments such as telomeric and centromeric DNA, replication origins, and intergenic DNA. Genomic nucleic acid may be obtained from the nucleus of a cell, or recombinantly produced. For purposes of analyzing or assessing genomic DNA, amplification techniques and / or capture techniques may be used.
[0129] Nucleic acid regions of interest (e.g. genomic DNA regions of interest or transcriptomic regions of interest) can be captured from the nucleic acid sample by means of (hybridization to) an oligonucleotide probe. Alternatively, a plurality of nucleic acid regions of interest, is captured by or captured by means of (hybridization to) a library of oligonucleotide probes (capture library / nucleic acid capture), or captured by or captured by means of (hybridization to) a plurality of oligonucleotide probes. In a particular embodiment thereto, such oligonucleotide probe(s) are attached to a solid support such as a sheet, bead, or plate well (of any suitable size or dimension). In a particular embodiment thereto, a spacer is introduced between the solid support and the oligonucleotide probe(s) to support efficient capture of the nucleic acid of interest from the sample such as biological sample. This process is also known as hybridization capture or target enrichment. The process may include prior shearing, such as RoRad / REPEAT / 848 mechanical shearing, of the input material. The nucleic acid may in principle be obtained from any type of biological sample comprising nucleic acids including fresh or processed (e.g. FFPE, fresh frozen) samples, nucleic acids isolated from blood or blood cells, nucleic acids isolated from plasma, nucleic acids isolated from an organ, nucleic acids isolated from cerebrospinal fluid (CSF), etc. In a particular embodiment, the genomic DNA region of interest is actively expressed and the analysis can be adapted accordingly and be performed on an expression product such as mRNA or the cDNA derived from it (i.e. analysis of transcriptome or analysis of a transcriptomic sample), alternatively or in addition to analysis of the genomic DNA.
[0130] As outlined in detail in the Examples, analyzing in the DNA of a genomic DNA sample obtained from a subject any of the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451, and / or the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608 contributes to the identification or diagnosis of a subject to have atypical frontotemporal lobar dementia with ubiquitinated inclusions (aFTLD-U), to be a carrier of aFTLD-U risk alleles, or to be at risk or susceptible of developing aFTLD-U. This is in particular the case when the DNA of the genomic DNA sample of / obtained from the subject is comprising or comprises, such as detected or determined following its analysis or assessment, one or more of: the minor allele of the single nucleotide variant rsl48687709, the minor allele of the single nucleotide variant rs549846383, a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1) or equivalent in an alternative human genome build version, and a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608 according to the human genome build GRCh38 (SEQ ID NO:2) or equivalent in an alternative human genome build version.
[0131] Alternatively or further contributing to the identification or diagnosis of a subject to have atypical frontotemporal lobar dementia with ubiquitinated inclusions (aFTLD-U), to be a carrier of aFTLD-U risk alleles, or to be at risk or susceptible of developing aFTLD-U is analyzing, assessing or determining the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1).
[0132] Thus, this disclosure likewise relates to (a) any of the above methods further comprising (a step of) diagnosing or identifying a subject to have atypical frontotemporal lobar dementia with ubiquitinated inclusions (aFTLD-U), to be a carrier of aFTLD-U risk alleles, or to be at risk or susceptible of developing aFTLD-U; or (b) methods of or for diagnosing or identifying a subject to have atypical frontotemporal RoRad / REPEAT / 848 lobar dementia with ubiquitinated inclusions (aFTLD-U), to be a carrier of aFTLD-U risk alleles, or to be at risk or susceptible of developing aFTLD-U; when the genomic DNA of the subject or the DNA in the genomic DNA sample obtained from the subject comprises one or more of (see supra for more extensive lists exemplifying "one or more of"), or is, as a result of the analysis or assessment, confirmed to comprise one or more of:
[0133] - the minor allele of the single nucleotide variant rsl48687709;
[0134] - the minor allele of the single nucleotide variant rs549846383;
[0135] - a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1 / relative or compared to SEQ ID NO:1) or equivalent in an alternative human genome build version; and
[0136] - a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608 according to the human genome build GRCh38 (SEQ ID NO:2 / relative or compared to SEQ ID NO:2) or equivalent in an alternative human genome build version. and / or when the transcriptome of the subject comprises (or is, as a result of the analysis or assessment, confirmed to comprise) a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1).
[0137] In one embodiment thereto, such methods are methods comprising diagnosing or identifying a subject to have aFTLD-U or to be at risk or susceptible of developing aFTLD-U when the genomic DNA of the subject or the DNA in the genomic DNA sample obtained from the subject comprises:
[0138] - the minor allele of the single nucleotide variant rs549846383, or the minor allele of the single nucleotide variant rs549846383 and the minor allele of the single nucleotide variant rsl48687709; and
[0139] - a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1 / relative or compared to SEQ ID NO:1) or equivalent in an alternative human genome build version, wherein the tandem repeat expansion comprises a long CT dimer repeat.
[0140] In another embodiment, such methods are methods comprising diagnosing or identifying the subject to be a carrier of aFTLD-U risk alleles and likely not to have aFTLD-U or not to be at risk or susceptible of developing aFTLD-U when the genomic DNA of the subject or the DNA in the genomic DNA sample obtained from the subject comprises: the minor allele of the single nucleotide variant rsl48687709, or the minor allele of the single nucleotide variant rsl48687709 and the minor allele of the single nucleotide variant rs549846383; and RoRad / REPEAT / 848
[0141] - a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1 / relative or compared to SEQ ID NO:l)or equivalent in an alternative human genome build version, wherein the tandem repeat expansion comprises a CCCT tetramer repeat flanked by a short CT dimer repeat.
[0142] In another embodiment, such methods are methods comprising diagnosing or identifying the subject likely not to have aFTLD-U or not to be at risk or susceptible of developing aFTLD-U when the genomic DNA of the subject or the DNA in the genomic DNA sample obtained from the subject: is not comprising the minor allele of the single nucleotide variant rs549846383, is not comprising the minor allele of the single nucleotide variant rsl48687709, or is not comprising the minor allele of the single nucleotide variant rsl48687709 and not comprising the minor allele of the single nucleotide variant rs549846383; and is comprising a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1 / relative or compared to SEQ ID NO:l)or equivalent in an alternative human genome build version, wherein the tandem repeat expansion comprises a CCTT tetramer repeat.
[0143] In particular, in any of the above methods, the minor allele of the single nucleotide variant rsl48687709 is having a cytosine (C) whereas the major allele is having a thymine (T) (Figure 2C). In particular, in any of the above methods, the minor allele of the single nucleotide variant rs549846383 is lacking a thymine compared to the major allele (Figure 2B)(as described supra).
[0144] Further in particular, in any of the above methods and if not already specified, the tandem repeat expansion in a genomic DNA relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451 (relative or compared to SEQ ID NO:1), or the tandem repeat expansion in a transcriptome relative or compared to a wild-type transcriptome corresponding to or covering the wildtype genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1) comprises one of:
[0145] - a CT dimer-rich expansion; or
[0146] - a CnT-rich sequence, optionally combined with or flanked by CT dimers; or
[0147] - a CCCCT pentamer repeat, optionally combined with or flanked by CT dimers; or
[0148] -a CCCTCT hexamer repeat combined with or flanked by CT dimers.
[0149] In one embodiment thereto, the tandem repeat expansion is a CT dimer-rich expansion.
[0150] In a further embodiment thereto, the tandem repeat expansion is comprising at least 300 basepairs (bp) (or at least 300 bases or nucleotides in case of an (m)RNA or cDNA) of CT dimers. The at least 300 bp (or bases) can be at least 310 bp (or bases), at least 320 bp (or bases), at least 330 bp (or bases), at least 340 bp (or bases), at least 350 bp (or bases), at least 360 bp (or bases), at least 370 bp (or bases), at least 380 RoRad / REPEAT / 848 bp (or bases), at least 390 bp (or bases), at least 400 bp (or bases) can be at least 450 bp (or bases), at least 500 bp (or bases), at least 550 bp (or bases), at least 600 bp (or bases), at least 650 bp (or bases), at least 700 bp (or bases), at least 750 bp (or bases), at least 800 bp (or bases), at least 850 bp (or bases), at least 900 bp (or bases), at least 950 bp (or bases), or at least 1000 bp (or bases). In particular, a tandem repeat expansion is identified or determined as comprising at least 300 bp (or bases) of CT dimers by removing from the expanded repeat all motifs that are CCCTCT (hexamer), CCCCT (pentamer), CCCT, CTTT, and CCTT (tetramers) followed by counting the number of CT dimers in the remaining sequence. The number of counted CT dimers is then multiplied by two (2) to arrive at the number of basepairs (or bases) consisting of CT dimers. For the sake of clarity, the CT dimers in the tandem repeat sequence from which the other motifs are removed do not have to occur as a contiguous or continuous uninterrupted stretch of CT dimers. Thus, alternatively phrased, the tandem repeat expansion comprises, relative or compared to SEQ ID NO:1, at least 150 CT dimers (or GA dimers in case of looking at the complement sequence, or AG dimers in case of looking at the reverse complement sequence), or at least 155, at least 160, at least 165, at least 170, at least 175, at least 180, at least 185, at least 190, at least 195, at least 200, at least 225, at least 250, at least 275, at least 300, at least 325, at least 350, at least 375, at least 400, at least 425, at least 450, at least 475 or at least 500 CT dimers in the expanded repeat from which CCCTCT, CCCCT, CCCT, CTTT and CCTT motifs (if present) have been omitted, removed or deleted / after omission, removal or deletion of CCCTCT, CCCCT, CCCT, CTTT and CCTT motifs (if present) from the expanded repeat.
[0151] In an alternative further embodiment thereto, the total length of the expanded repeat is at least 400 basepairs (bp) (or at least 400 bases or nucleotides in case of an (m)RNA or cDNA) and the expanded repeat comprises a CT dimer (or GA dimer in case of looking at the complement sequence, or AG dimer in case of looking at the reverse complement sequence) content of at least 80%. The at least 400 bp (or bases) can be at least 450 bp (or bases), at least 500 bp (or bases), at least 550 bp (or bases), at least 600 bp (or bases), at least 650 bp (or bases), at least 700 bp (or bases), at least 750 bp (or bases), at least 800 bp (or bases), at least 850 bp (or bases), at least 900 bp (or bases), at least 950 bp (or bases), or at least 1000 bp (or bases). The CT dimer (or GA dimer in case of looking at the complement sequence, or AG dimer in case of looking at the reverse complement sequence) content of at least 80% can be at least 85%, at least 90%, or at least 95%.
[0152] In a further embodiment, the consensus CnT-rich sequence is a C62T (n =62) sequence.
[0153] In case of transcription of a region comprising a tandem repeat expansion or an expanded repeat into (m)RNA, the direction of the transcription is determining the sequence of the (m)RNA, and any thymine (T) is appearing as a uridine (U). Thus when transcribed, and when performing any of the above methods starting from a transcriptome / (m)RNA for detection of e.g. CT dimers (relative or compared to SEQ ID RoRad / REPEAT / 848
[0154] NO:1), then the detection will be looking for CU dimer repeats or AG dimer repeats in the (m)RNA depending on the direction of the transcription.
[0155] Further in particular, the repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608 (SEQ ID NO:2) is a sequence polymorphism which is on average 6 basepairs (bp) longer compared to the wild-type genomic DNA region and does not comprise an actual expansion.
[0156] In particular, in any of the above methods, the detection of any of the minor allele of the single nucleotide variant rsl48687709, the minor allele of the single nucleotide variant rs549846383, the tandem repeat expansion (in a genomic DNA and / or in a transcriptome) relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451, or the repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608, comprises a nucleic acid fingerprinting or nucleic acid genotyping, such as a DNA fingerprinting step or DNA genotyping step.
[0157] With nucleic acid fingerprinting or nucleic acid genotyping is meant any available methodology that is in particular able to detect the presence of a single nucleotide variant and / or the presence of a tandem repeat expansion or of a repeat length polymorphism in a nucleic acid molecule of interest (and compared to a canonic wild-type nucleic acid molecule). The nucleic acid can be DNA (such as genomic DNA or cDNA) or RNA (such as mRNA). Examples of nucleic acid fingerprinting / genotyping methods or assays include allele-specific hybridization, allele-specific amplification, 5'-nuclease assays, single-strand conformation polymorphism analysis, melting curve analysis, assays based on molecular beacon probes, restriction fragment length polymorphism (RFLP) assays, amplified fragment length polymorphism (AFLP) assays, quantitative fluorescent polymerase chain reaction (QF-PCR) assays, PCR amplification of short tandem repeats (STRs), DNA sequencing (possibly combined with a PCR amplification step), repeat- primed PCR, etc. The QF-PCR assay uses fluorescent-labelled primers to amplify the DNA fragments followed by analysis of capillary electrophoresis.
[0158] In one particular embodiment, the sequencing is targeted (e.g. involving linear, non-linear or PCR amplification and / or capture by hybridization of a nucleic acid region of interest) or untargeted (e.g. whole genome, exome, transcriptome) sequencing. In general, sequencing methods may include, but are not limited to: high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, singlemolecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, RNA-Seq, Digital Gene Expression, Next generation sequencing, Single Molecule Sequencing by Synthesis (SMSS), massive parallel sequencing, Clonal Single Molecule Array, RoRad / REPEAT / 848 shotgun sequencing, Maxam-Gilbert or Sanger sequencing, primer walking, sequencing using PacBio, SOLID, Ion Torrent, or Nanopore platforms, short read sequencing, long read sequencing, massive parallel sequencing, and any other sequencing methods known in the art.
[0159] Any of the above methods in one particular embodiment are in vitro methods, or methods on biologicals samples having been obtained from a subject or patient. In any of the above, a subject or patient in general is a mammalian species. The mammalian species in general is a higher species including primates, cattle (e.g. cows, sheep, goats, pigs), horses, and pets (e.g. dogs, cats). In one embodiment the subject or patient is a human subject or patient.
[0160] The term "diagnosis / diagnosing" or "identification / identifying" means detecting or determining a disease state or condition, or detecting determining potential susceptibility, predisposition or risk to develop a disease state or condition, in a patient in such a way as to inform a health care provider as to the necessity or suitability of a treatment, or follow-up for the patient.
[0161] Kits
[0162] In view of all above, kits, such as diagnostic or prognostic kits or kits of parts, can be designed that are tailored to enable any of the methods / uses disclosed herein. Such kits can be comprising the tools to detect, such as by means of a suitable nucleic acid fingerprinting or genotyping assay, the haplotype of the single nucleotide variant rs549846383, the haplotype of the single nucleotide variant rsl48687709, the presence of a tandem repeat expansion (in a genomic DNA and / or in a transcriptome) relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451, and / or the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576-34480608. In particular, such kits comprise tools such as one or more oligonucleotides such as a capture probe or such as a primer or probe enabling the detection of a haplotype, repeat expansion or repeat length polymorphism such as when applied in a nucleic acid fingerprinting or genotyping method (as described above). The oligonucleotides can thus be primers and / or probes. In particular, the primers and / or probes are labelled primers and / or probes; primers and / or probes comprising non-naturally occurring nucleotides or comprising nucleotides with a non-naturally occurring modification (i.e. a non-naturally occurring or man-made oligonucleotide, primer or probe); hairpin or structurally locked primers and / or probes; or a combination thereof. In particular, the oligonucleotides comprise at least a sequence specifically hybridizing to a nucleic acid of interest or to part of a nucleic acid of interest. In particular, such oligonucleotide is comprising at least one modified or non-naturally occurring nucleotide. Further in particular, the oligonucleotide may be part of a primer and probe set, of which set at least one primer or probe is comprising a sequence specifically hybridizing to a nucleic RoRad / REPEAT / 848 acid of interest or to part of a nucleic acid of interest. Further in particular, the oligonucleotide(s) are specifically designed to enable amplification and / or detection of a repeat-containing nucleic acid sequence. Such kits / diagnostic kits can alternatively comprise a multi-membered set of oligonucleotides, wherein each member of the set comprises at least one modified or non-naturally occurring nucleotide and a sequence specifically hybridizing to a nucleic acid of interest or to part of a nucleic acid of interest. Such kits / diagnostic kits can alternatively comprise a plurality of separate primer and probe sets, wherein each set is comprising a primer or probe comprising of which at least one of the primer or probe is comprising a modified or non-naturally occurring nucleotide, and wherein each set comprises a primer or probe of which at least one of the primer or probe is comprising a sequence specifically hybridizing to a nucleic acid of interest or to part of a nucleic acid of interest. A non-naturally occurring nucleotide (such as a labelled nucleotide) may be a nucleotide that is chemically different from a nucleotide present in a living cell, or may be a chemically naturally occurring nucleotide but which is mutated relative to the natural target nucleic acid on which the oligonucleotide is specifically hybridizing.
[0163] Such kits may further comprise one or more reagent such as reagents required for extraction of DNA from cells, reagents for amplification of DNA, hybridization reagents, DNA intercalating dyes, reaction vials, and a kit instruction manual or a kit manual.
[0164] Thus, the current disclosure further relates to diagnostic kits comprising one or more tools, wherein the one or more tools enable the detection, in (the DNA of) a genomic DNA sample obtained from a subject, of one or more of (see supra for more extensive list):
[0165] - the haplotype of the single nucleotide variant rs549846383;
[0166] - the haplotype of the single nucleotide variant rsl48687709;
[0167] - the presence of a tandem repeat expansion relative or compared to the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1) according to the human genome build GRCh38 or equivalent in an alternative human genome build version; - the presence of a repeat length polymorphism relative or compared to the wild-type genomic DNA region of chrl5:34480576- 34480608 (SEQ ID NO:2) according to the human genome build GRCh38 or equivalent in an alternative human genome build version; and / or- wherein the one or more tools enable the detection, in a transcriptomic sample obtained from a subject, the presence of a tandem repeat expansion wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451 (SEQ ID NO:1).
[0168] In one embodiment thereto, the one or more tools is / are (an) oligonucleotide(s) enabling / enables the detection by means of a DNA fingerprinting or DNA genotyping method. As described above, such oligonucleotide can in particular embodiments either be a capture probe or an oligonucleotide enabling RoRad / REPEAT / 848 detection by means of PCR amplification and / or sequencing, by means of PCR amplification, or by means of PCR amplification and capillary electrophoresis. In a further particular embodiment, such oligonucleotide comprises one or more non-naturally occurring nucleotides or comprises one or more nucleotides with a non-naturally occurring modification. Exemplary oligonucleotides are listed in Tables 1 and 2 hereinafter.
[0169] In a further embodiment thereto, such diagnostic kits can in addition comprise one or more tools enabling the detection of FTLD-tau, FTLD-TDP, neuronal intermediate filament inclusion body disease (NIFID), and / or basophilic inclusion body disease (BIBD). In yet a further embodiment thereto, such diagnostic kits can in addition comprise one or more tools enabling the detection of a non-FTLD neurological disease or condition, such as e.g. Alzheimer disease, amyotrophic lateral sclerosis (ALS), Parkinson's disease, etc. More in particular, these tools are one or more further oligonucleotide(s) specific for detecting a biomarker of non-aFTLD-U.
[0170] Genetic screening for FTLD is usually involving detection of mutations in the C9orf72, MAPT and PGRN genes. In order to distinguish from other neurodegenerative or psychiatric disorders, serum and cerebrospinal fluid (CSF) biomarkers can be assessed, and furthermore imaging techniques (MRI, PET, ...) can be used in helping to diagnose FTLD (e.g. Gifford et al. 2023, Biomarkers Neuropsychiatry 8:100065). Recently, the plasma GFAP / NfL ratio has been suggested as biomarker to distinguish sporadic FTLD-tau and FTLD-TDP during life (Cousins et al. 2022, JAMA Neurol 79:1155-1164). Even more recently, analysis of the contents of plasma extracellular vesicles (EVs) revealed that a high ratio of 3-repeat (3R) to 4- repeat (4R) tau isoforms in plasma EVs is indicative for FTLD-tau; and high levels of TDP-43 in plasma EVs being indicative for FTLD-TDP-43 (Chatterjee et al. 2024, Nature Medicine 30:1771-1783).
[0171] The current disclosure further relates to use of such diagnostic kit in any of the above-described methods. In particular such diagnostic kit is comprising the one or more tools as described hereinabove.
[0172] Methods of screening or stratification of subjects
[0173] Any of the above-described methods and kits may be applied in a method of or for screening a population of subjects having or suspected to have a neurological disease or disorder, such methods comprising identification or diagnosing a subject as aFTLD-U patient or as suspected aFTLD-U patient, or as being at risk or susceptible of developing aFTLD-U, based on / using any of the above-described methods and kits.
[0174] Any of the above-described methods and kits may be applied in a method of or for stratifying a population of subjects having or suspected to have a neurological disease or disorder, such methods comprising identification or diagnosing a subject as aFTLD-U patient or as suspected aFTLD-U patient, or RoRad / REPEAT / 848 as being at risk or susceptible of developing aFTLD-U, based on / using any of the above-described methods and kits.
[0175] Any of the above-described methods and kits may be applied in a method of or for screening a population of subjects having or suspected to have FTLD, such method comprising identification or diagnosing a subject as aFTLD-U patient or as suspected aFTLD-U patient, or as being at risk or susceptible of developing aFTLD-U, based on / using any of the above-described methods and kits.
[0176] Any of the above-described methods and kits may be applied in a method of or for stratifying a population of subjects having or suspected to have FTLD, such method comprising identification or diagnosing a subject as aFTLD-U patient or as suspected aFTLD-U patient, or as being at risk or susceptible of developing aFTLD-U, based on / using any of the above-described methods and kits.
[0177] Correct stratification of the patients with different FTLD forms (and distinguishing them from other neurodegenerative or psychiatric disorders with sometimes overlapping features) is important when testing candidate therapeutic agents. Defining clinically homogenous groups for inclusion in a clinical trial is of utmost importance in order to increase the success rate of halting disease progression or of curing the disease or disorder.
[0178] Other Definitions
[0179] The present disclosure is described with respect to particular embodiments and with reference to certain drawings but the disclosure is not limited thereto but only by the claims. Any reference signs in the claims shall not be construed as limiting the scope. The drawings described are only schematic and are nonlimiting. In the drawings, the size of some of the elements may be exaggerated and not drawn on scale for illustrative purposes. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Where an indefinite or definite article is used when referring to a singular noun e.g. "a" or "an", "the", this includes a plural of that noun unless something else is specifically stated. Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the present invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., current Protocols in Molecular Biology (Supplement 100), John Wiley & Sons, New York (2012), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art. RoRad / REPEAT / 848
[0180] EXAMPLES
[0181] EXAMPLE 1. Association analysis
[0182] The relatively homogenous clinical and pathological disease presentation of atypical FTLD with ubiquitinated inclusions (aFTLD-U) patients prompted us to hypothesize that there may be common genetic factors contributing to the risk of developing this rare neurological disease. We thus performed a first phase single variant genome-wide association study (GWAS) comparing 22 neuropathologically confirmed aFTLD-U patients and 12 very young onset (<40 years) bvFTD (behavioural variant FTD) patients selected for their possible diagnosis of aFTLD-U to "'1157 control individuals. Using an additive disease risk model, we identified a single significant locus at chrl5ql4 with four significant variants (lead variant rs549846383, p-val= 9.29xl012, OR=44.80) (Figure 1). While highly encouraging, the small number of patients and uncertainty concerning the underlying pathology in the clinically diagnosed patients provided caution and a clear need to identify more aFTLD-U patients for the study.
[0183] In an international effort, our study cohort of aFTLD-U patients was systematically expanded in order to confirm these findings. We identified and genome sequenced 36 additional neuropathologically confirmed aFTLD-U patients and obtained publicly available and in-house generated short-read whole genome sequencing data from ~2000 neurologically normal controls. Together, this led to a second phase GWAS with 59 pathologically confirmed aFTLD-U patients (including 22 aFTLD-U + 1 young bvFTD with a confirmed aFTLD-U diagnosis from the first phase) and ~3150 controls (all 1150 first phase controls included). This analysis strengthened the chrl5ql4 association and confirmed rs549846383 as the lead variant (p-val = 2.52xl0-36, OR = 70.58). This variant was one of 77 genome-wide significant variants at chrl5ql4, and its minor allele was found in 49.15% (n=29 / 59) of our pathologically confirmed cases as compared to only 1.40% (n=44 / 3152) of control individuals (Figure 1). A similar low minor allele frequency was observed in our cohort of FTLD patients with a distinct neuropathology (FTLD-TDP) included in a recent GWAS (n=9 / 957; 0.94%) (Pottier et al. 2024, medRxiv doi.org / 10.1101 / 2024.06.24.24309088). Eight additional variants reached genome-wide significance; however, these had much less significant p-values (maximum p-val 9.93xl0-9) and were each located in a distinct locus without additional variants to support the signal.
[0184] Motivated by the large number of associated variants at chrl5ql4, we next performed a conditional analysis by excluding rs549846383 minor-allele carriers. This showed an independent effect at chrl5ql4 for rsl48687709 (p-val = 1.34xl0-9, OR = 5.85) with 40.00% of the remaining patients (n=12 / 30) and 5.98% (n=178 / 2976) of the remaining control individuals carrying the minor C-allele. The carriers of rs549846383 form a subset of those with rsl48687709, and rsl48687709 thus appears to tag the ancestral haplotype on which rs549846383 occurred. We will refer to the initially discovered haplotype tagged by the minor alleles of rs549846383 and rsl48687709 as haplotype A, and refer to the haplotype RoRad / REPEAT / 848 tagged exclusively by the minor allele of rsl48687709 (with major haplotype of rs549846383) as haplotype B. A second conditional analysis, leaving out haplotype A and B carriers, did not reveal any additional variants in this locus .
[0185] Notably, other genome-wide significant loci were found as part of these conditional analyses; however, these risk alleles were only found in only a few patients.
[0186] Sanger sequencing confirmed the rs549846383 and rsl48687709 genotypes observed in our aFTLD-U population and allowed screening of an additional control cohort from Mayo Clinic (n=1027), which confirmed the low frequency of haplotypes A and B in a neurologically healthy population. Sanger sequencing of an additional 28 aFTLD-U patients identified ten more carriers of the chrl5ql4 risk haplotypes (nine haplotype A and one haplotype B carrier). Together, in our combined cohort of 87 aFTLD-U patients with DNA available identified to date, 38 patients (43.7%) carried haplotype A, 13 patients (14.9%) carried haplotype B, while 36 patients (41.4%) carried neither of the chrl5ql4 risk haplotypes.
[0187] A schematic overview of the chrl5ql4 locus is given in Figure 2A. Information about the rs549846383 and rsl48687709 single nucleotide variants is depicted in Figure 2B and Figure 2C, respectively. The minor rs549846383 allele is lacking a "T" compared to the major allele. The minor rsl48687709 allele is having a "C" compared to a "T" in the major allele.
[0188] Exemplary primer sequences that can be used for genotyping the haplotype variants are provided in Table 1.
[0189] Table 1.
[0190] EXAMPLE 2. Investigation of the locus using short-read sequencing
[0191] The chrl5ql3-14 region is notoriously repetitive and characterized by tandem segmental duplications of GOLGA genes (Antonacci et al. 2014, Nat Genet 46:1293-1302; Pujana et al. 2002, Eur J Hum Genet 10:26-35), with rs549846383 upstream of GOLGA8B and rsl48687709 intronic in GOLGA8A (Figure 2). GOLGA8A and GOLGA8B have a pairwise sequence similarity of 98.9%, complicating most analyses with short-read genome sequencing data due to ambiguous alignments. However, the region between GOLGA8A and GOLGA8B is unique in the reference genome, which allowed us to determine RoRad / REPEAT / 848 the copy number of the intervening sequence in all aFTLD-U patients and controls. This led to the identification of deletion and duplication alleles, with 80.9% of the cohort having a normal copy number, 17.2% a heterozygous deletion, 1.0% a heterozygous duplication, and 0.9% a homozygous deletion of sequence between GOLGA8A and GOLGA8B. We assume the copy number variations (CNVs) are recurrent and mediated by the direct orientation of the GOLGA8A and GOLGA8B genes, forming a hybrid of both genes upon deletion. Carriers of a duplication on one haplotype with a deletion on the other are expected in a sufficiently large cohort as ours. However, the current method cannot distinguish them from a normal copy number. Importantly, this deletion was not associated with the disease (Fisher's Exact test p=1.0). Notably, the rsl48687709 position is included in the deletion, and, in line with this observation, we observed individuals carrying a deletion as well as the rs549846383 major risk allele but without the minor allele of rsl48687709, suggesting that those individuals have a deletion on their associated haplotype and are, therefore, hemizygous for the major allele of rsl48687709.
[0192] EXAMPLE 3. Long-read data analysis
[0193] The highly repetitive chrl5ql3-14 region and blind spots due to the low mapability of short reads prompted us to use long-read sequencing for the remainder of this study. As part of ongoing projects, we had already generated genome-wide long-read sequencing data for 283 subjects, mostly FTLD-TDP patients and control individuals. This cohort was found to include two carriers of haplotype A (one FTLD- TDP patient and one neuropathological normal control) and 14 haplotype B carriers by chance). To enrich for disease haplotypes and focus on aFTLD-U, we performed additional long-read sequencing in 53 aFTLD-U patients (22 haplotype A, 9 haplotype B, and 22 carrying neither haplotypes A or B) and in five non-aFTLD-U subjects (two FTLD-TDP patients and three neuropathological normal controls) selected to carry haplotype A based on the short-read genome sequencing data.
[0194] As a reminder, haplotype A comprises the minor alleles of rs549846383 and rsl48687709 whereas haplotype B comprises the minor allele of rsl48687709.
[0195] To identify a potential functional variant tagged by rs549846383, we performed haplotype-phasing of the long reads and, upon visual inspection (Thorvaldsdottir et al. 2013, Brief Bioinform 14:178-92), identified a tandem repeat expansion in cis at chrl5:34419425-34419451, part of an AluY interspersed repeat element. When we genotyped this tandem repeat in the in-house long-read cohort (N=341) with STRdust, we observed variation in repeat length, with longer alleles predominantly observed in aFTLD-U patients (Figure 3). We additionally performed a GWAS for tandem repeat lengths as a continuous variable for the disease, as determined in the long-read sequencing cohort with inquiSTR. This analysis compared 52 aFTLD-U patients with 283 non-aFTLD-U subjects, excluding the five haplotype-carrying non-aFTLD-U subjects specifically selected for long-read sequencing as well as one Asian aFTLD-U patient. 318299 loci passed call-rate filtering, resulting in two genome-wide significant loci in the RoRad / REPEAT / 848 chrl5ql4 locus and showing a strong association for the length of the GOLGA8A STR at chrl5:34419425- 34419451 (GRCh38) with aFTLD-U, with a p-value of 1.98.1013and odds ratio of 17.1. The other genomewide significant locus was an intergenic repeat length polymorphism at chrl5:34480576-34480608, which is, on average, 6bp longer on the associated haplotype but showed no actual expanded alleles. Information about the wild-type repeats at chrl5:34419425-34419451 and chrl5:34480576-34480608 (human genome build GRCh38) is provided in Figures 4A and 4B, respectively. The top-strand sequences are provided by SEQ ID NO:1 and SEQ ID NO:2, respectively (obviously, the respective bottom strand sequence are the (reverse) complements of the top-strand sequences:
[0196] - tandem repeat at chrl5:34419425-34419451 (GRCh38)
[0197] TTTCTTTCTTTCTTTCTTTCTTTCTTT ( SEQ ID NO : 1 )
[0198] - chrl5:34480576-34480608 (GRCh38)
[0199] GTGTGTACGTATGTGTGTGTGTGTGTGTGTGTG ( SEQ ID NO : 2 ) .
[0200] EXAMPLE 4. Tandem repeat motif composition
[0201] The associated repeat at chrl5:34419425-34419451 is annotated as a short tandem repeat (STR) with a TTTC motif by Tandem Repeat Finder (Benson 1999, Nucleic Acids Res 27:573-80) with 6.75 copies in the reference genome (GRCh38)(see Example 3; Figure 4A; SEQ ID NO:1). Surprisingly, analysis of the sequence composition of expanded alleles across the cohort additionally identified a CT dimer, a CCCT tetramer, CCTT tetramer, CCCTCT hexamer, and a CCCCT pentamer repeat motif, with the CT dimer motif exclusively seen for individuals carrying haplotypes A or B. In contrast, expansions of the CCTT tetramer motif were never observed on the associated haplotypes. As the CCCCT pentamer represented only a single patient, the repeat composition in the cohort was quantified and visualized for 12-mer units to simultaneously represent dimer, tetramer, and hexamer motifs (Figure 5). This illustrated that CT dimers are exclusively observed on haplotypes A and B and especially at high frequencies in aFTLD-U patients. In contrast, some haplotype A and B carriers have different compositions. Specifically, repeats composed primarily of CCCTCT hexamer motifs were found for three non-aFTLD-U subjects with haplotypes A and B. At the same time, two aFTLD-U patients carrying haplotype B were found with a mixed composition of CT and CCCTCT motifs, with most of their repeat sequence comprising CT dimers. Using a 12-mer heatmap we observed that CT-dimers are exclusively found on haplotypes A and B and especially at high frequencies in aFTLD-U patients. CCTT expansions are only observed in individuals without haplotype A or B, while CCCTCT hexamer expansions are found on haplotypes A and B, but more so in non-aFTLD-U subjects.
[0202] We additionally identified an individual with haplotype B carrying a CnT-rich sequence for which no clear repeat motif could be described. The repeat consensus sequence has up to 62 consecutive Cs (n=62), RoRad / REPEAT / 848 flanked by CT dimer motifs at the expansion ends. Interestingly, this patient is the only one for which there is a positive family history of aFTLD-U, but no DNA is available from his affected mother.
[0203] Multiple individuals also show a mixed composition, where, e.g., the hexamer sequence is combined with long CT stretches in the two aFTLD-U patients carrying haplotype B described above, and the pentamer motif seen in an aFTLD-U patient is additionally flanked with long CT fragments. We also identified a non-aFTLD-U subject with a 12-mer repeat motif at the 5' end, showing some motif interruptions, and a CT composition at the 3' half of the repeat. All non-aFTLD-U subjects carrying haplotype B showed a CCCT stretch at the 5' end of the repeat, followed by a short CT motif stretch.
[0204] We also observed several flanking motifs, which are variable in length but short (<20 units), including tetramers (CTTT and CCCT) and a pentamer (CCTTT). Most expanded alleles contained variable lengths of the flanking CCTTT pentamer motif at the 5' end, more specifically the reference CTTT units are followed by 2-6 copies of CCTTT before the sequence transitions into the expanded CT-dimer stretch. All non-aFTLD-U subjects carrying haplotype B showed a short CCCT stretch flanking the 5' end of the repeat (10-20 units), followed by a short CT-stretch.
[0205] EXAMPLE 5. Tandem repeat motif and length in haplotype A and B carriers
[0206] Focusing in on the associated haplotypes led us to hypothesize that long expansions predominantly composed of CT dimer motifs appear to drive the aFTLD-U risk. Notably, of the seven non-aFTLD-U subjects with haplotype A in this discovery cohort, one had a CCCTCT hexamer repeat composition, one had a 12-mer repeat, and the remaining five individuals had CT-rich repeat lengths ranging from only 149bp up to 1178bp (median expansion length 433bp; 71%). In stark contrast, all 22 aFTLD-U patients with haplotype A had long expansions ranging from 489bp to 2133bp (median 760bp; 100%). For the 14 non-aFTLD-U subjects with haplotype B, we observed two carriers of a CCCTCT hexamer repeat composition (14.3%) and twelve carriers of a repeat predominantly composed of CT dimers but of relatively short length ranging from 187bp to 235bp (median expansion length 214bp; 85.7%). In contrast, except for one aFTLD-U patient with haplotype B with the CnT-repeat, and one aFTLD-U patient with a short CT expansion reminiscent of the haplotype B seen in non-aFTLD-U subjects, the remaining 29 aFTLD-U patients carrying haplotypes A or B in our cohort were all found to have long expansions predominantly composed of CT dimer motifs (77.8%), with lengths ranging from 489 to 2133bp (median expansion length 760 bp), or, upon re-analysis, with lengths ranging from 484bp to 1245bp (median 834bp). In haplotype B carriers, we observed a clear separation based on repeat length between aFTLD- U patients and non-aFTLD-U subjects, apart from the one patient who may have inherited haplotype B by chance and likely has another (genetic) etiology. This separation by length is not the case for haplotype A, where, in line with our expectations of a genetic risk factor and the sporadic nature of the RoRad / REPEAT / 848 disease, a subset of the non-aFTLD-U subjects carry long CT-rich repeat expansion comparable to aFTLD- U patients.
[0207] Next, we studied the chrl5ql4 risk locus in several non-aFTLD-U populations to better characterize the frequency of the chrl5ql4 risk haplotypes and the tandem repeat composition in haplotype A carriers. First, short-read genome sequencing data from the Mayo Clinic Brain Bank derived from patients with tauopathies and a-synucleinopathies (LBD, PSP, and MSA; N=1747) was queried for the haplotype A- tagging variant rs549846383 showing 22 heterozygous carriers (0.013%). Six of those were selected for long-read sequencing. The same variant was also imputed in control individuals from the European Alzheimer's Disease (AD) DNA BioBank (EADB) Belgian cohort (N=1606) (the European Alzheimer's Disease Biobank (EADB) consists of >20.000 clinically diagnosed AD patients and >21.000 unrelated age- matched controls. This dataset has been put together thanks to a long-term collaboration between 15 different European countries), identifying 15 heterozygous carriers (0.9%). Two of those were selected for long-read sequencing. As such, the repeat length and composition could be evaluated in long-read sequencing data for an additional eight non-aFTLD-U subjects. This analysis showed that two of these individuals did not have an expansion, two had only a short expansion (137bp and 159bp), one individual had a hexamer motif expansion, and only three individuals had a long CT-rich expansion. Of note, immunohistochemical analyses were also performed on the six subjects from the Mayo Clinic brain bank, and none was found to have FUS or TAF15 pathology.
[0208] In summary, our combined long-read sequencing efforts generated data for 15 non-aFTLD-U subjects carrying haplotype A. Long CT-rich repeat expansions were observed for 8 of these 15 individuals (53%), while all sequenced 22 aFTLD-U patients with haplotype A showed long CT expansion. The presence of an expanded CT-repeat thus appears to be a better predictor of the aFTLD-U phenotype than the tagging variant of haplotype A alone. In support of this finding, repeat-primed PCR assays of the remaining 16 non-aFTLD-U subjects from the Mayo Clinic brain bank carrying haplotype A, for which no long-read sequencing data was generated, which confirmed the absence of a long CT-rich expansion in 8 / 16 or 50% of subjects. Moreover, since we cannot accurately determine the repeat length in the lower size range using this assay, it remains possible that a greater subset of haplotype A carriers does not carry the risk allele.
[0209] Next, using Sanger sequencing for the haplotype-tagging variants, we screened 17 NIFID and 11 BIBD patients, identifying one NIFID patient carrying haplotype A and two carrying haplotype B. None of the BIBD patients had one of the haplotype-tagging variants. Using repeat-primed PCR, the NIFID patient with haplotype A was found to be a carrier of a long CT-rich repeat. However, misdiagnosis cannot be excluded due to the high pathological similarity with aFTLD-U. No DNA of sufficient quality was available for long-read sequencing for haplotype B carriers. RoRad / REPEAT / 848
[0210] We additionally analyzed 1181 individuals from the 1000 Genomes Project for whom short- and long- read sequencing data was available to determine the frequency of the associated haplotypes, the GOLGA8A-B copy number, and the length and motif composition of the GOLGA8A tandem repeat (Schloissnig et al. 2024, bioRxiv 2024.04.18.590093; Noyvert et al. 2023; Gustafson et al. 2024, medRxiv 2024.03.05.24303792; Byrska-Bishop et al. 2021). We could again identify homozygous deletions (0.8%), heterozygous deletions (19.6%), normal copy number (72.4%), heterozygous duplications (6.7%), and homozygous duplications (0.5%). We identified four haplotype A carriers and ten haplotype B carriers.
[0211] We summarized the repeat genotypes of all 1530 individuals with the length of the repeat consensus sequence plotted against its CT-fraction.
[0212] The longest allele is per individual, with a minimal expansion length of 20bp were taken into account. This analysis showed a peak of expansions at 50% CT, and again confirmed that carriers of haplotype A have longer CT-rich expansions, which include a small group of non-aFTLD-U subjects that cannot be distinguished from aFTLD-U patient alleles. Out of the four carriers of haplotype A in the 1000 Genomes Project, one has a deletion of the repeat locus, two have modest repeats including hexamer stretches, and one has a long CT-rich repeat comparable to the aFTLD-U patients. There was a clear separation between aFTLD-U patients and non-aFTLD-U subjects for haplotype B. All ten carriers of haplotype B from the 1000 Genomes Project have a repeat shorter than the aFTLD-U patients, with a mean expansion length of 211bp and a maximum length of 261. Notable outliers were the aFTLD-U patients with rare motif compositions described above - i.e., the CnT-rich haplotype B carrier and the haplotype A carrier with a pentamer repeat.
[0213] Based on the current data from in-house and public cohorts, we propose that a repeat expansion of >~400 bp and >80% CT content is one possible predictor for aFTLD-U patients. An alternative predictor is the presence of at least 300 bp in or of CT dimers wherein the CT dimers do not need to be present as a contiguous stretch (i.e. are contiguously or non-contiguously present in the tandem repeat expansion). In particular, other repeat motifs (CCCTCT, CCCCT, CCCT, CTTT and CCTT; if present) are removed from the expanded repeat prior to counting the number of CT dimer repeats and determining the length of the CT dimer repeat sequence.
[0214] Alternatively, we propose that a repeat expansion of >450 bp and >80% CT content predicts aFTLD-U patients among haplotype carriers, with a precision of 0.80 (95% Cl 0.64-0.91) and recall of 0.90 (95% Cl 0.75-0.97). With an alternative classification, using a threshold of 190 CT-dimer motifs in haplotype carriers (after subtracting other repeat motifs), we obtain a precision of 0.78 (95% Cl: 0.62-0.89) and recall of 0.94 (95% CI:0.79-1.00) for prediction of aFTLD-U. RoRad / REPEAT / 848
[0215] We additionally calculated the association with aFTLD-U for each of the two repeat-based classifiers and compared this with the association with aFTLD-U of the tagging variants, using Fisher's exact tests in the 1,715 individuals in the long-read cohort. The p-value is 7.29x10-25 based on rs549846383 (tagging haplotype A), 2.01x10-29 based on rsl48687709 (tagging haplotype A and B), 5.77x10-40 for the classification using the double cut-off of >450bp expansion with >80% CT content, and 4.86x10-41 for the classification based on expansion alleles with >190 CT-dimer motifs.
[0216] Additional screening of other and / or future cohorts may be useful to refine those cutoffs further.
[0217] Interestingly, and distinct from other repeat expansion disorders, the GOLGA8A repeat is characterized by a degenerate motif, showing dimer, tetramer, pentamer and hexamer motifs composed of C and T nucleotides, of which some were found to expand, and others were flanking the expansion and remained stable in size. We even observed a nearly pure C-repeat in the only known inherited case of aFTLD-U, of which the relevance to disease may become clear in future functional studies. We further observed motif length switches and hybrid compositions within the same repeat allele for which the pathogenic role requires further observations in additional patients and / or controls. Variation in repeat motif composition in disease-associated repeats has been described before, typically involving pentamer repeat motifs, where only specific motifs are pathogenic if expanded (Rajan-Babu et al. 2024, Nat Rev Genet 25:476-499). However, the GOLGA8A repeat is unparalleled in the variation in repeat length, motif length, and motif sequence.
[0218] For a small set of patients, we additionally sequenced DNA extracted from other brain areas, such as the cerebellum, caudate and occipital cortex, and from lymphoblast cell line (LCL) cultures. This again identified variation in repeat length, and the identification of expanded alleles of the repeat in LCLs from patients demonstrates that the lack of inheritance of the disease is not just a consequence of somatic expansion exclusively in the brain.
[0219] EXAMPLE 6. Repeat-primed PCR of repeat expansion
[0220] The genomic region on chrl5ql4 containing the expanded alleles was amplified using a panel of three- primer repeat-primed PCR assays, each with one FAM-labeled primer flanking the repeat, one sequencespecific primer targeting each of the repeat motifs and one booster primer recognizing the tail of the sequence-specific primer to amplify the signal. A total of six primer sets (see Table 2) were designed based on observed repeat sequences, in particular to determine the presence of CT motifs on the left and right ends of the repeat, CCCTCT motifs on the left and right ends, CCCT motifs on the left, and CCCCT motifs on the left end of the repeat. Fragment lengths were determined with capillary electrophoresis on an ABI3730XL and visualized using in-house developed software. RoRad / REPEAT / 848
[0221] In a first stage, the repeat-primed PCR method was used to confirm / validate repeat sequence composition previously identified by long-read DNA sequencing. Representative images are provided in Figure 6 for 3 aFTLD-U patients and one non-aFTLD-U subject carrying distinct repeat compositions. aFTLD-U 1 showed characteristic stutter pattern indicative of a repeat-expansion using the CT-right assay but failed to show amplification of the CT-left assay due to the presence of a different motif, e.g. a CCCCT- expansion on this side. aFTLD-U 2 showed characteristic stutter patterns for the CT-left and CT-right assays and was negative for all other assays. aFTLD-U 3 showed the characteristic stutter-pattern using the CT-left assay, and a stutter pattern using the CCCTCT-right assay. Note that there was also amplification of a small stable fragment using the CT-right assay most likely due to the presence of small CT-dimer stretches within the CCCTCT repeat. Finally, non-aFTLD-U subject 1 carried the typical haplotype B seen in controls with a short motif of CCCT-tetramers and CT-dimers. Due to the small total size of the repeat both the CT-left and CT-right assays showed amplification; however, the pattern suggested a smaller (and largely stable size) of the CT-expansion. Positive amplification was also seen for the CCCT-left assay whereas all other primer sets were negative.
[0222] Subsequently, the repeat-primed PCR method was applied to screen presence of aFTLD-U associated repeat expansion in subjects for which no long-read DNA sequencing data are yet available. Genotyping of rs549846383 and rsl48687709 will allow the identification of Chrl5ql4 haplotype A and haplotype B carriers. In the populations that we studied, which were all of Caucasian origin, CT-rich repeat expansions associated with an increased risk for aFTLD-U were only identified on these haplotypes. However, not all haplotypes A and B carry expanded alleles, and in some instances the region housing the repeat can even be deleted. To determine the presence of aFTLD-U risk associated CT-rich expansions we used the right- CT and left-CT repeat-primed PCR assays, the results of which are depicted in Figure 7. In a population of 22 non-aFTLD-U subjects carrying haplotype A this allowed the identification 9 individuals in which there was no evidence for a repeat expansion. In 3 of the patients the CT-expansions indicated the amplification of a smaller repeat (samples 1, 21 and 22). Two of those 3 were also found to be positive for repeat-primed PCR assays targeting the expanded CCCTCT-hexamers (samples 1 and 22), while the other one (sample 21) was considered to have pure but smaller CT-repeats (confirmed by ONT sequencing as being <100bp of CT-dimers). In conclusion the combined use of repeat-primed PCR assays targeting the two most commonly observed repeat motifs e.g. expanded CT-dimers (carrying aFTLD-U risk) or expanded CCCTCT-hexamers (carrying no FTLD-U risk), led to the exclusion of disease-associated repeat expansions in 12 out of 22 non-aFTLD-U subjects. As we are unable to accurately determine the length or the repeat expansion using this assay alone, further analysis of the remaining 10 cases will require other methods. RoRad / REPEAT / 848
[0223] The chrl5:34419425-34419451 region in DNA isolated from blood obtained from an aFTLD-U patient was shown to comprise an CT dimer repeat expansion, as determined by nanopore sequencing. Electropherograms of repeat-primed PCR (RP-PCR) products starting from the DNA of the same blood sample confirmed the presence of the CT dimer repeat expansion. The PCR primers were adjusted to detect the presence of CT-repeats at the left or right end of the region of repeat. Results are depicted in Figure 8, and again (see Example 5) confirm that the presence of an expanded repeat in this chromosomal region is not limited to brain cells.
[0224] Table 2. Primer panel for repeat primed PCR assays. Of note one of the primers may e.g. not bind to a CT- or CCCTCT-repeat but to the reverse complement GA- or GGGAGA-sequence, respectively. RoRad / REPEAT / 848
[0225] EXAMPLE 7. Detection of tandem repeat motif in cDNA
[0226] We performed long-read sequencing (Oxford Nanopore PromethlON) from cDNA for which target enrichment was performed using oligonucleotide capture probes designed against the exons of the GOLGA8A gene, putative exons, and other overlapping transcripts. The RNA was extracted from human brain samples. As such, we obtained ultra-deep coverage of the transcripts in the locus of interest (directly surrounding the repeat). After alignment of the reads to a modified reference genome (masking the GOLGA8B gene with 'N's to get unambiguous alignments) and analysis using customized Python scripts, cDNA molecules containing the CT repeat were identified in a small fraction of the sequenced fragments. A similar analysis is performed on biofluid samples in order to detect the presence of the tandem repeat motif in biofluid RNA or cDNA derived from it.
[0227] EXAMPLE 8 . Methods
[0228] 8.1. aFTLD-U International Consortium
[0229] We established an international consortium to identify and bring together a sufficiently large patient population to assess this rare disorder systematically. aFTLD-U patients were identified through inquiries at brain banks focused on neurodegenerative disease research and by contacting authors of relevant publications, including FTLD-FET patients. A limited number of patients with BIBD and NIFID were identified and collected during these efforts. All patients provided consent to participate in research studies in accordance with the Declaration of Helsinki and local ethics review board standards. Our cohort continuously expands, with 108 aFTLD-U patients currently identified from 24 sites worldwide. Frozen brain tissue from the cerebellum and / or frontal cortex was obtained from 84 patients, while DNA extracted from blood was available from four additional patients. The remaining 20 patients only have fixed tissue available.
[0230] An experienced neuropathologist from one of the collaborating sites analyzed paraffin-embedded tissue sections for each patient to confirm the neuropathological diagnosis. aFTLD-U was diagnosed based on the presence of tau- and TDP-43 negative, FUS-positive neuronal cytoplasmic inclusions (NCI) and FUS- positive neuronal intranuclear inclusions (Nil). FUS immunostaining was performed at most sites using primary antibodies 11570-1-AP (Proteintech Group) and / or HPA008784 (Sigma Life Sciences), and occasionally A300-302A (Bethyl Laboratories) or NB100-565 (Novus). None of the patients showed basophilic inclusions (characteristic of BIBD) or other cellular inclusions, such as hyaline conglomerate inclusions (typical of NIFID), on H&E staining. The diagnosis of aFTLD-U was further supported by the presence of only limited FUS pathology in subcortical regions and limited variability in the morphology of NCIs. In those patients where a differential diagnosis of NIFID was considered, neurofilament or alpha- internexin (AIN) immunohistochemistry was performed to exclude a pathological diagnosis of NIFID. In a minority of patients, TAF15 immunohistochemistry was also performed, confirming TAF15 RoRad / REPEAT / 848 immunoreactivity of the inclusions. Basic demographic and clinical information were collected for all patients, including details on self-reported ethnicity, clinical diagnosis, family history of neurological diseases, age at disease onset, age at death, and sex.
[0231] As such, we identified 108 aFTLD-U patients, 34 (31.5%) female and 74 (68.5%) male. All patients were self-reported Caucasian except for one Asian patient. The mean age at onset in the full cohort was 44.3 years (median 43, standard deviation 10.4 years, range 21-73 years) with the mean age at death of 51.0 (median 51, standard deviation 10.0 years, range 30-77 years) and a mean disease duration of 6.8 years (median 6, standard deviation 3.4, range 2-19 years).
[0232] 8.2. Short-read genome sequencing and association analysis
[0233] Patient samples (22 neuropathologically confirmed aFTLD-U patients, 12 clinical patients with young onset bvFTD) collected between 2010 and date2 were included in the original phasel GWAS compared to "'1100 control individuals, with samples (59 neuropathologically confirmed aFTLD-U) collected after date2 with additional controls from the ADSP project (NG00067) forms the phase2 GWAS. Short-read genome sequencing for patients and control individuals (2x 150bp) was generated on the Illumina HiSeq X and Novaseq as previously described (Pottier et al. 2019) (ref Pottier 2024). Sequencing data was aligned using bwa mem (Li 2013), followed by joint variant calling with GATK HaplotypeCaller (Poplin et al. 2018). Variants were filtered on a minor allele frequency of <0.01, a call rate of 90% in patients and controls, and hardy-weinberg equilibrium in controls (variant removed if p<10-6). The second phase also included a batch filter to remove spurious associated variants due to differences in the library preparation method, implemented as a Fisher exact test comparing patients' calls in phase 1 against phase 2 (variant removed if p<0.05). The association test was performed in R.
[0234] We additionally performed a conditional GWAS analysis in a subset of the phase 2 cohort after removing carriers of the rs549846383 top hit, applying the same filters described above.
[0235] 8.3. Sanger sequencing validations
[0236] The results have to be interpreted as a tetrapioid region for rsl48687709 as no unique primers could be designed, and the paralogous sequence in GOLGA8B will also be amplified.
[0237] 8.4. Long-read genome sequencing
[0238] Long-read genome sequencing on the PromethlON P24 (Oxford Nanopore Technologies) was performed for 53 aFTLD-U patients and 5 non-aFTLD-U subjects and compared to an ongoing genome sequencing initiative of 283 additional non-aFTLD-U subjects, mostly FTLD-TDP patients neurologically normal controls. DNA was extracted from the frontal cortex using the Nanobind tissue kit (PacBio), followed by quality control using the Dropsense (Trinean), Qubit (Thermo Fisher Scientific), and Fragment Analyzer (Agilent) to asses purity, concentration, and fragment length. DNA was sheared using the Megaruptor (Diagenode) on speed Y, followed by removing short fragments with the Short Read Eliminator (PacBio) when considered appropriate. The library prep was generated using the SQK-LSK110 kit (ONT) according RoRad / REPEAT / 848 to the manufacturer's instructions, except for longer incubation times for enzymatic steps, before sequencing on an R9.4.1 flow cell for 72 hours. Four individuals were additionally sequenced on the PacBio Sequel II with libraries prepared according to the manufacturer's instructions.
[0239] 8.5. Long-read data analysis
[0240] The sequencing data was base called with guppy v6.7.3 using the HAc model (ONT), including cytosine methylation and hydroxymethylation inference. The data was processed using several snakemake workflows (Koster and Rahmann 2012). Reads were aligned to the GRCh38 reference genome (GCA_000001405.15_GRCh38_no_alt_analysis_set) with minimap2 (Li 2021), followed by sorting reads by coordinate and conversion to CRAM format with samtools (Li et al. 2009). The data quality was assessed with cramino, as was the concordance with the expected sex based on the normalized read depth of the sex chromosomes (De Coster and Rademakers 2022). Reads were phased with longshot (Edge and Bansal 2019). SVs were called using Sniffles2 (Smolka et al. 2024).
[0241] Tandem repeats of interest were genotyped with STRdust (De Coster et al. 2024), and the length of all human tandem repeats (English et al. 2023) was determined using inquiSTR, followed by association analysis for repeat length (using the largest length per individual) using str_assoc with a minimal call rate of 80% and Bonferroni correction for multiple testing. The repeat composition was assessed using aSTRonaut (De Coster et al. 2024) to visualize the sequence of specific repeat motifs per allele.
[0242] 8.6. Copy number variant analysis
[0243] The copy number of the region between GOLGA8A and GOLGA8B (chrl5:34438297-34524132) was quantified using the coverage obtained from mosdepth (Pedersen and Quinlan 2018), normalized to a copy-number neutral interval (chrl5:54033377-56279876) for both short- and long-read genome sequencing data. Visualization was performed in Python using plotly (Plotly Technologies Inc. 2015), and statistical analysis was performed for carriers of the deletion allele using a Fisher exact test as implemented in scipy (Jones et al. 2001).
[0244] 8.7. Repeat-primed PCR
[0245] Expanded repeats were detected using a combination of three-primer repeat-primed PCR assays with a fluorescent FAM label on a flanking primer. Primer sets were designed to determine the presence of CT motifs on the left and right ends of the repeat, CCCTCT motifs on the left and right ends, CCCT motifs on the left, and CCCCT motifs on the left. These assays were complemented by a spanning amplicon, using the LEFT and RIGHT flanking primers (of which only one was labelled). Fragment lengths were determined with capillary electrophoresis on an ABI3730XL.
[0246] 8.8. Tandem repeat analysis
[0247] Tandem repeats of interest were genotyped with STRdust (v0.11.7)( De Coster et al. 2024, Genome Res 34:2074-2080), either from local files as sequenced in-house or over FTP for the participants from the 1000 Genomes Project resequenced with ONT (Noyvert et al. 2023, medRxiv 2023.12.20.23300308; RoRad / REPEAT / 848
[0248] Schloissnig et al. 2024, bioRxiv 2024.04.18.590093; Gustafson et al. 2024, medRxiv 2024.03.05.24303792). STRdust was used in standard (phased) mode to establish that the repeat expansion is present on the associated haplotype. As read phasing by LongShot was found to be unreliable for this locus, resulting in the omission of a large proportion of the reads from the phased results due to ambiguous alignment and uncertain haplotype assignment, the unphased mode of STRdust was used to obtain the genotypes used in this manuscript, determining alleles by hierarchical clustering the extracted repeat sequence for each read. STRdust generates a consensus allele by Partial Overlap Alignment as implemented in rust-bio (Koster 2016, Bioinformatics 32, 444-446), ignoring length outliers. The observed length variation suggests that the consensus sequence can change substantially due to random sampling of sequenced fragments from the library, especially at low sequencing depth.
[0249] The length of all human tandem repeats (English et al. 2023, bioRxiv 2023.10.29.564632) was determined using inquiSTR (v0.13.0) (github.com / wdecoster / inquiSTR). We developed STR_regression.R (vl.6) (github.com / wdecoster / inquiSTR / scripts / STR_regression. R) for running association testing of tandem repeat lengths, which can fit generalized linear models using the output of inquiSTR repeat lengths and phenotypic information of multiple samples. STR_regression.R can run both logistic and linear regressions based on binary and continuous phenotypes (and optionally with covariates), and it outputs detailed statistics of repeat length associations. Moreover, it has multiple functionalities, including different repeat length processing modes (either considering mean, minimum, or maximum repeat length for a given tandem repeat), various run options (genome-wide, per chromosome, and a region of interest based on a chromosomal interval or a list of regions of interest based on a BED file), and it can also take into account provided cut-offs to define expanded alleles of tandem repeats. For this analysis, we compared 52 aFTLD-U patients with 283 non-aFTLD-U subjects, excluding one Asian aFTLD- U patient and the five haplotype-A-carrying non-aFTLD-U subjects specifically selected for long-read sequencing. We used the longest allele per individual for all human tandem repeats, with a binary phenotype (aFTLD-U or not), a minimal call rate of 80%, and Bonferroni correction for multiple testing.
[0250] The repeat composition was assessed using a k-mer heatmap, in which all 12-mers were quantified. As the CCCCT pentamer expansion was only found in a single patient, the repeat composition in the cohort was quantified and visualized using the least common multiple of 12-mer units to simultaneously represent dimer, tetramer, and hexamer motifs, i.e. the most commonly observed motifs. VCF files were parsed with cyvcf2 (v0.30.16) (Pedersen & Quinlan 2017, Bioinformatics 33:1867-1869), and each 12- mer in the repeat consensus sequences was counted. After counting, all motifs were rotated and represented by the lexicographical first, then collected in a pandas data frame (McKinney 2011, Python High Perform Sci Comput 1-9) before filtering motifs rarely observed, except if highly prevalent in one individual. Visualization was done using Plotly (v5.14.1) (Plotly Technologies Inc. Collaborative data RoRad / REPEAT / 848 science. (2015)). We also used aSTRonaut (vl.O) (De Coster et al. 2024, Genome Res 34:2074-2080) to visualize the sequence of the observed repeat motifs per allele (CT, CCTT, CTTT, CCCT, CCCTCT, CCCCT, CCTTT, and CCCCCC), replacing motifs by colored dots of the same length, substituting longer motifs first. We calculated the CT dimer count for each repeat allele by removing all occurrences of other repeat motifs in which CT is a substring (CCCTCT, CCCCT, CCTT, CCCT, and CTTT) from the consensus allele and counting the remaining CT units. Precision and recall of the proposed cut-offs (>190 CT dimers or >450bp repeat and >80% CT) was calculated using scikit-learn (vl.6.1) (Pedregosa et al. 2011, J Mach Learn Res 12:2825-2830) with confidence intervals calculated using bootstrapping as implemented in scipy (vl.15.1) (Jones et al. 2001, SciPy: Open source scientific tools for Python).
Claims
RoRad / REPEAT / 848CLAIMS1. A method of analyzing the genome of a subject, the method comprising detecting in the DNA of a genomic DNA sample obtained from the subject one or more of:- the haplotype of the single nucleotide variant rs549846383;- the haplotype of the single nucleotide variant rsl48687709;- the presence of a tandem repeat expansion relative to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1) or equivalent in an alternative human genome build version; and- the presence of a repeat length polymorphism relative to the wild-type genomic DNA region of chrl5:34480576-34480608 according to the human genome build GRCh38 (SEQ ID NO:2) or equivalent in an alternative human genome build version.
2. A method of analyzing the transcriptome of a subject, the method comprising detecting in a transcriptomic sample obtained from the subject the presence of a tandem repeat expansion relative to the wild-type transcriptomic sample corresponding to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1) or equivalent in an alternative human genome build version.
3. The method according to claim 1, further comprising diagnosing the subject to have atypical frontotemporal lobar dementia with ubiquitinated inclusions (aFTLD-U), to be a carrier of aFTLD-U risk alleles, or to be at risk or susceptible of developing aFTLD-U when the genomic DNA of the subject comprises one or more of:- the minor allele of the single nucleotide variant rsl48687709;- the minor allele of the single nucleotide variant rs549846383;- a tandem repeat expansion relative to the wild-type genomic DNA region of chrl5:34419425- 34419451 according to the human genome build GRCh38 or equivalent in an alternative human genome build version; and- a repeat length polymorphism relative to the wild-type genomic DNA region of chrl5:34480576- 34480608 according to the human genome build GRCh38 or equivalent in an alternative human genome build version.
4. The method according to claim 1, further comprising diagnosing the subject to have aFTLD-U or to be at risk or susceptible of developing aFTLD-U when the genomic DNA of the subject comprises:RoRad / REPEAT / 848- the minor allele of the single nucleotide variant rs549846383, or the minor allele of the single nucleotide variant rs549846383 and the minor allele of the single nucleotide variant rsl48687709; and- a tandem repeat expansion relative to the wild-type genomic DNA region of chrl5:34419425- 34419451 according to the human genome build GRCh38 or equivalent in an alternative human genome build version, wherein the tandem repeat expansion comprises a long CT dimer repeat.
5. The method according to any one of claims 1 or 3, wherein the tandem repeat expansion relative to the wild-type genomic DNA region of chrl5:34419425-34419451 comprises:- a CT dimer-rich expansion;- a CnT-rich sequence, optionally combined with or flanked by CT dimers;- a CCCCT pentamer repeat, optionally combined with or flanked by CT dimers; or -a CCCTCT hexamer repeat combined with or flanked by CT dimers.
6. The method according to claim 5, wherein the tandem repeat expansion comprises a CT dimer-rich expansion.
7. The method according to any one of claims 4 to 6 wherein the tandem repeat expansion comprises at least 150 CT dimers wherein the CT dimers are contiguously or non-contiguously present in the tandem repeat expansion and are counted after deletion of CCCTCT, CCCCT, CCCT, CTTT and CCTT motifs from the tandem repeat expansion.
8. The method according to any one of claims 4 to 6 wherein the total length of the expanded repeat is at least 400 bp and wherein the expanded repeat comprises a CT dimer content of at least 80%.
9. The method according to any one of the foregoing claims, wherein the detection of a single nucleotide variant haplotype, of the presence of a tandem repeat expansion and / or of the presence of the repeat length polymorphism comprises a DNA genotyping step.
10. The method according to claim 9 wherein the DNA genotyping step comprises a PCR amplification step and / or a DNA sequencing step, or comprises a PCR amplification step and capillary electrophoresis.
11. The method according to claim 9 wherein the DNA genotyping step comprises a repeat-primer based PCR amplification step.RoRad / REPEAT / 84812. Use of a diagnostic kit in a method according to any one of the foregoing claims, wherein the detecting comprises the use of an oligonucleotide in a DNA genotyping method.
13. A diagnostic kit comprising an oligonucleotide, wherein the oligonucleotide enables the detection of:- the haplotype of the single nucleotide variant rs549846383 in the DNA of a genomic DNA sample obtained from a subject;- the haplotype of the single nucleotide variant rsl48687709 in the DNA of a genomic DNA sample obtained from a subject;- the presence of a tandem repeat expansion in the DNA of a genomic DNA sample obtained from a subject relative to the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1) or equivalent in an alternative human genome build version;- the presence of a repeat length polymorphism in the DNA of a genomic DNA sample obtained from a subject relative to the wild-type genomic DNA region of chrl5:34480576-34480608 according to the human genome build GRCh38 (SEQ ID NO:2) or equivalent in an alternative human genome build version; or- the presence of a tandem repeat expansion in a transcriptomic sample obtained from the subject, wherein the tandem repeat expansion is relative or compared to a wild-type transcriptome corresponding to or covering the wild-type genomic DNA region of chrl5:34419425-34419451 according to the human genome build GRCh38 (SEQ ID NO:1) or equivalent in an alternative human genome build version; wherein the oligonucleotide enables the detection by means of a DNA genotyping method.
Citation Information
Patent Citations
Method for diagnosing a neurodegenerative disease
CA2846307A1
Methods for the diagnosis of amyotrophic lateral sclerosis and frontotemporal lobar degeneration
WO2013041577A1