Methods for detecting repeat expansion diseases
The method leverages long read sequencing and Adaptive Sampling to comprehensively diagnose repeat elongation diseases, overcoming the limitations of current techniques by providing a practical and effective diagnostic flow for repeat elongation diseases.
Patent Information
- Application Number
- JP2023184169
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-26
- Publication Date
- 2025-05-13
AI Technical Summary
Current methods for diagnosing repeat elongation diseases are time-consuming, technically challenging, and often incomplete due to the need for repeated experiments and the difficulty in analyzing GC-rich and long tandem repeats.
A method using long read sequencing by nanopore sequencing that employs Adaptive Sampling for target enrichment, allowing for comprehensive analysis of repeat elongation diseases without the need for advanced expertise, and includes a diagnostic flow that ranks and evaluates repeat regions for pathological extensions.
The method provides a practical and comprehensive approach to diagnosing repeat elongation diseases, achieving comparable or higher diagnostic performance than conventional methods, and allowing for easy expansion of analyzed genomic regions to cover new pathological repeat extensions.
Smart Images

Figure 2025073408000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a method for detecting repeat expansion diseases. [Background technology]
[0002] Tandem repeats are one of the common sequence alterations in the human genome (Non-Patent Document 1). Tandem repeat expansions can cause diseases, usually with neurological symptoms. To date, it has been reported that approximately 60 tandem repeat expansions are associated with more than 69 types of diseases (Non-Patent Document 2). Repeat expansions are the most common cause of inherited neuromuscular diseases (Non-Patent Document 3), and the clinical course of these diseases is characterized by progressive neuromuscular dysfunction and severe impairment of the ADL of patients, so there is an urgent need to develop and provide effective neurotherapies. In recent years, promising neurotherapeutic approaches have been reported for various neurodegenerative diseases, such as antisense oligonucleotides (Non-Patent Documents 4 and 5), small molecule compounds (Non-Patent Document 6), and antibodies (Non-Patent Document 7). For all of these approaches, accurate molecular diagnosis is required.
[0003] Molecular diagnosis of repeat expansion diseases is a difficult challenge for physicians and researchers. First, there is a high locus heterogeneity, which requires repeated experiments to investigate possible loci. Second, tandem repeats are often GC-rich and long, making PCR difficult. Conventional diagnostic methods mainly rely on PCR-based methods such as flanking PCR and fragment analysis of the expansion region, repeat-primed PCR (RP-PCR), or Southern blotting, and PCR conditions / primers or probes must be set up specifically for each locus. However, this is time-consuming, technically challenging, requires extensive optimization, and is sometimes not feasible. In reality, diagnosis is performed by selecting a few loci rather than all of them, which may result in incomplete screening.
[0004] Short-read sequencing technology, which is a so-called second-generation sequencing technology, has contributed to the discovery of new repeat expansions that cause diseases to some extent, but the read length of about 100 to 300 bases is often not enough to cover the entire repeat expansion, and its contribution has been limited. In contrast, long-read sequencing, which is a third- or fourth-generation sequencing technology represented by the Oxford Nanopore or Pacific Biosciences platforms, enables long-read sequencing of several thousand bases or more than 10 kb, which can cover the entire repeat expansion or overcome low-complexity GC-rich genomic regions. In recent years, these technologies have led to the discovery of repeat expansion diseases involving difficult sequences such as GC-rich repeat motifs (Non-Patent Documents 8 and 9) and repeat motifs that differ from the reference genome (Non-Patent Documents 1, 10-13). In addition, the recent availability of methods to capture regions of interest, such as PCR-based enrichment, non-PCR enrichment with Cas9 (Non-Patent Documents 14, 15), or software-based target enrichment methods that do not require prior sample preparation, such as Read Until (Non-Patent Documents 16-18), has enabled long-read sequencing of target regions and is beginning to be applied to human genetic research (Non-Patent Documents 19, 20).
[0005] The above two platforms are representative examples of long-read sequencing platforms currently in practical use. The third-generation technology of Pacific Biosciences uses a single molecule of DNA polymerase to detect fluorescently labeled nucleotides incorporated during DNA synthesis in real time. The fourth-generation technology is a single-molecule DNA sequencing technology that does not use fluorescent labels and directly reads DNA molecules using physical or chemical methods, and a representative example of this is nanopore sequencing. Nanopore sequencing is a technology that determines the sequence of genomic DNA molecules by the change in current when they pass through a nano-sized pore (nanopore). In addition to methods that use nanopores of biomolecules, such as the Oxford Nanopore Technologies platform that uses pores made of proteins, methods that use pores made of inorganic materials (solid nanopores) are also being developed (Non-Patent Documents 21, 22).
[0006] Two papers using different techniques have been published on comprehensive detection methods using long-read sequencing targeting pathological repeat expansion sequences that cause known repeat expansion diseases in humans (Non-Patent Documents 20 and 23). The detection method disclosed in Non-Patent Document 20 is a detection method in which the target region is sequenced by nanopore sequencing without concentrating the library, as in the present invention, but the flow leading to diagnosis after sequencing is not shown, and it is more of a research technology. In addition, in the method of Non-Patent Document 20, some causative genes that cause human neurological diseases to be searched are excluded from the genes to be analyzed. The detection method disclosed in Non-Patent Document 23 is a method that uses CRISPR / Cas9 to concentrate the target region, and sample processing using the CRISPR / Cas9 system is essential during library preparation. In this method, it is necessary to design guide RNAs according to the target region, which is cumbersome for comprehensive analysis of the entire genome. In fact, even in Non-Patent Document 23, the analysis is limited to 10 locations. Furthermore, in both of the methods described in Non-Patent Documents 20 and 23, it is difficult to determine whether or not a repeat expansion disease is present without advanced specialized knowledge of the disease, and the methods are not at a level suitable for practical clinical application as diagnostic tests. [Prior art documents] [Non-patent literature]
[0007] [Non-Patent Document 1] Depienne, C. & Mandel, JL 30 years of repeat expansion disorders: what have we learned and what are the remaining challenges? Am. J. Hum. Genet. 108, 764-785 (2021). [Non-Patent Document 2] Gall-Duncan, T., Sato, N., Yuen, RKC & Pearson, CE Advancing genomic technologies and clinical awareness accelerates discovery of disease-associated tandem repeat sequences. Genome Res. 32, 1-27 (2022). [Non-Patent Document 3] Lockhart, PJ Advancing the diagnosis of repeat expansion disorders. Lancet Neurol. 21, 205-207 (2022). [Non-Patent Document 4] Tran, H. et al. Suppression of mutant C9orf72 expression by a potent mixed backbone antisense oligonucleotide. Nat. Med. 28, 117-124 (2022). [Non-Patent Document 5] Ellerby, LM Repeat expansion disorders: mechanisms and therapeutics. Neurotherapeutics 16, 924-927 (2019). [Non-Patent Document 6] Nakamori, M. et al. A slipped-CAG DNA-binding small molecule induces trinucleotide-repeat contractions in vivo. Nat. Genet. 52, 146-159 (2020). [Non-Patent Document 7] Nguyen, L. et al. Antibody therapy targeting RAN proteins rescues C9 ALS / FTD phenotypes in C9orf72 mouse model. Neuron 105, 645-662.e611 (2020). [Non-Patent Document 8] Sone, J. et al. Long-read sequencing identifies GGC repeat expansions in NOTCH2NLC associated with neuronal intranuclear inclusion disease. Nat. Genet. 51, 1215-1221 (2019). [Non-Patent Document 9] Ishiura, H. et al. Noncoding CGG repeat expansions in neuronal intranuclear inclusion disease, oculopharyngodistal myopathy and an overlapping disease. Nat. Genet. 51, 1222-1232 (2019). [Non-Patent Document 10] Ishiura, H. et al. Expansions of intronic TTTCA and TTTTA repeats in benign adult familial myoclonic epilepsy. Nat. Genet. 50, 581-590 (2018). [Non-Patent Document 11] Florian, RT et al. Unstable TTTTA / TTTCA expansions in MARCH6 are associated with Familial Adult Myoclonic Epilepsy type 3. Nat. Commun. 10, 4919 (2019). [Non-Patent Document 12] Yeetong, P. et al. TTTCA repeat insertions in an intron of YEATS2 in benign adult familial myoclonic epilepsy type 4. Brain 142, 3360-3366 (2019). [Non-Patent Document 13] Corbett, MA et al. Intronic ATTTC repeat expansions in STARD7 in familial adult myoclonic epilepsy linked to chromosome 2. Nat. Commun. 10, 4920 (2019). [Non-Patent Document 14] Giesselmann, P. et al. Analysis of short tandem repeat expansions and their methylation state with nanopore sequencing. Nat. Biotechnol. 37, 1478-1481 (2019). [Non-Patent Document 15] Gilpatrick, T. et al. Targeted nanopore sequencing with Cas9-guided adapter ligation. Nat. Biotechnol. 38, 433-438 (2020). [Non-Patent Document 16] Payne, A. et al. Readfish enables targeted nanopore sequencing of gigabase-sized genomes. Nat. Biotechnol. 39, 442-450 (2021). [Non-Patent Document 17] Kovaka, S., Fan, Y., Ni, B., Timp, W. & Schatz, MC Targeted nanopore sequencing by real-time mapping of raw electrical signal with UNCALLED. Nat. Biotechnol. 39, 431-441 (2021). [Non-Patent Document 18] Loose, M., Malla, S. & Stout, M. Real-time selective sequencing using nanopore technology. Nat. Methods 13, 751-754 (2016). [Non-Patent Document 19] Miller, DE et al. Targeted long-read sequencing identifies missing disease-causing variation. Am. J. Hum. Genet. 108, 1436-1449 (2021). [Non-Patent Document 20] Stevanovski, I. et al. Comprehensive genetic diagnosis of tandem repeat expansion disorders with programmable targeted nanopore sequencing. Sci. Adv. 8, eabm5386 (2022). [Non-Patent Document 21] Nakamura, Current status and future of next generation sequencing technology-2020. Journal of Bioengineering, Vol. 99, No. 5, 242-245 [Non-Patent Document 22] Special WEB column: Applied physics learned from the COVID-19 pandemic "DNA sequencer" Kenichi Takeda, July 1, 2020, [Retrieved October 12, 2023], Internet<URL: https: / / www.jsap.or.jp / columns-covid19 / covid19_2-4-1> [Non-Patent Document 23] Erdmann, H. et al. Parallel in-depth analysis of repeat expansions in ataxia patients by long-read sequencing. Brain, Volume 146, Issue 5, May 2023, Pages 1831-1843. Summary of the Invention [Problem to be solved by the invention]
[0008] An object of the present invention is to provide a practical method for detecting repeat expansion diseases that allows comprehensive testing for repeat expansion diseases without the need for advanced specialized knowledge about repeat expansion diseases. [Means for solving the problem]
[0009] The inventors of the present application focused on a long-read sequencing method using nanopore sequencing from Oxford Nanopore Technologies, which uses Adaptive Sampling, a software-based targeted sequencing method that does not require sample pretreatment to concentrate the target region. After conducting a discovery study using 12 patients who had been genetically diagnosed with repeat expansion diseases as positive controls, and a validation study on 10 patients who had been clinically diagnosed with SCA or CANVAS but had not been genetically diagnosed, they have succeeded in developing an analysis and diagnostic flow that makes it possible to reach a diagnosis even without advanced specialized knowledge of repeat expansion diseases, and have completed the present invention, which encompasses the following aspects.
[0010] [1] A method for detecting a repeat expansion disease, comprising the steps of: detecting a repeat expansion disease in a subject determined to have a pathological repeat; A library preparation step in which a library for sequencing is prepared from genomic DNA isolated from the subject. A sequencing step in which long-read sequencing of the library is performed by nanopore sequencing using multiple disease-associated repeat regions and their surrounding genomic regions as target regions to obtain sequence data for each target region of the subject's genome. A consensus sequence construction process in which a consensus sequence is constructed from the sequence data obtained in the sequencing process. a ranking step of detecting the number of repeat motif repetitions in the plurality of disease-associated repeat regions as repeat numbers, and ranking the disease-associated repeat regions in the subject genome in order of likelihood of being pathological repeat expansions by comparing each repeat number data in the subject genome with repeat number data in each disease-associated repeat region in a group of healthy individuals. A determination step in which evaluation is started from the disease-associated repeat region ranked first according to steps S1 to S7 below, and whether or not the subject has a pathogenic repeat expansion is determined. S1: Check whether the repeat region contains an abnormal repeat expansion that exceeds the disease threshold. If it does, proceed to S2. If it does not, proceed to S7. S2: Confirm the inheritance type of the repeat expansion disorder associated with the repeat expansion. If the disorder is autosomal dominant or X-linked dominant and the patient is female, proceed to S3. If the disorder is autosomal recessive, X-linked recessive, or X-linked dominant and the patient is male, proceed to S5. S3: Check whether the repeat region is a region known to have a benign repeat expansion. If yes, proceed to S4. If no, proceed to decision 1. S4: Check whether the repeat expansion detected in the repeat region contains a pathological repeat motif. If it does, proceed to decision 1. If it does not, proceed to S7. S5: Check whether all alleles have the repeat expansion. If yes, proceed to S6. If no, proceed to S7. S6: Check whether the repeat region is the repeat region of the RFC1 locus. If it is the RFC1 locus, proceed to S4. If it is not the RFC1 locus, proceed to decision 1. S7: Check whether the evaluation of the disease-associated repeat region ranked second is complete. If not, return to S1 and start the evaluation of the disease-associated repeat region ranked second. If it is complete, proceed to decision 2. Decision 1: The subject is determined to have a pathogenic repeat expansion. Decision 2: The subject is determined to have no known pathogenic repeat expansion. [2] The disease-associated repeat regions are the repeat regions (1) to (46) below (the location of each repeat on the chromosome is shown in the human reference genome GRCh38): (1) A repeat region associated with neuronal intranuclear inclusion disease (ILD) in the 5' untranslated region of the NOTCH2NLC locus (positions 149390802-149390842 on chromosome 1) (2) A repeat region associated with familial adult myoclonus epilepsy type 2 (positions 96197066 to 96197124 on chromosome 2) located in an intron of the STARD7 gene locus (3) A repeat region associated with syndactyly in the coding region of the HOXD13 gene locus (positions 176093058 to 176093103 on chromosome 2) (4) A repeat region associated with early infantile epileptic encephalopathy type 71 (positions 190880872 to 190880920 on chromosome 2) in the 5' untranslated region of the GLS locus (5) A repeat region associated with spinocerebellar ataxia type 7 in the coding region of the ATXN7 gene locus (positions 63912685-63912715 on chromosome 3) (6) A repeat region associated with myotonic dystrophy type 2 (positions 129172576-129172656 on chromosome 3) located in an intron of the CNBP gene locus (7) A repeat region (positions 138946020 to 138946062 on chromosome 3) in the coding region of the FOXL2 gene locus that is associated with blepharophimosis, ptosis, and epicanthal inversion syndrome. (8) A repeat region associated with familial adult myoclonus epilepsy type 4 (positions 183712187 to 183712226 on chromosome 3) located in an intron of the YEATS2 gene locus (9) A repeat region associated with Huntington's disease in the coding region of the HTT gene locus (positions 3074876-3074939 on chromosome 4) (10) A repeat region associated with cerebellar ataxia-neuropathy-vestibular areflexia syndrome (CTA-VAS) in an intron of the RFC1 gene (positions 39348424-39348483 on chromosome 4) (11) A repeat region associated with congenital central hypoventilation syndrome (CCH) in the coding region of the PHOX2B gene locus (positions 41745971-41746031 on chromosome 4) (12) A repeat region associated with familial adult myoclonus epilepsy type 7 in an intron of the RAPGEF2 gene locus (positions 159342526 to 159342618 on chromosome 4) (13) A repeat region associated with familial adult myoclonus epilepsy type 3 (positions 10356339 to 10356411 on chromosome 5) in an intron of the MARCHF6 gene locus (14) A repeat region associated with spinocerebellar ataxia type 12, located in an intron of the PPP2R2B gene locus (positions 146878728-146878758 on chromosome 5) (15) A repeat region associated with spinocerebellar ataxia type 1 in the coding region of the ATXN1 gene locus (positions 16327635 to 16327722 on chromosome 6) (16) A repeat region associated with cleidocranial dysplasia in the coding region of the RUNX2 gene locus (positions 45422750 to 45422801 on chromosome 6) (17) A repeat region associated with spinocerebellar ataxia type 17 in the coding region of the TBP gene locus (positions 170561907 to 170562021 on chromosome 6) (18) A repeat region associated with hand-foot-genital syndrome in the coding region of the HOXA13 gene locus (positions 27199924-27199966 on chromosome 7) (19) A repeat region associated with oculopharyngeal distal myopathy in the 5' untranslated region of the LRP12 gene locus (positions 104588970 to 104588999 on chromosome 8) (20) A repeat region associated with familial adult myoclonus epilepsy type 1 (positions 118366815 to 118366918 on chromosome 8) located in an intron of the SAMD12 gene locus (21) A repeat region associated with frontotemporal dementia / amyotrophic lateral sclerosis (positions 27573528-27573546 on chromosome 9) located in an intron of the C9orf72 gene locus (22) A repeat region associated with Friedreich's ataxia in an intron of the FXN gene (positions 69037286 to 69037304 on chromosome 9) (23) A repeat region associated with oculopharyngeal distal myopathy, located in the exons of the LOC642361 and NUTM2B-AS1 loci (positions 79826383 to 79826404 on chromosome 10) (24) A repeat region associated with dentatorubral-pallidoluysian atrophy (positions 6936716-6936773 on chromosome 12) in the coding region of the ATN1 gene locus (25) A repeat region associated with spinocerebellar ataxia type 2 in the coding region of the ATXN2 gene locus (positions 111598950 to 111599019 on chromosome 12) (26) A repeat region associated with spinocerebellar ataxia type 8 in an exon of the ATXN8OS gene locus (positions 70139383 to 70139428 on chromosome 13) (27) A repeat region associated with holoprosencephaly type 5 in the coding region of the ZIC2 locus (positions 99985448-99985493 on chromosome 13) (28) A repeat region associated with oculopharyngeal muscular dystrophy in the coding region of the PABPN1 gene locus (positions 23321472 to 23321502 on chromosome 14) (29) A repeat region associated with spinocerebellar ataxia type 3 in the coding region of the ATXN3 gene locus (positions 92071010 to 92071040 on chromosome 14) (30) A repeat region (positions 24613438 to 24613532 on chromosome 16) associated with familial adult myoclonus epilepsy type 6 in an intron of the TNRC6A gene locus (31) A repeat region associated with spinocerebellar ataxia type 31 in an intron of the BEAN1 gene locus (positions 66490396 to 66490466 on chromosome 16) (32) A repeat region associated with Huntington's disease type 2, located in an intron of the JPH3 gene locus (positions 87604287 to 87604329 on chromosome 16) (33) A repeat region associated with Fuchs endothelial corneal dystrophy type 3, located in an intron of the TCF4 gene locus (positions 55586153 to 55586229 on chromosome 18) (34) A repeat region associated with spinocerebellar ataxia type 6 in the coding region of the CACNA1A gene locus (positions 13207858-13207897 on chromosome 19) (35) A repeat region associated with oculopharyngeal distal myopathy in the 5' untranslated region of the GIPC1 gene locus (positions 14496041-14496075 on chromosome 19) (36) A repeat region associated with myotonic dystrophy type 1 in the 3' untranslated region of the DMPK gene locus (positions 45770204 to 45770264 on chromosome 19) (37) A repeat region associated with spinocerebellar ataxia type 36 in an intron of the NOP56 gene locus (positions 2652733 to 2652757 on chromosome 20) (38) A repeat region associated with Unverricht-Lundborg disease / progressive myoclonus epilepsy type 1 in the promoter region of the CSTB gene (positions 43776443-43776479 on chromosome 21) (39) A repeat region associated with spinocerebellar ataxia type 10 in an intron of the ATXN10 gene locus (positions 45795354 to 45795424 on chromosome 22) (40) A repeat region associated with spinal-bulbar muscular atrophy (positions 67545317 to 67545386 on the X chromosome) in the coding region of the AR gene locus (41) A repeat region associated with early infantile epileptic encephalopathy type 1 (positions 25013649 to 25013697 on the X chromosome) in the coding region of the ARX gene locus (42) A repeat region (positions 140504316 to 140504361 on the X chromosome) in the coding region of the SOX3 gene locus that is associated with mental retardation associated with isolated growth hormone deficiency (43) A repeat region associated with fragile X-associated tremor / ataxia syndrome in the 5' untranslated region of the FMR1 locus (positions 147912050 to 147912110 on the X chromosome) (44) A repeat region associated with fragile XE syndrome in the 5' untranslated region of the AFF2 locus (positions 148500637 to 148500682 on the X chromosome) (45) A repeat region associated with spinocerebellar ataxia type 37 in an intron of the DAB1 gene locus (positions 57367043 to 57367125 on chromosome 1) (46) A repeat region associated with Baratela-Scott syndrome in the promoter region of the XYLT1 gene locus (positions 17470907-17470930 on chromosome 16) and the regions known to have benign repeat expansions are (2), (8), (12), (13), (20), (30), (31) and (45). [3] The method according to [1] or [2], wherein in the library preparation step, the library is prepared by shearing the genomic DNA to 35 to 45 kb. [4] The sequencing process comprises: Step 1, specifying a target region and initiating sequencing of each single DNA molecule in the library; step 2, mapping the sequence of the 5'-side 300 to 700 bases of the single DNA molecule to a reference human genome in real time; and Step 3: Continue sequencing the single DNA molecule if it is mapped to the specified target region, and stop sequencing the single DNA molecule if it is mapped outside the specified target region, eject the single DNA molecule from the pore, start sequencing another single DNA molecule, and return to step 2. The method according to any one of [1] to [3], comprising: [5] The method according to any one of [1] to [4], wherein the target region consists of each disease-associated repeat region and 30 kb to 70 kb of genomic regions upstream and downstream thereof. [6] The method according to any one of [1] to [5], wherein in the sequencing step, sequence data having an average read depth of 10x or more for the target region is obtained. [7] The method according to any one of [1] to [6], comprising: in a library preparation step, preparing multiple libraries derived from multiple subjects; in a sequencing step, after sequencing of one library is completed, washing the flow cell with nuclease, loading the next library and sequencing, thereby sequentially sequencing the multiple libraries in one flow cell, and performing a ranking step and an evaluation step for each subject. Effect of the Invention
[0011] The present invention provides a practical method for detecting repeat expansion diseases for the first time, which can comprehensively test for repeat expansion diseases without advanced specialized knowledge of repeat expansion diseases. Through a comparative study with conventional methods, it has been confirmed that the diagnostic performance is equal to or better than that of conventional methods (see the following examples), and the practicality is very high. According to the present invention, repeat regions not covered in Non-Patent Documents 20 and 23 can be covered, and the target genomic region to be analyzed can be easily expanded at any time, so that even if a new pathological repeat expansion is identified, it can be easily added and comprehensively analyzed. In addition, the procedure (diagnosis flow) performed in the judgment step is not disclosed in Non-Patent Documents 20 and 23, and is a flow independently developed by the present inventors. The present invention greatly contributes to the diagnosis of repeat expansion diseases. [Brief description of the drawings]
[0012] [Figure 1] FIG. 2 is a flow chart showing the procedure of the determination step in the method of the present invention. [Diagram 2]FIG. 1 shows a workflow of screening for repeat expansion diseases causing cerebellar ataxia by a conventional method, which was performed in the Examples. [Diagram 3] 1 is a diagram showing the flow and required time of data analysis according to the present invention performed in the Examples. In the sections on data analysis and additional analysis, the programs used are presented to the right of each analysis item. [Figure 4] Box plot showing coverage depth for 59 targeted loci in 22 patients. Vertical lines within boxes indicate median coverage, dots indicate outliers. [Diagram 5] This figure plots the distribution of repeat numbers in 54 alleles derived from 27 control samples for 46 loci that are particularly closely related to known repeat expansion diseases with Mendelian inheritance among the 59 target loci. The horizontal line in each plot indicates the median repeat number. In the graphs for loci related to autosomal dominant (dominant) genetic diseases with known benign repeat expansions, the vertical axis is in log10 scale. In the BEAN1 gene and SAMD12 gene, repeat expansions consisting only of benign sequences were found in the control samples (*), suggesting that it is not possible to distinguish between affected and unaffected individuals based on the repeat number alone. [Figure 6A] As an example of successful target enrichment in patient 9, a CANVAS patient, an integrative genomics viewer (IGV) image is shown showing successful capture of the entire RFC1 gene region by adaptive sampling. [Figure 6B] Coverage plots for patient 8 (left) and patient 5 (right). The top, middle and bottom panels show coverage plots across all chromosomes for all reads, on-target reads and off-target reads, respectively. [Figure 6C] The coverage per locus for "on-target" and coverage per 5000 bp of "off-target" reads were all plotted for patients 8 and 5. The average coverage depth for on-target and off-target is shown in the graph. [Figure 7-1]Histograms and waterfall plots of tandem-genotypes output for 12 positive control cases in the validation study. The x-axis of the histograms shows the copy number change compared to the repeat number in the human reference genome (HTT: 21, ATXN3: 10, CACNA1A: 13, DMPK: 20, ATXN8OS: 15, NOTCH2NLC: 13). Waterfall plots were generated in hac mode (patients 1, 2, 3) or sup mode (patients 4, 5, 6). [Figure 7-2] Histograms and waterfall plots of tandem-genotypes output for 12 positive control cases in the validation study. The x-axis of the histograms shows copy number changes relative to the repeat number in the human reference genome (20 for PHOX2B, 20 for SAMD12, 11 for RFC1, 13 for BEAN1, 4 for NOP56, and 3 for CSTB). Waterfall plots were generated in hac mode (patients 7, 8, and 9) or sup mode (patients 10, 11, and 12). As annotated in the graphs, the SAMD12 and BEAN1 loci are loci that can harbor benign repeat expansions in healthy individuals. [Figure 8] As an example of the diagnosis of a pathological repeat expansion, we present the case of patient 3. The repeat expansion in the TNRC6A locus, ranked first in patient 3, was determined to be a polymorphism by testing the consensus sequence constructed by our workflow. The repeat expansion in the CACNA1A gene, ranked second, was determined to be pathological in 13 / 22 instances. A PCR-based conventional method confirmed the presence of 22 pathogenic expansions in the CACNA1A gene. [Figure 9A]Analysis results of patient 13 in the discovery study. The results of the conventional method are shown on the left, and the results of our diagnostic method using GridION (T-LRS) are shown on the right. Posi: positive control, Nega: negative control, NTC: experimental condition control in which PCR reaction was performed without template DNA. The top left panel shows the results of flanking PCR for the CACNA1A locus. Patient 13 and six other patients (1–6) were tested on a 2% agarose gel. Only patient “2” was positive. The top right panel shows the results of targeted long-read sequencing (T-LRS) that detected the CACNA1A gene as the ranked 1st locus. The bottom right panel shows the results of flanking PCR and fragment analysis of the CACNA1A locus performed to confirm the findings, which confirmed the pathological repeat expansion of the CANCA1A locus. Flanking PCR was evaluated on a 2.5% agarose gel to improve resolution. [Figure 9B]Analysis results of patient 17 in the discovery study. The results of the conventional test are shown on the left, and the results of our diagnostic method using GridION are shown on the right. Posi is a positive control, Nega is a negative control, and NTC is an experimental condition control in which PCR was performed without adding template DNA. The left panel shows the results of flanking PCR for the ATXN8OS / ATXN8 locus, and patient 17 was judged to be positive. The right panel shows the results of T-LRS, which detected the repeat expansion of the ATXN8OS / ATXN8 locus as the number one ranked locus. The number of repeats was 66, and the pathogenicity was judged to be moderate (intermediate expansion that may not cause symptoms). Next, we constructed a consensus sequence and confirmed the sequence, and found that the benign polymorphic repeat sequence CTA, which is a benign polymorphic repeat sequence, was expanded (19 times) beyond the normal range (8-15 times) before the pathogenic sequence CTG, and that the expansion of the pathogenic sequence was relatively not that long. Although it cannot be completely ruled out that this locus has moderate pathogenicity, taking into consideration the total repeat length and the relatively small proportion of the pathological sequence therein, it was determined that this finding does not explain the patient's disease. The second ranked locus was ATXN7, but the repeat expansion of this locus did not reach a pathological level, so its pathogenicity was denied. [Figure 9C]Analysis results of patients 14 and 16 in the discovery study. The results of the conventional method are shown on the left, and the results of our diagnostic method using GridION are shown on the right. Posi: positive control, Nega: negative control, NTC: experimental condition control in which PCR was performed without template DNA, RP-PCR: repeat-primed PCR. The upper left panel shows electropherograms of flanking PCR of the BEAN1 locus in patients 14 and 16. In patients 14 and 16, abnormally extended amplicons were observed (arrows), as in the positive control. The center left panel shows the results of RP-PCR performed using primers that detect the pathogenic repeat sequence TGGAA repeat, and a sawtooth waveform was observed when the TGGAA sequence is present. The lower left panel shows the results of Sanger sequencing for a SNP that is frequently linked to SCA31, and one patient (patient 14) did not have the SNP. The right panel shows the results of T-LRS (histogram of copy number changes and consensus sequence in waterfall plot style) that detected the BEAN1 gene as the first ranked gene, and the genotyping of SCA31-linked SNPs displayed in the integrative genomics viewer. The results were consistent between the two methods, but our diagnostic method was able to obtain all the information in one experiment, whereas the conventional method required frequent experiments. It is known that there are rare SCA31-affected patients who do not have the SNPs linked to SCA31, and this was also the case in patient 14, suggesting that genotyping this SNP for diagnostic purposes may not be appropriate. [Figure 10A] Comparison of accuracy between hac and sup modes. The left panel shows the consensus sequences constructed in hac mode, and the right panel shows the consensus sequences constructed in sup mode. For patients 4, 5, 6, 10, 11, and 12, the accuracy of base calling increased when Guppy was used in sup mode. The consensus sequences were depicted in a waterfall plot. [Figure 10B]Comparison of accuracy between hac and sup modes. The left panel shows the consensus sequence constructed in hac mode, and the right panel shows the consensus sequence constructed in sup mode. In patient 19, who has an AAGGG repeat expansion in the RFC1 gene, base calling accuracy did not increase when the base was called in sup mode. The "other" sequences were mostly AAGG repeats. [Figure 10C] Comparison of accuracy between hac mode and sup mode. Patient A, who has an AAGGG repeat expansion in the RFC1 gene, was sequenced by T-LRS and high-accuracy long-read whole genome sequencing (HiFi LR-WGS) using PacBio Sequel II as previously reported (Miyatake, S. et al. 2022). The top left panel shows the results of T-LRS sequencing using Kit 110, and the bottom left panel shows the results of HiFi LR-WGS (Sequel II). HiFi LR-WGS showed almost no AAGG repeats, indicating that AAGG is an error sequence. The top right panel shows the tandem-genotypes histogram of T-LRS, showing that there is no strand bias in the read distribution. The bottom right panel shows the results of Southern blotting, showing that patient A has an AAGGG-specific repeat expansion. The read length appears to be shorter for T-LRS, which is likely due to AAGGG being recognized as AAGG. [Figure 11] Correlation analysis of repeat length between conventional and T-LRS methods. The top row shows the overall correlation of repeat lengths obtained by T-LRS and conventional methods, respectively. In the middle row, the left panel shows the correlation in short repeats, and the right panel shows the correlation in long repeats. The bottom right panel shows that a significant correlation was observed in long repeats after adding samples as previously reported (Miyatake, S. et al. 2022). [Figure 12A]Illustrative diagram of time-lag sampling. A single GridION flow cell was loaded with 45 fmol of library from sample 1 on day 1, and 55 fmol of library from sample 2 on day 2. Pathological repeat expansions were detected in both samples. [Figure 12B] An explanatory diagram of time-lag sampling. The results of time-lag sampling were performed by loading Sample 3 and Sample 4 onto one GridION flow cell. Sample 3 was loaded with 55 fmol of library on day 1, and Sample 4 was loaded with 55 fmol on day 2 and 15 fmol on day 3. Pathological repeat expansions were detected in both samples. For Sample 4, the sequencing output did not reach a sufficient level at the end of day 2, so the remaining library (15 fmol) was loaded on day 3. [Figure 12C] FIG. 1 is an illustration of time-lag sampling. The results of time-lag sampling using Cas9-mediated PCR-free enrichment libraries of sample 1 and sample 2. Sample 1 was sequenced for 7.5 hours and sample 2 for 24 hours. The sequence output of sample 2 was free of AAGGG repeat expansions, indicating that carryover DNA from sample 1 was not present. [Figure 13A] Figure 1 shows the results of downsampling fastq data of patient 10 at various ratios to estimate the recommended coverage depth of the method of the present invention. A coverage depth of 10x was required to resolve the 2 alleles. [Figure 13B] Figure 1 shows the results of downsampling the fastq data of patient 18 at various ratios to estimate the recommended coverage depth of the method of the present invention. A coverage depth of 14x was required to resolve the 2 alleles. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] In the present invention, the terms "dominant genes" and "recessive genes" recommended by the Japanese Medical Association are used to indicate the mode of inheritance. Dominant genes are synonymous with dominant genes, and recessive genes are synonymous with recessive genes. In this specification, the terms are also used in combination with the conventional terms "dominant genes (dominant genes)" and "recessive genes (recessive genes)."
[0014] For the long-read sequencing in the method of the present invention, nanopore sequencing, which is a fourth-generation sequencing technology, is used. Nanopore sequencing is a technology that determines the sequence by the change in current when one molecule of single-stranded DNA passes through a nano-sized pore structure (nanopore) arranged on a support. When single-stranded DNA is passed through a nanopore while a current is flowing through it, a characteristic disturbance in the current occurs for each nucleotide. An electrode from one nanopore is connected to one channel of the sensor, and this disturbance is measured for each pore and converted into base information (A, T, G, C), thereby reading the base sequence of the DNA molecule. The first and second generation sequencing require PCR amplification of the region to be sequenced, and the third generation single molecule real-time sequencing, which detects DNA synthesis in real time, also requires a polymerase reaction. In contrast, nanopore sequencing is a technique that directly reads DNA molecules, and by using the adaptive sampling method, it is possible to sequence only the desired genomic region from the entire genome library by specifying the target region or region of interest to the sequencer, without performing sample pretreatment such as concentrating the genomic region to be analyzed.
[0015] Nanopore sequencers such as Oxford Nanopore Technologies' GridION and PromethION use a sensor chip with a protein nanopore arranged on a lipid bilayer membrane, in which a nanopore is formed by a protein molecule, which is a biopolymer. The sensor chip is mounted on a flow cell that loads the library. Such a method using a biopolymer is called the bio-nanopore method. Bacterial proteins (CsgG, α-hemolysin, MspA, etc.) that have been appropriately modified are generally used as the proteins that make up the bio-nanopore. In contrast, a method using a nano-sized through hole formed in a thin film of an inorganic material as a nanopore (solid-state nanopore) is called the solid-state nanopore method. In both methods, the top and bottom of the membrane in which the nanopore is arranged are filled with an electrolyte solution, and an ionic current is generated by applying a voltage to the membrane. In contrast, the nanogap method is also known, which uses nanogap electrodes arranged on a support with a nano-sized distance between the electrodes, and converts the change in tunnel current characteristic of each base that occurs when a DNA molecule passes between the nanogap electrodes into base information to read the sequence.
[0016] The method for diagnosing or detecting a repeat expansion disease according to the present invention includes a library preparation step, a sequencing step, a consensus sequence construction step, a ranking step, and a determination step. The consensus sequence construction step and the ranking step may be performed between the sequencing step and the determination step, and either step may be performed first.
[0017] <Library preparation process> In the library preparation step, a library for sequencing is prepared from genomic DNA isolated from the subject. The subject is typically a human patient who has symptoms suggesting a repeat expansion disease and who needs to detect or diagnose whether or not he or she has a repeat expansion disease.
[0018] Genomic DNA may be extracted and purified by a conventional method from peripheral blood leukocytes collected from a subject or a lymphoblastoid cell line established from lymphocytes collected from a subject. Genomic DNA may be left unsheared, but in the present invention, it is preferable to shear the DNA to a size of about 35 to 45 kb to prepare a library. The expected extension size of the repeat extension targeted in the present invention is a maximum of 20 kb, and by shearing the DNA to the above size, the coverage depth of a single run can be increased while covering the entire length of the large repeat extension with a single read. An apparatus for fragmenting DNA to a predetermined size is known (for example, Megaruptor by Diagenode).
[0019] Unsheared or sheared genomic DNA is purified, end-repaired, or nicked, and then an adapter for nanopore sequencing is added to the 5' end of the double-stranded DNA to prepare a library. The adapter is a structure that allows single-stranded DNA molecules to pass through a nanopore, and in bio-nanopore platforms, an adapter bound to a motor protein such as helicase that moves on DNA is used to move the DNA molecule that has entered the nanopore little by little. When using a commercially available platform, kits for preparing a sequence library optimal for that platform are also available, so the library can be prepared using such a kit.
[0020] As described above, in nanopore sequencing, particularly in the adaptive sampling method, a target region or region of interest can be specified in the sequencer, allowing only a desired genomic region to be sequenced from a whole genome library. Therefore, in the present invention, pretreatment of the genomic DNA sample, such as concentration of the genomic region to be analyzed, is not required, and only shearing of the genomic DNA can be performed as desired.
[0021] <Sequencing process> In the sequencing step, a plurality of disease-related repeat regions and their surrounding genomic regions are designated as target regions, and long-read sequencing is performed by nanopore sequencing of the library to obtain sequence data of each target region of the subject's genome. The method of nanopore sequencing is not particularly limited, but the bio-nanopore method can be preferably used.
[0022] The disease-associated repeat regions to be analyzed include, for example, repeat regions (1) to (46) shown in Table 1. These 46 loci are the 46 loci that are particularly closely related to known repeat expansion diseases of Mendelian inheritance, out of the 59 loci related to repeat expansion diseases that were analyzed in the Examples below.
[0023] [Table 1-1]
[0024] [Table 1-2]
[0025] [Table 1-3]
[0026] [Table 1-4]
[0027] In Table 1, the repeat regions in which only the wild-type repeat motif is listed as the repeat motif are repeat regions whose pathological or non-pathological status is determined by the length of the repeat motif (number of repeats), such as (1), (3)-(7), (9), (11), (14)-(19), (21)-(29), (32)-(44), and (46).
[0028] A repeat region in which a mutant repeat motif is described as a repeat motif in addition to a wild-type repeat motif is a repeat region in which whether or not it is a pathological repeat expansion is determined not only by the length of the repeat expansion but also by the motif. A mutant repeat motif is a pathological repeat motif, and a repeat expansion that exceeds the onset threshold and contains a pathological repeat motif is a pathological repeat expansion. Nine such regions fall into this category: (2), (8), (10), (12), (13), (20), (30), (31), and (45).
[0029] In addition, all repeat motif notations in Table 1 are written in the sequence of the plus strand, but they may be written in the minus strand of the complementary strand or in a sequence shifted by one to several bases, and in such cases, they are considered to mean the same motif. Even if a motif consisting of a very small number of bases is expressed by shifting the bases, there is no particular effect on the number of repeats detected. For example, the fluctuation in motif notation between Table 1 and Reference iv or v is as follows, but all of them represent the repeat motif of the same locus related to the same disease. These are only examples, and the fluctuation in notation of each motif is not limited to these. The GGC motifs in (1), (23), (43), and (46) are represented as CGG with a shifted base. The ATTTC motifs in (2), (8), (12), (13), (20), and (30) are written as TTTCA with a shifted base. The GCA motif in (5) and (40) is represented as CAG with a shifted base. The CAGG motif in (6) is represented as CCTG on the minus strand. The GCT motif in (14) and (25) is represented as CAG with a shifted base on the minus strand. The TGC motif in (15) is represented as CAG with a shifted base on the minus strand. The GCA motif in (17) is represented as CAG with a shifted base. The CCG motif in (19) and (35) is represented as CGG on the minus strand. The GCCCCG motif in (21) is written as GGGGCC with a shifted base on the minus strand. (26) is also abbreviated as CAG in reference iv. The CTG motif in (29) is represented as CAG on the minus strand. The GCT motif in (32) is written as CAG with a base shift on the minus strand (reference iv) or as CTG with a base shift (reference v). The AGC motif in (33) is represented as CTG with a shifted base on the minus strand. The CAG motif in (36) is represented as CTG on the minus strand. The GGGCCT motif in (37) is written as GGCCTG with a shifted base. The GCC motif in (44) is represented as CCG with a shifted base.
[0030] In addition, those denoted as GCN in Table 1 are polyalanine diseases in which the repeat expansion portion codes for polyalanine (see References iv and v). Since N can be any base, it is denoted as GCN. Typical motifs in each locus include GCG in (3), GCAGCT in (7), GCC in (11), GCG in (16), GCC in (18), GCG in (27), GCG in (28), GCC in (41), and GCG in (42). In these loci, motifs in which N is a different base may be mixed, in which case the number of repeats is counted including the motif in which N is a different base. The same applies to GCN, as described above, with regard to the fluctuation in motif notation. Even if it is written as CNG or NGC by shifting one or two bases, or written on the complementary strand side, it means the same motif expressed as GCN in this specification.
[0031] In the sequencing step, multiple target regions are designated to a sequencer, and sequencing is performed to obtain sequence data for each target region of the subject's genome. Step 1, specifying a target region and initiating sequencing of each single DNA molecule in the library; step 2, mapping the sequence of the 5'-side 300 to 700 bases of the single DNA molecule to a reference human genome in real time; and Step 3: Continue sequencing the single DNA molecule if it is mapped to the specified target region, and stop sequencing the single DNA molecule if it is mapped outside the specified target region, eject the single DNA molecule from the pore, start sequencing another single DNA molecule, and return to step 2. By performing this procedure, it is possible to obtain enriched sequence data of the target region from a whole genome DNA library.
[0032] In step 1, for example, each disease-associated repeat region and genomic regions each about 30 kb to 70 kb (for example, about 30 kb to 60 kb) upstream and downstream thereof may be designated as target regions.
[0033] In step 2, once the base sequence of about 300 to 700 bases (for example, about 400 to 600 bases) on the 5' side of the single DNA molecule has been read, the base sequence of the 5' region is mapped in parallel to a reference human genome to determine whether or not it is a target region. "In real time" means that mapping is performed in parallel while sequencing is continued. For example, GRCh38 (or its latest successor version) or T2T-CHM13 (or its latest successor version) can be used as the reference human genome. In principle, this step can be performed by providing various reference genomes to the sequencer.
[0034] In step 3, if the base sequence of the 5' region is mapped to the specified target region, the sequencing of the single DNA molecule continues. If it is mapped outside the target region, the sequencing of the single DNA molecule is stopped, the single DNA molecule is ejected from the nanopore, and the sequencing of another single DNA molecule is started, and the process returns to step 2 to confirm by mapping the 5' region, and then proceeds to step 3 again. Thereafter, steps 2 and 3 are repeated for the DNA molecule that was mapped outside the target region. Software for specifying target regions and reading sequences using the above-mentioned procedure is known, for example, the open source software package Readfish integrated with the ONT Read Until API as used in Non-Patent Document 20 (A. Payne, N. Holmes, T. Clarke, R. Munro, BJ Debebe, M. Loose, Readfish enables targeted nanopore sequencing of gigabase-sized genomes. Nat. Biotechnol. 39, 442-450 (2021)).) and the adaptive sampling in the form of Read Until installed in the MinKNOW software of GridION MK1 from Oxford Nanopore Technologies, which was also used in the examples below (Loose, M., Malla, S. & Stout, M. Real-time selective sequencing using nanopore technology. Nat. Methods 13, 751-754 (2016)).
[0035] It is desirable to always use the latest base caller that converts the current change into base sequence information. When a sequencer has multiple base call modes, such as the high accuracy (hac) mode or the ultra high accuracy (sup) mode of the nanopore sequencer of Oxford Nanopore Technologies, a high accuracy mode may be appropriately selected and used as desired. When using a nanopore sequencer of Oxford Nanopore Technologies, it is desirable to use the hac mode or the sup mode.
[0036] In the present invention, if the average read depth of the target region is 10x or more, for example 12x or more, or 15x or more, both alleles can be separated and repeat expansion can be detected.
[0037] In the following examples, a nuclease flush is performed once or twice during the sequencing of one sequence library (for one subject), in which the flow cell of the sequencer is washed with nuclease, and the library is reloaded and sequenced, so that one library is sequenced two or three times, but such an operation is not essential when carrying out the method of the present invention. One library may be sequenced only once, or, as in the examples, a nuclease flush may be performed during the sequencing, the library may be reloaded, and sequencing may be carried out two or three times.
[0038] <Consensus sequence construction process> In this step, a consensus sequence is constructed from the sequence data obtained in the sequencing step. A consensus sequence is a single type of most likely sequence determined by calculating the most frequent base at each position of the sequence alignment in which the reads are aligned. Various tools for constructing consensus sequences are known, such as lamassemble (https: / / gitlab.com / mcfrith / lamassemble), and a consensus sequence can be constructed using such known tools. After constructing the consensus sequence, a waterfall plot such as that shown under each bar graph in Figure 7 may be created.
[0039] <Ranking process> In the ranking step, the repeat motif repeat number in a plurality of disease-associated repeat regions is detected as the repeat number. Next, the repeat number data of each subject genome is compared with the repeat number data of each disease-associated repeat region of a healthy subject group to rank the disease-associated repeat regions of the subject genome in order of likelihood of being pathological repeat expansions. In the following examples, the repeat number is detected and ranked using the tool tandem-genotypes (https: / / github.com / mcfrith / tandemgenotypes) (Mitsuhashi, S. et al. Tandem-genotypes: robust detection of tandem repeat expansions from long DNA reads. Genome Biology (2019) 20:58) available under an open source license.
[0040] Tandem-genotypes outputs a number for each repeat site indicating the increase or decrease in the number of repeats in the subject compared to the number of repeats in the reference human genome (copy number change). When focusing on one repeat site, the increase or decrease is output for each read that corresponds to this site, but since humans have two alleles, the multiple numbers output are distributed as a group with two peaks (see Figure 7). Two representative values are determined from this distribution, and the number of repeats for allele 1 is calculated as the representative value of peak 1, the number of repeats for allele 2 is calculated as the representative value of peak 2, and so on. The representative value of each peak is the increase or decrease in the number of repeats for that allele at that repeat site.
[0041] Next, the degree to which this copy number change number differs from the number of repeats in the reference human genome (the magnitude of the difference) is calculated as the priority score. When calculating the priority score, the representative value is not used, but the maximum and minimum values are excluded from the multiple copy number change numbers output, and the number with the largest change among the remaining numbers is used to calculate the value α using the following formula. α=(Length increase in bases) / (reference repeat length +30)
[0042] Next, α is multiplied according to the following rules to obtain the priority score. This process is performed so that even small changes in the number of repeats will have an effect on important regions. Among the 46 loci shown in Table 1, (5), (9), (15), (17), (24), (25), (29), (34), and (40) are polyglutamine diseases in which the repeat expansion codes for polyglutamine. Repeats of the coding region: x50 UTR repeats: x20 Promoter repeats: x15, Non-protein-coding RNA exon repeats: x15 Intron repeats: x5 In addition, if the repeat encodes polyglutamine or polyalanine, an additional x2 is performed.
[0043] In the method of the present invention, ranking is performed by comparing the repeat number data of the subject with the repeat number data of the healthy subject group. The comparison of repeat number data includes, as described above, calculating a score such as a priority score from the repeat number and comparing the scores. The healthy subject group is a group consisting of multiple healthy subjects who do not have repeat expansion diseases. The number of healthy subjects is not necessarily the more the better, and the inventors of the present application have found that a number of about 27 is sufficient (data omitted). Therefore, in the present invention, about 20 to 100 healthy subjects, for example about 25 to 35 healthy subjects, may be used as the healthy subject group. For the healthy subject group, sequence data of each healthy subject is prepared in the same manner as for the subject, and a priority score is calculated for each repeat site in the same manner and compared. The priority score of the subject is d, and the cube mean of the priority scores of each healthy subject in the healthy subject group is h. If d<0, the change in that site is considered to be small and is excluded from the calculation. If d ≥ 0, max(d-max(h, 0), 0) The joint priority score is calculated using this formula, and the repeat regions in the subject genome are ranked in descending order of this score and output.
[0044] <Judgment process> In the determination step, the evaluation is started from the disease-associated repeat region ranked first according to the following procedure or steps S1 to S7, and it is determined whether or not the subject has a pathological repeat expansion. In the subject determined to have a pathological repeat expansion, a repeat expansion disease is detected. These procedures are shown in a flow chart in FIG. 1. S1: Check whether the repeat region contains an abnormal repeat expansion that exceeds the disease threshold. If it does, proceed to S2. If it does not, proceed to S7. S2: Confirm the inheritance type of the repeat expansion disease associated with the repeat expansion. If the disease is autosomal dominant or X-linked dominant and the subject is female, proceed to S3. If the disease is autosomal recessive, X-linked recessive, or X-linked dominant and the subject is male, proceed to S5. S3: Check whether the repeat region is a region known to have a benign repeat expansion. If yes, proceed to S4. If no, proceed to decision 1. S4: Check whether the repeat expansion detected in the repeat region contains a pathological repeat motif. If it does, proceed to decision 1. If it does not, proceed to S7. S5: Check whether all alleles have the repeat expansion. If yes, proceed to S6. If no, proceed to S7. S6: Check whether the repeat region is the repeat region of the RFC1 locus. If it is the RFC1 locus, proceed to S4. If it is not the RFC1 locus, proceed to decision 1. S7: Check whether the evaluation of the disease-associated repeat region ranked second is complete. If not, return to S1 and start the evaluation of the disease-associated repeat region ranked second. If it is complete, proceed to decision 2. Decision 1: The subject is determined to have a pathogenic repeat expansion. Decision 2: The subject is determined to have no known pathogenic repeat expansion.
[0045] (Step S1) It is confirmed whether the repeat region under evaluation has an abnormal repeat expansion exceeding the onset threshold. The onset threshold can be set for each disease-associated repeat region from publicly known information. Table 1 above shows examples of onset thresholds X that can be set for the repeat regions (1) to (46). (1) can be set to a value within the range of, for example, 62 to 66, 63 to 66, or 64 to 66. (2) can be set to a value within the range of, for example, 140 to 150, 145 to 150, or 148 to 150. (3) can be set to a value within the range of, for example, 18 to 24, 20 to 24, or 20 to 22. (4) can be set to a value within the range of, for example, 50 to 680, 100 to 680, 200 to 680, 300 to 680, 400 to 680, 500 to 680, or 600 to 680. (5) can be set to 36 or 37. (6) can be set to a value within the range of, for example, 35 to 50, 40 to 50, or 45 to 50. (6) can be set to a value within the range of, for example, 35 to 50, 40 to 50, or 45 to 50. (7) can be set to a value within the range of, for example, 17 to 22, or 20 to 22. (8) can be set to a value within the range of, for example, 450 to 1000, 500 to 1000, 600 to 1000, 700 to 1000, 800 to 1000, 900 to 1000, or 950 to 1000. (10) can be set to a value within the range of, for example, 100 to 400, 200 to 400, 300 to 400, 350 to 400, or 480 to 400. (11) can be set to a value within the range of, for example, 22 to 25, 23 to 25, or 24 to 25. (12) can be set to a value within the range of, for example, 20 to 45, 20 to 40, or 25 to 40. (13) can be set to a value within the range of, for example, 200 to 660, 300 to 660, 400 to 660, 500 to 660, or 600 to 660. (14) can be set to a value within the range of, for example, 48 to 51, or 50 to 51. (15) can be set to a value of, for example, 39 or 40. (16) can be set to a value within the range of, for example, 20 to 27, 23 to 27, or 25 to 27. (17) can be set to a value within the range of, for example, 45 to 46. (18) can be set to a value within the range of, for example, 20 to 24, or 22 to 24. (19) can be set to a value within the range of, for example, 46 to 85, 46 to 80, 46 to 70, 46 to 65, or 46 to 60.(20) can be set to a value within the range of, for example, 80 to 105, 90 to 105, or 100 to 105. (21) can be set to a value within the range of, for example, 26 to 700, 25 to 600, 26 to 600, 25 to 500, 26 to 500, 25 to 400, 26 to 400, 25 to 300, 26 to 300, 25 to 200, 26 to 200, 25 to 100, 26 to 100, 25 to 50, 26 to 50, 25 to 40, or 26 to 40. (22) can be set to a value within the range of, for example, 40 to 66, 50 to 66, or 60 to 66. (23) can be set to a value within the range of, for example, 20 to 40, 20 to 39, 25 to 40, 25 to 39, 30 to 40, or 30 to 39. (24) can be set to a value within the range of, for example, 40 to 48, or 45 to 48. (26) can be set to a value within the range of, for example, 55 to 74, 60 to 74, or 65 to 74. (27) can be set to a value within the range of, for example, 18 to 25, 20 to 25, or 23 to 25. (28) can be set to a value within the range of, for example, 11 to 15, 12 to 16, or 12 to 15. (29) can be set to a value within the range of, for example, 50 to 55. (30) can be set to a value within the range of, for example, 18 to 500, 18 to 300, 18 to 100, 20 to 100, 25 to 100, or 25 to 50. (31) can be set to a value within the range of, for example, 100 to 500, 200 to 500, 300 to 500, or 400 to 500. (32) can be set to a value within the range of, for example, 32 to 41, 35 to 41, or 38 to 41. (33) can be set to a value within the range of, for example, 41 or 50. (34) can be set to a value within the range of, for example, 20 or 21. (35) can be set to a value within the range of, for example, 35 to 80, 35 to 70, or 40 to 60. (36) can be set to a value within the range of, for example, 40 to 51, or 45 to 51. (37) can be set to a value within the range of, for example, 200 to 650, 300 to 650, 400 to 650, 500 to 650, or 600 to 650. (38) can be set to a value within the range of, for example, 10 to 30, 15 to 30, 20 to 30, or 25 to 30. (39) can be set to a value within the range of, for example, 50 to 280, 100 to 280, 150 to 280, 200 to 280, or 250 to 280. (40) can be set to 37 or 38. (41) can be set to a value within the range of, for example, 20 to 23.(42) can be set to a value within the range of, for example, 20 to 26, or 23 to 26. (43) can be set to a value within the range of, for example, 53 to 55. (44) can be set to a value within the range of, for example, 70 to 200, 100 to 200, or 150 to 200. (45) can be set to a value within the range of, for example, 27 to 31. (46) can be set to a value within the range of, for example, 30 to 120, 50 to 120, 70 to 120, or 100 to 120.
[0046] When setting an abnormality threshold according to the output of tandem-genotypes, a numerical value can be set by subtracting the number of repeats in the reference human genome from the onset threshold numerical value in Table 1 and described above. Table 1 shows a specific example of an abnormality threshold according to the output of tandem-genotypes, which was set by the inventors of the present application from the number of repeats in 27 people or approximately 250 healthy individuals. However, the abnormality threshold according to the output of tandem-genotypes is not limited to this specific example.
[0047] If there is an abnormal repeat expansion exceeding the onset threshold, proceed to S2, otherwise proceed to S7. When performing the ranking step using tandem-genotypes, in this step S1, the judgment can be made first based on the output of tandem-genotypes, but the judgment may also be made by checking the consensus sequence constructed in the consensus sequence construction step or the waterfall plot.
[0048] (Step S2) If the repeat region being evaluated has an abnormal repeat expansion exceeding the disease threshold, proceed to this step to confirm the inheritance mode of the repeat expansion disease associated with that locus. Table 1 also shows the inheritance modes of the repeat expansion diseases associated with the repeat regions (1) to (46). If the inheritance pattern is autosomal dominant (or X-linked dominant (or X-linked dominant) and the subject is female, proceed to S3. If the inheritance pattern is autosomal recessive (or X-linked recessive (or X-linked recessive) and the subject is male, proceed to S5.
[0049] (Step S3) In the case where the inheritance pattern is autosomal dominant (dominant) or X-linked dominant (dominant) and the subject is female in S2, the procedure proceeds to step S3, where it is confirmed whether the region is known to have benign repeat expansion. In the repeat region shown in Table 1, of the nine locations mentioned above, eight locations (2), (8), (12), (13), (20), (30), (31), and (45) are known to have benign repeat expansion. Among autosomal dominant (dominant) pathological repeat expansions, there are cases in which the wild-type repeat motif, which is also seen in healthy individuals, is expanded as it is and exceeds the onset threshold, resulting in pathological repeat expansion, and cases in which the wild-type repeat motif, which is also seen in healthy individuals, is expanded and the repeat motif is mutated into a specific abnormal sequence (pathological motif) and expanded to include pathological repeat expansion. In the latter case, even if the wild-type repeat motif (normal motif) is extended to the same length as in patients with repeat expansion disease, it will not be a repeat expansion disease unless it contains an expansion of a pathological motif. More specifically, in a region known to have benign repeat expansion, in pathological repeat expansion, pathological motifs and normal motifs are generally mixed and extended (in the case of the above-mentioned 8 loci corresponding to Table 1, no cases have been reported in which only pathological motifs are extended, but even if only pathological motifs are extended, it can be said to be a pathological repeat expansion), and in the repeat region seen in a healthy person, both the repeat region not exceeding the onset threshold and the repeat expansion exceeding the onset threshold, only normal motifs are repeated, and pathological motifs are not mixed. Therefore, in step S3, it is confirmed whether the region is known to have benign repeat expansion, and the step to proceed to is branched to S4 and decision 1 depending on the result. In the above eight locations in Table 1, the wild-type repeat motif shown in the upper row is the normal motif, and the mutant repeat motif shown in the lower row is the pathological motif. If the answer is Yes, that is, if the region is known to have a benign repeat expansion, proceed to S4 and confirm whether the detected repeat expansion contains a pathological repeat motif.If the answer is No, i.e., if the region is not known to contain benign repeat expansions, the detected repeat expansion exceeding the onset threshold is deemed to be a pathological repeat expansion, so proceed to decision 1 and determine that the subject has a pathological repeat motif.
[0050] (Step S4) If the process proceeds to S4, the repeat expansion detected in the repeat region is checked for the presence of a pathological repeat motif using the consensus sequence constructed in the consensus sequence construction step. If a pathological repeat motif is present, the process proceeds to decision 1, and the subject is determined to have a pathological repeat motif. If a pathological repeat motif is not present, the repeat expansion is not a pathological repeat expansion, and the process proceeds to S7.
[0051] (Step S5) In S2, if the inheritance pattern of the abnormal repeat expansion is autosomal recessive (recessive) or X-linked recessive (recessive), and X-linked dominant (dominant) disease and the subject is male, proceed to this step S5. In this step, it is confirmed whether the repeat expansion is present in all alleles (two alleles in the case of autosomal recessive (recessive), one allele in the case of X-linked disease if male, and two alleles in the case of X-linked disease if female), in other words, whether there is no normal allele. If it is Yes, that is, if there is the repeat expansion in all alleles (no normal allele), proceed to S6. If it is No, that is, if there is a normal allele, the abnormal repeat expansion is not a pathological repeat (the abnormal repeat expansion does not cause a repeat expansion disease), so proceed to S7. Note that if the subject is male and proceeds to S5 because of X-linkage, the subject has only one allele on the X chromosome, so it automatically becomes Yes.
[0052] (Step S6) If the procedure proceeds from S5 to S6, it is confirmed whether the repeat region is a repeat region of the RFC1 locus. Repeat expansion disease CANVAS (Table 1 (10)) caused by repeat expansion of the RFC1 locus is an autosomal recessive (recessive) disease in which there is biallelic repeat expansion, and both repeat expansions contain pathological repeat motifs. In other words, it is the only disease known to have benign repeat expansion among known autosomal recessive (recessive) inherited repeat expansion diseases. Therefore, in S6, it is confirmed whether the repeat is of the RFC1 locus, and if Yes, proceed to S4 to confirm whether the pathological repeat motif is included. The normal variation of the expansion is generally recognized as AAAAG expansion and AAAGG expansion, while the pathological variation of the expansion is generally recognized as AAGGG expansion and ACAGG expansion. CANVAS patients have one of three patterns: AAGGG / AAGGG (both alleles have AAGGG expansion), AAGGG / ACAGG (one allele has AAGGG expansion and the other has ACAGG expansion), and ACAGG / ACAGG (both alleles have ACAGG expansion). If it is the RFC1 locus, proceed to S4 and check whether the pathogenic repeat motif (AAGGG, ACAGG) is included in the repeat expansion. If it is not the RFC1 locus, proceed to judgment 1 and determine that the subject has a pathogenic repeat motif.
[0053] (Step S7) In step S7, it is confirmed whether the evaluation of the disease-associated repeat region ranked second in the ranking step has been completed. If not, return to S1 and start the evaluation of the disease-associated repeat region ranked second. If completed, proceed to judgment 2 and judge that the subject does not have a known pathological repeat expansion. According to the method of the present invention, since the pathological repeat expansion of the subject is ranked within the second place in the ranking step, it is possible to detect whether the subject has a pathological repeat expansion, i.e., whether the subject suffers from a repeat expansion disease, by evaluating the repeat regions ranked first and second in the judgment step.
[0054] The steps S1 to S7 of the determination step may be evaluated sequentially by an implementer while referring to data and publicly known information, or a program for causing a computer to execute the steps S1 to S7 may be written and executed by the computer.
[0055] When detecting whether or not a plurality of subjects have repeat expansion disease by the method of the present invention, a nuclease flush may be performed to wash the flow cell of the sequencer with nuclease, and libraries for a plurality of subjects may be sequenced sequentially with one flow cell. That is, in the library preparation step, a plurality of libraries derived from a plurality of subjects are prepared, and in the sequencing step, after the sequence of one library is completed, the flow cell may be washed with nuclease, and the next library may be loaded and sequenced, thereby sequentially sequencing a plurality of libraries with one flow cell, and a ranking step and a judgment step may be performed for each subject. In this way, by performing a nuclease flush and sequencing libraries for a plurality of subjects with one flow cell, the cost of the sequencing step can be significantly reduced. In the case of a bio-nanopore type nanopore sequencer, the durability of the nanopore made of a biopolymer is lower than that of a solid nanopore, and the pore is easily damaged by reading long reads (Oxford Nanopore Technologies' nanopore sequencer is said to be capable of sequencing for up to 72 hours). Therefore, when using a bionanopore sequencer, it is desirable to limit the repeated use of the flow cell to about two or three people. EXAMPLES
[0056] The present invention will be described in more detail below with reference to examples, although the present invention is not limited to the following examples.
[0057] method patient We analyzed 22 patients with various neurological and neuromuscular diseases, including 12 patients genetically diagnosed with repeat expansion disease (positive controls) and 10 patients clinically diagnosed with SCA or CANVAS but without genetic diagnosis. Patients 8 and 12 (Mizuguchi, T. et al. 2021) and patient 9 (Miyatake, S. et al. 2022) have been reported previously. As a positive control for methylation analysis, we analyzed unaffected individual 1, who has a heterozygous hypermethylated repeat expansion in the NOTCH2NLC gene. This individual was recently reported as the asymptomatic father of a patient with NIID (Fukuda, H. et al. 2021) and is unrelated to patient 6 with NIID in our cohort. The experimental protocol was approved by the Ethics Committee of Yokohama City University School of Medicine. Written informed consent was obtained from all participants. Clinical information was collected from the patients' primary physicians. Genomic DNA was extracted by standard methods from peripheral blood leukocytes or lymphoblastoid cells established from the patient's lymphocytes.
[0058] Conventional repeat expansion screening For genetic diagnosis, we used flanking PCR, Sanger sequencing of the PHOX2B repeat region (Matera, I. et al. 2004), genotyping of a SNP linked to SCA31 located in the 5' untranslated region of the PLEKHG4 gene [rs886041026; PLEKHG4 gene (NM_001129729.3):c.-16C>T] (Ishikawa, K. et al. 2005), fragment analysis, RP-PCR, and / or Southern blotting. Fragment analysis of the HTT gene in patient 1 and Southern blotting of the DMPK gene in patient 4 were performed in a certified clinical laboratory. RP-PCR of the NOP56 gene in patient 11 was performed in another laboratory. Our conventional screening workflow for repeat expansion disorders causing cerebellar ataxia is shown in Figure 2.
[0059] Flanking PCR As a first-line screening test for patients suspected of SCA and / or CANVAS, the ATN1 gene (linked to dentatorubral-pallidoluysian atrophy) (Koide, R. et al. 1994), ATXN1 gene (linked to SCA1) (Orr, HT et al. 1993), TBP gene (linked to SCA17) (Nakamura, K. et al. 2001), PPP2R2B gene (linked to SCA12) (Holmes, SE et al. 1999), ATXN7 gene (linked to SCA7) (David, G. et al. 1997), ATXN8OS / ATXN8 genes (linked to SCA8) (Koob, MD et al. 1999), ATXN3 gene (linked to SCA3) (Kawaguchi, Y. et al. 1994), CACNA1A gene (linked to SCA6) (Zhuchenko, Flanking PCR encompassing the repeat sequences in the gene regions of the ATXN2 gene (linked to SCA2) (Pulst, SM et al. 1996), BEAN1 gene (linked to SCA31) (Sato, N. et al. 2009), and RFC1 gene (linked to CANVAS) (Cortese, A. et al. 2019) was performed as previously described. Primer sequences and PCR conditions are shown in Table 2.
[0060] Sanger sequencing Details of the primer sequences and PCR conditions used for Sanger sequencing of the PLEKHG4 and PHOX2B genes are shown in Table 2. Sanger sequencing was performed on an Applied Biosystems 3500xL Genetic Analyzer (Thermo Fisher Scientific) using the Big Dye Terminator cycle sequencing kit (v1.1 or v3.1) (Thermo Fisher Scientific Waltham, MA, USA).
[0061] Fragment Analysis Fragment analysis of the CACNA1A gene was performed by Eurofins Genomics (Tokyo, Japan) using the same primers and PCR settings as those used in the flanking PCR (Zhuchenko, O. et al. 1997), except that the forward primer was labeled with 6-carboxyfluorescein (6-FAM). The primer sequences and PCR conditions are shown in Table 2.
[0062] [Table 2]
[0063] RP-PCR RP-PCR to detect the GGC repeat in the NOTCH2NLC gene (linked to NIID) (Sone, J. et al. 2019), the TTTCA repeat in the SAMD12 gene (linked to benign adult familial myoclonic epilepsy) (Ishiura, H. et al. 2018; Mizuguchi, T. et al. 2021), the AAGGG repeat in the RFC1 gene (Cortese, A. et al. 2019), and the TGGAA repeat in the BEAN1 gene (linked to SCA31) (Ishige, T. et al. 2012) was performed as previously described (Sone, J. et al. 2019; Ishiura, H. et al. 2018; Mizuguchi, T. et al. 2021; Cortese, A. et al. 2019; Ishige, T. et al. 2012). The primer sequences and PCR conditions are shown in Table 2. RP-PCR products were resolved and visualized using an Applied Biosystems 3500xL Genetic Analyzer (Thermo Fischer Scientific) and analyzed using GeneMapper software (Thermo Fischer Scientific).
[0064] Southern blotting Southern blotting for repeat expansions in the SAMD12 gene and repeat expansions in the CSTB gene for patients 8 and 12 was performed as previously reported (Mizuguchi, T. et al. 2021). Southern blotting for repeats in the RFC1 gene intron was performed as previously reported (Cortese, A. et al. 2019). Specifically, 3 or 5 μg of genomic DNA was digested overnight with EcoRI-HF (New England Biolabs, Ipswich, Massachusetts, USA), and the digested DNA was separated on a 0.8% agarose gel in Tris-borate-EDTA buffer. The gel was depurinated with 0.25 M HCl for 8 min, followed by two DNA denaturations for 15 min with 0.5 M NaOH / 1.5 M NaCl. The digested DNA was transferred to a positively charged nylon membrane. DNA fragments were fixed using the automated cross-linking mode of a UV Stratalinker 2400 (Stratagene, La Jolla, CA, USA). Prehybridization was performed in DIG Easy Hyb buffer (Sigma-Aldrich, St. Louis, MO, USA) at 37°C for 60 min. DIG-labeled probes were amplified by PCR of genomic fragments using the forward primer 5'-ATTAGGTGTCTGGTGAGGGC-3' (SEQ ID NO: 48) and reverse primer 5'-GAAGAATGGCCCCAAAAGCA-3' (SEQ ID NO: 49) (Eurofins Genomics). PCR products were cloned into pCR4-TOPO Vector (Invitrogen, Waltham, MA, USA). DIG-labeled RFC1-PCR probes were generated using the PCR DIG Probe Synthesis Kit (Roche, Basel, Switzerland) according to the manufacturer's instructions. Hybridization was performed overnight at 37°C in DIG Easy Hyb buffer containing denatured RFC1-PCR labeled probe.After hybridization, the membrane was washed twice in 2x saline-sodium citrate (SSC) buffer containing 0.1% SDS for 5 min at room temperature, then twice in 0.5x SSC buffer containing 0.1% SDS for 15 min at 65°C. DIG-labeled probes were detected by chemiluminescence using alkaline phosphatase-conjugated anti-DIG antibody (Anti-Digoxigenin-AP, Fab fragment, Roche) and chemiluminescent substrate CDP-star (Roche). Briefly, the membrane was blocked in 1x blocking solution for 30 min, then incubated in antibody solution (75 mU / mL anti-DIG-AP) for 30 min, followed by washing twice in washing buffer (0.1 M maleic acid, 0.15 M NaCl, 0.3% Tween 20) for 15 min at room temperature. Signals were visualized with a ChemiDoc Touch (Bio-Rad, Hercules, CA, USA). The membrane was reprobed with a custom DIG-labeled probe (synthesized by Eurofins Genomics) to detect AAGGG repeats ([DIG]-AAGGGAAGGGAAGGGAAGGGAAGGGAAGGGAAGGGAAGGGAAGGG, SEQ ID NO: 50) and ACAGG repeats ([DIG]-ACAGGACAGGACAGGACAGGACAGGACAGGACAGGACAGGACAGG, SEQ ID NO: 51).
[0065] Real-time T-LRS on GridION using adaptive sampling Our workflow is shown in Figure 3. Approximately 2–3 μg of unsheared purified genomic DNA or DNA sheared to a target size of 40 kb with a Megaruptor 2 (Diagenode, Seraing, Belgium) was used to construct sequencing libraries using the Oxford Nanopore Ligation Sequencing Kit (SQK-LSK109) (Oxford Nanopore Technologies, Oxford, UK) according to the manufacturer's instructions, except that the enzyme incubation time was twice as long as indicated in the instructions and the final AMPure purification incubation lasted 10 min at 37°C. Approximately 30–50 fmol of the library was loaded onto the flow cell (FLO-MIN106D, R9.4.1) of a GridION sequencer (Oxford Nanopore Technologies). Using the adaptive sampling option of GridION Mk1 (Loose, M., Malla, S. & Stout, M. 2016), we enriched target regions comprising 0.2% of the whole genome using a bed file that specified 59 loci associated with repeat expansion diseases and their surrounding 100 kb regions, as well as the SCA31-linked SNP in the PLEKHG4 gene and its surrounding 40 kb region (Table 3) (Miyatake, S. et al. 2022). For technical validation of time-lag sampling, Cas9-mediated PCR-free enrichment was performed targeting the RFC1 repeat locus as previously reported (Fukuda, H. et al. 2021; Mizuguchi, T. et al. 2021; Nakamura, H. et al. 2020) according to the manufacturer's protocol (Nanopore protocol for Cas9-targeted sequencing, ENR_9084_v109_revO_04Dec2018, Oxford Nanopore Technologies).A mixture of four Alt-R® crRNAs (5'-GACAGUAACUGUACCACAAU-3' (SEQ ID NO: 52) and 5'-ACCACUAGCCAAUGCCUGUU-3' (SEQ ID NO: 53) for the positive strand; 5'-CUAUAUUCGUGGAACUAUCU-3' (SEQ ID NO: 54) and 5'-UAGGACAUUCGGAAAUUCUU-3' (SEQ ID NO: 55) for the negative strand) was used.
[0066] [Table 3-1]
[0067] [Table 3-2]
[0068] Sequencing was performed in hac mode for approximately 2 to 3 days. Once or twice during sequencing, the flow cell was nuclease flushed using a Flow Cell Wash Kit (EXP-WSH004) (Oxford Nanopore Technologies) and additional libraries were loaded. For time-lag sampling, a method for sequencing multiple samples using a single flow cell, sample 1 was loaded onto the GridION on the first day, and sample 2 was loaded after nuclease flushing on the second and third days. For time-lag sampling using Cas9-mediated PCR-free enrichment, sample 1 was loaded onto the GridION on the first day and sequenced for 7.5 hours. After the nuclease flush, sample 2 was loaded and sequenced for 24 hours.
[0069] 4. Data Analysis A data analysis pipeline was developed using the analysis tools described below, and the output data from the pipeline was evaluated according to the data evaluation workflow shown in Figure 1. Sequences were base called using Guppy v4.3.4, v5.0.11, or v5.1.13 in hac mode during the GridION run. Sequences were re-called using Guppy v6.0.6 in super accuracy (sup) mode for patients 4, 5, 6, 10, 11, 12, 14, and 16, because the super accuracy (sup) mode increases the accuracy of base calling according to Oxford Nanopore. Sequences were then aligned to the reference genome GRCh38 using either minimap2 v2.14 (https: / / github.com / lh3 / minimap2) or LAST 1132 (https: / / gitlab.com / mcfrith / last). Coverage depth was calculated using mosdepth v0.3.1 (https: / / github.com / brentp / mosdepth). The median and range of coverage depth for all 59 loci in 22 patients were calculated (Fig. 4). tandem-genotypes v1.3.0, v1.8.2, or v1.9.0 (https: / / github.com / mcfrith / tandem-genotypes) was used to detect tandem repeat length changes at selected loci. The tandem-genotypes-plot command was used to generate histograms of the output. In the histograms in Figs. 7–9 and 11, the x-axis “copy number change” indicates the difference from the number of repeats defined in the reference human genome, and the gray bars (red in the color output figures) indicate forward strand reads, and the dark gray bars (blue in the color output figures) indicate reverse strand reads. The -o2 option of tandem-genotypes was used to estimate the approximate number of repeats in each repeat expansion sequence present in the two alleles (data not shown).To rank the detected repeats in order of likelihood of pathological repeat expansion, tandem-genotypes-join was performed using data from 27 control subjects sequenced using PromethION (Oxford Nanopore Technologies). The 27 controls were healthy Japanese subjects without neurodegenerative diseases. Figure 5 shows the distribution of repeat unit numbers in 54 alleles from 27 control samples for 46 loci that are particularly associated with known repeat expansion diseases with Mendelian inheritance among the 59 target loci. Consensus sequences were constructed using lamassemble 1.4.2 (https: / / gitlab.com / mcfrith / lamassemble) or the tandem-genotypes merge option implemented in version 1.8.0 and later. When pathological repeat expansions were detected, additional analyses, including detailed repeat analysis and methylation analysis, were performed (ad hoc). Detailed repeat analysis, including the creation of waterfall plots from consensus fasta files of the regions of interest or coverage plots of mapped reads, was performed using RepeatAnalysisTools (https: / / github.com / PacificBiosciences / apps-scripts / tree / master / RepeatAnalysisTools). When abnormal methylation was suspected, methylation analysis was performed as previously reported (Fukuda, H. et al. 2021). However, the base calling for detecting 5-methylcytosine was performed using the Guppy v6.0.6 base caller with a config file named “dna_r9.4.1_450bps_modbases_5mc_hac.cfg”, and the downstream analysis using the methylstat-methylcall-ontsbisul program shown in Figure 3 was performed with the default settings.
[0070] To evaluate the accuracy and efficiency of target enrichment, reads of 1000 bp or less that were determined to be "off-target" by GridION and whose sequencing was interrupted by the principle of adaptive sampling were removed using Filtlong v0.2.1 (https: / / github.com / rrwick / Filtlong) to generate fastq files that passed the quality filter. Then, the coverage depth for both on-target and off-target regions in all selected reads was calculated using mosdepth. For on-target, the average "per-locus depth of coverage" was calculated for all loci. For off-target, a bam file was generated using the samtools view -L option, including reads outside the target region assigned by the bed file, and the coverage depth of each region included in the bam file was calculated using mosdepth by dividing the region into 5000 bp segments. The enrichment efficiency was evaluated by comparing the average coverage depth data for on-target and off-target.
[0071] For downsampling of fastq data, seqkit v0.11.0 was used with the “sample” command, and various ratios of downsampling were specified with the -p option to set the virtual depth of the target region to 4–5×, 7–8×, 10×, 14–15×, and 20–25× ( https: / / bioinf.shenwei.me / seqkit / usage / #sample ). Downsampling was performed on the fastq data of patients 10 and 18.
[0072] Correlation analysis between conventional and GridION T-LRS Repeat lengths detected by either conventional methods (fragment analysis, Sanger sequencing, or Southern blotting) were compared with those detected by T-LRS, and correlations between the two were statistically analyzed by linear regression model using GraphPad Prism v8.0.2 (GraphPad Software, San Diego, CA, USA) for patients 1, 3, 7, 9, 13, 15, 18, and 19. A P value of <0.05 was considered statistically significant.
[0073] result Assessing sequence quality For all 22 patients, 59 target loci were captured with an average coverage depth of 24.7 (Figure 4). Figure 6A shows an example of the capture of the RFC1 locus. The average depth of coverage for all 59 target loci was homogeneous except for the NOTCH2NLC gene, which may be due to the presence of the locus in a segmental duplication region or the presence of paralogous genes such as NOTCH2NLA, NOTCH2NLB, and NOTCH2NLR (Figure 4). For a more detailed evaluation, we plotted the coverage depth for each chromosome obtained in a single run for two patients, patient 8, who had a relatively high depth, and patient 5, who had a lower depth. In these samples, relatively uniform coverage for on-target reads was reproduced. Although several off-target loci were observed in common between the two patients with relatively high coverage depths, off-target reads were generally low across all chromosomes (Figure 6B). Visual inspection of the IGV output revealed that the majority of these off-target regions did not contain coding genes and were located within repeat regions or at centromeres (data not shown). Although there were several highly covered off-targets, the target regions appeared to be highly accurately enriched overall, as the average coverage depth per locus for on-targets (46.86× in patient 8 and 12.97× in patient 5) was approximately 1000-fold higher than that for off-targets (0.041× in patient 8 and 0.0087× in patient 5) across all selected reads (Figure 6C).
[0074] Repeat expansions precisely identified in validation studies All pathogenic repeat expansions, regardless of repeat sequence and length, were detected in 12 positive controls (patients 1–12, Table 4 and Figure 7) and ranked first in 10 patients and second in 2 patients by our prioritization workflow (Figure 3). The loci ranked as pathogenic in each patient in the validation and discovery studies are shown in Tables 5 and 6, respectively. In patients 3 and 4, the pathogenic locus was ranked second, while in patient 3, the TNRC6A gene repeat expansion polymorphism and in patient 4, the heterozygous RFC1 gene repeat expansion were ranked first. In patient 3, the repeat expansion of the TNRC6A gene was determined to be polymorphic by inspection of the consensus sequence constructed by our workflow (Figure 8). In patient 4, the repeat expansion of the RFC1 gene was heterozygous, so the patient was considered a carrier by calculating the number of repeat units of each of the two alleles using tandem-genotypes (Table 5).
[0075] [Table 4-1]
[0076] [Table 4-2]
[0077] [Table 5-1]
[0078] [Table 5-2]
[0079] [Table 5-3]
[0080] [Table 6-1]
[0081]
Table 6-2
[0082] A single targeted long-read sequencing (T-LRS) analysis provided comprehensive results including the entire repeat sequence, as well as the length / number of expanded repeats and their distribution. Conventional methods provide only partial information, such as the repeat sequence specific to one repeat locus (RP-PCR), the length of the expanded repeat (Southern blotting) or the number of repeats (fragment length analysis), or only imply that the locus is pathogenic without detailed data (flanking PCR). Furthermore, T-LRS analysis at nucleotide resolution provided detailed information about interrupted sequences that are different from the pathogenic motif near the pathogenic repeat, which may function as disease modifiers or markers. In patient 1, the expanded CAG repeat has one additional (CAACAG) at the end, indicating that this patient will have an average disease course and severity. The presence or absence of the CAACAG sequence affects the age of onset and severity (Wright, GEB et al. 2019) and may be a prognostic marker. In patient 2, a CGG associated with intergenerational instability was identified in the mutated allele. In patient 4, a CTA-ATG sequence was inserted at the 5' end of the expanded repeat. This sequence remained after rebase calling of the sequence in sup mode. Various interrupting sequences, such as CCG, CGG, CAG, and CTC, have been reported at the 3' end and, more rarely, at the 5' end of the DMPK gene repeat expansion with an estimated frequency of 3-5% and were associated with a milder phenotype (Peric, S., Pesovic, J., Savic-Pavicevic, D., Rakocevic Stojanovic, V. & Meola, G. 2021). Although the possibility of a sequencing error could not be excluded, the CTA-ATG sequence found in the patient is a previously unreported interrupting sequence and its clinical significance is unknown. In patient 5, a pathological CTG expansion was observed in both alleles, indicating that this patient has a biallelic expansion. The waterfall plot for this patient showed full repeat sequence content, including benign CTA repeats and pathogenic CTG repeats.In conventional methods, it was difficult to clarify the sequence content of all repeats, and their pathogenicity was determined based on the total repeat length, so even if the benign CTA repeat was relatively long and the pathological repeat was relatively short, it was judged to be a pathological repeat when viewed as a total repeat length, even if it appeared to be pathologically expanded. Our T-LRS can overcome this difficulty. In patient 6, the normal allele was interrupted by GGA, which is thought to reduce GGC repeat instability (Fukuda, H. et al. 2021). In patient 7, the color output waterfall plot clearly showed not only the repeat length abnormality but also the detailed sequence content of GCX (X is A / T / G / C). In patient 9, a homozygous AAGGG repeat expansion was detected. In patient 10, the pathological TGGAA, polymorphic TAAAA (a sequence commonly found in all ethnicities), and TAGAA (a sequence commonly found in Japanese) repeat sequences were all confirmed (Figure 7).
[0083] Methylation analysis was performed on patients 4, 6, and individual 1. Patient 4 was diagnosed with a mild form of adult-onset myotonic dystrophy with a relatively short repeat expansion (approximately 100 repeats). Based on a recent paper (Morales, F. et al. 2021) which reported that methylation abnormalities are mostly observed in congenital forms of myotonic dystrophy and that patients with longer expansion alleles are more likely to show methylation abnormalities, the expansion repeat in this patient would not be hypermethylated. Patient 6 was diagnosed with neuronal intranuclear inclusion disease (NIID), so the expansion repeat is expected to be unmethylated. Individual 1 is an unaffected father with an extremely long hypermethylated repeat expansion in the NOTCH2NLC gene (Fukuda, H. et al. 2021). As expected, the pathogenic repeat expansions in patients 4 and 6 were unmethylated, whereas the extremely long repeat expansion in asymptomatic individual 1 was hypermethylated (data not shown).
[0084] Repeat expansions identified in previously undiagnosed patients We tested whether our method could detect pathogenic repeat expansions in patients without a molecular diagnosis. We tested 10 patients without a molecular diagnosis who were clinically diagnosed with spinocerebellar ataxia (SCA) (n = 8) or cerebellar ataxia, neuropathy, vestibular areflexia syndrome (CANVAS) (n = 2). We divided the investigators into two groups and assigned all 10 patients to either the conventional analysis group or the T-LRS analysis group. The investigators assigned to each group were blinded to the results obtained from the other analysis group. Both the conventional and T-LRS methods provided molecular diagnoses in 6 of the 10 patients. In two patients (patients 13 and 17, see below), the results differed between the conventional and T-LRS methods. In both patients, the conventional diagnosis was revised and the T-LRS results were deemed correct (Table 4 and Figures 9A-C). In the remaining patients, the results were concordant.
[0085] Case studies patient 13 CAG expansion in the CACNA1A gene, which consists of 20-32 repeats, causes spinocerebellar ataxia type 6 (SCA6) (Zhuchenko, O. et al. 1997). Since the normal limit of repeat units is 18, close to the abnormal threshold of 20 repeat units, it can be difficult to diagnose SCA6 using only conventional methods such as flanking PCR.
[0086] As a conventional approach, flanking PCR targeting 11 cerebellar ataxia-associated loci was performed (Figure 2). The first screening was negative, so the patient was diagnosed as not having a pathological repeat expansion. T-LRS called the CACNA1A gene linked to SCA6 as the ranked 1st locus, which was consistent with the patient's phenotype. According to the data evaluation flow, this locus was judged to be pathogenic because 1) an abnormal number of repeat units was detected (21 repeat units), 2) it is known to follow an autosomal dominant inheritance pattern, and 3) it is known to have no benign sequence expansions. To confirm this, flanking PCR of the CACNA1A locus was performed again and the size of the PCR amplicon was verified by gel electrophoresis. As a result, it was confirmed that a normal allele with a normal upper limit size was separated from a slightly larger allele (pathological expansion allele). Then, 21 CAG repeat units were confirmed by fragment analysis (Figure 9A). Looking back at this case, it seems likely that this case could have been diagnosed using conventional methods if flanking PCR had been performed carefully at the outset, or if fragment analysis had been performed on all cases without prior screening by flanking PCR. However, this is difficult to do without great care or expertise in repeat expansion diseases, and this case demonstrated the clear advantage of T-LRS in that it is possible to detect subtle changes in repeat number that would have been overlooked by low-resolution conventional methods, even without such special care or expertise.
[0087] patient 17 Repeat expansions in the bidirectionally transcribed ATXN8OS / ATXN8 genes (where the repeat motif is CTG in the ATXN8OS gene and complementary CAG in the ATXN8 gene) result in mRNAs and polyglutamine proteins with expanded CUG repeats (Koob, MD et al. 1999; Moseley, ML et al. 2006) that cause SCA8. Reduced penetrance occurs in SCA8, but interruption of the pathogenic repeat expansion by may be a modifier of that penetrance (Perez, BA et al. 2021). Normal alleles usually have 15-50 repeats (in ATXN8ON, the polymorphic CTA preceded by the pathogenic CTG) and pathogenic alleles have 71-1300 repeats (Hu, Y. et al. 2017).
[0088] A large amplicon was detected by flanking PCR surrounding the SCA8 locus, and conventional methods indicated that this patient had SCA8. T-LRS also called this locus as the number one ranked locus. However, the following two points suggested that its pathogenicity was negative: the number of expanded repeats was 66, which is a so-called intermediate type that slightly deviates from the normal range (a repeat number between full expansion and the normal range, which may not cause disease); and when the entire sequence of the expanded repeat was confirmed, the number of CTA repeats, which is a benign polymorphic sequence, was relatively long at 19, whereas previous data showed that it was usually 8-15 (hence the relatively pathological CTG The second most interesting locus, the AXTN7 gene linked to SCA7, was also excluded as pathogenic because its repeat length did not exceed the threshold for pathogenic repeat length. Therefore, the patient was determined not to have a pathogenic repeat expansion by T-LRS. After comparing the results of the two methods, we finally diagnosed the patient as having "intermediate SCA8 repeat expansion but questionable pathogenicity." Although the repeat length generally remained within the intermediate range and the benign repeat sequence was assumed to be longer, suggesting a low possibility of disease onset, we cannot completely exclude the possibility of SCA8 having intermediate repeat expansion. This case demonstrated the advantage of T-LRS when both the repeat number and repeat sequence are important for pathogenicity determination (Figure 9B).
[0089] Patients 14 and 16 SCA31 is relatively common in Japan. The patient has the following mutation in the intron region of the BEAN1 and TK2 genes: (TGGAA) n , (TAGAA) n , (TAAAA) n Or (TAAAATAGAA) n A 2.5-3.8 kb long pentanucleotide repeat expansion consisting of the repeat TGGAA was found, but only TGGAA was linked to the disease (Ishiguro, T., Nagai, Y. & Ishikawa, K. 2021; Sato, N. et al. 2009). There is also a single nucleotide polymorphism (SNP) in the 5' untranslated region of the PLEKHG4 gene that is very strongly associated with the disease, and it was conventionally thought that a positive flanking PCR and disease-associated SNP would confirm the diagnosis of SCA31.
[0090] In patients 14 and 16, large amplicons were detected by flanking PCR linked to SCA31, and TGGAA repeats were detected by RP-PCR. Sanger sequencing confirmed the SCA31-linked SNPs in patient 14 and patient 16. The SCA31-linked BEAN1 gene was called as the first-ranked locus in both patients by T-LRS, and abnormal repeat numbers were detected (repeat lengths of 2756 and 2915 bp in patients 14 and 16, respectively, corresponding to approximately 551 and 583 repeats), and the TGGAA sequence (249 and 269 repeats in patients 14 and 16, respectively) was detected in the consensus repeat sequence, confirming SCA31. As the SCA31-linked SNPs were included in the target region of adaptive sampling, genotyping information for both patients was obtained by checking the SNPs in IGV without additional experiments. In the case of this case, multiple tests were required for diagnosis using conventional methods (flanking PCR > RP-PCR > Sanger sequencing), but T-LRS may be advantageous in that it can provide this information in a single sequence. Patient 14 also demonstrated that genotyping of SNPs linked to SCA31 cannot be used to exclude SCA31 (Figure 9C).
[0091] Sequencing Accuracy For nanopore sequencing, the accuracy of the sequence depends on the version of the library preparation kit used, the version of Guppy_basecaller used, and the base calling model. We used kit 109, and base calling was performed with Guppy v4.3.4, v5.0.11, or v5.1.13 using the base calling model in high accuracy (hac) mode. According to Oxford Nanopore, the accuracy of base calling using kit 109 is about 95% with Guppy v4.3.4 in hac mode, and 97.8% with Guppy v.5.0.11 or later in hac mode (https: / / nanoporetech.com / accuracy). When the sequences were re-base called with Guppy v6.0.6 in sup mode, the accuracy of base calling increased to 98.3%. Therefore, for some patients whose waterfall plots had many "other" sequences in the repeat sequences, we re-called them in sup mode because these sequences were probably erroneous sequences and could be removed by base calling in sup mode. Comparing the waterfall plots of the consensus sequences generated in hac mode with those in sup mode, sup mode improved the sequence accuracy and reduced the "other" sequences (Figure 10A). For the AAGGG repeat expansion in the RFC1 gene that gives rise to CANVAS, the sequence of patient 19 was re-called in sup mode. However, this resulted in more "other" sequences in the consensus sequences in a strand-specific manner. Visual inspection of the consensus sequences revealed that the majority of the "other" sequences were AAGG repeat sequences (Figure 10B).We previously encountered a phenomenon in which a CANVAS patient (patient A) with an AAGGG repeat expansion had a similar AAGG sequence in nanopore sequencing (Miyatake, S. et al. 2022). We sequenced patient A using high-accuracy long-read whole-genome sequencing (HiFi LR-WGS) with the PacBio Sequel II system (Pacific Biosciences, Menlo Park, CA, USA), and found that the AAGG repeats seen in nanopore sequencing were AAGGG by PacBio HiFi LR-WGS. Therefore, the “other” sequences in the waterfall plot of patient 19 are likely to be sequencing or base calling errors, as was seen in patient A (Figure 10C). Patient 9, another CANVAS patient with an AAGGG repeat expansion in the RFC1 gene, showed a “noisy” waterfall plot pattern similar to that of patient 19, which is also likely due to sequencing errors.
[0092] As an alternative approach to assessing sequencing accuracy, we examined the correlation between repeat lengths determined by conventional and T-LRS using data from patients 1, 3, 7, 9, 13, 15, 18, and 19. A significant correlation was observed between repeat lengths determined by conventional and T-LRS (P < 0.0001, r 2: 0.9822) (Figure 11 and Table 7). When the correlation was examined by dividing the relatively short repeat length (up to 150 bp) and the long repeat length, a significant correlation was observed between the short repeats (n = 9 alleles) (P < 0.0001, r2: 0.9940), but the correlation between the long repeats (n = 5 alleles) was not statistically significant. This is a reasonable result, since the longer the read, the more prone it is to errors. Alternatively, it may be partially due to the limited number of samples used for evaluation. Therefore, we increased the number of samples by adding data from previously reported samples (Miyatake, S. et al. 2022) and performed the data analysis again (n = 16 alleles), and the long repeats also reached a significant correlation (P < 0.0001, r2: 0.9436).
[0093] [Table 7]
[0094] Sensitivity and specificity of T-LRS as a diagnostic method Our repeat detection workflow outputs a prioritized list of loci in order of importance (i.e., the magnitude of copy number change in the patient) compared to loci in our 27 unaffected controls. This list does not tell the examiner which loci are pathogenic, but can be thought of as being ranked in order of likelihood of being a pathogenic repeat, and an efficient determination can be made by determining whether each repeat locus has a pathogenic repeat expansion in descending order of rank. If the examiner makes this determination in order of priority starting from the locus ranked 1, all pathogenic repeat expansions will be nominated to either 1st or 2nd place, making it possible to detect pathogenic repeat expansions easily and quickly.
[0095] As a diagnostic tool, it is important to provide sensitivity and specificity. However, this detection workflow does not call any locus as pathogenic, so specificity or sensitivity could not be calculated. As shown in the prioritized list of the top 20 loci of all patients in the validation study (Table 5), pathogenic repeat expansions were ranked first in 83.3% (10 / 12) of patients and first or second in 100% (12 / 12). This may be a proxy for the sensitivity of this detection workflow. No patients with multiple expanded repeats were detected in this study.
[0096] As a surrogate for specificity, we checked whether any SCA loci other than the true pathogenic repeat expansion were miscalled as false-positive pathogenic repeat expansions in six SCA patients in the validation study (patients 2, 3, 5, 9, 10, and 11). In these patients, all expanded repeats except for one true pathogenic expansion had been previously ruled out by conventional methods. As shown in Table 8, no patient had a miscalled SCA repeat expansion locus, and only true pathogenic loci were detected. Thus, the surrogate specificity was 100%.
[0097] [Table 8]
[0098] Reducing sequencing costs with time-lag sampling In Japan, it costs a minimum of 804 USD to perform one GridION run on one flow cell with two nuclease flushes and all the necessary reagents. The cost of the conventional method is approximately 6–26 USD depending on how many experiments are required. To reduce the cost of sequencing, we attempted “time-lag” sampling, using nuclease flush to sequence two different samples on one GridION flow cell (Fig. 12A, B). This reduced the cost by approximately half (452 USD per sample). We sequenced four CANVAS patients. Samples 2, 3, and 4 were previously sequenced on a HiFi LR-WGS (Miyatake, S. et al. 2022). Sample 1 was identical to patient 9. The two samples were loaded sequentially onto the same flow cell as described in Fig. 12A, B. The average coverage depth of the four samples was 15.0× (11.41–16.46×). In all samples, the pathogenic repeat locus (RFC1 gene) was accurately detected by time-lag sampling. The repeat sequence [(AAGGG) in samples 1 and 3] n / (AAGGG) n , and in sample 2 it is (ACAGG) n / (ACAGG) n , and in sample 4 it is (AAGGG). n / (ACAGG) n] and repeat length were consistent with previous results within 10% (Table 9). When sequencing two different libraries on the same flow cell, there is a risk that the first library loaded may remain on the flow cell after the nuclease flush and be mixed with the next library loaded. In this regard, Oxford Nanopore Technologies states that the wash removes 99.9% of the library, but this implies that some DNA may remain on the flow cell [Nanopore protocol Flow Cell Wash Kit (EXP-WSH004), version: WFC_9120_v1_revB_08Dec2020]. Assuming that 0.1% of the previous library is carried over, approximately 0.015× (15 × 0.001) of the time lag sampling depth may come from the remaining library loaded before the nuclease flush, but this can be ignored for practical purposes.
[0099] [Table 9]
[0100] Theoretically, the deeper the sequencing output (the more reads), the more likely it is that the output will contain sequence reads from residual libraries input before nuclease flush. To experimentally determine the carryover risk of residual libraries in our diagnostic system, we performed time-lag sampling of sample 1 (repeat unit AAGGG / AAGGG) and sample 2 (repeat unit ACAGG / ACAGG) with maximizing coverage depth by Cas9-mediated PCR-free enrichment of the RFC1 repeat locus (Figure 12C). To maximize the depth of sample 2, sample 1 was sequenced for 7.5 hours and sample 2 for 24 hours. The depths of samples 1 and 2 were 163.76× and 403.02×, respectively, and the number of reads containing the full length of the expanded repeat sequence was 94 and 375, respectively. Sample 2 had no AAGGG repeat expansions (repeat expansion sequences derived from sample 1) in the sequencing output (Figure 12C). Considering that our method using adaptive sampling does not go to such depths, we conclude that time-lag sampling can be used without any practical concerns about carryover of residual libraries loaded before the nuclease flush.
[0101] Consideration We have developed a method for diagnosing human repeat expansion diseases by performing real-time T-LRS using the adaptive sampling option on the GridION system. T-LRS is a time-saving and efficient method that can simultaneously screen a large number of disease-associated repeat loci of interest and obtain accurate and comprehensive data on the repeat sequence and repeat length / number. T-LRS can also be useful for genotyping disease-linked SNPs of interest to assist in diagnosis.
[0102] T-LRS has shown that it has much higher resolution than conventional methods and can better detect repeat expansions and determine intermediate-length expansions that have a low risk of disease development. This advantage improves the sensitivity and specificity of repeat disease diagnosis. Furthermore, recent studies on CAG expansions in several diseases, such as SCA and Huntington's disease, have shown that intermediate-length, non-completely pathological repeat expansions may be associated with susceptibility to other diseases other than the disease in question (Gall-Duncan, T., Sato, N., Yuen, RKC & Pearson, CE 2022; Tezenas du Montcel, S. et al. 2014). Such findings have also been reported in amyotrophic lateral sclerosis (Elden, AC et al. 2010; Conforti, FL et al. 2012), frontotemporal dementia (Fournier, C. et al. 2018), and Alzheimer's disease (Rosas, I. et al. 2020). Although there were no such patients in our limited cohort, to detect such findings broadly, analysis using T-LRS that can target multiple loci including those related to other diseases other than the disease of interest and has high enough resolution to examine the sequence in detail is desirable.
[0103] T-LRS can obtain specific pathogenic repeat unit sequences and their neighboring interrupting sequences. In many repeat expansion disease genes, such as FMR1 gene (Eichler, EE et al. 1994), ATXN1 gene (Chong, SS et al. 1995), HTT gene (Wright, GEB et al. 2019) (protective effect) and DMPK gene (Morales, F. et al. 2021; Braida, C. et al. 2010) (detrimental or protective effect), it has been recognized in recent years that the insertion of a sequence different from the pathogenic repeat sequence within the repeat expansion (interrupting sequence) can act as a modifier that can affect genetic instability and disease course protectively or detrimentally. Such data will be useful to predict the disease progression / prognosis of patients and to provide information to the patients' offspring for family planning. From these data, unknown pathological mechanisms and unexpected phenotype-genotype associations can also be inferred.
[0104] One advantage of our method is that it allows screening of as many targets as desired, regardless of GC content or expanded repeat length, and the output is very homogeneous with uniform coverage depth. A second advantage is that, unlike other target enrichment systems such as Cas9-based approaches (Giesselmann, P. et al. 2019), target loci can be specified without prior experimental preparation. Thus, newly identified loci can be added to the target region at any time without targeting optimization experiments. A third advantage is that methylation profiles can be obtained in parallel. Although the patients in this study were not suitable for this analysis, we have shown that methylation analysis can be added to our diagnostic workflow in the form of ad hoc analysis for patients in whom hypermethylation of expanded repeats and adjacent CpG islands is assumed to cause gene downregulation, thereby providing additional evidence to support the diagnosis (Chintalaphani, SR, Pineda, SS, Deveson, IW & Kumar, KR 2021).
[0105] We developed an analysis pipeline that automatically prioritizes the most pathogenic loci after sequencing, and outputs the most pathogenic loci. A data evaluation chart makes it easy to perform a diagnosis without specialized experience in repeat expansion diseases.
[0106] We also revealed a weakness of long-read sequencing in that the accuracy of base calling decreases for specific repeat unit sequences such as AAGGG. Regarding this issue, Tan et al. performed telomere sequencing on a nanopore sequencer using Guppy 5 and Bonito for base calling and reported that errors due to base calling occur widely in telomeric regions rich in repeat sequences. For example, (TT AGGG ) n 40-50% of repeat customers are (TTAAAA) n , (TT AAGG ) n , (TTAGAG) n , (TTGGGG) n , (CTTCTT) n or (CCCCTGG) n (Tan, K., Slevin, M., Meyerson, M. & Li, H. 2022). Tan et al. also reported that miscalls of other types of repeats occur in a strand-specific manner, and that these "other types of repeats" are not observed in the CHM13 reference genome or PacBio HiFi reads, indicating that the "other types of repeats" are artifacts of the nanopore sequencing or base calling process rather than true sequence changes (Tan, K., Slevin, M., Meyerson, M. & Li, H. 2022). Tan et al.'s results seem to be consistent with our results of sequencing errors in patient 19 with an AAGGG repeat expansion.
[0107] The key point of this system is that a sufficient depth of coverage is necessary to obtain reliable sequence data that can avoid errors in base calling and separate the two alleles for the diagnosis of autosomal recessive diseases. To estimate the recommended coverage depth of our method, we performed downsampling of fastq data of patients 10 and 18 at various ratios to obtain virtual depths of the target region of 4–5×, 7–8×, 10×, 14–15×, and 20–25× (Figure 13A, B). We concluded that a minimum depth of 10–15× is required to obtain two separated alleles. Although the amount of output data is highly dependent on the flow cell quality (Stevanovski, I. et al. 2022), one way to increase the coverage depth of a single run is to shear the high molecular weight DNA. To detect short tandem repeat or minisatellite expansions, it is desirable to shear DNA to approximately 40 kb, since the expected expansion size is a maximum of 20 kb (Gall-Duncan, T., Sato, N., Yuen, RKC & Pearson, CE 2022).
[0108] Although T-LRS is practical and cost-effective compared to whole genome sequencing using long-read sequencers, it is still relatively expensive compared to conventional methods. However, the various advantages of T-LRS over conventional methods are significant. Furthermore, we propose a time-lag sampling method to reduce the cost by approximately half.
[0109] In conclusion, we propose a diagnostic method for repeat expansion diseases that can meet the urgent need for rapid, accurate and comprehensive molecular diagnosis of repeat expansion diseases.
[0110] References Braida, C. et al. Variant CCG and GGC repeats within the CTG expansion dramatically modify mutational dynamics and likely contribute toward unusual symptoms in some myotonic dystrophy type 1 patients. Hum. Mol. Genet 19, 1399-1412 (2010). Chong, S. S. et al. Gametic and somatic tissue-specific heterogeneity of the expanded SCA1 CAG repeat in spinocerebellar ataxia type 1. Nat. Genet. 10, 344-350 (1995). Conforti, F. L. et al. Ataxin-1 and ataxin-2 intermediate-length PolyQ expansions in amyotrophic lateral sclerosis. Neurology 79, 2315-2320 (2012). Cortese, A. et al. Biallelic expansion of an intronic repeat in RFC1 is a common cause of late-onset ataxia. Nat. Genet. 51, 649-658 (2019). David, G. et al. Cloning of the SCA7 gene reveals a highly unstable CAG repeat expansion. Nat. Genet. 17, 65-70 (1997). Eichler, E. E. et al. Length of uninterrupted CGG repeats determines instability in the FMR1 gene. Nat. Genet. 8, 88-94 (1994). Elden, A. C. et al. Ataxin-2 intermediate-length polyglutamine expansions are associated with increased risk for ALS. Nature 466, 1069-1075 (2010). Fournier, C. et al. Interrupted CAG expansions in ATXN2 gene expand the genetic spectrum of frontotemporal dementias. Acta Neuropathol. Commun. 6, 41 (2018). Fukuda, H. et al. Father-to-offspring transmission of extremely long NOTCH2NLC repeat expansions with contractions: genetic and epigenetic profiling with longread sequencing. Clin. Epigenetics 13, 204 (2021). Gall-Duncan, T., Sato, N., Yuen, R. K. C. & Pearson, C. E. Advancing genomic technologies and clinical awareness accelerates discovery of disease-associated tandem repeat sequences. Genome Res. 32, 1-27 (2022). Giesselmann, P. et al. Analysis of short tandem repeat expansions and their methylation state with nanopore sequencing. Nat. Biotechnol. 37, 1478-1481 (2019). Holmes, S. E. et al. Expansion of a novel CAG trinucleotide repeat in the 5' region of PPP2R2B is associated with SCA12. Nat. Genet. 23, 391-392 (1999). Hu, Y. et al. Sequence configuration of spinocerebellar ataxia type 8 repeat expansions in a Japanese cohort of 797 ataxia subjects. J. Neurol. Sci. 382, 87-90 (2017). Ishige, T. et al. Pentanucleotide repeat-primed PCR for genetic diagnosis of spinocerebellar ataxia type 31. J. Hum. Genet. 57, 807-808 (2012). Ishiguro, T., Nagai, Y. & Ishikawa, K. Insight into spinocerebellar ataxia type 31 (SCA31) from Drosophila model. Front. Neurosci. 15, 648133 (2021). Ishikawa, K. et al. An autosomal dominant cerebellar ataxia linked to chromosome 16q22.1 is associated with a single-nucleotide substitution in the 5' untranslated region of the gene encoding a protein with spectrin repeat and Rho guanine-nucleotide exchange-factor domains. Am. J. Hum. Genet. 77, 280-296 (2005). Ishiura, H. et al. Expansions of intronic TTTCA and TTTTA repeats in benign adult familial myoclonic epilepsy. Nat. Genet. 50, 581-590 (2018). Kawaguchi, Y. et al. CAG expansions in a novel gene for Machado-Joseph disease at chromosome 14q32.1. Nat. Genet. 8, 221-228 (1994). Koide, R. et al. Unstable expansion of CAG repeat in hereditary dentatorubral-pallidoluysian atrophy (DRPLA). Nat. Genet. 6, 9-13 (1994). Koob, M. D. et al. An untranslated CTG expansion causes a novel form of spinocerebellar ataxia (SCA8). Nat. Genet. 21, 379-384 (1999). Loose, M., Malla, S. & Stout, M. Real-time selective sequencing using nanopore technology. Nat. Methods 13, 751-754 (2016). Matera, I. et al. PHOX2B mutations and polyalanine expansions correlate with the severity of the respiratory phenotype and associated symptoms in both congenital and late onset Central Hypoventilation syndrome. J. Med. Genet. 41, 373-380 (2004). Miyatake, S. et al. Repeat conformation heterogeneity in cerebellar ataxia, neuropathy, vestibular areflexia syndrome. Brain 145, 1139-1150 (2022). Mizuguchi, T. et al. Complete sequencing of expanded SAMD12 repeats by longread sequencing and Cas9-mediated enrichment. Brain 144, 1103-1117 (2021). Morales, F. et al. Myotonic dystrophy type 1 (DM1) clinical subtypes and CTCF site methylation status flanking the CTG expansion are mutant allele length-dependent. Hum. Mol. Genet. 31, 262-274 (2021). Moseley, M. L. et al. Bidirectional expression of CUG and CAG expansion transcripts and intranuclear polyglutamine inclusions in spinocerebellar ataxia type 8. Nat. Genet. 38, 758-769 (2006). Nakamura, H. et al. Long-read sequencing identifies the pathogenic nucleotide repeat expansion in RFC1 in a Japanese case of CANVAS. J. Hum. Genet. 65, 475-480 (2020). Nakamura, K. et al. SCA17, a novel autosomal dominant cerebellar ataxia caused by an expanded polyglutamine in TATA-binding protein. Hum. Mol. Genet 10, 1441-1448 (2001). Orr, H. T. et al. Expansion of an unstable trinucleotide CAG repeat in spinocerebellar ataxia type 1. Nat. Genet. 4, 221-226 (1993). Perez, B. A. et al. CCG*CGG interruptions in high-penetrance SCA8 families increase RAN translation and protein toxicity. EMBO Mol. Med. 13, e14095 (2021). Peric, S., Pesovic, J., Savic-Pavicevic, D., Rakocevic Stojanovic, V. & Meola, G. Molecular and clinical implications of variant repeats in myotonic dystrophy type 1. Int. J. Mol. Sci. 23, 354 (2021). Pulst, S. M. et al. Moderate expansion of a normally biallelic trinucleotide repeat in spinocerebellar ataxia type 2. Nat. Genet. 14, 269-276 (1996). Rosas, I. et al. Role for ATXN1, ATXN2, and HTT intermediate repeats in frontotemporal dementia and Alzheimer's disease. Neurobiol. Aging 87, 139.e131-139.e137 (2020). Sato, N. et al. Spinocerebellar ataxia type 31 is associated with "inserted" pentanucleotide repeats containing (TGGAA)n. Am. J. Hum. Genet. 85, 544-557 (2009). Sone, J. et al. Long-read sequencing identifies GGC repeat expansions in NOTCH2NLC associated with neuronal intranuclear inclusion disease. Nat. Genet. 51, 1215-1221 (2019). Stevanovski, I. et al. Comprehensive genetic diagnosis of tandem repeat expansion disorders with programmable targeted nanopore sequencing. Sci. Adv. 8, eabm5386 (2022). Tezenas du Montcel, S. et al. Modulation of the age at onset in spinocerebellar ataxia by CAG tracts in various genes. Brain 137, 2444-2455 (2014). Wright, G. E. B. et al. Length of uninterrupted CAG, independent of polyglutamine size, results in increased somatic instability, hastening onset of Huntington disease. Am. J. Hum. Genet. 104, 1116-1126 (2019). Zhuchenko, O. et al. Autosomal dominant cerebellar ataxia (SCA6) associated with small polyglutamine expansions in the alpha 1A-voltage-dependent calcium channel. Nat. Genet. 15, 62-69 (1997).
Claims
1. A method for detecting a repeat expansion disease, comprising the steps of: detecting a repeat expansion disease in a subject determined to have a pathogenic repeat; A library preparation step in which a library for sequencing is prepared from genomic DNA isolated from the subject. A sequencing step in which long-read sequencing of the library is performed by nanopore sequencing using multiple disease-associated repeat regions and their surrounding genomic regions as target regions to obtain sequence data for each target region of the subject's genome. A consensus sequence construction process in which a consensus sequence is constructed from the sequence data obtained in the sequencing process. A ranking step in which the repeat motif repetition numbers in the plurality of disease-associated repeat regions are detected as repeat numbers, and the repeat number data of each of the subject's genome is compared with the repeat number data of each of the disease-associated repeat regions in a group of healthy individuals to rank the disease-associated repeat regions in the subject's genome in order of likelihood that they are pathological repeat expansions. A determination step in which evaluation is started from the disease-associated repeat region ranked first according to steps S1 to S7 below, and whether or not the subject has a pathogenic repeat expansion is determined. S1: Check whether the repeat region contains an abnormal repeat expansion that exceeds the disease threshold. If it does, proceed to S2. If it does not, proceed to S7. S2: Confirm the inheritance type of the repeat expansion disorder associated with the repeat expansion. If the disorder is autosomal dominant or X-linked dominant and the patient is female, proceed to S3. If the disorder is autosomal recessive, X-linked recessive, or X-linked dominant and the patient is male, proceed to S5. S3: Check whether the repeat region is a region known to have a benign repeat expansion. If yes, proceed to S4. If no, proceed to decision 1. S4: Check whether the repeat expansion detected in the repeat region contains a pathological repeat motif. If it does, proceed to decision 1. If it does not, proceed to S7. S5: Check whether all alleles have the repeat expansion. If yes, proceed to S6. If no, proceed to S7. S6: Check whether the repeat region is the repeat region of the RFC1 locus. If it is the RFC1 locus, proceed to S4. If it is not the RFC1 locus, proceed to decision 1. S7: Check whether the evaluation of the disease-associated repeat region ranked second has been completed. If not, return to S1 and start the evaluation of the disease-associated repeat region ranked second. If completed, proceed to decision 2. Decision 1: The subject is determined to have a pathogenic repeat expansion. Decision 2: The subject is determined to have no known pathogenic repeat expansion.
2. The disease-associated repeat regions are the following repeat regions (1) to (46) (the position on each chromosome is shown by the position in the human reference genome GRCh38): (1) A repeat region associated with neuronal intranuclear inclusion disease (ILD) in the 5' untranslated region of the NOTCH2NLC gene locus (positions 149390802-149390842 on chromosome 1) (2) A repeat region associated with familial adult myoclonus epilepsy type 2 (positions 96197066-96197124 on chromosome 2) located in an intron of the STARD7 gene locus (3) A repeat region associated with syndactyly in the coding region of the HOXD13 gene locus (positions 176093058-176093103 on chromosome 2) (4) A repeat region associated with early infantile epileptic encephalopathy type 71 (positions 190880872-190880920 on chromosome 2) in the 5' untranslated region of the GLS locus (5) A repeat region associated with spinocerebellar ataxia type 7 in the coding region of the ATXN7 gene locus (positions 63912685-63912715 on chromosome 3) (6) A repeat region associated with myotonic dystrophy type 2, located in an intron of the CNBP gene locus (positions 129172576-129172656 on chromosome 3) (7) A repeat region associated with blepharoptosis, ptosis and epicanthal inversion syndrome in the coding region of the FOXL2 gene locus (positions 138946020-138946062 on chromosome 3) (8) A repeat region associated with familial adult myoclonus epilepsy type 4 (positions 183712187-183712226 on chromosome 3) located in an intron of the YEATS2 gene locus (9) A repeat region associated with Huntington's disease in the coding region of the HTT gene locus (positions 3074876-3074939 on chromosome 4) (10) A repeat region associated with cerebellar ataxia-neuropathy-vestibular areflexia syndrome (39348424-39348483 on chromosome 4) located in an intron of the RFC1 gene locus (11) A repeat region associated with congenital central hypoventilation syndrome (CCH) in the coding region of the PHOX2B gene locus (positions 41745971-41746031 on chromosome 4) (12) A repeat region associated with familial adult myoclonus epilepsy type 7 in an intron of the RAPGEF2 gene locus (positions 159342526-159342618 on chromosome 4) (13) A repeat region associated with familial adult myoclonus epilepsy type 3 (positions 10356339-10356411 on chromosome 5) located in an intron of the MARCHF6 gene locus (14) A repeat region associated with spinocerebellar ataxia type 12, located in an intron of the PPP2R2B gene locus (positions 146878728-146878758 on chromosome 5) (15) A repeat region associated with spinocerebellar ataxia type 1 in the coding region of the ATXN1 gene locus (positions 16327635-16327722 on chromosome 6) (16) A repeat region associated with cleidocranial dysplasia in the coding region of the RUNX2 gene locus (positions 45422750-45422801 on chromosome 6) (17) A repeat region associated with spinocerebellar ataxia type 17 in the coding region of the TBP gene locus (positions 170561907 to 170562021 on chromosome 6) (18) A repeat region associated with hand-foot-genital syndrome in the coding region of the HOXA13 gene locus (positions 27199924-27199966 on chromosome 7) (19) A repeat region associated with oculopharyngeal distal myopathy in the 5' untranslated region of the LRP12 gene locus (positions 104588970-104588999 on chromosome 8) (20) A repeat region associated with familial adult myoclonus epilepsy type 1 (positions 118366815-118366918 on chromosome 8) located in an intron of the SAMD12 gene locus (21) A repeat region associated with frontotemporal dementia / amyotrophic lateral sclerosis (positions 27573528-27573546 on chromosome 9) located in an intron of the C9orf72 gene locus (22) A repeat region associated with Friedreich's ataxia, located in an intron of the FXN gene (positions 69037286-69037304 on chromosome 9) (23) A repeat region associated with oculopharyngeal distal myopathy, located in the exons of the LOC642361 and NUTM2B-AS1 loci (positions 79826383-79826404 on chromosome 10) (24) A repeat region associated with dentatorubral-pallidoluysian atrophy (positions 6936716-6936773 on chromosome 12) in the coding region of the ATN1 gene locus (25) A repeat region associated with spinocerebellar ataxia type 2 in the coding region of the ATXN2 gene locus (positions 111598950-111599019 on chromosome 12) (26) A repeat region associated with spinocerebellar ataxia type 8, located in an exon of the ATXN8OS gene locus (positions 70139383-70139428 on chromosome 13) (27) A repeat region associated with holoprosencephaly type 5 in the coding region of the ZIC2 locus (positions 99985448-99985493 on chromosome 13) (28) A repeat region associated with oculopharyngeal muscular dystrophy in the coding region of the PABPN1 gene locus (positions 23321472-23321502 on chromosome 14) (29) A repeat region associated with spinocerebellar ataxia type 3 in the coding region of the ATXN3 gene locus (positions 92071010 to 92071040 on chromosome 14) (30) A repeat region associated with familial adult myoclonus epilepsy type 6 (positions 24613438 to 24613532 on chromosome 16) in an intron of the TNRC6A gene locus (31) A repeat region associated with spinocerebellar ataxia type 31, located in an intron of the BEAN1 gene locus (positions 66490396-66490466 on chromosome 16) (32) A repeat region associated with Huntington's disease type 2, located in an intron of the JPH3 locus (positions 87604287-87604329 on chromosome 16) (33) A repeat region associated with Fuchs endothelial corneal dystrophy type 3, located in an intron of the TCF4 gene locus (positions 55586153 to 55586229 on chromosome 18) (34) A repeat region associated with spinocerebellar ataxia type 6 in the coding region of the CACNA1A gene locus (positions 13207858-13207897 on chromosome 19) (35) A repeat region associated with oculopharyngeal distal myopathy in the 5' untranslated region of the GIPC1 gene locus (positions 14496041-14496075 on chromosome 19) (36) A repeat region associated with myotonic dystrophy type 1 in the 3' untranslated region of the DMPK gene locus (positions 45770204 to 45770264 on chromosome 19) (37) A repeat region associated with spinocerebellar ataxia type 36, located in an intron of the NOP56 gene locus (positions 2652733-2652757 on chromosome 20) (38) A repeat region associated with Unverricht-Lundborg disease / progressive myoclonus epilepsy type 1 in the promoter region of the CSTB locus (positions 43776443-43776479 on chromosome 21) (39) A repeat region associated with spinocerebellar ataxia type 10 in an intron of the ATXN10 gene locus (positions 45795354-45795424 on chromosome 22) (40) A repeat region associated with spinal-bulbar muscular atrophy (positions 67545317 to 67545386 on the X chromosome) in the coding region of the AR locus (41) A repeat region associated with early infantile epileptic encephalopathy type 1 (positions 25013649-25013697 on the X chromosome) in the coding region of the ARX gene locus (42) A repeat region in the coding region of the SOX3 gene locus associated with mental retardation associated with isolated growth hormone deficiency (positions 140504316 to 140504361 on the X chromosome) (43) A repeat region associated with fragile X-associated tremor / ataxia syndrome (positions 147912050-147912110 on the X chromosome) in the 5' untranslated region of the FMR1 locus (44) A repeat region associated with fragile XE syndrome in the 5' untranslated region of the AFF2 locus (positions 148500637-148500682 on the X chromosome) (45) A repeat region associated with spinocerebellar ataxia type 37 in an intron of the DAB1 locus (positions 57367043-57367125 on chromosome 1) (46) A repeat region associated with Baratela-Scott syndrome in the promoter region of the XYLT1 gene locus (positions 17470907-17470930 on chromosome 16) and the regions known to have benign repeat expansions are (2), (8), (12), (13), (20), (30), (31) and (45).
3. The method according to claim 1, wherein in the library preparation step, the library is prepared by shearing the genomic DNA to 35 to 45 kb.
4. The sequencing step comprises: Step 1: Specifying a target region and initiating sequencing of each single DNA molecule in the library; step 2, mapping the sequence of the 5'-side 300 to 700 bases of the single DNA molecule to a reference human genome in real time; and Step 3: Continue sequencing the single DNA molecule if it is mapped to the specified target region, and stop sequencing the single DNA molecule if it is mapped outside the specified target region, eject the single DNA molecule from the pore, start sequencing another single DNA molecule, and return to step 2. The method of claim 1 , comprising:
5. The method according to claim 1, wherein the target region consists of each disease-associated repeat region and genomic regions of 30 kb to 70 kb upstream and downstream thereof.
6. The method according to claim 1, wherein in the sequencing step, sequence data is obtained having an average read depth of 10x or more for the target region.
7. The method according to claim 1, comprising: in the library preparation step, preparing a plurality of libraries derived from a plurality of subjects; in the sequencing step, after the sequencing of one library is completed, washing the flow cell with nuclease, loading and sequencing the next library, thereby sequentially sequencing the plurality of libraries in one flow cell, and performing a ranking step and an evaluation step for each subject.