Novel assay for distal genomic locus phasing using conjugation analysis via long-read sequencing hybrid data analysis.
A hybrid analysis of genomic DNA and cDNA using long-read sequencing addresses the limitations of short-read sequencing by accurately phasing distal SNPs and STRs, enhancing diagnostic precision and therapeutic targeting for Huntington's disease.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- F HOFFMANN LA ROCHE & CO AG
- Filing Date
- 2024-03-28
- Publication Date
- 2026-04-23
AI Technical Summary
Current DNA sequencing methods, particularly short-read sequencing, struggle to accurately phase distal single nucleotide polymorphisms (SNPs) and short tandem repeats (STRs) due to high error rates, short read lengths, and complex genetic variants, leading to incomplete phasing information and labor-intensive, error-prone processes.
A hybrid analysis method using both genomic DNA and cDNA (mRNA) for spatial linkage, employing long-read sequencing to phase exon and intron SNPs and distal target SNPs through exon reference SNPs, utilizing oligonucleotide primers and reverse transcription to generate amplification products for precise haplotype determination.
Enables accurate and efficient phasing of distal SNPs and STRs, facilitating personalized treatment strategies for diseases like Huntington's disease by identifying patient-specific genetic markers, thereby improving diagnostic accuracy and therapeutic matching.
Smart Images

Figure 2026513264000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure provides, for example, novel molecular biology assays for preclinical or clinical biomarker characterization in the field of neuroscience (e.g., Huntington's disease, HD). The disclosed methods and kits may be used as companion diagnostic tools, and the identification of two or more paired loci is required to provide information essential for the safe and efficient stratification of patients receiving a particular therapy or drug treatment.
Background Art
[0002] Huntington's disease (HD) is a rare genetic disorder that causes progressive degeneration of brain cells. HD is caused by an abnormal expansion of short tandem repeats (STRs) in exon 1 of chromosome 4 of the huntingtin gene. These repeats consist of a sequence of cytosine-adenine-guanine (CAG) nucleotides. A normal huntingtin gene typically has 10 to 35 CAG repeats, but in individuals with Huntington's disease, the number of CAG repeats can range from 36 to over 100. CAG expansion can occur spontaneously or be inherited in an autosomal dominant manner. If a parent has a CAG expansion in their huntingtin gene, there is a 50% chance that they will pass it on to their offspring. Individuals with a CAG expansion (typically 40 or more repeat units) have a 100% risk of developing Huntington's disease at lifetime. As the number of CAG repeats increases, the age of symptom onset decreases and the severity of the disease increases. The exact mechanism by which CAG expansion causes Huntington's disease is not fully understood, but it is thought that expanded CAG repeats cause abnormal folding of the huntingtin protein, leading to the accumulation of toxic protein aggregates that damage brain cells and ultimately cause cell death. Irregular expansions of STR (such as CAG repeat expansion) are not unique to Huntington's disease (HD) and are associated with several other genetic disorders, including spinocerebellar ataxia, myotonic dystrophy, and some forms of muscular dystrophy (Hannan, 2018). Several SNPs have been identified within the huntingtin gene that alter the risk of developing Huntington's disease, as well as in relation to variations in age of onset and symptom severity (Claassen, DO et al., 2020; Becanovic, K. et al., 2015). For example, several SNPs within the huntingtin gene, known as rs13102260, rs362277, rs3025814, and rs2530596, have been associated with age of onset (Kartsaki, E., et al., 2006; Kay, C., et al., 2015; Ramos, E., et al., 2012).Furthermore, specific SNP alleles at a single polymorphic site (e.g., rs7685686, rs362331, rs6446723, rs6844859, rs363080s, rs363125, rs362307, rs362273) have been identified as enabling selective treatment of HD patients while simultaneously allowing for non-selective treatment of all other patients (Kay, C., et al. 2015; Shin et al. 2022). Such intronic and exonic SNPs can be targeted using antisense oligonucleotide (ASO)-based therapies, selectively targeting precursor mRNA (premRNA) or messenger RNA (mRNA) to alter the mRNA and thereby induce protein expression through various mechanisms.
[0003] For example, advances in long-read sequencing technologies, such as long-read sequencing platforms commercially available from Oxford Nanopore Technologies Ltd. (Oxford, UK) and Biosciences of California, Inc. (Menlo Park, USA), have enabled the accurate detection and characterization of SNPs in the huntingtin gene and other genes associated with Huntington's disease, as well as CAG repeat expansions. Identifying SNPs associated with Huntington's disease and related disorders is a critical area of research because it could enable the development of new approaches to disease management and treatment. For instance, SNPs may provide targets for the development of new drugs or therapies that can modulate the expression or activity of the huntingtin protein, or for the development of gene-editing approaches that can modify the huntingtin gene to prevent or reverse disease progression. Overall, the identification and characterization of SNPs associated with Huntington's disease is a crucial area of research with the potential to deepen our understanding of the disease and advance the development of new treatments and therapies. Huntington's disease manifests as the loss of GABAergic middle spine (GABAMS) neurons in the striatum, caused by an expansion of the CAG repeat in exon 1 of the huntingtin gene. Within cells, this wild-type protein may be involved in chemical signaling, substance transport, attachment (binding) to proteins and other structures, and protection of cells from self-destruction (apoptosis). Symptoms of Huntington's disease typically begin in middle age, but can appear earlier or later in life. Early symptoms include involuntary movements such as convulsions or monoconvulsions, and difficulty with coordination and balance. As the disease progresses, symptoms become more severe and may include cognitive impairment, mood swings, and behavioral changes. The time from the first symptom to death is often about 10–30 years. While there is no cure, treatments can alleviate symptoms, and support is available. The huntingtin protein is found in many body tissues (e.g., the liver) and has the highest levels of activity in the brain.Furthermore, genetic counseling and testing can help individuals and families understand their risk of developing diseases and make informed decisions about family planning.
[0004] DNA sequencing is a fundamental tool in biological and medical research and is particularly important for the paradigm of personalized medicine. Various new DNA sequencing methods have been studied with the aim of ultimately achieving the $1,000 genome goal; the dominant method is sequencing by synthesis (SBS), an approach that determines short DNA sequences during polymerase reactions (Slatko et al., 2018). PacBio sequencing, also known as SMRT (Single-Molecule Real-Time) sequencing, allows sequencing of very long fragments up to 30–50 kb. The SMRT method involves ligating an engineered DNA polymerase to the bottom of a Zero-Mode Waveguides (ZMW) well, and a DNA library ligated with an SMRT bell adapter is loaded onto the DNA polymerase. Four nucleotides are labeled with different phospho-linked fluorophores for differential detection. As nucleotides are incorporated into the growing strand, imaging is performed on a millisecond timescale because the correct fluorescently labeled nucleotide is incorporated into the complementary strain of the sequenced single-stranded DNA molecule. After each dNTP polymerization process, the phosphate-linked fluorescent moiety is released and dissipates from the detection region, becoming undetectable. The next nucleotide can then be incorporated. Hereinafter, imaging is timed to the rate of nucleotide incorporation so that each base is identified as it is incorporated into the growing DNA strand (Slatko et al., above). Nanopore-based DNA sequencing was first proposed in the late 1990s, and commercialization has recently been achieved by Oxford Nanopore Technologies Ltd. (Oxford, UK), where protein nanopores are embedded in an electrical resistance bilayer film, resulting in a characteristic change in current (picoamperes - pA) as each nucleotide passes through a detector that allows for short, long, and ultra-long read lengths from 0.1 kb to 1 Mb.More specifically, long dsDNA molecules are first ligated into motor enzymes and tether molecules, thereby precipitating the library onto a sequencing flow cell, increasing the proximity of molecules to the nanopores, and similarly increasing the amount of molecules available for analysis. When the guide strand containing the motor-complex encounters an available nanopore, a template of single-stranded DNA (ssDNA) enters the nanopore, disrupting the pA baseline readout via alteration. Herein, the translocation rate is regulated by the nucleotide sequence and its associated epigenetic modifications. The motor enzyme slows down the processability of DNA through the channel, allowing for improved raw data quality. Each nucleotide k-mer present in the nanopore is recorded in real time as a current disruption event (Slatko et al., above), subsequently providing a characteristic electronic pattern that calls the base to a standardized FASTQ / FASTA dataset.
[0005] Well-established short-read sequencing-by-synthesis (SBS) platforms (e.g., the MiSeq or NextSeq series by Illumina, Inc., San Diego, CA, USA, or Ion GeneStudio Systems by Thermo Fisher, Waltham, MA, USA) are routinely used for SNP genotyping, and several algorithms exist for resolving haplotypes based on SNP data. However, such approaches are limited in many ways because they are primarily designed for haplotyping of whole-genome assemblies and rely on existing population-based reference panels. This is a statistical approach and is not 100% accurate, and more complex genetic variants such as STRs are not included in the reference panel. Also, regions with little genetic variation can lead to the disruption of phasing information between distal loci. Novel sequencing platforms based on long-read technology (e.g., Oxford Nanopore Technologies Ltd., Oxford, UK, and Pacific Biosciences of California, Inc., Menlo Park, USA) may help directly resolve SNP phasing without statistical inference; however, the high error rate of Oxford Nanopore sequencing and the shorter average read length of PacBio (approximately 20kb) are limitations to deconvolution of distal regions >150kb at the single-nucleotide level in an accurate, low-cost, and rapid manner. To overcome the problem of phasing distal SNPs with CAG repeats, Asuragen, Inc. (Austin, TX, USA) is developing a companion diagnostic test that uses AmplideX® PCR technology to size and phase CAG repeats of HTT at two different SNPs targeted by Wave's WVE-120101 and WVE-120102 investigational therapy programs. However, those approaches that use RNA as a starting material to shorten the distance between CAG repeats and exon SNPs cannot be used for intron SNPs.An alternative method for phasing distal SNPs using CAG repeats is oligo-based hybridization capture of full-length genes via biotinylated probes aligned with the sequence of interest, followed by PCR amplification (e.g., available from Twist Biosciences, South San Francisco, US, e.g., no. 101341 Twist Custom Panel Plus or no. 102989 Twist Human Custom Comprehensive Exome). This method requires genomic DNA material as input and necessitates a concentration step following whole-genome amplification. However, extraction of very long genomic DNA fragments (e.g., 15–30 kb nt) remains challenging, and this approach can result in fragmentation of the concentrated material and loss of the desired long signal necessary for phasing distal loci, particularly in regions with fewer genetic mutations. Another alternative method for phasing two distal loci is provided in International Publication 2018 / 022473, which involves amplifying two genomic regions, each containing the locus of interest, using a primer set that yields sticky-end amplification products, and then subjecting them to a ligation procedure. In this specification, one type of nucleic acid is used as a template for determining the locus of interest (e.g., a chromosome or its fragment, genomic DNA, or mRNA / cDNA). Due to the random ligation of the amplification products, the resulting ligation products cannot be used directly for haplotyping. Therefore, the method disclosed in International Publication 2018 / 022473 requires the division of the nucleic acid template into droplets, resulting in only one nucleic acid molecular template per droplet. PCR amplification and sticky-end ligation of two different loci are performed in the droplets. Subsequently, a second amplification reaction of the properly ligated products is required before sequencing the amplified ligation products by next-generation sequencing. Therefore, this method is very labor-intensive, complex, and prone to fading errors and artifacts due to the many handling steps involved.International Publication No. 2016 / 191380 describes a general concept of phasing SNPs from long-read sequencing data using heterozygous SNPs as references for aligning sequences from one type of nucleic acid (i.e., DNA or RNA). This method relies on having several heterozygous SNPs within a locus when long distances need to be covered by the sequencing run. However, depending on the target sequence being analyzed and the location of the SNPs used for phasing, this may not always be possible.
[0006] Therefore, there is a need for novel methods to overcome such technical challenges associated with the need for high input, short read, custom oligonucleotide development, or high error profiles of genomic DNA. [Overview of the Initiative]
[0007] To address the shortcomings of current methods, this disclosure provides a hybrid analysis of long amplicons generated from genomic DNA and cDNA (mRNA) for distal information pairing via spatial linkage, i.e., phasing of two or more required loci. Such methods could bring about significant advances in applied diagnostics and are a prerequisite for allele-specific treatments for hereditary diseases such as Huntington's disease. Applying the hybrid approach disclosed herein using both DNA and cDNA (reverse transcribed from mature mRNA) as starting materials makes it possible to phase desired short tandem repeats or exon single nucleotide polymorphisms and distal target SNPs (introns or exons) via exon reference SNPs using only a very small number of amplicons (e.g., at least one DNA and at least one cDNA amplicon). In this specification, the method enables the analysis of heterozygous and homozygous exon or intron target SNPs. In particular, for optimal resolution of phasing information, if the target SNP is heterozygous, the exon reference SNP should be heterozygous.
[0008] The methods and kits are particularly applicable in the field of companion diagnostics, which involves identifying loci to provide crucial information about a patient's genomic state and can be used to safely and effectively match patients with specific treatments or drug therapies, thereby identifying and stratifying patients most likely to benefit from those treatments or therapies. Such applications have become increasingly important in recent years due to the development of more targeted therapies. Furthermore, the methods and kits of this disclosure can also be used in basic research that can help identify genetic variations associated with various diseases and conditions. This could lead to a better understanding of the underlying mechanisms of diseases and ultimately to the development of new therapies and treatments.
[0009] In one embodiment, a method for phasing at least one desired distal single nucleotide polymorphism and short tandem repeat within the same target locus of nucleic acid isolated from a biological sample, comprising: (a) carrying out an amplification step of contacting genomic DNA isolated from the biological sample containing the target locus with a first set of oligonucleotide primers to produce a first amplification product containing at least one desired single nucleotide polymorphism and at least one exon-reference single nucleotide polymorphism within the same locus; and (b) phasing mRNA isolated from the biological sample containing the target locus with a second set of oligonucleotide primers. A method is provided which includes: (a) performing a reverse transcription and amplification step, which involves contacting a set of to produce a second amplification product containing a short tandem repeat and the at least one exon-referenced single nucleotide polymorphism; (c) determining the nucleic acid sequences of the first and second amplification products; (d) aligning the nucleic acid sequences of the first and second amplification products determined in step (c) based on the location of the at least one exon-referenced single nucleotide polymorphism; and (e) determining the haplotype of the at least one distal single nucleotide polymorphism of interest and the short tandem repeat in the sample. In one embodiment, the at least one distal single nucleotide polymorphism of interest is an intron single nucleotide polymorphism. In a particular embodiment, the intron having the at least one distal single nucleotide polymorphism of interest is located adjacent to an exon having the at least one exon-referenced single nucleotide polymorphism within the same locus of isolated genomic DNA. In one embodiment, the amplification step (a) and the reverse transcription and amplification step (b) are carried out in separate reaction vessels. In another embodiment, the nucleic acid sequences of the first and second amplification products in step (c) are determined using long-read sequencing, such as PacBio sequencing, also known as SMRT (Single Molecule Real Time) sequencing, or nanopore-based DNA long-read sequencing, commercially available from Oxford Nanopore Technologies. In another embodiment, the biological sample is selected from the group consisting of tissues, tumor tissues, blood, saliva, and cell lines derived from organisms.In one embodiment, when aligning the nucleic acid sequences of the first and second amplification products in step (d), if at least one distal single nucleotide polymorphism of interest is heterozygous, then at least one exon-reference single nucleotide polymorphism is heterozygous. In another embodiment, before aligning the nucleic acid sequences of the first and second amplification products in step (d), the method further includes aligning the nucleic acid sequences of the first and second amplification products to the nucleic acid sequence of a reference gene at the target locus, or to the nucleic acid sequence of the complete human genome or a portion thereof containing the reference gene at the target locus. In another embodiment, the target locus is the huntingtin gene. In certain embodiments, the short tandem repeat is a CAG repeat. In some embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is an exon SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, at least one target single nucleotide polymorphism is the intron SNP rs7685686. In some embodiments, at least one exon reference single nucleotide polymorphism is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099.In this specification, an exon-reference single nucleotide polymorphism (e.g., rs362331, rs362273, rs34315806, and rs363099) shall not be identical to (i.e., not corresponding to) at least one single nucleotide polymorphism of interest. In some embodiments, at least one exon-reference single nucleotide polymorphism is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099, if at least one single nucleotide polymorphism of interest is an intronic SNP. In a particular embodiment, at least one exon-reference single nucleotide polymorphism is rs362331. In a particular embodiment, determining the haplotype of a short tandem repeat includes determining the number of short tandem repeat units. In a particular embodiment, mRNA is spliced mRNA (mature mRNA).
[0010] In another embodiment, an in vitro method for diagnosing whether an individual is at risk of developing Huntington's disease, comprising: (a) carrying out an amplification step comprising contacting genomic DNA isolated from the biological sample containing a target locus with a first set of oligonucleotide primers to produce a first amplification product containing at least one target single nucleotide polymorphism and at least one exon-referenced single nucleotide polymorphism within the same locus; carrying out a reverse transcription and amplification step comprising contacting mRNA isolated from the biological sample containing the target locus with a second set of oligonucleotide primers to produce a second amplification product containing a short tandem repeat and the at least one exon-referenced single nucleotide polymorphism; determining the nucleic acid sequences of the first and second amplification products; aligning the nucleic acid sequences of the first and second amplification products determined in the step based on the position of at least one exon-referenced single nucleotide polymorphism; and determining the haplotype of at least one target distal single nucleotide polymorphism and short tandem repeat in the sample, thereby obtaining in vitro results from the individual. An in vitro method is provided which includes (b) determining the haplotype of at least one target distal single nucleotide polymorphism and short tandem repeat in a vitro sample, and (b) determining the risk of an individual developing Huntington's disease based on the haplotype of at least one target distal single nucleotide polymorphism and CAG short tandem repeat determined in step (a). In one embodiment, the at least one target distal single nucleotide polymorphism is an intron single nucleotide polymorphism. In a particular embodiment, the intron having the at least one target distal single nucleotide polymorphism is located adjacent to an exon having at least one exon reference single nucleotide polymorphism within the same locus of isolated genomic DNA. In one embodiment, the amplification step and the reverse transcription and amplification step are carried out in separate reaction vessels.In another embodiment, the nucleic acid sequences of the first and second amplification products are determined using long-read sequencing, such as PacBio sequencing, also known as SMRT (Single Molecule Real Time) sequencing, or nanopore-based DNA long-read sequencing commercially available from Oxford Nanopore Technologies. In another embodiment, the biological sample is selected from the group consisting of tissues, tumor tissues, blood, saliva, and cell lines derived from an individual. In one embodiment, when aligning the nucleic acid sequences of the first and second amplification products, at least one exon-reference single nucleotide polymorphism is heterozygous if at least one distal single nucleotide polymorphism of interest is heterozygous. In another embodiment, before aligning the nucleic acid sequences of the first and second amplification products, the method further includes the step of aligning the nucleic acid sequences of the first and second amplification products to the nucleic acid sequence of a reference gene at the target locus, or to the nucleic acid sequence of the complete human genome or a portion thereof containing the reference gene at the target locus. In another embodiment, the target locus is the huntingtin gene. In certain embodiments, the short tandem repeat is a CAG repeat. In certain embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is an exon SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, at least one target single nucleotide polymorphism is the intron SNP rs7685686. In some embodiments, at least one exon reference single nucleotide polymorphism is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099.In this specification, an exon-reference single nucleotide polymorphism (e.g., rs362331, rs362273, rs34315806, and rs363099) shall not be identical to (i.e., not corresponding to) at least one single nucleotide polymorphism of interest. In some embodiments, at least one exon-reference single nucleotide polymorphism is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099, if at least one single nucleotide polymorphism of interest is an intronic SNP. In a particular embodiment, at least one exon-reference single nucleotide polymorphism is rs362331. In a particular embodiment, determining the haplotype of a short tandem repeat includes determining the number of short tandem repeat units. In a particular embodiment, mRNA is spliced mRNA (mature mRNA).
[0011] Identifying patients with Huntington's disease who are likely to respond to therapy targeting at least one target single nucleotide polymorphism in A vitro method comprising: (a) performing an amplification step, which includes contacting genomic DNA isolated from the biological sample containing the target gene locus with a first set of oligonucleotide primers to produce a first amplification product containing at least one target single nucleotide polymorphism and at least one exon-referenced single nucleotide polymorphism within the same locus; performing a reverse transcription and amplification step, which includes contacting mRNA isolated from the biological sample containing the target gene locus with a second set of oligonucleotide primers to produce a second amplification product containing a short tandem repeat and the at least one exon-referenced single nucleotide polymorphism; determining the nucleic acid sequences of the first and second amplification products; aligning the nucleic acid sequences of the first and second amplification products determined in the above step based on the position of at least one exon-referenced single nucleotide polymorphism; and determining the haplotype of at least one target distal single nucleotide polymorphism and short tandem repeat in the sample, thereby obtaining in from the individual. An in vitro method is provided which includes (b) determining the haplotype of at least one target distal single nucleotide polymorphism and short tandem repeat in a vitro sample, and (b) identifying a patient as likely to respond to therapy based on the haplotype determined in step (a) and the determination of the presence of a specific allele of at least one target single nucleotide polymorphism. In some embodiments, the therapy targeting at least one target single nucleotide polymorphism is an antisense oligonucleotide treatment that targets the suppression of an RNA molecule containing a specific allele of at least one target single nucleotide polymorphism. In one embodiment, the at least one target distal single nucleotide polymorphism is an intron single nucleotide polymorphism. In certain embodiments, the intron having at least one target distal single nucleotide polymorphism is located adjacent to an exon having at least one exon reference single nucleotide polymorphism within the same locus of isolated genomic DNA. In one embodiment, the amplification step (a) and the reverse transcription and amplification step (b) are carried out in separate reaction vessels.In another embodiment, the nucleic acid sequences of the first and second amplification products in step (c) are determined using long-read sequencing, such as PacBio sequencing, also known as SMRT (Single Molecule Real Time) sequencing, or nanopore-based DNA long-read sequencing, commercially available from Oxford Nanopore Technologies. In another embodiment, the biological sample is selected from the group consisting of tissues, tumor tissues, blood, saliva, and cell lines derived from an individual. In one embodiment, when aligning the nucleic acid sequences of the first and second amplification products, at least one exon-reference single nucleotide polymorphism is heterozygous if at least one distal single nucleotide polymorphism of interest is heterozygous. In another embodiment, before aligning the nucleic acid sequences of the first and second amplification products, the method further includes the step of aligning the nucleic acid sequences of the first and second amplification products to the nucleic acid sequence of a reference gene at a target locus, or the nucleic acid sequence of the complete human genome or a portion thereof containing the reference gene at a target locus. In another embodiment, the target locus is the huntingtin gene. In a particular embodiment, the short tandem repeat is a CAG repeat. In some embodiments, at least one single nucleotide polymorphism of interest is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In a particular embodiment, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804.In certain embodiments, at least one target single nucleotide polymorphism is an exon SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, at least one target single nucleotide polymorphism is the intron SNP rs7685686. In some embodiments, at least one exon reference single nucleotide polymorphism is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In this specification, an exon reference single nucleotide polymorphism (e.g., rs362331, rs362273, rs34315806, and rs363099) shall not be identical to (i.e., not corresponding to) at least one target single nucleotide polymorphism. In some embodiments, at least one exon-reference single nucleotide polymorphism (SNP) is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099, if at least one SNP of interest is an intronic SNP. In a particular embodiment, at least one exon-reference SNP is rs362331. In a particular embodiment, determining the haplotype of a short tandem repeat includes determining the number of short tandem repeat units. In a particular embodiment, mRNA is spliced mRNA (mature mRNA).
[0012] In another embodiment, a kit is provided for determining the nucleic acid sequences of at least one target distal single nucleotide polymorphism and short tandem repeats within the same target locus of nucleic acids isolated from a biological sample, comprising: a first set of oligonucleotide primers for generating a first amplification product containing at least one target single nucleotide polymorphism and at least one exon-reference single nucleotide polymorphism within the same locus; and a second set of oligonucleotide primers for generating a second amplification product containing at least one target exon-reference single nucleotide polymorphism and the at least one exon-reference single nucleotide polymorphism. In this specification, the kit is adapted to carry out any of the methods disclosed herein. In a particular embodiment, the first set of oligonucleotide primers comprises an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 1 and an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 2. In another particular embodiment, the second set of oligonucleotide primers comprises an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 3, an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 4 and an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 5. In some embodiments, the kit further comprises at least one of nucleoside triphosphate, nucleic acid polymerase, and buffer, which are necessary for the function of nucleic acid polymerase and / or reverse transcriptase. In some embodiments, the kit further comprises one of any of the amplification reagents suitable for reverse transcription and / or amplification, such as reverse transcriptase, DNA polymerase, dNTP, buffer, and / or other elements (e.g., cofactors or aptamers). Typically, the reagent mixture is concentrated and aliquots are added to the final reaction volume along with the sample (e.g., RNA or DNA), enzyme, and / or water. In some embodiments, the kit further comprises reverse transcriptase (or an enzyme having reverse transcriptase activity) and / or DNA polymerase (e.g., a thermostable DNA polymerase such as Taq, ZO5, and its derivatives).
[0013] In another embodiment, a method for phasing at least one desired distal single nucleotide polymorphism and at least one desired exon single nucleotide polymorphism within the same target locus of nucleic acid isolated from a biological sample, comprising: (a) carrying out an amplification step of contacting genomic DNA isolated from the biological sample containing the target locus with a first set of oligonucleotide primers to produce a first amplification product containing at least one desired single nucleotide polymorphism and at least one exon reference single nucleotide polymorphism within the same locus; and (b) contacting mRNA isolated from the biological sample containing the target locus with a second set of oligonucleotide primers A method is provided which includes: (c) performing a reverse transcription and amplification step, which involves contacting with a sample to produce a second amplification product containing at least one target exon single nucleotide polymorphism and the at least one exon reference single nucleotide polymorphism; (d) determining the nucleic acid sequences of the first and second amplification products; (e) aligning the nucleic acid sequences of the first and second amplification products determined in step (c) based on the location of the at least one exon reference single nucleotide polymorphism; and (e) determining the haplotypes of the at least one target distal single nucleotide polymorphism and the at least one target exon single nucleotide polymorphism in the sample. In one embodiment, the at least one target distal single nucleotide polymorphism is an intron single nucleotide polymorphism. In a particular embodiment, the intron having the at least one target distal single nucleotide polymorphism is located adjacent to an exon having the at least one exon reference single nucleotide polymorphism at the same locus in the isolated genomic DNA. In one embodiment, the amplification step (a) and the reverse transcription and amplification step (b) are carried out in separate reaction vessels. In another embodiment, the nucleic acid sequences of the first and second amplification products in step (c) are determined using long-read sequencing, such as PacBio sequencing, also known as SMRT (Single Molecule Real Time) sequencing, or nanopore-based DNA long-read sequencing, commercially available from Oxford Nanopore Technologies. In another embodiment, the biological sample is selected from the group consisting of tissues, tumor tissues, blood, saliva, and cell lines derived from organisms.In one embodiment, when aligning the nucleic acid sequences of the first and second amplification products in step (d), if at least one distal single nucleotide polymorphism of interest is heterozygous, then at least one exon-reference single nucleotide polymorphism is heterozygous. In another embodiment, before aligning the nucleic acid sequences of the first and second amplification products in step (d), the method further includes aligning the nucleic acid sequences of the first and second amplification products to the nucleic acid sequence of a reference gene at a target locus, or to the nucleic acid sequence of the complete human genome or a portion thereof containing the reference gene at a target locus. In another embodiment, the target locus is the huntingtin gene. In some embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is an exon SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, at least one target single nucleotide polymorphism is the intron SNP rs7685686. In some embodiments, at least one exon-reference single nucleotide polymorphism (SNP) is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In this specification, the exon-reference SNP (e.g., rs362331, rs362273, rs34315806, and rs363099) must not be identical to (i.e., not corresponding to) at least one SNP of interest.In some embodiments, at least one exon-reference single nucleotide polymorphism (SNP) is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099, where at least one SNP of interest is an intronic SNP. In a particular embodiment, at least one exon-reference SNP is rs362331. In a particular embodiment, the mRNA is spliced mRNA (mature mRNA).
[0014] In another embodiment, there is provided an antisense oligonucleotide that specifically hybridizes to at least one target distal single nucleotide polymorphism in the huntingtin (HTT) gene for use in treating a patient having Huntington's disease, the antisense oligonucleotide being selected for treatment when the patient determines the specific haplotype of at least one target distal single nucleotide polymorphism and CAG short tandem repeat detected in the patient's biological sample. In some embodiments, at least one target distal single nucleotide polymorphism is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is an exon SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, at least one target distal single nucleotide polymorphism is rs7685686. In some embodiments, the specific haplotype includes the presence of an A allele for rs7685686, and the number of CAG short tandem repeats is greater than 36. In other embodiments, the specific haplotype includes the presence of a G allele for rs7685686, and the number of CAG short tandem repeats exceeds 36. In some embodiments, the antisense oligonucleotide specifically hybridizes to a region within the huntingtin (HTT) gene containing an A allele for rs7685686.In some embodiments, antisense oligonucleotides are selected to specifically hybridize to a desired distal single nucleotide polymorphism in an allele-specific manner. In some embodiments, the specific haplotype is determined using a method disclosed herein.
[0015] In another embodiment, an in vitro use is provided for haplotyping of at least one target distal single nucleotide polymorphism and short tandem repeat detected in the huntingtin (HTT) gene determined in a biological sample of an individual for the purpose of diagnosing Huntington's disease, wherein the detection of a specific haplotype of at least one target distal single nucleotide polymorphism and short tandem repeat detected in a patient's biological sample indicates that the individual has Huntington's disease. In some embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is an exon SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, at least one target single nucleotide polymorphism is the intron SNP rs7685686. In some embodiments, the specific haplotype includes the presence of an A allele for rs7685686, and the number of CAG short tandem repeats is greater than 36. In other embodiments, the specific haplotype includes the presence of a G allele for rs7685686, and the number of CAG short tandem repeats exceeds 36. In some embodiments, the specific haplotype is determined using the methods disclosed herein.
[0016] In another embodiment, an in vitro use of haplotyping of at least one target distal single nucleotide polymorphism and short tandem repeat detected in the huntingtin (HTT) gene determined in a patient's biological sample, for determining whether a patient with Huntington's disease is likely to respond to a therapy comprising an antisense oligonucleotide that specifically hybridizes to at least one target distal single nucleotide polymorphism in the huntingtin (HTT) gene, wherein if the specific haplotype of at least one target distal single nucleotide polymorphism and short tandem repeat is detected in the patient's biological sample, the patient is identified as being more likely to respond to the therapy. In some embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In a particular embodiment, at least one target distal single nucleotide polymorphism is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In a particular embodiment, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In a particular embodiment, at least one target single nucleotide polymorphism is an exon SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099.In certain embodiments, at least one target single nucleotide polymorphism is the intronic SNP rs7685686. In some embodiments, the specific haplotype includes the presence of an A allele for rs7685686 and the number of CAG short tandem repeats is greater than 36. In some embodiments, the antisense oligonucleotide specifically hybridizes to a region in the huntingtin (HTT) gene containing an A allele for rs7685686. In other embodiments, the specific haplotype includes the presence of a G allele for rs7685686 and the number of CAG short tandem repeats is greater than 36. In some embodiments, the antisense oligonucleotide is selected to specifically hybridize to the target distal single nucleotide polymorphism in an allele-specific manner. In some embodiments, the specific haplotype is determined using a method disclosed herein.
[0017] A method for treating a patient with Huntington's disease is further disclosed, comprising: (a) determining the haplotype of at least one target distal single nucleotide polymorphism and short tandem repeat in the huntingtin (HTT) gene in a biological sample of the patient; and (b) administering an antisense oligonucleotide that specifically hybridizes to at least one target distal single nucleotide polymorphism in the huntingtin (HTT) gene. Here, at least one distal single nucleotide polymorphism of interest may be selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, at least one target single nucleotide polymorphism is an exon SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, at least one target single nucleotide polymorphism is the intron SNP rs7685686. Furthermore, the determined haplotype includes the presence of an A allele for rs7685686 and has more than 36 CAG short tandem repeats. It is also disclosed that the antisense oligonucleotide specifically hybridizes to a region within the huntingtin (HTT) gene containing the A allele for rs7685686. In other embodiments, the specific haplotype includes the presence of a G allele for rs7685686. In some embodiments, the antisense oligonucleotide is selected to specifically hybridize to the desired distal single nucleotide polymorphism in an allele-specific manner.In this specification, specific haplotypes may be determined using the methods disclosed herein.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention pertains. Methods and materials similar to or equivalent to those described herein may be used in carrying out or testing the present invention, but preferred methods and materials are described below. In addition, the materials, methods, and examples are illustrative and not intended to be limiting.
[0019] Details of one or more embodiments of the present invention are shown in the accompanying drawings and the following description. Other features, purposes and advantages of the present invention will become apparent from the drawings and the modes for carrying out the invention, as well as from the claims. [Brief explanation of the drawing]
[0020] [Figure 1] Figure 1 shows distal separation between a tandem repeat region (i.e., CAG) and the intron or exon SNP of interest. The distance is expressed at the genomic DNA level and can extend up to 200 kb of nucleotides. Introns and exon SNPs are separated from each other by only about 10 kb and can be amplified in a PCR assay (i.e., F1 and R1 primers). Transcription and mRNA maturation of the gene target involve splicing, which reduces the distance between the CAG tandem repeat and the exon reference SNP, allowing for reverse transcription and PCR of approximately 10 kb of amplicon from a single molecule (i.e., F2 and R2 primers). [Figure 2]Figure 2 shows two amplicons: the first is generated from genomic DNA and contains intron-targeted SNPs and exon-reference SNPs, and the second is generated from cDNA derived from mature mRNA and containing CAG tandem repeats and exon-reference SNPs. Hybrid analysis allows for the phasing of distal intron SNPs with several CAG tandem repeats by aligning the exon SNPs between the two amplicons. [Figure 3] Figure 3 shows the phasing results for sample GM04282. Figure 3A shows that the mutant HTT allele, with approximately 77.6 CAGs ("high CAG"), which account for approximately 22.7% of the reads, shows phasing with the exon T allele in approximately 95% of the reads. The wild-type HTT allele, with approximately 18.1 CAGs ("low CAG"), which account for approximately 77.2% of the reads, also shows approximately 95% of reads with the exon T allele. Both alleles have >4% of reads with an undefined nucleotide at the location of the SNP of interest, which may be due to sequencing errors such as deletions. Since the exon SNP is homozygous, it is not useful for hybrid analysis. Figure 3B shows a bar plot of the cDNA data, which shows the presence of two distributions (i.e., low CAG repeat count and high CAG repeat count) for the mean within the range of approximately 18 and approximately 77 repeats (Figure 3A). The exon SNP alleles in both distributions are identified as homozygous thymine (Figure 3A). Figure 3C shows that DNA amplicon data indicate that the exon reference SNP allele T is phased in 100% of the reads passed with the intron-targeted SNP allele A. The other alleles have not been detected as showing homozygosity of the intron-targeted SNP to high-CAG repeat regions and low-CAG repeat regions. Therefore, this patient would not be suitable for allele-specific treatment. [Figure 4-1]Figure 4 shows the phasing results for sample GM13503. The AcDNA data in Figure 4 show that the mutant HTT allele, containing approximately 46.3 CAGs ("high CAG"), comprising about 48.1% of the reads, exhibits phasing with approximately 99.3% of the reads having the exon T allele and only 0.7% having the C allele. Conversely, the wild-type HTT allele, containing 18 CAGs ("low CAG"), comprises approximately 51.9% of the reads, with approximately 99.7% of the reads having the exon C allele and only 0.16% having the T allele. Both alleles have approximately 0% of the reads with undefined nucleotides at the location of the target SNP. Since the exon SNP is heterozygous, this patient may be suitable for allele-specific treatment, but further identification of intron SNPs phasing with high CAG is needed. [Figure 4-2] Figure 4B shows a bar plot of cDNA data, illustrating two distributions (i.e., low CAG count and high CAG count) for mean values within the range of approximately 18 and 46 repeats (Figure 4A). Exon SNPs for high CAG counts are designated as thymine (Figure 4A), while low CAG alleles have the majority of reads assigned as cytosine, indicating heterozygosity of the exon reference SNP. Figure 4C shows that DNA amplicon data for exon reference SNP allele C is phased by >98% of reads by intron-targeted SNP allele G and only 1.33% by allele A. Exon reference SNP allele T, on the other hand, is phased by approximately 98.03% of reads by intron-targeted SNP allele A and only 1.97% by intron-targeted SNP allele G. The low percentage of outward phasing results (1.33% and 1.97%) may be due to PCR errors or chimeric molecules, which introduce small noise into the final dataset. In conclusion, the mutant CAG allele (high CAG) is phased at the exon-reference SNP allele T, and then at the intron-targeted SNP allele A, indicating that this patient is eligible for A-specific ASO treatment. [Figure 5]Figure 5 shows the phasing results for sample GM04724. Figure 5A cDNA data shows that the mutant HTT allele, with approximately 71.2 CAGs ("high CAG"), comprising approximately 30.7% of the reads, shows phasing with the reference exon SNP T allele in approximately 96.68% of the reads and only about 0.55% of the C allele. The wild-type HTT allele, with approximately 16 CAGs ("low CAG"), comprising approximately 69.3% of the reads, shows approximately 98.16% of reads with the exon C allele and only 0.25% with the T allele. The two alleles show 1.59% and 2.77% of reads with undefined nucleotides at the location of the SNP of interest; this may be due to errors such as insertions or deletions. Since the exon SNP is heterozygous, this patient may be suitable for allele-specific treatment, but further identification of intron SNPs phasing with high CAG is needed. Figure 5B shows a bar plot of cDNA data, illustrating two distributions (i.e., low CAG count and high CAG count) for mean values within the ranges of approximately 16 and 71.2 repeats (Figure 5A). Exon SNPs for high CAG counts are designated as thymine (Figure 5A), while low CAG alleles have the majority of reads assigned as cytosine, indicating heterozygosity of the exon reference SNP. Figure 5C shows that DNA amplicon data for exon reference SNP allele C is phased by 97.54% of reads by intron-targeted SNP allele G and only 2.46% by allele A. Exon reference SNP allele T, on the other hand, is phased by 98.8% of reads by intron-targeted SNP allele A and only 1.2% by intron-targeted SNP allele G. The low percentages of outward phasing results (2.46% and 1.2%) may be due to PCR errors or chimeric molecules, which introduce small noise into the final dataset. In conclusion, the mutant CAG allele (high CAG) is phased at the exon-reference SNP allele T, and then at the intron-targeted SNP allele A, and this patient is eligible for A-specific ASO treatment, as shown in the example in Figure 4. [Figure 6] Figure 6 shows the results of clustering reads with a similarity threshold of 1.0 (i.e., clusters formed by the same reads). The results represent two final contigs showing the presence or absence of variant alleles of intron and exon SNPs. Residual reads that did not cluster into the main haplotype group were omitted from the analysis due to the presence of nucleotides with low quality scores, which resulted in sporadic errors at various nucleotide positions and a lack of 100% similarity between core clusters and off reads. [Figure 7] Figure 7 represents the results from the Integrative Genomics Viewer (IGV, Broad Institute) software for two final DNA amplicon clusters, including exon and intron SNP positions of heterozygous samples. [Figure 8] Figure 8 shows a flowchart of the bioinformatics analysis starting from raw sequencing data to visualization of DNA / cDNA reads and fading of CAG numbers by intron SNPs. [Figure 9] Figure 9 shows a detailed flowchart of the bioinformatics analysis starting from raw sequencing data to visualization of DNA / cDNA reads and fading of CAG numbers by intron SNPs. [Figure 10] Figure 10 shows a scheme for using the phase information of the present method for stratification and selection of patients likely to respond to SNP-specific antisense oligonucleotide treatment.
Embodiments for Carrying Out the Invention
[0021] I. Introduction Fading short tandem repeats (STRs) together with distal single nucleotide polymorphisms (SNPs) within the same locus at the individual patient level is challenging. Standard methodologies such as PCR following short-read sequencing of genomic DNA are unsuitable for such analyses when gene variants are too far apart (more than tens of kb). To overcome this problem, long-range PCR, which uses mRNA as a starting material for PCR amplification followed by long-read sequencing, helps to bring exon SNPs closer to distal STRs. However, this approach is not suitable for intronic SNPs. Therefore, there is still a need in the art for reliable methods that provide haplotype information for STRs and distal SNPs, especially intronic SNPs, within the same locus. Furthermore, there is also still a need for reliable methods that provide haplotype information for one or more exonic SNPs and one or more distal SNPs, especially intronic SNPs, within the same locus.
[0022] To address this task, this disclosure provides a novel molecular biology assay for preclinical or clinical biomarker characterization, for example, in the field of neuroscience (e.g., Huntington's disease, HD). Such methods may be used as companion diagnostic tools, and the identification of two or more paired loci is required to provide essential information for the safe and efficient stratification of patients receiving specific therapies or drug treatments. More specifically, the method enables the precise assignment of the spatial relationship (i.e., phase / phasing) of single nucleotide polymorphisms (SNPs) with short tandem repeats (STRs) or of another SNP from a very distal region within the genome (e.g., >150kb) at the haploid genome level. The method is based on DNA and RNA extraction, as well as parallel amplification of DNA and cDNA obtained from a single sample (e.g., blood or tissue), after which hybrid analysis of the data is provided by long-read sequencing techniques such as those of Pacific Biosciences of California, Inc. or Oxford Nanopore Technologies Limited.
[0023] This specification provides a method for analyzing a sample to phase exon and / or intron SNP alleles together with STRs or other distal SNPs. In certain embodiments, a method is provided for analyzing a sample to phase exon and / or intron SNP alleles together with trinucleotide CAG repeat regions located at distances greater than 150 kb, and for using the phase information for patient stratification (e.g., for patient inclusion in clinical trials or for patients to be provided for therapeutic treatment). In certain embodiments, phasing information of at least one distal SNP (e.g., rs7685686) within the huntingtin gene may be used to determine whether a patient could benefit from antisense oligonucleotide (ASO) treatment targeting trinucleotide CAG repeat regions and specific distal SNPs of Huntington's disease.
[0024] The present invention offers several advantages over current methods for SNP and STR genotyping. Most importantly, this novel method enables the precise determination of spatial relationships between SNPs and STRs or other SNPs at the haploid genome level, which is not possible with standard SNP and STR genotyping methods, particularly those originating from very distal regions within the genome. This is especially important for allele-specific treatment approaches. For example, the goal of allele-specific treatment for Huntington's disease is to suppress mutant (extended CAG) RNA molecules and maintain the wild type using SNP allele-specific antisense oligonucleotides (ASOs). Therefore, the assay described herein is crucial for identifying which SNP alleles are in the same haplotype (i.e., phase) as the extended CAG repeat.
[0025] Comprehensive design of PCR primers is usually necessary to develop robust and accurate assays. A potential problem is nonspecific amplification of random gene fragments, which will increase signal-to-noise ratio in the final sequencing data. However, alignment to a human reference genome sequence allows for the consideration of target sequences only in the final analysis. Furthermore, the design of PCR primers and hybrid assays must take into account linkage disequilibrium (LD) between intron-target SNPs and exon-reference SNPs to maximize the informativeness of the fingerprint approach (for example, if either SNP is homozygous, allele-specific treatment is impossible). LD structures vary between different populations, and this must be considered in assay design.
[0026] The workflow for the most challenging scenario, where the target SNP is an intron and too far from the CAG repeat to be amplified from genomic DNA by PCR, is as follows: (1) Co-extraction of DNA and RNA from a single sample. (2) Using RNA as an input material, the exon reference SNP is brought closer to the CAG repeat by RT-PCR amplification of the region covering the CAG repeat and the reference exon SNP. Ideally, the reference exon SNP should be in a state of high ligation disequilibrium with the target intron SNP, or at least within the same allele-frequency range. (3) PCR amplification of the region covering the above-mentioned reference exon SNP and intron target SNP using genomic DNA as input material. (4) Sequencing of the amplicon using long-read techniques to ensure that phasing information between variants within the amplicon is preserved. (5) Hybrid data analysis in which DNA and cDNA amplicon sequences are aligned and intron-targeted SNP alleles are phased with CAG repeat count using exon-reference SNP fingerprints.
[0027] Figure 1 shows distal separation between a tandem repeat region (i.e., CAG) and the intron or exon SNP of interest. The distance is expressed at the genomic DNA level and can extend up to 200 kb of nucleotides. Because the target intron SNP is far from the CAG repeat (>150 kb), the locus containing both cannot be PCR amplified using genomic DNA as a template. However, the intron target SNP and exon reference SNP are separated from each other by only about 10 kb, and the locus can be amplified by a PCR assay using genomic DNA as a template (i.e., using the oligonucleotide primers F1 and R1). Transcription and mRNA maturation of the gene target are characterized by a splicing event that results in a reduction in the distance between the CAG tandem repeat and the exon reference SNP, which allows for reverse transcription and PCR of a 10 kb amplicon from a single molecule by using mRNA as a starting material (i.e., using the oligonucleotide primers F2 and R2).
[0028] Figure 2 shows the two obtained amplicons: a first amplicon generated from genomic DNA and containing intron target SNPs and exon reference SNPs, and a second amplicon generated from cDNA derived from mature mRNA and containing CAG tandem repeats and exon reference SNPs. Here, the exon reference SNPs in both amplicons are the same. Thus, hybrid analysis allows for the phasing of at least one distal intron SNP with several CAG tandem repeats by the alignment of exon reference SNPs between the two amplicons. In some embodiments, the phasing accuracy can be further improved by using two or more exon reference SNPs.
[0029] II. Definition A "companion diagnostic" is a diagnostic test used as a companion to a therapeutic agent or treatment to determine its applicability to a specific patient or patient group, thereby stratifying patients according to a specific genomic profile. A companion diagnostic refers to a diagnostic test developed in parallel with a specific drug or therapy (e.g., an antibody, small molecule, or antisense oligonucleotide) to select or exclude patients or patient groups for treatment with that particular drug based on their biological characteristics that determine responders and non-responders to that therapy. These tests often involve the identification of specific genetic biomarkers or mutations that are associated with the disease or condition being treated, or that are prospectively useful in predicting a potential response or severe toxicity.
[0030] The term "biomarker" can refer to any detectable marker used to distinguish individual samples, for example, cancer versus non-cancerous samples. Biomarkers include modifications (e.g., DNA methylation, protein phosphorylation), differential expression, and mutations or variants (e.g., single nucleotide mutations, insertions, deletions, splice variants, and fusion variants). Biomarkers can be detected in DNA, RNA, and / or protein samples.
[0031] The terms “nucleic acid,” “polynucleotide,” and “oligonucleotide” refer to polymers of nucleotides (e.g., ribonucleotides or deoxyribonucleotides), including naturally occurring (e.g., adenosine, guanidine, cytosine, uracil, and thymidine) and naturally occurring (human-modified) nucleic acids. The term is not limited to the length of the polymer (e.g., the number of monomers). Nucleoside triphosphates containing ribose as a sugar are conventionally abbreviated as NTP, and nucleoside triphosphates containing deoxyribose as a sugar are abbreviated as dNTP. “Nucleic acid” means any nucleic acid molecule, including but not limited to DNA, RNA, and their hybrids, unless otherwise specified. In one embodiment, the nucleic acid bases forming a nucleic acid molecule may be bases A, C, G, T, and U, as well as their derivatives (A-adenine; C-cytosine; G-guanine; T-thymine; U-uracil). Derivatives or analogues of these bases are well known in the art and are exemplified in PCR Systems, Reagents and Consumables (Perkin Elmer Catalog 1996–1997, Roche Molecular Systems, Inc., Branchburg, New Jersey, USA). Nucleic acids can be single-stranded or double-stranded and generally contain a 5'-3' phosphodiester bond, although nucleotide analogs may have other bonds. Monomers are typically referred to as nucleotides. The terms "unnatural nucleotide" or "modified nucleotide" refer to nucleotides that contain a modified nitrogen-containing base, sugar, or phosphate group, or incorporate an unnatural moiety into their structure. Examples of unnatural nucleotides include LNA, dideoxynucleotides, biotinylated nucleotides, aminated nucleotides, deaminated nucleotides, alkylated nucleotides, benzylated nucleotides, and fluoroform-labeled nucleotides. "LNA" refers to locked nucleic acid. LNA is a modified RNA nucleotide in which the ribose portion is modified with an extra crosslink connecting the 2' oxygen and 4' carbon.LNA nucleotides can be mixed with DNA or RNA residues at any position within the oligonucleotide and hybridize with DNA or RNA according to Watson-Crick base pairing rules. Locked ribose conformation improves hybridization properties (e.g., by increasing the melting temperature).
[0032] A "nucleotide residue" is a single nucleotide that exists after being incorporated into a polynucleotide and thereby becoming a monomer of the polynucleotide. Therefore, a nucleotide residue is a nucleotide monomer of a polynucleotide, such as DNA, in which (i) the 3' terminal nucleotide residue is bound to only one adjacent nucleotide monomer of the polynucleotide by a phosphodiester bond from its phosphate group, and (ii) the 5' terminal nucleotide residue is bound to only one adjacent nucleotide monomer of the polynucleotide by a phosphodiester bond from the 3' position of its sugar, except thatnucleotide residue is bound to one adjacent nucleotide monomer of the polynucleotide via a phosphodiester bond at the 3' position of its sugar.
[0033] Due to well-understood base pairing rules, the identity (base identity) of a dNTP analog (or rNTP analog) incorporated into a primer or DNA elongation product (or RNA elongation product) is determined by measuring the unique electrical signal of the tag translocating through a nanopore. The identity of the incorporated dNTP analog (or rNPP analog) then allows for the identification of the complementary nucleotide residue in the single-stranded polynucleotide into which the primer or DNA elongation product (or RNA elongation product) hybridizes. Therefore, if the incorporated dNPP analog contains adenine, thymine, cytosine, or guanine, the complementary nucleotide residue in the single-stranded DNA is identified as thymine, adenine, guanine, or cytosine, respectively. Purine adenine (A) pairs with pyrimidine thymine (T). Pyrimidine cytosine (C) pairs with purine guanine (G). Similarly, with respect to RNA, if the incorporated rNPP analog contains adenine, uracil, cytosine, or guanine, the complementary nucleotide residues in the single-stranded RNA are identified as uracil, adenine, guanine, or cytosine, respectively.
[0034] In the context of this disclosure, terms such as “cell-free nucleic acid,” “cell-free DNA,” and “cell-free RNA” refer to non-tissue samples derived from an individual that have been processed to remove most cells (e.g., liquid biopsies). Examples of non-tissue samples include blood and blood components, urine, saliva, tears, and mucus.
[0035] Pre-mRNA, or "precursor mRNA," is a primary molecule in the eukaryotic transcription process, produced from a DNA template within the cell nucleus. Pre-mRNA molecules contain both coding (exons) and non-coding (intron) sequences and undergo maturation, which includes a splicing process in which the intron regions are removed, resulting in the molecule becoming mRNA after processing.
[0036] Mature mRNA is a eukaryotic RNA transcript that has been spliced and processed, and is ready for translation during the process of protein synthesis. Unlike pre-transcribed eukaryotic RNA, known as precursor mRNA, mature mRNA consists only of exons, with all introns removed.
[0037] A single nucleotide polymorphism (SNP) is a germline substitution of a single nucleotide at a specific location within the genome and is present in a sufficiently large proportion (more than 1%) of the population. The genomic distribution of SNPs is not homogeneous; SNPs occur more frequently in non-coding regions (intron regions) than in coding regions (exon regions). As used herein, a "distal SNP" is located more than 10 kb away from another SNP or short tandem repeat on the same nucleic acid. In some examples, the distance may exceed 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, 150 kb, 160 kb, 170 kb, 180 kb, or 190 kb. In some examples, the distances can be between 10-200kb, 10-160kb, 20-200kb, 20-160kb, 50-200kb, 50-160kb, 100-200kb, 100-160kb, or 150-200kb.
[0038] An "exon-referenced single nucleotide polymorphism" or "exon-referenced SNP," as used herein, may be covered by an intron-free cDNA amplicon (reverse transcribed from mature mRNA) and a DNA (e.g., genomic DNA) amplicon, thus enabling long-range phasing. Thus, an "exon-referenced SNP" also enables phasing and analysis of heterozygous and homozygous exon-targeted SNPs or intron-targeted SNPs, although certain scenarios require the "exon-referenced SNP" to be heterozygous (e.g., if the target SNP is heterozygous). In particular, if the distal target single nucleotide polymorphism (SNP) (intron or exon) of interest is heterozygous in a given individual, the optimal "exon-referenced SNP" should be polymorphic and heterozygous. Thus, an "exon-referenced SNP" may be used to align the nucleic acid sequences of a first amplification product derived from the target DNA nucleic acid (e.g., genomic DNA) and a second amplification product derived from the target cDNA (reverse transcribed from mature mRNA) nucleic acid based on the position of the "exon-referenced SNP."
[0039] The term "short tandem repeat" (STR), also known as a "microsatellite," refers to a cluster of repeating DNA in which a specific DNA motif (ranging in length from 1 to 6 or more base pairs) is repeated typically 5 to 50 times. Microsatellites occur in thousands of locations within the genome of organisms, making up about 3% of the human genome, and are often found in introns and intergeneric regions, but also in exons. They have a higher mutation rate than other regions of DNA, resulting in high genetic diversity. STRs may be involved in the development of disorders such as Huntington's disease. In this specification, an abnormal expansion of a tandem repeat consisting of the cytosine-adenine-guanine (CAG) nucleotide sequence in exon 1 of chromosome 4 of the huntingtin gene is considered to be the cause of the disease. A normal huntingtin gene typically has 10 to 35 CAG repeats, while individuals with Huntington's disease exhibit a number of CAG repeats ranging from 36 to over 100.
[0040] The term "phasing" refers to the assignment of genetic variants to homologous chromosomes of origin. Humans have two copies of every chromosome, one inherited from the mother and one from the father. In the context of STR and SNP phasing, the goal is to identify which SNP alleles reside on the same chromosome (i.e., within the haploid genome) with a particular length of STR. In the context of phasing two or more SNPs, the goal is to identify which SNP alleles of the first and second (and third…) SNPs reside on the same chromosome.
[0041] The term “primer” refers to a short nucleic acid (oligonucleotide) that serves as a starting point for polynucleotide chain synthesis by nucleic acid polymerase under appropriate conditions. The polynucleotide synthesis and amplification reaction typically involves a suitable buffer, dNTPs and / or rNTPs, and one or more optional cofactors, and is carried out at a suitable temperature. The primer typically includes at least one target hybridization region that is at least substantially complementary to the target sequence (e.g., having 0, 1, or 2 mismatches). For the purposes of this disclosure, this region is typically about 4 to about 10 nucleotides long, e.g., 5 to 8 nucleotides. “Primer pair” refers to a forward primer and a reverse primer oriented in opposite directions to the target sequence, which produce an amplified product under amplification conditions. The terms “forward” and “reverse” are arbitrarily assigned. Those skilled in the art will understand that the forward primer and reverse primer (primer pair) define the boundary of the amplified product. In some embodiments, multiple primer pairs rely on a single common forward or reverse primer. For example, multiple allele-specific forward primers can be considered part of a primer pair that shares the same common reverse primer, for instance, when multiple alleles are in close proximity to each other. A “primer set” or “primer palette” can refer to a primer pair, or two or more primer pairs designed to function together in a single multiple reaction.
[0042] As used herein, “probe” means any molecule that can selectively bind to a specifically intended target biomolecule, such as a nucleic acid sequence of interest to hybridize to the probe. The probe is detected by labeling with at least one non-nucleotide moiety. In some embodiments, the probe is labeled with a fluorophore and a quencher.
[0043] The terms "complementary" or "complementarity" refer to the ability of a nucleic acid within a polynucleotide to form base pairs with another nucleic acid within a second polynucleotide. For example, the sequence AGT (AGU for RNA) is complementary to the sequence TCA (UCA for RNA). Complementarity may be partial, where only a portion of the nucleic acid matches according to base pairing, or it may be complete, where all of the nucleic acid matches according to base pairing. A probe or primer is considered "specific" to a target sequence if it is at least partially complementary to the target sequence. Depending on the circumstances, the degree of complementarity to the target sequence is typically higher for shorter nucleic acids, such as primers, than for longer sequences (e.g., above 80%, 90%, 95%, or 98%). In some embodiments, primers and / or probes are 100% complementary to the target sequence.
[0044] The term "specifically amplifies" indicates that a primer set amplifies the target sequence at a statistically significant level compared to the non-target sequence. The term "specifically detects" indicates that a probe detects more target sequences at a statistically significant level than the non-target sequence. As understood in the art, specific amplification and detection can be determined using a negative control, e.g., a sample containing the same nucleic acid as the test sample but without the target sequence, or a sample lacking the nucleic acid. For example, primers and probes that specifically amplify and detect a target sequence result in a Ct that is easily distinguishable from the background (non-target sequence), e.g., a Ct that is at least 2, 3, 4, 5, 5-10, 10-20, or 10-30 cycles shorter than the background. The term "allele-specific" PCR refers to the amplification of a target sequence using primers that specifically amplify a particular allele variant of the target sequence. Typically, a forward or reverse primer contains the exact complement of the allele variant at its position.
[0045] In the context of two or more nucleic acids or two or more polypeptides, the terms “identical” or “percent identity” refer to two or more sequences or subsequences that are identical or have a specific percentage of identical nucleotides or amino acids (e.g., approximately 60% identity, e.g., at least one of 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater identity across a particular region) when measured using the BLAST or BLAST2.0 sequence comparison algorithm with default parameters, or by manual alignment and visual inspection. See, for example, the NCBI website ncbi.nlm.nih.gov / BLAST. Such sequences are then said to be “substantially identical.” Since percentage identity is typically determined across optimally aligned sequences, this definition applies to sequences with deletions and / or additions, as well as sequences with substitutions. Algorithms commonly used in this art describe gaps, etc. Typically, identity exists over a region containing at least approximately 8 to 25 amino acids or nucleotides in length, or over a region of 50 to 100 amino acids or nucleotides in length, or over the entire length of the reference sequence.
[0046] The terms “isolate,” “separate,” and “purify,” and similar terms, are not intended to be absolute. For example, the isolation of DNA or genomic DNA does not require the removal of 100% of non-DNA molecules. Those skilled in the art will recognize an acceptable level of purity for a given situation.
[0047] The term "amplification product" refers to the product of an amplification reaction. Amplification products include the primers used to initiate each round of polynucleotide synthesis. "Amplicon" is the sequence targeted for amplification, and this term can also be used to refer to amplification products. The 5' and 3' boundaries of the amplicon are defined by forward and reverse primers. "Reverse transcript," "RT product," and similar terms refer to cDNA molecules produced by extending RT primers on an RNA template with a polymerase having reverse transcriptase activity.
[0048] The term "kit" refers to any manufactured product (e.g., package or container) containing at least one reagent, such as a nucleic acid probe or probe pool, for the specific amplification, capture, tagging / conversion, or detection of RNA or DNA as described herein.
[0049] The term "amplification conditions" refers to the conditions in a nucleic acid amplification reaction (e.g., PCR amplification) that enable primer hybridization and template-dependent extension. The term "amplicon" or "amplification product" refers to a nucleic acid molecule that contains all or a fragment of the target nucleic acid sequence and is formed as the product of in vitro amplification by any suitable amplification method. The term "produces amplification product," when applied to a primer, indicates that the primer produces a defined amplification product under appropriate conditions (e.g., in the presence of nucleotide polymerase and NTP). Various PCR conditions are described in PCR Strategies (Innis et al., 1995, Academic Press, San Diego, CA) at Chapter 14; PCR Protocols: A Guide to Methods and Applications (Innis et al., Academic Press, NY, 1990).
[0050] A "nanopore" is defined as a structure having a nanoscale channel that allows ions in a solution to pass from one side to the other. Examples of nanopores include protein nanopores (e.g., α-hemolysin and other multi-subunit porins), synthetic nanopores, and hybrid protein / synthetic nanopores. In relevant embodiments, these nanopores are inserted into natural or artificial membranes that would otherwise help prevent the passage of ions and other molecules. The width of the nanopore channel should typically allow polymers such as single-stranded DNA to pass through when a voltage gradient is applied across the membrane. During transport, the size, charge, or other properties reduce the ionic current at a particular voltage. A "nanopore" includes, for example, a structure comprising: (a) first and second compartments separated by a physical barrier, the barrier having, for example, at least one pore having a diameter of about 1 to 10 nm; and (b) means for applying an electric field across the barrier so that charged molecules such as DNA, nucleotides, nucleotide analogs, or tags can pass from the first compartment through the pore to the second compartment. Ideally, the nanopore further includes means for measuring the electronic signature of molecules passing through its barrier. The nanopore barrier may be partially synthetic or naturally occurring. Examples of barriers include lipid bilayers having α-hemocyanin in them, oligomeric protein channels such as porins, and synthetic peptides. The barrier may also include an inorganic plate having one or more holes of appropriate size. Herein, "nanopore," "nanopore barrier," and "pore" within a nanopore barrier may be used interchangeably.
[0051] Nanopore devices are known in the art, and nanopores and methods using them are disclosed in U.S. Patents Nos. 7,005,264, 7,846,738, 6,617,113, 6,746,594, 6,673,615, 6,627,067, 6,464,842, 6,362,002, 6,267,872, 6,015,714, 5,795,782, and U.S. Patent Publications 2004 / 0121525, 2003 / 0104428, and 2003 / 0104428, each of which is incorporated herein by reference in whole.
[0052] A "nanopore array" is a chip containing many individual nanopores at known locations, and each nanopore can be individually interrogated electronically (enabling single-molecule electron nanopore-based sequencing through synthesis).
[0053] A "nanopore detectable tag" (also called a "nanopore tag") is a molecule, usually a polymer, covalently bonded to a nucleotide in a nanopore SBS reaction. Different nanopore tags are typically attached to each nucleotide, A, C, G, and T (or U), so as to induce different ion current cutoff signals as they pass through the nanopore channel when a voltage gradient is applied across the membrane.
[0054] Synthetic nanopore sequencing (also known as "nanopore SBS") refers to the approach previously described by the inventors (Kumar et al. 2012; Fuller et al. 2016; Stranges et al. 2016), where tags bound to nucleotides can be distinguished by their effect on the ionic current passing through the nanopore as these modified nucleotides are attached to the growing DNA strand. Measurements can be performed while the tagged nucleotides are still part of the ternary complex, or after their tags have been released by polymerase reactions.
[0055] The terms “individual,” “subject,” and “patient” are used interchangeably herein. An individual may be pre-diagnosis, post-diagnosis, pre-treatment, treatment, or post-treatment. In the context of this disclosure, an individual is typically seeking medical treatment.
[0056] The term "sample" refers to any composition that contains or is presumed to contain nucleic acids. This term includes purified or isolated components of cells, tissues, or blood, e.g., DNA, RNA, proteins, cell-free portions, or cell lysates. A sample may be, for example, an FFPET derived from a tumor or metastatic lesion. A sample may also be derived from frozen or fresh tissue, or from liquid samples, e.g., blood or blood components (plasma or serum), urine, semen, saliva, sputum, mucus, semen, tears, lymph, cerebrospinal fluid, oral / throat rinses, bronchoalveolar lavage, or materials washed from swabs. A sample may also include components and constituents of in vitro cultures of cells obtained from an individual, including cell lines. A sample may also be partially processed from samples obtained directly from an individual, e.g., cell lysates or blood depleted of red blood cells. Tumor samples may include tumor-derived tissue or tumor-derived DNA, e.g., ctDNA in the blood of a cancer patient.
[0057] The term "obtaining a sample from an individual" means that a biological sample from that individual is provided for testing. The sample may be obtained directly from the individual or from a third party who has obtained a sample directly from the individual. The sample may be collected before, during, or after treatment. The sample may be collected from a patient who is likely to require treatment because they are suspected of having or diagnosed with disease X, or from a normal individual who is not suspected of having any disorder. Treatment regimen, (high / low / high-frequency / low-frequency) dose.
[0058] The term “evaluate disease X” is used to indicate that the methods disclosed herein assist a healthcare professional in evaluating, for example, whether a physician has disease X, is at risk of developing disease X, or is at risk of progressing through the course of disease X. The presence of biomarker Y, a combination of biomarkers Y;Z;..., or a ratio of biomarkers Y;Z;... indicates that an individual has disease X, is at risk of developing disease X, or is at risk of progressing through the course of disease X. In one embodiment, the term “evaluate disease X” is used to indicate that the methods according to the present invention assist a healthcare professional in evaluating whether an individual has disease X. In this embodiment, the presence of biomarker Y in a sample indicates that the individual has disease X, i.e., the presence of biomarker Y indicates that disease X is present in the individual at / above / below the reference level. In a particular embodiment, the term “at the reference level” refers to a level of biomarker in a sample derived from an individual or patient that is essentially identical to a level that is up to 1%, up to 2%, up to 3%, up to 4%, up to 5% different from the reference level.
[0059] The term “to provide a therapy to an individual” means that a therapy is prescribed, recommended, or made available to an individual. The therapy may, in practice, be administered to an individual by a third party (e.g., by injection to a hospitalized patient) or by the individual itself. The expression “to select a therapy,” as used herein, means identifying or selecting a therapy for a patient using generated information or data regarding the level or presence of biomarker Y in a patient sample. In some embodiments, the therapy may include drug D. In some embodiments, the expression “to identify / select a therapy” includes identifying a patient who requires an indication of an effective amount of drug D to be administered. In some embodiments, recommending a treatment includes recommending that the amount of drug D to be administered be adapted. As used herein, the expression “to recommend a treatment” may also mean using generated information or data to suggest or select a therapy containing drug D for patients identified or selected as likely or unlikely to respond to a therapy containing drug D. The information or data used or generated may be in written, oral, or electronic form. In some embodiments, using the generated information or data includes communication, presentation, reporting, storage, transmission, transfer, supply, outgoing, implementation, or a combination thereof. In some embodiments, communication, presentation, reporting, storage, transmission, transfer, supply, transmission, administration, or a combination thereof is performed by a computer device, an analyzer unit, or a combination thereof. In some further embodiments, communication, presentation, reporting, storage, transmission, transfer, supply, transmission, administration, or a combination thereof is performed by a researcher or medical professional. In some embodiments, the information or data includes comparing the level of biomarker Y to a reference level. In some embodiments, the information or data includes an indicator of whether biomarker Y is present or absent in the sample. In some embodiments, the information or data includes an indicator of whether a therapy including drug D is appropriate for the patient.
[0060] Terms such as "label," "tag," and "detectable portion" refer to compositions that can be detected by spectroscopic, photochemical, biochemical, immunochemical, chemical, or other physical means. For example, useful labels include fluorescent dyes (fluorophores), luminescent agents, and radioisotopes (e.g., 32 P, 3 Examples include electron-dense reagents or affinity-based moieties, such as poly-A (which interacts with poly-T) or poly-T tags (which interact with poly-A), His tags (which interact with Ni), or streptavidin tags (which are separable from biotin). Those skilled in the art will understand that detectable labels conjugated to nucleic acids do not exist naturally.
[0061] Unless otherwise defined, technical and scientific terms used herein have the same meanings as those generally understood by those skilled in the art. See, for example, Lackie, DICTIONARY OF CELL AND MOLECULAR BIOLOGY, Elsevier (4th ed. 2007); Sambrook et al., MOLECULAR CLONING, A LABORATORY MANUAL, Cold Springs Harbor Press (Cold Springs Harbor, NY 1989). The terms “a” or “an” are intended to mean “one or more.” The terms “comprise,” “comprises,” and “comprising” when preceding a list of processes or elements are intended to mean that the addition of further processes or elements is optional and not excluded.
[0062] S. Nucleic acid samples Samples for biomarker detection can be obtained from any suspected source containing substantial amounts of unfragmented nucleic acids or large fragments of nucleic acids (>10kb), e.g., tissue (including tumor tissue), blood (including cell-free nucleic acids such as cell-free DNA and cell-free RNA), skin, swabs (e.g., buccal, vaginal), urine, saliva, etc. Methods for isolating nucleic acids from biological samples are known, for example, as described by Sambrook, and several kits (e.g., High Pure RNA Isolation Kit, High Pure Viral Nucleic Acid Kit, and MagNA Pure LC Total Nucleic Acid Isolation Kit, DNA Isolation Kit for Cells and Tissues, DNA Isolation Kit for Mammalian Blood, and the High Pure FFPET DNA Isolation Kit available from Roche) are commercially available. In the context of the methods of this disclosure, genomic DNA and RNA can be collected and isolated.
[0063] IV. Detailed description of the workflow The following provides a detailed description of an exemplary laboratory workflow, using an exemplary subsequent bioinformatics workflow that includes sample preparation, amplification, library preparation and sequencing analysis, as well as analysis of sequencing data illustrating the implementation of the disclosed methods.
[0064] Laboratory Workflow The sample of interest (e.g., blood or tissue) is processed using a DNA and RNA co-extraction method (e.g., AllPrep DNA / RNA Micro Kit, Qiagen) according to the manufacturer's protocol. Alternatively, the sample of interest may be split into two parts, with DNA extracted from the first part (e.g., using the DNeasy Blood & Tissue Kit, Qiagen) and RNA extracted from the second part of the sample according to the manufacturer's protocol (e.g., using the RNeasy Kit, Qiagen). The isolated nucleic acid material is analyzed using a quality control workflow (e.g., NanoDrop® spectrophotometer, Qubit® fluorometer) according to the manufacturer's protocol, and subsequently subjected to amplification reactions based on polymerase chain reaction (PCR) or reverse transcription and polymerase chain reaction (RT-PCR). The aforementioned reactions (i.e., PCR and RT-PCR) are performed on genomic DNA and mRNA accordingly, utilizing standard reagents (i.e., polymerase, reaction buffer, primers, etc.) and standard equipment (e.g., pipettes, tips, laboratory tubing, PCR hood, and PCR thermocycler) for long-amplicon generation. Custom primer designs for performing PCR and RT-PCR amplification reactions to generate DNA and cDNA amplicons can be obtained using the open-source algorithms NCBI Primer Blast or Primer3 (National Library of Medicine, Bethesda, MD, USA). PCR and RT-PCR reactions are performed according to the manufacturer's protocol (e.g., Expand® High Fidelity PCR System, Roche). Here, the number of PCR and RT-PCR cycles depends on the quantity and quality of the isolated nucleic acid used as starting material.The target DNA and cDNA amplicons are subjected to a cleanup process according to standard methodologies and manufacturer documentation (e.g., Agincourt AMPure XP magnetic beads, Beckman Coulter; or Blue Pippin Prep, Sage Science) to remove excess unused amplification primers or short-off target molecules. The purified amplicons are tested in a quality control workflow (e.g., NanoDrop® spectrophotometer, Qubit® fluorometer, or Bioanalyzer, Agilent Technologies, Santa Clara, CA, USA) according to the manufacturer's protocol and recommendations. Two amplicons from a single sample (i.e., DNA and cDNA) are pooled at equimolar concentrations and ligated with a barcode index containing a sequencing adapter (e.g., Hairpin Loop for PacBio Sequencing using Library Preparation Kit, Pacific Biosciences of California, Inc., Menlo Park, USA) according to the manufacturer's protocol, and further combined into the final sequencing library pool. These generated samples were quantified using a Qubit® spectrophotometer (e.g., Broad Range DNA Kit, Thermo Fisher Scientific), loaded onto a sequencing device (e.g., Sequel System, Pacific Biosciences of California, Inc., Menlo Park, USA), and processed according to the manufacturer's instructions.
[0065] Bioinformatics Workflow Long-read sequencing devices (e.g., Sequel system, Pacific Biosciences of California, Inc., Menlo Park, USA; or GridION system, Oxford Nanopore Technologies Ltd., Oxford, UK) can be bioinformatically processed according to the flowchart shown in Figure 8. In this specification, the first step is to base-call the sequencing raw data (100) to FASTQ reads (110) using onboard software. Such generated data can be transferred to a local Linux®-based server for bioinformatics data analysis. The available reads are demultiplexed using open-source algorithms (e.g., LIMA or pbCCS) or a consensus called (120), and aligned to a whole human genome reference (e.g., BWA, STAR, or minimap2) depending on the presence or absence of introns (130). Alignment results allow for the extraction of single-molecule SNP calls (MUSCLE, SNVer, VarScan, Samtools / mpileup) and exon or intron statistics (131). The number of short tandem repeats from cDNA reads can be quantified using Repeat Genotyper, RepeatAnalysisTools, or DeepRepeat software (132). DNA and cDNA data undergo similarity clustering (e.g., UCLUST, VSEARCH, or pbAA) to enable better analysis of CAG tandem repeats or exon / intron SNPs. Information from DNA and cDNA data is combined into a single CSV file and binned according to the amount of CAG tandem repeats or exon / intron SNPs. The resulting bins for homozygotes and heterozygotes can be visualized using the Integrative Genomics Viewer (IGV) and ultimately plotted in a table for each separate step to show which target SNP alleles are in the same phase as the mutant CAG repeat (140).Statistical analysis may be performed to remove noisy signals or chimeric reads and improve the final statistical accuracy.
[0066] In the relevant workflow, raw data generated by one of the long-read sequencing devices (e.g., Sequel system, Pacific Biosciences of California, Inc., Menlo Park, USA) can be bioinformatically processed according to the flowchart shown in Figure 9. This method uses multiple steps of read filtering to generate the highest quality data for clustering and alignment. Statistical analysis is performed using validated open-source programs to align signals to the final phasing report and STR distribution. In the first step, the raw sequencing data is transferred to a server via a local area network (LAN) for post-processing and statistical data analysis. A consensus call is performed on FASTQ reads using pbCCS software to filter out shorter amplicons and allow reads with low consensus quality or incomplete sequencing paths. The generated consensus reads are demultiplexed using an open-source algorithm (i.e., LIMA) (200), and the data is stored separately for cDNA amplicons (202) and gDNA amplicons (201). Subsequently, high-quality reads are clustered using PacBio Amplicon Analysis (pbAA) (212, 211). The raw and clustered bins are aligned to a whole human genome reference, depending on the presence or absence of introns (222, 221). Read binning is performed to remove noisy signals (e.g., PCR chimeras) and improve the final statistical accuracy. High-quality data are used for extracting gene locations, cDNA containing exon SNPs and motifs of interest (e.g., CAG, CAA, CCG, CCA, CGG) (232), and gDNA containing exon and intron SNP coordinates (231, 232) (233). The alignment results enable single-molecule SNP calling and subsequent extraction of exon and intron statistical metrics at target locations (241). The number of short tandem repeats from cDNA reads is quantified using RepeatAnalysisTools (233).Reads of gDNA and cDNA subjected to similarity clustering (i.e., pbAA) allow for better analysis of STRs and exon / intron SNPs (242), while unclustered reads allow for better quantitative plotting of reads per haplotype. Information from both of the aforementioned workflows (i.e., error-corrected and uncorrected gDNA / cDNA) is combined into a single file and binned according to the amount of STRs or exon / intron SNPs. The resulting bins for homozygotes and heterozygotes are visualized in a waterfall image to represent STRs (252). The primary outcome of the overall analysis is to show which intron SNP alleles are synchronized with mutant or wild-type STRs (251).
[0067] Patient stratification workflow The generated results may be analyzed to identify the presence of a desired intronic SNP allele synchronized with the mutant STR allele (e.g., CAG repeat). Figure 10 visualizes the four final possible phasing outcomes (i.e., haplotypes consisting of the CAG repeat and intronic target SNP). Note that only combinations with heterozygous target SNPs enable allele-specific treatment using antisense oligonucleotide (ASO)-based therapies. For adenine-specific ASOs, the following haplotypes would be eligible, as shown in Figure 10: 1) Haplotype 1: Mutant STR+Adenine SNP and Haplotype 2: Wild-type STR+Guanine SNP Reverse SNP allele: 2) Haplotype 1: Mutant STR+Guanine SNP and Haplotype 2: Wild-type STR+Adenine SNP
[0068] This results in depletion of wild-type STR mRNA and conservation of mutant STR mRNA. Therefore, treatment of patients with reverse SNP alleles²) will require the use of Guanine-specific ASOs. Remaining homozygous combinations³ (homozygous SNP A / A) or⁴ (homozygous SNP G / G) will either result in depletion of wild-type and mutant STR mRNA or will not result in on-target reduction of either STR allele after ASO administration.
[0069] V. Kit Provided herein is a kit for determining the nucleic acid sequences of at least one target distal single nucleotide polymorphism and short tandem repeat within the same target locus of nucleic acids isolated from a biological sample, comprising: a first set of oligonucleotide primers for generating a first amplification product containing at least one target single nucleotide polymorphism and at least one exon-reference single nucleotide polymorphism within the same locus; and a second set of oligonucleotide primers for generating a second amplification product containing at least one target exon-reference single nucleotide polymorphism and the at least one exon-reference single nucleotide polymorphism. In this specification, the kit is adapted to carry out any of the methods disclosed herein. In certain embodiments, the first set of oligonucleotide primers includes an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 1 and an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 2. In another particular embodiment, the second set of oligonucleotide primers includes an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 3, an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 4 and an oligonucleotide primer containing the nucleic acid sequence of SEQ ID NO: 5. In some embodiments, the kit includes a sample collection container (e.g., a tube, vial, multiwell plate, or multivessel cartridge).
[0070] In some embodiments, the kit includes reagents and / or components for nucleic acid purification. For example, the kit may include lysis buffers (e.g., detergents, synergisticants, buffers, etc.), enzymes or reagents for denaturing proteins or other undesirable substances in the sample (e.g., protease K), and enzymes for preserving nucleic acids (e.g., DNase and / or RNase inhibitors). In some embodiments, the kit includes components for nucleic acid isolation, such as solid or semi-solid matrices including chromatography matrices, magnetic beads, magnetic glass beads, glass fibers, and silica filters. In some embodiments, the kit includes washing and / or elution buffers for the purification and release of nucleic acids from the solid or semi-solid matrices. For example, the kit may include components such as the MagNA Pure LC Total Nucleic Acid Isolation Kit, DNA Isolation Kit for Mammalian Blood, High Pure or MagNA Pure RNA Isolation Kit (Roche), DNeasy or RNeasy Kit (Qiagen), and PureLink DNA or RNA Isolation Kit (Thermo Fisher).
[0071] This kit may further include reagents for amplification, such as reverse transcriptase, DNA polymerase, dNTPs, buffer, and / or other elements suitable for reverse transcription and / or amplification (e.g., cofactors or aptamers). Typically, the reagent mixture is concentrated and aliquots are added to the final reaction volume along with the sample (e.g., RNA or DNA), enzyme, and / or water. In some embodiments, the kit further includes reverse transcriptase (or an enzyme having reverse transcriptase activity) and / or DNA polymerase (e.g., a thermostable DNA polymerase such as Taq, ZO5, and its derivatives).
[0072] In some embodiments, the kit further includes consumables, such as plates or tubes for nucleic acid preparation, tubes for sample collection, or plates, tubes, or microchips for PCR or qRT-PCR. In some embodiments, the kit further includes, for example, instructions for use, a website, or a reference to software for further processing of sequencing data. [Examples]
[0073] Example 1 Amplification of genomic DNA and reverse transcription of mRNA from the HTT gene Huntington's disease is a rare autosomal dominant genetic disorder caused by an abnormal expansion of a tandem repeat of cytosine-adenine-guanine (CAG) nucleotides in exon 1 of chromosome 4 of the huntingtin gene (HTT). While normal HTT typically has 10–35 CAG repeats, individuals with Huntington's disease may have 36–100 or more CAG repeats. Ideally, to treat the disease, only mutant HTT should be suppressed, leaving wild-type HTT intact. Targeting heterozygous SNPs near the CAG repeats would enable such an allele-specific approach. Several SNPs with high allele frequencies, including the intron 42 SNP rs7685686, have been identified in the HTT gene as candidate targets for treatment. However, assays to identify heterozygous patients with specific SNP alleles by phasing them against mutant CAG repeats are a prerequisite for such selective treatment. Fading of the CAG repeat in exon 1 and the SNP in intron 42 of HTT is impossible using standard methodologies due to the long distance of over 130kb between these genetic variants. Furthermore, using RNA alone as a starting material to shorten the distance between the CAG repeat and the distal SNP is not feasible for intronic SNPs. Therefore, a novel laboratory workflow utilizing both genomic DNA and RNA has been developed to enable fading analysis of STRs and distal, as well as intronic SNPs, within the same locus.
[0074] Thirteen samples from the NIGMS Human Genetic Cell Repository at the Coriell Institute for Medical Research were processed as described below. Genomic DNA and total RNA were isolated using co-extraction (AllPrep DNA / RNA Micro Kit, Qiagen) according to the manufacturer's protocol. The extracted nucleic acids were processed in a quality control workflow based on quantification of concentration and purity using the Qubit HS DNA / RNA Kit and nanodrop spectrophotometer accordingly. Samples characterized by high concentrations and correct 260 / 230, 260 / 280 quality and RNA Integrity Number (RIN) were subjected to PCR and RT-PCR.
[0075] For a DNA assay of an amplicon of approximately 10kb, a minimum of 100ng of high-quality material input was required. This was mixed with 10µl of 5xPrimeSTAR GXL Buffer, 4µl of dNTPS Mixture (200µM each), 0.7µl of 15µM forward and reverse primers, 1µl of PrimerSTAR GXL DNA Polymerase, and 13.6µl of nuclease-free water (i.e., 30µl of Master Mix). Input genomic DNA was added at 5ng / µl, resulting in a total volume of 20µl (50µl final reaction volume). The amplification reaction was carried out in a preheated thermocycler under the following conditions: pre-incubation at 98°C for 1 minute, followed by 30 cycles of 20 seconds at 98°C, followed by incubation at 58°C for 15 seconds, and then incubation at 68°C for 15 minutes. The reaction was stopped by infinitely decreasing the temperature to 4°C. The amplicons were purified using AMPure® PB Beads (100-265-900, Pacific Biosciences) at a 0.45-fold ratio according to the manufacturer's protocol. The amount of amplicons was measured using a Qubit BR DNA assay, and their size was verified using TapeStation and Genomic DNA Screen Tapes. The resulting amplicons were subjected to a second purification using AMPure® PB Beads (100-265-900, Pacific Biosciences) at a 3.1-fold ratio according to the manufacturer's protocol to remove any potentially remaining fragments of <7kb. After the second purification, the amount of amplicons was measured again using a Qubit BR DNA assay, and their size was verified using TapeStation and Genomic DNA Screen Tapes.
[0076] The cDNA assay of an amplicon of approximately 10kb required a minimum of 1000 ng of highly complete total RNA for the first step. The reaction involved gDNA removal by mixing 1 μL of 10x exDNase buffer with 1 μL of exDNase enzyme and 8 μL of total RNA input (minimum 1000 ng). The prepared sample was incubated at 37°C for 2 minutes, followed by the addition of 1 μL of 100 mM DTT to inactivate the enzyme (final volume 11 μL), and incubated at 55°C for 5 minutes. The next step involved annealing with gene-specific reverse transcription primers (2 μM concentration) and the addition of 10 mM dNTP mix (10 mM each) to 11 μL of RNA from the previous step. The reaction was heated at 65°C for 5 minutes and immediately placed on ice for at least 1 minute. The subsequent step involved mixing 5x SuperScript IV buffer with 100 mM DTT, Ribonucleotide Inhibitor, SuperScript IV RT enzyme (7 μL), and 13 μL of RNA with pre-annealed RT primers. Reverse transcription required incubation at 55°C for 10 minutes, followed by incubation at 80°C for 10 minutes, and finally holding at 4°C for indefinite time. The final step of reverse transcription involved removal of RNA from reaction with 1 μL of E. coli RNase H, followed by incubation at 37°C for 20 minutes. After cDNA synthesis, the cDNA was diluted 5-fold with water. The next step in PCR amplification used 5X PrimeSTAR GXL Buffer, dNTP Mixture, forward and reverse primers (15 μM), PrimerSTAR GXL DNA Polymerase, and water (total 40 μL), as well as 10 μL of cDNA diluted from the previous step (total volume 50 μL reaction). To increase yield, four PCR wells were used for each sample. The amplification reaction was carried out in a preheated thermocycler under the following conditions: pre-incubation at 98°C for 1 minute, followed by 33 cycles of 20 seconds at 98°C, 15 seconds at 60°C, followed by incubation at 68°C for 20 minutes. The reaction was stopped by allowing the temperature to decrease indefinitely to 4°C.Each amplicon / sample had four wells pooled together before initial purification at a 0.45-fold ratio using AMPure® PB Beads (100-265-900, Pacific Biosciences) according to the manufacturer's protocol. The amount of amplicons was measured using a Qubit BR DNA assay, and their size was validated using TapeStation. The generated amplicons were subjected to a second purification at a 3.1-fold ratio using AMPure® PB Beads (100-265-900, Pacific Biosciences) according to the manufacturer's protocol to remove any remaining <7kb potential fragments. After the second purification, the amount of amplicons was measured using a Qubit BR DNA assay, and their size was validated using TapeStation, Genomic DNA Screen Tapes.
[0077] DNA and cDNA amplicons from the same sample were normalized according to the Qubit value and pooled at equimolar concentrations. Each sample pool was ligated with a sequencing index adapter and loaded into a sequencing flow cell according to the manufacturer's protocol (Pacific Biosciences of California, Inc., Menlo Park, USA). Pooling of DNA and cDNA samples was performed to combine the results from each sample into the indexed bulk data.
[0078] [Table 1]
[0079] [Table 2]
[0080] Example 2 Bioinformatics analysis The readouts from both amplicons (i.e., DNA sequences derived from the genomic DNA amplification and sequencing process; as well as cDNA sequences derived from the mRNA reverse transcription, amplification, and sequencing process) are first aligned to a reference gene or a complete human genome and processed with open-source programs to quantify the number of tandem repeats per molecule and extract nucleotide signals at desired SNP locations. Alignment programs used include BWA, STAR, or minimap2, and programs for CAG quantification include Tandem Repeat Genotyper, RepeatAnalysisTools, or DeepRepeat software. Exemplary results for three different Coriell samples are shown in Figures 3–5 (cell lines GM04282, GM13503, and GM04724 were obtained from the NIGMS Human Genetic Cell Repository at the Coriell Institute for Medical Research). The results of the cDNA analysis are shown in Panel A (Figures 3-5) and visualized in Panel B (Figures 3-5), indicating which exon SNP alleles are in phase with high and low CAG repeat counts, and showing technical details such as mean CAG count, standard deviation, and percentage of reads containing high or low tandem repeat counts. The results of the DNA amplicon are shown in Panel C (Figures 3-5), showing the percentage of reads containing intron polymorphisms (i.e., Guanine or Adenine) versus exon polymorphisms (i.e., Cytosine or Thymine). In this specification, Figure 3 illustrating the results for sample GM04282 shows that, based on the cDNA data, the mutant HTT allele with approximately 77.6 CAGs ("high CAG"), comprising approximately 22.7% of the reads, exhibits a phasing with the exon T allele in approximately 95% of the reads. On the other hand, wild-type HTT alleles containing approximately 18.1 CAGs ("low CAG"), which comprise about 77.2% of the reads, also represent about 95% of the reads that contain the exon T allele.Both alleles have >4% of reads with undefined nucleotides at the exon reference SNP position, which may be due to errors such as insertions or deletions. DNA amplicon data (Figure 3C) shows that both exon and intron SNPs are homozygous, and therefore this patient would not be suitable for allele-specific treatment. Figures 4C and 5C show data for samples GM13503 and GM04742, which are heterozygous for both exon and intron SNPs, making both of these individuals suitable for allele-specific approaches.
[0081] Hybrid analysis of DNA and cDNA amplicon sequences required for phasing CAG tandem repeats with intronic SNPs is based on read tiling via exon reference SNPs, and the results for each amplicon that enable this are shown in Figures 3, 4, and 5 (A and C). The data binning and filtration process is based on sequence similarity (i.e., 100%) as shown in Figure 6, where error-prone reads based on QC metrics are discarded, and the signal with the highest readout volume is called the consensus. Individual haplotypes can be visualized using IGV software, as shown in the genomic DNA amplicons in Figure 7 or in the bar plots in Panel B (Figures 3–5).
[0082] References Becanovic, K., Norremolle, A., Neal, SJ, Kay, C., Collins, JA, Arenillas, D., Lilja, T., Gaudenzi, G., Manoharan, S., Doty, CNand Beck, J. (2015) A SNP in the HTT promoter alters NF-κB binding and is a bidirectional genetic modifier of Huntington disease.Nature neuroscience,18(6),pp.807-816. Carroll,J.B.,Warby,S.C.,Southwell,A.L.,Doty,C.N.,Greenlee,S.,Skotte,N.,Hung,G.,Bennett,C.F.,Freier,S.M.and Hayden,M.R.(2011)Potent and selective antisense oligonucleotides targeting single-nucleotide polymorphisms in the Huntington disease gene / allele-specific silencing of mutant huntingtin.Molecular Therapy,19(12),pp.2178-2185. Claassen D.O.,Corey-Bloom J.,Dorsey E.R.,Edmondson M.,Kostyk S.K.,LeDoux M.S.,Reilmann R.,Rosas H.D.,Walker F.,Wheelock V.,Svrzikapa N.,Longo K.A.,Goyal J.,Hung S.,Panzara M.A.(2020)Genotyping single nucleotide polymorphisms for allele-selective therapy in Huntington disease.Neurol Genet,6(3)e430. Flower M.,Lomeikaite V.,Ciosi M.,Cumming S.,Morales F.,Lo K.,Moss D.H.,Jones L.,Holmans P.,Monckton D.G.,Tabrizi S.J.,(2019)MSH3 modifies somatic instability and disease severity in Huntington’s and myotonic dystrophy type 1,Brain,142(7):1876-1886 Goold R.,Flower M.,Moss D.H.,Medway C.,Wood-Kaczmar A.,Andre R.,Farshim P.,Bates G.P.,Holmans P.,Jones L.,Tabrizi S.J.,(2019)FAN1 modifies Huntington’s disease progression by stabilizing the expanded HTT CAG repeat,Human Molecular Genetics,28(4):650-661. Hannan,A.(2018)Tandem repeats mediating genetic plasticity in health and disease.Nat Rev Genet 19,286-298. Kartsaki,E.,Spanaki,C.,Tzagournissakis,M.,Petsakou,A.,Moschonas,N.,MacDonald,M.,&Plaitakis,A.(2006)Late-onset and typical Huntington disease families from Crete have distinct genetic origins.International Journal of Molecular Medicine,17,335-346. Kay,C.,Collins,J.A.,Skotte,N.H.,Southwell,A.L.,Warby,S.C.,Caron,N.S.,Doty,C.N.,Nguyen,B.,Griguoli,A.,Ross,C.J.and Squitieri,F.(2015)Huntingtin haplotypes provide prioritized target panels for allele-specific silencing in Huntington disease patients of European ancestry.Molecular Therapy,23(11):1759-1771. Lee,J.M.,Gillis,T.,Mysore,J.S.,Ramos,E.M.,Myers,R.H.,Hayden,M.R.,Morrison,P.J.,Nance,M.,Ross,C.A.,Margolis,R.L.and Squitieri,F.(2012)Common SNP-based haplotype analysis of the 4p16.3 Huntington disease gene region.The American Journal of Human Genetics,90(3):434-444. Ramos,E.M.,Latourelle,J.C.,Lee,JH.et al.(2012)Population stratification may bias analysis of PGC-1α as a modifier of age at Huntington disease motor onset.Hum Genet 131:1833-1840. Shin,J.W.,Shin,A.,Park,S.S.and Lee,J.M.(2022)Haplotype-specific insertion-deletion variations for allele-specific targeting in Huntington’s disease.Molecular Therapy-Methods&Clinical Development,25:84-95. Skotte,N.H.,Southwell,A.L.,Ostergaard,M.E.,Carroll,J.B.,Warby,S.C.,Doty,C.N.,Petoukhov,E.,Vaid,K.,Kordasiewicz,H.,Watt,A.T.and Freier,S.M.(2014)Allele-specific suppression of mutant huntingtin using antisense oligonucleotides:providing a therapeutic option for all Huntington disease patients.PloS one,9(9),p.e107434. Slatko B.E.,Gardner A.F.,Ausubel F.M.(2018)Overview of Next-Generation Sequencing Technologies.Curr Protoc Mol Biol.122(1):e59. Warby,S.C.,Montpetit,A.,Hayden,A.R.,Carroll,J.B.,Butland,S.L.,Visscher,H.,Collins,J.A.,Semaka,A.,Hudson,T.J.and Hayden,M.R.(2009)CAG expansion in the Huntington disease gene is associated with a specific and targetable predisposing haplogroup.The American Journal of Human Genetics,84(3):351-366.
Claims
1. A method for phasing at least one desired distal single nucleotide polymorphism and short tandem repeat within the same target gene locus of nucleic acids isolated from a biological sample, a. An amplification step is performed, which includes contacting genomic DNA isolated from the biological sample containing the target gene locus with a first set of oligonucleotide primers to produce a first amplification product containing at least one target single nucleotide polymorphism and at least one exon-reference single nucleotide polymorphism within the same gene locus, b. A reverse transcription and amplification step is carried out, which includes contacting mRNA isolated from the biological sample containing the target gene locus with a second set of oligonucleotide primers to produce a second amplification product containing the short tandem repeat and the at least one exon-reference single nucleotide polymorphism, c. Determining the nucleic acid sequences of the first amplification product and the second amplification product, d. Aligning the nucleic acid sequences of the first amplification product and the second amplification product determined in step (c) based on the position of the at least one exon-reference single nucleotide polymorphism, e. To determine the haplotype of at least one target distal single nucleotide polymorphism and the short tandem repeat in the sample. Methods that include...
2. A method for phasing at least one target distal single nucleotide polymorphism and at least one target exonal single nucleotide polymorphism within the same target gene locus of nucleic acids isolated from a biological sample, a. An amplification step is performed, which includes contacting genomic DNA isolated from the biological sample containing the target gene locus with a first set of oligonucleotide primers to produce a first amplification product containing at least one target single nucleotide polymorphism and at least one exon-reference single nucleotide polymorphism within the same gene locus, b. A reverse transcription and amplification step is performed, which includes contacting mRNA isolated from the biological sample containing the target gene locus with a second set of oligonucleotide primers to produce a second amplification product containing the at least one target exon single nucleotide polymorphism and the at least one exon reference single nucleotide polymorphism. c. Determining the nucleic acid sequences of the first amplification product and the second amplification product, d. Aligning the nucleic acid sequences of the first amplification product and the second amplification product determined in step (c) based on the position of the at least one exon-reference single nucleotide polymorphism, e. Determining the haplotypes of the at least one target distal single nucleotide polymorphism and the at least one target exon single nucleotide polymorphism in the sample. Methods that include...
3. The method according to any one of claims 1 and 2, wherein the at least one target distal single nucleotide polymorphism is an intron single nucleotide polymorphism.
4. The method according to claim 3, wherein the intron having the at least one target distal single nucleotide polymorphism is located adjacent to the exon having the at least one exon reference single nucleotide polymorphism within the same locus of the isolated genomic DNA.
5. The method according to any one of claims 1 to 4, wherein the amplification step (a) and the reverse transcription and amplification step (b) are carried out in separate reaction vessels.
6. The method according to any one of claims 1 to 5, wherein the nucleic acid sequences of the first amplification product and the second amplification product in step (c) are determined using long-read sequencing.
7. The method according to any one of claims 1 to 6, wherein the biological sample is selected from the group consisting of tissue, tumor tissue, blood, saliva, and cell lines derived from an individual.
8. The method according to any one of claims 1 to 7, further comprising the step of aligning the nucleic acid sequences of the first amplification product and the second amplification product to the nucleic acid sequence of the reference gene of the target locus, or the nucleic acid sequence of the complete human genome or a portion thereof containing the reference gene of the target locus, before aligning the nucleic acid sequences of the first amplification product and the second amplification product in step (d).
9. The method according to any one of claims 1, 3 to 8, wherein the target gene locus is the huntingtin gene.
10. The method according to claim 9, wherein the short tandem repeat is a CAG repeat.
11. The method according to any one of claims 9 to 10, wherein the single nucleotide polymorphism for at least one purpose is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804.
12. An in vitro method for diagnosing whether an individual is at risk of developing Huntington's disease, a. Determining the haplotype of at least one target distal single nucleotide polymorphism and short tandem repeat in an in vitro sample obtained from the individual using the method according to any one of claims 1, 3 to 11, b. Determining the risk of the individual developing Huntington's disease based on the haplotype of the at least one target distal single nucleotide polymorphism and CAG short tandem repeat determined in step (a). In vitro methods, including [specific methods].
13. An in vitro method for identifying patients with Huntington's disease who are likely to respond to therapy targeting at least one target single nucleotide polymorphism, a. Determining the haplotype of at least one target distal single nucleotide polymorphism and short tandem repeat in an in vitro sample obtained from an individual using the method according to any one of claims 1, 3 to 11, b. Identifying the patient as likely to respond to the therapy based on the haplotype determined in step (a) and the determination of the presence of a specific allele of the at least one target single nucleotide polymorphism. In vitro methods, including [specific methods].
14. The method according to claim 13, wherein the therapy targeting the at least one target single nucleotide polymorphism is an antisense oligonucleotide treatment aimed at silencing an RNA molecule containing the specific allele of the at least one target single nucleotide polymorphism.
15. A kit for determining the nucleic acid sequences of at least one target distal single nucleotide polymorphism and short tandem repeat within the same target gene locus of nucleic acids isolated from a biological sample, - A first set of oligonucleotide primers for generating a first amplification product comprising the at least one target single nucleotide polymorphism and at least one exon-reference single nucleotide polymorphism within the same locus; - A second set of oligonucleotide primers that generate a second amplification product comprising the at least one target exon single nucleotide polymorphism and the at least one exon reference single nucleotide polymorphism. A kit that includes this.
16. Antisense oligonucleotides that specifically hybridize to at least one target distal single nucleotide polymorphism in the huntingtin (HTT) gene for use in treating patients having Huntington's disease, wherein the patient is selected for treatment when determining a specific haplotype of the at least one target distal single nucleotide polymorphism and short tandem repeat detected in a biological sample of the patient, and the specific haplotype is determined using the method according to any one of claims 1, 3 to 11.
17. In vitro use of haplotype determination of at least one target distal single nucleotide polymorphism and short tandem repeat detected in the huntingtin (HTT) gene determined in a biological sample of an individual for the diagnosis of Huntington's disease, wherein the detection of a specific haplotype of the at least one target distal single nucleotide polymorphism and the short tandem repeat detected in a biological sample of a patient indicates that the individual has Huntington's disease, and the specific haplotype is determined using the method according to any one of claims 1, 3 to 11.
18. In vitro use of haplotype determination of at least one target distal single nucleotide polymorphism and a short tandem repeat detected in the huntingtin (HTT) gene determined in a biological sample of the patient, for determining whether a patient with Huntington's disease is likely to respond to a therapy comprising an antisense oligonucleotide that specifically hybridizes to at least one target distal single nucleotide polymorphism in the huntingtin (HTT) gene, wherein if the specific haplotype of the at least one target distal single nucleotide polymorphism and the short tandem repeat is detected in the biological sample of the patient, the patient is identified as being more likely to respond to the therapy, and the specific haplotype is determined using the method of any one of claims 1, 3 to 11.