New assays for phasing remote genomic loci with zygotic resolution via long read length sequencing mixed data analysis
By combining genomic DNA and cDNA analysis with exon reference SNPs, long-read sequencing technology was used to phase distant loci of Huntington's disease-related genes, solving the problem of inaccurate phasing in existing technologies and enabling accurate diagnosis and personalized treatment of hereditary diseases.
Patent Information
- Application Number
- CN202480022247.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-31
- Filing Date
- 2024-03-28
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies struggle to accurately phase distant loci at the single nucleotide level, particularly in the localization of CAG repeat sequences and SNPs in Huntington's disease-related genes. This leads to interruptions in phasing information in regions with low genetic variation, and existing methods are complex and prone to errors.
A hybrid analysis approach was adopted, which involved the mixed analysis of long amplicon generated from genomic DNA and cDNA. Using DNA and cDNA as starting materials, and combining exon reference SNPs, phase determination of exon and intron target SNPs was performed. Nucleic acid sequences were then determined using long-read sequencing technologies such as PacBio and Oxford Nanopore Technologies' nanopore sequencing.
It has enabled accurate diagnosis and personalized treatment of genetic diseases such as Huntington's disease, provided safe and effective treatment options, improved our understanding of disease mechanisms, and promoted the development of new therapies.
Smart Images

Figure CN120917149A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure provides novel molecular biology assays for preclinical or clinical biomarker characterization, for example, in the field of neuroscience (e.g., Huntington's Disease, HD). The disclosed methods and kits can be used as companion diagnostic tools, where two or more paired loci need to be identified to provide essential information for safe and effective stratification of patients receiving a particular therapy or drug treatment. BACKGROUND
[0002] Huntington's disease (HD) is a rare genetic disorder that causes progressive degeneration of brain cells. The HD disorder is caused by an abnormal expansion of short tandem repeats (STRs) in exon 1 of the Huntingtin gene on chromosome 4. These repeat sequences are composed of cytosine-adenine-guanine (CAG) nucleotide sequences. A normal Huntingtin gene typically has 10 to 35 CAG repeat sequences, but in individuals with Huntington's disease, the number of CAG repeat sequences can range from 36 to over 100. CAG expansions can occur spontaneously or be inherited in an autosomal dominant manner. If a parent has a CAG expansion in the Huntingtin gene, there is a 50% chance that they will pass it on to their children. Individuals with CAG expansions, typically over 40 repeat sequence units, have a 100% risk of developing Huntington's disease during their lifetime. As the number of CAG repeat sequences increases, the age of symptom onset decreases, and the severity of the disease increases. The exact mechanism by which CAG expansions cause Huntington's disease is not fully understood, but it is believed that the expanded CAG repeat sequences cause the Huntingtin protein to misfold, leading to the accumulation of toxic protein aggregates that damage brain cells, causing cell death. Irregular expansions of STRs, such as CAG repeat expansions, are not unique to Huntington's disease (HD) and have been associated with several other genetic disorders, including spinocerebellar ataxias, myotonic dystrophy, and several forms of muscular dystrophy (Hannan, 2018). Several SNPs have been identified in the Huntingtin gene that alter the risk of developing Huntington's disease, as well as causing changes in age of onset and severity of symptoms (Claassen, D.O., et al., 2020; Bečanović, K., et al., 2015). For example, several SNPs in the Huntingtin gene, known as rs13102260, rs362277, rs3025814, rs2530596, have been associated with age of onset (Kartsaki, E., et al., 2006; Kay, C., et al., 2015; Ramos, E.M., et al., 2012). In addition, certain SNP alleles at a polymorphic site, such as, for example, rs7685686, rs362331, rs6446723, rs6844859, rs363080s rs363125, rs362307, rs362273, have been identified to enable selective treatment of HD patients while also allowing non-selective treatment of all remaining patients (Kay, C., et al. 2015; Shin et al. 2022).Such intronic and exonic SNPs can be targeted using antisense oligonucleotide (ASO)-based therapies to selectively target pre-mRNA or messenger RNA (mRNA), thereby altering the mRNA and, in turn, protein expression through various mechanisms.
[0003] Advances in long-read sequencing technologies, such as sequencing platforms based on long-read technology (e.g., commercially available from Oxford Nanopore Technologies Ltd. (Oxford, UK) and Pacific Biosciences of California, Inc. (Menlo Park, USA)), have made it possible to accurately detect and characterize SNPs in the huntingtin gene and other genes associated with Huntington’s disease, as well as CAG repeat expansions. Identifying SNPs associated with Huntington’s disease and related conditions is an important area of research, as it can enable the development of new disease management and treatment methods. For example, SNPs can provide targets for developing new drugs or therapies that can modulate the expression or activity of huntingtin protein, or for developing gene editing methods that can modify the huntingtin gene to prevent or reverse disease progression. Overall, identifying and characterizing SNPs associated with Huntington’s disease is an important area of research that has the potential to improve our understanding of the disease and drive the development of new treatments and therapies. Huntington’s disease itself manifests as the loss of GABAergic medium spiny (GABA MS) neurons in the striatum and is caused by a CAG repeat expansion in the first exon of the huntingtin gene. Within cells, this wild-type protein can be involved in chemical signaling, transporting substances, attaching (binding) proteins, and other structures, and protecting cells from self-destruction (apoptosis). Symptoms of Huntington’s disease typically begin in middle age, but they can also occur early or late in life. Early symptoms include involuntary movements, such as jerking or twitching, as well as difficulty with coordination and balance. As the disease progresses, symptoms become more severe and can include cognitive impairment, mood swings, and behavioral changes. The time from initial symptoms to death is typically about 10 to 30 years. There is currently no cure, but treatments can alleviate symptoms and provide support. Huntingtin protein is present in many body tissues (e.g., liver), with the highest level of activity in the brain. Additionally, genetic counseling and testing can help individuals and families understand their risk of developing the disease and make informed decisions about family planning.
[0004] DNA sequencing is a fundamental tool in biological and medical research and is particularly important for the personalized medicine paradigm. To ultimately achieve the $1,000 genome goal, various new DNA sequencing methods have been investigated; the primary method is sequencing by synthesis (SBS), which is a method of determining short DNA sequences during a polymerase reaction (Slatko et al., 2018). PacBio sequencing, also known as SMRT (single molecule real time) sequencing, enables sequencing of ultra-long fragments up to 30 to 50 kb. The SMRT method involves binding an engineered DNA polymerase to the bottom of a zero-mode waveguide (ZMW) well, where a DNA library linked with SMRT-bell adaptors is loaded onto the DNA polymerase. The four nucleotides are labeled with different phosphonate-linked fluorophores for differential detection. Imaging occurs on a millisecond time scale when the correct fluorescently labeled nucleotide is incorporated into the complementary strand of the single-stranded DNA molecule being sequenced as the nucleotide is incorporated into the growing strand. After each dNTP incorporation, the phosphonate-linked fluorescent moiety is released and dissipates from the detection area and cannot be detected again. The next nucleotide can then be incorporated. Here, the rate of nucleotide incorporation is timed with imaging so that each base is identified as it is incorporated into the growing DNA strand (Slatko et al., supra). Nanopore-based DNA sequencing was first proposed in the late 1990s and has recently been commercialized by Oxford Nanopore Technologies Ltd. (Oxford, UK), where a protein nanopore is embedded in a resistive bilayer membrane, through which the current (pico-amp, pA) undergoes characteristic changes as each nucleotide passes through the detector, allowing for short, long, and ultra-long read lengths from 0.1 kb up to 1 Mb. More specifically, long dsDNA molecules are first attached to a motor enzyme and tether molecule, which allows the library to be precipitated onto the sequencing flow cell and increases the proximity of the molecules to the nanopore and simultaneously increases the amount of molecules available for analysis. When the guide strand containing the motor complex encounters an available nanopore, the template of single-stranded DNA (ssDNA) enters the nanopore and disrupts the pA baseline reading via changes. Here, the translocation rate is modulated by the nucleotide sequence and accompanying epigenetic modifications. The motor enzyme enables the DNA to slow the progression through the channel and improves the quality of the raw data. Each nucleotide k-mer present in the nanopore provides a characteristic electronic pattern recorded in real time as a current disruption event (Slatko et al., supra), and is subsequently called to a standardized FASTQ / FASTA dataset.
[0005] Well-established short-read sequencing by synthesis (SBS) platforms (such as, for example, the MiSeq or NextSeq Series by Illumina, Inc., San Diego, CA, USA or the Ion GeneStudio Systems by Thermo Fisher, Waltham, MA, USA) are commonly used for SNP genotyping, and several algorithms exist for resolving haplotypes based on SNP data. However, such methods are limited in many ways as they are primarily designed for haplotype phasing of whole genome assemblies and rely on existing population-based reference panels. This is a statistical approach that is never 100% accurate and more complex genetic variants such as STRs are not included in the reference panels. Furthermore, regions of small genetic variation will cause breaks in phasing information between distant loci. New sequencing platforms based on long-read technology (e.g., available from Oxford Nanopore Technologies Ltd., Oxford, UK and Pacific Biosciences of California, Inc., Menlo Park, USA) can help directly resolve SNP phasing without the need for statistical inference, but the high error rate of Oxford Nanopore sequencing and the shorter average read length of PacBio (~20 kb) have limitations in deconvoluting > 150 kb of distant regions at the single nucleotide level in an accurate, low-cost, and fast manner. To overcome the problem of phasing distant SNPs as well as CAG repeat sequences, Asuragen, Inc. (Austin, TX, USA) is developing a companion diagnostic test using AmplideX® PCR technology that size adjusts and phases the HTT CAG repeat sequence as well as two different SNPs targeted by Wave’s WVE-120101 and WVE-120102 research treatment plans. However, their approach of using RNA as starting material to reduce the distance between the CAG repeat sequence and the exonic SNPs cannot be used for intronic SNPs.Alternative methods for phasing distant SNPs and CAG repeat sequences can include oligonucleotide-based hybrid capture along the sequence of interest followed by PCR amplification of the full-length gene via biotinylated probes (such as e.g. #101341 Twist Custom Panel Plus or #102989 Twist Human Custom Comprehensive Exome, available from Twist Biosciences, South San Francisco, US). This method requires genomic DNA material as input and an enrichment step after whole genome amplification. However, the extraction of very long genomic DNA fragments (e.g. 15-30 kb nt) is still difficult and this method can lead to fragmentation of the enriched material and loss of the desired long signals required for phasing distant loci, especially for regions with low genetic variation. Another alternative method for phasing two distant loci is provided in WO2018 / 022473, which comprises amplifying two genomic regions each comprising the loci of interest using primer sets, thereby generating sticky-end amplification products, followed by a ligation procedure. Here, one type of nucleic acid is used as template to determine the loci of interest (e.g. chromosome or fragment thereof, genomic DNA or mRNA / cDNA). Due to the random ligation of the amplification products, the resulting ligated products cannot be directly used for haplotyping. Therefore, the method disclosed in WO2018 / 022473 requires the distribution of the nucleic acid templates into droplets such that only one nucleic acid molecule template is present in each droplet. PCR amplification and sticky-end ligation of the two different loci is performed in the droplets. Subsequently, a second amplification reaction is required for the correctly ligated products before the amplified ligated products are sequenced by next generation sequencing. Thus, this method is very laborious and complex and prone to phasing errors and artifacts due to the numerous processing steps. WO 2016 / 191380 describes the general idea of using heterozygous SNPs as reference to align sequences from one type of nucleic acid (i.e. DNA or RNA) for phasing SNPs from long read sequencing data. This method relies on having several heterozygous SNPs in one locus if the sequencing run needs to cover long distances. However, this can not always be possible depending on the target sequence to be analyzed and the position of the SNPs used for phasing.
[0006] Therefore, there is a need for a new method that overcomes such technical problems related to the necessity of high input genomic DNA, short reads, custom oligonucleotide development or high error profiles. SUMMARY
[0007] To address the shortcomings of current methods, the present disclosure provides mixed analysis of long amplicons generated from genomic DNA and cDNA (mRNA) for pairing long distance information via spatial linkage, i.e. phasing of two or more loci of interest. These methods can provide significant advances for applications diagnostics and are a prerequisite for allele-specific treatments for genetic diseases such as Huntington’s disease. Using both DNA and cDNA (reverse transcribed from mature mRNA) as starting material, using the mixed approach as disclosed herein, it is possible to phase short tandem repeat sequences or exonic single nucleotide polymorphisms and long distance target SNPs (intronic or exonic) via an exonic reference SNP using very few amplicons (e.g. at least one DNA and at least one cDNA amplicon). Herein, the method allows for analysis of heterozygous and homozygous exonic or intronic target SNPs. In particular, to best resolve phasing information, the exonic reference SNP should be heterozygous if the target SNPs are heterozygous.
[0008] These methods and kits are particularly suitable for the field of companion diagnostics, which involves the identification of loci to provide key information about the genomic status of a patient that can be used to safely and effectively match patients with specific treatments or drug therapies, thereby identifying and stratifying patients most likely to benefit from a particular drug treatment or therapy. In recent years, such applications have become increasingly important with the development of more targeted therapies. Furthermore, the methods and kits of the present disclosure can also be used for basic research, where it can help to identify genetic variations associated with various diseases and conditions. This can lead to a better understanding of the underlying mechanisms of diseases, which ultimately can lead to the development of new therapies and treatments.
[0009] In one aspect, a method of phasing at least one distal single nucleotide polymorphism of interest and a short tandem repeat sequence within the same target locus of nucleic acid isolated from a biological sample is provided, the method comprising: (a) performing an amplification step comprising contacting genomic DNA comprising the target locus isolated from the biological sample with a first set of oligonucleotide primers to produce a first amplification product comprising at least one single nucleotide polymorphism of interest and at least one exon reference single nucleotide polymorphism in the same locus; (b) performing a reverse transcription and amplification step comprising contacting mRNA comprising the target locus isolated from the biological sample with a second set of oligonucleotide primers to produce a second amplification product comprising the short tandem repeat sequence and the at least one exon reference single nucleotide polymorphism; (c) determining the nucleic acid sequence of the first amplification product and the second amplification product; (d) aligning the nucleic acid sequences of the first amplification product and the second amplification product determined in step (c) based on the location of the at least one exon reference single nucleotide polymorphism; and (e) determining the haplotype of the at least one distal single nucleotide polymorphism of interest and the short tandem repeat sequence in the sample. In one embodiment, the at least one distal single nucleotide polymorphism of interest is an intronic single nucleotide polymorphism. In particular embodiments, the intron carrying the at least one distal single nucleotide polymorphism of interest is adjacent to an exon carrying the at least one exon reference single nucleotide polymorphism within the same locus of the isolated genomic DNA. In one embodiment, the amplification step (a) and the reverse transcription and amplification step (b) are performed in separate reaction vessels. In another embodiment, the nucleic acid sequences of the first amplification product and the second amplification product in step (c) are determined using long-read sequencing, such as PacBio sequencing, also known as SMRT (single molecule real-time) sequencing, or nanopore-based DNA long-read sequencing commercialized by Oxford Nanopore Technologies. In another embodiment, the biological sample is selected from the group consisting of tissue, tumor tissue, blood, saliva, and a cell line derived from an individual. In one embodiment, when the nucleic acid sequences of the first amplification product and the second amplification product are aligned in step (d), if the at least one distal single nucleotide polymorphism of interest is heterozygous, then the at least one exon reference single nucleotide polymorphism is heterozygous. In another embodiment, prior to aligning the nucleic acid sequences of the first amplification product and the second amplification product in step (d), the method further comprises a step of aligning the nucleic acid sequences of the first amplification product and the second amplification product to the nucleic acid sequence of a reference gene of the target locus or to the nucleic acid sequence of a complete human genome or to a portion thereof comprising the reference gene of the target locus. In another embodiment, the target locus is the Huntingtin gene.In particular embodiments, the short tandem repeat sequence is a CAG repeat sequence. In some embodiments, the at least one polymorphic site of interest is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphic site of interest is an intronic SNP selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphic site of interest is an exonic SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In particular embodiments, the at least one polymorphic site of interest is the intronic SNP rs7685686. In some embodiments, the at least one exonic reference SNP is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. Herein, an exonic reference SNP (such as, for example, rs362331, rs362273, rs34315806, and rs363099) is not the same as (i.e., does not correspond to) the at least one polymorphic site of interest. In some embodiments, if the at least one polymorphic site of interest is an intronic SNP, the at least one exonic reference SNP is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, the at least one exonic reference SNP is rs362331. In certain embodiments, determining the haplotype of the short tandem repeat sequence comprises determining the number of short tandem repeat sequence units. In certain embodiments, the mRNA is a spliced mRNA (mature mRNA).
[0010] In another aspect, there is provided an in vitro method for diagnosing whether an individual is at risk of developing Huntington’s disease, the method comprising: (a) determining the haplotype of at least one polymorphic site of interest and a short tandem repeat sequence in an in vitro sample obtained from the individual by performing an amplification step comprising contacting genomic DNA comprising the target locus isolated from the biological sample with a first set of oligonucleotide primers to produce a first amplification product comprising at least one polymorphic site of interest and at least one exon reference single nucleotide polymorphism in the same locus; performing a reverse transcription and amplification step comprising contacting mRNA comprising the target locus isolated from the biological sample with a second set of oligonucleotide primers to produce a second amplification product comprising the short tandem repeat sequence and the at least one exon reference single nucleotide polymorphism; determining the nucleic acid sequences of the first amplification product and the second amplification product; aligning the nucleic acid sequences of the first amplification product and the second amplification product determined in the previous step based on the position of the at least one exon reference single nucleotide polymorphism; and determining the haplotype of at least one polymorphic site of interest and a short tandem repeat sequence in the sample; (b) determining the risk of the individual developing Huntington’s disease based on the haplotype of at least one polymorphic site of interest and a CAG short tandem repeat sequence determined in step (a). In one embodiment, the at least one polymorphic site of interest is an intronic single nucleotide polymorphism. In a particular embodiment, the intron carrying the at least one polymorphic site of interest is adjacent to an exon carrying the at least one exon reference single nucleotide polymorphism within the same locus of the isolated genomic DNA. In one embodiment, the amplification step and the reverse transcription and amplification step are performed in separate reaction vessels. In another embodiment, the nucleic acid sequences of the first amplification product and the second amplification product are determined using long-read sequencing, such as PacBio sequencing, also known as SMRT (single molecule real-time) sequencing, or nanopore-based DNA long-read sequencing commercialized by Oxford Nanopore Technologies. In another embodiment, the biological sample is selected from the group consisting of tissue, tumor tissue, blood, saliva, and a cell line derived from the individual. In one embodiment, when aligning the nucleic acid sequences of the first amplification product and the second amplification product, if the at least one polymorphic site of interest is heterozygous, then the at least one exon reference single nucleotide polymorphism is heterozygous. In another embodiment, prior to aligning the nucleic acid sequences of the first amplification product and the second amplification product, the method further comprises a step of aligning the nucleic acid sequences of the first amplification product and the second amplification product to the nucleic acid sequence of the reference gene of the target locus or to the nucleic acid sequence of the complete human genome or to the portion thereof comprising the reference gene of the target locus.In another embodiment, the target locus is the Huntington gene. In particular embodiments, the short tandem repeat sequence is a CAG repeat sequence. In certain embodiments, the at least one polymorphic site of interest is an intronic SNP selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphic site of interest is an exonic SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In particular embodiments, the at least one polymorphic site of interest is the intronic SNP rs7685686. In some embodiments, the at least one exonic reference SNP is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. Herein, an exonic reference SNP (such as, for example, rs362331, rs362273, rs34315806, and rs363099) is not the same as (i.e., does not correspond to) the at least one polymorphic site of interest. In some embodiments, if the at least one polymorphic site of interest is an intronic SNP, the at least one exonic reference SNP is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, the at least one exonic reference SNP is rs362331. In certain embodiments, determining the haplotype of the short tandem repeat sequence comprises determining the number of units of the short tandem repeat sequence. In certain embodiments, the mRNA is a spliced mRNA (mature mRNA).
[0011] In another aspect, an in vitro method of identifying a patient having Huntington’s disease as likely to respond to a therapy targeting at least one polymorphic site of interest, the method comprising: (a) determining a haplotype of at least one polymorphic site of interest and a short tandem repeat sequence in an in vitro sample obtained from the individual by performing an amplification step comprising contacting genomic DNA isolated from the biological sample comprising the target locus with a first set of oligonucleotide primers to produce a first amplification product comprising at least one polymorphic site of interest and at least one exonic reference single nucleotide polymorphism in the same locus; performing a reverse transcription and amplification step comprising contacting mRNA isolated from the biological sample comprising the target locus with a second set of oligonucleotide primers to produce a second amplification product comprising the short tandem repeat sequence and the at least one exonic reference single nucleotide polymorphism; determining the nucleic acid sequence of the first amplification product and the second amplification product; aligning the nucleic acid sequences of the first amplification product and the second amplification product determined in the previous step based on the location of the at least one exonic reference single nucleotide polymorphism; and determining the haplotype of at least one polymorphic site of interest and a short tandem repeat sequence in the sample; (b) identifying the patient as likely to respond to the therapy based on the haplotype determined in step (a) and a determination of the presence of a particular allele of at least one polymorphic site of interest. In some embodiments, the therapy targeting at least one polymorphic site of interest is an antisense oligonucleotide therapy intended to inhibit an RNA molecule comprising a particular allele of at least one polymorphic site of interest. In one embodiment, the at least one polymorphic site of interest is an intronic single nucleotide polymorphism. In particular embodiments, the intron carrying the at least one polymorphic site of interest is adjacent to an exon carrying the at least one exonic reference single nucleotide polymorphism within the same locus of the isolated genomic DNA. In one embodiment, the amplification step (a) and the reverse transcription and amplification step (b) are performed in separate reaction vessels. In another embodiment, long-read sequencing, such as PacBio sequencing, also known as SMRT (single molecule real-time) sequencing, or nanopore-based DNA long-read sequencing commercialized by Oxford Nanopore Technologies, is used to determine the nucleic acid sequences of the first amplification product and the second amplification product in step (c). In another embodiment, the biological sample is selected from the group consisting of tissue, tumor tissue, blood, saliva, and a cell line derived from the individual. In one embodiment, when aligning the nucleic acid sequences of the first amplification product and the second amplification product, if the at least one polymorphic site of interest is heterozygous, then the at least one exonic reference single nucleotide polymorphism is heterozygous.In another embodiment, prior to aligning the nucleic acid sequences of the first and second amplification products, the method further comprises the step of aligning the nucleic acid sequences of the first and second amplification products to the nucleic acid sequence of a reference gene of the target locus or to the nucleic acid sequence of the complete human genome or to a portion thereof that contains the reference gene of the target locus. In another embodiment, the target locus is the Huntingtin gene. In particular embodiments, the short tandem repeat sequence is a CAG repeat sequence. In some embodiments, the at least one polymorphic single nucleotide of interest is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphic single nucleotide of interest is an intronic SNP selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphic single nucleotide of interest is an exonic SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In particular embodiments, the at least one polymorphic single nucleotide of interest is the intronic SNP rs7685686. In some embodiments, the at least one exonic reference single nucleotide polymorphism is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. Herein, an exonic reference single nucleotide polymorphism (such as, for example, rs362331, rs362273, rs34315806, and rs363099) is not the same (i.e., does not correspond) to the at least one polymorphic single nucleotide of interest. In some embodiments, if the at least one polymorphic single nucleotide of interest is an intronic SNP, the at least one exonic reference single nucleotide polymorphism is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, the at least one exonic reference single nucleotide polymorphism is rs362331. In certain embodiments, determining the haplotype of the short tandem repeat sequence comprises determining the number of units of the short tandem repeat sequence. In certain embodiments, the mRNA is a spliced mRNA (mature mRNA).
[0012] In another aspect, a kit for determining nucleic acid sequences of at least one distant single nucleotide polymorphism and short tandem repeat sequence of interest within the same target locus of a nucleic acid isolated from a biological sample is provided, the kit comprising a first set of oligonucleotide primers for generating a first amplification product comprising at least one single nucleotide polymorphism of interest and at least one exon reference single nucleotide polymorphism in the same locus; and a second set of oligonucleotide primers for generating a second amplification product comprising at least one exon single nucleotide polymorphism of interest and the at least one exon reference single nucleotide polymorphism. Herein, the kit is suitable for use in carrying out any of the methods disclosed herein. In a particular embodiment, the first set of oligonucleotide primers comprises an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 1 and an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 2. In another particular embodiment, the second set of oligonucleotide primers comprises an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 3, an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 4, and an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 5. In some embodiments, the kit further includes at least one of nucleoside triphosphates, a nucleic acid polymerase, and buffers necessary for nucleic acid polymerase and / or reverse transcriptase function. In some embodiments, the kit further includes any of the reagents for amplification, such as reverse transcriptase, DNA polymerase, dNTPs, buffers, and / or other components suitable for reverse transcription and / or amplification (e.g., cofactors or aptamers). Typically, the reagent mixtures are concentrated so that aliquots are added to the final reaction volume along with the sample (e.g., RNA or DNA), enzymes, and / or water. In some embodiments, the kit further includes a reverse transcriptase (or enzyme with reverse transcriptase activity) and / or a DNA polymerase (e.g., a thermostable DNA polymerase such as Taq, Z05, and derivatives thereof).
[0013] In another aspect, a method of phased at least one distal single nucleotide polymorphism of interest and at least one exonic single nucleotide polymorphism of interest within a same target locus of a nucleic acid isolated from a biological sample, the method comprising: (a) performing an amplification step comprising contacting genomic DNA comprising the target locus isolated from the biological sample with a first set of oligonucleotide primers to produce a first amplification product comprising at least one single nucleotide polymorphism of interest and at least one exonic reference single nucleotide polymorphism in the same locus; (b) performing a reverse transcription and amplification step comprising contacting mRNA comprising the target locus isolated from the biological sample with a second set of oligonucleotide primers to produce a second amplification product comprising at least one exonic single nucleotide polymorphism of interest and the at least one exonic reference single nucleotide polymorphism; (c) determining the nucleic acid sequence of the first amplification product and the second amplification product; (d) aligning the nucleic acid sequences of the first amplification product and the second amplification product determined in step (c) based on the location of the at least one exonic reference single nucleotide polymorphism; and (e) determining the haplotype of at least one distal single nucleotide polymorphism of interest and at least one exonic single nucleotide polymorphism of interest in the sample. In one embodiment, the at least one distal single nucleotide polymorphism of interest is an intronic single nucleotide polymorphism. In a particular embodiment, the intron carrying the at least one distal single nucleotide polymorphism of interest is adjacent to an exon carrying the at least one exonic reference single nucleotide polymorphism within the same locus of the isolated genomic DNA. In one embodiment, the amplification step (a) and the reverse transcription and amplification step (b) are performed in separate reaction vessels. In another embodiment, the nucleic acid sequences of the first amplification product and the second amplification product in step (c) are determined using long-read sequencing, such as PacBio sequencing, also known as SMRT (single molecule real-time) sequencing, or nanopore-based DNA long-read sequencing commercialized by Oxford Nanopore Technologies. In another embodiment, the biological sample is selected from the group consisting of tissue, tumor tissue, blood, saliva, and a cell line derived from an individual. In one embodiment, when aligning the nucleic acid sequences of the first amplification product and the second amplification product in step (d), if the at least one distal single nucleotide polymorphism of interest is heterozygous, then the at least one exonic reference single nucleotide polymorphism is heterozygous. In another embodiment, prior to aligning the nucleic acid sequences of the first amplification product and the second amplification product in step (d), the method further comprises a step of aligning the nucleic acid sequences of the first amplification product and the second amplification product to the nucleic acid sequence of a reference gene of the target locus or to the nucleic acid sequence of the complete human genome or to a portion thereof comprising the reference gene of the target locus. In another embodiment, the target locus is the Huntingtin gene.In some embodiments, the at least one SNP of interest is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, the at least one SNP of interest is an intronic SNP selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, the at least one SNP of interest is an exonic SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In particular embodiments, the at least one SNP of interest is the intronic SNP rs7685686. In some embodiments, the at least one exonic reference SNP is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. Herein, an exonic reference SNP (such as, for example, rs362331, rs362273, rs34315806, and rs363099) is not the same as (i.e., does not correspond to) the at least one SNP of interest. In some embodiments, if the at least one SNP of interest is an intronic SNP, the at least one exonic reference SNP is selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In certain embodiments, the at least one exonic reference SNP is rs362331. In certain embodiments, the mRNA is a spliced mRNA (mature mRNA).
[0014] In another aspect, an antisense oligonucleotide is provided for use in treating a patient having Huntington’s disease, the antisense oligonucleotide specifically hybridizing to at least one polymorphism of interest distant single nucleotide polymorphism in the Huntingtin (HTT) gene, wherein the patient is selected for treatment when a particular haplotype of the at least one polymorphism of interest distant single nucleotide polymorphism and a CAG short tandem repeat sequence detected in a biological sample of the patient is determined. In some embodiments, the at least one polymorphism of interest distant single nucleotide polymorphism is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphism of interest single nucleotide polymorphism is an intronic SNP selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphism of interest single nucleotide polymorphism is an exonic SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In particular embodiments, the at least one polymorphism of interest distant single nucleotide polymorphism is rs7685686. In some embodiments, the particular haplotype comprises the presence of an A allele for rs7685686, and a number of CAG short tandem repeat sequences higher than 36. In other embodiments, the particular haplotype comprises the presence of a G allele for rs7685686, and a number of CAG short tandem repeat sequences higher than 36. In some embodiments, the antisense oligonucleotide specifically hybridizes to a region in the Huntingtin (HTT) gene comprising an A allele for rs7685686. In some embodiments, the antisense oligonucleotide is selected to specifically hybridize to the polymorphism of interest distant single nucleotide polymorphism in an allele-specific manner. In some embodiments, the particular haplotype is determined using the methods disclosed herein.
[0015] In another aspect, there is provided an in vitro use of at least one polymorphism of interest detected in the Huntingtin (HTT) gene and a haplotype of short tandem repeat sequences determined in a biological sample of an individual for diagnosing Huntington disease, wherein detection of a particular haplotype of at least one polymorphism of interest detected in the biological sample of the patient is indicative of the individual having Huntington disease. In some embodiments, the at least one polymorphism of interest is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphism of interest is an intronic SNP selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphism of interest is an exonic SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In particular embodiments, the at least one polymorphism of interest is the intronic SNP rs7685686. In some embodiments, the particular haplotype comprises the presence of an A allele for rs7685686 and a number of CAG short tandem repeat sequences higher than 36. In other embodiments, the particular haplotype comprises the presence of a G allele for rs7685686 and a number of CAG short tandem repeat sequences higher than 36. In some embodiments, the particular haplotype is determined using the methods disclosed herein.
[0016] In another aspect, detecting at least one polymorphism of interest in the Huntington (HTT) gene and haplotype of short tandem repeat sequences determined in a biological sample of a patient having Huntington's disease determines an in vitro use of the patient to identify the patient as likely to respond to a therapy comprising an antisense oligonucleotide that specifically hybridizes to at least one polymorphism of interest in the Huntington (HTT) gene, wherein the patient is identified as more likely to respond to the therapy when a particular haplotype of at least one polymorphism of interest and short tandem repeat sequences is detected in the patient's biological sample. In some embodiments, the at least one polymorphism of interest is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphism of interest is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphism of interest is an intronic SNP selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, the at least one polymorphism of interest is an exonic SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In particular embodiments, the at least one polymorphism of interest is the intronic SNP rs7685686. In some embodiments, the particular haplotype comprises the presence of an A allele for rs7685686, and the number of CAG short tandem repeat sequences is higher than 36. In some embodiments, the antisense oligonucleotide specifically hybridizes to a region in the Huntington (HTT) gene comprising an A allele for rs7685686.In other embodiments, the particular haplotype includes the presence of the G allele for rs7685686 and the number of CAG short tandem repeat sequences is higher than 36. In some embodiments, the antisense oligonucleotide is selected to specifically hybridize to the distal single nucleotide polymorphism of interest in an allele-specific manner. In some embodiments, the particular haplotype is determined using the methods disclosed herein.
[0017] Further disclosed is a method of treating a patient having Huntington's disease, the method comprising: (a) determining a haplotype of at least one distal single nucleotide polymorphism of interest and a short tandem repeat sequence in the Huntingtin (HTT) gene in a biological sample of the patient; and (b) administering an antisense oligonucleotide that specifically hybridizes to the at least one distal single nucleotide polymorphism of interest in the Huntingtin (HTT) gene. Herein, the at least one distal single nucleotide polymorphism of interest can be selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804. In certain embodiments, the at least one single nucleotide polymorphism of interest is an intronic SNP selected from the group consisting of rs7685686, rs363088, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs2298967, rs362271, rs3121419, and rs16843804. In certain embodiments, the at least one single nucleotide polymorphism of interest is an exonic SNP selected from the group consisting of rs362331, rs362273, rs34315806, and rs363099. In particular embodiments, the at least one single nucleotide polymorphism of interest is the intronic SNP rs7685686. Also disclosed is the determined haplotype includes the presence of the A allele for rs7685686 and the number of CAG short tandem repeat sequences is higher than 36. Also disclosed is the antisense oligonucleotide specifically hybridizes to a region in the Huntingtin (HTT) gene that includes the A allele for rs7685686. In other embodiments, the particular haplotype includes the presence of the G allele for rs7685686. In some embodiments, the antisense oligonucleotide is selected to specifically hybridize to the distal single nucleotide polymorphism of interest in an allele-specific manner. Herein, the particular haplotype can be determined using the methods disclosed herein.
[0018] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present application, suitable methods and materials are described below. Furthermore, the materials, methods, and examples are illustrative only and not intended to be limiting.
[0019] The details of one or more embodiments of the application are set forth in the accompanying drawings and the detailed description below. Other features, objects, and advantages of the application will be apparent from the drawings and detailed description, and from the claims. BRIEF DESCRIPTION OF DRAWINGS
[0020] FIG. 1 indicates the long distance separation between the tandem repeat sequence region (i.e., CAG) and the intronic SNP or exonic SNP of interest. The distance is represented on the genomic DNA level and can span up to 200 kb of nucleotides. The intronic SNP and the exonic SNP are separated from each other by only about 10 kb and can be amplified with PCR (i.e., F1 and R1 primers). The transcription and mRNA maturation of the gene target involves splicing and thereby reduces the distance between the CAG tandem repeat sequence and the exonic reference SNP, which allows reverse transcription and PCR (i.e., F2 and R2 primers) of about 10 kb amplicons from a single molecule.
[0021] FIG. 2 depicts two amplicons, a first amplicon generated from genomic DNA and containing the intronic target SNP and the exonic reference SNP; a second amplicon generated from cDNA derived from mature mRNA and containing the CAG tandem repeat sequence and the exonic reference SNP. The mixed analysis allows phasing of the long distance intronic SNP to the number of CAG tandem repeat sequences through alignment of the exonic SNP between the two amplicons.
[0022] Figure 3 shows phasing results for sample GM04282. Figure 3A cDNA data indicates that phasing of the mutant HTT allele with about 77.6 CAG ("high CAG") of the reads that account for about 22.7% shows exon T allele in about 95% of the reads. The wild type HTT allele with about 18.1 CAG ("low CAG") of the reads that account for about 77.2% also shows about 95% of the reads with exon T allele. Both alleles have >4% of the reads with undefined nucleotides at the SNP position of interest, which can be due to sequencing errors such as deletions. Since the exon SNP is homozygous, this has no reference value for the mixed analysis. The bar graph of the cDNA data of Figure 3B indicates that there are two distributions (i.e., low and high CAG repeat number), with mean values in the range of about 18 and about 77 repeats (Figure 3A). The exon SNP allele of both distributions is designated as thymine (Figure 3A), indicating homozygosity. Figure 3C DNA amplicon data indicates that the exon reference SNP allele T is in phase with the intron target SNP allele A in 100% of the reads. No other alleles have been detected, indicating homozygosity for the intron target SNP for both high and low CAG repeat regions. Therefore, this patient is not suitable for allele-specific treatment.
[0023] Figure 4 shows the phasing results for sample GM13503. Figure 4A cDNA data indicates that the mutant HTT allele with approximately 46.3 CAGs ("high CAG") that accounts for approximately 48.1% of reads shows exon T allele in approximately 99.3% of reads and C allele in only 0.7% of reads. The wild-type HTT allele with 18 CAGs ("low CAG") accounts for approximately 51.9% of reads and shows exon C allele in approximately 99.7% of reads and T allele in only 0.16% of reads. Both alleles have near 0% of reads with undefined nucleotides at the SNP position of interest. Since the exon SNP is heterozygous, this patient can be eligible for allele-specific treatment, but further identification of intronic SNPs with high CAG phasing is needed. Figure 4B bar graph of cDNA data indicates the two distributions (i.e., low and high CAG numbers), with mean values in the range of approximately 18 and approximately 46 repeats (Figure 4A). The exon SNP for high CAG number is assigned as thymine (Figure 4A), while the majority of reads for the low CAG allele are assigned as cytosine, indicating heterozygosity of the exon reference SNP. Figure 4C DNA amplicon data indicates that in >98% of reads, the exon reference SNP allele C is in phase with the intronic target SNP allele G, and in only 1.33% of reads with allele A. While in approximately 98.03% of reads, the exon reference SNP allele T is in phase with the intronic target SNP allele A, and in only 1.97% of reads with the intronic allele G. The low percentages (1.33% and 1.97%) of unrelated phasing results can be the result of PCR errors or chimeric molecules, which in turn create small noise in the final dataset. In summary, the mutant CAG allele (high CAG) is in phase with the exon reference SNP allele T, which in turn is in phase with the intronic target SNP allele A, making this patient eligible for A-specific ASO treatment.
[0024] Figure 5 shows phasing results for sample GM04724. Figure 5A cDNA data indicates that the mutant HTT allele with about 71.2 CAGs ("high CAG") that accounts for about 30.7% of reads shows the reference exon SNP T allele in about 96.68% of reads and the C allele in only about 0.55% of reads. The wild-type HTT allele with about 16 CAGs ("low CAG") that accounts for about 69.3% of reads shows about 98.16% of reads with the exon C allele and only 0.25% of reads with the T allele. Both alleles indicate 1.59% and 2.77% of reads at the SNP of interest position have undefined nucleotides; this can be due to errors such as insertions or deletions. Since the exon SNP is heterozygous, this patient can be eligible for allele-specific treatment, but further identification of intronic SNPs with the high CAG phasing is needed. Figure 5B bar graph of cDNA data indicates the two distributions (i.e., low and high CAG number), with mean values in the range of about 16 and about 71.2 repeats (Figure 5A). The exon SNP for the high CAG number is assigned as thymine (Figure 5A), while the majority of reads for the low CAG allele are assigned as cytosine, indicating heterozygosity of the exon reference SNP. Figure 5C DNA amplicon data indicates that in 97.54% of reads, the exon reference SNP allele C is in phase with the intronic target SNP allele G, and in only 2.46% of reads with the allele A. In 98.8% of reads, the exon reference SNP allele T is in phase with the intronic target SNP allele A, and in only 1.2% of reads with the intronic allele G. The low percentages (2.46% and 1.2%) of unrelated phasing results can be a result of PCR errors or chimeric molecules, which in turn create small noise in the final dataset. In summary, the mutant CAG allele (high CAG) is in phase with the exon reference SNP allele T, which in turn is in phase with the intronic target SNP allele A, making this patient eligible for A-specific ASO treatment, as shown in the example in Figure 4.
[0025] Figure 6 depicts the results of clustering reads with a similarity threshold of 1.0 (i.e., clusters formed from identical reads). The results represent two final contigs showing the presence or absence of variant alleles for intronic and exonic SNPs. Residual reads that did not cluster to the major haplotype group were omitted from the analysis due to the presence of low quality nucleotides, which resulted in sporadic errors at various nucleotide positions and lack of 100% similarity between core cluster closing reads.
[0026] Figure 7 represents the results of two final DNA amplicon clusters from the Integrative Genomics Viewer (IGV, Broad Institute) software, including exonic and intronic SNP positions for a heterozygous sample.
[0027] Figure 8 shows a flowchart of the bioinformatics analysis, starting from raw sequencing data until visualization of DNA / cDNA reads and phasing of CAG number with intronic SNPs.
[0028] Figure 9 shows a detailed flowchart of the bioinformatics analysis, starting from raw sequencing data until visualization of DNA / cDNA reads and phasing of CAG number with intronic SNPs.
[0029] Figure 10 shows a scheme for using the phase information from the present method to stratify and select patients that can be responsive to SNP-specific antisense oligonucleotide therapy. DETAILED DESCRIPTION
[0030] I. INTRODUCTION
[0031] Phasing short tandem repeats (STRs) within the same locus together with distant single nucleotide polymorphisms (SNPs) at the individual patient level is challenging. Standard methods such as PCR followed by short-read sequencing of genomic DNA are not suitable for such analysis if the genetic variations are too far apart (more than tens of kb). To overcome this problem, long-range PCR using mRNA as starting material for PCR amplification followed by long-read sequencing helps to bring exonic SNPs closer to distant STRs. However, this approach is not suitable for intronic SNPs. Therefore, there remains a need in the art for a reliable method to provide haplotype information for STRs and distant SNPs, particularly intronic SNPs, within the same locus. Furthermore, there remains a need for a reliable method to provide haplotype information for one or more exonic SNPs and one or more distant SNPs, particularly intronic SNPs, within the same locus.
[0032] To accomplish this task, the present disclosure provides a novel molecular biology assay for preclinical or clinical biomarker characterization, for example, in the field of neuroscience (e.g., Huntington’s Disease, HD). Such a method can be used as a companion diagnostic tool where the identification of two or more paired loci is required to provide essential information for safe and effective stratification of patients receiving a particular therapy or drug treatment. More particularly, the method allows for accurate determination of the spatial relationship (i.e., phasing) of a single nucleotide polymorphism (SNP) to a short tandem repeat sequence (STR) or another SNP from a very distant region in the genome (e.g., > 150 kb) at the haplotype genome level. The method is based on extraction of DNA and RNA and parallel amplification of DNA and cDNA obtained from a single sample (e.g., blood or tissue), followed by mixed analysis of data provided by long read sequencing technologies such as Pacific Biosciences of California, Inc. or Oxford Nanopore Technologies Limited.
[0033] Provided herein are methods of analyzing a sample to phase exon and / or intron SNP alleles with STRs or other distant SNPs. In certain aspects, provided are methods of analyzing a sample to phase exon and / or intron SNP alleles with trinucleotide CAG repeat sequences more than 150 kb away in location and using the phasing information for patient stratification (e.g., for inclusion of a patient in a clinical trial or for subjecting a patient to a therapeutic treatment). In particular aspects, the phasing information of a trinucleotide CAG repeat sequence region within the Huntingtin gene and at least one distant SNP such as, for example, rs7685686 can be used to determine whether a patient can benefit from antisense oligonucleotide (ASO) treatment targeting a particular distant SNP of Huntington’s Disease.
[0034] The present invention provides several advantages over current methods for SNP and STR genotyping. Most importantly, this new method allows for accurate determination of the spatial relationship between a SNP and a STR or other SNP from a distant region in the genome at the haplotype genome level, which is not possible in standard SNP and STR genotyping methods. This is particularly important for allele-specific therapeutic approaches. The goal of allele-specific therapies, for example, Huntington’s Disease, is to use SNP allele-specific antisense oligonucleotides (ASOs) to suppress the mutant (expanded CAG) RNA molecule and preserve the wild-type intact. Thus, the assays described herein are critical for identifying which SNP alleles are in the same haplotype (i.e., in phase) with the expanded CAG repeat sequence.
[0035] Overall design of PCR primers is generally required to develop robust and accurate assays. Non-specific amplification of random gene fragments can be problematic, which will increase the signal-to-noise ratio in the final sequencing data. However, alignment to human reference genome sequences only allows for consideration of target sequences in the final analysis. In addition, design of PCR primers and multiplexed assays must take into account linkage disequilibrium (LD) between intronic target SNPs and exonic reference SNPs to maximize the information content of the fingerprinting method (e.g. allele-specific therapy would not be possible if either of the SNPs were homozygous). LD structure is different between different populations and must be taken into account in assay design.
[0036] For the most difficult scenario where the target SNP is intronic and too far from the CAG repeat sequence to be amplified from genomic DNA by PCR, the workflow is as follows:
[0037] (1) Co-extract DNA and RNA from a single sample.
[0038] (2) Perform RT-PCR amplification using RNA as input material to cover the region of the CAG repeat sequence and a reference exonic SNP closer to the CAG repeat sequence. The reference exonic SNP should ideally be in high linkage disequilibrium with the target intronic SNP, or at least within the same allelic frequency range.
[0039] (3) Perform PCR amplification using genomic DNA as input material to cover the region of the above reference exonic SNP and the intronic target SNP.
[0040] (4) Sequence the amplicons using long read technology to ensure phasing information between variants within the amplicon is maintained.
[0041] (5) Perform mixed data analysis via alignment of DNA and cDNA amplicon sequences using the exonic reference SNP fingerprint and phasing of the intronic target SNP allele with the CAG repeat number.
[0042] Figure 1 indicates the long distance separation between the tandem repeat sequence region (i.e., CAG) and the intron SNP or exon SNP of interest. The distance is represented on the genomic DNA level and can span up to 200 kb of nucleotides. Since the target intron SNP is far away from the CAG repeat sequence (>150 kb), the locus including both cannot be PCR amplified using genomic DNA as a template. However, the intron target SNP and the exon reference SNP are only about 10 kb apart from each other and the locus can be amplified by PCR using genomic DNA as a template (i.e., using the set of oligonucleotide primers Fl and Rl). The transcription and mRNA maturation of the gene target is characterized by a splicing event that causes the distance between the CAG tandem repeat sequence and the exon reference SNP to decrease, which allows reverse transcription and PCR (i.e., using the set of oligonucleotide primers F2 and R2) of the 10 kb amplicon from a single molecule using mRNA as starting material.
[0043] Figure 2 depicts the two resulting amplicons, i.e., a first amplicon generated from genomic DNA and containing the intron target SNP and the exon reference SNP; and a second amplicon generated from cDNA derived from mature mRNA and containing the CAG tandem repeat sequence and the exon reference SNP. Herein, the exon reference SNP in both amplicons is identical. Thus, the mixed analysis allows the phasing of at least one long distance intron SNP to the number of CAG tandem repeat sequences by alignment of the exon reference SNP between the two amplicons. In some aspects, more than one exon reference SNP can be used to further improve the phasing accuracy.
[0044] II. DEFINITIONS
[0045] "Companion diagnostics" is a diagnostic test used as a companion to a therapeutic drug or treatment to determine its suitability for a particular patient or group of patients, and thereby stratify patients according to a certain genomic profile. Companion diagnostics refers to diagnostic tests developed alongside a particular drug or therapy (e.g., an antibody, small molecule, or antisense oligonucleotide) to select or exclude patients or groups of patients for treatment with that particular drug based on determining biological characteristics of responders and non-responders to the therapy. These tests often involve the identification of specific genetic biomarkers or mutations that are associated with the disease or condition being treated or that prospectively help predict a likely response or severe toxicity.
[0046] The term“biomarker” can refer to any detectable marker used to distinguish between individual samples, e.g., cancer versus non-cancer samples. Biomarkers include modifications (e.g., methylation of DNA, phosphorylation of proteins), differential expression, and mutations or variants (e.g., single nucleotide variations, insertions, deletions, splice variants, and fusion variants). Biomarkers can be detected in DNA, RNA, and / or protein samples.
[0047] The terms“nucleic acid,”“polynucleotide,” and“oligonucleotide” refer to polymers of nucleotides (e.g., ribonucleotides or deoxyribonucleotides) and include naturally occurring (e.g., adenosine, guanine, cytosine, uracil, and thymidine) and non-naturally occurring (human modified) nucleic acids. The term is not limited by the length of the polymer (e.g., the number of monomers) The triphosphonucleosides containing ribose as the sugar are generally abbreviated NTP, while the triphosphonucleosides containing deoxyribose as the sugar are abbreviated dNTP. Unless otherwise indicated, “nucleic acid” refers to any nucleic acid molecule, including but not limited to DNA, RNA, and hybrids thereof. In one embodiment, the nucleic acid bases forming the nucleic acid molecule can be the bases A, C, G, T, and U, as well as derivatives thereof (A - adenine; C - cytosine; G - guanine; T - thymine; U - uracil). “Derivatives” or“analogs” of these bases are well known in the art and are exemplified in PCR Systems, Reagents and Consumables (Perkin Elmer Catalog 1996-1997, Roche Molecular Systems, Inc., Branchburg, New Jersey, USA). Nucleic acids can be single-stranded or double-stranded and will typically contain 5’-3’ phosphodiester bonds, although in some cases nucleotide analogs can have other linkages. The monomers are often referred to as nucleotides. The term“non-natural nucleotide” or“modified nucleotide” refers to a nucleotide that contains a modified nitrogenous base, sugar, or phosphate group or incorporates a non-natural moiety in its structure. Examples of non-natural nucleotides include LNA, dideoxynucleotides, biotinylated nucleotides, aminated nucleotides, deaminated nucleotides, alkylated nucleotides, benzylated nucleotides, and fluorescently labeled nucleotides.“LNA” refers to locked nucleic acid. LNA is a modified RNA nucleotide in which the ribose moiety is modified by an extra bridge connecting the 2’ oxygen and the 4’ carbon. LNA nucleotides can be mixed with DNA or RNA residues at any position in an oligonucleotide and hybridize to DNA or RNA according to the Watson-Crick base pairing rules. The locked ribose conformation enhances hybridization properties (e.g., increases melting temperature).
[0048] A "nucleotide residue" is a single nucleotide in its state of existence after being incorporated and thereby becoming a monomer of a polynucleotide. Thus, a nucleotide residue is a nucleotide monomer of a polynucleotide (e.g., DNA) that is bound to an adjacent nucleotide monomer of the polynucleotide by a phosphodiester bond at the 3' position of its sugar, and to a second adjacent nucleotide monomer by its phosphate group, with the exception that (i) a 3' terminal nucleotide residue is bound to only one adjacent nucleotide monomer of the polynucleotide by a phosphodiester bond of its phosphate group, and (ii) a 5' terminal nucleotide residue is bound to only one adjacent nucleotide monomer of the polynucleotide by a phosphodiester bond of the 3' position of its sugar.
[0049] Due to well-known base pairing rules, determining the (base) identity of a dNTP analog (or rNTP analog) that is incorporated into a primer or DNA extension product (or RNA extension product) by measuring a unique electrical signal of a tag translocating through a nanopore, and thereby determining the identity of the incorporated dNTP analog (or rNPP analog), allows identification of a complementary nucleotide residue in a single-stranded polynucleotide hybridized to the primer or DNA extension product (or RNA extension product). Thus, if the incorporated dNPP analog comprises adenine, thymine, cytosine, or guanine, the complementary nucleotide residue in the single-stranded DNA is identified as thymine, adenine, guanine, or cytosine, respectively. Purine adenine (A) pairs with pyrimidine thymine (T). Pyrimidine cytosine (C) pairs with purine guanine (G). Similarly, with respect to RNA, if the incorporated rNPP analog comprises adenine, uracil, cytosine, or guanine, the complementary nucleotide residue in the single-stranded RNA is identified as uracil, adenine, guanine, or cytosine, respectively.
[0050] In the context of the present disclosure, the terms "cell-free nucleic acid", "cell-free DNA", "cell-free RNA", and like terms refer to a non-tissue sample (e.g., a liquid biopsy) from an individual that has been treated to substantially remove cells. Examples of non-tissue samples include blood and blood components, urine, saliva, tears, mucus, and the like.
[0051] A "pre-mRNA" or pre-mRNA is the primary molecule in the process of eukaryotic transcription, produced from a DNA template within the nucleus. The pre-mRNA molecule has both coding (exon) and non-coding (intron) sequences, and is undergoing maturation including a splicing step in which intron regions are removed and the molecule becomes mRNA after processing.
[0052] A“mature mRNA” is a eukaryotic RNA transcript that has been spliced and processed and is ready for translation in the process of protein synthesis. Unlike eukaryotic RNA immediately after transcription, which is referred to as pre-mRNA, a mature mRNA consists of exons only and has all introns removed.
[0053] A“single nucleotide polymorphism” (SNP) is a germline substitution of a single nucleotide at a particular position in the genome and is present at a population fraction large enough (1% or more). The genomic distribution of SNPs is not uniform; SNPs occur more frequently in non-coding regions (intron regions) than in coding regions (exon regions). As used herein, a“distant single nucleotide polymorphism” is located at a distance greater than 10 kb on the same nucleic acid relative to another single nucleotide polymorphism or short tandem repeat. In some cases, the distance can be greater than 20 kb, 30 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 110 kb, 120 kb, 130 kb, 140 kb, 150 kb, 160 kb, 170 kb, 180 kb, or 190 kb. In some cases, the distance can be between 10 and 200 kb, between 10 and 160 kb, between 20 and 200 kb, between 20 and 160 kb, between 50 and 200 kb, between 50 and 160 kb, between 100 and 200 kb, between 100 and 160 kb, or between 150 and 200 kb.
[0054] As all throughout this document, an“exonic reference single nucleotide polymorphism” or“exonic reference SNP” enables long-range phasing because it can be covered by both cDNA amplicons (reverse transcribed from mature mRNA) and DNA (e.g., genomic DNA) amplicons that do not include introns. Thus, an“exonic reference SNP” also allows phasing and analysis of heterozygous and homozygous exonic or intronic target SNPs, while certain scenarios require the“exonic reference SNP” to be heterozygous (e.g., when the target SNP is heterozygous). In particular, if a distant target single nucleotide polymorphism (SNP) of interest (either intronic or exonic) in a given individual is heterozygous, the optimal“exonic reference SNP” should be polymorphic and heterozygous. Thus, an“exonic reference SNP” can be used to align nucleic acid sequences of a first amplification product derived from a target DNA nucleic acid (e.g., genomic DNA) and a second amplification product derived from a target cDNA (reverse transcribed from mature mRNA) nucleic acid based on the location of the“exonic reference SNP”.
[0055] The term "short tandem repeat sequence" (STR) can also be referred to as "microsatellite" and denotes a repetitive DNA segment in which certain DNA motifs (ranging in length from one to six or more base pairs) are repeated typically 5 to 50 times. Microsatellites exist at thousands of locations in the genomes of organisms, accounting for about 3% of the human genome, and are usually found in introns and intergenic regions, but also in exons. They have a higher mutation rate than other regions of DNA, resulting in higher genetic diversity. STRs can be involved in the onset of a disorder, such as e.g. Huntington's disease. In this context, an abnormal amplification of the tandem repeat sequence consisting of cytosine-adenine-guanine (CAG) nucleotide sequences in the number 1 exon of the Huntington gene on chromosome 4 is considered to be the cause of the onset of the disease. While a normal Huntington gene usually has 10 to 35 CAG repeat sequences, individuals suffering from Huntington's disease exhibit a range of CAG repeat sequence numbers from 36 to more than 100.
[0056] The term "phasing" refers to the assignment of genetic variants to their originating homologous chromosomes. Each chromosome of a human has two copies, one inherited maternally and the other paternally. In the context of phasing STRs and SNPs, the goal is to identify which SNP alleles are on the same chromosome (i.e. in the haploid genome) as a STR of a particular length. In the context of phasing two or more SNPs, the goal is to determine which SNP alleles of the first and second (and third...) SNPs are on the same chromosome.
[0057] The term "primer" refers to a short nucleic acid (oligonucleotide) that serves as a starting point for polynucleotide chain synthesis by a nucleic acid polymerase under suitable conditions. Polynucleotide synthesis and amplification reactions typically include a suitable buffer, dNTPs and / or rNTPs, and one or more optional cofactors, and are carried out at a suitable temperature. Primers typically include at least one target-hybridizing region that is at least substantially complementary (e.g., has 0, 1, or 2 mismatches) to a target sequence. For purposes of the present disclosure, this region is typically about 4 to about 10 nucleotides in length, e.g., 5 to 8 nucleotides. A "primer pair" refers to a forward and reverse primer that are oriented in opposite directions with respect to a target sequence and that generate an amplification product under amplification conditions. The terms "forward" and "reverse" are arbitrarily assigned. One of ordinary skill in the art will understand that forward and reverse primers (primer pairs) define the boundaries of an amplification product. In some embodiments, multiple primer pairs rely on a single common forward or reverse primer. For example, multiple allele-specific forward primers can be considered part of a primer pair with the same common reverse primer, e.g., if the multiple alleles are in close proximity to one another. A "set" or "panel" of primers can refer to one primer pair, or more than one primer pair designed to work together in a single multiplex reaction.
[0058] As used herein, "probe" means any molecule capable of selectively binding to a particular intended target biomolecule, e.g., a nucleic acid sequence of interest that hybridizes to the probe. The probe is detectably labeled with at least one non-nucleotide moiety. In some embodiments, the probe is labeled with a fluorophore and a quencher.
[0059] The term "complementary" or "complementarity" refers to the ability of a nucleic acid in a polynucleotide to form a base pair with another nucleic acid in a second polynucleotide. For example, the sequence A-G-T (A-G-U for RNA) is complementary to the sequence T-C-A (U-C-A for RNA). Complementarity can be partial, in which only some of the nucleic acids match by base pairing, or complete, in which all of the nucleic acids match by base pairing. A probe or primer is said to be "specific" for a target sequence if it is at least partially complementary to the target sequence. Depending on the conditions, the degree of complementarity (e.g., greater than 80%, 90%, 95%, or 98%) to a target sequence is typically higher for shorter nucleic acids such as primers than for longer sequences. In some embodiments, a primer and / or probe is 100% complementary to a target sequence.
[0060] The term "specific amplification" indicates that the target sequence amplified by the primer set is, at a statistically significant level, more than the non-target sequence. The term "specific detection" indicates that the target sequence detected by the probe will, at a statistically significant level, be more than the non-target sequence. As will be appreciated in the art, negative controls (e.g., samples including the same nucleic acids as the test sample but not including the target sequence or samples lacking nucleic acids) can be used to determine specific amplification and detection. For example, primers and probes that specifically amplify and detect a target sequence produce a Ct that is readily distinguishable from background (non-target sequence), e.g., at least 2, 3, 4, 5, 5-10, 10-20, or 10-30 cycles less than background. The term "allele-specific" PCR refers to amplification of a target sequence using primers that specifically amplify a particular allelic variant of the target sequence. Typically, the forward or reverse primer includes the exact complement of the allelic variant at that position.
[0061] In the context of two or more nucleic acids or two or more polypeptides, the terms "identical" or percent "identity" mean that two or more sequences or subsequences are identical or have a specified percentage of amino acid residues or nucleotides that are the same (e.g., at least any one of about 60% identity, e.g., 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher identity over a specified region, when compared and aligned for maximum correspondence over the comparison window or designated region) as measured using a BLAST or BLAST 2.0 sequence comparison algorithms with default parameters or by manual alignment and visual inspection, see, e.g., ncbi.nlm.nih.gov / BLAST. Such sequences are then termed "substantially identical." Percent identity is typically determined over the best aligned sequences, so this definition applies to sequences with deletions and / or additions as well as sequences with substitutions. Algorithms commonly used in the art consider gaps and the like. Typically, identity exists over a region that includes a sequence of at least about 8 to 25 amino acids or nucleotides in length, or over a region of 50 to 100 amino acids or nucleotides in length, or over the entire length of the reference sequence.
[0062] The terms "isolated," "separated," "purified," and the like are not intended to be absolute. For example, isolation of DNA or genomic DNA need not be 100% removal of non-DNA molecules. One of skill in the art will recognize the acceptable level of purity for a given situation.
[0063] The term "amplification product" refers to the product of an amplification reaction. The amplification product includes primers used to initiate each round of polynucleotide synthesis. An "amplicon" is a sequence targeted for amplification, and the term can also be used to refer to the amplification product. The 5' and 3' boundaries of an amplicon are defined by the forward and reverse primers. The terms "reverse transcription product," "RT product," and the like refer to a cDNA molecule produced by elongation of an RT primer on an RNA template by a polymerase with reverse transcriptase activity.
[0064] The term "kit" refers to any article of manufacture (e.g., a package or container) that includes at least one reagent such as a nucleic acid probe or probe pool, etc., for specifically amplifying, capturing, labeling / converting, or detecting an RNA or DNA described herein.
[0065] The term "amplification conditions" refers to conditions in a nucleic acid amplification reaction (e.g., PCR amplification) that allow for hybridization and template-dependent extension of primers. The term "amplicon" or "amplification product" refers to a nucleic acid molecule that contains all or a fragment of a target nucleic acid sequence and is formed as an in vitro amplification product by any suitable amplification method. When applied to a primer, the term "generates an amplification product" indicates that the primer will produce the defined amplification product under appropriate conditions (e.g., in the presence of a nucleotide polymerase and NTPs). Various PCR conditions are described in PCR Strategies (Innis et al., 1995, Academic Press, San Diego, CA) Chapter 14; PCR Protocols: A Guide to Methods and Applications (Innis et al., Academic Press, NY, 1990).
[0066] A "nanopore" is defined as a structure with a nanoscale channel that can pass ions from a solution from one side to the other. Examples of nanopores are protein nanopores (e.g., a-hemolysin and other multi-subunit pore proteins), synthetic nanopores, and hybrid protein / synthetic nanopores. In related embodiments, these nanopores are inserted into a natural or artificial membrane that would otherwise serve to block the passage of ions and other molecules. The width of the nanopore channel should generally be such that it allows passage of polymers such as single-stranded DNA when a voltage gradient is applied across the membrane. During their transit, they will reduce the ionic current at a given voltage due to their size, charge, or other properties. "Nanopore" includes, for example, structures comprising (a) first and second compartments separated by a physical barrier having at least one pore, e.g., about 1 to 10 nm in diameter, and (b) means for applying an electric field across the barrier so that a charged molecule such as DNA, a nucleotide, a nucleotide analog, or a tag can pass from the first compartment through the pore to the second compartment. Ideally, the nanopore further comprises means for measuring electronic signatures of molecules passing through its barrier. The nanopore barrier can be partially synthetic or naturally occurring. The barrier can include, for example, a lipid bilayer with a-hemolysin in it, oligomeric protein channels such as porins and synthetic peptides, and the like. The barrier can also include an inorganic slab with one or more holes of suitable dimensions. In this document, "nanopore," "nanopore barrier," and "pore" in nanopore barrier are sometimes used equivalently.
[0067] Nanopore devices are known in the art, and nanopores and methods of using them are disclosed in U.S. Patent Nos. 7,005,264; 7,846,738; 6,617,113; 6,746,594; 6,673,615; 6,627,067; 6,464,842; 6,362,002; 6,267,872; 6,015,714; 5,795,782 and U.S. Publication Nos. 2004 / 0121525, 2003 / 0104428, and 2003 / 0104428, each of which is hereby incorporated by reference in its entirety.
[0068] A "nanopore array" is a chip containing many individual nanopores at known locations; each nanopore can be queried electronically individually (allowing single-molecule electronic nanopore-based sequencing to be performed).
[0069] A "nanopore-detectable tag" (also called a "nanopore tag") is a molecule, usually a polymer, that is covalently attached to a nucleotide in a nanopore SBS reaction. Different nanopore tags are typically attached to each nucleotide A, C, G, and T (or U) so as to elicit different ionic current blockage signals as they pass through the channel of a nanopore when a voltage gradient is applied across the membrane.
[0070] Nanopore sequencing by synthesis (also referred to as“nanopore SBS”) refers to our previously described method (Kumar et al. 2012; Fuller et al. 2016; Stranges et al. 2016) in which tags attached to nucleotides can be distinguished by their effect on the ionic current through a nanopore as these modified nucleotides are added to a growing DNA chain. Measurements can be made while the tagged nucleotides are still part of a ternary complex, or after their tags are released by the polymerase reaction.
[0071] The terms“individual,”“subject,” and“patient” are used interchangeably herein. An individual can be pre-diagnostic, post-diagnostic but pre-treatment, receiving treatment, or post-treatment. In the context of the present disclosure, the individual is typically seeking medical care.
[0072] The term“sample” or“biological sample” refers to any composition containing or presumed to contain nucleic acids. The term includes purified or isolated components of cells, tissues, or blood, e.g., DNA, RNA, proteins, acellular fractions, or cell lysates. The sample can be FFPET, e.g., from a tumor or metastatic lesion. The sample can also be from frozen or fresh tissue, or from a liquid sample, e.g., blood or a blood component (plasma or serum), urine, semen, saliva, sputum, mucus, semen, tears, lymphatic fluid, cerebrospinal fluid, mouth / throat rinse, bronchial alveolar lavage, materials washed from swabs, etc. The sample can also include components and fractions of in vitro cultures of cells obtained from an individual, including cell lines. The sample can also be partially processed from a sample obtained directly from an individual, e.g., a cell lysate or red blood cell-depleted blood. A tumor sample can include tissue from a tumor, or a sample including DNA from a tumor, e.g., ctDNA in the blood of a cancer patient.
[0073] The term“obtaining a sample from an individual” means providing a biological sample from an individual for testing. It can be obtained directly from the individual, or it can be obtained from a third party who obtained the sample directly from the individual. The sample can be taken pre-treatment, during treatment, or post-treatment. The sample can be taken from a patient suspected of having or diagnosed with disease X, and thus can be in need of treatment, or it can be taken from a normal individual not suspected of having any ailment. The treatment regimen, (higher / lower / more frequent / less frequent) dosage.
[0074] The term "assessing disease X" is used to indicate that the methods disclosed herein will aid a medical professional, including for example a physician, in assessing whether an individual has disease X or is at risk of having disease X or predicting a course of disease X. The presence of biomarker Y, a combination of biomarker Y; Z;... or a ratio of biomarker Y; Z;... in a sample indicates that the individual has disease X or that the individual is at risk of having disease X or predicting a course of disease X. In one embodiment, the term assessing disease X is used to indicate that the methods according to the present application will aid a medical professional in assessing whether an individual has disease X. In this embodiment, the presence of biomarker Y in a sample indicates that the individual has disease X, i.e. the presence of biomarker Y indicates the presence of disease X in the individual is at / above / below a reference level. In certain embodiments, the term "at a reference level" means that the level of the biomarker in a sample from the individual or patient is substantially the same as the reference level or that the level thereof differs from the reference level by at most 1%, at most 2%, at most 3%, at most 4%, at most 5%.
[0075] The term "providing a therapy to an individual" means prescribing, recommending, or offering a therapy to an individual. The therapy can actually be administered to the individual by a third party (e.g., an inpatient injection) or by the individual themselves. As used herein, the phrase "selecting a therapy" means using information or data generated that correlates with the level or presence of biomarker Y in a patient sample to identify or select a therapy for the patient. In some embodiments, the therapy can include drug D. In some embodiments, the phrase "identifying / selecting a therapy" includes identifying a patient in need of adjusting the effective amount of drug D being administered. In some embodiments, recommending a therapy includes recommending adjusting the amount of drug D being administered. As used herein, the phrase "recommending a therapy" can also mean using information or data generated for a patient that is determined or selected as more likely or less likely to respond to a therapy comprising drug D to suggest or select a therapy comprising drug D. The information or data used or generated can be in any form, written, oral, or electronic. In some embodiments, using the information or data generated includes communicating, presenting, reporting, storing, sending, transferring, supplying, transmitting, distributing, or a combination thereof. In some embodiments, the communicating, presenting, reporting, storing, sending, transferring, supplying, transmitting, distributing, or a combination thereof is performed by a computing device, an analyzer unit, or a combination thereof. In some further embodiments, the communicating, presenting, reporting, storing, sending, transferring, supplying, transmitting, distributing, or a combination thereof is performed by a laboratory or medical professional. In some embodiments, the information or data includes a comparison of the level of biomarker Y to a reference level. In some embodiments, the information or data includes an indication of the presence or absence of biomarker Y in a sample. In some embodiments, the information or data includes an indication that a therapy comprising drug D is appropriate for the patient.
[0076] The terms "marker," "tag," "detectable moiety," and the like refer to a composition that can be detected by spectroscopic, photochemical, biochemical, immunochemical, chemical or other physical means. For example, useful markers include fluorescent dyes (fluorophores), luminescent agents, radioactive isotopes, electron-dense reagents, or affinity agents, e.g., biotin or streptavidin. Those of ordinary skill will appreciate that a detectable marker conjugated to a nucleic acid is not naturally occurring. 32 P, 3 H), electron-dense reagents, or affinity-based moieties, e.g., poly A (interacts with poly T) or poly T tags (interact with poly A), His tags (interact with Ni) or streptavidin tags (separable from biotin). Those of ordinary skill will appreciate that a detectable marker conjugated to a nucleic acid is not naturally occurring.
[0077] Unless otherwise defined, scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. See, e.g., Lackie, DICTIONARY OF CELL AND MOLECULAR BIOLOGY, Elsevier (4th ed. 2007); Sambrook et al., MOLECULAR CLONING, A LABORATORY MANUAL, Cold Springs Harbor Press (Cold Springs Harbor, N.Y. 1989). The terms "a" or "an," are intended to mean "one or more" The terms "comprise," "comprises," and "comprising" are intended to mean that the steps or elements (features) that follow are optional additions to the steps or elements (features) that precede them, and are not meant to be exclusive.
[0078] III. Nucleic Acid Sample
[0079] Samples for biomarker detection can be obtained from any source suspected of containing large amounts of non-fragmented nucleic acids or large fragment (>10 kb) nucleic acids, e.g., tissue (including tumor tissue), blood (including cell-free nucleic acids such as cell-free DNA and cell-free RNA), skin, swabs (e.g., oral, vaginal), urine, saliva, etc. Methods for isolating nucleic acids from biological samples are known, e.g., as described in Sambrook, and several kits are commercially available (e.g., High Pure RNA Isolation Kit, High Pure Viral Nucleic Acid Kit, and MagNA Pure LC Total Nucleic Acid Isolation Kit, DNA Isolation Kit for Cells and Tissues, DNA Isolation Kit for Mammalian Blood, High Pure FFPET DNA Isolation Kit, available from Roche). In the context of the methods of the present disclosure, genomic DNA and RNA can be collected and isolated.
[0080] IV. Detailed description of the workflow
[0081] In the following, a detailed description of an exemplary laboratory workflow using sample preparation, amplification, library preparation, and sequencing analysis is provided, as well as an exemplary subsequent bioinformatics workflow, including analysis of sequencing data exemplifying the practice of the disclosed methods.
[0082] Laboratory workflow
[0083] A DNA and RNA co-extraction method (e.g., AllPrep DNA / RNA Micro Kit, QIAGEN) is used to process the sample of interest (e.g., blood or tissue) according to the manufacturer’s protocol. Alternatively, the sample of interest can be split into two parts, while DNA can be extracted from the first part (e.g., using DNeasy Blood and Tissue Kit, Qiagen) and RNA can be extracted from the second part of the sample (e.g., using RNeasy Kit, Qiagen) according to the manufacturer’s protocol. The isolated nucleic acid material is analyzed using quality control workflows (e.g., NanoDrop™ spectrophotometer, Qubit™ fluorometer) according to the manufacturer’s protocol, and subsequently a polymerase chain reaction (PCR)-based amplification reaction or a reverse transcription and polymerase chain reaction (RT-PCR)-based amplification reaction is performed. The above reactions (i.e., PCR and RT-PCR) are performed on genomic DNA and mRNA, respectively, and utilize standard reagents (i.e., polymerase, reaction buffer, primers, etc.) and standard equipment (e.g., pipettes, tips, lab tubes, PCR capping, and PCR thermocycler) for long amplicon generation. The design of custom primers for performing PCR and RT-PCR amplification reactions to generate DNA and cDNA amplicons can be obtained using NCBI Primer Blast or Primer3 open source algorithm (National Library of Medicine, Bethesda, MD, USA). The PCR and RT-PCR reactions are performed according to the manufacturer’s protocol (e.g., Expand™ High Fidelity PCR System, Roche). The amount of PCR and RT-PCR cycles depends on the amount and quality of the isolated nucleic acid as starting material. A cleanup step is performed on the DNA and cDNA amplicons of interest according to standard methods and manufacturer’s documentation (e.g., Agencourt AMPure XP magnetic beads, Beckman Coulter; or Blue Pippin Prep, Sage Science) to remove excess unused amplification primers or short off-target molecules. The purified amplicons are tested using quality control workflows (e.g., NanoDrop™ spectrophotometer, Qubit™ fluorometer, or bioanalyzer, Agilent Technologies, Santa Clara, CA, USA) according to the manufacturer’s protocol and recommendations.According to the manufacturer's protocol, both amplicons (i.e. DNA and cDNA) from a single sample are combined in equimolar concentrations and barcoded indexed with sequencing adaptors (e.g. hairpin loop for PacBio sequencing using the Library Preparation Kit, Pacific Biosciences of California, Inc., Menlo Park, USA) and further combined into a final sequencing library pool. This generated sample is quantified using a Qubit™ spectrophotometer (e.g. Broad Range DNA Kit, ThermoFisher Scientific) and loaded onto a sequencing device (e.g. Sequel system, Pacific Biosciences of California, Inc., Menlo Park, USA) and processed according to the manufacturer's device instructions.
[0084] Bioinformatics workflow
[0085] Raw data generated using one of the long-read sequencing devices (e.g., Sequel System, Pacific Biosciences of California, Inc., Menlo Park, USA; or GridION System, Oxford Nanopore Technologies Ldt., Oxford, UK) can be bioinformatically processed according to the flowchart shown in FIG. 8. Here, in a first step, sequencing raw data (100) can be base-called into FASTQ reads (110) using onboard software. This generated data can be transferred onto a Linux-based local server for bioinformatics data analysis. Available reads are demultiplexed or called (120) with open-source algorithms (e.g., LIMA or pbCCS) and aligned to the human genome reference (e.g., BWA, STAR, or minimap2) (130) according to the presence or absence of introns. The alignment results allow for single molecule SNP calling (MUSCLE, SNVer, VarScan, Samtools / mpileup) and extraction of exon or intron statistics (131). The number of short tandem repeats from cDNA reads can be quantified (132) using Tandem Repeat Genotyper, RepeatAnalysisTools, or DeepRepeat software. Similarity clustering (e.g., UCLUST, VSEARCH, or pbAA) can be performed on DNA and cDNA data, allowing for better resolution of CAG tandem repeats or exon / intron SNPs. Information from DNA and cDNA data is merged into a single CSV file and binned according to the amount of CAG tandem repeats or exon / intron SNPs. The homozygous and heterozygous result bins can be visualized using Integrative Genomics Viewer (IGV) and plotted in a table for each individual step to finally indicate which target SNP allele is in phase with the mutant CAG repeat (140). Statistical analysis can be performed to remove noise signals or chimeric reads and improve the final statistical accuracy.
[0086] In a related workflow, raw data generated using one of the long read sequencing devices (e.g., Sequel System, Pacific Biosciences of California, Inc., Menlo Park, USA) can be bioinformatically processed according to the flowchart shown in FIG. 9. This method uses multiple read filtering steps to generate the highest quality data for clustering and alignment. Statistical analysis is performed using validated open source programs and signals are merged into a final phasing report and STR profile. In the first step, raw sequencing data is transferred via a local area network (LAN) onto a server for post-processing and statistical data analysis. Consensus calling is performed into FASTQ reads using pbCCS software, allowing filtering out shorter amplicons and reads with lower consensus quality or incomplete sequencing pass. Generated consensus reads are demultiplexed (200) using open source algorithm (i.e., LIMA) and data is saved separately for cDNA amplicons (202) and gDNA amplicons (201). Subsequently, high quality reads are clustered (212, 211) using PacBio Amplicon Analysis (pbAA). Raw and clustered bins are aligned (222, 221) to the human reference genome according to the presence or absence of introns. Read binning is performed to remove noise signals (e.g., PCR chimeras) and improve final statistical accuracy. High quality data is used to extract genetic positions, including exon SNPs and motifs of interest (e.g., CAG, CAA, CCG, CCA, CGG) for cDNA (232) and exon and intron SNP coordinates (233) for gDNA (231, 232). Alignment results allow single molecule SNP calling and subsequent extraction of exon and intron statistical indicators (241) for targeted positions. Number of short tandem repeats from cDNA reads is quantified using RepeatAnalysisTools (233). gDNA and cDNA reads subjected to similarity clustering (i.e., pbAA) allow better resolution of STRs and exon / intron SNPs (242), while non-clustered reads allow better quantification of reads per haplotype. Information from both workflows (i.e., error corrected and non-corrected gDNA / cDNA) is merged into a single file and binned according to the amount of STR or exon / intron SNPs. Homozygous and heterozygous result bins are visualized with waterfall images and STRs (252) are presented. Main results of the global analysis indicate which intron SNP alleles are in phase with mutant or wild-type STRs (251).
[0087] Patient stratification workflow
[0088] The generated results can be analyzed to identify the presence of the desired intronic SNP allele in phase with the mutant STR allele (e.g., CAG repeat). Figure 10 visualizes the final four possible phasing results (i.e., haplotypes consisting of CAG repeat and intronic target SNP). Notably, only the combination with the heterozygous target SNP allows for allele-specific treatment using antisense oligonucleotide (ASO)-based therapy. As shown in Figure 10, for an adenine-specific ASO, the following haplotypes qualify:
[0089] 1) Haplotype 1: mutant STR + adenine SNP, and
[0090] Haplotype 2: wild-type STR + guanine SNP
[0091] Reverse SNP allele:
[0092] 2) Haplotype 1: mutant STR + guanine SNP, and
[0093] Haplotype 2: wild-type STR + adenine SNP
[0094] will cause depletion of wild-type STR mRNA and preservation of mutant STR mRNA. Therefore, treatment of patient 2) with a reverse SNP allele requires use of a guanine-specific ASO. The remaining homozygous combinations 3 (homozygous SNP A / A) or 4 (homozygous SNP G / G) will cause depletion of both wild-type and mutant STR mRNA, or will not cause on-target reduction of either of the STR alleles upon ASO administration.
[0095] V. Kits
[0096] Provided herein are kits for determining nucleic acid sequences of at least one distant single nucleotide polymorphism and short tandem repeat sequence of interest within the same target locus of a nucleic acid isolated from a biological sample, the kit comprising a first set of oligonucleotide primers for generating a first amplification product comprising at least one single nucleotide polymorphism of interest and at least one exon reference single nucleotide polymorphism in the same locus; and a second set of oligonucleotide primers for generating a second amplification product comprising at least one single nucleotide polymorphism of interest and the at least one exon reference single nucleotide polymorphism. In this context, the kit is suitable for use in carrying out any of the methods disclosed herein. In a particular embodiment, the first set of oligonucleotide primers comprises an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 1 and an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 2. In another particular embodiment, the second set of oligonucleotide primers comprises an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 3, an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 4, and an oligonucleotide primer comprising the nucleic acid sequence of SEQ ID NO: 5. In some embodiments, the kit includes a sample collection container (e.g., a tube, a vial, a multi-well plate, or a multi-container cartridge).
[0097] In some embodiments, the kit includes reagents and / or components for nucleic acid purification. For example, the kit can include a lysis buffer (e.g., including a detergent, a chaotropic agent, a buffering agent, etc.), an enzyme or reagent for denaturing proteins or other unwanted materials in the sample (e.g., proteinase K), an enzyme for preserving nucleic acids (e.g., a DNase and / or RNase inhibitor). In some embodiments, the kit includes components for nucleic acid isolation, e.g., a solid or semi-solid matrix such as a chromatographic matrix, magnetic beads, magnetic glass beads, glass fibers, silica filters, etc. In some embodiments, the kit includes wash and / or elution buffers for purifying and releasing nucleic acids from a solid or semi-solid matrix. For example, the kit can include components from a MagNA Pure LC Total Nucleic Acid Isolation Kit, a DNA Isolation Kit for Mammalian Blood, a High Pure or MagNA Pure RNA Isolation Kit (Roche), a DNeasy or RNeasy Kit (Qiagen), a PureLink DNA or RNA Isolation Kit (Thermo Fisher), etc.
[0098] The kit can further include reagents for amplification, e.g., reverse transcriptase, DNA polymerase, dNTPs, buffers, and / or other elements suitable for reverse transcription and / or amplification (e.g., cofactors or aptamers). Typically, the reagent mixtures are concentrated so that aliquots are added to the final reaction volume along with the sample (e.g., RNA or DNA), enzymes, and / or water. In some embodiments, the kit further includes a reverse transcriptase (or an enzyme with reverse transcriptase activity) and / or a DNA polymerase (e.g., a thermostable DNA polymerase such as Taq, Z05, and derivatives thereof).
[0099] In some embodiments, the kit further includes consumables, e.g., plates or tubes for nucleic acid preparation, tubes for sample collection, or plates, tubes, or microchips for PCR or qRT-PCR. In some embodiments, the kit further includes instructions, references to websites, or software, e.g., for further processing of sequencing data.
[0100] Examples
[0101] Example 1
[0102] Genomic DNA amplification and reverse transcription of mRNA from the HTT gene
[0103] Huntington's disease is a rare autosomal dominant genetic disorder caused by abnormal expansion of a tandem repeat sequence composed of cytosine-adenine-guanine (CAG) nucleotides in exon 1 of the huntingtin gene (HTT) on chromosome 4. Normal HTT usually has 10 to 35 CAG repeats, but in individuals with Huntington's disease, the number of CAG repeats can range from 36 to over 100. To treat this disease, ideally only mutant HTT is inhibited, while wild-type HTT is left intact. Targeting a heterozygous SNP near the CAG repeat would allow this allele-specific approach. Several SNPs in the HTT gene with high allele frequencies, including SNP rs7685686 in intron 42, have been identified as candidates for targeted therapy. However, the assay to identify heterozygous patients with specific SNP alleles in phase with the mutant CAG repeat is a prerequisite for this selective therapy. Due to the long distance of over 130 kb between these genetic variants, it is not possible to phase the CAG repeat in exon 1 and the SNP in intron 42 using standard methods. Moreover, using RNA only as starting material to reduce the distance between the CAG repeat and the distant SNP is not possible for intronic SNPs. Therefore, a new laboratory workflow was developed that utilizes both genomic DNA and RNA to enable phasing analysis of STRs and distant (even intronic) SNPs within the same locus.
[0104] Thirteen samples from the NIGMS Human Genetic Cell Repository at Coriell Institute for Medical Research were processed as described below. Genomic DNA and total RNA were isolated using a co-extraction method (AllPrep DNA / RNA Micro Kit, QIAGEN) according to the manufacturer's protocol. Therefore, the extracted nucleic acids were processed with a quality control workflow based on quantification of concentration and purity using the Qubit HS DNA / RNA Kit and a nanodrop spectrophotometer. Samples characterized by high concentration and correct 260 / 230, 260 / 280 mass, and RNA integrity values (RIN) were subjected to PCR and RT-PCR.
[0105] DNA assay of -10kb amplicons required a minimum of 100ng of high quality material to be inputted and mixed with 10ul 5x PrimeSTAR GXL buffer, 4uL dNTPS mix (200uM each), 0.7uL 15uM forward and reverse primers, 1uL PrimerSTAR GXL DNA polymerase and 13.6uL nuclease free water (i.e. 30uL master mix) while adding input genomic DNA at 5ng / uL and total 20uL volume (50uL final reaction volume). The amplification reaction was performed in a pre-heated thermocycler with the following conditions: 1 minute pre-incubation at 98°C, then 30 cycles of 20 seconds at 98°C, 15 seconds at 58°C, then 15 minutes incubation at 68°C. The reaction was stopped by lowering the temperature to 4°C for infinite time. Amplicons were purified with AMPure® PB Beads (100-265-900, Pacific Biosciences) at a 0.45x ratio according to the manufacturer's protocol. The quantity of amplicons was measured with Qubit BR DNA assay while the size of these amplicons was verified with TapeStation, Genomic DNA Screen Tapes. The generated amplicons were second purified with AMPure® PB Beads (100-265-900, Pacific Biosciences) at a 3.1x ratio according to the manufacturer's protocol to remove potential <7kb fragments still present. After the second purification, the quantity of amplicons was measured again with Qubit BR DNA assay while the size of these amplicons was verified with TapeStation, Genomic DNA Screen Tapes.
[0106] For the main step, cDNA assay of ~10 kb amplicon requires a minimum of 1000 ng of high integrity total RNA. The reaction includes gDNA removal by mixing 1 uL of 10x exDNase buffer with 1 uL of exDNase enzyme and 8 uL of total RNA input (minimum of 1000 ng). This prepared sample is incubated at 37 °C for 2 minutes, followed by the addition of 1 uL of 100 mM DTT to inactivate the enzyme (final volume 11 uL) and incubated at 55 °C for 5 minutes. The next step includes annealing of gene-specific reverse transcription primers (2 uM concentration) adding 10 mM dNTP mix (10 mM each) to 11 uL of RNA from the previous step. The reaction is heated at 65 °C for 5 minutes and immediately placed on ice for at least 1 minute. The subsequent step is processed by mixing 5x Superscript IV buffer with 100 mM DTT, ribonucleotide inhibitor, Superscript IV RT enzyme (7 uL) and 13 uL of RNA with pre-annealed RT primers. Reverse transcription requires incubation at 55 °C for 10 minutes, followed by incubation at 80 °C for 10 minutes and finally at 4 °C for an infinite amount of time. The last step of reverse transcription includes removing RNA from the reaction with 1 uL of E. coli RNase H and incubation at 37 °C for 20 minutes. After cDNA synthesis, the cDNA is diluted 5-fold with water. In the next step of PCR amplification,
[0107] Using 5X PrimeSTAR GXL Buffer, dNTP Mix, forward and reverse primers (15 uM), PrimerSTAR GXL DNA Polymerase, and water (40 uL total) and 10 uL of diluted cDNA from the previous step (reaction total volume 50 uL). Four PCR wells were used per sample to improve yield. The amplification reaction was performed in a pre-heated thermocycler with the following conditions: pre-incubate at 98°C for 1 minute, then cycle 33 times at 98°C for 20 seconds, at 60°C for 15 seconds, then at 68°C for 20 minutes. The reaction was stopped by lowering the temperature to 4°C indefinitely. Four wells were pooled per amplicon / sample, then purified for the first time with AMPure® PB Beads (100-265-900, Pacific Biosciences) at a 0.45x ratio according to the manufacturer’s protocol. The quantity of amplicons was measured with Qubit BR DNA Assay while the size of these amplicons was verified with TapeStation, and the resulting amplicons were purified for the second time with AMPure® PB Beads (100-265-900, Pacific Biosciences) at a 3.1x ratio according to the manufacturer’s protocol to remove residual potential <7 kb fragments. After the second purification, the quantity of amplicons was measured with Qubit BR DNA Assay while the size of these amplicons was verified with TapeStation, Genomic DNA Screen Tapes.
[0108] DNA and cDNA amplicons from the same sample were normalized according to Qubit values and pooled at equimolar concentrations. Each sample pool was ligated with sequencing index adaptors and loaded onto a sequencing flow cell according to the manufacturer’s protocol (Pacific Biosciences of California, Inc., Menlo Park, USA). Pools of DNA and cDNA samples were performed to bin the results of each sample into indexed batch data.
[0109] PCR primers for the HTT gene from genomic DNA (i.e. Exon SNV + Intron SNP)
[0110]
[0111] Table 1
[0112] RT-PCR primers for the HTT gene from mRNA / cDNA (i.e. CAG repeat sequence + exon SNV)
[0113]
[0114] Table 2
[0115] Example 2
[0116] Bioinformatic analysis
[0117] The two amplicon reads (i.e., DNA sequences derived from genomic DNA amplification and sequencing steps; and cDNA sequences derived from mRNA reverse transcription, amplification, and sequencing steps) are first aligned to a reference gene or the complete human genome and processed with open source programs to quantify the number of tandem repeat sequences for each molecule and extract the nucleotide signal at the desired SNP location. The calibration programs used include BWA, STAR, or minimap2, while the quantification programs for CAG include Tandem Repeat Genotyper, RepeatAnalysisTools, or DeepRepeat software. Example results for three different Coriell samples are shown in Figures 3-5 (cell lines GM04282, GM13503, and GM04724 are obtained from the Coriell Institute for Medical Research NIGMS Human Genetic Cell Repository). The results of the cDNA analysis are shown in Figure A (Figures 3-5) and visualized in Figure B (Figures 3-5), which indicates which exon SNP allele is in phase with high and low CAG repeat number and shows technical details such as the average CAG count, standard deviation, percentage of reads containing high or low tandem repeat number, etc. The results of the DNA amplicon are shown in Figure C (Figures 3-5) and indicate the percentage of reads containing intronic polymorphism (i.e., guanine or adenine) for the exon polymorphism (i.e., cytosine or thymine). In Figure 3, which describes the results for sample GM04282, it is shown that based on the cDNA data, the phasing of the mutant HTT allele with about 77.6 CAGs (“high CAG”) that accounts for about 22.7% of the reads shows an exon T allele in about 95% of the reads. The wild-type HTT allele with about 18.1 CAGs (“low CAG”) that accounts for about 77.2% of the reads also shows an exon T allele in about 95% of the reads. Both alleles have >4% of reads with undefined nucleotides at the exon reference SNP location, which can be due to errors such as insertions or deletions. The DNA amplicon data (Figure 3C) shows that both the exon and intron SNPs are homozygous and therefore this patient is not suitable for allele-specific treatment. Figures 4C and 5C show the data for samples GM13503 and GM04742, where both the exon and intron SNPs are heterozygous, which would make both of these individuals eligible for allele-specific approaches.
[0118] The mixed analysis of DNA and cDNA amplicon sequences required for phasing CAG trinucleotide repeat sequences with intronic SNPs was based on read paving via the exonic reference SNP - the results for each amplicon allowing for this operation are shown in Figures 3, 4 and 5 (A and C). The binning and filtering steps of the data were based on sequence similarity (i.e. 100%) represented in Figure 6, where error-prone reads based on QC indicators were discarded, while the signal with the highest read number was called as the consensus. Individual haplotypes could be visualized with the IGV software, as shown in genomic DNA amplicons in Figure 7 or bar graphs in Figure B (Figures 3 to 5)
[0119] References
[0120] Bečanović, K., Nørremølle, A., Neal, S.J., Kay, C., Collins, J.A., Arenillas, D., Lilja, T., Gaudenzi, G., Manoharan, S., Doty, C.N. and Beck, J. (2015) A SNP in the HTT promoter alters NF-κB binding and is a bidirectional genetic modifier of Huntington disease. Nature neuroscience, 18(6), pp.807-816.
[0121] Carroll, J.B., Warby, S.C., Southwell, A.L., Doty, C.N., Greenlee, S., Skotte, N., Hung, G., Bennett, C.F., Freier, S.M. and Hayden, M.R. (2011) Potent and selective antisense oligonucleotides targeting single-nucleotide polymorphisms in the Huntington disease gene / allele-specific silencing of mutant huntingtin. Molecular Therapy, 19(12), pp.2178-2185.
[0122] Claassen D.O., Corey-Bloom J., Dorsey E.R., Edmondson M., Kostyk S.K., LeDoux M.S., Reilmann R., Rosas H.D., Walker F., Wheelock V., Svrzikapa N., Longo K.A., Goyal J., Hung S., Panzara M.A. (2020) Genotyping single nucleotide polymorphisms for allele-selective therapy in Huntington disease. Neurol Genet, 6(3) e430.
[0123] Flower M., Lomeikaite V., Ciosi M., Cumming S., Morales F., Lo K., Moss D.H., Jones L., Holmans P., Monckton D.G., Tabrizi S.J., (2019) MSH3 modifies somatic instability and disease severity in Huntington’s and myotonic dystrophy type 1, Brain, 142(7): 1876-1886
[0124] Goold R., Flower M., Moss D.H., Medway C., Wood-Kaczmar A., Andre R., Farshim P., Bates G.P., Holmans P., Jones L., Tabrizi S.J., (2019) FAN1 modifies Huntington’s disease progression by stabilizing the expanded HTT CAG repeat, Human Molecular Genetics, 28(4):650-661.
[0125] Hannan, A. (2018) Tandem repeats mediating genetic plasticity in health and disease. Nat Rev Genet 19, 286-298.
[0126] Kartsaki, E., Spanaki, C., Tzagournissakis, M., Petsakou, A., Moschonas, N., MacDonald, M., & Plaitakis, A. (2006) Late-onset and typical Huntington disease families from Crete have distinct genetic origins. International Journal of Molecular Medicine, 17, 335-346.
[0127] Kay, C., Collins, J.A., Skotte, N.H., Southwell, A.L., Warby, S.C., Caron, N.S., Doty, C.N., Nguyen, B., Griguoli, A., Ross, C.J. and Squitieri, F. (2015) Huntingtin haplotypes provide prioritized target panels for allele-specific silencing in Huntington disease patients of European ancestry. Molecular Therapy, 23(11): 1759-1771.
[0128] Lee, J.M., Gillis, T., Mysore, J.S., Ramos, E.M., Myers, R.H., Hayden, M.R., Morrison, P.J., Nance, M., Ross, C.A., Margolis, R.L., and Squitieri, F. (2012) Common SNP-based haplotype analysis of the 4p16.3 Huntington disease gene region. The American Journal of Human Genetics, 90(3):434-444.
[0129] Ramos, E.M., Latourelle, J.C., Lee, JH., et al. (2012) Population stratification may bias analysis of PGC-1a as a modifier of age at Huntington disease motor onset. Hum Genet 131 : 1833-1840.
[0130] Shin, J.W., Shin, A., Park, S.S., and Lee, J.M. (2022) Haplotype-specific insertion-deletion variations for allele-specific targeting in Huntington's disease. Molecular Therapy-Methods & Clinical Development, 25:84-95.
[0131] Skotte, N.H., Southwell, A.L., Ostergaard, M.E., Carroll, J.B., Warby, S.C., Doty, C.N., Petoukhov, E., Vaid, K., Kordasiewicz, H., Watt, A.T., and Freier, S.M. (2014) Allele-specific suppression of mutant huntingtin using antisense oligonucleotides: providing a therapeutic option for all Huntington disease patients. PloS one, 9(9), p.e107434.
[0132] Slatko B.E., Gardner A.F., Ausubel F.M. (2018) Overview of Next- Generation Sequencing Technologies. Curr Protoc Mol Biol. 122(1):e59.
[0133] Warby, S.C., Montpetit, A., Hayden, A.R., Carroll, J.B., Butland, S.L., Visscher, H., Collins, J.A., Semaka, A., Hudson, T.J., and Hayden, M.R. (2009) CAG expansion in the Huntington disease gene is associated with a specific and targetable predisposing haplogroup. The American Journal of Human Genetics, 84(3):351-366.
Claims
1. A method of phasing at least one distal single nucleotide polymorphism of interest and a short tandem repeat sequence within the same target locus of nucleic acid isolated from a biological sample, the method comprising: a. performing an amplification step comprising contacting genomic DNA isolated from the biological sample comprising the target locus with a first set of oligonucleotide primers to produce a first amplification product comprising at least one single nucleotide polymorphism of interest and at least one exon reference single nucleotide polymorphism in the same locus; b. performing a reverse transcription and amplification step comprising contacting mRNA isolated from the biological sample comprising the target locus with a second set of oligonucleotide primers to produce a second amplification product comprising the short tandem repeat sequence and the at least one exon reference single nucleotide polymorphism; c. determining the nucleic acid sequence of the first amplification product and the second amplification product; d. aligning the nucleic acid sequences determined in step (c) of the first amplification product and the second amplification product based on the location of the at least one exon reference single nucleotide polymorphism; e. determining the haplotype of the at least one distal single nucleotide polymorphism of interest and the short tandem repeat sequence in the sample.
2. A method of phasing at least one distal single nucleotide polymorphism of interest and at least one exon single nucleotide polymorphism of interest within the same target locus of nucleic acid isolated from a biological sample, the method comprising: a. performing an amplification step comprising contacting genomic DNA isolated from the biological sample comprising the target locus with a first set of oligonucleotide primers to produce a first amplification product comprising at least one single nucleotide polymorphism of interest and at least one exon reference single nucleotide polymorphism in the same locus; b. performing a reverse transcription and amplification step comprising contacting mRNA isolated from the biological sample comprising the target locus with a second set of oligonucleotide primers to produce a second amplification product comprising the at least one exon single nucleotide polymorphism of interest and the at least one exon reference single nucleotide polymorphism; c. determining the nucleic acid sequence of the first amplification product and the second amplification product; d. aligning the nucleic acid sequences determined in step (c) of the first amplification product and the second amplification product based on the location of the at least one exon reference single nucleotide polymorphism; e. determining the haplotype of the at least one distal single nucleotide polymorphism of interest and the at least one exon single nucleotide polymorphism of interest in the sample.
3. The method of any one of claims 1 and 2, wherein the at least one distal single nucleotide polymorphism of interest is an intronic single nucleotide polymorphism.
4. The method of claim 3, wherein the position of the intron carrying the at least one distant single nucleotide polymorphism of interest is adjacent to an exon carrying the at least one exon reference single nucleotide polymorphism within the same locus of the isolated genomic DNA.
5. The method of any one of claims 1 to 4, wherein the amplification step (a) and the reverse transcription and amplification step (b) are performed in separate reaction vessels.
6. The method of any one of claims 1 to 5, wherein the nucleic acid sequences of the first amplification product and the second amplification product in step (c) are determined using long read sequencing.
7. The method of any one of claims 1 to 6, wherein the biological sample is selected from the group consisting of tissue, tumor tissue, blood, saliva, and a cell line derived from an individual.
8. The method of any one of claims 1 to 7, wherein prior to aligning the nucleic acid sequences of the first amplification product and the second amplification product in step (d), the method further comprises a step of aligning the nucleic acid sequences of the first amplification product and the second amplification product to a nucleic acid sequence of a reference gene of the target locus or to a nucleic acid sequence of a complete human genome or to a portion of the complete human genome comprising the reference gene of the target locus.
9. The method of any one of claims 1, 3 to 8, wherein the target locus is the Huntingtin gene.
10. The method of claim 9, wherein the short tandem repeat sequence is a CAG repeat sequence.
11. The method of any one of claims 9 to 10, wherein the at least one single nucleotide polymorphism of interest is selected from the group consisting of rs7685686, rs362331, rs363088, rs362273, rs2024115, rs6446723, rs363064, rs2285086, rs6844859, rs363080, rs34315806, rs2298967, rs362271, rs363099, rs3121419, and rs16843804.
12. An in vitro method for diagnosing whether an individual is at risk of developing Huntington’s disease, the method comprising: a. determining the haplotype of at least one distant single nucleotide polymorphism of interest and short tandem repeat sequence in an in vitro sample obtained from the individual using the method of any one of claims 1, 3 to 11; b. determining the risk of the individual developing Huntington’s disease based on the haplotype of the at least one distant single nucleotide polymorphism of interest and CAG short tandem repeat sequence determined in step (a).
13. An in vitro method for identifying a patient having Huntington’s disease as likely to respond to a therapy targeting at least one single nucleotide polymorphism of interest, the method comprising: a. determining the haplotype of at least one distant single nucleotide polymorphism of interest and short tandem repeat sequence in an in vitro sample obtained from the patient using the method of any one of claims 1, 3 to 11; b. identifying the patient as likely to respond to a therapy targeting at least one single nucleotide polymorphism of interest based on the haplotype of the at least one distant single nucleotide polymorphism of interest and CAG short tandem repeat sequence determined in step (a). a. determining in an in vitro sample obtained from the individual at least one distant single nucleotide polymorphism and haplotype of short tandem repeat sequence of interest using the method according to any one of claims 1, 3 to 11; b. identifying the patient as likely to respond to the therapy based on the haplotype determined in step (a) and on a determination of the presence of a specific allele of the at least one single nucleotide polymorphism of interest.
14. The method according to claim 13, wherein the therapy targeting the at least one single nucleotide polymorphism of interest is an antisense oligonucleotide treatment aimed at silencing an RNA molecule comprising the specific allele of the at least one single nucleotide polymorphism of interest.
15. A kit for determining the nucleic acid sequence of at least one distant single nucleotide polymorphism and haplotype of short tandem repeat sequence of interest within the same target locus of a nucleic acid isolated from a biological sample, the kit comprising: - a first set of oligonucleotide primers for generating a first amplification product comprising at least one single nucleotide polymorphism of interest and at least one exon reference single nucleotide polymorphism in the same locus; - a second set of oligonucleotide primers for generating a second amplification product comprising at least one exon single nucleotide polymorphism of interest and the at least one exon reference single nucleotide polymorphism.
16. An antisense oligonucleotide specifically hybridizing to at least one distant single nucleotide polymorphism of interest in the Huntingtin (HTT) gene for use in the treatment of a patient suffering from Huntington’s disease, wherein the patient is selected for treatment when a specific haplotype of the at least one distant single nucleotide polymorphism and haplotype of short tandem repeat sequence of interest detected in a biological sample of the patient is determined, and wherein the specific haplotype is determined using the method according to any one of claims 1, 3 to 11.
17. Use of a haplotype of at least one distant single nucleotide polymorphism and haplotype of short tandem repeat sequence of interest detected in the Huntingtin (HTT) gene determined in a biological sample of an individual for diagnosing Huntington’s disease, wherein the detection of a specific haplotype of the at least one distant single nucleotide polymorphism and haplotype of short tandem repeat sequence of interest detected in a biological sample of a patient is indicative of the individual suffering from Huntington’s disease, and wherein the specific haplotype is determined using the method according to any one of claims 1, 3 to 11.
18. Use of the detection of at least one polymorphic site of interest and haplotype of short tandem repeat sequence detected in the Huntingtin (HTT) gene determined in a biological sample of a patient suffering from Huntington disease for determining in vitro the patient as likely to respond to a therapy comprising an antisense oligonucleotide specifically hybridizing to at least one polymorphic site of interest in the Huntingtin (HTT) gene, wherein the patient is identified as more likely to respond to the therapy when a specific haplotype of the at least one polymorphic site of interest and the short tandem repeat sequence is detected in the biological sample of the patient, and wherein the specific haplotype is determined using the method according to any one of claims 1, 3 to 11.
Citation Information
Patent Citations
Method for characterization of nucleic acid molecules
US20030104428A1
System with nano-scale conductor and nano-opening
US20040121525A1
Characterization of individual polymer molecules based on monomer-interface interactions
US5795782A
Characterization of individual polymer molecules based on monomer-interface interactions
US6015714A
Miniature support for thin films containing single channels or nanopores and methods for using same
US6267872B1