Method and system for visualizing short reads within repeating regions of the genome

JP7923185B2Active Publication Date: 2026-09-17ILLUMINA INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022580202
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-11
Filing Date
2021-12-10
Publication Date
2026-09-17
Estimated Expiration
2041-12-10

Smart Images

  • Figure 0007923185000005
    Figure 0007923185000005
  • Figure 0007923185000006
    Figure 0007923185000006
  • Figure 0007923185000007
    Figure 0007923185000007
Patent Text Reader

Abstract

Disclosed embodiments relate to methods, devices, systems, and computer program products for genotyping and visualizing repetitive sequences, such as medically significant short tandem repeats (STRs). Some embodiments can be used to genotype and visualize repetitive sequences, each comprising two or more repetitive subsequences. Some embodiments provide a computer tool that generates sequence read pileups for visualizing repetitive sequences of samples having different genotypes of the repetitive sequences, each sequence pileup comprising reads aligned to two or more different haplotypes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Built-in by reference PCT application forms are filed concurrently with this application as part of this application. Each application that this application claims the benefit or priority of what is specified in a concurrently filed PCT application form is incorporated herein by reference in its entirety. [Background technology]

[0002] A tandem repeat (TR) is a tract of repeating DNA in which a specific DNA motif is repeated. In certain practice, if the repeating motif, also referred to herein as a repeat unit, contains fewer than 10 base pairs, a TR is considered a short tandem repeat (STR) or microsatellite. If the repeating motif is in the range of 10 to 60 base pairs, a TR is considered a minisatellite.

[0003] Short tandem repeats (STRs) are ubiquitous throughout the human genome. While our understanding of STR biology is complete, emerging evidence suggests that STRs play a crucial role in fundamental cellular processes.

[0004] Repeat elongation is a condition in which the trigonometric sequence (TR) of an organism has more repeat motifs than its reference sequence. Repeat elongation is also known as dynamic variation, resulting from the instability of STRs when they elongate beyond a certain size. STR elongation is a major cause of numerous severe neurological disorders, including amyotrophic lateral sclerosis (ALS), Friedreich's ataxia (FRDA), Huntington's disease (HD), and fragile X syndrome (FXS).

[0005] Identifying repeat extensions is crucial in the diagnosis and treatment of these genetic diseases. However, determining TRs, particularly STRs, presents many technical challenges. Some of these challenges require the use of short sequence reads that do not fully traverse the repetitive sequences, a technique commonly used in large-scale parallel sequencing. Therefore, it is desirable to develop methods using short sequence reads to detect and genotype medically relevant repeat extensions.

[0006] Due to the numerous technical challenges in detecting repeat elongations and determining genotypes, there is a great need for computer tools to visualize the genotype of STR and the sequence read data used to determine the genotype. Such tools can help validate genotype calls and understand the clinically and biologically important genetic features associated with STR. Various embodiments disclosed herein are designed to detect repeat elongations and determine genotypes, as well as to visualize the genotype of STR and the sequence data used to determine the genotype. [Overview of the project] [Means for solving the problem]

[0007] The disclosed embodiments relate to methods, apparatus, systems, and computer program products for sequencing and graphically visualizing genomic loci containing repetitive sequences, such as short tandem repeat sequences, which may be associated with genetic diseases. The visualization method generates a sequence pileup, which includes a graphical representation of sequence reads aligned to multiple haplotypes, particularly those containing repetitive sequences.

[0008] A first aspect of the present disclosure provides a computer-implemented method for generating computer graphics, each graphic representing sequence reads aligned to a plurality of haplotypes of a genomic region, including, for example, tandem repeats or structural variants. The method is performed using a computer including one or more processors and a system memory. The method comprises: (a) aligning, using one or more processors, a plurality of sequence reads to a set of alignment positions on a plurality of haplotype sequences corresponding to a plurality of haplotypes of a genomic region, wherein the plurality of sequence reads are obtained from the genomic region of a nucleic acid sample; (b) estimating, by the one or more processors, an alignment score for the set of alignment positions; (c) repeating (a) to (b) in a plurality of iterations to obtain a plurality of alignment scores for a plurality of different sets of alignment positions; (d) selecting, by the one or more processors, a set of alignment positions from the plurality of different sets of alignment positions based on the plurality of alignment scores; and (e) generating, using the one or more processors, computer graphics representing the plurality of sequence reads and the plurality of haplotypes, wherein the plurality of sequence reads are aligned to the plurality of haplotypes in the set of alignment positions selected in (d).

[0009] In some embodiments, the alignment score indicates how uniformly the plurality of sequence reads are distributed across the plurality of haplotype sequences.

[0010] In some embodiments, the genomic region comprises one or more tandem repeats.

[0011] In some embodiments, at least one haplotype among the plurality of haplotypes comprises a repeat expansion. In some embodiments, each haplotype comprises an allele. In some embodiments, the plurality of haplotypes comprises two haplotypes.

[0012] In some embodiments, the selected set of alignment positions has the best alignment score among a plurality of different sets of alignment positions. In some embodiments, the selected set of alignment positions has an alignment score that exceeds a selection criterion.

[0013] In some embodiments, at least one haplotype among the plurality of haplotypes comprises a structural variant. In some embodiments, the structural variant is longer than 50 bp and is selected from the group consisting of deletions, duplications, copy number variants, insertions, inversions, translocations, and any combination thereof. In some embodiments, the structural variant comprises a variant shorter than 50 bp. In some embodiments, the variant shorter than 50 bp comprises a single nucleotide polymorphism (SNP).

[0014] In some embodiments, (a) comprises: (i) determining possible alignment positions of each read for each haplotype, wherein the plurality of sequence reads comprises read pairs obtained by paired-end sequencing; (ii) generating a constrained alignment position for each read pair from the alignment positions of the constituent reads, such that (A) both reads of the read pair align to the same haplotype, and (B) the corresponding fragment length of the read pair is as close as possible to the average fragment length; and (iii) randomly selecting an alignment position for each read pair from the constrained alignment positions.

[0015] In some embodiments, the alignment score comprises a root-mean-square deviation from the average of the distances between the start positions of two consecutive reads.

[0016] In some embodiments, the alignment score is estimated using a probabilistic model that assumes read pairs are uniformly distributed across multiple haplotype sequences. In some embodiments, the alignment score includes probabilities of multiple sequence reads derived from a set of alignment positions, taking the probabilistic model into account. In some embodiments, the multiple sequence reads include paired-end reads obtained from nucleic acid fragments, and the probabilistic model is configured to accept the mean fragment length as input. In some embodiments, the probabilistic model is configured to accept the haplotype length as input.

[0017] In some embodiments, p (k) The probability x of the individual alignment position of the k-th read pair from the start of the haplotype, as shown by, can be modeled as follows:

[0018]

number

[0019]

number

[0020] In some embodiments, the alignment score for a set of alignment positions is estimated as the product of the probabilities of the individual alignment positions.

[0021] In some embodiments, the method further includes estimating one or more sequencing metrics for multiple sequence reads aligned to multiple haplotype sequences in a selected set of alignment positions. In some embodiments, one or more sequencing metrics include sequence coverage. In some embodiments, one or more sequencing metrics include sequence coverage for each alignment position. In some embodiments, one or more sequencing metrics include alignment quality scores. In some embodiments, one or more sequencing metrics include alignment quality scores for each alignment position. In some embodiments, one or more sequencing metrics include mapping quality scores.

[0022] In some embodiments, the plurality of sequence reads include at least 100 sequence reads.

[0023] In some embodiments, the above method further includes performing operation (a) on different genomic regions using different sets of sequence reads. In some embodiments, the different genomic regions include at least 100 different genomic regions.

[0024] In some embodiments, the above method further comprises, prior to operation (a), aligning a first number of sequence reads to one or more sequence graphs corresponding to genomic regions to obtain a plurality of sequence reads and / or a plurality of haplotypes. In some embodiments, aligning a first number of sequence reads to a sequence graph comprises (i) providing a first number of sequence reads of a nucleic acid sample; (ii) aligning a first number of sequence reads to one or more repetitive sequences represented by a sequence graph, wherein the sequence graph has a directed graph data structure having vertices representing nucleic acid sequences and directed edges connecting the vertices, the sequence graph includes one or more self-loops, each self-loop representing a repetitive subsequence, and each repetitive subsequence containing a repeat of one or more repeating units of nucleotides; (iii) determining one or more genotypes of one or more repetitive sequences; and (iv) providing a first number of sequence reads as a plurality of sequence reads in (a) and / or providing one or more genotypes of one or more repetitive sequences.

[0025] In some embodiments, the method further comprises phasing one or more genotypes of one or more repetitive sequences to determine multiple haplotypes of (b). In some embodiments, the method further comprises first aligning a second number of sequence reads into a genome to provide a first number of sequence reads, the second number of sequence reads comprising at least 10,000 sequence reads.

[0026] Another aspect of the present disclosure provides a system for generating computer graphics, each graphic representing sequence reads aligned to multiple haplotypes of a genomic region.

[0027] In some embodiments, the system also includes a sequencer for sequencing the nucleic acids of a test sample.

[0028] In some embodiments, one or more processors are configured to perform various methods described herein.

[0029] Another aspect of the present disclosure provides a computer program product that includes a non-temporary, machine-readable medium for storing program code that causes a computer system to perform the above method for generating computer graphics when executed by one or more processors of the computer system, wherein each graphic represents sequence reads aligned to multiple haplotypes of a genomic region.

[0030] In some embodiments, the program code includes code for performing the operation of the method described herein. [Brief explanation of the drawing]

[0031] [Figure 1A] This is a schematic diagram illustrating the difficulties in aligning sequence reads with respect to repetitive sequences in a reference sequence. [Figure 1B] To overcome the difficulties shown in Figure 1A, a schematic diagram illustrating the alignment of sequence leads using paired-end leads according to a specific disclosed embodiment is provided. [Figure 1C] This diagram shows tandem repetition using the CAG motif. [Figure 1D] The diagram shows paired reads generated by sequencing tandem repetitions longer than the read length. [Figure 2A] This illustrates a scenario where aligning the leads to the TR region is difficult, even when using paired-end leads. [Figure 2B] This illustrates a scenario where aligning the leads to the TR region is difficult, even when using paired-end leads. [Figure 3A] A schematic representation of a conventional lead pile-up is shown below. [Figure 3B] Several embodiments of lead pile-up are schematically shown. [Figure 4] The following outlines a general workflow for generating a read pile-up using several embodiments. [Figure 5]A flowchart of process 50 for generating computer graphics representing sequence reads aligned to the haplotypes of genomic regions is shown. [Figure 6] A flowchart of process 600 for generating computer graphics representing a sequence read pileup containing multiple haplotypes is shown. [Figure 7] A flowchart of Process 700 for aligning sequence reads to a set of alignment positions is shown. [Figure 8] A flowchart illustrating the process for genotyping genomic loci containing repetitive sequences, according to several embodiments, is shown. [Figure 9] The first sequence graph representing the first genomic locus is shown. [Figure 10] The second sequence graph representing the second genomic locus is shown. [Figure 11] This shows a third sequence graph representing the third genomic locus. [Figure 12] A schematic diagram of the process for determining the genotype of a variant at the HTT locus containing two STR sequences, according to several embodiments, is shown. [Figure 13] Schematic diagrams of processes for determining the genotype of variants at the Lynch I locus, including SNVs and STRs, according to several embodiments are shown. The left panel of Figure 12 shows a schematic diagram of a general process for targeted genotyping, and the right panel shows the application of this process to genotyping variants at loci associated with Lynch I syndrome. [Figure 14] This flowchart provides a high-level description of one example of a method for determining whether or not repetitive sequences in a sample are elongated. [Figure 15] This flowchart illustrates an example of a method for detecting repeat extensions using paired-end leads. [Figure 16] This flowchart illustrates an example of a method for detecting repeat extensions using paired-end leads. [Figure 17]This is a flowchart illustrating a method for determining repeat extension using unaligned reads that are not related to any given iterative sequence. [Figure 18] This is a block diagram of a distributed system for processing test samples. [Figure 19] This shows a read pileup of ATXN3 iterations performed according to several embodiments. [Figure 20] The read pile-up of DMPK iterations, performed according to several embodiments, is shown. [Figure 21A] The read pile-up of the HTT locus, carried out according to several embodiments, is shown. [Figure 21B] This shows the read pileup of the HTT gene locus generated by conventional methods. [Figure 22] The following describes a read pileup involving an incorrectly called extension of C9ORF72 iterations, according to several embodiments. [Figure 23] Several embodiments demonstrate read pileups including incorrectly called extensions of FMR1 iterations. [Modes for carrying out the invention]

[0032] This disclosure relates to methods, apparatus, systems, and computer program products for identifying and visualizing target repeat elongations, such as medically significant repeat sequence elongations. Examples of repeat elongations include, but are not limited to, those associated with genetic diseases such as fragile X syndrome, ALS, Huntington's disease, Friedreich's ataxia, spinocerebellar degeneration, bulbar spinal muscular atrophy, myotonic dystrophy, Machado Joseph disease, and dentatorubral-pallidoluysian atrophy.

[0033] Unless otherwise specified, the methods and systems disclosed herein include conventional techniques and apparatus commonly used in molecular biology, microbiology, protein purification, protein engineering, protein and DNA sequencing, as well as the field of recombinant DNA within the scope of the art. Such techniques and apparatus are known to those skilled in the art and are described in numerous texts and reference studies (see, for example, Sambrook et al., "Molecular Cloning: A Laboratory Manual," Third Edition (Cold Spring Harbor),

[2001] , and Ausubel et al., "Current Protocols in Molecular Biology"

[1987] ).

[0034] A numerical range includes the number that defines that range. All maximum numerical limits given throughout this specification are intended to include any lower numerical limits, as if such lower numerical limits were explicitly stated herein. All minimum numerical limits given throughout this specification include any higher numerical limits, as if such higher numerical limits were explicitly stated herein. All numerical ranges given throughout this specification include any narrower numerical ranges that fall within such wider numerical ranges, as if all such narrower numerical ranges were explicitly stated herein.

[0035] The headings provided herein are not intended to limit this disclosure.

[0036] The examples herein relate to humans, and the language is primarily human; however, the concepts described herein are applicable to genomes from any plant or animal. These and other purposes and features of the Disclosure may be more fully apparent from the following description and the appended claims, or may be learned by the practices of the Disclosure described below.

[0037] Unless otherwise defined herein, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art. Various scientific dictionaries containing the terms included herein are known and available in the art. While it has been found that any methods and materials similar to or equivalent to those described herein can be used in carrying out or testing the embodiments disclosed herein, only a few methods and materials are described herein.

[0038] The terms defined below are more fully described by referring to the specification as a whole. It should be understood that this disclosure may be modified depending on the context in which it is used by those skilled in the art, and is therefore not limited to the specific methodologies, protocols, and reagents described.

[0039] definition In the present invention, the singular forms "a," "an," and "the" include multiple references unless the context explicitly indicates otherwise.

[0040] Unless otherwise specified, nucleic acids are written from left to right with a 5' to 3' orientation, and amino acid sequences are written from left to right with an amino-to-carboxyl orientation.

[0041] The term "plural" means two or more elements. For example, as used herein, the term refers to a number of nucleic acid molecules or sequence reads sufficient to identify significant differences in repeat extension between a test sample and a control sample using the methods disclosed herein.

[0042] The term “haplotype” is used herein to refer to a set of alleles within a cluster of linked genes on a chromosome. In various embodiments herein, a haplotype includes the TR allele.

[0043] The term "haplotype sequence" refers to a sequence of consecutive genes containing a set of alleles on a chromosome. For example, a haplotype sequence may contain two flanking regions and an STR sequence (e.g., Figure 20), or it may contain two flanking regions and two nearby STR sequences flanking an intervening sequence (e.g., Figures 21A and 21B).

[0044] The term “repetitive sequence” means a nucleic acid sequence that contains the repetitive occurrence of shorter sequences. These shorter sequences are referred to herein as “repetitive units” or “repetitive motifs,” or simply “motifs.” The repetitive occurrence of repeating units is referred to as “repeat” or “replica” of the repeating unit. In many contexts, the location of a repeating sequence is associated with a protein-coding gene. In other situations, repeating sequences may be located within non-coding regions. Repeating units may occur in repeating sequences with or without breaks between repeating units. For example, in a normal sample, the FMR1 gene tends to contain AGG breaks in CGG repeats, such as (CGG)10+(AGG)+(CGG)9. Samples without breaks, as well as longer repeating sequences with some breaks, are prone to repeated repeat elongation of the associated gene, which can lead to genetic disease when the repeats elongate beyond a certain number. In various embodiments of this disclosure, the number of repeats, with or without breaks, is counted as in-frame repeats. Methods for estimating in-frame repeats are further described below.

[0045] In various embodiments, each repeating unit contains 1 to 100 nucleotides. Many widely studied repeating units are trinucleotide or hexanenucleotide units. Some other repeating units that have been well studied and are applicable to the embodiments disclosed herein include, but are not limited to, units of 4, 5, 6, 8, 12, 33, or 42 nucleotides. See, for example, Richards (2001) Human Molecular Genetics, 10, No. 20, 2187-2194. The applications of this disclosure are not limited to the above-mentioned specific numbers of nucleotide bases, as long as they are relatively short compared to repeat sequences having multiple repeats or copies of repeating units. For example, repeating units may contain at least 3, 6, 8, 10, 15, 20, 30, 40, or 50 nucleotides. Alternatively or additionally, repeating units may contain up to approximately 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 6, or 3 nucleotides.

[0046] Repetitive sequences can be elongated under evolutionary, developmental, and mutational conditions, allowing for the creation of more copies of the same repeating unit. This is referred to in the field as “repeat elongation.” This process is also called “dynamic mutation,” due to the unstable nature of repeating unit elongation. Some repeat elongations have been shown to be associated with genetic diseases and pathological symptoms. Other repeat elongations are not well understood or studied. Both known and novel repeat elongations can be identified using the methods disclosed herein. In some embodiments, repetitive sequences with repeat elongation are longer than about 100, 150, 300, or 500 base pairs (bp). In some embodiments, repetitive sequences with repeat elongation are longer than about 1000 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, or 10000 bp, etc.

[0047] In graph theory, vertices and edges are the two fundamental units upon which a graph is constructed. A vertex or node is one of the points on a graph that can be defined and connected by edges. In graph diagrams, vertices can be represented by shapes with labels, and edges can be represented by lines (undirected edges) or arrows (directed edges) extending from one vertex to another.

[0048] Two vertices connected by an edge are said to be endpoints of the edge. If a graph contains an edge (x, y), then vertex x is said to be adjacent to another vertex y.

[0049] An undirected graph consists of a set of vertices and a set of undirected edges (connecting irregular pairs of vertices), while a directed graph consists of a set of vertices and a set of directed edges (connecting regular pairs of vertices).

[0050] In graph theory, each edge has two (or more in hypergraphs) attached vertices called its endpoints. Edges can be directed or undirected; undirected edges are also called lines, while directed edges are also called arcs or arrows.

[0051] A directed edge is an edge that connects an upstream vertex to a downstream vertex, where the upstream vertex appears before the directed edge and the downstream vertex appears after the directed edge.

[0052] An undirected edge is an edge that connects two vertices, and either vertex can appear before the other in the graph path.

[0053] Loops, self-loops, and single-node loops are used interchangeably in this specification. A loop has one node and ends that are connected to one node at both ends.

[0054] A cycle is a path containing two or more vertices, and the path of a cycle begins and ends at the same vertex. A simple cycle is one that has no repeating vertices or edges other than the start and end vertices.

[0055] A ring graph is a graph that contains at least one cycle.

[0056] An acyclic graph is a graph that does not contain any cycles or self-loops.

[0057] A directed acyclic graph (DAG) is a directed graph that does not have any arbitrary cycles or self-loops.

[0058] A graph path is an array of vertices and edges, where the endpoints of an edge appear adjacent to an edge in the array. In a directed graph, a graph path has upstream vertices that appear before a directed edge (or arc or arrow) and downstream vertices that appear after the directed edge.

[0059] The Poisson distribution is a discrete probability distribution that represents the probability of a given number of events occurring at a given time or spatial interval, given that these events occur at a known constant rate and independently of the time elapsed since the last event.

[0060] The fully specified base abbreviations include G, A, T, and C, which represent guanine, adenine, thymine, and cytosine, respectively.

[0061] Nucleic acid nomenclature that is not fully specified includes, among others: Purines (adenine or guanine): R Pyrimidine (thymine or cytosine): Y Adenine or thymine: W Guanine or cytosine: S Adenine or cytosine: M Guanine or thymine: K Adenine or thymine or cytosine:H Guanine, cytosine, or thymine: B Guanine, adenine, or cytosine: V Guanine, adenine, or thymine: D Guanine or adenine or thymine or cytosine: N

[0062] The term “paired-end read” refers to reads obtained from paired-end sequencing, where one read is obtained from each end of a nuclear fragment. Paired-end sequencing involves fragmenting DNA into sequences called inserts. In some protocols used by Illumina, reads from shorter inserts (e.g., about 10 to several hundred bp) are referred to as short-insert paired-end reads, or simply paired-end reads. In contrast, reads from longer inserts (e.g., about several thousand bp) are referred to as mate-pair reads. In this disclosure, both short-insert paired-end reads and long-insert mate-pair reads may be used and are not distinguished with respect to the process for analyzing repeat extension. Thus, the term “paired-end read” may also mean both short-insert paired-end reads and long-insert mate-pair reads, which will be further described later herein. In some embodiments, paired-end reads include reads of about 20 bp to 1000 bp. In some embodiments, paired-end reads include reads of approximately 50 bp to 500 bp, approximately 80 bp to 150 bp, or approximately 100 bp. It will be understood that the two reads of a paired-end do not need to be located at the very end of the fragment being sequenced. Rather, one or both reads can be close to the end of the fragment. Furthermore, the methods illustrated herein in the context of paired-end reads can be performed with any of the various paired reads, regardless of whether the reads are derived from the end portion of the fragment or other parts of the fragment.

[0063] In the present invention, the terms “alignment” and “aligned” mean the process of comparing a read to a reference sequence to determine whether the reference sequence contains the read sequence. The alignment process attempts to determine whether a read can be located in the reference sequence, but a read is not always aligned to the reference sequence. If the reference sequence contains a read, the read may be located in the reference sequence, or, in certain other embodiments, may be mapped to a specific location within the reference sequence. In some cases, alignment simply indicates whether a read is a member of a particular reference sequence (i.e., whether the read is present or not in the reference sequence). For example, alignment of a read to a reference sequence for human chromosome 13 indicates whether the read is present in the reference sequence for chromosome 13. The tool that provides this information may be called a set membership tester. In some cases, alignment further indicates the location within the reference sequence where the read map exists. For example, if the reference sequence is the entire human genome sequence, alignment may indicate that the read is present on chromosome 13, and further may indicate that the read is located on a specific strand and / or region of chromosome 13.

[0064] Aligned reads are one or more sequences identified as being consistent with respect to the order of their nucleic acid molecules relative to a known reference sequence, such as a reference genome. Aligned reads on a reference sequence and their determined positions constitute a sequence tag. Alignment can be performed manually, but is typically carried out by computer algorithms because it is impossible to align reads at a reasonable time interval to implement the methods disclosed herein. An example of a sequence alignment algorithm is the efficient local alignment of the Distributed Nucleotide Data (ELAND) computer program as part of the Illumina Genomics Analysis pipeline. Alternatively, reads can be aligned to a reference genome using a Bloom filter or a similar set membership tester. See U.S. Patent Application No. 14 / 354,528 (filed April 25, 2014), which is incorporated herein by reference in its entirety. Sequence read consistency may be 100% sequence consistency or less than 100% consistency (i.e., incomplete consistency).

[0065] As used herein, the term "mapping" means assigning read sequences to a larger sequence, such as a reference genome, through alignment.

[0066] In some cases, one end of two paired-end reads aligns to a repeating sequence in the reference sequence, while the other end of the paired-end reads is not. In such cases, the paired-end read aligned to the repeating sequence in the reference sequence is called the “anchor read.” A paired-end read that is not aligned to a repeating sequence but is paired with the anchor read is called an anchor-type read. Thus, unaligned reads can be anchored to and associated with repeating sequences. In some embodiments, unaligned reads include both reads that cannot be aligned to the reference sequence and reads that are poorly aligned to the reference sequence. A read is considered poorly aligned if it aligns to a reference sequence that has more mismatched bases than a certain criterion. For example, in various embodiments, a read is considered poorly aligned if it aligns to at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mismatches. In some cases, both pairs of reads align to the reference sequence. In such cases, in various embodiments, both reads may be analyzed as “anchor reads.”

[0067] The terms "polynucleotide," "nucleic acid," and "nucleic acid molecule" are used interchangeably and refer to a covalent-like sequence of nucleotides in which the 3' position of one nucleotide pentose is bonded to the 5' position of the next pentose by a phosphodiester group (i.e., ribonucleotides in the case of RNA, and deoxyribonucleotides in the case of DNA). Nucleotides include, but are not limited to, RNA molecules and DNA molecules such as cell-free DNA (cfDNA) molecules, as well as any form of nucleic acid sequence. The term "polynucleotide" includes, but is not limited to, single-stranded and double-stranded polynucleotides.

[0068] In this specification, the term “test sample” means a sample, typically derived from a biological fluid, cell, tissue, organ, or organism, comprising nucleic acids or mixtures of nucleic acids having at least one nucleic acid sequence to be screened for copy number variation. In certain embodiments, the sample has at least one nucleic acid sequence whose replication number is thought to be mutated. Such samples include, but are not limited to, sputum / oral fluid, amniotic fluid, blood, blood fractions, or microneedle biopsy samples, urine, peritoneal fluid, pleural fluid, etc. Samples are often taken from human subjects (e.g., patients), but the assay can be used for copy number variation (CNV) in samples of any mammal, including, but not limited to, dogs, cats, horses, goats, sheep, cattle, pigs, etc. Samples may be used directly from the biological source or after pretreatment to modify the characteristics of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids, etc. Pretreatment methods may include, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, addition of reagents, and dissolution. When such pretreatment methods are employed on a sample, the target nucleic acid typically remains in the test sample, sometimes at a concentration proportional to that in the untreated test sample (e.g., a sample not subjected to such pretreatment). Such “treated” or “processed” samples are still considered biological “test” samples with respect to the methods described herein.

[0069] The control sample may be a negative control sample or a positive control sample. A "negative control sample" or "unaffected sample" means a sample containing nucleic acids known or expected to have a repetitive sequence with a large number of repeats within a non-pathogenic range. A "positive control sample" or "affected sample" is known or expected to have a repetitive sequence with a large number of repeats within a pathogenic range. The repeats in the repetitive sequence in the negative control sample are usually not extended beyond the normal range, while the repeats in the repetitive sequence in the positive control sample are usually extended beyond the normal range. Therefore, nucleic acids in the test sample can be compared to one or more control samples.

[0070] The term "target sequence" as used herein means a nucleic acid sequence related to differences in sequence representation between healthy individuals and individuals with disease. The target sequence may be a repetitive sequence on a chromosome that is extended due to disease or a genetic condition. The target sequence may be part of a chromosome, gene, coding, or non-coding sequence.

[0071] In this specification, the term “next-generation sequencing (NGS)” refers to sequencing methods that enable large-scale parallel sequencing of clonely amplified molecules and single nucleic acid molecules. Non-exclusive examples of NGS include sequencing-by-synthesis and sequencing-by-ligation using reversible diterminators.

[0072] In this specification, the term "parameter" means a numerical value that characterizes a physical property. Often, parameters numerically characterize numerical relationships between quantitative datasets and / or quantitative datasets. For example, the ratio (or function of the ratio) of the number of sequence tags located on a chromosome to the length of the chromosome to which the tags are mapped is a parameter.

[0073] In this specification, the term “call criterion” means any number or amount used as a cutoff for characterizing a sample, such as a test sample containing nucleic acids, from an organism suspected of having a medical condition. By comparing this threshold to a parameter value, it may be determined whether a sample producing such a parameter value suggests that the organism has a medical condition. In certain embodiments, the threshold is calculated using a control dataset and serves as a diagnostic limit for repeat elongation in the organism. In some embodiments, if the results obtained from the methods disclosed herein exceed the threshold, the subject may be diagnosed with repeat elongation. An appropriate threshold for the methods described herein may be identified by analyzing values ​​calculated for a training set of samples or control samples. Thresholds can also be calculated from empirical parameters such as sequencing depth, read length, and repeat sequence length. Alternatively, affected samples known to have repeat elongation can be used to verify that the selected threshold is useful in distinguishing affected samples from unaffected samples in a test set. The selection of a threshold depends on the level of confidence the user desires to make the classification. In some embodiments, the training set used to identify an appropriate threshold includes at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 2000, at least 3000, at least 4000, or more eligible samples. Using a larger set of eligible samples may be advantageous in improving the diagnostic utility of the threshold.

[0074] The term "read" refers to a sequence read from a portion of a nucleic acid sample. Typically, but not always, a read represents a short sequence of consecutive base pairs in the sample. A read may be symbolically represented by the base-pair sequence (ATCG) of the sample portion. Reads may be stored in a memory device and appropriately processed to determine whether they match a reference sequence or meet other criteria. Reads may be obtained directly from a sequencing instrument or indirectly from stored sequence information about the sample. In some cases, they may be DNA sequences of sufficient length (e.g., at least about 25 bp) to identify larger sequences or regions that can be aligned and positioned on a chromosome or genomic region or gene.

[0075] The term "genome read" is used to refer to a read of any segment within the entire genome of an individual.

[0076] The term "site" refers to a specific location on the reference genome (e.g., chromosome ID, chromosome position, and orientation). In some embodiments, the site may be the location of a residue, a sequence tag, or a segment on a sequence.

[0077] As used in this invention, the terms “reference genome” or “reference sequence” refer to a specific known genome sequence, either partial or complete, of any organism or virus that can be used to reference a specified sequence from a subject. For example, reference genomes used for human subjects, as well as many other organisms, can be found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. “Genome” means the complete genetic information of an organism or virus as expressed in a nucleic acid sequence.

[0078] In various embodiments, the reference sequence is significantly larger than the reads aligned to it. For example, it may be at least about 100 times larger, or at least about 1,000 times larger, or at least about 10,000 times larger, or at least about 10 5Twice as large, or at least about 10 6 Twice as large, or at least about 10 7 It can be twice as large.

[0079] In one embodiment, the reference sequence is from the full-length human genome. Such sequences are sometimes called genome reference sequences. In another embodiment, the reference sequence is limited to a specific human chromosome, such as chromosome 13. In some embodiments, the reference Y chromosome is the Y chromosome sequence from the human genome version hg19. Such sequences are sometimes called chromosome reference sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, subchromosome regions (such as strands) of any species.

[0080] In some embodiments, the reference sequence for alignment may have a sequence length of about 1 to about 100 times the length of the read. In such embodiments, alignment and sequencing are considered as target alignment or sequencing rather than the entire genome alignment or sequencing. In these embodiments, the reference sequence typically includes a gene and / or a target repeat sequence.

[0081] In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, for specific applications, the reference sequence may be taken from a specific individual.

[0082] In this specification, the term “clinically relevant sequence” means a nucleic acid sequence that is known or suspected to be related to or suggestive of a genetic or pathological condition. Determining the presence or absence of clinically relevant sequences may be useful in determining a diagnosis, confirming a diagnosis of a medical condition, or providing a prognosis for the onset of a disease.

[0083] The term "derived from," when used in the context of nucleic acids or mixtures of nucleic acids, means herein the means by which nucleic acids are obtained from a source from which they originate. For example, in one embodiment, a mixture of nucleic acids derived from two different genomes means that the nucleic acids, e.g., cfDNA, were spontaneously released by cells through spontaneous processes such as necrosis or apoptosis. In another embodiment, a mixture of nucleic acids derived from two different genomes means that the nucleic acids were extracted from two different types of cells from a subject.

[0084] When the term "based on" is used in the context of obtaining a specific quantitative value, it means using another quantity as input to calculate a specific quantitative value as output.

[0085] In this specification, the term “patient sample” means a biological sample obtained from a patient, i.e., a recipient of medical attention, care, or treatment. A patient sample may be any of the samples described herein. In certain embodiments, a patient sample may be obtained by a non-invasive procedure, e.g., a peripheral blood sample or a fecal sample. The methods described herein are not limited to humans. Therefore, a variety of veterinary applications can be conceived in which the patient sample may be a sample from a non-human mammal (e.g., a cat, pig, horse, or cow).

[0086] In this specification, the term “biological fluid” means a liquid taken from a biological source, including, for example, blood, serum, plasma, sputum, lavage fluid, cerebrospinal fluid, urine, semen, sweat, tears, saliva, etc. As used in the present invention, the terms “blood,” “plasma,” and “serum” expressly include their fractions or processed portions. Similarly, when a sample is taken from a biopsy, swab, smear, etc., “sample” expressly includes the processed fraction or portion obtained from the biopsy, swab, smear, etc.

[0087] In the present invention, the term "corresponding to" may mean a nucleic acid sequence present in the genome of different subjects, such as a gene or chromosome, which does not necessarily have the same sequence in all genomes, but serves to provide identity rather than genetic information of the target sequence (e.g., a gene or chromosome).

[0088] As used in the present invention, the term "chromosome" means a gene carrier in a living cell having the efficacy of the present invention, derived from chromatin strands containing DNA and protein components (particularly histones). Conventional, internationally recognized individual human genome chromosome numbering systems are used herein.

[0089] As used in this invention, the term "polynucleotide length" means the absolute number of nucleic acid monomer subunits (nucleotides) in a sequence or within a region of a reference genome. The term "chromosome length" refers to the known length of a chromosome given in base pairs, as provided in NCBI36 / hg18 of human chromosomes, which can be found, for example, on the World Wide Web at |genome|||ucsc|||edu / cgi-bin / hgTracks?hgsid=167155613&chromInfoPage=.

[0090] In this specification, the term "subject" refers not only to human subjects but also to non-human subjects such as mammals, invertebrates, vertebrates, fungi, yeasts, bacteria, and viruses. While the examples and language of this specification relate primarily to humans, the concepts disclosed herein are applicable to genomes from any plant or animal and are useful in veterinary medicine, animal science, laboratories, and related fields.

[0091] In the present invention, the term "primer" means an isolated oligonucleotide that can act as a starting point for synthesis when placed under conditions inducible to the synthesis of the extension product (for example, conditions including nucleotides, an inducer such as DNA polymerase, and suitable temperature and pH). The primer may preferably be single-stranded or double-stranded for maximum efficiency in amplification. If double-stranded, the primer is first treated to separate its strands before being used to prepare the extension product. The primer may also be an oligodeoxyribonucleotide. The primer is of sufficient length to prime the synthesis of the extension product in the presence of an inducer. The exact length of the primer depends on many factors, including temperature, primer source, method used, and parameters used in primer design.

[0092] Introduction Sequences consisting of repeats of relatively short pieces of DNA, known as tandem repeats (TRs), occur throughout the genome (e.g., Figure 1C). TR mutation rates can be 10 to 1000 times higher than in other genomic regions, and as a result, TRs can be a major source of human genetic variation. TRs mutate on a large scale through "slippage," where the number of repeats increases or decreases between generations. Accumulating evidence indicates that TRs play a role in fundamental cellular processes, and large elongations of tandem repeats are associated with a variety of neurological disorders, including amyotrophic lateral sclerosis (ALS), fragile X syndromes, and various forms of ataxia.

[0093] Sequencing a region containing TRs generates a set of reads that partially or completely overlap with the repetitive sequence (Figure 1D). By stitching together the alignments of these reads, the length of the repeats on each haplotype can be determined. The inventors have developed several methods for both targeted and whole-genome TR analysis. Below, the applicant describes ExpansionHunter, a method for targeted analysis of a region containing one or more adjacent TRs that can estimate the size of both repeats shorter and longer than the read length. Also described are methods, systems, and computer program products for visualizing and displaying the alignments of reads generated by ExpansionHunter.

[0094] TR genotyping is an extremely difficult problem, and even the best methods can produce incorrect genotype calls. Therefore, it is crucial to have a robust visualization method for examining the alignment of reads used to genotype the repeat in question. Furthermore, such a visualization method would enable the detection of changes in the repeat motif (e.g., interruptions) that could have clinically significant effects. Standard data visualization pipelines are typically limited to displaying read alignment relative to a reference genome and are therefore insufficient for repeats that are extended relative to the reference or repeats with alleles of different lengths. To address these issues, the inventors of this disclosure have developed Repeat Expansion Viewer (REViewer), a tool for visualizing graph-realigned reads output by ExpansionHunter. REViewer determines haplotype sequences by phasing adjacent repeats and then distributes the read alignment to these haplotypes. The resulting static images allow for a visual evaluation of the accuracy of a given genotype call and identification of whether the repeat sequence contains any interruptions.

[0095] Repeat extensions have biological and medical significance, but detecting them and genotyping STRs using short reads is difficult, as described below. There is a great need to develop techniques for detecting and genotyping repeat extensions, as well as computer-implemented tools for visualizing sequence read data and the genotypes determined from that data. Such tools can help validate genotype calls and understand the clinically and biologically important genetic features associated with STRs.

[0096] STR elongation is a major cause of many severe neurological disorders. Table 1 illustrates a small number of pathogenic repeat elongations that differ from the repeat sequences in normal samples. The columns show the gene associated with the repeat sequence, the nucleic acid sequence of the repeat unit, the exemplary number of repeats in the repeat unit of normal and pathogenic sequences (different repeat cutoffs may be used for different applications), and the disease associated with the repeat elongation.

[0097] [Table 1]

[0098] Genetic disorders involving repeat elongation are heterogeneous in many respects. The size, degree of elongation, location relative to the affected gene, and pathogenic mechanism can vary from disease to disease. For example, ALS involves elongation of the hexane nucleotide repeat GGGGCC in the C9orf72 gene, located on the short arm of the open reading frame 72 on chromosome 9. In contrast, fragile X syndrome is associated with elongation of a triplet repeat of CGG affecting the fragile X chromosome 1 (FMR1) gene on the X chromosome. The elongation of the CGG repeat prevented the expression of fragile X chromosome malformation protein (FMRP), which is necessary for normal neurodevelopment. Depending on the length of the CGG repeat, the allele can be classified as normal (not affected by the syndrome), pre-mutant (at risk of fragile X chromosome-related disorders), or complete mutation (usually affected by the syndrome). According to various estimates, the FMR1 gene mutation causing fragile X syndrome in affected individuals has 230–4000 CGG repeats, compared to 60–230 repeats in carriers with a tendency towards ataxia and 5–54 repeats in unaffected individuals. Repeat elongation in the FMR1 gene is a cause of autism because approximately 5% of autistic individuals have FMR1 repeat elongation. McLennan, et al. (2011), Fragile X Syndrome, Current Genomics 12(3):216-224. Definitive diagnosis of fragile X syndrome involves genetic testing to determine the number of CGG repeats.

[0099] Several studies have identified various common characteristics of repeat elongation-related disorders. Repeat elongation, or dynamic mutation, typically manifests as an increase in the number of repeats, and the mutation rate is related to the number of repeats. Rare events, such as loss of repeat interruptions, can result in alleles with increased elongation potential; such events are known as founder events. A relationship may exist between the number of repeats in a repeat sequence and the severity and / or onset of disorders caused by repeat elongation.

[0100] Therefore, identifying and calling repeat extensions is crucial in the diagnosis and treatment of various diseases. However, identifying repeat sequences, especially using reads that do not completely traverse them, presents several challenges. Firstly, aligning repeats to reference sequences is difficult because there is no clear one-to-one mapping between reads and the reference genome. In addition, even when reads are aligned to reference sequences, they are often too short to fully cover medically relevant repeat sequences. For example, a read may be about 100 bp. In comparison, repeat extensions can span hundreds to thousands of base pairs. In fragile X syndromes, for example, the FMR1 gene can have well-defined repeats spanning over 3000 bp, exceeding 1000. Therefore, a 100 bp read cannot map the entire length of the repeat extension. Furthermore, assembling short reads to longer sequences may not overcome the short-read vs. long-repetition problem, because ambiguous alignment between repeats in one read and repeats on another makes it difficult to assemble short reads to longer sequences.

[0101] Alignment is a primary cause of information loss, resulting from either incompleteness of the reference sequence, nonspecific correspondence between reads and sites on the reference sequence, or significant deviations from the reference sequence. Systematic sequencing errors and other issues affecting read accuracy are secondary causes of failure in detecting repeat sequences. In some experimental protocols, approximately 7% of reads are unaligned or have a MAPQ score of 0. Even when researchers work to improve sequencing techniques and analytical tools, a considerable number of unaligned or poorly aligned reads may always exist. Embodiments of the method described herein rely on unaligned or poorly aligned reads to identify repeat extensions.

[0102] Methods using long reads to detect repeat extensions have their own challenges. In next-generation sequencing, currently available techniques using longer reads tend to be slower and more error-prone than techniques using shorter reads. Furthermore, long reads are not viable for some applications, such as sequencing cell-free DNA. Cell-free DNA obtained from maternal blood can be used for prenatal genetic diagnostics. Cell-free DNA typically exists as fragments shorter than 200 bp. Therefore, methods using long reads are not viable for prenatal genetic diagnostics using cell-free DNA. Embodiments of the methods described herein use short reads to identify medically relevant repeat extensions.

[0103] Furthermore, conventional methods are not designed to handle complex loci with multiple repeats. Important examples of such loci include the CAG repeat, which causes HD and is located lateral to the CCG repeat; the GAA repeat, which causes FRDA and is located lateral to the adenosine homopolymer; and the CAG repeat, which causes spinocerebellar degeneration type 8 (SCA8) and is located lateral to the CT repeat. A more extreme example is the CCTG repeat in the CNBP gene, whose elongation causes myotonic dystrophy type 2 (DM2). This repeat is adjacent to polymorphic TG and TCTG repeats (JELee and Cooper 2009), making precise alignment of reads to this locus particularly difficult. Another type of complex repeat is the polyalanine repeat (Shoubridge and Gecz 2012), which is associated with at least nine diseases. Polyalanine repeats consist of repeats of the α-amino acid codons GCA, GCC, GCG, or GCT.

[0104] Variant clustering can affect the accuracy of alignment and genotyping (Lincoln et al. 2019). Variants adjacent to less complex polymorphic sequences can be even more problematic because methods for variant discovery may output clusters of consistently represented or false variant calls in such genomic regions. This is partly due to an increased error rate in such regions in sequencing data (Benjamini and Speed ​​2012; Dolzhenko et al. 2017). One example is a single nucleotide variant (SNV) adjacent to an adenosine homopolymer in MSH2 that causes Lynch syndrome I (Froggatt et al. 1999).

[0105] The embodiments disclosed herein can handle complex gene loci such as those described above. They use sequence graphs as a general and flexible model for each target gene locus.

[0106] In some embodiments, the disclosed method addresses the aforementioned challenges in identifying and calling repeat extensions by utilizing paired-end sequencing. Paired-end sequencing involves fragmenting DNA into sequences called inserts. In some protocols used by Illumina, reads from shorter inserts (e.g., about 10 to several hundred bp) are called short-insert paired-end reads, or simply paired-end reads. In contrast, reads from longer inserts (e.g., about several thousand bp) are called mate-pair reads. As described above, both short-insert paired-end reads and long-insert mate-pair reads may be used in various embodiments of the method disclosed herein.

[0107] Figure 1A is a schematic diagram illustrating the specific difficulties in aligning sequence reads to repeat sequences on a reference sequence, particularly when aligning sequence reads obtained from samples with long repeat sequences that have repeat extensions. At the bottom of Figure 1A is a reference sequence 101 with a relatively short repeat sequence 103, indicated by vertical hatch lines. In the middle of the figure is a hypothetical sequence 105 from a patient sample with a long repeat sequence 107 that has repeat extensions, also indicated by vertical hatch lines. At the top of the figure are sequence reads 109 and 111, indicated at the corresponding site locations of sample sequence 105. In some of these sequence reads, such as read 111, several base pairs originate from the long repeat sequence 107 and are highlighted with circles, also indicated by vertical hatch lines. Reads 111 with these repeats are potentially difficult to align to reference sequence 101 because the repeats do not have clear corresponding positions on reference sequence 101. These potentially misaligned reads cannot be clearly associated with the repeat sequence 103 in the reference sequence 101, making it difficult to obtain information about the repeat sequence and its extension from these potentially misaligned reads 111. Furthermore, since these reads tend to be shorter than the longer repeat sequence 107 with repeat extensions, they cannot directly provide clear information about the identity or location of the repeat sequence 107. In addition, the repeats within the reads 111 are difficult to assemble due to their ambiguous corresponding positions on the reference sequence 101 and the ambiguous relationships between the reads 111. Reads partially arising from the long repeat sequence 107 in the sample, shown as semi-hatched and semi-solid black, may be aligned by bases arising from outside the repeat sequence 107. If there are too few base pairs in the read outside the repeat sequence 107, the read may be poorly aligned or not aligned at all. Therefore, some of these reads with partial repeats may be analyzed as anchored reads, and other reads may be analyzed as anchored reads, as further described below.

[0108] Figure 1B is a schematic diagram illustrating how paired-end reads may be used in some disclosed embodiments to overcome the difficulties shown in Figure 1A. In paired-end sequencing, sequencing is performed from both ends of a nucleic acid fragment in a test sample. Shown at the bottom of Figure 1B are the reference sequence 101 and sample sequence 105, as well as reads 109 and 111, which are equivalent to those shown in Figure 1A. Shown at the top of Figure 1B are the fragment 125 derived from the test sample sequence 105, as well as the primer regions 131 for read 1 and 133 for read 2, for obtaining the two reads 135 and 137 of the paired-end reads. Fragment 125 is also called the insert for the paired-end reads. In some embodiments, the insert may be amplified with or without PCR. Some repeat sequences, such as those containing a large number of GC or GCC repeats, cannot be well sequenced by conventional methods, including PCR amplification. For such sequences, amplification may not involve PCR. For other sequences, amplification may be performed by PCR.

[0109] The insert 125 shown in Figure 1B corresponds to, or is derived from, a section of the sample sequence 105 located to the sides of the two vertical arrows shown in the lower half of the figure. Specifically, the insert 125 has a repeat section 127 that corresponds to a portion of a long repeat 107 in the sample sequence 105. The length of the insert may be adjusted for various applications. In some embodiments, the insert may be somewhat shorter than the repeat sequence or repeat sequence with repeat extensions of the target. In other embodiments, the insert may have a length similar to the repeat sequence or repeat sequence with repeat extensions. Still in further embodiments, the insert may be somewhat longer than the repeat sequence or repeat sequence with repeat extensions. Such an insert may be a long insert for mate-pair sequencing in some embodiments further described below. Typically, the reads obtained from the insert are shorter than the repeat sequences. Because the insert is longer than the reads, paired-end reads can capture signals better from longer extensions of the repeat sequences in the sample than single-end reads.

[0110] The illustrated insert 125 has two lead primer regions 131 and 133 at its two ends. In some embodiments, the lead primer regions are specific to the insert. In other embodiments, the primer regions are introduced into the insert by ligation or extension. The left end of the insert shows the lead 1 primer region 131, which allows hybridization of the lead 1 primer 132 into the insert 125. Extension of the lead 1 primer 132 produces a first lead or lead 1, labeled as 135. The right end of the insert 125 shows the lead 2 primer region 133, which allows hybridization of the lead 2 primer 134 into the insert 125, which initiates a second lead or lead 2, labeled as 137. In some embodiments, the insert 125 may also include an index barcode region (not shown here) which may provide a mechanism for identifying different samples in a multiplex sequencing process. In some embodiments, paired-end reads 135 and 137 can be obtained by sequencing Illumina using a synthesis platform. An example of a sequencing process performed on such a platform will be further described later in the Sequencing Methods section, which produces two paired-end reads and two index reads.

[0111] Next, as shown in Figure 1B, the obtained pair-end reads may be aligned to a reference sequence 101 having a relatively short repeat sequence 103. In this way, the relative position and orientation of the pair of reads are known. This allows unaligned or poorly aligned reads, such as those shown in circle 111, to be indirectly associated with a relatively long repeat sequence 107 in the sample sequence 105 through the corresponding paired read 109 of the read, as seen at the bottom of Figure 1B. In the exemplary embodiment, the reads obtained from pair-end sequencing are approximately 100 bp, and the insert is approximately 500 bp. In this exemplary setting, the relative position of the two pair-end reads is approximately 300 base pairs from their 3' ends, and they have opposite directions. The relationship between the read pairs allows one read to be better associated with the repeat region. In some cases, the first read of the pair aligns with a non-repeating sequence located on the side of the repeat region on the reference sequence, and the second read of the pair does not align well with the reference sequence. For example, referring to the pair of leads 109a and 111a shown in the lower half of Figure 1B, the left lead 109a is the first lead, and the right lead 111a is the second lead. Considering the pair of two leads 109a and 111a, the second lead 111a can be associated with a repeating region 107 in the sample sequence 105, despite the fact that the second lead 111a cannot be aligned to the reference sequence 101. By knowing the distance and direction of the second lead 111a relative to the first lead 109a, the position of the second lead 111a within the long repeating region 107 can be further determined. If there is a break between repeats in the second lead 111a, the position of the break relative to the reference sequence 101 can also be determined. Leads such as the left lead 109a aligned to the reference are referred to as anchor leads in this disclosure. Leads such as the right lead 111a that are not aligned to the reference sequence but are paired with an anchor lead are referred to as anchor-type leads. Therefore, non-aligned sequences can be anchored to and associated with repeat extensions. In this way, short reads can be used to detect long repeat extensions.The challenge of detecting repeat extensions typically increases with the length of the extension due to sequencing difficulties, but the method disclosed herein can detect higher signals from longer repeat extension sequences than from shorter repeat extension sequences. This is because, as the repeat sequence or repeat extension becomes longer, more reads are locked into the extension region, more reads enter the repeat region completely, and more repeats may occur per read.

[0112] Figures 2A and 2B illustrate scenarios where it is difficult to align reads to the TR region even when using paired-end reads. This is because sequence reads originating from the TR region may align to different genomic locations within the TR region, or to one of the two alleles.

[0113] Figure 2A shows two alleles of a repetitive region, including the repetitive sequence indicated by a hatch pattern and two flanking regions. Allele 1 is shown above, and allele 2 is shown below. Allele 1 has a shorter TR sequence than allele 2. A pair of sequence reads (20) can be uniquely aligned to one position on each of the two alleles. Figure 2B shows the two alleles and a pair of reads (22) derived from the TR sequence. Both reads of this pair can be aligned to different locations on the repetitive sequence. Even with constraints on the relative positions of the two reads, they can still be aligned to multiple locations on the repetitive sequence. They can also be aligned to either of the alleles. Given the ambiguity of the alignment positions of read pairs, it is difficult or impossible to determine the location of the genomic region from which the read pair actually originates. This also makes it difficult to visualize the alignment of reads to alleles.

[0114] As explained above, due to the technical challenges in genotyping TRs (especially STRs) using short reads, it is desirable to develop computer-implemented tools for visualizing sequence read data and the genotypes determined from that data. Such tools can help validate genotype calls and understand clinically and biologically significant genetic features associated with STRs. For example, such visualization tools could enable the detection of changes in repetitive motifs (e.g., interruptions) that may have clinically significant effects.

[0115] Visualization of array read pileup for STR Because it is difficult to align sequence reads to repeat regions, it is important to develop computer-implemented tools to visualize sequence reads aligned to tandem repeat regions, inspect the quality of the alignment, and verify the genotype of the repeat regions. However, conventional visualization tools align sequence reads to a standard reference sequence. Figure 3A schematically shows a conventional graphical representation of sequence reads aligned to a reference sequence containing an STR sequence. A graphical representation of reference and sequence reads aligned to a reference sequence is called a sequence read "pile-up".

[0116] Conventional visualization tools, such as those shown in Figure 3A, use standard reference sequences that are not customized for individual samples or subjects. This approach has various limitations in visualizing tandem repeat regions with repeat extensions. Furthermore, it cannot effectively reflect the actual length and detail of tandem repeat sequences in individual samples. Sequence reads containing repeat motifs not present in the reference sequence may be truncated. See, for example, sequence read 32. If the repeat sequences in individual samples are shorter than those in the reference sequence (not shown in this embodiment), sequence reads derived from TR sequences may produce uneven coverage.

[0117] Some embodiments of this disclosure provide a computer implementation tool that generates computer graphics for visualizing tandem repeat regions. The tool generates sequence read pileups, each pileup containing multiple haplotypes specific to the sample. In the example shown in Figure 3B, the sample has two different haplotypes. A first haplotype 34 is shown above, which has a shorter tandem repeat region than the second haplotype 36 shown below. Sequence reads are aligned to each of the two haplotypes. When sequence reads can be aligned to multiple locations on a haplotype, often within a tandem repeat region indicated by a hatch pattern, the sequence reads are uniformly distributed on the haplotype, giving uniform coverage across the haplotype.

[0118] In some embodiments, the haplotype may comprise a single repeat sequence, as shown herein. In other embodiments, the haplotype may comprise multiple repeat sequences. These can be used to visualize short indels even when genotyping tools for determining the genotype of a repeat do not effectively detect this variant type. While the various embodiments described herein visualize TR regions, they can also be used to visualize other types of variants having different genotypes on different haplotypes.

[0119] In various embodiments, each sequence pile-up includes a customized and individualized haplotype for the sample. This allows for better visualization of the length of repeat regions and sequences. These plots can be used to detect interruptions within repeat sequences and within sequences directly surrounding repeats. It also allows for examination of alignment characteristics to the haplotype, providing a means to verify the genotype of repeat sequences within genomic regions. As shown in the experimental data below, when the provided haplotype is correct, sequence reads tend to be uniformly distributed on the haplotype, and different genomic locations tend to have similar coverage.

[0120] Furthermore, some embodiments allow for the visualization of motifs and nucleotides that have biological or clinical significance. In some embodiments, a haplotype may include multiple TR sequences. In such applications, sequence data may need to be phased and combined into haplotypes of two or more chromosomes. The genotype of a TR sequence can be determined using various techniques, such as the sequence graphing technique and paired-end read anchoring technique described below. In some embodiments, sequence read data from the entire genome may be preprocessed using the techniques described herein to provide a subset of sequence reads.

[0121] Figure 4 shows schematic workflows for several embodiments using sequence graph alignment techniques to obtain sequence reads and haplotypes used to visualize repetitive regions. Panel 1 of Figure 4 shows that sequence reads are obtained from a target region containing repetitive sequences. These reads are paired-end reads. They can be obtained, for example, by aligning genome-wide reads to the genome using conventional alignment methods and selecting reads aligned to or near the target region.

[0122] Panel 2 of Figure 4 shows that after sequence reads are obtained for a target region, the sequence reads are aligned to a sequence graph representing the target region. The repeat region represented by this sequence graph includes, from left to right, a left franking region, a CAG tandem repeat sequence, a CAACAG intercalating sequence, a CCG tandem repeat sequence, and a right flank region.

[0123] Read alignment into the sequence graph provides realigned sequence reads as shown in Panel 3. Further details regarding aligning sequence reads into the sequence graph to obtain realigned reads are described below with reference to Figures 8 to 13.

[0124] Read alignment into a sequence graph also determines the genotype of the STR sequence within the repeat region. Alignment of sequence reads into a sequence graph determines that one allele of the CAG STR contains 4 repeat units and the other allele contains 78 repeat units. Sequence graph alignment also determines that one allele of the CCG STR contains 7 repeat units and the other allele contains 10 repeat units. Considering the determined genotypes, two pairs of possible haplotype sequences are shown in panel 5 of Figure 4. Some embodiments include phasing the genotypes to determine the haplotype pair that best matches the realigned reads. As shown in panel 6, the best haplotype pair has a first haplotype containing 4 CAG repeat units and 7 CCG repeats, and a second haplotype containing 78 CAG repeats and 10 CCG repeats.

[0125] Next, the illustrated embodiment visualizes the pile-up of repeating regions using the realigned reads and the best haplotype pairs. In some embodiments, different sequence read pairs may have different possible alignment scenarios. For example, sequence read pair "a" is aligned to one location on haplotype 1 and one identical location on haplotype 2. Sequence read pair "b" is aligned to multiple locations on haplotype 2. Sequence read pair "c" is aligned to a single location on haplotype 1. For example, sequence read pair "d" is aligned to one location on haplotype 1 and a corresponding location on haplotype 2.

[0126] In some embodiments, all possible alignment positions are determined for each pair of reads. Both reads in a pair are aligned on the same haplotype. Then, random positions are selected for all read pairs to determine a set of alignment positions. The same random selection is repeated to obtain multiple sets of alignment positions. In various embodiments, at least 1,000, 5,000, 10,000, 50,000, or 100,000 sets of alignment positions are obtained. To generate a pileup including two haplotypes and pairs of sequence reads aligned to the two haplotypes, as shown in Panel 8, a set of alignment positions having the most uniform distribution on the two haplotypes is selected.

[0127] Figure 5 shows a flowchart of process 50 for generating a computer graphic representing sequence reads aligned to haplotypes in a genomic region. The graphic includes a sequence read pileup as described above. Process 50 includes determining a set of multiple sequence positions for multiple sequence reads aligned to multiple haplotype sequences corresponding to multiple haplotypes in a genomic region. See box 52. The multiple sequence reads are obtained from a genomic region of a nucleic acid sample. Process 50 further includes selecting a set of alignment positions that are more uniformly distributed over the multiple haplotypes than other sets of alignment positions within the set of multiple alignment positions. See block 54. Process 50 further includes generating a computer graphic representing the multiple sequence reads and multiple haplotypes. The multiple sequence reads are located in the selected set of alignment positions. In some embodiments, process 50 may include features described in process 600 shown in Figure 6.

[0128] Figure 6 shows a flowchart of process 600 for generating a computer graphic representing a pileup of sequence reads containing multiple haplotypes. Process 600 includes aligning multiple sequence reads to a set of alignment positions on multiple haplotype sequences corresponding to multiple haplotypes in a genomic region. See box 602. In some embodiments, the multiple sequence reads include at least 100, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 sequence reads.

[0129] In some embodiments, at least one of the multiple haplotypes includes a repeat extension. In some embodiments, the multiple haplotypes include two haplotypes in a genomic region on a chromosome pair. In various embodiments, the multiple haplotypes include at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 haplotypes. In various embodiments, the genomic region includes at least 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 bp.

[0130] In some embodiments, at least one of the multiple haplotypes includes a structural variant. In some embodiments, the structural variant is longer than 50 bp. The structural variant may be a deletion, duplication, copy number variant, insertion, inversion, transposition, etc.

[0131] In some embodiments, the structural variant is shorter than 50 bp. In some embodiments, the structural variant shorter than 50 bp includes a single nucleotide polymorphism (SNP).

[0132] Figure 7 shows a flowchart of process 700 for aligning sequence reads to a set of alignment positions. In some embodiments, operation 602 of process 600 may be implemented according to process 700. Process 700 includes determining possible alignment positions for each read for each haplotype, where a plurality of sequence reads include read pairs obtained by pair-end sequencing.

[0133] Process 700 further includes creating a constrained alignment position for each lead pair from the alignment positions of the constituent lead such that (A) both lead pairs of the lead pair align with the same haplotype, and (B) the corresponding fragment length of the lead pair is as close as possible to the average fragment length.

[0134] Process 700 also includes randomly selecting the alignment position of each read pair from the constrained alignment positions. In some embodiments, non-restored random sampling techniques are used to select reads from the constrained alignment positions. These techniques can cover the entire position space more quickly. After all positions have been sampled, all samples can be restored. In some embodiments, restored random sampling techniques are used, which do not require restoration at the end and may, in some cases, obtain the desired combination of positions faster than non-restored methods. This latter approach may save time if a predetermined convergence criterion (e.g., a desired alignment score) is used instead of a fixed number of iterations to stop the search for alignment positions.

[0135] Returning to Figure 6, in some embodiments, process 600 includes aligning different sets of sequence reads to different genomic regions. In some various embodiments, the different genomic regions include at least 100, 200, 300, 500, 600, 700, 800, 900, 1,000, 5,000, or 10,000 regions.

[0136] In some embodiments, multiple haplotypes can be obtained using the sequence graph alignment techniques described herein. In other embodiments, multiple sequence reads and / or multiple haplotypes can be obtained using the paired-end read anchoring techniques described below.

[0137] In some embodiments, process 600 includes aligning a first number of sequence reads to one or more sequence graphs corresponding to genomic regions to obtain a plurality of sequence reads and / or a plurality of haplotypes. In some embodiments, aligning a first number of sequence reads to a sequence graph includes providing a first number of sequence reads of a nucleic acid sample and aligning the first number of sequence reads to one or more repetitive sequences each represented by the sequence graph. The sequence graph has a directed graph data structure having vertices representing nucleic acid sequences and directed edges connecting the vertices. The sequence graph includes one or more self-loops, each self-loop representing a repetitive subsequence, each repetitive subsequence containing repeats of one or more nucleotide repeating units. Aligning a first number of sequence reads to a sequence graph also includes determining one or more genotypes of one or more repetitive sequences and providing a first number of sequence reads as a plurality of sequence reads in (a) and / or providing one or more genotypes of one or more repetitive sequences.

[0138] In some embodiments, process 600 further includes phasing one or more genotypes to determine multiple haplotypes. In some embodiments, the process further includes first aligning a second number of sequence reads to the genome to provide a first number of sequence reads. The second number of sequence reads may be genome-wide reads and may include at least 10,000, 100,000, or 1,000,000 sequence reads.

[0139] Process 600 further includes estimating an alignment score for a set of alignment positions. See block 604. Process 600 then loops back to operation 602 to repeat the process for several different sets of alignment positions. In some embodiments, the process may loop for a definite number of iterations. In various embodiments, the process obtains at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 20,000, 50,000, 100,000, or 500,000 sets of different alignment positions. In other embodiments, the process repeats the iterations until the alignment score meets a criterion. In other embodiments, other alignment metrics for alignment positions may be used to set a criterion for terminating the loop. For example, an alignment quality score, a mapping quality score, or coverage may be used to set a criterion for terminating the loop.

[0140] In some embodiments, the alignment score indicates how uniformly multiple sequence reads are distributed across multiple haplotype sequences corresponding to multiple haplotypes. A more uniform distribution of reads results in a more consistent coverage level across haplotypes. Conceptually, assuming that DNA fragments of the same length are used to generate reads, and that these DNA fragments are uniformly distributed across the genome, the most uniform distribution of reads would be such that any two consecutive non-overlapping reads have exactly the same distance. The less uniformly distributed the reads are, the more each individual consecutive read deviates from the mean of that distance. Therefore, in some embodiments, the alignment score includes the root mean square difference from the mean of the distance between the starting positions of two consecutive reads. A smaller alignment score indicates a more uniform distribution of sequence reads on a haplotype, and a better alignment score.

[0141] In some embodiments, the alignment score is estimated using a probabilistic model, assuming that read pairs are uniformly distributed across a plurality of haplotype sequences. In some embodiments, the alignment score is the probability of a plurality of sequence reads derived from a set of alignment positions taking into account the probabilistic model. In some embodiments, the plurality of sequence reads comprise paired-end reads obtained from nucleic acid fragments, and the probabilistic model is configured to receive an average fragment length as an input. In some embodiments, the probabilistic model is configured to receive a haplotype length as an input.

[0142] In some embodiments, p (k) represents the probability x of the individual alignment position of the k-th read pair from the start of the haplotype, which is modeled by the following formula:

[0143] [Formula]

[0144] wherein i is the haplotype to which the read pair is aligned, H i is the length of haplotype i, L is the average fragment length, n i is the number of read pairs aligned to haplotype i. The alignment score for the set of alignment positions is estimated as a product of the probabilities of individual alignment positions.

[0145] Process 600 includes selecting a set of alignment positions from a set of several different sets of alignment positions based on a set of alignment scores. In some embodiments, the selected set of alignment positions has the best lime score among the set of several different sets of alignment positions. In some embodiments, the selected set of alignment positions has an alignment score that exceeds a selection criterion. In some embodiments, the selection criterion may be the top 1st, 2nd, 3rd, 4th, 5th, 10th, or 20th percentile of the alignment score. This may allow a combination of the alignment score and one or more other metrics (e.g., coverage, mapping quality, alignment quality) to be considered when selecting the final set of alignment positions.

[0146] In some embodiments, process 600 optionally includes generating computer graphics representing multiple sequence reads and multiple haplotypes, where the multiple sequence reads are located in a selected set of alignment positions. See block 608.

[0147] In some embodiments, process 600 does not require operation 608. Instead, sequence reads can be assigned to locations in genomic regions, and these assigned locations can be used for other downstream processing, whether or not computer graphics are generated.

[0148] Some embodiments involve estimating one or more sequencing metrics for multiple sequence reads aligned to multiple haplotype sequences in a selected set of alignment locations. In some embodiments, one or more sequencing metrics include sequence coverage. In some embodiments, one or more sequencing metrics include sequence coverage for each alignment location. In some embodiments, one or more sequencing metrics include alignment quality scores, which indicate the quality of the match between the read sequence and the reference sequence. In some embodiments, one or more sequencing metrics include alignment quality scores for each alignment location. In some embodiments, one or more sequencing metrics include mapping quality scores, which indicate the reliability that the read is correctly mapped to genomic coordinates. For example, a read may be mapped to several genomic locations to match almost perfectly at all locations. In that case, the alignment score would be high, but the mapping quality would be low.

[0149] Sequencing quality metrics can provide crucial information about the accuracy of each step in the process, including library preparation, base calling, read alignment, and variant calling. Base calling accuracy, measured by the Phred quality score (Q score), is a common metric used to assess the accuracy of a sequencing platform. It indicates the probability that a given base is incorrectly called by the sequencer. Figure 24 shows read mapping quality scores for several embodiments of a genomic region containing C9ORF72 repeats. The upper panel shows haplotypes with short repeats, and the lower panel shows haplotypes with long repeats. The horizontal axis shows the bins on the haplotypes. The vertical bars, similar to histograms, show the read coverage in the bins. Q scores have been determined for reads assigned to the bins of haplotypes in several embodiments. Reads with a Q score greater than 11 are reflected at the bottom of each bar, and reads with a Q score of 11 or less are reflected at the top of each bar. 98% of reads aligned to short haplotypes in the upper panel have a Q score greater than 11. 97% of reads aligned to long haplotypes in the lower panel have a Q score greater than 11. Coverage for each bin is determined according to several embodiments. The variance of coverage can be determined, which provides a measure of the uniformity of the read distribution. The mean coverage for long repeat haplotypes is 26, and for short repeats it is 18. Overall, the reads are relatively uniformly distributed within and across haplotypes. Using these sequence metrics and derivation measures, the quality of read alignment can be examined and the effectiveness of allelic genotypes in the sequence can be inferred, such as those in Examples 1–5 described below.

[0150] Genotyping of variants at repetitive loci using sequence graphs Figure 8 shows a flowchart illustrating a process 140 for genotyping a genomic locus containing repetitive sequences, according to several embodiments. Some embodiments provide a method for targeted type analysis of a region containing one or more adjacent TRs, which can estimate the size of both repetitions shorter and longer than the read length. In some embodiments, the locus is predefined in a variant catalog that includes the genomic location and the structure of the locus at the genomic location. Figures 9, 10, and 11 show three different sequence graphs according to several embodiments.

[0151] Figure 12 shows a schematic diagram of the process for determining the genotype of variants at the HTT locus containing two STR sequences, according to several embodiments. Panel (a) of Figure 12 shows a portion of the variant catalog, including genomic loci and their structures as locus specifications. For example, ignoring repeats, the sequence at the HTT locus is CAGCAACAGCGG (SEQ ID NO: 2), and the sequence at the CNBP locus is CAGGCAGACA (SEQ ID NO: 3).

[0152] Figure 13 shows a schematic diagram of the process for determining the genotype of variants at the Lynch I locus, including SNVs and STRs, according to several embodiments. Box 162 in Figure 13 shows the general structure of the locus specification, and box 163 shows a specific example of the Lynch I(MSH2) locus specification.

[0153] In the variant catalog, locus structures are specified using a restricted subset of regular expression syntax. For example, a repeating region bound to HD indicates that it has a variable number of CAG and CCG repeats separated by a CAACAG interruption in expression (CAG). * CAACAG(CGG) * Alternatively, it can be defined by Sequence ID No. 2 (Ignore Repetitions). The region bound to the FRDA region is expressed (A) * (GAA) * The region that corresponds to and is joined to SCA8 is (CTA) * (CTG)* Corresponding to this, the DM2 iteration region consisting of three adjacent iterations is (CAGG) * (CAGA) * (CA) * Alternatively, as defined by Sequence ID No. 3 (repeated neglect), MSH2 SNVs adjacent to the A homopolymer causing Lynch syndrome I are (A|T)(A) * It corresponds to.

[0154] Furthermore, regular expressions can include multiple alleles or "degenerate" base abbreviations that can be identified using the International Union of Pure and Applied Chemistry (IUPAC) notation ("Nomenclature for Incompletely Specified Bases in Nucleic Acid Sequences. Recommendations 1984. Nomenclature Committee of the International Union of Biochemistry (NC-IUB)" 1986).

[0155] Incompletely identified bases corresponding to bases in degenerate codons are referred to herein as degenerate bases. Degenerate bases can represent certain types of incomplete DNA repeats, for example, different bases may occur at the same position. This notation is used for expression (GCN) * Polyalanine repeats can be encoded by polyglutamine repeats, and polyglutamine repeats can be expressed by (CAR) * It can be coded by [this method].

[0156] In some embodiments, the repeat sequences contained within the genomic locus include short tandem repeat (STR) sequences. In some embodiments, FTR elongation is associated with fragile X syndrome, amyotrophic lateral sclerosis (ALS), Huntington's disease, Friedreich's ataxia, spinocerebellar degeneration, bulbar spinal muscular atrophy, myotonic dystrophy, Machado-Joseph disease, or dentatorubral-pallidoluysian atrophy.

[0157] Process 140 includes collecting nucleic acid sequence reads of a test sample from a database. See block 142. In some embodiments, nucleic acid sequence reads are first aligned to a reference genome, but the processes herein realign the sequence reads to the target genomic locus, as described below. In alternative embodiments, the reads may be aligned directly to the sequence graph without first being aligned to a reference genome.

[0158] Process 140 includes aligning sequence reads to sequences for genomic loci containing one or more repeat sequences. See block 144. Sequences for genomic loci are represented by data stored in system memory, which has a sequence graph data structure. The sequence graph includes a directed graph having vertices representing nucleic acid sequences and directed edges connecting the vertices. Nucleic acid sequences within a vertex contain one or more nucleic acid bases. The sequence graph includes one or more self-loops. Each self-loop represents a repeat sequence of one or more repeat sequences. Each repeat sequence contains repeats of one or more nucleotide repeat units.

[0159] In some embodiments, sequence reads are first aligned to a reference genome to determine the genomic coordinates of the reads before a subset of the initially aligned reads are aligned to one or more sequence graphs representing one or more target sequences. In some embodiments, the initially aligned reads determine repeat extensions in tens to thousands of regions (each region corresponding to a sequence graph). The total number of initially aligned reads that are realigned to a sequence graph during each implementation of the embodiment may range from thousands to millions of reads.

[0160] In some embodiments, reads that are initially aligned to or near the target sequence or locus are selected as a subset of reads, which are then aligned to repetitive sequences represented by a sequence graph, each having one or more self-loops representing one or more repetitive sequences. In various embodiments, reads within approximately 10, 50, 100, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 base pairs from the target sequence or locus are considered to be near the target sequence or locus. In some embodiments, reads within approximately 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, or 10,000 bases of the target locus are located near the target locus. Some of the raw reads may have poor initial alignment, for example, because they contain repetitive sequences that are difficult to align without ambiguity. In some embodiments, reads that are poorly initially aligned (as measured, for example, by alignment scores) but are paired with reads aligned to or near the target locus (paired-end read pairs) are aligned in the sequence graph. In some embodiments, reads that are initially aligned to off-target regions, which are known hotspots for read mismatch, are aligned in the sequence graph.

[0161] Figures 9, 10, and 11 show three different sequence graphs according to several embodiments. Figure 9 shows a first sequence graph 1100 representing a first genomic locus containing a repeat sequence having a trinucleotide repeat unit CAG. The first sequence graph 1100 includes vertices 1102 and 1112, each representing two flanking sequences. The first sequence graph also includes vertex 1106, representing a repeat sequence containing a trinucleotide repeat unit CAG. The first sequence graph includes a directed edge 1104 connecting vertex 1102 (flanking sequence) and vertex 1106 (CAG repeat sequence), with the direction progressing from vertex 1102 to vertex 1106. The direction of the edge indicates the relative position of the two nucleic acid sequences. The first sequence graph also includes a directed edge 1104 connecting vertex 1102 (flanking sequence) and vertex 1106 (CAG repeat sequence), with the direction progressing from vertex 1102 to vertex 1106. The first sequence graph also includes a directed edge 1110 connecting vertex 1106 (CAG repeat sequence) and vertex 1112 (flanking sequence), with the direction progressing from vertex 1106 to vertex 1112. The first sequence graph also includes a self-loop 1108 indicating that the repeat sequence contains one or more repeating units CAG (shown at vertex 1106). The path from the start vertex to the end vertex of the sequence graph represents the sequence of a genomic locus, which may contain nucleotides near repeat sequences such as flanking sequences.

[0162] Figure 10 shows a second sequence graph 1200 representing a second genomic locus. The second sequence graph 1200 includes vertices 1202 and 1224, each representing two flanking sequences. The second sequence graph also includes vertices 1206 and 1216, each representing a repeat sequence containing a trinucleotide repeat unit CAG and a repeat sequence containing a trinucleotide repeat unit CCG, respectively. The second sequence graph further includes vertex 1212, representing a non-repeat sequence CAACAG. The second sequence graph includes directed edges 1204, 1210, 1214, and 1220. These directed edges connect vertices 1202, 1206, 1212, 1216, and 1224 in a directed manner, as shown in the figure. The second sequence graph also includes a self-loop 1208, indicating that the repeat sequence contains one or more repeating repeat units CAG (shown at vertex 1206). The second array graph also includes a self-loop 1218, which indicates that the iterative array contains a repeating unit CCG (shown at vertex 1216) that repeats one or more times.

[0163] Figure 11 shows a third sequence graph 1300 representing a third genomic locus. The third sequence graph 1300 is similar to the second sequence graph 1200, but includes two alternative pathways representing two alleles, CAC and CAT. The two alleles may be alleles of an SNV or an SNP. Directed edges 1310, vertex 1312, and directed edge 1314 represent the first allele of CAC. Directed edges 1316, vertex 1318, and directed edge 1320 represent the second allele of CAT. The third sequence graph includes elements similar to those in the second sequence graph, including vertices 1302, 1306, 1322, and 1328. It also includes self-loops 1308 and 1324 representing CAG and CCG repeats of repetitive sequences. It further includes directed edges 1304 and 1326.

[0164] In some embodiments, sequence reads are aligned to a sequence graph using the techniques described below. 1. The Kmer index is constructed over the entire graph, taking into account the kmer from the array, so that all graph nodes on which such kmer begin or end can be enumerated. In some cases, a kmer may begin on one node and end on another. 2. For each graph hit, extract two subgraphs, one in the forward direction of the kmer and the other in the reverse direction. The subgraphs expand repeat extensions to the remaining read length, but do not include any nodes that are further away from the kmer hit than the remaining read length, assuming no repeat extensions. The procedure is a breadth-first search and generates a data structure containing the following: - Concatenation of all node arrays (including expanded iterations) within the subgraph. - Node indexes that make it easier to obtain node IDs from offsets in the array when backtracking on the Smith-Waterman process - An offset array of node ends containing edges, with respect to the start offset of each node. - An index for each node that makes it easy to indicate whether a base is present or absent at the beginning of the node, and that enumerates all terminal offsets of the preceding node. 3. Align - Supports affine gaps. - Given the above information and penalty matrix, find the best score sort(s) for the sequence.

[0165] Two differential interfaces are available. - The best alignment score and the second best alignment score are reported. - The entire array of the best alignment score and the second best alignment score.

[0166] Alignment is a global alignment that penalizes any gap between a candidate kmer and the start of the aligned array. Some embodiments fine-tune the compile-time parameters.

[0167] The current algorithm for matrix filling is available in two embodiments. -N * A continuous loop with complexity M. -By defaulting to 16 for the fixed-length compilation time parameter, which is a series of fixed-size loops, gcc automatically recognizes and converts SSE or AVX vector instructions on the CPU.

[0168] In some embodiments, a particular repeating unit of one or more repeating sequences contains at least one incompletely identified nucleotide. In some embodiments, a particular repeating unit contains a degenerate codon.

[0169] In some embodiments, one or more self-loops include two or more self-loops that represent two or more iterative arrays. See, for example, panel (b) of Figures 10, 11, and 12.

[0170] In some embodiments, the sequence graph further includes two or more alternative pathways for two or more alleles. See, for example, reference numerals 1312 and 1318 in Figure 11. Also, referring to Figure 13, reference numerals 165 and 167a for the locus Lynch I (MSH2) include the upper pathway with nucleotide base A as the vertex and the lower pathway with nucleotide base T as the vertex.

[0171] In some embodiments, two or more alleles include indels or substitutions. In some embodiments, substitutions include single nucleotide variants (SNVs) or single nucleotide polymorphisms (SNPs). See, for example, reference numerals 1312 and 1318 in Figure 11.

[0172] In some embodiments, aligning sequence reads to a sequence graph includes finding KMER alignment between the sequence reads and the path in the sequence graph, and then extending this path to a complete alignment. In some embodiments, the alignment includes extracting subgraphs around the path, non-rolling any loops in the subgraphs to obtain a directed aringal graph, and performing Smith-Waterman alignment of the sequence reads to the directed aringal graph.

[0173] In some embodiments, aligning sequence reads to a sequence graph involves reducing the graph by removing ends with low confidence in the alignment. After the reads are aligned to the graph, the method searches for other similar alternative alignments. This is done by rearranging the original reads to paths through the graph that overlap with the paths of the original alignment. This makes it possible to detect, for example, cases where one or both ends of the initial alignment have low confidence, indicating that they may have been aligned in a different way. The ability to detect high-confidence and low-confidence regions of the alignment makes it possible to accurately determine which gene variants the reads support.

[0174] In some embodiments, aligning sequence reads to a sequence graph includes aligning subsequences of reads to the sequence graph, and aligning and merging by merging the sequences of the subsequences to form a complete alignment of sequence reads.

[0175] In some embodiments, the process also includes generating a sequence graph based on a locus specification, which includes the locus structure of the genomic locus. In some embodiments, the locus specification is defined in a variant catalog as described above.

[0176] See also panels (b)–(d) of Figure 12 for a schematic diagram of read alignment into the sequence graph of the HTT locus. Figure 13 schematically shows the locus analyzer 164 for performing read alignment into the sequence graph, such as the Lynch I (165) locus.

[0177] Process 140 further includes determining one or more genotypes of one or more repeat sequences using sequence reads aligned in a sequence graph. See block 140. See also panel(e) of Figure 12, which shows the determination of two STRs (CAG and CCG) at the HTT locus. The sequence on the left containing the CAG repeat is CAGCAGCAGCAGCAG (Sequence ID 4). The sequence on the left containing the CCG repeat is CCGCCGCCGCCGCCG (Sequence ID 5).

[0178] Figure 13 shows a variant genotype determination module (168) for determining variants at the Lynch I locus, including SNVs having A / T alleles (169a) and A monomer repeats (169b). Figure 13 also shows a variant analyzer module (166) for curating sequence-aligned data and providing them to the variant genotype determination module (168), as well as an embodiment of the variant analyzer for SNVs having A / T alleles (167a) and A monomer repeats (167b). The locus results obtained from the genotype determination module are shown in box 170 of Figure 13, specifically as the genotype of an SNV having A / T alleles (171a) and A monomer repeats (171b).

[0179] In some embodiments, the sequence graph includes two alternative pathways for two alleles, and the method further includes genotyping two or more alleles using sequence reads aligned to two or more alternative pathways. In some embodiments, genotyping two or more alleles includes providing the coverage of two or more alternative pathways to a probabilistic model in order to determine the probabilities of two or more alleles. In some embodiments, the probabilistic model simulates the probabilities of alleles as a function of allele coverage, the function being selected from a Poisson distribution, a negative binomial distribution, a binomial distribution, or a beta-binomial distribution.

[0180] In some embodiments, the probability function is a Poisson distribution, and its rate parameter is estimated from the read length and mean depth observed at the genomic locus.

[0181] In a Poisson system model, the probabilities of alleles are expressed as follows: P(Y=y)=(C y ×e -C ) / y! •y is the read coverage of the base. • C is the average depth at the genomic locus.

[0182] In some embodiments, the average depth C is It is estimated that C = LN / G. G represents the length of the genomic locus. • L is the lead length. • N is the total number of reads.

[0183] Graph Tool Library In some embodiments, the basic functionality of array graphs is achieved by applying the GraphTools library. The library implements core graph extraction (the graph itself, graph paths, and graph alignment), their behavior, and algorithms for aligning linear arrays into graphs.

[0184] In some embodiments, the sequence graph consists of nodes and directed edges. The graph may contain self-loops (edges connecting nodes to themselves) but not other cycles. The nodes contain sequences consisting of core bases and IUPAC degenerate base codes.

[0185] A graph path is defined by an array of nodes that the path passes through, along with its starting position on the first node and its ending position on the last node. Positions are determined using a zero-based, semi-open coordinate system. The library defines several operations on paths, including path stretching and shrinking, duplicate checking, and path merging.

[0186] Graph alignment codes how linear query arrays (typically array-determined reads) are aligned in a graph. In some embodiments, graph alignment includes a graph path and an array of linear alignments that define the alignment of query arrays to the nodes of the graph path. Using corresponding actions on the path, graph alignments can be reduced or merged with other graph alignments. Path reduction provides a mechanism for removing unreliable alignment ends, while alignment merging is used by graph alignment algorithms to stitch together a complete alignment of query arrays from a sub-array (e.g., kmer) alignment. In some embodiments, the alignment algorithm works by finding kmer alignments between the query arrays and the graph, and then extending this alignment into a complete alignment. In some embodiments, the alignment includes extracting subgraphs around the path corresponding to kmer alignments (non-rolling any loops in the process). It then performs Smith-Waterman alignment on the resulting directed aring graph. In some embodiments, the algorithm is written using constant-length loops to support affine gap penalties and to allow the compiler to generate SIMD code.

[0187] In some embodiments, a graph path can be obtained using a search algorithm, which involves extending or shrinking the path by increasing or decreasing the number of iterations of a self-loop-based iteration unit until the sorting reaches a search criterion or convergence (e.g., the sorting score is maximized).

[0188] In some embodiments, multiple graph paths are generated from an array graph, and each graph path represents a specific number of iterations of a repeating unit represented by a self-loop. A query array is sorted into multiple graph paths, and then paths that satisfy the sorting criteria are selected for graph sorting.

[0189] Application Architecture Several embodiments are designed as general tools for targeted variant genotyping (Figure 13). During each run, the program uses a set of variants listed in a variant catalog file. An attempt is made to determine the genotype. Variants located in close proximity to each other are grouped into the same locus. The locus structure is specified using a restricted subset of regular expression (RE) syntax. Res contains a sequence above an alphabet consisting of core base abbreviations and IUPAC degenerate base codes, separated, in some cases by sequence interruptions, one or more expressions (<sequence>), (<sequence a>|<sequence b>), (<sequence>). * The locus must contain (<sequence>)+. These expressions correspond to insertions / deletions, substitutions, sequences that repeat zero or more times, and sequences that repeat at least once. Furthermore, each locus description includes a set of reference regions for the locus and reference coordinates of each constituent variant.

[0190] The work bulk is orchestrated by LocusAnalyzer class objects, which synthesize sequence graphs representing loci from corresponding REs during initialization. After initialization, the locus analyzers process the relevant reads by aligning them into a graph and then passing the resulting alignment through a Variant Analyze, which is defined for each variant contained in the locus. The Variant Analyze extracts the relevant information for genotyping the relevant variants and passes it through a genotype determiner that performs the actual genotyping. The output VCF file is then created using the results output by each genotype determiner.

[0191] For example, the LocusAnalyzer responsible for processing loci containing pathogenic variants associated with Lynch I syndrome utilizes SNV and STR analyzers (right panel of Figure S1).

[0192] Indel genotype determiner Some STRs may have small insertions or deletions (indels) nearby. Such indels are modeled as additional subgraphs in the flanking sequence of the STR. The number of reads located at each allele (or graph pathway) is modeled using a Poisson distribution, and its rate parameter is estimated from the average depth and read length observed at the locus. Genotype likelihoods are calculated under a Bayesian framework.

[0193] Identifying the growth of repeat purchases The embodiments disclosed herein allow for the determination of various genetic conditions associated with repeat elongations with higher efficiency, sensitivity, and / or selectivity compared to conventional methods. Some embodiments of the present invention provide methods for identifying and calling medically relevant repeat elongations, such as the CGG repeat elongation causing intellectual disability in fragile X syndrome, using sequence reads that do not completely traverse the repetitive sequence. Short reads, such as 100 bp reads, are not long enough to sequence through many repeat elongations. However, when analyzed by the disclosed methods, samples with repeat elongations exhibit a statistically significant excess of reads containing a large number of repetitive sequences. In addition, very large repeat elongations include unaligned read pairs where both reads are composed entirely or nearly entirely of the repetitive sequence. A standard sample is used to identify background expectations.

[0194] The conventional wisdom is that repeat elongation cannot be detected without reads that extend across the entire repeat. The previous approach to detecting repeat elongation used targeted sequencing with longer reads if the reads were abnormal because they were not long enough to extend across the repeat sequence. The results of some disclosed embodiments have been surprisingly satisfied by using normal (non-targeted) sequence data and read lengths of only about 100 bp, yet yielding very high sensitivity with respect to detecting repeat elongation. The method described herein can detect the number of repeating units during repeat elongation using paired reads with insert lengths shorter than the entire repeat sequence (i.e., two sequence reads and an intervening sequence).

[0195] Referring to details of methods for determining the presence of repeat extension according to several embodiments, Figure 14 shows a flowchart providing a high-level depiction of an embodiment for determining the presence or absence of repeat extension in a repetitive sequence in a sample. A repetitive sequence is a nucleic acid sequence that contains a repeating appearance of short sequences called repeat units. Table 1 above provides examples of repeat units, the number of repeats of repeat units in repetitive sequences of normal and pathogenic sequences, genes associated with the repetitive sequence, and examples of diseases associated with repeat extension. Process 200 in Figure 14 begins by obtaining paired-end reads from a test sample. See Block 202. The paired-end reads are processed to align with a reference sequence containing the repetitive sequence of interest. In some syntax, the alignment process is also called the mapping process. The test sample contains nucleic acids and may also be in the form of bodily fluids, tissues, etc., as further described in the Samples section below. Sequence reads undergo an alignment process to be positioned on a reference sequence. Various alignment tools and algorithms may be used to attempt to align the reads to the reference sequence, as described elsewhere in this disclosure. Typically, in a sorting algorithm, some reads are successfully sorted to the reference sequence, while others may not be successfully sorted to the reference sequence, or may not be sorted at all. Reads that are continuously sorted to the reference sequence are associated with a site on the reference sequence. Sorted reads and their associated sites are also called sequence tags. As mentioned above, some sequence reads containing a large number of repeats tend to be more difficult to sort to the reference sequence. Read sorting is considered poor if a read is sorted to a reference sequence that has more mismatched bases than a certain criterion. In various embodiments, read sorting is considered poor if it is sorted to at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mismatches. In other embodiments, read sorting is considered poor if it is sorted to at least about 5% mismatches. In other embodiments, read sorting is considered poor if it is sorted to at least about 10%, 15, or 20% mismatched bases.

[0196] As shown in Figure 2, process 200 proceeds to identify the anchor read and anchor-type read within the paired-end reads. See block 204. The anchor read is the read between the paired-end reads that is aligned with or near the target repetitive sequence. For example, the anchor read can be aligned to a position on the reference sequence separated from the repetitive sequence by a sequence length shorter than the sequence length of the insert. The separation length may be even shorter. For example, the anchor read can be aligned to a position on the reference sequence separated from the repetitive sequence by a sequence length shorter than the sequence length of the anchor read, or by a sequence length less than the combined sequence length of the anchor read and the sequence connecting the anchor read to the anchor-type read (i.e., the length of the insert minus the length of the anchor-type read). In some embodiments, the target repetitive sequence may be a repetitive sequence in the FMR1 gene containing repeats of the repeating unit CGG. In a typical reference sequence, the repetitive sequence in the FMR1 gene contains approximately 6 to 32 repeats of the repeating unit CGG. When repeats are extended beyond 200 copies, the repeat extension becomes pathogenic and tends to cause fragile X chromosome syndrome. In some embodiments, reads are considered aligned near the target sequence if they are aligned within 1000 bp of the target repeat sequence. In other embodiments, this parameter may be adjusted to ranges such as approximately 100 bp, 200 bp, 300 bp, 400 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1500 bp, 2000 bp, 3000 bp, 5000 bp, etc. In addition, this process also identifies anchor-type reads that are paired with anchor reads but are poorly aligned with or unable to align with their reference sequences. Further details of poorly aligned reads are as described above.

[0197] Process 200 further includes determining whether repeat extensions of repetitive sequences may be present in the test sample, at least partially based on the identified anchor-type leads. See block 206. This determination step may include various preferred analyses and calculations, as further described below. In some embodiments, the process uses the identified anchor leads, as well as the anchor-type leads, to determine whether repeat extensions may be present. In some embodiments, the number of repeats in the identified anchor and anchor-type leads is analyzed and compared to one or more criteria derived from theoretically derived or empirical data from affected control samples.

[0198] In the various embodiments described herein, the repetitions are obtained as intraframe repetitions, where two repetitions of the same repeating unit enter the same reading frame. A reading frame is a method of dividing a nucleotide sequence in a nucleic acid (DNA or RNA) molecule into a set of consecutive non-overlapping triplets. During translation, the triplets code for amino acids and are called codons. Thus, any particular sequence has three possible reading frames. In some embodiments, the repetitions are counted according to three different reading frames, and the maximum of the three counts is determined to be the number of corresponding repetitions to be read.

[0199] An example of a process involving additional operations and analyses is shown in Figure 3. Figure 15 shows a flow chart illustrating process 300 for detecting repeat extension using paired-end reads with a large number of repeats. Process 300 includes additional upstream processes for processing the test sample. This process is initiated by sequencing the test sample containing nucleic acids to obtain paired-end reads. See Block 302. In some embodiments, the test sample may be obtained and prepared in various ways, as further described in the Samples section below. For example, the test sample may be a biological fluid, e.g., plasma, or any preferred sample described below. The sample may be obtained using non-invasive procedures such as simple blood draw. In some embodiments, the test sample contains a mixture of nucleic acid molecules, e.g., cfDNA molecules. In some embodiments, the test sample is a maternal plasma sample containing a mixture of fetal and maternal cfDNA molecules.

[0200] Before sequencing, nucleic acids are extracted from the sample. Suitable extraction processes and apparatus are described elsewhere in this specification. In some embodiments, the apparatus processes DNA from a total of multiple samples to provide a multiplexed library and sequence data. In some embodiments, the apparatus processes DNA from eight or more test samples in parallel. As described below, the sequencing system can process the extracted DNA to generate a library of encoded (e.g., barcoded) DNA fragments.

[0201] In some embodiments, the nucleic acids in the test sample may be further processed to prepare a sequencing library for multiplex or singleplex sequencing, as further described in the “Preparation of Sequencing Library” section below. After the sample has been processed and prepared, sequencing of the nucleic acids may be performed by various methods. In some embodiments, various next-generation sequencing platforms and protocols may be employed, as further described in the “Sequencing Methods” section below.

[0202] Regardless of the specific sequencing platform and protocol, in block 302, at least a portion of the nucleic acids contained in the sample are sequenced to generate tens of thousands, hundreds of thousands, or millions of sequence reads (e.g., 100 bp reads). In some embodiments, the reads include paired-end reads. In other embodiments, such as those described below with respect to Figure 5, in addition to paired-end reads, single-end long reads containing hundreds, thousands, or hundreds of thousands of bases may be used to determine repetitive sequences. In some embodiments, sequence reads include approximately 20 bp, 25 bp, 30 bp, 35 bp, 36 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 110 bp, 120 bp, 130, 140 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, or 500 bp. Technological advancements are expected to enable single-ended reads exceeding 500 bp, and, when paired-ended reads are generated, enable reads exceeding approximately 1000 bp.

[0203] Process 300 proceeds to align paired-end reads obtained from block 302 to a reference sequence containing a repeat sequence. See block 304. In some embodiments, repeat sequences are prone to elongation. In some embodiments, repeat elongation is known to be associated with genetic diseases. In other embodiments, repeat elongation of repeat sequences has not been studied to date to establish an association with genetic diseases. The methods disclosed herein enable the detection of repeat sequences and repeat elongation regardless of any associated pathology. In some embodiments, reads are aligned to a reference genome such as hg18. In other embodiments, reads are aligned to a reference genome, such as a chromosome or a portion of a chromosomal segment. Reads that are uniquely mapped to a reference genome are known as sequence tags. In one embodiment, at least about 3 × 10⁻¹⁶ 6 A qualified array of tags, at least approximately 5 x 10 6 1 qualified array tag, at least approximately 8 x 10 6 10 qualified array tags, at least approximately 10 x 10 6 Each qualified array tag, at least approximately 15 x 10 6 A qualified array of tags, at least approximately 20 x 10 6 Each qualified array tag, at least approximately 30 x 10 6 A qualified array of tags, at least approximately 40 x 10 6 10 qualified array tags, or at least approximately 50 x 10 6 Each qualified sequence tag is obtained from a read that uniquely maps to the reference genome.

[0204] In some embodiments, the process may filter sequence reads before alignment. In some embodiments, read filtering is a quality filtering process enabled by a software program implemented in the sequencer to filter out erroneous and low-quality reads. For example, Illumina's Sequencing Control Software (SCS) and Consensus Assessment of Sequence and Variation software programs remove erroneous and low-quality reads by filtering by converting the raw image data generated by the sequencing reaction into an additional format that provides intensity scores, base calls, quality score alignment, and biologically relevant information for downstream analysis.

[0205] In certain embodiments, reads generated by the sequencing device are provided in electronic form. Alignment is achieved using a computing device such as those described below. Each read is compared to a reference genome, which is often enormous (millions of base pairs) to identify the site in the reference genome that the read specifically corresponds to. In some embodiments, the alignment procedure can limit inconsistencies between the read and the reference genome. In some cases, one, two, three or more base pairs in a read allow for inconsistencies in the corresponding base pairs in the reference genome, and mapping is still performed. In some embodiments, a read is considered aligned if it is aligned to a reference sequence having one, two, three, or four or fewer base pairs. Correspondingly, an unaligned read is one that cannot be aligned or is poorly aligned. A poorly aligned read is one that has more inconsistencies than an aligned read. In some embodiments, a read is considered aligned if it is aligned to a reference sequence having 1%, 2%, 3%, 4%, 5%, or 10% or fewer base pairs.

[0206] After aligning paired-end reads to a reference sequence containing the target repetition sequence, process 300 identifies anchor reads and anchor-type reads among the paired-end reads. See block 306. As described above, an anchor read is an end read aligned to or near the repetition sequence. In some embodiments, the anchor read is a paired-end read aligned within 1kb of the repetition sequence. The anchor read is paired with an anchor-type read, but as described above, it cannot be aligned to the reference sequence or is poorly aligned to the reference sequence.

[0207] Process 300 analyzes the number of repetitions in the repeating units within the identified anchor read and / or anchor-type read to determine whether the repeating sequence is elongated. More specifically, Process 300 includes using the number of repetitions within the read to obtain a large number of high-count reads within the anchor read and / or anchor-type read. A high-count read is a read that has more repetitions than a threshold. In some embodiments, high-count reads are obtained only from anchor-type reads. In other embodiments, high-count reads are obtained from both anchor reads and anchor-type reads. In some embodiments, a read is considered a high-count read if the number of repetitions is close to the maximum number of repetitions that can be read. For example, if the read is 100 bp and the repeating unit under consideration is 3 bp, the maximum number of repetitions is 33. In other words, the maximum value is calculated from the length of the paired-end read and the length of the repeating unit. Specifically, the maximum number of repetitions may be obtained by dividing the read length by the length of the repeating unit and rounding down the fractional part of the number. In this embodiment, various embodiments may identify a 100 bp read having at least about 28, 29, 30, 31, 32, or 33 repeats as a high-count read. The number of repeats may be adjusted upward or downward with respect to high-count reads based on empirical factors and considerations. In various embodiments, the threshold for high-count reads is at least about 80%, 85%, 90%, or 95% of the maximum number of repeats.

[0208] Next, process 300 determines, based on the number of high-count reads, whether there is a high probability of repeat extension in the repetitive sequence. See block 310. In some embodiments, the analysis compares the acquired high-count reads to a calling criterion and determines that there is a high probability of repeat extension if the criterion is exceeded. In some embodiments, the calling criterion is obtained from the distribution of high-count reads in a control sample. For example, several control samples known to have or be suspected of having a normal repetitive sequence are analyzed, and high-count reads are obtained for the control samples as described above. The distribution of high-count reads in the control samples can be obtained, and the probability of unaffected samples having more high-count reads than a certain value can be estimated. This probability, given a calling criterion set to this specific value, allows for the determination of sensitivity and selectivity. In some embodiments, the calling criterion is set to a threshold, thereby ensuring that the probability of unaffected samples having more high-count reads than the threshold is less than 5%. In other words, the p-value is less than 0.05. In these embodiments, as the repeat sequence is extended, the repeat sequence becomes longer and can occur entirely within the repeat sequence, allowing for more high-count reads to be obtained for the sample. In various alternative embodiments, more conservative calling criteria may be selected such that the probability of an unaffected sample having more high-count reads than the threshold is less than approximately 1%, 0.1%, 0.01%, 0.001%, 0.0001%, etc. It will be understood that the calling criteria can be adjusted upward or downward based on various factors and the need to increase the sensitivity or selectivity of the test.

[0209] In some embodiments, instead of empirically obtaining a calling criterion of the number of high-count reads from a control sample, or in addition to that, a calling criterion may be theoretically obtained to determine repeat extensions. Considering a number of parameters including paired-end read length, sequence length with repeat extensions, and sequencing depth, it is possible to calculate the predicted number of reads that are fully present within a repeat. For example, sequencing depth can be used to calculate the average spacing between reads in an aligned genome. If individual samples are sequenced to a depth of 30x, the total number of sequenced bases is equal to the genome size multiplied by the depth. For humans, this is approximately 3x10⁻⁶. 9 x30 = 9 x 10 10 This would result in a total of 9x10 if each read is 100bp long. 8 There are 10 reads. Since the genome is diploid, half of these reads sequence one chromosome / haplotype and the other half sequence the other chromosome / haplotype. 4.5 × 10 per haplotype 8 There are 10^ 9 / 4.5×10 8(=1 read every 6.7 bp on average). This number can be used to estimate the number of reads that can be complete within a repeat sequence, based on the size of that repeat sequence in a particular individual. If the total size of the repeat sequence is 300 bp, any read starting within the first 200 bp of that repeat sequence can be complete within the repeat sequence (any read starting within the last 100 bp will be at least partially outside the repeat sequence, based on a read length of 100 bp). Since reads are expected to align every 6.7 bp, 200 bp / (6.7 bp / reads) = 30 reads are expected to be completely aligned within the repeat sequence. There is some variability around this number, but this makes it possible to estimate all reads that can be complete within a repeat sequence for any given elongation size. The length of the repeat sequence and the corresponding expected number of completely aligned reads within the repeat sequence calculated according to this method are shown in Table 2 of Example 1 below.

[0210] In some embodiments, the calling criterion is calculated from the distance between the first and last observations of a repeating sequence in the read, thus allowing for variations and sequencing errors within the repeating sequence.

[0211] In some embodiments, the process may further include diagnosing individuals from whom test samples are obtained that show an increased risk of genetic diseases such as fragile X syndrome, ALS, Huntington's disease, Friedreich's ataxia, spinocerebellar degeneration, bulbar spinal muscular atrophy, myotonic dystrophy, Machado-Joseph disease, dentatorubral-pallidoluysian atrophy, etc. Such a diagnosis may be based on the determination that repeat elongation is likely to be present in the test sample and is likely to be present on the gene and the repeat sequence with repeat elongation. In other embodiments, if the genetic disease is unknown, some embodiments can detect an abnormally high number of repeats to identify a new genetic cause of the disease.

[0212] Figure 16 is a flowchart illustrating another process for detecting repeat extensions according to several embodiments. Process 400 determines the presence of repeat extensions using the number of repetitions in paired-end reads of a test sample rather than high-count reads. Process 400 begins by sequencing a test sample containing nucleic acids to obtain paired-end reads. See block 402, equivalent to block 302 in process 300. Process 400 continues by aligning the paired-end reads to a reference sequence containing the repetition sequence. See block 404, equivalent to block 304 in process 300. This process proceeds by identifying anchor reads and anchor-type reads in the paired-end reads, where an anchor read is a read that aligns to or near the repetition sequence, and an anchor-type read is an unaligned read paired with the anchor read. In some embodiments, unaligned reads include both reads that cannot be aligned to the reference sequence and reads that are poorly aligned to the reference sequence.

[0213] After identifying the anchor leads and anchor-type leads, process 400 obtains the number of repeats within the anchor leads and / or anchor-type leads from the test sample. See block 408. Next, the process obtains the distribution of repeats for all anchor leads and / or anchor-type leads obtained from the test sample. In some embodiments, only the number of repeats from anchor-type leads is analyzed. In other embodiments, repeats from both anchor leads and anchor-type leads are analyzed. Next, the distribution of repeats in the test sample is compared to the distribution of one or more control samples. See block 410. In some embodiments, the process determines that repeat extension of the repeat sequence is present in the test sample if the distribution in the test sample differs statistically significantly from the distribution in the control samples. See block 412. Process 400 analyzes the number of repeats of leads including both high-count and low-count leads, unlike the process that analyzes only high-count leads as described above with respect to process 300.

[0214] In some embodiments, comparing the distribution of the test sample with that of the control sample involves using the Mann-Whitney Rank test to determine whether the two distributions are significantly different. In some embodiments, the analysis determines that repeat extension is likely to exist in the test sample if the distribution of the test sample is more skewed towards a higher number of repeats than that of the control sample, and the p-value of the Mann-Whitney Rank test is less than approximately 0.0001 or 0.00001. The p-value may be adjusted as necessary to improve the selectivity or sensitivity of the test.

[0215] The process for detecting repeat elongation described above, as shown in Figures 2-4, uses anchored leads, which are unaligned leads paired with leads aligned to the target repetitive sequence. Variations of these processes may include searching through unaligned leads in lead pairs, both composed almost entirely of several types of repetitive sequences, to discover novel, previously unidentified repeat elongations that may be medically relevant. While this method does not quantify the exact number of repeats, it is powerful for identifying extreme repeat elongations or outliers that should be flagged for further quantification. By combining this method with longer reads, it is possible to both identify and quantify repeats of up to 200 bp or more in total length.

[0216] Figure 17 shows a flowchart of process 500, which uses unaligned reads unrelated to any target repetitive sequence to identify repeat elongation. Process 500 may use unaligned reads from the entire genome to detect repeat elongation. This process is initiated by sequencing a test sample containing nucleic acids to obtain paired-end reads. See Block 502. Process 500 proceeds by aligning the paired-end reads to a reference genome. See Block 504. Next, this process identifies unaligned reads from the entire genome. Unaligned reads include paired-end reads that cannot be aligned to a reference sequence or are poorly aligned to a reference sequence. In some embodiments, poorly aligned reads include reads that are aligned to a reference sequence with an alignment quality score or mapping score below a certain threshold. In some embodiments, poorly aligned reads include reads that are aligned but have some inconsistent, inserted, or deleted bases. See Block 506. Next, this process analyzes the number of repeats of repeating units in the unaligned reads to determine whether repeat extensions are likely to be present in the test sample. This analysis cannot determine anything for any particular repeating sequence. The analysis can be applied to various potential repeating units, and the number of repeats of different repeating units from the test sample can be compared to those of multiple control samples. The comparison technique between the test sample and control samples described above can be applied to this analysis. If the comparison indicates that the test sample has an unusually large number of repeats of the repeating unit, additional analysis may be performed to determine whether the test sample contains repeat extensions of the specific repeating sequence in question. See Block 510.

[0217] In some embodiments, the additional analysis may involve very long sequence reads, which may extend to long repeat sequences with medically relevant repeat extensions. The reads in this additional analysis are longer than paired-end reads. In some embodiments, long reads are obtained using single-molecule sequencing or synthetic long-read sequencing. In some embodiments, the relationship between repeat extensions and genetic diseases is known in the art. However, in other embodiments, the relationship between repeat extensions and genetic diseases does not need to be established in the art.

[0218] In some embodiments, analyzing the number of repetitions in a repeating unit in unaligned leads in operation 510 includes a high-count analysis equivalent to operation 308 in Figure 3. The analysis includes obtaining the number of high-count leads, which are unaligned leads with more repetitions than a threshold, and comparing the number of high-count leads in the test sample to a calling criterion. In some embodiments, the threshold for high-count leads is at least about 80% of the maximum number of repetitions, and this maximum is calculated as the ratio of the length of the paired-end leads to the length of the repeating unit. In some embodiments, high-count leads also include leads that are paired with unaligned leads and also have more repeats than the threshold.

[0219] In some embodiments, prior to further analysis of operation 510, the process further includes (a) identifying paired-end reads that are paired with unaligned reads and aligned to or near a repetitive sequence on a reference genome, and (b) providing a repetitive sequence as a specific repetitive sequence to be targeted for operation 510. Further analysis of the target repetitive sequence can then employ any of the methods described above in relation to Figures 2 to 4.

[0220] sample The sample used to determine repeat extension may include a sample taken from any cell, fluid, tissue, or organ containing nucleic acids whose repeat extension of one or more repetitive sequences of interest is determined. In some embodiments involving fetal diagnosis, it is advantageous to obtain cell-free nucleic acids, such as cell-free DNA (cfDNA), from maternal fluids. Cell-free nucleic acids, including cell-free DNA, can be obtained from biological samples, including but not limited to plasma, serum, and urine, by a variety of methods known in the art (see, for example, Fan et al., Proc Natl Acad Sci 105:16266-16271

[2008] ; Koide et al., Prenatal Diagnosis 25:604-607

[2005] ; Chen et al., Nature Med.2:1033-1035

[1996] ; Lo et al., Lancet 350:485-487

[1997] ; Botezatu et al., Clin Chem.46:1078-1084, 2000; and Su et al., J Mol.Diagn.6:101-107

[2004] ).

[0221] In various embodiments, nucleic acids (e.g., DNA or RNA) present in a sample can be enriched specifically or nonspecificly before use (e.g., before preparing a sequencing library). DNA is used as an example of nucleic acid in the following exemplary examples. Nonspecific enrichment of sample DNA means the entire genomic amplification of the sample's genomic DNA fragments, which can be used to increase the level of sample DNA before preparing a cfDNA sequencing library. Whole-genome amplification methods are known in the art. Degenerate oligonucleotide-primed PCR (DOP), primer extension PCR technique (PEP), and multiple displacement amplification (MDA) are examples of whole-genome amplification methods. In some embodiments, the sample is not enriched with respect to DNA.

[0222] Samples containing nucleic acids to which the methods described herein are applied typically include biological samples ("test samples") as described above. In some embodiments, nucleic acids to be screened for repeat extension are purified or isolated by one of numerous known methods.

[0223] Therefore, in certain embodiments, the sample may contain or be essentially composed of purified or isolated polynucleotides, or the sample may include tissue samples, biological fluid samples, cell samples, etc. Suitable biological fluid samples include, but are not limited to, blood, plasma, serum, sweat, tears, sputum, urine, ear fluid, lymph, saliva, cerebrospinal fluid, lavage, bone marrow suspension, vaginal fluid, cervical fluid, femoral neck fluid, cerebral fluid, ascites, milk, secretions from the airways, intestines and urogenital tract, amniotic fluid, milk, and leukocyte phlebotomy samples. In some embodiments, the sample is a sample that can be easily obtained by non-invasive procedures, such as blood, plasma, serum, sweat, tears, sputum, urine, ear fluid, saliva, or feces. In certain embodiments, the sample is a peripheral blood sample, or the plasma and / or serous fraction of a peripheral blood sample. In other embodiments, the biological sample is a swab or smear, a biopsy specimen, or a cell culture. In yet another embodiment, the sample is a mixture of two or more biological samples, for example, the biological sample may include two or more of a biological fluid sample, a tissue sample, and a cell culture sample. As used in the present invention, the terms “blood,” “plasma,” and “serum” expressly include their fractions or processed portions. Similarly, if the sample is taken from a biopsy, swab, smear, etc., “sample” expressly includes the processed fraction or portion obtained from the biopsy, swab, smear, etc.

[0224] In certain embodiments, samples include, but are not limited to, samples from different individuals, samples from the same or different individuals at different developmental stages, samples from individuals with different diseases (e.g., individuals suspected of having a genetic disease), normal individuals, samples obtained at different stages of disease in an individual, samples obtained from individuals that have received different treatments for the disease, samples from individuals that have been subjected to different environmental factors, samples from individuals predisposed to the disease, and samples from individuals exposed to infectious agents.

[0225] In one exemplary but non-limiting embodiment, the sample is a maternal sample obtained from a pregnant woman, e.g., a pregnant woman. In this case, the sample can be analyzed using the methods described herein to provide an early diagnosis of potential chromosomal abnormalities in the fetus. The maternal sample may be a tissue sample, a biological fluid sample, or a cell sample. Examples of biological fluids include, but are not limited to, blood, plasma, serum, sweat, tears, sputum, urine, ear fluid, lymph, saliva, cerebrospinal fluid, lavage, bone marrow suspension, vaginal fluid, cervical fluid, femoral neck fluid, cerebral fluid, ascites, milk, secretions from the airways, intestines and urogenital tract, and leukocyte phlebotomy samples.

[0226] In certain embodiments, samples may also be obtained from in vitro cultured tissues, cells, or other polynucleotide-containing sources. Cultured samples may be taken from sources including, but not limited to, cultures (e.g., tissues or cells) maintained in different media and conditions (e.g., pH, pressure, or temperature), cultures (e.g., tissues or cells) maintained for different periods, cultures (e.g., tissues or cells) treated with different elements or reagents (e.g., drug candidates or modifiers), or cultures of different types of tissues and / or cells.

[0227] Methods for isolating nucleic acids from biological sources are known and vary depending on the nature of the source. Those skilled in the art can easily separate nucleic acids from sources as required by the methods described herein. In some cases, it may be advantageous to fragment the nucleic acid molecules in a nucleic acid sample. Fragmentation may be random or specific, for example, achieved by using restricted endonuclease digestion. Methods for random fragmentation are known in the art and include, for example, restricted DNAse digestion, alkaline treatment, and physical shearing.

[0228] Preparation of sequencing libraries In various embodiments, sequencing may be performed on various sequencing platforms that require the preparation of a sequencing library. Preparation typically includes fragmenting the DNA (sonication, atomization, or shearing), followed by DNA repair and end polishing (blunt end or A-overhang), and platform-specific adapter ligation. In one embodiment, the methods described herein can utilize next-generation sequencing technology (NGS), thereby enabling the sequencing of multiple samples individually as genomic molecules (i.e., singleplex sequencing) or individually as a pooled sample containing indexed genomic molecules on a single sequencing run (e.g., multiplex sequencing). These methods can generate up to several million DNA sequence reads. In various embodiments, genomic nucleic acid sequences and / or sequences of indexed genomic nucleic acids can be determined using, for example, next-generation sequencing technology (NGS) as described herein. In various embodiments, analysis of large amounts of sequence data obtained using NGS can be performed using one or more processors as described herein.

[0229] In various embodiments, the use of such sequencing techniques does not involve the preparation of a sequencing library.

[0230] However, in certain embodiments, the sequencing methods contemplated herein include the preparation of a sequencing library. In one exemplary approach, the preparation of a sequencing library includes the generation of a random collection of adapter-modified DNA fragments (e.g., polynucleotides) ready for sequencing. A polynucleotide sequencing library can be prepared from DNA or RNA, including equivalents or analogues of either DNA or cDNA, such as DNA or cDNA, which are complementary DNA or copy DNA generated from an RNA template by the action of reverse transcriptase. The polynucleotides may originate from a double-stranded form (e.g., genomic DNA fragments, cDNA, dsDNA such as PCR amplification products), or, in certain embodiments, the polynucleotides may originate from a single-stranded form (e.g., ssDNA, RNA, etc.) and be converted to dsDNA form. Exemplarily, in certain embodiments, a single-stranded mRNA molecule can be copied to a double-stranded cDNA suitable for use in the preparation of a sequencing library. The exact sequence of the primary polynucleotide molecule is generally not important to the method of library preparation and may be known or unknown. In one embodiment, the polynucleotide molecule is a DNA molecule. More specifically, in a particular embodiment, the polynucleotide molecule represents the entire or substantially entire genetic complement of an organism and is a genomic DNA molecule (e.g., cellular DNA, cell-free DNA (cfDNA), etc.), but typically includes intron and exon sequences (coding sequences), as well as non-coding regulatory sequences such as promoter and enhancer sequences. In a particular embodiment, the primary polynucleotide molecule includes a human genomic DNA molecule, for example, a cfDNA molecule present in the peripheral blood of a pregnant subject.

[0231] The preparation of sequencing libraries for some NGS sequencing platforms is facilitated by the use of polynucleotides containing a specific range of fragment sizes. Such library preparation typically involves fragmenting large polynucleotides (e.g., cellular genomic DNA) to obtain polynucleotides within a desired size range.

[0232] Paired-end leads are used in the methods and systems disclosed herein for determining repeat elongation. The length of the fragment or insert is longer than the lead length, and typically longer than the sum of the lengths of the two leads.

[0233] In some exemplary embodiments, the sample nucleic acid(s) is obtained as genomic DNA, which is fragmented into fragments of approximately 100+, 200+, 300+, 400+, or 500+ base pairs, allowing for easy application of NGS. In some embodiments, paired-end reads are obtained from inserts of approximately 100-5000 bp. In some embodiments, the inserts are approximately 100-1000 bp in length. These may be executed as paired-end reads for typical short inserts. In some embodiments, the inserts are approximately 1000-5000 bp in length. These may be executed as mate-paired reads for long inserts, as described above.

[0234] In some embodiments, long inserts are designed to evaluate very long, elongated repeat sequences. In some embodiments, mate-pair reads may be applied to obtain reads separated by thousands of base pairs. In these runs, the insert or fragment ranges from several hundred to several thousand base pairs, and there are two biotin-binding adapters on two ends of the insert. The biotin-binding adapters then join the two ends of the insert to form a circularized molecule, which is further fragmented. The fine fragments containing the biotin-binding adapters, and the two ends of the original insert, are selected for sequencing on a platform designed to sequence shorter fragments.

[0235] Fragmentation can be achieved by any of the many methods known to those skilled in the art. For example, fragmentation can be achieved by mechanical means, including but not limited to atomization, sonication, and hydroshearing. However, mechanical fragmentation typically cleaves the DNA backbone at CO, PO, and CC bonds, resulting in a heterogeneous mixture of blunt ends and 3'- and 5'-overhang ends with broken CO, PO, and CC bonds (see, e.g., Alnemri and Liwack, J Biol. Chem 265:17323-17333

[1990] ; Richards and Boyer, J Mol Biol 11:327-240

[1965] ), which may require repair due to the lack of 5' phosphates essential for subsequent enzymatic reactions, such as the ligation of sequencing adapters required to prepare DNA for sequencing.

[0236] In contrast, cfDNA typically exists as fragments of less than approximately 300 base pairs, and consequently, fragmentation is typically not necessary to generate sequencing libraries using cfDNA samples.

[0237] Typically, whether polynucleotides are forcibly fragmented (e.g., fragmented in vitro) or exist naturally as fragments, they are converted to blunt-terminated DNA having 5'-phosphate and 3'-hydroxyl. Standard protocols, such as those for sequencing using the Illumina platform as described elsewhere herein, instruct the user to purify the end-repaired product of the end-repaired sample DNA before dA-tailing and to purify the dA-tailing product before the adapter-ligating step of library preparation.

[0238] Various embodiments of the method for preparing a sequence library described herein eliminate the need to perform one or more of the steps typically mandated by standard protocols to obtain a modified DNA product that can be sequenced by NGS. The abbreviated method (ABB method), the 1-step method, and the 2-step method are examples of methods for preparing a sequencing library, which can be found in Patent Application No. 13 / 555,037 (filed July 20, 2012), which is incorporated herein by reference in its entirety.

[0239] Sequencing method As mentioned above, the prepared sample (e.g., a sequencing library) is sequenced as part of a procedure for identifying copy number variation. Any of a number of sequencing technologies can be used.

[0240] Some sequencing technologies are commercially available, including, as described below, the sequencing-by-hybridization platform from Affymetrix Inc. (Sunnyvale, Calif.), the sequencing-by-synthesis platforms from 454 Life Sciences (Bradford, Conn.), Illumina / Solexa (San Diego, Calif.), and Helicos Biosciences (Cambridge, Mass.), and the sequencing-by-ligation platform from Applied Biosystems (Foster City, Calif.). In addition to single molecule sequencing performed using sequencing-by-synthesis from Helicos Biosciences, other single molecule sequencing technologies include, but are not limited to, the SMRT™ technology from Pacific Biosciences, and the nanopore sequencing method developed by Oxford Nanopore Technologies.

[0241] The automated Sanger method is considered a "first-generation" technology, but Sanger sequencing including automated Sanger sequencing can also be employed in the methods described herein. Additional suitable sequencing methods include, but are not limited to, nucleic acid imaging techniques such as atomic force microscopy (AFM) or transmission electron microscopy (TEM). Exemplary sequencing techniques are described in further detail below.

[0242] In some embodiments, the disclosed method involves obtaining sequence information about nucleic acids in a test sample by large-scale parallel sequencing of millions of DNA fragments using sequencing chemistry based on Illumina sequencing bisynthesis and reversible terminators (as described, for example, in Bentley et al., Nature 6:53-59

[2009] ). The template DNA may be genomic DNA, e.g., cellular DNA or cfDNA. In some embodiments, genomic DNA from isolated cells is used as a template and fragmented into pieces of several hundred base pairs in length. In other embodiments, cfDNA is used as a template, but fragmentation is not required because cfDNA exists as short fragments. For example, fetal cfDNA circulates in the bloodstream as fragments of about 170 base pairs (bp) in length (Fan et al., Clin Chem 56:1279-1286

[2010] ), and does not require DNA fragmentation before sequencing. Illumina's sequencing technology relies on the attachment of fragmented genomic DNA to a planar, optically transparent surface to which oligonucleotide anchors are bound. Template DNA is repaired at the ends to produce 5' phosphorylated blunt ends, and a single A base is added to the 3' end of the blunt-phosphorylated DNA fragment using the polymerase activity of the Krenow fragment. This addition prepares the DNA fragment for ligation to oligonucleotide adapters, which have a single T base overhang at their 3' ends to enhance ligation efficiency. The adapter oligonucleotides are complementary to the anchor oligo in the flow cell (not to be confused with anchor reads / anchor-type reads in repeat extension analysis). Under restriction dilution conditions, adapter-modified single-strand template DNA is added to the flow cell and immobilized by hybridization to the anchor oligo. The attached DNA fragment is extended and the bridges are amplified to create ultra-high-density sequencing flow cells with hundreds of millions of clusters, each containing approximately 1,000 copies of the same template.In one embodiment, randomly fragmented genomic DNA is amplified using PCR before undergoing cluster amplification. Alternatively, unamplified genomic library preparation is used, and randomly fragmented genomic DNA is enriched using cluster amplification only (Kozarewa et al., Nature Methods 6:291-295

[2009] ). The template is sequenced using robust four-color DNA sequencing-by-synthesis technology with a reversible terminator having a removable fluorescent dye. High-sensitivity fluorescence detection is achieved using laser excitation and internal total internal reflection optical elements. Short sequence reads of approximately tens to hundreds of base pairs are aligned relative to a reference genome, and the unique mapping of the short sequence reads relative to the reference genome is identified using specially developed data analysis pipeline software. After the first read is completed, the template can be regenerated in situ to allow a second read from the opposite end of the fragment. Thus, either single-end sequencing or pair-end sequencing of the DNA fragment can be used.

[0243] Various embodiments of this disclosure may utilize sequencing bisynthesis to enable paired-end sequencing. In some embodiments, sequencing by an Illumina synthesis platform involves clustered fragments. Clustering is a process in which each fragment molecule is isothermally amplified. In some embodiments, as an example described herein, the fragment has two different adapters attached to the two ends of the fragment, which allow the fragment to hybridize with two different oligos on the surface of the flow cell lane. The fragment further includes, or is attached to, two index sequences at the two ends of the fragment, which provide labels for identifying different samples in multiplex sequencing. In some sequencing platforms, the fragment to be sequenced is also called an insert.

[0244] In some embodiments, the flow cell for clustering on the Illumina platform is a glass slide with lanes. Each lane is a glass channel coated with a community of two types of oligos. Hybridization is enabled by the first of the two oligos on the surface. This oligo is complementary to a first adapter at one end of the fragment. Polymerase generates the complementary strand of the hybridized fragment. The double-stranded molecule denatures and washes away the original template strand. The remaining strand is clonally amplified by bridging in parallel with many other remaining strands.

[0245] In bridge amplification, a second adapter region on the second end of the chain hybridizes with a second type of oligo on the flow cell surface. Polymerase generates a complementary chain, forming a double-stranded crosslink molecule. This double-stranded molecule denatures, resulting in two single-stranded molecules linked to the flow cell via two different oligos. This process is then repeated across millions of clusters, occurring simultaneously and resulting in clonal amplification of all fragments. After bridge amplification, the reverse strand is cleaved and washed away, leaving only the forward strand. The 3' end is blocked to prevent undesirable priming.

[0246] After clustering, sequencing begins by extending the first sequencing primer to generate the first read. In each cycle, fluorescently labeled nucleotides compete to be added to the growing chain. Only one is incorporated based on the template sequence. After the addition of each nucleotide, the cluster is excited by a light source, emitting a characteristic fluorescence signal. The number of cycles determines the read length. The emission wavelength and signal intensity determine the base call. For a given cluster, all identical chains are read simultaneously. Hundreds of millions of clusters are sequenced in a large-scale parallel manner. Upon completion of the first read, the read product is washed away.

[0247] In the next step of the protocol, which includes two index primers, the index 1 primer is introduced and mixed into the index 1 region on the template. The index region provides identification of fragments useful for demultiplexing the sample in the multiplex sequencing process. Index 1 reads are generated in the same manner as the first reads. After the index 1 reads are complete, the read product is washed away to deprotect the 3' end of the chain. The template chain is then folded over the second oligo on the flow cell and bound to the second oligo. The index 2 sequence is read in the same manner as index 1. The index 2 read product is then washed away at the end of the step.

[0248] After reading the two indices, polymerase is used to initiate read 2, extending the second flow cell oligo to form a double-stranded bridge. This double-stranded DNA is denatured, and the 3' end is cut off. The original forward strand is cleaved and washed away, leaving the reverse strand. Read 2 is initiated with the introduction of the read 2 sequencing primer. Similar to read 1, the sequencing process is repeated until the desired length is achieved. The product of read 2 is washed away. This entire process generates millions of reads representing all fragments. Sequences from the pooled sample library are separated based on the unique indices introduced during sample preparation. For each sample, reads of similar extension base calls are locally clustered. Forward and reverse reads are paired to create a contiguous sequence. These contiguous sequences are aligned to a reference genome for variant identification.

[0249] The above example of sequencing bisynthesis involves paired-end reads, which are used in many embodiments of the disclosed method. A paired-end sequence contains two reads from two ends of a fragment. Paired-end reads are used to resolve ambiguous alignments. Paired-end sequencing allows a user to select the length of an insert (or fragment to be sequenced), sequence either end of the insert, and generate high-quality, alignable sequence data. Because the distance between each paired read is known, the alignment algorithm can use this information to more accurately position the reads on repetitive regions. This results in better alignment of reads, particularly across repetitive regions of the genome that are difficult to sequence. Paired-end sequencing can detect realignments, including insertions and deletions (indels) as well as inversions.

[0250] Paired-end reads may use inserts of different lengths (i.e., different fragment sizes to be sequenced). In the default sense of this disclosure, paired-end reads are used to mean reads obtained from various insert lengths. In some cases, to distinguish paired-end reads from paired-end reads from paired-end reads from long inserts, the latter are specifically referred to as mate-paired reads. In some embodiments involving mate-paired reads, two biotin-binding adapters are first attached to the two ends of a relatively long insert (e.g., several kb). The biotin-binding adapters then link the two ends of the insert to form a circulating molecule. Next, a fine fragment containing the biotin-binding adapters can be obtained by further fragmenting the circulating molecule. Then, a fine fragment containing the two ends of the original fragment in the reverse order can be sequenced by the same procedure as the paired-end sequencing of short inserts described above. Further details on mate-pair sequencing using the Illumina platform are presented in an online publication at the following address, which is incorporated herein by reference in its entirety: res.illumina.com / documents / products / technotes / technote_nextera_matepair_data_processing.pdf

[0251] After sequencing of DNA fragments, sequence reads of a predetermined length (e.g., 100 bp) are mapped or aligned to a known reference genome. The mapped or aligned reads and their corresponding positions on the reference sequence are also called tags. Many of the analyses in the embodiments disclosed herein for determining repeat extension use reads that are poorly aligned or cannot be aligned, as well as aligned reads (tags). In one embodiment, the reference genome sequence is the NCBI36 / hg18 sequence, which is available on the World Wide Web at genome.ucsc.edu / cgi-bin / hgGateway?org=Human&db=hg18&hgsid=166260105. Alternatively, the reference genome sequence is GRCh37 / hg19, which is available on the World Wide Web (www) at genome.ucsc.edu / cgi-bin / hgGateway. Other sources of publicly available sequence information include GenBank, dbEST, dbSTS, EMBL (the European Molecular Biology Laboratory), and DDBJ (the Japanese DNA database). Numerous computer algorithms are available for sequence alignment, including, but not limited to, BLAST (Altschul et al., 1990), Blitz (MPsrch) (Sturrock & Collins, 1993), FASTA (Person & Lipman, 1988), BOWTIE (Langmead et al., Genome Biology 10:R25.1~R25.10

[2009] ), or ELAND (Illumina, Inc., San Diego, California, USA). In one embodiment, one end of a clonal extension copy of a plasma cfDNA molecule is sequenced and processed by genetic information alignment analysis on the Illumina Genome Analyzer using Efficient Large-Scale Alignment of Nucleotide Databases (ELAND) software.

[0252] In one exemplary but non-limiting embodiment, the method described herein involves obtaining sequence information about nucleic acids in a test sample using single-molecule sequencing technology of Helicos True Single Molecule Sequencing (tSMS) technology (as described, for example, in Harris TD et al., Science 320:106-109

[2008] ). In tSMS technology, a DNA sample is cleaved into strands of approximately 100-200 nucleotides, and a poly(A) sequence is added to the 3' end of each DNA strand. Each strand is labeled by the addition of a fluorescently labeled adenosine nucleotide. The DNA strands are then mixed into a flow cell containing millions of oligo-T capture sites immobilized on the surface of the flow cell. In a particular embodiment, the template is approximately 100 million templates / cm². 2 The density can be as follows. The flow cell is then added to an instrument, such as a HeliScope® sequencer, and a laser is shone on the surface of the flow cell to reveal the position of each template. A CCD camera can position the template on the surface of the flow cell. Next, the template fluorescent label is cleaved and washed away. The sequencing reaction is initiated by introducing DNA polymerase and fluorescently labeled nucleotides. Oligo-T nucleic acids act as primers. The polymerase incorporates the labeled nucleotides into the primers in a template-inducing manner. The polymerase and unincorporated nucleotides are removed. Templates with induced incorporation of fluorescently labeled nucleotides are identified by imaging the surface of the flow cell. After imaging, the cleavage step removes the fluorescent label. This process is repeated with other fluorescently labeled nucleotides until the desired read length is achieved. Sequence information is collected in each nucleotide addition step. Whole-genome sequencing by single-molecule sequencing technology excludes or typically eliminates the amplification of PCR systems in the preparation of sequencing libraries, and this method allows for the direct measurement of the sample rather than the measurement of a copy of the sample.

[0253] In another exemplary but non-limiting embodiment, the method described herein includes obtaining sequence information about nucleic acids in a test sample using 454 sequencing (Roche) (as described, for example, in Margulies, M. et al. Nature 437:376-380

[2005] ). 454 sequencing typically involves two steps. In the first step, the DNA is sheared into fragments of approximately 300–800 base pairs, with the fragments having blunt ends. Oligonucleotide adapters are then ligated to the ends of the fragments. The adapters serve as primers for amplification and sequencing of the fragments. The fragments can be attached to DNA capture beads, for example, beads coated with streptavidin, using adapter B containing a 5' biotin tag, for example. The fragments attached to the beads are PCR amplified in droplets of an oil-in-water emulsion. The result is multiple copies of the clonely amplified DNA fragment on each bead. In the second step, the beads are trapped in wells (e.g., picoliter-sized wells). Pyrosequencing is performed in parallel for each DNA fragment. The addition of one or more nucleotides generates an optical signal recorded by a CCD camera in the sequencing apparatus. The signal intensity is proportional to the number of nucleotides incorporated. Pyrosequencing uses pyrophosphate (PPi) released upon nucleotide addition. PPi is converted to ATP by ATP sulfurylase in the presence of adenosine 5'-phosphosulfate. Luciferase uses ATP to convert luciferin to oxyluciferin, a reaction that generates the light to be measured and analyzed.

[0254] In another exemplary but non-limiting embodiment, the method described herein includes obtaining sequence information about nucleic acids in a test sample using SOLiD® technology (Applied Biosystems). SOLiD® sequencing biligation involves shearing genomic DNA into fragments and attaching adapters to the 5' and 3' ends of the fragments to generate a fragment library. Alternatively, internal adapters may be introduced by ligating the adapters to the 5' and 3' ends of the fragments, circulating the fragments, digesting the circularized fragments to generate internal adapters, and attaching the adapters to the 5' and 3' ends of the resulting fragments to generate a mate-pair library. Next, a clonal bead population is prepared in a microreactor containing beads, primers, a template, and PCR components. After PCR, the template is denatured, and the beads are concentrated to separate the beads from the extended template. The template on selected beads undergoes 3' modification to enable binding to a glass slide. The sequence can be determined by sequential hybridization and ligation of partially random oligonucleotides using a central determination base (or base pair) identified by a specific phosphor. After the color is recorded, the ligated oligonucleotides are cleaved and removed, and then the process is repeated.

[0255] In another exemplary but non-limiting embodiment, the method described herein includes obtaining sequence information about nucleic acids in a test sample using Pacific Biosciences' single-molecule, real-time (SMRT®) sequencing technology. In SMRT sequencing, the sequential incorporation of dye-labeled nucleotides is imaged during DNA synthesis. A single DNA polymerase molecule is mounted on the bottom of individual zero-mode wavelength detectors (ZMW detectors) that acquire sequence information, while a phospho-binding nucleotide is incorporated into a growing primer chain. The ZMW detector includes a constraint structure that allows observation of single-base incorporation by DNA polymerase against a background of fluorescent nucleotides rapidly diffusing out of the ZMW (e.g., in microseconds). It typically takes several milliseconds to incorporate the nucleotide into the growing chain. At this time, the fluorescent label is excited, producing a fluorescent signal, and the fluorescent label is cleaved. Measurement of the corresponding fluorescence of the dye indicates which base has been incorporated. This process is repeated to provide the sequence.

[0256] In another exemplary but non-limiting embodiment, the methods described herein include obtaining sequence information about nucleic acids in a test sample using nanopore sequencing (for example, as described in Soni GV and Meller A. Clin Chem 53:1996-2001

[2007] ). Nanopore sequencing DNA analysis techniques have been developed by numerous companies, including, for example, Oxford Nanopore Technologies (Oxford, United Kingdom), Sequenom, NABsys, and others. Nanopore sequencing is a single-molecule sequencing technique in which a single molecule of DNA is directly sequenced as it passes through a nanopore. A nanopore is a small, ordinal pore, typically 1 nanometer in diameter. When a nanopore is immersed in a conductive fluid and an electric potential (voltage) is applied, a small current is generated due to the conduction of ions through the nanopore. The amount of current that flows is sensitive to the size and shape of the nanopore. As DNA molecules pass through nanopores, each nucleotide on the DNA molecule obstructs the nanopores to a different degree, causing variations in the magnitude of the current passing through the nanopores. Therefore, this change in the current as DNA molecules pass through nanopores results in DNA sequence reads.

[0257] In another exemplary but non-limiting embodiment, the method described herein includes obtaining sequence information about nucleic acids in a test sample using a chemosensitive field-effect transistor (chemFET) array (as described, for example, in U.S. Patent Publication No. 2009 / 0026082). In one example of the technique, a DNA molecule can be placed in a reaction chamber, and a template molecule can be mixed with a sequencing primer conjugated to polymerase. The incorporation of one or more triphosphates into a new nucleic acid chain at the 3' end of the sequencing primer can be identified as a change in current by the chemFET. The array may have multiple chemFET sensors. In another embodiment, a single nucleic acid can be attached to a bead, the nucleic acid can be amplified on the bead, and the individual beads can be transferred to individual reaction chambers on a chemFET array, each chamber having a chemFET sensor, and the nucleic acid can be sequenced.

[0258] In another embodiment, the DNA sequencing technology is Ion Torrent single-molecule sequencing, which pairs semiconductor technology with a single sequencing chemistry to directly translate chemically encoded information (A, C, G, T) into digital information (0, 1) on a semiconductor chip. Essentially, when nucleotides are incorporated into the DNA strand by polymerase, hydrogen ions are released as a byproduct. Ion Torrent carries out this biochemical process in a large-scale parallel manner using a high-density array of micro-machined wells. Each well holds a different DNA molecule. Beneath the wells is an ion-sensitive layer, but beneath the ion sensor. When a nucleotide, e.g., C, is added to the DNA template and then incorporated into the DNA strand, hydrogen ions are released. The charge from these ions can change the pH of the solution, which can be detected by Ion Torrent's ion sensor. The sequencer (essentially the world's smallest solid-state pH meter) calls the bases, going directly from chemical information to digital information. Next, the Ion Personal Genome Machine (PGM®) sequencer immerses the chip with nucleotides one after another. If the next nucleotide flowing across the chip does not match, no voltage change is recorded, and no base is called. If two identical bases exist on the DNA strand, the voltage may be duplicated, and the chip records the two identical bases that were called. Direct detection allows for recording of nucleotide incorporation in seconds.

[0259] In another embodiment, the method includes obtaining sequence information of nucleic acids in a test sample using sequencing by hybridization. Sequencing bihybridization involves contacting multiple polynucleotide sequences with multiple polynucleotide probes, each of which may optionally be attached to a substrate. The substrate may be a flat surface containing an array of known nucleotide sequences. The pattern of hybridization to the array can be used to determine the polynucleotide sequences present in the sample. In other embodiments, each probe is attached to a bead, such as a magnetic bead. Hybridization to the beads can be determined and used to identify multiple polynucleotide sequences in the sample.

[0260] In some embodiments of the methods described herein, sequence reads are approximately 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 110 bp, 120 bp, 130, 140 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, or 500 bp. Technological advancements are expected to enable single-ended reads exceeding 500 bp, and, when paired-ended reads are generated, enable reads exceeding approximately 1000 bp. In some embodiments, paired-end reads are used to determine repeat elongation, including sequence reads that are approximately 20 bp–1000 bp, approximately 50 bp–500 bp, or 80 bp–150 bp. In various embodiments, paired-end reads are used to evaluate sequences that have repeat elongation. Sequences with repeat elongation are longer than the read. In some embodiments, sequences with repeat elongation are longer than approximately 100 bp, 500 bp, 1000 bp, or 4000 bp. Sequence read mapping is achieved by comparing the sequence of the read to a reference sequence to determine the chromosomal origin of the sequenced nucleic acid molecule, and specific gene sequence information is not required. Minor mismatches (0–2 mismatches per read) can account for trace polymorphisms that may exist between the reference genome and the genome in the mixed sample. In some embodiments, reads aligned to the reference sequence are used as anchor reads, and reads that pair with the anchor read but cannot align to the reference sequence or align poorly to the reference sequence are used as anchor-type reads. In some embodiments, poorly aligned leads may have a relatively high mismatch rate per lead, for example, at least about 5%, at least about 10%, at least about 15%, or at least about 20% mismatch per lead.

[0261] Multiple sequence tags (i.e., reads aligned to a reference sequence) are typically obtained per sample. In some embodiments, for example, 100 bp, at least 3 × 10⁶ 6 A sequence of tags, at least approximately 5 x 10 6 A sequence of tags, at least approximately 8 x 10 6 A number of array tags, at least approximately 10 x 10 6 A sequence of tags, at least approximately 15 x 10 6 A sequence of tags, at least approximately 20 x 10 6 A sequence of tags, at least approximately 30 x 10 6 A sequence of tags, at least approximately 40 x 10 6 A sequence of tags, or at least approximately 50 x 10 6 A sequence tag is obtained from the mapping of reads to a reference genome per sample. In some embodiments, all sequence reads are located across the entire region of the reference genome, providing genome-wide reads. In other embodiments, they are located on the target sequence, such as a chromosome, a chromosomal segment, or a target repetitive sequence.

[0262] Apparatus and system for determining repeat extension Analysis of sequencing data and diagnostics obtained therefrom is typically performed using various computer-executed algorithms and programs. Accordingly, certain embodiments employ processes involving data stored in or transferred through one or more computer systems or other processing systems. Embodiments disclosed herein also relate to apparatus for performing these operations. The apparatus may be specially constructed for the required purposes, or it may be a general-purpose computer (or a group of computers) selectively activated or reconfigured by a computer program and / or data structure stored in the computer. In some embodiments, a group of processors collaboratively (e.g., via a network or cloud computing) and / or in parallel performs some or all of the recited analytical operations. A processor or group of processors for performing the methods described herein may be of various types including microcontrollers and microprocessors such as programmable devices (e.g., CPLDs and FPGAs), and non-programmable devices such as gate array ASICs or general-purpose microprocessors.

[0263] One embodiment provides a system for use in determining the genotype of a variant at a genomic locus containing repetitive sequences, the system comprising: a sequencer for receiving a nucleic acid sample and providing nucleic acid sequence information from the sample; a processor; and a machine-readable storage medium for genotyping a variant, wherein the variant is genetically determined by (a) collecting nucleic acid sequence reads of a test sample from a database; (b) aligning the sequence reads to one or more repetitive sequences represented by a sequence graph, the sequence graph having a directed graph data structure with vertices representing nucleic acid sequences and directed edges connecting the vertices, the sequence graph comprising one or more self-loops, each self-loop representing a repetitive subsequence, and each repetitive subsequence comprising repetitions of one or more repeating units of nucleotides; and (c) determining one or more genotypes relating to one or more repetitive sequences using the sequence reads aligned to one or more repetitive sequences.

[0264] In some embodiments of any of the systems provided herein, the sequencer is configured to perform next-generation sequencing (NGS). In some embodiments, the sequencer is configured to perform large-scale parallel sequencing using sequencing bisynthesis with a reversible diterminator. In other embodiments, the sequencer is configured to perform sequencing biligation. In yet another embodiment, the sequencer is configured to perform single-molecule sequencing.

[0265] In addition, certain embodiments relate to tangible and / or non-temporary computer-readable media or computer program products containing program instructions and / or data (including data structures) for performing various computer implementation operations. Examples of computer-readable media include, but are not limited to, semiconductor memory devices, magnetic media such as disk drives, magnetic tapes, optical media (such as CDs and magneto-optical media), and hardware devices specifically configured to store and execute program instructions, such as read-only memory devices (ROM) and random access memory (RAM). The computer-readable media may be directly controlled by the end user, or the media may be indirectly controlled by the end user. An example of directly controlled media is media located on a medium not shared with user facilities and / or other components. An example of indirectly controlled media is media that is indirectly accessible to the user via an external network and / or via a service that provides shared resources such as a “cloud.” Examples of program instructions include both machine code, such as that generated by a compiler, and files containing higher-level code than that which can be executed by a computer using an interpreter.

[0266] In various embodiments, the data or information used in the disclosed methods and apparatus is provided in electronic form. Such data or information may include reads and tags derived from nucleic acid samples, reference sequences (including single or primarily polymorphic reference sequences), repeat extension calls, counseling recommendations, diagnostic calls, etc. When used in the present invention, the data or other information provided in electronic form is available for machine storage and machine-to-machine transmission. Conventionally, data in electronic form is provided digitally and may be stored as bits and / or bytes in various data structures, lists, databases, etc. The data may be embodied electronically, optically, etc.

[0267] One embodiment provides a computer program product for generating an output indicating the presence or absence of repeat extensions in a test sample. The computer product may include instructions for performing any one or more of the above methods for determining repeat extensions. As described, the computer product may include a non-temporary and / or tangible computer-readable medium having computer-executable or compilable logic (e.g., instructions) recorded thereon, thereby enabling a processor to determine whether or not anchor reads and repetitions, and repeat extensions, are present in anchor-type reads. In one embodiment, the computer product includes a computer-readable medium having computer-executable or compilable logic (e.g., instructions) recorded thereon, enabling a processor to diagnose repeat extensions, which includes a receiving procedure for receiving sequencing data from at least a portion of nucleic acid molecules from a biological sample, wherein the sequencing data includes paired-end reads that have been aligned into a repeating sequence; computer-aided logic for analyzing repeat extensions from the received data; and an output procedure for generating an output indicating the presence or type of repeat extensions.

[0268] Sequence information from the sample under consideration can be positioned in a chromosomal reference sequence to identify paired-end reads aligned to or anchored to the target repetitive sequence, and to identify repeat extensions of the repetitive sequence. In various embodiments, the reference sequence is stored in a database such as a relational database or an object database.

[0269] It should be understood that it is impractical, or in most cases even impossible, for humans to perform the computational operations of the methods disclosed herein without assistance. For example, mapping a single 30 bp read from a sample to any one of the human chromosomes can require considerable effort without the assistance of a computing device. Naturally, the problem is complicated because reliable repeat extension calls generally require mapping thousands (e.g., at least about 10,000) or even millions of reads to one or more chromosomes.

[0270] In various embodiments, raw sequence reads are aligned to one or more sequence graphs representing one or more target sequences. In various embodiments, at least 10,000, 100,000, 500,000, 1,000,000, 5,000,000, or 10,000,000 reads are aligned to one or more sequence graphs. In various embodiments, one or more sequence graphs include at least 1, 2, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, or 50,000 sequence graphs.

[0271] In some embodiments, raw sequence reads are first aligned to a reference genome to determine the genomic coordinates of the reads before a subset of the initially aligned reads are aligned to one or more sequence graphs representing one or more target sequences. In various embodiments, at least 10,000, 100,000, 500,000, 1,000,000, 5,000,000, 10,000,000, or 100,000,000 reads are initially aligned to the reference genome. In some embodiments, the initially aligned reads are re-aligned to sequence graphs to determine repeat extensions in numerous regions (each region corresponding to a sequence graph). The total number of reads re-aligned to sequence graphs during each embodiment may range from several thousand to several million reads. In various embodiments, at least 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000, or 10,000,000 reads are realigned into each array graph. In various embodiments, one or more array graphs include at least 1, 2, 5, 10, 50, 100, 500, 1,000, 5,000, 10,000, or 50,000 array graphs.

[0272] The methods disclosed herein may be carried out using a system for determining the genotype of variants at genomic loci containing repetitive sequences. The system may include (a) a sequencer for receiving nucleic acids from a test sample that provides nucleic acid sequence information from the sample; (b) a processor; and (c) one or more computer-readable storage media having instructions stored thereon for performing genotyping of variants at genomic loci containing repetitive sequences on the processor. In some embodiments, the method is directed by a computer-readable medium having computer-readable instructions stored thereon for performing a method for identifying any repeat extensions. Thus, one embodiment provides a computer program product including a non-temporary machine-readable medium storing program code when executed by one or more processors of a computer system, causing the computer system to perform a method for identifying repeat extensions of repetitive sequences in a test sample containing nucleic acids, wherein the repetitive sequence includes repeats of nucleotide repeating units. The program code may include (a) code for collecting sequence reads of a test sample from a database; (b) code for aligning the sequence reads to one or more repeat sequences represented by a sequence graph, wherein the sequence graph has a directed graph data structure with vertices representing nucleic acid sequences and directed edges connecting the vertices, the sequence graph includes one or more self-loops, each self-loop representing a repeating subsequence, and each repeating subsequence including repeats of one or more nucleotide repeating units; and (c) code for determining one or more genotypes relating to one or more repeat sequences using the sequence reads aligned to the one or more repeat sequences.

[0273] In some embodiments, the instructions may further include automatically recording information related to the method, such as repeating leads and anchored leads, and whether or not there is repeat extension in the medical records of the patient of the human subject providing the test sample. The patient's medical records may be managed, for example, by a laboratory, a doctor's office, a hospital, a healthcare facility, an insurance company, or a personal medical record website. Furthermore, based on the results of the analysis performed by the processor, the method may further include directing, initiating, and / or modifying the treatment of the human subject from whom the test sample was taken. This may include performing one or more additional tests or analyses on additional samples taken from the subject.

[0274] The disclosed methods may also be carried out using a computer processing system adapted or configured to perform a method for identifying any repeat extensions. One embodiment provides a computer processing system adapted or configured to perform the methods described herein. In one embodiment, the apparatus includes a sequencing device adapted or configured to sequence at least a portion of nucleic acid molecules in a sample to obtain the types of sequence information described elsewhere in this specification. The apparatus may also include components for processing the sample. Such components are described elsewhere in this specification.

[0275] Sequences or other data may be input into a computer or stored directly or indirectly on a computer-readable medium. In one embodiment, the computer system is directly connected to a sequencing device that reads and / or analyzes nucleic acid sequences from a sample. Sequences or other information from such a tool are provided through an interface within the computer system. Alternatively, sequences processed by the system are provided from a sequence storage source such as a database or other repository. Once the processing device is available, a memory device or mass storage device buffers or stores the nucleic acid sequences at least temporarily. In addition, the memory device may store a number of tags for various chromosomes or genomes, etc. The memory may also store various routines and / or programs for analyzing the presentation of sequences or mapped data. Such programs / routines may include programs for performing statistical analysis, etc.

[0276] In one embodiment, a user provides a sample to a sequencing device. The data is collected and / or analyzed by a sequencing device connected to a computer. Software on the computer enables the data collection and / or analysis. The data may be stored, displayed (via a monitor or other similar device), and / or transmitted to another location. The computer may be connected to the Internet, which is used to transmit the data to a handheld device used by a remote user (e.g., a physician, scientist, or analyst). It is understood that the data may be stored and / or analyzed before transmission. In some embodiments, raw data is collected and transmitted to a remote user or device that analyzes and / or stores the data. Transmission may be via the Internet, but may also be via satellite or other connections. Alternatively, the data may be stored on a computer-readable medium, which may be delivered to an end user (e.g., via email). Remote users may be located in the same or different geographical locations, including but not limited to buildings, cities, states, countries, or continents.

[0277] In some embodiments, the method also includes collecting data relating to multiple polynucleotide sequences (e.g., reads, tags, and / or reference chromosome sequences) and transmitting the data to a computer or other computing system. For example, the computer may be connected to laboratory equipment, such as a sample collection device, a nucleotide amplification device, a nucleotide sequencing device, or a hybridization device. The computer can then collect the applicable data collected by the laboratory device. The data may be stored on the computer at any step, for example, during real-time collection, before, during, or in connection with the transmission, or after the transmission. The data may be stored on a computer-readable medium that can be extracted from the computer. The collected or stored data may be transmitted from the computer to a remote location, for example, via a local network or a wide-area network such as the Internet. At the remote location, various operations can be performed on the transmitted data, as described below.

[0278] The types of electronically formatted data that can be stored, transmitted, analyzed, and / or manipulated by the systems, apparatus, and methods disclosed herein include, among others: Reads obtained by sequencing nucleic acids in test samples Tags obtained by aligning reads to a reference genome or one or more other reference sequences. Reference genome or sequence Locus specifications indicate the identity, location, and structure of a gene locus. Lead coverage Variant genotype Array graph Graph Path Graph sorting information Actual Calls for Repeat Customer Growth Diagnosis (clinical condition related to the call) Recommendations for further testing derived from calls and / or diagnoses Treatment and / or monitoring plan derived from call and / or diagnosis

[0279] These various types of data may be acquired, stored, transmitted, analyzed, and / or manipulated at one or more locations using separate devices. Processing options span a broad spectrum. At one end of the spectrum, all or much of this information is stored and used at the location where the test sample is processed, e.g., a physician's office or other clinical setting. In other extreme cases, the sample is acquired at one location, processed at different locations, sequenced as desired, the reads aligned, calls made at one or more different locations, and diagnostic, recommendation, and / or planning prepared at yet another location (which may be where the sample was obtained).

[0280] In various embodiments, reads are generated by a sequencing device and then sent to a remote site where they are processed to generate repeat extension calls. At this remote location, for example, the reads are aligned into a reference array to generate anchor reads and anchor-type reads. The processing operations that may be employed at individual locations are as follows: Sample collection Preliminary sample processing for sequencing Sequence determination Analyze array data and derive repeat expansion calls. diagnosis Report the diagnosis and / or call to the patient or healthcare provider. Develop a plan for further processing, testing, and / or monitoring.

[0281] Any one or more of these operations can be automated as described elsewhere in this specification. Typically, sequencing and analysis of sequence data and derivation of repeat extension calls can be performed computationally. Other operations may be performed manually or automatically.

[0282] Figure 18 shows one embodiment of a distributed system for generating calls or diagnoses from test samples. A sample collection location 01 is used to collect a test sample from a patient. The sample is then provided to a processing and sequencing location 03, which can process and sequence the test sample as described above. Location 03 includes equipment for processing the sample, as well as equipment for sequencing the processed sample. The sequencing results are typically provided in electronic format, as described elsewhere in this specification, and are a collection of reads provided to a network such as the Internet, shown in Figure 18 by reference no. 05.

[0283] Sequence data is provided to a remote location 07 where analysis and call generation are performed. This location may include one or more powerful computing devices, such as a computer or processor. After the computing resources at location 07 complete their analysis and generate a call from the received sequence information, the call is relayed back to network 05. In some embodiments, not only the call generated at location 07 but also associated diagnostics are generated. The call and / or diagnostics are then transmitted across the network and returned to the sample collection location 01, as shown in Figure 18. As will be described, this is one of many variations in how the various operations related to generating a call or diagnostic can be divided among the various locations. One common variant involves providing sample collection and processing and sequencing at a single location. Another variant involves providing processing and sequencing at the same location as the analysis and call generation.

[0284] experiment Example 1: Short repetitions Examples 1-3 visualize correctly genotyped repeat regions in several embodiments. We consider read pileups of ATXN3 repeats where the alleles are shorter than the read lengths shown in Figure 19. Figure 19 shows read pileups of ATXN3 repeats with genotype 20 / 20, having 20 motifs on repeat region 1902 on both haplotypes. Sequence interruptions correspond to positions where most of the read alignment is mismatched.

[0285] Each panel in this plot corresponds to a haplotype. Haplotype sequences and reads are colored according to their overlap with repeat 1902 (orange) or surrounding flanking sequences (blue). All mismatched bases within the reads are shown, and the locations where alignment is clipped are indicated by jagged edges.

[0286] This pile-up plot shows that the genotype call is well supported by the reads, as each allele is supported by many spanning reads (reads that extend across the entire repeat) and there are no reads with mismatched alignments. (Mismatched alignments mean that the read does not match either of the two haplotypes. For example, a read with 40 repeats does not match genotype 20 / 20.) There is clear evidence of interruption in the repeat sequences. For example, cytosine in the third to last motif is mutated to thymine.

[0287] Example 2: Extended Repeat Figure 20 shows a DMPK repeat with a normal-sized allele 2002 and an elongated allele 2204. The elongated repeat is well supported by reads so that embodiments distribute the reads throughout the repeat to achieve similar read coverage across the haplotype. Note that the alignment of reads within the repeat is randomly selected. The short allele is also well supported by a large number of spanning reads.

[0288] Example 3: Loci with two adjacent repeats To demonstrate more complex applications of some embodiments, they are used to visualize an HTT repeat region containing two adjacent repeats, namely a pathogenic CAG repeat and a nearby “harmful” CCG repeat. The former repeat is genotyped as 14 / 17, and the latter repeat is genotyped as 9 / 12. Figure 21A shows a read pileup of the HTT locus containing two nearby repeats, specifically the CAG repeat 2104 and the CCG repeat 2108. This pileup also includes the left flank 2102, the intervening sequence 2106CAACAG, and the right flank 2110.

[0289] As a result, one of the haplotypes shown in Figure 21A contains repeats of sizes 14 and 12, respectively, while the other haplotype contains repeats of sizes 17 and 9. It is clear that both haplotypes are well supported by reads. Furthermore, this pile-up plot shows that there is a G-to-A mutation in the second copy of the CCG repeat motif on both haplotypes. In particular, the coverage level is relatively uniform across locations on both haplotypes.

[0290] For comparison, Figure 21B shows a sequence pileup of the HTT region using conventional tools and the same sequence read data. This pileup includes only one strand reference sequence instead of two individualized haplotypes. The repeat region includes two nearby repeats, specifically the CAG repeat (2124) and the CCG repeat (2128). This pileup also includes the left flank (2122), the intervening sequence CAACAG (2126), and the right flank (2130). Note that the sequence reads are not uniformly distributed across the reference sequence. Coverage in the repeat region 2128 is low, and numerous reads are split so that they extend through sections of the repeat region with low or no coverage. This is a sign that the data does not match the genotype of the reference in this region, but the pileup does not clearly indicate the true genotype of the sample.

[0291] Example 4: Overestimated iteration size Examples 4 and 5 visualize mistyped repeat regions. To provide an example of a pile-up corresponding to a false-positive repeat extension call, Example 4 shows a simulated read from a C9ORF72 repeat region with homozygous genotype 10 / 10. The practitioner performed several embodiments of spiking in C homopolymer reads that closely resemble the repeat sequence and forcing the repeat genotype to 10 / 30 instead of 10 / 10. Figure 22 shows a read pile-up containing a mistyped extension of a C9ORF72 repeat in a 2204 repeat region on one haplotype using simulated data. The repeat region 2202 on another haplotype is not extended. As expected, this pile-up also shows that all but one of the reads placed on the haplotype with the longer repeat also match the shorter haplotype. Only one poorly aligned read supports the extension. In practice, this would likely be considered a false-positive call caused by a single low-quality read.

[0292] Example 5: Underestimated iteration size To generate an example of a false-negative repeat extension call, the researchers simulated an FMR1 repeat with genotype 15 / 55 and then forced the generation of a read pileup corresponding to (incorrect) genotype 15 / 30. Figure 23 shows the FMR1 repeat pileup corresponding to the genotype, where the size of the longest allele at 2304 is underestimated, while the size of the shortest allele at 2302 is correct. To match the reads that occur within a repeat of size 55, some embodiments clipped the ends of the alignment to the size of the longest allele. Because there are an excess of reads overlapping with repeats having 30 motifs, and all of these reads consist of repeat sequences, the researchers can conclude that the repeat size is likely underestimated.

[0293] This disclosure may be embodied in other specific forms without departing from its spirit or essential features. The embodiments described are to be considered in all respects only and not limiting. Accordingly, the scope of this disclosure is indicated by the appended claims rather than by the foregoing description. All modifications included in the meaning of the claims and equivalents are encompassed within those scopes.

Claims

1. A system for generating computer graphics representing sequence reads aligned to haplotypes of genomic regions, wherein the system includes one or more processors and system memory, and the one or more processors (a) Aligning a plurality of sequence reads to a set of alignment positions on a plurality of haplotype sequences corresponding to a plurality of haplotypes in the genomic region, wherein the plurality of sequence reads are obtained from the genomic region of a nucleic acid sample, and each alignment position in the set of alignment positions represents a possible alignment position of one of the sequence reads on the plurality of haplotype sequences, (b) Estimating an alignment score for the set of alignment positions, wherein the alignment score indicates how uniformly the plurality of sequence reads are distributed on the plurality of haplotype sequences, (c) Repeating (a) to (b) over multiple iterations to obtain multiple sort scores for multiple different sets of sort positions, (d) Selecting a set of alignment positions from a plurality of different sets of alignment positions based on the plurality of alignment scores, wherein the selection is based on the alignment scores associated with the selected set of alignment positions that satisfy the criteria for uniformity of read distribution on the plurality of haplotype sequences, (e) A system configured to generate computer graphics representing the plurality of sequence reads and the plurality of haplotypes, wherein the plurality of sequence reads are aligned to the plurality of haplotypes in the set of alignment positions selected in (d).

2. The system according to claim 1, wherein the set of multiple different alignment positions includes at least 1,000 sets of alignment positions.

3. The system according to claim 1 or 2, wherein the genomic region includes one or more tandem repeats.

4. The system according to any one of claims 1 to 3, wherein at least one of the plurality of haplotypes includes repeat extension.

5. The system according to any one of claims 1 to 4, wherein each haplotype includes an allele.

6. The system according to any one of claims 1 to 5, wherein the plurality of haplotypes include two haplotypes.

7. The system according to any one of claims 1 to 6, wherein the selected set of alignment positions has the best alignment score among the set of multiple different alignment positions.

8. The system according to any one of claims 1 to 7, wherein the selected set of alignment positions has an alignment score that exceeds the selection criteria.

9. The system according to any one of claims 1 to 8, wherein at least one of the plurality of haplotypes includes a structural variant.

10. The system according to claim 9, wherein the structural variant is longer than 50 bp and is selected from the group consisting of deletions, duplications, copy number variants, insertions, inversions, translocations, and any combination thereof.

11. The system according to claim 9, wherein the structural variant includes a variant shorter than 50 bp.

12. The system according to claim 11, wherein the variant shorter than 50 bp includes a single nucleotide polymorphism (SNP).

13. (a) is, (i) Determining possible alignment positions for each read for each haplotype, wherein the plurality of sequence reads include read pairs obtained by pair-end sequencing, (ii) (A) Create a constrained alignment position for each lead pair from the alignment position of the constituent leads such that both leads of the lead pair are aligned on the same haplotype, and (B) the corresponding fragment length of the lead pair is as close as possible to the average fragment length, (iii) The system according to any one of claims 1 to 12, further comprising randomly selecting the alignment position of each lead pair from the constraint alignment position.

14. The system according to any one of claims 1 to 13, wherein the alignment score includes the root mean square difference from the mean of the distances between the starting positions of two consecutive leads.

15. The system according to any one of claims 1 to 14, wherein the alignment score is estimated using a probabilistic model that assumes read pairs are uniformly distributed on the plurality of haplotype sequences.

16. The system according to any one of claims 1 to 15, wherein one or more processors are configured to align a first number of sequence reads into one or more sequence graphs corresponding to the genomic regions before operation (a) to obtain the plurality of sequence reads and / or the plurality of haplotypes.

17. A method for generating computer graphics, which is implemented using a computer including one or more processors and system memory, wherein the method is (a) Aligning a plurality of sequence reads to a set of alignment positions on a plurality of haplotype sequences corresponding to a plurality of haplotypes in a genomic region using one or more processors, wherein the plurality of sequence reads are obtained from a genomic region of a nucleic acid sample, and each alignment position in the set of alignment positions represents a possible alignment position of one of the sequence reads on the plurality of haplotype sequences, (b) Estimating an alignment score for the set of alignment positions using one or more processors, wherein the alignment score indicates how uniformly the plurality of sequence reads are distributed on the plurality of haplotype sequences, (c) Repeating (a) to (b) over multiple iterations to obtain multiple sort scores for multiple different sets of sort positions, (d) One or more processors select a set of alignment positions from a plurality of different sets of alignment positions based on the plurality of alignment scores, wherein the selection is based on the alignment scores associated with the selected set of alignment positions that satisfy the criteria for uniformity of read distribution on the plurality of haplotype sequences, (e) A method comprising generating computer graphics representing the plurality of array reads and the plurality of haplotypes using one or more processors, wherein the plurality of array reads are aligned to the plurality of haplotypes in the set of alignment positions selected in (d).

18. The method according to claim 17, wherein the set of multiple different alignment positions includes at least 1,000 sets of alignment positions.

19. When executed by one or more processors of a computer system, the computer system (a) Aligning multiple sequence reads to a set of alignment positions on multiple haplotype sequences corresponding to multiple haplotypes in a genomic region, wherein the multiple sequence reads are obtained from a genomic region of a nucleic acid sample, and each alignment position in the set of alignment positions represents a possible alignment position of one of the multiple sequence reads on the multiple haplotype sequences, (b) Estimating an alignment score for the set of alignment positions, wherein the alignment score indicates how uniformly the plurality of sequence reads are distributed on the plurality of haplotype sequences, (c) Repeating (a) to (b) over multiple iterations to obtain multiple sort scores for multiple different sets of sort positions, (d) Selecting a set of alignment positions from a plurality of different sets of alignment positions based on the plurality of alignment scores, wherein the selection is based on the alignment scores associated with the selected set of alignment positions that satisfy the criteria for uniformity of read distribution on the plurality of haplotype sequences, (e) A computer-readable storage medium storing computer executable instructions for causing the plurality of array reads and the plurality of haplotypes to be generated, wherein the plurality of array reads are aligned to the plurality of haplotypes in the set of alignment positions selected in (d).

20. A method for aligning sequence reads to a genomic region, which is implemented using a system including one or more processors and system memory, wherein the method is (a) Aligning a plurality of sequence reads to a set of alignment positions on a plurality of haplotype sequences corresponding to a plurality of haplotypes in the genomic region, wherein the plurality of sequence reads are obtained from the genomic region of a nucleic acid sample, and each alignment position in the set of alignment positions represents a possible alignment position of one of the sequence reads on the plurality of haplotype sequences, (b) Estimating an alignment score for the set of alignment positions, wherein the alignment score indicates how uniformly the plurality of sequence reads are distributed on the plurality of haplotype sequences, (c) Repeating (a) to (b) over multiple iterations to obtain multiple sort scores for multiple different sets of sort positions, (d) Selecting a set of alignment positions from a plurality of different sets of alignment positions based on the plurality of alignment scores, wherein the selection is based on the alignment scores associated with the selected set of alignment positions that satisfy the criteria for uniformity of read distribution on the plurality of haplotype sequences, A method comprising determining the final alignment positions of the plurality of sequence reads such that they are the set of alignment positions selected in (e) and (d).

Citation Information

Patent Citations

  • Method and system for determining haplotypes and haplotype phasing.

    JP2015522289A

  • Methods, systems and processes for de novo assembly of sequencing reads

    JP2018500625A

  • Detecting repeat expansions with short read sequencing data

    US20170249421A1

  • Sequence-graph based tool for determining variation in short tandem repeat regions

    US20200286586A1

  • Detecting repeat expansions with short read sequencing data

    US20200335178A1