Systems and methods for tandem repeat mapping
By generating a graph for sequence reads and mapping using the longest path, the problem in the prior art is solved that it is difficult to accurately map sequence reads to genomic regions containing tandem repeats in the art, and accurate identification of genomic region status and effective diagnosis of disease are achieved.
Patent Information
- Application Number
- CN202380072115.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-22
- Filing Date
- 2023-09-22
- Publication Date
- 2025-05-16
AI Technical Summary
It is difficult for prior art to accurately map sequence reads to genomic regions containing tandem repeats, especially if the genomic regions have undergone a certain degree of genomic amplification.
A system and method are provided to generate corresponding graphs for mapping by obtaining multiple sequence reads and repeat sequence definitions for genomic regions. The method includes segmenting the sequence reads using repeating sequence definitions and mapping the sequence reads to the genomic region through the longest path.
Accurate mapping of genomic regions containing tandem repeats is achieved, and the status, stage, presence or absence of the disease can be determined, thereby providing a basis for subsequent treatment.
Smart Images

Figure CN120019440A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 376,733, entitled “SYSTEMS AND METHODS FOR TANK REPEAT SEQUENCE MAPPING,” filed on September 22, 2022, which is hereby incorporated by reference in its entirety for all purposes. Background Art
[0003] It is well known that sequencing long stretches of repetitive nucleotides is difficult, but it is clinically significant because the length and structure of the repeat region are diagnostic markers associated with a variety of serious human diseases (La Spada and Taylor, 2010, "Repeat expansion disease: Progress and puzzles in disease pathogenesis," Nature Reviews Genetics 11(4), pp. 247-258; Lopez et al., 2010 "Repeat instability as the basis for human diseases and as a potential target for therapy," Nature Reviews Molecular Cell Biology 11(3), pp. 165-170), each of which is hereby incorporated by reference into this article. Sequence reads of genomic regions containing tandem repeat sequences are particularly difficult to map back to the corresponding genomic regions because such regions vary greatly between different organisms. For example, it is well known that such regions undergo repeat expansion, in which the short tandem repeat sequences in such genomic regions in some organisms of a given species are increased (amplified) relative to other organisms. Because short tandem repeat sequences are unstable when they expand beyond a certain length, this expansion is also called dynamic mutation. Figure 4 As shown in the journal Nature, there are more than 1 million tandem repeats in the human genome. In addition, tandem repeats have been linked to changes in gene expression, genomic instability in cancer, more than 50 neurological diseases including amyotrophic lateral sclerosis (ALS), fragile X syndrome (FXS) and ataxia, and autism spectrum disorders.
[0004] Tandem repeat diseases (TRDs) include a class of neuropathological diseases associated with the accumulation of short tandem repeats (STRs; repetitive DNA sequences of 2 to 6 base pairs in length). TRD occurs with an expansion of the number of STRs from a normal state to a pathological state, with the number varying depending on the disease. TRDs are responsible for more than 20 heritable neuropathological diseases, including Huntington's disease, Kennedy's disease, myotonic dystrophy, fragile X syndrome, and several spinocerebellar ataxias. See Ellegren, 2004, "Microsatellites: simple sequences with complex evolution: Nat Rev. Genet. 5:435-445, which is incorporated herein by reference.
[0005] In addition, different amplification states (number of repeat sequences) in these regions can be associated with different states of such diseases. However, it is difficult to identify the amplification state of genomic repeat sequences using sequence reads derived from such genomic repeat sequences because there are a large number of different ways to locate sequence reads to genomic regions containing tandem repeat sequences, especially when the genomic region has undergone a certain degree of genomic amplification. In fact, such genomic regions with repeat sequences can exceed 1000 base pairs in length, resulting in an exponential increase in the number of possible methods for mapping sequence reads to such regions. Figure 5 As shown, tandem repeats in the human genome are disproportionately represented in known human genomic variation.
[0006] Therefore, what are needed in the art are systems and methods that can accurately map sequence reads to genomic regions containing tandem repeat sequences. Summary of the invention
[0007] The present disclosure provides, among other things, systems, computer-readable media, methods, and computer-executable programs for mapping multiple sequence reads to genomic regions containing tandem repeats. These systems, computer-readable media, methods, and computer-executable programs are particularly useful for determining the state, stage, presence, or absence of any of the above-mentioned diseases. For individuals found to suffer from such diseases by the systems, computer-readable media, methods, and computer-executable programs of the present disclosure, treatment for the disease may then be provided.
[0008] Using a repeat sequence definition. In some embodiments, a method for mapping a plurality of sequence reads to a genomic region is provided. In some embodiments, the method includes obtaining a plurality of sequence reads mapped to a genomic region in electronic form.
[0009] In some embodiments, the plurality of sequence reads has an average length of at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, or 2000 residues. In some embodiments, the plurality of sequence reads comprises 1000, 2000, 5000, 10,000 sequence reads, 20,000 sequence reads, 50,000 sequence reads, 100,000 sequence reads, or 1 x 10 6 sequence reads.
[0010] In some embodiments, the plurality of sequence reads are generated in a single molecule sequencing-by-synthesis reaction. In some embodiments, the single molecule sequencing-by-synthesis reaction is a single molecule real-time (SMRT) sequencing reaction.
[0011] In some embodiments, a repetitive sequence definition for a genomic region is obtained. In these embodiments, the repetitive region includes at least: (i) a first region including a first variable number of repeats of a first repetitive sequence; (ii) a second region including a second variable number of repeats of a second repetitive sequence; and (iii) a fixed interrupt sequence located between the first region and the second region.
[0012] In some embodiments, the repeat sequence definition specifies that the first repeat sequence is repeated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times, and the second repeat sequence is repeated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times. In some embodiments, the repeat sequence definition specifies that the first repeat sequence is repeated at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 times, and the second repeat sequence is repeated at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 times.
[0013] In some embodiments, the first repeat sequence has a length of between 2 and 100 residues, the fixed break sequence has a length of between 2 and 100 residues, and the second repeat sequence has a length of between 2 and 100 residues.
[0014] In some embodiments, for each corresponding sequence in a plurality of sequences, a program is executed, the program including using a repetitive sequence definition to generate a corresponding graph for the corresponding sequence read. The corresponding graph includes a corresponding plurality of nodes and a corresponding plurality of edges. The graph is scanned from the first end to the second end to completely match each of the corresponding plurality of motifs in the repetitive sequence definition. Each of the nodes represents a motif in the plurality of motifs. The plurality of motifs include at least a first instance of a first repetitive sequence, a first instance of a second repetitive sequence, an instance of a fixed interrupt sequence, and a second instance of a first repetitive sequence or a second repetitive sequence. In the plurality of motifs observed to be continuous in the corresponding sequence read, each of the plurality of edges connects the corresponding node of the first motif and the corresponding node of the second motif. The corresponding graph has one or more branch points. The program further includes identifying the longest path through the corresponding graph as a candidate segmentation for the corresponding sequence read. In the program, the corresponding sequence read is mapped to a genomic region using the longest path in the corresponding graph.
[0015] In some embodiments, mapping using the longest path includes generating a corresponding plurality of segmentations according to the longest path and the repetitive sequence definition, selecting a corresponding first segmentation having the best score in the corresponding plurality of segmentations as the segmentation for the corresponding sequence read, and mapping the corresponding sequence read to the genomic region using the corresponding first segmentation. In some embodiments, the corresponding plurality of segmentations includes 100, 500, 1000, 2000, 3000, 4000, 5000, 10000, 100000, or 1 x 10 6 A different split.
[0016] Another aspect of the present disclosure provides a system for mapping multiple sequence reads to genomic regions. The system includes a memory, an input / output terminal, and a processor coupled to the memory. The system is configured to perform a method, which includes obtaining multiple sequence reads in electronic form. The method further includes obtaining a repetitive sequence definition for a genomic region. The repetitive region includes at least: (i) a first region, which includes a first variable number of repetitive sequences of a first repetitive sequence, (ii) a second region, which includes a second variable number of repetitive sequences of a second repetitive sequence, and (iii) a fixed interrupt sequence between the first region and the second region. The method further includes, for each corresponding sequence in a plurality of sequences, executing a program, the program including using the repetitive sequence definition to generate a corresponding graph for the corresponding sequence read. The corresponding graph includes a corresponding plurality of nodes and a corresponding plurality of edges. The corresponding graph is constructed by scanning the corresponding sequence read from the first end to the second end to fully match each of the corresponding plurality of motifs in the repetitive sequence definition. Each of the respective nodes represents a motif in the plurality of motifs. The plurality of motifs include at least a first instance of a first repeat sequence, a first instance of a second repeat sequence, an instance of a fixed interrupt sequence, and a second instance of the first repeat sequence or the second repeat sequence. In the plurality of motifs observed to be continuous in the corresponding sequence reads, each of the plurality of edges connects a corresponding node of the first motif and a corresponding node of the second motif. The corresponding graph has one or more branch points. The program further includes identifying the longest path through the corresponding graph as a candidate split for the corresponding sequence read. The program further includes mapping the corresponding sequence read to a genomic region using the longest path in the corresponding graph.
[0017] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium. A non-transitory computer-readable storage medium stores instructions that, when executed by a computer system, cause the computer system to perform a method for mapping multiple sequence reads to a genomic region. The method includes obtaining multiple sequence reads in electronic form. The method further includes obtaining a repetitive sequence definition for a genomic region. The repetitive region includes at least: (i) a first region including a first variable number of repetitive sequences of a first repetitive sequence, (ii) a second region including a second variable number of repetitive sequences of a second repetitive sequence, and (iii) a fixed interrupt sequence between the first region and the second region. The method further includes, for each corresponding sequence read in a plurality of sequences, executing a program. The program generates a corresponding graph for the corresponding sequence read using the repetitive sequence definition. The corresponding graph includes a corresponding plurality of nodes and a corresponding plurality of edges. The corresponding graph is generated by scanning the corresponding sequence read from the first end to the second end to fully match each of the corresponding plurality of motifs in the repetitive sequence definition. Each of the individual nodes represents a motif in the plurality of motifs. The plurality of motifs include at least a first instance of a first repeat sequence, a first instance of a second repeat sequence, an instance of a fixed interrupt sequence, and a second instance of the first repeat sequence or the second repeat sequence. In the plurality of motifs observed to be continuous in the corresponding sequence reads, each of the plurality of edges connects a corresponding node of the first motif and a corresponding node of the second motif. The corresponding graph has one or more branch points. The program further includes identifying the longest path through the corresponding graph as a candidate split for the corresponding sequence read. The program maps the corresponding sequence read to a genomic region using the longest path in the corresponding graph.
[0018] Using a Markov model. In some embodiments, a method of mapping multiple sequence reads to a genomic region is provided, the method using a computer system including one or more processors and a system memory. In some embodiments, the genomic region has a length between 200 and 5000 residues, between 1000 and 8000 residues, or between 2000 and 10,000 residues. In some embodiments, the method includes obtaining multiple sequence reads mapped to the genomic region in electronic form.
[0019] In some embodiments, the average length of the plurality of sequence reads is at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, or 2000 residues. In some embodiments, the plurality of sequence reads comprises 1000, 2000, 5000, or 10,000 sequence reads.
[0020] In some embodiments, the plurality of sequence reads are generated in a single molecule sequencing-by-synthesis reaction. In some embodiments, the single molecule sequencing-by-synthesis reaction is a single molecule real-time (SMRT) sequencing reaction.
[0021] In some embodiments, the method includes obtaining an initial Markov model for a genomic region. The initial Markov model includes at least: (i) a first repeat sequence for a first repeat region, (ii) a second repeat sequence for a second repeat region, and (iii) an intermediate region connecting the first repeat sequence to the second repeat sequence. In some embodiments, the first region includes one or more instances of a first repeat sequence, the first repeat sequence having a length between 2 and 100 residues, the intermediate region having a length between 2 and 100 residues, and the second region includes one or more instances of a second repeat sequence, the second repeat sequence having a length between 2 and 100 residues. In some embodiments, the first region further includes one or more residues in addition to the first repeat sequence, and the second region further includes one or more residues in addition to the second repeat sequence.
[0022] In some embodiments, the method includes optimizing the initial Markov model using multiple sequence reads to obtain an optimized Markov model. In some embodiments, for each corresponding sequence read in a plurality of sequences, the method includes executing a program. The program uses the corresponding sequence reads to find the highest probability path through the Markov model. The program then maps the corresponding sequence reads to the genomic region using the highest probability path. In some embodiments, the mapping includes generating a corresponding plurality of segmentations, each segmentation being an arrangement of the highest probability path, selecting a corresponding first segmentation with the best score in the corresponding plurality of segmentations as a segmentation for the corresponding sequence read, and mapping the corresponding sequence reads to the genomic region using the corresponding first segmentation. In some embodiments, the corresponding plurality of segmentations include 100, 500, 1000, 2000, 3000, 4000, 5000, 10000, 100000 or 1x 10 6 A different split.
[0023] Another aspect of the present disclosure provides a system for mapping multiple sequence reads to genomic regions. The system includes a memory, an input / output terminal, and a processor coupled to the memory. The system is configured to execute a method. The method includes obtaining multiple sequence reads in electronic form. The method further obtains an initial Markov model for the genomic region. The initial Markov model includes at least: (i) a first repeating sequence for a first repeating region, (ii) a second repeating sequence for a second repeating region, and (iii) an intermediate region connecting the first repeating sequence to the second repeating sequence. The method optimizes the initial Markov model using multiple sequence reads to obtain an optimized Markov model. For each corresponding sequence read in a plurality of sequences, the method executes a program. The program includes finding the highest probability path through the Markov model using the corresponding sequence read. The program maps the corresponding sequence read to the genomic region using the highest probability path.
[0024] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium. A non-transitory computer-readable storage medium stores instructions that, when executed by a computer system, cause the computer system to execute a method for mapping multiple sequence reads to a genomic region. The method includes obtaining multiple sequence reads in electronic form. The method further includes obtaining an initial Markov model for a genomic region. The initial Markov model includes at least: (i) a first repeating sequence for a first repeating region, (ii) a second repeating sequence for a second repeating region, and (iii) connecting the first repeating sequence to an intermediate region of the second repeating sequence. The method includes optimizing the initial Markov model using multiple sequence reads to obtain an optimized Markov model. The method further includes executing a program for each corresponding sequence read in a plurality of sequences. The program includes finding the highest probability path through a Markov model using the corresponding sequence reads. The program further includes mapping the corresponding sequence reads to a genomic region using the highest probability path. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A system for mapping a plurality of sequence reads to a genomic region having tandem repeat sequences according to some embodiments of the present disclosure is presented.
[0026] Figure 2A and Figure 2B A method of mapping a plurality of sequence reads to a genomic region using a repetitive sequence definition for a genomic region according to some embodiments of the present disclosure is shown, wherein optional steps are represented by dashed boxes.
[0027] Figure 3A and Figure 3BA method of mapping a plurality of sequence reads to a genomic region using a Markov model for the genomic region according to some embodiments of the present disclosure is shown, wherein optional steps are represented by dashed boxes.
[0028] Figure 4 A genomic region with a tandem repeat motif flanked by flanking regions is shown.
[0029] Figure 5 It shows that although tandem repeats account for less than 4% of the human genome, a disproportionate number of variants occur in genomic regions with tandem repeats.
[0030] Figure 6 shows how functional variation in tandemly repeated genomic regions is complex, resulting in high variability in size of alleles in such regions.
[0031] Figure 7 demonstrated how many genomic tandem repeat regions are of high structural complexity, general indel detection tools are insufficient for tandem repeat analysis, and accurate tandem repeat analysis requires new bioinformatics tools.
[0032] Figure 8 Bioinformatics tools for analyzing tandem repeat regions of the genome are summarized, including tandem repeat genotyping tools, tandem repeat visualization tools, and a genome-wide tandem repeat sequence catalog according to some embodiments of the present disclosure, which has annotations for tandem repeat sequences, including population distribution of size and methylation status.
[0033] Fig. 9 and Fig.10 It is demonstrated that according to an embodiment of the present disclosure, a repeat sequence definition is used for a genotype region having tandem repeat sequences to assist in genotyping sequence reads mapped to the genotype region.
[0034] Fig.11 Sequence reads mapped to the HTT gene containing tandem repeat sequences are shown using the systems and methods of the present disclosure.
[0035] Fig.12 It is demonstrated that according to an embodiment of the present disclosure, an initial segmentation is determined for an input sequence mapped to a genomic region having tandem repeat sequences according to a definition of a repeat sequence in the genomic region.
[0036] Fig.13A , 13B, 13C, 13D and 13E show using a repetitive sequence definition for a genomic region, by scanning the corresponding sequence read from a first end to a second end to fully match each of the corresponding multiple motifs in the repetitive sequence definition, generating a corresponding graph for the corresponding sequence read to be mapped to the genomic region, the corresponding graph comprising a corresponding plurality of nodes and a corresponding plurality of edges, wherein each of the corresponding plurality of nodes represents a motif in the multiple motifs, the multiple motifs comprising at least a first instance of a first repetitive sequence, a first instance of a second repetitive sequence, an instance of a fixed interrupt sequence, and a second instance of a first repetitive sequence or a second repetitive sequence, each of the multiple edges connecting a corresponding node of an instance of a first motif in the multiple motifs observed to be continuous in the corresponding sequence read and a corresponding node of an instance of a second motif, and the corresponding graph having one or more branch points; (ii) determining the longest path through the corresponding graph as a candidate split of the corresponding sequence read; and (iii) mapping the corresponding sequence read to the genomic region using the longest path in the corresponding graph.
[0037] Fig.14 It is demonstrated that dynamic programming is used to find suitable splits for sequence reads according to an embodiment of the present disclosure.
[0038] Fig.15 Sequence reads mapped to a copy of the FMR1 gene having 31 copies of CGG repeats are shown using the systems and methods of the present disclosure.
[0039] Fig.16 Sequence reads that have been mapped to a CNBP gene copy with three adjacent repeat sequences using the systems and methods of the present disclosure are shown.
[0040] Fig.17 Sequence reads mapped to RFC1 gene copies having a non-reference AAGAG motif using the systems and methods of the present disclosure are shown.
[0041] Fig.18 Embodiments according to the present disclosure are presented that use Mendelian consistency as a measure of accuracy.
[0042] Fig.19 It is demonstrated how the types of repeat sequences generated using the disclosed systems and methods have a high Mendelian consistency.
[0043] Fig. 20 demonstrated how, in a given genomic region with repetitive sequences, polymorphic tandem repeats can have a wide range of repeat lengths.
[0044] Fig.21 showed that the methylation profile of genomic regions harboring tandem repeats is broadly similar to that of the rest of the human genome.
[0045] Fig. 22 demonstrated that methylation in genomic regions with tandem repeats can exhibit a bimodal methylation pattern.
[0046] Fig.23 It is demonstrated how the use of the systems and methods of the present disclosure discovered a methylated mosaic FMR1 expansion between 386 and 519 CGGs, an ATXN8 expansion spanning 577 CTGs, and seven biallelic RFC1 repeat expansions with 186 to 1647 AAGGGs.
[0047] Fig.24 The problematic KCNMB2 repeat locus is shown, which is annotated as an overlapping cluster of AT repeat sequences.
[0048] Fig.25 Shown Fig.24 The KCNMB2 repeat loci in question are composed of low-complexity motifs with identical structures ((CT) n STR, AAGAGG core sequence and (AT) n STR), where each n is an independent integer.
[0049] Fig.26 In accordance with an embodiment of the present disclosure, an initial unoptimized hidden Markov model is used to define the KCNMB2 repeat locus, the model comprising (i) a first repeat sequence of a first repeat region (CT repeat sequence); (ii) a second repeat sequence of a second repeat region (AT repeat sequence); and (iii) an intermediate region connecting the first repeat sequence to the second repeat sequence (VNTR core sequence).
[0050] Fig. 27 Demonstrates how the disclosed systems and methods can be used Fig.26 The initial hidden Markov model in maps sequence reads to the KCNMB2 repeat locus.
[0051] Fig.28 shows how the KCNMB2 VNTR is modestly polymorphic for the samples analyzed, with an average motif length of 27 to 30 base pairs.
[0052] Fig.29 It was revealed that the expansion of repeat sequences in genomic RFC1 causes cerebellar ataxia, neuropathy, and vestibular areflexia syndrome.
[0053] Fig.30 Defining the RFC1 repeat locus using an initial unoptimized hidden Markov model according to an embodiment of the present disclosure is demonstrated.
[0054] Fig.31 ,32 , 33 and 34 show how the systems and methods of the present disclosure can be used Fig.30 The initial hidden Markov model in maps sequence reads to the RFC1 repeat locus.
[0055] Fig.35 It shows how the AAAAG motif is the most common RFC1 motif in the aligned sequence reads.
[0056] Fig.36 It shows how the AAAGGG motif is the second most common RFC1 motif in the aligned sequence reads but is present in small numbers in most alleles.
[0057] Fig.37 A command line interface for the alignment and visualization tool of the present disclosure is presented.
[0058] Fig.38 and 39 It shows how VCF describes allele sequences and tandem repeat sequences contained therein according to an embodiment of the present disclosure.
[0059] Fig.40 It is shown how the genotype field contains the coordinates of the haplotype length and tandem repeat sequence according to some embodiments of the present disclosure.
[0060] Fig.41A It is shown how the Allele Length (AL) field contains the length of each repeat allele according to some embodiments of the present disclosure.
[0061] Fig.41B and 41C It is shown how the Motif Span (FS) field contains the span of each tandem repeat sequence on each allele according to some embodiments of the present disclosure. DETAILED DESCRIPTION
[0062] The current disclosure provides, in particular, an improved process for mapping sequence reads to genomic regions with tandem repeats. In a first approach, each sequence read is segmented according to a definition of a repeat sequence for a genomic region. That is, for each corresponding sequence read, a segmentation is constructed using the sequence of the corresponding sequence read and the definition of a repeat sequence for a genomic region. In this way, each sequence read receives its own segmentation. Each such segmentation is optimized against the sequence of its corresponding sequence read, thereby enabling the mapping of sequence reads to genomic regions. For more complex genomic regions, an initial Markov model of the genomic region is defined, and then optimized against multiple sequences. The Markov model is used to provide segmentation for each corresponding sequence read in multiple sequence reads based on the sequence of the corresponding sequence read. Each such segmentation is optimized against the sequence of its corresponding sequence read, thereby enabling the mapping of sequence reads to genomic regions.
[0063] The disclosed system and method can accurately quantify the number of repeats at a specific genomic site. Tandem repeats (TR) are repetitive sequences consisting of two or more base pairs, which are adjacent to each other and exist in large quantities throughout the genome. Due to their repetitive properties, they are highly mutagenic and play a key role in human health and disease. See Madsen et al., 2008, "Short tandem repeats in human exons: a target for disease mutations," BMC genomics, 9, 410, incorporated herein by reference. An increase in repeat length within a certain range, usually a longer repeat sequence, may be pathogenic. It is known that more than 50 diseases are caused by the amplification of tandem repeats (TR), and further research may reveal its association with more rare diseases whose causes are not yet clear. The disclosed system and method can be used for practical applications, including accurately quantifying the number of repeats as a genomic position, identifying interrupted sequences at genomic positions, determining allele phases, and determining methylation profiles. In some embodiments, multiple tandem repeat sequence catalogs are provided so that analysis can be performed and simplified. In some embodiments, for any given target genetic region (e.g., locus), the disclosed systems and methods identify sequence reads spanning the region, assign them to haplotypes, and determine the structure of the resulting repeat alleles. In some embodiments, multiple tandem repeat sequence catalogs contain tandem repeat sequence maps of variable number of tandem repeat sequences associated with diseases such as Alzheimer's disease, autism, epilepsy, and amyotrophic lateral sclerosis (ALS). See, Ryan, 2019, "Tandem repeat disorders," Evolution, Medicine, and Public Health (1), 17; and Paulson, 2018, "Repeat expansion diseases," Handbook of clinical neurology 147, 105–123, each of which is hereby incorporated by reference.
[0064] Reference will now be made in detail to the embodiments, examples of which are shown in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that one of ordinary skill in the art can practice the present disclosure without these specific details. In other cases, well-known methods, processes, components, circuits, and networks are not described in detail in order to avoid unnecessarily obscuring aspects of the embodiments.
[0065] It should also be understood that although the terms "first", "second", etc. can be used to describe various elements in this article, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the scope of the present disclosure, the first subject can be referred to as the second subject, and similarly, the second subject can be referred to as the first subject. Although the first subject and the second subject are both subjects, these subjects are not the same subject.
[0066] The terms used in this disclosure are only used for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the description of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "one", "a kind of" and "said" are intended to include plural forms as well. It should also be understood that, as used herein, the term "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should be further understood that, when used in this specification, the terms "comprises" and / or "comprising" specify the presence of stated features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof.
[0067] When using ranges in this article to describe, for example, physical or chemical properties such as molecular weight or chemical formula, all combinations and subcombinations of ranges and specific embodiments therein are intended to be included. When referring to a number or a numerical range, the use of the term "about" means that the number or numerical range mentioned is an approximation within the experimental variability (or within the statistical experimental error), and therefore the number or numerical range can vary. The variation is usually 0% to 15%, or 0% to 10%, or 0% to 5% of the quantity or numerical range. The term "comprising" (and related terms, such as "comprise / comprises" or "having" or "including") includes those embodiments, such as "consisting of" or "essentially consisting of" the embodiment of any composition of the material, method or process of the feature composition.
[0068] definition.
[0069] As used herein, the term "about" means that dimensions, sizes, formulations, parameters, shapes, and other quantities and characteristics are not and need not be exact, but may be approximate and / or larger or smaller as desired, thereby reflecting tolerances, conversion factors, rounding, measurement errors, etc. and other factors known to those skilled in the art. In general, dimensions, sizes, formulations, parameters, shapes, or other quantities or characteristics are "about" or "approximate", whether or not expressly stated as such. Note that embodiments of very different sizes, shapes, and dimensions may employ the described arrangements.
[0070] As used herein, the term "allele" refers to a specific sequence of one or more nucleotides at a chromosomal locus.
[0071] The transition terms "comprising," "consisting essentially of," and "consisting of," when used in the appended claims, in both original and amended forms, limit the scope of the claims relative to the exclusion of unrecited additional claim elements or steps, if any, from the scope of the claims. The term "comprising" is intended to be inclusive or open-ended, and does not exclude any additional, unrecited elements, methods, steps, or materials. The term "consisting of" does not include any elements, steps, or materials other than those specified in the claim, and in the latter case, does not include impurities normally associated with the specified material. The term "consisting essentially of" limits the scope of the claim to the specified elements, steps, or materials, and those elements, steps, or materials that do not materially affect the basic and novel characteristics of the claimed invention. Alternatively, all embodiments of the present invention may be more specifically defined by any of the transition terms "comprising," "consisting essentially of," and "consisting of."
[0072] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [a stated condition or event] is detected" may be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [a stated condition or event]" or "in response to detecting [a stated condition or event]," depending on the context.
[0073] As described herein, the term "locus" or "site" refers to a position in a genome, for example, on a specific chromosome and / or a position with a specific orientation. In some embodiments, a locus refers to a residue, a sequence tag, or a position of a segment on a reference sequence. In some embodiments, a locus refers to a single nucleotide position in a genome, for example, on a specific chromosome. In some embodiments, a locus refers to a small group of nucleotide positions in a genome, for example, defined by a mutation (e.g., substitution, insertion, or deletion) of consecutive nucleotides in a cancer genome. Since normal mammalian cells have a diploid genome, a normal mammalian genome (e.g., a human genome) typically has two copies at each locus in the genome, or at least two copies of each locus are located on autosomes, for example, one copy on a maternal autosome, and one copy on a paternal autosome.
[0074] As used herein, the term "mapping" refers to assigning a read sequence to a larger sequence, such as a reference genome. In some embodiments, mapping is performed by alignment. For example, mapping of a sequence read to a reference genome determines the locus in the reference genome that best matches the sequence of the sequence read.
[0075] As described herein, the term "nucleotide" can be used to refer to a natural nucleotide or an analog thereof. Examples include, but are not limited to, nucleotide triphosphates (NTPs), such as ribonucleotide triphosphates (rNTPs), deoxyribonucleotide triphosphates (dNTPs) or non-natural analogs thereof, such as dioxy nucleotide methylphosphates (ddNTPs) or reversibly terminated nucleotide triphosphates (rtNTPs).
[0076] As used interchangeably herein, the terms "polynucleotide," "nucleic acid," and "nucleic acid molecule" refer to a covalent sequence of nucleotides (e.g., ribonucleotides for RNA and deoxynucleotides for DNA) in which the 3' position of the pentose of one nucleotide is linked to the 5' position of the pentose of the next nucleotide via a phosphodiester group. In some embodiments, nucleotides include sequences of any form of nucleic acid, including but not limited to RNA and DNA molecules, such as cell-free DNA (cfDNA) molecules. The term "polynucleotide" includes but is not limited to single-stranded polynucleotides.
[0077] As used herein, the term "repetitive sequence" refers to a longer nucleic acid sequence that contains repeated occurrences of a shorter sequence. The shorter sequence is referred to herein as a "repeat unit." The repeated occurrences of a repeat unit are referred to as a "count," "repeat," or "copy" of the repeat unit. In many cases, repeat sequences are associated with genes that encode proteins. In other cases, repeat sequences are in non-coding regions. In some embodiments, repeat units occur in repeat sequences with or without breaks between repeat units. For example, in normal samples, the FMR1 gene tends to include AGG breaks in the CGG repeat sequence, such as (CGG) 5 +(AGG)+(CGG) 4. The term "tandem repeat sequence" as used herein refers to a repeat sequence in which the repeating unit is continuous. Repeating sequences without interruptions, as well as long repeating sequences with few interruptions, are prone to repeated expansion of related genes, and in some cases, as repeated expansion exceeds a specific number, it will cause genetic diseases. In different embodiments, the repeating unit comprises 2 to 100 nucleotides. Many widely studied repeating units are trinucleotide or hexanucleotide units. Other repeating units that have been studied in depth and are suitable for the embodiments disclosed herein include, but are not limited to, units consisting of 4, 5, 6, 8, 12, 33 or 42 nucleotides. See, for example, 2001, Richards, Human Molecular Genetics, 10:20, 2187-2194. The disclosed applications are not limited to the specific number of the above-mentioned nucleotide bases, as long as they are relatively short compared to the repeating sequences with multiple repeating sequences or copies of the repeating unit. For example, in some embodiments, the repeating unit includes at least 2, 3, 6, 8, 10, 15, 20, 30, 40 or 50 nucleotides. Alternatively or additionally, in some embodiments, the repeat unit includes at most about 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 6 or 3 nucleotides. In some embodiments, the repetitive sequence forms polymorphisms through evolution, development or mutagenesis conditions, thereby forming more or less copies of the same repetitive unit. Due to the unstable nature of the number of repetitive units, this process is also referred to as "dynamic mutation". Some repetitive polymorphisms have been shown to be associated with genetic diseases and pathological symptoms. Other repetitive polymorphisms are not yet well understood or studied. In some embodiments, the methods disclosed herein are used to identify previously known and new unknown repetitive polymorphisms. In some embodiments, the repetitive sequence polymorphism is longer than about 5 base pairs (bp), about 10bp, about 20bp, about 50bp, about 100bp, about 200bp, about 500bp or about 1000bp. In some embodiments, the repetitive sequence polymorphism is longer than about 1000bp, 2000bp, 3000bp, 4000bp, 5000bp or more. In some embodiments, the repetitive sequence polymorphism is no longer than about 10,000 bp, about 5000 bp, about 2000 bp, about 1000 bp, about 500 bp, about 100 bp, about 50 bp, about 20 bp, about 10 bp or less.
[0078] As used herein, the terms "sequencing", "sequencing", etc. refer to any and all biochemical processes that can be used to determine the order of biological macromolecules such as nucleic acids or proteins. For example, in some embodiments, sequencing data includes all or part of the nucleotide bases in a nucleic acid molecule such as an mRNA transcript or a genomic locus.
[0079] As described herein, the term "sequence read" or "read" refers to a sequence read from a portion of a nucleic acid sample. Typically, although not necessarily, a read represents a short sequence of consecutive bases in a sample. In some embodiments, the read is symbolically represented by a base pair sequence (in ATCG) of a sample portion. In some cases, the read is stored in a memory device and properly processed to determine whether it matches a reference sequence or meets other criteria. In some cases, the read is obtained directly from a sequencing device or not directly from the stored sequence information of the relevant sample. In some cases, the read is a sufficiently long DNA sequence (e.g., at least about 25bp), which can be used to identify a larger sequence or region, for example, it can be aligned and mapped to a chromosome or genomic region or gene. Sequence reads can be obtained in a variety of ways, for example, using sequencing technology or using probes, such as in a hybridization array or capture probe; or amplification technology, such as polymerase chain reaction (PCR) or linear amplification using a single primer, or isothermal amplification. In some embodiments, sequence reads are generated by any sequencing process described herein or known in the art. In some cases, reads are generated from one end of a nucleic acid fragment ("single-end reads") or from both ends of a nucleic acid fragment (e.g., paired-end reads, double-end reads). The length of a sequence read is often associated with a particular sequencing technology. For example, high-throughput methods can provide sequence reads that can vary in size from tens to hundreds of base pairs (bp).
[0080] In some embodiments, the sequence reads are HiFi sequence reads. HiFi reads are generated using circular consensus sequencing (CCS) mode on the PacBio long-read system. See Wenger et al., 2019, "Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome," Nature Biotechnology, 37, 1155–1162, hereby incorporated by reference herein.
[0081] As used herein, the term "subject" refers to human subjects as well as non-human subjects, such as mammals, invertebrates, vertebrates, fungi, yeast, bacteria, and viruses. Although the examples herein relate to humans and the language is primarily directed to human-related issues, the concepts disclosed herein are applicable to genomes from any plant or animal and are useful in veterinary medicine, animal science, research laboratories, and the like.
[0082] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. All patents and publications mentioned herein are incorporated by reference in their entirety.
[0083] Exemplary system embodiments.
[0084] Now that an overview of certain aspects of the disclosure has been provided, as well as certain definitions used in the disclosure, Figure 1 In conjunction with details describing an exemplary system, a computer system 100 for mapping a plurality of sequence reads to genomic regions is presented.
[0085] See also Figure 1 In a typical embodiment, computer system 100 includes one or more computers. Figure 1 For purposes of illustration, computer system 100 is represented as a single computer that includes all of the functionality of the disclosed computer system 100. However, the present disclosure is not limited thereto. The functionality of computer system 100 may be distributed across any number of networked computers and / or reside on each of a number of networked computers and / or virtual machines. Those skilled in the art will appreciate that a variety of different computer topologies are possible for computer system 100, and all such topologies are within the scope of the present disclosure.
[0086] Considering the above content, Figure 1 , the computer system 100 includes one or more processors (CPU) 59, a network or other communication interface 84, a user interface 78 (e.g., including an optional display 82 and an optional keyboard 80 or other form of input device), a memory 92 (e.g., random access memory, persistent memory, or a combination thereof), one or more disk storage and / or persistent devices 90 optionally accessed by one or more controllers 88, one or more communication buses 12 for interconnecting the above components, and a power supply 79 for powering the above components. To the extent that components of the memory 92 are not persistent, the data in the memory 92 can be seamlessly shared with the non-volatile memory 90 or the non-volatile / persistent portion of the memory 92 using known computing techniques such as caching. The memory 92 and / or storage 90 may include a large capacity memory remotely located relative to the central processor 59. In other words, some of the data stored in the memory 92 and / or storage 90 may actually be hosted on a computer external to the computer system 100, but may be electronically accessed by the computer system 100 using the network interface 84 over the Internet, an intranet, or other form of network or electronic cable. In some embodiments, computer system 100 utilizes models running from memory associated with one or more graphics processors to improve system speed and performance. In some alternative embodiments, computer system 100 utilizes models running from memory 92 rather than memory associated with a graphics processor.
[0087] The memory 92 of the computer system 100 stores:
[0088] An optional operating system 100, including programs that handle various basic system services;
[0089] An alignment module 101 for mapping multiple sequence reads to genomic regions;
[0090] Data 102 for a plurality of sequence reads 102, comprising, for each corresponding sequence read 104 (e.g., 104-1, ..., 104-M, where M is a positive integer greater than or equal to 3): a sequence read sequence 106, an optional corresponding graph 108 (the graph comprising a corresponding plurality of nodes 110 (e.g., 110-1-1-1, ..., 110-1-1-P, where P is a positive integer)) and edges 112 (e.g., 112-1-1-1, ..., 112-1-1-Q, where Q is a positive integer), a candidate partition 114 (e.g., 114-1-1), and (to a genomic region)
[0091] Sequence read mapping 116 (e.g., 116-1-1);
[0092] a repetitive sequence definition data repository 118, which contains, for each genomic region under consideration, a repetitive sequence definition 120 (e.g., 120-1, 120-2, ..., 120-Z), each repetitive sequence definition including a corresponding plurality of motifs 122;
[0093] An initial Markov model 124 for segmenting sequence reads; and
[0094] • Optimized Markov model 126 for mapping sequence reads.
[0095] In some implementations, one or more of the above-identified data elements or modules of the computer system 100 are stored in one or more of the previously mentioned memory devices and correspond to a set of instructions for performing the above-mentioned functions. The above-identified data, modules, or programs (e.g., instruction sets) need not be implemented as separate software programs, processes, or modules, and thus various subsets of these modules may be combined or otherwise rearranged in various implementations. In some embodiments, the memory 92 and / or 90 optionally stores a subset of the above-identified modules and data structures. In addition, in some embodiments, the memory 92 and / or 90 stores additional modules and data structures not described above.
[0096] Now that a system for mapping multiple sequence reads to genomic regions has been disclosed, a method for performing such mapping will be described in detail with reference to Figures 2 and 3 discussed below.
[0097] Directed graph.
[0098] refer to Figure 2A Module 4300 in, in some embodiments, at a computer system comprising one or more processors and a system memory, a method for mapping multiple sequence reads to genomic regions is provided.
[0099] Referring to module 4302, in some embodiments, the method includes obtaining in electronic form a plurality of sequence reads mapped to a genomic region.
[0100] Referring to module 4304, in some embodiments, the plurality of sequence reads have an average length of at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600 or 2000 residues. In some embodiments, a plurality of sequence reads have a mean, median, or average length of about 5,000 bp to 50,000 bp (e.g., about 5,000 bp, about 7,500 bp, about 10,000 bp, about 12,500 bp, about 15,000 bp, about 20,000 bp, about 25,000 bp, about 30,000 bp, about 35,000 bp, about 40,000 bp, about 45,000 bp, about 50,000 bp, about 55,000 bp, about 60,000 bp, about 65,000 bp, about 70,000 bp, about 75,000 bp, or about 80,000 bp). In some embodiments, the plurality of sequence reads have a mean, median, or average length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more.
[0101] Referring to module 4306, in some embodiments, the plurality of sequence reads comprises 1000, 2000, 5000, or 10,000 sequence reads. In some embodiments, the plurality of sequence reads comprises at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, at least 5 million, at least 6 million, at least 7 million, at least 8 million, at least 9 million, or more sequence reads. In some embodiments, the plurality of sequence reads comprises at least 1 x 10 7 At least 2x 10 7 At least 3x 10 7 At least 4x 10 7 At least 5x 10 7 At least 6x 10 7At least 7x 10 7 At least 8x 10 7 At least 9x 10 7 , at least 1x 10 8 At least 2x 10 8 At least 3x 10 8 , at least 4, x, 10 8 At least 5x 10 8 At least 6x 10 8 At least 7x 10 8 At least 8x 10 8 At least 9x 10 8 , at least 1x 10 9 In some embodiments, the plurality of sequence reads comprises no more than 5 x 10 7 No more than 1x 10 7 No more than 5 x 10 6 No more than 4 x 10 6 No more than 3 x 10 6 No more than 2 x 10 6 No more than 1x 10 6 no more than 500,000, no more than 100,000, no more than 50,000, no more than 30,000, no more than 20,000, no more than 10,000, no more than 9000, no more than 8000, no more than 7000, no more than 6000, no more than 5000, no more than 4000, no more than 3000, no more than 2000, no more than 1000, or fewer sequence reads.
[0102] In some embodiments, multiple sequence reads can be obtained in a variety of ways, for example, using sequencing techniques or using probes, for example, in hybridization arrays or capture probes; or amplification techniques, such as polymerase chain reaction (PCR) or linear amplification using a single primer, or isothermal amplification.
[0103] Figure 6 Demonstrated in targeting Figure 6In the example studied, the FRM1 genomic region with an 87 base pair allele with two AGG interruptions can be up to 1200 base pairs in length. Therefore, for this and similar sized or larger genomic repetitive regions, it is desirable to have long sequence reads, for example, sequence reads with an average length of at least 1000 base pairs, such as disclosed in Rhoads, 2015, "PacBio Sequencing and Its Applications," Genomics, Proteomics & Bioinformatics 13(5), pp. 278-289, which can cover the entire genomic repetitive region, which is hereby incorporated by reference. References Figure 7 , sequence reads that contain the entire genome repetitive region are desirable because such sequence reads reduce the computational complexity of mapping to the genome repetitive region. Figure 7 As mentioned above, due to the high structural complexity of many genomic tandem repeat regions, conventional indel (insertion and deletion) analysis tools are insufficient for tandem repeat sequence analysis.
[0104] Modules 4308-4310. Referring to module 4308, in some embodiments, the plurality of sequence reads are generated in a single molecule sequencing-by-synthesis reaction. Referring to module 4310, in some embodiments, the single molecule sequencing-by-synthesis reaction is a single molecule real-time (SMRT) sequencing reaction. In some embodiments, the plurality of sequence reads are generated in a single molecule nanopore sequencing reaction. In some embodiments, the single molecule sequencing-by-synthesis reaction is a single molecule real-time (SMRT) sequencing reaction from Pacific Biosciences. Sequencing The present invention relates to sequencing a polynucleotide substrate, or sequencing a genomic fragment used in a nanopore sequencing platform (e.g., provided by Oxford Nanopore Technologies, Genia, etc.), or sequencing on any other suitable single molecule sequencing platform. In some embodiments, examples of single molecule sequencing platforms and methods that can be used to generate sequence reads used by the systems and methods of the present disclosure can be found in the following U.S. patents and U.S. patent application publications, which are hereby incorporated by reference: US8324914, US2013 / 0244340, US2015 / 0119259, US2010 / 0196203, US2011 / 0229877, US2016 / 0162634, US7315019, US2009 / 0087850, and US2018 / 0023134.
[0105] refer to Figure 2AModule 4312 in and Figure 1 In the system 100, in some embodiments, a repetitive sequence definition 120 is obtained for the genomic region. In some embodiments, the repetitive region includes at least: (i) a first region including a first variable number of repeats of a first repetitive sequence, (ii) a second region including a second variable number of repeats of a second repetitive sequence, and (iii) a fixed interrupt sequence between the first region and the second region. In some embodiments, a sequence read in the plurality of sequence reads has at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 repeats of the first repetitive sequence, followed by a fixed interrupt sequence, and then at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 repeats of the second repetitive sequence. In some embodiments, at least 5 sequence reads, at least 10 sequence reads, at least 15 sequence reads, at least 20 sequence reads, at least 50 sequence reads, at least 100 sequence reads, at least 250 sequence reads, at least 500 sequence reads, at least 1000 sequence reads, at least 5000 sequence reads in the plurality of sequence reads have at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 repeats of a first repeated sequence followed by a fixed break sequence and then at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 repeats of a second repeated sequence. Fig. 9 The definition of the repeat sequence in the genomic region is shown: (CAG) n CAACAG(CCG) n . Here each instance of "n" is the same or different positive integer. In this example, (CAG) n is a motif 122 of a repeat sequence definition 120 and is a first region including a first variable number of first repeat sequences; (CCG) nis another motif 122 of the repeat sequence definition 120, and is a second region of a second variable number of repeat sequences; and CAACAG is a fixed interrupt sequence between the first region and the second region. Therefore, in one sequence, the first instance of "n" is 2, and the second instance of "n" is 3, that is, (CAG CAG)CAACAG(CCG CCG CCG) (Seq. Id. No. 16), which is included in this particular repeat sequence definition; similarly, in another sequence, the first instance of "n" is 4, and the second instance of "n" is 2, that is, (CAG CAG CAG CAG)CAACAG(CCG CCG) (Seq. Id. No. 17). Fig. 9 The tandem repeat genotyper disclosed in Figure 1 In an embodiment of the alignment module 101 in FIG. 1 , a repeat sequence definition 120 is used to map sequence reads to genomic regions represented by the repeat sequence definition.
[0106] However, in some embodiments, the repeat sequence definition 120 has at least: (i) a first region comprising a first variable number of repeats of a first repeat sequence, (ii) a second region comprising a second variable number of repeats of a second repeat sequence, and (iii) a fixed interrupt sequence between the first region and the second region, but the present disclosure is not limited thereto. The repeat sequence definition may include more than two repeat regions, nor may it include only a single fixed interrupt sequence. In some embodiments, the repeat sequence definition 120 includes 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more motifs 122, wherein each motif 122 is either a repeat sequence or a fixed interrupt sequence between two other motifs in the repeat sequence definition. For example, a repeat definition 120 having five motifs 122 is a motif consisting of (i) a first region (motif 1) comprising a first variable number of repeats of a first repeating sequence, (ii) a second region (motif 2) comprising a second variable number of repeats of a second repeating sequence, (iii) a first fixed interrupt sequence (motif 3) located between the first region and the second region, (iv) a third region (motif 4) comprising a third variable number of repeats of a third repeating sequence, and (v) a second fixed interrupt sequence (motif 5) located between the second region and the third region. In some embodiments, the repeat definition 120 includes between 3 and 100 motifs 122.
[0107] In some embodiments, the repeat region comprises three different adjacent repeat regions without a fixed interrupting sequence. Fig.17An example is shown for the CNBP region in , which includes the corresponding adjacent CAGG, CAGA, and CA repeat regions.
[0108] In some embodiments, the repeat region comprises 3, 4, 5, 6, 7, 8 or 9 different adjacent repeat regions without fixed interruption sequences between them. In some embodiments, the repeat region comprises three different consecutive repeat regions, followed by an interruption sequence motif, and then a fourth repeat region.
[0109] Referring to module 4314, in some embodiments, the repeating sequence definition specifies that the first repeating sequence is repeated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times and the second repeating sequence is repeated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times.
[0110] Referring to module 4316, in some embodiments, the repeating sequence definition specifies that the first repeating sequence is repeated at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 times and the second repeating sequence is repeated at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 times.
[0111] Referring to module 4318, in some embodiments, the first repeat sequence has a length between 2 and 100 residues, the fixed interrupt sequence has a length between 2 and 100 residues, and the second repeat sequence has a length between 2 and 100 residues.
[0112] Referring to module 4320, in some embodiments, for each corresponding sequence read in the plurality of sequences, a program is executed to determine an appropriate form of a repetitive sequence definition for a genomic region to which the corresponding sequence read is mapped. The general method to module 4320 is described in Fig.12 Generate a set of reasonable partitions for the repetitive sequence definition 120. For example, consider the repetitive sequence definition Fig. 9 Situation shown: (CAG) n CAACAG(CCG) n Here, each instance of "n" can be the same or different positive integers. (CAG) n CAACAG(CCG) n A reasonable split sets the first instance of "n" to 2 and the second instance of "n" to 3: (CAG CAG)CAACAG(CCG CCG CCG) (Seq. Id. No. 16). (CAG) n CAACAG(CCG) nAnother reasonable split of sets the first instance of "n" to 4 and the second instance of "n" to 2: (CAG CAG CAGCAG)CAACAG(CCG CCG) (Seq. Id. No. 17). Fig.12 The input sequence of the sequence reads to be mapped to the genomic region is then scored for each possible split of the repetitive sequence definition, and the repetitive sequence definition with the highest score for the sequence read is selected as the final split for the sequence read. Fig.12 The procedure outlined in is useful for simple repetitive regions, but in practice there are so many possible partitions of the repetitive sequence definition 120 that such an approach is not computationally feasible.
[0113] In some embodiments, the method shown in FIG. 13 is used to narrow the partition search space defined for the repeated sequence. Fig.13A The problem is outlined. Sequence reads with the sequence CAGCAGCAGCAGCCGCAGCAGCAACAGCCGCCGCAGCCG (Seq.Id.No.:1) need to be aligned with the repeat sequence definition (CAG) n CAACAG(CCG) n The repeat sequence definition 120 is used to generate a corresponding graph 108 for the corresponding sequence read 104. The corresponding graph 108 includes a corresponding plurality of nodes 110 and a corresponding plurality of edges 112. Fig. 13B As shown in , to begin building the graph, the sequence 106 of the corresponding sequence read 104 is scanned from the first end to the second end to fully match each motif 122 of the corresponding plurality of motifs in the repeat sequence definition 120 . Fig. 13B The repeat sequence definition 120 consists of three motifs: CAG (122-1), CAAACAG (122-2), and CCG (122-3). Therefore, in the corresponding graph 108, each position of each of these motifs in the sequence 106 of the corresponding sequence read is used as a node 110. In other words, Fig. 13B As shown, each node 110 in the corresponding plurality of nodes represents an instance of a motif 122 in the plurality of motifs. Fig. 13B As shown, the plurality of motifs include at least a first instance of a first repeating sequence (CAG) 122-1, a first instance of a second repeating sequence (CCG) 122-3, an instance of a fixed interrupt sequence (CAACAG) 122-2, and a second instance of a first (CAG) or second (CCG) repeating sequence. Fig. 13C, among the plurality of motifs observed to be continuous in the corresponding sequence reads, each edge 112 of the plurality of edges connects a corresponding node 110 of a first motif with a corresponding node 110 of a second motif. Fig. 13C As shown, due to the high repetitiveness of the genomic repetitive region from which the sequence 106 of the sequence read 104 originates, the corresponding graph has one or more branch points. For example, the node 110-4 branches to the node 110-6 through the edge 112-4, and branches to the node 110-5 through the edge 112-5.
[0114] In some embodiments, the graph 108 is directional (e.g., from the 5' to the 3' end of the sequence 106 of the corresponding sequence read 104, or from the 3' to the 5' end of the sequence 106 of the corresponding sequence read 104). In addition, each node 110 in the plurality of nodes is connected to at least one other node in the plurality of nodes by an edge 112.
[0115] In some embodiments, the graph 108 is a directed graph. In some embodiments, the directed graph is an acyclic graph (DAG), which has both directionality and no cycles. That is, the graph consists of a finite number of nodes and edges, each edge pointing from one node to another, so that starting from any node v, it is impossible to eventually loop back to v along the sequence of edges that are always directed. Equivalently, the DAG is a directed graph with a topological order, that is, the order of the vertices is such that each edge points from the front to the back in the sequence 106 of the corresponding sequence read segment 104.
[0116] exist Fig. 13C In Figure 1, we can see that edge 112-1 is labeled with a value of "3", while edge 119-9 is labeled with a value of "15". Each of these labels, and Fig. 13C The labels used for other edges in , all indicate the position of the relative starting point of the target node in sequence 106 relative to the starting point of the source node in sequence 106 in terms of nucleotides. For example, in the case of edge 112-1, the source node is node 110-1, and the target node is 110-2. The label "3" on edge 112-1 between these two nodes indicates that, in the sequence 106 of the corresponding sequence read 104, the starting position of motif 122 of target node 110-2 is offset by three residues from the starting position of motif 122 of source node 110-1. Fig. 13C In the case of , the directed graph is in the 5' to 3' direction of the sequence 106, and therefore the label "3" on the edge 112-1 between the two nodes indicates that the starting position of the motif 122 of the target node 110-2 is three residues downstream of the starting position of the motif 122 of the source node 110-1 in the sequence 106. Therefore, according to the edge 112-1, if the motif 110-1 starts at the 1st position of the sequence 106, then the motif 110-2 starts at the 4th position of the sequence 106.
[0117] In some embodiments, for a corresponding sequence read in the plurality of sequence reads, there are 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 1000, 10,000, or 1 x 10 6 In some embodiments, for each corresponding sequence read in the plurality of sequence reads, there are 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 50, 100, 1000, 10,000, 1 x 10 6 or more paths.
[0118] In some embodiments, the corresponding graph for a corresponding sequence read in the plurality of sequence reads comprises 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nodes, and 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more edges. In some embodiments, the corresponding graph for each corresponding sequence read in the plurality of sequence reads comprises 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nodes, and 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more edges.
[0119] like Fig. 13C As shown, after constructing a graph 108 for the sequence 106 of the sequence read 104 using the motif 122 found in the repetitive sequence definition 120 of the genomic region to which the sequence read 104 is to be mapped, attention turns to determining which path in the graph should be used as a segmentation of the repetitive sequence definition 120 for the sequence read 104. Fig. 13C As shown, there are multiple branch points in the graph, and therefore there are multiple paths in the graph, each path representing a traversal between position 1 and position 34 of the sequence 106 of the corresponding sequence read 104. Each such path represents a potential partitioning of the repeat sequence definition 120 based on the sequence 106 of the corresponding sequence read 106. For example, since node 110-7 represents a branch point in the graph, one set of paths passes through edge 112-8, while another set of paths passes through edge 112-7. Fig.13D and Fig.13E One such path is shown in the figure. Note that this path does not pass through nodes 110-9 or 110-12. Fig.13D The path shown in Fig. 13C The longest path in the corresponding graph in , and therefore, according to Figure 2B In module 4320 in the corresponding graph, the path is designated as a candidate segmentation 114 for the corresponding sequence read 104. Then, the longest path in the corresponding graph is used to map the corresponding sequence read to a genomic region. In some embodiments, the graph includes 10 or more paths, 100 or more paths, 1000 or more paths, 10,000 or more paths, 100,000 or more paths, or 1 x 10 6 or more paths, each of which is a possible segmentation for a corresponding sequence read. Thus, in such an embodiment, the length of each of these paths is evaluated to determine which path is the longest path.
[0120] Referring to module 4322, in some embodiments, using candidate segmentation 114 (such as Fig.13E The method comprises generating a plurality of corresponding segmentations according to the longest path and the repetitive sequence definition, selecting a corresponding first segmentation with the best score in the corresponding plurality of segmentations as the segmentation for the corresponding sequence read, and mapping the corresponding sequence read to the genomic region using the corresponding first segmentation. For example, a limited number of instances of the motif 122 specified by the repetitive sequence definition 120 are added and the repetitive sequence definition is used to generate a segmentation based on the repetitive sequence definition. Fig.13E Thus, referring to module 4324, in some embodiments, the corresponding plurality of partitions to be considered based on the longest path in the corresponding graph 108 include 100, 500, 1000, 2000, 3000, 4000, 5000, 10000, 100000, 1 x 10 6 or more different partitions.
[0121] The above example shows how mapping sequence reads to repetitive regions of the genome cannot be performed by the brain alone. Fig.12The method generally outlined in will take several days of computation on a high-speed computer for a repeat sequence definition that includes at least (i) a first region that includes a first variable number of repeats of a first repeat sequence, (ii) a second region that includes a second variable number of repeats of a second repeat sequence, and (iii) a fixed interrupt sequence between the first region and the second region (e.g., where the repeat sequence definition specifies that the first repeat sequence is repeated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times, and the second repeat sequence is repeated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times). Given a repeat sequence definition 120 for a sequence 106 of a given sequence read 104, such computation will determine the best segmentation. Although the longest path through the correspondence graph 108 reduces the astronomical number of possible splits considered by the brute force approach by several orders of magnitude as shown in FIG13 , there is still a need to optimize the split given by the longest path, which results in that for each sequence read, 100, 500, 1000, 2000, 3000, 4000, 5000, 10,000, 100,000, 1 x 10 6 or more different segmentations. Each such calculation requires scoring the sequence 106 of the sequence read 104 with the sequence of the candidate segmentation to find the best score. Each such comparison requires matching the sequence 104 of the sequence read with the sequence of the candidate sequence. In some embodiments, in order to map the sequence reads to the genomic region, the segmentation method of the longest path with deletions, insertions and gaps introduced will also be considered, which adds more complexity to the mapping. In the case of typical practical applications, it is further necessary to map 10, 100, 500, 1000, 2000, 5000, 10,000 or more sequence reads to the genomic region. In some embodiments, according to module 4320, a graph 108 is constructed for each such sequence read, which further increases the complexity of the related tasks and makes it impossible to perform it mentally.
[0122] In some embodiments, it may be difficult to resolve the variation in tandem repeat (TR) regions based solely on the repetitive sequences. One example is measuring the methylation of homozygous repetitive sequences: if the repetitive sequences are homozygous, the reads and their methylation levels cannot be assigned to alleles based solely on the repetitive sequences. Another example is genotyping repetitive sequences with mosaic alleles. These alleles produce reads that support a range of repeat lengths, which makes it difficult to determine their allelic origin. In these embodiments, the alignment module 101 uses single nucleotide polymorphisms (SNPs) around the repetitive sequences to solve these problems. These flanking SNPs provide independent evidence that can assign sequence reads to alleles, thereby genotyping the repetitive sequences and determining their allele-specific methylation.
[0123] In some embodiments, for modeling purposes, each sequence read 104r spanning a repetitive sequence is associated with a vector of 1s and 0s that represents the presence or absence of each single nucleotide polymorphism covered by the sequence read. That is, if the sequence read r contains the kth SNP, then r[k]=1; otherwise, r[k]=0. The local haplotype is also defined as a vector of 0s and 1s. The genotype is a pair of local haplotypes G=(H 1 ,H 2 ). The posterior probability of genotype G given a set of observed sequence reads is estimated according to the following model for SNP genotyping:
[0124] P(G|R)~P(R|G)·P(G),
[0125] where P(R|G) is the probability of observing read R given genotype G, and P(G) is the prior probability of genotype G. In addition,
[0126]
[0127] In this P(r|H i )=∏P(k|r,H i ), where if r[k]=H i [k],P(k|r,H i )=p, otherwise P(k|r,H i)=1–p. The genotype probability P(G) can be estimated by genotyping the repeated sequences in the control cohort. This genotyping model is described in Li et al., 2009, “SNP detection for massively parallel whole-genomeresequencing,” Genome Research 19:1124-132, which is incorporated herein by reference. Using this model, in some embodiments, the alignment module 101 determines the most likely genotype G=(H 1 ,H 2 ), and each sequence read r to the corresponding H 1 or H 2 Finally, in these embodiments, the consensus sequence for each repeat allele is calculated based on the reads assigned to the corresponding local haplotype. In some embodiments, Figure 2A and Figure 2B The method shown maps a sequence read having a non-reference motif to a genomic region containing the non-reference motif. This occurs when the source subject of the sequence read has an insertion in the genomic region that is not recorded in the reference for the genomic region, or the insertion is otherwise rare such that the motif is not included in the repetitive sequence definition 120 for the genomic region. For example, Fig.17 Shows examples, based on Figure 2A and Figure 2B In the method shown, sequence reads containing a non-reference AAGAG motif are successfully mapped to the RFC1 genomic region, even though the AAGAG motif is not included in the repetitive sequence definition 120 used. In some embodiments, the plurality of sequence reads include 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more motifs that are not present in the repetitive sequence definition, wherein each such motif is between 1 residue and 20 residues in length and is repeated between 1 and 100 times in at least some of the plurality of sequence reads. In some embodiments, between 5% and 40% of the sequences of at least 10% of the sequence reads in the plurality of sequence reads are from motifs that are not present in the repetitive sequence definition used to map the sequence reads to the genomic region from which they originated.
[0128] Markov model.
[0129] Although the above combination Figure 2A and Figure 2B The described methods are applicable to a variety of genomic regions where repeat expansion has occurred, but in some embodiments, for those genomic regions where repeat expansion has occurred that are difficult to describe using the repeat sequence definition 120, the alignment module 101 uses different techniques. Figure 3A Module 4400 and Figure 1 In some embodiments, a method of mapping a plurality of sequence reads to genomic regions is provided, the method using a computer system comprising one or more processors and a system memory encoding an initial Markov model 126 .
[0130] Referring to module 4402, in some embodiments, the genomic region in which repeat expansion has occurred has a length between 200 and 5000 residues, between 1000 and 8000 residues, or between 2000 and 10,000 residues.
[0131] Reference module 4404, as described above in conjunction with Figure 2A and Figure 2B As is the case with the disclosed methods, in some embodiments, the method includes obtaining in electronic form a plurality of sequence reads mapped to a genomic region.
[0132] Referring to module 4406, in some embodiments, the plurality of sequence reads have an average length of at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600 or 2000 residues. In some embodiments, a plurality of sequence reads have a mean, median, or average length of about 5,000 bp to 50,000 bp (e.g., about 5,000 bp, about 7,500 bp, about 10,000 bp, about 12,500 bp, about 15,000 bp, about 20,000 bp, about 25,000 bp, about 30,000 bp, about 35,000 bp, about 40,000 bp, about 45,000 bp, about 50,000 bp, about 55,000 bp, about 60,000 bp, about 65,000 bp, about 70,000 bp, about 75,000 bp, or about 80,000 bp). In some embodiments, the plurality of sequence reads have a mean, median, or average length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more.
[0133] Referring to module 4408, in some embodiments, the plurality of sequence reads comprises 1000, 2000, 5000, or 10,000 sequence reads. In some embodiments, the plurality of sequence reads comprises at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 million, at least 2 million, at least 3 million, at least 4 million, at least 5 million, at least 6 million, at least 7 million, at least 8 million, at least 9 million, or more sequence reads. In some embodiments, the plurality of sequence reads comprises at least 1 x 10 7 At least 2x 10 7 At least 3x 10 7 At least 4x 10 7 At least 5x 10 7 At least 6x 10 7 At least 7x 10 7 At least 8x 10 7 At least 9x 10 7 , at least 1x 10 8 At least 2x 10 8 At least 3x 10 8 At least 4x 10 8 At least 5x 10 8 At least 6x 10 8 At least 7x 10 8 At least 8x 10 8 At least 9x 10 8 , at least 1x 10 9 In some embodiments, the plurality of sequence reads comprises no more than 5 x 10 7 No more than 1x 10 7 No more than 5 x 10 6 No more than 4 x 10 6 No more than 3 x 10 6 No more than 2 x 10 6 No more than 1x 10 6no more than 500,000, no more than 100,000, no more than 50,000, no more than 30,000, no more than 20,000, no more than 10,000, no more than 9000, no more than 8000, no more than 7000, no more than 6000, no more than 5000, no more than 4000, no more than 3000, no more than 2000, no more than 1000, or fewer sequence reads.
[0134] In some embodiments, multiple sequence reads can be obtained in a variety of ways, for example, using sequencing techniques or using probes, for example, in hybridization arrays or capture probes; or amplification techniques, such as polymerase chain reaction (PCR) or linear amplification using a single primer, or isothermal amplification.
[0135] Referring to module 4410, in some embodiments, the plurality of sequence reads are generated in a single molecule sequencing by synthesis reaction. Referring to module 4412, in some embodiments, the single molecule sequencing by synthesis reaction is a single molecule real-time (SMRT) sequencing reaction. In some embodiments, the plurality of sequence reads are generated in a single molecule nanopore sequencing reaction. In some embodiments, the single molecule sequencing by synthesis reaction is a single molecule real-time (SMRT) sequencing reaction from Pacific Biosciences. Sequencing The present invention relates to sequencing a polynucleotide substrate, or sequencing a genomic fragment used in a nanopore sequencing platform (e.g., provided by Oxford Nanopore Technologies, Genia, etc.), or sequencing on any other suitable single molecule sequencing platform. In some embodiments, examples of single molecule sequencing platforms and methods that can be used to generate sequence reads used by the systems and methods of the present disclosure can be found in the following U.S. patents and U.S. patent application publications, which are hereby incorporated by reference: US8324914, US2013 / 0244340, US2015 / 0119259, US2010 / 0196203, US2011 / 0229877, US2016 / 0162634, US7315019, US2009 / 0087850, and US2018 / 0023134.
[0136] Referring to module 4414, in some embodiments, the method includes obtaining an initial Markov model for a genomic region. In the Markov model, the distribution of nucleic acids at each position in a set of sequence reads can be used to determine the transition probabilities between states of a hidden Markov model (HMM), and then the HMM is trained. For example, a hidden Markov model is described in Schliep et al., 2003, Bioinformatics 19(1):i255-i263, which is incorporated herein by reference.
[0137] In some embodiments, regions known to undergo repeat expansion require more complex Markov models. For example, Fig.24 An exemplary sequence read that was mapped to the KCNMB2 repeat locus by conventional mapping tools is shown. The KCNMB2 repeat locus is a notoriously difficult region to map sequence reads into, such as Fig.24 As indicated by the overlapping and internally consistent reference annotations for this region shown at the bottom for the KCNMB2 repeat locus. Fig.25 As shown, the KCNMB2 repeat locus includes low-complexity motifs with the same structure ((CT)nSTR, AAGAG core and (AT)nSTR, where each n is the same or different and is a positive integer). Figure 2A and Figure 2B Unlike the genome shown, these repeat regions are not perfect. For example, in the (CT)n region, there are sequences other than CT, such as CC and AC; and in the (AT)n region, there are sequences other than AT, such as AC and AAT.
[0138] To process images Fig.24 and Fig.25In one aspect of the present disclosure, an initial Markov model 124 is provided for a genomic region in which a complex repeat expansion (such as a KCNMB2 repeat locus) has occurred, and the model includes multiple states and multiple transition attributes, encoding at least the following: (i) a first repeat sequence for a first repeat region; (ii) a second repeat sequence for a second repeat region; and (iii) an intermediate region connecting the first repeat sequence to the second repeat sequence. In some embodiments, a sequence read in the plurality of sequence reads has at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 repeats of the first repeat sequence, followed by a fixed interrupt sequence, and then at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 repeats of the second repeat sequence. In some embodiments, at least 5 sequence reads, at least 10 sequence reads, at least 15 sequence reads, at least 20 sequence reads, at least 50 sequence reads, at least 100 sequence reads, at least 250 sequence reads, at least 500 sequence reads, at least 1000 sequence reads, at least 5000 sequence reads in the plurality of sequence reads have at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 repeats of a first repeated sequence followed by a fixed break sequence and then at least 2, 3, 4, 5, 6, 7, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 repeats of a second repeated sequence.
[0139] Fig.26 Demonstrated. Fig.26 In the Fig.26 In the example of the first repeat region (CT) n, the first repeat sequence; the AT repeat sequence constitutes Fig.26 In the example of the second repeat region (AT)n, the second repeat sequence of the second repeat region (AT)n; and the VNTR core sequence constitutes the middle region connecting the first repeat sequence to the second repeat sequence. In this model, arrow 2602 will contain the probability of repeating when the CT repeat region is given a C / T base; the VNTR core sequence will encode multiple probabilities across the core to cover all possible sequences in the multiple sequences; and arrow 2604 will contain the probability of repeating when the AT repeat region is given an A / T base. Multiple sequences can be arranged according to Fig.25 As shown, the AAGAGG core sequence is used for alignment, and then these aligned sequences can be used for training Fig.26 The transition probabilities of the Markov model (eg, transitions 2602 and 2604).
[0140] Referring to module 4418, in some embodiments, the first region further includes one or more residues other than the first repeating sequence, and the second region further includes one or more residues other than the second repeating sequence. Fig.26 A possible Markov model for the KCNMB2 repeat locus is shown, but the model is shown by example to demonstrate important features of the model, such as at least two repeat transition probabilities for two different repeat regions (arrows 2602 and 2604). However, in practical applications, more complex Markov models are needed to encode rarer states. For example, in (CT) n Region, taking sequences other than CT, such as CC and AC, as (CT) of the Markov model n Part of the state, and give the necessary transition probability; and in (AT) n region, taking sequences other than AT, such as AC and AAT, as the (AT) n The state of the part and assign necessary transition probabilities.
[0141] Referring to module 4416, in some embodiments, the first region includes one or more instances of a first repeating sequence having a length between 2 and 100 residues, the middle region has a length between 2 and 100 residues, and the second region includes one or more instances of a second repeating sequence having a length between 2 and 100 residues.
[0142] Referring to module 4420, in some embodiments, the method includes optimizing the initial Markov model using multiple sequence reads to obtain an optimized Markov model. For example, as described above, the sequence reads mapped to KCNMB2 can be aligned with the AAGAGG core sequence and then used to train Fig.26 Transition probabilities for the Markov model shown.
[0143] Referring to module 4420, in some embodiments, for each corresponding sequence read in a plurality of sequences, the method includes executing a procedure including: (i) using the corresponding sequence read to find the highest probability path through a Markov model; and (ii) using the highest probability path to map the corresponding sequence read to a genomic region. Therefore, after the Markov model training is completed, the sequence 104 of the corresponding sequence read 106 is input into the Markov model, thereby obtaining the highest probability path for the corresponding sequence read 106 through the Markov model. This highest probability path represents the segmentation for the corresponding sequence read. As described above in conjunction with Figure 2A and Figure 2B As with the examples of the described methods, this segmentation is subsequently used to map the sequence reads to genomic regions.
[0144] Referring to module 4422, in some embodiments, using the highest probability path to map the corresponding sequence read to the genomic region includes generating a corresponding plurality of segmentations, each segmentation being a permutation of the highest probability path, selecting a corresponding first segmentation having the best score in the corresponding plurality of segmentations as the segmentation for the corresponding sequence read, and mapping the corresponding sequence read to the genomic region using the corresponding first segmentation. Although the highest probability path through the optimized Markov model 126 reduces the astronomical number of possible segmentations considered by the brute force approach by several orders of magnitude, there is still a situation where the segmentation given by the most likely path needs to be optimized, which results in each corresponding sequence read of the plurality of sequence reads needing to evaluate 100, 500, 1000, 2000, 3000, 4000, 5000, 10,000, 100,000, 1 x 10 6 or more different segmentations. Each such calculation requires scoring the sequence 106 of the sequence read 104 with the sequence of the candidate segmentation to find the best score. Each such comparison requires matching the sequence 106 of the sequence read with the sequence of the candidate sequence. In some embodiments, in order to map the sequence reads to the genomic regions, the segmentation method with the maximum possible path of deletions, insertions and gaps introduced is also considered, which adds more complexity to the mapping. Therefore, referring to module 4424, in some embodiments, the corresponding multiple segmentations include 100, 500, 1000, 2000, 3000, 4000, 5000, 10,000, 100,000 or 1x 10 6 Different segmentations are performed to read corresponding sequence reads in the plurality of sequence reads. In typical practical applications, 10, 100, 500, 1000, 2000, 5000, 10000 or more sequence reads need to be mapped to a specific genomic region.
[0145] Fig. 27 Shown with Fig.24 Using the same sequence reads Fig.24 Compared with the traditional mapping method, the improvement achieved when mapping the sequence to KCNMB2 according to the method disclosed in Figure 3. Fig.28 Analysis of the mapped sequences is provided.
[0146] In some embodiments, genotyped SNPs are used to resolve repetitive sequences that some Markov models cannot satisfactorily resolve using the techniques described above in conjunction with module 4322.
[0147] Example.
[0148] Example 1. Fig.15 Shows an alignment of a set of sequence reads that map to genomic locations containing a portion of the FMR1 expansion according to the FMR1 repeat definition (CAG)nCAACAG(CCG)n. Figure 2A and Figure 2B The disclosed method has successfully mapped sequence reads to a genome even though the genome contains 31 consecutive copies of the CGG motif.
[0149] Example 2. Fig.16 shows an arrangement graph of a set of sequence reads, according to Figure 2A and Figure 2B According to the disclosed method, the sequence reads are mapped to genomic locations containing CNBP expansions according to the CNBP repeat sequence definition (which includes three adjacent different repeat sequences CAGG, CAGA and CA).
[0150] Example 3. Fig.17 It is shown that even if a motif is present in a genomic region that is not included in the repetitive sequence definition 120, Figure 2A and Figure 2B The method shown is powerful enough to map sequence reads to genomic regions with repetitive sequences. Fig.17 middle, Figure 2A and Figure 2B The method shown has successfully mapped sequence reads to the RFC1 genomic region for a subject containing a non-reference AAGAG motif. That is, the AAGAG motif is not in the repetitive sequence definition 120 for RFC1.
[0151] Example 4. Fig.29 Details of another genomic region that has undergone repeat expansion that is amenable to the mapping approach described above in conjunction with Figure 3 are shown. This genomic region encodes RFC1, which is associated with cerebellar ataxia, neuropathy, vestibular anesthetic syndrome (CANVAS). Previous studies revealed a range of different possible RFC1 motifs: AAAAG, AAAGG, AAGGG, AAGAG, AGAGG, AACGG, ACGGG, and AAAGGG, with an expansion of one of the motifs, (AAGGG) n , which is associated with tardive ataxia. Fig.30 The Markov model defined for the genomic region according to the above method in combination with FIG. 3 is shown. Fig.31 , Fig.32 , Fig.33 and Fig.34 It shows how the Markov model can map multiple sequence reads from a control sample to the RFC1 region using the approach described in Figure 3. Fig.35 and Fig.36 The statistics of the genotypes represented by these mapped sequence reads are detailed. Fig.37 A command line interface for the alignment and visualization tool of the present disclosure is presented. Fig.38 and 39 It shows how VCF describes allele sequences and tandem repeat sequences contained therein according to an embodiment of the present disclosure. Fig.40 It is shown how the genotype field contains the coordinates of the haplotype length and tandem repeat sequence according to some embodiments of the present disclosure. Fig.41A It is shown how the Allele Length (AL) field contains the length of each repeat allele according to some embodiments of the present disclosure. Fig.41B and 41C It is shown how the Motif Span (FS) field contains the span of each tandem repeat sequence on each allele according to some embodiments of the present disclosure. Fig.23 It is demonstrated how the use of the systems and methods of the present disclosure discovered a methylated mosaic FMR1 expansion between 386 and 519 CGGs, an ATXN8 expansion spanning 577 CTGs, and seven biallelic RFC1 repeat expansions with 186 to 1647 AAGGGs.
[0152] Cited References and Alternative Examples
[0153] All publications, patents, patent applications, and information available on the Internet and mentioned in this specification are incorporated herein by reference to the same extent as if each individual publication, patent, patent application, or item of information was specifically and individually indicated to be incorporated by reference. To the extent an incorporated by reference publication, patent, patent application, or item of information contradicts the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.
[0154] The present invention may be implemented as a computer program product comprising a computer program mechanism embedded in a non-transitory computer-readable storage medium. For example, a computer program product may comprise Figure 1 Shown and / or Figure 2A , 2B, 3A and / or 3B. These program modules can be stored on a CD-ROM, a DVD, a disk storage product, a USB key or any other non-transitory computer-readable data or program storage product.
[0155] It will be apparent to those skilled in the art that many modifications and variations may be made thereto without departing from the spirit and scope of the invention. The specific embodiments described herein are provided as examples only. Such embodiments are selected and described in order to best explain the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize the invention and various embodiments with various modifications suitable for the specific purposes covered. The invention is limited only by the terms of the appended claims and the full scope of equivalents to which such claims are entitled.
Claims
1. A method for mapping a plurality of sequence reads to a genomic region, the method comprising: At a computer system including one or more processors and system memory: a) obtaining the plurality of sequence reads in electronic form, wherein each sequence read in the plurality of sequence reads overlaps with the genomic region; b) obtaining a repetitive sequence definition for the genomic region, wherein the repetitive region comprises at least: (i) a first region comprising a first variable number of repeats of a first repetitive sequence, (ii) a second region comprising a second variable number of repeats of a second repetitive sequence, and (iii) a fixed interrupt sequence between the first region and the second region; c) for each corresponding sequence read in the plurality of sequences, performing a procedure comprising: (i) using the repetitive sequence definition, generating a corresponding graph for the corresponding sequence reads by scanning the corresponding sequence reads from a first end to a second end to fully match each of a plurality of corresponding motifs in the repetitive sequence definition, the corresponding graph comprising a plurality of corresponding nodes and a plurality of corresponding edges, wherein Each node in the corresponding plurality of nodes represents a motif in the plurality of motifs, The plurality of motifs includes at least a first instance of the first repeating sequence, a first instance of the second repeating sequence, an instance of the fixed interrupt sequence, and a second instance of the first repeating sequence or the second repeating sequence, Among the plurality of motifs observed to be contiguous in the corresponding sequence reads, each edge of the plurality of edges connects a corresponding node of a first motif and a corresponding node of a second motif, and The corresponding graph has one or more branch points, (ii) identifying the longest path through the corresponding graph as a candidate split for the corresponding sequence read, and (iii) mapping the corresponding sequence reads to the genomic regions using the longest path in the corresponding graph.
2. The method of claim 1, wherein the repeat sequence definition specifies that the first repeat sequence is repeated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times and the second repeat sequence is repeated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 times.
3. The method of claim 1, wherein the repeat sequence definition specifies that the first repeat sequence is repeated at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 times and the second repeat sequence is repeated at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 times.
4. The method according to any one of claims 1 to 3, wherein the first repeating sequence has a length between 2 and 100 residues, the fixed interrupting sequence has a length between 2 and 100 residues, and the second repeating sequence has a length between 2 and 100 residues.
5. The method of any one of claims 1 to 4, wherein the plurality of sequence reads have an average length of at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, or 2000 residues.
6. The method of any one of claims 1 to 5, wherein the plurality of sequence reads comprises 1000, 2000, 5000 or 10,000 sequence reads.
7. The method according to any one of claims 1 to 6, wherein the use (iii) comprises: Generate corresponding multiple segmentations according to the longest path and the repeated sequence definition, selecting a respective first segmentation of the respective plurality of segmentations having the best score as the segmentation for the respective sequence read, and The corresponding sequence reads are mapped to the genomic regions using the corresponding first partitioning.
8. The method of claim 7, wherein the corresponding plurality of partitions comprises 100, 500, 1000, 2000, 3000, 4000, 5000, 10,000, 100,000, or 1×10 6 A different split.
9. The method of any one of claims 1 to 6, wherein the plurality of sequence reads are generated in a single molecule sequencing-by-synthesis reaction.
10. The method of claim 9, wherein the single molecule sequencing-by-synthesis reaction is a single molecule real-time (SMRT) sequencing reaction.
11. The method of any one of claims 1 to 10, wherein the genomic region is in a genome.
12. The method of claim 11, wherein the genome is a human genome.
13. A method according to claim 12, wherein the multiple sequence reads are derived from a subject, the genomic region is associated with a disease, and the corresponding sequence reads are mapped to the genomic region using the longest path in the corresponding graph to identify the state, stage, presence or absence of the disease in the subject.
14. The method of claim 13, wherein the disease is tandem repeat disease, Alzheimer's disease, autism spectrum disorder, fragile X syndrome, epilepsy, amyotrophic lateral sclerosis, Huntington's disease, Kennedy's disease, myotonic dystrophy, or spinocerebellar ataxia.
15. The method according to any one of claims 1 to 14, wherein said obtaining the repetitive sequence definition for the genomic region comprises identifying the repetitive sequence definition from a plurality of repetitive sequence definitions based on an identification of the genomic region.
16. The method of claim 15, wherein the plurality of repeat sequence definitions comprises 10 or more repeat sequence definitions, 100 or more repeat sequence definitions, 1000 or more repeat sequence definitions, 100000 or more repeat sequence definitions, or 1×10 6 or more repeat sequence definitions.
17. The method of any one of claims 1 to 14, wherein the plurality of sequence reads originate from a subject, and the method further comprises phasing the genomic region using the mapping of the plurality of sequence reads.
18. The method of any one of claims 1 to 14, wherein the plurality of sequence reads originate from a subject, and the method further comprises determining a status of a genetic disease associated with the genomic region in the subject using the mapping of the plurality of sequence reads.
19. A method for mapping a plurality of sequence reads to genomic regions, the method comprising: At a computer system including one or more processors and system memory: a) obtaining the plurality of sequence reads in electronic form, wherein each sequence read in the plurality of sequence reads overlaps with the genomic region; b) obtaining an initial Markov model for the genomic region, wherein the initial Markov model includes at least: (i) a first repeat sequence for a first repeat region, (ii) a second repeat sequence for a second repeat region, and (iii) an intermediate region connecting the first repeat sequence to the second repeat sequence; c) optimizing the initial Markov model using the plurality of sequence reads to obtain an optimized Markov model; and d) for each corresponding sequence read in the plurality of sequences, performing a procedure comprising: (i) using the corresponding sequence reads to find the highest probability path through the Markov model, and (ii) mapping the corresponding sequence reads to the genomic region using the highest probability path.
20. The method of claim 19, wherein the first region comprises one or more instances of a first repeating sequence having a length between 2 and 100 residues, the middle region has a length between 2 and 100 residues, and the second region comprises one or more instances of a second repeating sequence having a length between 2 and 100 residues.
21. The method according to claim 20, wherein The first region further comprises one or more residues other than the first repeat sequence, and The second region further includes one or more residues other than the second repeat sequence.
22. The method of any one of claims 19 to 21, wherein the genomic region has a length between 200 and 5000 residues.
23. The method of any one of claims 19 to 21, wherein the genomic region has a length between 1000 and 8000 residues.
24. The method of any one of claims 19 to 21, wherein the genomic region has a length of between 2000 and 10,000 residues.
25. The method of any one of claims 19 to 24, wherein the plurality of sequence reads have an average length of at least 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, or 2000 residues.
26. The method of any one of claims 19 to 25, wherein the plurality of sequence reads comprises 1000, 2000, 5000 or 10,000 sequence reads.
27. The method according to any one of claims 19 to 26, wherein the using (ii) comprises: Generate a corresponding plurality of partitions, each partition being a permutation of the highest probability path, selecting a respective first segmentation of the respective plurality of segmentations having the best score as the segmentation for the respective sequence read, and The corresponding sequence reads are mapped to the genomic regions using the corresponding first partitioning.
28. The method of claim 27, wherein the corresponding plurality of partitions comprises 100, 500, 1000, 2000, 3000, 4000, 5000, 10,000, 100,000, or 1×10 6 A different split.
29. The method of any one of claims 19 to 28, wherein the plurality of sequence reads are each generated in a single molecule sequencing-by-synthesis reaction.
30. The method of claim 29, wherein the single molecule sequencing-by-synthesis reaction is a single molecule real-time (SMRT) sequencing reaction.
31. The method of any one of claims 19 to 30, wherein the genomic region is in a genome.
32. The method of claim 19, wherein the genome is a human genome.
33. A method according to claim 32, wherein the multiple sequence reads are derived from a subject, the genomic region is associated with a disease, and the status of the disease in the subject is determined by mapping the corresponding sequence reads to the genomic region using the highest probability path.
34. The method of claim 33, wherein the disease is Alzheimer's disease, autism, epilepsy, or ALS.
35. The method of any one of claims 19 to 34, wherein said obtaining the repetitive sequence definition for the genomic region comprises identifying the repetitive sequence definition from a plurality of repetitive sequence definitions based on an identification of the genomic region.
36. The method of claim 35, wherein the plurality of repeat sequence definitions comprises 10 or more repeat sequence definitions, 100 or more repeat sequence definitions, 1000 or more repeat sequence definitions, 100,000 or more repeat sequence definitions, or 1×10 6 or more repeat sequence definitions.
37. A method according to any one of claims 19 to 33, wherein the plurality of sequence reads originate from a subject, and the method further comprises mapping the corresponding sequence reads to the genomic region using the highest probability path to perform a phase analysis on the genomic region.
38. A method according to any one of claims 19 to 33, wherein the plurality of sequence reads originate from a subject, and the method further comprises using the mapping of the plurality of sequence reads to determine the status of a genetic disease associated with the genomic region in the subject.
39. A system for mapping a plurality of sequence reads to genomic regions, comprising: Memory; Input / output terminal; as well as a processor coupled to the memory, wherein the system is configured to perform a method comprising: a) obtaining the plurality of sequence reads in electronic form, wherein each sequence read in the plurality of sequence reads overlaps with the genomic region; b) obtaining a repetitive sequence definition for the genomic region, wherein the repetitive region comprises at least: (i) a first region comprising a first variable number of repeats of a first repetitive sequence, (ii) a second region comprising a second variable number of repeats of a second repetitive sequence, and (iii) a fixed interrupt sequence between the first region and the second region; c) for each corresponding sequence read in the plurality of sequences, performing a procedure comprising: (i) using the repetitive sequence definition, generating a corresponding graph for the corresponding sequence reads by scanning the corresponding sequence reads from a first end to a second end to fully match each of a plurality of corresponding motifs in the repetitive sequence definition, the corresponding graph comprising a plurality of corresponding nodes and a plurality of corresponding edges, wherein Each node in the corresponding plurality of nodes represents a motif in the plurality of motifs, The plurality of motifs includes at least a first instance of the first repeating sequence, a first instance of the second repeating sequence, an instance of the fixed interrupt sequence, and a second instance of the first repeating sequence or the second repeating sequence, Among the plurality of motifs observed to be contiguous in the corresponding sequence reads, each edge of the plurality of edges connects a corresponding node of a first motif and a corresponding node of a second motif, and The corresponding graph has one or more branch points, (ii) identifying the longest path through the corresponding graph as a candidate split for the corresponding sequence read, and (iii) mapping the corresponding sequence reads to the genomic regions using the longest path in the corresponding graph.
40. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores instructions that, when executed by a computer system, cause the computer system to perform a method for mapping a plurality of sequence reads to genomic regions, the method comprising: a) obtaining the plurality of sequence reads in electronic form, wherein each sequence read in the plurality of sequence reads overlaps with the genomic region; b) obtaining a repetitive sequence definition for the genomic region, wherein the repetitive region comprises at least: (i) a first region comprising a first variable number of repeats of a first repetitive sequence, (ii) a second region comprising a second variable number of repeats of a second repetitive sequence, and (iii) a fixed interrupt sequence between the first region and the second region; c) for each corresponding sequence read in the plurality of sequences, performing a procedure comprising: (i) using the repetitive sequence definition, generating a corresponding graph for the corresponding sequence reads by scanning the corresponding sequence reads from a first end to a second end to fully match each of a plurality of corresponding motifs in the repetitive sequence definition, the corresponding graph comprising a plurality of corresponding nodes and a plurality of corresponding edges, wherein Each node in the corresponding plurality of nodes represents a motif in the plurality of motifs, The plurality of motifs includes at least a first instance of the first repeating sequence, a first instance of the second repeating sequence, an instance of the fixed interrupt sequence, and a second instance of the first repeating sequence or the second repeating sequence, Among the plurality of motifs observed to be contiguous in the corresponding sequence reads, each edge of the plurality of edges connects a corresponding node of a first motif and a corresponding node of a second motif, and The corresponding graph has one or more branch points, (ii) identifying the longest path through the corresponding graph as a candidate split for the corresponding sequence read, and (iii) mapping the corresponding sequence reads to the genomic regions using the longest path in the corresponding graph.
41. A system for mapping a plurality of sequence reads to genomic regions, comprising: Memory; Input / output terminal; as well as a processor coupled to the memory, wherein the system is configured to perform a method comprising: a) obtaining the plurality of sequence reads in electronic form, wherein each sequence read in the plurality of sequence reads overlaps with the genomic region; b) obtaining an initial Markov model for the genomic region, wherein the initial Markov model includes at least: (i) a first repeat sequence for a first repeat region, (ii) a second repeat sequence for a second repeat region, and (iii) an intermediate region connecting the first repeat sequence to the second repeat sequence; c) optimizing the initial Markov model using the plurality of sequence reads to obtain an optimized Markov model; and d) for each corresponding sequence read in the plurality of sequences, performing a procedure comprising: (i) using the corresponding sequence reads to find the highest probability path through the Markov model, and (ii) mapping the corresponding sequence reads to the genomic region using the highest probability path.
42. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores instructions that, when executed by a computer system, cause the computer system to perform a method for mapping a plurality of sequence reads to genomic regions, the method comprising: a) obtaining the plurality of sequence reads in electronic form, wherein each sequence read in the plurality of sequence reads overlaps with the genomic region; b) obtaining an initial Markov model for the genomic region, wherein the initial Markov model includes at least: (i) a first repeat sequence for a first repeat region, (ii) a second repeat sequence for a second repeat region, and (iii) an intermediate region connecting the first repeat sequence to the second repeat sequence; c) optimizing the initial Markov model using the plurality of sequence reads to obtain an optimized Markov model; and d) for each corresponding sequence read in the plurality of sequences, executing a procedure comprising: (i) finding a highest probability path through the Markov model using the corresponding sequence read, and (ii) mapping the corresponding sequence read to the genomic region using the highest probability path.
Citation Information
Patent Citations
Nucleic acid sequencing methods and systems
US20090087850A1
Formation of Lipid Bilayers
US20100196203A1
Enzyme-pore constructs
US20110229877A1
Nanopore Based Molecular Detection and Sequencing
US20130244340A1
Nucleic acid sequencing by nanopore detection of tag molecules
US20150119259A1