Determination of pathogenic RFC1 expansions from sequencing data
Patent Information
- Application Number
- JP2023576188
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-06-11
- Filing Date
- 2022-06-10
- Publication Date
- 2025-06-19
AI Technical Summary
Current diagnostic methods for biallelic intronic AAGGG repeat expansions in the RFC1 gene, which cause familial cerebellar ataxia and vestibular reflex syndrome (CANVAS), are time-consuming and inaccurate, especially in GC-high genomic regions, making it difficult to identify pathogenic expansions from sequencing data.
A method using sequence graph alignment and pentamer filtering to automatically and accurately identify pathogenic RFC1 repeat expansions by determining the frequency and quality of repetitive sequences in sequencing data, allowing for rapid screening of large populations.
The method significantly improves the accuracy and speed of identifying pathogenic RFC1 repeat expansions, distinguishing between pathogenic, carrier, and benign statuses with high precision, even in challenging genomic regions.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Related Applications This application claims priority under 35 USC § 119(e) to U.S. Provisional Application No. 63 / 209,608, filed June 11, 2021. The contents of the related application are incorporated herein by reference in their entirety.
[0002] Sequence Listing Reference This application has been filed with a sequence listing in electronic format. The sequence listing is provided as a file entitled Sequence_Listing_47CX-311978-WO.txt, created May 29, 2022, which is 10 kilobytes in size. The information in the electronic format of the sequence listing is incorporated herein by reference in its entirety.
[0003] The present disclosure relates generally to the field of processing sequence data, and more specifically to determining repeats. [Background technology]
[0004] A biallelic intronic AAGGG repeat expansion in the replication factor C subunit (RFC1) gene can cause familial cerebellar ataxia, neuropathy, and vestibular absent reflex syndrome (CANVAS) and tardive ataxia. Current diagnostic methods for this pathogenic AAGGG expansion are time-consuming and cannot be performed on a large scale. One method is clinical whole genome sequencing (cWGS) screening by manual inspection of reads from alignments. However, in this GC-rich genomic region, the base quality of paired-end reads is reduced, making manual inspection more difficult and error-prone, especially in the presence of pathogenic AAGGG or other high GC repeat patterns. There is a need for an automated and accurate method to identify this pathogenic expansion from sequencing data, including cWGS data. Summary of the Invention
[0005] Disclosed herein is a method for determining replication factor C subunit 1 (RFC1) repeat expansion status. In some embodiments, the method for determining RFC1 repeat expansion status is under the control of a processor (such as a hardware processor or a virtual processor) and includes (a) receiving a plurality of sequence reads generated from a sample obtained from a subject. The method may include (b) aligning the plurality of sequence reads to a sequence graph to generate a plurality of aligned sequence reads. The sequence graph may represent the locus of RFC1. The sequence graph may include a representation of a repeat sequence adjacent to a non-repetitive sequence of the locus of RFC1. The plurality of aligned sequence reads may include the plurality of sequence reads and an alignment of the plurality of sequence reads to the sequence graph. The method may include (c) determining the occurrence of the plurality of repeat sequences in the aligned sequence reads of the plurality of aligned sequence reads using a first occurrence threshold and a first quality threshold. The method may include (d) determining a frequency representation of the occurrence of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences. The method may include (e) determining the repeat expansion status at the subject's RFC1 locus using a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of multiple repeat sequences.
[0006] In some embodiments, the method includes determining that the subject has 0, 1, or 2 alleles with a repeat expansion at the locus of RFC1 using a plurality of aligned sequence reads. In some embodiments, the repeat expansion at the locus of RFC1 is at about chr4:39348424 of hg38, or a corresponding position in another reference genome sequence.
[0007] In some embodiments, the status of the repeat expansion at the RFC1 locus is pathogenic status, carrier status, or benign status. In some embodiments, the repeat expansion is associated with or causes a disease. The disease may be cerebellar ataxia, neuropathy, and vestibulo-ocular syndrome (CANVAS). In some embodiments, the method includes confirming the status of the repeat expansion at the RFC1 locus of the subject using one or more diagnostic methods. The one or more diagnostic methods may include polymerase chain reaction (PCR) and Sanger sequencing, Southern blot, and linkage analysis.
[0008] In some embodiments, the subject has two alleles with a repeat expansion at the RFC1 locus. In some embodiments, determining the status of the repeat expansion at the RFC1 locus comprises determining that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is greater than a first status threshold. Determining the status of the repeat expansion at the RFC1 locus may comprise determining the status of the repeat expansion at the RFC1 locus as pathogenic status.
[0009] In some embodiments, determining the status of the repeat expansion at the RFC1 locus comprises determining a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is less than or equal to a first status threshold and greater than or equal to a second status threshold. Determining the status of the repeat expansion at the RFC1 locus may comprise determining the status of the repeat expansion at the RFC1 locus as a carrier status.
[0010] In some embodiments, determining the status of the repeat expansion at the RFC1 locus includes determining a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is less than or equal to a first status threshold and greater than or equal to a second status threshold. Determining the status of the repeat expansion at the RFC1 locus may include determining a frequency representation of the number of occurrences of the pathogenic repeat sequence and determining that a sequence with high sequence similarity to the pathogenic repeat sequence is greater than a third status threshold. For example, the pathogenic repeat sequence and a sequence similar to the pathogenic repeat sequence may differ by one base. Determining the status of the repeat expansion at the RFC1 locus may include determining a frequency representation of the number of occurrences of the pathogenic sequence is greater than a fourth status threshold. Determining the status of the repeat expansion at the RFC1 locus may include determining the status of the repeat expansion at the RFC1 locus as a carrier status.
[0011] In some embodiments, determining the status of the repeat expansion at the RFC1 locus comprises determining that a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to a total number of occurrences of the plurality of repeat sequences is less than a second status threshold. Determining the status of the repeat expansion at the RFC1 locus may comprise determining the status of the repeat expansion at the RFC1 locus as a benign status.
[0012] In some embodiments, the subject has one allele of RFC1 with a repeat expansion at the RFC1 locus. In some embodiments, determining the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of sequence reads comprises determining the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of sequence reads from one allele with a repeat expansion at the RFC1 locus. Determining the status of the repeat expansion at the RFC1 locus may comprise determining that the frequency representation of the occurrence number of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is greater than a first status threshold. Determining the status of the repeat expansion at the RFC1 locus may comprise determining the status of the repeat expansion at the RFC1 locus as pathogenic status.
[0013] In some embodiments, determining the number of occurrences of the plurality of repeat sequences in the aligned sequence reads of the plurality of sequence reads comprises determining the number of occurrences of the plurality of repeat sequences in the aligned sequence reads of the plurality of sequence reads from one allele having a repeat expansion at the RFC1 locus. Determining the status of the repeat expansion at the RFC1 locus may comprise determining that a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is equal to or less than a first status threshold. Determining the status of the repeat expansion at the RFC1 locus may comprise determining the status of the repeat expansion at the RFC1 locus as a carrier status.
[0014] In some embodiments, the method includes selecting aligned sequence reads derived from one allele of RFC1 having a repeat expansion at the locus of RFC1 based on an alignment of the aligned sequence reads. In some embodiments, the aligned sequence reads derived from one allele of RFC1 having a repeat expansion at the locus of RFC1 include aligned sequence reads that are (i) in-repeat reads or (ii) adjacent reads, each of which has an overlap to a repeat expansion at the locus of RFC1 greater than a repeat expansion overlap threshold. The repeat expansion overlap threshold may be about 60 base pairs. The aligned sequence reads derived from one allele of RFC1 having a repeat expansion at the locus of RFC1 may include (i) all aligned sequence reads that are in-repeat reads and (ii) less than all aligned sequence reads that are adjacent reads.
[0015] In some embodiments, the subject has 0 alleles with a repeat expansion at the RFC1 locus. The status of the repeat expansion at the RFC1 locus can be benign.
[0016] In some embodiments, the repeat expansion at the RFC1 locus comprises one or more threshold total copies of the repeat sequence. The threshold total copy number of the one or more repeat sequences may be 20 total copies of the one or more repeat sequences. In some embodiments, each of the multiple repeat sequences has a number of occurrences equal to or greater than a first occurrence threshold, and each occurrence has a number of bases each having a quality score equal to or greater than a first quality threshold. The first occurrence threshold may be 2. The first quality threshold may be about 20. The number of bases of the repeat sequence each having a quality score greater than the first quality threshold may be 5. In some embodiments, each of the occurrences has a number of bases each having a quality score greater than a second quality threshold. The second quality threshold may be about 20. The number of bases of each occurrence having a quality score greater than the second quality threshold may be 3.
[0017] In some embodiments, the repeat sequence representation is degenerate. The repeat sequence representation can be AARRG. The repeat sequence representation can be at least 5 bases in length. The repeat sequence representation can be 5 bases in length. The repeat sequence representation can be 6 bases in length. Each of the plurality of repeat sequences can be at least 5 bases in length. Each of the plurality of repeat sequences can be 5 bases in length. Each of the plurality of repeat sequences can be 6 bases in length. The pathogenic repeat sequence can be AAGGG or ACAGG. The pathogenic repeat sequence can have a GC content of at least 60%.
[0018] In some embodiments, the frequency representation of the number of occurrences of a pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences is a percentage of the number of occurrences of the pathogenic repeat sequence from the total number of occurrences of the multiple repeat sequences. The frequency representation of the number of occurrences of a pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences can be the ratio of the number of occurrences of a pathogenic repeat sequence to the total number of occurrences of the multiple repeat sequences.
[0019] In some embodiments, the plurality of sequence reads are aligned to the RFC1 locus. In some embodiments, receiving the plurality of sequence reads generated from the obtained sample includes aligning a second plurality of sequence reads comprising the plurality of sequence reads to a reference genome sequence. Receiving the plurality of sequence reads generated from the obtained sample may include selecting a plurality of sequence reads from the second plurality of sequence reads, where the plurality of sequence reads are aligned to the RFC1 locus.
[0020] In some embodiments, the plurality of sequence reads comprises sequence reads each of which is between about 100 base pairs and about 1000 base pairs in length. The plurality of sequence reads may comprise paired-end sequence reads. The plurality of sequence reads may comprise single-end sequence reads. The plurality of sequence reads may be generated by targeted sequencing and / or whole genome sequencing (WGS), optionally where the WGS is clinical WGS (cWGS). In some embodiments, the sample comprises cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, a blood sample, a biopsy sample, or a combination thereof. In some embodiments, the reference genome sequence comprises a reference human genome sequence. The subject may be a human subject.
[0021] Disclosed herein is a system for determining the repeat expansion status of a gene of interest. In some embodiments, the system for determining the repeat expansion status of a gene of interest includes a non-transitory memory configured to store executable instructions, and a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory. The processor may be programmed by the executable instructions to: (a) receive a plurality of sequence reads generated from a sample obtained from a subject. The processor may be programmed by the executable instructions to: (b) align the plurality of sequence reads to a sequence graph to generate a plurality of aligned sequence reads. The sequence graph may represent a locus of the gene of interest. The sequence graph may include a representation of a repeat sequence adjacent to a non-repeat sequence of the locus of the gene. The plurality of aligned sequence reads may include a plurality of sequence reads. The plurality of aligned sequence reads may include an alignment of the plurality of sequence reads to the sequence graph. The processor may be programmed with executable instructions to (c) determine a number of occurrences of the plurality of repetitive sequences in the aligned sequence reads of the plurality of aligned sequence reads using a first occurrence threshold and a first quality threshold. The processor may be programmed with executable instructions to (d) determine a frequency representation of the number of occurrences of the pathogenic repetitive sequence relative to a total number of occurrences of the plurality of repetitive sequences. The processor may be programmed with executable instructions to (e) determine a repeat expansion status at a locus of a gene of interest of a subject using a frequency representation of the number of occurrences of the pathogenic repetitive sequence relative to a total number of occurrences of the plurality of repetitive sequences.
[0022] In some embodiments, the gene of interest is replication factor C subunit 1 (RFC1). The locus of the gene of interest with the repeat expansion is at about chr4:39348424 of hg38 or the corresponding position of another reference genome sequence. The repeat expansion may be associated with or cause a disease. The disease may be cerebellar ataxia, neuropathy, and vestibulo-ocular reflex syndrome (CANVAS). The repeat expression may be AARRG. The pathogenic repeat may be AAGGG or ACAGG.
[0023] In some embodiments, the processor is programmed with executable instructions to use the plurality of aligned sequence reads to determine that the subject has zero, one or two alleles with a repeat expansion at the locus of the gene of interest. In some embodiments, the status of the repeat expansion at the locus of the gene of interest is a pathogenic status, a carrier status or a benign status. The repeat expansion may be associated with or cause a disease. The disease may be a neurological disease. In some embodiments, the processor is programmed with executable instructions to receive a conformation of the status of the repeat expansion at the locus of the gene of interest of the subject determined using one or more diagnostic systems. The one or more diagnostic systems may include polymerase chain reaction (PCR) and Sanger sequencing, Southern blot, and linkage analysis.
[0024] In some embodiments, the subject has two alleles with repeat expansion at the locus of the gene of interest.In some embodiments, determining the status of the repeat expansion at the locus of the gene of interest comprises determining that the frequency representation of the number of occurrences of pathogenic repeat sequences relative to the total number of occurrences of multiple repeat sequences is greater than a first status threshold.Determining the status of the repeat expansion at the locus of the gene of interest can comprise determining the status of the repeat expansion at the locus of the gene of interest as pathogenic status.
[0025] In some embodiments, determining the status of the repeat expansion at the locus of the gene of interest comprises determining a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is less than or equal to a first status threshold and greater than or equal to a second status threshold. Determining the status of the repeat expansion at the locus of the gene of interest may comprise determining the status of the repeat expansion at the locus of the gene of interest as a carrier status.
[0026] In some embodiments, determining the status of the repeat expansion at the locus of the gene of interest comprises determining a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is less than or equal to a first status threshold and greater than or equal to a second status threshold. Determining the status of the repeat expansion at the locus of the gene of interest may comprise determining a frequency representation of the number of occurrences of the pathogenic repeat sequence and determining that a sequence with high sequence similarity to the pathogenic repeat sequence is greater than a third status threshold. For example, the pathogenic repeat sequence and a sequence similar to the pathogenic repeat sequence differ by one base. Determining the status of the repeat expansion at the locus of the gene of interest may comprise determining a frequency representation of the number of occurrences of the pathogenic sequence is greater than a fourth status threshold, and determining the status of the repeat expansion at the locus of the gene of interest may comprise determining the status of the repeat expansion at the locus of the gene of interest as a carrier status.
[0027] In some embodiments, determining the status of the repeat expansion at the locus of the gene of interest comprises determining that a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is less than a second status threshold. Determining the status of the repeat expansion at the locus of the gene of interest may comprise determining the status of the repeat expansion at the locus of the gene of interest as a benign status.
[0028] In some embodiments, the subject has one allele of the gene of interest with a repeat expansion at the locus of the gene of interest.In some embodiments, determining the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of sequence reads comprises determining the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of aligned sequence reads derived from one allele of the gene of interest with a repeat expansion at the locus of the gene of interest.Determining the status of the repeat expansion at the locus of the gene of interest may comprise determining that the frequency representation of the occurrence number of the pathogenic repeat sequence relative to the total number of occurrence number of the plurality of repeat sequences is greater than a first status threshold.Determining the status of the repeat expansion at the locus of the gene of interest may comprise determining the status of the repeat expansion at the locus of the gene of interest as pathogenic status.
[0029] In some embodiments, determining the occurrence number of a plurality of repeat sequences in the aligned sequence reads of the plurality of sequence reads comprises determining the occurrence number of a plurality of repeat sequences in the aligned sequence reads of the plurality of aligned sequence reads derived from one allele of a gene of interest having a repeat expansion at a locus of the gene of interest. Determining the status of the repeat expansion at the locus of the gene of interest may comprise determining that a frequency representation of the occurrence number of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is equal to or less than a first status threshold. Determining the status of the repeat expansion at the locus of the gene of interest may comprise determining the status of the repeat expansion at the locus of the gene of interest as a carrier status.
[0030] In some embodiments, the processor is programmed with executable instructions to select aligned sequence reads derived from one allele of a gene of interest having a repeat expansion at the locus of the gene of interest based on the alignment of the aligned sequence reads. In some embodiments, the aligned sequence reads derived from one allele of a gene of interest having a repeat expansion at the locus of the gene of interest include aligned sequence reads that are (i) in-repeat reads or (ii) adjacent reads, each of which has an overlap with a repeat expansion at the locus of the gene of interest greater than a repeat expansion overlap threshold, optionally the repeat expansion overlap threshold being about 60 base pairs. The aligned sequence reads derived from one allele of a gene of interest having a repeat expansion at the locus of the gene of interest may include (i) all aligned sequence reads that are in-repeat reads, and (ii) less than all aligned sequence reads that are adjacent reads.
[0031] In some embodiments, the subject has zero alleles with the repeat expansion at the locus of the gene of interest. The status of the repeat expansion at the locus of the gene of interest may be benign.
[0032] In some embodiments, the repeat expansion at the locus of the gene of interest comprises one or more threshold total copies of the repeat sequence. The threshold total copy number of the one or more repeat sequences may be 20 total copies of the one or more repeat sequences. In some embodiments, each of the multiple repeat sequences has a number of occurrences equal to or greater than a first occurrence threshold, and each occurrence has a number of bases each having a quality score equal to or greater than a first quality threshold. The first occurrence threshold may be 2. The first quality threshold may be about 20. The number of bases of the repeat sequence each having a quality score greater than the first quality threshold may be 5. In some embodiments, each of the occurrences has a number of bases each having a quality score greater than or equal to a second quality threshold. The second quality threshold may be about 20. The number of bases of each occurrence having a quality score greater than the second quality threshold may be 3.
[0033] In some embodiments, the repetitive sequence representation is degenerate. The repetitive sequence representation and / or each of the plurality of repetitive sequences may be at least 5 bases in length. The repetitive sequence representation and / or each of the plurality of repetitive sequences may be 5 bases in length. The repetitive sequence representation and / or each of the plurality of repetitive sequences may be 6 bases in length. In some embodiments, the pathogenic repetitive sequence has a GC content of at least 60%.
[0034] In some embodiments, the frequency representation of the number of occurrences of a pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is the percentage of the number of occurrences of a pathogenic repeat sequence out of the total number of occurrences of the plurality of repeat sequences, or the ratio of the number of occurrences of a pathogenic repeat sequence to the total number of occurrences of the plurality of repeat sequences.
[0035] In some embodiments, the plurality of sequence reads are aligned to the locus of the gene of interest.In some embodiments, receiving the plurality of sequence reads generated from the obtained sample comprises aligning a second plurality of sequence reads comprising the plurality of sequence reads to a reference genome sequence.Receiving the plurality of sequence reads generated from the obtained sample may comprise selecting a plurality of sequence reads from the second plurality of sequence reads, wherein the plurality of sequence reads are aligned to the locus of the gene of interest.
[0036] In some embodiments, the plurality of sequence reads comprises sequence reads each of about 100 base pairs to about 1000 base pairs in length. The plurality of sequence reads comprises paired-end sequence reads. The plurality of sequence reads comprises single-end sequence reads. In some embodiments, the plurality of sequence reads are generated by targeted sequencing and / or whole genome sequencing (WGS), optionally where the WGS is clinical WGS (cWGS). In some embodiments, the sample comprises cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, a blood sample, a biopsy sample, or a combination thereof. In some embodiments, the reference genome sequence comprises a reference human genome sequence. The subject may be a human subject.
[0037] The details of one or more implementations of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the specification, drawings, and claims. Neither this summary nor the following detailed description is intended to define or limit the scope of the inventive subject matter. [Brief description of the drawings]
[0038] [Figure 1] FIG. 1 shows a non-limiting, exemplary schematic diagram of a sequence graph representing the RFC1 locus and classification of reads overlapping with a repeat as spanning reads, in-repeat reads and flanking reads. [Diagram 2]
[0023] Figure 1 shows a non-limiting exemplary visualization of patient sample reads overlapping the repeat. The patient was confirmed to have two expanded AAGGG haplotypes by PCR. [Diagram 3] Figure 1 shows a non-limiting exemplary visualization of patient sample reads overlapping with a repeat. The human patient was a putative patient. However, PCR did not validate the putative. PCR suggested an AAGAG repeat instead of an AAGGG repeat. [Figure 4A] FIG. 1 is a schematic diagram of a non-limiting exemplary 5mer filtering and counting. [Figure 4B] FIG. 1 is a schematic diagram of a non-limiting exemplary 5mer filtering and counting. [Figure 5A] FIG. 1 shows that carriers require better separation for the two haplotypes. [Figure 5B] FIG. 1 shows that carriers require better separation for the two haplotypes. [Figure 6A] FIG. 1 is a non-limiting, exemplary schematic diagram of sequence reads derived from an extended allele. [Figure 6B] FIG. 1 is a non-limiting, exemplary schematic diagram of sequence reads derived from an extended allele. [Figure 7]FIG. 1 shows the proportion of 5-mer repeats in the expanded allele for samples with one short allele and one expanded allele in a cohort of 75 patients. [Figure 8] FIG. 1 shows the proportion of 5-mer repeats in the expanded allele for samples with one short allele and one expanded allele for the Polaris population of 150 unrelated healthy individuals. [Figure 9] FIG. 11 is a flow diagram illustrating an exemplary method for determining RFC1 repeat extension status. [Figure 10] FIG. 1 is a flow diagram showing an exemplary method for determining the repeat expansion status of a gene of interest (e.g., RFC1). [Figure 11] FIG. 1 is a block diagram of an exemplary computing system configured to perform a determination of the repeat expansion status of a gene of interest (e.g., RFC1).
[0039] Throughout the drawings, reference numbers may be reused to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0040] In the following detailed description, reference is made to the accompanying drawings, which form a part of this specification. In the drawings, like symbols typically identify like components unless the context dictates otherwise. The exemplary embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It is readily understood that aspects of the present disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein and make part of the disclosure herein.
[0041] All patents, published patent applications, other publications, and sequences from GenBank and other databases referenced herein are hereby incorporated by reference in their entirety with respect to the relevant art.
[0042] Disclosed herein is a method for determining replication factor C subunit 1 (RFC1) repeat expansion status. In some embodiments, the method for determining RFC1 repeat expansion status is under the control of a processor (such as a hardware processor or a virtual processor) and includes (a) receiving a plurality of sequence reads generated from a sample obtained from a subject. The method may include (b) aligning the plurality of sequence reads to a sequence graph to generate a plurality of aligned sequence reads. The sequence graph may represent the locus of RFC1. The sequence graph may include a representation of a repeat sequence adjacent to a non-repetitive sequence of the locus of RFC1. The plurality of aligned sequence reads may include the plurality of sequence reads and an alignment of the plurality of sequence reads to the sequence graph. The method may include (c) determining the occurrence of the plurality of repeat sequences in the aligned sequence reads of the plurality of aligned sequence reads using a first occurrence threshold and a first quality threshold. The method may include (d) determining a frequency representation of the occurrence of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences. The method may include (e) determining the repeat expansion status at the subject's RFC1 locus using a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of multiple repeat sequences.
[0043] Disclosed herein is a system for determining the repeat expansion status of a gene of interest. In some embodiments, the system for determining the repeat expansion status of a gene of interest includes a non-transitory memory configured to store executable instructions, and a processor (e.g., a hardware processor or a virtual processor) in communication with the non-transitory memory. The processor may be programmed by the executable instructions to: (a) receive a plurality of sequence reads generated from a sample obtained from a subject. The processor may be programmed by the executable instructions to: (b) align the plurality of sequence reads to a sequence graph to generate a plurality of aligned sequence reads. The sequence graph may represent a locus of the gene of interest. The sequence graph may include a representation of a repeat sequence adjacent to a non-repeat sequence of the locus of the gene. The plurality of aligned sequence reads may include a plurality of sequence reads. The plurality of aligned sequence reads may include an alignment of the plurality of sequence reads to the sequence graph. The processor may be programmed with executable instructions to (c) determine a number of occurrences of the plurality of repetitive sequences in the aligned sequence reads of the plurality of aligned sequence reads using a first occurrence threshold and a first quality threshold. The processor may be programmed with executable instructions to (d) determine a frequency representation of the number of occurrences of the pathogenic repetitive sequence relative to a total number of occurrences of the plurality of repetitive sequences. The processor may be programmed with executable instructions to (e) determine a repeat expansion status at a locus of a gene of interest of a subject using a frequency representation of the number of occurrences of the pathogenic repetitive sequence relative to a total number of occurrences of the plurality of repetitive sequences.
[0044] Determination of pathogenic RFC1 expansions A biallelic intronic AAGGG repeat expansion in the replication factor C subunit (RFC1) gene has been recently discovered as a cause of familial cerebellar ataxia, neuropathy, and vestibular absent reflex syndrome (CANVAS) and a frequent cause of late-onset ataxia. In the general population, this repeat expansion is highly variable, with many different benign repeat patterns. Current diagnostic methods for this pathogenic AAGGG expansion include polymerase chain reaction (PCR) and Sanger sequencing, Southern blot and linkage analysis, but all of these methods are time-consuming and cannot be performed on a large scale. Compared to these methods, clinical whole genome sequencing (cWGS) screening allows this expansion to be identified in a massively parallel manner. However, there is no automated and accurate method to identify this pathogenic expansion from cWGS data. Currently, the pathogenic expansions are identified by manual examination of reads from alignments, such as alignments by ExpansionHunter (Dolzhenko, E. et al. ExpansionHunter Denovo: a computational method for locating known and novel repeat expansions in short-read sequencing data. Genome Biol 21, 102 (2020). https: / / doi.org / 10.1186 / s13059-020-02017-z, the contents of which are incorporated herein by reference in their entirety). When sample sizes are large, this manual screening can be very time-consuming.
[0045] Furthermore, in this GC-rich genomic region, the base quality of the paired-end reads is reduced. This quality reduction correlates with what the repeat pattern is. This makes manual screening more difficult and error-prone, especially in the presence of pathogenic AAGGG or other high GC repeat patterns. Therefore, there was an opinion that it is not possible to accurately diagnose this pathogenic repeat from paired-end cWGS data.
[0046] Disclosed herein is a method for automatically and accurately screening this pathogenic expansion from sequencing data, such as cWGS data. There are no other methods for screening this pathogenic repeat from sequencing data. The method implements the logic described herein to minimize artifacts due to poor data quality in this area. For example, for one cWGS sample, conventional methods determined more than 80 different repeat patterns, with only about 64% AAGGG percentage. In comparison, for the same sample, the disclosed method determined only a few different repeat patterns, with more than 90% AAGGG percentage. The method implements rules to separate actual pathogenic repeats from benign repeats. The method can rapidly screen a large population and output patients with double expanded AAGGG repeats, carriers with expanded AAGGG repeats on one haplotype and benign repeats on the other haplotype, and normal people without expanded AAGGG on either haplotype. For one 30× cWGS sample, the method took less than 1 second to output screening results (after ExpansionHunter was run).
[0047] A repeat expansion in RFC1 has been identified as the cause of CANVAS. In GRCh38, RFC1 is located at chr4:39348424-39348479 (AAAAG) 11 Has repetition. (AAAAG) exp and in some cases (AAAGG) exp There are many other benign expansion patterns, such as: Homozygous (AAGGG) exp and (ACAGG) exphas been identified as pathogenic. Diagnostic methods for pathogenic RFC1 expansions include PCR and Sanger sequencing, Southern blot, linkage analysis, and clinical whole genome sequencing (cWGS) screening. PCR and Sanger sequencing utilize different primers for mutants and wild type. However, long-range PCR is more error-prone in repeat regions. Southern blots are more difficult than PCR. Southern blots can confirm the presence of biallelic large expansions. Southern blots do not show what kind of expansions they are. With regard to linkage analysis, family-based studies have identified several linkage disequilibrium (LD) single nucleotide polymorphisms (SNPs) for mutants. However, some lineages do not guarantee that the SNPs are in LD with the pathogenic repeats.
[0048] This pathogenic expansion can be identified by cWGS screening, for example by manual inspection of reads from ExpansionHunter (EH) alignments. Screening is performed by genotyping the reference AAAAG and pathogenic AAGGG (AARRG) sequences simultaneously. n (R is A or G) (Figure 1) with the constructed repeat graph, it can be performed by ExpansionHunter. In the graph alignment, the reads overlapping with the repeat can be classified as spanning reads, in-repeat reads and flanking reads as shown in Figure 1. Spanning reads can occur when the repeat is shorter than the read length such that the repeat includes the repeat and the flanking regions at both ends of the repeat. When the repeat is longer than the read length, in-repeat reads can occur such that the entire read (whether single-end or paired-end sequencing reads) or one read of the paired-end sequencing read includes only a part of the repeat and not the flanking regions of the repeat. Flanking reads include a part of the repeat and the flanking regions at one end of the repeat.
[0049] To identify pathogenic expansions in cWGS screening by manual inspection, all reads overlapping the repeat can be visualized by the EH visualization tool. Reads are examined to determine whether the read is pathogenic (AAGGG) n It can be examined whether the repeats contain a pathogenic repeat. Figure 2 shows the visualization of the reads of a human patient sample overlapping with the repeat. Human patient FAM-092-D12 was confirmed by PCR to have two extended AAGGG haplotypes. As shown in Figure 2, the reads were either in-repeat reads and flanking reads containing mainly AAGGG. No spanning reads were observed because the repeats were longer than the read length. However, manual inspection can be difficult and inaccurate. For example, it can be difficult to distinguish between two different haplotypes. Patterns similar to pathogenic repeats (e.g., AAAGG) can result in false positive calls. Figure 3 shows the visualization of the reads of a patient sample overlapping with the repeat. Human patient FAM-062-F08 was a putative patient. However, PCR did not validate the putative. As shown in Figure 3, it is difficult to tell with the naked eye what repeats these are. PCR suggested AAGAG repeats instead of AAGGG repeats.
[0050] Disclosed herein is a method to automatically and accurately screen for this pathogenic expansion from sequencing data, such as cWGS data. In graph-aligned reads, it is possible to count how many times each pentamer is observed in the repeat overlap. In some embodiments, the following criteria can be used to minimize noise from reduced data quality: each type of pentamer has two occurrences where all five bases have a quality (Q) equal to or greater than a quality threshold (e.g., 20). With reference to FIG. 4A, bases with a quality of 20 or greater are shown in uppercase. For example, of the four AAAAG / AAaAg / AaaaG pentamers shown in FIG. 4A, two had all five bases with a quality of 20 or greater (pentamers shown as AAAAG), one had four bases with a quality of 20 or greater (pentamers shown as AAaAG), and one had two bases with a quality of 20 or greater (pentamers shown as AaaaG). Four AAAAG / AAaAg / AaaaG pentamers meet this criteria and are counted. As another example, of the three AAGGG / AAGgG pentamers shown in FIG. 4A, two had all five bases with a quality of 20 or greater (pentamers shown as AAGGG) and one had four bases with a quality of 20 or greater (pentamer shown as AAGgG). All three AAGGG / AAGgG pentamers meet this criterion and are counted. As a further example, of the two remaining pentamers shown in FIG. 4A, one had all five bases with a quality of 20 or greater (pentamer shown as AGGGG) and one had four bases with a quality of 20 or greater (pentamer shown as AAGaG). These two pentamers do not meet the criterion that each type of pentamers has two occurrences with all five bases with a quality (Q) greater than or equal to the quality threshold of 20, since each type of pentamers has one occurrence. These two pentamers are filtered and not counted.
[0051] In some embodiments, the following criteria can be used to minimize noise from reduced data quality: each type of pentamer has two occurrences where all five bases have a quality (Q) equal to or greater than a quality threshold (e.g., 20), and each pentamer counted has at least three bases with a quality equal to or greater than a quality threshold (e.g., 20). With reference to FIG. 4B, bases with a quality equal to or greater than 20 are shown in capital letters. For example, of the four AAAAG / AAaAg / AaaaG pentamers shown in FIG. 4B, two had all five bases with a quality equal to or greater than 20 (pentamers shown as AAAAG), one had four bases with a quality equal to or greater than 20 (pentamer shown as AAaAG), and one had two bases with a quality equal to or greater than 20 (pentamer shown as AaaaG). The AAAAG / AAaAg pentamers meet the criteria and are counted. The AaaaG pentamer does not meet the criteria (as this pentamer has only two bases with a quality equal to or greater than 20) and is filtered out and not counted. As another example, of the three AAGGG / AAGgG pentamers shown in FIG. 4B, two had all five bases with a quality of 20 or greater (pentamers shown as AAGGG) and one had four bases with a quality of 20 or greater (pentamer shown as AAGgG). All three AAGGG / AAGgG pentamers meet the criteria and are counted. As a further example, of the two remaining pentamers shown in FIG. 4B, one had all five bases with a quality of 20 or greater (pentamer shown as AGGGG) and one had four bases with a quality of 20 or greater (pentamer shown as AAGaG). These two pentamers do not meet the criteria that each type of pentamer has two occurrences with all five bases with a quality (Q) equal to or greater than the quality threshold of 20, and are therefore filtered and not counted, since each type of pentamer has one occurrence.
[0052] Sequence reads from patient samples were processed and 5mer counts were performed using the standard. The actual patient had nearly 100% AAGGG (pathogenic repeats). With 5mer filtering and counting rules, 93% AAGGG was observed in PCR-validated patient FAM-062-F08 (see Figure 2 and accompanying legend). Other pentamers were AAAGG (2%) and GGGGG (5%). These other pentamers are most likely due to sequencing errors / artifacts.
[0053] As shown in Figures 5A-5B, carriers require better separation for the two haplotypes. In the reads from the extended allele, the carrier had nearly 100% AAGGG. For the patient who is a carrier with one extended allele, the sequence reads of the patient's sample were processed and 5mer filtered and counted using the criteria. By 5mer filtering the counting criteria, the patient was observed to have 75% AAGGG (Figure 5A). By 5mer filtering using the criteria and counting the reads from the extended allele, the patient was observed to have 93% AAGGG. With reference to Figures 6A-6B, the spanned reads are from the short allele. For adjacent reads with short overlaps with the repeat, it is uncertain which allele the read is from. The remaining flanking reads and in-repeat reads shown in Figure 6B are from the extended allele.
[0054] Samples from 77 patients showing signs of repeat expansion such as expected (earlier onset or more severe phenotype in subsequent generations) and / or neuropathic symptoms were subjected to cWGS. These patients were suspected to have some pathogenic RFC1 repeat expansion. None of them had clinical documentation of a known expansion. The failure to identify repeat expansions in these samples may be due to insufficient testing or associated with previously unknown pathogenic repeat expansions. A general ExpensionHunter screen identified RFC1 as "expanded" in a small number of patients. Previously, actual RFC1 patients were identified by manually examining sequence reads aligned one by one to the repeat. The disclosed method can be used for automated screening of RFC1 status (pathogenic, carrier, or benign and / or number of expansions such as 0, 1, or 2). Patterns of RFC1 expansions in the 77 patient cohort were examined. When screening for repeats in RFC1, we used 85% or more for AAAAG for the reference allele and 85% or more for AAGGG for the pathogenic allele. Repeat patterns were identified by looking at the expanded alleles. Figure 7 shows the percentage of 5mer repeats in the expanded allele for samples with one short allele and one expanded allele in the 75 patient cohort. Table 1 shows the results of screening the patient cohort using the following criteria: With two expanded alleles and greater than 85% AAGGG repeats, it is pathogenic. A person with one expanded allele and more than 85% AAGGG repeats or two expanded alleles and 30% of -85% AAGGG repeats is a carrier. No expanded alleles, one expanded allele and less than 85% AAGGG repeats, or two expanded alleles and less than 30% AAGGG repeats are benign.
[0055] [Table 1]
[0056] The patterns of RFC1 repeats were examined in the Polaris population of 150 unrelated healthy individuals. The Polaris population is a diverse group of European, East Asian, and African peoples. Figure 8 shows the proportion of 5mer repeats in the expanded allele for samples with one short allele and one expanded allele for this Polaris population. The analysis shows a high frequency for AAGGG and many other patterns. Table 1 shows the screening results for the Polaris population using the following criteria: With two expanded alleles and greater than 85% AAGGG repeats, it is pathogenic. A person with one expanded allele and more than 85% AAGGG repeats or two expanded alleles and 30% of -85% AAGGG repeats is a carrier. No expanded alleles, one expanded allele and less than 85% AAGGG repeats, or two expanded alleles and less than 30% AAGGG repeats are benign.
[0057] [Table 2]
[0058] Determining RFC1 repeat expansion status FIG. 9 is a flow diagram illustrating an exemplary method 900 for determining replication factor C subunit 1 (RFC1) repeat extension status. The method 900 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives of a computing system. For example, a computing system 1100, illustrated in FIG. 11 and described in more detail below, may execute a set of executable program instructions to perform the method 900. When the method 900 is initiated, the executable program instructions may be loaded into a memory, such as a RAM, and executed by one or more processors of the computing system 1100. Although the method 900 is described with respect to the computing system 1100 illustrated in FIG. 11, the description is merely exemplary and not intended to be limiting. In some embodiments, the method 900 or portions thereof may be executed by multiple computing systems, either serially or in parallel.
[0059] After the method 900 begins at block 904, the method 900 proceeds to block 908, where a computing system (e.g., computing system 1100 described with reference to FIG. 11) receives a plurality of sequence reads generated from a sample obtained from a subject. The sequence reads can be, for example, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1250, 1500, 1750, 2000 or more base pairs (bps) in length, respectively. For example, the sequence reads are each about 100 base pairs to about 1000 base pairs in length. The sequence reads can include paired-end sequence reads. The sequence reads can include single-end sequence reads. The sequence reads can be generated by targeted sequencing. The sequence reads can be generated by whole genome sequencing (WGS). The sequence reads can be generated by whole genome sequencing (WGS). The WGS can be clinical WGS (cWGS). The sample can include cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, blood samples, biopsy samples, or combinations thereof. The subject can be a human subject.
[0060] The received sequence reads may include only reads from the locus of RFC1. For example, the plurality of sequence reads may be aligned to the locus of RFC1. Alternatively, the received sequence reads may include reads from RFC1 and loci elsewhere. The computing system may align a second plurality of sequence reads including the plurality of sequence reads aligned to the locus of RFC1 to a reference genome sequence. The reference genome sequence may include a reference human genome sequence, such as hg19 or hg38. The computing system may select the plurality of sequence reads aligned to the locus of RFC1 from the second plurality of sequence reads.
[0061] The computing system can store the sequence reads in memory. The computing system can load the sequence reads into memory. The sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by ligation, or sequencing by ligation. The sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).
[0062] The method 900 proceeds from block 908 to block 912, where the computing system aligns the plurality of sequence reads to a sequence graph to generate a plurality of aligned sequence reads. The sequence graph can represent the locus of RFC1. The sequence graph can include a representation of a repeat sequence flanked by non-repeat sequences of the locus of RFC1 (e.g., AARRG, where R is A or G). The plurality of aligned sequence reads can include or be associated with the plurality of sequence reads and an alignment of the plurality of sequence reads to the sequence graph.
[0063] Method 900 proceeds from block 912 to block 916, where the computing system determines the number of occurrences of a plurality of repeat sequences (e.g., AAGGG, AAAAG, AAAGG, AAGAG, AACGG, and ACGGG) in the aligned sequence reads of the plurality of aligned sequence reads using a first occurrence threshold (e.g., 2) and a first quality threshold (e.g., 20). Determining or counting the number of occurrences of a plurality of repeat sequences in the aligned sequence reads of the plurality of aligned sequence reads using a first occurrence threshold and a first quality threshold is referred to herein as 5-mer filtering and counting. The 5-mer filtering and counting can be for both alleles or extended alleles.
[0064] The repeat expansion at the RFC1 locus can comprise one or more threshold total copies of the repeat sequence (e.g., 20 copies total). The threshold total copies of the repeat sequence(s) can be, for example, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more or less total copies of the repeat sequence(s).
[0065] In some embodiments, each of the multiple repeat sequences has a number of occurrences equal to or greater than a first occurrence threshold, and each occurrence has a number of bases each having a quality score equal to or greater than a first quality threshold. The first occurrence threshold can be, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. The first quality threshold can be, for example, about 10, 15, 20, 25, 30, 35, 40, 45, 50, or more or less. The number of bases of the repeat sequence each having a quality score greater than (or greater than) the first quality threshold can be, for example, 4 or 5. For example, each type of pentamer has two occurrences with all five bases having a quality (Q) equal to or greater than a quality threshold (e.g., the first quality threshold) of 20. In some embodiments, each occurrence has several bases each having a quality score equal to or greater than a second quality threshold. The number of bases each having a quality score greater than (or greater than) the second quality threshold can be, for example, 2, 3, 4, or 5. The second quality threshold can be, for example, about 10, 15, 20, 25, 30, 35, 40, 45, 50, or more or less. The number of bases in each occurrence having a quality score greater than the second quality threshold can be, for example, 2, 3, 4, or 5. For example, each 5-mer counted has at least 3 bases with a quality greater than or equal to a quality threshold (e.g., the second quality threshold) of 20. In some embodiments, the first quality threshold and the second quality threshold are the same.
[0066] The repeat sequence representation may be degenerate. The repeat sequence representation may be AARRG, where R is A or G. The repeat sequence representation may be at least 5 bases in length. The repeat sequence representation may be 5 bases in length. The repeat sequence representation may be 6 bases in length. Each of the plurality of repeat sequences may be at least 5 bases in length. Each of the plurality of repeat sequences may be 5 bases in length. Each of the plurality of repeat sequences may be 6 bases in length. The repeat sequence representation and the repeat sequence may have the same length.
[0067] From block 916, the method 900 proceeds to block 920, where a computing system determines a frequency representation of the pathogenic repeat sequence (e.g., AAGGG or ACAGG) or the number of occurrences of one or more pathogenic repeat sequences relative to the total number of occurrences of the plurality of repeat sequences. The pathogenic repeat sequence may be AAGGG or ACAGG. The pathogenic repeat sequence may have a GC content of at least 50%, 55%, 60%, 65%, 70%, 75%, 80% or more or less. The pathogenic repeat sequence may be at least 5 bases in length. The pathogenic repeat sequence may be 5 bases in length. The pathogenic repeat sequence may be 6 bases in length.
[0068] A frequency representation of the number of occurrences of a pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences can be a percentage of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences. A frequency representation of the number of occurrences of a pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences can be a ratio of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences.
[0069] Method 900 proceeds from block 920 to block 924, where the computing system determines the repeat expansion status (or repeat expansion status) at the RFC1 locus of the subject using a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences. The repeat expansion status at the RFC1 locus can be a pathogenic status, a carrier status, or a benign status. The computing system can determine that the subject has 0, 1, or 2 alleles with a repeat expansion at the RFC1 locus using the plurality of aligned sequence reads. The repeat expansion at the RFC1 locus can be at or near chr4:39348424 of hg38, or a corresponding location in another reference genome sequence. The repeat expansion can be associated with or cause a disease. The disease can be cerebellar ataxia, neuropathy, and vestibulo-ocular syndrome (CANVAS) or ataxia.
[0070] If both alleles of two expanded alleles are expanded, 5-mer counts of both expanded alleles can be performed. In some embodiments, a subject can have two alleles, each with a repeat expansion at the RFC1 locus.
[0071] Two expansion alleles - pathogenic status. To determine the repeat expansion status at the RFC1 locus, the computing system can determine that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is greater than (or equal to) a first status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or greater or less than 60%. The computing system can determine the repeat expansion status at the RFC1 locus as pathogenic status.
[0072] Two expansion alleles - carrier status. To determine the repeat expansion status at the RFC1 locus, the computing system can determine that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is equal to or less than a first status threshold and equal to or greater than a second status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more or less. The second status threshold can be, for example, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more or less. The computing system can determine the repeat expansion status at the RFC1 locus as carrier status.
[0073] Two expansion alleles-carrier status. To determine the repeat expansion status at the RFC1 locus, the computing system can determine that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is less than or equal to a first status threshold and is greater than or equal to a second status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more or less. The second status threshold can be, for example, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more or less.
[0074] The computing system can determine a frequency representation of (1) the pathogenic repeat sequence and (2) the number of occurrences of sequences with high sequence similarity to the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences. The sequence similarity of the pathogenic repeat sequence and the sequence with high sequence similarity to the pathogenic repeat sequence can be, for example, (about) 70%, 75%, 80%, 85%, 90%, 95% or more or less, or can be less. The pathogenic repeat sequence and the sequence similar to the pathogenic repeat sequence differ by one or more bases, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more bases. The sequence with high sequence similarity to the pathogenic repeat sequence can be non-pathogenic and / or associated (e.g., associated) with benign expansions.
[0075] A frequency representation of the number of occurrences of (1) pathogenic repetitive sequences and sequences with high sequence similarity to (2) pathogenic repetitive sequences relative to the total number of occurrences of the multiple repetitive sequences may be a percentage of the number of occurrences of (1) pathogenic repetitive sequences and sequences with high sequence similarity to (2) pathogenic repetitive sequences among the total number of occurrences of the multiple repetitive sequences. A frequency representation of the number of occurrences of (1) pathogenic repetitive sequences and sequences with high sequence similarity to (2) pathogenic repetitive sequences relative to the total number of occurrences of the multiple repetitive sequences may be a ratio of the number of occurrences of (1) pathogenic repetitive sequences and sequences with high sequence similarity to (2) pathogenic repetitive sequences among the total number of occurrences of the multiple repetitive sequences.
[0076] The computing system can determine a frequency representation of the number of occurrences of the pathogenic repeat sequence, where the number of sequences with high sequence similarity to the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is greater than (or equal to) a third status threshold, which can be, for example, 70%, 75%, 80%, 85%, 90%, 95%, or more or less than that.
[0077] The computing system can determine that the frequency representation of the number of occurrences of the pathogenic sequence is greater than (or equal to) a fourth status threshold. The fourth status threshold can be, for example, 5%, 10%, 15%, 20%, 25%, 30%, or greater or less than. Alternatively or additionally, the computing system can determine a frequency representation of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence. The frequency representation of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence relative to the total number of occurrences of the plurality of repeat sequences can be a percentage of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence among the total number of occurrences of the plurality of repeat sequences. The frequency representation of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence relative to the total number of occurrences of the plurality of repeat sequences can be a ratio of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence to the total number of occurrences of the plurality of repeat sequences. The computing system can determine that the frequency representation of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence is less than (or equal to) a fifth status threshold. The fifth status threshold can be, for example, 70%, 75%, 80%, 85%, 90%, 95%, or more or less. The computing system can determine the repeat expansion status at the RFC1 locus as the carrier status.
[0078] In some embodiments, the computing system determines that a frequency representation of the occurrence of the pathogenic sequence is less than or equal to a fourth status threshold. Alternatively or additionally, the computing system can determine that a frequency representation of the occurrence of sequences with high sequence similarity to the pathogenic sequence is greater than or equal to a fifth status threshold. The computing system can determine the repeat expansion status at the RFC1 locus as benign status.
[0079] For example, the pathogenic repeat sequence is AAGGG, and the sequence with high sequence similarity to the pathogenic repeat sequence is AAAGG. The two sequences have 80% sequence identity. The two sequences differ by one base. The proportion of AAGGG repeat sequences and AAAGG repeat sequences among all repeat sequences may be greater than 80%. The percentage of AAGGG repeat sequences among all repeat sequences may be greater than 10%. Alternatively, or in addition, the percentage of AAAGG repeat sequences among all repeat sequences may be less than or equal to 90%. The computing system can determine the repeat expansion status at the RFC1 locus as carrier status. If the proportion of AAGGG repeat sequences among all repeat sequences is less than or equal to 10% and / or the proportion of AAAGG repeat sequences among all repeat sequences is greater than 90%, the computing system can determine the repeat expansion status at the RFC1 locus as benign status.
[0080] Two expanded alleles - benign status. To determine the repeat expansion status at the RFC1 locus, the computing system can determine that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is less than (or equal to) a second status threshold. The second status threshold can be, for example, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more or less than that. The computing system can determine the repeat expansion status at the RFC1 locus as benign status.
[0081] One extended allele: If one allele is extended and one allele is not extended, 5mer counts of the extended allele can be performed. In some embodiments, the subject has one allele of RFC1 with a repeat expansion at the RFC1 locus.
[0082] One expansion allele - pathogenic status. To determine the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of sequence reads, the computing system can determine the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of aligned sequence reads derived from one allele of RFC1 having a repeat expansion at the locus of RFC1. To determine the repeat expansion status at the locus of RFC1, the computing system can determine that the frequency representation of the occurrence number of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is greater than (or equal to) a first status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more or less than. The computing system can determine the repeat expansion status at the locus of RFC1 as pathogenic status.
[0083] One expansion allele-carrier status. To determine the occurrence number of a plurality of repeat sequences in the aligned sequence reads of a plurality of sequence reads, the computing system can determine the occurrence number of a plurality of repeat sequences in the aligned sequence reads of a plurality of aligned sequence reads derived from one allele of RFC1 having a repeat expansion at the locus of RFC1. To determine the repeat expansion status at the locus of RFC1, the computing system can determine that the frequency representation of the occurrence number of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is equal to or less than a first status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more or less than that. The computing system can determine the repeat expansion status at the locus of RFC1 as a carrier status.
[0084] To determine the repeat expansion status at the locus of RFC1 as pathogenic status or carrier status when one expansion allele is present, the computing system can select aligned sequence reads derived from one allele of RFC1 that have a repeat expansion at the locus of RFC1 based on the alignment of the aligned sequence reads. The aligned sequence reads derived from one allele of RFC1 that have a repeat expansion at the locus of RFC1 can include aligned sequence reads that are (i) in-repeat reads or (ii) adjacent reads each having an overlap to a repeat expansion at the locus of RFC1 that is greater than a repeat expansion overlap threshold. The repeat expansion overlap threshold can be, for example, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90 or more or less base pairs. Aligned sequence reads from one allele of RFC1 with a repeat expansion at the RFC1 locus may include (i) all aligned sequence reads that are in-repeat reads, and (ii) sequence reads that are not all aligned sequence reads that are flanking reads. If the repeat is longer than the read length, the in-repeat reads may occur such that all reads (whether single-end or paired-end sequencing reads) or one read of the paired-end sequencing reads include only a portion of the repeat, but not the flanking region of the repeat. The flanking reads include a portion of the repeat and the flanking region at one end of the repeat.
[0085] No expansion allele If both alleles are not expanded, 5mer counts of both alleles can be performed.In some embodiments, the subject has 0 alleles with repeat expansion at the RFC1 locus.The repeat expansion status at the RFC1 locus can be benign status.
[0086] In some embodiments, a computing system can generate a user interface (UI), such as a graphical user interface, that includes or represents any results (including intermediate results) of method 900. For example, the UI can include or represent an iteration extension status. The UI can include, for example, a dashboard. The UI can include one or more UI elements. The UI element can include or represent an iteration extension status. The UI element can be a window (e.g., a container window, a browser window, a text terminal, a child window, or a message window), a menu (e.g., a menu bar, a context menu, or an add menu), an icon, or a tab. The UI element can be for an input control (e.g., a check box, a radio button, a drop-down list, a list box, a button, a toggle, a text field, or a date field). The UI element can be for navigation (e.g., a breadcrumb, a slider, a search field, a pagination, a slider, a tag, an icon). The UI element can provide information (e.g., a tooltip, an icon, a progress bar, a notification, a message box, or a modal window). The UI element can be a container (e.g., an accordion). In some embodiments, the computing system can generate a report that includes or represents any results (including intermediate results) of the method 900, such as the iterative extension status.
[0087] The computing system can cause one or more diagnostic methods to be performed to confirm the repeat expansion status at the RFC1 locus of the subject. The UI or report can include or indicate that one or more diagnostic methods should be performed to confirm the repeat expansion status at the RFC1 locus of the subject. The computing system can receive confirmation of the repeat expansion status at the RFC1 locus of the subject determined using the one or more diagnostic methods. The one or more diagnostic methods can include polymerase chain reaction (PCR) and Sanger sequencing, Southern blot, and linkage analysis.
[0088] In some embodiments, any threshold of method 900 (e.g., an occurrence threshold (e.g., a first occurrence threshold), a quality threshold (e.g., a first quality threshold, or a second quality threshold), a threshold total copies, a status threshold (e.g., a first status threshold, a second status threshold, a third status threshold, a fourth status threshold, or a fifth status threshold), or a repeat extension overlap threshold) can be determined using a large number of samples, such as 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, 5000, or more or less samples.
[0089] The method 900 ends at block 928.
[0090] Determining the repeat expansion status of a gene of interest FIG. 10 is a flow diagram illustrating an exemplary method 1000 of determining a repeat expansion status (also referred to herein as repeat expansion status) of a gene of interest. Method 1000 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives of a computing system. For example, a computing system 1100, shown in FIG. 11 and described in more detail below, may execute a set of executable program instructions to perform method 1000. When method 1000 is initiated, the executable program instructions may be loaded into a memory, such as a RAM, and executed by one or more processors of computing system 1100. Although method 1000 is described with respect to computing system 1100 shown in FIG. 11, the description is merely exemplary and not intended to be limiting. In some embodiments, method 1000 or portions thereof may be performed serially or in parallel by multiple computing systems.
[0091] After method 1000 begins at block 1004, method 1000 proceeds to block 1008, where a computing system (e.g., computing system 1100 described with reference to FIG. 11) receives a plurality of sequence reads generated from a sample obtained from a subject. The sequence reads can be, for example, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1250, 1500, 1750, 2000 or more base pairs (bps) in length, respectively. For example, the sequence reads are each about 100 base pairs to about 1000 base pairs in length. The sequence reads can include paired-end sequence reads. The sequence reads can include single-end sequence reads. The sequence reads can be generated by targeted sequencing. The sequence reads can be generated by whole genome sequencing (WGS). The sequence reads can be generated by whole genome sequencing (WGS). The WGS can be clinical WGS (cWGS). The sample can include cells, cell-free DNA, cell-free fetal DNA, amniotic fluid, blood samples, biopsy samples, or combinations thereof. The subject can be a human subject.
[0092] The received sequence reads may include only reads from the locus of the gene of interest. For example, the plurality of sequence reads may be aligned to the locus of the gene of interest. Alternatively, the received sequence reads may include reads from the locus of the gene of interest and other locations. The computing system may align the second plurality of sequence reads, including the plurality of sequence reads aligned to the locus of the gene of interest, to a reference genome sequence. The reference genome sequence may include a reference human genome sequence, such as hg19 or hg38. The computing system may select the plurality of sequence reads aligned to the locus of the gene of interest from the second plurality of sequence reads.
[0093] The computing system can store the sequence reads in memory. The computing system can load the sequence reads into memory. The sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by ligation, or sequencing by ligation. The sequence reads can be generated using instruments such as MINSEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).
[0094] Method 1000 proceeds from block 1008 to block 1012, where the computing system aligns the plurality of sequence reads to a sequence graph to generate a plurality of aligned sequence reads. The sequence graph can represent the locus of a gene of interest (e.g., RFC1). The sequence graph can include a representation of a repeat sequence flanked by non-repeated sequences of the locus of the gene (e.g., AARRG, where R is A or G if the gene of interest is RFC1). The plurality of aligned sequence reads can include a plurality of sequence reads. The plurality of aligned sequence reads can include or be associated with an alignment of the plurality of sequence reads to a sequence graph.
[0095] Method 1000 proceeds from block 1012 to block 1016, where the computing system determines the number of occurrences of a plurality of repeat sequences (e.g., AAGGG, AAAAG, AAAGG, AAGAG, AACGG, and ACGGG if the gene of interest is RFC1) in the aligned sequence reads of the aligned sequence reads using a first occurrence threshold (e.g., 2) and a first quality threshold (e.g., 20). Determining or counting the number of occurrences of a plurality of repeat sequences in the aligned sequence reads of the aligned sequence reads using a first occurrence threshold and a first quality threshold is referred to herein as n-mer filtering and counting, where n is the length of the repeat sequence. The n-mer filtering and counting can be for both alleles or extended alleles.
[0096] A repeat expansion at the locus of a gene of interest can comprise one or more threshold total copies of the repeat sequence. The threshold total copies of the repeat sequence(s) can be, for example, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more or less total copies of the repeat sequence(s).
[0097] In some embodiments, each of the multiple repeat sequences has a number of occurrences equal to or greater than a first occurrence threshold, and each occurrence has a number of bases each having a quality score equal to or greater than a first quality threshold. The first occurrence threshold can be, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. The first quality threshold can be, for example, about 10, 15, 20, 25, 30, 35, 40, 45, 50, or more or less. The number of bases of the repeat sequence each having a quality score greater than or equal to the first quality threshold can be, for example, 4, 5, 6, 7, 9, 9, 10 or more. For example, each type of pentamer has two occurrences with all five bases having a quality (Q) equal to or greater than a quality threshold (e.g., the first quality threshold) of 20. In some embodiments, each occurrence has several bases each having a quality score greater than or equal to a second quality threshold. The number of bases each having a quality score equal to or greater than the second quality threshold can be, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more. The second quality threshold can be, for example, about 10, 15, 20, 25, 30, 35, 40, 45, 50 or more or less. The number of bases in each occurrence having a quality score greater than the second quality threshold can be, for example, 2, 3, 4, or 5. For example, each pentamer counted has at least 3 bases with a quality equal to or greater than a quality threshold of 20 (e.g., the second quality threshold). In some embodiments, the first quality threshold and the second quality threshold are the same.
[0098] The repeat sequence representation can be degenerate. The repeat sequence representation can be AARRG, where R is A or G when the gene of the insert is RFC1. The repeat sequence representation can be 5, 6, 7, 8, 9, 10 or more bases in length. Each of the multiple repeat sequences can be 5, 6, 7, 8, 9, 10 or more bases in length. The repeat sequence representation and the repeat sequence can have the same length.
[0099] From block 1016, method 1000 proceeds to block 1020, where a computing system determines a frequency representation of the pathogenic repeat sequence (e.g., AAGGG or ACAGG if the gene of interest is RFC1) or the number of occurrences of one or more pathogenic repeat sequences relative to the total number of occurrences of the plurality of repeat sequences. The pathogenic repeat sequence can have a GC content of at least 50%, 55%, 60%, 65%, 70%, 75%, 80% or more or less. The pathogenic repeat sequence can be 5, 6, 7, 8, 9, 10 or more bases in length.
[0100] A frequency representation of the number of occurrences of a pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences can be a percentage of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences. A frequency representation of the number of occurrences of a pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences can be a ratio of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the multiple repeat sequences.
[0101] Method 1000 proceeds from block 1020 to block 1024, where the computing system determines a repeat expansion status at the locus of interest of the subject using a frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences. The computing system may determine that the subject has 0, 1, or 2 alleles with a repeat expansion at the locus of the gene of interest using the plurality of aligned sequence reads. The repeat expansion status at the locus of the gene of interest is a pathogenic status, a carrier status, or a benign status. The repeat expansion may be associated with or cause a disease. The disease may be, for example, tardive ataxia, cerebellar ataxia, sensory neuronopathy, bilateral vestibular disorder, or cerebellar ataxia, neuropathy, and vestibular anesthesia syndrome (CANVAS). The disease may be cancer, a non-cancerous disease, a neurological disease, a neurodegenerative disease, an autoimmune disease, Alzheimer's disease, Parkinson's disease, dementia, rheumatoid arthritis, or inflammation. The disease can be Huntington's disease, spinal and bulbar muscular atrophy, dentatorubral-pallidoluysian atrophy spinocerebellar ataxia, fragile X, fragile X tremor ataxia syndrome, other fragile sites, myotonic dystrophy type 1, Huntington's disease-like 2, spinocerebellar ataxia type 8, Fuchs corneal dystrophy, Friedreich's ataxia, FRAXE mental retardation, oculopharyngeal muscular dystrophy, myotonic dystrophy type 1, spinocerebellar ataxia type 10, spinocerebellar ataxia type 31, spinocerebellar ataxia type 36, frontotemporal dementia / amyotrophic lateral sclerosis, or EPM1 (myoclonus epilepsy). The gene of interest can be huntingtin, the androgen receptor (AR) gene, ATN1, ATXN1, ATXN2, ATXN3, ATXN10, CACNA1A, ATXN7, the TBP gene, PPP2R2B, TK2, BEAN, NOP56, JPH3, FRDA, CSTB, PABP2, TCF4, or C9ORF72.
[0102] In some embodiments, the gene of interest is replication factor C subunit 1 (RFC1). The locus of the gene of interest with the repeat expansion can be at or near chr4:39348424 of hg38, or the corresponding location of another reference genome sequence. The repeat expansion can be associated with or cause a disease. The disease can be cerebellar ataxia, neuropathy, and vestibulo-ocular reflex syndrome (CANVAS) or ataxia. The repeat expression can be AARRG. The pathogenic repeat can be AAGGG or ACAGG.
[0103] When two expanded alleles, both alleles, are expanded, 5-mer counts of both expanded alleles can be performed. In some embodiments, a subject may have two alleles, each with a repeat expansion at the locus of a gene of interest.
[0104] Two expansion alleles - pathogenic status. To determine the repeat expansion status at the locus of the gene of interest, the computing system can determine that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is greater than (or equal to) a first status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or greater or less than 60%. The computing system can determine the repeat expansion status at the locus of the gene of interest as pathogenic status.
[0105] Two expansion alleles - carrier status. To determine the repeat expansion status at the locus of the gene of interest, the computing system can determine that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is equal to or less than a first status threshold and equal to or greater than a second status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more or less. The second status threshold can be, for example, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more or less. The computing system can determine the repeat expansion status at the locus of the gene of interest as carrier status.
[0106] Two expansion alleles-carrier status. To determine the repeat expansion status at the locus of a gene of interest, the computing system can determine that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is less than or equal to a first status threshold and is greater than or equal to a second status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more or less. The second status threshold can be, for example, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more or less.
[0107] The computing system can determine a frequency representation of (1) the pathogenic repeat sequence and (2) the number of occurrences of sequences with high sequence similarity to the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences. The sequence similarity of the pathogenic repeat sequence and the sequence with high sequence similarity to the pathogenic repeat sequence can be, for example, (about) 70%, 75%, 80%, 85%, 90%, 95% or more or less, or can be less. The pathogenic repeat sequence and the sequence similar to the pathogenic repeat sequence differ by one or more bases, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more bases. The sequence with high sequence similarity to the pathogenic repeat sequence can be non-pathogenic and / or associated (e.g., associated) with benign expansions.
[0108] A frequency representation of the number of occurrences of (1) pathogenic repetitive sequences and sequences with high sequence similarity to (2) pathogenic repetitive sequences relative to the total number of occurrences of the multiple repetitive sequences may be a percentage of the number of occurrences of (1) pathogenic repetitive sequences and sequences with high sequence similarity to (2) pathogenic repetitive sequences among the total number of occurrences of the multiple repetitive sequences. A frequency representation of the number of occurrences of (1) pathogenic repetitive sequences and sequences with high sequence similarity to (2) pathogenic repetitive sequences relative to the total number of occurrences of the multiple repetitive sequences may be a ratio of the number of occurrences of (1) pathogenic repetitive sequences and sequences with high sequence similarity to (2) pathogenic repetitive sequences among the total number of occurrences of the multiple repetitive sequences.
[0109] The computing system can determine a frequency representation of the number of occurrences of the pathogenic repeat sequence, where the number of sequences with high sequence similarity to the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is greater than (or equal to) a third status threshold, which can be, for example, 70%, 75%, 80%, 85%, 90%, 95%, or more or less than that.
[0110] The computing system can determine that the frequency representation of the number of occurrences of the pathogenic sequence is greater than (or equal to) a fourth status threshold. The fourth status threshold can be, for example, 5%, 10%, 15%, 20%, 25%, 30%, or greater or less than. Alternatively or additionally, the computing system can determine a frequency representation of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence. The frequency representation of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence relative to the total number of occurrences of the plurality of repeat sequences can be a percentage of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence among the total number of occurrences of the plurality of repeat sequences. The frequency representation of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence relative to the total number of occurrences of the plurality of repeat sequences can be a ratio of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence to the total number of occurrences of the plurality of repeat sequences. The computing system can determine that the frequency representation of the number of occurrences of sequences with high sequence similarity to the pathogenic sequence is less than (or equal to) a fifth status threshold. The fifth status threshold can be, for example, 70%, 75%, 80%, 85%, 90%, 95%, or more or less. The computing system can determine the repeat expansion status at the locus of the gene of interest as the carrier status.
[0111] In some embodiments, the computing system determines that a frequency representation of the occurrence of the pathogenic sequence is less than or equal to a fourth status threshold. Alternatively or additionally, the computing system can determine that a frequency representation of the occurrence of sequences with high sequence similarity to the pathogenic sequence is greater than or equal to a fifth status threshold. The computing system can determine the repeat expansion status at the RFC1 locus as benign status.
[0112] Two expanded alleles - benign status. To determine the repeat expansion status at the locus of the gene of interest, the computing system can determine that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is less than (or equal to) a second status threshold. The second status threshold can be, for example, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, or more or less than that. The computing system can determine the repeat expansion status at the locus of the gene of interest as benign status.
[0113] 2. One extended allele If one allele is extended and one allele is not extended, 5mer counts of the extended allele can be performed. In some embodiments, the subject has one allele of the gene of interest that has a repeat expansion at the locus of the gene of interest.
[0114] One expansion allele - pathogenic status. To determine the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of sequence reads, the computing system can determine the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of aligned sequence reads derived from one allele of the gene of interest having a repeat expansion at the locus of RFC1. To determine the repeat expansion status at the locus of the gene of interest, the computing system can determine that the frequency representation of the occurrence number of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is greater than (or is greater than) a first status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more or less than. The computing system can determine the repeat expansion status at the locus of the gene of interest as pathogenic status.
[0115] One expansion allele-carrier status. To determine the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of sequence reads, the computing system can determine the occurrence number of the plurality of repeat sequences in the aligned sequence reads of the plurality of aligned sequence reads derived from one allele of the gene of interest having a repeat expansion at the locus of RFC1. To determine the repeat expansion status at the locus of the gene of interest, the computing system can determine that the frequency representation of the occurrence number of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is equal to or less than a first status threshold. The first status threshold can be, for example, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more or less than that. The computing system can determine the repeat expansion status at the locus of the gene of interest as a carrier status.
[0116] To determine the repeat expansion status at the locus of the gene of interest as pathogenic status or carrier status when one expansion allele is present, the computing system can select aligned sequence reads derived from one allele of the gene of interest having a repeat expansion at the locus of the gene of interest based on the alignment of the aligned sequence reads. The aligned sequence reads derived from one allele of the gene of interest having a repeat expansion at the locus of the gene of interest can include aligned sequence reads that are (i) in-repeat reads or (ii) adjacent reads each having an overlap to a repeat expansion at the locus of the gene of interest that is greater than a repeat expansion overlap threshold. The repeat expansion overlap threshold can be, for example, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90 or more or less base pairs. Aligned sequence reads from one allele of a gene of interest having a repeat expansion at the locus of the gene of interest may include (i) all aligned sequence reads that are in-repeat reads, and (ii) sequence reads that are not all aligned sequence reads that are flanking reads. If the repeat is longer than the read length, the in-repeat reads may occur such that all reads (whether single-end or paired-end sequencing reads) or one read of paired-end sequencing reads only include a part of the repeat, but not the flanking region of the repeat. Flanking reads include a part of the repeat and the flanking region at one end of the repeat.
[0117] No expansion alleles When both alleles are not expanded, 5mer counts of both alleles can be performed.In some embodiments, the subject has 0 alleles with repeat expansion at the locus of the gene of interest.The repeat expansion status at the locus of the gene of interest can be benign status.
[0118] In some embodiments, a computing system can generate a user interface (UI), such as a graphical user interface, that includes or represents any results (including intermediate results) of method 1000. For example, the UI can include or represent an iteration extension status. The UI can include, for example, a dashboard. The UI can include one or more UI elements. The UI element can include or represent an iteration extension status. The UI element can be a window (e.g., a container window, a browser window, a text terminal, a child window, or a message window), a menu (e.g., a menu bar, a context menu, or an add menu), an icon, or a tab. The UI element can be for an input control (e.g., a check box, a radio button, a drop-down list, a list box, a button, a toggle, a text field, or a date field). The UI element can be for navigation (e.g., a breadcrumb, a slider, a search field, a pagination, a slider, a tag, an icon). The UI element can provide information (e.g., a tooltip, an icon, a progress bar, a notification, a message box, or a modal window). The UI element can be a container (e.g., an accordion). In some embodiments, the computing system can generate a report that includes or represents any results (including intermediate results) of the method 900, such as the iterative extension status.
[0119] The computing system can cause one or more diagnostic methods to be performed to ascertain the repeat expansion status at the locus of the gene of interest of the subject. The UI or report can include or indicate that one or more diagnostic methods should be performed to ascertain the repeat expansion status at the locus of the gene of interest of the subject. The computing system can receive the conformation of the repeat expansion status at the locus of the gene of interest of the subject determined using one or more diagnostic systems. The one or more diagnostic systems can include polymerase chain reaction (PCR) and Sanger sequencing, Southern blot, and linkage analysis.
[0120] In some embodiments, any threshold of method 1000 (e.g., an occurrence threshold (e.g., a first occurrence threshold), a quality threshold (e.g., a first quality threshold, or a second quality threshold), a threshold total copies, a status threshold (e.g., a first status threshold, a second status threshold, a third status threshold, a fourth status threshold, or a fifth status threshold), or a repeat extension overlap threshold) can be determined using a large number of samples, such as 100, 200, 300, 400, 500, 1000, 2000, 3000, 4000, 5000, or more or less samples.
[0121] The method 1000 ends at block 1028.
[0122] Execution environment FIG. 11 illustrates a general architecture of an exemplary computing device 1100 configured to determine the repeat expansion status of a gene of interest (e.g., RFC1). The general architecture of the computing device 1100 illustrated in FIG. 11 includes an arrangement of computer hardware and software components. The computing device 1100 may include more (or less) elements than those illustrated in FIG. 11. However, not all of these general conventional elements need to be illustrated to provide a useful disclosure. As illustrated, the computing device 1100 includes a processing unit 1110, a network interface 1120, a computer-readable medium drive 1130, an input / output device interface 1140, a display 1150, and an input device 1160, all of which may communicate with each other via a communication bus. The network interface 1120 may provide connectivity to one or more networks or computing systems. The processing unit 1110 may therefore receive information and instructions from other computing systems or services via a network. The processing unit 1110 may also communicate with memory 1170 and further provide output information for an optional display 1150 via an input / output device interface 1140. The input / output device interface 1140 may also accept input from any input device 1160, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, game pad, accelerometer, gyroscope, or other input device.
[0123] Memory 1170 may include computer program instructions (grouped into modules or components in some embodiments) that processing unit 1110 executes to implement one or more embodiments. Memory 1170 generally includes RAM, ROM, and / or other persistent, secondary, or non-transitory computer-readable media. Memory 1170 may store an operating system 1172 that provides computer program instructions for use by processing unit 1110 in the overall management and operation of computing device 1100. Memory 1170 may further include computer program instructions and other information for implementing aspects of the disclosure.
[0124] For example, in one embodiment, memory 1170 includes a repeat extension status determination module 1174 for determining a repeat expansion status (e.g., pathogenic, carrier, or benign), such as method 900 described with reference to Figure 9 and method 1000 described with reference to Figure 10. Additionally, memory 1170 may include or be in communication with data store 1190 and / or one or more other data stores that store sequence reads being processed and the determined extension status, repeat extension status (and any intermediate results thereof).
[0125] Additional Considerations In at least some of the foregoing embodiments, one or more elements used in one embodiment may be used interchangeably in another embodiment, except where such an exchange is not technically feasible. Those skilled in the art will appreciate that various other omissions, additions, and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and variations are intended to be included within the scope of the subject matter, as defined by the appended claims.
[0126] Those skilled in the art will appreciate that, for this and other processes and methods disclosed herein, the functions performed in the processes and methods may be performed in differing orders. Moreover, the outlined steps and operations are provided only as examples, and some of the steps and operations may be optional, may be combined into fewer steps and operations, or may be expanded to additional steps and operations without departing from the essence of the disclosed embodiments.
[0127] For the use of substantially any plural and / or singular term herein, one of ordinary skill in the art may substitute the plural for the singular and / or the singular for the plural as appropriate to the context and / or application. For clarity, various singular / plural permutations may be expressly set forth herein. As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, phrases such as "a device configured to" are intended to include one or more enumerated devices. Such one or more enumerated devices may also be collectively configured to execute the recited detailed description. For example, "a processor configured to execute detailed description A, B, and C" may include a first processor configured to execute detailed description A and to perform operations in conjunction with a second processor configured to execute detailed description B and C. Any reference to "or" herein is intended to encompass "and / or" unless otherwise indicated.
[0128] Generally, the terms used herein, and particularly the terms used in the appended claims (e.g., the body of the appended claims), are generally intended as "open" terms (e.g., the term "including" should be interpreted as "including but not limited to", the term "having" should be interpreted as "having at least", the term "includes" should be interpreted as "includes but is not limited to", etc.). Those skilled in the art will further understand that where a specific number of introduced claim recitations are intended, such intention will be expressly recited in the claim, and in the absence of such recitation, no such intention exists. For example, to aid in understanding, the following appended claims may include the use of the introductory phrases "at least one" and "one or more" to introduce the claim recitations. However, the use of such phrases should not be construed as limiting any particular claim that includes such an introduced claim statement to embodiments that include only one of such statements, even when the same claim includes "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted to mean "at least one" or "one or more"), and the same applies to the use of indefinite articles used to introduce claim statements. Moreover, even when a specific number of introduced claim statements is explicitly recited, one of skill in the art will recognize that such a recitation should be interpreted in the sense of at least the recited number (e.g., the unmodified recitation of "two statements" means, in the absence of other modifications, at least two statements or more than two statements).Furthermore, when phrases similar to "at least one of A, B, and C, etc." are used, generally such structures are intended in the sense that one of ordinary skill in the art would understand the phrase (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems having only A, only B, only C, both A and B, both A and C, both B and C, and / or both A, B, and C, etc.). When phrases similar to "at least one of A, B, or C, etc." are used, generally such structures are intended in the sense that one of ordinary skill in the art would understand the phrase (e.g., "a system having at least one of A, B, or C" includes, but is not limited to, systems having only A, only B, only C, both A and B, both A and C, both B and C, and / or both A, B, and C, etc.). Those skilled in the art will further appreciate that virtually any disjunction and / or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."
[0129] Additionally, where features or aspects of the disclosure are described in terms of a Markush group, one of skill in the art will thereby recognize that the disclosure is also described in terms of any individual members or subgroups of the Markush group members.
[0130] As will be understood by those skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of those subranges. Any recited range can be readily recognized as fully descriptive and allowing the same range to be broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range described herein can be readily broken down into a lower third, a middle third, and an upper third, etc. Also, as will be understood by those skilled in the art, all language such as "up to," "at least," "greater than," "less than" and the like refers to a range that includes the recited numbers and can then be broken down into subranges as described above. Finally, as will be understood by those skilled in the art, a range includes each individual component. Thus, for example, a group having 1 to 3 items means a group having 1, 2, or 3 items. Similarly, a group having 1-5 items means a group having 1, 2, 3, 4, or 5 items, etc.
[0131] Various embodiments of the present disclosure have been described herein for illustrative purposes, and it will be understood that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various embodiments disclosed herein are not intended to be limiting, the true scope and spirit of which is set forth by the following claims.
[0132] It should be understood that not all objectives or advantages may necessarily be achieved in accordance with any particular embodiment described herein. Thus, for example, those skilled in the art will recognize that a particular embodiment may be configured to operate in a manner that achieves or optimizes an advantage or set of advantages as taught herein, without necessarily achieving other objectives or advantages that may be taught or suggested herein.
[0133] All of the processes described herein may be embodied in, and may be fully automated via, software code modules executed by a computing system including one or more computers or processors. The code modules may be stored on any type of non-transitory computer-readable medium or other computer storage device. Some or all of the methods may be embodied in dedicated computer hardware.
[0134] Many other variations beyond those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain acts, events, or functions of any of the algorithms described herein may be performed in a different order, added, combined, or omitted entirely (e.g., not all described acts or events are necessary to implement the algorithm). Furthermore, in certain embodiments, acts or events may be performed simultaneously rather than sequentially, e.g., via multi-threading, interrupt processing, or multiple processors or processor cores, or on other parallel systems. In addition, different tasks or processes may be performed by different machines and / or computing systems that can function together.
[0135] The various example logic blocks and modules described in connection with the embodiments disclosed herein may be implemented or performed by mechanical devices such as processing units or processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor may be a microprocessor, but alternatively the processor may be a controller, microcontroller, or state machine, combinations thereof, and the like. The processor may include electrical circuitry configured to process computer-executable instructions. In another embodiment, the processor includes an FPGA or other programmable device that performs logical operations without processing computer-executable instructions. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in association with a DSP core, or any other such configuration. Although primarily digital technology is described herein, the processor may also include primarily analog components. For example, some or all of the signal processing algorithms described herein may be implemented in analog circuitry or mixed analog and digital circuitry. The computing environment may include any type of computer system, including, but not limited to, a computer system based on a computing engine within a microprocessor, mainframe computer, digital signal processor, portable computing device, device controller, or appliance, to name a few.
[0136] Any process illustrations, elements or blocks in the flow diagrams described herein and / or shown in the accompanying drawings should be understood as potentially representing modules, segments or portions of code that include one or more executable instructions for implementing a particular logical function or element in the process. Alternate embodiments are included within the scope of the embodiments described herein, and depending on the functionality involved, elements or functions may be omitted or removed from the order from that shown or discussed, including substantially simultaneously or in reverse order, as will be understood by those skilled in the art.
[0137] It should be emphasized that many variations and modifications may be made to the above-described embodiments, and that the elements thereof should be understood to be among the other acceptable examples. All such modifications and variations are intended to be included herein within the scope of this disclosure and protected by the following claims.
Claims
1. A system for determining the repeat expansion status of a target gene, comprising: A non-transitory memory configured to store executable instructions; and A hardware processor communicating with the non-transitory memory, the hardware processor programmed by the executable instructions, wherein the hardware processor: (a) receives a plurality of sequence reads generated from a sample obtained from a subject; (b) aligns the plurality of sequence reads to an array graph to generate a plurality of aligned sequence reads, the array graph representing a locus of a target gene, the non-repetitive sequence of the locus of the gene including a repetitive sequence representation adjacent thereto, and the plurality of aligned sequence reads including the plurality of sequence reads and the alignment of the plurality of sequence reads to the array graph; (c) using a first occurrence threshold and a first quality threshold, determines the number of occurrences of a plurality of repetitive sequences in the aligned sequence reads among the plurality of aligned sequence reads; (d) determines a frequency representation of the number of occurrences of a pathogenic repetitive sequence relative to the total number of occurrences of the plurality of repetitive sequences; and (e) determines the repeat expansion status at the locus of the target gene of the subject using the frequency representation of the number of occurrences of the pathogenic repetitive sequence relative to the total number of occurrences of the plurality of repetitive sequences, a hardware processor communicating with the non-transitory memory.
2. The system according to claim 1, wherein the target gene is replication factor C subunit 1 (RFC1).
3. The system according to claim 2, wherein the repeat expansion is associated with or causes a disease, and optionally, the disease is cerebellar ataxia, neuropathy, and vanishing reflex syndrome (CANVAS).
4. The system according to claim 2, wherein the repetitive sequence representation is AARRGG.
5. The system according to claim 2, wherein the pathogenic repeat sequence is AAGGG or ACAGG.
6. The system according to claim 1, wherein the hardware processor is programmed by the executable instructions to use the plurality of aligned sequence reads to determine that the subject has zero, one, or two alleles having a repeat expansion at the locus of the gene of interest.
7. The system according to claim 1, wherein the subject has two alleles having a repeat expansion at the locus of the gene of interest.
8. Determining the repeat expansion status at the locus of the gene of interest comprises determining that the frequency representation of the number of occurrences of the pathogenic repeat sequence relative to the total number of occurrences of the plurality of repeat sequences is greater than a first status threshold; determining the repeat expansion status at the locus of the gene of interest as a pathogenic status; The system according to claim 7, comprising:
9. The system according to claim 1, wherein the subject has one allele of the gene of interest having a repeat expansion at the locus of the gene of interest.
10. The system according to claim 1, wherein the subject has zero alleles having a repeat expansion at the locus of the gene of interest, and the repeat expansion status at the locus of the gene of interest is a benign status.
11. The system according to claim 1, wherein the repeat expansion status at the locus of the gene of interest is a pathogenic status, a carrier status, or a benign status.
12. Each of the plurality of repetitive arrays has an occurrence number equal to or greater than the first occurrence threshold, and each occurrence has a number of bases each having a quality score equal to or greater than the first quality threshold, the system according to claim 1.
13. Each of the occurrences has a number of bases each having a quality score equal to or greater than a second quality threshold, the system according to claim 1.
14. The repetitive array representation is degenerate, the system according to any one of claims 40 to 66.
15. Each of the repetitive array representation and / or the plurality of repetitive arrays is at least 5 bases in length, and optionally, each of the repetitive array representation and / or the plurality of repetitive arrays is 6 bases in length, the system according to claim 1.
16. The pathogenic repetitive array has a GC content of at least 60%, the system according to claim 1.
17. The frequency representation of the occurrence number of the pathogenic repetitive array with respect to the total occurrence number of the plurality of repetitive arrays is the percentage of the occurrence number of the pathogenic repetitive array among the total occurrence number of the plurality of repetitive arrays, or the ratio of the occurrence number of the pathogenic repetitive array to the total occurrence number of the plurality of repetitive arrays, the system according to claim 1.
18. The plurality of sequence reads are aligned to the locus of the target gene, the system according to claim 1.
19. The step of receiving the plurality of sequence reads generated from the obtained sample aligning a second plurality of sequence reads including the plurality of sequence reads to a reference genome sequence selecting the plurality of sequence reads from the second plurality of sequence reads, wherein the plurality of sequence reads are aligned to the locus of the target gene comprising, the system according to claim 1.
20. The executable instructions are programmed to cause the hardware processor to receive a conformation of the repeat extension status at the locus of the target's target gene determined using one or more diagnostic systems, and optionally, the one or more diagnostic systems include polymerase chain reaction (PCR) and Sanger sequencing, Southern blot, and linkage analysis, the system of claim 1.