Fragile x syndrome AGG interruption genotyping

The FXS AGG PCR assay with targeted primer and capillary electrophoresis accurately detects AGG breaks and CGG repeats, addressing inefficiencies in conventional methods and enhancing genotyping precision for Fragile X Syndrome risk assessment.

JP2025134816APending Publication Date: 2025-09-17LABORATORY CORPORATION OF AMERICA HOLDINGS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025100116
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-01-30
Filing Date
2025-06-16
Publication Date
2025-09-17

AI Technical Summary

Technical Problem

Conventional PCR methods struggle to accurately detect and analyze AGG interruptions in CGG repeat expansions, leading to reduced PCR efficiency and high error rates in genotyping for Fragile X Syndrome, particularly in determining the risk of late-onset neurodegenerative disorders and transmission of full mutant alleles.

Method used

The FXS AGG PCR assay uses a primer with a 3' 'A' nucleotide to target CGG repeats and detect AGG breaks, combined with capillary electrophoresis for peak separation, followed by a computer-implemented method to iteratively search for AGG peaks and calculate CGG repeat numbers, providing a more accurate AGG genotype and risk assessment.

Benefits of technology

This approach enhances the accuracy of Fragile X Syndrome genotyping by precisely identifying AGG breaks and CGG repeat numbers, reducing errors and improving the prediction of neurodegenerative disease risk and allele transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025134816000001_ABST
    Figure 2025134816000001_ABST
Patent Text Reader

Abstract

To provide Fragile X Syndrome (FXS) clinical testing, in particular to provide a FXS AGG interruption polymerase chain reaction (PCR) assay and an AGG interruption genotyping algorithm for implementation into clinical testing.SOLUTION: An AGG genotyping method comprises obtaining raw data from a FXS assay performed on a sample, iteratively searching the raw data and identifying one or more AGG peaks on a first allele using a first set of search spaces determined based on an expected AGG peak size, determining a number of CGG repeats downstream of a final AGG interruption and a number of CGG repeats preceding a first AGG interruption on the first allele based on the one or more AGG peaks, and generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the final AGG interruption and the number of CGG repeats preceding the first AGG interruption.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Priority claim This application claims the benefit of and priority to U.S. Provisional Application No. 62 / 967,792, filed January 30, 2020, which is incorporated herein by reference in its entirety for all purposes.

[0002] Field The present disclosure relates to Fragile X Syndrome (FXS or FRAX) clinical trials, particularly the FRAX AGG break polymerase chain reaction (PCR) assay and AGG break genotyping algorithm for implementation in clinical trials. [Background technology]

[0003] background PCR and related amplification techniques can be used for analytical purposes. Typical analytical applications of PCR include the diagnosis of genotype status or determination of genotypes involving polymorphic loci. An example of a locus that exhibits a medically relevant polymorphism is the 5' untranslated region (UTR) of the human FMR1 gene on the X chromosome. Normal individuals typically have 5-44 CGG repeats at this locus. In contrast, alleles at this locus containing large CGG repeat expansions (>200 repeats, full mutant alleles) disrupt FMR1 gene expression and cause FXS. Furthermore, individuals with the premutation (PM) allele (50-200 repeats) are at risk for developing the late-onset neurodegenerative disorders Fragile X-associated tremor / ataxia syndrome (FXTAS) or Fragile X-associated primary ovarian insufficiency (FXPOI). Female PM allele carriers (pan-ethnic frequency 1 in 201) are at risk for transmitting the full mutant allele to their offspring. The risk depends on the size of the CGG repeat, as measured by the fragile X PCR assay, and the number of AGG interruptions between the CGG repeats (usually occurring every 9–11 CGG repeats).

[0004] When a PM allele with 55-90 repeats is transmitted by a carrier female, the risk of expansion to a full mutation in offspring decreases with increasing number of AGG interruptions in the CGG repeat sequence. For PM alleles with more than 90 repeats, the expansion risk is greater than 50% regardless of the number of AGG interruptions. Therefore, the AGG interruption PCR assay may have the greatest clinical utility in characterizing AGG interruptions for PM alleles with 55-90 CGG repeats. Summary of the Invention [Means for solving the problem]

[0005] Abstract In various embodiments, a method includes obtaining raw data from a Fragile X Syndrome (FXS) assay performed on a sample, the raw data including gene-specific (GS) polymerase chain reaction (PCR) data and AGG break PCR data separated by capillary electrophoresis; determining an expected AGG peak size for a first allele identified in the raw data; iteratively searching the raw data to identify one or more AGG peaks on the first allele using a first set of search spaces determined based on the expected AGG peak sizes; determining the number of CGG repeats downstream of the last AGG break on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; A computer-implemented method is provided that includes determining the number of CGG repeats preceding the first AGG break on the first allele based on the PCR data and one or more AGG peaks identified on the first allele; generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break; and providing the AGG genotype for the first allele.

[0006] In some embodiments, the method further includes determining the number of CGG repeats separated by any two adjacent AGG breaks on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, and an AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG break, the number of CGG repeats separated by any two adjacent AGG breaks, and the number of CGG repeats preceding the first AGG break.

[0007] In some embodiments, the method includes iteratively searching the raw data to identify one or more AGG peaks on the second allele using a second set of search spaces determined based on expected AGG peak sizes; determining the number of CGG repeats downstream of the last AGG break on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; and determining the number of CGG repeats preceding the first AGG break on the second allele based on the PCR data and the one or more AGG peaks identified on the second allele; generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break; and providing an AGG genotype for the second allele, wherein (i) the first allele is a normal allele and the second allele is a pre-mutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a pre-mutation allele and the second allele is a different pre-mutation allele.

[0008] In some embodiments, the method further includes determining a risk score for the subject associated with the sample based on the AGG genotype generated for the first allele, the second allele, or both the first and second alleles, wherein the risk score identifies the subject's risk of developing a late-onset neurodegenerative disease, Fragile X-associated Tremor / Ataxia Syndrome (FXTAS) or Fragile X-associated Primary Ovarian Insufficiency (FXPOI), or the risk of transmitting the full mutant allele to the subject's offspring, or any combination thereof.

[0009] In some embodiments, iteratively searching the raw data to identify one or more AGG peaks on the first allele comprises determining a first search space of a first set of search spaces for the first allele based on an expected AGG peak size; searching the raw data and identifying an initial AGG peak on the first allele using the first search space; determining a second search space for the first allele based on a peak size of the initial AGG peak on the first allele; and iteratively searching the raw data and identifying one or more additional AGG peaks on the first allele using the second search space, wherein the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break are determined for the first allele based on the GS PCR data, the initial peak identified on the first allele, and the one or more additional peaks identified on the first allele.

[0010] In some embodiments, iteratively searching the raw data to identify one or more AGG peaks on the second allele comprises determining a first search space of a second set of search spaces for the second allele based on an expected AGG peak size; searching the raw data and identifying an initial AGG peak on the second allele using the second search space; determining a second search space for the second allele based on a peak size of the initial AGG peak on the second allele; and iteratively searching the raw data and identifying one or more additional AGG peaks on the second allele using the second search space, wherein the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break are determined for the second allele based on the GS PCR data, the initial peak identified on the second allele, and the one or more additional peaks identified on the second allele.

[0011] In some embodiments, once one or more AGG peaks on a first allele are identified, the raw data is iteratively searched to remove one or more AGG peaks on the first allele from the raw data before identifying one or more AGG peaks on a second allele.

[0012] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods or processes disclosed herein.

[0013] In some embodiments, a computer program product is provided that is tangibly embodied in a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of one or more methods disclosed herein.

[0014] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium including instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium including instructions configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0015] The terms and expressions which have been employed are used as terms of description rather than of limitation, and there is no intention in the use of such terms and expressions to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention as claimed. Thus, while the present invention has been specifically disclosed by embodiments and optional features, it will be understood that variations and modifications of the concepts disclosed herein may be employed by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0016] The invention will be better understood in view of the following non-limiting figures. [Brief explanation of the drawings]

[0017] [Figure 1A] FIG. 1A shows a schematic diagram of the FRAX AGG and gene-specific (GS) PCR assays according to various embodiments and the locations of the GS and AGG PCR primers.

[0018] [Figure 1B]FIG. 1B shows a representative electropherogram from a FRAX AGG PCR assay, with each peak corresponding to an AGG break in the sample according to various embodiments.

[0019] [Figure 2] FIG. 2 shows a block diagram of an FXS AGG PCR assay platform for disruption testing of PM allele carriers according to various embodiments.

[0020] [Figure 3] FIG. 3 shows an exemplary flow for AGG genotyping using the FXS AGG PCR assay platform and genotyping technology, according to various embodiments.

[0021] [Figure 4A] 4A-4H show an exemplary flow chart for AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques according to various embodiments. [Figure 4B] 4A-4H show an exemplary flow chart for AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques according to various embodiments. [Figure 4C] 4A-4H show an exemplary flow chart for AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques according to various embodiments. [Figure 4D] 4A-4H show an exemplary flow chart for AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques according to various embodiments. [Figure 4E] 4A-4H show an exemplary flow chart for AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques according to various embodiments. [Figure 4F]4A-4H show an exemplary flow chart for AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques according to various embodiments. [Figure 4G] 4A-4H show an exemplary flow chart for AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques according to various embodiments. [Figure 4H] 4A-4H show an exemplary flow chart for AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques according to various embodiments.

[0022] [Figure 5] FIG. 5 illustrates an exemplary computing device in accordance with various embodiments.

[0023] [Figure 6A] Figures 6A and 6B show FRAX AGG PCR CE long-injection electropherograms for samples S1-S5. The numbers following each sample name indicate the total number of repeats for each allele. Green and red peak size annotations represent AGG peaks assigned to the normal and PM alleles, respectively, according to various embodiments. [Figure 6B] Figures 6A and 6B show FRAX AGG PCR CE long-injection electropherograms for samples S1-S5. The numbers following each sample name indicate the total number of repeats for each allele. Green and red peak size annotations represent AGG peaks assigned to the normal and PM alleles, respectively, according to various embodiments.

[0024] [Figure 7A-1] FIG. 7A shows the FRAX TRP-PCR assay CE electropherograms of samples S2-S4, demonstrating two AGG breaks in the normal allele of each sample according to various embodiments. [Figure 7A-2]FIG. 7A shows the FRAX TRP-PCR assay CE electropherograms of samples S2-S4, demonstrating two AGG breaks in the normal allele of each sample according to various embodiments.

[0025] [Figure 7B] FIG. 7B shows a CE electropherogram of an AGG PCR assay using gel-purified GS-PCR products from each allele of sample S2, according to various embodiments.

[0026] [Figure 8A] FIG. 8A shows a FRAX AGG PCR assay in which two and five AGG breaks were identified in the normal and PM alleles, according to various embodiments.

[0027] [Figure 8B] FIG. 8B shows a TRP-PCR CE electropherogram indicating that the 159 bp AGG peak is present in both alleles, according to various embodiments.

[0028] [Figure 8C] FIG. 8C shows a CE electropherogram of a FRAX AGG PCR analyzing the separated normal and PM alleles of sample S17 and identifying two and six AGG breaks in the normal and PM alleles, respectively, according to various embodiments.

[0029] [Figure 9A] FIG. 9A shows a FRAX AGG PCR assay in which three and four AGG breaks were identified in the normal and PM alleles, according to various embodiments.

[0030] [Figure 9B] FIG. 9B shows a TRP-PCR CE electropherogram showing that the 159 bp AGG peak is present in either the normal or PM allele, according to various embodiments.

[0031] [Figure 9C] Figure 9C shows a CE electropherogram of FRAX AGG PCR using gel-purified GS-PCR products from each allele of sample S23, identifying two and five AGG breaks in the normal and PM alleles, respectively, according to various embodiments.

[0032] [Figure 10] FIG. 10 shows a FRAX TRP-PCR CE electropherogram of an AGG-free male PM carrier sample, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0033] In the accompanying drawings, similar components and / or features may have the same reference label. Furthermore, various components of the same type may be distinguished by following the reference label with a dash and a second label that distinguishes between the similar components. When only a first reference label is used herein, the description is applicable to any similar component having the same first reference label, regardless of the second reference label.

[0034] Detailed Description The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It will be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the appended claims.

[0035] In the following description, specific details are set forth to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

[0036] Also, it should be noted that particular embodiments may be described as a process that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. While a flowchart or diagram may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Moreover, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to a calling function or a main function. I. Introduction

[0037] Conventional systems use CGG repeat-primed (TRP)-PCR, separated by CE, to identify alleles of loci containing large CGG repeat expansions. In the presence of AGG interruptions, the affinity of the CGG repeat primer for the target sequence decreases, thereby reducing PCR efficiency and resulting in a "dip" in the CE electropherogram. Based on TRP-PCR CE data, it may be possible to indirectly calculate the number and location of each AGG. However, it is very difficult to detect AGG interruptions located at similar distances from the start of the CGG repeat region in a stepwise and accurate manner. To overcome these drawbacks, some conventional systems utilize GGA-primed PCR for direct AGG detection, followed by manual genotyping analysis. However, GGA-primed PCR can generate very high background and typically does not meet more stringent quality standards. Furthermore, conventional manual genotyping analysis performs poorly and is highly error-prone.

[0038] To address these limitations and issues, the FXS AGG PCR assay described herein targets the CGG repeat region of the 5' UTR of the FMR1 gene and detects all AGG breaks within the CGG repeats on all alleles. The FXS AGG PCR assay utilizes a primer containing both the CGG repeat and an "A" nucleotide at its 3' end that specifically binds to the AGG break in the CGG repeat, as well as a reverse primer for detection (see, e.g., Figure 1A). Capillary electrophoresis (CE) can be used to separate these amplicons as individual peaks corresponding to the position of the AGG break on the electropherogram (see, e.g., Figure 1B). In some instances, one or more end primers (e.g., forward and / or reverse primers) can be labeled with a fluorescent tag (e.g., fluorescein amidite (FAM)) to aid in separation by CE. In other instances, fluorescent labeling can be provided by adding a labeled base that is incorporated during PCR extension. In other examples, fluorescent labels are not used, and PCR products are generated with untagged primers that can be separated by high-resolution CE. Because AGGs typically occur with a periodicity of approximately 30 bp within repeat regions when located on the same allele, phasing of AGGs is possible using the results of both FXS AGG PCR and gene-specific (GS) PCR assays. Peak identification techniques can be used to identify peaks in the raw data of FXS GS and AGG PCR CE. The raw data may include the size and abundance of each analyzed PCR amplicon as peak size (base pairs) and height (e.g., relative fluorescence units (RFU)), respectively. To facilitate data analysis, the FXS AGG genotyping technique described herein processes the raw data of FXS GS and AGG PCR CE to determine the AGG genotype of each allele.

[0039] The FXS AGG genotyping technique first checks the sizing standard and peak height to ensure that the raw data quality of the FXS GS and AGG PCR CE is acceptable for analysis. The FXS AGG genotyping technique then uses the FXS GS and AGG PCR CE raw data to calculate the number of CGG repeats for each allele (e.g., normal and PM alleles). To determine the number and location of AGG breaks for each allele, the technique further includes iteratively searching for and identifying AGG peaks on alleles with fewer repeats, which usually correspond to normal alleles. After removing all AGG peaks assigned to this allele from the peak output table, the technique further includes repeating the process for the PM allele. In this way, the same peak is prevented from being assigned to both alleles, which could otherwise underestimate the elongation risk of the PM allele. Once peaks are identified and assigned to alleles, the technique further includes evaluating the AGG genotype on each allele. In some embodiments, the number of CGG repeats downstream of the last AGG break on each allele is calculated using the following formula: (P - 132) ÷ 3 - 1 (P: peak size of the smallest AGG break on each allele). In some embodiments, the number of CGG repeats separated by any two adjacent AGG breaks (if present) is calculated using the following formula: (P n -P n-1 )÷3-1(P n -P n-1 (The size of the peak of adjacent AGG breaks is used to denote the size of the peak of adjacent AGG breaks.) In some embodiments, the number of CGG repeats preceding the first AGG break is determined by subtracting the number of CGG repeats downstream of the final AGG, the total number of CGG repeats separating adjacent AGG breaks, and the total number of AGG breaks from the total number of CGG repeats of the allele as determined by PCR assay.

[0040] As used herein, the terms "substantially," "approximately," and "about" are defined as that which is largely, but not necessarily fully specified (and including that which is fully specified), as would be understood by one of ordinary skill in the art. In any disclosed embodiment, the terms "substantially," "approximately," or "about" can be replaced with "within [a percentage]" of what is specified, where percentage includes 0.1, 1, 5, and 10%. As used herein, when an action is "based on" something, this means that the action is at least partially based on at least a portion of something.

[0041] The FXS AGG genotyping techniques disclosed herein provide a more accurate prediction of other types of FXS compared to the FXS GS and AGG PCR CE raw data specifically described herein. It will be understood that the PCR CE raw data can be evaluated. It will also be understood that other types of FXS PCR are contemplated to identify CGG repeats and AGG breaks. For example, alternatively or additionally, non-anchor primed PCR can be used to identify CGG repeats and / or AGG breaks. II. FXS AGG PCR Assay Technology

[0042] One or more embodiments described herein can be implemented using program modules, engines, or components. A program module, engine, or component can include a program, subroutine, portion of a program, or software or hardware component that can perform one or more described tasks or functions. As used herein, a module or component can exist on a hardware component independently of other modules or components. Alternatively, a module or component can be a shared element or process of other modules, programs, or machines. Figure 2 shows a block diagram of an FXS AGG PCR assay platform 200 for disruption testing of PM allele carriers (e.g., samples with 55-90 CGG repeat expansions) and illustrates modules, engines, or components (e.g., programs, code, or instructions) executable by one or more processors that can be used to implement various subsystems of an analyzer system 205 according to various embodiments. The analyzer system 205 is a genetic analyzer (also known as a DNA sequencer), an automated system that can sequence DNA and analyze fragments for various applications. In some examples, the genetic analyzer is a capillary electrophoresis-based system in which probe-bound DNA fragments migrate through a polymer and fluorescence emission is measured. An array of multiple capillaries allows for sample loading in a multiwell microplate format. In other examples, the genetic analyzer is a pyrosequencing technology-based system for rapid sequencing and analysis. Pyrosequencing is a DNA sequencing method based on the "sequencing-by-synthesis" principle, in which sequencing is performed by detecting nucleotides incorporated by DNA polymerase. Pyrosequencing relies on optical detection of pyrophosphate release, which is a chain reaction. The module, engine, or component may be stored on a non-transitory computer medium.As needed, one or more of the modules, engines, or components can be loaded into system memory (e.g., RAM) and executed by one or more processors of the analyzer system 205. In the example shown in Figure 2, modules, engines, or components for implementing a genetic mapper subsystem 210 and an AGG genotyping subsystem 215 are shown.

[0043] 2 also shows a wet lab subsystem 220 that includes a laboratory where chemicals, drugs, or other materials or biological substances are tested and analyzed requiring water, direct ventilation, and dedicated plumbing utilities. The FXS AGG PCR assay platform 200 includes obtaining one or more samples 225 in the wet lab subsystem 220 in block 227. In some examples, batches of samples 225 are processed simultaneously, e.g., more than 20, more than 50, or more than 100 samples can be processed simultaneously using the FXS AGG PCR assay platform 200. For example, multiple well assay plates, such as 96-well plates or 384-well plates, can be processed simultaneously using the FXS AGG PCR assay platform 200. The AGG PCR assay platform 200 can be used to simultaneously analyze more than 20, more than 50, or more than 100 samples. In certain examples, multiple assay plates can be constructed and then arranged in a plate stacker of a capillary electrophoresis apparatus, allowing the assay plates to be sequentially loaded. In some examples, the sample 225 contains nucleic acid. In some examples, the sample 225 contains nucleic acid obtained from a female patient. In some examples, the sample 225 is whole blood or amniotic fluid containing nucleic acid obtained from a female patient. In certain examples, the FRAX AGG PCR assay is a reflex assay from a fragile X PCR diagnostic assay, and the reflex assay is triggered when the fragile X PCR diagnostic assay identifies one or more samples 225 having at least 45-100 CGG repeats (e.g., 55-90 CGG repeats) in the PM allele of the FMR1 gene. In certain examples, the reflex assay is triggered when the fragile X PCR diagnostic assay identifies one or more samples 225 having 55-90 repeats in the PM allele of the FMR1 gene.

[0044] In block 230 within the wet lab subsystem 220, FXS AGG PCR assays are performed, including GS PCR 235 and AGG interrupted PCR 240 (AGG PCR). The assays are performed by diluting the sample 225 (e.g., using Tris-HCl), subjecting the diluted sample 225 to GS PCR in the assay system, and then performing the GS PCR in the assay system. Mixing reagents / primers for GS PCR 235 and AGG PCR 240 (e.g., PCR tubes or well plates), GS PCR 235 and AGG PCR This may include PCR amplification (e.g., reaction cycling), post-PCR cleanup (e.g., purification of PCR products), and PCR product separation (e.g., capillary electrophoresis or a high-resolution gel such as Lonza MetaPhor Agarose) for both GS PCR and AGG PCR. In some instances, the FXS AGG PCR assay is performed in a PCR plate (e.g., a 96-well PCR assay plate). 240 may be performed in parallel on the same PCR plate, and controls may be loaded onto the PCR plate for quality control. GS PCR 235 is used to amplify the CGG repeat-containing region of the FMR1 gene and determine the length of the repeated CGG trinucleotide sequence within the FMR1 gene. Due to the high GC content of this region, a GC-rich PCR system can be prepared and used for reliable and robust GS PCR amplification. In some instances, the GC-rich PCR system includes a GC-rich PCR reaction buffer and a GC-rich resolution solution. The GC-rich PCR system may also include an enzyme blend of thermostable Taq DNA polymerase and Tgo DNA polymerase, which are thermostable enzymes with proofreading (3'-5' exonuclease) activity. The GC-rich PCR reaction buffer, GC-rich resolution solution, and dimethyl sulfoxide (DMSO) as an additive enable robust amplification of this difficult CGG repeat region. AGG PCR 240 targets the CGG repeat region and detects all AGG interruptions within it on all alleles. The AGG PCR system is based on the GC-rich PCR system, except that Taq DNA polymerase and Tgo DNA polymerase can be replaced by other DNA polymerases, such as KAPA2G Robust HotStart DNA polymerase, for robust amplification of AGG-specific sequences with minimal background.

[0045] DNA polymerases such as Taq DNA polymerase, Tgo DNA polymerase, and KAPA2G Robust HotStart DNA polymerase can only generate DNA when given a primer, a short nucleotide sequence that provides a starting point for DNA synthesis. In some examples, GS PCR 235 uses the forward primer FRAX-F1 and the reverse primer FRAX-R-6 to flank the target region (the CGG repeat region to be copied). In some examples, GS PCR 235 uses a fluorescently labeled reverse primer, such as FRAX-R-6FAM. In some examples, AGG PCR 240 uses the forward primer FXS-AGG-Forw, a chimeric primer containing three CGG repeats and an "A" nucleotide at the 3' end. The forward primer FXS-AGG-Forw specifically binds to the AGG break within the CGG repeat, thereby producing amplicons of various sizes. In some examples, AGG PCR 240 uses the same fluorescently labeled reverse primer, such as FRAX-R-6FAM, as GC PCR 235. In another example, the AGG PCR 240 uses a different fluorescently labeled reverse primer than that of the GC PCR 235, for example, FRAX-R-*FAM.

[0046] In some instances, post-PCR cleanup is performed using GS PCR 235 and AGG This includes mixing the PCR 240 products with magnetic beads, washing with a wash solution such as 70% ethanol, air-drying, and eluting the purified PCR products 245 to increase the signal-to-noise ratio. After amplification and post-PCR cleanup, the purified PCR products 245 can be loaded into an analyzer system 205 (e.g., a fluorescence-based separation instrument system), and amplification products such as CGG repeats and AGG interruptions are determined in block 250 and output as GS and AGG PCR CE raw data 255. In some examples, GS PCR 235 is optimized to detect large CGG repeats; for example, up to 300 CGG repeats (e.g., 262 repeats) can be consistently amplified and detected by GS PCR 235. In certain examples, the purified PCR products are detected or separated using CE, and the amplicons are visualized as "stutter" peaks, each separated by one CGG repeat. In certain embodiments, a first parameter, such as a long injection time or increased voltage, is used to detect larger size alleles (e.g., greater than 69 CGG), while a second parameter, such as a short injection time or decreased voltage, is used for accurate sizing of smaller size alleles (e.g., 69 CGG or less). Expanded alleles may be detected when the stutter extends beyond 55 CGG repeats, and PM alleles may be detected when the stutter extends between 45 and 200 CGG repeats.

[0047] In block 260 within the analyzer system 205, a separation module 265 (e.g., any software and / or hardware capable of acquiring raw CE data and determining fragment size and abundance, such as GeneMapper™) is used to identify peaks in the GS and AGG PCR CE raw data 255 based on previously established QC metrics. In some examples, both the AGG PCR assay and the GS PCR assay are identified using the same parameters (e.g., the same long injection or increased voltage) of the CE analysis method in the genetic mapper module 265 to identify peaks in the GS and AGG PCR CE raw data 255. The genetic mapper module 265 outputs the size and abundance of each analyzed PCR amplicon as peak size and height, respectively, which are then processed by the AGG genotyping module 275 of the AGG genotyping subsystem 215 in block 270. The AGG genotyping module 260 first checks sizing criteria and peak heights to ensure that the quality of the GS and AGG PCR CE raw data 255 is acceptable for analysis. The AGG genotyping module 275 then uses the GS-PCR peak data to calculate the number of CGG repeats for each allele. To determine the number and location of AGG breaks for each allele, the AGG genotyping module 275 begins by iteratively searching for and identifying AGG peaks on alleles with fewer repeats, which usually correspond to normal alleles. After removing all AGG peaks assigned to this allele from the peak output table of the gene mapper module 265, the AGG genotyping module 275 repeats the process for one or more PM alleles. In this way, the same peak is prevented from being assigned to both normal and PM alleles, which could otherwise underestimate the risk of extension of one or more PM alleles. The AGG genotyping module 275 then determines the AGG genotype of each allele using the number of CGG repeats for each allele and the number and location of AGG breaks for each allele.The AGG genotype for each allele and optional risk result for each sample is output by the analyzer system 205 as the final result 280. In some examples, all thresholds and QC parameters used by the AGG genotyping module 275 are maintained in a separate configuration file and can be used across any number of FXS AGG PCR assays. III. AGG genotyping techniques

[0048] FIG. 3 illustrates a process 300 for AGG genotyping using an FXS AGG PCR assay platform and genotyping technology (e.g., the FXS AGG PCR assay platform 200 described with reference to FIG. 2). Process 300 begins at block 305, where raw data is obtained from an FXS assay performed on a sample. In some examples, the raw data includes GS PCR data and AGG disruption PCR data separated by CE. In block 310, an expected AGG peak size is determined for a first allele identified in the raw data. In some examples, the first allele is the allele with the fewest number of repeats or a normal allele. In certain examples, the expected AGG peak size is calculated using equation (1). Expected AGG peak size (base pairs) = ((NL)*3)+A Formula (1) where N is the number of CGG repeats in the allele (e.g., the allele with the fewest number of repeats), L is the expected position (in number of CGG repeats) of the first AGG break, which may be between 8 and 10 CGG repeats, for example 9 CGG repeats, 3 is the number of base pairs for each CGG repeat, and A is the size of the PCR amplicon without the repeats; in some examples, the PCR amplicon has 132 base pairs, so A=132.

[0049] In blocks 315 and 320, the raw data is iteratively searched to identify one or more AGG peaks on the first allele using a first set of search spaces determined based on expected AGG peak sizes. The first set of search spaces may include one or more search spaces determined based on expected AGG peak sizes. In some examples, the searching and identifying includes (i) in block 315, determining a first search space of the first set of search spaces for the first allele based on expected AGG peak sizes; (ii) in block 315, searching the raw data and identifying an initial AGG peak on the first allele using the first search space; and (iii) in block 320, determining a second search space for the first allele based on the peak size of the initial AGG peak on the first allele and iteratively searching the raw data and identifying one or more additional AGG peaks on the first allele using the second search space. If the search does not result in identifying an initial AGG peak on the first allele using the first search space in block 315, then the first search space may be modified in block 325 (e.g., 30 bps may be subtracted from the expected AGG peak size), and blocks 315 and 320 may be performed using the modified first search space.

[0050] The first search space can be determined based on a first peak size range. In some examples, the first peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the expected AGG peak size. In particular examples, the predetermined amount or size range adjustment threshold is + / - 12 or 15 base pairs of the expected initial AGG peak size. The second search space can be determined based on a new or second peak size range. In some examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the peak size of the identified initial peak. In particular examples, the predetermined amount or size range adjustment threshold is + / - 12 or 15 base pairs of the peak size of the identified initial peak. In other examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the revised peak size (the peak size adjustment threshold is subtracted from the peak size of the identified initial peak). In some embodiments, the peak size adjustment threshold is 15-50 bps (e.g., 30 bps). In particular examples, the predetermined amount or size range adjustment threshold is + / - 12 or 15 base pairs of the modified peak size (24 or 30 bps is subtracted from the peak size of the identified initial peak).

[0051] The process of blocks 315, 320, and 325 may be repeated for each additional allele (e.g., one or more PM alleles) until all peak data and alleles in the raw data have been evaluated. For example, the raw data may be iteratively searched to identify one or more AGG peaks on a second allele using a second set of search spaces determined based on expected AGG peak sizes. The second set of search spaces may include one or more search spaces determined based on expected AGG peak sizes. In certain examples, some or all of the one or more search spaces in the second set of search spaces are the same as the one or more search spaces in the first set of search spaces. In some examples, the searching and identifying includes: (i) determining a first search space of a second set of search spaces for the second allele based on the expected AGG peak size; (ii) searching the raw data and identifying an initial AGG peak on the second allele using the second search space; (iii) determining a second search space for the second allele based on the peak size of the initial AGG peak on the second allele; and (iv) iteratively searching the raw data and identifying one or more additional AGG peaks on the second allele using the second search space. If the search results in no initial AGG peak on the second allele being identified using the first search space, the first search space may be modified (e.g., 24 bps may be subtracted from the expected AGG peak size), and processing may continue with the modified first search space.

[0052] The second search space can be determined based on a second peak size range. In some examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the expected AGG peak size. In particular examples, the predetermined amount or size range adjustment threshold is + / - 12 or 15 base pairs of the expected initial AGG peak size. The second search space can be determined based on a new or second peak size range. In some examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the peak size of the identified initial peak. In particular examples, the predetermined amount or size range adjustment threshold is + / - 12 or 15 base pairs of the peak size of the identified initial peak. In other examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the corrected peak size (the peak size adjustment threshold is subtracted from the peak size of the identified initial peak). In some embodiments, the peak size adjustment threshold is 15-50 bps (e.g., 24 bps). In particular examples, the predetermined amount or size range adjustment threshold is + / - 12 or 15 base pairs of the modified peak size (24 or 30 bps is subtracted from the peak size of the identified initial peak). In some examples, once one or more AGG peaks on a first allele or any other previous allele are identified, the raw data can be iteratively searched to remove one or more AGG peaks on the first allele or any other previous allele from the raw data before identifying one or more AGG peaks on a second allele or another subsequent allele.

[0053] In block 330, the genotype of each allele evaluated in blocks 315, 320, and 325 is AGG genotyped. In some examples, genotyping includes (i) determining the number of CGG repeats downstream of the last AGG break on the first allele, the second allele, or any other allele based on the GS PCR data and one or more AGG peaks identified in the corresponding allele, (ii) determining the number of CGG repeats (if present) separated by any two adjacent AGG breaks on the first allele, the second allele, or any other allele based on the GS PCR data and one or more AGG peaks identified in the corresponding allele, and (iii) determining the number of CGG repeats preceding the first AGG break on the first allele, the second allele, or any other allele based on the GS PCR data and one or more AGG peaks identified in the corresponding allele. In some examples, genotyping further includes generating an AGG genotype of the first allele, the second allele, or any other allele based on the number of CGG repeats downstream of the last AGG break, the number of CGG repeats separated by any two adjacent AGG breaks (if present), and the number of CGG repeats preceding the first AGG break.

[0054] Optionally, in block 335, the genotypes of one or more alleles determined in block 330 (e.g., the first allele, the second allele, or both the first and second alleles) are used to determine a subject risk score associated with the sample. The risk score identifies the subject's risk of developing the late-onset neurodegenerative diseases Fragile X-associated Tremor / Ataxia Syndrome (FXTAS) or Fragile X-associated Primary Ovarian Insufficiency (FXPOI), or the risk of transmitting the full mutant allele to the subject's offspring, or any combination thereof. In block 340, the AGG genotypes determined for one or more alleles determined in block 330 can be output. Outputting the AGG genotypes and optional risk scores determined for one or more alleles determined in block 330 can include providing the output to an end user and / or recording the output to a storage device (e.g., displaying the output on a user interface and / or storing the output in a database results file).

[0055] 4A-4H show a simplified flowchart 400 illustrating an example of a process for AGG genotyping using an FXS AGG PCR assay platform and genotyping technology (e.g., the FXS AGG PCR assay platform 200 and the technology described with respect to FIGS. 2 and 3). Process 400 is divided into the following sections: (i) data preparation (FIG. 4A), (ii) algorithm flow for the shortest or normal allele (FIGS. 4B and 4C), (iii) algorithm flow for the PM allele (FIGS. 4C, 4D, 4E, 4F, and 4G), and (iv) AGG genotyping flow for each allele (FIG. 4H). Data preparation

[0056] FIG. 4A shows that, in block 402, GS and AGG PCR CE raw data, including the size and abundance (e.g., peak size and height) of each PCR amplicon, are obtained for one or more samples. In some examples, initial CGG datasets (portions of the GS and AGG PCR CE raw data) are created for each of a first parameter (e.g., short injection time) analysis and a second parameter (e.g., long injection time) analysis. The first parameter analysis can return up to three alleles, and the second parameter analysis can return up to two alleles. A final CGG dataset can be created that includes a combination of the initial CGG datasets created for each of the first parameter analysis and the second parameter analysis. For example, (i) peaks for short CGGs (0-59 repeats) may be checked to ensure that duplicates are not selected; (ii) peaks for long CGGs (60-90 repeats) may be checked to ensure that long alleles do not overlap with short alleles (e.g., a short call of 58 repeats and a long call of 60 repeats); and (iii) peaks for long CGGs may be checked to ensure that duplicates are not selected. In certain examples, a peak buffer of + / - 1-5 CGG repeats, e.g., 2 CGG repeats, may be used during the check for overlapping peaks of short and long CGGs. Once overlaps and duplications are checked and removed, a final CGG dataset containing short and long CGGs may be created. In some examples, a final AGG dataset (a portion of the GS and AGG PCR CE raw data) is created that includes a combination of the first and second parameter analyses.

[0057] In block 404, the sizing criteria and peak heights in the final CGG and AGG datasets are checked to confirm that the raw data quality of the GS and AGG PCR CEs is acceptable for analysis. In some examples, the check can include comparing the sizing criteria and peak heights in the final CGG and AGG datasets to one or more quality control (QC) parameters or metrics to determine whether the raw data quality of the GS and AGG PCR CEs is acceptable for analysis. In particular examples, the QC parameters or metrics include a QC check threshold of a red dye (size marker) height ratio of 125 / 1000 greater than 2 (>2). Additionally or alternatively, the AGG-specific peak height QC parameter or metric can include a minimum RFU threshold for screening AGG-specific peaks. Normal alleles can have an RFU threshold greater than 1000 RFU (>1000). On the other hand, if desired, the PM allele may have an RFU threshold of >10% of the lowest AGG peak in the normal allele, or >200 if no peak is detected in the normal allele. If the sizing criteria and peak heights in the final CGG and AGG datasets satisfy the QC parameters or metrics, the final CGG dataset is used to calculate the number of CGG repeats for each allele (e.g., the normal or short allele and the PM or long allele) in block 406. If the sizing criteria and peak heights in the final CGG and AGG datasets do not satisfy the QC parameters or metrics, the sample's assay, the GS and AGG PCR CE raw data, and / or the size and abundance (e.g., peak size and height) of each PCR amplicon from the GS and AGG PCR CE raw data are rejected in block 408, and the process ends.

[0058] At block 406, a determination is made as to whether the calculated number of CGG repeats from each allele is within a predetermined range. In some examples, the FXS AGG PCR assay is a reflex assay from a fragile X PCR diagnostic assay, and the reflex assay is triggered when the fragile X PCR diagnostic assay identifies one or more samples having at least 45-100 CGG repeats (e.g., 55-90 CGG repeats) in the PM allele of the FMR1 gene. In particular examples, the reflex assay is triggered when the fragile X PCR diagnostic assay identifies one or more samples having at least 45-100 CGG repeats (e.g., 55-90 CGG repeats) in the PM allele of the FMR1 gene. The PCR diagnostic assay is initiated upon identification of one or more samples 225 having 55-90 repeats in the PM allele of the FMR1 gene. Thus, a minimum CGG threshold (MinCGG) of 40-60 repeats and a maximum CGG threshold (MaxCGG) of 75-110 repeats can be set to ensure AGG genotyping is performed only on one or more samples having at least 40-110 CGG repeats (e.g., 55-90 CGG repeats) in the PM allele of the FMR1 gene. In certain embodiments, the MinCGG threshold is set at 55 repeats and the MaxCGG threshold is set at 90 repeats, with the predetermined range for the number of CGG repeats being between 55 and 90 repeats (the expansion risk for PM alleles with more than 90 repeats is greater than 50% regardless of the number of AGG interruptions). If the number of CGG repeats identified in the final CGG dataset is within a predetermined range, the AGG genotyping process continues with the normal or short allele algorithm flow shown in FIG. 4B. If the number of CGG repeats identified in the final CGG dataset is outside the predetermined range, the number of CGG repeats is output at block 410 and the process ends. At block 410, the number of CGG repeats can be output along with a message stating that AGG genotyping was not performed. Outputting the number of CGG repeats and optional message can include providing the output to an end user and / or recording the output to a storage device (e.g., displaying the output on a user interface and / or storing the output in a database results file). Normal Allele Algorithm Flow

[0059] To determine the number and location of AGG breaks for each allele, the process begins by iteratively searching for and identifying AGG peaks on alleles with fewer repeats, which usually correspond to normal alleles. As shown in Figure 4B, in block 412, the expected AGG peak size is calculated for the allele. The expected AGG peak size may be calculated for the allele with the fewest number of CGG repeats (e.g., normal allele). In some examples, the expected AGG peak size is calculated using formula (1): Expected AGG peak size (base pairs) = ((NL)*3)+A Formula (1) where N is the number of CGG repeats in the allele (e.g., the allele with the fewest number of repeats), L is the expected position (in number of CGG repeats) of the first AGG break, which may be between 8 and 10 CGG repeats, for example 9 CGG repeats, 3 is the number of base pairs for each CGG repeat, and A is the size of the PCR amplicon without repeats; in some examples, the PCR amplicon has 132 base pairs without repeats, so A=132.

[0060] In block 414, an initial peak is identified for the allele (e.g., the allele with the smallest number of repeats). Identifying the initial peak includes determining an initial peak size search space for the allele (peak size search space in base pairs) in block 416. The initial peak size search space can be determined based on an initial peak size range. In some examples, the initial peak size range is a predetermined amount of base pairs or size range adjustment threshold (+ / -) around the expected AGG peak size (range = + / size range adjustment threshold). In particular examples, the predetermined amount or size range adjustment threshold is + / - 15 base pairs around the expected AGG peak size. Once the initial peak size search space is determined, in block 418, the initial peak size search space can be used to search the final AGG dataset to identify the highest peak within the initial peak size search space based on a highest peak criterion. In some examples, the criteria for identifying the highest peak in the initial peak size search space include (i) a peak size greater than a minimum peak size threshold (in some embodiments, the minimum peak size threshold is set at 132 base pairs), (ii) a peak height greater than a minimum peak height threshold (in some embodiments, the minimum peak height threshold is set at 200 RFU), (iii) a peak height greater than a background peak height threshold (in some embodiments, the background peak height threshold is set at 1000 RFU), or (iv) any combination thereof. The criteria for identifying the highest peak in the initial peak size search space provide various technical advantages, including: (i) the minimum peak size threshold is the size of a repeat-free PCR amplicon and can be used as a criterion for preventing false peak calling; (ii) the minimum peak height threshold can be used to prevent background low-level peak calling; and (iii) the background peak height threshold can be used as the minimum peak height required for an AGG peak in a normal allele and can also be used to prevent background peak calling.

[0061] If a peak that meets the highest peak criteria is found within the initial peak size search space, the found peak is identified as an initial peak in block 420. Furthermore, in block 420, the height of the initial peak is set to (i) the height of the previous peak height threshold, (ii) the height of the lowest normal height threshold, and (iii) the height of the first normal peak threshold. This setting provides various technical advantages, including preventing the calling of background or stutter (false positive) peaks. Once an initial peak is identified for an allele (e.g., the allele with the fewest number of repeats), the peak data associated with the initial peak can be removed from the final AGG dataset. After removing the AGG peak assigned to this allele from the output table of the final AGG dataset, the process can either repeat for another search space within the same allele or start a new process for another allele (e.g., the PM allele). Removing discovered AGG peaks (in this block or any other block described herein) assigned to alleles from the output table of the final AGG dataset achieves the technical advantage of preventing the same peak in the output table of the final AGG dataset from being discovered twice and / or assigned to multiple alleles, which may otherwise underestimate the elongation risk of PM alleles. Removal of data from the output table also reduces the complexity and time of downstream processing.

[0062] If no peaks meeting the highest peak criteria are found within the initial peak size search space, then in block 422, a peak size adjustment threshold is subtracted from the expected AGG peak size determined for the allele in block 422 (or a previously calculated new expected AGG peak size) to obtain a new expected AGG peak size. The new expected AGG peak size is used to identify initial peaks by repeating block 414. This adjustment shifts / expands the initial peak size search space to help find initial peaks within the final AGG dataset. In some embodiments, the peak size adjustment threshold is 15-50 bps (e.g., 30 bps).

[0063] Once an initial peak for an allele is identified in block 414, additional AGG peaks may be identified for the allele (e.g., the allele with the fewest number of repeats). In block 424, additional peaks are identified for the allele. Identifying the additional peaks includes subtracting a peak size adjustment threshold from the peak size of the initial peak found in block 414 to obtain a next expected AGG peak size in block 426. In block 428, a next peak size search space (peak size search space in base pairs) is then determined for the allele based on the expected AGG peak size. The next peak size search space can be determined based on a next peak size range around the next expected AGG peak size. In some examples, the next peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the next expected AGG peak size. In a particular example, the predetermined amount or size range adjustment threshold is + / - 15 base pairs of the next expected AGG peak size.

[0064] Once the next peak size search space is determined, the second search space can be used to search the final AGG data set to identify the highest peak in the next peak size search space based on highest peak criteria in block 430. In some examples, the criteria for identifying the highest peak in the next peak size search space include: (i) the peak size is greater than a minimum peak size threshold (in some embodiments, the minimum peak size threshold is set to 132 base pairs), (ii) the peak height is greater than a minimum peak height threshold (in some embodiments, the minimum peak height threshold is set to 200 RFU), (iii) the peak height is greater than a background peak height threshold (in some embodiments, the background peak height threshold is set to 1000 RFU), (iv) if an initial peak is found in the initial peak size search space, the peak height of the additional found peak is the previous peak height threshold (set in block 420) * a secondary normal height percent threshold (e.g., 50%), or (v) any combination thereof. The criteria for identifying the highest peak in the following peak size search space offer various technical advantages, including: (i) a minimum peak size threshold, which is the size of a repeat-free PCR amplicon, can be used as a criterion to prevent calling false peaks; (ii) a minimum peak height threshold can be used to prevent calling background low-level peaks; and (iii) a background peak height threshold can be used as the minimum peak height required for an AGG peak in a normal allele.

[0065] If a peak is found within the next peak size search space that meets the highest peak criterion, the found peak is identified as an additional peak in block 432. Furthermore, the height of the additional peak is (i) set to the height of the previous peak height threshold; or (ii) if the height of the found peak is smaller than the height of the initial peak, the height of the found peak is set to the height of the minimum normal height threshold. This setting provides various technical advantages, including preventing the calling of background or stutter (false positive) peaks. Once an additional peak is identified for an allele (e.g., the allele with the fewest number of repeats), the peak data associated with the additional peak can be removed from the final AGG dataset. After removing the AGG peak assigned to this allele from the output table of the final AGG dataset, the process can either repeat for another search space within the same allele or begin a new process for another allele (e.g., the PM allele).

[0066] In block 434, a determination is made as to whether the final AGG dataset should continue to be searched for additional peaks (whether the output table of the final AGG dataset has more peaks for the short allele). In some embodiments, the decision to continue searching the final AGG dataset is made based on whether the PCR amplicon size limit (e.g., 132 bp) for the short / normal allele has been reached. If it is determined that there are additional peaks in the final AGG dataset for the allele, block 424 is repeated. In block 426, a peak size adjustment threshold is subtracted from the previously calculated next predicted AGG peak size to obtain the next predicted AGG peak size (essentially the third, fourth, fifth, etc. expected AGG peak size). This adjustment shifts / expands the additional peak size search space to help find additional peak(s) in the final AGG dataset. Blocks 424 and 434 are repeated until all peaks of the allele (e.g., the allele with the fewest number of repeats) have been identified in the final AGG dataset. Once all peaks for alleles (eg, alleles with the fewest number of repeats) have been identified in the final AGG dataset, the AGG genotyping process continues for each of the remaining alleles at block 440. PM Allele Algorithm Flow

[0067] In block 440, an initial peak is identified for another allele in the final AGG dataset (e.g., a PM allele having a number of repeats between 55 and 90). Identifying the initial peak includes determining an initial peak size search space (peak size search space in base pairs) for the another allele in block 442. The initial peak size search space can be determined based on an initial peak size range. In some examples, the initial peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the expected AGG peak size determined in block 412. In particular examples, the predetermined amount or size range adjustment threshold is + / - 12 base pairs of the expected AGG peak size. In particular examples, for another allele (e.g., a PM allele) having a number of CGG repeats greater than 65: if the lower end of the initial peak size range is less than (less than) the first normal peak threshold set in block 420, the lower end of the initial peak size range is set to the first normal peak threshold.

[0068] Once the initial peak size search space is determined, the final AGG data set can be searched using the initial peak size search space to identify a highest peak in the initial peak size search space based on highest peak criteria in block 444. In some examples, the criteria for identifying a highest peak in the initial peak size search space include (i) a peak size greater than a minimum peak size threshold (in some embodiments, the minimum peak size threshold is set at 132 base pairs), (ii) a peak height greater than a minimum peak height threshold (in some embodiments, the minimum peak height threshold is set at 200 RFU), (iii) a peak height greater than a minimum normal height threshold (set in block 420 or 432) * a minimum percentage PM threshold (in some embodiments, the minimum percentage PM threshold is set at 10%), or (iv) any combination thereof. The criteria for identifying the highest peak in the initial peak size search space offer various technical advantages, including: (i) a minimum peak size threshold, the size of a repeat-free PCR amplicon, can be used as a criterion to prevent false peak calling; (ii) a minimum peak height threshold can be used to prevent background low-level peak calling; and (iii) using the previously identified peak height as the threshold for AGG peak search in PM alleles provides built-in dynamic DNA input control for AGG peak height in PM alleles. More specifically, AGG peak size in PM alleles is highly variable due to the 55-90 repeat range, and the corresponding AGG peak height also varies dramatically, which is also greatly influenced by variable amounts of DNA input. Lower DNA input results in lower peak height in normal alleles, and vice versa. Instead of using a fixed peak height threshold screening for AGG peaks in PM alleles, this dynamic threshold allows the process to more accurately identify AGG peaks.

[0069] If a peak is found within the initial peak size search space that meets the highest peak criteria, the found peak is identified as an initial peak. Furthermore, the height of the initial peak is set as (i) the height of the previous peak height threshold and (ii) the temporary previous peak height (which is used when a secondary peak is found to determine which peak to compare it to). This setting provides various technical advantages, including preventing the calling of background or stutter (false positive) peaks. If an initial peak is identified for another allele (e.g., a PM allele with a number of repeats between 55 and 90), the peak data associated with the initial peak can be removed from the final AGG dataset. After removing the AGG peak assigned to this allele from the output table of the final AGG dataset, the process can either repeat for another search space within the same allele or start a new process for another allele (e.g., another PM allele).

[0070] In block 450, the final AGG data set is searched using the initial peak size search space to identify the next highest peak or secondary peak in the initial peak size search space based on next highest peak criteria. In some examples, the criteria for identifying the next highest peak in the initial peak size search space include (i) the peak size of the secondary peak being greater or less than a secondary peak threshold (e.g., 7) from the peak size of the initial peak identified in block 440, (ii) comparing the size of the secondary peak to the initial found peak, and if the peak size of the secondary peak is less than the peak size of the initial peak, the height of the secondary peak should be greater than a set previous peak height threshold * secondary PM peak height threshold (in some embodiments, the secondary PM peak height threshold is set to 80%) called in the same search space in block 446, or if the peak size of the secondary peak is greater than the peak size of the initial peak, the height of the secondary peak should be greater than a set previous peak height threshold * secondary PM peak height threshold (in some embodiments, the secondary PM peak height threshold is set to 80%) from the previous search space in block 432, or (iii) any combination thereof. When the process is in the initial peak size search space, the threshold value has not yet been established, so the threshold value used is 0. After the initial peak size search space, the previous peak height threshold value has already been established. The criteria for identifying the next highest or secondary peak within the initial peak size search space provides various technical advantages, including avoiding the missing of additional AGG peaks within the same search window after the first peak has been identified.

[0071] If a peak that meets the criteria for identifying the next highest peak is found within the initial peak size search space, the found peak is identified as a secondary peak in block 454. In some examples, if a peak that meets the criteria for identifying the next highest peak as a secondary peak is found within the initial peak size search space and the peak size of the secondary peak is smaller than the peak size of the initial peak, the height of the secondary peak is set as the previous peak height threshold. This setting provides various technical advantages, including preventing the calling of background or stutter (false positive) peaks. If a secondary peak is identified for another allele (e.g., a PM allele with a number of repeats between 55 and 90), the peak data associated with the secondary peak may be removed from the final AGG dataset. At block 456, the initial and / or secondary peaks for the other alleles are reported (e.g., the processor provides the initial and / or secondary peaks on a display for the user to view via a user interface, a storage device for recordation purposes, another component of the computing device for future processing, or an output device such as a printer for a hard copy of the report) as follows: (i) if no peaks are found (no initial and secondary peaks), report that no peaks are found; (ii) if one peak is found (initial or secondary), report the found peak; or (iii) if both peaks are found (initial and secondary), report the peaks in descending order since the algorithm begins with the CGG end. For example, if initial peak size 1 < secondary peak size 2: return peak 2, peak 1; whereas, if secondary peak size 2 > initial peak size 2: return peak 1, peak 2. The AGG genotyping process continues at block 460 for each of the remaining peaks of the other alleles.

[0072] If no peaks meeting the next highest peak criteria are found within the initial peak size search space, then in block 458, a peak size adjustment threshold is subtracted from the expected AGG peak size determined in block 412 (or a previously calculated new expected AGG peak size) to obtain a new expected AGG peak size. The new expected AGG peak size is used to identify an initial peak by repeating block 440. This adjustment shifts / expands the initial peak size search space to help find the initial peak in the final AGG data set. In some embodiments, the peak size adjustment threshold is 15-50 bps (e.g., 24 bps).

[0073] Once an initial peak is identified for an allele in block 440 / 450, additional AGG peaks may be identified for the allele (e.g., an allele with a greater number of repeats). In block 460, additional peaks are identified for the allele. Identifying the additional peaks includes subtracting a peak size adjustment threshold from the peak size of the initial peak found in block 440 / 450 to obtain a next expected AGG peak size in block 462. In block 464, a next peak size search space (peak size search space in base pairs) is then determined for the allele based on the expected AGG peak size. The next peak size search space may be determined based on a next peak size range. In some examples, the next peak size range is a predetermined amount or size range adjustment threshold (+ / -) of base pairs around the peak size of the initial peak found in block 440 / 450. In a particular example, the predetermined amount or size range adjustment threshold is + / - 12 base pairs of the peak size of the initial peak. In certain examples, if the lower end of the next peak size range is less than the first normal peak threshold, the lower end of the next peak size range is set to the first normal peak threshold.

[0074] Once the second search space is determined, the next peak size search space can be used to search the final AGG data set to identify the highest peak in the next peak size search space based on highest peak criteria in block 464. In some examples, the criteria for identifying the highest peak in the next peak size search space include: (i) the peak size is greater than a minimum peak size threshold (in some embodiments, the minimum peak size threshold is set to 132 base pairs), (ii) the peak height is greater than a minimum peak height threshold (in some embodiments, the minimum peak height threshold is set to 200 RFU), (iii) the peak height is greater than the minimum normal height threshold (set in block 420 or 432) * a minimum percentage PM threshold (in some embodiments, the minimum percentage PM threshold is set to 10%), (iv) the peak height is greater than the previous peak height threshold (set in block 446 or 454) * a secondary PM peak height percent threshold (e.g., 80%), or (v) any combination thereof. The criteria for identifying the highest peak in the following peak size search space offer various technical advantages, including: (i) a minimum peak size threshold, which is the size of a repeat-free PCR amplicon, can be used as a criterion to prevent calling false peaks; (ii) a minimum peak height threshold can be used to prevent calling background low-level peaks; and (iii) using the height of a previously identified peak as the threshold for AGG peak search in PM alleles provides built-in dynamic DNA input control for the height of the AGG peak in PM alleles.

[0075] If a peak is found within the next peak size search space that meets the highest peak criteria, the found peak is identified as an additional peak in block 468. Additionally, the size of the additional peak is set as the temporary previous peak height (which is used when a secondary peak is found to determine which peak to compare the secondary peak to). If an additional peak is identified for another allele, the peak data associated with the additional peak may be removed from the final AGG dataset.

[0076] In block 470, the final AGG data set is searched using the next peak size search space to identify a secondary peak or the next highest peak within the next peak size search space. In some examples, the criteria for identifying a secondary peak within the next peak size search space include (i) the peak size of the secondary peak is greater than or less than a secondary peak threshold (e.g., 7) from the peak size of the additional peak size identified in block 460, (ii) comparing the size of the secondary peak with the additional discovered peak, and if the peak size of the secondary peak is less than the peak size of the additional peak, the height of the secondary peak should be greater than the temporary previous peak height * secondary PM peak height threshold set in block 468 (in some embodiments, the secondary PM peak height threshold is set to 80%), or if the peak size of the secondary peak is greater than the peak size of the additional peak, the height of the secondary peak should be greater than block 454 or 468 * secondary PM peak height threshold (in some embodiments, the threshold is set to 80%), or (iii) any combination thereof. The next peak size criterion for identifying the secondary peak or next highest peak within the search space provides various technical advantages, including (i) preventing background calling or stutter (false positive) peaks, and (ii) avoiding missing additional AGG peaks within the same search window after the first peak has been identified.

[0077] If a peak that meets the criteria for a secondary peak is found within the second search space, the found peak is identified as a secondary peak. In some instances, a peak that meets the criteria for identifying the next highest peak as a secondary peak is found within the next peak size search space. If the peak size of the secondary peak is smaller than the peak size of the additional peak, the height of the secondary peak is set as the previous peak height threshold. This setting provides various technical advantages, including preventing the calling of background or stutter (false positive) peaks. If a secondary peak is identified for another allele (e.g., a PM allele with a number of repeats between 55 and 90), the peak data associated with the secondary peak can be removed from the final AGG dataset. At block 456, the initial and / or secondary peaks for the different alleles are reported (e.g., the processor provides the initial and / or secondary peaks on a display for the user to view via a user interface, a storage device for recordation purposes, another component of the computing device for future processing, or an output device such as a printer for a hard copy of the report) as follows: (i) if no peak is found (no initial and secondary peak), report that no peak is found; (ii) if one peak is found (initial or secondary), report the found peak; or (iii) if both peaks are found (initial and secondary), report the peaks in descending order since the algorithm starts at the end of the CGG. For example, if initial peak size 1 < secondary peak size 2: return peak 2, peak 1; whereas, if secondary peak size 2 > initial peak size 2: return peak 1, peak 2.

[0078] In block 478, a determination is made as to whether the final AGG dataset should continue to be searched for additional peaks (i.e., whether the output table of the final AGG dataset has more peaks for the long allele). In some embodiments, for PM alleles with a repeat size less than 65, the decision to continue searching the final AGG dataset is made based on whether the size limit of the repeat-free PCR amplicon (e.g., 132 bp) has been reached for the long / PM allele. In other embodiments, for PM alleles with a repeat size greater than 65, the decision to continue searching the final AGG dataset is made based on whether the maximum size AGG peak seen in the normal / short allele has been reached for the long / PM allele. If it is determined that there are more peaks in the final AGG dataset for the allele, blocks 460 and 470 are repeated. In block 460, a peak size adjustment threshold is subtracted from the previously calculated next predicted AGG peak size to obtain the next expected AGG peak size (essentially the third, fourth, fifth, etc. expected AGG peak size). This adjustment shifts / expands the search space for additional peak sizes to help find additional peak(s) in the final AGG dataset. Blocks 460 and 470 are repeated until all peaks of alleles (e.g., alleles with a greater number of repeats) have been identified in the final AGG dataset. Once all peaks of alleles (e.g., alleles with the fewest number of repeats) have been identified in the final AGG dataset, a decision is made in block 480 as to whether the final AGG dataset should continue to be searched for additional alleles (does the output table of the final AGG dataset have peaks for other alleles?). If it is determined that more alleles exist, the AGG genotyping process continues to block 440 for each remaining allele. If it is determined that no more alleles exist, the process continues to block 490, where alleles are genotyped based on the AGG interruption data obtained in blocks 412-476. AGG genotype flow

[0079] In block 490, the number of CGG repeats downstream of the last AGG break on the allele (e.g., normal allele or PM allele) is calculated using the CGG repeats identified in blocks 402-410 and the AGG peak identified in blocks 412-480. In some embodiments, the number of CGG repeats downstream of the last AGG break on the allele is calculated using equation (2). Number of CGG repeats downstream of the last AGG break (using CCG#) = ((P n -A) / 3)-1 Formula (2) where P n A is the peak size corresponding to the last AGG break on the allele (e.g., the smallest AGG break on each allele), A is the size of the PCR amplicon without repeats; in some examples, the PCR amplicon has 132 base pairs, so A = 132, 3 is the number of base pairs for each CGG repeat, and 1 indicates a single AGG break. The position of the last AGG break can be used to start the resulting string of alleles (constructed backwards). AGG(CGG) n where "n" is the result of equation (1) (i.e., the expected first AGG peak size).

[0080] The remaining AGG peaks identified in blocks 412-480 are then cycled through until all have been processed for alleles. In some embodiments, the number of CGG repeats separated by any two adjacent AGG breaks (if present) is calculated using equation (3). The number of CGG repeats (using CCG#) separated by any two adjacent AGG breaks = ((P n -P n-1 ) / 3)-1 Formula (3) where P n and P n-1where A is the peak size corresponding to adjacent AGG breaks on the allele, 3 is the number of base pairs for each CGG repeat, and 1 indicates a single AGG break. After processing the number of CGG repeats separated by any two adjacent AGG breaks, the resulting string may be updated by prepending a string of the same format (constructed backwards): AGG(CGG) n where "n" is the result of equation (1) (i.e., the expected first AGG peak size).

[0081] The number of CGG repeats preceding the first AGG break can be calculated by subtracting (i) the number of CGG repeats downstream of the last AGG break calculated by equation (2), (ii) the total number of CGG repeats separating adjacent AGG breaks calculated by equation (3), and (iii) the total number of AGG breaks determined by the AGG peaks identified in blocks 412-480 from the total number of CGG repeats of the allele identified by the GS-PCR assay and identified in blocks 402-410. After processing the number of CGG repeats preceding the first AGG break, the resulting string may be updated (constructed backwards) by prepending AGG(CGG) to a string of the same format. n where "n" is the result of equation (1) (i.e., the expected first AGG peak size). The final result string, after processing the number of CGG repeats preceding the first AGG break, is the AGG genotype of the allele. The process of block 490 can be repeated for each allele identified for the sample.

[0082] In optional block 492, an expansion risk score can be calculated based on the AGG genotype determined for each allele of the sample. The risk score can be calculated using the number of CGG repeats and the AGG interruptions therein, as shown in Table 1. In some examples, the risk score identifies the patient's risk of developing the late-onset neurodegenerative disease Fragile X-associated Tremor / Ataxia Syndrome (FXTAS) or Fragile X-associated Primary Ovarian Insufficiency (FXPOI), or the risk of transmitting the full mutant allele to the subject's offspring, or any combination thereof. In a specific example, if there are more than five AGG interruptions, the number of AGG interruptions can be set to 5 for calculating the risk score(s), because five interruptions and more than five interruptions generally have the same clinical outcome. [Table 1]

[0083] In block 494, the determined AGG genotype for each allele and optional risk score may be output. Outputting the determined AGG genotype for each allele and optional risk score may include providing the output to an end user and / or recording the output to a storage device (e.g., displaying the output on a user interface and / or storing the output in a database results file). As will be appreciated, the process flow for AGG genotyping using the FXS AGG PCR assay platform and genotyping technology described with respect to Figures 4A-4H allows the user to process all samples in parallel in batches without human error and seamlessly incorporate the results into a report summary, significantly reducing turnaround time. With respect to the report summary, the AGG genotype and optional risk score inform the clinician / geneticist on how to guide the subject (e.g., patient) through family planning options (e.g., in vitro diagnostics (IVD)) for current or future pregnancies, or provide / prepare options for the current pregnancy, if the mother is determined to be at high risk for PM elongation. Thus, as a final step, the clinician / genetics specialist provides advice to the subject or patient based on the AGG genotype and optional risk score. The advice may include subsequent medical diagnostic testing (e.g., IVD), a detailed discussion of the inheritance patterns of FXD, the clinical presentation of all three conditions (FXS, FXPOI, FXTAS), reproductive options if appropriate, guidance on conversations with children and long-term at-risk family members, considerations for testing asymptomatic children, research opportunities, family support, and family planning options such as referrals to medical, developmental, and psychological providers as indicated.

[0084] FIG. 5 illustrates an exemplary computing device 500 suitable for use in systems and methods for AGG genotyping using FXS AGG PCR assay platform and genotyping techniques according to the present disclosure. The exemplary computing device 500 includes a processor 505 that communicates with memory 510 and other components of the computing device 500 using one or more communication buses 515. The processor 505 is configured to execute processor-executable instructions stored in the memory 510 to perform one or more methods for locating and identifying AGG peaks present in raw data, determining AGG genotypes of alleles, and / or determining a patient's risk score according to different examples, such as some or all of the exemplary processes 300 or 400 described above with respect to FIGS. 3 and 4A-4H. In this example, the memory 510 stores processor-executable instructions that provide AGG peak analysis 520 and AGG genotyping 525, as described above with respect to FIGS. 2, 3, and 4A-4H.

[0085] Computing device 500 also includes one or more user input devices 530, such as a keyboard, mouse, touchscreen, microphone, etc., for accepting user input in this example. Computing device 500 also includes a display 535 for providing visual output to a user, such as a user interface. Computing device 500 also includes a communications interface 540. In some examples, communications interface 540 may enable communication using one or more networks, including a local area network ("LAN"), a wide area network ("WAN") such as the Internet, a metropolitan area network ("MAN"), point-to-point or peer-to-peer connections, etc. Communications with other devices may be achieved using any suitable network protocol. For example, one suitable network protocol may include Internet Protocol ("IP"), Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), or a combination thereof, such as TCP / IP or UDP / IP. [Example]

[0086] IV. Working Examples The systems and methods implemented in various embodiments may be better understood with reference to the following examples. Example 1: FRAX AGG Interruption PCR Assay and AGG Genotyping Algorithm Specimens used and reproducibility

[0087] Ten blood samples (S1–S10) were extracted at the first external testing laboratory, including five samples that underwent external testing for AGG disruption. Five samples (S11–S15) were extracted from blood at the first clinical laboratory and previously subjected to a fragile X PCR assay. These five samples are expected to represent samples that may be received for clinical trials. Twenty-three samples were extracted and tested for AGG disruption at the second external testing laboratory (S16–S38). The AGG genotypes of these 23 samples were blinded to the operator prior to use in this example. In addition to the samples listed above, 25 samples used in assay development were also tested to compare the automated FRAX AGG PCR genotyping algorithm calls with those from manual genotyping. For intra-assay reproducibility, five samples, including three with known AGG genotype results, were analyzed in triplicate. These samples were also tested in two additional runs in a single batch for inter-assay reproducibility, one run performed by a different operator using a different lot of KAPA2G Robust Hotstart Enzyme mix and AGG PCR primer mix. Lot and expiration date information is listed in Table 2. Two models of thermal cyclers and two ABI 3730xl instruments were also included to assess inter-instrument reproducibility. [Table 2]

[0088] Genotyping calls by the FRAX AGG genotyping algorithm were reproducible across all five samples (Table 3). Note that slight genotypic differences between repeats are expected due to experimental variability, especially for the highly repetitive GC-rich regions isolated by CE. Furthermore, the resolution of CGG repeat numbers determined by GS-PCR was previously confirmed to vary by 1–4 repeats depending on repeat length. Therefore, given the inherent inter-run variability, a difference of 1–2 repeats (i.e., 3–6 bases) between repeats is within the expected range of variation for normal or PM alleles, or for the number of CGGs interrupted by AGGs. [Table 3-1] [Table 3-2] Analytical sensitivity and specificity

[0089] Analytical sensitivity and specificity were first evaluated using five blood DNA samples (samples S1–S5) whose AGG genotypes had previously been reported by an external testing laboratory using the validated platform. CE electropherograms of the AGG PCR products are shown in Figures 6A and 6B. The AGG genotypes obtained from the external laboratory and the FRAX AGG PCR assay are listed in Table 4. All controls passed, and no false positives or false negatives were called. The FRAX AGG PCR assay identified the same number of AGG breaks in all samples as the expected results, four of which showed concordant genotypes (Table 4). As previously mentioned, differences in the location of the AGG breaks and / or total repeat number by 1–2 repeats between the results from this example and those from the external testing were expected and considered concordant. [Table 4]

[0090] For sample S2, the total number of AGG interruptions identified by the FRAX AGG PCR assay was consistent with the expected results. However, the allele phasing of the AGG interruptions was inconsistent with the external laboratory results (Table 3). While testing by the external laboratory assigned only one interruption to the normal allele (allele 1), the FRAX AGG PCR assay identified two AGG-specific peaks (159 and 189 bp, Figure 6A). Notably, two other samples, both of which had 29 total repeats in their normal alleles and had AGG peaks of the same size as sample S2, were consistent for the two AGG interruptions when compared to the expected results (Figure 6A and Figure 6B, samples S2–S4). As mentioned above, the FRAX AGG PCR genotyping algorithm was initially developed to assign AGG peaks to normal alleles, so the genotyping difference observed for sample S2 may be due to which allele the AGG peak was initially assigned to.

[0091] To resolve the issue of genotyping discrepancies, two approaches are used. As part of the FRAX PCR assay, the TRP-PCR assay screens for expanded FRAX alleles by using a CGG repeat primer paired with a GS primer to amplify the CGG repeat region. In the presence of an AGG interruption, the affinity of the CGG repeat primer for the target sequence decreases, thereby reducing PCR efficiency and resulting in a "dip" in the CE electropherogram (Figure 7A). If one allele has more repeats than the other (e.g., PM vs. normal allele), or if the AGG interruption is identical in both alleles, the signal intensity of these dips will be close to baseline. Thus, for sample S2, if there were two AGG interruptions in the normal allele, the TRP-PCR electropherogram would show two AGG dips corresponding to the interruptions in the repeat range of the normal allele.

[0092] As shown in Figure 7A, the TRP-PCR results support the AGG phasing results of the FRAX AGG PCR assay for sample S2. Similar profiles were observed for samples S3 and S4. Interestingly, upon closer inspection, the signal intensity for one of the AGG dips in the normal allele for S2 was observed to have dropped to near baseline levels, indicating that the AGG interruption likely existed at the same location in both alleles (the 189-bp AGG dip in the S2 panel, Figure 7A). To confirm this, the normal and PM alleles were amplified, gel-purified, and analyzed by the FRAX AGG PCR assay to separately determine the number of AGG interruptions on each allele. As expected, the FRAX AGG PCR CE electropherogram clearly identified two AGGs in the normal allele and four AGGs in the PM allele. Importantly, neither the FRAX AGG PCR assay nor the results from the external laboratory testing identified overlapping AGG breaks in both alleles, although it is worth noting that this difference did not alter the risk score for expansion (<1%, see Table 1).

[0093] To further evaluate the analytical sensitivity, specificity, and accuracy of the FRAX AGG PCR assay and the FRAX AGG PCR genotyping algorithm, an additional 23 samples of known AGG genotype from an external laboratory were tested in replicates (Table 5). Samples were anonymized to the operator before testing, and all calls were manually analyzed using the FRAX AGG PCR genotyping algorithm. [Table 5-1] [Table 5-2] [Table 5-3] [Table 5-4]

[0094] As shown in Table 5, at least one AGG disruption was detected in all samples, and FRAX The AGG genotype calls from the AGG PCR assay were concordant in 21 of 23 samples. No controls failed, and no false-negative or false-positive calls were made. Samples S25 and S37 were not genotyped by the FRAX AGG PCR assay because the number of repeats in the PM allele exceeded the 90-repeat limit specified for the FRAX AGG PCR genotyping algorithm; however, the number of AGG interruptions per normal allele was consistent with the expected results.

[0095] The FRAX AGG PCR assay did not match the expected results for samples S17 and S23, likely caused by either an AGG break at the same position in both alleles or the order in which the alleles were assigned for the break. In sample S17, the expected result for an AGG break in the PM allele was six, but the FRAX AGG PCR assay identified five AGGs (Table 5 and Figure 8A). As with sample S2 above, review of the TRP-PCR CE data indicated that one of the AGG dips in the normal allele size range could potentially be located at the same position in both the normal and premutation alleles (159 bp AGG dip, Figure 8B). Using gel-purified normal and PM allele PCR products, the FRAX AGG PCR further confirmed that the AGG break, represented by the 159-bp peak, occurred in both alleles (Fig. 8C ).

[0096] For sample S23, the AGG break represented by the 159-bp AGG PCR peak could be assigned to either the normal or PM allele based on two nearby AGG peaks at 186 bp and 192 bp, both approximately 30 bp or approximately 10 repeats from the 159-bp AGG break (Figure 9A). While an external laboratory assigned two and five AGG breaks to the normal and PM alleles, respectively, the FRAX AGG PCR genotyping algorithm was developed to first assign AGG break peaks to shorter alleles, so three AGGs were assigned to the normal allele and four AGGs to the premutation allele (Table 5 and Figure 9A). TRP-PCR analysis showed that there was only one AGG break corresponding to the 159-bp peak and that it could be located in either the normal or PM allele (Figure 9B). Analysis of each allele by FRAX AGG PCR clearly demonstrated two and five AGG interruptions in the normal and PM alleles, respectively, in agreement with the genotypes from the external testing laboratory (Figure 9C).

[0097] In summary, the FRAX AGG PCR assay was concordant for the total number of AGG interruptions in 27 of the 28 samples. The one sample (S17) where there was a discrepancy between the expected result and the number of AGG interruptions was most likely due to the order of the AGG peak data assigned by the FRAX AGG PCR genotyping algorithm compared to that from the external test. Furthermore, we identified limitations in both the FRAX AGG PCR assay and the external laboratory test when the AGG interruption occurs at the same position in both the normal and PM alleles. This was not surprising, given that CE and GeneMapper only detect and output a single AGG PCR peak. However, as shown in Table 1, this difference does not affect the expansion risk score when the number of interruptions is two or more and the repeat size is less than 65. In cases with more than 65 repeats and at least one AGG interruption, the genotypes were concordant. The AGG PCR fragments corresponding to interruptions in these longer repeats are much larger than those in the normal allele, allowing for unambiguous allele phasing by the FRAX AGG PCR genotyping algorithm. FRAX AGG PCR genotyping algorithm test

[0098] Results from the FRAX AGG genotyping algorithm were in 100% agreement with results from a manual genotyping analysis based on the 38 samples performed in this example. To further test agreement, 25 samples used during assay development were also genotyped manually and by the FRAX AGG genotyping algorithm, and results were in 100% agreement between the two methods. Not only did the algorithm perform comparably to the manual analysis for genotyping, but it also performed well in identifying any samples where the quality of the input CE data varied or the genotype output was outside the normal range and flagging them for "manual review."

[0099] As will be appreciated, it is possible to encounter samples that meet the testing criteria (e.g., female samples with PM alleles in the range of 55–90 total repeats) but do not contain an AGG interruption. For example, a sample's normal allele may have 21 total repeats and contain no AGG interruption. Its PM allele may also lack an AGG, as in samples S1 or S4 (Table 4). The absence of an AGG PCR product can be difficult to distinguish from poor assay performance. To ensure an accurate call, in addition to repeating the assay and confirming a "no AGG" result, the sample's FRAX TRP-PCR assay CE electropherogram can also be reviewed. In contrast to what is observed in Figures 7A, 8B, and 9B, a smooth, gradual decrease in CGG stutter peak height without a sudden drop in signal supports a "no AGG" call (Figure 10). Given that typically only female PM allele carriers will be tested, the possibility of a sample not containing an AGG interruption would be rare, based on a total of approximately 95 samples tested during assay development and validation in which no samples without an AGG interruption were encountered. conclusion

[0100] The technical performance of the FRAX AGG interruption PCR assay was determined to be reproducible and robust for the generation of AGG-specific PCR products and for the automated results of AGG genotypes by the FRAX AGG genotyping algorithm.

[0101] Of 28 samples with known AGG genotypes, 27 were consistent with the expected results. For one sample, the total number of AGG interruptions matched the expected number, but the AGG genotype differed due to the allele order in which AGG was assigned by the FRAX AGG genotyping algorithm. However, this difference would not have had a clinical impact on this sample's final risk score. Inter- and intra-assay reproducibility was also 100% for AGG genotyping of all five samples, including three samples with known AGG genotypes. The FRAX AGG interruption PCR assay can tolerate 10–160 ng of input DNA extracted from blood specimens, and assay performance remained robust with gDNA extracted from blood specimens. Furthermore, the FRAX AGG genotyping algorithm showed 100% agreement with manual genotyping analysis in identifying AGG-specific peaks and determining the AGG genotype for all samples tested. V. Further Considerations

[0102] In the above description, specific details are set forth to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits may be shown in block diagrams in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

[0103] The implementation of the techniques, blocks, steps, and means described above can be done in various ways. For example, these techniques, blocks, steps, and means can be implemented in hardware, software, or a combination thereof. In the case of a hardware implementation, the processing unit can be one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLCs), or other devices. The functions described above may be implemented in a programmable logic device (PLD), a field programmable gate array (FPGA), a processor, a controller, a microcontroller, a microprocessor, other electronic unit designed to perform the functions described above, and / or combinations thereof.

[0104] Also, it should be noted that the embodiments may be described as a process that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. While a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Moreover, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination corresponds to a return of the function to the calling function or the main function.

[0105] Furthermore, embodiments may be implemented by hardware, software, scripting languages, firmware, middleware, microcode, hardware description languages, and / or any combination thereof. When implemented in software, firmware, middleware, scripting languages, and / or microcode, the program code or code segments to perform the necessary tasks may be stored in a machine-readable medium such as a storage medium. A code segment or machine-executable instructions may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a script, a class, or any combination of instructions, data structures, and / or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, and / or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, ticket passing, network transmission, etc.

[0106] For a firmware and / or software implementation, the methodologies may be implemented with modules (e.g., procedures, functions, and so on) that perform the functions described herein. Any machine-readable medium tangibly embodying instructions may be used in practicing the methodologies described herein. For example, software code may be stored in memory. The memory may be implemented within the processor or external to the processor. As used herein, the term "memory" refers to any type of long-term, short-term, volatile, non-volatile, or other storage medium and is not limited to any particular type or number of memories or the type of medium on which the memory is stored.

[0107] Additionally, as disclosed herein, the terms "storage medium," "storage," or "memory" can refer to one or more memories for storing data, including read-only memory (ROM), random-access memory (RAM), magnetic RAM, core memory, magnetic disk storage media, optical storage media, flash memory devices, and / or other machine-readable media for storing information. The term "machine-readable medium" includes, but is not limited to, portable or non-removable storage devices, optical storage devices, wireless channels, and / or various other storage media capable of storing instructions and / or data that contain or carry the instructions and / or data.

[0108] While the principles of the present disclosure have been described above in connection with specific apparatus and methods, it is to be clearly understood that this description is made only by way of example and not as a limitation on the scope of the disclosure. The present invention provides, for example, the following items. (Item 1) obtaining raw data from a Fragile X Syndrome (FXS) assay performed on a sample using a genetic analyzer, the raw data including gene-specific (GS) polymerase chain reaction (PCR) data and AGG-interruption PCR data separated by capillary electrophoresis; determining an expected AGG peak size for a first allele identified in the raw data; and iteratively searching the raw data to identify one or more AGG peaks on the first allele using a first set of search spaces determined based on the expected AGG peak sizes; determining the number of CGG repeats downstream of the last AGG break on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; determining the number of CGG repeats preceding the first AGG break on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break; providing said AGG genotype for said first allele; A method comprising: (Item 2) 2. The method of claim 1, further comprising determining the number of CGG repeats separated by any two adjacent AGG breaks on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, wherein the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG break, the number of CGG repeats separated by any two adjacent AGG breaks, and the number of CGG repeats preceding the first AGG break. (Item 3) iteratively searching the raw data to identify one or more AGG peaks on a second allele using a second set of search spaces determined based on the expected AGG peak sizes; determining the number of CGG repeats downstream of the last AGG break on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; determining the number of CGG repeats preceding the first AGG break on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the final AGG break and the number of CGG repeats preceding the first AGG break; providing said AGG genotype for said second allele; 3. The method of claim 1 or 2, wherein (i) the first allele is a normal allele and the second allele is a premutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a premutation allele and the second allele is a different premutation allele. (Item 4) 4. The method of claim 1, further comprising determining a risk score for the subject associated with the sample based on the AGG genotype generated for the first allele, the second allele, or both the first allele and the second allele, wherein the risk score identifies the subject's risk of developing a late-onset neurodegenerative disease, Fragile X-associated Tremor / Ataxia Syndrome (FXTAS) or Fragile X-associated Primary Ovarian Insufficiency (FXPOI), or the risk of transmitting a full mutant allele to the subject's offspring, or any combination thereof. (Item 5) iteratively searching the raw data to identify one or more AGG peaks on the first allele; determining a first search space of the first set of search spaces for the first allele based on the expected AGG peak size; searching the raw data and identifying an initial AGG peak on the first allele using a first search space; determining a second search space for the first allele based on a peak size of the initial AGG peak on the first allele; iteratively searching the raw data to identify one or more additional AGG peaks on the first allele using the second search space; 4. The method of claim 1, wherein the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break are determined for the first allele based on the GS PCR data, the initial peak identified in the first allele, and the one or more additional peaks identified in the first allele. (Item 6) iteratively searching the raw data to identify one or more AGG peaks on the second allele; determining a first search space of the second set of search spaces for the second allele based on the expected AGG peak size; searching the raw data and identifying an initial AGG peak on the second allele using a second search space; determining a second search space for the second allele based on a peak size of the initial AGG peak on the second allele; iteratively searching the raw data to identify one or more additional AGG peaks on the second allele using the second search space; 6. The method of claim 3 or 5, wherein the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break are determined for the second allele based on the GS PCR data, the initial peak identified in the second allele, and the one or more additional peaks identified in the second allele. (Item 7) 6. The method of claim 3 or 5, wherein once the one or more AGG peaks on the first allele are identified, the one or more AGG peaks on the first allele are removed from the raw data before iteratively searching the raw data to identify the one or more AGG peaks on the second allele. (Item 8) one or more data processors; When executed on said one or more data processors, it causes said one or more data processors to: obtaining raw data from a Fragile X Syndrome (FXS) assay performed on a sample using a genetic analyzer, the raw data including gene-specific (GS) polymerase chain reaction (PCR) data and AGG-interruption PCR data separated by capillary electrophoresis; determining an expected AGG peak size for a first allele identified in the raw data; and iteratively searching the raw data to identify one or more AGG peaks on the first allele using a first set of search spaces determined based on the expected AGG peak sizes; determining the number of CGG repeats downstream of the last AGG break on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; determining the number of CGG repeats preceding the first AGG break on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break; providing the AGG genotype for the first allele; and a non-transitory computer-readable storage medium containing instructions for causing an action to be performed, the action including: (Item 9) 9. The system of claim 8, wherein the action further comprises determining the number of CGG repeats separated by any two adjacent AGG breaks on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, and wherein the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG break, the number of CGG repeats separated by any two adjacent AGG breaks, and the number of CGG repeats preceding the first AGG break. (Item 10) the actions include iteratively searching the raw data to identify one or more AGG peaks on a second allele using a second set of search space determined based on the expected AGG peak sizes; determining the number of CGG repeats downstream of the last AGG break on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; determining the number of CGG repeats preceding the first AGG break on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the final AGG break and the number of CGG repeats preceding the first AGG break; providing said AGG genotype for said second allele; 10. The system of claim 8 or 9, wherein (i) the first allele is a normal allele and the second allele is a premutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a premutation allele and the second allele is a different premutation allele. (Item 11) 11. The system of claim 8, wherein the action further comprises determining a risk score for the subject associated with the sample based on the AGG genotype generated for the first allele, the second allele, or both the first allele and the second allele, wherein the risk score identifies the subject's risk of developing a late-onset neurodegenerative disease, Fragile X-associated Tremor / Ataxia Syndrome (FXTAS) or Fragile X-associated Primary Ovarian Insufficiency (FXPOI), or the risk of transmitting a full mutant allele to the subject's offspring, or any combination thereof. (Item 12) iteratively searching the raw data to identify one or more AGG peaks on the first allele; determining a first search space of the first set of search spaces for the first allele based on the expected AGG peak size; searching the raw data and identifying an initial AGG peak on the first allele using a first search space; determining a second search space for the first allele based on a peak size of the initial AGG peak on the first allele; iteratively searching the raw data to identify one or more additional AGG peaks on the first allele using the second search space; 11. The system of claim 8, wherein the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break are determined for the first allele based on the GS PCR data, the initial peak identified in the first allele, and the one or more additional peaks identified in the first allele. (Item 13) iteratively searching the raw data to identify one or more AGG peaks on the second allele; determining a first search space of the second set of search spaces for the second allele based on the expected AGG peak size; searching the raw data and identifying an initial AGG peak on the second allele using a second search space; determining a second search space for the second allele based on a peak size of the initial AGG peak on the second allele; iteratively searching the raw data to identify one or more additional AGG peaks on the second allele using the second search space; 13. The system of claim 10 or 12, wherein the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break are determined for the second allele based on the GS PCR data, the initial peak identified in the second allele, and the one or more additional peaks identified in the second allele. (Item 14) 13. The system of claim 10, wherein once the one or more AGG peaks on the first allele are identified, the raw data is iteratively searched and the one or more AGG peaks on the first allele are removed from the raw data before identifying the one or more AGG peaks on the second allele. (Item 15) A computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product causing one or more data processors to: obtaining raw data from a Fragile X Syndrome (FXS) assay performed on a sample using a genetic analyzer, the raw data including gene-specific (GS) polymerase chain reaction (PCR) data and AGG-interruption PCR data separated by capillary electrophoresis; determining an expected AGG peak size for a first allele identified in the raw data; and iteratively searching the raw data to identify one or more AGG peaks on the first allele using a first set of search spaces determined based on the expected AGG peak sizes; determining the number of CGG repeats downstream of the last AGG break on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; determining the number of CGG repeats preceding the first AGG break on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break; providing the AGG genotype for the first allele; 1. A computer program product comprising instructions configured to cause the computer to perform actions including: (Item 16) 16. The computer program product of item 15, wherein the action further comprises determining the number of CGG repeats separated by any two adjacent AGG breaks on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, and wherein the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG break, the number of CGG repeats separated by any two adjacent AGG breaks, and the number of CGG repeats preceding the first AGG break. (Item 17) the actions include iteratively searching the raw data to identify one or more AGG peaks on a second allele using a second set of search space determined based on the expected AGG peak sizes; determining the number of CGG repeats downstream of the last AGG break on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; determining the number of CGG repeats preceding the first AGG break on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the final AGG break and the number of CGG repeats preceding the first AGG break; providing said AGG genotype for said second allele; 17. The computer program product of claim 15 or 16, wherein (i) the first allele is a normal allele and the second allele is a premutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a premutation allele and the second allele is a different premutation allele. (Item 18) 18. The computer program product of item 15, 16, or 17, wherein the actions further include determining a risk score for a subject associated with the sample based on the AGG genotype generated for the first allele, the second allele, or both the first allele and the second allele, wherein the risk score identifies a risk of the subject developing a late-onset neurodegenerative disease, Fragile X-associated Tremor / Ataxia Syndrome (FXTAS) or Fragile X-associated Primary Ovarian Insufficiency (FXPOI), or a risk of transmitting a full mutant allele to the subject's offspring, or any combination thereof. (Item 19) iteratively searching the raw data to identify one or more AGG peaks on the first allele; determining a first search space of the first set of search spaces for the first allele based on the expected AGG peak size; searching the raw data and identifying an initial AGG peak on the first allele using a first search space; determining a second search space for the first allele based on a peak size of the initial AGG peak on the first allele; iteratively searching the raw data to identify one or more additional AGG peaks on the first allele using the second search space; 18. The computer program product of claim 15, wherein the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break are determined for the first allele based on the GS PCR data, the initial peak identified in the first allele, and the one or more additional peaks identified in the first allele. (Item 20) iteratively searching the raw data to identify one or more AGG peaks on the second allele; determining a first search space of the second set of search spaces for the second allele based on the expected AGG peak size; searching the raw data and identifying an initial AGG peak on the second allele using a second search space; determining a second search space for the second allele based on a peak size of the initial AGG peak on the second allele; iteratively searching the raw data to identify one or more additional AGG peaks on the second allele using the second search space; 20. The computer program product of claim 17 or 19, wherein the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break are determined for the second allele based on the GS PCR data, the initial peak identified in the second allele, and the one or more additional peaks identified in the second allele.

Claims

1. A method comprising: assaying a sample obtained from a subject using a fragile X (FRAX) polymerase chain reaction (PCR) assay to obtain fragile X (FRAX) polymerase chain reaction (PCR) raw data, wherein the sample comprises nucleic acid, and the FRAX PCR assay utilizes a primer comprising both a CGG repeat and an "A" nucleotide at its 3' end, and a reverse primer.

2. The assaying performing gene-specific (GS) PCR and AGG-interrupted PCR (AGG PCR) on the sample to obtain a PCR product; and 2. The method of claim 1, comprising separating the PCR products using capillary electrophoresis or a high-resolution gel to obtain the FRAX PCR raw data.

3. The method of claim 2, wherein at least one primer of the FRAX PCR assay is labeled with a fluorescent tag; 3. The method of claim 2, wherein the PCR products are separated using capillary electrophoresis, the capillary electrophoresis comprising the PCR products migrating through a polymer and measuring fluorescent emissions from the at least one primer labeled with the fluorescent tag.

4. The method of claim 1, wherein the assaying comprises adding labeled bases during PCR extension.

5. The method of claim 1, wherein the FRAX PCR assay targets the CGG repeat region in the 5' untranslated region (UTR) of the FMR1 gene.

6. The method of claim 1, further comprising obtaining a batch of samples, wherein the batch of samples is assayed simultaneously using the FRAX PCR assay.

7. The method of claim 6, wherein the batch of samples is assayed using a multiple well assay plate.

8. The method of claim 1, wherein the FRAX PCR assay comprises gene-specific (GS) PCR and AGG-interrupted PCR (AGG PCR).

9. The assaying diluting the sample; mixing the diluted sample with reagents and primers for the GS PCR and the AGG PCR, wherein the primer for the AGG PCR is the primer containing both the CGG repeat and the "A" nucleotide at the 3' end; conducting PCR amplification for both the GS PCR and the AGG PCR; and 9. The method of claim 8, further comprising performing a post-PCR cleanup.

10. The method of claim 9, wherein the GS PCR amplification is performed using a GC-rich PCR system comprising a GC-rich PCR reaction buffer and a GC-rich resolution solution.

11. The method of claim 8, wherein the GS PCR uses a forward primer FRAX-F1 and a reverse primer FRAX-R-6.

12. The method of claim 11, wherein the AGG PCR uses the same reverse primer FRAX-R-6.

13. The method of claim 1, wherein the FRAX PCR raw data includes the number of CGG repeats on each allele.

14. The method of claim 13, further comprising determining an AGG genotype for at least one of the alleles based on the FRAX PCR raw data.

15. A GS PCR subsystem using forward primer FRAX-F1 and reverse primer FRAX-R-6; and A FRAX PCR assay system comprising a primer containing both a CGG repeat and an "A" nucleotide at the 3' end, and an AGG PCR subsystem that uses a reverse primer.

16. The assay system of claim 15, wherein the reverse primer of the AGG PCR subsystem is the same reverse primer FRAX-R-6.

17. The assay system of claim 15, wherein the GS PCR subsystem and the AGG PCR subsystem target a CGG repeat region in the 5' untranslated region (UTR) of the FMR1 gene.

18. The assay system of claim 15, wherein the assay system uses a multiple well assay plate.

19. The assay system of claim 15, wherein the GS PCR subsystem comprises a GC-rich PCR subsystem comprising a GC-rich PCR reaction buffer and a GC-rich resolution solution.

20. The assay system of claim 19, wherein the GC-rich PCR subsystem further comprises an enzyme blend of thermostable Taq DNA polymerase and Tgo DNA polymerase.

Citation Information

Patent Citations

  • Automated nucleic acid repeat count calling methods

    US20150134267A1