Genotyping of Fragile X Syndrome AGG Interruptions

A computer-implemented method for processing Fragile X Syndrome assay data improves the detection of AGG interruptions, leading to accurate genotyping and risk assessments for Fragile X-associated disorders.

JP7699137B2Active Publication Date: 2025-06-26LABORATORY CORPORATION OF AMERICA HOLDINGS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022546022
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-30
Filing Date
2021-01-29
Publication Date
2025-06-26
Estimated Expiration
2041-01-29

AI Technical Summary

Technical Problem

Current methods for detecting AGG interruptions in the CGG repeat region of the FMR1 gene are inefficient, particularly in distinguishing between normal and premutation alleles, leading to inaccurate risk assessments for Fragile X Syndrome and associated disorders.

Method used

A computer-implemented method that processes raw data from Fragile X Syndrome assays using capillary electrophoresis, iteratively searches for AGG peaks, and generates an AGG genotype based on the number of CGG repeats downstream and preceding AGG interruptions, thereby improving the accuracy of risk assessments.

Benefits of technology

The method enhances the detection of AGG interruptions, leading to more accurate genotyping and risk scoring for Fragile X-associated tremor/ataxia syndrome, Fragile X-associated primary ovarian insufficiency, and the risk of transmitting full mutation alleles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007699137000010
    Figure 0007699137000010
  • Figure 0007699137000011
    Figure 0007699137000011
  • Figure 0007699137000012
    Figure 0007699137000012
Patent Text Reader

Abstract

The present disclosure relates to a fragile X syndrome (FXS) clinical trial, and in particular to an FXS AGG break polymerase chain reaction (PCR) assay and AGG break genotyping algorithm for implementation in the clinical trial. In particular, an embodiment relates to obtaining raw data from an FXS assay performed on a sample, iteratively searching the raw data, identifying one or more AGG peaks on a first allele using a first set of search spaces determined based on expected AGG peak sizes, determining the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break on the first allele based on the one or more AGG peaks, and generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG break and the number of CGG repeats preceding the first AGG break.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Priority Claim This application claims the benefit and priority of U.S. Provisional Application No. 62 / 967,792, filed on January 30, 2020, which is hereby incorporated by reference in its entirety for all purposes.

[0002] Field The present disclosure relates to Fragile X Syndrome (FXS or FRAX) clinical trials, particularly the FRAX AGG interrupted polymerase chain reaction (PCR) assay for implementation in clinical trials and the genotyping algorithm for AGG interruption.

Background Art

[0003] Background PCR and related amplification techniques can be used for analytical purposes. Typical analytical uses of PCR include the diagnosis of the status or determination of genotypes containing loci with polymorphisms. Examples of loci showing medically relevant polymorphisms are the 5' untranslated region (UTR) of the human FMR1 gene on the X chromosome. Normal individuals typically have 5 - 44 CGG repeats at this locus. In contrast, alleles at this locus containing large CGG repeat expansions (over 200 repeats, full mutation alleles) disrupt FMR1 gene expression and cause FXS. Furthermore, individuals with premutation (PM) alleles (50 - 200 repeats) are at risk of developing late - onset neurodegenerative diseases, Fragile X - associated tremor / ataxia syndrome (FXTAS) or Fragile X - associated primary ovarian insufficiency (FXPOI). Female PM allele carriers (a pan - ethnic frequency of 1 in 201) are at risk of transmitting full mutation alleles to their offspring. This risk depends on the size of the CGG repeat measured by the Fragile X PCR assay and the number of AGG interruptions between CGG repeats (which typically occur every 9 - 11 CGG repeats).

[0004] When PM alleles with 55 to 90 repeats are transmitted by carrier females, the risk of expansion to full mutations in descendants decreases with an increase in the number of AGG interruptions in the CGG repeat sequence. For PM alleles with more than 90 repeats, the expansion risk is more than 50% regardless of the number of AGG interruptions. Thus, the AGG interruption PCR assay can have the greatest clinical utility in characterizing AGG interruptions for PM alleles with 55 to 90 CGG repeats.

Summary of the Invention

Means for Solving the Problems

[0005] Abstract In various embodiments, obtaining raw data from a Fragile X Syndrome (FXS) assay performed on a sample, wherein the raw data includes gene-specific (GS) polymerase chain reaction (PCR) data and AGG interruption PCR data separated by capillary electrophoresis, determining an expected AGG peak size for a first allele identified in the raw data, iteratively searching the raw data and using a first set of search spaces determined based on the expected AGG peak size to identify one or more AGG peaks on the first allele, determining the number of CGG repeats downstream of the last AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, determining the number of CGG repeats preceding the first AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption, and providing the AGG genotype for the first allele. A computer-implemented method is provided.

[0006] In some embodiments, the method further includes determining the number of CGG repeats separated by any two adjacent AGG interruptions on the first allele based on the GS PCR data and one or more AGG peaks identified on the first allele, wherein the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG interruption, the number of CGG repeats separated by any two adjacent AGG interruptions, and the number of CGG repeats preceding the first AGG interruption.

[0007] In some embodiments, the method further includes iteratively searching the raw data to identify one or more AGG peaks on the second allele using a second set of search spaces determined based on the predicted AGG peak sizes, determining the number of CGG repeats downstream of the last AGG interruption on the second allele based on the GS PCR data and one or more AGG peaks identified on the second allele, determining the number of CGG repeats preceding the first AGG interruption on the second allele based on the GS PCR data and one or more AGG peaks identified on the second allele, generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption, and providing the AGG genotype for the second allele, wherein (i) the first allele is a normal allele and the second allele is a premutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a premutation allele and the second allele is a different premutation allele.

[0008] In some embodiments, the method further comprises determining a risk score for a subject related to a sample based on an AGG genotype generated for a first allele, a second allele, or both the first allele and the second allele, wherein the risk score identifies the risk that the subject will develop a late-onset neurodegenerative disease, fragile X-associated tremor / ataxia syndrome (FXTAS) or fragile X-associated primary ovarian insufficiency (FXPOI), or the risk of transmitting a full mutation allele to the subject's offspring, or any combination thereof.

[0009] In some embodiments, iteratively searching raw data to identify one or more AGG peaks on a first allele comprises determining a first search space of a first set of search spaces for the first allele based on an expected AGG peak size, searching the raw data and using the first search space to identify an initial AGG peak on the first allele, determining a second search space for the first allele based on the peak size of the initial AGG peak on the first allele, and iteratively searching the raw data and using the second search space to identify one or more additional AGG peaks on the first allele, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the first allele based on the GS PCR data, the initial peak identified on the first allele, and the one or more additional peaks identified on the first allele.

[0010] In some embodiments, repeatedly searching raw data to identify one or more AGG peaks on a second allele, based on an expected AGG peak size, determines a first search space of a second set of search spaces for the second allele, searching the raw data to identify an initial AGG peak on the second allele using the second search space, determining a second search space for the second allele based on the peak size of the initial AGG peak on the second allele, and repeatedly searching the raw data to identify one or more additional AGG peaks on the second allele using the second search space, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the second allele based on the GS PCR data, the initial peak identified on the second allele, and the one or more additional peaks identified on the second allele.

[0011] In some embodiments, when one or more AGG peaks on a first allele are identified, the raw data is repeatedly searched and the one or more AGG peaks on the first allele are removed from the raw data before searching for one or more AGG peaks on the second allele.

[0012] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium that includes instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of the one or more methods or processes disclosed herein.

[0013] In some embodiments, a computer program product is provided that is tangibly embodied on a non-transitory machine-readable storage medium and includes instructions configured to cause one or more data processors to perform some or all of the one or more methods disclosed herein.

[0014] Some embodiments of the present disclosure include a system that includes one or more data processors. In some embodiments, the system, when executed on one or more data processors, includes a non-transitory computer-readable storage medium that causes the one or more data processors to execute some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product tangibly embodied in a non-transitory machine-readable storage medium that includes instructions configured to cause one or more data processors to execute some or all of one or more methods disclosed herein and / or some or all of one or more processes disclosed herein.

[0015] The terms and expressions used are used as terms of description and not of limitation, and in the use of such terms and expressions, there is no intention to exclude equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the claimed invention. Accordingly, although the invention has been specifically disclosed by way of embodiments and optional features, variations and modifications of the concepts disclosed herein may be used by those skilled in the art, and such modifications and variations are considered to be within the scope of the invention as defined by the appended claims.

[0016] The present invention will be better understood in view of the following non-limiting figures.

Brief Description of the Drawings

[0017]

FIG. 1A

[0018]

FIG. 1B

[0019]

FIG. 2

[0020]

FIG. 3

[0021]

FIG. 4A

FIG. 4B

FIG. 4C

FIG. 4D

FIG. 4E

FIG. 4F

FIG. 4G

FIG. 4H

[0022]

FIG. 5

[0023]

FIG. 6A

FIG. 6B

[0024]

FIG. 7A-1

FIG. 7A-2

[0025]

FIG. 7B

[0026]

FIG. 8A

[0027]

FIG. 8B

[0028]

FIG. 8C

[0029]

FIG. 9A

[0030]

FIG. 9B

[0031]

FIG. 9C

[0032]

FIG. 10

[0033] In the accompanying drawings, similar components and / or features can have the same reference labels. Further, various components of the same type can be distinguished by following the reference label with a dash and a second label that distinguishes similar components. When only the first reference label is used herein, the description is applicable to any of the similar components having the same first reference label, regardless of the second reference label.

[0034] DETAILED DESCRIPTION The following description provides only preferred exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of the preferred exemplary embodiments provides those skilled in the art with a possible description for implementing various embodiments. It is understood that various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope recited in the appended claims.

[0035] In the following description, specific details are set forth in order to provide a thorough understanding of the embodiments. It will be understood, however, that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

[0036] Also, note that individual embodiments may be described as a process shown as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. A flowchart or diagram can describe the operations as a sequential process, but many of the operations may be performed in parallel or simultaneously. Further, the order of the operations may be rearranged. A process ends when its operations are completed, but it can have additional steps not included in the figure. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its end can correspond to the return of the function to the calling function or the main function. I. Introduction

[0037] Conventional systems use CGG repeat primer (TRP)-PCR separated by CE to identify alleles at loci containing large CGG repeat expansions. In the presence of AGG interruptions, the affinity of the CGG repeat primer for its target sequence is reduced, thereby reducing PCR efficiency and causing a "dip" in the CE electropherogram. Based on the TRP-PCR CE data, it may be possible to indirectly calculate the number and position of each AGG. However, it is very difficult to detect AGG interruptions located at a similar distance from the start of the CGG repeat region in a stepwise and accurate manner. To overcome these drawbacks, some conventional systems utilize GGA primer PCR for direct AGG detection, followed by manual genotyping analysis. However, GGA primer PCR can generate a very high background and typically does not meet more stringent quality criteria. Furthermore, conventional manual genotyping analysis has poor performance and is very error-prone.

[0038] To address these limitations and problems, the FXS AGG PCR assay described herein targets the CGG repeat region of the 5’UTR of the FMR1 gene and detects all AGG interruptions within the CGG repeats on all alleles. The FXS AGG PCR assay utilizes a primer that contains both the CGG repeat and a 3’-terminal “A” nucleotide that specifically binds to the AGG interruption of the CGG repeat, as well as a reverse primer for detection (see, e.g., FIG. 1A). Capillary electrophoresis (CE) can be used to separate these amplicons as individual peaks corresponding to the positions of the AGG interruptions on the electropherogram (see, e.g., FIG. 1B). In some examples, one or more end primers (e.g., forward and / or reverse primers) can be labeled with a fluorescent tag (e.g., fluorescein amidite (FAM)) to assist in separation by CE. In other examples, the fluorescent label can be provided by adding a labeled base that is incorporated during PCR extension. In other examples, no fluorescent label is used and the PCR products are generated with untagged primers that can be separated by high-resolution CE. Since AGG typically occurs with a periodicity of approximately 30 bp within the repeat region when located on the same allele, phasing of AGG is possible using the results of both FXS AGG PCR and gene-specific (GS) PCR assays. Peak identification techniques can be used to identify the peaks in the raw data of FXS GS and AGG PCR CE. The raw data can include the size and abundance of each PCR amplicon analyzed as peak size (base pairs) and height (e.g., relative fluorescence units (RFU)), respectively. To facilitate data analysis, the FXS AGG genotyping technique described herein processes the raw data of FXS GS and AGG PCR CE and determines the AGG genotype of each allele.

[0039] The FXS AGG genotyping technique first checks sizing standards and peak heights to ensure that the raw data quality of FXS GS and AGG PCR CE is acceptable for analysis. Next, the FXS AGG genotyping technique uses the FXS GS and AGG PCR CE raw data to calculate the CGG repeat numbers for each allele (e.g., normal allele and PM allele). To determine the number and position of AGG interruptions for each allele, the technique further includes repeatedly searching for and identifying AGG peaks on the allele with fewer repeats, which usually corresponds to the normal allele. After removing all AGG peaks assigned to this allele from the peak output table, the technique further includes repeating the process for the PM allele. In this way, the same peak is prevented from being assigned to both alleles, which could otherwise underestimate the PM allele's elongation risk. Once the peaks are identified and assigned to the alleles, the technique further includes evaluating the AGG genotype on each allele. In some embodiments, the number of CGG repeats downstream of the last AGG interruption on each allele is calculated using the following formula: (P - 132) ÷ 3 - 1 (P: peak size of the minimum AGG interruption on each allele). In some embodiments, the number of CGG repeats separated by any two adjacent AGG interruptions (if present) is calculated using the following formula: (P n - P n-1 ) ÷ 3 - 1 (P n - P n-1 : peak sizes of adjacent AGG interruptions). In some embodiments, the number of CGG repeats preceding the first AGG interruption is determined by subtracting the number of CGG repeats downstream of the final AGG, the total number of CGG repeats separating adjacent AGG interruptions, and the total number of AGG interruptions from the total CGG repeat number of the allele determined by the GS-PCR assay.

[0040] As used herein, the terms "substantially", "approximately", and "about" are defined as being mostly, but not necessarily fully, specified (including being fully specified) as would be understood by one of ordinary skill in the art. In any disclosed embodiment, the terms "substantially", "approximately", or "about" can be replaced by "within [percentage]" of the specified amount, where the percentages include 0.1, 1, 5, and 10%. As used herein, when an action is "based on" something, this means that the action is at least partially based on at least a portion of that something.

[0041] It will be understood that the FXS genotyping technology disclosed herein can be applied to evaluate other types of FXS PCR CE raw data as compared to the FXS GS and AGG PCR CE raw data specifically described herein. It will also be understood that other types of FXS PCR are contemplated for identifying CGG repeats and AGG interruptions. For example, alternatively or additionally, non-anchor primer PCR may be used to identify CGG repeats and / or AGG interruptions. II. FXS AGG PCR Assay Technology

[0042] One or more embodiments described herein can be implemented using program modules, engines, or components. Program modules, engines, or components can include software components or hardware components that can execute a program, subroutine, part of a program, or one or more of the described tasks or functions. As used herein, a module or component can exist on a hardware component independently of other modules or components. Alternatively, a module or component can be a shared element or process of other modules, programs, or machines. FIG. 2 shows a block diagram of an FXS AGG PCR assay platform 200 for interrupted testing of PM alleles (e.g., samples having CGG repeat extensions of 55 to 90), and shows modules, engines, or components (e.g., programs, code, or instructions) executable by one or more processors for implementing various subsystems of an analyzer system 205 according to various embodiments. Analyzer system 205 is a genetic analyzer (also known as a DNA sequencer), an automated system that can sequence DNA and analyze fragments for various applications. In some examples, the genetic analyzer is a capillary electrophoresis-based system in which DNA fragments bound to probes move through a polymer and fluorescence emission is measured. An array of multiple capillaries enables sample loading in a multiwell microplate format. In other examples, the genetic analyzer is a pyrosequencing technology-based system for rapid sequencing and analysis. Pyrosequencing is a method of DNA sequencing based on the principle of "sequencing by synthesis," where sequencing is performed by detecting nucleotides incorporated by DNA polymerase. Pyrosequencing relies on light detection based on a chain reaction when pyrophosphate is released. Modules, engines, or components can be stored on a non-transitory computer medium.If necessary, one or more of the modules, engines, or components can be loaded into the system memory (e.g., RAM) and executed by one or more processors of the analyzer system 205. In the example shown in FIG. 2, modules, engines, or components for implementing the gene mapper subsystem 210 and the AGG genotyping subsystem 215 are shown.

[0043] FIG. 2 also shows a wet laboratory subsystem 220 that includes a laboratory where chemical substances, drugs, or other materials or biological substances are tested and analyzed requiring water, direct ventilation, and dedicated plumbing utilities. The FXS AGG PCR assay platform 200 includes obtaining one or more samples 225 within the wet laboratory subsystem 220 at block 227. In some examples, batches of samples 225 are processed simultaneously. For example, more than 20, more than 50, or more than 100 samples can be processed simultaneously using the FXS AGG PCR assay platform 200. For example, multiple well assay plates, such as 96 well plates or 384 well plates, can be used with the FXS AGG PCR assay platform 200 to analyze more than 20, more than 50, or more than 100 samples simultaneously. In certain examples, multiple assay plates can be constructed and then lined up in a plate stacker of a capillary electrophoresis device to continuously inject the assay plates. In some examples, the sample 225 includes nucleic acids. In some examples, the sample 225 includes nucleic acids obtained from a female patient. In some examples, the sample 225 is whole blood or amniotic fluid that includes nucleic acids obtained from a female patient. In a particular example, the FRAX AGG PCR assay is a reflex assay from a fragile X PCR diagnostic assay, and the reflex assay is initiated when the fragile X PCR diagnostic assay identifies one or more samples 225 having at least 45 - 100 CGG repeats (e.g., 55 - 90 CGG repeats) in the PM allele of the FMR1 gene. In a particular example, the reflex assay is initiated when the fragile X PCR diagnostic assay identifies one or more samples 225 having 55 - 90 repeats in the PM allele of the FMR1 gene.

[0044] In block 230 within the wet laboratory subsystem 220, an FXS AGG PCR assay including GS PCR 235 and AGG interrupted PCR 240 (AGG PCR) is performed. Performing the assay can include diluting the sample 225 (e.g., using Tris-HCl), mixing the diluted sample 225 with reagents / primers for GS PCR 235 and AGG PCR 240 in the assay system (e.g., PCR test tubes or well plates), PCR amplification (e.g., reaction cycles) for both GS PCR 235 and AGG PCR 240, post-PCR cleanup (e.g., purification of the PCR product), and separation of the PCR product (e.g., capillary electrophoresis or a high-resolution gel such as Lonza MetaPhor Agarose). In some examples, the FXS AGG PCR assay is performed in a PCR plate (e.g., a 96-well PCR assay plate). GS PCR 235 and AGG PCR 240 may be performed in parallel on the same PCR plate, and controls may be loaded on the PCR plate for quality control. GS PCR 235 amplifies the region containing the CGG repeats of the FMR1 gene and is used to determine the length of the repeated CGG trinucleotide sequence within the FMR1 gene. Due to the high GC content of this region, a GC-rich PCR system can be prepared and used for reliable and robust GS PCR amplification. In some examples, the GC-rich PCR system includes a GC-rich PCR reaction buffer and a GC-rich resolving solution. The GC-rich PCR system may also include an enzyme blend of thermostable Taq DNA polymerase and Tgo DNA polymerase, which are thermostable enzymes with proofreading (3'-5' exonuclease) activity. The GC-rich PCR reaction buffer, the GC-rich resolving solution, and dimethyl sulfoxide (DMSO) as an additive enable robust amplification of this difficult CGG repeat region. AGG PCR 240 targets the CGG repeat region and detects all AGG interruptions within it on all alleles.The AGG PCR system is based on the GC rich PCR system, except that Taq DNA polymerase and Tgo DNA polymerase can be replaced by other DNA polymerases, such as KAPA2G Robust HotStart DNA polymerase, for robust amplification of AGG specific sequences with minimal background.

[0045] DNA polymerases such as Taq DNA polymerase, Tgo DNA polymerase, and KAPA2G Robust HotStart DNA polymerase can only make DNA when given a primer, which is a short nucleotide sequence that provides a starting point for DNA synthesis. In some examples, GS PCR 235 uses forward primer FRAX-F1 and reverse primer FRAX-R-6 adjacent to the target region (the CGG repeat region to be copied). In some examples, GS PCR 235 uses a fluorescently labeled reverse primer such as FRAX-R-6FAM. In some examples, AGG PCR 240 uses forward primer FXS-AGG-Forw, which is a chimeric primer containing three CGG repeats and an "A" nucleotide at the 3' end. Forward primer FXS-AGG-Forw specifically binds to the AGG interruption within the CGG repeat, thereby producing amplicons of various sizes. In some examples, AGG PCR 240 uses a fluorescently labeled reverse primer such as FRAX-R-6FAM, the same as GC PCR 235. In other examples, AGG PCR 240 uses a fluorescently labeled reverse primer different from that of GC PCR 235, such as FRAX-R-*FAM.

[0046] In some examples, post-PCR cleanup includes mixing the GS PCR 235 and AGG PCR 240 products with magnetic beads, washing with a wash solution such as 70% ethanol, air drying, and eluting the purified PCR product 245 to increase the signal-to-noise ratio. After amplification and post-PCR cleanup, the purified PCR product 245 can be loaded into an analyzer system 205 (e.g., a fluorescence-based separation instrument system), and amplification products such as CGG repeats and AGG interruptions are determined at block 250 and output as GS and AGG PCR CE raw data 255. In some examples, GS PCR 235 is optimized to detect large CGG repeats, for example, up to 300 CGG repeats (e.g., 262 repeats) can be consistently amplified and detected by GS PCR 235. In certain examples, the purified PCR product is detected or separated using CE, the amplicons are visualized as "stutter" peaks, and each peak is separated by one CGG repeat. In certain embodiments, a first parameter such as a long injection time or increased voltage is used to detect larger-sized alleles (e.g., greater than 69 CGGs), while a second parameter such as a short injection time or decreased voltage is used for accurate sizing of smaller-sized alleles (e.g., 69 CGGs or less). Extended alleles may be detected if the stutter extends beyond 55 CGG repeats, and PM alleles may be detected if the stutter extends between 45 and 200 CGG repeats.

[0047] In block 260 within analysis device system 205, a separation module 265 (e.g., any software and / or hardware capable of acquiring raw CE data and determining fragment sizes and amounts such as GeneMapper™) is used to identify peaks in the GS and AGG PCR CE raw data 255 based on previously established QC metrics. In some examples, both the AGG PCR assay and the GS PCR assay use the same parameter (e.g., the same long injection or increased voltage) CE analysis method of the gene mapper module 265 to identify peaks in the GS and AGG PCR CE raw data 255. The gene mapper module 265 outputs the size and abundance of each analyzed PCR amplicon as peak size and height, respectively, and then at block 270, it is processed by the AGG genotyping module 275 of the AGG genotyping subsystem 215. The AGG genotyping module 260 first checks the sizing criteria and peak height to ensure that the quality of the GS and AGG PCR CE raw data 255 is acceptable for analysis. Next, the AGG genotyping module 275 uses the GS-PCR peak data to calculate the number of CGG repeats for each allele. To determine the number and location of AGG interruptions for each allele, the AGG genotyping module 275 begins by iteratively searching for and identifying AGG peaks on the allele with fewer repeats, which typically corresponds to the normal allele. After removing all AGG peaks assigned to this allele from the peak output table of the gene mapper module 265, the AGG genotyping module 275 repeats the process for one or more PM alleles. In this way, it is prevented that the same peak is assigned to both the normal allele and the PM allele, which otherwise might underestimate the extension risk of one or more PM alleles. Next, the AGG genotyping module 275 uses the number of CGG repeats for each allele and the number and location of AGG interruptions for each allele to determine the AGG genotype of each allele.The AGG genotype of each allele and the optionally selected risk result of each sample are output as the final result 280 by the analyzer system 205. In some examples, all the thresholds and QC parameters used by the AGG genotyping module 275 are maintained in a separate configuration file and can be used across any number of FXS AGG PCR assays. III. AGG Genotyping Techniques

[0048] Figure 3 shows a process 300 for AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques (e.g., the FXS AGG PCR assay platform 200 described with respect to Figure 2). The process 300 starts at block 305, and the raw data is obtained from the FXS assay performed on the sample. In some examples, the raw data includes GS PCR data and AGG interrupted PCR data separated by CE. At block 310, for the first allele identified in the raw data, the expected AGG peak size is determined. In some examples, the first allele is the allele with the fewest number of repeats or the normal allele. In a particular example, the expected AGG peak size is calculated using Equation (1). Expected AGG peak size (base pairs) = ((N - L) * 3) + A Equation (1) Where N is the number of CGG repeats in the allele (e.g., the allele with the fewest number of repeats), L is the expected position (in terms of the number of CGG repeats) of the first AGG interruption, which may be between 8 and 10 CGG repeats, e.g., 9 CGG repeats, 3 is the number of base pairs per CGG repeat, A is the size of the PCR amplicon without repeats, and in some examples, the PCR amplicon has 132 base pairs, and thus A = 132.

[0049] In blocks 315 and 320, raw data is iteratively searched to identify one or more AGG peaks on the first allele using a first set of search spaces determined based on the predicted AGG peak size. The first set of search spaces can include one or more search spaces determined based on the predicted AGG peak size. In some examples, the search and identification include: (i) in block 315, determining a first search space of the first set of search spaces for the first allele based on the predicted AGG peak size; (ii) in block 315, searching the raw data and identifying an initial AGG peak on the first allele using the first search space; and (iii) in block 320, determining a second search space for the first allele based on the peak size of the initial AGG peak on the first allele, and in block 320, iteratively searching the raw data and identifying one or more additional AGG peaks on the first allele using the second search space. If the search does not result in identifying an initial AGG peak on the first allele using the first search space in block 315, the first search space may be modified (e.g., 30 bps may be subtracted from the predicted AGG peak size) in block 325, and blocks 315 and 320 may be executed using the modified first search space.

[0050] The first search space can be determined based on the first peak size range. In some examples, the first peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the predicted AGG peak size. In a particular example, the predetermined amount or size range adjustment threshold is + / -12 or 15 base pairs of the predicted first AGG peak size. The second search space can be determined based on a new or second peak size range. In some examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) near the peak size of the identified initial peak. In a particular example, the predetermined amount or size range adjustment threshold is + / -12 or 15 base pairs of the peak size of the identified initial peak. In other examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the modified peak size (the peak size adjustment threshold is subtracted from the peak size of the identified initial peak). In some embodiments, the peak size adjustment threshold is 15 to 50 bps (e.g., 30 bps). In a particular example, the predetermined amount or size range adjustment threshold is + / -12 or 15 base pairs of the modified peak size (24 or 30 bps is subtracted from the peak size of the identified initial peak).

[0051] The processes of blocks 315, 320, and 325 can be repeated for each additional allele (e.g., 1 or more PM alleles) until all peak data and alleles within the raw data are evaluated. For example, the raw data can be searched iteratively to identify 1 or more AGG peaks on the second allele using a second set of search spaces determined based on the predicted AGG peak size. The second set of search spaces can include 1 or more search spaces determined based on the predicted AGG peak size. In certain examples, some or all of the 1 or more search spaces of the second set of search spaces are the same as the 1 or more search spaces of the first set of search spaces. In some examples, the searching and identification include: (i) determining a first search space of a second set of search spaces for the second allele based on the predicted AGG peak size; (ii) searching the raw data and identifying an initial AGG peak on the second allele using the second search space; (iii) determining a second search space for the second allele based on the peak size of the initial AGG peak on the second allele; and (iv) iteratively searching the raw data and identifying 1 or more additional AGG peaks on the second allele using the second search space. As a result of the search, if no initial AGG peak is identified on the second allele using the first search space, the first search space may be modified (e.g., 24 bps can be subtracted from the predicted AGG peak size), and the process can continue with the modified first search space.

[0052] The second search space can be determined based on the second peak size range. In some examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the predicted AGG peak size. In a specific example, the predetermined amount or size range adjustment threshold is + / -12 or 15 base pairs of the predicted first AGG peak size. The second search space can be determined based on the new or second peak size range. In some examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) near the peak size of the identified initial peak. In a specific example, the predetermined amount or size range adjustment threshold is + / -12 or 15 base pairs of the peak size of the identified initial peak. In other examples, the second peak size range is a predetermined amount of base pairs or a size range adjustment threshold (+ / -) around the modified peak size (the peak size adjustment threshold is subtracted from the peak size of the identified initial peak). In some embodiments, the peak size adjustment threshold is 15 to 50 bps (e.g., 24 bps). In a specific example, the predetermined amount or size range adjustment threshold is + / -12 or 15 base pairs of the modified peak size (24 or 30 bps is subtracted from the peak size of the identified initial peak). In some examples, when one or more AGG peaks on the first allele or any other previous allele are identified, the raw data is searched iteratively, and one or more AGG peaks on the first allele or any other previous allele are removed from the raw data before one or more AGG peaks on the second allele or another subsequent allele are identified.

[0053] In block 330, the genotype of each allele evaluated in blocks 315, 320, and 325 is determined as the AGG genotype. In some examples, the genotyping comprises: (i) determining the number of CGG repeats downstream of the last AGG interruption on the first allele, the second allele, or any other allele based on the GS PCR data and one or more AGG peaks identified with the corresponding allele; (ii) determining the number of CGG repeats separated by any two adjacent AGG interruptions on the first allele, the second allele, or any other allele based on the GS PCR data and one or more AGG peaks identified with the corresponding allele, if any; and (iii) determining the number of CGG repeats preceding the first AGG interruption on the first allele, the second allele, or any other allele based on the GS PCR data and one or more AGG peaks identified with the corresponding allele. In some examples, the genotyping further comprises generating the AGG genotype of the first allele, the second allele, or any other allele based on the number of CGG repeats downstream of the last AGG interruption, the number of CGG repeats separated by any two adjacent AGG interruptions, if any, and the number of CGG repeats preceding the first AGG interruption.

[0054] Optionally, in block 335, using the genotype of one or more alleles determined at block 330 (e.g., the first allele, the second allele, or both the first allele and the second allele), a risk score for the subject associated with the sample is determined. The risk score identifies the risk that the subject will develop Fragile X-associated tremor / ataxia syndrome (FXTAS) or Fragile X-associated primary ovarian insufficiency (FXPOI), which are late-onset neurodegenerative diseases, or the risk of transmitting a full mutation allele to the subject's offspring, or any combination thereof. In block 340, the AGG genotype determined for one or more alleles determined at block 330 may be output. Output of the AGG genotype determined for one or more alleles determined at block 330 and an optional risk score may include providing the output to an end user and / or recording the output on a storage device (e.g., displaying the output on a user interface and / or storing the output in a results file of a database).

[0055] Figures 4A-4H show a simplified flowchart 400 illustrating an example of a process for AGG genotyping using an FXS AGG PCR assay platform and genotyping techniques (e.g., FXS AGG PCR assay platform 200 and the techniques described with respect to FIGS. 2 and 3). Process 400 is divided into the following sections: (i) data preparation (FIG. 4A), (ii) shortest or normal allele algorithm flow (FIGS. 4B and 4C), (iii) PM allele algorithm flow (FIGS. 4C, 4D, 4E, 4F, 4G), and (iv) AGG genotype flow for each allele (FIG. 4H). Data Preparation

[0056] Figure 4A shows that in block 402, GS and AGG PCR CE raw data including the size and abundance (e.g., peak size and height) of each PCR amplicon are obtained for one or more samples. In some examples, an initial CGG data set (part of the GS and AGG PCR CE raw data) is created for each of a first parameter (e.g., short injection time) analysis and a second parameter (e.g., long injection time) analysis. The first parameter analysis can return up to three alleles, and the second parameter analysis can return up to two alleles. A final CGG data set can be created that includes combinations of the initial CGG data sets created for each of the first parameter analysis and the second parameter analysis. For example, (i) peaks for short CGG (0 - 59 repeats) can be checked to ensure no duplicates are selected, (ii) a check for long CGG (60 - 90 repeats) can be performed to ensure that long alleles do not overlap with short alleles (e.g., a 58 - repeat short call and a 60 - repeat long call), and (iii) peaks for long CGG can be checked to ensure no duplicates are selected. In certain examples, a peak buffer of + / - 1 - 5 CGG repeats, e.g., 2 CGG repeats, can be used during the check for overlapping peaks of short and long CGG. Once duplicates and overlaps are checked and removed, a final CGG data set including short and long CGG can be created. In some examples, a final AGG data set (part of the GS and AGG PCR CE raw data) is created that includes combinations of the first parameter analysis and the second parameter analysis.

[0057] In block 404, check the sizing criteria and peak heights in the final CGG dataset and the AGG dataset to confirm that the raw data quality of GS and AGG PCR CE is acceptable for analysis. In some examples, the check can include comparing the sizing criteria and peak heights in the final CGG dataset and the AGG dataset to one or more quality control (QC) parameters or metrics to determine whether the raw data quality of GS and AGG PCR CE is acceptable for analysis. In a particular example, the QC parameter or metric includes a QC check threshold for the height ratio of the red dye (size marker) of 125 / 1000 that is greater than 2 (>2). Additionally or alternatively, the AGG-specific peak height QC parameter or metric can include a minimum RFU threshold for screening AGG-specific peaks. Normal alleles can have an RFU threshold greater than 1000 (>1000) RFU. On the other hand, if necessary, the PM allele can have an RFU threshold greater than 10% (>10%) of the lowest AGG peak in the normal allele, or greater than 200 (>200) RFU if no peak is detected in the normal allele. If the sizing criteria and peak heights in the final CGG dataset and the AGG dataset meet the QC parameter or metric, use the final CGG dataset to calculate the CGG repeat number of each allele (e.g., normal or short allele, and PM or long allele) in block 406. If the sizing criteria and peak heights in the final CGG dataset and the AGG dataset do not meet the QC parameter or metric, in block 408, the assay of the sample, the GS and AGG PCR CE raw data, and / or the size and abundance (e.g., peak size and height) of each PCR amplicon from the GS and AGG PCR CE raw data are rejected and the process ends.

[0058] In block 406, a determination is made as to whether the calculated number of CGG repeats from each allele is within a predetermined range. In some examples, the FXS AGG PCR assay is a reflex assay from a Fragile X PCR diagnostic assay, and the reflex assay is initiated when the Fragile X PCR diagnostic assay identifies one or more samples having at least 45-100 CGG repeats (e.g., 55-90 CGG repeats) in the PM allele of the FMR1 gene. In a particular example, the reflex assay is initiated when the Fragile X PCR diagnostic assay identifies one or more samples 225 having 55-90 repeats in the PM allele of the FMR1 gene. Thus, a minimum CGG threshold (MinCGG) of 40-60 repeats and a maximum CGG threshold (MaxCGG) of 75-110 repeats can be set so that AGG genotyping is reliably performed only on one or more samples having at least 40-110 CGG repeats (e.g., 55-90 CGG repeats) in the PM allele of the FMR1 gene. In a particular embodiment, the MinCGG threshold is set at 55 repeats, the MaxCGG threshold is set at 90 repeats, and the predetermined range of the number of CGG repeats is between 55-90 repeats (the risk of expansion of the PM allele having more than 90 repeats is greater than 50% regardless of the number of AGG interruptions). If the number of CGG repeats identified in the final CGG dataset is within the predetermined range, the AGG genotyping process follows the normal or short allele algorithm flow shown in FIG. 4B. If the number of CGG repeats identified in the final CGG dataset is outside the predetermined range, the number of CGG repeats is output in block 410 and the process ends. In block 410, the number of CGG repeats can be output along with a message stating that AGG genotyping was not performed. Output of the number of CGG repeats and an optional message can include providing the output to an end user and / or recording the output in a storage device (e.g., displaying the output on a user interface and / or storing the output in a results file of a database). Normal Allele Algorithm Flow

[0059] To determine the number and position of AGG interruptions for each allele, the process begins by repeatedly searching for and identifying AGG peaks on alleles that have a smaller number of repeats, which typically correspond to normal alleles. As shown in Figure 4B, in block 412, the expected AGG peak size is calculated for the allele. The expected AGG peak size can be calculated for the allele with the fewest number of CGG repeats (e.g., the normal allele). In some examples, the expected AGG peak size is calculated using Equation (1). Expected AGG peak size (base pairs) = ((N - L) * 3) + A Equation (1) Where N is the number of CGG repeats in the allele (e.g., the allele with the fewest number of repeats), L is the expected position (in terms of the number of CGG repeats) of the first AGG interruption, which may be between 8 and 10 CGG repeats, e.g., 9 CGG repeats, 3 is the number of base pairs per CGG repeat, and A is the size of the PCR amplicon without repeats. In some examples, the PCR amplicon has 132 base pairs without repeats, and thus A = 132.

[0060] In block 414, an initial peak is identified for an allele (e.g., the allele having the fewest number of repeats). Identifying the initial peak involves, in block 416, determining an initial peak size search space (peak size search space in base pairs) for the allele. The initial peak size search space can be determined based on an initial peak size range. In some examples, the initial peak size range is a predetermined amount or size range adjustment threshold (+ / -) of base pairs around the predicted AGG peak size (range = + / size range adjustment threshold). In a particular example, the predetermined amount or size range adjustment threshold is + / - 15 base pairs around the predicted AGG peak size. Once the initial peak size search space is determined, in block 418, the initial peak size search space can be used to search a final AGG data set and identify the highest peak within the initial peak size search space based on criteria for the highest peak. In some examples, the criteria for identifying the highest peak within the initial peak size search space include (i) the peak size being greater than a minimum peak size threshold (in some embodiments, the minimum peak size threshold is set to 132 base pairs), (ii) the peak height being greater than a minimum peak height threshold (in some embodiments, the minimum peak height threshold is set to 200 RFU), (iii) the peak height being greater than a background peak height threshold (in some embodiments, the background peak height threshold is set to 1000 RFU), or (iv) any combination thereof. The criteria for identifying the highest peak in the initial peak size search space provide various technical advantages including (i) the minimum peak size threshold is the size of a PCR amplicon without repeats and can be used as a criterion to prevent false peak calls, (ii) the minimum peak height threshold can be used to prevent the calling of background low-level peaks, and (iii) the background peak height threshold can be used as the minimum peak height required for an AGG peak in a normal allele and can also be used to prevent the calling of background peaks.

[0061] If a peak meeting the criteria of the highest peak is found within the initial peak size search space, at block 420, the found peak is identified as the initial peak. Further, at block 420, the height of the initial peak is set as (i) the height of the previous peak height threshold, (ii) the height of the lowest normal height threshold, and (iii) the height of the first normal peak threshold. This setting provides various technical advantages including preventing the calling of background or stutter (false positive) peaks. Once the initial peak is identified for an allele (e.g., the allele having the fewest number of repeats), the peak data associated with the initial peak can be removed from the final AGG data set. After removing the AGG peak assigned to this allele from the output table of the final AGG data set, the process repeats for another search space within the same allele or starts a new process for another allele (e.g., the PM allele). Removing the discovered AGG peak assigned to an allele (in this block or any other block described herein) from the output table of the final AGG data set prevents the same peak within the output table of the final AGG data set from being discovered repeatedly and / or being assigned to multiple alleles, which otherwise could result in an underestimation of the PM allele extension risk, achieving the technical advantage. The removal of data from the output table also reduces the complexity and time of downstream processing.

[0062] If no peak meeting the criteria of the highest peak is found within the initial peak size search space, in block 422, the peak size adjustment threshold is subtracted from the predicted AGG peak size (or the newly calculated predicted AGG peak size) determined for the allele in block 422 to obtain a new predicted AGG peak size. The new predicted AGG peak size is used to identify the initial peak by repeating block 414. This adjustment helps shift / enlarge the initial peak size search space to find the initial peak within the final AGG data set. In some embodiments, the peak size adjustment threshold is 15 - 50 bps (e.g., 30 bps).

[0063] Once an initial peak for an allele is identified in block 414, additional AGG peaks for the allele (e.g., the allele with the fewest number of repeats) can be identified. In block 424, additional peaks for the allele are identified. Identifying the additional peaks involves, in block 426, subtracting the peak size adjustment threshold from the peak size of the initial peak found in block 414 to obtain the next predicted AGG peak size. In block 428, based on the next predicted AGG peak size, the next peak size search space (base pair peak size search space) for the allele is determined. The next peak size search space can be determined based on the next peak size range around the next predicted AGG peak size. In some examples, the next peak size range is a predetermined amount or size range adjustment threshold (+ / -) of base pairs around the next predicted AGG peak size. In a particular example, the predetermined amount or size range adjustment threshold is + / - 15 base pairs of the next predicted AGG peak size.

[0064] Once the next peak size search space is determined, at block 430, the second search space can be used to search the final AGG data set to identify the highest peak within the next peak size search space based on the criteria of the highest peak. In some examples, the criteria for identifying the highest peak within the next peak size search space are: (i) the peak size is greater than a minimum peak size threshold (in some embodiments, the minimum peak size threshold is set to 132 base pairs), (ii) the peak height is greater than a minimum peak height threshold (in some embodiments, the minimum peak height threshold is set to 200 RFU), (iii) the peak height is greater than a background peak height threshold (in some embodiments, the background peak height threshold is set to 1000 RFU), (iv) if an initial peak is found within the initial peak size search space, the peak height of the additional found peak is the previous peak height threshold (set at block 420) * a secondary normal height percentage threshold (e.g., 50%), or (v) any combination thereof. The criteria for identifying the highest peak in the next peak size search space provide various technical advantages, including: (i) the minimum peak size threshold is the size of a non-repeating PCR amplicon and can be used as a criterion to prevent false peak calls, (ii) the minimum peak height threshold can be used to prevent the calling of background low-level peaks, and (iii) the background peak height threshold can be used as the minimum peak height required for an AGG peak in a normal allele.

[0065] If a peak is found within the next peak size search space that meets the criteria of the highest peak, in block 432, the found peak is identified as an additional peak. Further, the height of the additional peak is set as (i) the height of the previous peak height threshold, and (ii) if the height of the found peak is less than the height of the initial peak, the height of the found peak is set as the height of the minimum normal height threshold. This setting provides various technical advantages including preventing the calling of background or stutter (false positive) peaks. When an additional peak is identified for an allele (e.g., the allele with the fewest number of repeats), the peak data associated with the additional peak can be removed from the final AGG data set. After removing the AGG peak assigned to this allele from the output table of the final AGG data set, the process repeats for another search space within the same allele or starts a new process for another allele (e.g., the PM allele).

[0066] In block 434, a determination is made as to whether the final AGG data set should continue to be searched for additional peaks (whether the output table of the final AGG data set has more peaks for short alleles). In some embodiments, the determination of whether to continue searching the final AGG data set is based on whether the size limit of the PCR amplicon (e.g., 132 bp) has been reached for the short / normal alleles. If it is determined that there are additional peaks in the final AGG data set for an allele, block 424 is repeated. In block 426, the peak size adjustment threshold is subtracted from the previously calculated next predicted AGG peak size in order to obtain the next predicted AGG peak size (essentially, the third, fourth, fifth, etc. predicted AGG peak sizes). This adjustment shifts / expands the additional peak size search space to help find additional peak(s) within the final AGG data set. Blocks 424 and 434 are repeated until all peaks for an allele (e.g., the allele with the fewest number of repeats) are identified within the final AGG data set. Once all peaks for an allele (e.g., the allele with the fewest number of repeats) are identified within the final AGG data set, the AGG genotyping process continues for each of the remaining alleles in block 440. PM Allele Algorithm Flow

[0067] In block 440, an initial peak is identified for another allele within the final AGG data set (e.g., a PM allele having a number of repeats between 55 and 90). Identifying the initial peak includes, in block 442, determining an initial peak size search space (peak size search space in base pairs) for the other allele. The initial peak size search space can be determined based on an initial peak size range. In some examples, the initial peak size range is a predetermined amount or size range adjustment threshold (+ / -) of base pairs around the predicted AGG peak size determined in block 412. In a particular example, the predetermined amount or size range adjustment threshold is + / - 12 base pairs of the predicted AGG peak size. In a particular example, for another allele (e.g., a PM allele) having a number of CGG repeats greater than 65: if the lower end of the initial peak size range is less than (less than) the first normal peak threshold set in block 420, set the lower end of the initial peak size range to the first normal peak threshold.

[0068] Once the initial peak size search space is determined, in block 444, the initial peak size search space can be used to search the final AGG data set, and the highest peak within the initial peak size search space can be identified based on the highest peak criterion. In some examples, the criteria for identifying the highest peak within the initial peak size search space include (i) the peak size being greater than a minimum peak size threshold (in some embodiments, the minimum peak size threshold is set to 132 base pairs), (ii) the peak height being greater than a minimum peak height threshold (in some embodiments, the minimum peak height threshold is set to 200 RFU), (iii) the peak height being greater than the lowest normal height threshold (set in block 420 or 432) * the minimum percentage PM threshold (in some embodiments, the minimum percentage PM threshold is set to 10%), or (iv) any combination thereof. The criteria for identifying the highest peak in the initial peak size search space include (i) the minimum peak size threshold being the size of a PCR amplicon without repeats and can be used as a criterion to prevent false peak calls, (ii) the minimum peak height threshold can be used to prevent the calling of background low-level peaks, and (iii) by using the height of the peak previously identified as a threshold for AGG peak search in the PM allele, a dynamic DNA input amount control incorporated for the AGG peak height in the PM allele is provided, including various technical advantages. More specifically, the AGG peak size of the PM allele is very variable for the 55 - 90 repeat range, and the corresponding AGG peak height also varies dramatically, which is also greatly affected by variable amounts of DNA input. Lower DNA input results in lower peak heights in normal alleles and vice versa. Instead of using a fixed peak height threshold screening for AGG peaks in the PM allele, this dynamic threshold allows the process to more accurately identify AGG peaks.

[0069] If a peak is found within the initial peak size search space that meets the criteria of the highest peak, the peak found is identified as the initial peak. Further, the height of the initial peak is set to (i) the height of the previous peak height threshold, and (ii) the temporarily set previous peak height (which is used if a secondary peak is found to determine which peak to compare the secondary peak to). This setting provides various technical advantages, including preventing the calling of background or stutter (false positive) peaks. When an initial peak is identified for another allele (e.g., a PM allele having a number of repeats between 55 and 90), the peak data associated with the initial peak can be removed from the final AGG data set. After removing the AGG peak assigned to this allele from the output table of the final AGG data set, the process repeats for another search space within the same allele or starts a new process for another allele (e.g., another PM allele).

[0070] In block 450, the final AGG data set is searched using the initial peak size search space, and the next highest peak or secondary peak within the initial peak size search space is identified based on the criteria of the next highest peak. In some examples, the criteria for identifying the next highest peak within the initial peak size search space are: (i) whether the peak size of the secondary peak is greater than or less than the secondary peak threshold (e.g., 7) from the peak size of the initial peak identified in block 440; (ii) comparing the size of the secondary peak with the initially discovered peak, and if the peak size of the secondary peak is less than the peak size of the initial peak, the height of the secondary peak should be greater than the set previous peak height threshold * secondary PM peak height threshold (in some embodiments, the secondary PM peak height threshold is set to 80%) called in the same search space in block 446, or if the peak size of the secondary peak is greater than the peak size of the initial peak, the height of the secondary peak should be greater than the set previous peak height threshold * secondary PM peak height threshold (in some embodiments, the secondary PM peak height threshold is set to 80%) from the previous search space in block 432, or (iii) any combination thereof. When the process is in the initial peak size search space, since the threshold has not yet been established, the threshold used is 0. After the initial peak size search space, the previous peak height threshold has already been established. The criteria for identifying the next highest peak or secondary peak within the initial peak size search space provide various technical advantages, including avoiding the omission of additional AGG peaks within the same search window after the first peak is identified.

[0071] If a peak that meets the criteria of the next highest peak is found within the initial peak size search space, at block 454, the found peak is identified as a secondary peak. In some examples, if a peak that meets the criteria for identifying the next highest peak as a secondary peak is found within the initial peak size search space and the peak size of the secondary peak is smaller than the peak size of the initial peak, the height of the secondary peak is set as the previous peak height threshold. This setting provides various technical advantages including preventing the calling of background or stutter (false positive) peaks. When a secondary peak is identified for another allele (e.g., a PM allele having a number of repeats between 55 - 90), the peak data associated with the secondary peak can be removed from the final AGG data set. At block 456, the initial peak and / or secondary peak for another allele is reported as follows (e.g., the processor provides the initial peak and / or secondary peak to a display for the user to view via a user interface, a storage device for recording purposes, another component of a computing device for future processing, or an output device such as a printer for a hard copy of the report): (i) if no peak is found (no initial peak and no secondary peak), report that no peak was found, (ii) if one peak is found (either an initial peak or a secondary peak), report the found peak, or (iii) if both peaks are found (an initial peak and a secondary peak), report the peaks in descending order since the algorithm starts at the end of CGG. For example, if initial peak size 1 < secondary peak size 2: return peak 2, peak 1; while if secondary peak size 2 > initial peak size 2: return peak 1, peak 2. The AGG genotyping process continues for each of the remaining peaks of the other allele at block 460.

[0072] If no peak that meets the criteria of the next highest peak is found within the initial peak size search space, in block 458, the peak size adjustment threshold is subtracted from the predicted AGG peak size (or the newly calculated predicted AGG peak size previously calculated) determined in block 412 to obtain a new predicted AGG peak size. The new predicted AGG peak size is used to identify the initial peak by repeating block 440. This adjustment helps to shift / enlarge the initial peak size search space to find the initial peak within the final AGG dataset. In some embodiments, the peak size adjustment threshold is between 15 and 50 bps (e.g., 24 bps).

[0073] When an initial peak is identified for an allele in block 440 / 450, additional AGG peaks may be identified for the allele (e.g., the allele with a greater number of repeats). In block 460, additional peaks are identified for the allele. Identifying the additional peaks includes subtracting the peak size adjustment threshold from the peak size of the initial peak found in block 440 / 450 to obtain the next predicted AGG peak size in block 462. In block 464, based on the next predicted AGG peak size, the next peak size search space (base pair peak size search space) is determined for the allele. The next peak size search space may be determined based on the next peak size range. In some examples, the next peak size range is a predetermined amount or size range adjustment threshold (+ / -) of base pairs near the peak size of the initial peak found in block 440 / 450. In a specific example, the predetermined amount or size range adjustment threshold is + / - 12 base pairs of the peak size of the initial peak. In a specific example, if the lower end of the next peak size range is less than the first normal peak threshold, the lower end of the next peak size range is set to the first normal peak threshold.

[0074] Once the second search space is determined, in block 464, the final AGG data set can be searched using the next peak size search space to identify the highest peak within the next peak size search space based on the highest peak criterion. In some examples, the criteria for identifying the highest peak within the next peak size search space are: (i) the peak size is greater than the minimum peak size threshold (in some embodiments, the minimum peak size threshold is set to 132 base pairs), (ii) the peak height is greater than the minimum peak height threshold (in some embodiments, the minimum peak height threshold is set to 200 RFU), (iii) the peak height is greater than the lowest normal height threshold (set in block 420 or 432) * the minimum percentage PM threshold (in some embodiments, the minimum percentage PM threshold is set to 10%), (iv) the peak height is greater than the previous peak height threshold (set in block 446 or 454) * the secondary PM peak height percentage threshold (e.g., 80%), or (v) any combination thereof. The criteria for identifying the highest peak in the next peak size search space include: (i) the minimum peak size threshold is the size of the PCR amplicon without repeats and can be used as a criterion to prevent false peak calls, (ii) the minimum peak height threshold can be used to prevent the calling of background low-level peaks, and (iii) by using the height of the peak previously identified as a threshold for AGG peak search in the PM allele, dynamic DNA input amount control incorporated with respect to the height of the AGG peak in the PM allele is provided, providing various technical advantages.

[0075] If a peak is found within the next peak size search space that meets the criteria of the highest peak, at block 468, the found peak is identified as an additional peak. Further, the size of the additional peak is set as the temporary previous peak height (which is used if a secondary peak is found to determine which peak to compare the secondary peak to). When an additional peak is identified for another allele, the peak data associated with the additional peak can be removed from the final AGG data set.

[0076] At block 470, the final AGG data set is searched using the next peak size search space to identify a secondary peak or the next highest peak within the next peak size search space. In some examples, the criteria for identifying a secondary peak within the next peak size search space are: (i) whether the peak size of the secondary peak is greater than or less than the secondary peak threshold (e.g., 7) from the peak size of the additional peak size identified at block 460, (ii) comparing the size of the secondary peak to the additional peak(s) found, and if the peak size of the secondary peak is less than the peak size of the additional peak, the height of the secondary peak should be greater than the temporary previous peak height * secondary PM peak height threshold (in some embodiments, the secondary PM peak height threshold is set to 80%) set at block 468, or if the peak size of the secondary peak is greater than the peak size of the additional peak, the height of the secondary peak should be greater than the block 454 or 468 * secondary PM peak height threshold (set to the threshold set at 80% in some embodiments), or (iii) any combination thereof. The criteria for identifying a secondary peak or the next highest peak within the next peak size search space include various technical advantages such as (i) preventing call-background or stutter (false positive) peaks, and (ii) avoiding the omission of additional AGG peaks within the same search window after the first peak is identified.

[0077] If a peak that meets the criteria for a secondary peak is found within the second search space, the peak found is identified as the secondary peak. In some examples, a peak that meets the criteria for identifying the next highest peak as the secondary peak is found within the next peak size search space, and if the peak size of the secondary peak is smaller than the peak size of the additional peak, the height of the secondary peak is set as the previous peak height threshold. This setting provides various technical advantages, including preventing the calling of background or stutter (false positive) peaks. When a secondary peak is identified for another allele (e.g., a PM allele having a number of repeats between 55 and 90), the peak data associated with the secondary peak can be removed from the final AGG data set. At block 456, the initial peak and / or secondary peak for another allele is reported as follows (e.g., the processor provides the initial peak and / or secondary peak on a display for the user to view via a user interface, a storage device for recording purposes, another component of a computing device for future processing, or an output device such as a printer for a hard copy of the report): (i) if no peak is found (no initial peak and no secondary peak), report that no peak was found, (ii) if one peak is found (initial peak or secondary peak), report the peak found, or (iii) if both peaks are found (initial peak and secondary peak), report the peaks in descending order since the algorithm starts at the end of the CGG. For example, if initial peak size 1 < secondary peak size 2: return peak 2, peak 1; while if secondary peak size 2 > initial peak size 2: return peak 1, peak 2.

[0078] In block 478, a decision is made as to whether the final AGG data set should continue to be searched for additional peaks (i.e., whether the output table of the final AGG data set has more peaks for long alleles). In some embodiments, for PM alleles having a repeat size less than 65, the decision of whether to continue searching the final AGG data set is based on whether the size limit of the PCR amplicon without repeats for the long / PM allele (e.g., 132 bp) has been reached. In other embodiments, for PM alleles having a repeat size greater than 65, the decision of whether to continue searching the final AGG data set is based on whether the maximum size AGG peak found in the normal / short allele for the long / PM allele has been reached. If it is determined that there are more peaks in the final AGG data set for an allele, blocks 460 and 470 are repeated. In block 460, the peak size adjustment threshold is subtracted from the previously calculated next predicted AGG peak size to obtain the next predicted AGG peak size (essentially, the third, fourth, fifth, etc. predicted AGG peak sizes). This adjustment shifts / enlarges the additional peak size search space to help find additional peak(s) within the final AGG data set. Blocks 460 and 470 are repeated until all peaks of the allele (e.g., an allele having a greater number of repeats) are identified within the final AGG data set. When all peaks of an allele (e.g., an allele having the fewest number of repeats) are identified within the final AGG data set, in block 480, a decision is made as to whether the final AGG data set should continue to be searched for additional alleles (i.e., whether the output table of the final AGG data set has peaks for other alleles). If it is determined that there are more alleles, the AGG genotyping process continues with block 440 for each of the remaining alleles. If it is determined that there are no more alleles, the process continues to block 490, where the alleles are genotyped based on the AGG break data obtained in blocks 412 - 476. AGG genotype flow

[0079] In block 490, the number of CGG repeats downstream of the last AGG interruption on the allele (e.g., the normal allele or the PM allele) is calculated using the CGG repeats identified in blocks 402 - 410 and the AGG peaks identified in blocks 412 - 480. In some embodiments, the number of CGG repeats downstream of the last AGG interruption on the allele is calculated using Equation (2). Number of CGG repeats downstream of the last AGG interruption (using CCG#) = ((P n - A) / 3)-1 Equation (2) where P n is the peak size corresponding to the last AGG interruption on the allele (e.g., the minimum AGG interruption on each allele), A is the size of the PCR amplicon without repeats, and in some examples, the PCR amplicon has 132 base pairs and thus A = 132, 3 is the number of base pairs per CGG repeat, and 1 represents a single AGG interruption. The position of the last AGG interruption can be used to start the resulting string of the allele (constructed backwards). AGG(CGG) n - where "n" in the formula is the result of Equation (1) (i.e., the expected first AGG peak size).

[0080] Thereafter, the remaining AGG peaks identified in blocks 412 - 480 are cycled until all are processed for the allele. In some embodiments, the number of CGG repeats separated by any two adjacent AGG interruptions (if present) is calculated using Equation (3). Number of CGG repeats separated by any two adjacent AGG interruptions (using CCG#) = ((P n - P n-1 ) / 3)-1 Equation (3) where P n and P n-1is the peak size corresponding to adjacent AGG interruptions on the allele, 3 is the number of base pairs for each CGG repeat, and 1 indicates a single AGG interruption. After processing the number of CGG repeats separated by any two adjacent AGG interruptions, the resulting string may be updated by adding it to the beginning of a string of the same format (constructed in reverse). AGG(CGG) n - where "n" is the result of formula (1) (i.e., the expected first AGG peak size).

[0081] The number of CGG repeats preceding the first AGG interruption can be calculated by subtracting the total number of AGG interruptions determined by the AGG peaks identified in blocks 412 - 480, the total number of CGG repeats separating adjacent AGG interruptions calculated by formula (3), and the number of CGG repeats downstream of the last AGG interruption calculated by formula (2), from the total number of CGG repeats of the allele identified by the GS-PCR assay and identified in blocks 402 - 410. After processing the number of CGG repeats preceding the first AGG interruption, the resulting string may be updated by adding it to the beginning of a string of the same format (constructed in reverse). AGG(CGG) n - where "n" is the result of formula (1) (i.e., the expected first AGG peak size). The final resulting string after processing the number of CGG repeats preceding the first AGG interruption is the AGG genotype of the allele. The process of block 490 can be repeated for each allele identified for the sample.

[0082] In an optional block 492, an elongation risk score can be calculated based on the AGG genotype determined for each allele of the sample. The risk score can be calculated using the number of CGG repeats and the AGG interruptions therein, as shown in Table 1. In some examples, the risk score identifies the risk that a patient will develop a late-onset neurodegenerative disease, fragile X-associated tremor / ataxia syndrome (FXTAS) or fragile X-associated primary ovarian insufficiency (FXPOI), or the risk of transmitting a full mutation allele to an offspring of the subject, or any combination thereof. In a particular example, if there are more than five AGG interruptions, the number of AGG interruptions can be set to five for the calculation of the risk score(s) because five interruptions and more than five interruptions generally have the same clinical outcome.

Table 1

[0083] In block 494, the determined AGG genotypes for each allele can be output. The output of the determined AGG genotypes for each allele and optional risk scores can include providing the output to an end user and / or recording the output on a storage device (e.g., displaying the output on a user interface and / or storing the output in a results file of a database). As will be appreciated, the process flow of AGG genotyping using the FXS AGG PCR assay platform and genotyping techniques described with respect to FIGS. 4A-4H enables a user to process all samples in parallel in a batch without human error and seamlessly incorporate the results into a reporting summary that significantly reduces turn-around time. With respect to the reporting summary, the AGG genotype and optional risk scores inform a clinician / geneticist how to counsel a subject (e.g., a patient) through current or future pregnancies, e.g., family planning options (e.g., in vitro diagnostics (IVD)) for a future pregnancy, or what options to offer / prepare for a current pregnancy. Thus, as a final step, the clinician / geneticist provides advice to the subject or patient based on the AGG genotype and optional risk scores. The advice can include guidance regarding subsequent medical diagnostic tests (e.g., IVD), a detailed discussion of the genetic pattern of FXD, clinical findings for all three phenotypes (FXS, FXPOI, FXTAS), reproductive options if appropriate, conversations with children and at-risk family members over the long term, considerations for testing asymptomatic children, research opportunities, family support, and family planning options such as referrals to medical, developmental, and as indicated psychological providers.

[0084] FIG. 5 shows an exemplary computing device 500 suitable for use in a system and method for AGG genotyping using the FXS AGG PCR assay platform and genotyping technology according to the present disclosure. The exemplary computing device 500 includes a processor 505 that communicates with a memory 510 and other components of the computing device 500 using one or more communication buses 515. The processor 505 is configured to execute processor-executable instructions stored in the memory 510 to search for and identify AGG peaks present in raw data, determine an allelic AGG genotype, and / or execute one or more methods for determining a patient's risk score, according to different examples such as some or all of the exemplary processes 300 or 400 described above with respect to FIGS. 3 and 4A-4H. In this example, the memory 510 stores processor-executable instructions that provide AGG peak analysis 520 and AGG genotyping 525, as described above with respect to FIGS. 2, 3, and 4A-4H.

[0085] In this example, computing device 500 also includes one or more user input devices 530, such as a keyboard, mouse, touch screen, microphone, etc., to receive user input. Computing device 500 also includes a display 535 to provide visual output to the user, such as a user interface. Computing device 500 also includes a communication interface 540. In some examples, communication interface 540 can enable communication using one or more networks, including local area networks ("LAN"), wide area networks ("WAN") such as the Internet, metropolitan area networks ("MAN"), point-to-point or peer-to-peer connections, etc. Communication with other devices can be achieved using any suitable network protocol. For example, one suitable network protocol can include Internet Protocol ("IP"), Transmission Control Protocol ("TCP"), User Datagram Protocol ("UDP"), or combinations thereof such as TCP / IP or UDP / IP.

Example

[0086] IV. Example The systems and methods implemented in various embodiments can be better understood by referring to the following examples. Example 1: FRAX AGG Interrupted PCR Assay and AGG Genotyping Algorithm Specimens Used and Reproducibility

[0087] Ten blood samples (S1 - S10) were extracted in a first external test laboratory containing five samples that had been externally tested for AGG interruption. Five samples (S11 - S15) were extracted from blood in a first clinical laboratory and had previously undergone a fragile X PCR assay. These five samples are expected to represent samples that can be received for clinical trials. Twenty-three samples were extracted and tested for AGG interruption in a second external test laboratory (S16 - S38). The AGG genotypes of these twenty-three samples were blinded to the operator prior to use in this example. In addition to the above samples, twenty-five samples used in assay development were also tested to compare the automated FRAX AGG PCR genotype determination algorithm calls from those of manual genotyping. For intra-assay reproducibility, five samples including three with known AGG genotype results were analyzed in triplicate. These samples were also tested in two additional runs once for inter-assay reproducibility, one of which was performed by a different operator using different lots of KAPA2G Robust Hotstart Enzyme mix and AGG PCR primer mix. Lot and expiration date information is listed in Table 2. Two models of thermal cyclers and two ABI 3730xl instruments were also included to evaluate reproducibility between instruments.

Table 2

[0088] Genotype calls by the FRAX AGG genotype determination algorithm were reproducible in all five samples (Table 3). It should be noted that for the highly repetitive GC-rich regions separated by CE, minor genotype differences between replicates are expected due to experimental variation. Additionally, it has previously been confirmed that the resolution of CGG repeat numbers determined by GS-PCR varies by only 1 - 4 repeats depending on the length of the repeat. Therefore, considering the inherent variation between runs, a difference of 1 - 2 repeats (i.e., 3 - 6 bases) between replicates is within the range of variation expected for normal or PM alleles, or for the number of CGGs interrupted by AGG.

Table 3-1

Table 3-2

[0089] Analysis sensitivity and specificity were first evaluated using five blood DNA samples (Samples S1 - S5) from an external testing laboratory with previously reported AGG genotypes using a validated platform. CE electropherograms of the AGG PCR products are shown in Figures 6A and 6B. The AGG genotypes obtained from the external laboratory and the FRAX AGG PCR assay are listed in Table 4. All controls passed and no false positives or false negatives were called. The FRAX AGG PCR assay identified the same number of AGG interruptions in all samples as expected, and four of them showed matching genotypes (Table 4). As described above, the differences in the position and / or total number of repeats of the AGG interruption by 1 - 2 repeats between the results from this example and those from external tests were as expected and considered to be in agreement.

Table 4

[0090] For sample S2, the total number of AGG interruptions identified by the FRAX AGG PCR assay was consistent with the expected results. However, the allelic phasing of the AGG interruptions did not match the results of the external laboratory (Table 3). In the tests by the external laboratory, only one interruption was assigned to the normal allele (allele 1), while in the FRAX AGG PCR assay, two AGG-specific peaks were identified (159 and 189 bp, Figure 6A). Notably, both of them had 29 total repeats in their normal alleles, and two other samples with AGG peaks of the same size as sample S2 were consistent for the two AGG interruptions when compared with the expected results (Figure 6A and Figure 6B, samples S2 - S4). As described above, since the FRAX AGG PCR genotyping algorithm was initially developed to first assign AGG peaks to normal alleles, the genotyping differences observed for sample S2 may be due to which allele the AGG peaks are first assigned to.

[0091] To address the issue of genotyping inconsistencies, two approaches were used. As part of the FRAX PCR assay, the TRP-PCR assay screens for expanded FRAX alleles by using CGG repeat primers paired with GS primers to amplify the CGG repeat region. In the presence of an AGG interruption, the affinity of the CGG repeat primer for its target sequence decreases, thereby reducing PCR efficiency and resulting in a "dip" in the CE electropherogram (Figure 7A). If one allele has more repeats than the other (e.g., PM vs. normal allele), or if the AGG interruptions coincide at the same position in both alleles, the signal intensity of these dips will be close to the baseline. Thus, for sample S2, if there were two AGG interruptions in the normal allele, the TRP-PCR electropherogram would show two AGG dips corresponding to the interruptions in the repeat range of the normal allele.

[0092] As shown in Figure 7A, the results of the TRP-PCR supported the AGG phasing results of the FRAX AGG PCR assay for sample S2. Similar profiles were also observed for samples S3 and S4. Interestingly, upon further examination, it was found that the signal intensity for one of the AGG dips in the normal allele for S2 decreased near the baseline level, indicating that the AGG interruption was likely present at the same position in both alleles (the 189 bp AGG dip in the S2 panel, Figure 7A). To confirm this, the normal and PM alleles were amplified, gel purified, and analyzed by the FRAX AGG PCR assay to separately determine the number of AGG interruptions on each allele. As expected, the FRAX AGG PCR CE electropherograms clearly identified two AGGs for the normal allele and four AGGs for the PM allele. Importantly, neither the FRAX AGG PCR assay nor the results from external laboratory tests identified overlapping AGG interruptions in both alleles, but it is noteworthy that this difference did not change the elongation risk score (<1%, see Table 1).

[0093] To further evaluate the analytical sensitivity, specificity, and accuracy of the FRAX AGG PCR assay and the FRAX AGG PCR genotyping algorithm, an additional 23 samples with known AGG genotypes from external laboratories were tested in replicates (Table 5). The samples were anonymized to the operator prior to testing, and all calls were manually analyzed using the FRAX AGG PCR genotyping algorithm.

Table 5-1

Table 5-2

Table 5-3

Table 5-4

[0094] As shown in Table 5, at least one AGG interruption was detected in all samples, and the AGG genotype calls from the FRAX AGG PCR assay matched in 21 out of 23 samples. The controls did not fail, and no false-negative or false-positive calls were made. Samples S25 and S37 were not genotyped by the FRAX AGG PCR assay because the number of repeats in the PM allele exceeded the 90-repeat limit specified for the FRAX AGG PCR genotyping algorithm, but the number of AGG interruptions per normal allele was consistent with the expected results.

[0095] The FRAX AGG PCR assay did not match the expected results for samples S17 and S23, which were likely caused by AGG interruptions at the same position in both alleles or by the order in which the alleles were assigned due to interruptions. In sample S17, the expected result for the AGG interruption in the PM allele was six, but the FRAX AGG PCR assay identified five AGGs (Table 5 and Figure 8A). Similar to sample S2 above, review of the TRP-PCR CE data showed that one of the AGG dips in the normal allele size range could potentially be located at the same position in both the normal and premutation alleles (159bp AGG dip, Figure 8B). FRAX AGG PCR using gel-purified normal and PM allele PCR products further confirmed that the AGG interruption represented by the 159bp peak occurred in both alleles (Figure 8C).

[0096] For sample S23, the AGG interruption represented by the 159 bp AGG PCR peak could be assigned to either the normal or PM allele based on two adjacent AGG peaks at 186 bp and 192 bp, which were both approximately 30 bp or approximately 10 repeats from the 159 bp AGG interruption (Figure 9A). In the external laboratory, 2 and 5 AGG interruptions were assigned to the normal and PM alleles, respectively. However, since the FRAX AGG PCR genotyping algorithm was developed to first assign the AGG interruption peak to the shorter allele, 3 AGGs were assigned to the normal allele and 4 AGGs were assigned to the pre-mutation allele (Table 5 and Figure 9A). TRP-PCR analysis showed that there was only one AGG interruption corresponding to the 159 bp peak and that it could be located in either the normal or PM allele (Figure 9B). Analysis of each allele by FRAX AGG PCR clearly showed 2 and 5 AGG interruptions in the normal and PM alleles, respectively, consistent with the genotype from the external testing laboratory (Figure 9C).

[0097] In summary, the FRAX AGG PCR assay agreed on the total number of AGG interruptions in 27 out of 28 samples. The one sample (S17) that differed from the expected result in terms of the number of AGGs was most likely due to the order of the AGG peak data assigned by the FRAX AGG PCR genotyping algorithm compared to that from external testing. Additionally, it was identified that both the FRAX AGG PCR assay and the external laboratory testing have limitations when the AGG interruption occurs at the same position in both the normal and PM alleles. Given that CE and GeneMapper detect and output only one AGG PCR peak, this was not surprising. However, as shown in Table 1, when the number of interruptions is 2 or more and the repeat size is less than 65, this difference does not affect the elongation risk score. For repeats over 65 and at least one AGG interruption, the genotypes were consistent. The AGG PCR fragments corresponding to interruptions in these longer repeats are much larger than those in the normal allele that allow for clear allele phasing by the FRAX AGG PCR genotyping algorithm. FRAX AGG PCR Genotyping Algorithm Test

[0098] The results from the FRAX AGG genotyping algorithm were 100% consistent with the results from manual genotyping analysis based on 38 samples run in this example. To further test the concordance, the 25 samples used during assay development were also genotyped manually and by the FRAX AGG genotyping algorithm, and the results were 100% concordant between the two methods. The algorithm not only functioned equivalently to manual analysis for genotyping but also performed well in identifying any samples where the quality of the input CE data varied or the genotype output was outside the normal range and flagging them for "manual review."

[0099] As will be understood, it is possible to encounter samples that meet the test criteria (e.g., samples of women with PM alleles in the range of a total repeat of 55 - 90), but do not contain an AGG interruption. For example, the normal allele of a sample may have 21 total repeats and no AGG interruption. That PM allele may also be without AGG, similar to sample S1 or S4 (Table 4). The absence of an AGG PCR product can be difficult to distinguish from low assay performance. To ensure accurate calling, in addition to repeating the assay and confirming the "no AGG" result, the FRAX TRP-PCR assay CE electropherogram of the sample can also be reviewed. In contrast to what is observed in FIGS. 7A, 8B, and 9B, a smooth and stepwise decrease in CGG stutter peak height without a sharp decrease in signal supports a "no AGG" call (FIG. 10). Considering that typically only female PM allele carriers will be tested, based on a total of approximately 95 samples tested during assay development and validation that did not encounter samples without an AGG interruption, the likelihood of a sample without an AGG interruption would be rare. Conclusion

[0100] The technical performance of the FRAX AGG interruption PCR assay was determined to be reproducible and robust with respect to the generation of AGG-specific PCR products and the automated results of AGG genotyping by the FRAX AGG genotyping algorithm.

[0101] Of the 28 samples with known AGG genotypes, 27 were consistent with the expected results. For one sample, the total number of AGG disruptions was consistent with the expected number, but the AGG genotype differed due to the order of alleles assigned by the FRAX AGG genotyping algorithm. However, this difference would not have had a clinical impact on the final risk score of this sample. Inter- and intra-assay reproducibility was also 100% for the determination of the AGG genotype in all five samples, including three samples with known AGG genotypes. The FRAX AGG disruption PCR assay was able to withstand input DNA from 10 to 160 ng extracted from blood specimens, and assay performance remained robust for gDNA extracted from blood specimens. Additionally, the FRAX AGG genotyping algorithm identified AGG-specific peaks and was 100% consistent with manual genotyping analysis when determining the AGG genotype for all samples tested. V. Further Considerations

[0102] In the above description, specific details are shown to provide a complete understanding of the embodiments. However, it is understood that the embodiments can be practiced without these specific details. For example, circuits can be shown in block diagrams so as not to unnecessarily obscure the embodiments in excessive detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques can be shown without unnecessary detail so as not to obscure the embodiments.

[0103] The implementation of the above-described technologies, blocks, processes, and means can be carried out in various ways. For example, these techniques, blocks, processes, and means can be implemented in hardware, software, or a combination thereof. In the case of hardware implementation, the processing unit can be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, other electronic units designed to perform the above-described functions, and / or combinations thereof.

[0104] It should also be noted that embodiments can be described as a process shown as a flowchart, a flow diagram, a data flow diagram, a structural diagram, or a block diagram. A flowchart can describe operations as a sequential process, but many of the operations can be performed in parallel or simultaneously. Additionally, the order of the operations can be rearranged. A process ends when its operations are completed, but it can have additional steps not included in the figure. A process can correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its end corresponds to the return of the function to the calling function or the main function.

[0105] Furthermore, embodiments can be implemented by hardware, software, script language, firmware, middleware, microcode, hardware description language, and / or any combination thereof. When implemented in software, firmware, middleware, script language, and / or microcode, the program code or code segments for performing the necessary tasks can be stored in a machine-readable medium such as a storage medium. Code segments or machine-executable instructions can represent procedures, functions, subprograms, programs, routines, subroutines, modules, software packages, scripts, classes, or any combination of instructions, data structures, and / or program statements. A code segment can be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, and / or memory contents. Information, arguments, parameters, data, etc. can be passed, transferred, or transmitted via any suitable means including memory sharing, message passing, ticket passing, network transmission, etc.

[0106] In the case of firmware and / or software implementation, the methodology can be implemented with modules (e.g., procedures, functions, etc.) that perform the functions described herein. When implementing the methodology described herein, any machine-readable medium that tangibly embodies instructions can be used. For example, software code can be stored in memory. The memory can be implemented within or external to the processor. As used herein, the term "memory" refers to any type of long-term, short-term, volatile, non-volatile, or other storage medium and is not limited to any particular type of memory or number of memories, or the type of medium in which the memory is stored.

[0107] Furthermore, as disclosed herein, the terms "memory medium", "memory", or "memory" can represent one or more memories for storing data, including read-only memory (ROM), random access memory (RAM), magnetic RAM, core memory, magnetic disk storage media, optical storage media, flash memory devices, and / or other machine-readable media for storing information. The term "machine-readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, wireless channels, and / or various other storage media that can store or carry instructions and / or data.

[0108] The principles of the present disclosure have been described above in connection with specific devices and methods, but it should be clearly understood that this description is for illustrative purposes only and not a limitation on the scope of the present disclosure. The present invention provides, for example, the following items. (Item 1) Obtaining raw data from a Fragile X Syndrome (FXS) assay performed on a sample using a gene analysis apparatus, wherein the raw data includes gene-specific (GS) polymerase chain reaction (PCR) data and AGG interrupted PCR data separated by capillary electrophoresis, Determining an expected AGG peak size for a first allele identified in the raw data, Repeatedly searching the raw data and identifying one or more AGG peaks on the first allele using a first set of search spaces determined based on the expected AGG peak size, Determining the number of CGG repeats downstream of the last AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, Determining the number of CGG repeats preceding the first AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, Generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption, Providing the AGG genotype for the first allele A method comprising. (Item 2) The method according to item 1, further comprising determining the number of CGG repeats separated by any two adjacent AGG interruptions on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, wherein the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG interruption, the number of CGG repeats separated by any two adjacent AGG interruptions, and the number of CGG repeats preceding the first AGG interruption. (Item 3) Repeatedly search the raw data and identify one or more AGG peaks on the second allele using a second set of search spaces determined based on the predicted AGG peak size, Determine the number of CGG repeats downstream of the last AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele, Determine the number of CGG repeats preceding the first AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele, Generate an AGG genotype for the second allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption, Further comprising providing the AGG genotype for the second allele, The method according to item 1 or 2, wherein (i) the first allele is a normal allele and the second allele is a premutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a premutation allele and the second allele is a different premutation allele. (Item 4) Further comprising determining a risk score for the subject associated with the sample based on the AGG genotype generated for the first allele, the second allele, or both the first allele and the second allele, the risk score identifying the risk that the subject will develop a late-onset neurodegenerative disease, fragile X-associated tremor / ataxia syndrome (FXTAS) or fragile X-associated primary ovarian insufficiency (FXPOI), or the risk of transmitting a full mutation allele to the subject's offspring or any combination thereof. The method according to item 1, 2 or 3. (Item 5) Repeatedly searching the raw data and identifying one or more AGG peaks on the first allele, Determining a first search space of a first set of search spaces for the first allele based on the predicted AGG peak size, Searching the raw data to identify an initial AGG peak on the first allele using a first search space; Determining a second search space for the first allele based on the peak size of the initial AGG peak on the first allele; Repeatedly searching the raw data to identify one or more additional AGG peaks on the first allele using the second search space, comprising: The method according to item 1, 2, or 3, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the first allele based on the GS PCR data, the initial peak identified on the first allele, and the one or more additional peaks identified on the first allele. (Item 6) Repeatedly searching the raw data to identify one or more AGG peaks on the second allele, comprising: Determining a first search space of a second set of the search space for the second allele based on the predicted AGG peak size; Searching the raw data to identify an initial AGG peak on the second allele using a second search space; Determining a second search space for the second allele based on the peak size of the initial AGG peak on the second allele; Repeatedly searching the raw data to identify one or more additional AGG peaks on the second allele using the second search space; The method according to item 3 or 5, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the second allele based on the GS PCR data, the initial peak identified on the second allele, and the one or more additional peaks identified on the second allele. (Item 7) The method according to item 3 or 5, wherein when the one or more AGG peaks on the first allele are identified, the raw data is repeatedly searched and the one or more AGG peaks on the first allele are removed from the raw data before identifying the one or more AGG peaks on the second allele. (Item 8) One or more data processors, and when executed on the one or more data processors, cause the one or more data processors to obtain raw data from a Fragile X Syndrome (FXS) assay performed on a sample using a genetic analyzer, wherein the raw data includes gene-specific (GS) polymerase chain reaction (PCR) data and AGG interrupted PCR data separated by capillary electrophoresis, determine an expected AGG peak size for a first allele identified in the raw data, iteratively search the raw data and identify one or more AGG peaks on the first allele using a first set of search spaces determined based on the expected AGG peak size, determine the number of CGG repeats downstream of the last AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, determine the number of CGG repeats preceding the first AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, generate an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption, provide the AGG genotype for the first allele, A non-transitory computer-readable storage medium including instructions that cause an action including the above to be executed, and a system including the same. (Item 9) The system according to item 8, wherein the action further includes determining the number of CGG repeats separated by any two adjacent AGG interruptions on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, and the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG interruption, the number of CGG repeats separated by any two adjacent AGG interruptions, and the number of CGG repeats preceding the first AGG interruption. (Item 10) The action includes repeatedly searching the raw data and identifying one or more AGG peaks on a second allele using a second set of search spaces determined based on the predicted AGG peak size, determining the number of CGG repeats downstream of the last AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele, determining the number of CGG repeats preceding the first AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele, generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption, and further including providing the AGG genotype for the second allele, (i) the first allele is a normal allele and the second allele is a pre-mutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a pre-mutation allele and the second allele is a different pre-mutation allele, the system according to item 8 or 9. (Item 11) The action further includes determining a risk score for a subject associated with the sample based on the AGG genotype generated for the first allele, the second allele, or both the first allele and the second allele, the risk score identifying the risk that the subject will develop a late-onset neurodegenerative disease, fragile X-associated tremor / ataxia syndrome (FXTAS) or fragile X-associated primary ovarian insufficiency (FXPOI), or the risk of transmitting a full mutation allele to the subject's offspring or any combination thereof, the system according to item 8, 9 or 10. (Item 12) Repeatedly searching the raw data and identifying one or more AGG peaks on the first allele includes determining a first search space of a first set of search spaces for the first allele based on the predicted AGG peak size, ​ Search the raw data and identify an initial AGG peak on the first allele using a first search space; Determine a second search space for the first allele based on the peak size of the initial AGG peak on the first allele; Iteratively search the raw data and identify one or more additional AGG peaks on the first allele using the second search space, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the first allele based on the GS PCR data, the initial peak identified in the first allele, and the one or more additional peaks identified in the first allele, the system according to item 8, 9, or 10. (Item 13) Iteratively searching the raw data to identify one or more AGG peaks on the second allele, Determine a first search space of a second set of the search space for the second allele based on the predicted AGG peak size; Search the raw data and identify an initial AGG peak on the second allele using a second search space; Determine a second search space for the second allele based on the peak size of the initial AGG peak on the second allele; Iteratively search the raw data and identify one or more additional AGG peaks on the second allele using the second search space, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the second allele based on the GS PCR data, the initial peak identified in the second allele, and the one or more additional peaks identified in the second allele, the system according to item 10 or 12. (Item 14) When the one or more AGG peaks on the first allele are identified, the raw data is repeatedly searched, and the one or more AGG peaks on the first allele are removed from the raw data before identifying the one or more AGG peaks on the second allele. The system according to item 10 or 12. (Item 15) A computer program product tangibly embodied in a non-transitory machine-readable storage medium, the computer program product comprising instructions configured to cause one or more data processors to obtain raw data from a fragile X syndrome (FXS) assay performed on a sample using a gene analyzer, wherein the raw data includes gene-specific (GS) polymerase chain reaction (PCR) data and AGG interrupted PCR data separated by capillary electrophoresis; determine an expected AGG peak size for a first allele identified in the raw data; repeatedly search the raw data and identify one or more AGG peaks on the first allele using a first set of search spaces determined based on the expected AGG peak size; determine the number of CGG repeats downstream of the last AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; determine the number of CGG repeats preceding the first AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; generate an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption; provide the AGG genotype for the first allele; A computer program product comprising instructions configured to cause an action including the above to be performed. (Item 16) The action further includes determining the number of CGG repeats separated by any two adjacent AGG interruptions on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, wherein the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG interruption, the number of CGG repeats separated by any two adjacent AGG interruptions, and the number of CGG repeats preceding the first AGG interruption. The computer program product according to item 15. (Item 17) The action includes iteratively searching the raw data and identifying one or more AGG peaks on a second allele using a second set of search spaces determined based on the predicted AGG peak sizes, determining the number of CGG repeats downstream of the last AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele, determining the number of CGG repeats preceding the first AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele, generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption, and providing the AGG genotype for the second allele. The computer program product according to item 15 or 16 further includes: (i) the first allele is a normal allele and the second allele is a pre-mutation allele; (ii) the first allele is a normal allele and the second allele is a different normal allele; or (iii) the first allele is a pre-mutation allele and the second allele is a different pre-mutation allele. The computer program product according to item 15 or 16. (Item 18) The action further includes determining a risk score for a subject related to the sample based on the AGG genotype generated for the first allele, the second allele, or both the first allele and the second allele, the risk score identifying the risk that the subject will develop a late-onset neurodegenerative disease, Fragile X-associated tremor / ataxia syndrome ( FXTAS) or Fragile X-associated primary ovarian insufficiency (FXPOI), or the risk of transmitting a full mutation allele to the subject's offspring, or any combination thereof, the computer program product of item 15, 16, or 17. (Item 19) Repeatedly searching the raw data to identify one or more AGG peaks on the first allele, Determining a first search space of a first set of the search spaces for the first allele based on the predicted AGG peak size, Searching the raw data and using the first search space to identify an initial AGG peak on the first allele, Determining a second search space for the first allele based on the peak size of the initial AGG peak on the first allele, Repeatedly searching the raw data and using the second search space to identify one or more additional AGG peaks on the first allele, including The number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the first allele based on the GS PCR data, the initial peak identified on the first allele, and the one or more additional peaks identified on the first allele, the computer program product of item 15, 16, or 17. (Item 20) Repeatedly searching the raw data to identify one or more AGG peaks on the second allele, Determining a first search space of a second set of the search spaces for the second allele based on the predicted AGG peak size, Searching the raw data and using the second search space to identify an initial AGG peak on the second allele, Determining a second search space for the second allele based on the peak size of the initial AGG peak on the second allele; iteratively searching the raw data and using the second search space to identify one or more additional AGG peaks on the second allele, The computer program product according to item 17 or 19, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the second allele based on the GS PCR data, the initial peak identified in the second allele, and the one or more additional peaks identified in the second allele.

Claims

Claim 1 Obtaining raw data from a fragile X syndrome (FXS) assay performed on a sample using a gene analysis device, wherein the raw data includes gene-specific (GS) polymerase chain reaction (PCR) data and AGG interrupted PCR data separated by capillary electrophoresis, Determining an expected AGG peak size for a first allele identified in the raw data, Iteratively searching the raw data and identifying one or more AGG peaks on the first allele using a first set of search spaces determined based on the expected AGG peak size, wherein the search comprises: Determining a first search space of the first set of search spaces for the first allele based on the expected AGG peak size, Searching the raw data and identifying an initial AGG peak on the first allele using the first search space, Determining a second search space for the first allele based on the peak size of the initial AGG peak on the first allele, Iteratively searching the raw data and identifying additional AGG peaks on the first allele using the second search space, Including: Determining the number of CGG repeats downstream of the last AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, Determining the number of CGG repeats preceding the first AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, Generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption, Providing the AGG genotype for the first allele, A method comprising. Claim 2 Further comprising determining the number of CGG repeats separated by any two adjacent AGG interruptions on the first allele based on the GSP CR data and the one or more AGG peaks identified on the first allele, wherein the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG interruption, the number of CGG repeats separated by any two adjacent AGG interruptions, and the number of CGG repeats preceding the first AGG interruption, the method according to claim 1.

3. Repeatedly searching the raw data and identifying one or more AGG peaks on a second allele using a second set of search spaces determined based on the predicted AGG peak sizes; Determining the number of CGG repeats downstream of the last AGG interruption on the second allele based on the GSP CR data and the one or more AGG peaks identified on the second allele; Determining the number of CGG repeats preceding the first AGG interruption on the second allele based on the GSP CR data and the one or more AGG peaks identified on the second allele; Generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption; Further comprising providing the AGG genotype for the second allele, (i) the first allele is a normal allele and the second allele is a premutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a premutation allele and the second allele is a different premutation allele, the method according to claim 1 or 2.

4. Further comprising determining a risk score for a subject related to the sample based on the AGG genotype generated for the first allele, the second allele, or both the first allele and the second allele, wherein the risk score identifies the risk that the subject will develop a late-onset neurodegenerative disease, Fragile X-associated tremor / ataxia syndrome (FXTAS) or Fragile X-associated primary ovarian insufficiency (FXPOI), or the risk of transmitting a full mutation allele to the subject's offspring, or any combination thereof, the method according to claim 3.

5. Further comprising capturing the nucleic acid of the sample using the FXS assay to obtain the raw data, wherein the FXS assay targets the CGG repeat region in the 5'UTR of the target gene and utilizes primers containing the CGG repeat and an "A" nucleotide at the 3' end. The method according to claim 1, 2 or 3, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the first allele based on the GS PCR data, the initial AGG peak identified by the first allele, and the one or more additional AGG peaks identified by the first allele.

6. Iteratively searching the raw data to identify one or more AGG peaks on the second allele. Determining a first search space of a second set of the search space for the second allele based on the predicted AGG peak size. Searching the raw data and using the first search space of the second set of the search space to identify an initial AGG peak on the second allele. Determining a second search space for the second allele based on the peak size of the initial AGG peak on the second allele. Including iteratively searching the raw data and using the second search space to identify one or more additional AGG peaks on the second allele. The number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the second allele based on the GS PCR data, the initial AGG peak identified in the second allele, and the one or more additional AGG peaks identified in the second allele, the method according to claim 3 or 4.

7. Once the one or more AGG peaks on the first allele are identified, the raw data is searched iteratively, and the one or more AGG peaks on the first allele are removed from the raw data before identifying the one or more AGG peaks on the second allele, the method according to claim 3, 4 or 6.

8. One or more data processors, When executed on the one or more data processors, cause the one or more data processors to obtain raw data from a fragile X syndrome (FXS) assay performed on a sample using a genetic analyzer, the raw data including gene-specific (GS) polymerase chain reaction (PCR) data and AGG interruption PCR data separated by capillary electrophoresis, determine an expected AGG peak size for a first allele identified in the raw data, search the raw data iteratively and identify one or more AGG peaks on the first allele using a first set of search spaces determined based on the expected AGG peak size, the searching comprising determining a first search space of the first set of search spaces for the first allele based on the expected AGG peak size, searching the raw data and identifying an initial AGG peak on the first allele using the first search space, determining a second search space for the first allele based on the peak size of the initial AGG peak on the first allele, searching the raw data iteratively and identifying additional AGG peaks on the first allele using the second search space and comprising. Determining the number of CGG repeats downstream of the last AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; Determining the number of CGG repeats preceding the first AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; Generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption; Providing the AGG genotype for the first allele; A non-transitory computer-readable storage medium including instructions for causing an action including the above, and a system including the same. **Claim 9** The system according to claim 8, wherein the action further includes determining the number of CGG repeats separated by any two adjacent AGG interruptions on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, and the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG interruption, the number of CGG repeats separated by any two adjacent AGG interruptions, and the number of CGG repeats preceding the first AGG interruption. **Claim 10** The action includes iteratively searching the raw data and identifying one or more AGG peaks on a second allele using a second set of search spaces determined based on the predicted AGG peak sizes; Determining the number of CGG repeats downstream of the last AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; Determining the number of CGG repeats preceding the first AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption; further comprising providing the AGG genotype for the second allele; The system according to claim 8 or 9, wherein (i) the first allele is a normal allele and the second allele is a premutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a premutation allele and the second allele is a different premutation allele. **Claim 11** The action further comprises determining a risk score for a subject associated with the sample based on the AGG genotype generated for the first allele, the second allele, or both the first allele and the second allele, the risk score identifying the risk that the subject will develop a late-onset neurodegenerative disease, Fragile X-associated tremor / ataxia syndrome (FXTAS) or Fragile X-associated primary ovarian insufficiency (FXPOI), or the risk of transmitting a full mutation allele to the subject's offspring, or any combination thereof. The system according to claim 10. **Claim 12** The action further comprises capturing the nucleic acid of the sample using the FXS assay to obtain the raw data, the FXS assay targeting the CGG repeat region in the 5'UTR of the target gene and utilizing a primer containing the CGG repeat and an "A" nucleotide at the 3' end. The system according to claim 8, 9 or 10, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the first allele based on the GS PCR data, the initial AGG peak identified in the first allele, and the one or more additional AGG peaks identified in the first allele. **Claim 13** identifying one or more AGG peaks on the second allele by repeatedly searching the raw data; Based on the predicted AGG peak size, determining a first search space of a second set of the search space for the second allele; Searching the raw data and identifying an initial AGG peak on the second allele using the first search space of the second set of the search space; Based on the peak size of the initial AGG peak on the second allele, determining a second search space for the second allele; Iteratively searching the raw data and identifying one or more additional AGG peaks on the second allele using the second search space, including: The system according to claim 10 or 11, wherein the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the second allele based on the GS PCR data, the initial AGG peak identified in the second allele, and the one or more additional AGG peaks identified in the second allele. **Claim 14** When one or more AGG peaks on the first allele are identified, iteratively searching the raw data and removing the one or more AGG peaks on the first allele from the raw data before identifying the one or more AGG peaks on the second allele, the system according to claim 10, 11 or 13. **Claim 15** A computer program product tangibly embodied in a non-transitory machine-readable storage medium, comprising: one or more data processors, Obtaining raw data from a fragile X syndrome (FXS) assay performed on a sample using a gene analyzer, wherein the raw data includes gene-specific (GS) polymerase chain reaction (PCR) data and AGG interruption PCR data separated by capillary electrophoresis; Determining a predicted AGG peak size for a first allele identified in the raw data; Iteratively searching the raw data and identifying one or more AGG peaks on the first allele using a first set of a search space determined based on the predicted AGG peak size, wherein the search Determining a first search space of a first set of the search space for the first allele based on the predicted AGG peak size; Searching the raw data and identifying an initial AGG peak on the first allele using the first search space; Determining a second search space for the first allele based on the peak size of the initial AGG peak on the first allele; Iteratively searching the raw data and identifying additional AGG peaks on the first allele using the second search space; Including; Determining the number of CGG repeats downstream of the last AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; Determining the number of CGG repeats preceding the first AGG interruption on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele; Generating an AGG genotype for the first allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption; Providing the AGG genotype for the first allele; A computer program product comprising instructions configured to cause an action including.

16. The action further includes determining the number of CGG repeats separated by any two adjacent AGG interruptions on the first allele based on the GS PCR data and the one or more AGG peaks identified on the first allele, and the AGG genotype of the first allele is generated based on the number of CGG repeats downstream of the last AGG interruption, the number of CGG repeats separated by any two adjacent AGG interruptions, and the number of CGG repeats preceding the first AGG interruption. The computer program product according to claim 15.

17. The action includes iteratively searching the raw data and identifying one or more AGG peaks on a second allele using a second set of the search space determined based on the predicted AGG peak size; Determining the number of CGG repeats downstream of the last AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; Determining the number of CGG repeats preceding the first AGG interruption on the second allele based on the GS PCR data and the one or more AGG peaks identified on the second allele; Generating an AGG genotype for the second allele based on the number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption; Further comprising providing the AGG genotype for the second allele, The computer program product according to claim 15 or 16, wherein (i) the first allele is a normal allele and the second allele is a premutation allele, (ii) the first allele is a normal allele and the second allele is a different normal allele, or (iii) the first allele is a premutation allele and the second allele is a different premutation allele.

18. The action further comprises determining a risk score for a subject associated with the sample based on the AGG genotype generated for the first allele, the second allele, or both the first allele and the second allele, the risk score identifying the risk that the subject will develop a late-onset neurodegenerative disease, fragile X-associated tremor / ataxia syndrome (FXTAS) or fragile X-associated primary ovarian insufficiency (FXPOI), or the risk of transmitting a full mutation allele to the subject's offspring, or any combination thereof, the computer program product according to claim 17.

19. The action further comprises capturing nucleic acid of the sample using the FXS assay to obtain the raw data, the FXS assay targeting a CGG repeat region in the 5'UTR of a target gene and utilizing a primer that includes a CGG repeat and an "A" nucleotide at the 3' end. The number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the first allele based on the GS PCR data, the initial AGG peak identified in the first allele, and the one or more additional AGG peaks identified in the first allele, the computer program product according to claim 15, 16 or 17.

20. Iteratively searching the raw data to identify one or more AGG peaks on the second allele, Determining a first search space of a second set of the search space for the second allele based on the predicted AGG peak size, Searching the raw data to identify an initial AGG peak on the second allele using the first search space of the second set of the search space, Determining a second search space for the second allele based on the peak size of the initial AGG peak on the second allele, Iteratively searching the raw data to identify one or more additional AGG peaks on the second allele using the second search space, including, The number of CGG repeats downstream of the last AGG interruption and the number of CGG repeats preceding the first AGG interruption are determined for the second allele based on the GS PCR data, the initial AGG peak identified in the second allele, and the one or more additional AGG peaks identified in the second allele, the computer program product according to claim 17 or 18.

Citation Information

Patent Citations

  • Automated nucleic acid repeat count calling methods

    US20150134267A1