Genotyping for tandem repeats

The method and system address the challenge of accurately detecting VNTRs by employing wrap-around alignment and Bayesian modeling to classify reads and estimate fragment distributions, enhancing the precision of VNTR length determination and genetic variant analysis.

WO2025250322A1PCT designated stage Publication Date: 2025-12-04ILLUMINA INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/027914
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-26
Filing Date
2025-05-06
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing nucleic acid sequencing technologies face challenges in accurately detecting and characterizing tandem repeats, particularly large variable number tandem repeats (VNTRs), due to their repetitive nature and low complexity, which complicates alignment and variant detection, leading to low precision and inefficiencies in identifying genetic variations associated with diseases.

Method used

A method and system for determining tandem repeat lengths using paired-end sequencing reads, employing wrap-around alignment and Bayesian likelihood modeling to classify reads into various classes, estimate fragment distributions, and calculate posterior probabilities for haplotype lengths, enabling accurate diploid genotyping of VNTRs.

Benefits of technology

Enables precise and efficient detection of VNTR lengths, improving the accuracy of genetic variant calling and disease association analysis by leveraging fragment information and probabilistic genotyping, even for haplotypes longer than read lengths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025027914_04122025_PF_FP_ABST
    Figure US2025027914_04122025_PF_FP_ABST
Patent Text Reader

Abstract

In one aspect, the disclosed technology relates to systems and methods for determining a length of a tandem repeat region in each haplotype of a diploid genome. In some embodiments, the method may include obtaining paired-end sequencing reads of the diploid genome; aligning the paired-end sequencing reads to a tandem repeat region in a reference genome sequence; classifying each of the paired-end sequencing reads that overlaps the tandem repeat region based on the alignments into a plurality of classes and counting the number of paired-end sequencing reads in each class; providing a set of hypotheses of a length of the tandem repeat region in each haplotype of the diploid genome; and evaluating which hypothesis has the highest likelihood of generating the counted or observed number of paired-end sequencing reads in the plurality of classes to determine the length of the tandem repeat region in each haplotype of the diploid genome.
Need to check novelty before this filing date? Find Prior Art

Description

GENOTYPING FOR TANDEM REPEATSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 654,774, filed May 31, 2024, and U.S. Provisional Application No. 63 / 664,556, filed June 26, 2024, the content of each of which is incorporated by reference in its entirety.SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled ILLINC.835. xml, created and last saved on April 14, 2025, which is 16 kilobytes in size. The information in the electronic format of the Sequence Listing is hereby incorporated by reference in its entirety.BACKGROUNDField

[0003] The disclosed technology relates to the field of nucleic acid sequencing. More particularly, the disclosed technology relates to detecting and identifying tandem repeats (TRs) — for example variable number tandem repeats (VNTRs) — in a sample nucleic acid and determining the lengths of the TRs.Description of the Related Art

[0004] Accurate detection of TRs has long been complicated by the low- complexity nature of TR regions. In some cases, the large size of the repetitive sequences in TRs further complicates the detection process. There exists a continuing need for improving the detection and characterization of TRs in nucleic acid sequencing technologies, especially since VNTRs account for a significant proportion of between-genome variations in humans.SUMMARY

[0005] In one aspect, the disclosed technology relates to method of determining a length of a tandem repeat region in each haplotype of a diploid genome in a sample, the methodcomprising obtaining paired-end sequencing reads of the diploid genome; aligning the paired- end sequencing reads to a tandem repeat region in a reference genome sequence; classifying paired-end sequencing reads that overlap the tandem repeat region based on the alignments into a plurality of classes and counting the number of paired-end sequencing reads in each class; providing a set of hypotheses of a length of the tandem repeat region in each haplotype of the diploid genome; and evaluating which hypothesis from the set of hypotheses has the highest probability of generating the counted or observed number of paired-end sequencing reads in the plurality of classes to determine the length of the tandem repeat region in each haplotype of the diploid genome.

[0006] In some embodiments, the probability of a hypothesis generating the counted or observed number of paired-end sequencing reads in the plurality of classes comprises a genotype prior and a likelihood of observing the counted or observed number of paired-end sequencing reads in the plurality of classes given the hypothesis.

[0007] In some embodiments, the likelihood of observing the counted or observed number of paired-end sequencing reads in the plurality of classes given the hypothesis depends on the distribution of the size of nucleic acid fragments associated with the paired-end sequencing reads that overlap the tandem repeat region.

[0008] In some embodiments, the disclosed method further comprises computing the distribution of the size of nucleic acid fragments associated with the paired-end sequencing reads that overlap the tandem repeat region based on estimating the distribution of the size of nucleic acid fragments overlapping fixed non-repetitive regions in the diploid genome in the sample.

[0009] In some embodiments, the genotype prior is determined based on population frequencies of known lengths of the tandem repeat region.

[0010] In some embodiments, the genotype prior comprises unequal priors for each haplotype.

[0011] In some embodiments, wherein a class in the plurality of classes consists of paired-end sequencing reads wherein at least one read spans the tandem repeat region in its entirety and overlaps with both flanks.

[0012] In some embodiments, a class in the plurality of classes consists of paired- end sequencing reads wherein a first read overlaps with a flank of the tandem repeat region and a second read overlaps with the other flank of the tandem repeat region.

[0013] In some embodiments, wherein the first read and / or the second read also partially overlap with the tandem repeat region.

[0014] In some embodiments, a class in the plurality of classes consists of paired- end sequencing reads wherein a first read overlaps with both a flank of the tandem repeat region and the tandem repeat region, and a second read overlaps with the same flank of the tandem repeat region.

[0015] In some embodiments, the second read also overlaps with the tandem repeat region.

[0016] In some embodiments, a class in the plurality of classes consists of paired- end sequencing reads wherein a first read overlaps with a flank of the tandem repeat region and a second read is contained within the tandem repeat region.

[0017] In some embodiments, the first read also overlaps with the tandem repeat region.

[0018] In some embodiments, a class in the plurality of classes consists of paired- end sequencing reads wherein both reads are contained within the tandem repeat region.

[0019] In some embodiments, obtaining paired-end sequencing reads of the diploid genome comprises performing paired-end sequencing of the sample.

[0020] In some embodiments, aligning the paired-end sequencing reads uses wraparound alignment.

[0021] In some embodiments, the sample is extracted from cells, a cell-free DNA sample, an amniotic fluid, a blood sample, a biopsy sample, or any combination thereof, of a subject.

[0022] In some embodiments, the tandem repeat region is a variable number tandem repeat (VNTR) locus.

[0023] In some embodiments, each nucleic acid fragment is about 250 base pairs to about 1000 base pairs in length.

[0024] In some embodiments, the paired-end sequencing reads are generated by whole genome sequencing (WGS) of the sample.

[0025] In some embodiments, the paired-end sequencing reads are generated by a next generation sequencing reaction.

[0026] In another aspect, the disclosed technology relates to a system for determining a length of a tandem repeat region in each haplotype of a diploid genome from a sample, the system comprising: a nucleic acid sequencer; non-transitory memory configured to store executable instructions; and a hardware processor in communication with the nucleic acid sequencer and the non-transitory memory, the hardware processor programmed by the executable instructions to perform the method of any of claims 1 to 21.

[0027] In some embodiments, the hardware processor is configured to receive paired-end sequencing reads from the nucleic acid sequencer.

[0028] In some embodiments, the hardware processor is configured to control the nucleic acid sequencer to perform sequencing.

[0029] In some embodiments, the hardware processor is configured to output, on a display, the length of the tandem repeat region in each haplotype of the diploid genome.BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear.

[0031] FIG. 1A shows a non-limiting exemplary illustration of a VNTR in a genomic reference sequence. The nucleotide sequence of the repeat unit (SEQ ID NO: 1) is shown with bolded nucleotides being the differing nucleotides in different repeat units.

[0032] FIG. IB is a non-limiting exemplary illustration showing an alignment between sequences of different repeat units.

[0033] FIG. 1C shows a non-limiting exemplary illustration of a VNTR in a reference sequence, with a repeat unit layout between different copy number variants of human subjects from different regions including Africa (AFR), Europe (EUR) and Eurasian Countries (EAS).

[0034] FIG. 2 is a flow chart that illustrates a method of genotyping in the presence of VNTRs according to some embodiments of the disclosed technology.

[0035] FIG. 3 schematically illustrates the classification of reads overlapping a VNTR region according to some embodiments of the disclosed technology.

[0036] FIG. 4 schematically illustrates the classification of fragments overlapping a VNTR region according to some embodiments of the disclosed technology.

[0037] FIG. 5 is a flow chart that illustrates a method of classifying paired-end sequencing reads and evaluating various hypotheses of VNTR length according to some embodiments of the disclosed technology.

[0038] FIG. 6A is a block diagram illustrating that a VNTR detection system may be employed within various different types of systems, and connected through a network to perform the disclosed methods.

[0039] FIG. 6B is a block diagram of an exemplary computing device that may be used in connection with the exemplary sequencing system of FIG. 6A.DETAILED DESCRIPTION

[0040] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entireties for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.Overview

[0041] Tandem repeats (TRs) are regions in the genome which have repetitions of a "pattern sequence" (or a "repeat unit"). Short tandem repeats (STRs) are TRs that have short repeat units (e.g., 1-20 bp) repeated many times (e.g., hundreds or thousands). In STRs, variation within the repeat unit tends to be low, and the number of repeats may be adequately described as an integer. Variable number tandem repeats (VNTRs) are TRs that may differ inthe number ("copy number") of repeat units across the population. In VNTRs, repeat units can be long (e.g., more than 10 bp, and can be tens or hundreds of base pairs or more), different instances of the repeat unit can be substantially different from each other, and the number of repeats need not be an integer.

[0042] The changes in the copy numbers of TRs in the human genome have been linked to gene silencing, differences in gene expression, genetic variations and various diseases. Large TRs have been linked to various phenotypes and diseases, such as susceptibility to human type 1 diabetes, chronic obstructive pulmonary disease (COPD), epilepsy, Parkinson disease, etc. Examples of the large TRs of interest may include mini-satellites and macrosatellites, which are sometimes defined as TRs having repeat units of size larger than about 10 bp and about 100 bp, respectively. In some cases, mini-satellites may be defined as TRs having repeat units of size larger than about 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 15 bp, 20 bp, etc. In some cases, macro-satellites may be defined as TRs having repeat units of size larger than about 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 150 bp, 200 bp, etc. However, very few large TRs (e.g., those having an array size larger than 1000 bp) have been studied in large cohorts due to the lack of reliable and efficient tools in the literature.

[0043] While methods exist for detecting, identifying or estimating the copy numbers of smaller TRs (e.g., those having an array size of less than a few hundred base pairs) in the human genome in conjunction with sequencing by synthesis (SBS) technologies, reliable and efficient methods for detecting, identifying or estimating larger TRs (e.g., those that are hundreds of base pairs or longer) have been challenging for short-read SBS technologies. The lack of a reliable and efficient method for detecting, identifying or estimating larger TRs in short-read SBS technologies is mainly due to the difficulty of performing secondary analysis (i.e., alignment and / or assembly of nucleic acid sequencing reads, and / or determination of genetic variants) on tandem repeats. On the one hand, short SBS reads cannot span large TRs; on the other hand, the repetitive nature of the TRs may result in unreliable alignments of the SBS reads in the TR regions.

[0044] As used herein, a “fragment” refers to a polynucleotide molecule of which the two ends are sequenced. As used herein, a “read” refers to the sequencing data gathered from one end of the polynucleotide molecule. A “read pair” refers to the combination of the two reads which are initiated (primed) at both the 3’ and 5’ ends of the molecule. For example,a fragment may be about 500 bp long, and each read starting on the 3’ end and the 5’ end of the fragment may be about 150 bp long. Aspects of the disclosed technology enable diploid calling of tandem repeat regions, such as STRs and VNTRs, using short-read sequencing data. The sequence caller detailed herein was found to accurately call diploid genotypes containing any combination of short (i.e., less than read length), medium (i.e., between read length and fragment length) and long (i.e., above fragment length) tandem repeated regions of various haplotypes. The disclosed method may perform probabilistic (likelihood-based) genotyping in a Bayesian framework, making use of information from fragments, rather than the individual reads. The use of fragment information enables the calling of haplotypes longer than the read length. Embodiments of the disclosed method can estimate the size of the haplotypes in each region and produce variant calls, including the number of copies of a repeat pattern within a nucleotide fragment or genomic sequence for the sample in question. In some embodiments, the disclosed method further provides a score of the quality of the call for candidate haplotype lengths.

[0045] Briefly, there are three ways a read can overlap a repeat region: spanning, flanking, and contained. This gives rise to five fragment classes: “fixed-sized spanning” (also called “fixed spanning” or “fixed read spanning”), “variable-sized spanning” (also called “spanning” or “variable spanning”), “fixed-sized flanking” (also called “fixed flanking”), “variable-sized flanking” (also called “variable flanking”), and “contained”, each of which has its own expected relative sequencing coverage as a function of the haplotype lengths. Aspects of the invention include systems and methods which construct a likelihood model wherein the expected coverage for each of the five classes is derived as a function of the two haplotypes and of the fragment length. Given that the fragment length is unknown, the expected coverage for each class then becomes a function of the fragment length distribution, which may be sample-specific. Therefore, by observing the relative sequencing coverages of the five fragment classes, the disclosed systems and methods can determine the most likely diploid genotype in each repeated region of interest. The simultaneous use of classes representing all of the possible fragment types enables the calling of the full range of haplotype sizes in the same calling process, which is useful for calling mixed genotypes (i.e., genotypes containing one short and one long haplotype).

[0046] The disclosed systems and methods may calculate the distribution of fragment lengths overlapping a haplotype of a given size as a function of the distribution of fragment lengths overlapping a haplotype of a different size. In some embodiments, the resulting distribution is a function of the haplotype size. An accurate estimate of the fragment length distribution is important for accurately estimating haplotype lengths that are longer than the read length. The disclosed method may make use of accurate per-sample estimation of the statistical distribution of fragment lengths.

[0047] When evaluating the genotype prior, the disclosed method may also use unequal priors for different alleles. For example, when a diploid genotype contains one short and one long haplotype, the probability of sampling a fragment overlapping the long haplotype is larger than that of sampling a fragment overlapping the short haplotype, and this needs to be taken into account during diploid calling.Variable Number Tandem Repeats

[0048] Variable number tandem repeats (VNTRs) are a class of structural variants that include tandem repeats of patterns, for example patterns larger than 10 base pairs (bps), and that differ in copy number among the genomes of individuals of a species. While VNTRs cover <5% of the human genome, about 50% of all structural variants (variants greater than 50 bp) are VNTRs. In some cases, a VNTR can have fewer than 20% mismatches for an exact repeat. In some cases, VNTRs can have small variants, such as SNPs and indels in the repetitive sequences. On average one person has about 2.2 mega base pairs (Mbps) of deleted sequence and about 5.7 Mbps of inserted sequence in VNTRs. Variations in VNTRs can depend on the populations within a species.

[0049] Some VNTRs are known to be associated with genetic diseases, such as bipolar disorder, MCKD1, stroke, CAD, FSHD, ADHD, Parkinson’s, diffuse panbronchiolitis (DPB), monogenic diabetes, T1D, T2D, obesity, OCD, osteochondritis dissecans, Kawasaki, ATF in stroke, BPSD, Alzheimer’s, anxiety, schizophrenia, metastatic colorectal cancer, Kawasaki, or progressive myoclonic epilepsy 1A. A VNTR can be present in the coding region or non-coding region. Moreover, a VNTR can be present in the 5’ untranslated region (UTR), promoter, intron, or 3’ UTR. The gene that includes, or is affected by, the VNTR canbe, for example, PER3, MUC1, IL1RN, DUX4, DAT1, MUC21, CEL, INS, DRD4, ACAN, ZFHX3, GP1BA, SERT, SERT, HIC1, MMP9, CSTB, or MAO A.

[0050] FIG. 1A, FIG. IB and FIG. 1C show a non-limiting exemplary illustration of a VNTR in a reference sequence. FIG. 1A shows that a VNTR in the reference human genome GRCh38 is at chrl:3428147-3428340. The repeat unit has a length of 48 bps. The reference sequence of the repeat unit isACCCCGAGCTAGGGTGCAGCCCGGCCGCACTGCAGGAGACCCACCAGG (SEQ ID NO: 1) in GRCh38. Different copies of the repeat unit in the VNTR (within a haplotype or across haplotypes) can vary, in particular at the three bases bolded and underlined. FIG. IB shows that the three bases can be G, G, and A, respectively, in a first type or sequence of the repeat unit; G, G, and G, respectively, in a second type or sequence of the repeat unit; A, G, and A, respectively, in a third type or sequence of the repeat unit; and G, A, and G, respectively, in a fourth type or sequence of the repeat unit. FIG. 1C shows that the VNTR includes four copies of the repeat unit in GRCh38. The four copies include two copies of the first type followed by two copies of the second type. The five samples shown in FIG. 1C included three, five, seven, seven, and ten copies of the repeat unit, respectively. For sample NA19240 of a subject who is African, the VNTR included one copy of the first type followed by two copies of the second type. For sample NA12878 of a subject who is European, the VNTR included one copy of the first type, three copies of the second type, and one copy of the first type. For sample NA24385 of a subject who is European, the VNTR included one copy of the first type, one copy of the second type, two copies of the first type, two copies of the second type, and one copy of the third type. For sample HG00597 of a subject who is Eastern Asian, the VNTR included three copies of the second type, one copy of the first type, and three copies of the second type. For sample HGO3453 of a subject who is African, the VNTR included one copy of the first type, two copies of the second type, one copy of the fourth type, one copy of the first type, one copy of the second type, one copy of the fourth type, and three copies of the second type. The examples discussed in connection with FIG. 1 A, FIG. IB and FIG. 1C pertain to homozygous variants where both alleles include the VNTR locus.

[0051] The difficulty of detecting VNTRs is multi-dimensional. The nature of the tandem repeats causes low mappability and high sequencing errors. Existing sequencing techniques (including, for example, using population haplotypes in the genome graph) sufferfrom low precision in detecting VNTRs due to the repetitive nature of VNTRs. Short-read sequencing technologies have a higher throughput compared to long-read sequencing technologies, but short sequencing reads often cannot cover the full length of most VNTRs. For example, around 29% of the VNTRs have additional repeats with total length greater than or equal to 150 bps in one individual. Due to the repetitive nature of VNTRs, correctly rebuilding VNTRs’ haplotypes from short reads is difficult. With short sequencing reads, methods of detecting VNTRs may utilize the read sequences and some form of circular alignment (or wrap-around alignment) to infer the copy number changes in tandem repeats; however, these methods only allow for identification of small VNTRs (i.e., smaller than the read length). The abnormal fragment size of a read pair that maps beyond the normal distribution have been used in the prior art to infer some classes of large structural variants such as large changes in VNTRs; however, some VNTRs may not be accurately detected by this approach if the VNTRs are shorter compared to the variance in the insert size of the sequencing reads. For example, VNTRs may not be accurately detected with paired-end sequencing reads, which have a high variance in insert size. Moreover, using existing methods, local reassembly of the VNTR sequences is difficult and often fails. Therefore, there is a need for improved methods for detecting and identifying VNTRs.Embodiments of Processes for Detecting Lengths of Tandem Repeats

[0052] In some aspects, the disclosed technology relates to methods of detecting or estimating the lengths of tandem repeats (TRs) in a sample using SBS technologies, such as short-read SBS technologies. One example is the short- read SBS sequencing technologies from Illumina, Inc. (San Diego, CA).

[0053] FIG. 2 is a flow chart that illustrates a method according to some embodiments of the disclosed technology. The method may take as input a set of aligned / mapped reads from the sample in question and a VNTR catalog file. In some embodiments, the VNTR catalog is a file specifying the TR regions which may be found in the type of sample being analyzed. Each region in the file may include the start and the end of a tandem repeat sequence, and the sequence of the repeat unit / pattern.

[0054] The method may start with read pair collection at block 201. In some embodiments, read pairs, rather than the individual reads, are collected. For example, to obtainall of the relevant read pairs for each TR region, all of the reads that overlap the TR region are found, and then all of their mates are collected as well. Due to the repetitive nature of TR regions, existing read-alignments may be unreliable. Therefore, in some embodiments, spanning reads, unmapped reads, and reads with soft-clips may also undergo a specialized wrap-around alignment process. The wrap-around alignment allows for a read to align to the same pattern sequence multiple times without penalty, mirroring the structure of the tandem repeat. The wrap-around alignment may produce more reliable alignments of the read to the TR region. Details of the wrap-around alignment process may be found in “Benson, G., 2005. Tandem cyclic alignment. Discrete applied mathematics, 146(2), pp.124-133”, which is incorporated herein by reference.

[0055] By using the information of fragments (i. e. , both reads from a read pair, plus information on how far they are typically apart), the disclosed method can call haplotypes longer than the read length. As fixed flanking fragments serve as a proxy for local coverage, there is no need for local coverage estimation. The disclosed method can make diploid calls when the fragments span only one of the haplotypes, since contained and / or variable flanking fragments give information about the spanned haplotype. Using all information from different fragment types at the same time, the method is able to call heterozygous genotypes (e.g., one healthy and one expanded haplotype).

[0056] Once reliable alignments of the read pairs have been obtained, the method may move to block 202 to classify each fragment associated with a read pair. FIG. 3 schematically illustrates some examples of the classification of individual reads according to some embodiments of the disclosed technology. Relative to a reference genomic sequence that includes a TR region (302) and its flanks (301 and 303), reads may be classified as nonoverlapping (304), flanking (305), spanning (306), and contained (307).

[0057] FIG. 4 schematically illustrates some examples of the classification of fragments associated with the read pairs according to some embodiments of the disclosed technology. Relative to a reference genomic sequence that includes a TR region (402) and its flanks (401 and 403), fragments overlapping a TR can be classified into the five classes mentioned above. For example, the five classes may include: “fixed spanning”, “variable spanning”, “fixed flanking”, “variable flanking”, and “contained”, as shown in FIG. 4. The “fixed spanning” class consists of read pairs wherein one read is spanning the TR region andthe other read is on a flank or spanning the TR region. The “variable spanning” or “spanning” class consists of read pairs wherein one read is on a flank (i.e., overlaps with a flank) and the other read is on the other flank, both of which may also overlap with the TR region. The “fixed flanking” class consists of read pairs wherein one read is on the flank and overlapping with the TR region, and the other read is on the same flank (and may overlap with the TR region). The “variable flanking” class consists of read pairs wherein one read is on a flank (and may overlap with the TR region) and the other read is contained within the TR region. The “contained” class consists of read pairs wherein both of the reads are contained within the TR region. Here, “fixed” versus “variable” refers to whether the fragment length is known (while the reads mapping is known, the TR length is unknown). The output of block 202 may be the set of all read pairs in each TR region, re-aligned as necessary, with each read pair given a classification. This collection of read pairs is referred to as a “pileup”.

[0058] Returning to the flowchart of Fig. 2, in parallel to blocks 201 and 202, the method may execute block 203 to compute the fragment size / length distribution from the sample. The distribution of fragment lengths across the sample is not the same as the distribution of fragment lengths overlapping a particular point or region (because larger fragments are more likely to overlap any given point), and it depends on the length of the region. The disclosed method estimates the sample-wide fragment length distribution for fragments overlapping “fixed non-repetitive regions”, and then transforms it into the distribution of fragment lengths for fragments overlapping a TR as a function of the haplotype lengths under consideration. The “fixed non-repetitive regions” are predetermined regions in the genome that are evolutionarily conserved, and the distribution of their lengths can be directly estimated.

[0059] In some embodiments, the number, or proportion, of fragments in each class acts as evidence for the haplotype lengths of the TR region. A Bayesian likelihood model may be used to evaluate what pair of haplotypes has the highest likelihood of generating the observation of these fragment class counts.

[0060] Once the collection of read pairs (pileup) is obtained, the method may move to block 204 to compute fragment constraints. Based on existing knowledge about the fragments, certain constraints can be applied to the repeat region. For example, when there are contained sequencing reads, there must also be flanking sequencing reads with respect to therepeat region. Another example is if the number of contained reads is below a minimum threshold, then the number of spanning reads must be above a minimum threshold for the method to be applied and vice versa. A further example is if there are fewer than three spanning reads supporting a specific candidate haplotype length, then those fragments are excluded from consideration.

[0061] Once the fragment constraints have been computed at block 204, the method may move to block 205 to generate genotype hypotheses. For example, a set of diploid genotype hypotheses may be generated. Each diploid genotype is a haplotype pair, where a haplotype is characterized by its length h, in base pairs. The haplotype lengths to consider are generated by parameterizing them as h = p(n+f), where p is the pattern length, n is an integer denoting the number of repeats, and f is some fraction. Appropriate ranges of values for n and f then generate a range of haplotype lengths which in turn generate diploid genotype hypotheses. For example, p may be about 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 150 bp, 200 bp, 300 bp, 500 bp, 1000 bp or longer, n may be about 2, 5, 10, 20, 50, 100, 500 or more, and f may be between about 0 and %, between about % and ’A, between about A and %, or between about % and 1. In some embodiments, the set of hypotheses may also include the size of the TR in the reference as well as sizes estimated by spanning reads using the wrap-around alignment in 201.

[0062] Once the genotype hypotheses have been generated, the method may move to block 206 to compute the pileup likelihoods for each genotype hypothesis, P(R|G). The pileup likelihood may be calculated as the likelihood of observing the fragment class counts given the candidate diploid genotype, based on an underlying model for how fragments are generated from a TR region haplotype of a given length. The likelihood of a pileup is the accumulation of all likelihoods of the individual fragments. In turn, the likelihood of a fragment depends on its class. The likelihood of a fragment is proportional to the expected density of the fragment of that class given the two haplotypes. In some embodiments, the likelihood of each fragment is computed from class-specific coverage distributions and the sample-specific fragment size distribution that is obtained at block 203.

[0063] Once the pileup likelihoods for the hypotheses have been computed, the method may move to block 207 to construct the diploid copy number call. The method calculates the posterior probability of each candidate diploid genotype hypothesis based on thegenotype prior and the pileup likelihood. The posterior probability is made up of the genotype prior and the pileup likelihood. Specifically, the posterior probability for genotype G given the observed read-pair pileup R is P(Cr|7?) oc P(R|G)P(G), where P(R|G) is the pileup likelihood and P(G) is the genotype prior. In some embodiments, the disclosed method assigns the genotype prior P(G) based on population frequencies of known per-locus haplotypes. In some embodiments, the disclosed method uses unequal allele priors, which are a function of haplotype lengths. For example, when the two haplotypes are of unequal length, the probability of sampling the longer one is higher than the probability of sampling the shorter one.

[0064] In some embodiments, the diploid genotype hypothesis with the highest posterior probability is chosen as the resulting call for the TR region. In some embodiments, the output of block 207 is the diploid copy number call of the TR. In cases where insufficient information is available for reliable diploid calls, the disclosed method can fall back on reporting "total" calls instead of diploid calls. In other words, instead of reporting two (less reliable) haplotype lengths, the method reports the sum of the two haplotype lengths (which is more reliable in such cases).

[0065] FIG. 5 is a flow chart that illustrates a method 500 according to some embodiments of the disclosed technology. The disclosed method 500 can determine a length of a tandem repeat region in each haplotype of a diploid genome in a sample. The method may start from block 501 to obtain paired-end sequencing reads of the diploid genome. In some embodiments, obtaining paired-end sequencing reads of the diploid genome comprises performing paired-end sequencing of the sample. In some embodiments, the sample is extracted from cells, a cell-free DNA sample, an amniotic fluid, a blood sample, a biopsy sample, or any combination thereof, of a subject. In some embodiments, each fragment associated with the paired-end sequencing reads is about 250 base pairs to about 1000 base pairs in length. In some embodiments, the paired-end sequencing reads are generated by whole genome sequencing (WGS) of the sample. In some embodiments, the paired-end sequencing reads are generated by a next generation sequencing reaction.

[0066] The method 500 may then move to block 503 to align the paired-end sequencing reads to a tandem repeat region in a reference genome sequence. In some embodiments, the tandem repeat region is a variable number tandem repeat (VNTR) locus. In some embodiments, aligning the paired-end sequencing reads uses wrap-around alignment. Insome embodiments, the disclosed method further comprises computing the distribution of the size of fragments associated with the paired-end sequencing reads that overlap the tandem repeat region based on estimating the distribution of the size of fragments overlapping fixed non-repetitive regions in the diploid genome in the sample.

[0067] The method 500 may then move to block 505 to classify paired-end sequencing reads that overlap the tandem repeat region based on the alignments into a plurality of classes and count the number of paired-end sequencing reads in each class. In some embodiments, a class in the plurality of classes includes paired-end sequencing reads wherein at least one read spans the tandem repeat region in its entirety and overlaps with both flanks. In some embodiments, a class in the plurality of classes includes paired-end sequencing reads wherein a first read overlaps with a flank of the tandem repeat region and a second read overlaps with the other flank of the tandem repeat region. In some embodiments, wherein the first read and / or the second read also partially overlap with the tandem repeat region. In some embodiments, a class in the plurality of classes consists of paired-end sequencing reads wherein a first read overlaps with both a flank of the tandem repeat region and the tandem repeat region, and a second read overlaps with the same flank of the tandem repeat region. In some embodiments, the second read also overlaps with the tandem repeat region. In some embodiments, a class in the plurality of classes consists of paired-end sequencing reads wherein a first read overlaps with a flank of the tandem repeat region and a second read is contained within the tandem repeat region. In some embodiments, the first read also overlaps with the tandem repeat region. In some embodiments, a class in the plurality of classes consists of paired-end sequencing reads wherein both reads are contained within the tandem repeat region.

[0068] The method 500 may then move to block 507 to provide a set of hypotheses of a length of the tandem repeat region in each haplotype of the diploid genome. For example, a set of diploid genotype hypotheses may be generated. Each diploid genotype is a haplotype pair, where a haplotype is characterized by its length h, in base pairs. The haplotype lengths to consider may be generated by parameterizing them as h = p(n+f), where p is the pattern length, n is an integer denoting the number of repeats, and f is some fraction. Appropriate ranges of values for n and f then generate a range of haplotype lengths which in turn generate diploid genotype hypotheses.

[0069] The method 500 may then move to block 509 to evaluate which hypothesis from the set of hypotheses has the highest probability of generating the counted or observed number of paired-end sequencing reads in the plurality of classes to determine the length of the tandem repeat region in each haplotype of the diploid genome. In some embodiments, the probability of a hypothesis generating the counted or observed number of paired-end sequencing reads in the plurality of classes comprises a genotype prior and a likelihood of observing the counted or observed number of paired-end sequencing reads in the plurality of classes given the hypothesis. In some embodiments, the likelihood of observing the counted or observed number of paired-end sequencing reads in the plurality of classes given the hypothesis depends on the distribution of the size of fragments associated with the paired-end sequencing reads that overlap the tandem repeat region. In some embodiments, the genotype prior is determined based on population frequencies of known lengths of the tandem repeat region. In some embodiments, the genotype prior comprises unequal priors for each haplotype.Embodiments of Sequencing Systems

[0070] FIG. 6A is a block diagram of an exemplary sequencing system 6000 that may be used to perform or implement the disclosed technology, such as the example method described in connection with FIG. 2, or the example method described in connection with FIG. 5. For example, the sequencing system 6000 can be configured to determine the lengths of TR regions in a sample nucleic acid. The illustrative sequencing system 6000 may include a nucleic acid sequencer 6001, a non-transitory memory 6003 configured to store executable instructions, and a hardware processor 6005 in communication with the nucleic acid sequencer 6001 and the non-transitory memory 6003. The hardware processor 6005 may be programmed by the executable instructions to perform the methods disclosed herein.

[0071] In some embodiments, the non-transitory memory 6003 is configured to store the reference sequence. In some embodiments, the hardware processor 6005 is configured to obtain the reference sequence from an external database. In some embodiments, the hardware processor 6005 is configured to receive paired-end sequence reads from the nucleic acid sequencer 6001. In some embodiments, the hardware processor 6005 is configured to control the nucleic acid sequencer 6001 to perform sequencing of the sample nucleic acid. Insome embodiments, the hardware processor 6005 is configured to output, on a display, the most likely lengths of TR regions in the sample nucleic acid.

[0072] FIG. 6B is a block diagram of an exemplary computing device 600 that may be used in connection with the illustrative sequencing system 6000 of FIG. 6A. The computing device 600 may be configured to determine a VNTR status, such as identifying a VNTR. The general architecture of the computing device 600 depicted in FIG. 6B includes an arrangement of computer hardware and software components. The computing device 600 may include many more (or fewer) elements than those shown in FIG. 6B. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the computing device 600 includes a processing unit 610, a network interface 620, a computer readable medium drive 630, an input / output device interface 640, a display 650, and an input device 660, all of which may communicate with one another by way of a communication bus. The network interface 620 may provide connectivity to one or more networks or computing systems. The processing unit 610 may thus receive information and instructions from other computing systems or services via a network. The processing unit 610 may also communicate to and from memory 670 and further provide output information for an optional display 650 via the input / output device interface 640. The input / output device interface 640 may also accept input from the optional input device 660, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0073] The memory 670 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 610 executes in order to implement one or more embodiments. The memory 670 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer-readable media. The memory 670 may store an operating system 672 that provides computer program instructions for use by the processing unit 610 in the general administration and operation of the computing device 600. The memory 670 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0074] For example, in one embodiment, the memory 670 includes a VNTR status determination module 674 for determining a VNTR status. The VNTR status determination module 674 can perform the methods disclosed herein. In addition, memory 670 may includeor communicate with the data store 690 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of determining a VNTR status of the present disclosure, such the long reads, the short reads, and the VNTR status determined.

[0075] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.

[0076] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.

[0077] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.

[0078] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly or are indirectly held on a computer readable medium and loaded as needed onto acomputer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third party service provider.

[0079] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (i.e., PDA, Blackberry, iPhone), a tablet computer (e.g., iPAD), a hard drive, a server, a memory stick, a flash drive and the like.

[0080] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by theindividual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.

[0081] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.

[0082] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.

[0083] In some embodiments, a hardware platform for providing a computational environment comprises a processor (i.e., CPU) wherein processor time and memory layout such as random access memory (i.e., RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described hereincomprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.

[0084] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (i.e., grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations.Samples

[0085] In some embodiments, the sample comprises or consists of a purified or isolated polynucleotide derived from a tissue sample, a biological fluid sample, a cell sample, and the like. Suitable biological fluid samples include, but are not limited to blood, plasma, serum, sweat, tears, sputum, urine, sputum, ear flow, lymph, saliva, cerebrospinal fluid, ravages, bone marrow suspension, vaginal flow, trans-cervical lavage, brain fluid, ascites, milk, secretions of the respiratory, intestinal and genitourinary tracts, amniotic fluid, milk, and leukophoresis samples. In some embodiments, the sample is a sample that is easily obtainable by non-invasive procedures, e.g., blood, plasma, serum, sweat, tears, sputum, urine, sputum, ear flow, saliva or feces. In certain embodiments the sample is a peripheral blood sample, or the plasma and / or serum fractions of a peripheral blood sample. In other embodiments, the biological sample is a swab or smear, a biopsy specimen, or a cell culture. In another embodiment, the sample is a mixture of two or more biological samples, e.g., a biological sample can comprise two or more of a biological fluid sample, a tissue sample, and a cell culture sample. As used herein, the terms “blood,” “plasma” and “serum” expressly encompass fractions or processed portions thereof. Similarly, where a sample is taken from a biopsy, swab, smear, etc., the “sample” expressly encompasses a processed fraction or portion derived from the biopsy, swab, smear, etc.

[0086] In certain embodiments, samples can be obtained from sources, including, but not limited to, samples from different individuals, samples from different developmentalstages of the same or different individuals, samples from different diseased individuals (e.g., individuals with cancer or suspected of having a genetic disorder), normal individuals, samples obtained at different stages of a disease in an individual, samples obtained from an individual subjected to different treatments for a disease, samples from individuals subjected to different environmental factors, samples from individuals with predisposition to a pathology, samples individuals with exposure to an infectious disease agent, and the like.

[0087] In one illustrative, but non-limiting embodiment, the sample is a maternal sample that is obtained from a pregnant female, for example a pregnant woman. The maternal sample can be a tissue sample, a biological fluid sample, or a cell sample. In another illustrative, but non-limiting embodiment, the maternal sample is a mixture of two or more biological samples, e.g., the biological sample can comprise two or more of a biological fluid sample, a tissue sample, and a cell culture sample.

[0088] In certain embodiments samples can also be obtained from in vitro cultured tissues, cells, or other polynucleotide-containing sources. The cultured samples can be taken from sources including, but not limited to, cultures (e.g., tissue or cells) maintained in different media and conditions (e.g., pH, pressure, or temperature), cultures (e.g., tissue or cells) maintained for different periods of length, cultures (e.g., tissue or cells) treated with different factors or reagents (e.g., a drug candidate, or a modulator), or cultures of different types of tissue and / or cells.

[0089] In some embodiments, the use of the disclosed sequencing technology does not involve the preparation of sequencing libraries. In other embodiments, the sequencing technology contemplated herein involve the preparation of sequencing libraries. In one illustrative approach, sequencing library preparation involves the production of a random collection of adapter-modified DNA fragments (e.g., polynucleotides) that are ready to be sequenced.

[0090] Sequencing libraries of polynucleotides can be prepared from DNA or RNA, including equivalents, analogs of either DNA or cDNA, for example, DNA or cDNA that is complementary or copy DNA produced from an RNA template, by the action of reverse transcriptase. The polynucleotides may originate in double-stranded form (e.g., dsDNA such as genomic DNA fragments, cDNA, PCR amplification products, and the like) or, in certain embodiments, the polynucleotides may originate in single-stranded form (e.g., ssDNA, RNA,etc.) and have been converted to dsDNA form. By way of illustration, in certain embodiments, single stranded mRNA molecules may be copied into double-stranded cDNAs suitable for use in preparing a sequencing library. The precise sequence of the primary polynucleotide molecules is generally not material to the method of library preparation, and may be known or unknown. In one embodiment, the polynucleotide molecules are DNA molecules. More particularly, in certain embodiments, the polynucleotide molecules represent the entire genetic complement of an organism or substantially the entire genetic complement of an organism, and are genomic DNA molecules (e.g., cellular DNA, cell free DNA (cfDNA), etc.), that typically include both intron sequence and exon sequence (coding sequence), as well as noncoding regulatory sequences such as promoter and enhancer sequences. In certain embodiments, the primary polynucleotide molecules comprise human genomic DNA molecules, e.g., cfDNA molecules present in peripheral blood of a pregnant subject.

[0091] Methods of isolating nucleic acids from biological sources may differ depending upon the nature of the source. One of skill in the art can readily isolate nucleic acids from a source as needed for the method described herein. In some instances, it can be advantageous to fragment large nucleic acid molecules (e.g. cellular genomic DNA) in the nucleic acid sample to obtain polynucleotides in the desired size range. Fragmentation can be random, or it can be specific, as achieved, for example, using restriction endonuclease digestion. Methods for random fragmentation may include, for example, limited DNase digestion, alkali treatment and physical shearing. Fragmentation can also be achieved by any of a number of methods known to those of skill in the art. For example, fragmentation can be achieved by mechanical means including, but not limited to nebulization, sonication and hydroshear.

[0092] In some embodiments, sample nucleic acids are obtained from as cfDNA, which is not subjected to fragmentation. For example, cfDNA, typically exists as fragments of less than about 300 base pairs and consequently, fragmentation is not typically necessary for generating a sequencing library using cfDNA samples.

[0093] Typically, whether polynucleotides are forcibly fragmented (e.g., fragmented in vitro), or naturally exist as fragments, they are converted to blunt-ended DNA having 5 ’-phosphates and 3 ’-hydroxyl. Protocols for sequencing may instruct users to end-repair sample DNA, to purify the end-repaired products prior to dA-tailing, and to purify the dA-tailing products prior to the adaptor- ligating steps of the library preparation.

[0094] In various embodiments, verification of the integrity of the samples and sample tracking can be accomplished by sequencing mixtures of sample genomic nucleic acids, e.g., cfDNA, and accompanying marker nucleic acids that have been introduced into the samples, e.g., prior to processing.Definitions

[0095] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See, e.g. Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of the present disclosure, the following terms are defined below.

[0096] As used herein, a “nucleotide” includes a nitrogen containing heterocyclic base, a sugar, and one or more phosphate groups. Nucleotides are monomeric units of a nucleic acid sequence. Examples of nucleotides include, for example, ribonucleotides or deoxyribonucleotides. In ribonucleotides (RNA), the sugar is a ribose, and in deoxyribonucleotides (DNA), the sugar is a deoxyribose, i.e., a sugar lacking a hydroxyl group that is present at the 2' position in ribose. The nitrogen containing heterocyclic base can be a purine base or a pyrimidine base. Purine bases include adenine (A) and guanine (G), and modified derivatives or analogs thereof. Pyrimidine bases include cytosine (C), thymine (T), and uracil (U), and modified derivatives or analogs thereof. The C-l atom of deoxyribose is bonded to N-l of a pyrimidine or N-9 of a purine. The phosphate groups may be in the mono- , di-, or tri-phosphate form. These nucleotides may be natural nucleotides, but it is to be further understood that non-natural nucleotides, modified nucleotides or analogs of the aforementioned nucleotides can also be used.

[0097] As used herein, “nucleobase” is a heterocyclic base such as adenine, guanine, cytosine, thymine, uracil, inosine, xanthine, hypoxanthine, or a heterocyclic derivative, analog, or tautomer thereof. A nucleobase can be naturally occurring or synthetic. Non-limiting examples of nucleobases are adenine, guanine, thymine, cytosine, uracil,xanthine, hypoxanthine, 8-azapurine, purines substituted at the 8 position with methyl or bromine, 9-oxo-N6-methyladenine, 2-aminoadenine, 7-deazaxanthine, 7-deazaguanine, 7- deaza-adenine, N4-ethanocytosine, 2,6- diaminopurine, N6-ethano-2,6-diaminopurine, 5- methylcytosine, 5-(C3-C6)- alkynylcytosine, 5-fluorouracil, 5-bromouracil, thiouracil, pseudoisocytosine, 2-hydroxy-5-methyl-4-triazolopyridine, isocytosine, isoguanine, inosine, 7,8-dimethylalloxazine, 6-dihydrothymine, 5,6-dihydrouracil, 4-methyl-indole, ethenoadenine and the non-naturally occurring nucleobases described in U.S. Pat. Nos. 5,432,272 and 6,150,510 and PCT applications WO 92 / 002258, WO 93 / 10820, WO 94 / 22892, and WO 94 / 24144, and Fasman (“Practical Handbook of Biochemistry and Molecular Biology”, pp. 385-394, 1989, CRC Press, Boca Raton, LO), all herein incorporated by reference in their entireties.

[0098] The term “nucleic acid” or “polynucleotide” refers to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, diTP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.

[0099] The term “primer,” as used herein refers to an isolated oligonucleotide that is capable of acting as a point of initiation of synthesis when placed under conditions inductive to synthesis of an extension product (e.g., the conditions include nucleotides, an inducing agent such as DNA polymerase, and a suitable temperature and pH). The primer is preferably single stranded for maximum efficiency in amplification, but may alternatively be double stranded. If double stranded, the primer is first treated to separate its strands before being used to prepare extension products. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be sufficiently long to prime the synthesis of extension products in the presence of the inducingagent. The exact lengths of the primers will depend on many factors, including temperature, source of primer, use of the method, and the parameters used for primer design.

[0100] As used herein the term “chromosome” refers to the heredity-bearing gene carrier of a living cell, which is derived from chromatin strands comprising DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is employed herein.

[0101] A “genome” refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences.

[0102] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. For example, the reference sequence can be a reference human genome sequence, such as hgl9 or hg38. In another example, the reference sequence is limited to a specific human chromosome such as chromosome 13. In some embodiments, a reference Y chromosome is the Y chromosome sequence from human genome version hgl9. Such sequences may be referred to as chromosome reference sequences. Other examples of reference sequences include genomes of other species, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual.

[0103] The term “nucleic acid sample” herein refers to a sample, typically derived from a biological fluid, cell, tissue, organ, or organism, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence that is to be screened for copy number variation. In certain embodiments the nucleic acid sample comprises at least onenucleic acid sequence whose copy number is suspected of having undergone variation. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (e.g., surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (e.g., patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (e.g., namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein.

[0104] The term “subject” herein refers to a human subject as well as a non-human subject such as a mammal, an invertebrate, a vertebrate, a fungus, a yeast, a bacterium, and a virus. Although the examples herein concern humans and the language is primarily directed to human concerns, the concepts disclosed herein are applicable to genomes from any plant or animal, and are useful in the fields of veterinary medicine, animal sciences, research laboratories and such.

[0105] The term “condition” or “medical condition” is used herein as a broad term that includes all diseases and disorders, but can include injuries and normal health situations, such as pregnancy, that might affect a person’s health, benefit from medical assistance, or have implications for medical treatments.

[0106] As used herein, the term “cluster” or “clump” refers to a group of molecules, e.g., a group of DNA, or a group of signals. In some embodiments, the signals of a cluster are derived from different features. In some embodiments, a signal clump represents a physical region covered by one amplified oligonucleotide. Each signal clump could be ideally observedas several signals. Accordingly, duplicate signals could be detected from the same clump of signals. In some embodiments, a cluster or clump of signals can comprise one or more signals or spots that correspond to a particular feature. When used in connection with microarray devices or other molecular analytical devices, a cluster can comprise one or more signals that together occupy the physical region occupied by an amplified oligonucleotide (or other polynucleotide or polypeptide with a same or similar sequence). For example, where a feature is an amplified oligonucleotide, a cluster can be the physical region covered by one amplified oligonucleotide. In other embodiments, a cluster or clump of signals need not strictly correspond to a feature. For example, spurious noise signals may be included in a signal cluster but not necessarily be within the feature area. For example, a cluster of signals from four cycles of a sequencing reaction could comprise at least four signals.

[0107] The term “next generation sequencing (NGS)” herein refers to sequencing methods that allow for massively parallel sequencing of clonally amplified molecules and of single nucleic acid molecules. Non-limiting examples of NGS include sequencing-by- synthesis using reversible dye terminators, and sequencing-by-ligation.

[0108] The term “read” or “sequence read” (or sequencing reads) refer to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (e.g., at least about 25 bp) that can be used to identify a larger sequence or region, e.g., that can be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (e.g., 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. A sequence read may be obtained in a variety of ways, e.g., using sequencing techniques or using probes, e.g., in hybridization arrays or capture probes, or amplificationtechniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by binding, or sequencing by ligation. Sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).

[0109] The term “sequencing depth,” as used herein, generally refers to the number of times a locus is covered by a sequence read aligned to the locus. The locus may be as small as a nucleotide, or as large as a chromosome arm, or as large as the entire genome. Sequencing depth can be expressed as 50 , 100 , etc., where “x” refers to the number of times a locus is covered with a sequence read. Sequencing depth can also be applied to multiple loci, or the whole genome, in which case x can refer to the mean number of times the loci or the haploid genome, or the whole genome, respectively, is sequenced. When a mean depth is quoted, the actual depth for different loci included in the dataset spans over a range of values. Ultra-deep sequencing can refer to at least 100 / in sequencing depth.

[0110] The term “coverage” refers to the abundance of sequence tags mapped to a defined sequence. Coverage can be quantitatively indicated by sequence tag density (or count of sequence tags), sequence tag density ratio, normalized coverage amount, adjusted coverage values, etc. In some cases, “effective read coverage” of a chromosome is defined as the actual amount of bases covered by reads. Sequencing depth, which refers to the expected coverage of nucleotides by reads, is computed based on the assumption that reads are synthesized uniformly across chromosomes. In reality, read coverage across genomes is not uniform. Although a coverage of lOx, for example, means a nucleotide is covered 10 times on average, in certain parts of a genome, nucleotides are covered much more or much less. One factor that influences coverage is the ability of a read aligner to align reads to genomes. If a part of a genome is complex, e.g. having many repeats, aligners might have troubles aligning reads to that region, resulting in low coverage.

[0111] As used herein, the terms “aligned,” “alignment,” or “aligning” refer to the process of comparing a read or tag to a reference sequence and thereby determining the likelihood of the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. For example, the alignmentof a read to the reference sequence for human chromosome 13 will tell the likelihood of the read is present in the reference sequence for chromosome 13. In some cases, an alignment additionally indicates a location where the read or tag maps to in the reference sequence. For example, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13. A “site” may be a unique position on a polynucleotide sequence or a reference genome (i.e. chromosome ID, chromosome position and orientation). In some embodiments, a site may provide a position for a residue, a sequence tag, or a segment on a sequence.

[0112] Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period for implementing the methods disclosed herein. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).

[0113] Alignment may be performed by modifications and / or combinations of methods such as Burrows-Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.

[0114] The term “mapping” used herein refers to specifically assigning a sequence read to a larger sequence, e.g., a reference genome, by alignment.

[0115] A “genetic variation” or “genetic alteration” refers to a particular genotype present in certain individuals, and often a genetic variation is present in a statistically significant sub-population of individuals. The presence or absence of a genetic variance can be determined using a method or apparatus described herein. In certain embodiments, the presence or absence of one or more genetic variations is determined according to an outcomeprovided by methods and apparatuses described herein. In some embodiments, a genetic variation is a chromosome abnormality (e.g., aneuploidy), partial chromosome abnormality or mosaicism, each of which is described in greater detail herein. Non-limiting examples of genetic variations include one or more deletions (e.g., micro-deletions), duplications (e.g., micro-duplications), insertions, mutations, polymorphisms (e.g., single-nucleotide polymorphisms), fusions, repeats (e.g., short tandem repeats), distinct methylation sites, distinct methylation patterns, the like and combinations thereof. An insertion, repeat, deletion, duplication, mutation or polymorphism can be of any length, and in some embodiments, is about 1 base or base pair (bp) to about 250 megabases (Mb) in length. In some embodiments, an insertion, repeat, deletion, duplication, mutation or polymorphism is about 1 base or base pair (bp) to about 1,000 kilobases (kb) in length (e.g., about 10 bp, 50 bp, 100 bp, 500 bp, 1 kb, 5 kb, 10 kb, 50 kb, 100 kb, 500 kb, or 1000 kb in length).

[0116] A genetic variation is sometimes a deletion. In certain embodiments a deletion is a mutation (e.g., a genetic aberration) in which a part of a chromosome or a sequence of DNA is missing. A deletion is often the loss of genetic material. Any number of nucleotides can be deleted. A deletion can comprise the deletion of one or more entire chromosomes, a segment of a chromosome, an allele, a gene, an intron, an exon, any non-coding region, any coding region, a segment thereof or combination thereof. A deletion can comprise a microdeletion. A deletion can comprise the deletion of a single base.

[0117] A genetic variation is sometimes a genetic duplication. In certain embodiments a duplication is a mutation (e.g., a genetic aberration) in which a part of a chromosome or a sequence of DNA is copied and inserted back into the genome. In certain embodiments a genetic duplication (i.e. duplication) is any duplication of a region of DNA. In some embodiments a duplication is a nucleic acid sequence that is repeated, often in tandem, within a genome or chromosome. In some embodiments a duplication can comprise a copy of one or more entire chromosomes, a segment of a chromosome, an allele, a gene, an intron, an exon, any non-coding region, any coding region, segment thereof or combination thereof. A duplication can comprise a microduplication. A duplication sometimes comprises one or more copies of a duplicated nucleic acid. A duplication sometimes is characterized as a genetic region repeated one or more times (e.g., repeated 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 times). Duplications can range from small regions (thousands of base pairs) to whole chromosomes insome instances. Duplications frequently occur as the result of an error in homologous recombination or due to a retrotransposon event. Duplications have been associated with certain types of proliferative diseases. Duplications can be characterized using genomic microarrays or comparative genetic hybridization (CGH).

[0118] A genetic variation is sometimes an insertion. An insertion is sometimes the addition of one or more nucleotide base pairs into a nucleic acid sequence. An insertion is sometimes a microinsertion. In certain embodiments an insertion comprises the addition of a segment of a chromosome into a genome, chromosome, or segment thereof. In certain embodiments an insertion comprises the addition of an allele, a gene, an intron, an exon, any non-coding region, any coding region, segment thereof or combination thereof into a genome or segment thereof. In certain embodiments an insertion comprises the addition (i.e., insertion) of nucleic acid of unknown origin into a genome, chromosome, or segment thereof. In certain embodiments an insertion comprises the addition (i.e. insertion) of a single base.

[0119] A genetic variation sometimes includes copy number variations, i.e., variations in the number of copies of a nucleic acid sequence present in a test sample in comparison with the copy number of the nucleic acid sequence present in a reference sample. In certain embodiments, the nucleic acid sequence is 1 kb or larger. In some cases, the nucleic acid sequence is a whole chromosome or significant portion thereof. A copy number variant may refer to the sequence of nucleic acid in which copy-number differences are found by comparison of a nucleic acid sequence of interest in test sample with an expected level of the nucleic acid sequence of interest. For example, the level of the nucleic acid sequence of interest in the test sample is compared to that present in a qualified sample. Copy number variants / variations may include deletions, including microdeletions, insertions, including microinsertions, duplications, multiplications, and translocations. CNVs encompass chromosomal aneuploidies and partial aneuploidies.

[0120] As used herein, the term “array” may refer to a sequence of given size in the genome. In some examples, an array may comprise the total length of a VNTR. In some examples, an array may include all of the repeat copies of a VNTR. In some examples, an array may further comprise another target region. In some embodiments, an array may include the whole VNTR as well as some non-repetitive elements in the genome.

[0121] As used herein, the term “consensus pattern motif (logo)” refers to the consensus sequence of the VNTR pattern describing the frequency at which different bases occur at each position.

[0122] As used herein, the term “copy number” refers to the number of times (e.g., 0, 1, 1.5, 2, 3.5, 5, etc.) the repeat unit is repeated for a given VNTR. The change in copy number for a VNTR can be represented as the difference in copy number relative to the reference (e.g., -1, 0, +1, +2, etc.).

[0123] As used herein, the term “fragment size” refers to the length of the original nucleic acid sequence used to generate paired-end reads, calculated based on where those reads are mapped.

[0124] As used herein, the term “indels” refers to small insertions or deletions less than 50 base pairs in length in a nucleic acid sequence.

[0125] As used herein, the term “paired-end reads” or “paired end reads” refers to paired reads generated from sequencing the forward and reverse ends of a larger nucleic acid fragment. In some examples, the forward and reverse ends of a larger nucleic acid fragment may share the same name. The paired-end reads may be generated from paired end sequencing that obtains one read from each end of a nucleic acid fragment.

[0126] As used herein, the term “pattern” refers to the sequence of a repeat unit of the tandem repeat.

[0127] As used herein, the term “mate” or “mate of a read” refers to the pair of the read in question; i.e., the other read generated from the same nucleic acid fragment.

[0128] As used herein, the term “repeat unit” refers to the sequence of a single copy that is repeated multiple times in a VNTR.

[0129] As used herein, the term “single nucleotide variants” or “SNVs” refers to single base substitutions in a nucleic acid sequence.

[0130] As used herein, the term “small variant event” refers to a collection of adjacent SNVs or indels that occurs in the same haplotype of the VNTR array within a maximal distance of each other (for example, a maximal distance of 10 base-pairs).

[0131] As used herein, the term “structural variation” or “SV” refers to a large nucleic acid variant greater than 50 base pairs corresponding to either a duplication, deletion, insertion, inversion, or translocation.

[0132] As used herein, the term “tandem repeat” or “TR” refers to a nucleic acid sequence with a repeat unit of at least 10 base pairs, where the repeat unit is repeated at least 1.6 times with a similarity score of at least 1.7, consistent with the definitions in “Benson, Gary. ‘Tandem repeats finder: a program to analyze DNA sequences.’ Nucleic acids research 27.2 (1999): 573-580, the disclosures of which are incorporated herein by reference in their entirety.

[0133] As used herein, the term “variable number tandem repeat” or “VNTR” refers to a tandem repeat that has been observed to vary in the number of copies (e.g., duplicated copies or deleted copies) in the population of a species.

[0134] As used herein, the term “VNTR array” refers to the sequence covering the entire length of a VNTR. The VNTR array includes all of the copies of the repeat units.

[0135] In some cases, two haplotypes of a VNTR may comprise different numbers of copies of the repeat unit. In some cases, two haplotypes of the VNTR may comprise an identical number of copies of the repeat unit. The repeat units in each of the two haplotypes can include differentiating bases. A sequence of the repeat unit of one of the two haplotypes and a sequence of the repeat unit of the other one of the two haplotypes can be different at one or more differentiating positions; these sequences can have (or can have at least) 70%, 75%, 80%, 85%, 90%, 95%, 99%, or more, sequence identity. A sequence of the repeat unit of one of the two haplotypes and a sequence of the repeat unit of the other one of the two haplotypes can be identical in some examples.

[0136] Each haplotype of a VNTR can comprise a plurality of copies of a repeat unit. The repeat unit can be (or be at least or be more than) 6 bps, 7 bps, 8 bps, 9 bps, 10 bps, 11 bps, 12 bps, 13 bps, 14 bps, 15 bps, 16 bps, 17 bps, 18 bps, 19 bps, 20 bps, or more in length. The number of the plurality of copies can be (or be at least or be more than) 1.6, or more. The pathogenic copy number can be equal to, more than, or less than, the copy number in the reference sequence.

[0137] Two copies of a repeat unit of a haplotype can include differentiating bases. For example, sequences of two copies of the repeat unit of a haplotype can be different at one or more differentiating positions (e.g., 2, 3, 4, 5, 10, 20, or more, positions). The sequences of the two copies of the repeat unit of a haplotype may have (or may have at least) 70%, 75%,80%, 85%, 90%, 95%, 99%, or more, sequence identity. Sequences of two copies of the repeat unit of a haplotype can be identical in some examples.

[0138] As used herein, the set of all fragments overlapping a TR is called a “pileup”.

[0139] As used herein, “fragment collection” refers to the processes of finding all fragments overlapping each TR and aggregating them into pileups.Additional Notes

[0140] The embodiments described herein are exemplary. Modifications, rearrangements, substitute processes, etc. may be made to these embodiments and still be encompassed within the teachings set forth herein. One or more of the steps, processes, or methods described herein may be carried out by one or more processing and / or digital devices, suitably programmed.

[0141] The various illustrative imaging or data processing techniques described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. The described functionality can be implemented in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the disclosure.

[0142] The various illustrative detection systems described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a processor configured with specific instructions, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, aplurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. For example, systems described herein may be implemented using a discrete memory chip, a portion of memory in a microprocessor, flash, EPROM, or other types of memory.

[0143] The elements of a method, process, or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of computer-readable storage medium known in the art. An exemplary storage medium can be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. A software module can comprise computer-executable instructions which cause a hardware processor to execute the computerexecutable instructions.

[0144] Conditional language used herein, such as, among others, “can,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, such conditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” “involving,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list.

[0145] Disjunctive language such as the phrase “at least one of X, Y or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y or Z, or any combination thereof (e.g., X,Y and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y or at least one of Z to each be present.

[0146] The terms “about” or “approximate” and the like are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, where the range can be ±20%, ±15%, ±10%, ±5%, or ±1%. The term “substantially” is used to indicate that a result (e.g., measurement value) is close to a targeted value, where close can mean, for example, the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value.

[0147] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items. Accordingly, phrases such as “a device configured to” or “a device to” are intended to include one or more recited devices. Such one or more recited devices can also be collectively configured to carry out the stated recitations. For example, “a processor to carry out recitations A, B and C” can include a first processor configured to carry out recitation A working in conjunction with a second processor configured to carry out recitations B and C.

[0148] While the above detailed description has shown, described, and pointed out novel features as applied to illustrative embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.

[0149] It should be appreciated that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.

Claims

WHAT IS CLAIMED IS:

1. A method of determining a length of a tandem repeat region in each haplotype of a diploid genome in a sample, the method comprising: obtaining paired-end sequencing reads of the diploid genome; aligning the paired-end sequencing reads to a tandem repeat region in a reference genome sequence; classifying paired-end sequencing reads that overlap the tandem repeat region based on the alignments into a plurality of classes and counting the number of paired- end sequencing reads in each class; providing a set of hypotheses of a length of the tandem repeat region in each haplotype of the diploid genome; and evaluating which hypothesis from the set of hypotheses has the highest probability of generating the observed number of paired-end sequencing reads in the plurality of classes to determine the length of the tandem repeat region in each haplotype of the diploid genome.

2. The method of claim 1, wherein the probability of a hypothesis generating the observed number of paired-end sequencing reads in the plurality of classes comprises a genotype prior and a likelihood of observing the observed number of paired-end sequencing reads in the plurality of classes given the hypothesis.

3. The method of claim 2, wherein the likelihood of observing the observed number of paired-end sequencing reads in the plurality of classes given the hypothesis depends on the distribution of the size of nucleic acid fragments associated with the paired-end sequencing reads that overlap the tandem repeat region.

4. The method of claim 3, further comprising computing the distribution of the size of nucleic acid fragments associated with the paired-end sequencing reads that overlap the tandem repeat region based on estimating the distribution of the size of nucleic acid fragments overlapping fixed non-repetitive regions in the diploid genome in the sample.

5. The method of any of claims 1-4, wherein the genotype prior is determined based on population frequencies of known lengths of the tandem repeat region.

6. The method of any of claims 1-5, wherein the genotype prior comprises unequal priors for each haplotype.

7. The method of any of claims 1-6, wherein a class in the plurality of classes includes paired-end sequencing reads and wherein at least one read spans the tandem repeat region in its entirety and overlaps with both flanks.

8. The method of any of claims 1-7, wherein a class in the plurality of classes includes paired-end sequencing reads and wherein a first read overlaps with a flank of the tandem repeat region and a second read overlaps with the other flank of the tandem repeat region.

9. The method of claim 8, wherein the first read and / or the second read also partially overlap with the tandem repeat region.

10. The method of any of claims 1-9, wherein a class in the plurality of classes includes paired-end sequencing reads and wherein a first read overlaps with both a flank of the tandem repeat region and the tandem repeat region, and a second read overlaps with the same flank of the tandem repeat region.

11. The method of claim 10, wherein the second read also overlaps with the tandem repeat region.

12. The method of any of claims 1-11, wherein a class in the plurality of classes includes paired-end sequencing reads and wherein a first read overlaps with a flank of the tandem repeat region and a second read is contained within the tandem repeat region.

13. The method of claim 12, wherein the first read also overlaps with the tandem repeat region.

14. The method of any of claims 1-13, wherein a class in the plurality of classes includes paired-end sequencing reads and wherein both reads are contained within the tandem repeat region.

15. The method of any of claims 1-14, wherein obtaining paired-end sequencing reads of the diploid genome comprises performing paired-end sequencing of the sample.

16. The method of any of claims 1-15, wherein aligning the paired-end sequencing reads uses wrap-around alignment.

17. The method of any of claims 1-16, wherein the sample is extracted from cells, a cell-free DNA sample, an amniotic fluid, a blood sample, a biopsy sample, or any combination thereof, of a subject.

18. The method of any of claims 1-17, wherein the tandem repeat region is a variable number tandem repeat (VNTR) locus.

19. The method of any of claims 1-18, wherein each nucleic acid fragment associated with the paired-end sequencing reads is about 250 base pairs to about 1000 base pairs in length.

20. The method of any of claims 1-19, wherein the paired-end sequencing reads are generated by whole genome sequencing (WGS) of the sample.

21. The method of any of claims 1 -20, wherein the paired-end sequencing reads are generated by a next generation sequencing reaction.

22. A system for determining a length of a tandem repeat region in each haplotype of a diploid genome from a sample, the system comprising: a nucleic acid sequencer; non-transitory memory configured to store executable instructions; anda hardware processor in communication with the nucleic acid sequencer and the non-transitory memory, the hardware processor programmed by the executable instructions to perform the method of any of claims 1 to 21.

23. The system of claim 22, wherein the hardware processor is configured to receive paired-end sequencing reads from the nucleic acid sequencer.

24. The system of claim 22 or claim 23, wherein the hardware processor is configured to control the nucleic acid sequencer to perform sequencing.

25. The system of any of claims 22-24, wherein the hardware processor is configured to output, on a display, the length of the tandem repeat region in each haplotype of the diploid genome.

Citation Information

Patent Citations

  • Method for incorporating into a DNA or RNA oligonucleotide using nucleotides bearing heterocyclic bases

    US5432272A

  • Modified oligonucleotides, their preparation and their use

    US6150510A

  • Nuclease resistant, pyrimidine modified oligonucleotides that detect and modulate gene expression

    WO1992002258A1

  • Enhanced triple-helix and double-helix formation with oligomers containing modified pyrimidines

    WO1993010820A1

  • 7-deazapurine modified oligonucleotides

    WO1994022892A1