NGS-based aneuploidy detection using HMM method with support for mosaicism
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-09
- Publication Date
- 2026-08-13
Smart Images

Figure IMGF000011_0001_TABLE 
Figure IMGF000012_0001_TABLE 
Figure IMGF000013_0001_TABLE
Abstract
Description
Attorney Docket No. 206979-713601NGS-BASED ANEUPLOIDY DETECTION USING HMM METHOD WITH SUPPORT FOR MOSAICISMCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims priority and the benefit of U.S. Provisional Application No.63 / 756,647, filed February 10, 2025, the content of which is incorporated by reference in its entirety.BACKGROUND
[0002] There is a need for accurately predicting mosaic aneuploidies (corresponding to noninteger chromosome numbers) to inform many diseases, disorders and conditions. Current methods only model a small number of fixed, discrete states, while mosaic aneuploidies present as having variable, continuous numbers of chromosomes, depending on the mosaicism percentage. The present disclosure relates generally to methods of detecting and predicting mosaic aneuploidies.SUMMARY
[0003] Provided herein are methods for detecting at least one mosaic aneuploidies in an embryonic sample, the method comprising: obtaining an embryonic chromosome data file from the embryonic sample; obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file; estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that reflects the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample. In some embodiments, the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, or neural network methodologies. In some embodiments, the at least one statistical model is an HMM. In some embodiments, the at least one estimated mosaic aneuploidy states are added to the HMM. In some embodiments, the HMM is capable of generating a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions. In some embodiments,Attorney Docket No. 206979-713601the identifying of the at least one candidate chromosomal regions for mosaic aneuploidy is determined based on the phased parental chromosome data file, and / or sequence read counts comprised in the embryonic chromosome data file. In some embodiments, the estimated mosaicism percentage is estimated based on sequence read counts at each of the at least one candidate chromosomal regions comprised in the embryonic chromosome data file relative to the corresponding chromosomal region of the allele balance signal. In some embodiments, the sequence read counts comprised in the embryonic chromosome data file comprises read counts for a reference allele and an alternate allele. In some embodiments, the phased parental chromosome data file comprises a phased maternal chromosome datafile, a phased paternal chromosome data file, or a combination of both. In some embodiments, the phased maternal chromosome data file is obtained based on a sequence alignment file of maternal chromosome information and / or the phased paternal chromosome data file is obtained based on a sequence alignment file of paternal chromosome information. In some embodiments, the phased maternal chromosome data file and / or the phased paternal chromosome data file is in VCF format. In some embodiments, the at least one estimated mosaic aneuploidy states are generated based on a matrix of euploidy and aneuploidy probabilities at each of the at least one candidate chromosomal regions. In some embodiments, a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions comprises one of 64 possible states. In some embodiments, the 64 possible states comprises 4 possible euploidy states and 60 possible aneuploidy states. In some embodiments, the at least one estimated mosaic aneuploidy states comprises an aggregate of euploidy and aneuploidy probabilities at each of the at least one candidate chromosomal regions based on a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions. In some embodiments, identifying the at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file comprises: dividing each chromosome from the embryonic chromosome data file, the phased parental chromosome data file and / or the allele balance signal into a fixed number of bins; obtaining read depth signals from the embryonic chromosome data file and the allele balance signal in each bin; determining whether the read depth signals and allele balance signals in each bin deviate from an expected integer chromosome state; and identifying a bin in which both the read depth signals and the allele balance signals deviate more than a threshold from the expected integer chromosome state as a candidate chromosomal region for mosaic aneuploidy. In some embodiments, the expected integer chromosome state comprises euploidy expected integer chromosome state (2), monosomy expected integer chromosome state (1), or trisomy expected integer chromosome state (3). In some embodiments, the read depth signals from the embryonic chromosome data file is normalized. In some embodiments, comprising identifying at least two adjacent bins in which both the read-depthAtorney Docket No. 206979-713601signals and the allele balance signals deviate more than a threshold from the integer chromosome states as the candidate chromosomal region for mosaic aneuploidy. In some embodiments, the estimated mosaicism percentage at each of the at least one candidate chromosomal regions is based on a normalized and averaged read depth in each of the at least one candidate chromosomal regions and a chromosome copy number.- In some embodiments, the chromosome copy number comprises a euploidy expected integer chromosome state (2). In some embodiments, the mosaicism percentage at each of the at least one candidate chromosomal regions is estimated according to a formula: m = d - 2, wherein d > 2; and m = 2 - d, wherein d < 2, wherein the m is the estimated mosaicism percentage at each of the at least one candidate chromosomal regions, and the d is a normalized and averaged read depth in each of the at least one candidate chromosomal regions. In some embodiments, the d corresponds to a chromosome copy number at the at least one candidate chromosomal regions. In some embodiments, the d is a non-integer. In some embodiments, the method comprises determining a probability of a mosaic aneuploidy in the embryonic sample, wherein the probability of the mosaic aneuploidy in the embryonic sample is based on a percentage of cells in the embryonic sample having an estimated mosaicism percentage above a predetermined threshold in at least one of the candidate chromosomal regions. In some embodiments, the at least one estimated mosaic aneuploidy states reflects that one parent is considered to have contributed to an embryo correlating the embryonic sample a continuous number of chromosomes d-1, wherein d is the normalized and averaged read depth in each of the at least one candidate chromosomal regions. In some embodiments, the methods further comprise selecting an embryo correlating to the embryonic sample for implantation based on a determination using the probability of the mosaic aneuploidy in the embryonic sample. Also provided herein are apparatuses for detecting at least one mosaic aneuploidies in an embryonic sample, the apparatus comprising a processor and a memory storing software instructions that, when executed by the processor, cause the apparatuses to perform the methods disclosed herein. Also provided herein are computer program products for detecting at least one mosaic aneuploidies in an embryonic sample, the computer program products comprising at least one non-transitory computer-readable storage medium storing software instructions that, when executed by an apparatus, cause the apparatus to perform the methods disclosed herein.
[0004] Provided herein are apparatuses for performing a method for detecting at least one mosaic aneuploidies in an embryonic sample, the apparatus comprising: a processor; a memory for receiving a plurality of data files and for storing software instruction that, when executed by the processor, cause the apparatus to perform the method using the plurality of data files, wherein said plurality of data files comprise an embryonic chromosome data file, a phased parental chromosome data file, and wherein said method comprises: obtaining an embryonic chromosome data file fromAtorney Docket No. 206979-713601the embryonic sample; obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file; estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that correlates to the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample. In some embodiments, the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, or neural network methodologies. In some embodiments, the at least one statistical model is an HMM.
[0005] Also provided herein are systems comprising an apparatus and software instruction that, when executed by the apparatus, cause the apparatus to perform a method for detecting at least one mosaic aneuploidies in an embryonic sample, wherein said method comprises: obtaining an embryonic chromosome data file from the embryonic sample; obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file; estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that correlates to the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample. In some embodiments, the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, or neural network methodologies. In some embodiments, the at least one statistical model is an HMM.INCORPORATION BY REFERENCE
[0006] All publications, patents and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.Atorney Docket No. 206979-713601BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:
[0008] FIG. 1 shows an illustrative graph depicting HMM aneuploidy detection without support for mosaicism. The X-axis represents the position on chromosome 13, and the Y-axis represents the posterior probability.
[0009] FIG. 2 shows an illustrative graph depicting HMM aneuploidy detection with support for mosaicism. The X-axis represents the position on chromosome 13, and the Y-axis represents the posterior probability.
[0010] FIG. 3 shows a graph illustrating the copy number of chromosome 21 corresponding to the level of mosaicism in terms of percentage from 10-90%. The X-axis represents the position in the human genome (chromosome 1-22), and the Y-axis represents chromosome copy number.DETAILED DESCRIPTION
[0011] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only, and are not restrictive of the disclosure.
[0012] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0013] All documents, or portions of documents, cited in this application, including, but not limited to, patents, patent applications, articles, books and treatises, are hereby expressly incorporated by reference in their entirety for any purpose.Definitions
[0014] Unless defined otherwise, all terms of art, notations and other technical and scientific terms or terminology used herein are intended to have the same meaning as is commonly understood by one of ordinary skill in the art to which the claimed subject matter pertains. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over what is generally understood in the art.
[0015] Throughout this application, various embodiments may be presented in a range format. The description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such asAtorney Docket No. 206979-713601from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.
[0016] The singular forms “a,” “an,” and “the” are used herein to include plural references unless the context clearly dictates otherwise. Accordingly, unless the contrary is indicated, the numerical parameters set forth in this application are approximations that can vary depending upon the desired properties sought to be obtained.
[0017] The terms “determining,” “measuring,” “evaluating,” “assessing,” “assaying,” and “analyzing” are often used interchangeably herein to refer to forms of measurement and include determining if an element may be present or not (for example, detection), or the amount of an element. These terms can include quantitative and qualitative determinations. Assessing can be alternatively relative or absolute. “Detecting the presence of’ includes determining the amount of something present, as well as determining whether it may be present or absent.
[0018] The terms, “or” and “and / or,” as used herein, include any and all combinations of one or more of the associated listed items.
[0019] The terms, “including,” “includes,” “included,” and other forms, are not limiting.
[0020] The terms, “comprise” and its grammatical equivalents, as used herein, specify the presence of stated features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0021] The term, “about,” as used herein in reference to a number or range of numbers, is understood to mean the stated number and numbers + / - 10% thereof, or 10% below the lower listed limit and 10% above the higher listed limit for the values listed for a range.
[0022] As used herein, the term “nucleic acids” includes, but is not limited to, any polymer or oligomer of pyrimidine and purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively. Indeed, the present disclosure contemplates any deoxyribonucleotide, ribonucleotide or peptide nucleic acid component, and any chemical variants thereof, such as methylated, hydroxymethylated or glucosylated forms of these bases, and the like. The polymers or oligomers may be heterogeneous or homogeneous in composition, and may be isolated from naturally occurring sources or may be artificially or synthetically produced. In addition, the nucleic acids may be DNA or RNA, or a mixture thereof, and may exist permanently or transitionally in single-stranded or double-stranded form, including homoduplex, heteroduplex, and hybrid states. An oligonucleotide or polynucleotide is a nucleic acid ranging from at least 2, preferably at least 8, 15 or 20 nucleotides in length, but may be up to 50, 100, 1000, or 5000 nucleotides long or a compound that specifically hybridizes to a polynucleotide. Polynucleotides of the presentAtorney Docket No. 206979-713601disclosure include sequences of deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) or mimetics thereof which may be isolated from natural sources, recombinantly produced or artificially synthesized. A further example of a polynucleotide of the present disclosure may be a peptide nucleic acid (PNA). In some embodiments, the nucleic acids disclosed herein also encompasses situations in which there is a nontraditional base pairing such as Hoogsteen base pairing which has been identified in certain tRNA molecules and postulated to exist in a triple helix. “Polynucleotide” and “oligonucleotide” are used interchangeably in this application.
[0023] As used herein, the term fragment refers to a portion of a larger DNA polynucleotide or DNA. A polynucleotide, for example, can be broken up, or fragmented into, a plurality of fragments. Various methods of fragmenting nucleic acid such as either chemical or physical exist. Chemical fragmentation may include partial degradation with a DNase; partial depurination with acid; the use of restriction enzymes; intron-encoded endonucleases; DNA-based cleavage methods, such as triplex and hybrid formation methods, that rely on the specific hybridization of a nucleic acid segment to localize a cleavage agent to a specific location in the nucleic acid molecule; or other enzymes or compounds which cleave DNA at known or unknown locations. Physical fragmentation methods may involve subjecting the DNA to a high shear rate. High shear rates may be produced, for example, by moving DNA through a chamber or channel with pits or spikes, or forcing the DNA sample through a restricted size flow passage, e.g., an aperture having a cross sectional dimension in the micron or submicron scale. Other physical methods include sonication and nebulization. Combinations of physical and chemical fragmentation methods may likewise be employed such as fragmentation by heat and ion-mediated hydrolysis. These methods can be optimized to digest a nucleic acid into fragments of a selected size range. Useful size ranges may be from 100, 200, 400, 700 or 1000 to 500, 800, 1500, 2000, 4000 or 10,000 base pairs. However, larger size ranges such as 4000, 10,000 or 20,000 to 10,000, 20,000 or 500,000 base pairs may also be useful.
[0024] As used herein, the term “Genome” designates or denotes the complete, single-copy set of genetic instructions for an organism as coded into the DNA of the organism. A genome may be multi-chromosomal such that the DNA is cellularly distributed among a plurality of individual chromosomes. For example, in human there are 22 pairs of chromosomes plus a gender associated XX or XY pair.
[0025] The term “chromosome” refers to the heredity-bearing gene carrier of a living cell which is derived from chromatin and which comprises DNA and protein components (especially histones). The conventional internationally recognized individual human genome chromosome numbering system is employed herein. The size of an individual chromosome can vary from one type to another with a given multi-chromosomal genome and from one genome to another. In theAtorney Docket No. 206979-713601case of the human genome, the entire DNA mass of a given chromosome is usually greater than about 100,000,000 bp. For example, the size of the entire human genome is about 3^109 bp. The largest chromosome, chromosome no. 1, contains about 2.4x108 bp while the smallest chromosome, chromosome no. 22, contains about 5.3x107 bp.
[0026] A “chromosomal region” is a portion of a chromosome. The actual physical size or extent of any individual chromosomal region can vary greatly. The term “region” is not necessarily definitive of a particular one or more genes because a region need not take into specific account the particular coding segments (exons) of an individual gene.
[0027] As used herein, the term “allele” refers to one specific form of a genetic sequence (such as a gene) within a cell, an individual or within a population, the specific form differing from other forms of the same gene in the sequence of at least one, and frequently more than one, variant sites within the sequence of the gene. The sequences at these variant sites that differ between different alleles are termed “variances,” “polymorphisms,” or “mutations.” At each autosomal specific chromosomal location or “locus” an individual possesses two alleles, one inherited from one parent and one from the other parent, for example one from the mother and one from the father. An individual is “heterozygous” at a locus if it has two different alleles at that locus. An individual is “homozygous” at a locus if it has two identical alleles at that locus.
[0028] The “polymorphism” refers to the occurrence of two or more genetically determined alternative sequences or alleles in a population. A polymorphic marker or site is the locus at which divergence occurs. Preferred markers have at least two alleles, each occurring at frequency of preferably greater than 1%, and more preferably greater than 10% or 20% of a selected population. A polymorphism may comprise one or more base changes, an insertion, a repeat, or a deletion. A polymorphic locus may be as small as one base pair. Polymorphic markers include, but is not limited to, restriction fragment length polymorphisms, variable number of tandem repeats (VNTR's), hypervariable regions, mini satellites, dinucleotide repeats, trinucleotide repeats, tetranucleotide repeats, simple sequence repeats, and insertion elements such as Alu. In some cases, the first identified allelic form is arbitrarily designated as the reference form and other allelic forms are designated as alternative or variant alleles. In some embodiments, the allelic form occurring most frequently in a selected population is sometimes referred to as the wildtype form. In some embodiments, diploid organisms are homozygous or heterozygous for allelic forms. In some embodiments, a diallelic polymorphism has two forms. In some embodiments, a triallelic polymorphism has three forms. In some embodiments, a polymorphism between two nucleic acids can occur naturally, or be caused by exposure to or contact with chemicals, enzymes, or other agents, or exposure to agents that cause damage to nucleic acids, for example, ultraviolet radiation, mutagens or carcinogens. In some embodiments, the polymorphismAtorney Docket No. 206979-713601is a single nucleotide polymorphism (SNP). In some embodiments, the SNP is at a position at which two alternative bases occur at appreciable frequency (>1%) in the human population. In some embodiments, the SNP is the most common type of human genetic variation.
[0029] As used herein, the term “read-depth” in DNA sequencing refers to the number of reads that include a given nucleotide in the reconstructed sequence. Coverage histograms are commonly used to depict the range and uniformity of sequencing coverage for an entire data set. They illustrate the overall coverage distribution by displaying the number of reference bases that are covered by mapped sequencing reads at various depths. Mapped “read depth” refers to the total number of bases sequenced and aligned at a given reference base position. Typically, in a sequencing coverage histogram, the read depths are binned and displayed on the x-axis, while the total numbers of reference bases that occupy each read depth bin are displayed on the y-axis. These can also be written as percentages of reference bases.
[0030] As used herein “depth coverage” refers to the number of unique reads that their mapping overlaps a specific chromosomal region or a genome coordinate.
[0031] The term subset or representative subset refers to a fraction of a genome. The subset may be 0.1, 1, 3, 5, 10, 25, 50 or 75% of the genome. The partitioning of fragments into subsets may be done according to a variety of physical characteristics of individual fragments. For example, fragments may be divided into subsets according to size, according to the particular combination of restriction sites at the ends of the fragment, or based on the presence or absence of one or more particular sequences.
[0032] As used herein, the term “aneuploidy” refers to the presence or absence of an entire chromosome, as well as the presence of partial chromosomal duplications or deletions or kilobase or greater size, as opposed to genetic mutations or polymorphisms where sequence differences exist.Overview
[0033] Disclosed herein are methods for detecting at least one mosaic aneuploidies in an embryonic sample. The methods as described herein can be used during in-vitro fertilization (IVF) to screen embryos for aneuploidies, or abnormal chromosome numbers. The methods described herein can be used for a pre-implantation testing for aneuploidy (PGT-A). The screening for aneuploidies is often based on HMMs (Hidden Markov Models) which take microarray or sequencing data as input, and which return probabilities for different euploid and aneuploid states of each chromosome as output. Hidden Markov Models are effective at predicting non-mosaic aneuploidies (corresponding to an integer numbers of chromosomes), however they often inaccurately predict mosaic aneuploidies (corresponding to non-integer chromosome numbers).Atorney Docket No. 206979-713601This is because HMMs can only model a small number of fixed, discrete states, while mosaic aneuploidies may present as having variable, continuous numbers of chromosomes, depending on the mosaicism percentage. The methods according to the present disclosure solves this problem and provides a method for detecting mosaic aneuploidies. The present methods have benefits over general chromosome abnormality detection as said methods are able to facilitate downstream embryo selection based on mosaic probability rather than just ploidy calling.Methods
[0034] Provided herein are methods for detecting at least one mosaic aneuploidies in an embryonic sample. In some embodiments, the method comprises obtaining an embryonic chromosome data file from the embryonic sample. In some embodiments, the method comprises obtaining a phased parental chromosome data file. In some embodiments, the method comprises generating an allele balance signal from the parental chromosome data file and embryonic chromosome data file. In some embodiments, the method comprises identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file. In some embodiments, the method comprises estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal. In some embodiments, the method comprises implementing a Hidden Markov Model (HMM), wherein the HMM comprises at least one estimated mosaic aneuploidy states that reflects the estimated mosaicism percentage, thereby detecting at least one mosaic aneuploidies in the embryonic sample. The method described herein is surprisingly able to detect at least one mosaic aneuploidies without being limited to integer copy-number hypotheses evaluated via statistical likelihood or allelic ratios as utilized by prior methods. Instead, the methods disclosed herein model non-integer chromosomal dosage, estimate a mosaic fraction per region, and / or construct model states that incorporate fractional aneuploidies. Such a method can produce a continuous-valued chromosomal dosage enabling calling partial mosaic monosomy or trisomy and / or low-level or segmental mosaicism, which prior methods could not compute because they lack continuous-dosage state handling.
[0035] As used herein, the term “mosaic aneuploidy” refers to a condition where an individual has a mixture of cells with different chromosome numbers, meaning some cells have an abnormal number of chromosomes (aneuploidy), while others have a normal number, creating a "mosaic" effect. The term “mosaicism” as used herein refers to aneuploidy in some cells, but not all cells, of a subject. Certain chromosome abnormalities can exist as mosaic and non-mosaic chromosome abnormalities. For example, certain trisomy 21 individuals have mosaic Down syndrome and some have non-mosaic Down syndrome. Different mechanisms can lead to mosaicism. For example, (i)Atorney Docket No. 206979-713601an initial zygote may have three 21 st chromosomes, which normally would result in simple trisomy 21, but during the course of cell division one or more cell lines lost one of the 21st chromosomes; and (ii) an initial zygote may have two 21st chromosomes, but during the course of cell division one of the 21st chromosomes were duplicated. Somatic mosaicism most likely occurs through mechanisms distinct from those typically associated with genetic syndromes involving complete or mosaic aneuploidy. In some cases, mosaicism comprise somatic mosaicism. Somatic has been identified in certain types of cancers and in neurons, for example, genetic syndromes in which an individual is predisposed to breakage of chromosomes (chromosome instability syndromes) are frequently associated with increased risk for several types of cancer, thus highlighting the role of somatic aneuploidy in carcinogenesis. Methods and protocols described herein can identify presence or absence of non-mosaic and mosaic chromosome abnormalities. In some embodiments, the methods disclosed herein improves the Hidden Markov Model (HMM) ability to detect nonmosaic and mosaic chromosome abnormalities. A non-limiting list of mosaic aneuploidies or chromosome abnormalities that can be identified using the methods disclosed herein include, but is not limited to the list provided in TABLE 1.TABLE 1: Exemplary list of chromosome abnormalitiesAtorney Docket No. 206979-713601Atorney Docket No. 206979-713601
[0036] In some embodiments, the embryonic sample comprises cell-free fetal DNA and RNA circulating in maternal blood. In some embodiments, the embryonic sample comprises circulating fetal DNA. In some embodiments, the embryonic sample is derived in the maternal bloodstream during pregnancy. In some embodiments, the embryonic sample can comprise cfDNA, or fragments of cell-free fetal RNA (cfRNA). In some embodiments, the fetal cfDNA is distinguished from the maternal cfDNA to enable the determination of a mosaic aneuploidy. In some embodiments, the embryonic sample does not comprise cell-free fetal DNA, cell-free fetal RNA, and / or and RNA circulating in maternal blood. In some embodiments, the embryonic sample comprises at least one cell obtained from an embryo. In some embodiments, methods described herein comprise a step of obtaining an embryonic sample, for example, drawing blood from a pregnant female subject, or from removing at least one cell from an embryo.
[0037] In some embodiments, the mosaic aneuploidy is a complete chromosomal trisomy or monosomy, or a partial trisomy or monosomy. In some embodiments, mosaic aneuploidies are caused by a loss-of or gain-of chromosome, and encompass chromosomal imbalances resulting from unbalanced translocations, unbalanced inversions, deletions and insertions. Other aneuploidies with known clinical significance include Edward syndrome (trisomy 18) and Patau Syndrome (trisomy 13), which are frequently fatal within the first few months of life. Abnormalities associated with the number of sex chromosomes are also known and include monosomy X e.g., Turner syndrome (XO), and triple X syndrome (XXX) in female births and Kleinefelter syndrome (XXY) and XYY syndrome in male births, which are all associated with various phenotypes including sterility and reduction in intellectual skills. The method of the present disclosure can be used to diagnose these and other chromosomal abnormalities prenatally. In some embodiments, the trisomy determined by the present disclosure include, but is not limited to, trisomy 21 (T21; Down Syndrome), trisomy 18 (T18; Edward's Syndrome), trisomy 16 (T16), trisomy 22 (T22; Cat Eye Syndrome), trisomy 15 (T15; Prader Willi Syndrome), trisomy 13 (T13; Patau Syndrome), trisomy 8 (T8; Warkany Syndrome) and the XXY (Kleinefelter Syndrome), XYY, XXX trisomies, partial trisomy lq32-44, trisomy 9 p, trisomy 4 mosaicism, trisomy 17p, partial trisomy 4q26-qter, trisomy 9, partial 2p trisomy, partial trisomy Iq, and / or partial trisomy 6p / monosomy 6q.Atorney Docket No. 206979-713601
[0038] In some embodiments, the method of the present disclosure determines chromosomal monosomy X, and partial monosomies such as, monosomy 13, monosomy 15, monosomy 16, monosomy 21, and monosomy 22, which are known to be involved in pregnancy miscarriage. Partial monosomy of chromosomes typically involved in complete aneuploidy can also be determined by the method according to the present disclosure. The method according to the present disclosure detects monosomy 18p. Monosomy 18 is a rare chromosomal disorder in which all or part of the short arm (p) of chromosome 18 is deleted (monosomic). The method according to the present disclosure detects conditions caused by changes in the structure or number of copies of chromosome 15 including Angelman Syndrome and Prader-Willi Syndrome. Angelman Syndrome and Prader-Willi Syndrome involve a loss of gene activity in the same part of chromosome 15, the 15qll-ql3 region. The method according to the present disclosure detects partial monosomy 13 q. Partial monosomy is a rare chromosomal disorder that results when a piece of the long arm (q) of chromosome 13 is missing (monosomic). The method according to the present disclosure detects 22qll.2 deletion syndrome, also known as DiGeorge syndrome. DiGeorge syndrome is caused by the deletion of a small piece of chromosome 22. In some embodiments, the method of the present disclosure is used to determine partial monosomies including, but not limited to, monosomy 18p, partial monosomy of chromosome 15 ( 15ql 1 -q 13), partial monosomy 13q, and partial monosomy of chromosome 22 can also be determined using the method. In some embodiments, the method of the present disclosure is used to determine any aneuploidy if one of the parents is a known carrier of such abnormality. These include, but not limited to, mosaic for a small supernumerary marker chromosome (SMC); t(l 1 ; 14)(p 15;p 13) translocations; unbalanced translocation t(8; 11)(p23.2;pl5.5); llq23 microdeletion; Smith-Magenis syndrome 17pl 1.2 deletion; 22ql3.3 deletion; Xp22.3 microdeletion; 10pl4 deletion; 20p microdeletion, DiGeorge syndrome [del(22)(ql 1.2ql 1.23)], Williams syndrome (7ql 1.23 and 7q36 deletions); lp36 deletion; 2p microdeletion; neurofibromatosis type 1 (17qll.2 microdeletion), Yq deletion ; Wolf-Hirschhorn syndrome (WHS, 4pl6.3 microdeletion); lp36.2 microdeletion; llql4 deletion; 19ql3.2 microdeletion; Rubinstein-Taybi (16 pl3.3 microdeletion); 7p21 microdeletion; Miller-Dieker syndrome (17pl3.3), 17pl 1.2 deletion; and 2q37 microdeletion.
[0039] The method as disclosed herein can utilize results from sequencing methodologies, including, but not limited to, Next Generation Sequencing (NGS), Sanger sequencing, Capillary sequencing, Bisulfite Sequencing, 454-Pyrosequencing, RNA sequencing, whole genome sequencing, targeted sequencing, short-read sequencing, or long-read sequencing. Sequencing of polynucleotides in some instances is verified by high-throughput sequencing such as by next generation sequencing. Sequencing of a polynucleotide library can be performed with anyAtorney Docket No. 206979-713601appropriate sequencing technology, including but not limited to single-molecule real-time (SMRT) sequencing, polony sequencing, sequencing by ligation, reversible terminator sequencing, proton detection sequencing, ion semiconductor sequencing, nanopore sequencing, electronic sequencing, pyrosequencing, Maxam-Gilbert sequencing, chain termination (e.g., Sanger) sequencing, +S sequencing, or sequencing by synthesis. In some embodiments, the output of sequencing is further analyzed by secondary analysis platform. In some embodiments, the output of sequencing and / or secondary analysis comprises whole genome sequencing files or whole genome nucleotide files. In some embodiments, the whole genome nucleotide files as disclosed herein, comprise sequence alignment files. In some embodiments, sequence alignment files, as disclosed herein, are in SAM, BAM or CRAM format. Noninformative reads that match the reference genome may also be used.
[0040] The number of times a single nucleotide or polynucleotide is identified or “read” is defined as the sequencing depth or read depth. In some embodiments, sequencing is performed at a read depth of at least 1, at least 5, at least 10, at least 50, at least 100, at least 500, at least 1000, at least 5000 at least 10,000, at least 20,000, at least 30,000, at least 50,000, at least 75,000, at least 100,000 unique reads per base. In some cases, the read depth is referred to as a fold coverage, for example, 55-fold (or 55*) coverage, optionally describing a percentage of bases. In some embodiments, the methods provided herein are used for aneuploidy detection using at least one signals from NGS. In some embodiments, the at least one signals comprise read counts, read density, read depth signals, or sequencing coverage. In some embodiments, sequencing results comprise reads from about 1.2 billion to about 6.5 billion base pairs. In some embodiments, the method described herein to detect aneuploidy from NGS data comprise two kinds of signals. In some embodiments, the two kinds of signals comprise (1) the overall read-depth, or coverage, and (2) the allele balance, or fraction or alleles inherited from the mother vs the father.
[0041] In some embodiments, a signal comprises a sequence read. In some embodiments, the sequence reads mapped are quantified to determine the number of reads that map to a region or portion of a reference genome. In some embodiments, a read that maps to a reference genome, a reference chromosome, or a region, portion, or segment thereof, is called a read count. In some embodiments, a read count is comprised of a value. In some embodiments, a count value is determined through a mathematical process. In some cases, a read count is determined by a suitable mathematical method, operation, or procedure. In some embodiments, a read count is weighted, removed, filtered, normalized, adjusted, averaged, added or subtracted, or processed by a combination thereof. In some embodiments, a read count is derived from a sequence read that is processed or manipulated by an appropriate mathematical method, operation, or process described herein or known in the art. For example, a read count is normalized and / or weighted based on atAtorney Docket No. 206979-713601least one biases associated with a sequence read. In some embodiments, a read count is normalized and / or weighted based on the GC bias associated with a sequence read. In some embodiments, a read count is derived from a raw sequence read and / or filtered sequence reads. In some embodiments, at least one read counts are not manipulated mathematically. The term “raw count” and “raw counts” as used herein refers to at least one read counts that have not been mathematically manipulated. In some embodiments, a read count is determined for some or all of the sequence reads mapped to a reference genome, a reference chromosome, or a region, portion, or segment thereof. In some embodiments, read counts are determined from a predefined subset of mapped sequence reads. In some cases, the predefined subsets (e.g., selected subsets) of mapped sequence reads can be defined or selected with the use of any suitable characteristic or variable. In some embodiments, the predefined subset of mapped sequence reads include, but is not limited to, 1 - n sequence reads, where n represents a number equal to the sum of all sequence reads generated from a test subject sample, a reference subject ample or a reference sample. In some cases, read counts are often derived from sequence reads obtained from an embryo (for example, an embryonic sample). In some cases, the read counts are derived from sequence readings obtained from a nucleic acid sample from a pregnant woman carrying a fetus. Nucleic acid sequence read counts are often representative read counts of both a fetus and a mother of a fetus (e.g, for a pregnant female subject). In some embodiments, where a subject is a pregnant woman, some read counts are derived from a fetal genome and some read counts are derived from a maternal genome.
[0042] In some cases, a sequence read count (e.g, weighted counts) is represented as a read density. A read density is determined and / or generated for at least one portions of a genome. In some embodiments, a read density is determined and / or generated for at least one chromosomes. In some embodiments, a read density is determined and / or generated for at least one chromosomal regions. In some embodiments, a read density comprises a quantitative measure of counts of sequence reads mapped to a portion of a reference genome. In some embodiments, a read density is determined by a suitable distribution and / or a suitable distribution function. Non-limiting examples of a distribution function include a probability function, probability distribution function, probability density function (PDF), a kernel density function (kernel density estimate), a cumulative distribution function, probability mass function, discrete probability distribution, an absolutely continuous univariate distribution, the like, any suitable distribution, or combinations thereof. Non-limiting examples of a kernel density function for generating a local genome bias estimate include a uniform kernel density function (uniform kernel), a Gaussian kernel density function (Gaussian kernel), a triangular core density function (triangular core), a biweight core density function (biweight core), a tricube core density function (tricube core), a triponderal core density function (triponderal core), cosine kernel functions (cosine kernel), an EpanechnikovAtorney Docket No. 206979-713601kernel density function (Epanechnikov kernel), a normal kernel density function (normal kernel), or a combination thereof. In some cases, a read density is a density estimate derived from an appropriate probability density function. In some embodiments, the density estimate is the construction of an estimate, based on observed data, of an underlying probability density function. In some embodiments, a read density comprises a density estimate (e.g., a probability density estimate, a kernel density estimate). In some cases, a density estimate often comprises a core density estimate. In some embodiments, a read density is an estimate of core density, determined based on a core density function. A read density is often generated according to a process that comprises generating a density estimate for each of the at least one portions of a genome where each portion comprises counts of sequence reads. A read density is often generated for normalized and / or weighted counts mapped to a slice. In some embodiments, each read mapped to a chunk often contributes to a read density, a value (e.g., a read count) equal to its weight obtained from a normalization process.
[0043] In some embodiments a read density profile is determined. The term “read density profile,” as used herein, refers to a product of a mathematical and / or statistical manipulation of read densities that can facilitate the identification of patterns and / or correlations in large quantities of sequence read data. In some embodiments, a read density profile comprises at least one read density, and often comprises two or more read densities (e.g. , a read density profile often comprises multiple read densities). In some embodiments, a read density profile comprises an appropriate quantitative value (e.g., a mean, median, Z score, or the like). A read density profile often comprises values resulting from at least one read densities. A read density profile sometimes comprises values resulting from at least one manipulations of read densities based on at least one adjustments (e.g., normalizations). In some embodiments, a read density profile comprises unmanipulated read densities. In some embodiments, at least one read density profiles are generated from various aspects of a data set comprising read densities, or a derivation thereof (e.g., product of at least one data processing steps). In some embodiments, a read density profile comprises normalized read densities. In some embodiments, a read density profile comprises adjusted read densities. In some embodiments, a read density profile comprises raw (e.g., unmanipulated, unadjusted, or normalized) read densities, normalized read densities, weighted read densities, filtered slice read densities, density z scores of reading densities, p values of reading densities, integral values of reading densities (e.g., area under the curve), average, mean or median reading densities, principal components, the like, or combinations thereof. In some cases, the read densities of a read density profile and / or a read density profile are associated with an uncertainty measure (e.g., a DAM). In some embodiments, a read density profile is determined for a reference (e.g., a reference sample, a training set). A read density profile for a reference is sometimes referredAtorney Docket No. 206979-713601to herein as a reference profile. In some embodiments, a reference profile comprises read densities obtained from at least one references (e.g., reference sequences, reference samples). In some embodiments, a reference profile comprises read densities determined for at least one (e.g, a set of) known euploid samples. In some embodiments, a reference profile comprises read densities of filtered portions. In some embodiments, a reference profile comprises adjusted read densities based on the at least one principal components. In some embodiments, a read density profile comprises a median distribution of read densities. In some embodiments, a read density profile comprises a ratio (e.g, a fitted ratio, a regression, or the like) of a plurality of read densities. For example, sometimes a read density profile comprises a relationship between read densities (e.g., read densities value) and genomic locations (e.g., chunks, chunk locations). In some embodiments, a read density profile is generated using a static window procedure. In some embodiments, a read density profile is generated using a sliding window procedure.
[0044] In some embodiments, a read density profile comprises multiple data points, where each data point represents a quantitative value of at least one read densities. Any suitable number of data points can be included in a read density profile depending on the nature and / or complexity of a data set. In some embodiments, read density profiles may include 2 or more data points, 3 or more data points, 5 or more data points, 10 or more data points, 24 or more data points, 25 or more data points, 50 or more data points, 100 or more data points, 500 or more data points, 1000 or more data points, 5000 or more data points, 10,000 or more data points, 100,000 or more points of data or 1,000,000 or more data points. In some embodiments, a data point is a quantitative value and / or estimate of counts of sequence reads mapped to or associated with at least one slices. In some embodiments, a data point in a read density profile comprises the results of a data manipulation of counts mapped to at least one slices. In some embodiments, a data point is often a quantitative value and / or estimate of at least one read densities (e.g., an average read density). A read density profile often comprises multiple read densities associated and / or mapped to multiple portions of a reference genome. In some embodiments, a read density profile comprises read densities of 2 to about 1,000,000 slices. In some embodiments, read densities of 2 to about 500,000, 2 to about 100,000, 2 to about 50,000, 2 to about 40,000, 2 to about 30,000, 2 to about 20,000, 2 to about 10,000, from 2 to about 5000, from 2 to about 2500, from 2 to about 1250, from 2 to about 1000, from 2 to about 500, from 2 to about 250, from 2 to about 100 or from 2 to about 60 servings determine a read density profile. In some embodiments, read densities of about 10 to about 50 slices determine a read density profile.
[0045] In some embodiments, the at least one signals can be calculated along a sliding window across each chromosome. As used herein, the term “sliding window” refers to a contiguous, nonoverlapping regions spreading across a chromosome. In some embodiments, the slidingAtorney Docket No. 206979-713601window approach comprises a predetermined read length and averaging the resulting read-level mappability. In some embodiments, a sliding window comprise any suitable number of bases determined by a slice length. In some embodiments, a genome, or segments thereof, is divided into a plurality of sliding windows. The at least one chromosomal regions comprise spanning regions of a genome that overlap. In some embodiments, the at least one chromosomal regions comprise spanning regions of a genome that do not overlap. In some cases, the sliding windows are placed at equal distances from each other. In some embodiments, the sliding windows are placed at different distances from each other. In some embodiments, a genome, a chromosome, or segment thereof is divided into a plurality of sliding windows, where a window gradually slides across a genome, a chromosome, or segment thereof, where each window at each increment represents a chromosomal region. A sliding window can slide across a genome or a chromosome in any suitable increment or according to any mathematically defined numerical pattern or sequence. In some embodiments, the slide windows slide across a genome, a chromosome or a segment thereof, in an increment of about 100,000 bp or less, about 50,000 bp or less, about 25,000 bp or less, about 10,000 bp or less, about 5,000 bp or less, approximately 1,000 bp or less, approximately 500 bp or less, or approximately 100 bp or less. In some cases, a sliding window comprises approximately 100,000 bp and may slide across a genome in increments of 50,000 bp. In other cases, a sliding window comprises approximately 50,000 bp and may slide across a genome in increments of 25,000 bp. In some embodiments, a sliding window comprises approximately 20,000 bp and may slide across a genome in increments of 10,000 bp. In some embodiments, a sliding window comprises approximately 10,000 bp and may slide across a genome in increments of 5,000 bp.
[0046] As used herein, the term “read depth signals” refers to the number of sequencing reads mapped to a given locus during at least one sequencing runs. In some embodiments, the read depth signal (or depth signal) is normalized over the total number of reads. The read depth can be expressed in a variety of different ways, including, but not limited to, the absolute number of reads mapped to a particular locus by a sequencer or the percentage or proportion of reads mapped to that locus. In some instances, the greater the read depth of a locus, the closer the allelic balance signal of that locus is to the true allelic balance in the original genetic sample. In some embodiments, the loci may be filtered for minimum read depth for inclusion in the read depth data. In some cases, the read depth of a particular variation, particularly when normalized to the total number of reads, is indicative of the relative copy number of that variation compared to other variations. In some instances, comparing the relative copy number of the variation to at least one benchmarks of known copy numbers (e.g., from a reference genetic code) indicates, for example, whether an amplification or deletion event has occurred on one of the chromosomal homologs (e.g., in all or at least a portion of the cells from which the genetic sample was derived). In someAtorney Docket No. 206979-713601embodiments, the relative copy number is compared to a baseline copy number. In some embodiments, the relative copy number of about: 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5.0. In some embodiments, the baseline copy number of about: 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9, or 5.0.
[0047] In some embodiments, the chromosomal regions comprise a chromosomal segment of a chromosome of interest, including, but not limited to a chromosome where a genetic variation is evaluated (for example, an aneuploidy of chromosomes 13, 18 and / or 21 or a sex chromosome). In some embodiments, a chromosomal region is not limited to a single chromosome. In some embodiments, a chromosomal region is limited to a single chromosome. In some cases, at least one chromosomal regions comprise all or part of one chromosome or all or part of two or more chromosomes. In some embodiments, the at least one chromosomal can encompass one, two, or more entire chromosomes. In some embodiments, at least one chromosomal regions can encompass attached or disjointed regions of multiple chromosomes. In some embodiments, the at least one chromosomal regions can comprise genes, gene fragments, regulatory sequences, introns, exons or the like. In some embodiments, the at least one chromosomal regions can comprise a gene, a gene fragment, a regulatory sequence, an intron, or an exons.
[0048] In some cases, the methods as described herein can be used to focus on some regions within the human genome according to the present methods in order to identify partial monosomies and partial trisomies. In some embodiments, the present methods involve analyzing sequence data in a defined chromosomal sliding window. In some embodiments, at least one estimated mosaic aneuploidy states are generated based on a matrix of euploidy and aneuploidy probabilities at each of the at least one candidate chromosomal regions. In some embodiments, a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions comprises one of 64 possible states. The 64 possible states represent the pairwise combinations of 8 maternal, and 8 paternal haplotype sets. In some embodiments, each parental sample has at least 8 possible haplotypes. In some embodiments, each parental sample has 8 possible haplotypes. In some embodiments, each parental sample comprises a maternal sample, or a paternal sample. For example, the 8 possible haplotypes from the maternal sample can be combine with the 8 possible haplotypes from the paternal family to derive 64 possible states. In some embodiments, 64 possible states can be grouped into seven clinically meaningful groups. In some embodiments, the seven clinically meaningful groups comprise euploidy, trisomy, nullsomy, maternal trisomy, monosomy, disomy, uniparental disomy, matching trisomy, unmatching trisomy, paternal trisomy, tetrasomy,Atorney Docket No. 206979-713601balanced (2:2) tetrasomy, unbalanced (3:1) tetrasomy, other aneuploidy, or any combinations thereof. In some embodiments, 64 possible states comprises 4 possible euploidy states and 60 possible aneuploidy states. In some embodiments, the 4 possible euploidy states comprise euploidy, diploidy, haploidy, or polyploidy. In some instances, the 4 out of 64 possible states can represent euploidy (two haplotypes, one from each parent), for example, AA, AB, BA, BB. In some embodiments, the 60 possible aneuploidy states comprise nullsomy, monosomy, disomy, uniparental disomy, trisomy, matching trisomy, unmatching trisomy, maternal trisomy, paternal trisomy, partial trisomy, partial monosomy, partial tetrasomy, tetrasomy, balanced (2:2) tetrasomy, unbalanced (3:1) tetrasomy, pentasomy, hexasomy, mosaic monosomy, mosaic aneuploidy, other aneuploidy, or combinations thereof. In some embodiments, the 60 possible aneuploidy states comprise mixed or partial aneuploidy. In some embodiments, the mixed or partial aneuploidy comprise unbalanced translocations, balanced translocations, Robertsonian translocations, recombinations, deletions, insertions, crossovers, or combinations thereof. In some embodiments, the at least one estimated mosaic aneuploidy states comprises an aggregate of euploidy and aneuploidy probabilities at each of the at least one candidate chromosomal regions based on a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions.
[0049] In some embodiments, the at least one new states in the chromosomal regions can be detected, for example, using circular binary segmentation techniques or by using a hidden Markov model search algorithm. In some embodiments, the methods disclosed herein introduces at least one new states to a Hidden Markov Model (HMM). In some embodiments, the at least one new states improves the function of the HMM. The improvement results in accurately modeling variable, continuous numbers of chromosomes, depending on the mosaicism percentage, as compared to a Hidden Markov Model without the at least one new states. In some embodiments, the at least one new states comprise Ax, or Bx. In some cases, the “x” in the at least one new states comprise a variable that is calculated from the estimated mosaic fraction at each chromosome with suspected mosaicism. For example, a sliding window along a chromosome can select a putative microdeletion and the chromosome dosage can be measured within the selected window (for example, a reads-per-bin distribution within any given window). The measured chromosome dosage of the putative microdeletion is compared to an expected dosage, and a value of likelihood of a microdeletion or a value of statistical significance can be determined, as further explained below. In some embodiments, the microdeletion is about 500,000 bases to about 15 million bases in length (for example, about 1 million to about 2 million bases in length, about 2 million to about 4 million bases in length, about 4 million to about 6 million bases in length, about 6 million to about 8 million bases in length, about 8 million to about 10 million bases in length, about 10Atorney Docket No. 206979-713601million to about 12 million bases in length, or about 12 million bases to about 15 million bases in length). In some embodiments, the microdeletion is more than about 15 million bases in length.
[0050] In some embodiments, methods described herein result in higher confidence calls of mosaicism, in particular, in ambiguous regions where standard discrete-state models produce mixed probability outputs. Said methods also provide improved calling where mosaicism produces intermediate or ambiguous signals that can be misinterpreted by discrete- state models.
[0051] Provided herein are methods for detecting at least one mosaic aneuploidies in a sample. In some embodiments, the sample is an embryonic sample. In some cases, the embryonic sample is obtained directly from an organism or from a biological sample obtained from an organism. In some embodiments, the embryonic sample is derived from blood, serum, plasma, urine, cerebrospinal fluid, saliva, stool, lymph fluid, synovial fluid, cystic fluid, ascites, pleural effusion, amniotic fluid, chorionic villus sample, fluid from a preimplantation embryo, a placental sample, lavage and cervical vaginal fluid, interstitial fluid, a buccal swab sample, sputum, bronchial lavage, a Pap smear sample, or ocular fluid. In some embodiments, the embryonic sample is derived from tissue or body fluid. In some embodiments, the embryonic sample is isolated from a cell culture, for example a primary cell culture. In some embodiments, the embryonic sample comprises total nucleic acids extracted from a bodily fluid, a tissue, a cell culture derived from an embryonic cell or an embryonic tissue. A sample can also be total RNA extracted from a biological specimen, a cDNA library, viral, or genomic DNA. A sample may also be isolated DNA from a non-cellular origin.
[0052] In some embodiments, the sample can be from a subject or reference and sometimes is an aliquot from a subject or reference. In some cases, a sample sometimes comprises nucleic acid from a subject or reference. Nucleic acid utilized in methods described herein often is obtained and isolated from a subject or reference. A subject or reference can be any living or non-living source, including, but not limited to, a human, an animal, a plant, a bacterium, a fungus, a protist. Any human or animal can be selected, including but not limited, non-human, mammal, reptile, cattle, cat, dog, goat, swine, pig, monkey, ape, gorilla, bull, cow, bear, horse, sheep, poultry, mouse, rat, fish, dolphin, whale, and shark. In some embodiments, a sample sometimes comprises nucleic acid isolated from any type of suitable biological specimen. Example of specimens can be fluid or tissue from a subject, including, without limitation, umbilical cord blood, chorionic villi, amniotic fluid, cerbrospinal fluid, spinal fluid, lavage fluid (e.g., bronchoalveolar, gastric, peritoneal, ductal, ear, athroscopic), biopsy sample (e.g., from pre-implantation embryo), celocentesis sample, fetal nucleated cells or fetal cellular remnants, washings of female reproductive tract, urine, feces, sputum, saliva, nasal mucous, prostate fluid, lavage, semen, lymphatic fluid, bile, tears, sweat, breast milk, breast fluid, embryonic cells and fetal cells (e.g.1Atorney Docket No. 206979-713601placental cells). In some embodiments, a biological sample may be blood, and sometimes plasma. As used herein, the term “blood” encompasses whole blood or any fractions of blood, such as serum and plasma as conventionally defined. Blood plasma refers to the fraction of whole blood resulting from centrifugation of blood treated with anticoagulants. Blood serum refers to the watery portion of fluid remaining after a blood sample has coagulated. Fluid or tissue samples often are collected in accordance with standard protocols hospitals or clinics generally follow. For blood, an appropriate amount of peripheral blood (e.g., between 3-40 milliliters) often is collected and can be stored according to standard procedures prior to further preparation in such embodiments. A fluid or tissue sample from which nucleic acid is extracted may be acellular. In some embodiments, a fluid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, the sample comprises fetal cells.
[0053] The sample may be heterogeneous, which means that more than one type of nucleic acid species is present in the sample. For example, heterogeneous nucleic acid can include, but is not limited to, (i) fetally-derived and maternally-derived nucleic acid, (ii) cancer and non-cancer nucleic acid, (iii) pathogen and host nucleic acid, and more generally, (iv) mutated and wild-type nucleic acid. A sample may be heterogeneous because more than one cell type is present, such as a fetal cell and a maternal cell, a cancer and non-cancer cell, or a pathogenic and host cell. In some embodiments, a minority nucleic acid species and a majority nucleic acid species is present. In some embodiments, fluid or tissue samples may be collected from a female at a gestational age suitable for testing, or from a female who is being assessed for possible pregnancy. Suitable gestational age may vary depending on the prenatal test being performed. In some embodiments, a pregnant female subject sometimes is in the first trimester of pregnancy, at times in the second trimester of pregnancy, or sometimes in the third trimester of pregnancy. In certain embodiments, a fluid or tissue is collected from a pregnant woman at 1-4, 4-8, 8-12, 12-16, 16-20, 20-24, 24-28, 28-32, 32-36, 36-40, or 40-44 weeks of fetal gestation, and sometimes between 5-28 weeks of fetal gestation.
[0054] Provided herein are methods for detecting at least one mosaic aneuploidies in an embryonic sample. In some embodiments, the method for detecting at least one mosaic aneuploidies in the embryonic sample comprises obtaining an embryonic chromosome data file from the embryonic sample. In some embodiments, an embryonic chromosome data file, as disclosed herein, is a high-coverage bam file. In some embodiments, an embryonic chromosome data file, as disclosed herein, is a high-coverage sam file. In some embodiments, an embryonic chromosome data file, as disclosed herein, is a high-coverage cram file. In some embodiments, an embryonic chromosome data file, as disclosed herein, is a low-coverage bam file. In some embodiments, an embryonic chromosome data file, as disclosed herein, is a low-coverage sam file. In some embodiments, anAtorney Docket No. 206979-713601embryonic chromosome data file, as disclosed herein, is a low-coverage cram file. Im some embodiments, the embryonic chromosome data file has low read coverage with a sequencing depth of less than O.lx, less than lx, or less than 5 .
[0055] In some embodiments, an embryonic chromosome data file, as disclosed herein, comprises at least one mutations. In some embodiments, the at least one mutations are selected from the group consisting of an aneuploidy, a deletion, a duplication, an unbalanced translocation, a haploid, a polyploidy, and / or a mosaicism. In some embodiments, the at least one mutations are located in a region of nucleotides 1-48129895 in chromosome 21, nucleotides 1-78077248 of chromosome 18, nucleotides 1-115169878 of chromosome 13, nucleotides 23810184-28525505 of chromosome 15, nucleotides 22816713-28530182 of chromosome 15, nucleotides 18892575-21460220 of chromosome 22, nucleotides 151736-11411700 of chromosome 5, nucleotides 823964-6828363 of chromosome 1, nucleotides 19040000-21470000 of chromosome 22, nucleotides 1-155270560 of chromosome X, nucleotides 1-59373566 of chromosome Y, between chromosomes 13 and 14, or between chromosomes 14 and 21 in a human genome. In some embodiments, the human genome is Genome Reference Consortium Human Build 37 (GRCh37), Genome Reference Consortium Human Build 38 (GRCh38) or Telomere 2 telomere (T2T). In some embodiments, the at least one mutations are associated with Trisomy 21 (Down syndrome), Trisomy 18 (Edwards Syndrome), Trisomy 13, 15qll-ql3 / DEL (Prader-Willi syndrome), 15ql 1-ql 3 DEL (Angelman syndrome), 22del (DiGeorge syndrome), 5p deletion syndrome (Cru-di-chat syndrome), Monosomy lp36, 22qll.2del (Di George syndrome), 22qll.2dup, Monosomy X (Turner syndrome), XXX syndrome (Jacob’s syndrome), XYY syndrome, XXY syndrome (Klinefelter syndrome), or Robertsonian translocation.
[0056] Provided herein are methods for detecting at least one mosaic aneuploidies in an embryonic sample comprising obtaining a parental chromosome data file. In some embodiments, the parental chromosome data file comprises a phase chromosome data file. In some embodiments, parental chromosome data file comprises a maternal chromosome data file, a paternal chromosome data file, or a combination thereof. In some embodiments, parental chromosome data file is a high-coverage bam file. In some embodiments, parental chromosome data file is a high-coverage sam file. In some embodiments, parental chromosome data file is a high-coverage cram file. In some embodiments, parental chromosome data file is a low-coverage bam file. In some embodiments, parental chromosome data file is a low-coverage sam file. In some embodiments, parental chromosome data file is a low-coverage cram file. Im some embodiments, the parental chromosome data file has low read coverage with a sequencing depth of less than O.lx, less than 1X, or less than 5x.Atorney Docket No. 206979-713601
[0057] In some embodiments, a parental chromosome data file comprises at least one mutations. In some embodiments, the parental chromosome data file is a phased parental chromosome data file. In some embodiments, the at least one mutations are selected from the group consisting of an aneuploidy, a deletion, a duplication, an unbalanced translocation, a haploid, a polyploidy, and / or a mosaicism. In some embodiments, the at least one mutations are located in a region of nucleotides 1-48129895 in chromosome 21, nucleotides 1-78077248 of chromosome 18, nucleotides 1-115169878 of chromosome 13, nucleotides 23810184-28525505 of chromosome 15, nucleotides 22816713-28530182 of chromosome 15, nucleotides 18892575-21460220 of chromosome 22, nucleotides 151736-11411700 of chromosome 5, nucleotides 823964-6828363 of chromosome 1, nucleotides 19040000-21470000 of chromosome 22, nucleotides 1-155270560 of chromosome X, nucleotides 1-59373566 of chromosome Y, between chromosomes 13 and 14, or between chromosomes 14 and 21 in a human genome. In some embodiments, the human genome is Genome Reference Consortium Human Build 37 (GRCh37), Genome Reference Consortium Human Build 38 (GRCh38) or Telomere 2 telomere (T2T). In some embodiments, the at least one mutations are associated with Trisomy 21 (Down syndrome), Trisomy 18 (Edwards Syndrome), Trisomy 13, 15ql l-ql3 / DEL (Prader-Willi syndrome), 15qll-ql3 DEL (Angelman syndrome), 22del (DiGeorge syndrome), 5p deletion syndrome (Cru-di-chat syndrome), Monosomy lp36, 22qll.2del (Di George syndrome), 22qll.2dup, Monosomy X (Turner syndrome), XXX syndrome (Jacob’s syndrome), XYY syndrome, XXY syndrome (Klinefelter syndrome), or Robertsonian translocation.
[0058] In some embodiments, parental chromosome data files are phased to generate phased parental chromosome data files. In some embodiments, parental chromosome data files are phased using Glimpse, shapeit5, duohmm, or a combination thereof to generate phased parental chromosome data files. In some embodiments, parental chromosome data files are phased using Glimpse to generate phased parental chromosome data files, wherein the parental chromosome data files are low-coverage bam, sam or cram file. In some embodiments, parental chromosome data files are phased by: imputating with GLIMPSE2 for low-coverage samples, and then phasing with SHAPEIT that is pedigree-aware or with a reference population.
[0059] In some embodiments, a maternal chromosome data file is phased to generate a phased maternal chromosome data file. In some embodiments, phased maternal chromosome data files are obtained based on sequence alignment files of maternal chromosome information. In some embodiments, sequence alignment files of maternal chromosome information is in SAM, BAM, or CRAM format. In some embodiments, phased maternal chromosome data files are obtained based on a population-based phasing method and / or a molecular based phasing method. In some embodiments, phased maternal chromosome data files are obtained from a benchmark genome. InAtorney Docket No. 206979-713601some embodiments, phased maternal chromosome data files are obtained using at least one phasing software. In some embodiments, the at least one phasing software comprise Glimpse, shapeit5 and duohmm.
[0060] In some embodiments, a paternal chromosome data file is phased to generate a phased paternal chromosome data file. In some embodiments, phased paternal chromosome data files are obtained based on sequence alignment files of paternal chromosome information. In some embodiments, sequence alignment files of paternal chromosome information is in SAM, BAM, or CRAM format. In some embodiments, phased paternal chromosome data files are obtained based on a population-based phasing method and / or a molecular based phasing method. In some embodiments, phased paternal chromosome data files are obtained from a benchmark genome. In some embodiments, phased paternal chromosome data files are obtained using at least one phasing software. In some embodiments, the at least one phasing software comprise Glimpse, shapeit5 and duohmm.
[0061] The term “allele balance signal” as used herein refers to a fraction of alleles inherited from the mother relative to the fraction of alleles inherited from the father. In some cases, an allele balance signal is generated from a phase parental chromosome data file. In other cases, an allele balance signal is generated from an embryonic chromosome data file. In some embodiments, an allele balance signal and the plurality of reads can originate from the same genetic material sample. In some embodiments, the genetic material sample comprise a body fluid sample (e.g., blood sample, saliva sample) or a tissue biopsy sample. The allelic balance signal and the plurality of reads may originate from the same cell population. The allelic balance signal may be derived from cell-free DNA, and the plurality of reads are derived from cellular DNA. The cellular DNA may be from cells found in body fluids (e.g., blood or saliva).
[0062] The reference genetic code may be derived from sequencing used to generate an allelic balance signal. The reference genetic code may be derived, at least in part, from sequencing normal tissue in a subject for which the allelic balance signal is obtained; derived from sequencing of germline tissue in the subject; or from sequencing genetic material from at least one genetic relatives of the subject. The at least one relatives may be the mother and / or father of the subject. The reference genetic code may be derived, at least in part, from sequencing the germline of the at least one genetic relatives.
[0063] The reference genetic code may be derived at least in part from whole genome shotgun sequencing of the subject. The allele balance signal may be derived from the whole genome shotgun sequencing. In either case, whole genome shotgun sequencing can be performed on cell-free DNA in a bodily fluid sample (e.g., a blood sample or saliva sample). Non-error propagation techniques may require single cell sequencing. The method may further entail collecting a sampleAtorney Docket No. 206979-713601of genetic material from which the allelic balance signal is obtained and / or collecting a sample of genetic material from which the plurality of reads is obtained.
[0064] Correcting allele balance data may require correcting conversion errors in the reference genetic code that have been at least partially phased. The allelic balance signal may be averaged over a plurality of binned variations over a region of about, at least about, or no greater than about 50,000, 100,000, 200,000, 300,000, 400,000, 500,000, 750,000, 1,000,000, 50,000,000, or 100,000,000 bp. The allele balance signal may be averaged over at least one haplotype blocks. The at least one haplotype blocks can be determined by dilution pool sequencing. The allelic balance signal may result from the same sequencing used to determine the at least one haplotype blocks. In some cases, the allele-balancing signal is filtered for a minimum read depth (e.g., a minimum read depth of 5, 10, 15, 20, or 25 reads).
[0065] Provided herein are methods for identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file. As used herein, “at least one candidate chromosomal regions” can comprise any genomic region from which genetic information is desired. In some embodiments, the at least one candidate chromosomal regions can comprise a segment of a chromosome. In some embodiments, the at least one candidate chromosomal regions can comprise a whole chromosome. In some cases, the chromosome can be a diploid chromosome. In a human genome, for example, the diploid chromosome can be any of chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23. In some cases, the chromosome can be an X or Y chromosome. In some cases, the at least one candidate chromosomal regions comprises a portion of a chromosome. In some embodiments, the at least one candidate chromosomal regions can be of any length. The at least one candidate chromosomal regions can have a length that is between, e.g, about 1 to about 10 bases, about 5 to about 50 bases, about 10 to about 100 bases, about 70 to about 300 bases, about 200 bases to about 1000 bases (1 kb), about 700 bases to about 2000 bases, about 1 kb to about 10 kb, about 5 kb to about 50 kb, about 20 kb to about 100 kb, about 50 kb to about 500 kb, about 100 kb to about 2000 kb (2 Mb), about 1 Mb to about 50 Mb, about 10 Mb to about 100 Mb, about 50 Mb to about 300 Mb. For example, the at least one candidate chromosomal regions can be over 1 base, over 10 bases, over 20 bases, over 50 bases, over 100 bases, over 200 bases, over 400 bases, over 600 bases, over 800 bases, over 1000 bases (1 kb), over 1.5 kb, over 2 kb, over 3 kb, over 4 kb, over 5 kb, over 10 kb, over 20 kb, over 30 kb, over 40 kb, over 50 kb, over 60 kb, over 70 kb, over 80 kb, over 90 kb, over 100 kb, over 200 kb, over 300 kb, over 400 kb, over 500 kb, over 600 kb, over 700 kb, over 800 kb, over 900 kb, over 1000 kb (1 Mb), over 2 Mb, over 3 Mb, over 4 Mb, over 5 Mb, over 6 Mb, over 7 Mb, over 8 Mb, over 9 Mb, over 10 Mb, over 20 Mb, over 30 Mb, over 40 Mb, over 50 Mb, over 60 Mb, over 70 Mb, over 80 Mb, over 90 Mb, over 100 Mb, or over 200 Mb. The atAtorney Docket No. 206979-713601least one candidate chromosomal regions can comprise at least one informative loci. An informative locus can be a polymorphic locus, e.g., comprising two or more alleles. In some cases, the two or more alleles comprise a minor allele. In some embodiments, because methods described herein require candidate mosaic regions to be gated by both depth and allele balance deviation thereby improving discrimination of mosaicism from noise.
[0066] In some embodiments, the method of the disclosure further comprises implementing at least one statistical model. In some embodiments, the at least one statistical model can include, but is not limited to, Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, or neural network methodologies. In some embodiments, the method comprises utilizing a t-test and / or a jackknife resampling method to determine whether the normalized read depth signals and the allele balance signals in each bin deviate from the expected integer chromosome states. In some embodiments, the use of at least one statistical model in methods described herein are not arbitrary. Said use can improve robustness to noise and improve region calls relative to discrete state ploidy calling.
[0067] In some embodiments, the ploidy state of a chromosome or chromosome segment is determined from a reference genetic code. In some embodiments, the reference genetic code corresponds to the entire genome of the subject, at least one entire chromosomes of the subject, or at least one chromosome segments of the subject (on the same or different chromosomes). In some embodiments, the reference genetic code is obtained directly or indirectly from a subject whose genetic material is analyzed according to the methods disclosed herein. For example, in some cases, the reference genetic code results from sequencing normal genetic material (e.g., normal cells) from a subject. In some embodiments, the normal genetic material may be genetic material known as an euploid or aneuploidy having known properties previously identified. In some embodiments, the reference genetic code is obtained from somatic and / or germ line cell sequencing of the subject. In some cases, the reference genetic code is obtained by reconstructing the genetic code from sequencing of at least one parents or other genetic relatives of the subject whose genetic material is being analyzed according to routine methods, particularly if the subject is an embryo or fetus. In some embodiments, the construction of the reference genetic code may involve sampling somatic tissue and / or germline tissue of at least one genetic relatives. In some embodiments, constructing the reference genetic code may involve sampling a subject (e.g, embryo or fetus), even if only sparse genetic information is obtained. In some embodiments, the constructing the reference genetic code involves sequencing cells obtained from the subject. In some embodiments, constructing the reference genetic code involves sequencing cell-free DNA (cfDNA), for example by sampling DNA fragments in the subject's blood, in cell culture mediumAtorney Docket No. 206979-713601(in the case of embryos) or in the subject's maternal blood (in the case of fetuses). In some embodiments, the genome of the subject, or at least the genome of normal cells of the subject, as a reference genetic code, is be compared to determine a ploidy state (e.g., abnormal cells, such as tumor cells). In some embodiments, the subject's simulated or expected genome (i.e., a genome consisting of a particular chromosome inherited from the subject's parent, without any de novo change in ploidy status, e.g, from head amplification or deletion event) serves as a reference genetic code that can be compared to determine de novo change in ploidy status in the subject.
[0068] In some embodiments, the reference genetic code may not be phased. In some embodiments, the reference genetic code is fully phased or at least partially phased. In some embodiments, the reference genetic code may be phased by as error propagation phasing methods. In some embodiments, the genetic code is phased by computational techniques involving reference population groups. In some embodiments, the genetic code is phased by molecular techniques, such as dilution pool sequencing. In some embodiments, the genetic code is phased by sequencing germ line cells of the subject and / or at least one genetic relatives e.g., mother and father) of the subject.
[0069] In some embodiments, methods provided herein further comprise the step of embryo selection or exclusion based on results from aneuploidy detection using methods described herein. Such methods can be useful in the context of embryo implantation, such as In Vitro Fertilization (IVF).Systems
[0070] Disclosed herein is a system comprising an apparatus and software instructions, that when executed by the apparatus, cause the apparatus to perform the methods described herein. Specifically, in some embodiments, a system disclosed herein comprises an apparatus and software instruction that, when executed by the apparatus, cause the apparatus to perform a method for detecting at least one mosaic aneuploidies in an embryonic sample, wherein said method comprises: obtaining an embryonic chromosome data file from the embryonic sample; obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file; estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and implementing a Hidden Markov Model (HMM), wherein the HMM comprises at least one estimated mosaic aneuploidy states that correlates to the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample.Atorney Docket No. 206979-713601Apparatus
[0071] Disclosed herein are apparatuses for generating a simulated embryonic genetic profile. In some embodiments, an apparatus comprises a processor and a memory storing software instructions that, when executed by the processor, cause the apparatus to perform the methods described herein. Specifically, in some embodiments, an apparatus disclosed herein apparatus for detecting at least one mosaic aneuploidies in an embryonic sample, the apparatus comprising a processor and a memory storing software instructions that, when executed by the processor, cause the apparatus to perform the method disclosed herein. Alternatively, in some embodiments, an apparatus for performing a method for detecting at least one mosaic aneuploidies in an embryonic sample, the apparatus comprising: a processor; a memory for receiving a plurality of data files and for storing software instruction that, when executed by the processor, cause the apparatus to perform the method using the plurality of data files, wherein said plurality of data files comprise an embryonic chromosome data file, a phased parental chromosome data file, and wherein said method comprises: obtaining an embryonic chromosome data file from the embryonic sample; obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file; estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and implementing a Hidden Markov Model (HMM), wherein the HMM comprises at least one estimated mosaic aneuploidy states that correlates to the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample. Alternatively, provided herein are computer program products for detecting at least one mosaic aneuploidies in an embryonic sample, the computer program product comprising at least one non-transitory computer-readable storage medium storing software instructions that, when executed by an apparatus, cause the apparatus to perform a method. In some embodiments, the method comprises: obtaining an embryonic chromosome data file from the embryonic sample; obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file; estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and implementing a Hidden Markov Model (HMM), wherein the HMM comprises at least one estimated mosaic aneuploidy states that correlates to the estimatedAtorney Docket No. 206979-713601mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample.EXAMPLES
[0072] The following examples are included for illustrative purposes only and are not intended to limit the scope of the invention.Example 1: HMM aneuploidy detection without support for mosaicism
[0073] This example describes methods for detecting mosaic aneuploidy in an embryonic sample using an unmodified HMM model. Next generation sequencing (NGS) was performed on samples derived from an embryonic sample S13, a maternal sample S85, and a paternal sample S73. The genomic DNA of the embryo sample was amplified using Whole Genome Amplification (WGA) method, followed by NGS library preparation and sequencing on Illumina Novaseq platform. The genomic DNA of parental samples were extracted and resulting DNA underwent NGS library prep and genome sequencing. The fastq files obtained from Illumina platform were used for data processing and genomic alignments. The sequence read counts from the embryonic sample was used as the embryonic chromosome data file. The results from the maternal sample were used to generate a phased maternal chromosome data file. Similarly, the results from the paternal sample were used to generate a phased paternal chromosome data file. Both the phased maternal chromosome data file and the phased paternal chromosome data file were used to generate a phased parental chromosome data file. The allele balance signal correlates to the fraction of alleles inherited from the mother relative to the fraction of alleles inherited from the father.
[0074] Using the NGS data generated, two signals were calculated along sliding windows across each chromosome. The first signal calculated the overall read-depth, or coverage. The second signal calculated was the allele balance, which correlates to the fraction of alleles inherited from the mother vs the fraction of alleles inherited from the father. Both signals had defined expectations for euploidy (2 chromosomes, 1 from each parent), and for simple forms of aneuploidy, such as monosomy (1 chromosome from 1 parent, 0 from the other parent) or trisomy (1 chromosome from 1 parent, 2 from the other parent).
[0075] The read-depth signal was then normalized, and was directly proportional to the chromosome number (euploidy: 2, monosomy: 1, trisomy: 3), while the allele balance signal is the ratio of maternally to paternally inherited alleles (euploidy: 1:1, monosomy: 1:0 or 0:1, trisomy: 2:1 or 1:2). FIG. 1 shows HMM aneuploidy detection without support for mosaicism. The first row shows average read depth along the chromosome. The second row shows allele balance. The third row (largest panel) shows the euploidy or aneuploidy probabilities inferred by the HMM.Atorney Docket No. 206979-713601The first half of the chromosome is predicted to be euploid (blue line). The second half of the chromosome is uncertain, as the HMM assigns similar probabilities to euploidy and to aneuploidy. As shown, the model was unable to determine mosaicism, and instead resulted as mixtures of euploid cells and aneuploid cells, resulting in intermediate read-depth and allele balance signalsExample 2: Method for detecting mosaic aneuploidy in an embryonic sample using HMM with support for mosaicism
[0076] This example describes methods for detecting mosaic aneuploidy in an embryonic sample using an HMM model comprising at least one estimated mosaic aneuploidy states. As described in Example 1, NGS was performed on samples derived from an embryonic sample, a maternal sample, and a paternal sample, and the two signals were calculated as previously described.
[0077] Each chromosome was then divided into a fixed, small number of bins to identify regions suggestive of mosaic aneuploidy. Then, each bin was evaluated to determine whether the read depth and allele balance signals deviate from the expectations under euploidy, monosomy, and trisomy (integer chromosome states). The bins in which both the read-depth and the allele balance signals deviated significantly from these three integer chromosome states were flagged as candidate regions for mosaic aneuploidy. To correct for false positive mosaic aneuploidy regions, at least two adjacent bins were required to be flagged in order to be characterized as a mosaic aneuploidy region. To evaluate deviations from the three integer state chromosome states in each bin, an adaptation of regular t-tests combining it with the jackknife resampling method was used. This adaptation accounted for the non-independence of read depth and allele balance signal along adjacent sites in the genome, which would otherwise (in regular t-tests) lead to inflated test statistics. After regions with possible mosaic aneuploidies were identified, the average read depth in each of those regions was utilized to estimate the mosaicism percentage (m) (which represents the fraction of aneuploid cells). The mosaicism percentage was also estimated from the normalized read depth, d. For example, regions where d > 2, it was assumed that the sample contained a mix of euploid cells (d = 2) and cells with trisomy (d = 3). Therefore, the m was estimated as m = d-2. In regions where d < 2, it was assumed that the sample contained a mix of euploid cells (d = 2) and cells with monosomy (d = 1). In that case, m was estimated as m = 2 - d.
[0078] Using the mosaicism percentage estimated, at least one mosaic aneuploidy HMM states were created for the estimated mosaic regions, and these newly created states . In cases with multiple mosaic regions per chromosome we limit this to the largest mosaic region. While the regular HMM states as described in Example 1, assume that each parent contributed either 1, 2, or 3 chromosomes. The at least one mosaic aneuploidy states added to the HMM allows for the possibility that a single parent contributed a continuous number of chromosomes in the rangeAtorney Docket No. 206979-713601from 1 to 3, and this was estimated as d - 1. As shown in FIG. 2, the resulting HMM results now assign a high probability of mosaic monosomy to the second half of the chromosome. The were 8 possible haplotype sets for each parent as shown in the bottom right legend of FIG. 2. There was one haplotype set representing 0 haplotypes inherited from one parent ('0'), two haplotype sets representing 1 haplotype inherited from one parent (A, or B), three haplotype sets representing 2 haplotypes inherited from one parent (A2, B2, A+B), and 2 haplotype sets representing a variable number of haplotypes inherited from one parent (Ax, Bx; those are the added mosaic states). FIG.2 shows HMM aneuploidy detection with support for mosaicism: The shaded regions indicate significant deviations from the expectation under euploidy for depth (expectation = 2) and allele balance (expectation = 1:1). The averaged and normalized allele depth in the second half of the chromosome is 1.17, which is consistent with a mosaicism percentage of 2 - 1.17 = 83%, assuming a mix of cells with euploidy and monosomy. New states accommodating the estimated mosaicism percentage were added to the HMM.Example 3: Accurate prediction of the level of mosaicism of chromosome 21 aneuploidy from synthetic mosaic samples using HMM
[0079] This example describes the capability of the method disclosed herein utilizing the HMM comprising at least one estimated mosaic aneuploidy states for prediction of aneuploidy mosaicism level in chromosome 21 (FIG. 3). To evaluate mosaicism detection accuracy, synthetic genomic DNA of trisomy 21 cell samples with different levels of mosaicism (0% normal control, 10%, 20%, 30%, 50%, 70%, 90% mosaic) were constructed in the lab. The synthetic DNA samples underwent whole-genome amplification and next generation sequencing and each percentage of mosaic sample was tested in triplicates. The data of each sample with known percentage of mosaicism (for example 90% trisomy 21) was analyzed using the methods disclosed herein. All other chromosomes except for chromosome 21 had the baseline copy number of 2. At 10% mosaic level, the predicted copy number was about 2.1 which was distinguishable from the baseline. Increasing mosaic level such as 90% is equivalent to copy number value of 2.9. This is different from Example 2 that picked up monosomy mosaicism. Overall, this example shows the accuracy of methods disclosed herein in the detection of mosaicism for aneuploidy trisomy at different mosaic levels. Based on these results, above 10% mosaic can be accurately called using this model.
[0080] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments described herein can be employed. It is intended that the following claims defineAtorney Docket No. 206979-713601the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.ILLUSTRATIVE EMBODIMENTS
[0081] Embodiment 1. A method for detecting at least one mosaic aneuploidies in an embryonic sample, the method comprising:
[0082] (a) obtaining an embryonic chromosome data file from the embryonic sample;
[0083] (b) obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file;
[0084] (c) identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file;
[0085] (d) estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and
[0086] (e) implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that reflects the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample.
[0087] Embodiment 2. The method of Embodiment 1, wherein the at least one estimated mosaic aneuploidy states are added to the at least one statistical model.
[0088] Embodiment 3. The method of Embodiment 1 or 2, wherein the at least one statistical model is capable of generating a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions.
[0089] Embodiment 4. The method of any one of Embodiments 1-3, wherein the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, or neural network methodologies.
[0090] Embodiment 5. The method of any one of Embodiments 1-4, wherein the at least one statistical model is an HMM.
[0091] Embodiment 6. The method of any one of Embodiments 1-5, wherein the identifying of the at least one candidate chromosomal regions for mosaic aneuploidy is determined based on the phased parental chromosome data file, and / or sequence read counts comprised in the embryonic chromosome data file.
[0092] Embodiment 7. The method of any one of Embodiments 1-6, wherein the estimated mosaicism percentage is estimated based on sequence read counts at each of the at least oneAtorney Docket No. 206979-713601candidate chromosomal regions comprised in the embryonic chromosome data file relative to the corresponding chromosomal region of the allele balance signal.
[0093] Embodiment 8. The method of any one of Embodiments 6 or 7, wherein the sequence read counts comprised in the embryonic chromosome data file comprises read counts for a reference allele and an alternate allele.
[0094] Embodiment 9. The method of any one of Embodiments 1-8, wherein the phased parental chromosome data file comprises a phased maternal chromosome datafile, a phased paternal chromosome data file, or a combination of both.
[0095] Embodiment 10. The method of Embodiment 9, wherein the phased maternal chromosome data file is obtained based on a sequence alignment file of maternal chromosome information and / or the phased paternal chromosome data file is obtained based on a sequence alignment file of paternal chromosome information.
[0096] Embodiment 11. The method of any one of Embodiments 9 or 10, wherein the phased maternal chromosome data file and / or the phased paternal chromosome data file is in VCF format.
[0097] Embodiment 12. The method of Embodiment 6, wherein the at least one estimated mosaic aneuploidy states are generated based on a matrix of euploidy and aneuploidy probabilities at each of the at least one candidate chromosomal regions.
[0098] Embodiment 13. The method of Embodiment 6, wherein a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions comprises one of 64 possible states.
[0099] Embodiment 14. The method of Embodiment 13, wherein the 64 possible states comprise 4 possible euploidy states and 60 possible aneuploidy states.
[0100] Embodiment 15. The method of any one of Embodiments 6-14, wherein the at least one estimated mosaic aneuploidy states comprises an aggregate of euploidy and aneuploidy probabilities at each of the at least one candidate chromosomal regions based on a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions.
[0101] Embodiment 16. The method of any one of Embodiments 1-15, wherein identifying the at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file comprises:
[0102] (a) dividing each chromosome from the embryonic chromosome data file, the phased parental chromosome data file and / or the allele balance signal into a fixed number of bins;
[0103] (b) obtaining read depth signals from the embryonic chromosome data file and the allele balance signal in each bin;
[0104] (c) determining whether the read depth signals and allele balance signals in each bin deviate from an expected integer chromosome state; andAtorney Docket No. 206979-713601
[0105] (d) identifying a bin in which both the read depth signals and the allele balance signals deviate more than a threshold from the expected integer chromosome state as a candidate chromosomal region for mosaic aneuploidy.
[0106] Embodiment 17. The method of Embodiment 16, wherein the expected integer chromosome state comprises euploidy expected integer chromosome state (2), monosomy expected integer chromosome state (1), or trisomy expected integer chromosome state (3).
[0107] Embodiment 18. The method of Embodiment 16, wherein the read depth signals from the embryonic chromosome data file is normalized.
[0108] Embodiment 19. The method of Embodiment 16, further comprising identifying at least two adjacent bins in which both the read-depth signals and the allele balance signals deviate more than a threshold from the integer chromosome states as the candidate chromosomal region for mosaic aneuploidy.
[0109] Embodiment 20. The method of any one of Embodiments 1-19, wherein the estimated mosaicism percentage at each of the at least one candidate chromosomal regions is based on a normalized and averaged read depth in each of the at least one candidate chromosomal regions and a chromosome copy number.
[0110] Embodiment 21. The method of Embodiment 20, wherein the chromosome copy number comprises a euploidy expected integer chromosome state (2).[OHl] Embodiment 22. The method of Embodiment 1-20, wherein the mosaicism percentage at each of the at least one candidate chromosomal regions is estimated according to a formula:
[0112] m = d - 2, wherein d > 2; and
[0113] m = 2 - d, wherein d < 2,
[0114] wherein the m is the estimated mosaicism percentage at each of the at least one candidate chromosomal regions, and the d is a normalized and averaged read depth in each of the at least one candidate chromosomal regions.
[0115] Embodiment 23. The method of Embodiment 22, wherein the d corresponds to a chromosome copy number at the at least one candidate chromosomal regions.
[0116] Embodiment 24. The method of Embodiment 22 or 23, wherein the d is a non-integer.
[0117] Embodiment 2. The method of any one of Embodiments 1-24, wherein the method comprises determining a probability of a mosaic aneuploidy in the embryonic sample, wherein the probability of the mosaic aneuploidy in the embryonic sample is based on a percentage of cells in the embryonic sample having an estimated mosaicism percentage above a predetermined threshold in at least one of the candidate chromosomal regions.
[0118] Embodiment 26. The method of any one of Embodiments 1-25, wherein the at least one estimated mosaic aneuploidy states reflects that one parent is considered to have contributed to anAttorney Docket No. 206979-713601embryo correlating the embryonic sample a continuous number of chromosomes d-1, wherein d is the normalized and averaged read depth in each of the at least one candidate chromosomal regions.
[0119] Embodiment 27. The method of any one of Embodiments 1-25, further comprising:
[0120] selecting an embryo correlating to the embryonic sample for implantation based on a determination using the probability of the mosaic aneuploidy in the embryonic sample.
[0121] Embodiment 28. An apparatus for detecting at least one mosaic aneuploidies in an embryonic sample, the apparatus comprising a processor and a memory storing software instructions that, when executed by the processor, cause the apparatus to perform the method of any one of Embodiments 1-27.
[0122] Embodiment 29. A computer program product for detecting at least one mosaic aneuploidies in an embryonic sample, the computer program product comprising at least one non-transitory computer-readable storage medium storing software instructions that, when executed by an apparatus, cause the apparatus to perform the method of any one of Embodiments 1-27.
[0123] Embodiment 30. An apparatus for performing a method for detecting at least one mosaic aneuploidies in an embryonic sample, the apparatus comprising:
[0124] a processor; and
[0125] a memory for receiving a plurality of data files and for storing software instruction that, when executed by the processor, cause the apparatus to perform the method using the plurality of data files,
[0126] wherein said plurality of data files comprise an embryonic chromosome data file, a phased parental chromosome data file, and
[0127] wherein said method comprises:
[0128] (a) obtaining an embryonic chromosome data file from the embryonic sample;
[0129] (b) obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file;
[0130] (c) identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file;
[0131] (d) estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and
[0132] (e) implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that correlates to the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample.Atorney Docket No. 206979-713601
[0133] Embodiment 31. The apparatus of Embodiment 30, wherein the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, or neural network methodologies.
[0134] Embodiment 32. The apparatus of Embodiment 30 or 31, wherein the at least one statistical model is an HMM.
[0135] Embodiment 33. A system comprising an apparatus and software instruction that, when executed by the apparatus, cause the apparatus to perform a method for detecting at least one mosaic aneuploidies in an embryonic sample, wherein said method comprises:
[0136] (a) obtaining an embryonic chromosome data file from the embryonic sample;
[0137] (b) obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file;
[0138] (c) identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file;
[0139] (d) estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and
[0140] (e) implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that correlates to the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample.
[0141] Embodiment 34. The system of Embodiment 33, wherein the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, or neural network methodologies.
[0142] Embodiment 35. The system of Embodiment 33 or 34, wherein the at least one statistical model is an HMM.
[0143] Embodiment 36. A system for detecting at least one mosaic aneuploidies in an embryonic sample, the system comprising:
[0144] (a) obtaining an embryonic chromosome data file from the embryonic sample;
[0145] (b) obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file;
[0146] (c) identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file;Atorney Docket No. 206979-713601
[0147] (d) estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal;
[0148] (e) implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that reflects the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample; and
[0149] (f) selecting an embryo correlating to the embryonic sample for implantation based on a determination using the probability of the mosaic aneuploidy in the embryonic sample.
[0150] Embodiment 37. The system of Embodiment 36, wherein the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, or neural network methodologies.
[0151] Embodiment 38. The system of Embodiment 36 or 37, wherein the at least one statistical model is an HMM.
Claims
Attorney Docket No. 206979-713601CLAIMSWhat is claimed is:
1. A method for detecting at least one mosaic aneuploidies in an embryonic sample, the method comprising:(a) obtaining an embryonic chromosome data file from the embryonic sample;(b) obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; (c) identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file;(d) estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and(e) implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that reflects the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample.
2. The method of claim 1, wherein the at least one estimated mosaic aneuploidy states are added to the at least one statistical model.
3. The method of claim 1 or 2, wherein the at least one statistical model is capable of generating a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions.
4. The method of any one of claims 1-3, wherein the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, neural network methodologies, or a combination thereof.
5. The method of any one of claims 1-4, wherein the at least one statistical model is an HMM.
6. The method of any one of claims 1-5, wherein the identifying of the at least one candidate chromosomal regions for mosaic aneuploidy is determined based on the phasedAtorney Docket No. 206979-713601parental chromosome data file, and / or sequence read counts comprised in the embryonic chromosome data file.
7. The method of any one of claims 1-6, wherein the estimated mosaicism percentage is estimated based on sequence read counts at each of the at least one candidate chromosomal regions comprised in the embryonic chromosome data file relative to the corresponding chromosomal region of the allele balance signal.
8. The method of any one of claims 6 or 7, wherein the sequence read counts comprised in the embryonic chromosome data file comprises read counts for a reference allele and an alternate allele.
9. The method of any one of claims 1-8, wherein the phased parental chromosome data file comprises a phased maternal chromosome datafile, a phased paternal chromosome data file, or a combination of both.
10. The method of claim 9, wherein the phased maternal chromosome data file is obtained based on a sequence alignment file of maternal chromosome information and / or the phased paternal chromosome data file is obtained based on a sequence alignment file of paternal chromosome information.
11. The method of any one of claims 9 or 10, wherein the phased maternal chromosome data file and / or the phased paternal chromosome data file is in VCF format.
12. The method of claim 6, wherein the at least one estimated mosaic aneuploidy states are generated based on a matrix of euploidy and aneuploidy probabilities at each of the at least one candidate chromosomal regions.
13. The method of claim 6, wherein a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions comprises one of 64 possible states.
14. The method of claim 13, wherein the 64 possible states comprise 4 possible euploidy states and 60 possible aneuploidy states.
15. The method of any one of claims 6-14, wherein the at least one estimated mosaic aneuploidy states comprises an aggregate of euploidy and aneuploidy probabilities at each of theAtorney Docket No. 206979-713601at least one candidate chromosomal regions based on a matrix of euploidy and aneuploidy probabilities for each of the at least one candidate chromosomal regions.
16. The method of any one of claims 1-15, wherein identifying the at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file comprises:(a) dividing each chromosome from the embryonic chromosome data file, the phased parental chromosome data file and / or the allele balance signal into a fixed number of bins;(b) obtaining read depth signals from the embryonic chromosome data file and the allele balance signal in each bin;(c) determining whether the read depth signals and allele balance signals in each bin deviate from an expected integer chromosome state; and(d) identifying a bin in which both the read depth signals and the allele balance signals deviate more than a threshold from the expected integer chromosome state as a candidate chromosomal region for mosaic aneuploidy.
17. The method of claim 16, wherein the expected integer chromosome state comprises euploidy expected integer chromosome state (2), monosomy expected integer chromosome state (1), or trisomy expected integer chromosome state (3).
18. The method of claim 16, wherein the read depth signals from the embryonic chromosome data file is normalized.
19. The method of claim 16, further comprising identifying at least two adjacent bins in which both the read depth signals and the allele balance signals deviate more than a threshold from the integer chromosome states as the candidate chromosomal region for mosaic aneuploidy.
20. The method of any one of claims 1-19, wherein the estimated mosaicism percentage at each of the at least one candidate chromosomal regions is based on a normalized and averaged read depth in each of the at least one candidate chromosomal regions and a chromosome copy number.
21. The method of claim 20, wherein the chromosome copy number comprises a euploidy expected integer chromosome state (2).Atorney Docket No. 206979-71360122. The method of claim 1-20, wherein the mosaicism percentage at each of the at least one candidate chromosomal regions is estimated according to a formula:m = d - 2, wherein d > 2; andm = 2 - d. wherein d < 2,wherein the m is the estimated mosaicism percentage at each of the at least one candidate chromosomal regions, and the is a normalized and averaged read depth in each of the at least one candidate chromosomal regions.
23. The method of claim 22, wherein the d corresponds to a chromosome copy number at the at least one candidate chromosomal regions.
24. The method of claim 22 or 23, wherein the is a non-integer.
25. The method of any one of claims 1-24, wherein the method comprises determining a probability of a mosaic aneuploidy in the embryonic sample, wherein the probability of the mosaic aneuploidy in the embryonic sample is based on a percentage of cells in the embryonic sample having an estimated mosaicism percentage above a predetermined threshold in at least one of the candidate chromosomal regions.
26. The method of any one of claims 1-25, wherein the at least one estimated mosaic aneuploidy states reflects that one parent is considered to have contributed to an embryo correlating the embryonic sample a continuous number of chromosomes < / -l, wherein is a normalized and averaged read depth in each of the at least one candidate chromosomal regions.
27. The method of any one of claims 1-25, further comprising:selecting an embryo correlating to the embryonic sample for implantation based on a determination using a probability of the mosaic aneuploidy in the embryonic sample.
28. An apparatus for detecting at least one mosaic aneuploidies in an embryonic sample, the apparatus comprising a processor and a memory storing software instructions that, when executed by the processor, cause the apparatus to perform the method of any one of claims 1-27.
29. A computer program product for detecting at least one mosaic aneuploidies in an embryonic sample, the computer program product comprising at least one non-transitory computer-readable storage medium storing software instructions that, when executed by an apparatus, cause the apparatus to perform the method of any one of claims 1-27.Attorney Docket No. 206979-71360130. An apparatus for performing a method for detecting at least one mosaic aneuploidies in an embryonic sample, the apparatus comprising:a processor; anda memory for receiving a plurality of data files and for storing software instruction that, when executed by the processor, cause the apparatus to perform the method using the plurality of data files,wherein said plurality of data files comprise an embryonic chromosome data file, a phased parental chromosome data file, andwherein said method comprises:(a) obtaining an embryonic chromosome data file from the embryonic sample;(b) obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; (c) identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file;(d) estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and(e) implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that correlates to the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample.
31. The apparatus of claim 30, wherein the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, neural network methodologies, or a combination thereof.
32. The apparatus of claim 30 or 31, wherein the at least one statistical model is an HMM.
33. A system comprising an apparatus and software instruction that, when executed by the apparatus, cause the apparatus to perform a method for detecting at least one mosaic aneuploidies in an embryonic sample, wherein said method comprises:(a) obtaining an embryonic chromosome data file from the embryonic sample;Atorney Docket No. 206979-713601(b) obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; (c) identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file;(d) estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal; and(e) implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that correlates to the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample.
34. The system of claim 33, wherein the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, neural network methodologies, or a combination thereof.
35. The system of claim 33 or 34, wherein the at least one statistical model is an HMM.
36. A system for detecting at least one mosaic aneuploidies in an embryonic sample, the system comprising:(a) obtaining an embryonic chromosome data file from the embryonic sample;(b) obtaining a phased parental chromosome data file, and generating an allele balance signal from the phased parental chromosome data file and the embryonic chromosome data file; (c) identifying at least one candidate chromosomal regions for mosaic aneuploidy in the embryonic chromosome data file;(d) estimating a mosaicism percentage at each of the at least one candidate chromosomal regions from the embryonic chromosome data file relative to a corresponding chromosomal region of the allele balance signal;(e) implementing at least one statistical model, wherein the at least one statistical model comprises at least one estimated mosaic aneuploidy states that reflects the estimated mosaicism percentage, thereby detecting the at least one mosaic aneuploidies in the embryonic sample; and(f) selecting an embryo correlating to the embryonic sample for implantation based on a determination using the probability of the mosaic aneuploidy in the embryonic sample.Attorney Docket No. 206979-71360137. The system of claim 33, wherein the at least one statistical model is a Hidden Markov Model (HMM), dynamic programming, support vector machine, Bayesian network, trellis decoding, Viterbi decoding, expectation maximization, Kalman filtering, neural network methodologies, or a combination thereof.
38. The system of claim 33 or 37, wherein the at least one statistical model is an HMM.