Application of mosaicism ratio in multifetal pregnancies and personalized risk assessment

Mosaicism ratios in circulating cell-free nucleic acids enable accurate classification of fetal genetic conditions and gender in multiple pregnancies, addressing false positives in NIPT and enhancing prenatal care.

JP2026004475APending Publication Date: 2026-01-14SEQUENOM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025166348
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-06-24
Filing Date
2025-10-02
Publication Date
2026-01-14

AI Technical Summary

Technical Problem

Existing non-invasive prenatal testing (NIPT) methods face challenges with false-positive results due to placental localized mosaicism, leading to unnecessary invasive procedures and patient reluctance, as they cannot accurately differentiate between placental and fetal genetic variations.

Method used

The use of mosaicism ratios derived from sequencing circulating cell-free nucleic acids to determine the fraction of nucleic acids with copy number variations and fetal nucleic acids, allowing classification of genetic mosaicism and providing personalized risk assessments for aneuploidies and fetal sex in multiple pregnancies.

Benefits of technology

Enables accurate prediction of fetal genetic conditions and gender through non-invasive means, reducing the need for invasive procedures and improving prenatal care by providing reliable, personalized risk assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026004475000021
    Figure 2026004475000021
  • Figure 2026004475000022
    Figure 2026004475000022
  • Figure 2026004475000023
    Figure 2026004475000023
Patent Text Reader

Abstract

Methods of identifying genetic mutations and / or genetic alterations are provided.SOLUTION: Methods for classifying the presence or absence of genetic mosaicism for a copy number variation in one or more fetuses (e.g., predicting whether a fetus or more than one fetus is affected by a copy number variation) are provided. The sample nucleic acid is subjected to a sequencing process and the resulting sequence reads are analyzed to identify genetic copy number variation regions. Genetic mosaicism for a copy number variation region is classified for a fetus or more than one fetus based on (i) a mosaicism ratio of a fraction of nucleic acid having the copy number variation region to a fraction of fetal nucleic acid and (ii) a chromosome having the genetic copy number variation region (e.g., an identified type of aneuploidy) or (ii) a number of fetuses carried by a pregnant female.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Field The technology provided herein relates in part to a technique for non-invasive classification of mosaic copy number variation (CNV) for one or more fetuses.The technology provided herein is useful for classifying mosaic CNV for samples, for example, as part of non-invasive prenatal testing (NIPT) and oncology testing. [Background technology]

[0002] background The genetic information of living organisms (e.g., animals, plants, and microorganisms) and other forms that replicate genetic information (e.g., viruses) is encoded in deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a sequence of nucleotides or modified nucleotides that exhibits a chemical or hypothetical nucleic acid primary structure. In humans, the complete genome consists of approximately 300 genes located on 24 chromosomes (i.e., 22 autosomes, 1 X chromosome, and 1 Y chromosome; see The Human Genome, T. Strachan, BIOS Scientific Publishers, 1992). The human genome contains 10,000 genes. Each gene encodes a specific protein, whose expression, via transcription and translation, then carries out a specific biochemical function within living cells.

[0003] Many medical conditions are caused by one or more genetic mutations and / or genetic alterations. Certain genetic mutations and / or genetic alterations cause medical conditions, including, for example, hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease, and cystic fibrosis (CF) (Human Genome Mutations, DN Cooper and M. Krawczak, BIOS Publishers, 1993). Such genetic diseases can result from the addition, substitution, or deletion of a single nucleotide in the DNA of a specific gene. Certain birth defects are caused by chromosomal abnormalities, also known as aneuploidies, such as trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), monosomy X (Turner syndrome), and certain sex chromosome aneuploidies, such as Klinefelter syndrome (XXY). Another genetic variation is the sex of the fetus, which can often be determined based on the sex chromosomes X and Y. Some genetic variations may predispose an individual to or cause any of several diseases, such as, for example, diabetes, arteriosclerosis, obesity, various autoimmune diseases and cell proliferative disorders, e.g., cancer, tumor, neoplasia, metastatic disease, etc., or a combination thereof. The cancer, tumor, neoplasia, or metastatic disease may be a disorder or condition of the liver, lung, spleen, pancreas, colon, skin, bladder, eye, brain, esophagus, head, neck, ovaries, testes, prostate, etc., or a combination thereof. Identifying one or more genetic mutations and / or genetic alterations (e.g., copy number alterations, copy number variations, single nucleotide alterations, single nucleotide alterations, chromosomal alterations, translocations, deletions, insertions, etc.) or differences can result in the diagnosis of a particular medical condition or the determination of a predisposition to such a medical condition. Identifying genetic differences can lead to facilitated medical decisions and / or the adoption of beneficial medical procedures. In certain embodiments, identifying one or more genetic mutations and / or genetic alterations involves the analysis of circulating cell-free nucleic acids. Circulating cell-free nucleic acids (CCF-NA), such as cell-free DNA (CCF-DNA), are composed of DNA fragments originating from cell death and circulating in peripheral blood. High concentrations of CCF-DNA can indicate certain clinical conditions, such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infection, and other diseases. Furthermore, cell-free fetal DNA (CFF-DNA) can be detected in the maternal bloodstream and used for various non-invasive prenatal diagnostic methods. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] he Human Genome, T. Strachan, BIOS Scientific Publishers, 1992 [Non-patent document 2] Human Genome Mutations, DN Cooper and M. Krawczak, BIOS Publishers, 1993 Summary of the Invention [Means for solving the problem]

[0005] overview In various embodiments, a computer-implemented method is provided that includes the steps of: identifying, by a computing device, a region of genetic copy number variation in a sample comprising circulating cell-free nucleic acid from a pregnant female subject having a multiple fetal pregnancy, wherein the region of genetic copy number variation comprises a copy number variation, and the circulating cell-free nucleic acid comprises maternal nucleic acid and fetal nucleic acid; determining, by the computing device, a fraction of nucleic acid in the circulating cell-free nucleic acid that has the copy number variation; determining, by the computing device, a fraction of fetal nucleic acid in the circulating cell-free nucleic acid; generating, by the computing device, a mosaicism ratio, wherein the mosaicism ratio is the fraction of nucleic acid in the circulating cell-free nucleic acid that has the copy number variation divided by the fraction of fetal nucleic acid in the circulating cell-free nucleic acid; and classifying, by the computing device, the presence or absence of genetic mosaicism for the region of copy number variation according to the mosaicism ratio based on the mosaicism ratio and the number of fetuses carried by the pregnant female subject.

[0006] In some embodiments, the fraction of nucleic acid in the circulating cell-free nucleic acid that has copy number variation is determined for regions of copy number variation.

[0007] In some embodiments, the fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is determined according to sequencing-based fraction estimation.

[0008] In some embodiments, the fraction of nucleic acids in the circulating cell-free nucleic acids that have copy number variations is determined according to the allelic ratio of the polymorphic sequence.

[0009] In some embodiments, the fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is determined according to quantification of differentially methylated nucleic acids.

[0010] In some embodiments, the fraction of nucleic acids in the circulating cell-free nucleic acid that have copy number variations is the fetal fraction determined for the region of copy number variation.

[0011] In some embodiments, the fetal fraction of nucleic acids with copy number variations in the circulating cell-free nucleic acids is determined according to sequencing-based fetal fraction estimation.

[0012] In some embodiments, the fetal fraction of nucleic acids with copy number variations in the circulating cell-free nucleic acids is determined according to the allelic ratios of polymorphic sequences in the fetal nucleic acids and the maternal nucleic acids.

[0013] In some embodiments, the fetal fraction of nucleic acids with copy number variations in the circulating cell-free nucleic acids is determined according to quantification of differentially methylated fetal and maternal nucleic acids.

[0014] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined for a genomic region that is larger than the region of copy number variation.

[0015] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined for a genomic region that is distinct from the region of copy number variation.

[0016] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined according to a sequencing-based fetal fraction estimation.

[0017] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined according to the allelic ratio of the polymorphic sequence in the fetal nucleic acid and the maternal nucleic acid.

[0018] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined according to quantification of differentially methylated fetal and maternal nucleic acids.

[0019] In some embodiments, the mosaicism ratio is the fraction of nucleic acids with copy number variations in the circulating cell-free nucleic acid divided by the fraction of fetal nucleic acid in the circulating cell-free nucleic acid.

[0020] In some embodiments, the method further comprises providing, by the computing system, no classification if the mosaicism ratio is below a minimum threshold.

[0021] In some embodiments, the minimum threshold is about 0.1.

[0022] In some embodiments, the method further comprises providing, by the computing system, no classification if the mosaicism ratio is greater than a maximum threshold.

[0023] In some embodiments, the maximum threshold is about 1.7.

[0024] In some embodiments, the method further includes obtaining, by a computing system, a positive screening result from a non-invasive prenatal test (NIPT) for the presence of one or more aneuploidies in a sample comprising circulating cell-free nucleic acids from the pregnant female subject.

[0025] In some embodiments, the method further includes providing, by the computing system, an interpretation of a positive screening result from the NIPT as a negative result, or the absence of one or more aneuploidies, if no classification is provided and the mosaicism ratio is below a minimum threshold.

[0026] In some embodiments, the method further includes providing, by the computing system, an interpretation of a positive screening result from the NIPT as excessive or indeterminate if no classification is provided and the mosaicism ratio is greater than a maximum threshold.

[0027] In some embodiments, the method further includes providing, by the computing system, an interpretation of a positive screening result from the NIPT as positive, along with a comment regarding the likelihood of mosaicism, if the presence of genetic mosaicism is classified for the copy number variation region.

[0028] In some embodiments, the method further includes providing, by the computing system, an interpretation of a positive screening result from the NIPT as positive if the absence of genetic mosaicism is classified for the region of copy number variation.

[0029] In various embodiments, a method is provided for classifying the sex of fetuses in a multiple pregnancy, the method comprising the steps of: determining, by a computing device, a fraction of nucleic acids having a Y chromosome or a region of a Y chromosome in a sample comprising circulating cell-free nucleic acid from a pregnant female subject having a multiple pregnancy, wherein the circulating cell-free nucleic acid comprises maternal nucleic acid and fetal nucleic acid; determining, by the computing device, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid; generating, by the computing device, a mosaicism ratio, wherein the mosaicism ratio is the fraction of nucleic acid having a Y chromosome or a region of a Y chromosome in the circulating cell-free nucleic acid divided by the fraction of fetal nucleic acid in the circulating cell-free nucleic acid; and classifying, by the computing device, the sex of the fetuses based on the mosaicism ratio and the number of fetuses carried by the pregnant female subject.

[0030] In some embodiments, the fraction of nucleic acids bearing a Y chromosome or regions of a Y chromosome in circulating cell-free nucleic acid is determined according to a sequencing-based fraction estimate.

[0031] In some embodiments, the fraction of nucleic acids in the circulating cell-free nucleic acids that have a Y chromosome or regions of a Y chromosome is determined according to the allelic ratio of the polymorphic sequence.

[0032] In some embodiments, the fraction of nucleic acids bearing a Y chromosome or regions of a Y chromosome in circulating cell-free nucleic acids is determined according to quantification of differentially methylated nucleic acids.

[0033] In some embodiments, the fraction of nucleic acids in the circulating cell-free nucleic acids that have a Y chromosome or a region of a Y chromosome is the determined fetal fraction for the Y chromosome or a region of a Y chromosome.

[0034] In some embodiments, the fetal fraction of nucleic acids having a Y chromosome or regions of a Y chromosome in circulating cell-free nucleic acids is determined according to sequencing-based fetal fraction estimation.

[0035] In some embodiments, the fetal fraction of nucleic acids having a Y chromosome or a region of a Y chromosome in the circulating cell-free nucleic acids is determined according to the allelic ratio of the polymorphic sequence in the fetal nucleic acid and the maternal nucleic acid.

[0036] In some embodiments, the fetal fraction of nucleic acids having a Y chromosome or regions of a Y chromosome in circulating cell-free nucleic acids is determined according to quantification of differentially methylated fetal and maternal nucleic acids.

[0037] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined for the Y chromosome or a genomic region larger than the region of the Y chromosome.

[0038] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined for the Y chromosome or a genomic region distinct from a region of the Y chromosome.

[0039] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined according to a sequencing-based fetal fraction estimation.

[0040] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined according to the allelic ratio of the polymorphic sequence in the fetal nucleic acid and the maternal nucleic acid.

[0041] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined according to quantification of differentially methylated fetal and maternal nucleic acids.

[0042] In some embodiments, the mosaicism ratio is the fraction of nucleic acid in the circulating cell-free nucleic acid that has a Y chromosome or a region of a Y chromosome divided by the fraction of fetal nucleic acid in the circulating cell-free nucleic acid.

[0043] In some embodiments, the method further comprises obtaining, by the computing system, a positive screening result from a non-invasive prenatal test (NIPT) for the presence of one or more aneuploidies in the sample.

[0044] In various embodiments, the method includes obtaining a positive screening result from a non-invasive prenatal test (NIPT) for the presence of aneuploidy in a first sample comprising circulating cell-free nucleic acid from a pregnant female subject, wherein the positive screening result comprises a type of aneuploidy detected in the first sample; identifying a genetic copy number variation region associated with aneuploidy in a second sample comprising circulating cell-free nucleic acid from the pregnant female subject, wherein the genetic copy number variation region comprises a copy number variation and the circulating cell-free nucleic acid comprises maternal nucleic acid and fetal nucleic acid; determining the fraction of nucleic acid in the circulating cell-free nucleic acid that has a copy number variation; determining the fraction of fetal nucleic acid in the circulating cell-free nucleic acid; generating a mosaicism ratio, where the mosaicism ratio is the fraction of nucleic acid with copy number variation in the circulating cell-free nucleic acid divided by the fraction of fetal nucleic acid in the circulating cell-free nucleic acid; classifying the presence or absence of genetic mosaicism for the copy number variation region based on the mosaicism ratio; and providing a personalized risk assessment for a pregnant female subject's fetus with aneuploidy based on a positive screening result from NIPT, the mosaicism ratio, and the type of aneuploidy.

[0045] In some embodiments, the fraction of nucleic acid in the circulating cell-free nucleic acid that has copy number variation is determined for regions of copy number variation.

[0046] In some embodiments, the fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is determined according to sequencing-based fraction estimation.

[0047] In some embodiments, the fraction of nucleic acids in the circulating cell-free nucleic acids that have copy number variations is determined according to the allelic ratio of the polymorphic sequence.

[0048] In some embodiments, the fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is determined according to quantification of differentially methylated nucleic acids.

[0049] In some embodiments, the fraction of nucleic acids in the circulating cell-free nucleic acid that have copy number variations is the fetal fraction determined for the region of copy number variation.

[0050] In some embodiments, the fetal fraction of nucleic acids with copy number variations in the circulating cell-free nucleic acids is determined according to sequencing-based fetal fraction estimation.

[0051] In some embodiments, the fetal fraction of nucleic acids with copy number variations in the circulating cell-free nucleic acids is determined according to the allelic ratios of polymorphic sequences in the fetal nucleic acids and the maternal nucleic acids.

[0052] In some embodiments, the fetal fraction of nucleic acids with copy number variations in the circulating cell-free nucleic acids is determined according to quantification of differentially methylated fetal and maternal nucleic acids.

[0053] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined for a genomic region that is larger than the region of copy number variation.

[0054] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined for a genomic region that is distinct from the region of copy number variation.

[0055] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined according to a sequencing-based fetal fraction estimation.

[0056] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined according to the allelic ratio of the polymorphic sequence in the fetal nucleic acid and the maternal nucleic acid.

[0057] In some embodiments, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined according to quantification of differentially methylated fetal and maternal nucleic acids.

[0058] In some embodiments, the mosaicism ratio is the fraction of nucleic acids with copy number variations in the circulating cell-free nucleic acid divided by the fraction of fetal nucleic acid in the circulating cell-free nucleic acid.

[0059] In some embodiments, the type of aneuploidy is trisomy 13, trisomy 18, or trisomy 21.

[0060] In some embodiments, the method further comprises providing, by the computing system, no classification if the mosaicism ratio is equal to or less than a minimum threshold value.

[0061] In some embodiments, the minimum threshold is about 0.2.

[0062] In some embodiments, the method further comprises providing, by the computing system, no classification if the mosaicism ratio is equal to or greater than a maximum threshold value.

[0063] In some embodiments, the maximum threshold is about 1.7.

[0064] In some embodiments, the first sample and the second sample are the same sample.

[0065] In some embodiments, the presence of genetic mosaicism is classified for a region of copy number variation if the mosaicism ratio is between 0.2 and 0.7, and the absence of genetic mosaicism is classified for a region of copy number variation if the mosaicism ratio is equal to or greater than 0.7.

[0066] In some embodiments, providing a personalized risk assessment includes providing an interpretation of a positive screening result from NIPT as excessive or indeterminate if no classification is provided and the mosaicism ratio is greater than a maximum threshold.

[0067] In some embodiments, providing a personalized risk assessment includes providing an interpretation of a positive screening result from NIPT with the comment that if the absence of genetic mosaicism is classified for the region of copy number variation, the mosaicism ratio suggests a non-mosaic form of aneuploidy.

[0068] In some embodiments, providing a personalized risk assessment includes providing an interpretation of a positive screening result from NIPT, with the comment that if the presence of genetic mosaicism is classified for regions of copy number variation, the mosaicism ratio suggests that the aneuploidy is a mosaic form.

[0069] In some embodiments, the method further comprises classifying, by the computing system, the presence of genetic mosaicism as "low mosaic" for the region of copy number variation if the mosaicism ratio is between 0.2 and 0.49, or classifying, by the computing system, the presence of genetic mosaicism as "high mosaic" for the region of copy number variation if the mosaicism ratio is between 0.5 and 0.69.

[0070] In some embodiments, providing a personalized risk assessment includes providing an interpretation of a positive screening result from NIPT, with the comment that if the presence of genetic mosaicism is classified as "high mosaic" for the region of copy number variation, the mosaicism ratio strongly suggests that the aneuploidy is a mosaic form.

[0071] In some embodiments, providing a personalized risk assessment includes providing an interpretation of a positive screening result from NIPT when the presence of genetic mosaicism is classified as "low mosaic" for a region of copy number variation, with the comment that the mosaicism ratio weakly suggests that the aneuploidy is a mosaic form.

[0072] In some embodiments, providing a personalized risk assessment includes providing an interpretation of a positive screening result from NIPT when the presence of genetic mosaicism is classified as "high mosaic" for the region of copy number variation and the type of aneuploidy is trisomy 13, with the comment that the mosaicism ratio is weakly suggestive of a mosaic form of aneuploidy.

[0073] In some embodiments, providing a personalized risk assessment includes providing an interpretation of a positive screening result from NIPT when the presence of genetic mosaicism is classified as "low mosaic" for the region of copy number variation and the type of aneuploidy is trisomy 13, with the comment that the mosaicism ratio weakly suggests that the aneuploidy is a mosaic form.

[0074] In some embodiments, providing a personalized risk assessment includes providing an interpretation of a positive screening result from NIPT, where the presence of genetic mosaicism is classified as "high mosaic" for the region of copy number variation and the type of aneuploidy is trisomy 18 or trisomy 21, with the comment that the mosaicism ratio strongly suggests that the aneuploidy is a mosaic form.

[0075] In some embodiments, providing a personalized risk assessment includes providing an interpretation of a positive screening result from NIPT when the presence of genetic mosaicism is classified as "low mosaic" for the region of copy number variation and the type of aneuploidy is trisomy 18 or trisomy 21, with the comment that the mosaicism ratio weakly suggests that the aneuploidy is a mosaic form.

[0076] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods disclosed herein.

[0077] In some embodiments, a computer program product is provided that includes instructions tangibly embodied in a non-transitory machine-readable storage medium and configured to cause one or more data processors to perform some or all of one or more of the methods disclosed herein.

[0078] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer-readable storage medium containing instructions that, when executed on the one or more data processors, cause the one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer program product including instructions tangibly embodied in a non-transitory machine-readable storage medium configured to cause one or more data processors to perform some or all of one or more methods and / or some or all of one or more processes disclosed herein.

[0079] The terms and expressions which have been employed are used as terms of description rather than limitation, and the use of such terms and expressions to exclude any equivalents of the features shown and described or portions thereof is not intended, but it is recognized that various modifications are possible within the scope of the claimed invention. Thus, while the claimed invention has been specifically disclosed by embodiments and optional features, it should be understood that modifications and variations of the concepts disclosed herein may be resorted to by those skilled in the art, and that such modifications and variations are deemed to be within the scope of the invention as defined by the appended claims.

[0080] Various embodiments are further described in the following description, examples, claims and drawings.

[0081] The drawings illustrate, but do not limit, certain embodiments of the present technology. For clarity and ease of illustration, the drawings are not to scale and in some cases various aspects may be shown exaggerated or enlarged to facilitate an understanding of particular embodiments. [Brief explanation of the drawings]

[0082] [Figure 1] Figure 1 shows the cell lineages early after conception (Figure adapted from Thomas, D, et al. (1994, July 10) Trisomy 22, placenta; World Wide Web URL sonoworld.com / Fetus / page.aspx?id=182). The majority of cells develop into placental trophoblast / chorionic ectoderm (direct chorionic villus sampling (CVS) preparation, NIPT). A small minority of cells develop into chorionic villi / mesoderm (CVS culture cells). Two cells in this image go on to form the embryo and amniotic tissue (amniocentesis).

[0083] [Figure 2]FIG. 2 illustrates a process flow for classifying the presence or absence of genetic mosaicism in one or more fetuses for a biological sample, according to various embodiments.

[0084] [Figure 3] FIG. 3 illustrates a process flow for classifying a biological sample as having the presence or absence of genetic mosaicism and providing clinical interpretation and / or diagnostic follow-up information, according to various embodiments.

[0085] [Figure 4] FIG. 4 illustrates an alternative process flow for classifying the presence or absence of genetic mosaicism in a biological sample and providing clinical interpretation and / or diagnostic follow-up information, according to various embodiments.

[0086] [Figure 5] FIG. 5 illustrates a process flow for classifying the sex of one or more fetuses for a biological sample, according to various embodiments.

[0087] [Figure 6] FIG. 6 illustrates an exemplary embodiment of a system in which various embodiments of the present technology may be implemented.

[0088] [Figure 7] Figure 7 shows the composition of the samples included in the aneuploid cohort [Aneuploid cohort: clinical + research samples].

[0089] [Figure 8] Figure 8 shows the composition of the samples included in the Y cohort [Y cohort: clinical + research samples].

[0090] [Figure 9] Figure 9 shows the distribution of mosaicism ratios for the aneuploid chromosome in affected singletons versus twins with one trisomy [Aneuploid cohort: clinical + research samples].

[0091] [Figure 10] FIG. 10 shows the distribution of Y chromosome mosaicism ratios in XX / XY and XY / XY twin pregnancies [Y cohort: clinical + research samples].

[0092] [Figure 11] Figure 11 shows the distribution of Y MR for XX / XY and XY / XY in euploid pregnancies [Y cohort: clinical + research samples].

[0093] [Figure 12] FIG. 12 shows the effect of mosaicism ratio on positive predictive value, according to various embodiments.

[0094] [Figure 13-1] 13A and 13B show 50 kb traces suggesting (A) non-mosaic trisomy 13 data and (B) mosaic trisomy 13 data from prenatal cfDNA screening specimens, according to various embodiments. [Figure 13-2] 13A and 13B show 50 kb traces suggesting (A) non-mosaic trisomy 13 data and (B) mosaic trisomy 13 data from prenatal cfDNA screening specimens, according to various embodiments.

[0095] [Figure 14] FIG. 14 shows the distribution of mosaicism ratios by MR group and aneuploidy across the positive screening cohort, according to various embodiments.

[0096] [Figure 15-1] 15A-15C show graphs of PPV by MR (within 0.1 range) with upper and lower 95th percentile confidence intervals according to various embodiments - (A) Trisomy 13, (B) Trisomy 18, (C) Trisomy 21. [Figure 15-2]15A-15C show graphs of PPV by MR (within 0.1 range) with upper and lower 95th percentile confidence intervals according to various embodiments - (A) Trisomy 13, (B) Trisomy 18, (C) Trisomy 21.

[0097] [Figure 16-1] 16A-16C show graphs of PPV by MR group with upper and lower 95th percentile confidence intervals—(A) trisomy 13, (B) trisomy 18, (C) trisomy 21, according to various embodiments. [Figure 16-2] 16A-16C show graphs of PPV by MR group with upper and lower 95th percentile confidence intervals—(A) trisomy 13, (B) trisomy 18, (C) trisomy 21, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0098] Detailed Description Systems and methods for non-invasive classification of mosaic copy number variations (CNVs) for one or more fetuses are provided herein. In various embodiments, bioinformatics tools and processes are used to classify the presence or absence of genetic mosaicism for copy number variations in one or more fetuses (i.e., predict whether one or more fetuses in a multifetal pregnancy are affected by a copy number variation). The methods herein can be utilized for various polynucleotides, including, for example, fragmented or cleaved nucleic acids, nucleic acid templates, cellular nucleic acids, and / or cell-free nucleic acids. In some embodiments, the sample nucleic acids subjected to the sequencing process and the resulting sequence reads are further analyzed to identify the level of genetic copy number variations and / or Y chromosomes (e.g., at one or more genomic segment levels, profile levels) in a sample containing circulating cell-free nucleic acids from a pregnant female subject with a multifetal pregnancy. The sample nucleic acid may include maternal nucleic acids and fetal nucleic acids from multiple fetuses. In some embodiments, the fraction of maternal nucleic acids in the sample nucleic acid is determined, and / or the fraction of fetal nucleic acids in the sample nucleic acid is determined. In some embodiments, the fraction of maternal nucleic acids with copy number variations in the sample nucleic acid is determined, and / or the fraction of fetal nucleic acids with copy number variations in the sample nucleic acid is determined, wherein the polymorphic sequence of the maternal nucleic acid is different from the polymorphic sequence of the fetal nucleic acid.

[0099] In some embodiments, a genetic copy number variation region is identified in a sample containing circulating cell-free nucleic acids from a pregnant female subject with a multiple pregnancy. The genetic copy number variation region comprises copy number variation, and the circulating cell-free nucleic acids comprise maternal nucleic acids and fetal nucleic acids. The fraction (e.g., minority fraction or fetal fraction) of nucleic acids with copy number variation in the sample nucleic acids is determined, and the fraction of fetal nucleic acids in the sample nucleic acids is determined. The fraction of nucleic acids with copy number variation is compared with the fraction of fetal nucleic acids, thereby providing a comparison and generating a mosaicism ratio. In some embodiments, genetic mosaicism for the copy number variation region is classified for one fetus or more than one fetus based on: (i) the mosaicism ratio of the fraction of nucleic acids with copy number variation to the fraction of fetal nucleic acids, and (ii) the number of fetuses carried by the pregnant female. In other words, the mosaicism ratio can be interpreted taking into account the number of fetuses carried by the pregnant female to predict whether one or more fetuses are affected by copy number variation (e.g., aneuploidy).

[0100] In some embodiments, the fraction of nucleic acids having a Y chromosome or a Y chromosome region (e.g., minority fraction or fetal fraction) in a sample containing circulating cell-free nucleic acids from a pregnant female subject with a multiple pregnancy is determined. In certain embodiments, the fraction of nucleic acids having a Y chromosome or a Y chromosome region is determined in part according to the level of the Y chromosome or the Y chromosome region (e.g., at one or more genomic segment levels, profile levels). The circulating cell-free nucleic acids include maternal nucleic acids and fetal nucleic acids, and the fraction of fetal nucleic acids in the circulating cell-free nucleic acids is determined. The fraction of nucleic acids having a Y chromosome or a Y chromosome region is compared with the fraction of fetal nucleic acids, thereby providing a comparison and generating a mosaicism ratio. In some embodiments, the gender of the fetus is classified based on: (i) the mosaicism ratio of the fraction of nucleic acids having a Y chromosome or a Y chromosome region to the fraction of fetal nucleic acids, and (ii) the number of fetuses carried by the pregnant female. In other words, the mosaicism ratio can be interpreted taking into account the number of fetuses carried by the pregnant female to predict the gender of one or more fetuses.

[0101] In some embodiments, systems, machines and computer program products for implementing the methods or portions of the methods described herein are also provided.

[0102] As used herein, when an action, such as a decision, is "induced by," "follows from," or "based on" something, this means that the action is induced by, follows from, or is based, at least in part, on that something. Classification of genetic mosaicism for a particular copy number variation can provide useful information about copy number variations to medical professionals and patients.

[0103] As used herein, the terms "substantially," "approximately," and "about" (unless otherwise defined herein) are defined as roughly, but not necessarily (and including) the fully specified, as would be understood by one of ordinary skill in the art. In any disclosed embodiment, the term "substantially," "approximately," or "about" may be substituted with "within [a percentage] of" the specified percentage, which percentage includes 0.1, 1, 5, and 10 percent. introduction

[0104] The detection of cell-free nucleic acids in fluid samples, particularly samples from pregnant subjects, offers great potential for use in non-invasive prenatal testing. Cell-free nucleic acid screening, or non-invasive prenatal testing (NIPT), is a screening test that utilizes bioinformatics tools and processes and next-generation sequencing of DNA fragments in maternal serum to determine the probability of a specific chromosomal condition during pregnancy. Every individual has their own cell-free DNA in their bloodstream. During pregnancy, cell-free fetal DNA from the placenta (mainly trophoblast cells) also enters the maternal bloodstream and mixes with maternal cell-free DNA. The DNA of trophoblast cells typically reflects the chromosomal makeup of the fetus. Cell-free nucleic acids are routinely screened for trisomy 21, trisomy 18, and trisomy 13. Screening for other conditions, such as fetal sex, sex chromosome aneuploidy, other aneuploidies, triploidy, and certain microdeletion conditions, is also available. Abnormal results typically indicate an increased risk of the identified condition. However, abnormal results are not diagnostic, and patients should be offered confirmatory testing via diagnostic procedures such as amniocentesis. Abnormal results may indicate an affected fetus, but may also represent a false-positive result in an unaffected pregnancy, placental-confined mosaicism, placental and fetal mosaicism, vanishing twins, an unrecognized maternal condition, or other unknown biological entity.

[0105] In particular, with prenatal cell-free DNA testing, there can be a discrepancy between analytical performance, sensitivity, specificity, clinical performance, and positive predictive value (PPV), which has posed challenges in interpreting positive NIPT results. One of the major causes underlying this discrepancy or discordant results is the difference in genetic makeup between the placenta and fetus. Chromosomal abnormalities restricted to the placenta are often mosaic and can be localized to the placenta. For example, in most pregnancies, the chromosomal complement detected in the fetus is also present in the placenta. Both the fetus and placenta share the same Because they arise from the zygote, detection of identical chromosome sets in both is expected. However, in approximately 2% of live pregnancies studied by chorionic villus sampling (CVS) at 9-11 weeks of gestation, cytogenetic abnormalities, most often trisomies, can be localized to the placenta (see, e.g., Kalousek DK, Vekemans M. Confined placental mosaicism. Journal of Medical Genetics. 1996;33(7):529-533). This phenomenon is This is known as placental localized mosaicism (CPM).Contrary to placental and fetal mosaicism, which is characterized by the existence of two or more karyotypically different cell lines in both fetus and placenta, CPM shows the discrepancy between the chromosomal constitution of the cells in placenta and the cells in fetus.As a result, CPM is usually accompanied by normal fetal outcome (for example, when CPM is found, it most commonly shows trisomic cell line in placenta and normal diploid chromosome set in baby), but it may be misunderstood from a diagnostic point of view (i.e., false positive in NIPT).

[0106] Given that NIPT can produce false positives, positive NIPT results are typically confirmed using invasive testing, such as CVS and / or amniocentesis. For example, prenatal care is typically not a discrete event but rather a 40-week continuum of care for patients. Therefore, each data point collected throughout pregnancy should provide clinicians with ample clinically relevant information to contextualize all available information. Ideally, clinical data, including CVS and / or amniocentesis analysis, for all positive NIPT results helps mitigate concerns about false positives before making irreversible treatment decisions (e.g., abortion). However, CPM can also cause false-positive results with CVS. Therefore, conventional practice is to proceed with CVS and test all cell lines using both uncultured samples or short-term cultures and long-term cultures of samples using fluorescence in situ hybridization (FISH). If all results indicate aneuploidy, the results are reported to the patient. Otherwise, if these results are also mosaic, amniocentesis is recommended and analyzed by both FISH and karyotyping. Nevertheless, a real-world limitation to conventional practice is that not all women agree to invasive diagnostic testing, especially in the first trimester.

[0107] To address these false-positive problems and the reluctance of many women to consent to invasive diagnostic testing, various embodiments described herein introduce the application of mosaicism ratios (a metric that can be obtained from prenatal cell-free DNA testing, as described in detail herein) to identify patients with multiple pregnancies in which aneuploidy may be present in a mosaic form (e.g., CPM). As shown in FIG. 1 , the majority of cells develop from the zygote into placental trophoblast / chorionic ectoderm 105, a very small minority of cells develop into chorionic villi / mesoderm 110, and only two cells proceed to form the embryo and amniotic tissue 115. If errors in cell division occur at different levels in this chain, different levels of fetal or placental (or both) mosaicism can result, which may have fundamentally different clinical implications. If this is the case, not all cell-free trophoblast DNA in maternal plasma is affected. This observation can be used to calculate the mosaicism ratio (MR) of affected cell-free DNA (e.g., fraction with copy number variation or Y chromosome) and total cell-free DNA (e.g., total fraction of fetal cell-free DNA).

[0108] In various embodiments, the MR is calculated by: (a) determining the fraction of nucleic acids having copy number variations and / or Y chromosome levels (e.g., at one or more genomic segment levels, profile levels) in the sample nucleic acid; (b) determining the fraction of minority nucleic acids in the sample nucleic acid (e.g., fetal fraction); and (c) comparing the fraction in (a) with the fraction in (b) to generate a ratio of (a:):(b). Furthermore, it has been discovered that the MR can be used to predict whether one or more fetuses in a singleton or multipleton pregnancy subject are affected by copy number variations (e.g., aneuploidy). Furthermore, it has been discovered that the MR can be used to predict the gender of one or more fetuses in a singleton or multipleton pregnancy subject. In some embodiments, the MR is used to 1) predict whether one or more fetuses are affected by aneuploidy and / or 2) provide information about the expected gender of one or more fetuses. For example, the mosaicism ratio can be interpreted in consideration of the number of fetuses carried by a pregnant female to determine whether one or more fetuses are affected by copy number variations (e.g., aneuploidy) and / or to predict the gender of one or more fetuses. In certain embodiments, the presence of genetic mosaicism is classified for copy number variations based on: (i) the mosaicism ratio of the fraction of nucleic acids with copy number variations to the fraction of fetal nucleic acids, and (ii) the chromosomes carrying genetic copy number variations (e.g., identified aneuploidy types). In certain embodiments, the presence of genetic mosaicism is classified for copy number variations based on: (i) the mosaicism ratio of the fraction of nucleic acids with copy number variations to the fraction of fetal nucleic acids, and (ii) the number of fetuses carried by a pregnant female. In certain embodiments, the gender of a fetus is classified based on: (i) the mosaicism ratio of the fraction of nucleic acids with Y chromosomes or regions of Y chromosomes to the fraction of fetal nucleic acids, and (ii) the number of fetuses carried by a pregnant female.The use of mosaicism ratios in such situations has many advantages over traditional processes for confirming positive NIPT results, including a non-invasive approach for confirming a positive NIPT result for one or more fetuses and providing information about the predicted gender of the one or more fetuses.

[0109] Furthermore, knowledge of the presence or absence of mosaicism can be used by physicians and genetic counselors to better interpret positive NIPT results, which can lead to improved post-test counseling and overall prenatal care. For example, the presence of a genetic mosaicism classification in a fetus of a singleton pregnancy subject for a copy number variation region (e.g., 20% to 70% MR) can be interpreted as a non-standard positive NIPT result, along with a mosaic comment. Alternatively, the presence of a genetic mosaicism classification in one fetus of a multiton pregnancy subject for a copy number variation region (e.g., 20% to 60% MR) can be interpreted as a non-standard positive NIPT result, along with a mosaic comment. The presence of a genetic mosaicism classification in more than one fetus of a multiton pregnancy subject for a copy number variation region (e.g., 60% to 130% MR) can be interpreted as a non-standard positive NIPT result, along with a mosaic comment. The absence of a genetic mosaicism classification for a copy number variation region in a multifetal pregnancy subject (e.g., an MR greater than 130%) can be interpreted as a standard positive NIPT result (e.g., a positive result for fetal copy number variation), one or more affected fetuses, fetal copy number variation, full copy number variation, true copy number variation, complete copy number variation, etc. No classification (e.g., no call, no clinical relevance) can be provided if the MR value in a multifetal pregnancy subject is below a certain threshold for the copy number variation region (e.g., an MR less than 20%), which can be interpreted as a negative NIPT result for fetal copy number variation for all fetuses. Genetic mosaicism classification for fetuses in singleton pregnancies

[0110] Provided herein is a method for classifying the presence or absence of genetic mosaicism (e.g., CPM) in sample (e.g., biological sample; test sample).In various embodiments, the presence or absence of genetic mosaicism is classified based on copy number variation.Copy number variation, which can be referred to as copy number alteration, can include aneuploidy (e.g., chromosome trisomy, chromosome monosomy), deletion (e.g., microdeletion; subchromosome deletion) and duplication (e.g., microduplication, subchromosome duplication), and will be described in more detail herein.

[0111] The presence or absence of genetic mosaicism can be classified for copy number variation regions (for example, trisomic cell lines confined to the placenta).Copy number variation regions refer to genomic regions (for example, chromosomes, parts of chromosomes) for which copy number variation is identified.Copy number variation regions can refer to specific chromosomes or chromosomal locations (for example, regions spanning a specific genomic coordinate).Copy number variation regions can be identified using any suitable method for identifying copy number variation in the art or as described herein.

[0112] In some embodiments, the methods herein include determining the fraction of nucleic acids having copy number variations in sample nucleic acids. Determining the fraction of nucleic acids refers to quantifying nucleic acids of a particular species in a nucleic acid mixture. For example, determining the fraction of nucleic acids may refer to quantifying minority nucleic acid species, fetal nucleic acids, cancer nucleic acids, etc. Determining the fraction of nucleic acids having copy number variations refers to quantifying a subset of nucleic acids (e.g., a subset of nucleic acid fragments, a subset of sequence reads) for which copy number variations are identified. In some embodiments, determining the fraction of nucleic acids having copy number variations refers to quantifying a subset of nucleic acids (e.g., a subset of nucleic acid fragments, a subset of sequence reads) from a region (e.g., a genomic region) for which copy number variations are identified. In some embodiments, determining the fraction of nucleic acids having copy number variations refers to quantifying a subset of nucleic acids for a species (e.g., a subset of nucleic acid fragments for a species, a subset of sequence reads for a species) from a region (e.g., a genomic region) for which copy number variations are identified. For example, for a sample containing maternal and fetal nucleic acids, if the fetal nucleic acid is identified as having trisomy of chromosome 21, determining the fraction of nucleic acids having copy number variations refers to determining the fetal fraction based on information from or associated with chromosome 21 or a portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequences, differentially methylated sequences).

[0113] In some embodiments, the methods herein include determining a fraction for a region (e.g., a genomic region). In some embodiments, the methods herein include determining a fraction for a copy number variation region. The fraction for a copy number variation region may be referred to as an affected fraction or a fraction for an affected region. As discussed above, the fraction for a copy number variation region may be determined according to information (e.g., sequence information, epigenetic information) obtained for a region (e.g., a genomic region) identified as having a copy number variation. The fraction for a copy number variation region may be determined using any suitable method for quantifying nucleic acid species in a nucleic acid mixture. For example, the fraction for a copy number variation region may be determined according to sequencing-based fraction estimation. Methods for determining nucleic acid fractions according to sequencing-based fraction estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815. and U.S. Pat. Appl. Pub. No. 2011 / 0224087, each of which is hereby incorporated by reference herein. Sequencing-based fraction estimation may be referred to as bin-based fraction estimation and / or portion-specific fraction estimation. In some embodiments, the fraction for a copy number variation region may be determined according to the allele ratio of a polymorphic sequence. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining nucleic acid fractions according to the allele ratio of a polymorphic sequence are described herein and in U.S. Pat. Appl. Pub. No. 2011 / 0224087, which is hereby incorporated by reference herein. In some embodiments, the fraction for a copy number variation region may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated nucleic acids). Methods for determining nucleic acid fractions according to quantification of differentially methylated nucleic acids are described, for example, herein and in U.S. Pat. Appl. Pub. No. 2010 / 0105049, which is hereby incorporated by reference herein.

[0114] In some embodiments, the sample nucleic acid comprises a majority nucleic acid and a minority nucleic acid. In some embodiments, the majority nucleic acid comprises a maternal nucleic acid, and the minority nucleic acid comprises a fetal nucleic acid. Thus, in some embodiments, the methods herein comprise determining a fetal fraction. In some embodiments, the methods herein comprise determining a fetal fraction for a region (e.g., a genomic region). In some embodiments, the methods herein comprise determining a fetal fraction for a copy number variation region. The fetal fraction for a copy number variation region may be referred to as an affected fraction, an affected fetal fraction, and / or a fetal fraction for an affected region. As discussed above, the fetal fraction for a copy number variation region may be determined according to information (e.g., sequence information, epigenetic information) obtained for a region (e.g., a genomic region) identified as having a fetal copy number variation. The fetal fraction for a copy number variation region may be determined using any suitable method for quantifying fetal nucleic acid in a mixture of maternal and fetal nucleic acids. For example, the fetal fraction for a copy number variation region may be determined according to sequencing-based fetal fraction (SeqFF) estimation. Methods for determining fetal fraction according to sequencing-based fetal fraction (SeqFF) estimation are described herein, as well as in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is hereby incorporated by reference herein. The sequencing-based fetal fraction (SeqFF) estimation may be referred to as bin-based fetal fraction (BFF) estimation and / or part-specific fetal fraction estimation. In some embodiments, the fetal fraction for a copy number variation region may be determined according to the allele ratio of a polymorphic sequence in fetal nucleic acid and maternal nucleic acid. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining the fetal fraction according to the allele ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference. In some embodiments, the fetal fraction for a copy number variation region may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated fetal nucleic acid and maternal nucleic acid). Methods for determining the fetal fraction according to the quantification of differentially methylated fetal nucleic acid and maternal nucleic acid are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference.

[0115] In some embodiments, the methods herein include determining the fraction of minority nucleic acids in the sample nucleic acid. Determining the fraction of minority nucleic acids in the sample nucleic acid generally involves methods, such as those described above, that quantify nucleic acid species based on information about regions identified as having copy number variations. Rather, determining the fraction of minority nucleic acids in the sample nucleic acid may include methods that quantify minority nucleic acids according to information from regions across the genome and / or regions different from the regions identified as having copy number variations. In some embodiments, the fraction of minority nucleic acids is determined for genomic regions that are larger than the copy number variation region. For example, the fraction of minority nucleic acids may be determined for genomic regions that contain more genomic content (e.g., base pairs, kilobases, megabases) than the regions identified as having copy number variations. For example, for a sample in which minority nucleic acids are identified as having trisomy 21, the fraction of minority nucleic acids may be determined according to information from or associated with multiple chromosomes (e.g., sequence information, sequence read quantification, polymorphism sequences, differentially methylated sequences). In this example, the plurality of chromosomes may include all chromosomes, all autosomes, a subset of chromosomes, a subset of autosomes, a subset of chromosomes including chromosome 21, a subset of autosomes including chromosome 21, a subset of chromosomes excluding chromosome 21, a subset of autosomes excluding chromosome 21, or a portion thereof. In some embodiments, the fraction of minority nucleic acids is determined for a genomic region different from the copy number variation region. For example, for a sample in which minority nucleic acids are identified as having trisomy of chromosome 21, the fraction of minority nucleic acids can be determined according to information from or related to chromosomes other than chromosome 21 (for example, sequence information, sequence read quantification, polymorphism sequence, differentially methylated sequence).

[0116] The minority nucleic acid fraction in sample nucleic acid can be determined by any suitable method for quantifying the nucleic acid species in nucleic acid mixture.For example, the minority nucleic acid fraction can be determined according to sequencing-based fraction estimation.The method for determining minority nucleic acid fraction according to sequencing-based fraction estimation is described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, and these are each hereby incorporated by reference herein. Sequencing-based fraction estimation may be referred to as bin-based fraction estimation and / or portion-specific fraction estimation. In some embodiments, the fraction of minority nucleic acids may be determined according to the allelic ratio of polymorphic sequences. Polymorphic sequences may include, for example, single nucleotide polymorphisms (SNPs). Methods for determining the minority nucleic acid fraction according to the allelic ratio of polymorphic sequences are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference herein. In some embodiments, the fraction of minority nucleic acids may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated nucleic acids). Methods for determining the minority nucleic acid fraction according to the quantification of differentially methylated nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference herein.

[0117] In some embodiments, the minority nucleic acid comprises a fetal nucleic acid. Thus, in some embodiments, the method herein comprises determining a fetal fraction. The fetal fraction can be determined using any suitable method for quantifying fetal nucleic acid in a mixture of maternal and fetal nucleic acids. For example, the fetal fraction can be determined according to sequencing-based fetal fraction (SeqFF) estimation. Methods for determining fetal fraction according to sequencing-based fetal fraction (SeqFF) estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is incorporated herein by reference. The present invention relates to a method for determining the fetal fraction of a fetal gene by using a sequencing-based fetal fraction (SeqFF) estimation method. The method may be referred to as a bin-based fetal fraction (BFF) estimation method and / or a site-specific fetal fraction estimation method. In some embodiments, the fetal fraction may be determined according to the allele ratio of a polymorphic sequence in fetal and maternal nucleic acids. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining the fetal fraction according to the allele ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference. In some embodiments, the fetal fraction may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated fetal and maternal nucleic acids). Methods for determining the fetal fraction according to the quantification of differentially methylated fetal and maternal nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference. In some embodiments, fetal fraction can be determined according to a Y chromosome assay. Methods for determining fetal fraction according to a Y chromosome assay are described herein and in Lo YM, et al. (1998) Am J Hum Genet 62:768-775. .

[0118] In some embodiments, the fraction of copy number variation region and the fraction of minority nucleic acid are determined using the same methodology.For example, the fraction of copy number variation region and the fraction of minority nucleic acid can be determined according to sequencing-based fraction estimation.In some embodiments, the fraction of copy number variation region and the fraction of minority nucleic acid are determined using different methodologies.For example, the fraction of copy number variation region can be determined according to the allele ratio of polymorphic sequence, and the fraction of minority nucleic acid can be determined according to differential epigenetic biomarkers.

[0119] In some embodiments, the fetal fraction of copy number variation region and the fetal fraction of nucleic acid sample are determined using the same methodology.For example, the fetal fraction of copy number variation region and the fetal fraction of nucleic acid sample can be determined according to sequencing-based fetal fraction estimation.In some embodiments, the fetal fraction of copy number variation region and the fetal fraction of nucleic acid sample are determined using different methodologies.For example, the fetal fraction of copy number variation region can be determined according to the allele ratio of polymorphism sequence, and the fetal fraction of nucleic acid sample can be determined according to Y chromosome assay.

[0120] In some embodiments, the fraction for copy number variation (e.g., copy number variation region) is determined for a chromosome or a portion thereof. The fraction for copy number variation determined for a chromosome or a portion thereof refers to quantification of nucleic acid species based on information from or associated with the chromosome or a portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequence, differentially methylated sequence). In some embodiments, the fraction for copy number variation (e.g., copy number variation region) is determined for chromosome 13, chromosome 18, or chromosome 21. In some embodiments, the fraction of minority nucleic acids is determined for a chromosome or a portion thereof different from the chromosome or portion thereof used to determine the fraction for copy number variation. In some embodiments, the fraction of minority nucleic acids is determined for multiple chromosomes or multiple portions of chromosomes. In some embodiments, the fraction of minority nucleic acids is determined for multiple autosomes or multiple portions of autosomes. In some embodiments, the fraction of minority nucleic acids is determined for multiple regions (e.g., genomic regions). In some embodiments, the fraction of minority nucleic acids is determined for multiple regions (e.g., genomic regions) genome-wide.

[0121] In some embodiments, the fetal fraction for a copy number variation (e.g., a copy number variation region) is determined for a chromosome or a portion thereof. The fetal fraction for a copy number variation determined for a chromosome or a portion thereof refers to quantification of fetal nucleic acid based on information from or associated with the chromosome or a portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequence, differentially methylated sequence). In some embodiments, the fetal fraction for a copy number variation (e.g., a copy number variation region) is determined for chromosome 13, chromosome 18, or chromosome 21. In some embodiments, the fetal fraction for the sample nucleic acid is determined for a chromosome or a portion thereof that is different from the chromosome or portion thereof used to determine the fetal fraction for the copy number variation. In some embodiments, the fetal fraction for the sample nucleic acid is determined for multiple chromosomes or multiple portions of chromosomes. In some embodiments, the fetal fraction for the sample nucleic acid is determined for multiple autosomes or multiple portions of autosomes. In some embodiments, the fetal fraction for the sample nucleic acid is determined for multiple regions (e.g., genomic regions). In some embodiments, the fetal fraction for the sample nucleic acid is determined for multiple regions genome-wide (e.g., genomic regions).

[0122] In some embodiments, the methods herein include comparing the fraction for copy number variation with the fraction of minority nucleic acids. In some embodiments, comparing the fraction for copy number variation with the fraction of minority nucleic acids includes generating a mosaicism ratio. For example, the mosaicism ratio can be the fraction for copy number variation divided by the fraction of minority nucleic acids.

[0123] In some embodiments, the methods herein include comparing the fetal fraction for the copy number variation with the fetal fraction for the sample nucleic acid. In some embodiments, comparing the fetal fraction for the copy number variation with the fetal fraction for the sample nucleic acid includes generating a ratio. For example, the mosaicism ratio can be the fetal fraction for the copy number variation divided by the fetal fraction for the sample nucleic acid.

[0124] In some embodiments, the method herein comprises classifying the presence or absence of genetic mosaicism for the copy number variation region. The presence or absence of genetic mosaicism for the copy number variation region can be classified according to the comparison. For example, the presence or absence of genetic mosaicism for the copy number variation region can be classified according to the comparison of the fraction for copy number variation and the fraction of minority nucleic acid. In some embodiments, the presence or absence of genetic mosaicism for the copy number variation region can be classified according to the comparison of the fetal fraction for copy number variation and the fetal fraction for the sample nucleic acid. The presence or absence of genetic mosaicism for the copy number variation region can be classified according to the ratio. For example, the presence or absence of genetic mosaicism for the copy number variation region can be classified according to the mosaicism ratio of the fraction for copy number variation to the fraction of minority nucleic acid (for example, the fraction for copy number variation divided by the fraction of minority nucleic acid). In some embodiments, the presence or absence of genetic mosaicism for the copy number variation region may be classified according to the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid (e.g., the fetal fraction for the copy number variation divided by the fetal fraction for the sample nucleic acid).

[0125] In some embodiments, the presence of genetic mosaicism is classified for copy number variation regions. The presence of genetic mosaicism classification for copy number variation regions can be interpreted as mosaic copy number variation, affected fetus, unaffected fetus, partially affected fetus, fetal copy number variation, partial fetal copy number variation, partial copy number variation, placental copy number variation, partial placental copy number variation, incomplete copy number variation, placental mosaicism, placental limited mosaicism (CPM), etc.

[0126] In some embodiments, the presence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the copy number variation fraction to the minority nucleic acid fraction is less than 1. For example, the presence of genetic mosaicism can be classified for a copy number variation region when the mosaicism ratio of the copy number variation fraction to the minority nucleic acid fraction is about 0.1 to about 0.9, or about 0.1 to about 0.8, or about 0.1 to about 0.7, or about 0.1 to about 0.6, or about 0.2 to about 0.9, or about 0.2 to about 0.8, or about 0.2 to about 0.7, or about 0.2 to about 0.6. In certain embodiments, the presence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the copy number variation fraction to the minority nucleic acid fraction is between 0.2 and 0.7. For example, the presence of genetic mosaicism can be classified for a copy number variation region when the mosaicism ratio of the fraction for copy number variation to the fraction for the minority nucleic acid is about 0.2, 0.3, 0.4, 0.5, 0.6, or 0.7. As used herein, the terms "substantially," "approximately," and "about" (unless otherwise defined herein) are defined as generally, but not necessarily (and including) fully specified, as understood by those skilled in the art. In any disclosed embodiment, the terms "substantially," "approximately," or "about" can be substituted with "within [a percentage] of" the specified percentage, including 0.1, 1, 5, and 10 percent.

[0127] In some embodiments, the presence of genetic mosaicism is further classified as "low mosaic" for a copy number variation region when the mosaicism ratio value of the fraction for copy number variation to the fraction for minority nucleic acid is between 0.2 and 0.49, and the presence of genetic mosaicism is classified as "high mosaic" for a copy number variation region when the mosaicism ratio value of the fraction for copy number variation to the fraction for minority nucleic acid is between 0.5 and 0.69.

[0128] In some embodiments, the presence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is within a range of values ​​less than 1. For example, the presence of genetic mosaicism can be classified for a copy number variation region when the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is about 0.1 to about 0.9, or about 0.1 to about 0.8, or about 0.1 to about 0.7, or about 0.1 to about 0.6, or about 0.2 to about 0.9, or about 0.2 to about 0.8, or about 0.2 to about 0.7, or about 0.2 to about 0.6. In some embodiments, the presence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is between 0.2 and 0.7. For example, the presence of genetic mosaicism can be classified for a copy number variation region if the mosaicism ratio value of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is about 0.2, 0.3, 0.4, 0.5, 0.6, or 0.7.

[0129] In some embodiments, the presence of genetic mosaicism is further classified as "low mosaic" for a copy number variation region when the mosaicism ratio value of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is between 0.2 and 0.49, and the presence of genetic mosaicism is classified as "high mosaic" for a copy number variation region when the mosaicism ratio value of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is between 0.5 and 0.69.

[0130] In some embodiments, the absence of genetic mosaicism is classified for the copy number variation region. The absence of genetic mosaicism classification for the copy number variation region can be interpreted as a standard positive result (e.g., a positive result for fetal copy number variation), an affected fetus, a fetal copy number variation, a complete copy number variation, a true copy number variation, a complete copy number variation, etc.

[0131] In some embodiments, the absence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the copy number variation fraction to the minority nucleic acid fraction is greater than 0.6. For example, the absence of genetic mosaicism can be classified for a copy number variation region when the mosaicism ratio of the copy number variation fraction to the minority nucleic acid fraction is between about 0.7 and about 1.5, or between about 0.7 and about 1.3, or between about 0.7 and about 1.1, or between about 0.8 and about 1.1, or between about 0.8 and about 1.0, or between about 0.8 and about 0.9. In some embodiments, the absence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the copy number variation fraction to the minority nucleic acid fraction is between about 0.71 and about 1.3. For example, the absence of genetic mosaicism can be classified for a copy number variation region when the mosaicism ratio of the fraction for copy number variation to the fraction for minority nucleic acid is about 0.71, 0.8, 0.9, 1.0, 1.1, 1.2, or 1.3. In other embodiments, the absence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the fraction for copy number variation to the fraction for minority nucleic acid is equal to or greater than 0.7.

[0132] In some embodiments, the absence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is greater than 0.6. For example, the absence of genetic mosaicism can be classified for a copy number variation region when the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is between about 0.7 and about 1.5, or between about 0.7 and about 1.3, or between about 0.7 and about 1.1, or between about 0.8 and about 1.1, or between about 0.8 and about 1.0, or between about 0.8 and about 0.9. In some embodiments, the absence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is between about 0.71 and about 1.3. For example, the absence of genetic mosaicism can be classified for a copy number variation region when the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is about 0.71, 0.8, 0.9, 1.0, 1.1, 1.2, or 1.3. In other embodiments, the absence of genetic mosaicism is classified for a copy number variation region when the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is equal to or greater than 0.7.

[0133] In some embodiments, no classification is provided. For example, if the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is below a certain threshold, no classification (e.g., no call, no clinical relevance) may be provided. In some embodiments, if the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is about 0.3 or less, no classification is provided. In some embodiments, if the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is about 0.2 or less, no classification is provided. In some embodiments, if the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is about 0.1 or less, no classification is provided.

[0134] In some embodiments, when the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is above a certain threshold, no classification is provided.For example, when the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is about 0.9, 1.0, 1.1, 1.2 or 1.3 or more, no classification can be provided.In some embodiments, when the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is about 1.3 or more, no classification is provided.A value above a certain threshold (for example, above 1.3) can indicate copy number variation (for example, maternal copy number variation) present in majority nucleic acid.

[0135] In some embodiments, no classification (e.g., no call, no clinical relevance) may be provided if the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is below a certain threshold. In some embodiments, no classification is provided if the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is about 0.3 or less. In some embodiments, no classification is provided if the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is about 0.2 or less. In some embodiments, no classification is provided if the mosaicism ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid is about 0.1 or less.

[0136] In some embodiments, when the mosaicism ratio value of the fetal fraction of copy number variation to the fetal fraction of sample nucleic acid is above a certain threshold, no classification is provided.For example, when the mosaicism ratio value of the fetal fraction of copy number variation to the fetal fraction of sample nucleic acid is about 0.9, 1.0, 1.1, 1.2 or 1.3 or more, no classification can be provided.In some embodiments, when the mosaicism ratio value of the fetal fraction of copy number variation to the fetal fraction of sample nucleic acid is about 1.3 or more, no classification is provided.

[0137] FIG. 2 shows a process 200 for classifying the presence or absence of genetic mosaicism in one or more fetuses for a biological sample, according to various embodiments. A set of sequence reads is provided in block 205. The sequence reads may be obtained from circulating cell-free sample nucleic acids from a test sample obtained from a multifetal pregnant subject (e.g., a pregnant female subject having multiple fetuses). Additionally, the number of fetuses carried by the multifetal pregnant subject is obtained. The circulating cell-free nucleic acids may include maternal nucleic acids and fetal nucleic acids. The circulating cell-free sample nucleic acids may be captured by probe oligonucleotides under hybridization conditions. In block 210, genetic copy number variation regions are identified in the circulating cell-free nucleic acids from the set of sequence reads. The fraction of circulating cell-free nucleic acids having copy number variations in the sample nucleic acids is determined in block 215. The fraction may be a fetal fraction determined for the copy number variation region. The fraction of fetal nucleic acids in the circulating cell-free sample nucleic acids is determined in block 220. The fraction of circulating cell-free nucleic acid with copy number variation is compared to the fraction of fetal nucleic acid in block 225 to generate a mosaicism ratio of the fraction of circulating cell-free nucleic acid with copy number variation to the fraction of fetal nucleic acid. The presence or absence of genetic mosaicism for the copy number variation region in one or more fetuses is classified in block 230 according to the mosaicism ratio and the number of fetuses carried by the multiple-fetal pregnancy subject.

[0138] FIG. 3 shows a process 300 for classifying the presence or absence of genetic mosaicism for a biological sample and providing clinical interpretation and / or diagnostic follow-up information, according to various embodiments. A set of sequence reads is provided, and a screening test for a genetic condition (e.g., NIPT) is obtained from the set of sequence reads in step 305. The sequence reads may be obtained from circulating cell-free sample nucleic acids from a test sample obtained from a test subject (e.g., a pregnant female subject). The test sample may be the same or a different sample from the sample used to generate the mosaicism ratio. The circulating cell-free nucleic acids may include maternal and fetal nucleic acids. The circulating cell-free sample nucleic acids may be captured by probe oligonucleotides under hybridization conditions. In various embodiments, the genetic condition being screened for includes the presence of one or more aneuploidies, e.g., copy number variations. The presence (flag as positive) or absence (flag as negative) of one or more aneuploidies may be identified in the circulating cell-free nucleic acids from the set of sequence reads in step 310 or 315 based on a z-score. If the absence of one or more aneuploidies (flagged as negative) is identified, no further testing may be performed in step 320, or diagnostic testing may be performed in step 325. If the presence of one or more aneuploidies (flagged as positive) is identified, a mosaicism ratio is generated as described with respect to FIG. 2, and the mosaicism ratio value is used to classify the presence or absence of genetic mosaicism and provide improved interpretation of NIPT results. The mosaicism ratio can be used to identify patients with a higher likelihood of discordant positive results due to mosaicism (e.g., CPM).

[0139] The presence of genetic mosaicism may be classified in step 330 for a copy number variation region if the mosaicism ratio value is between 0.2 and 0.7. The absence of genetic mosaicism may be classified 335 for a copy number variation region if the mosaicism ratio value is equal to or greater than 0.7. Furthermore, no classification may be provided in steps 340 / 345 for a copy number variation region if the mosaicism ratio value is equal to or greater than about 1.3 or equal to or less than about 0.2. If no classification is provided and the mosaicism ratio value is greater than about 1.3, the positive NIPT result may be interpreted as possibly excessive or indeterminate in step 350, and diagnostic follow-up including amniocentesis, CVS, maternal testing, and / or other testing may be recommended in step 355, depending on the consensus decision between the genetic counselor and the physician. If no classification is provided and the mosaicism ratio value is less than about 0.2, the positive NIPT result may be interpreted in step 360 as a negative result or the absence of one or more aneuploidies, and no diagnostic follow-up may be invoked in step 365.

[0140] If the presence of genetic mosaicism is classified (e.g., the mosaicism ratio is between 0.2 and 0.7), the positive NIPT result may be interpreted as positive in step 370 with a mosaic comment (e.g., the understanding that the mosaicism ratio suggests that the aneuploidy is in a mosaic form), and diagnostic follow-up including amniocentesis and / or CVS may be recommended in step 375, depending on the consensus decision between the genetic counselor and the physician. If the absence of genetic mosaicism is classified (e.g., the mosaicism ratio is greater than or equal to 0.7 but less than about 1.3), the positive NIPT result may be interpreted as positive in step 380 with a mosaic comment (e.g., the understanding that the mosaicism ratio suggests that the aneuploidy is in a non-mosaic form), and diagnostic follow-up including amniocentesis and / or CVS may be recommended in step 385 for confirmation.

[0141] In various embodiments, step 370 may include further analysis using finer-grained classification of genetic mosaicism and interpretation that takes into account the type of aneuploidy detected in NIPT. In some cases, step 370 may further include classifying the presence of genetic mosaicism as "low mosaic" for the copy number variation region if the mosaicism ratio is between 0.2 and 0.49, or classifying the presence of genetic mosaicism as "high mosaic" for the copy number variation region if the mosaicism ratio is between 0.5 and 0.69. If the presence of genetic mosaicism is classified as low mosaic (e.g., the mosaicism ratio is between 0.2 and 0.49), the positive NIPT result may be interpreted as positive in step 370 with a mosaic comment (e.g., the understanding that the mosaicism ratio weakly suggests that the aneuploidy is in a mosaic form, particularly if the type of aneuploidy is trisomy 13, trisomy 18, or trisomy 21), and diagnostic follow-up including amniocentesis and / or CVS may be recommended in step 375, depending on the consensus decision between the genetic counselor and the physician. If the presence of genetic mosaicism is classified as high mosaicism (e.g., the mosaicism ratio is between 0.5 and 0.69), the positive NIPT result may be interpreted as positive in step 370 with a mosaic comment (e.g., the understanding that the mosaicism ratio weakly suggests that the aneuploidy is in a mosaic form, particularly if the type of aneuploidy is trisomy 13; or the understanding that the mosaicism ratio strongly suggests that the aneuploidy is in a mosaic form, particularly if the type of aneuploidy is trisomy 18 or trisomy 21), and diagnostic follow-up including amniocentesis and / or CVS may be recommended in step 375, depending on the consensus decision between the genetic counselor and the physician. Genetic mosaicism classification for one or more fetuses in a multiple pregnancy

[0142] Provided herein is a method for classifying the presence or absence of genetic mosaicism (e.g., CPM) in one or more fetuses for a sample (e.g., biological sample; test sample).In various embodiments, the presence or absence of genetic mosaicism in one or more fetuses is classified for copy number variation (i.e., predict whether one fetus or more than one fetus in a multifetal pregnancy is affected by copy number variation).Copy number variation, which can be referred to as copy number alteration, can include aneuploidy (e.g., chromosome trisomy, chromosome monosomy), deletion (e.g., microdeletion; subchromosome deletion) and duplication (e.g., microduplication, subchromosome duplication), and will be described in more detail herein.

[0143] The presence or absence of genetic mosaicism in one or more fetuses can be classified according to copy number variation region (for example, trisomy cell lineage localized in placenta).Copy number variation region refers to the genomic region (for example, chromosome, part of chromosome) for which copy number variation is identified.Copy number variation region can refer to a specific chromosome, or can refer to a location on a chromosome (for example, a region that spans a specific genomic coordinate).Copy number variation region can be identified using any suitable method for identifying copy number variation in the art or as described herein.

[0144] In some embodiments, the methods herein include determining the fraction of nucleic acids having copy number variations in a sample of nucleic acids from a subject with a multiple pregnancy. Determining the fraction of nucleic acids refers to quantifying nucleic acids of a particular species in a nucleic acid mixture. For example, determining the fraction of nucleic acids can refer to quantifying minority nucleic acid species, quantifying fetal nucleic acids, quantifying cancer nucleic acids, etc. Determining the fraction of nucleic acids having copy number variations refers to quantifying a subset of nucleic acids (e.g., a subset of nucleic acid fragments, a subset of sequence reads) for which copy number variations are identified. In some embodiments, determining the fraction of nucleic acids having copy number variations refers to quantifying a subset of nucleic acids (e.g., a subset of nucleic acid fragments, a subset of sequence reads) from a region (e.g., a genomic region) for which copy number variations are identified. In some embodiments, determining the fraction of nucleic acids having copy number variations refers to quantifying a subset of nucleic acids for a species (e.g., a subset of nucleic acid fragments for a species, a subset of sequence reads for a species) from a region (e.g., a genomic region) for which copy number variations are identified. For example, for a sample containing maternal and fetal nucleic acids from a subject with a multiple pregnancy, if the fetal nucleic acid is identified as having trisomy of chromosome 21, determining the fraction of nucleic acids having copy number variations refers to determining the fetal fraction based on information from or associated with chromosome 21 or a portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequences, differentially methylated sequences).

[0145] In some embodiments, the methods herein include determining a fraction for a region (e.g., a genomic region). In some embodiments, the methods herein include determining a fraction for a copy number variation region. The fraction for a copy number variation region may be referred to as an affected fraction or a fraction for an affected region. As discussed above, the fraction for a copy number variation region may be determined according to information (e.g., sequence information, epigenetic information) obtained for a region (e.g., a genomic region) identified as having a copy number variation. The fraction for a copy number variation region may be determined using any suitable method for quantifying nucleic acid species in a nucleic acid mixture. For example, the fraction for a copy number variation region may be determined according to sequencing-based fraction estimation. Methods for determining nucleic acid fractions according to sequencing-based fraction estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815. and U.S. Pat. Appl. Pub. No. 2011 / 0224087, each of which is hereby incorporated by reference herein. Sequencing-based fraction estimation may be referred to as bin-based fraction estimation and / or portion-specific fraction estimation. In some embodiments, the fraction for a copy number variation region may be determined according to the allele ratio of a polymorphic sequence. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining nucleic acid fractions according to the allele ratio of a polymorphic sequence are described herein and in U.S. Pat. Appl. Pub. No. 2011 / 0224087, which is hereby incorporated by reference herein. In some embodiments, the fraction for a copy number variation region may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated nucleic acids). Methods for determining nucleic acid fractions according to quantification of differentially methylated nucleic acids are described, for example, herein and in U.S. Pat. Appl. Pub. No. 2010 / 0105049, which is hereby incorporated by reference herein.

[0146] In some embodiments, the nucleic acid sample includes majority nucleic acids (e.g., more than minority nucleic acids) and minority nucleic acids (e.g., less than majority nucleic acids). In some embodiments, the majority nucleic acids include maternal nucleic acids, and the minority nucleic acids include fetal nucleic acids. Thus, in some embodiments, the methods herein include determining a fetal fraction. In some embodiments, the methods herein include determining a fetal fraction for a region (e.g., a genomic region). In some embodiments, the methods herein include determining a fetal fraction for a copy number variation region. The fetal fraction for a copy number variation region may be referred to as an affected fraction, an affected fetal fraction, and / or a fetal fraction for an affected region. As discussed above, the fetal fraction for a copy number variation region may be determined according to information (e.g., sequence information, epigenetic information) obtained for a region (e.g., a genomic region) identified as having a fetal copy number variation. The fetal fraction for a copy number variation region may be determined using any suitable method for quantifying fetal nucleic acids in a mixture of maternal and fetal nucleic acids. For example, the fetal fraction for the copy number variation region can be determined according to sequencing-based fetal fraction (SeqFF) estimation. Methods for determining fetal fraction according to sequencing-based fetal fraction (SeqFF) estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is hereby incorporated by reference. Sequencing-based fetal fraction (SeqFF) estimation may be referred to as bin-based fetal fraction (BFF) estimation and / or part-specific fetal fraction estimation. In some embodiments, the fetal fraction for a copy number variation region may be determined according to the allele ratio of a polymorphic sequence in fetal nucleic acid and maternal nucleic acid. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining the fetal fraction according to the allele ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference. In some embodiments, the fetal fraction for a copy number variation region may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated fetal nucleic acid and maternal nucleic acid). Methods for determining the fetal fraction according to the quantification of differentially methylated fetal nucleic acid and maternal nucleic acid are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference.

[0147] In some embodiments, the methods herein include determining the fraction of minority nucleic acids in the sample nucleic acid. Determining the fraction of minority nucleic acids in the sample nucleic acid generally involves methods, such as those described above, that quantify nucleic acid species based on information about regions identified as having copy number variations. Rather, determining the fraction of minority nucleic acids in the sample nucleic acid may include methods that quantify minority nucleic acids according to information from regions across the genome and / or regions different from the regions identified as having copy number variations. In some embodiments, the fraction of minority nucleic acids is determined for genomic regions that are larger than the copy number variation region. For example, the fraction of minority nucleic acids may be determined for genomic regions that contain more genomic content (e.g., base pairs, kilobases, megabases) than the regions identified as having copy number variations. For example, for a sample in which minority nucleic acids are identified as having trisomy 21, the fraction of minority nucleic acids may be determined according to information from or associated with multiple chromosomes (e.g., sequence information, sequence read quantification, polymorphism sequences, differentially methylated sequences). In this example, the plurality of chromosomes may include all chromosomes, all autosomes, a subset of chromosomes, a subset of autosomes, a subset of chromosomes including chromosome 21, a subset of autosomes including chromosome 21, a subset of chromosomes excluding chromosome 21, a subset of autosomes excluding chromosome 21, or a portion thereof. In some embodiments, the fraction of minority nucleic acids is determined for a genomic region different from the copy number variation region. For example, for a sample in which minority nucleic acids are identified as having trisomy of chromosome 21, the fraction of minority nucleic acids can be determined according to information from or related to chromosomes other than chromosome 21 (for example, sequence information, sequence read quantification, polymorphism sequence, differentially methylated sequence).

[0148] The minority nucleic acid fraction in sample nucleic acid can be determined by any suitable method for quantifying the nucleic acid species in nucleic acid mixture.For example, the minority nucleic acid fraction can be determined according to sequencing-based fraction estimation.The method for determining minority nucleic acid fraction according to sequencing-based fraction estimation is described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, and these are each hereby incorporated by reference herein. Sequencing-based fraction estimation may be referred to as bin-based fraction estimation and / or portion-specific fraction estimation. In some embodiments, the fraction of minority nucleic acids may be determined according to the allelic ratio of polymorphic sequences. Polymorphic sequences may include, for example, single nucleotide polymorphisms (SNPs). Methods for determining the minority nucleic acid fraction according to the allelic ratio of polymorphic sequences are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference herein. In some embodiments, the fraction of minority nucleic acids may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated nucleic acids). Methods for determining the minority nucleic acid fraction according to the quantification of differentially methylated nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference herein.

[0149] In some embodiments, the minority nucleic acid comprises a fetal nucleic acid. Thus, in some embodiments, the method herein comprises determining a fetal fraction. The fetal fraction can be determined using any suitable method for quantifying fetal nucleic acid in a mixture of maternal and fetal nucleic acids. For example, the fetal fraction can be determined according to sequencing-based fetal fraction (SeqFF) estimation. Methods for determining fetal fraction according to sequencing-based fetal fraction (SeqFF) estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is incorporated herein by reference. The present invention relates to a method for determining the fetal fraction of a fetal gene by using a sequencing-based fetal fraction (SeqFF) estimation method. The method may be referred to as a bin-based fetal fraction (BFF) estimation method and / or a site-specific fetal fraction estimation method. In some embodiments, the fetal fraction may be determined according to the allele ratio of a polymorphic sequence in fetal and maternal nucleic acids. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining the fetal fraction according to the allele ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference. In some embodiments, the fetal fraction may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated fetal and maternal nucleic acids). Methods for determining the fetal fraction according to the quantification of differentially methylated fetal and maternal nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference. In some embodiments, fetal fraction can be determined according to a Y chromosome assay. Methods for determining fetal fraction according to a Y chromosome assay are described herein and in Lo YM, et al. (1998) Am J Hum Genet 62:768-775. .

[0150] In some embodiments, the fraction of copy number variation region and the fraction of minority nucleic acid are determined using the same methodology.For example, the fraction of copy number variation region and the fraction of minority nucleic acid can be determined according to sequencing-based fraction estimation.In some embodiments, the fraction of copy number variation region and the fraction of minority nucleic acid are determined using different methodologies.For example, the fraction of copy number variation region can be determined according to the allele ratio of polymorphic sequence, and the fraction of minority nucleic acid can be determined according to differential epigenetic biomarkers.

[0151] In some embodiments, the fetal fraction of copy number variation region and the fetal fraction of nucleic acid sample are determined using the same methodology.For example, the fetal fraction of copy number variation region and the fetal fraction of nucleic acid sample can be determined according to sequencing-based fetal fraction estimation.In some embodiments, the fetal fraction of copy number variation region and the fetal fraction of nucleic acid sample are determined using different methodologies.For example, the fetal fraction of copy number variation region can be determined according to the allele ratio of polymorphism sequence, and the fetal fraction of nucleic acid sample can be determined according to Y chromosome assay.

[0152] In some embodiments, the fraction for copy number variation (e.g., copy number variation region) is determined for a chromosome or a portion thereof. The fraction for copy number variation determined for a chromosome or a portion thereof refers to quantification of nucleic acid species based on information from or associated with the chromosome or a portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequence, differentially methylated sequence). In some embodiments, the fraction for copy number variation (e.g., copy number variation region) is determined for chromosome 13, chromosome 18, or chromosome 21. In some embodiments, the fraction of minority nucleic acids is determined for a chromosome or a portion thereof different from the chromosome or portion thereof used to determine the fraction for copy number variation. In some embodiments, the fraction of minority nucleic acids is determined for multiple chromosomes or multiple portions of chromosomes. In some embodiments, the fraction of minority nucleic acids is determined for multiple autosomes or multiple portions of autosomes. In some embodiments, the fraction of minority nucleic acids is determined for multiple regions (e.g., genomic regions). In some embodiments, the fraction of minority nucleic acids is determined for multiple regions (e.g., genomic regions) genome-wide.

[0153] In some embodiments, the fetal fraction for a copy number variation (e.g., a copy number variation region) is determined for a chromosome or a portion thereof. The fetal fraction for a copy number variation determined for a chromosome or a portion thereof refers to quantification of fetal nucleic acid based on information from or associated with the chromosome or a portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequence, differentially methylated sequence). In some embodiments, the fetal fraction for a copy number variation (e.g., a copy number variation region) is determined for chromosome 13, chromosome 18, or chromosome 21. In some embodiments, the fetal fraction for the sample nucleic acid is determined for a chromosome or a portion thereof that is different from the chromosome or portion thereof used to determine the fetal fraction for the copy number variation. In some embodiments, the fetal fraction for the sample nucleic acid is determined for multiple chromosomes or multiple portions of chromosomes. In some embodiments, the fetal fraction for the sample nucleic acid is determined for multiple autosomes or multiple portions of autosomes. In some embodiments, the fetal fraction for the sample nucleic acid is determined for multiple regions (e.g., genomic regions). In some embodiments, the fetal fraction for the sample nucleic acid is determined for multiple regions genome-wide (e.g., genomic regions).

[0154] In some embodiments, the method herein comprises comparing the fraction of copy number variation with the fraction of minority nucleic acid. In some embodiments, comparing the fraction of copy number variation with the fraction of minority nucleic acid comprises generating a ratio. For example, the ratio can be the fraction of nucleic acid with copy number variation divided by the fraction of minority nucleic acid.

[0155] In some embodiments, the methods herein include comparing the fetal fraction for the copy number variation with the fetal fraction for the sample nucleic acid. In some embodiments, comparing the fetal fraction for the copy number variation with the fetal fraction for the sample nucleic acid includes generating a ratio. For example, the ratio can be the fetal fraction for the copy number variation divided by the fetal fraction for the sample nucleic acid.

[0156] In some embodiments, the method herein comprises classifying the presence or absence of genetic mosaicism for copy number variation regions in one or more fetuses.The presence or absence of genetic mosaicism for copy number variation regions in one or more fetuses can be classified according to the comparison.For example, the presence or absence of genetic mosaicism for copy number variation regions in one or more fetuses can be classified according to the comparison of the fraction for copy number variation and the fraction of minority nucleic acid.In some embodiments, the presence or absence of genetic mosaicism for copy number variation regions in one or more fetuses can be classified according to the comparison of the fetal fraction for copy number variation and the fetal fraction for sample nucleic acid.The presence or absence of genetic mosaicism for copy number variation regions in one or more fetuses can be classified according to the ratio.For example, the presence or absence of genetic mosaicism for copy number variation regions in one or more fetuses can be classified according to the ratio of the fraction for copy number variation to the fraction of minority nucleic acid (for example, the fraction for copy number variation divided by the fraction of minority nucleic acid). In some embodiments, the presence or absence of genetic mosaicism for the copy number variation region in one or more fetuses may be classified according to the ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid (e.g., the fetal fraction for the copy number variation divided by the fetal fraction for the sample nucleic acid).

[0157] In some embodiments, the presence of genetic mosaicism is classified for copy number variation regions in one or more fetuses.The presence of genetic mosaicism classification for copy number variation regions in one or more fetuses can be interpreted as mosaic copy number variation, affected fetus, unaffected fetus, partially affected fetus, fetal copy number variation, partial fetal copy number variation, partial copy number variation, placental copy number variation, partial placental copy number variation, incomplete copy number variation, placental mosaicism, placental limited mosaicism (CPM) etc.

[0158] In some embodiments, the presence of genetic mosaicism is classified for a copy number variation region in one or more fetuses of a multiple pregnancy based on: (i) the mosaicism ratio value of the fraction of copy number variation (e.g., fetal fraction) to the fraction of minority nucleic acid (e.g., fetal nucleic acid), and (ii) the number of fetuses carried by the pregnant female. For example, the presence of genetic mosaicism can be classified for a copy number variation region in one fetus of a pregnant female carrying twins if the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is less than about 0.7, e.g., 0.54, 0.44, or 0.6. Alternatively, the presence of genetic mosaicism can be classified for a copy number variation region in both fetuses of a pregnant female carrying twins if the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is greater than about 0.9, e.g., 1.17. Alternatively, the presence of genetic mosaicism can be classified for a copy number variation region in one fetus of a pregnant woman carrying triplets when the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is less than about 0.4, for example, 0.33. Alternatively, the presence of genetic mosaicism can be classified for a copy number variation region in two fetuses of a pregnant woman carrying triplets when the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is between about 0.4 and about 0.8, for example, 0.62. It should be understood that the mosaic ratio value needs to be interpreted taking into account the number of fetuses carried by the pregnant woman.

[0159] In some embodiments, the absence of genetic mosaicism is classified for the copy number variation region. The absence of genetic mosaicism classification for the copy number variation region can be interpreted as a standard positive result (e.g., a positive result for fetal copy number variation), an affected fetus, a fetal copy number variation, a complete copy number variation, a true copy number variation, a complete copy number variation, etc.

[0160] In some embodiments, the absence of genetic mosaicism is classified for a copy number variation region in one or more fetuses if the mosaicism ratio of the fraction for copy number variation (e.g., fetal fraction) to the fraction for minority nucleic acid (e.g., fetal nucleic acid) is greater than 1.3. For example, the absence of genetic mosaicism can be classified for a copy number variation region in one or more fetuses if the mosaicism ratio of the fraction for copy number variation to the fraction for minority nucleic acid is between about 1.3 and about 1.7, or between about 1.3 and about 1.5. In some embodiments, the absence of genetic mosaicism is classified for a copy number variation region in one or more fetuses if the mosaicism ratio of the fraction for copy number variation to the fraction for minority nucleic acid is between about 1.31 and about 1.7. For example, the absence of genetic mosaicism can be classified for a region of copy number variation in one or more fetuses if the mosaicism ratio value of the fraction for copy number variation to the fraction for the minority nucleic acid is about 1.31, 1.4, 1.5, 1.6, or 1.7.

[0161] In some embodiments, no classification is provided. For example, if the mosaicism ratio value of the fraction of copy number variation (e.g., fetal fraction) to the fraction of minority nucleic acid (e.g., fetal nucleic acid) is below a certain threshold, no classification (e.g., no call, no clinical relevance) can be provided. In some embodiments, if the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is about 0.1 or less, no classification is provided. In some embodiments, if the mosaicism ratio value of the fraction of copy number variation to the fraction of minority nucleic acid is about 0.1 or less, no classification is provided.

[0162] In some embodiments, if the mosaicism ratio value of the copy number variation fraction (e.g., fetal fraction) to the fraction of minority nucleic acid (e.g., fetal nucleic acid) is greater than a certain threshold, no classification is provided.For example, if the mosaicism ratio value of the copy number variation fraction to the fraction of minority nucleic acid is about 1.7, 1.8, 1.9, 2.0, or 2.5 or greater, no classification can be provided.In some embodiments, if the mosaicism ratio value of the copy number variation fraction to the fraction of minority nucleic acid is about 1.7 or greater, no classification is provided.A value greater than a certain threshold (e.g., greater than 1.7) can indicate copy number variation (e.g., maternal copy number variation) present in the majority nucleic acid.

[0163] FIG. 4 shows a process 400 for classifying the presence or absence of genetic mosaicism for a biological sample and providing clinical interpretation and / or diagnostic follow-up information, according to various embodiments. A set of sequence reads is provided, and a screening test for a genetic condition (e.g., NIPT) is obtained from the set of sequence reads in block 405. The sequence reads may be obtained from circulating cell-free sample nucleic acids from a test sample obtained from a multifetal pregnant subject (e.g., a pregnant female subject having multiple fetuses). The circulating cell-free sample nucleic acids may include maternal nucleic acids and fetal nucleic acids. The circulating cell-free sample nucleic acids may be captured by probe oligonucleotides under hybridization conditions. In various embodiments, the genetic condition being screened for includes the presence of one or more aneuploidies, e.g., copy number variations. Additionally, the number of fetuses carried by the multifetal pregnant subject is obtained.

[0164] The presence (flag as positive) or absence (flag as negative) of one or more aneuploidies can be identified in the circulating cell-free nucleic acids from the set of sequence reads based on the z-score in block 410 or 415. If the absence (flag as negative) of one or more aneuploidies is identified, no further testing can be performed in block 420, or diagnostic testing can be performed in block 425. If the presence (flag as positive) of one or more aneuploidies is identified, a mosaicism ratio is generated as described with respect to FIG. 2, and the mosaicism ratio value is used to classify the presence or absence of genetic mosaicism in one or more fetuses and provide improved interpretation of NIPT results. The mosaicism ratio can be used to identify patients with a higher likelihood of discordant positive results due to mosaicism (e.g., CPM).

[0165] The presence or absence of genetic mosaicism can be classified in blocks 430 and 435 for copy number variation regions in one or more fetuses of a multiple pregnancy based on: (i) the mosaicism ratio value of the fraction for copy number variation (e.g., fetal fraction) to the fraction of minority nucleic acid (e.g., fetal nucleic acid), and (ii) the number of fetuses carried by the pregnant female. Furthermore, if the mosaicism ratio value is greater than about 1.7 or less than about 0.1, no classification can be provided for the copy number variation region in blocks 440 and 445. If no classification is provided and the mosaicism ratio value is greater than about 1.7, the positive NIPT result can be interpreted as possibly excessive or indeterminate in block 450, and diagnostic follow-up including amniocentesis, CVS, maternal testing, and / or other tests can be recommended in block 455, depending on the consensus decision between the genetic counselor and the physician. If no classification is provided and the mosaicism ratio value is less than about 0.1, the positive NIPT result may be interpreted as a negative result or the absence of one or more aneuploidies in block 460, and diagnostic follow-up may not be invoked in block 465. If the presence of genetic mosaicism is classified for one or more fetuses (e.g., the mosaicism ratio is between about 0.1 and about 1.7, depending on the number of fetuses), the positive NIPT result may be interpreted as positive in block 470, with the understanding that a mosaic comment, a possible mosaic presentation for one or more fetuses, exists, and diagnostic follow-up including amniocentesis and / or CVS may be recommended in block 475, depending on the consensus decision between the genetic counselor and the physician. If the absence of genetic mosaicism is classified (e.g., the mosaicism ratio is greater than about 1.0, depending on the number of fetuses), the positive NIPT result may be interpreted as positive in block 480, and diagnostic follow-up including amniocentesis and / or CVS may be recommended for confirmation in block 485. Sex classification for one or more fetuses

[0166] Provided herein are methods for classifying the sex of one or more fetuses in a sample (e.g., a biological sample; a test sample). In various embodiments, the sex of one or more fetuses is classified according to the level of the Y chromosome (e.g., at one or more genomic region levels, at the profile level) and the number of fetuses carried by the pregnant female. In some embodiments, the methods herein include determining the fraction of nucleic acids bearing a Y chromosome in a sample of nucleic acids from a subject with a multifetal pregnancy. Determining the fraction of nucleic acids refers to quantifying a specific species of nucleic acid in a nucleic acid mixture. For example, determining the fraction of nucleic acids can refer to quantifying minority nucleic acid species, quantifying fetal nucleic acids, quantifying cancer nucleic acids, etc. Determining the fraction of nucleic acids bearing a Y chromosome refers to quantifying a subset of nucleic acids (e.g., a subset of nucleic acid fragments, a subset of sequence reads) for which a Y chromosome is identified. In some embodiments, the fraction of nucleic acids bearing a Y chromosome is determined in part according to the level of the Y chromosome (e.g., at one or more genomic region levels, at the profile level). In some embodiments, determining the fraction of nucleic acids bearing a Y chromosome refers to quantifying a subset of nucleic acids (e.g., a subset of nucleic acid fragments, a subset of sequence reads) from a region (e.g., a genomic region) for which a Y chromosome is identified. In some embodiments, determining the fraction of nucleic acids bearing a Y chromosome refers to quantifying a subset of nucleic acids for a species (e.g., a subset of nucleic acid fragments for a species, a subset of sequence reads for a species) from a region (e.g., a genomic region) for which a Y chromosome is identified. For example, for a sample containing maternal and fetal nucleic acids from a subject with a multiple pregnancy, if the fetal nucleic acid is identified as bearing a Y chromosome, determining the fraction of nucleic acids bearing a Y chromosome refers to determining the fetal fraction based on information from or associated with the Y chromosome or a portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequences, differentially methylated sequences).

[0167] In some embodiments, the methods herein include determining a fraction for a region (e.g., a genomic region). In some embodiments, the methods herein include determining a fraction for a region of a Y chromosome. The fraction for a region determined for a Y chromosome or portion thereof refers to quantification of nucleic acid species based on information from or associated with the Y chromosome or portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequences, differentially methylated sequences). In some embodiments, the fraction for a region is determined for the Y chromosome. In some embodiments, the fraction of a minority nucleic acid is determined for a Y chromosome or portion thereof that is different from the X chromosome or portion thereof used to determine the fraction for the region associated with the Y chromosome.

[0168] As discussed above, the fraction of a region of the Y chromosome can be determined according to information (e.g., sequence information, epigenetic information) obtained about a region (e.g., a genomic region) identified as being associated with the Y chromosome. The fraction of a region of the Y chromosome can be determined using any suitable method for quantifying nucleic acid species in a nucleic acid mixture. For example, the fraction of a region of the Y chromosome can be determined according to sequencing-based fraction estimation. Methods for determining nucleic acid fraction according to sequencing-based fraction estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is hereby incorporated by reference. The sequencing-based fraction estimation may be referred to as bin-based fraction estimation and / or part-specific fraction estimation. In some embodiments, the fraction for a region of the Y chromosome may be determined according to the allelic ratio of a polymorphic sequence. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining nucleic acid fractions according to the allelic ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference. In some embodiments, the fraction for a region of the Y chromosome may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated nucleic acids). Methods for determining nucleic acid fractions according to quantification of differentially methylated nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference.

[0169] In some embodiments, the nucleic acid sample includes a majority nucleic acid (e.g., more than the minority nucleic acid) and a minority nucleic acid (e.g., less than the majority nucleic acid). In some embodiments, the majority nucleic acid includes maternal nucleic acid, and the minority nucleic acid includes fetal nucleic acid. Thus, in some embodiments, the methods herein include determining a fetal fraction. In some embodiments, the methods herein include determining a fetal fraction for a region (e.g., a genomic region) identified as associated with the Y chromosome. In some embodiments, the methods herein include determining a fetal fraction for a region of the Y chromosome. As discussed above, the fetal fraction for a region of the Y chromosome can be determined according to information (e.g., sequence information, epigenetic information) obtained for the region (e.g., a genomic region) identified as associated with the Y chromosome. The fetal fraction for a region of the Y chromosome can be determined using any suitable method for quantifying fetal nucleic acid in a mixture of maternal and fetal nucleic acids. For example, the fetal fraction for a region of the Y chromosome can be determined according to sequencing-based fetal fraction (SeqFF) estimation. Methods for determining fetal fraction according to sequencing-based fetal fraction (SeqFF) estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815. are each hereby incorporated by reference herein. Sequencing-based fetal fraction (SeqFF) estimation may be referred to as bin-based fetal fraction (BFF) estimation and / or region-specific fetal fraction estimation. In some embodiments, the fetal fraction for a region of the Y chromosome may be determined according to the allele ratio of a polymorphic sequence in fetal and maternal nucleic acids. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining the fetal fraction according to the allele ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference herein. In some embodiments, the fetal fraction for a region of the Y chromosome may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated fetal and maternal nucleic acids). Methods for determining fetal fraction according to quantification of differentially methylated fetal and maternal nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference herein.

[0170] In some embodiments, the methods herein include determining the fraction of minority nucleic acids in the sample nucleic acid. Determining the fraction of minority nucleic acids in the sample nucleic acid is generally not limited to methods, such as those described above, that quantify nucleic acid species based on information about regions identified as associated with the Y chromosome. Rather, determining the fraction of minority nucleic acids in the sample nucleic acid may include methods that quantify minority nucleic acids according to information from regions across the genome and / or regions different from the regions identified as associated with the Y chromosome. In some embodiments, the fraction of minority nucleic acids is determined for a genomic region that is larger than the region of the Y chromosome. For example, the fraction of minority nucleic acids can be determined for a genomic region that contains more genomic content (e.g., base pairs, kilobases, megabases) than the region of the Y chromosome. For example, for a sample identified as having a minority nucleic acid with trisomy 21 of chromosome 21, the fraction of minority nucleic acids can be determined according to information from or associated with multiple chromosomes (e.g., sequence information, sequence read quantification, polymorphism sequences, differentially methylated sequences). In this example, such a plurality of chromosomes can include all chromosomes, all autosomes, a subset of chromosomes, a subset of autosomes, a subset of chromosomes that includes chromosome Y, a subset of autosomes, a subset of chromosomes that excludes chromosome Y, a subset of chromosomes that includes chromosome X, or a portion thereof. In some embodiments, the fraction of minority nucleic acids is determined for a genomic region that is different from the region of chromosome Y. For example, for a sample that is identified as having minority nucleic acids in the region of chromosome Y, the fraction of minority nucleic acids can be determined according to information from or related to chromosomes other than chromosome Y (for example, sequence information, sequence read quantification, polymorphism sequence, differentially methylated sequence).

[0171] The minority nucleic acid fraction in sample nucleic acid can be determined by any suitable method for quantifying the nucleic acid species in nucleic acid mixture.For example, the minority nucleic acid fraction can be determined according to sequencing-based fraction estimation.The method for determining minority nucleic acid fraction according to sequencing-based fraction estimation is described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, and these are each hereby incorporated by reference herein. Sequencing-based fraction estimation may be referred to as bin-based fraction estimation and / or portion-specific fraction estimation. In some embodiments, the fraction of minority nucleic acids may be determined according to the allelic ratio of polymorphic sequences. Polymorphic sequences may include, for example, single nucleotide polymorphisms (SNPs). Methods for determining the minority nucleic acid fraction according to the allelic ratio of polymorphic sequences are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference herein. In some embodiments, the fraction of minority nucleic acids may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated nucleic acids). Methods for determining the minority nucleic acid fraction according to the quantification of differentially methylated nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference herein.

[0172] In some embodiments, the minority nucleic acid comprises a fetal nucleic acid. Thus, in some embodiments, the method herein comprises determining a fetal fraction. The fetal fraction can be determined using any suitable method for quantifying fetal nucleic acid in a mixture of maternal and fetal nucleic acids. For example, the fetal fraction can be determined according to sequencing-based fetal fraction (SeqFF) estimation. Methods for determining fetal fraction according to sequencing-based fetal fraction (SeqFF) estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is incorporated herein by reference. The present invention relates to a method for determining the fetal fraction of a fetal gene by using a sequencing-based fetal fraction (SeqFF) estimation method. The method may be referred to as a bin-based fetal fraction (BFF) estimation method and / or a site-specific fetal fraction estimation method. In some embodiments, the fetal fraction may be determined according to the allele ratio of a polymorphic sequence in fetal and maternal nucleic acids. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining the fetal fraction according to the allele ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference. In some embodiments, the fetal fraction may be determined according to differential epigenetic biomarkers (e.g., quantification of differentially methylated fetal and maternal nucleic acids). Methods for determining the fetal fraction according to the quantification of differentially methylated fetal and maternal nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference. In some embodiments, fetal fraction can be determined according to a Y chromosome assay. Methods for determining fetal fraction according to a Y chromosome assay are described herein and in Lo YM, et al. (1998) Am J Hum Genet 62:768-775. .

[0173] In some embodiments, the fraction of Y chromosome (or Y chromosome region) and the fraction of minority nucleic acid are determined using the same methodology.For example, the fraction of Y chromosome (or Y chromosome region) and the fraction of minority nucleic acid can be determined according to sequencing-based fraction estimation, respectively.In some embodiments, the fraction of Y chromosome (or Y chromosome region) and the fraction of minority nucleic acid are determined using different methodologies.For example, the fraction of Y chromosome (or Y chromosome region) can be determined according to the allele ratio of polymorphic sequence, and the fraction of minority nucleic acid can be determined according to differential epigenetic biomarkers.

[0174] In some embodiments, the fetal fraction for Y chromosome (or Y chromosome region) and the fetal fraction for nucleic acid sample are determined using the same methodology.For example, the fetal fraction for Y chromosome (or Y chromosome region) and the fetal fraction for nucleic acid sample can each be determined according to sequencing-based fetal fraction estimation.In some embodiments, the fetal fraction for Y chromosome (or Y chromosome region) and the fetal fraction for nucleic acid sample are determined using different methodologies.For example, the fetal fraction for Y chromosome (or Y chromosome region) can be determined according to the allele ratio of polymorphic sequence, and the fetal fraction for nucleic acid sample can be determined according to Y chromosome assay.

[0175] In some embodiments, the method herein comprises comparing the fraction of Y chromosome (or region of Y chromosome) with the fraction of minority nucleic acid. In some embodiments, comparing the fraction of Y chromosome (or region of Y chromosome) with the fraction of minority nucleic acid comprises generating a ratio. For example, the ratio can be the fraction of nucleic acid with Y chromosome (or region of Y chromosome) divided by the fraction of minority nucleic acid.

[0176] In some embodiments, the methods herein include comparing the fetal fraction for the Y chromosome (or a region of the Y chromosome) with the fetal fraction for the sample nucleic acid. In some embodiments, comparing the fetal fraction for the Y chromosome (or a region of the Y chromosome) with the fetal fraction for the sample nucleic acid includes generating a ratio. For example, the ratio can be the fetal fraction for the Y chromosome (or a region of the Y chromosome) divided by the fetal fraction for the sample nucleic acid.

[0177] In some embodiments, the methods herein include classifying the sex of one or more fetuses based on the mosaicism ratio. The fraction of nucleic acids with a Y chromosome in the sample nucleic acid can be compared with the fraction of fetal nucleic acid in the sample nucleic acid, thereby providing a comparison and generating a mosaicism ratio. In some embodiments, the sex is classified for a fetus based on the mosaicism ratio of the fraction of nucleic acids with a Y chromosome to the fraction of fetal nucleic acid. For example, the sex of one or more fetuses can be classified according to the ratio of the fraction for the Y chromosome (or a region of the Y chromosome) to the fraction of the minority nucleic acid (e.g., the fraction for the Y chromosome (or a region of the Y chromosome) divided by the fraction of the minority nucleic acid). In some embodiments, the sex of one or more fetuses can be classified according to the ratio of the fetal fraction for the Y chromosome (or a region of the Y chromosome) to the fetal fraction for the sample nucleic acid (e.g., the fetal fraction for the Y chromosome (or a region of the Y chromosome) divided by the fetal fraction for the sample nucleic acid).

[0178] In some embodiments, the sex of the fetuses in a multiple pregnancy is classified based on: (i) the mosaicism ratio of the fraction (e.g., fetal fraction) for the Y chromosome (or region of the Y chromosome) to the fraction of minority nucleic acids (e.g., fetal nucleic acids), and (ii) the number of fetuses carried by the pregnant female.

[0179] For example, the sex of the fetuses in a multifetal pregnancy can be classified as one male and one female for a pregnant woman carrying twins if the mosaicism ratio of circulating cell-free nucleic acid having a Y chromosome (or a region of the Y chromosome) to the fraction of fetal nucleic acid is between about 0.4 and 0.7. Alternatively, the sex of the fetuses in a multifetal pregnancy can be classified as both female for a pregnant woman carrying twins if the mosaicism ratio of circulating cell-free nucleic acid having a Y chromosome (or a region of the Y chromosome) to the fraction of fetal nucleic acid is less than about 0.2. Alternatively, the sex of the fetuses in a multifetal pregnancy can be classified as both male for a pregnant woman carrying twins if the mosaicism ratio of circulating cell-free nucleic acid having a Y chromosome (or a region of the Y chromosome) to the fraction of fetal nucleic acid is greater than about 1.0. Alternatively, the fetal sex of a multiple pregnancy can be classified as one male and two female for a pregnant female carrying triplets if the mosaicism ratio value of circulating cell-free nucleic acid having a Y chromosome (or a region of a Y chromosome) to the fraction of fetal nucleic acid is between about 0.12 and about 0.4. Alternatively, the fetal sex of a multiple pregnancy can be classified as three female for a pregnant female carrying triplets if the mosaicism ratio value of circulating cell-free nucleic acid having a Y chromosome (or a region of a Y chromosome) to the fraction of fetal nucleic acid is less than about 0.1.

[0180] FIG. 5 illustrates a process 500 for classifying the sex of one or more fetuses for a biological sample, according to various embodiments. A set of sequence reads is provided in block 505. The sequence reads may be obtained from circulating cell-free sample nucleic acids from a test sample obtained from a multifetal pregnant subject (e.g., a pregnant female subject having multiple fetuses). The circulating cell-free nucleic acids may include maternal nucleic acids and fetal nucleic acids. The circulating cell-free sample nucleic acids may be captured by probe oligonucleotides under hybridization conditions. Furthermore, the number of fetuses carried by the multifetal pregnant subject is obtained. In block 510, a Y chromosome (or a region of the Y chromosome) is identified in the circulating cell-free nucleic acids from the set of sequence reads. The fraction of circulating cell-free nucleic acids having a Y chromosome (or a region of the Y chromosome) in the sample nucleic acid is determined in block 515. The fraction may be a fetal fraction determined for the Y chromosome (or a region of the Y chromosome). The fraction of fetal nucleic acids in the circulating cell-free sample nucleic acid is determined in block 520. The fraction of circulating cell-free nucleic acid having a Y chromosome (or a region of a Y chromosome) is compared to the fraction of fetal nucleic acid in block 525 to generate a mosaicism ratio of the fraction of circulating cell-free nucleic acid having a Y chromosome (or a region of a Y chromosome) to the fraction of fetal nucleic acid. The sex of one or more fetuses is classified according to the mosaicism ratio in block 530 to obtain the number of fetuses carried by the multiple-fetal pregnancy subject. sample

[0181] Provided herein are systems, methods and products for analyzing nucleic acid.In some embodiments, the nucleic acid fragment in a mixture of nucleic acid fragments is analyzed.Nucleic acid fragments can be called nucleic acid templates, and these terms can be used interchangeably herein.Nucleic acid mixtures can comprise two or more nucleic acid fragment species with the same or different nucleotide sequences, different fragment lengths, different origins (for example, genomic origin, fetal origin vs. maternal origin, cell or tissue origin, cancer vs. non-cancer origin, tumor vs. non-tumor origin, sample origin, subject origin, etc.), or combinations thereof.

[0182] Nucleic acids or mixtures of nucleic acids utilized in the systems, methods, and products described herein are often isolated from samples obtained from subjects (e.g., test subjects), including humans, non-human animals, plants, bacteria, fungi, protozoa, or pathogens. The subject may be any living or non-living organism, including, but not limited to, any human or non-human animal. Any human or non-human animal may be selected, including, for example, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, cattle (e.g., cows), equines (e.g., horses), goats and ovines (e.g., sheep, goats), swine (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursinians (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The subject may be male or female (e.g., female, pregnant female). The subject may be of any age (e.g., embryo, fetus, infant, child, adult). The subject may be a cancer patient, a patient suspected of having cancer, a patient in remission, a patient with a family history of cancer, and / or a subject undergoing cancer screening. In some embodiments, the test subject is female. In some embodiments, the test subject is a human female having a multiple pregnancy. In some embodiments, the test subject is male. In some embodiments, the test subject is a human male.

[0183] Nucleic acids can be isolated from any type of suitable biological specimen or sample (e.g., test sample). The sample or test sample can be any specimen isolated or obtained from a subject or a part thereof (e.g., a human subject, a pregnant female, a cancer patient, a fetus, a tumor). The sample is sometimes derived from a pregnant female subject having a fetus at any stage of pregnancy (e.g., for a human subject, the first, second, or third trimester), and sometimes from a postnatal subject. The sample is sometimes derived from a pregnant subject having one or more fetuses that are euploid for all chromosomes, and sometimes from a pregnant subject having one or more fetuses with chromosomal aneuploidy (e.g., one, three (i.e., trisomy (e.g., T21, T18, T13)) or four copies of a chromosome) or other genetic mutation. Non-limiting examples of specimens include fluids or tissues from a subject, including, but not limited to, blood or blood products (e.g., serum, plasma, etc.), umbilical cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., bronchoalveolar, stomach, peritoneal cavity, duct, ear, arthroscope), biopsy samples (e.g., from preimplantation embryos; cancer biopsies), celiac cavity aspirates (celocentesis samples), cells (blood cells, placental cells, embryonic or fetal cells, nucleated fetal cells or fetal cellular remnants, normal cells, abnormal cells (e.g., cancer cells)) or portions thereof (e.g., mitochondria, nuclei, extracts, etc.), female reproductive tract washings, urine, feces, sputum, saliva, nasal mucosa, prostatic fluid, lavage, semen, lymph, bile, tears, sweat, breast milk, milk, etc., or combinations thereof. In some embodiments, the biological sample is a cervical swab from a subject. The fluid or tissue sample from which nucleic acids are extracted can be acellular (e.g., cell-free). In some embodiments, the fluid or tissue sample can contain cellular elements or cellular remnants. In some embodiments, fetal cells or cancer cells can be included in the sample.

[0184] The sample may be a liquid sample. The liquid sample may contain extracellular nucleic acids (e.g., circulating cell-free DNA). Non-limiting examples of liquid samples include blood or blood products (e.g., serum, plasma, etc.), urine, biopsy samples (e.g., liquid biopsy for cancer detection), the above liquid samples, etc., or combinations thereof. In certain embodiments, the sample is a liquid biopsy, which generally refers to the evaluation of a liquid sample from a subject for the presence, absence, progression, or remission of a disease (e.g., cancer). Liquid biopsy can be used in conjunction with or as an alternative to solid biopsy (e.g., tumor biopsy). In certain cases, extracellular nucleic acids in a liquid biopsy are analyzed.

[0185] In some embodiments, the biological sample may be blood, plasma, or serum. The term "blood" encompasses conventionally defined whole blood, blood products, or any fraction of blood, such as serum, plasma, or buffy coat. Blood or fractions thereof often contain nucleosomes. Nucleosomes contain nucleic acids and are sometimes acellular or intracellular. Blood also includes buffy coats, which are sometimes isolated using a Ficoll gradient. Buffy coats may contain white blood cells (e.g., leukocytes, T cells, B cells, platelets, etc.). Plasma refers to the fraction of whole blood obtained by centrifugation of anticoagulant-treated blood. Serum refers to the aqueous portion of the fluid remaining after a blood sample has clotted. Fluid or tissue samples are often collected according to standard protocols commonly followed by hospitals or clinics. For blood, an appropriate volume of peripheral blood (e.g., between 3 and 40 milliliters, between 5 and 50 milliliters) is often collected and can be stored according to standard procedures before or after preparation.

[0186] Analysis of nucleic acids found in a subject's blood can be performed using, for example, whole blood, serum, or plasma. For example, analysis of fetal DNA found in maternal blood can be performed using, for example, whole blood, serum, or plasma. For example, analysis of tumor DNA found in a patient's blood can be performed using, for example, whole blood, serum, or plasma. Methods for preparing serum or plasma from blood obtained from a subject (e.g., a maternal subject; a cancer patient) are known. For example, the subject's blood (e.g., the blood of a pregnant woman; the blood of a cancer patient) can be placed in a tube containing EDTA to prevent clotting, or in a specialized commercial product, such as Vacutainer SST (Becton Dickinson, Franklin Lakes, NJ), and plasma can then be obtained from the whole blood via centrifugation. Serum can be obtained with or without centrifugation after clotting. When centrifugation is used, it is typically, but not exclusively, performed at an appropriate speed, for example, 1,500 to 3,000 × g. The plasma or serum may be subjected to an additional centrifugation step before being transferred to a new tube for nucleic acid extraction. In addition to the acellular portion of whole blood, nucleic acids may also be recovered from the cellular fraction and may be enriched in the buffy coat portion that can be obtained after centrifugation of a whole blood sample from a subject and removal of plasma.

[0187] A sample may be heterogeneous. For example, a sample may contain more than one cell type and / or one or more nucleic acid species. In some cases, a sample may contain (i) fetal and maternal cells, (ii) cancer and non-cancer cells, and / or (iii) pathogenic and host cells. In some cases, a sample may contain (i) cancer and non-cancer nucleic acids, (ii) pathogen and host nucleic acids, (iii) fetal-derived and maternal-derived nucleic acids, and / or more generally, (iv) mutated and wild-type nucleic acids. In some cases, a sample may contain minority and majority nucleic acid species, as described in more detail below. In some cases, a sample may contain cells and / or nucleic acids from a single subject, or may contain cells and / or nucleic acids from multiple subjects. cell type

[0188] As used herein, "cell type" refers to a type of cell that can be distinguished from another type of cell. Extracellular nucleic acids can include nucleic acids from several different cell types. Non-limiting examples of cell types that can provide nucleic acids to circulating cell-free nucleic acids include liver cells (e.g., hepatocytes), lung cells, spleen cells, pancreatic cells, colon cells, skin cells, bladder cells, eye cells, brain cells, esophageal cells, head cells, cervical cells, ovarian cells, testicular cells, prostate cells, placental cells, epithelial cells, endothelial cells, adipocytes, kidney / renal cells, cardiac cells, muscle cells, blood cells (e.g., white blood cells), central nervous system (CNS) cells, etc., and combinations thereof. In some embodiments, the cell types that provide nucleic acids to the circulating cell-free nucleic acids being analyzed include white blood cells, endothelial cells, and hepatic liver cells. Different cell types can be screened as part of identifying and selecting nucleic acid loci where the marker status is the same or substantially the same for cell types in subjects with a medical condition and cell types in subjects without a medical condition, as described in further detail herein.

[0189] A particular cell type sometimes remains the same or substantially the same in a subject with a medical condition and a subject without a medical condition. In a non-limiting example, the number of live or viable cells of a particular cell type may be reduced in a cytopathic state in a subject with a medical condition, and the live, viable cells are not altered or not significantly altered.

[0190] A particular cell type is sometimes modified as part of a medical condition and has one or more properties that differ from its original state. In non-limiting examples, a particular cell type may proliferate at a faster than normal rate, transform into a cell with a different morphology, transform into a cell expressing one or more different cell surface markers, and / or become part of a tumor as part of a cancerous condition. In embodiments in which a particular cell type (i.e., a progenitor cell) is modified as part of a medical condition, the marker state for each of one or more markers assayed is often the same or substantially the same for a particular cell type in a subject with a medical condition and a particular cell type in a subject without a medical condition. Thus, the term "cell type" sometimes refers to a type of cell in a subject without a medical condition and to modified versions of the cell in a subject with a medical condition. In some embodiments, a "cell type" refers only to progenitor cells, not modified versions arising from progenitor cells. A "cell type" sometimes refers to progenitor cells and modified cells arising from progenitor cells. In such embodiments, the marker state for the markers being analyzed is often the same or substantially the same for cell types in subjects with the medical condition and cell types in subjects without the medical condition.

[0191] In certain embodiments, the cell type is a cancer cell. Certain cancer cell types include, for example, leukemia cells (e.g., acute myeloid leukemia, acute lymphocytic leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia); cancerous kidney / renal cells (e.g., renal cell carcinoma (clear cell, papillary 1, papillary 2, chromophobe, oncocytic, collecting duct), renal adenocarcinoma, Grawitz tumor, Wilms tumor, transitional cell carcinoma); brain tumor cells (e.g., acoustic neuroma, astrocytoma (Grade I; pilocytic astrocytoma, Grade I; Grade I: low-grade astrocytoma, Grade III: anaplastic astrocytoma, Grade IV: glioblastoma (GBM), chordoma, CNS lymphoma, craniopharyngioma, glioma (brain stem glioma, ependymoma, mixed glioma, optic nerve glioma, subependymoma), medulloblastoma, meningioma, metastatic brain tumor, oligodendroglioma, pituitary tumor, primitive neuroectodermal tumor (PNET), schwannoma, juvenile pilocytic astrocytoma (JPA), pineal tumor, rhabdoid tumor).

[0192] Different cell types can be distinguished by any suitable characteristics, including but not limited to, one or more different cell surface markers, one or more different morphological features, one or more different functions, one or more different protein (e.g., histone) modifications, and one or more different nucleic acid markers.Non-limiting examples of nucleic acid markers include single nucleotide polymorphisms (SNPs), the methylation status of nucleic acid loci, short tandem repeats, insertions (e.g., microinsertions), deletions (microdeletions), etc., and combinations thereof.Non-limiting examples of protein (e.g., histone) modifications include acetylation, methylation, ubiquitination, phosphorylation, sumoylation, etc., and combinations thereof.

[0193] As used herein, the term "related cell type" refers to a cell type that has multiple characteristics in common with another cell type, where 75% or more of the cell surface markers are sometimes common to the related cell types (e.g., about 80%, 85%, 90%, or 95% or more of the cell surface markers are common to the related cell types). nucleic acid

[0194] A method for analyzing nucleic acids is provided herein. The terms "nucleic acid," "nucleic acid molecule," "nucleic acid fragment," and "nucleic acid template" can be used interchangeably throughout this disclosure. These terms refer to nucleic acids of any composition, including DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA)), RNA (e.g., message RNA (mRNA), small inhibitory RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed by fetus or placenta), and / or DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and / or non-native backbones), RNA / DNA hybrids, and polyamide nucleic acids (PNAs), all of which can be in single-stranded or double-stranded form, and can include known analogs of natural nucleotides that can function in a manner similar to naturally occurring nucleotides, unless otherwise limited. In certain embodiments, a nucleic acid may be or may be derived from a plasmid, phage, virus, bacterium, autonomously replicating sequence (ARS), mitochondrion, centromere, artificial chromosome, chromosome, or other nucleic acid capable of replicating or being replicated in vitro or in a host cell, cell, cell nucleus, or cell cytoplasm. In some embodiments, the template nucleic acid may be derived from a single chromosome (e.g., a nucleic acid sample may be derived from one chromosome of a sample obtained from a diploid organism). Unless otherwise specified, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses its conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences, as well as the explicitly indicated sequence. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. The term nucleic acid is used interchangeably with locus, gene, cDNA, and mRNA encoded by a gene.The term may also include RNA or DNA equivalents, derivatives, variants, and analogs synthesized from nucleotide analogs, single-stranded ("sense" or "antisense," "plus" or "minus" strand, "forward" or "reverse" reading frame) and double-stranded polynucleotides. The term "gene" refers to a segment of DNA involved in producing a polypeptide chain, and generally includes regions preceding and following the coding region (leader and trailer), as well as intervening sequences (introns) between individual coding regions (exons), which are involved in the transcription / translation of the gene product and the regulation of transcription / translation. Nucleotides or bases generally refer to the purine and pyrimidine molecular units of nucleic acids (e.g., adenine (A), thymine (T), guanine (G), and cytosine (C)). For RNA, the base thymine is replaced by uracil. The length or size of a nucleic acid may be expressed as the number of bases.

[0195] The nucleic acid can be single-stranded or double-stranded. For example, single-stranded DNA can be generated by denaturing double-stranded DNA, for example, by heating or treating with alkali. In certain embodiments, the nucleic acid is a D-loop structure formed by the strand invasion of double-stranded DNA molecules by oligonucleotides or DNA-like molecules, such as peptide nucleic acids (PNAs). D-loop formation can be promoted, for example, by adding E. coli RecA protein and / or changing salt concentration, using methods known in the art.

[0196] Nucleic acids provided for the processes described herein can contain nucleic acids from one sample or from two or more samples (e.g., from 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, or 20 or more samples).

[0197] Nucleic acids can be derived from one or more sources (e.g., biological samples, blood, cells, serum, plasma, buffy coat, urine, lymph, skin, soil, etc.) by methods known in the art. Any suitable method can be used to isolate, extract, and / or purify DNA from a biological sample (e.g., from blood or blood products), non-limiting examples of which include methods for DNA preparation (e.g., as described by Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3d ed., 2001), various commercially available reagents or kits, and the like. Examples of suitable kits include Qiagen's QIAamp Circulating Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiagen, Hilden, Germany), GenomicPrep™ Blood DNA Isolation Kit (Promega, Madison, Wis.), and GFX™ Genomic Blood DNA Purification Kit (Amersham, Piscataway, NJ), or a combination thereof.

[0198] In some embodiments, nucleic acids are extracted from cells using a cell lysis procedure. Cell lysis procedures and reagents are known in the art and can generally be carried out by chemical methods (e.g., detergents, hypotonic solutions, enzymatic procedures, etc., or a combination thereof), physical methods (e.g., French press, sonication, etc.), or electrolytic lysis methods. Any suitable lysis procedure can be used. For example, chemical methods generally employ a lysis agent to disrupt cells and extract nucleic acids from the cells, followed by treatment with chaotropic salts. Physical methods, such as freeze / thaw followed by crushing, using a cell press, etc., are also useful. In some cases, high salt and / or alkaline lysis procedures can be used.

[0199] Nucleic acids include extracellular nucleic acids in certain embodiments. As used herein, the term "extracellular nucleic acid" may refer to nucleic acids isolated from a source substantially free of cells, and is also referred to as "cell-free" nucleic acids, "circulating cell-free nucleic acids" (e.g., CCF fragments, ccf DNA), and / or "cell-free circulating nucleic acids." Extracellular nucleic acids may be present in or obtained from blood (e.g., the blood of a human subject). Extracellular nucleic acids often do not contain detectable cells and may contain cellular elements or cellular remnants. Non-limiting examples of cell-free sources for extracellular nucleic acids are blood, plasma, serum, and urine. As used herein, the term "obtaining cell-free circulating sample nucleic acids" includes obtaining a sample directly (e.g., collecting a sample, e.g., a test sample) or obtaining a sample from another person who collected the sample. Without being bound by theory, extracellular nucleic acids may be the product of cellular apoptosis and cell destruction, which explains why extracellular nucleic acids often have a range of lengths across a spectrum (e.g., a "ladder"). In some embodiments, the sample nucleic acid from the test subject is circulating cell-free nucleic acid. In some embodiments, the circulating cell-free nucleic acid is from plasma or serum from the test subject.

[0200] Extracellular nucleic acids may contain different nucleic acid species, and therefore, in certain embodiments, are referred to herein as "heterogeneous." For example, serum or plasma from a person with cancer may contain nucleic acids from cancer cells (e.g., tumors, neoplasms) and nucleic acids from non-cancer cells. In another example, serum or plasma from a pregnant female may contain maternal nucleic acids and fetal nucleic acids. In some cases, the cancer or fetal nucleic acid is sometimes about 5% to about 50% of the total nucleic acid (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49% of the total nucleic acid is cancer or fetal nucleic acid).

[0201] At least two different nucleic acid species may be present in different amounts in extracellular nucleic acids, sometimes referred to as minor and major species. In certain cases, the minor nucleic acid is derived from a diseased cell type (e.g., cancer cells, wasted cells, cells attacked by the immune system). In certain cases, the minor nucleic acid is derived from apoptotic cells (e.g., circulating cell-free fetal nucleic acid from apoptotic placental cells). In certain embodiments, genetic mutations or alterations (e.g., copy number alterations, copy number variations, single nucleotide alterations, single nucleotide alterations, chromosomal alterations, and / or translocations) are determined for the minor nucleic acid species. In certain embodiments, genetic mutations or alterations are determined for the majority nucleic acid species. In general, the terms "minority" and "majority" are not intended to be strictly defined in any way. In one aspect, for example, a nucleic acid considered "minor" may have an abundance of at least about 0.1% of the total nucleic acids in the sample to less than 50% of the total nucleic acids in the sample. In some embodiments, the minority nucleic acid may have an abundance of at least about 1% of the total nucleic acids in the sample to about 40% of the total nucleic acids in the sample. In some embodiments, the minority nucleic acid may have an abundance of at least about 2% of the total nucleic acids in the sample to about 30% of the total nucleic acids in the sample. In some embodiments, the minority nucleic acid may have an abundance of at least about 3% of the total nucleic acids in the sample to about 25% of the total nucleic acids in the sample. For example, the minority nucleic acid may have an abundance of about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, or 30% of the total nucleic acids in the sample. In some cases, the minority species of extracellular nucleic acid is sometimes about 1% to about 40% of the total nucleic acid (e.g., about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleic acid is minor species nucleic acid). In some embodiments, the minority nucleic acid is extracellular DNA.In some embodiments, the minority nucleic acid is extracellular DNA derived from apoptotic tissue. In some embodiments, the minority nucleic acid is extracellular DNA derived from tissue affected by a cell proliferative disorder. In some embodiments, the minority nucleic acid is extracellular DNA derived from tumor cells. In some embodiments, the minority nucleic acid is extracellular fetal DNA.

[0202] In another aspect, for example, a nucleic acid considered to be the "majority" may have an abundance of greater than 50% of the total nucleic acids in the sample to about 99.9% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acid may have an abundance of at least about 60% of the total nucleic acids in the sample to about 99% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acid may have an abundance of at least about 70% of the total nucleic acids in the sample to about 98% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acid may have an abundance of at least about 75% of the total nucleic acids in the sample to about 97% of the total nucleic acids in the sample. For example, the majority nucleic acid may have an abundance of at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acid is extracellular DNA. In some embodiments, the majority nucleic acid is extracellular maternal DNA. In some embodiments, the majority nucleic acid is DNA derived from healthy tissue. In some embodiments, the majority nucleic acid is DNA derived from non-tumor cells.

[0203] In some embodiments, the minority species of extracellular nucleic acids are about 500 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species nucleic acids are about 500 base pairs or less in length), In some embodiments, the minority species of extracellular nucleic acids are about 300 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species nucleic acids are about 300 base pairs or less in length). In some embodiments, the minority species of extracellular nucleic acids are about 250 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species nucleic acids are about 250 base pairs or less in length), In some embodiments, the minority species of extracellular nucleic acids are about 200 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species nucleic acids are about 200 base pairs or less in length). In some embodiments, the minority species of extracellular nucleic acids are about 150 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species nucleic acids are about 150 base pairs or less in length), In some embodiments, the minority species of extracellular nucleic acids are about 100 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species nucleic acids are about 100 base pairs or less in length). In some embodiments, the minority species of extracellular nucleic acids are about 50 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species nucleic acids are about 50 base pairs or less in length).

[0204] Nucleic acids may be provided for performing the methods described herein with or without processing a sample(s) containing the nucleic acids. In some embodiments, nucleic acids are provided for performing the methods described herein after processing a sample(s) containing the nucleic acids. For example, nucleic acids may be extracted, isolated, purified, partially purified, or amplified from a sample(s). The term "isolated," as used herein, refers to a nucleic acid that has been removed from its original environment (e.g., the natural environment if it is naturally occurring, or a host cell if it is exogenously expressed) and thus has been altered from its original environment by human intervention (e.g., "by the hand of man"). The term "isolated nucleic acid," as used herein, may refer to a nucleic acid that has been removed from a subject (e.g., a human subject). Isolated nucleic acids may be provided with fewer non-nucleic acid components (e.g., proteins, lipids) than the amount of those components present in the source sample. A composition comprising isolated nucleic acids may be about 50% to greater than 99% free of non-nucleic acid components. A composition comprising an isolated nucleic acid may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of non-nucleic acid components. The term "purified," as used herein, may refer to a nucleic acid, provided that it contains fewer non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than the amount of non-nucleic acid components present before the nucleic acid is subjected to a purification procedure. A composition comprising a purified nucleic acid may be about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other non-nucleic acid components. The term "purified," as used herein, may refer to a nucleic acid, provided that it contains fewer nucleic acid species than in the sample source from which the nucleic acid was derived. A composition containing purified nucleic acid can be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other nucleic acid species. For example, fetal nucleic acid can be purified from a mixture containing maternal and fetal nucleic acids.In certain examples, small fragments of fetal nucleic acid (e.g., 30-500 bp fragments) can be purified or partially purified from a mixture containing both fetal and maternal nucleic acid fragments. In certain examples, nucleosomes containing smaller fragments of fetal nucleic acid can be purified from a mixture of larger nucleosome complexes containing larger fragments of maternal nucleic acid. In certain examples, cancer cell nucleic acid can be purified from a mixture containing cancer cell and non-cancer cell nucleic acid. In certain examples, nucleosomes containing small fragments of cancer cell nucleic acid can be purified from a mixture of larger nucleosome complexes containing larger fragments of non-cancer nucleic acid. In some embodiments, nucleic acid is provided for performing the methods described herein without prior processing of the sample(s) containing the nucleic acid. For example, nucleic acid can be analyzed directly from the sample without prior extraction, purification, partial purification, and / or amplification.

[0205] In some embodiments, nucleic acids, such as cellular nucleic acids, are sheared or cleaved before, during, or after the methods described herein. The terms "shearing" or "cleavage" generally refer to procedures or conditions under which a nucleic acid molecule, such as a nucleic acid template gene molecule or its amplified product, can be separated into two (or more) smaller nucleic acid molecules. Such shearing or cleavage can be sequence-specific, base-specific, or non-specific, and can be achieved by any of a variety of methods, reagents, or conditions, including, for example, chemical, enzymatic, or physical shearing (e.g., physical fragmentation). The sheared or cleaved nucleic acids can have a nominal, average, or mean length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to about 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs.

[0206] Sheared or cleaved nucleic acids can be produced by suitable methods, non-limiting examples of which include physical methods (e.g., shearing, e.g., sonication, French press, heat, UV irradiation, etc.), enzymatic processes (e.g., enzymatic cleavage agents (e.g., suitable nucleases, suitable restriction enzymes, suitable methylation-sensitive restriction enzymes)), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, base hydrolysis, heat, etc., or combinations thereof), processes described in U.S. Patent Application Publication No. 2005 / 0112590, etc., or combinations thereof. The average, mean, or nominal length of the resulting nucleic acid fragments can be controlled by selecting an appropriate fragment production method.

[0207] The term "amplified," as used herein, refers to subjecting a target nucleic acid in a sample to a process that linearly or exponentially produces amplicon nucleic acids having the same or substantially the same nucleotide sequence as the target nucleic acid or a portion thereof. In certain embodiments, the term "amplified" refers to a method that includes polymerase chain reaction (PCR). In certain cases, the amplified product may contain one or more nucleotides more than the amplified nucleotide region of the nucleic acid template sequence (e.g., a primer may contain "extra" nucleotides, such as a transcription initiation sequence, in addition to nucleotides complementary to the nucleic acid template gene molecule, resulting in an amplified product containing "extra" nucleotides, or nucleotides that do not correspond to the amplified nucleotide region of the nucleic acid template gene molecule).

[0208] Nucleic acid can also be exposed to a process that modifies certain nucleotides in nucleic acid before it is provided to the method described herein.For example, a process that selectively modifies nucleic acid based on the methylation state of the nucleotides therein can be applied to nucleic acid.In addition, conditions such as high temperature, ultraviolet radiation, x-ray radiation, etc. can induce changes in the sequence of nucleic acid molecules.Nucleic acid can be provided in any suitable form that is useful for performing sequence analysis. Enriching nucleic acids

[0209] In some embodiments, nucleic acids (e.g., extracellular nucleic acids) are enriched or relatively enriched for a subpopulation or species of nucleic acid. Nucleic acid subpopulations can include, for example, fetal nucleic acids, maternal nucleic acids, cancer nucleic acids, patient nucleic acids, nucleic acids comprising fragments of a particular length or length range, or nucleic acids derived from a particular genomic region (e.g., a single chromosome, a set of chromosomes, and / or a particular chromosomal region). Such enriched samples can be used in conjunction with the methods provided herein. Thus, in certain embodiments, the methods of the technology include an additional step of enriching for a subpopulation of nucleic acids in the sample, such as cancer or fetal nucleic acids. In certain embodiments, methods for determining the fraction or fetal fraction of cancer cell nucleic acids can also be used to enrich for cancer or fetal nucleic acids. In certain embodiments, nucleic acids derived from normal tissue (e.g., non-cancer cells) are selectively removed (partially, substantially, almost completely, or completely) from the sample. In certain embodiments, maternal nucleic acids are selectively removed (partially, substantially, almost completely, or completely) from the sample. In certain embodiments, enriching for specific low copy number nucleic acid (for example, cancer or fetal nucleic acid) can improve quantitative sensitivity.The method for enriching sample for specific nucleic acid is described in, for example, United States Patent No. 6,927,028, International Patent Application Publication No. WO2007 / 140417, International Patent Application Publication No. WO2007 / 147063, International Patent Application Publication No. WO2009 / 032779, International Patent Application Publication No. WO2009 / 032781, International Patent Application Publication No. WO2010 / 033639, International Patent Application Publication No. WO2011 / 034631, International Patent Application Publication No. WO2006 / 056480 and International Patent Application Publication No. WO2011 / 143659, the entire contents of each of which, including all text, tables, equations and figures, are hereby incorporated by reference herein.

[0210] In some embodiments, nucleic acid is enriched for a specific target fragment type and / or reference fragment type.In certain embodiments, nucleic acid is enriched for a specific nucleic acid fragment length or fragment length range using one or more length-based separation methods described below.In certain embodiments, nucleic acid is enriched for fragments from selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein and / or known in the art.

[0211] Non-limiting examples of methods for enriching for nucleic acid subpopulations in a sample include methods that utilize epigenetic differences between nucleic acid species (e.g., the methylation-based fetal nucleic acid enrichment method described in U.S. Patent Application Publication No. 2010 / 0105049, hereby incorporated by reference); restriction endonuclease-enhanced polymorphic sequence approaches (e.g., such as those described in U.S. Patent Application Publication No. 2009 / 0317818, hereby incorporated by reference); selective enzymatic degradation approaches; massively parallel signature sequencing (MPSS) approaches; amplification (e.g., PCR)-based approaches (e.g., locus-specific amplification methods, multiplex SNP allele PCR approaches; universal amplification methods); pull-down approaches (e.g., biotinylated ultramer pull-down methods); extension and ligation-based methods (e.g., molecular inversion probe (MIP) extension and ligation); and combinations thereof.

[0212] In some embodiments, nucleic acids are enriched for fragments from selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein. Sequence-based separation is generally based on nucleotide sequences that are present in the fragments of interest (e.g., target and / or reference fragments) and are substantially absent or present in only trace amounts (e.g., 5% or less) in other fragments of the sample. In some embodiments, sequence-based separation can produce separated target fragments and / or separated reference fragments. The separated target fragments and / or separated reference fragments are often isolated from the remaining fragments in the nucleic acid sample. In certain embodiments, the separated target fragments and the separated reference fragments are also isolated from each other (e.g., isolated in separate assay compartments). In certain embodiments, the separated target fragments and the separated reference fragments are isolated together (e.g., isolated in the same assay compartment). In some embodiments, unbound fragments can be differentially removed, degraded, or digested.

[0213] In some embodiments, selective nucleic acid capture process is used to separate target and / or reference fragment from nucleic acid sample.Commercially available nucleic acid capture system includes, for example, Nimblegen sequence capture system (Roche NimbleGen, Madison, WI); Illumina BEADARRAY platform (Illumina, San Diego, CA); Affymetrix GENECHIP platform (Affymetrix, Santa Clara, CA); Agilent SureSelect Target Enrichment System (Agilent Technologies, Santa Clara, CA); and related platforms.This method typically involves the hybridization of capture oligonucleotide to part or all of the nucleotide sequence of target or reference fragment, which can include the use of solid phase (for example, solid phase array) and / or solution-based platform. Capture oligonucleotides (sometimes called "baits") can be selected or designed so that they preferentially hybridize to nucleic acid fragments from a selected genomic region or locus (e.g., one of chromosomes 21, 18, 13, X, or Y, or a reference chromosome). In certain embodiments, hybridization-based methods (e.g., using oligonucleotide arrays) can be used to enrich for nucleic acid sequences from a particular chromosome (e.g., a potentially aneuploid chromosome, a reference chromosome, or another chromosome of interest), gene, or region thereof of interest. Thus, in some embodiments, a nucleic acid sample is optionally enriched, for example, by capturing a subset of fragments using capture oligonucleotides complementary to selected genes in the sample nucleic acid. In certain cases, the captured fragments are amplified. For example, adapter-containing captured fragments can be amplified using primers complementary to the adapter oligonucleotides to form a collection of amplified fragments indexed according to the adapter sequence.In some embodiments, nucleic acids are enriched for fragments derived from selected genomic regions (e.g., chromosomes, genes) by amplification of one or more regions of interest using oligonucleotides (e.g., PCR primers) complementary to sequences in fragments containing the region(s) of interest or portion(s) thereof.

[0214] In some embodiments, nucleic acids are enriched for specific nucleic acid fragment lengths, length ranges, or lengths below or above a specific threshold or cutoff using one or more length-based separation methods. Nucleic acid fragment length typically refers to the number of nucleotides in a fragment. Nucleic acid fragment length is sometimes also referred to as nucleic acid fragment size. In some embodiments, length-based separation methods are performed without measuring the length of individual fragments. In some embodiments, length-based separation methods are performed in conjunction with a method for determining the length of individual fragments. In some embodiments, length-based separation refers to a size fractionation procedure in which all or a portion of the fractionated pool can be isolated (e.g., retained) and / or analyzed. Size fractionation procedures are known in the art (e.g., separation on an array, separation by molecular sieve, separation by gel electrophoresis, separation by column chromatography (e.g., size exclusion column), and microfluidic-based approaches). In certain cases, length-based separation approaches may include, for example, selective sequence tagging approaches, fragment circularization, chemical treatments (e.g., formaldehyde, polyethylene glycol (PEG) precipitation), mass spectrometry, and / or size-specific nucleic acid amplification. Nucleic acid quantification

[0215] The amount (e.g., concentration, relative amount, absolute amount, copy number, etc.) of nucleic acid in a sample can be determined. In some embodiments, the amount (e.g., concentration, relative amount, absolute amount, copy number, etc.) of minority nucleic acid in nucleic acid is determined. In certain embodiments, the amount of minority nucleic acid species in a sample is referred to as the "minority species fraction." In some embodiments, the "minority species fraction" refers to the fraction of minority nucleic acid species in circulating cell-free nucleic acid in a sample (e.g., blood sample, serum sample, plasma sample, urine sample) obtained from a subject.

[0216] The amount of minority nucleic acids in extracellular nucleic acids can be quantified and used in conjunction with the methods provided herein. Thus, in certain embodiments, the methods described herein include the additional step of determining the amount of minority nucleic acids. The amount of minority nucleic acids can be determined in a sample from a subject before or after processing to prepare sample nucleic acids. In certain embodiments, the amount of minority nucleic acids is determined in a sample after the sample nucleic acids are processed and prepared, and this amount is used for further evaluation. In some embodiments, the outcome includes factoring the minority species fraction in the sample nucleic acids (e.g., adjusting counts, removing samples, making calls, or not making calls).

[0217] The determination of the minority species fraction can be performed before, during, or at any point during the methods described herein, or after certain methods described herein (e.g., detecting genetic mutations or genetic alterations). For example, to perform a genetic mutation / genetic alteration determination method with a certain sensitivity or specificity, a minority nucleic acid quantification method can be performed before, during, or after the genetic mutation / genetic alteration determination to identify samples with more than about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25% or more minority nucleic acids. In some embodiments, samples determined to have a certain threshold amount of minority nucleic acids (e.g., about 15% or more minority nucleic acids; about 4% or more minority nucleic acids) are further analyzed, for example, for genetic mutations / alterations, or the presence or absence of genetic mutations / alterations. In certain embodiments, for example, determination of genetic mutations or genetic alterations is selected (e.g., selected and contacted with patients) only for samples having a certain threshold amount of minority nucleic acids (e.g., about 15% or more minority nucleic acids; about 4% or more minority nucleic acids).

[0218] In some embodiments, the amount of cancer cell nucleic acid in the nucleic acid (e.g., concentration, relative amount, absolute amount, copy number, etc.) is determined. In certain cases, the amount of cancer cell nucleic acid in a sample is referred to as the "fraction of cancer cell nucleic acid," sometimes referred to as the "cancer fraction" or "tumor fraction." In some embodiments, the "fraction of cancer cell nucleic acid" refers to the fraction of cancer cell nucleic acid in the circulating cell-free nucleic acid in a sample (e.g., a blood sample, a serum sample, a plasma sample, a urine sample) obtained from a subject.

[0219] In some embodiments, the amount (e.g., concentration, relative amount, absolute amount, copy number, etc.) of fetal nucleic acid in nucleic acid is determined. In certain embodiments, the amount of fetal nucleic acid in a sample is referred to as the "fetal fraction." In some embodiments, "fetal fraction" refers to the fraction of fetal nucleic acid in circulating cell-free nucleic acid in a sample (e.g., a blood sample, a serum sample, a plasma sample, a urine sample) obtained from a pregnant female. Certain methods for determining fetal fraction described herein or known in the art can be used to determine the fraction and / or minority species fraction of cancer cell nucleic acid.

[0220] In some embodiments, a fraction for the copy number variation region is determined. In some embodiments, a fetal fraction for the copy number variation region is determined. In some embodiments, a fraction of the minority nucleic acid is determined. In some embodiments, a fetal fraction for the sample nucleic acid is determined. The fraction may be determined according to the methods for fraction (e.g., fetal fraction) estimation or determination described below.

[0221] In certain cases, fetal fraction can be determined according to markers specific to male fetuses (e.g., Y chromosome STR markers (e.g., DYS19, DYS385, DYS392 markers); RhD markers in RhD-negative females), according to the allelic ratio of polymorphic sequences, or according to one or more markers specific to fetal nucleic acids but not maternal nucleic acids (e.g., differential epigenetic biomarkers (e.g., methylation) between mother and fetus), or fetal RNA markers in maternal plasma (e.g., Lo (2005) Journal of Histochemistry and Cytochemistry 53 (3): 293-296). In some embodiments, the fetal fraction can be determined according to an appropriate assay of the Y chromosome (e.g., by comparing the abundance of a fetal-specific locus (e.g., the SRY locus on the Y chromosome in male pregnancies) with the abundance of any autosomal locus common to both the mother and the fetus by using quantitative real-time PCR (see, e.g., Lo YM, et al. (1998) Am J Hum Genet 62:768-775)).

[0222] Determination of fetal fraction is sometimes performed using a fetal quantifier assay (FQA), for example, as described in U.S. Patent Application Publication No. 2010 / 0105049, which is hereby incorporated by reference. This type of assay allows for the detection and quantification of fetal nucleic acid in a maternal sample based on the methylation status of the nucleic acid in the sample. In certain embodiments, the amount of fetal nucleic acid from the maternal sample can be determined relative to the total amount of nucleic acid present, thereby providing the percentage of fetal nucleic acid in the sample. In certain embodiments, the copy number of fetal nucleic acid can be determined in the maternal sample. In certain embodiments, the amount of fetal nucleic acid can be determined in a sequence-specific (or segment-specific) manner, sometimes with sufficient sensitivity to enable accurate chromosomal dosage analysis (e.g., to detect the presence or absence of fetal aneuploidy).

[0223] A fetal quantity assay (FQA) can be performed in conjunction with any of the methods described herein. Such an assay can be performed by any method known in the art and / or described in U.S. Patent Application Publication No. 2010 / 0105049, such as a method that can distinguish maternal nucleic acids from fetal nucleic acids based on differential methylation status and quantify (i.e., determine the amount of) fetal nucleic acids. Methods for distinguishing nucleic acids based on methylation status include, but are not limited to, methylation-sensitive capture, e.g., using the MBD2-Fc fragment (MBD-FC), in which the methyl-binding domain of MBD2 is fused to the Fc fragment of an antibody (Gebhard et al. (2006) Cancer Res. 66(12):6118-28); methylation-specific antibodies; bisulfite conversion methods, e.g., MSP (methylation-sensitive PCR), COBRA, methylation-sensitive single-nucleotide primer extension (Ms-SNuPE), or Sequenom MassCLEAVE™ technology; and the use of methylation-sensitive restriction enzymes (e.g., enriching for fetal nucleic acids by digesting maternal nucleic acids in a maternal sample with one or more methylation-sensitive restriction enzymes). Methyl-sensitive enzymes can also be used to distinguish nucleic acids based on methylation status, e.g., enzymes that can preferentially or substantially cleave or digest at their DNA recognition sequences when the nucleic acid is unmethylated. Thus, unmethylated DNA samples are cleaved into smaller fragments than methylated DNA samples, and hypermethylated DNA samples are not cleaved. Except as expressly stated, any method for distinguishing nucleic acids based on methylation status can be used with the methods of the compositions and techniques herein. The amount of fetal nucleic acid can be determined, for example, by introducing one or more competitors at known concentrations during the amplification reaction. Determining the amount of fetal nucleic acid can also be performed, for example, by RT-PCR, primer extension, sequencing, and / or counting. In certain cases, the amount of nucleic acid can be determined using the BEAMing technology described in U.S. Patent Application Publication No. 2007 / 0065823.In certain embodiments, the restriction efficiency can be determined, and the efficiency ratio is used to further determine the amount of fetal nucleic acid.

[0224] In certain embodiments, the minority fraction can be determined based on the allele ratio of a polymorphic sequence (e.g., a single nucleotide polymorphism (SNP)), for example, using a method such as that described in U.S. Patent Application Publication No. 2011 / 0224087, which is hereby incorporated by reference. In such a method for determining fetal fraction, for example, nucleotide sequence reads are obtained for a maternal sample, and the fetal fraction is determined by comparing the total number of nucleotide sequence reads that map to a first allele and the total number of nucleotide sequence reads that map to a second allele at an informative polymorphic site (e.g., SNP) in a reference genome. In certain embodiments, fetal alleles are distinguished by their relatively minor contribution to the mixture of fetal and maternal nucleic acids in the sample, for example, compared to the major contribution of maternal nucleic acids to the mixture. Thus, the relative abundance of fetal nucleic acids in a maternal sample can be determined as a parameter of the total number of unique sequence reads that map to a target nucleic acid sequence on a reference genome for each of the two alleles at a polymorphic site.

[0225] Minority species fractions, in some embodiments, may be determined using methods incorporating information derived from chromosomal abnormalities, e.g., as described in International Patent Application Publication No. WO2014 / 055774, which is hereby incorporated by reference. Minority species fractions, in some embodiments, may be determined using methods incorporating information derived from sex chromosomes, e.g., as described in U.S. Patent Application Publication Nos. 2013 / 0288244 and 2013 / 0338933, each of which is hereby incorporated by reference.

[0226] The minority fraction can, in some embodiments, be determined using methods that incorporate fragment length information (e.g., fragment length ratio (FLR) analysis, fetal ratio statistic (FRS) analysis, as described in International Patent Application Publication No. WO2013 / 177086, which is hereby incorporated by reference). Cell-free fetal nucleic acid fragments are generally shorter than maternal nucleic acid fragments (see, e.g., Chan et al. (2004) Clin. Chem. 50:88-92; Lo et al. (2010) Sci. Transl. Med. 2:61ra91). Thus, the fetal ..., e.g., International Patent Application Publication No. WO2013 / 177086, which is hereby incorporated by reference). can be determined by counting fragments below a certain length threshold and comparing the counts to, for example, counts from fragments above a certain length threshold and / or the total amount of nucleic acid in the sample. Methods for counting nucleic acid fragments of a certain length are described in further detail in International Patent Application Publication No. WO2013 / 177086.

[0227] In certain embodiments, the FLR or FRS is determined in part according to the amount of reads mapped to portions derived from CCF fragments having lengths less than a selected fragment length. In some embodiments, the FLR or FRS value is often a ratio of X to Y, where X is the amount of reads derived from CCF fragments having lengths less than a first selected fragment length, and Y is the amount of reads derived from CCF fragments having lengths less than a second selected fragment length. The first selected fragment length is often selected independently of the second selected fragment length, and vice versa, the second selected fragment length is typically longer than the first selected fragment length. The first selected fragment length can be about 200 bases or less to about 30 bases or less. In some embodiments, the first selected fragment length is about 200, 190, 180, 170, 160, 155, 150, 145, 140, 135, 130, 125, 120, 115, 110, 105, 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, or 50 bases. In some embodiments, the first selected fragment length is about 170 to about 130 bases, and sometimes about 160 to about 140 bases. In some embodiments, the second selected fragment length is about 2000 bases to about 200 bases. In certain embodiments, the second selected fragment length is about 1000, 950, 800, 850, 800, 750, 700, 650, 600, 550, 500, 450, 400, 350, 300, or 250 bases. In some embodiments, the first selected fragment length is about 140 to about 160 bases (e.g., about 150 bases) and the second selected fragment length is about 500 to about 700 bases (e.g., about 600 bases). In some embodiments, the first selected fragment length is about 150 bases and the second selected fragment length is about 600 bases.

[0228] In some embodiments, the minority fraction can be determined according to the level. For example, the fetal fraction can be determined according to the level (e.g., the level for the affected region; the level for the copy number variation). Determining the fetal fraction according to the level can include determining the absolute value of the deviation of the level from the expected level and multiplying the absolute value of the deviation by 2. The expected level can be given a value of 1, and the deviation of the first or second level can be negative (e.g., for deletion or microdeletion; a level less than 1) or positive (e.g., for duplication or microduplication; a level greater than 1). The magnitude of the deviation can depend on the fetal fraction in certain cases.

[0229] In some embodiments, determining the minority fraction (e.g., cancer cell nucleic acid fraction; fetal fraction) is not required or necessary to identify the presence or absence of a genetic mutation or genetic alteration. In some embodiments, identifying the presence or absence of a genetic mutation or genetic alteration does not require sequence discrimination of minority versus majority nucleic acids. In certain embodiments, this is because the combined contributions of both minority and majority sequences in a particular chromosome, chromosomal segment, or portion thereof are analyzed. In some embodiments, identifying the presence or absence of a genetic mutation or genetic alteration does not rely on a priori sequence information to distinguish minority nucleic acids from majority nucleic acids. Partial specific fraction estimation

[0230] In some embodiments, the minority fraction may be determined according to a fraction-specific fraction estimate (e.g., as described in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is hereby incorporated by reference). For example, in some embodiments, the fetal fraction (e.g., for a sample) may be determined according to a fraction-specific fraction estimate (e.g., as described in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is hereby incorporated by reference). All of the fragments can be determined according to a fragment-specific fetal fraction estimation. Without being bound by theory, the amount of reads from fetal circulating cell-free (CCF) fragments (e.g., fragments of a particular length or length range) often maps to a fragment with a range of frequencies (e.g., within the same sample, e.g., within the same sequencing run). Also, without being bound by theory, a particular fragment tends to have a similar representation of reads from fetal CCF fragments (e.g., fragments of a particular length or length range) when compared across multiple samples, and this representation correlates with the fragment-specific fetal fraction (e.g., the relative amount, percentage, or ratio of CCF fragments originating from the fetus). The fetal fraction estimated according to fragment-specific fraction estimation may be referred to herein as a sequencing-based fetal fraction (e.g., SeqFF) and / or a bin-based fetal fraction (BFF).

[0231] Part-specific fetal fraction estimation is generally determined according to part-specific parameters and their relationship with fetal fraction.Part-specific parameters can be any suitable parameter that reflects (e.g., correlates with) the amount or proportion of CCF fragment lengths of a specific size (e.g., size range) in a part.Part-specific parameters can be the average, mean, or median of part-specific parameters determined for multiple samples.Any suitable part-specific parameter can be used. Non-limiting examples of portion-specific parameters include counts (e.g., counts of sequence reads mapped to portions; counts of sequence reads mapped to portions in a reference genome), normalized counts (e.g., normalized counts of sequence reads mapped to portions; normalized counts of sequence reads mapped to portions in a reference genome), fragment length ratios (FLRs), fetal ratio statistics (FRSs), the amount of reads with lengths less than a selected fragment length, genome coverage (i.e., coverage), mappability, DNase I sensitivity, methylation status, acetylation, histone distribution, guanine-cytosine (GC) content, chromatin structure, and the like, or combinations thereof. In some embodiments, the portion-specific parameters can be any suitable parameter that correlates with FLR and / or FRS in a portion-specific manner. In some embodiments, some or all portion-specific parameters are direct or indirect indications of FLR for the portion. In some embodiments, the portion-specific parameter is not guanine-cytosine (GC) content.

[0232] In some embodiments, the segment-specific parameter is any suitable value that represents, correlates with, or is proportional to the amount of reads from CCF fragments, and the reads mapped to the segment have a length less than the selected fragment length. In certain embodiments, the segment-specific parameter represents the amount of reads derived from relatively short CCF fragments (e.g., about 200 base pairs or less, about 150 base pairs or less) that map to the segment. CCF fragments having a length less than the selected fragment length are often relatively short CCF fragments, and sometimes the selected fragment length is about 200 base pairs or less (e.g., CCF fragments that are about 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, or 80 bases long). The length of the CCF fragment or the reads derived from the CCF fragment can be determined (e.g., inferred or inferred) by any suitable method (e.g., sequencing method, hybridization approach). In some embodiments, the length of the CCF fragment is determined (e.g., inferred or inferred) by the reads obtained from the paired-end sequencing method. In certain embodiments, the length of the CCF fragment template is determined directly from the length of the reads (e.g., single-end reads) derived from the CCF fragment.

[0233] The part-specific parameters may be weighted, adjusted, or transformed by one or more weighting factors. In some embodiments, the weighted, adjusted, or transformed part-specific parameters can provide a part-specific fetal fraction estimate for a sample (e.g., a test sample). In some embodiments, the weighting or adjustment generally converts part counts (e.g., reads mapped to parts) or another part-specific parameter into a part-specific fetal fraction estimate, and such a conversion is sometimes considered a transformation.

[0234] In some embodiments, the weighting factor is a coefficient or constant that partially describes and / or defines the relationship between the fetal fraction (e.g., fetal fraction determined from the plurality of samples) and the part-specific parameters for a plurality of samples (e.g., a training set). In some embodiments, the weighting factor is determined according to the relationship for the plurality of fetal fraction determinations and the plurality of part-specific parameters. The relationship may be defined by one or more weighting factors, and the one or more weighting factors may be determined from the relationship. In some embodiments, the weighting factor (e.g., one or more weighting factors) is determined from a fitted relationship for the parts according to (i) the fraction of fetal nucleic acid determined for each of the plurality of samples (e.g., the plurality of samples in the training set) and (ii) the part-specific parameters for the plurality of samples (e.g., the plurality of samples in the training set).

[0235] The weighting coefficients may be any suitable coefficients, estimated coefficients, or constants derived from a suitable relationship (e.g., a suitable mathematical relationship, algebraic relationship, fitted relationship, regression, regression analysis, regression model). The weighting coefficients may be determined according to, derived from, or estimated from a suitable relationship. In some embodiments, the weighting coefficients are estimated coefficients from a fitted relationship. Fitting a relationship for multiple samples is sometimes referred to herein as training a model. Any suitable model and / or method for fitting a relationship (e.g., training a model on a training set) may be used. Non-limiting examples of suitable models that may be used include regression models, linear regression models, simple regression models, ordinary least squares regression models, multiple regression models, general multiple regression models, polynomial regression models, general linear models, generalized linear models, discrete choice regression models, logistic regression models, multinomial logit models, mixed logit models, probit models, multinomial probit models, ordered logit models, ordered probit models, Poisson models, multivariate response regression models, multilevel models, fixed effects models, random effects models, mixed models, nonlinear regression models, nonparametric models, semiparametric models, robust models, quantile models, isotonic models, principal component models, least angle models, local models, segmented models, and errors in variables models. In some embodiments, the fitted relationship is not a regression model. In some embodiments, the fitted relationship is selected from a decision tree model, a support vector machine model, and a neural network model. The result of training a model (e.g., a regression model, a relationship) is often a relationship that can be described mathematically, which relationship includes one or more coefficients (e.g., weighting coefficients). For example, for a linear least squares model, a general multiple regression model can be trained using fetal fraction values ​​and part-specific parameters (e.g., coverage, see, e.g., Example 4), resulting in the relationship described by equation (1), where the weighting coefficient β is further defined in equations (2), (3), and (4).More complex multivariate models may determine one, two, three, or more weighting coefficients. In some embodiments, the model is trained according to fetal fractions obtained from multiple samples and two or more part-specific parameters (e.g., coefficients) (e.g., fitted relationships fitted to the multiple samples, e.g., by a matrix).

[0236] The weighting coefficients may be derived from any suitable relationship (e.g., any suitable mathematical relationship, algebraic relationship, fitted relationship, regression, regression analysis, regression model) by any suitable method. In some embodiments, the fitted relationship is fitted by estimation, non-limiting examples of which include least squares, ordinary least squares, linear, partial, total, generalized, weighted, nonlinear, iterative reweighting, ridge regression, least absolute deviation, Bayes, Bayesian multivariate, reduced rank, LASSO, Weighted Rank Selection Criteria (WRSC), Rank Selection Criteria (RSC), elastic net estimators (e.g., elastic net regression), and combinations thereof.

[0237] The weighting factor may have any suitable value. In some embodiments, the weighting factor is about -1×10 -2 and approximately 1 x 10 -2 Between about -1 x 10 -3 and approximately 1 x 10 -3 Between about -5 x 10 -4 and about 5 × 10 -4 Between or about -1 x 10 -4 and approximately 1 x 10 -4In some embodiments, the distribution of weighting factors for the plurality of samples is substantially symmetric. Sometimes, the distribution of weighting factors for the plurality of samples is a normal distribution. Sometimes, the distribution of weighting factors for the plurality of samples is not a normal distribution. In some embodiments, the width of the distribution of weighting factors depends on the amount of reads from CCF fetal nucleic acid fragments. In some embodiments, fractions with higher fetal nucleic acid content generate larger coefficients (e.g., positive or negative, see, e.g., FIG. 19 ). A weighting factor may be zero, or the weighting factor may be greater than zero. In some embodiments, about 70% or more, about 75% or more, about 80% or more, about 85% or more, about 90% or more, about 95% or more, or about 98% or more of the weighting factors for the fractions are greater than zero.

[0238] The weighting factors may be determined for or associated with any suitable portion of the genome. The weighting factors may be determined for or associated with any suitable portion of any suitable chromosome. In some embodiments, the weighting factors are determined for or associated with some or all portions in the genome. In some embodiments, the weighting factors are determined for or associated with some or all chromosome portions in the genome. The weighting factors are sometimes determined for or associated with selected chromosome portions. The weighting factors may be determined for or associated with one or more autosome portions. The weighting factors may be determined for or associated with portions in a plurality of portions, including portions in autosomes or a subset thereof. In some embodiments, the weighting factors are determined for or associated with portions of sex chromosomes (e.g., the X and / or Y chromosomes). The weighting factors may be determined for or associated with portions of one or more autosomes and one or more sex chromosomes. In certain embodiments, the weighting factors are determined for or associated with portions in all autosomes and the X and Y chromosomes. The weighting factor may be determined for or related to a portion of the plurality of portions that does not include a portion in chromosome X and / or Y. In certain embodiments, the weighting factor is determined for or related to a portion of a chromosome where the chromosome includes aneuploidy (e.g., whole chromosome aneuploidy). In certain embodiments, the weighting factor is determined for or related to only a portion of a chromosome where the chromosome is not aneuploid (e.g., euploid chromosome). The weighting factor may be determined for or related to a portion of the plurality of portions that does not include a portion in chromosome 13, chromosome 18, and / or chromosome 21.

[0239] In some embodiments, weighting factors are determined for portions according to one or more samples (e.g., a training set of samples). Weighting factors are often portion-specific. In some embodiments, one or more weighting factors are independently assigned to portions. In some embodiments, weighting factors are determined according to fetal fraction determinations for multiple samples (e.g., sample-specific fetal fraction determinations) and relationships for portion-specific parameters determined according to the multiple samples. Weighting factors are often determined from multiple samples, e.g., from about 20 to about 100,000 or more, about 100 to about 100,000 or more, about 500 to about 100,000 or more, about 1,000 to about 100,000 or more, or about 10,000 to about 100,000 or more. Weighting factors can be determined from euploid samples (e.g., samples from subjects containing euploid fetuses, e.g., samples in which no aneuploid chromosomes are present). In some embodiments, the weighting factor is obtained from a sample containing aneuploid chromosomes (e.g., a sample from a subject with a euploid fetus). In some embodiments, the weighting factor is determined from a plurality of samples from a subject with a euploid fetus and a subject with a trisomic fetus. The weighting factor can be derived from a plurality of samples, and these samples are from male fetuses and / or female fetuses.

[0240] The fetal fraction is often determined for one or more samples of the training set from which the weighting factors are derived. The fetal fraction from which the weighting factors are determined is sometimes a sample-specific fetal fraction determination. The fetal fraction from which the weighting factors are determined can be determined by any suitable method described herein or known in the art. In some embodiments, determination of fetal nucleic acid content (e.g., fetal fraction) is performed using a suitable fetal quantity assay (FQA) described herein or known in the art, non-limiting examples of which include fetal fraction determination according to markers specific for male fetuses, based on allelic ratios of polymorphic sequences, according to one or more markers specific for fetal nucleic acids but not maternal nucleic acids, by using methylation-based DNA discrimination (e.g., A. Nygren, et al., (2010) Clinical Chemistry 56(10):1627-1635), by mass spectrometry methods and / or systems using competitive PCR approaches, by methods described in U.S. Patent Application Publication No. 2010 / 0105049, hereby incorporated by reference, or the like, or by combinations thereof. In certain instances, fetal fraction is determined in part according to the Y chromosome level (e.g., at one or more genomic segment levels, profile levels). In some embodiments, fetal fraction is determined by a suitable assay of the Y chromosome (e.g., comparing the abundance of a fetal-specific locus (e.g., the SRY locus on the Y chromosome in male pregnancies) to the abundance of any autosomal locus common to both the mother and fetus by using quantitative real-time PCR (e.g., Lo YM, et al. (1998) Am J Hum Genet 62:768-775).

[0241] The part-specific parameters (e.g., for a test sample) can be weighted, adjusted, or transformed by one or more weighting factors (e.g., weighting factors derived from a training set). For example, weighting factors can be derived for parts according to the relationship between the part-specific parameters and fetal fraction determination for a training set of multiple samples. The part-specific parameters of the test sample can then be adjusted and / or weighted according to the weighting factors derived from the training set. In some embodiments, the part-specific parameters from which the weighting factors are derived are the same as the part-specific parameters (e.g., of the test sample) being adjusted or weighted (e.g., both parameters are FLR). In certain embodiments, the part-specific parameters from which the weighting factors are derived are different from the part-specific parameters (e.g., of the test sample) being adjusted or weighted. For example, a weighting factor can be determined from the relationship between coverage (i.e., part-specific parameters) and fetal fraction for a training set of samples, and the FLR (i.e., another part-specific parameter) for the part of the test sample can be adjusted according to the weighting factor derived from the coverage. Without being bound by theory, the part-specific parameters (e.g., for a test sample) may sometimes be adjusted and / or weighted and / or transformed by weighting factors derived from different part-specific parameters (e.g., of a training set) due to the relationship and / or correlation between each part-specific parameter and a common part-specific FLR.

[0242] A portion-specific fetal fraction estimate can be determined for a sample (e.g., a test sample) by weighting, adjusting, or transforming a portion-specific parameter (e.g., the count of sequence reads mapped to a portion of a reference genome) by a weighting factor determined for that portion. Weighting can include adjusting, transforming, and / or transforming the portion-specific parameter (e.g., the count of sequence reads mapped to a portion of a reference genome) according to the weighting factor by applying any suitable mathematical operation, non-limiting examples of which include multiplication, division, addition, subtraction, integration, symbolic calculation, algebraic calculation, algorithm, trigonometric or geometric function, transformation (e.g., Fourier transform), etc., or a combination thereof. Weighting can include adjusting, transforming, and / or transforming the portion-specific parameter (e.g., the count of sequence reads mapped to a portion of a reference genome) according to the weighting factor, an appropriate mathematical model (e.g., the model shown in Example 4).

[0243] In some embodiments, the fetal fraction is determined for a sample according to one or more part-specific fetal fraction estimates. In some embodiments, the fetal fraction is determined (e.g., estimated) for a sample (e.g., a test sample) according to weighting, adjustment, or transformation of part-specific parameters for one or more parts (e.g., the count of sequence reads mapped to a part of a reference genome). In certain embodiments, the fetal nucleic acid fraction for a test sample is estimated based on the adjusted count or an adjusted subset of counts. In certain embodiments, the fetal nucleic acid fraction for a test sample is estimated based on the adjusted FLR, adjusted FRS, adjusted coverage, and / or adjusted mappability for the part. In some embodiments, about 1 to about 500,000, about 100 to about 300,000, about 500 to about 200,000, about 1000 to about 200,000, about 1500 to about 200,000, or about 1500 to about 50,000 part-specific parameters are weighted or adjusted.

[0244] The fetal fraction (e.g., for a test sample) may be determined according to multiple part-specific fetal fraction estimates (e.g., for the same test sample) by any suitable method. In some embodiments, a method for increasing the accuracy of an estimate of the fraction of fetal nucleic acid in a test sample from a pregnant female includes determining one or more part-specific fetal fraction estimates, wherein the estimate of the fetal fraction for the sample is determined according to the one or more part-specific fetal fraction estimates. In some embodiments, estimating or determining the fraction of fetal nucleic acid for a sample (e.g., a test sample) includes summing the one or more part-specific fetal fraction estimates. The summing may include determining an average, mean, median, AUC, or integral according to the multiple part-specific fetal fraction estimates.

[0245] In some embodiments, a method for increasing the accuracy of an estimate of the fraction of fetal nucleic acid in a test sample from a pregnant female includes obtaining counts of sequence reads mapped to portions of a reference genome, the sequence reads being reads of circulating cell-free nucleic acid from the test sample from the pregnant female, and at least a subset of the obtained counts are derived from regions of the genome that contribute a greater number of counts derived from fetal nucleic acid compared to the total counts from the region than the counts of fetal nucleic acid compared to the total counts of another region of the genome. In some embodiments, the estimate of the fraction of fetal nucleic acid is determined according to a subset of portions, the subset of portions being selected according to portions to which the counts derived from fetal nucleic acid compared to non-fetal nucleic acid are mapped that are greater than the counts of fetal nucleic acid compared to non-fetal nucleic acid in another portion. The counts mapped to all or a subset of portions can be weighted, adjusted, or transformed to provide weighted, adjusted, or transformed counts. The weighted, adjusted, or transformed counts can be utilized to estimate the fraction of fetal nucleic acids, and the counts can be weighted, adjusted, or transformed according to portions where the counts derived from fetal nucleic acids map to a greater number than the counts of fetal nucleic acids in another portion. In some embodiments, the counts are weighted according to portions where the counts derived from fetal nucleic acids compared to non-fetal nucleic acids map to a greater number than the counts of fetal nucleic acids compared to non-fetal nucleic acids in another portion.

[0246] The fetal fraction can be determined for a sample (e.g., a test sample) according to a plurality of part-specific fetal fraction estimates for the sample, where the part-specific estimates are from portions of any suitable region or segment of the genome. The part-specific fetal fraction estimates can be determined for one or more portions of suitable chromosomes (e.g., one or more selected chromosomes, one or more autosomes, sex chromosomes (e.g., X and / or Y chromosomes), aneuploid chromosomes, euploid chromosomes, etc., or combinations thereof). In some embodiments, the fetal fraction can be determined for a sample (e.g., a test sample) according to a plurality of part-specific fetal fraction estimates for the sample, where the part-specific estimates are from chromosomes or portions thereof classified as having copy number variations (e.g., aneuploidies, microduplications, microdeletions). The fetal fraction determined according to a plurality of part-specific fetal fraction estimates for a sample, where the part-specific estimates are from chromosomes or portions thereof classified as having copy number variations, can be referred to herein as the affected fraction (AF).

[0247] The portion-specific parameters (e.g., counts of sequence reads mapped to portions of a reference genome), weighting factors, portion-specific fetal fraction estimations, and / or fetal fraction determinations may be determined by a suitable system, machine, apparatus, non-transitory computer-readable storage medium (e.g., having an executable program stored thereon), etc., or combinations thereof. In certain embodiments, the portion-specific parameters (e.g., counts of sequence reads mapped to portions of a reference genome), weighting factors, portion-specific fetal fraction estimations, and / or fetal fraction determinations are determined (e.g., in part) by a system or machine including one or more microprocessors and memory. In some embodiments, the portion-specific parameters (e.g., counts of sequence reads mapped to portions of a reference genome), weighting factors, portion-specific fetal fraction estimations, and / or fetal fraction determinations are determined (e.g., in part) by a non-transitory computer-readable storage medium having an executable program stored thereon, which program instructs the microprocessor to perform the determinations.

[0248] In some embodiments, a fraction for a region of copy number variation is determined. In some embodiments, a fetal fraction for a region of copy number variation is determined. In some embodiments, a fraction of a minority nucleic acid is determined. In some embodiments, a fetal fraction for a sample nucleic acid is determined. The fraction may be determined according to sequencing-based fetal fraction estimation described herein. In some embodiments, a sequencing-based fraction (e.g., fetal fraction) estimate is generated according to a method including: (i) obtaining counts of sequence reads mapped to portions of a reference genome, wherein the sequence reads are obtained from sample nucleic acid derived from a subject; (ii) converting the counts of sequence reads mapped to each portion into a portion-specific fraction of nucleic acid (e.g., fetal nucleic acid) according to a weighting factor independently associated with each portion, thereby providing a portion-specific fraction estimate (e.g., fetal fraction estimate) for the sample nucleic acid derived from the subject according to the weighting factor, wherein each of the weighting factors is determined from a fitted relationship, for each portion, between (1) the fraction of nucleic acid (e.g., fetal nucleic acid) for each of a plurality of samples in a training set and (2) the counts of sequence reads mapped to each portion for the plurality of samples; and (iii) estimating the fraction of nucleic acid (e.g., fetal nucleic acid) for the sample nucleic acid derived from the subject based on the portion-specific fraction estimates (e.g., fetal fraction estimate).

[0249] For the step of determining a fraction for a copy number variation region, a part-specific fraction estimate is provided by converting the count of sequence reads mapped to each part in the copy number variation region into a part-specific fraction of nucleic acid according to a weighting factor independently associated with each part in the copy number variation region. For the step of determining a fetal fraction for a copy number variation region, a part-specific fetal fraction estimate is provided by converting the count of sequence reads mapped to each part in the copy number variation region into a part-specific fetal fraction of nucleic acid according to a weighting factor independently associated with each part in the copy number variation region.

[0250] For the step of determining the fraction of minority nucleic acids, a part-specific fraction estimate is provided by converting the count of sequence reads mapped to each part in a plurality of regions (e.g., regions not limited to the copy number variation regions; regions across the genome) into a part-specific fraction of nucleic acid according to a weighting factor independently associated with each part. For the step of determining the fetal fraction for the sample nucleic acid, a part-specific fetal fraction estimate is provided by converting the count of sequence reads mapped to each part in a plurality of regions (e.g., regions not limited to the copy number variation regions; regions across the genome) into a part-specific fraction of fetal nucleic acid according to a weighting factor independently associated with each part. Nucleic Acid Library

[0251] In some embodiments, a nucleic acid library is a plurality of polynucleotide molecules (e.g., a sample of nucleic acids) that have been prepared, assembled, and / or modified for a particular process, non-limiting examples of which include immobilization on a solid phase (e.g., a solid support, a flow cell, a bead), enrichment, amplification, cloning, detection, and / or nucleic acid sequencing. In certain embodiments, the nucleic acid library is prepared before or during the sequencing process. Nucleic acid libraries (e.g., sequencing libraries) can be prepared by any suitable method known in the art. Nucleic acid libraries can be prepared by targeted or non-targeted preparation processes.

[0252] In some embodiments, a library of nucleic acids is modified to include chemical moieties (e.g., functional groups) configured for immobilization of the nucleic acids to a solid support. In some embodiments, a library of nucleic acids is modified to include biomolecules (e.g., functional groups) and / or members of binding pairs configured for immobilization of the library to a solid support, non-limiting examples of which include thyroxine-binding globulin, steroid-binding proteins, antibodies, antigens, haptens, enzymes, lectins, nucleic acids, repressors, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding proteins, receptors, carbohydrates, oligonucleotides, polynucleotides, complementary nucleic acid sequences, and the like, and combinations thereof. Examples of portions of specific binding pairs include, but are not limited to, the following: an avidin moiety and a biotin moiety; an antigenic epitope and an antibody or immunologically reactive fragment thereof; an antibody and a hapten; a digoxigen moiety and an anti-digoxigen antibody; a fluorescein moiety and an anti-fluorescein antibody; an operator and a repressor; a nuclease and a nucleotide; a lectin and a polysaccharide; a steroid and a steroid-binding protein; an active compound and an active compound receptor; a hormone and a hormone receptor; an enzyme and a substrate; an immunoglobulin and protein A; an oligonucleotide or polynucleotide and its corresponding complement; and the like, or combinations thereof.

[0253] In some embodiments, a library of nucleic acids is modified to include one or more polynucleotides of known composition, non-limiting examples of which include identifiers (e.g., tags, indexing tags), capture sequences, labels, adapters, restriction enzyme sites, promoters, enhancers, origins of replication, stem loops, complimentary sequences (e.g., primers, These may include a binding site, an annealing site, an appropriate integration site (e.g., a transposon, a viral integration site), modified nucleotides, or a combination thereof. The polynucleotide of known sequence may be added at any suitable position, for example, on the 5' end, the 3' end, or within the nucleic acid sequence. The polynucleotides of known sequence may be the same or different sequences. In some embodiments, the polynucleotide of known sequence is configured to hybridize to one or more oligonucleotides immobilized on a surface (e.g., a surface in a flow cell). For example, a nucleic acid molecule containing a known sequence on the 5' side may hybridize to a first plurality of oligonucleotides, while a known sequence on the 3' side may hybridize to a second plurality of oligonucleotides. In some embodiments, the library of nucleic acids may include chromosome-specific tags, capture sequences, labels, and / or adapters. In some embodiments, the library of nucleic acids includes one or more detectable labels. In some embodiments, one or more detectable labels may be incorporated into the nucleic acid library at the 5' end, the 3' end, and / or at any nucleotide position within the nucleic acids in the library. In some embodiments, the library of nucleic acids comprises hybridized oligonucleotides. In certain embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, the library of nucleic acids comprises hybridized oligonucleotide probes prior to immobilization on the solid phase.

[0254] In some embodiments, the polynucleotide of known sequence comprises a universal sequence. A universal sequence is a specific nucleotide sequence incorporated into two or more nucleic acid molecules or two or more subsets of nucleic acid molecules, and the universal sequence is the same for all molecules or subsets of molecules into which it is incorporated. Universal sequences are often designed to hybridize to and / or amplify multiple different sequences using a single universal primer complementary to the universal sequence. In some embodiments, two (e.g., a pair) or more universal sequences and / or universal primers are used. The universal primer often comprises a universal sequence. In some embodiments, an adapter (e.g., a universal adapter) comprises a universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple species or subsets of nucleic acids.

[0255] In certain embodiments where nucleic acid libraries are prepared (e.g., by synthetic procedures, for certain sequencing), the nucleic acids are size-selected and / or fragmented to lengths of a few hundred base pairs or less (e.g., in preparation for library generation). In some embodiments, library preparation is performed without fragmentation (e.g., when using cell-free DNA).

[0256] In certain embodiments, ligation-based library preparation methods are used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego, CA). Ligation-based library preparation methods often use adapter (e.g., methylated adapter) designs that can incorporate index sequences (e.g., sample index sequences to identify the origin of the sample for the nucleic acid sequence) in the initial ligation step, and can often be used to prepare samples for single-read sequencing, paired-end sequencing, and multiplexed sequencing. For example, nucleic acids (e.g., fragmented nucleic acids or cell-free DNA) can be end-repaired by a fill-in reaction, an exonuclease reaction, or a combination thereof. Then, in some embodiments, the resulting blunt-end-repaired nucleic acid can be extended by a single nucleotide complementary to the single-nucleotide overhang on the 3' end of the adapter / primer. Any nucleotide can be used for the extension / overhang nucleotide.

[0257] In some embodiments, nucleic acid library preparation involves ligating adapter oligonucleotides (e.g., to sample nucleic acids, to sample nucleic acid fragments, to template nucleic acids). Adapter oligonucleotides are often complementary to flow cell anchors and are sometimes used to immobilize nucleic acid libraries to a solid support, such as the inner surface of a flow cell. In some embodiments, the adapter oligonucleotides comprise an identifier, one or more sequencing primer hybridization sites (e.g., sequences complementary to universal sequencing primers, single-end sequencing primers, paired-end sequencing primers, multiplexed sequencing primers, etc.), or combinations thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing). In some embodiments, the adapter oligonucleotides comprise one or more of a primer annealing polynucleotide (e.g., for annealing to flow cell-bound oligonucleotides and / or free amplification primers), an index polynucleotide (e.g., a sample index sequence for tracking nucleic acids from different samples; also referred to as a sample ID), and a barcode polynucleotide (e.g., a single molecule barcode (SMB) for tracking individual molecules of amplified sample nucleic acid prior to sequencing; also referred to as a molecular barcode). In some embodiments, the primer annealing component of the adapter oligonucleotide comprises one or more universal sequences (e.g., sequences complementary to one or more universal amplification primers). In some embodiments, an index polynucleotide (e.g., sample index; sample ID) is a component of an adapter oligonucleotide. In some embodiments, an index polynucleotide (e.g., sample index; sample ID) is a component of a universal amplification primer sequence.

[0258] In some embodiments, when adapter oligonucleotides are used in combination with designed amplification primers (e.g., universal amplification primers), they generate library constructs that include one or more of a universal sequence, a molecular barcode, a sample ID sequence, a spacer sequence, and a sample nucleic acid sequence. In some embodiments, when adapter oligonucleotides are used in combination with designed universal amplification primers, they generate library constructs that include an ordered combination of one or more of a universal sequence, a molecular barcode, a sample ID sequence, a spacer sequence, and a sample nucleic acid sequence. For example, a library construct may include a first universal sequence, followed by a second universal sequence, followed by a first molecular barcode, followed by a spacer sequence, followed by a template sequence (e.g., a sample nucleic acid sequence), followed by a spacer sequence, followed by a second molecular barcode, followed by a third universal sequence, followed by a sample ID, followed by a fourth universal sequence. In some embodiments, when adapter oligonucleotides are used in combination with designed amplification primers (e.g., universal amplification primers), they generate library constructs for each strand of a template molecule (e.g., a sample nucleic acid molecule). In some embodiments, the adapter oligonucleotide is a double-stranded adapter oligonucleotide.

[0259] The identifier may be a suitable detectable label incorporated into or attached to a nucleic acid (e.g., a polynucleotide) that allows for detection and / or identification of the nucleic acid containing the identifier. In some embodiments, the identifier is incorporated into or attached to a nucleic acid during a sequencing method (e.g., by a polymerase). Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indexes, or barcodes, radiolabels (e.g., isotopes), metal labels, fluorescent labels, chemiluminescent labels, phosphorescent labels, fluorophore quenchers, dyes, proteins (e.g., enzymes, antibodies or portions thereof, linkers, members of binding pairs), etc., or combinations thereof. In some embodiments, the identifier (e.g., nucleic acid index or barcode) is a unique, known, and / or identifiable sequence of nucleotides or nucleotide analogs. In some embodiments, the identifier is six or more consecutive nucleotides. Numerous fluorophores with a variety of different excitation and emission spectra are available. Any suitable type and / or number of fluorophores can be used as identifiers. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, or fifty or more different identifiers are utilized in the methods described herein (e.g., nucleic acid detection and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are linked to each nucleic acid in the library.Detection and / or quantification of the identifiers may be performed by any suitable method, device or machine, non-limiting examples of which include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, a luminometer, a fluorometer, a spectrophotometer, suitable gene chip or microarray analysis, Western blot, mass spectrometry, chromatography, cytofluorometry, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electric field suspension, suitable nucleic acid sequencing methods and / or devices, etc., and combinations thereof.

[0260] In some embodiments, transposon-based library preparation methods are used (e.g., EPICENTRE NEXTERA, Epicentre, Madison, WI). Transposon-based methods typically use in vitro transposition to simultaneously fragment and tag DNA in a single-tube reaction (often allowing for the incorporation of platform-specific tags and optional barcodes) and prepare sequencer-ready libraries.

[0261] In some embodiments, the nucleic acid library, or a portion thereof, is amplified (e.g., amplified by a PCR-based method). In some embodiments, the sequencing method includes amplification of the nucleic acid library. The nucleic acid library can be amplified before or after immobilization on a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification includes a process of amplifying or increasing the number of nucleic acid templates and / or their complements present (e.g., in a nucleic acid library) by producing one or more copies of the templates and / or their complements. Amplification can be performed by any suitable method. The nucleic acid library can be amplified by thermocycling or isothermal amplification methods. In some embodiments, a rolling circle amplification method is used. In some embodiments, amplification is performed on a solid support (e.g., in a flow cell) to which the nucleic acid library, or a portion thereof, is immobilized. In certain sequencing methods, the nucleic acid library is added to a flow cell and immobilized by hybridization to anchors under appropriate conditions. This type of nucleic acid amplification is often referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or a portion of the amplified product is synthesized by extension initiated from an immobilized primer. Solid-phase amplification reactions are similar to standard liquid-phase amplification, except that at least one of the amplification oligonucleotides (e.g., primers) is immobilized on a solid support. In some embodiments, modified nucleic acids (e.g., nucleic acids modified by the addition of adapters) are amplified.

[0262] In some embodiments, solid-phase amplification includes nucleic acid amplification reactions that include only one type of oligonucleotide primer immobilized on a surface. In certain embodiments, solid-phase amplification includes multiple different immobilized oligonucleotide primer species. In some embodiments, solid-phase amplification can include nucleic acid amplification reactions that include one type of oligonucleotide primer immobilized on a solid surface and a second, different oligonucleotide primer species in solution. Multiple different types of immobilized or solution-based primers can be used. Non-limiting examples of solid-phase nucleic acid amplification reactions include interface amplification, bridge amplification, emulsion PCR, WildFire amplification (e.g., U.S. Patent Application Publication No. 2013 / 0012399), etc., or combinations thereof. nucleic acid capture

[0263] In some embodiments, the sample nucleic acid (or sample nucleic acid library) is subjected to a target capture process. Generally, the target capture process is carried out by contacting the sample nucleic acid (or sample nucleic acid library) with a set of probe oligonucleotides under hybridization conditions. The set of probe oligonucleotides (e.g., capture oligonucleotides) generally comprises a plurality of probe oligonucleotides having sequences complementary or substantially complementary to sequences in the sample nucleic acid. The plurality of probe oligonucleotides may comprise about 10 probe oligonucleotide species, about 50 probe oligonucleotide species, about 100 probe oligonucleotide species, about 500 probe oligonucleotide species, about 1,000 probe oligonucleotide species, 2,000 probe oligonucleotide species, 3,000 probe oligonucleotide species, 4,000 probe oligonucleotide species, 5,000 probe oligonucleotide species, 10,000 probe oligonucleotide species, or more. Generally, a first probe oligonucleotide species has a different nucleotide sequence from a second probe oligonucleotide species, and each of the different species of probe oligonucleotides in the set has a different nucleotide sequence.

[0264] A probe oligonucleotide typically comprises a nucleotide sequence capable of hybridizing or annealing to a nucleic acid fragment (e.g., a target fragment) or portion thereof of interest. Probe oligonucleotides can be naturally occurring or synthetic and can be DNA or RNA-based. Probe oligonucleotides can, for example, enable specific separation of a target fragment from other fragments in a nucleic acid sample. The terms "specific" or "specificity," as used herein, refer to the binding or hybridization of one molecule to another molecule, e.g., an oligonucleotide to a target polynucleotide. "Specific" or "specificity" refers to the recognition, contact, and formation of a stable complex between two molecules, compared to significantly less recognition, contact, or complex formation between either of the two molecules and another molecule. As used herein, the terms "annealing" and "hybridizing" refer to the formation of a stable complex between two molecules. The terms "probe," "probe oligonucleotide," "capture probe," "capture oligonucleotide," "capture oligo," "oligo," or "oligonucleotide" can be used interchangeably throughout the document when referring to a probe oligonucleotide.

[0265] Probe oligonucleotides can be designed and synthesized using any suitable process and can be of any length suitable for hybridizing to a nucleotide sequence of interest and performing the separation and / or analysis processes described herein. Oligonucleotides can be designed based on the nucleotide sequence of interest (e.g., a target fragment sequence, a genomic sequence, or a gene sequence). Oligonucleotides (e.g., probe oligonucleotides) can, in some embodiments, be about 10 to about 300 nucleotides, about 50 to about 200 nucleotides, about 75 to about 150 nucleotides, about 110 to about 130 nucleotides, or about 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, or 129 nucleotides in length. Oligonucleotides can be composed of naturally occurring and / or non-naturally occurring nucleotides (e.g., labeled nucleotides), or mixtures thereof. Oligonucleotides suitable for use with the embodiments described herein can be synthesized and labeled using known techniques. Oligonucleotides can be chemically synthesized using an automated synthesizer according to the solid-phase phosphoramidite triester method first described by Beaucage and Caruthers (1981) Tetrahedron Letts. 22:1859-1862 and / or described by Needham-VanDevanter et al. (1984) Nucleic Acids Res. 12:6159-6168. Purification of oligonucleotides can be achieved by native acrylamide gel electrophoresis or anion-exchange high-performance liquid chromatography (HPLC), for example, as described in Pearson and Regnier (1983) J. Chrom. 255:137-149.

[0266] In some embodiments, all or part of the probe oligonucleotide sequence (naturally occurring or synthetic) can be substantially complementary to the target sequence or a portion thereof. As referred to herein, "substantially complementary" with respect to a sequence refers to nucleotide sequences that hybridize with each other. The stringency of hybridization conditions can be varied to allow for varying amounts of sequence mismatch. 55% or higher, 56% or higher, 57% or higher, 58% or higher, 59% or higher, 60% or higher, 61% or higher, 62% or higher, 63% or higher, 64% or higher, 65% or higher, 66% or higher, 67% or higher, 68% or higher, 69% or higher, 70% or higher, 71% or higher, 72% or higher, 73% or higher, 74% or higher, 75% or higher, 76% or higher, 77% or higher, 78% or higher of each other or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more complementary target and oligonucleotide sequences.

[0267] A probe oligonucleotide that is substantially complementary to a nucleotide sequence of interest (e.g., a target sequence) or a portion thereof will also be substantially similar to the complement of the target sequence or a relevant portion thereof (e.g., substantially similar to the antisense strand of a nucleic acid). One test for determining whether two nucleotide sequences are substantially similar is to determine the percent of identical nucleotide sequences shared. As referred to herein, "substantially similar" with respect to sequences means 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more of each other. , 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more identical nucleotide sequence.

[0268] Hybridization conditions (e.g., annealing conditions) can be determined and / or adjusted depending on the characteristics of the oligonucleotides used in the assay. The sequence and / or length of the oligonucleotide can sometimes affect hybridization to the target nucleic acid sequence. Depending on the degree of mismatch between the oligonucleotide and the target nucleic acid, low, medium, or high stringency conditions can be used to achieve annealing. As used herein, the term "stringent conditions" refers to the conditions for hybridization and washing. Methods for optimizing the temperature conditions of hybridization reactions are known in the art and can be found in Current Protocols in Molecular Biology, John Wiley & Sons, NY, 6.3.1-6.3.6 (1989). Aqueous and non-aqueous methods are described in this reference, and either can be used. A non-limiting example of stringent hybridization conditions is hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% SDS at 50° C. Another example of stringent hybridization conditions is hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% SDS at 55° C. A further example of stringent hybridization conditions is hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% SDS at 60° C. Often, stringent hybridization conditions are hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% SDS at 65° C. More often, stringent conditions are 0.5 M sodium phosphate, 7% SDS at 65° C., followed by one or more washes in 0.2×SSC, 1% SDS at 65° C.Stringent hybridization temperatures can also be altered (i.e., lowered) by the addition of certain organic solvents, such as formamide, which reduce the thermal stability of double-stranded polynucleotides, such that hybridization can be performed at lower temperatures while still maintaining stringent conditions and extending the useful life of nucleic acids that may be thermolabile.

[0269] In some embodiments, one or more probe oligonucleotides are associated with an affinity ligand, such as a member of a binding pair (e.g., biotin) or an antigen, that can bind to a capture agent, such as avidin, streptavidin, an antibody, or a receptor. For example, the probe oligonucleotide can be biotinylated so that it can be captured on streptavidin-coated beads.

[0270] In some embodiments, one or more probe oligonucleotides and / or capture agents are operatively linked to a solid support or substrate. The solid support or substrate can be any physically separable solid to which the probe oligonucleotides can be directly or indirectly bound, including, but not limited to, microarrays and wells, and surfaces provided by particles, such as beads (e.g., paramagnetic beads, magnetic beads, microbeads, nanobeads), microparticles, and nanoparticles. Solid supports can include, for example, chips, columns, optical fibers, wipes, filters (e.g., flat surface filters), one or more capillaries, glass, and modified or functionalized glass (e.g., controlled-pore glass). glass (CPG), quartz, mica, diazotized membranes (paper or nylon), polyformaldehyde, cellulose, cellulose acetate, paper, ceramic, metals, semi-metals, semiconductor materials, quantum dots, coated beads or particles, other chromatographic materials, magnetic particles; plastics (including acrylics, polystyrene, copolymers of styrene or other materials, polybutylene, polyurethane, TEFLON®, polyethylene, polypropylene, polyamide, polyester, polyvinylidene fluoride (PVDF), etc.), polysaccharides, nylon or nitrocellulose, resins, silica, or silicone, silica gel and silica-based materials, including modified silicon, Sephadex®, Sepharose®, carbon, metals (e.g., steel, gold, silver, aluminum, silicon, and copper), inorganic glasses, conductive polymers (including polymers such as polypyrrole and polyindole); micro- or nanostructured surfaces, e.g., surfaces modified with nucleic acid tiling arrays, nanotubes, nanowires, or nanoparticulates; or porous surfaces or gels, e.g., methacrylates, acrylamides, sugar polymers, cellulose, silicates, or other fibrous or stranded polymers.In some embodiments, the solid support or substrate may be coated using a passivating coating or a chemically derivatized coating with several materials, including polymers such as dextran, acrylamide, gelatin, or agarose. The beads and / or particles may be free or associated with one another (e.g., sintered). In some embodiments, the solid phase may be a collection of particles. In some embodiments, the particles may comprise silica, which may comprise silicon dioxide. In some embodiments, the silica may be porous, and in certain embodiments, the silica may be non-porous. In some embodiments, the particles further comprise an agent that imparts paramagnetic properties to the particles. In certain embodiments, the agent comprises a metal, and in certain embodiments, the agent is a metal oxide (e.g., iron or iron oxide, where the iron oxide is Fe). 2+ and Fe 3+ (containing a mixture of). The probe oligonucleotide may be linked to the solid support by covalent or non-covalent interactions, and may be linked to the solid support directly or indirectly (e.g., via an intermediary, such as a spacer molecule or biotin). The probe oligonucleotide may be linked to the solid support before, during, or after nucleic acid capture.

[0271] Modified nucleic acids, for example, nucleic acids modified by the addition of an adapter sequence as described herein, can be captured. In some embodiments, unmodified nucleic acids are captured. In some embodiments, the nucleic acid can be amplified before and / or after capture by an amplification process such as PCR. The term "captured nucleic acid" generally includes nucleic acids that have been captured, including nucleic acids that have been captured and amplified. In some embodiments, the captured nucleic acid can be subjected to additional rounds of capture and amplification. The captured nucleic acid can be sequenced, such as by a sequencing process as described herein. Nucleic Acid Sequencing and Processing

[0272] The methods provided herein generally involve the sequencing and analysis of nucleic acids. In some embodiments, nucleic acids are sequenced, and the sequencing products (e.g., a collection of sequence reads) are processed prior to or in conjunction with the analysis of the sequenced nucleic acids. For example, sequence reads may be processed according to one or more of the following: aligning, mapping, portion filtering, portion selection, counting, normalization, weighting, profile generation, etc., and combinations thereof. Certain processing steps may be performed in any order, and certain processing steps may be repeated. For example, portions may be filtered, followed by sequence read count normalization; in certain embodiments, sequence read counts may be normalized, followed by partial filtering. In some embodiments, a partial filtering step is followed by sequence read count normalization, followed by a further partial filtering step. Certain sequencing methods and processing steps are described in further detail below. Sequencing

[0273] In some embodiments, nucleic acids (e.g., nucleic acid fragments, sample nucleic acids, cell-free nucleic acids) are sequenced. In certain cases, complete or substantially complete sequences are obtained, and sometimes partial sequences are obtained. Nucleic acid sequencing generally generates a collection of sequence reads. As used herein, a "read" (e.g., "read," "sequence read") is a short nucleotide sequence generated by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read"), and sometimes from both ends of a nucleic acid fragment (e.g., a paired-end read, a double-end read).

[0274] The length of a sequence read is often related to a particular sequencing technology. High-throughput methods, for example, provide sequence reads that can vary in size from tens of base pairs (bp) to hundreds of base pairs (bp). Nanopore sequencing, for example, can provide sequence reads that can vary in size from tens of base pairs to hundreds to thousands of base pairs. In some embodiments, sequence reads are about 15 bp to about 900 bp in mean, median, average, or absolute length. In certain embodiments, sequence reads are about 1000 bp or longer in mean, median, average, or absolute length. In some embodiments, sequence reads are about 1500, 2000, 2500, 3000, 3500, 4000, 4500, or 5000 bp or longer in mean, median, average, or absolute length. In some embodiments, sequence reads are of a mean, median, average, or absolute length of about 100 bp to about 200 bp. In some embodiments, sequence reads are of a mean, median, average, or absolute length of about 140 bp to about 160 bp. For example, sequence reads can be of a mean, median, average, or absolute length of about 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, or 160 bp.

[0275] In some embodiments, the nominal, average, mean, or absolute length of a single-end read is sometimes from about 10 contiguous nucleotides to about 250 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 200 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 150 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 125 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 100 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 75 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 60 or more contiguous nucleotides, from 15 contiguous nucleotides to about 50 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 40 or more contiguous nucleotides, and sometimes about 15 contiguous nucleotides or about 36 or more contiguous nucleotides. In certain embodiments, the nominal, average, mean, or absolute length of a single-end read is about 20 to about 30 bases, or about 24 to about 28 bases in length. In certain embodiments, the nominal, average, mean, or absolute length of a single-end read is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 28, or about 29 bases in length or longer. In certain embodiments, the nominal, average, mean, or absolute length of a single-end read is about 20 to about 200 bases, about 100 to about 200 bases, or about 140 to about 160 bases in length. In certain embodiments, the nominal, average, mean or absolute length of the single-end reads is about 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190 or about 200 bases in length or longer.In certain embodiments, the nominal, average, mean, or absolute length of a paired-end read is sometimes about 10 contiguous nucleotides to about 25 contiguous nucleotides or longer (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length or longer), sometimes about 15 contiguous nucleotides to about 20 contiguous nucleotides or longer, and sometimes about 17 or about 18 contiguous nucleotides. In certain embodiments, the nominal, average, mean, or absolute length of a paired-end read is sometimes about 25 contiguous nucleotides to about 400 contiguous nucleotides or longer (e.g., about 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400 nucleotides in length or longer), about 50 contiguous nucleotides to about 400 contiguous nucleotides or longer (e.g., about 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400 nucleotides in length or longer), or about 50 contiguous nucleotides to about 400 contiguous nucleotides or longer. The length may be from about 350 contiguous nucleotides or longer, from about 100 contiguous nucleotides to about 325 contiguous nucleotides, from about 150 contiguous nucleotides to about 325 contiguous nucleotides, from about 200 contiguous nucleotides to about 325 contiguous nucleotides, from about 275 contiguous nucleotides to about 310 contiguous nucleotides, from about 100 contiguous nucleotides to about 200 contiguous nucleotides, from about 100 contiguous nucleotides to about 175 contiguous nucleotides, from about 125 contiguous nucleotides to about 175 contiguous nucleotides, and sometimes from about 140 contiguous nucleotides to about 160 contiguous nucleotides. In certain embodiments, the nominal, average, mean, or absolute length of a paired-end read is about 150 contiguous nucleotides, and sometimes 150 contiguous nucleotides.

[0276] In some embodiments, the nucleotide sequence read obtained from a sample is a partial nucleotide sequence read. As used herein, "partial nucleotide sequence read" refers to a sequence read of any length that has incomplete sequence information, also known as sequence ambiguity. A partial nucleotide sequence read may lack information about nucleic acid base identity and / or the position or order of nucleic acid bases. A partial nucleotide sequence read generally does not include a sequence read in which the only incomplete sequence information is due to an inadvertent or unintentional sequencing error (or in which less than all bases are sequenced or determined). Such sequencing errors may be inherent to a particular sequencing process, including, for example, an inaccurate call about nucleic acid base identity and missing or extra nucleic acid bases. Therefore, for the partial nucleotide sequence read herein, certain information about the sequence is often intentionally excluded. That is, sequence information about less than all nucleic acid bases, or that may otherwise be characterized as or may be a sequencing error, is intentionally obtained. In some embodiments, a partial nucleotide sequence read may span a portion of a nucleic acid fragment. In some embodiments, the partial nucleotide sequence reads can span the entire length of the nucleic acid fragment. Partial nucleotide sequence reads are described, for example, in International Patent Application Publication No. WO2013 / 052907, the entire contents of which are hereby incorporated by reference herein, including all text, tables, equations, and figures.

[0277] A read is generally a representation of a nucleotide sequence in a physical nucleic acid. For example, in a read containing an ATGC representation of a sequence, "A" represents an adenine nucleotide, "T" represents a thymine nucleotide, "G" represents a guanine nucleotide, and "C" represents a cytosine nucleotide in the physical nucleic acid. A sequence read obtained from a sample from a subject can be a read from a mixture of minority and majority nucleic acids. For example, a sequence read obtained from the blood of a cancer patient can be a read from a mixture of cancer nucleic acids and non-cancer nucleic acids. In another example, a sequence read obtained from the blood of a pregnant woman can be a read from a mixture of fetal nucleic acids and maternal nucleic acids. A mixture of relatively short reads can be transformed by the process described herein into a representation of genomic nucleic acids present in a subject and / or a representation of genomic nucleic acids present in a tumor or fetus. In certain cases, a mixture of relatively short reads can be transformed into a representation of, for example, copy number alterations, genetic mutations / alterations, or aneuploidy. In one example, a read from a mixture of cancer and non-cancer nucleic acids can be transformed into a representation of a composite chromosome or a portion thereof that includes one or both features of cancer cell chromosomes and non-cancer cell chromosomes. In another example, a read of a mixture of maternal and fetal nucleic acids can be transformed into a representation of a composite chromosome or portion thereof that contains features of one or both maternal and fetal chromosomes.

[0278] In some cases, circulating cell-free nucleic acid fragments (CCF fragments) obtained from a cancer patient include nucleic acid fragments originating from normal cells (i.e., non-cancer fragments) and nucleic acid fragments originating from cancer cells (i.e., cancer fragments). Sequence reads derived from CCF fragments originating from normal cells (i.e., non-cancerous cells) are referred to herein as "non-cancer reads." Sequence reads derived from CCF fragments originating from cancer cells are referred to herein as "cancer reads." CCF fragments from which non-cancer reads are derived may be referred to herein as non-cancer templates, and CCF fragments from which cancer reads are derived may be referred to herein as cancer templates.

[0279] In some cases, circulating cell-free nucleic acid fragments (CCF fragments) obtained from a pregnant female include nucleic acid fragments originating from fetal cells (i.e., fetal fragments) and nucleic acid fragments originating from maternal cells (i.e., maternal fragments). Sequence reads derived from CCF fragments originating from a fetus are referred to herein as "fetal reads." Sequence reads derived from CCF fragments originating from the genome of a pregnant female (e.g., a mother) bearing a fetus are referred to herein as "maternal reads." The CCF fragments from which fetal reads are derived are referred to herein as fetal templates, and the CCF fragments from which maternal reads are derived are referred to herein as maternal templates.

[0280] In certain embodiments, "obtaining" nucleic acid sequence reads for a sample from a subject and / or "obtaining" nucleic acid sequence reads for one or more reference human biological specimens may involve directly sequencing the nucleic acid to obtain sequence information. In some embodiments, "obtaining" may involve receiving sequence information obtained directly from the nucleic acid by another.

[0281] In some embodiments, some or all of the nucleic acids in a sample are enriched and / or amplified (e.g., non-specifically, such as by PCR-based methods) before or during sequencing. In certain embodiments, specific nucleic acid species or subsets in a sample are enriched and / or amplified before or during sequencing. In some embodiments, species or subsets of a preselected pool of nucleic acids are randomly sequenced. In some embodiments, nucleic acids in a sample are not enriched and / or amplified before or after sequencing.

[0282] In some embodiments, a representative fraction of the genome is sequenced, sometimes referred to as "coverage" or "fold coverage." For example, 1-fold coverage indicates that approximately 100% of the nucleotide sequence of the genome is represented by reads. In some cases, fold coverage refers to (and is directly proportional to) "sequencing depth." In some embodiments, "fold coverage" is a relative term that refers to a previous sequencing run as a reference. For example, a second sequencing run may have half the coverage of the first sequencing run. In some embodiments, the genome is sequenced with redundancy, and a given region of the genome may be covered by two or more reads or overlapping reads (e.g., a "fold coverage" greater than 1, e.g., 2-fold coverage). In some embodiments, the genome (e.g., the entire genome) is sequenced at about 0.01x to about 100x coverage, about 0.1x to 20x coverage, or about 0.1x to about 1x coverage (e.g., about 0.015, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90x or greater coverage). In some embodiments, a specific portion of the genome (e.g., a genome portion from a targeted method and / or a probe-based method) is sequenced, and the fold coverage value generally refers to the fraction of the specific genome portion sequenced (i.e., the fold coverage value does not refer to the entire genome). In some cases, the specific genome portion is sequenced at 1000-fold or higher coverage. For example, the specific genome portion may be sequenced at 2000-fold, 5,000-fold, 10,000-fold, 20,000-fold, 30,000-fold, 40,000-fold, or 50,000-fold coverage. In some embodiments, the sequencing is at about 1,000-fold to about 100,000-fold coverage. In some embodiments, the sequencing is at about 10,000-fold to about 70,000-fold coverage.In some embodiments, the sequencing is at about 20,000-fold to about 60,000-fold coverage, hi some embodiments, the sequencing is at about 30,000-fold to about 50,000-fold coverage.

[0283] In some embodiments, one nucleic acid sample from one individual is sequenced.In certain embodiments, the nucleic acid from each of two or more samples is sequenced, and these samples are from one individual or from different individuals.In certain embodiments, the nucleic acid samples from two or more biological samples are pooled, and each biological sample is from one individual or two or more individuals, and the pool is sequenced.In the latter embodiment, the nucleic acid sample from each biological sample is often identified by one or more unique identifiers.

[0284] In some embodiments, the sequencing method utilizes identifiers that allow multiplexing of sequence reactions in the sequencing process.The more unique identifiers there are, the more samples and / or chromosomes that can be multiplexed in the sequencing process for detection.The sequencing process can be carried out using any suitable number of unique identifiers (for example, 4, 8, 12, 24, 48, 96 or more).

[0285] The sequencing process sometimes uses a solid phase, sometimes including a flow cell onto which nucleic acids from a library can be bound and through which reagents can be flowed and contacted with the bound nucleic acids. Flow cells sometimes include flow cell lanes, and the use of identifiers facilitates the analysis of several reagents in each lane. Flow cells are often solid supports that can be configured to hold and / or allow the orderly passage of reagent solutions over bound analytes. Flow cells are frequently planar in shape, optically transparent, generally millimeter or submillimeter scale, and often have channels or lanes within which analyte / reagent interactions occur. In some embodiments, the number of samples analyzed in a given flow cell lane depends on the number of unique identifiers utilized during library preparation and / or probe design. Multiplexing using 12 identifiers, for example, allows for the simultaneous analysis of 96 samples (e.g., equivalent to the number of wells in a 96-well microwell plate) in an 8-lane flow cell. Similarly, multiplexing using 48 identifiers allows for simultaneous analysis of, for example, 384 samples in an 8-lane flow cell (e.g., equivalent to the number of wells in a 384-well microwell plate). Non-limiting examples of commercially available multiplex sequencing kits include Illumina's Multiplexed Sample Preparation Oligonucleotide Kit and Multiplexed Sequencing Primer and PhiX Control Kit (e.g., Illumina catalog numbers PE-400-1001 and PE-400-1002, respectively).

[0286] Any suitable method for sequencing nucleic acids may be used, non-limiting examples of which include Maxim & Gilbert, chain termination methods, sequencing by synthesis, sequencing by ligation, sequencing by mass spectrometry, microscopy-based techniques, and the like, or a combination thereof. In some embodiments, first-generation technologies, such as Sanger sequencing methods, including automated Sanger sequencing methods, including microfluidic Sanger sequencing, may be used in the methods provided herein. In some embodiments, sequencing techniques involving the use of nucleic acid imaging techniques (e.g., transmission electron microscopy (TEM) and atomic force microscopy (AFM)) may be used. In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods generally involve clonally amplified DNA templates or single DNA molecules that are sequenced in a massively parallel format, sometimes within a flow cell. Next-generation (e.g., second- and third-generation) sequencing techniques capable of sequencing DNA in a massively parallel format may be used for the methods described herein, and are collectively referred to herein as "massively parallel sequencing" (MPS). In some embodiments, MPS sequencing methods utilize a targeted approach, where a specific chromosome, gene, or region of interest is sequenced. In certain embodiments, a non-targeted approach is used, where most or all of the nucleic acids in a sample are sequenced, amplified, and / or randomly captured.

[0287] In some embodiments, targeted enrichment, amplification, and / or sequencing approaches are used. Targeted approaches often involve the use of sequence-specific oligonucleotides to isolate, select, and / or enrich a subset of nucleic acids in a sample for further processing. In some embodiments, a library of sequence-specific oligonucleotides is utilized to target (e.g., hybridize to) one or more sets of nucleic acids in a sample. The sequence-specific oligonucleotides and / or primers are often selective for specific sequences (e.g., unique nucleic acid sequences) present in one or more chromosomes, genes, exons, introns, and / or regulatory regions of interest. Any suitable method or combination of methods can be used to enrich, amplify, and / or sequence one or more subsets of targeted nucleic acids. In some embodiments, targeted sequences are isolated and / or enriched by capture onto a solid phase (e.g., flow cell, bead) using one or more sequence-specific anchors. In some embodiments, targeted sequences are enriched and / or amplified by polymerase-based methods (e.g., PCR-based methods with any suitable polymerase-based extension) using sequence-specific primers and / or primer sets. Often, a sequence-specific anchor can be used as a sequence-specific primer.

[0288] MPS sequencing sometimes uses sequencing by synthesis and certain imaging processes.The nucleic acid sequencing technology that can be used in the method described herein is sequencing by synthesis and reversible terminator-based sequencing (for example, Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 2500 (Illumina, San Diego CA)).Using this technology, millions of nucleic acid (e.g., DNA) fragments can be sequenced in parallel.One example of this type of sequencing technology uses a flow cell that contains an optically transparent slide with eight individual lanes that have oligonucleotide anchors (e.g., adapter primers) attached to its surface.

[0289] Sequencing by synthesis is generally performed by iteratively adding (e.g., by covalent addition) nucleotides to a primer or an existing nucleic acid strand in a template-dependent manner. Each iterative addition of a nucleotide is detected, and this process is repeated multiple times until the sequence of the nucleic acid strand is obtained. The length of the resulting sequence depends in part on the number of addition and detection steps performed. In some embodiments of sequencing by synthesis, one, two, three, or more nucleotides of the same type (e.g., A, G, C, or T) are added and detected in a single round of nucleotide addition. Nucleotides can be added by any suitable method (e.g., enzymatically or chemically). For example, in some embodiments, a polymerase or ligase adds nucleotides to a primer or an existing nucleic acid strand in a template-dependent manner. In some embodiments of sequencing by synthesis, different types of nucleotides, nucleotide analogs, and / or identifiers are used. In some embodiments, reversible terminators and / or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In certain embodiments, sequencing by synthesis includes a cleavage (e.g., cleavage and removal of the identifier) ​​and / or a washing step. In some embodiments, the addition of one or more nucleotides is detected by a suitable method described herein or known in the art, non-limiting examples of which include any suitable imaging device, suitable camera, digital camera, CCD (Charge Couple Device)-based imaging device (e.g., CCD camera), CMOS (Complementary Metal Oxide Silicon)-based imaging device (e.g., CMOS camera), photodiode (e.g., photomultiplier tube), electron microscope, field effect transistor (e.g., DNA field effect transistor), ISFET ion sensor (e.g., CHEMFET sensor), etc., or a combination thereof.

[0290] Any suitable MPS method, system or technology platform for carrying out the methods described herein can be used to obtain nucleic acid sequence reads. Non-limiting examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, Helicos True Single Molecule Sequencing, Ion Torrent and Ion semiconductor-based sequencing (e.g., developed by Life Technologies), WildFire, 5500, 5500xl W and / or 5500xl W Genetic Analyzer-based technologies (e.g., developed and sold by Life Technologies, U.S. Patent Application Publication No. 2013 / 0012399); Polony sequencing, pyrosequencing, massively parallel signature sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen systems and methods, Nanopore-based platforms, chemical-sensitive field effect transistor (CHEMFET) arrays, electron microscope-based sequencing (e.g., ZS Other sequencing methods that can be used to practice the methods herein include digital PCR, sequencing by hybridization, nanopore sequencing, chromosome-specific sequencing (e.g., using DANSR (digital analysis of selected regions) technology), etc.

[0291] In some embodiments, sequence reads are generated, obtained, collected, assembled, manipulated, transformed, processed, and / or provided by a sequence module. A machine including a sequence module can be a suitable machine and / or device that determines the sequence of a nucleic acid using sequencing techniques known in the art. In some embodiments, the sequence module can align, assemble, fragment, complement, reverse complement, and / or error check (e.g., error-correct sequence reads). Mapping leads

[0292] The sequence reads can be mapped, and the number of reads that map to a specified nucleic acid region (e.g., a chromosome or portion thereof) is referred to as a count. Any suitable mapping method (e.g., a process, algorithm, program, software, module, etc., or a combination thereof) can be used. Certain aspects of the mapping process are described herein below.

[0293] Mapping nucleotide sequence reads (i.e., sequence information from fragments whose physical genomic location is unknown) can be performed in several ways and often involves aligning the resulting sequence reads with matching sequences in a reference genome. In such alignments, sequence reads are generally aligned to a reference sequence, and the alignment is referred to as being "mapped," "mapped sequence read," or "mapped read." In certain embodiments, mapped sequence reads are referred to as "hits" or "counts." In some embodiments, mapped sequence reads are grouped together and assigned to specific genome portions according to various parameters, which are discussed in more detail below.

[0294] The terms "aligned," "alignment," or "aligning" generally refer to two or more nucleic acid sequences that can be identified as a match (e.g., 100% identity) or a partial match. Alignment can be performed manually or by computer (e.g., software, program, module, or algorithm), a non-limiting example of which is the Efficient Local Alignment algorithm distributed as part of the Illumina Genomics Analysis Pipeline. Examples of computer programs include the Evolution of Nucleotide Data (ELAND) computer program. The alignment of sequence reads can be a 100% sequence match. In some cases, the alignment is less than 100% sequence match (i.e., a non-perfect match, a partial match, or a partial alignment). In some embodiments, the alignment is about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% match. In some embodiments, the alignment includes mismatches. In some embodiments, the alignment includes 1, 2, 3, 4, or 5 mismatches. Two or more sequences can be aligned using either strand (e.g., the sense or antisense strand). In certain embodiments, a nucleic acid sequence is aligned with the reverse complement of another nucleic acid sequence.

[0295] Various computational methods can be used to map each sequence read to a part.Non-limiting examples of computer algorithms that can be used to align sequences include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE1, BOWTIE2, ELAND, MAQ, PROBEMATCH, SOAP, BWA or SEQMAP, or variations or combinations thereof.In some embodiments, sequence reads can be aligned with the sequence in a reference genome.In some embodiments, sequence reads can be found and / or aligned with sequences in nucleic acid databases known in the art, including, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory) and DDBJ (DNA Databank of Japan).BLAST or similar tools can be used to search identified sequences against sequence databases.Then, search hits can be used, for example, to sort identified sequences into appropriate parts (described herein below).

[0296] In some embodiments, a read may be uniquely or non-uniquely mapped to a portion in a reference genome. A read is considered "uniquely mapped" if it aligns with a single sequence in the reference genome. A read is considered "non-uniquely mapped" if it aligns with two or more sequences in the reference genome. In some embodiments, non-uniquely mapped reads are excluded from further analysis (e.g., quantification). In certain embodiments, a certain small degree of mismatch (0-1) may correspond to a single nucleotide polymorphism that may exist between the reference genome and the reads from the individual samples being mapped. In some embodiments, no degree of mismatch allows a read to be mapped to a reference sequence.

[0297] As used herein, the term "reference genome" can refer to any specific known, sequenced, or characterized genome of any organism or virus, whether partial or complete, which can be used to refer to an identified sequence from a subject. For example, reference genomes used for human subjects and many other organisms can be found on the World Wide Web at the National Center for Biotechnology Information at the URL ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus, expressed in nucleic acid sequences. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from one or more individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. In some embodiments, a reference genome includes sequences assigned to chromosomes.

[0298] In certain embodiments, mappability is evaluated for a genomic region (e.g., a portion, a genome portion).Mappability is typically the ability to uniquely align a nucleotide sequence read to a portion of a reference genome up to a specified number of mismatches, including, for example, 0, 1, 2 or more mismatches.For a given genomic region, the expected mappability can be estimated by using a sliding window approach with a preset read length and averaging the mappability values ​​at the read level obtained.Genomic regions that contain a stretch of unique nucleotide sequence sometimes have a high mappability value.

[0299] For paired-end sequencing, reads can be mapped to a reference genome by using an appropriate mapping and / or alignment program, non-limiting examples of which include BWA (Li H. and Durbin R. (2009) Bioinformatics 25, 1754-60), Novoalign[Novocraft (2010)]、Bowtie(Langmead B, et al., (2009) Genome Biol. 10:R25)、SOAP2(Li R, et al., (2009) Bioinformatics 25, 1966-67)、BFAST(Homer, ONE, ONE et al., 2009). 4, e7767)、GASSST(Rizk, G. and Lavenier, D. (2010) Bioinformatics 26, 2534-2540) and MPscan (Rivals E., et al. (2009) Lecture Notes in Computer Science 5724, 246-260). Paired-end reads can be mapped and / or aligned using an appropriate short read alignment program. Non-limiting examples of short read alignment programs include BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, BWA, CASHX, CUDA-EC, CUSHAW, CUSHAW2, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP, and Geneious. Examples of suitable sequencing methods include Assembler, iSAAC, LAST, MAQ, mrFAST, mrsFAST, MOSAIK, MPscan, Novoalign, NovoalignCS, Novocraft, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RTG, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3, SOCS, SSAHA, SSAHA2, Stampy, SToRM, Subread, Subjunc, Taipan, UGENE, VelociMapper, TimeLogic, XpressAlign, ZOOM, etc., or combinations thereof. Paired-end reads are often mapped to opposite ends of the same polynucleotide fragment according to a reference genome. In some embodiments, readmates are independently mapped. In some embodiments, information from both sequence reads (i.e., from each end) is factored in the mapping process. A reference genome is often used to determine and / or infer the sequence of nucleic acids located between paired-end readmates. The term "discordant read pair," as used herein, refers to a paired-end read that includes a pair of readmates, where one or both readmates cannot be uniquely mapped to the same region of the reference genome defined in part by a segment of contiguous nucleotides.In some embodiments, discordant read pairs are paired end read mates that map to unexpected locations in reference genome.Non-limiting examples of unexpected locations in reference genome include: (i) two different chromosomes; (ii) locations that are separated by a distance longer than a predetermined fragment size (for example, longer than 300bp, longer than 500bp, longer than 1000bp, longer than 5000bp, or longer than 10,000bp); (iii) orientations that do not match reference sequence (for example, opposite orientations), etc., or combinations thereof.In some embodiments, discordant read mates are identified according to the length (for example, average length, predetermined fragment size) or expected length of template polynucleotide fragments in sample.For example, read mates that map to locations that are separated by a distance longer than the average length or expected length of polynucleotide fragments in sample are sometimes identified as discordant read pairs. Read pairs that map in opposite orientations are sometimes determined by choosing the reverse complement of one of the reads and comparing the alignment of both reads using the same strand of the reference sequence. Discordant read pairs can be identified by any suitable method and / or algorithm known in the art or described herein (e.g., SVDetect, Lumpy, BreakDancer, BreakDancerMax, CREST, DELLY, etc., or a combination thereof). portion

[0300] In some embodiments, the mapped sequence reads are grouped together according to various parameters and assigned to specific genome portions (e.g., portions of the reference genome). A "portion" may also be referred to herein as a "genomic region," "bin," "distribution," "portion of the reference genome," "portion of a chromosome," or "genomic portion."

[0301] The portions are often defined by the distribution of the genome according to one or more characteristics. Non-limiting examples of certain distributional characteristics include length (e.g., fixed length, non-fixed length) and other structural characteristics. Genomic portions sometimes include one or more of the following characteristics: fixed length, non-fixed length, random length, non-random length, equal length, unequal length (e.g., at least two of the genome portions are of unequal length), non-overlapping (e.g., the 3' end of a genome portion sometimes abuts the 5' end of an adjacent genome portion), overlapping (e.g., at least two of the genome portions overlap), contiguous, consecutive, non-contiguous, and non-contiguous. The genome portion is sometimes about 1 to about 1,000 kilobases in length (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900 kilobases in length), about 5 to about 500 kilobases in length, about 10 to about 100 kilobases in length, or about 40 to about 60 kilobases in length.

[0302] The distribution is sometimes based on, or in part based on, certain information features, such as information content and information acquisition. Non-limiting examples of certain information features include alignment speed and / or ease, variability in sequencing coverage, GC content (e.g., stratified GC content, specific GC content, high or low GC content), GC content uniformity, other measures of sequence content (e.g., fraction of individual nucleotides, fraction of pyrimidines or purines, fraction of natural versus non-natural nucleic acids, fraction of methylated nucleotides, and CpG content), methylation status, duplex melting temperature, amenability to sequencing or PCR, uncertainty values ​​assigned to individual parts of the reference genome, and / or targeted searches for specific features. In some embodiments, information content can be quantified using p-value profiles that measure the significance of specific genomic locations for distinguishing between groups of confirmed normal and abnormal subjects (e.g., euploid and trisomic subjects, respectively).

[0303] In some embodiments, splitting the genome may exclude similar regions (e.g., identical or homologous regions or sequences) across the genome and retain only unique regions. The regions removed during splitting may be within a single chromosome, one or more chromosomes, or may span multiple chromosomes. In some embodiments, the split genome is reduced and optimized for faster alignment, often focusing on uniquely identifiable sequences.

[0304] In some embodiments, genome portions result from a non-overlapping, fixed-size partitioning that results in contiguous, non-overlapping portions of fixed length. Such portions are often shorter than chromosomes, and often shorter than regions of copy number variation (or copy number alteration) (e.g., duplicated or deleted regions), which may be referred to as segments. A "segment" or "genomic segment" often includes two or more fixed-length genome portions, and often includes two or more contiguous, fixed-length portions (e.g., about 2 to about 100 such portions (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 such portions)).

[0305] A plurality of parts are sometimes analyzed in groups, and sometimes the reads mapped to parts are quantified according to a specific group of genome parts. When parts are distributed according to structural features and correspond to regions in the genome, parts are sometimes grouped into one or more segments and / or one or more regions. Non-limiting examples of regions include subchromosomes (i.e., shorter than chromosomes), chromosomes, autosomes, sex chromosomes, and combinations thereof. One or more subchromosomal regions are sometimes genes, gene fragments, regulatory sequences, introns, exons, segments (e.g., segments spanning copy number alteration regions; segments spanning copy number variation regions), microduplications, microdeletions, etc. A region is sometimes smaller than or the same size as a target chromosome, and sometimes smaller than or the same size as a reference chromosome. Filtering and / or selecting parts

[0306] In some embodiments, the one or more processing steps may include one or more portion filtering and / or portion selection steps. The term "filtering," as used herein, refers to removing a portion or portions of a reference genome from consideration. In certain embodiments, one or more portions are filtered (e.g., subjected to a filtering process), thereby providing filtered portions. In some embodiments, the filtering process removes certain portions and retains portions (e.g., a subset of portions). After the filtering process, the retained portions are often referred to herein as filtered portions.

[0307] Portions of the reference genome may be selected for removal based on any suitable criteria, including, but not limited to, redundant data (e.g., redundant or overlapping mapped reads), non-informative data (e.g., portions of the reference genome with a median count of zero), portions of the reference genome with over- or under-represented sequences, noisy data, etc., or combinations of the above. The filtering process often involves removing one or more portions of the reference genome from consideration and subtracting the counts in one or more portions of the reference genome selected for removal from the counted or totaled counts for the portion, chromosome(s), or genome of the reference genome being considered. In some embodiments, portions of the reference genome may be removed sequentially (e.g., one at a time to allow for evaluation of the impact of removing each individual portion), and in certain embodiments, all portions of the reference genome marked for removal may be removed simultaneously. In some embodiments, portions of the reference genome characterized by variance above or below a certain level are removed, which is sometimes referred to herein as filtering "noisy" portions of the reference genome. In certain embodiments, the filtering process comprises obtaining data points from the dataset that deviate from the mean profile level of the portion, chromosome, or portion of a chromosome by a predetermined multiple profile variance, and in certain embodiments, the filtering process comprises removing data points from the dataset that do not deviate from the mean profile level of the portion, chromosome, or portion of a chromosome by a predetermined multiple profile variance. In some embodiments, the filtering process is utilized to reduce the number of candidate portions of the reference genome to be analyzed for the presence or absence of genetic mutations / alterations and / or copy number alterations (e.g., aneuploidies, microdeletions, microduplications).Reducing the number of candidate portions of a reference genome that are analyzed for the presence or absence of genetic variants / alterations and / or copy number alterations often reduces the complexity and / or dimensionality of the dataset, sometimes increasing the speed at which genetic variants / alterations and / or copy number alterations can be searched for and / or identified by two orders of magnitude or more.

[0308] Portions can be processed (e.g., filtered and / or selected) by any suitable method and according to any suitable parameters. Non-limiting examples of features and / or parameters that can be used to filter and / or select portions include redundant data (e.g., redundant or overlapping mapped reads), non-informative data (e.g., portions of the reference genome with a zero mapped count), portions of the reference genome with over- or under-represented sequences, noisy data, counts, count variability, coverage, mappability, variability, reproducibility measures, read density, read density variability, level of uncertainty, guanine-cytosine (GC) content, CCF fragment length and / or read length (e.g., fragment length ratio (FLR), fetal ratio statistic (FRS)), DNase I sensitivity, methylation status, acetylation, histone distribution, chromatin structure, percent repeats, etc., or combinations thereof. Portions can be filtered and / or selected according to any suitable feature or parameter that correlates with the features or parameters listed or described herein. Portions may be filtered and / or selected according to portion-specific features or parameters (e.g., determined for a single portion according to multiple samples) and / or sample-specific features or parameters (e.g., determined for multiple portions within a sample). In some embodiments, portions are filtered and / or removed according to relatively low mappability, relatively high variability, high level of uncertainty, relatively long CCF fragment lengths (e.g., low FRS, low FLR), relatively large fraction of repetitive sequences, high GC content, low GC content, low counts, zero counts, high counts, etc., or combinations thereof. In some embodiments, portions (e.g., subsets of portions) are selected according to an appropriate level of mappability, variability, level of uncertainty, fraction of repetitive sequences, counts, GC content, etc., or combinations thereof. In some embodiments, portions (e.g., subsets of portions) are selected according to relatively short CCF fragment lengths (e.g., high FRS, high FLR).The counts and / or reads mapped to the portions are sometimes processed (e.g., normalized) before and / or after filtering or selecting the portions (e.g., subsets of the portions). In some embodiments, the counts and / or reads mapped to the portions are not processed before and / or after filtering or selecting the portions (e.g., subsets of the portions).

[0309] In some embodiments, the portions may be filtered according to an error measure (e.g., standard deviation, standard error, calculated variance, p-value, mean absolute error (MAE), average absolute deviation, and / or mean absolute deviation (MAD)). In certain cases, the error measure may be referred to as count variability. In some embodiments, the portions are filtered according to count variability. In certain embodiments, the count variability is a measure of error determined for counts mapped to portions (i.e., portions) of a reference genome for multiple samples (e.g., multiple samples obtained from multiple subjects, e.g., 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more, or 10,000 or more subjects). In some embodiments, portions having count variability above the upper limit of a predetermined range are filtered (e.g., removed from consideration). In some embodiments, portions having count variability below the lower limit of a predetermined range are filtered (e.g., removed from consideration). In some embodiments, fractions with count variability outside a predetermined range are filtered (e.g., eliminated from consideration). In some embodiments, fractions with count variability within a predetermined range are selected (e.g., used to determine the presence or absence of copy number alterations). In some embodiments, the count variability of the fractions exhibits a distribution (e.g., a normal distribution). In some embodiments, fractions within a quantile of the distribution are selected. In some embodiments, fractions within the 99% quantile of the distribution of count variability are selected.

[0310] Sequence reads from any suitable number of samples can be used to identify a subset of portions that meet one or more criteria, parameters and / or features described herein. Sometimes, sequence reads from a group of samples from multiple subjects are used. In some embodiments, the multiple subjects include pregnant females. In some embodiments, the multiple subjects include healthy subjects. In some embodiments, the multiple subjects include cancer patients. One or more samples from each of a plurality of subjects may be performed (e.g., 1 to about 20 samples from each subject (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 samples)), and any suitable number of subjects may be performed (e.g., about 2 to about 10,000 subjects (e.g., about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000 subjects)). In some embodiments, sequence reads from the same test sample(s) from the same subject are mapped to portions in the reference genome and used to generate a subset of portions.

[0311] The portions may be selected and / or filtered by any suitable method. In some embodiments, the portions are selected according to visual inspection of the data, graphs, plots, and / or charts. In certain embodiments, the portions are selected and / or filtered (e.g., in part) by a system or machine including one or more microprocessors and a memory. In some embodiments, the portions are selected and / or filtered (e.g., in part) by a non-transitory computer-readable storage medium having an executable program stored thereon, the program instructing the ...

Claims

[Claim 1] A method for classifying the presence or absence of genetic mosaicism in one or more fetuses, comprising: identifying, by a computing device, regions of genetic copy number variation in a sample comprising circulating cell-free nucleic acid from a pregnant female subject having a multiple pregnancy, wherein the regions of genetic copy number variation comprise copy number variation and the circulating cell-free nucleic acid comprises maternal nucleic acid and fetal nucleic acid; determining, by the computing device, the fraction of nucleic acid in the circulating cell-free nucleic acid that has the copy number variation; determining, by the computing device, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid; generating, by the computing device, a mosaicism ratio, the mosaicism ratio being the fraction of nucleic acids in the circulating cell-free nucleic acids that have the copy number variation divided by the fraction of the fetal nucleic acids in the circulating cell-free nucleic acids; and classifying, by said computing device, the presence or absence of genetic mosaicism for said copy number variation region according to said mosaicism ratio and said mosaicism ratio based on the number of fetuses carried by said pregnant female subject. A method comprising: