Methods and processes for genetic mosaicism
The MR method addresses false-positive NIPT results by classifying genetic mosaicism, enhancing prenatal care through accurate interpretation of NIPT outcomes and reducing invasive testing.
Patent Information
- Application Number
- JP2023120117
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-03-17
- Filing Date
- 2023-07-24
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2038-03-19
AI Technical Summary
Current non-invasive prenatal testing (NIPT) methods face challenges with false-positive results due to placental confined mosaicism, leading to unnecessary invasive procedures and uncertainty in clinical decision-making.
A method to classify the presence or absence of genetic mosaicism using a mosaicism ratio (MR) derived from the fraction of nucleic acids with copy number variations in maternal and fetal DNA, allowing for a non-invasive approach to confirm NIPT results.
The MR enables accurate classification of genetic mosaicism, reducing false-positive NIPT results and improving prenatal care by providing clearer post-test counseling and reducing the need for invasive procedures.
Smart Images

Figure 0007746338000018 
Figure 0007746338000019 
Figure 0007746338000020
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 473,074, filed March 17, 2017, the entire contents of which are incorporated herein by reference in their entirety.
[0002] Field The technology provided herein relates in part to the method, system, machine and computer program product for non-invasively classifying the mosaic copy number variation (CNV) of test sample.The technology provided herein is useful for classifying the mosaic CNV of sample, for example, as part of non-invasive prenatal testing (NIPT) and oncology testing. [Background technology]
[0003] (background) The genetic information of living organisms (e.g., animals, plants, and microorganisms) and other forms that replicate genetic information (e.g., viruses) is encoded in deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a sequence of nucleotides or modified nucleotides that represent the chemical or hypothetical primary structure of nucleic acids. In humans, the complete genome contains approximately 30,000 genes located on 24 chromosomes (i.e., 22 autosomes, an X chromosome, and a Y chromosome; see *The Human Genome*, T. Strachan, BIOS Scientific Publishers, 1992). Each gene encodes a specific protein, which, after expression via transcription and translation within living cells, performs a specific biochemical function. Many medical conditions are caused by one or more gene variations and / or gene changes.Specific gene variations and / or gene changes cause medical conditions, such as hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease, and cystic fibrosis (CF) (Human Genome Mutations, DN Cooper and M. Krawczak, BIOS Publishers, 1993).Such genetic diseases can result from the addition, substitution, or deletion of a single nucleotide in the DNA of a specific gene.For example, certain congenital defects are caused by chromosomal abnormalities, also known as aneuploidies, such as trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), monosomy X (Turner syndrome), and certain sex chromosome aneuploidies, such as Klinefelter syndrome (XXY). Another genetic variation is the sex of the fetus, which can often be determined based on the sex chromosomes X and Y. Some genetic variations can predispose an individual to or develop any of several diseases, such as diabetes, arteriosclerosis, obesity, various autoimmune diseases, and cell proliferation disorders such as cancer, tumors, neoplasia, metastatic disease, etc., or a combination thereof. The cancer, tumor, neoplasia, or metastatic disease can be a disorder or condition of the liver, lung, spleen, pancreas, colon, skin, bladder, eye, brain, esophagus, head, neck, ovaries, testes, prostate, etc., or a combination thereof. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] The Human Genome, T. Strachan, BIOS Scientific Publishers, 1992 [Non-patent document 2] Human Genome Mutations, D.N. Cooper and M. Krawczak, BIOS Publishers, 1993 Summary of the Invention [Means for solving the problem]
[0005] Identification of one or more genetic variations and / or alterations (e.g., copy number alterations, copy number variations, single nucleotide alterations, single nucleotide variations, chromosomal alterations, translocations, deletions, insertions, etc.) or variances can lead to the diagnosis of a particular medical condition or the determination of a predisposition to such a condition. Identification of genetic variances can facilitate medical decisions and / or lead to the utilization of useful medical procedures. In certain embodiments, the identification of one or more genetic variations and / or alterations involves the analysis of circulating cell-free nucleic acids. Circulating cell-free nucleic acids (CCF-NAs), such as cell-free DNA (CCF-DNA), are composed of DNA fragments that arise, for example, from cell death and circulate in the peripheral blood. High concentrations of CCF-DNA can be indicative of certain clinical conditions, such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infection, and other diseases. Furthermore, cell-free fetal DNA (CFF-DNA) can be detected in the maternal bloodstream and used for various non-invasive prenatal diagnostic methods.
[0006] One or more computer systems can be configured to perform specific operations or actions by having software, firmware, hardware, or a combination thereof installed on the system that causes an action or causes the system to perform an action during operation. One or more computer programs can be configured to perform specific operations or actions by including instructions that, when executed by a data processing device, cause the device to perform an action. One general aspect includes a method for classifying the presence or absence of genetic mosaicism in a biological sample, the method including: (a) identifying regions of genetic copy number variation in sample nucleic acid from a subject, the sample nucleic acid including abundant and rare nucleic acids; (b) determining the fraction of nucleic acids having copy number variation in the sample nucleic acid; (c) determining the fraction of rare nucleic acids in the sample nucleic acid; (d) comparing the fraction in (b) with the fraction in (c), thereby providing a comparison; and (e) classifying the presence or absence of genetic mosaicism for the regions of copy number variation according to the comparison.
[0007] Various aspects include a method for classifying the presence or absence of genetic mosaicism in a biological sample, the method including the steps of: identifying, by a computing device, regions of genetic copy number variation in a sample comprising circulating cell-free nucleic acid from a pregnant female subject, wherein the regions of genetic copy number variation comprise copy number variation, and the circulating cell-free nucleic acid comprises maternal nucleic acid and fetal nucleic acid; determining, by the computing device, a fraction of nucleic acids having copy number variation in the circulating cell-free nucleic acid; determining, by the computing device, a fraction of fetal nucleic acid in the circulating cell-free nucleic acid; comparing, by the computing device, the fraction of nucleic acids having copy number variation in the circulating cell-free nucleic acid with the fraction of fetal nucleic acid in the circulating cell-free nucleic acid, thereby producing a comparison and generating a mosaicism ratio; and classifying, by the computing device, the presence or absence of genetic mosaicism for the regions of copy number variation according to the comparison and the mosaicism ratio. A mosaicism ratio between about 0.2 and about 0.7 classifies the presence of genetic mosaicism for regions of copy number variation, and a ratio between about 0.71 and about 1.3 classifies the absence of genetic mosaicism for regions of copy number variation.
[0008] Implementations may include one or more of the following features: the method, wherein the fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is determined for regions of copy number variation; the method, wherein the fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is determined according to a sequencing-based fraction estimation; the method, wherein the fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is determined according to the allelic ratio of a polymorphic sequence; the method, wherein the fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is determined according to quantification of methylation-variable nucleic acids; the method, wherein the fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is a fetal fraction determined for regions of copy number variation; the method, wherein the fetal fraction of nucleic acids with copy number variations in circulating cell-free nucleic acids is determined according to a sequencing-based fetal fraction estimation.
[0009] Implementations may also include one or more of the following features: a method in which the fetal fraction of nucleic acids having copy number variations in circulating cell-free nucleic acids is determined according to the ratio of alleles of polymorphic sequences in fetal and maternal nucleic acids; a method in which the fetal fraction of nucleic acids having copy number variations in circulating cell-free nucleic acids is determined according to quantification of methylation-variable fetal and maternal nucleic acids; a method in which the fraction of fetal nucleic acids in circulating cell-free nucleic acids is determined for genomic regions that are larger than the region of copy number variation; a method in which the fraction of fetal nucleic acids in circulating cell-free nucleic acids is determined for genomic regions that are different from the region of copy number variation; a method in which the fraction of fetal nucleic acids in circulating cell-free nucleic acids is determined according to a fetal fraction estimation based on sequencing; a method in which the fraction of fetal nucleic acids in circulating cell-free nucleic acids is determined according to the ratio of alleles of polymorphic sequences in fetal and maternal nucleic acids; a method in which the fraction of fetal nucleic acids in circulating cell-free nucleic acids is determined according to quantification of methylation-variable fetal and maternal nucleic acids. The method, wherein the mosaicism ratio is the fraction of nucleic acids with copy number variations in the circulating cell-free nucleic acid divided by the fraction of fetal nucleic acid in the circulating cell-free nucleic acid.
[0010] Implementations may also include one or more of the following features: the method further comprising the step of providing, by the computing system, a no classification if the mosaicism ratio is less than a minimum threshold; the method wherein the minimum threshold is about 0.2; the method further comprising the step of providing, by the computing system, a no classification if the mosaicism ratio is greater than a maximum threshold; the maximum threshold is about 1.3; the method further comprising the step of obtaining, by the computing system, a positive screening result from non-invasive prenatal testing (NIPT) for the presence of one or more aneuploidies in a sample comprising circulating cell-free nucleic acid from a pregnant female subject; the method further comprising the step of providing, by the computing system, interpreting the positive screening result from the NIPT as a negative result or absence of one or more aneuploidies if the mosaicism ratio is less than the minimum threshold; the method further comprising the step of providing, by the computing system, a no classification and interpreting the positive screening result from the NIPT as excessive or indeterminate if the mosaicism ratio is greater than the maximum threshold. The method further comprises a step of providing that if the computing system classifies the presence of genetic mosaicism for the region of copy number variation, the positive screening result from the NIPT is interpreted as positive with a comment regarding the likelihood of mosaic presentation.
[0011] Other embodiments of these aspects include corresponding computer systems, apparatus and computer programs stored on one or more computer storage devices, each configured to perform the functions of the method.
[0012] Various embodiments are further described in the following description, examples, claims, and drawings.
[0013] The drawings illustrate, but are not limiting, certain embodiments of the present technology. For clarity and ease of description, the drawings have not been made to scale and in some instances various aspects may be shown exaggerated or enlarged to facilitate an understanding of particular embodiments. [Brief explanation of the drawings]
[0014] [Figure 1] Figure 1 shows the early cell lineages after conception (Figure from Thomas, D. et al. (July 10, 1994) Trisomy 22, placenta; World Wide Web URL sonoworld.com / Fetus / page.aspx?id=182). The majority of cells develop into placental trophoblast / chorionic ectoderm (direct chorionic villus sampling (CVS) preparation, NIPT). A small minority of cells develop into chorionic villi / mesoderm (CVS culture cells). Two cells in this image go on to form the embryo and amniotic tissue (amniocentesis).
[0015] [Figure 2] FIG. 2 is a diagram illustrating a process flow consistent with various embodiments.
[0016] [Figure 3] FIG. 3 is a diagram illustrating a process flow consistent with various embodiments.
[0017] [Figure 4] FIG. 4 illustrates an exemplary embodiment of a system in which various embodiments of the technology can be implemented.
[0018] [Figure 5] Figure 5 shows the distribution of risk indices in the study population based on the information provided in the sample request form by the ordering physician for each study. AMA - advanced maternal age, US - abnormal ultrasound findings, AS - abnormal serum screening results, HIST - personal and / or family history, "other" - other reasons. The inner circle shows the risk indices for patients using MaterniT21® PLUS (n>500,000), and the outer circle shows the risk indices from MaterniT® GENOME (n>10,000).
[0019] [Figure 6]Figure 6 illustrates the positivity rate and type of finding by risk index. The left panel shows the positivity rate stratified by risk index and grouped by type of positive finding. The positivity rate graph reflects the positivity rate by index: the top bar is for "GENOME-only findings," the second / middle bar is for sex chromosome aneuploidies (SCAs), and the bottom bar is for core trisomies (13, 18, 21). The right panel shows the contribution of each type of positive finding to the positive cohort per risk group. The percent positive graph shows that 30% of "GENOME-only" findings occur, and there is a higher rate of "these unique results" in patients with "AMA-only," whereas there is a lower rate in patients with ultrasound findings (USF) or (serum biochemistry screening) SBS-marked patients. Risk indicators include: AMA - older maternal age; US - abnormal ultrasound findings; AS - abnormal serum screening; HIST - family history. Finding stratification includes (top to bottom in each bar graph): GENOME - genome-wide, SCA - sex chromosome aneuploidy, 13 / 18 / 21 - trisomy 13 / 18 / 21. The study cohort average genome-wide contribution of 30% is indicated by the 0.7 line.
[0020] [Figure 7] Figure 7 shows the agreement between fetal fraction based on SeqFF (x-axis) and fetal fraction estimates based on deviation of affected chromosomes from the population median (affected fraction (AF); y-axis). Parallel lines in the graph highlight the 95% confidence interval of the regression line describing the relationship between the two fetal fraction estimates.
[0021] [Figure 8] Figure 8 shows a histogram showing the prevalence of copy number variations (CNVs) in each size group among positive samples. Size groups are in megabases.
[0022] [Figure 9] Figure 9 shows the mosaicism ratio of cfDNA positive aneuploidy results.
[0023] [Figure 10] FIG. 10 shows the conflicting results as a function of mosaicism ratio.
[0024] [Figure 11] FIG. 11 shows the effect of mosaicism ratio on positive predictive value.
[0025] [Figure 12] FIG. 12 shows a portion of a MaterniT® GENOME report with detailed comments and ideograms of predicted events.
[0026] [Figure 13] FIG. 13 shows the ideogram of chromosome 12 of case A.
[0027] [Figure 14] FIG. 14 shows the chromosome 12 ideogram of case B.
[0028] [Figure 15] FIG. 15 shows the chromosome 12 ideogram of case C.
[0029] [Figure 16] FIG. 16 shows the whole genome profile of 12p duplication suggesting iso(12p) in case C.
[0030] [Figure 17] FIG. 17 shows the correlation (R=0.81, RMedSE=1.5) of predicted fetal fraction percentages for 19,312 test samples derived from a bin-based fetal fraction (BFF; also referred to herein as sequencing-based fetal fraction (SeqFF)) model based on 6000 training samples (x-axis) compared to fetal fraction percentages determined from the chromosome Y level (ChrFF, y-axis).
[0031] [Figure 18] Figure 18 shows the relative prediction error (x-axis) for bins (i.e., portions) with high fetal fraction content (distribution shown on the left) and bins with low fetal fraction content (distribution shown on the right) based on the fetal ratio statistic (FRS). Bins with high fetal content have better performance and lower error. Prediction scores are based on an elastic net regression procedure and use bootstrapping to obtain density profiles.
[0032] [Figure 19] Figure 19 illustrates four distributions of model coefficients (x-axis) determined using the elastic net regression procedure on subsets of bins separated according to fetal fraction content (e.g., low, medium-low, medium-high, high). Bins (i.e., portions) with higher fetal fraction content tend to yield higher coefficients (positive or negative).
[0033] [Figure 20] 20 shows two distributions of fetal fraction estimates (x-axis) determined using the bin-based fetal fraction (BFF; also referred to herein as sequencing-based fetal fraction (SeqFF)) method for female and male test samples. The two distributions substantially overlap. Male and female fetuses showed no difference in the distribution of fetal fraction (KS test P=0.49).
[0034] [Figure 21] FIG. 21 shows a four-group Venn diagram of samples with high-risk indicators detailing the co-occurrence of high-risk indicators.
[0035] [Figure 22]Figure 22 shows a bar graph of high-risk indicators: AMA: samples with advanced maternal age as a high-risk indicator, US: samples with ultrasound high-risk indicator, AS: samples with abnormal serum screening as a high-risk indicator; HIST: samples with personal or family history, Other: samples with other high-risk indicators or no high-risk indicators.
[0036] [Figure 23] Figure 23 shows each sample as a column and high-risk indicators as rows. AMA: samples with advanced maternal age as a high-risk indicator; US: samples with ultrasound high-risk indicators; AS: samples with abnormal serum screening as a high-risk indicator; HIST: samples with personal or family history; Other indicators: samples with other high-risk indicators or no high-risk indicators. Dark areas indicate that this indicator was not marked on the test request form. Light areas indicate that this indicator was marked on the test request form. DETAILED DESCRIPTION OF THE INVENTION
[0037] Detailed Description Provided herein are systems and methods for classifying the presence or absence of genetic mosaicism in biological samples. In various embodiments, bioinformatics tools and processes are used to classify the presence or absence of genetic mosaicism for copy number variations. The methods described herein can be applied to a variety of polynucleotides, including, for example, fragmented or cleaved nucleic acids, nucleic acid templates, cellular nucleic acids, and / or cell-free nucleic acids. In some embodiments, the sample nucleic acids subjected to the sequencing process and the resulting sequence reads are further analyzed to identify genetic copy number variations in a sample containing circulating cell-free nucleic acids from a pregnant female subject. The sample nucleic acids can include maternal nucleic acids and fetal nucleic acids. In some embodiments, the fraction of maternal nucleic acids with copy number variations in the sample nucleic acids is determined, and the fraction of fetal nucleic acids with copy number variations in the sample nucleic acids is determined. The polymorphic sequence of the maternal nucleic acid is different from the polymorphic sequence of the fetal nucleic acid. In some embodiments, the fraction of maternal nucleic acids with copy number variations is compared with the fraction of fetal nucleic acids with copy number variations to obtain a ratio of the fraction of maternal nucleic acids with copy number variations to the fraction of fetal nucleic acids with copy number variations. In some embodiments, genetic mosaicism is classified based on the ratio of the fraction of maternal nucleic acids with copy number variations to the fraction of fetal nucleic acids with copy number variations. In certain embodiments, a ratio of about 0.2 to about 0.7 classifies the presence of genetic mosaicism for a copy number variation, and a ratio of about 0.6 to about 1.0 classifies the absence of genetic mosaicism for a copy number variation. As used herein, when an action, such as a determination, is "induced by," "follows," or "based on" something, this means that the action is at least somewhat induced, follows, or is based, at least in part, on something. Classifying genetic mosaicism for a particular copy number variation can provide useful information about the copy number variation to health care professionals and patients.
[0038] In some embodiments, systems, machines and computer program products are also provided that implement the methods or parts of the methods described herein.
[0039] introduction The detection of cell-free nucleic acids in fluid samples, particularly those derived from pregnant subjects, offers great potential for use in noninvasive prenatal testing. Cell-free nucleic acid screening, or noninvasive prenatal testing (NIPT), is a screening test that utilizes bioinformatics tools and processes and next-generation sequencing of DNA fragments in maternal serum to determine the likelihood of certain chromosomal conditions during pregnancy. Every individual has their own cell-free DNA in their bloodstream. During pregnancy, cell-free fetal DNA derived from the placenta (mainly trophoblast cells) also enters the maternal bloodstream and mixes with maternal cell-free DNA. The DNA of trophoblast cells normally reflects the chromosomal makeup of the fetus. Cell-free nucleic acids are routinely screened for trisomy 21, trisomy 18, and trisomy 13. Screening for other conditions, such as fetal sex, sex chromosome aneuploidy, other aneuploidies, triploidy, and certain microdeletion conditions, is also available. Abnormal results usually indicate an increased risk of a particular condition. However, abnormal results are not diagnostic, and the patient must be offered confirmatory testing by a diagnostic procedure such as amniocentesis. Abnormal results may indicate an affected fetus, but may also represent a false-positive result in an unaffected pregnancy, localized placental mosaicism, placental and fetal mosaicism, vanishing twin, unrecognized maternal condition, or other unknown in vivo conditions.
[0040] In particular, in prenatal cell-free DNA testing, there may be discrepancies between analytical performance, sensitivity, specificity, clinical performance, and positive predictive value (PPV), which has caused challenges in interpreting positive NIPT results. One of the main underlying causes of this discrepancy or discordant results is the difference between the genetic makeup of the placenta and the fetus. Chromosomal abnormalities restricted to the placenta are often mosaic and may be localized to the placenta. For example, in most pregnancies, the chromosome set detected in the fetus is also present in the placenta. Since both arise from the same zygote, detection of the same chromosome set in both the fetus and the placenta is expected. However, in approximately 2% of viable pregnancies studied by chorionic villus sampling (CVS) at 9-11 weeks of gestation, the cytogenetic abnormality, most often a trisomy, can be confined to the placenta (see, e.g., Kalousek DK, Vekemans M. Confined placental mosaicism. Journal of Medical Genetics. 1996, 33(7):529-533). This phenomenon is known as placental confined mosaicism (CPM). In contrast to placental and fetal mosaicism, which is characterized by the presence of two or more karyotypically distinct cell lineages in both the fetus and placenta, CPM represents a discrepancy between the chromosomal makeup of cells in the placenta and cells in the fetus. As a result, although CPM is usually associated with a normal fetal outcome (e.g., most commonly, CPM represents a trisomic cell lineage in the placenta and a normal diploid chromosome set in the fetus), it can be misinterpreted from a diagnostic perspective (i.e., a false-positive result in NIPT).
[0041] Given that NIPT can yield false positives, positive NIPT results are typically confirmed using invasive testing such as CVS and / or amniocentesis. For example, prenatal care is typically not a separate event but rather a continuous 40-week period of care for a patient. Therefore, each data point collected throughout pregnancy should provide significant clinically relevant information, allowing clinicians to contextualize all available information. Ideally, clinical data including CVS and / or amniocentesis analysis on all positive NIPT results would help mitigate concerns about false positives before making irreversible treatment decisions (such as terminating the pregnancy). However, CPM can also cause false-positive results in CVS. Therefore, conventional practice is to proceed with CVS and examine all cell lineages using both uncultured samples or short-term and long-term cultures of the samples using fluorescence in situ hybridization (FISH). If all results indicate aneuploidy, the results are reported to the patient. Otherwise, if the results also indicate mosaicism, amniocentesis is recommended and analyzed by both FISH and karyotype. Nevertheless, a real-world limitation to conventional practice is that not all women consent to invasive diagnostic testing, especially in the first trimester.
[0042] To address these false-positive problems and the unwillingness of many women to undergo invasive diagnostic testing, various embodiments described herein use the mosaicism ratio (a newly discovered metric derived from prenatal cell-free DNA testing, described in detail herein) to identify patients in whom aneuploidy may be present in a mosaic form (e.g., CPM). As shown in FIG. 1, the majority of cells develop from the zygote into placental trophoblast cells / chorionic ectoderm 105, a small minority of cells develop into chorionic villi / mesoderm 110, and only two cells proceed to form the embryo and amniotic tissue 115. Errors in cell division at different levels in this chain can lead to different levels of fetal or placental (or both) mosaicism, which can have fundamentally different clinical implications. In this case, less than all cell-free trophoblast DNA in maternal plasma is affected. Using this knowledge, the mosaicism ratio (MR) of the affected cell-free DNA and total cell-free DNA can be calculated. In various embodiments, the MR is calculated by: (a) determining the fraction of nucleic acids in the sample nucleic acid that have copy number variation; (b) determining the fraction of low-abundance nucleic acids (e.g., fetal fraction) in the sample nucleic acid; and (c) comparing the fraction in (a) with the fraction in (b) to generate a ratio of (a:):(b). It has further been discovered that the MR ratio can be used to identify patients with a higher chance of a discordant positive result due to mosaicism (e.g., CPM). For example, the MR can be used to classify the presence or absence of genetic mosaicism for a region of copy number variation. In certain embodiments, an MR value between about 0.2 and about 0.7 classifies the presence of genetic mosaicism for a region of copy number variation. In certain embodiments, an MR value greater than 0.7 classifies the absence of genetic mosaicism for a region of copy number variation. The use of the mosaicism ratio in such situations has numerous advantages over conventional processes for confirming positive NIPT results, including a non-invasive approach for confirming a positive NIPT result.
[0043] Furthermore, knowledge of the presence or absence of mosaicism can then be used to better interpret positive NIPT results by physicians and genetic counselors, which can lead to improved post-test counseling and overall prenatal care. For example, the presence of a genetic mosaicism classification for a copy number variation region (e.g., an MR of 20% to 70%) can be interpreted as a non-standard positive NIPT result with a mosaic comment. The absence of a genetic mosaicism classification for a copy number variation region (e.g., an MR greater than 70%) can be interpreted as a standard positive NIPT result (e.g., a positive result for fetal copy number variation), affected fetus, fetal copy number variation, total copy number variation, true copy number variation, complete copy number variation, etc. For copy number variation regions, if the MR value is below a certain threshold (e.g., an MR less than 20%), no classification (e.g., no call, no clinical relevance) can be provided, which can be interpreted as a negative NIPT result for fetal copy number variation.
[0044] Genetic mosaicism classification Provided herein is a method for classifying the presence or absence of genetic mosaicism (e.g., CPM) of sample (e.g., biological sample, test sample).In various embodiments, the presence or absence of genetic mosaicism for copy number variation is classified.Copy number variation, sometimes referred to as copy number alteration, can include aneuploidy (e.g., chromosome trisomy, chromosome monosomy), deletion (e.g., microdeletion, partial chromosome deletion) and duplication (e.g., microduplication, partial chromosome duplication), and are described in more detail herein.
[0045] The presence or absence of genetic mosaicism can be classified for copy number variation regions (for example, trisomic cell lineages that are confined to placenta).Copy number variation regions refer to the genomic region (for example, chromosome, part of chromosome) where copy number variation is identified.Copy number variation regions can refer to specific chromosomes, or chromosomal locations (for example, regions that span a specific genomic coordinate).Copy number variation regions can be identified using any suitable method for identifying copy number variation in the art or as described herein.
[0046] In some embodiments, the methods herein include determining the fraction of nucleic acids having copy number variations in a sample nucleic acid. Determining the fraction of nucleic acids refers to quantifying a specific species of nucleic acid in a nucleic acid mixture. For example, determining the fraction of nucleic acids can refer to quantifying a low-abundance nucleic acid species, quantifying fetal nucleic acids, quantifying cancer nucleic acids, etc. Determining the fraction of nucleic acids having copy number variations refers to quantifying a subset of nucleic acids (e.g., a subset of nucleic acid fragments, a subset of sequence reads) in which copy number variations are identified. In some embodiments, determining the fraction of nucleic acids having copy number variations refers to quantifying a subset of nucleic acids (e.g., a subset of nucleic acid fragments, a subset of sequence reads) from a region (e.g., a genomic region) in which copy number variations are identified. In some embodiments, determining the fraction of nucleic acids having copy number variations refers to quantifying a subset of certain nucleic acids (e.g., a subset of certain nucleic acid fragments, a subset of certain sequence reads) from a region (e.g., a genomic region) in which copy number variations are identified. For example, for a sample containing maternal and fetal nucleic acids, if the fetal nucleic acid is identified as having trisomy of chromosome 21, determining the fraction of nucleic acids having copy number variations refers to determining the fetal fraction based on information derived from or associated with chromosome 21 or a portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequences, methylation variable sequences).
[0047] In some embodiments, the methods herein include determining a fraction for a region (e.g., a genomic region). In some embodiments, the methods herein include determining a fraction for a copy number variation region. The fraction for a copy number variation region may also be referred to as an affected fraction or a fraction for an affected region. As discussed above, the fraction for a copy number variation region can be determined according to information (e.g., sequence information, epigenetic information) obtained for a region (e.g., a genomic region) identified as having copy number variation. The fraction for a copy number variation region can be determined using any suitable method for quantifying certain nucleic acids in a nucleic acid mixture. For example, the fraction for a copy number variation region can be determined according to sequencing-based fraction estimation. Methods for determining nucleic acid fractions according to sequencing-based fraction estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is incorporated herein by reference. Sequencing-based fraction estimation is sometimes referred to as bin-based fraction estimation and / or portion-specific fraction estimation. In some embodiments, the fraction for a copy number variation region can be determined according to the allelic ratio of a polymorphic sequence. The polymorphic sequence can include, for example, a single nucleotide polymorphism (SNP). Methods for determining nucleic acid fractions according to the allelic ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is incorporated herein by reference. In some embodiments, the fraction for a copy number variation region can be determined according to various epigenetic biomarkers (e.g., quantification of methylation-variable nucleic acids). Methods for determining nucleic acid fractions according to quantification of methylation-variable nucleic acids are described herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference.
[0048] In some embodiments, the sample nucleic acid comprises a major amount of nucleic acid and a minor amount of nucleic acid. In some embodiments, the major amount of nucleic acid comprises maternal nucleic acid, and the minor amount of nucleic acid comprises fetal nucleic acid. Thus, in some embodiments, the methods herein comprise determining a fetal fraction. In some embodiments, the methods herein comprise determining a fetal fraction for a region (e.g., a genomic region). In some embodiments, the methods herein comprise determining a fetal fraction for a region of copy number variation. The fetal fraction for a region of copy number variation may also be referred to as an affected fraction, an affected fetal fraction, and / or a fetal fraction for an affected region. As discussed above, the fetal fraction for a region of copy number variation can be determined according to information (e.g., sequence information, epigenetic information) obtained for a region (e.g., a genomic region) identified as having a fetal copy number variation. The fetal fraction for a region of copy number variation can be determined using any suitable method for quantifying fetal nucleic acid in a mixture of maternal and fetal nucleic acids. For example, the fetal fraction for a copy number variation region can be determined according to sequencing-based fetal fraction (SeqFF) estimation. Methods for determining the fetal fraction according to sequencing-based fetal fraction (SeqFF) estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is incorporated herein by reference. Sequencing-based fetal fraction (SeqFF) estimation is also sometimes referred to as bin-based fetal fraction (BFF) estimation and / or site-specific fetal fraction estimation. In some embodiments, the fetal fraction for a copy number variation region can be determined according to the allele ratio of a polymorphic sequence in fetal nucleic acid and maternal nucleic acid. The polymorphic sequence can include, for example, a single nucleotide polymorphism (SNP). Methods for determining fetal fraction according to the allelic ratio of polymorphic sequences are described herein and in US Patent Application Publication No. 2011 / 0224087, which is incorporated herein by reference.In some embodiments, the fetal fraction for the copy number variation region can be determined according to various epigenetic biomarkers (e.g., quantification of methylation-variable fetal nucleic acids and maternal nucleic acids).Methods for determining the fetal fraction according to quantification of methylation-variable fetal nucleic acids and maternal nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference.
[0049] In some embodiments, the methods herein include determining the fraction of low-abundance nucleic acids in a sample nucleic acid. Determining the fraction of low-abundance nucleic acids in a sample nucleic acid is generally not limited to methods, such as those described above, that quantify nucleic acid species based on information about regions identified as having copy number variations. Rather, determining the fraction of low-abundance nucleic acids in a sample nucleic acid may include methods that quantify low-abundance nucleic acids according to information derived from regions across the genome and / or regions different from the regions identified as having copy number variations. In some embodiments, the fraction of low-abundance nucleic acids is determined for genomic regions that are larger than the regions of copy number variation. For example, the fraction of low-abundance nucleic acids can be determined for genomic regions that contain more genomic content (e.g., base pairs, kilobases, megabases) than the regions identified as having copy number variations. For example, for a sample in which low-abundance nucleic acids are identified as having trisomy 21, the fraction of low-abundance nucleic acids can be determined according to information derived from or associated with multiple chromosomes (e.g., sequence information, sequence read quantification, polymorphism sequences, methylation variable sequences). In this example, such a plurality of chromosomes may include all chromosomes, autosomes, a subset of chromosomes, a subset of autosomes, a subset of chromosomes that includes chromosome 21, a subset of autosomes that includes chromosome 21, a subset of chromosomes that does not include chromosome 21, a subset of autosomes that does not include chromosome 21, or portions thereof. In some embodiments, the fraction of low abundance nucleic acids is determined for genomic regions that are distinct from regions of copy number variation. For example, for a sample in which low abundance nucleic acids are identified as having trisomy of chromosome 21, the fraction of low abundance nucleic acids can be determined according to information (e.g., sequence information, sequence read quantification, polymorphic sequences, methylation variable sequences) derived from or associated with chromosomes other than chromosome 21.
[0050] The fraction of low-abundance nucleic acids in a sample nucleic acid can be determined using any suitable method for quantifying certain nucleic acids in a nucleic acid mixture. For example, the fraction of low-abundance nucleic acids can be determined according to sequencing-based fraction estimation. Methods for determining low-abundance nucleic acid fractions according to sequencing-based fraction estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is incorporated herein by reference. Sequencing-based fraction estimation is also referred to as bin-based fraction estimation and / or site-specific fraction estimation. In some embodiments, the fraction of low-abundance nucleic acids can be determined according to the allele ratio of a polymorphic sequence. The polymorphic sequence can include, for example, a single nucleotide polymorphism (SNP). Methods for determining low-abundance nucleic acid fractions according to the allele ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, incorporated herein by reference. In some embodiments, the fraction of low abundance nucleic acids can be determined according to various epigenetic biomarkers (e.g., quantification of methylation variable nucleic acids). Methods for determining the fraction of low abundance nucleic acids according to quantification of methylation variable nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference.
[0051] In some embodiments, the small amount of nucleic acid comprises fetal nucleic acid. Thus, in some embodiments, the method herein comprises determining the fetal fraction. The fetal fraction can be determined using any suitable method for quantifying fetal nucleic acid in a mixture of maternal and fetal nucleic acids. For example, the fetal fraction can be determined according to sequencing-based fetal fraction (SeqFF) estimation. Methods for determining the fetal fraction according to sequencing-based fetal fraction (SeqFF) estimation are described herein and in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is incorporated herein by reference. Sequencing-based fetal fraction (SeqFF) estimation is also sometimes referred to as bin-based fetal fraction (BFF) estimation and / or site-specific fetal fraction estimation. In some embodiments, the fetal fraction can be determined according to the allele ratio of a polymorphic sequence in fetal nucleic acid and maternal nucleic acid. The polymorphic sequence may include, for example, a single nucleotide polymorphism (SNP). Methods for determining the fetal fraction according to the allele ratio of a polymorphic sequence are described herein and in U.S. Patent Application Publication No. 2011 / 0224087, which is incorporated herein by reference. In some embodiments, the fetal fraction can be determined according to various epigenetic biomarkers (e.g., quantification of methylation-variable fetal nucleic acids and maternal nucleic acids). Methods for determining the fetal fraction according to quantification of methylation-variable fetal nucleic acids and maternal nucleic acids are described, for example, herein and in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference. In some embodiments, the fetal fraction can be determined according to chromosome Y assay. Methods for determining the fetal fraction according to chromosome Y assay are described herein and in Lo YM et al. (1998) Am J Hum Genet 62:768-775.
[0052] In some embodiments, the fraction of copy number variation region and the fraction of low amount of nucleic acid are determined using the same methodology.For example, the fraction of copy number variation region and the fraction of low amount of nucleic acid can be determined according to the fraction estimation based on sequencing.In some embodiments, the fraction of copy number variation region and the fraction of low amount of nucleic acid are determined using different methodologies.For example, the fraction of copy number variation region can be determined according to the allele ratio of polymorphic sequence, and the fraction of low amount of nucleic acid can be determined according to different epigenetic biomarkers.
[0053] In some embodiments, the fetal fraction of copy number variation region and the fetal fraction of nucleic acid sample are determined using the same methodology.For example, the fetal fraction of copy number variation region and the fetal fraction of nucleic acid sample can be determined according to the fetal fraction estimation based on sequencing.In some embodiments, the fetal fraction of copy number variation region and the fetal fraction of nucleic acid sample are determined using different methodologies.For example, the fetal fraction of copy number variation region can be determined according to the allele ratio of polymorphic sequence, and the fetal fraction of nucleic acid sample can be determined according to chromosome Y assay.
[0054] In some embodiments, the fraction of copy number variation (e.g., copy number variation regions) is determined for a chromosome or portion thereof. The fraction of copy number variation determined for a chromosome or portion thereof refers to the quantification of nucleic acid species based on information derived from or associated with the chromosome or portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequences, methylation variable sequences). In some embodiments, the fraction of copy number variation (e.g., copy number variation regions) is determined for chromosome 13, chromosome 18, or chromosome 21. In some embodiments, the fraction of low-abundance nucleic acids is determined for a chromosome or portion thereof different from the chromosome or portion thereof used to determine the fraction of copy number variation. In some embodiments, the fraction of low-abundance nucleic acids is determined for multiple chromosomes or multiple portions of chromosomes. In some embodiments, the fraction of low-abundance nucleic acids is determined for multiple autosomes or multiple portions of autosomes. In some embodiments, the fraction of low-abundance nucleic acids is determined for multiple regions (e.g., genomic regions). In some embodiments, the fraction of low-abundance nucleic acids is determined for multiple regions (e.g., genomic regions) genome-wide.
[0055] In some embodiments, the fetal fraction for copy number variation (e.g., copy number variation region) is determined for a chromosome or portion thereof. The fetal fraction for copy number variation determined for a chromosome or portion thereof refers to quantification of fetal nucleic acid based on information derived from or associated with the chromosome or portion thereof (e.g., sequence information, sequence read quantification, polymorphic sequence, methylation variable sequence). In some embodiments, the fetal fraction for copy number variation (e.g., copy number variation region) is determined for chromosome 13, chromosome 18, or chromosome 21. In some embodiments, the fetal fraction of the sample nucleic acid is determined for a chromosome or portion thereof different from the chromosome or portion thereof used to determine the fetal fraction for copy number variation. In some embodiments, the fetal fraction of the sample nucleic acid is determined for multiple chromosomes or multiple portions of chromosomes. In some embodiments, the fetal fraction of the sample nucleic acid is determined for multiple autosomes or multiple portions of autosomes. In some embodiments, the fetal fraction of the sample nucleic acid is determined for multiple regions (e.g., genomic regions). In some embodiments, the fetal fraction of the sample nucleic acid is determined for multiple regions genome-wide (eg, genomic regions).
[0056] In some embodiments, the methods herein include comparing the fraction of copy number variation to the fraction of low abundance nucleic acid. In some embodiments, comparing the fraction of copy number variation to the fraction of low abundance nucleic acid includes generating a ratio. For example, the ratio can be the fraction of copy number variation divided by the fraction of low abundance nucleic acid.
[0057] In some embodiments, the methods herein include comparing the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid. In some embodiments, comparing the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid includes generating a ratio. For example, the ratio can be the fetal fraction of copy number variation divided by the fetal fraction of the sample nucleic acid.
[0058] In some embodiments, the method herein comprises classifying the presence or absence of genetic mosaicism for copy number variation regions. The presence or absence of genetic mosaicism for copy number variation regions can be classified according to comparison. For example, the presence or absence of genetic mosaicism for copy number variation regions can be classified according to the comparison of the fraction of copy number variation and the fraction of low-abundance nucleic acid. In some embodiments, the presence or absence of genetic mosaicism for copy number variation regions can be classified according to the comparison of the fetal fraction of copy number variation and the fetal fraction of sample nucleic acid. The presence or absence of genetic mosaicism for copy number variation regions can be classified according to the ratio. For example, the presence or absence of genetic mosaicism for copy number variation regions can be classified according to the ratio of the fraction of copy number variation to the fraction of low-abundance nucleic acid (for example, the fraction of copy number variation divided by the fraction of low-abundance nucleic acid). In some embodiments, the presence or absence of genetic mosaicism for a region of copy number variation can be classified according to the ratio of the fetal fraction for the copy number variation to the fetal fraction for the sample nucleic acid (e.g., the fetal fraction for the copy number variation divided by the fetal fraction for the sample nucleic acid).
[0059] In some embodiments, the presence of genetic mosaicism for the region of copy number variation is classified, which can be interpreted as mosaic copy number variation, affected fetus, unaffected fetus, partially affected fetus, fetal copy number variation, partial fetal copy number variation, partial copy number variation, placental copy number variation, partial placental copy number variation, incomplete copy number variation, placental mosaicism, limited placental mosaicism (CPM), etc.
[0060] In some embodiments, the presence of genetic mosaicism is classified for a region of copy number variation when the ratio of the fraction of copy number variation to the fraction of the minor nucleic acid is less than 1. For example, the presence of genetic mosaicism can be classified for a region of copy number variation when the ratio of the fraction of copy number variation to the fraction of the minor nucleic acid is between about 0.1 and about 0.9, or about 0.1 and about 0.8, or about 0.1 and about 0.7, or about 0.1 and about 0.6, or about 0.2 and about 0.9, or about 0.2 and about 0.8, or about 0.2 and about 0.7, or about 0.2 and about 0.6. In certain embodiments, the presence of genetic mosaicism is classified for a region of copy number variation when the ratio of the fraction of copy number variation to the fraction of the minor nucleic acid is between about 0.2 and about 0.7. For example, the presence of genetic mosaicism can be classified for a region of copy number variation when the ratio of the fraction of copy number variation to the fraction of low-abundance nucleic acid is about 0.2, 0.3, 0.4, 0.5, 0.6, or 0.7. As used herein, the terms "substantially," "approximately," and "about" (unless otherwise defined herein) are defined as largely, but not necessarily entirely, of what is specified, as understood by those skilled in the art (including entirely of what is specified). In any disclosed embodiment, the terms "substantially," "approximately," or "about" may be substituted for "within a percentage of" what is specified, where the percentage includes 0.1, 1, 5, and 10 percent.
[0061] In some embodiments, the presence of genetic mosaicism is classified for a region of copy number variation when the ratio of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is within a range of values less than 1. For example, the presence of genetic mosaicism can be classified for a region of copy number variation when the ratio of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is between about 0.1 and about 0.9, or between about 0.1 and about 0.8, or between about 0.1 and about 0.7, or between about 0.1 and about 0.6, or between about 0.2 and about 0.9, or between about 0.2 and about 0.8, or between about 0.2 and about 0.7, or between about 0.2 and about 0.6. In some embodiments, the presence of genetic mosaicism is classified for a region of copy number variation when the ratio of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is between about 0.2 and about 0.7. For example, the presence of genetic mosaicism for a region of copy number variation can be classified if the ratio of the fetal fraction of copy number variation to the fetal fraction of sample nucleic acid is about 0.2, 0.3, 0.4, 0.5, 0.6, or 0.7.
[0062] In some embodiments, the absence of genetic mosaicism for the region of copy number variation is classified. Classification of the absence of genetic mosaicism for the region of copy number variation can be interpreted as a standard positive result (e.g., a positive result for fetal copy number variation), affected fetus, fetal copy number variation, total copy number variation, true copy number variation, full copy number variation, etc.
[0063] In some embodiments, the absence of genetic mosaicism is classified for a region of copy number variation when the ratio of the fraction of copy number variation to the fraction of the minor nucleic acid is greater than 0.6. For example, the absence of genetic mosaicism can be classified for a region of copy number variation when the ratio of the fraction of copy number variation to the fraction of the minor nucleic acid is between about 0.7 and about 1.5, or about 0.7 and about 1.3, or about 0.7 and about 1.1, or about 0.8 and about 1.1, or about 0.8 and about 1.0, or about 0.8 and about 0.9. In some embodiments, the absence of genetic mosaicism is classified for a region of copy number variation when the ratio of the fraction of copy number variation to the fraction of the minor nucleic acid is between about 0.71 and about 1.3. For example, the absence of genetic mosaicism can be classified for a region of copy number variation when the ratio of the fraction of copy number variation to the fraction of low abundance nucleic acid is about 0.71, 0.8, 0.9, 1.0, 1.1, 1.2, or 1.3. In other embodiments, the absence of genetic mosaicism can be classified for a region of copy number variation when the ratio of the fraction of copy number variation to the fraction of low abundance nucleic acid is greater than 0.7.
[0064] In some embodiments, the absence of genetic mosaicism is classified for a region of copy number variation when the ratio of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is greater than 0.6. For example, the absence of genetic mosaicism can be classified for a region of copy number variation when the ratio of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is between about 0.7 and about 1.5, or about 0.7 and about 1.3, or about 0.7 and about 1.1, or about 0.8 and about 1.1, or about 0.8 and about 1.0, or about 0.8 and about 0.9. In some embodiments, the absence of genetic mosaicism is classified for a region of copy number variation when the ratio of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is between about 0.71 and about 1.3. For example, the absence of genetic mosaicism for a region of copy number variation can be classified if the ratio of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is about 0.71, 0.8, 0.9, 1.0, 1.1, 1.2, or 1.3. In other embodiments, the absence of genetic mosaicism for a region of copy number variation is classified if the ratio of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is greater than 0.7.
[0065] In some embodiments, no classification is provided. For example, if the ratio value of the fraction of copy number variation for the low abundance nucleic acid fraction is below a certain threshold, no classification (e.g., no call, no clinical relevance) can be provided. In some embodiments, if the ratio value of the fraction of copy number variation for the low abundance nucleic acid fraction is about 0.3 or less, no classification is provided. In some embodiments, if the ratio value of the fraction of copy number variation for the low abundance nucleic acid fraction is about 0.2 or less, no classification is provided. In some embodiments, if the ratio value of the fraction of copy number variation for the low abundance nucleic acid fraction is about 0.1 or less, no classification is provided.
[0066] In some embodiments, if the ratio value of the fraction of copy number variation for the fraction of low abundance nucleic acid is above a certain threshold, no classification is provided. For example, if the ratio value of the fraction of copy number variation for the fraction of low abundance nucleic acid is about 0.9, 1.0, 1.1, 1.2, or 1.3 or more, no classification may be provided. In some embodiments, if the ratio value of the fraction of copy number variation for the fraction of low abundance nucleic acid is about 1.3 or more, no classification is provided. A value above a certain threshold (e.g., above 1.3) may indicate copy number variation (e.g., maternal copy number variation) present in the abundant nucleic acid.
[0067] In some embodiments, no classification (e.g., no call, no clinical relevance) can be provided when the ratio value of the fetal fraction of copy number variations to the fetal fraction of the sample nucleic acid is below a certain threshold. In some embodiments, no classification is provided when the ratio value of the fetal fraction of copy number variations to the fetal fraction of the sample nucleic acid is about 0.3 or less. In some embodiments, no classification is provided when the ratio value of the fetal fraction of copy number variations to the fetal fraction of the sample nucleic acid is about 0.2 or less. In some embodiments, no classification is provided when the ratio value of the fetal fraction of copy number variations to the fetal fraction of the sample nucleic acid is about 0.1 or less.
[0068] In some embodiments, if the ratio value of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is above a certain threshold, no classification is provided. For example, if the ratio value of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is about 0.9, 1.0, 1.1, 1.2, or 1.3 or more, no classification can be provided. In some embodiments, if the ratio value of the fetal fraction of copy number variation to the fetal fraction of the sample nucleic acid is about 1.3 or more, no classification is provided.
[0069] FIG. 2 illustrates a process 200 for classifying the presence or absence of genetic mosaicism in a biological sample, according to various embodiments. A set of sequence reads is provided 205. The sequence reads can be obtained from circulating cell-free sample nucleic acid obtained from a test sample derived from a test subject (e.g., a pregnant female subject). The circulating cell-free nucleic acid can include maternal nucleic acid and fetal nucleic acid. The circulating cell-free sample nucleic acid can be captured by a probe oligonucleotide under hybridization conditions. Regions of gene copy number variation in circulating cellular nucleic acid are identified from the set of sequence reads 210. The fraction of circulating cell-free nucleic acid in the sample nucleic acid having copy number variation is determined 215. The fraction can be a fetal fraction determined for the region of copy number variation. The fraction of fetal nucleic acid in the circulating cell-free sample nucleic acid is determined 220. The fraction of circulating cell-free nucleic acid having copy number variations is compared to the fraction of fetal nucleic acid 225 to provide a comparison and generate a mosaicism ratio of the fraction of circulating cell-free nucleic acid having copy number variations to the fraction of fetal nucleic acid. According to the comparison and the mosaicism ratio, the presence or absence of genetic mosaicism for the region of copy number variation is classified 230.
[0070] 3 illustrates a process 300 for classifying the presence or absence of genetic mosaicism in a biological sample and providing clinical interpretation and / or diagnostic follow-up information consistent with various embodiments. A set of sequence reads is provided, and a screening test for a genetic condition (e.g., NIPT) is obtained from the set of sequence reads 305. The sequence reads can be obtained from circulating cell-free sample nucleic acid obtained from a test sample obtained from a test subject (e.g., a pregnant female subject). The circulating cell-free sample nucleic acid can include maternal nucleic acid and fetal nucleic acid. The circulating cell-free sample nucleic acid can be captured by a probe oligonucleotide under hybridization conditions. In various embodiments, the genetic condition being screened for includes the presence of one or more aneuploidies, such as copy number variations. The presence (flag as positive) or absence (flag as negative) of one or more aneuploidies can be identified in the circulating cell-free nucleic acids from the set of sequence reads based on the z-score 310 or 315. If the absence (flag as negative) of one or more aneuploidies is identified, no further testing may be performed 320, or diagnostic testing may be performed 325. If the presence (flag as positive) of one or more aneuploidies is identified, the mosaicism ratio is described with reference to FIG. 2, and the mosaicism ratio value is used to classify the presence or absence of genetic mosaicism and provide enhanced interpretation of the NIPT results. The mosaicism ratio can be used to identify patients with a higher chance of a discordant positive result due to mosaicism (e.g., CPM).
[0071] A mosaicism ratio value between about 0.2 and about 0.7 may classify the presence of genetic mosaicism for the copy number variation region 330. A mosaicism ratio value greater than 0.7 may classify the absence of genetic mosaicism for the copy number variation region 335. Furthermore, no classification may be provided for the copy number variation region 340 / 345 when the mosaicism ratio value is greater than about 1.3 or less than about 0.2. If no classification is provided and the mosaicism ratio value is greater than about 1.3, the positive NIPT result may be interpreted as possibly excessive or indeterminate 350, and diagnostic follow-up 355, including amniocentesis, CVS, maternal testing, and / or other testing, may be recommended depending on a consensus decision between the genetic counselor and the physician. If no classification is provided and the mosaicism ratio value is less than about 0.2, the positive NIPT result may be interpreted as a negative result or the absence of one or more aneuploidies 360, and diagnostic follow-up 365 may not be required. If the presence of genetic mosaicism is classified (e.g., the mosaicism ratio is between about 0.2 and about 0.7), the positive NIPT result may be interpreted as positive with a mosaic comment (e.g., an understanding that mosaic presentation is possible) 370, and diagnostic follow-up including amniocentesis and / or CVS may be recommended depending on the consensus decision between the genetic counselor and the physician 375. If the absence of genetic mosaicism is classified (e.g., the mosaicism ratio is greater than about 0.7 but less than about 1.3), the positive NIPT result may be interpreted as positive 380, and diagnostic follow-up including amniocentesis and / or CVS may be recommended for confirmation 385.
[0072] sample This specification provides systems, methods, and products for analyzing nucleic acids. In some embodiments, nucleic acid fragments in a mixture of nucleic acid fragments are analyzed. Nucleic acid fragments may also be referred to as nucleic acid templates, and these terms may be used interchangeably herein. A mixture of nucleic acids may contain two or more nucleic acid fragment species with the same or different nucleotide sequences, different fragment lengths, different origins (e.g., genomic origin, fetal origin vs. maternal origin, cellular or tissue origin, cancer vs. non-cancer origin, tumor vs. non-tumor origin, sample origin, subject origin, etc.), or combinations thereof.
[0073] Often, nucleic acids or nucleic acid mixtures utilized in the systems, methods, and products described herein are isolated from samples obtained from a subject (e.g., a test subject). The subject can be any living or non-living organism, including, but not limited to, humans, non-human animals, plants, bacteria, fungi, protozoa, or pathogens. Any human or non-human animal can be selected, including, for example, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cows), equines (e.g., horses), caprines and ovines (e.g., sheep, goats), swine (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursidae (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The subject can be male or female (e.g., female, pregnant). The subject can be of any age (e.g., embryo, fetus, infant, child, adult). The subject can be a cancer patient, a patient suspected of having cancer, a patient in remission, a patient with a family history of cancer, and / or a subject undergoing cancer screening. In some embodiments, the test subject is female. In some embodiments, the test subject is a female human. In some embodiments, the test subject is male. In some embodiments, the test subject is a male human.
[0074] Nucleic acids can be isolated from any type of suitable biological specimen or sample (e.g., test sample). Sample or test sample can be any specimen isolated or obtained from a subject or its part (e.g., human subject, pregnant female, cancer patient, fetus, tumor). Samples are sometimes derived from pregnant female subjects with fetuses at any stage of pregnancy (e.g., first, second, or third trimester of human subjects), and sometimes from postnatal subjects. Samples are sometimes derived from pregnant subjects with fetuses that are euploid for all chromosomes, and sometimes from pregnant subjects with fetuses that have chromosomal aneuploidy (e.g., 1, 3 (i.e., trisomy (e.g., T21, T18, T13)) or 4 copies of chromosomes) or other genetic variations. Non-limiting examples of specimens include bodily fluids or tissues obtained from a subject, including, but not limited to, blood or blood products (e.g., serum, plasma, etc.), umbilical cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluid (e.g., from the bronchoalveolar, stomach, peritoneal cavity, ducts, ear, arthroscopy), biopsy samples (e.g., samples obtained from preimplantation embryo biopsies, cancer biopsies), peritoneal aspirate samples, cells (blood cells, placental cells, embryonic or fetal cells, fetal nucleated cells or fetal cell remnants, normal cells, abnormal cells (e.g., cancer cells)) or parts thereof (e.g., mitochondria, nuclei, extracts, etc.), female reproductive tract washings, urine, feces, sputum, saliva, nasal mucus, prostatic fluid, lavage fluid, semen, lymph, bile, tears, sweat, milk, mammary fluid, etc., or combinations thereof. In some embodiments, the biological sample is a cervical swab obtained from a subject. The bodily fluid or tissue sample from which nucleic acids are extracted may be free of cells (e.g., acellular). In some embodiments, the bodily fluid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, fetal cells or cancerous cells may be included in the sample.
[0075] The sample may be a liquid sample. The liquid sample may contain extracellular nucleic acids (e.g., circulating cell-free DNA). Non-limiting examples of liquid samples include blood or blood products (e.g., serum, plasma, etc.), urine, biopsy samples (e.g., liquid biopsy for cancer detection), the above liquid samples, etc., or combinations thereof. In certain embodiments, the sample is a liquid biopsy, which generally refers to the evaluation of a liquid sample from a subject for the presence, absence, progression, or remission of disease (e.g., cancer). A liquid biopsy can be used in conjunction with or as an alternative to a solid biopsy (e.g., tumor biopsy). In certain cases, extracellular nucleic acids are analyzed in a liquid biopsy.
[0076] In some embodiments, the biological sample may be blood, plasma, or serum. The term "blood" encompasses whole blood, blood products, or any fraction of blood, including conventionally defined serum, plasma, buffy coat, etc. Blood or fractions thereof often contain nucleosomes. Nucleosomes contain nucleic acids and are sometimes acellular or intracellular nucleosomes. Blood also includes buffy coats, which are sometimes isolated using a Ficoll gradient. Buffy coats can contain white blood cells (e.g., leukocytes, T cells, B cells, platelets, etc.). Plasma refers to the fraction of whole blood obtained by centrifugation of anticoagulant-treated blood. Serum refers to the aqueous liquid portion remaining after a blood sample has clotted. Body fluid or tissue samples are often collected according to standard protocols commonly followed in hospitals or outpatient clinics. In the case of blood, an appropriate volume of peripheral blood (e.g., 3-40 milliliters, 5-50 milliliters) is often collected and can be stored according to standard procedures before or after preparation.
[0077] Analysis of nucleic acids found in a subject's blood can be performed using, for example, whole blood, serum, or plasma. For example, analysis of fetal DNA found in maternal blood can be performed using, for example, whole blood, serum, or plasma. For example, analysis of tumor DNA found in a patient's blood can be performed using, for example, whole blood, serum, or plasma. Methods for preparing serum or plasma from blood derived from a subject (e.g., a maternal subject, a cancer patient) are known. For example, the subject's blood (e.g., a pregnant woman's blood; a cancer patient's blood) can be placed in a tube containing EDTA or a specialized commercially available product, such as Vacutainer SST (Becton Dickinson, Franklin Lakes, NJ), to prevent blood clotting, and plasma can then be obtained from the whole blood by centrifugation. Serum can be obtained with or without centrifugation after blood clotting. When centrifugation is used, it is typically, but not necessarily, performed at an appropriate speed, e.g., 1,500 to 3,000 times g. The plasma or serum may be subjected to an additional centrifugation step before being transferred to a new tube for nucleic acid extraction. In addition to the cell-free portion of whole blood, nucleic acids can also be recovered from the cellular fraction and concentrated in the buffy coat portion, which can be obtained by centrifuging a whole blood sample obtained from a subject and removing the plasma.
[0078] A sample can be heterogeneous. For example, a sample can contain more than one cell type and / or one or more nucleic acid species. In some cases, a sample can contain (i) fetal cells and maternal cells, (ii) cancerous cells and non-cancerous cells, and / or (iii) pathogenic cells and host cells. In some cases, a sample can contain (i) cancerous nucleic acids and non-cancerous nucleic acids, (ii) pathogenic nucleic acids and host nucleic acids, (iii) fetal-derived and maternal-derived nucleic acids, and / or more generally, (iv) mutated nucleic acids and wild-type nucleic acids. In some cases, a sample can contain minor and major nucleic acid species, as described in more detail below. In some cases, a sample can contain cells and / or nucleic acids from a single subject, or can contain cells and / or nucleic acids from multiple subjects. cell type
[0079] As used herein, "cell type" refers to a type of cell that can be distinguished from other types of cells. Extracellular nucleic acids can include nucleic acids from several different cell types. Non-limiting examples of cell types that can contribute nucleic acids to circulating cell-free nucleic acids include liver cells (e.g., hepatocytes), lung cells, spleen cells, pancreatic cells, colon cells, skin cells, bladder cells, eye cells, brain cells, esophageal cells, head cells, cervical cells, ovarian cells, testicular cells, prostate cells, placental cells, epithelial cells, endothelial cells, adipocytes, kidney / renal cells, cardiac cells, muscle cells, blood cells (e.g., leukocytes), central nervous system (CNS) cells, etc., and combinations of the above. In some embodiments, cell types that contribute nucleic acids to circulating cell-free nucleic acids that are analyzed include leukocytes, endothelial cells, and hepatocyte / liver cells. As described in more detail below, different cell types can be screened as part of identifying and selecting nucleic acid loci with identical or substantially identical marker status for cell types in subjects with a medical condition and for cell types in subjects without a medical condition.
[0080] A particular cell type sometimes remains the same or substantially the same in a subject with a medical condition and in a subject without a medical condition. In a non-limiting example, in a cytopathic condition, the number of live or viable cells of a particular cell type may be reduced, while in a subject with a medical condition, the live or viable cells are not modified or are not significantly modified.
[0081] A particular cell type is sometimes modified as part of a medical condition and has one or more characteristics that differ from its original state. In a non-limiting example, a particular cell type may proliferate at a faster than normal rate, may transform into a cell with a different morphology, may transform into a cell expressing one or more different cell surface markers, and / or may become part of a tumor as part of a cancerous condition. In embodiments in which a particular cell type (i.e., a precursor cell) is modified as part of a medical condition, the marker state for each of the one or more markers assayed is often the same or substantially the same for a particular cell type in a subject with a medical condition and for a particular cell type in a subject without the medical condition. Thus, the term "cell type" sometimes refers to the type of cell in a subject without a medical condition and to the modified version of the cell in a subject with a medical condition. In some embodiments, "cell type" refers only to precursor cells, not modified versions resulting from the precursor cells. "Cell type" sometimes refers to precursor cells and modified cells resulting from the precursor cells. In such embodiments, the marker states of the markers analyzed are often the same or substantially the same for cell types in subjects with the medical condition and for cell types in subjects without the medical condition.
[0082] In certain embodiments, the cell type is a cancer cell. Certain cancer cell types include, for example, leukemia cells (e.g., acute myeloid leukemia, acute lymphocytic leukemia, chronic myeloid leukemia, chronic lymphocytic leukemia); cancerous kidney / renal cells (e.g., renal cell carcinoma (clear cell, type 1 papillary, type 2 papillary, chromophobe, oncocytic, collecting duct), renal adenocarcinoma, Grawitz tumor, Wilms tumor, transitional cell carcinoma); brain tumor cells (e.g., acoustic neuroma, astrocytoma (grade I: pilocytic astrocytoma, grade II: pilocytic astrocytoma, grade III ... Grade II: low-grade astrocytoma, Grade III: anaplastic astrocytoma, Grade IV: glioblastoma (GBM), chordoma, CNS lymphoma, craniopharyngioma, glioma (brainstem glioma, ependymoma, mixed glioma, acoustic neuroglioma, subependymoma), medulloblastoma, meningioma, metastatic brain tumor, oligodendroglioma, pituitary tumor, primitive neuroectodermal tumor (PNET), schwannoma, juvenile pilocytic astrocytoma (JPA), pineal tumor, rhabdoid tumor).
[0083] Various cell types can be distinguished by any suitable characteristics, including, but not limited to, one or more different cell surface markers, one or more different morphological features, one or more different functions, one or more different protein (e.g., histone) modifications, and one or more different nucleic acid markers.Non-limiting examples of nucleic acid markers include single nucleotide polymorphisms (SNPs), methylation status of nucleic acid loci, short tandem repeats, insertions (e.g., microinsertions), deletions (microdeletions), etc., and combinations thereof.Non-limiting examples of protein (e.g., histone) modifications include acetylation, methylation, ubiquitination, phosphorylation, sumoylation, etc., and combinations thereof.
[0084] As used herein, the term "related cell type" refers to a cell type that has multiple characteristics in common with another cell type. In related cell types, sometimes 75% or more of the cell surface markers are common to the cell types (e.g., about 80%, 85%, 90%, or 95% or more of the cell surface markers are common to the related cell types).
[0085] nucleic acidA method for analyzing nucleic acids is provided herein. The terms "nucleic acid," "nucleic acid molecule," "nucleic acid fragment," and "nucleic acid template" can be used interchangeably throughout this disclosure. These terms refer to nucleic acids of any composition, including DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA)), RNA (e.g., messenger RNA (mRNA), small interfering RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed by fetuses or placentas), and / or DNA or RNA analogs (e.g., those containing base analogs, sugar analogs, and / or exogenously added backbones), RNA / DNA hybrids, and polyamide nucleic acids (PNAs), all of which can be in single-stranded or double-stranded form, and can include known analogs of natural nucleotides that can function in a manner similar to naturally occurring nucleotides, unless otherwise specified. In certain embodiments, the nucleic acid may be or may be derived from a plasmid, phage, virus, bacterium, autonomously replicating sequence (ARS), mitochondrion, centromere, artificial chromosome, chromosome, or other nucleic acid capable of replicating or being replicated in vitro or in a host cell, cell, cell nucleus, or cell cytoplasm. In some embodiments, the template nucleic acid may be derived from a single chromosome (e.g., a nucleic acid sample may be derived from one chromosome of a sample obtained from a diploid organism). Unless otherwise specified, the term encompasses nucleic acids that have similar binding properties to the reference nucleic acid and contain known analogs of natural nucleotides that are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise specified, a particular nucleic acid sequence implicitly encompasses not only the sequence explicitly indicated, but also its conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences. Specifically, degenerate codon substitutions can be obtained by generating sequences in which the third position of one or more selected (or all) codons is substituted with a mixed-base residue and / or a deoxyinosine residue. The term nucleic acid is used interchangeably with locus, gene, cDNA, and mRNA encoded by a gene.The term can also include, as equivalents, RNA or DNA derivatives, variants, and analogs synthesized from nucleotide analogs, single-stranded ("sense" or "antisense" strand, "plus" or "minus" strand, "forward" or "reverse" reading frame), and double-stranded polynucleotides. The term "gene" refers to a segment of DNA involved in producing a polypeptide chain and generally includes regions preceding and following the coding region (leader and trailer), which are involved in the transcription / translation of the gene product and the regulation of transcription / translation, as well as intervening sequences (introns) between individual coding regions (exons). Nucleotide or base generally refers to the purine and pyrimidine molecular units of nucleic acids (e.g., adenine (A), thymine (T), guanine (G), and cytosine (C)). For RNA, the base thymine is substituted with uracil. The length or size of a nucleic acid can be expressed as the number of bases.
[0086] The nucleic acid may be single-stranded or double-stranded. For example, single-stranded DNA can be generated by denaturing double-stranded DNA, for example, by treatment with heat or alkali. In certain embodiments, the nucleic acid adopts a D-loop structure formed by intercalating an oligonucleotide into the strand of a double-stranded DNA molecule, or is a DNA-like molecule, for example, a peptide nucleic acid (PNA). The formation of a D-loop is observed in E. coli. This can be facilitated by adding RecA protein and / or varying salt concentration, for example, using methods known in the art.
[0087] Nucleic acids provided for the processes described herein can contain nucleic acids from one sample or from two or more samples (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more, nineteen or more, or twenty or more samples).
[0088] Nucleic acids can be obtained from one or more sources (e.g., biological samples, blood cells, serum, plasma, buffy coat, urine, lymph, skin, soil, etc.) by methods known in the art. Any suitable method can be used to isolate, extract, and / or purify DNA from a biological sample (e.g., blood or blood products), including, but not limited to, methods for preparing DNA (e.g., as described by Sambrook and Russell, Molecular Cloning: A Laboratory Manual, 3rd Edition, 2001), various commercially available reagents or kits, such as Qiagen's QIAamp Circulating Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiagen, Hilden, Germany), GenomicPrep™ Blood DNA Isolation Kit (Promega, Madison, Wis.), and GFX™ Genomic Blood DNA Purification Kit (Amersham, Piscataway, NJ), or a combination thereof.
[0089] In some embodiments, nucleic acids are extracted from cells using a cell lysis procedure. Cell lysis procedures and reagents are known in the art and can generally be performed by chemical methods (e.g., detergents, hypotonic solutions, enzymatic procedures, etc., or a combination thereof), physical methods (e.g., French press, sonication, etc.), or electrolyte lysis methods. Any suitable lysis procedure can be used. For example, chemical methods generally utilize a lysis agent to disrupt cells and extract nucleic acids from the cells, followed by treatment with chaotropic salts. Physical methods, such as freeze / thaw followed by crushing; use of a cell press, etc., are also useful. In some cases, high salt and / or alkaline lysis procedures can be used.
[0090] In certain embodiments, nucleic acids can include extracellular nucleic acids. As used herein, the term "extracellular nucleic acid" can refer to nucleic acids isolated from a source that is substantially cell-free, and is also referred to as "cell-free" nucleic acid, "circulating cell-free nucleic acid" (e.g., CCF fragments, ccf DNA), and / or "cell-free circulating nucleic acid." Extracellular nucleic acids are present in and can be obtained from blood (e.g., the blood of a human subject). Extracellular nucleic acids often do not contain detectable cells and may contain cellular elements or cellular remnants. Non-limiting examples of cell-free sources from which to obtain extracellular nucleic acids are blood, plasma, serum, and urine. As used herein, the term "obtaining cell-free circulating sample nucleic acids" includes obtaining a sample directly (e.g., collecting a sample, e.g., a test sample) or obtaining a sample from another person who has collected the sample. Without being limited by theory, extracellular nucleic acids can be products of cellular apoptosis and cell degradation, which often result in extracellular nucleic acids with a range of lengths spanning a spectrum (e.g., a "ladder"). In some embodiments, the sample nucleic acid from the test subject is circulating cell-free nucleic acid. In some embodiments, the circulating cell-free nucleic acid is derived from plasma or serum from the test subject.
[0091] In certain embodiments, extracellular nucleic acids can contain different nucleic acid species, and are therefore referred to herein as "heterogeneous." For example, serum or plasma obtained from a person with cancer may contain nucleic acids derived from cancerous cells (e.g., tumors, neoplasms) and nucleic acids derived from non-cancerous cells. In another example, serum or plasma obtained from a pregnant female may contain maternal nucleic acids and fetal nucleic acids. In some cases, cancer or fetal nucleic acids are sometimes about 5% to about 50% of the total nucleic acids (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49% of all nucleic acids are cancer or fetal nucleic acids).
[0092] At least two different nucleic acid species may be present in different amounts in extracellular nucleic acid, sometimes referred to as a minor species and a major species. In certain cases, the minor nucleic acid species originates from an affected cell type (e.g., cancer cells, exhausted cells, cells attacked by the immune system). In certain cases, the minor nucleic acid species originates from apoptotic cells (e.g., circulating cell-free fetal nucleic acid from apoptotic placental cells). In certain embodiments, genetic variations or genetic alterations (e.g., copy number alterations, copy number variations, single nucleotide alterations, single nucleotide variations, chromosomal alterations, and / or rearrangements) are determined for minor nucleic acid species. In certain embodiments, genetic variations or genetic alterations are determined for major nucleic acid species. In general, the terms "minor" or "major" are not intended to be rigidly defined in any respect. In one aspect, a nucleic acid considered "minor" may have an amount of, for example, at least about 0.1% of the total nucleic acid in the sample to less than 50% of the total nucleic acid in the sample. In some embodiments, the low-abundance nucleic acids may comprise an amount of at least about 1% to about 40% of the total nucleic acids in the sample. In some embodiments, the low-abundance nucleic acids may comprise an amount of at least about 2% to about 30% of the total nucleic acids in the sample. In some embodiments, the low-abundance nucleic acids may comprise an amount of at least about 3% to about 25% of the total nucleic acids in the sample. For example, the low-abundance nucleic acids may comprise an amount of about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, or 30% of the total nucleic acids in the sample. In some cases, the minor species of extracellular nucleic acid is sometimes about 1% to about 40% of the total nucleic acid (e.g., about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, or 40% of the nucleic acid is minor species nucleic acid). In some embodiments, the minor nucleic acid is extracellular DNA.In some embodiments, the small amount of nucleic acid is extracellular DNA derived from apoptotic tissue. In some embodiments, the small amount of nucleic acid is extracellular DNA derived from tissue affected by a cell proliferation disorder. In some embodiments, the small amount of nucleic acid is extracellular DNA derived from tumor cells. In some embodiments, the small amount of nucleic acid is extracellular fetal DNA.
[0093] In another aspect, a nucleic acid considered "abundant" can have an amount of, for example, greater than 50% of the total nucleic acids in a sample to about 99.9% of the total nucleic acids in a sample. In some embodiments, an abundant nucleic acid can have an amount of at least about 60% of the total nucleic acids in a sample to about 99% of the total nucleic acids in a sample. In some embodiments, an abundant nucleic acid can have an amount of at least about 70% of the total nucleic acids in a sample to about 98% of the total nucleic acids in a sample. In some embodiments, an abundant nucleic acid can have an amount of at least about 75% of the total nucleic acids in a sample to about 97% of the total nucleic acids in a sample. For example, the abundant nucleic acid can have an amount of at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the total nucleic acid in the sample. In some embodiments, the abundant nucleic acid is extracellular DNA. In some embodiments, the abundant nucleic acid is extracellular maternal DNA. In some embodiments, the abundant nucleic acid is DNA derived from healthy tissue. In some embodiments, the abundant nucleic acid is DNA derived from non-tumor cells.
[0094] In some embodiments, the minor species of extracellular nucleic acid is about 500 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minor species nucleic acid is about 500 base pairs or less in length), In some embodiments, the minor species of extracellular nucleic acid is about 300 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minor species nucleic acid is about 300 base pairs or less in length). In some embodiments, the minor species of extracellular nucleic acid is about 250 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minor species nucleic acids are about 250 base pairs or less in length), In some embodiments, the minor species of extracellular nucleic acid is about 200 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minor species nucleic acids are about 200 base pairs or less in length). In some embodiments, the minor species of extracellular nucleic acid is about 150 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minor species nucleic acids are about 150 base pairs or less in length). In some embodiments, the minor species of extracellular nucleic acid is about 100 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minor species nucleic acids are about 100 base pairs or less in length). In some embodiments, the minor species of extracellular nucleic acid is about 50 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minor species nucleic acid is about 50 base pairs or less in length).
[0095] A sample containing nucleic acid may be treated or untreated to provide the nucleic acid and perform the methods described herein. In some embodiments, a sample containing nucleic acid is treated before providing the nucleic acid and performing the methods described herein. For example, the nucleic acid may be extracted, isolated, purified, partially purified, or amplified from the sample. The term "isolated," as used herein, refers to removing a nucleic acid from its original environment (e.g., the natural environment if naturally occurring, or a host cell if exogenously expressed); thus, the nucleic acid is altered in that it has been removed from its original environment by human intervention (e.g., "by the hand of man"). The term "isolated nucleic acid," as used herein, can refer to a nucleic acid removed from a subject (e.g., a human subject). An isolated nucleic acid may be provided with less non-nucleic acid components (e.g., proteins, lipids) than the amount of those components present in the source sample. A composition containing isolated nucleic acid may be about 50% to more than 99% free of non-nucleic acid components. A composition comprising an isolated nucleic acid may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of non-nucleic acid components. The term "purified," as used herein, can refer to providing a nucleic acid that contains fewer non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than were present before the nucleic acid was subjected to a purification procedure. A composition comprising a purified nucleic acid may be about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other non-nucleic acid components. The term "purified," as used herein, can refer to providing a nucleic acid that contains fewer nucleic acid species than in the sample source from which the nucleic acid was derived. A composition containing purified nucleic acid can be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other nucleic acid species. For example, fetal nucleic acid can be purified from a mixture containing maternal and fetal nucleic acids.In certain examples, small fragments of fetal nucleic acid (e.g., 30-500 bp fragments) can be purified or partially purified from a mixture containing both fetal and maternal nucleic acid fragments. In certain examples, nucleosomes containing smaller fragments of fetal nucleic acid can be purified from a mixture of larger nucleosome complexes containing larger fragments of maternal nucleic acid. In certain examples, cancer cell nucleic acid can be purified from a mixture containing cancer cell and non-cancer cell nucleic acid. In certain examples, nucleosomes containing small fragments of cancer cell nucleic acid can be purified from a mixture of larger nucleosome complexes containing larger fragments of non-cancerous nucleic acid. In some embodiments, nucleic acids are provided for performing the methods described herein without prior processing of the sample(s) containing the nucleic acid. For example, nucleic acids can be analyzed directly from the sample without prior extraction, purification, partial purification, and / or amplification.
[0096] In some embodiments, nucleic acids, such as cellular nucleic acids, are sheared or cleaved before, during, or after the methods described herein. The terms "shearing" or "cleavage" generally refer to a procedure or condition in which a nucleic acid molecule, such as a nucleic acid template gene molecule or its amplification product, can be cleaved into two (or more) smaller nucleic acid molecules. Such shearing or cleavage can be sequence-specific, base-specific, or non-specific, and can be achieved by any of a variety of methods, reagents, or conditions, including, for example, chemical, enzymatic, or physical shearing (e.g., physical fragmentation). The sheared or cleaved nucleic acids can have a nominal, average, or mean length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to about 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs.
[0097] Sheared or cleaved nucleic acids can be generated by any suitable method, non-limiting examples of which include physical methods (e.g., shearing, e.g., sonication, French press, heating, UV irradiation, etc.), enzymatic treatment (e.g., enzymatic cleavage agents (e.g., suitable nucleases, suitable restriction enzymes, suitable methylation-sensitive restriction enzymes)), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, base hydrolysis, heating, etc., or combinations thereof), treatments described in U.S. Patent Application Publication No. 2005 / 0112590, etc., or combinations thereof. The average, mean, or nominal length of the resulting nucleic acid fragments can be controlled by selecting an appropriate fragment generation method.
[0098] The term "amplification," as used herein, refers to subjecting a target nucleic acid in a sample to a process that linearly or exponentially produces amplicon nucleic acids having the same or substantially the same nucleotide sequence as the target nucleic acid or a portion thereof. In certain embodiments, the term "amplification" refers to a method that includes the polymerase chain reaction (PCR). In certain embodiments, the amplification product can contain one or more more nucleotides than the amplified nucleotide region of the nucleic acid template sequence (e.g., a primer can contain "extra" nucleotides, such as a transcription initiation sequence, in addition to nucleotides complementary to the nucleic acid template gene molecule, resulting in an amplification product that contains "extra" nucleotides, or nucleotides that do not correspond to the amplified nucleotide region of the nucleic acid template gene molecule).
[0099] Furthermore, before providing nucleic acid to the method described herein, nucleic acid can be exposed to a treatment that modifies specific nucleotides in nucleic acid. For example, nucleic acid can be subjected to a treatment that selectively modifies nucleic acid based on the methylation status of nucleotides therein. In addition, conditions such as high temperature, ultraviolet radiation, and X-ray radiation can cause changes in the sequence of nucleic acid molecules. Nucleic acid can be provided in any form that is useful for performing appropriate sequence analysis.
[0100] Nucleic acid concentration In some embodiments, nucleic acids (e.g., extracellular nucleic acids) are enriched or relatively enriched to obtain a subpopulation or species of nucleic acids. Subpopulations of nucleic acids can include, for example, fetal nucleic acids, maternal nucleic acids, cancer nucleic acids, parental nucleic acids, nucleic acids comprising fragments of a particular length or range of lengths, or nucleic acids derived from a particular genomic region (e.g., a single chromosome, a set of chromosomes, and / or a particular chromosomal region). Such enriched samples can be used in conjunction with the methods provided herein. Thus, in certain embodiments, the methods of the present technology include an additional step of enriching for a subpopulation of nucleic acids in a sample, such as cancer or fetal nucleic acids. In certain embodiments, methods for determining the fraction of cancer cell nucleic acids or the fetal fraction can also be used to enrich for cancer or fetal nucleic acids. In certain embodiments, maternal nucleic acids are selectively removed (partially, substantially, almost completely, or completely) from the sample. In certain embodiments, maternal nucleic acids are selectively removed (partially, substantially, almost completely, or completely) from the sample. In certain embodiments, enrichment for a particular low copy number species of nucleic acid (e.g., fetal nucleic acids) can improve quantitative sensitivity. Methods for enriching a sample for a particular species of nucleic acid are described, for example, in U.S. Pat. No. 6,927,028, International Patent Application Publication No. WO2007 / 140417, International Patent Application Publication No. WO2007 / 147063, International Patent Application Publication No. WO2009 / 032779, International Patent Application Publication No. WO2009 / 032781, International Patent Application Publication No. WO2010 / 033639, International Patent Application Publication No. WO2011 / 034631, International Patent Application Publication No. WO2006 / 056480, and International Patent Application Publication No. WO2011 / 143659, the entire contents of each of which are incorporated herein by reference, including all descriptions, tables, equations, and figures.
[0101] In some embodiments, nucleic acids are enriched to obtain specific target and / or reference fragment species. In certain embodiments, nucleic acids are enriched to obtain specific nucleic acid fragment lengths or ranges of fragment lengths using one or more length-based separation methods described below. In certain embodiments, nucleic acids are enriched to obtain fragments derived from selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein and / or known in the art.
[0102] Non-limiting examples of methods for enriching nucleic acid subpopulations in a sample include methods that exploit epigenetic differences between nucleic acid species (e.g., the methylation-based fetal nucleic acid enrichment methods described in U.S. Patent Application Publication No. 2010 / 0105049, incorporated herein by reference), approaches that enhance polymorphic sequences with restriction endonucleases (e.g., methods described in U.S. Patent Application Publication No. 2009 / 0317818, incorporated herein by reference), selective enzymatic degradation approaches, massively parallel signature sequencing (MPSS) approaches, amplification (e.g., PCR)-based approaches (e.g., locus-specific amplification, multiplex SNP allele PCR, universal amplification), pull-down approaches (e.g., biotinylated ultramer pull-down), extension and ligation-based methods (e.g., molecular inversion probe (MIP) extension and ligation), and combinations thereof.
[0103] In some embodiments, nucleic acids are enriched to obtain fragments derived from selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein. Sequence-based separation is generally based on the presence of nucleotide sequences in fragments of interest (e.g., target and / or reference fragments) and that are substantially absent or present in only trace amounts (e.g., 5% or less) in other fragments of the sample. In some embodiments, sequence-based separation can separate target fragments and / or reference fragments. The separated target fragments and / or separated reference fragments are often isolated and removed from the remaining fragments in the nucleic acid sample. In certain embodiments, the separated target fragments and the separated reference fragments are also isolated and removed from each other (e.g., as separate assay compartments). In certain embodiments, the separated target fragments and the separated reference fragments are isolated together (e.g., as the same assay compartment). In some embodiments, unbound fragments can be differentially removed, degraded, or digested.
[0104] In some embodiments, selective nucleic acid capture process is used to separate and extract target fragments and / or reference fragments from nucleic acid samples.Commercially available nucleic acid capture systems include, for example, Nimblegen sequence capture system (Roche NimbleGen, Madison, WI); Illumina BEADARRAY platform (Illumina, San Diego, CA); Affymetrix GENECHIP platform (Affymetrix, Santa Clara, CA); Agilent SureSelect Target Enrichment System (Agilent Technologies, Santa Clara, CA); and related platforms.This method typically involves the hybridization of capture oligonucleotides with part or all of the nucleotide sequence of target fragments or reference fragments, and can include the use of solid phase (for example, solid phase array) and / or solution-based platform. Capture oligonucleotides (sometimes referred to as "baits") are selected or designed to preferentially hybridize to nucleic acid fragments derived from a selected genomic region or locus (e.g., one of chromosomes 21, 18, 13, X, or Y, or a reference chromosome). In certain embodiments, hybridization-based methods (e.g., using oligonucleotide arrays) can be used to enrich for nucleic acid sequences derived from specific chromosomes (e.g., potentially aneuploid chromosomes, reference chromosomes, or other chromosomes of interest), or their genes or regions of interest. Thus, in some embodiments, a nucleic acid sample is optionally enriched by, for example, capturing a subset of fragments using capture oligonucleotides complementary to selected genes in the sample nucleic acid. In certain cases, the captured fragments are amplified. For example, adapter-containing captured fragments can be amplified using primers complementary to the adapter oligonucleotides to form a collection of amplified fragments indexed according to the adapter sequence.In some embodiments, nucleic acids are enriched for fragments derived from selected genomic regions (e.g., chromosomes, genes) by amplification of one or more regions of interest using oligonucleotides (e.g., PCR primers) that are complementary to sequences in fragments containing the region(s) of interest or portions thereof.
[0105] In some embodiments, one or more length-based separation methods are used to enrich nucleic acids for a particular nucleic acid fragment length, a range of lengths, or lengths below or above a particular threshold or cutoff. Nucleic acid fragment length typically refers to the number of nucleotides in the fragment. Nucleic acid fragment length is also sometimes referred to as nucleic acid fragment size. In some embodiments, length-based separation methods are performed without measuring the length of individual fragments. In some embodiments, length-based separation methods are performed in conjunction with methods for determining the length of individual fragments. In some embodiments, length-based separation refers to a size fractionation procedure, and all or a portion of the fractionated pool can be isolated (e.g., retained) and / or analyzed. Size fractionation procedures are known in the art (e.g., separation on an array, separation by molecular sieving, separation by gel electrophoresis, separation by column chromatography (e.g., molecular sieving column), and microfluidic technology-based approaches). In particular examples, length-based separation approaches can include, for example, selective tagging approaches, fragment circularization, chemical treatment (e.g., formaldehyde, polyethylene glycol (PEG) precipitation), mass spectrometry, and / or size-specific nucleic acid amplification.
[0106] Nucleic acid quantification The amount (e.g., concentration, relative amount, absolute amount, copy number, etc.) of nucleic acid in a sample can be determined. The amount (e.g., concentration, relative amount, absolute amount, copy number, etc.) of a low-abundance nucleic acid is determined. In certain embodiments, the amount of a low-abundance nucleic acid species in a sample is referred to as the "low-abundance species fraction." In some embodiments, the "low-abundance species fraction" refers to the fraction of low-abundance nucleic acid species in the circulating cell-free nucleic acids in a sample (e.g., blood sample, serum sample, plasma sample, urine sample) obtained from a subject.
[0107] The amount of low-abundance nucleic acids in extracellular nucleic acid can be quantified and used with the method provided herein.Therefore, in certain embodiments, the method described herein comprises the additional step of determining the amount of low-abundance nucleic acids.The amount of low-abundance nucleic acids in the sample from a subject can be determined before or after processing to prepare sample nucleic acid.In certain embodiments, the amount of low-abundance nucleic acids in the sample is determined after processing and preparing sample nucleic acid, and this amount is used for further evaluation.In some embodiments, the outcome includes adjusting the degree of contribution of low-abundance species fraction in sample nucleic acid (for example, adjusting count number, removing sample, making call, or not making call).
[0108] Determining the minor species fraction can be performed before, during, or at any point within the methods described herein, or after certain methods described herein (e.g., detecting genetic variation or genetic alteration). For example, to perform a genetic variation / genetic alteration determination method with a certain sensitivity or specificity, a quantification method of minor nucleic acids can be performed before, during, or after the genetic variation / genetic alteration determination to identify samples with more than about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25% or more minor nucleic acids. In some embodiments, samples determined to have a certain threshold amount of low-abundance nucleic acids (e.g., about 15% or more low-abundance nucleic acids, about 4% or more low-abundance nucleic acids) are further analyzed, for example, for the presence or absence of genetic variation / alteration or genetic variation / alteration. In certain embodiments, for example, only samples having a certain threshold amount of low-abundance nucleic acids (e.g., about 15% or more low-abundance nucleic acids, about 4% or more low-abundance nucleic acids) are selected for genetic variation or genetic alteration determination (e.g., selected and patients contacted).
[0109] In some embodiments, the amount (e.g., concentration, relative amount, absolute amount, copy number, etc.) of cancer cell nucleic acid in the nucleic acid is determined. In certain cases, the amount of cancer cell nucleic acid in a sample refers to the "fraction of cancer cell nucleic acid," sometimes referred to as the "cancer fraction" or "tumor fraction." In some embodiments, the "fraction of cancer cell nucleic acid" refers to the fraction of cancer cell nucleic acid in circulating cell-free nucleic acid in a sample (e.g., a blood sample, a serum sample, a plasma sample, a urine sample) obtained from a subject.
[0110] In some embodiments, the amount (e.g., concentration, relative amount, absolute amount, copy number, etc.) of fetal nucleic acid in nucleic acid is determined. In certain embodiments, the amount of fetal nucleic acid in a sample is referred to as the "fetal fraction." In some embodiments, "fetal fraction" refers to the fraction of fetal nucleic acid in circulating cell-free nucleic acid in a sample (e.g., blood sample, serum sample, plasma sample, urine sample) obtained from a pregnant female. Certain methods described herein or known in the art for determining the fetal fraction can be used to determine cancer cell nucleic acid and / or minor species fractions.
[0111] In some embodiments, the fraction is determined for regions of copy number variation. In some embodiments, the fetal fraction is determined for regions of copy number variation. In some embodiments, the fraction of a low-abundance nucleic acid is determined. In some embodiments, the fetal fraction is determined for sample nucleic acid. The fraction can be determined according to the methods for estimating or determining fractions (e.g., fetal fractions) described below.
[0112] In certain examples, the fetal fraction can be determined according to markers specific to male fetuses (e.g., Y chromosome STR markers (e.g., DYS19, DYS385, DYS392 markers); RhD markers in RhD-negative females), the allelic ratio of polymorphic sequences, or according to one or more markers specific to fetal nucleic acid but not maternal nucleic acid (e.g., differences in epigenetic biomarkers (e.g., methylation) between the mother and fetus, or fetal RNA markers in maternal plasma (see, e.g., Lo, 2005, Journal of Histochemistry and Cytochemistry, 53(3):293-296)). In some embodiments, the fetal fraction is determined according to a suitable assay of the Y chromosome (e.g., by comparing the abundance of a fetal-specific locus (e.g., the SRY locus on chromosome Y in male pregnancies) with that of any autosomal locus common to both the mother and the fetus, e.g., by using quantitative real-time PCR (e.g., Lo YM et al. (1998) Am J Hum Genet 62:768-775).
[0113] The fetal fraction is sometimes determined using a fetal quantifier assay (FQA), for example, as described in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference. This type of assay allows fetal nucleic acids in a maternal sample to be detected and quantified based on the methylation status of the nucleic acids in the sample. In certain embodiments, the amount of fetal nucleic acid from the maternal sample can be determined relative to the total amount of nucleic acid present, thereby providing the percentage of fetal nucleic acid in the sample. In certain embodiments, the copy number of fetal nucleic acid in the maternal sample can be determined. In certain embodiments, the amount of fetal nucleic acid can be determined in a sequence-specific (or segment-specific) manner, sometimes with sufficient sensitivity to allow accurate chromosome dosage analysis (e.g., to detect the presence or absence of fetal aneuploidy).
[0114] A fetal quantification assay (FQA) can be performed in conjunction with any of the methods described herein. Such an assay can be performed by any method known in the art and / or described in U.S. Patent Application Publication No. 2010 / 0105049, such as a method that can distinguish maternal nucleic acid from fetal nucleic acid based on differences in methylation status and quantify (i.e., determine the amount of) fetal nucleic acid. Methods for differentiating nucleic acids based on methylation status include, but are not limited to, methylation-sensitive capture, e.g., using the MBD2-Fc fragment (in which the methyl-binding domain of MBD2 is fused to the Fc fragment of an antibody (MBD-FC)) (Gebhard et al. (2006) Cancer Res. 66(12):6118-28); methylation-specific antibodies; bisulfite conversion methods, e.g., MSP (methylation-sensitive PCR), COBRA, methylation-sensitive single nucleotide primer extension (Ms-SNuPE), or Sequenom MassCLEAVE™ technology; and the use of methylation-sensitive restriction enzymes (e.g., digesting maternal nucleic acids in a maternal sample with one or more methylation-sensitive restriction enzymes, thereby enriching for fetal nucleic acids). Methyl-sensitive enzymes can also be used to differentiate nucleic acids based on methylation status; these enzymes can preferentially or substantially cleave or digest at their DNA recognition sequences, e.g., when the latter are unmethylated. Thus, unmethylated DNA samples are cut into smaller fragments than methylated DNA samples, and hypermethylated DNA samples are not cut. Unless explicitly stated, any method for differentiating nucleic acids based on methylation status can be used with the compositions and methods of the technology herein. The amount of fetal nucleic acid can be determined, for example, by introducing one or more competitors at known concentrations during the amplification reaction. The amount of fetal nucleic acid can also be determined, for example, by RT-PCR, primer extension, sequencing, and / or counting. In certain cases, the amount of nucleic acid can be determined using BEAMing technology as described in U.S. Patent Application Publication No. 2007 / 0065823.In certain embodiments, the restriction efficiency can be determined, and the ratio of efficiencies is used to further determine the amount of fetal nucleic acid.
[0115] In certain embodiments, the minor species fraction can be determined based on the ratio of alleles of a polymorphic sequence (e.g., a single nucleotide polymorphism (SNP)), for example, using a method such as that described in U.S. Patent Application Publication No. 2011 / 0224087, which is incorporated herein by reference. In such methods for determining the fetal fraction, for example, nucleotide sequence reads are obtained for a maternal sample, and the fetal fraction is determined by comparing the total number of nucleotide sequence reads that map to a first allele with the total number of nucleotide sequence reads that map to a second allele at a reference polymorphic site (e.g., an SNP) in a reference genome. In certain embodiments, for example, the fetal allele is distinguished from the mixture of fetal and maternal nucleic acids in a sample by the maternal nucleic acid making a large contribution to the mixture, compared with the relatively small contribution of the fetal allele. Thus, the relative abundance of fetal nucleic acid in the maternal sample can be determined as a parameter of the total number of unique sequence reads that mapped to the target nucleic acid sequence on the reference genome for each of the two alleles at the polymorphic site.
[0116] In some embodiments, the minor species fraction can be determined using a method that incorporates information derived from chromosomal abnormalities, such as that described in International Patent Application Publication No. WO2014 / 055774, which is incorporated herein by reference. In some embodiments, the minor species fraction can be determined using a method that incorporates information derived from sex chromosomes, such as that described in U.S. Patent Application Publication Nos. 2013 / 0288244 and 2013 / 0338933, each of which is incorporated herein by reference.
[0117] In some embodiments, the minor species fraction can be determined using methods that incorporate fragment length information (e.g., fragment length ratio (FLR) analysis, fetal ratio statistic (FRS) analysis, as described in International Patent Application Publication No. WO2013 / 177086, incorporated herein by reference). Fragments of cell-free fetal nucleic acid are generally shorter than fragments of maternal nucleic acid (see, e.g., Chan et al. (2004) Clin. Chem. 50:88-92; Lo et al. (2010) Sci. Transl. Med. 2:61ra91). Thus, in some embodiments, the fetal fraction can be determined by counting fragments below a certain length threshold and comparing these counts to, for example, the counts obtained from fragments above a certain length threshold and / or the amount of total nucleic acid in the sample. Methods for counting nucleic acid fragments of a particular length are described in further detail in International Patent Application Publication No. WO2013 / 177086.
[0118] In certain embodiments, the FLR or FRS is determined in part according to the amount of reads mapped to portions derived from CCF fragments having lengths less than a selected fragment length. In some embodiments, the FLR or FRS value is often the ratio of X to Y, where X is the amount of reads derived from CCF fragments having lengths less than a first selected fragment length, and Y is the amount of reads derived from CCF fragments having lengths less than a second selected fragment length. The first selected fragment length is often selected independently of the second selected fragment length, and vice versa, and the second selected fragment length is usually longer than the first selected fragment length. The first selected fragment length can be about 200 bases or less to about 30 bases or less. In some embodiments, the first selected fragment length is about 200, 190, 180, 170, 160, 155, 150, 145, 140, 135, 130, 125, 120, 115, 110, 105, 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, or 50 bases. In some embodiments, the first selected fragment length is about 170 to about 130 bases, and sometimes about 160 to about 140 bases. In some embodiments, the second selected fragment length is about 2000 bases to about 200 bases. In certain embodiments, the second selected fragment length is about 1000, 950, 800, 850, 800, 750, 700, 650, 600, 550, 500, 450, 400, 350, 300, or 250 bases. In some embodiments, the first selected fragment length is about 140 to about 160 bases (e.g., about 150 bases), and the second selected fragment length is about 500 to about 700 bases (e.g., about 600 bases). In some embodiments, the first selected fragment length is about 150 bases, and the second selected fragment length is about 600 bases.
[0119] In some embodiments, the minor species fraction can be determined according to the level. For example, the fetal fraction can be determined according to the level (for example, the level for the affected region, the level for the copy number variation). Determining the fetal fraction according to the level can include determining the absolute value of the deviation of the level from the expected level and multiplying the absolute value of the deviation by 2. The expected level can be given a value of 1, and the deviation of the first or second level can be negative (for example, for deletion or microdeletion, the level is less than 1) or positive (for example, for duplication or microduplication, the level is greater than 1). In certain cases, the magnitude of deviation can vary depending on the fetal fraction.
[0120] In some embodiments, determining the fraction of low abundance species (e.g., cancer cell nucleic acid fraction, fetal fraction) is not necessary or is required to identify the presence or absence of a genetic variation or genetic alteration. In some embodiments, identifying the presence or absence of a genetic variation or genetic alteration does not require sequence discrimination between low abundance nucleic acids and high abundance nucleic acids. In certain embodiments, this is because the combined contribution of both low abundance and high abundance sequences in a particular chromosome, chromosomal portion, or part thereof is analyzed. In some embodiments, identifying the presence or absence of a genetic variation or genetic alteration does not rely on prior sequence information that would distinguish low abundance nucleic acids from high abundance nucleic acids.
[0121] Partially specific fraction estimates In some embodiments, the low-abundance species fraction can be determined according to a portion-specific fraction estimate (e.g., as described in International Patent Application Publication No. WO2014 / 205401 and Kim et al. (2015) Prenatal Diagnosis 35:810-815, each of which is incorporated herein by reference). For example, in some embodiments, the fetal fraction (e.g., for a sample) can be determined according to a portion-specific fetal fraction estimate. Without being limited by theory, the amount of reads obtained from fetal circulating cell-free (CCF) fragments (e.g., fragments of a particular length or range of lengths) is often mapped using a frequency range for the portion (e.g., within the same sample, e.g., within the same sequencing run). Also, without being limited by theory, a particular portion, when compared across multiple samples, shows a similar representation of reads obtained from fetal CCF fragments (e.g., fragments of a particular length or range of lengths), which representation tends to correlate with the portion-specific fetal fraction (e.g., the relative amount, percentage, or ratio of CCF fragments originating from the fetus). Fetal fractions estimated according to portion-specific fraction estimates are sometimes referred to herein as sequencing-based fetal fractions (e.g., SeqFF) and / or bin-based fetal fractions (BFF).
[0122] In some embodiments, estimates of the part-specific fetal fraction are determined based, in part, on part-specific parameters and their relationship to the fetal fraction. The part-specific parameters can be any suitable parameter that reflects (e.g., correlates with) the amount or proportion of reads obtained from CCF fragment lengths of a particular size (e.g., size range) in the part. The part-specific parameters can be the average, mean, or median of the part-specific parameters determined for multiple samples. Any suitable part-specific parameter can be used. Non-limiting examples of portion-specific parameters include counts (e.g., counts of sequence reads mapped to a portion, counts of sequence reads mapped to a portion in a reference genome), normalized counts (e.g., normalized counts of sequence reads mapped to a portion, normalized counts of sequence reads mapped to a portion in a reference genome), fragment length ratio (FLR), fetal ratio statistic (FRS), amount of reads having a length less than a selected fragment length, genome coverage (i.e., coverage), mappability, DNase I sensitivity, methylation status, acetylation, histone distribution, guanine-cytosine (GC) content, chromatin structure, etc., or combinations thereof. In some embodiments, the portion-specific parameter can be any suitable parameter that correlates with FLR and / or FRS in a portion-specific manner. In some embodiments, some or all of the portion-specific parameters are direct or indirect indications of FLR for the portion. In some embodiments, the portion-specific parameter is not guanine-cytosine (GC) content.
[0123] In some embodiments, the portion-specific parameter is any suitable value that indicates, correlates with, or is proportional to the amount of reads obtained from a CCF fragment, where the reads mapped to the portion have a length less than the selected fragment length. In certain embodiments, the portion-specific parameter is an indication of the amount of reads obtained from a relatively short CCF fragment (e.g., about 200 base pairs or less, about 150 base pairs or less) mapped to the portion. CCF fragments having a length less than the selected fragment length are often relatively short CCF fragments, and sometimes the selected fragment length is about 200 base pairs or less (e.g., CCF fragments that are about 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, or 80 bases in length). The length of the CCF fragment or the reads obtained from the CCF fragment can be determined (e.g., estimated or inferred) by any suitable method (e.g., sequencing, hybridization approach). In some embodiments, the length of the CCF fragment is determined (e.g., estimated or inferred) from reads obtained from paired-end sequencing. In certain embodiments, the length of the CCF fragment template is determined directly from the length of reads obtained from the CCF fragment (e.g., reads from a single end).
[0124] The part-specific parameters can be weighted, adjusted, or transformed by one or more weighting factors. In some embodiments, the weighted, adjusted, or transformed part-specific parameters can provide an estimate of the part-specific fetal fraction for a sample (e.g., a test sample). In some embodiments, the weighting or adjustment generally converts part counts (e.g., reads mapped to parts) or another part-specific parameter into an estimate of the part-specific fetal fraction, and such a conversion is sometimes considered a translation.
[0125] In some embodiments, the weighting coefficient is a coefficient or constant that describes and / or defines, in part, the relationship between the fetal fraction (e.g., the fetal fraction determined from a plurality of samples) and the part-specific parameters for a plurality of samples (e.g., a training set). In some embodiments, the weighting coefficient is determined according to the relationship between the determination results of the fetal fraction and the part-specific parameters. A relationship can be defined by one or more weighting coefficients, and one or more weighting coefficients can be determined from a relationship. In some embodiments, the weighting coefficient (e.g., one or more weighting coefficients) is determined from a relationship for the parts that is fitted according to (i) the fraction of fetal nucleic acid determined for each of the plurality of samples (e.g., a plurality of samples in a training set) and (ii) the part-specific parameters for the plurality of samples (e.g., a plurality of samples in a training set).
[0126] The weighting coefficients can be any suitable coefficients, estimated coefficients, or constants obtained from a suitable relationship (e.g., a suitable mathematical relationship, algebraic relationship, fitted relationship, regression, regression analysis, regression model). The weighting coefficients can be determined according to, derived from, or estimated from a suitable relationship. In some embodiments, the weighting coefficients are estimated coefficients from the fitted relationship. Fitting a relationship for a plurality of samples is sometimes referred to herein as training a model. Any suitable model and / or method for fitting a relationship (e.g., training a model to obtain a training set) can be used. Non-limiting examples of suitable models that can be used include regression models, linear regression models, simple regression models, ordinary least squares regression models, multiple regression models, general multiple regression models, polynomial regression models, general linear models, generalized linear models, discrete choice regression models, logistic regression models, multinomial logit models, mixed logit models, probit models, multinomial probit models, ordered logit models, ordered probit models, Poisson models, multivariate response regression models, multilevel models, fixed effects models, random effects models, mixed models, nonlinear regression models, nonparametric models, semiparametric models, robust models, quantile models, isotonic models, principal component models, least angle models, local models, segmented models, and errors in variables models. In some embodiments, the fitted relationship is not a regression model. In some embodiments, the fitted relationship is selected from a decision tree model, a support vector machine model, and a neural network model. The result of training a model (e.g., a regression model, a relationship) is often a relationship that can be described mathematically, which includes one or more coefficients (e.g., weighting coefficients). For example, for a linear least squares model, a general multiple regression model can be trained using fetal fraction values and part-specific parameters (e.g., coverage, see, e.g., Example 4), resulting in the relationship described by equation (1), where the weighting coefficient β is further defined in equations (2), (3), and (4).More complex multivariate models can determine one, two, three, or more weighting coefficients. In some embodiments, the model is trained according to fetal fractions obtained from multiple samples and two or more site-specific parameters (e.g., coefficients) (e.g., fitting relationships to multiple samples, e.g., by a matrix).
[0127] The weighting coefficients may be obtained from any suitable relationship (e.g., any suitable mathematical relationship, algebraic relationship, fitted relationship, regression, regression analysis, regression model) by any suitable method. In some embodiments, the fitted relationship is fitted by estimation, non-limiting examples of which include least squares, ordinary least squares, linear regression, partial regression, full regression, generalized regression, weighted regression, nonlinear regression, iteratively weighted regression, ridge regression, least absolute deviation, Bayes, Bayesian multivariate, reduced rank, LASSO, Weighted Rank Selection Criteria (WRSC), Rank Selection Criteria (RSC), elastic net estimation methods (e.g., elastic net regression), and combinations thereof.
[0128] The weighting factors may have any suitable value. In some embodiments, the weighting factors are about -1×10 -2 and approximately 1 x 10 -2 Between, approximately -1×10 -3 and approximately 1 x 10 -3 Between, approximately -5×10 -4 and about 5 × 10 -4 Between or about -1×10 -4 and approximately 1 x 10 -4In some embodiments, the distribution of weighting coefficients for the plurality of samples is substantially symmetric. Sometimes, the distribution of weighting coefficients for the plurality of samples is a normal distribution. Sometimes, the distribution of weighting coefficients for the plurality of samples is not a normal distribution. In some embodiments, the width of the distribution of weighting coefficients varies depending on the amount of reads derived from CCF fetal nucleic acid fragments. In some embodiments, fractions containing higher fetal nucleic acid content generate larger coefficients (e.g., positive or negative, see, e.g., FIG. 19 ). A weighting coefficient can be zero, or the weighting coefficient can be greater than zero. In some embodiments, about 70% or more, about 75% or more, about 80% or more, about 85% or more, about 90% or more, about 95% or more, or about 98% or more of the weighting coefficients of the fractions are greater than zero.
[0129] Weighting coefficients can be determined for or associated with any suitable portion of a genome. Weighting coefficients can be determined for or associated with any suitable portion of any suitable chromosome. In some embodiments, weighting coefficients are determined for or associated with some or all portions in a genome. In some embodiments, weighting coefficients are determined for or associated with portions of some or all chromosomes in a genome. Weighting coefficients are sometimes determined for or associated with selected chromosome portions. Weighting coefficients can be determined for or associated with portions of one or more autosomes. Weighting coefficients can be determined for or associated with portions of a plurality of portions, including portions of autosomes or a subset thereof. In some embodiments, weighting coefficients are determined for or associated with portions of sex chromosomes (e.g., ChrX and / or ChrY). Weighting coefficients can be determined for or associated with portions of one or more autosomes and one or more sex chromosomes. In certain embodiments, weighting coefficients are determined for or associated with the portions of the plurality of portions in all autosomes and chromosomes X and Y. Weighting coefficients can be determined for or associated with the portions of the plurality of portions that do not include the portions in chromosomes X and / or Y. In certain embodiments, weighting coefficients are determined for or associated with the portions of a chromosome, and this chromosome includes aneuploidy (e.g., whole chromosome aneuploidy). In certain embodiments, weighting coefficients are determined for or associated only with the portions of a chromosome, and this chromosome is not aneuploid (e.g., is a euploid chromosome). Weighting coefficients can be determined for or associated with the portions of the plurality of portions that do not include the portions in chromosomes 13, 18, and / or 21.
[0130] In some embodiments, weighting coefficients are determined for portions according to one or more samples (e.g., samples from a training set). The weighting coefficients are often portion-specific. In some embodiments, one or more weighting coefficients are independently assigned to portions. In some embodiments, weighting coefficients are determined according to a relationship between fetal fraction determinations for multiple samples (e.g., sample-specific fetal fraction determinations) and portion-specific parameters determined according to multiple samples. Weighting coefficients are often determined from multiple samples, for example, from about 20 to about 100,000 or more, from about 100 to about 100,000 or more, from about 500 to about 100,000 or more, from about 1,000 to about 100,000 or more, or from about 10,000 to about 100,000 or more samples. The weighting factor can be determined from a sample that is euploid (e.g., a sample obtained from a subject with a euploid fetus, e.g., a sample in which no aneuploid chromosomes are present). In some embodiments, the weighting factor is obtained from a sample that contains aneuploid chromosomes (e.g., a sample obtained from a subject with a euploid fetus). In some embodiments, the weighting factor is determined from multiple samples obtained from a subject with a euploid fetus and a subject with a trisomic fetus. The weighting factor can be obtained from multiple samples, and these samples are obtained from subjects with male and / or female fetuses.
[0131] The fetal fraction is often determined for one or more samples of the training set, from which weighting factors are derived. The fetal fraction from which the weighting factors are determined is sometimes the result of a sample-specific fetal fraction determination. The fetal fraction from which the weighting factors are determined can be determined by any suitable method described herein or known in the art. In some embodiments, the determination of fetal nucleic acid content (e.g., fetal fraction) is performed using a suitable fetal quantification assay (FQA) described herein or known in the art, including, but not limited to, determination according to markers specific to male fetuses, determination based on the allelic ratio of polymorphic sequences, determination according to one or more markers specific to fetal nucleic acids but not maternal nucleic acids, determination by using methylation-based DNA discrimination (e.g., A. Nygren et al. (2010) Clinical Chemistry, 56(10):1627-1635), competitive PCR approaches, and the like. Examples of methods include determination by mass spectrometry methods and / or systems using a genomic DNA analysis approach, determination by the methods described in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference, or a combination thereof. In certain examples, the fetal fraction is determined, in part, at the level of the Y chromosome (e.g., at the level of one or more genomic segments; profile level). In some embodiments, the fetal fraction is determined according to an appropriate assay of the Y chromosome (e.g., using quantitative real-time PCR to compare the amount of a fetal-specific locus (e.g., the SRY locus on the Y chromosome in the case of a male fetus) with the amount of a locus on any autosome common to both the mother and the fetus (e.g., Lo YM et al. (1998) Am J Hum Genet, 62:768-775)).
[0132] The part-specific parameters (e.g., for the test sample) can be weighted, adjusted, or transformed by one or more weighting factors (e.g., weighting factors derived from a training set). For example, weighting factors can be derived for parts according to the relationship between the part-specific parameters and the fetal fraction determination results for a training set of multiple samples. The part-specific parameters of the test sample can then be adjusted and / or weighted according to the weighting factors derived from the training set. In some embodiments, the part-specific parameters from which the weighting factors are derived are the same as the part-specific parameters (e.g., for the test sample) that are adjusted or weighted (e.g., both parameters are FLR). In certain embodiments, the part-specific parameters from which the weighting factors are derived are different from the part-specific parameters (e.g., for the test sample) that are adjusted or weighted. For example, weighting factors can be determined from the relationship between coverage (i.e., part-specific parameters) and fetal fraction for the samples in the training set, and the FLR (i.e., another part-specific parameter) for the parts of the test sample can be adjusted according to the weighting factor derived from coverage. Without being limited by theory, the part-specific parameters (e.g., for a test sample) can sometimes be adjusted and / or weighted and / or transformed by weighting factors derived from different part-specific parameters (e.g., of a training set) due to the relationship and / or correlation between each part-specific parameter and the common part-specific FLR.
[0133] An estimate of the portion-specific fetal fraction can be determined for a sample (e.g., a test sample) by weighting, adjusting, or transforming a portion-specific parameter (e.g., the count of sequence reads mapped to a portion of the reference genome) by a weighting coefficient determined for that portion. Weighting can include adjusting, transforming, and / or converting a portion-specific parameter (e.g., the count of sequence reads mapped to a portion of the reference genome) by a weighting coefficient by applying any suitable mathematical operation, non-limiting examples of which include multiplication, division, addition, subtraction, integration, symbolic calculation, algebraic calculation, algorithm, trigonometric or geometric function, transformation (e.g., Fourier transform), etc., or a combination thereof. Weighting can include adjusting, transforming, and / or converting a portion-specific parameter (e.g., the count of sequence reads mapped to a portion of the reference genome) by a weighting coefficient by an appropriate mathematical model (e.g., the model presented in Example 4).
[0134] In some embodiments, the fetal fraction is determined for a sample according to one or more part-specific fetal fraction estimates.In some embodiments, the fetal fraction is determined (e.g., estimated) for a sample (e.g., test sample) according to weighting, adjustment, or transformation of part-specific parameters (e.g., the count of sequence reads that map to a part of a reference genome) for one or more parts.In certain embodiments, the fetal nucleic acid fraction for a test sample is estimated based on adjusted counts or adjusted subset counts.In certain embodiments, the fetal nucleic acid fraction for a test sample is estimated based on adjusted FLR, adjusted FRS, adjusted coverage, and / or adjusted mappability for a part. In some embodiments, between about 1 and about 500,000, between about 100 and about 300,000, between about 500 and about 200,000, between about 1000 and about 200,000, between about 1500 and about 200,000, or between about 1500 and about 50,000 part-specific parameters are weighted or adjusted.
[0135] The fetal fraction (e.g., for a test sample) is determined according to multiple part-specific fetal fraction estimates (e.g., for the same test sample) by any suitable method. In some embodiments, a method for improving the accuracy of estimating the fraction of fetal nucleic acid in a test sample obtained from a pregnant female includes determining one or more part-specific fetal fraction estimates, and the fetal fraction estimate for the sample is determined according to the one or more part-specific fetal fraction estimates. In some embodiments, estimating or determining the fraction of fetal nucleic acid for a sample (e.g., a test sample) includes a substep of summing one or more part-specific fetal fraction estimates. The summing substep can include determining the mean, average, median, AUC, or integral value according to the multiple part-specific fetal fraction estimates.
[0136] In some embodiments, a method for improving the accuracy of estimating the fraction of fetal nucleic acid in a test sample obtained from a pregnant female includes obtaining counts of sequence reads mapped to portions of a reference genome, where these sequence reads are reads of circulating cell-free nucleic acid obtained from a test sample from a pregnant female, and at least a subset of the obtained counts are obtained from a region of the genome, where the region provides a higher number of counts obtained from fetal nucleic acid compared to the total number of counts from this region than the total number of counts obtained from fetal nucleic acid compared to another region of the genome. In some embodiments, the estimate of the fraction of fetal nucleic acid is determined according to a subset of the portions, where the subset of portions is selected according to a portion to which the number of counts obtained from fetal nucleic acid compared to non-fetal nucleic acid is mapped to a higher number than the number of counts obtained from fetal nucleic acid compared to non-fetal nucleic acid in another portion. In some embodiments, the subset of portions is selected according to a portion to which the number of counts obtained from fetal nucleic acid compared to non-fetal nucleic acid is mapped to a higher number than the number of counts obtained from fetal nucleic acid compared to non-fetal nucleic acid in another portion. The counts mapped to all or a subset of the portions can be weighted, adjusted, or converted to obtain weighted, adjusted, or converted counts. The weighted, adjusted, or converted counts can be used to estimate the fraction of fetal nucleic acid, and the counts can be weighted, adjusted, or converted according to the portion to which the counts obtained from fetal nucleic acid are mapped that are greater than the counts of fetal nucleic acid in another portion. In some embodiments, the counts are weighted according to the portion to which the counts obtained from fetal nucleic acid relative to non-fetal nucleic acid are mapped that are greater than the counts of fetal nucleic acid relative to non-fetal nucleic acid in another portion.
[0137] A fetal fraction can be determined for a sample (e.g., a test sample) according to a plurality of part-specific fetal fraction estimates for the sample, where the part-specific estimates are obtained from parts of any suitable region or segment of the genome. A part-specific fetal fraction estimate can be determined for one or more parts of suitable chromosomes (e.g., one or more selected chromosomes, one or more autosomes, sex chromosomes (e.g., ChrX and / or ChrY), aneuploid chromosomes, euploid chromosomes, etc., or a combination thereof). In some embodiments, a fetal fraction can be determined for a sample (e.g., a test sample) according to a plurality of part-specific fetal fraction estimates for the sample, where the part-specific estimates are obtained from parts or portions of chromosomes classified as having copy number variations (e.g., aneuploidy, microduplication, microdeletion). A fetal fraction determined according to a plurality of part-specific fetal fraction estimates for a sample, where the part-specific estimates are obtained from parts or portions of chromosomes classified as having copy number variations, is sometimes referred to herein as an affected fraction (AF).
[0138] The portion-specific parameters (e.g., counts of sequence reads mapped to portions of the reference genome), weighting factors, portion-specific fetal fraction estimates and / or fetal fraction determinations can be determined by a suitable system, machine, apparatus, non-transitory computer-readable storage medium (e.g., having an executable program stored thereon), etc., or combinations thereof. In certain embodiments, the portion-specific parameters (e.g., counts of sequence reads mapped to portions of the reference genome), weighting factors, portion-specific fetal fraction estimates and / or fetal fraction determinations are determined (e.g., in part) by a system or machine including one or more microprocessors and memory. In some embodiments, the portion-specific parameters (e.g., counts of sequence reads mapped to portions of the reference genome), weighting factors, portion-specific fetal fraction estimates and / or fetal fraction determinations are determined (e.g., in part) by a non-transitory computer-readable storage medium having an executable program stored thereon, where the program instructs the microprocessor to perform the determinations.
[0139] In some embodiments, the fraction is determined for regions of copy number variation. In some embodiments, the fetal fraction is determined for regions of copy number variation. In some embodiments, the fraction of a low abundance nucleic acid is determined. In some embodiments, the fetal fraction of the sample nucleic acid is determined. The fraction can be determined according to the sequencing-based fetal fraction estimation described herein. In some embodiments, a sequencing-based fraction (e.g., fetal fraction) estimate is generated according to a method that includes: (i) obtaining counts of sequence reads mapped to portions of a reference genome, wherein the sequence reads are obtained from sample nucleic acid derived from a subject; (ii) converting the counts of sequence reads mapped to each portion into a portion-specific fraction of nucleic acid (e.g., fetal nucleic acid) according to a weighting coefficient independently associated with each portion, thereby providing a portion-specific fraction estimate (e.g., fetal fraction estimate) for the sample nucleic acid derived from the subject according to the weighting coefficients, wherein each of the weighting coefficients is determined from a fitted relationship for each portion between (1) the fraction of nucleic acid (e.g., fetal nucleic acid) for each of a plurality of samples in a training set and (2) the counts of sequence reads mapped to each portion for the plurality of samples; and (iii) estimating the fraction of nucleic acid (e.g., fetal nucleic acid) for the sample nucleic acid derived from the subject based on the portion-specific fraction estimates (e.g., fetal fraction estimates).
[0140] To determine the fraction for a region of copy number variation, the counts of sequence reads mapped to each portion in the region of copy number variation are converted to a portion-specific fraction of nucleic acid according to a weighting factor independently associated with each portion in the region of copy number variation, thereby providing a portion-specific fraction estimate. To determine the fetal fraction for a region of copy number variation, the counts of sequence reads mapped to each portion in the region of copy number variation are converted to a portion-specific fetal fraction of nucleic acid according to a weighting factor independently associated with each portion in the region of copy number variation, thereby providing a portion-specific fetal fraction estimate.
[0141] To determine the fraction of a low-abundance nucleic acid, the counts of sequence reads mapped to each portion in a plurality of regions (e.g., regions not restricted to the copy number variation regions described above, regions spanning the genome) are converted to a portion-specific fraction of nucleic acid according to a weighting factor independently associated with each portion, thereby providing a portion-specific fraction estimate. To determine the fetal fraction for a sample nucleic acid, the counts of sequence reads mapped to each portion in a plurality of regions (e.g., regions not restricted to the copy number variation regions described above, regions spanning the genome) are converted to a portion-specific fraction of fetal nucleic acid according to a weighting factor independently associated with each portion, thereby providing a portion-specific fetal fraction estimate.
[0142] Nucleic Acid Library In some embodiments, a nucleic acid library is a plurality of polynucleotide molecules (e.g., a sample of nucleic acids) that are prepared, collected, and / or modified for a particular process (non-limiting examples of which include immobilization on a solid phase (e.g., a solid support, a flow cell, beads), enrichment, amplification, cloning, detection) and / or for nucleic acid sequencing. In certain embodiments, the nucleic acid library is prepared before or during the sequencing process. Nucleic acid libraries (e.g., sequencing libraries) can be prepared by any suitable method known in the art. Nucleic acid libraries can be prepared by targeted or non-targeted preparation processes.
[0143] In some embodiments, libraries of nucleic acids are modified to include chemical moieties (e.g., functional groups) configured for immobilization of the nucleic acids to a solid support. In some embodiments, libraries of nucleic acids are modified to include biological molecules (e.g., functional groups) and / or members of binding pairs configured for immobilization of the library to a solid support, non-limiting examples of which include thyroxine-binding globulin, steroid-binding proteins, antibodies, antigens, haptens, enzymes, lectins, nucleic acids, repressors, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding proteins, receptors, carbohydrates, oligonucleotides, polynucleotides, complementary nucleic acid sequences, and the like, and combinations thereof. Some examples of specific binding pairs include, but are not limited to, an avidin moiety and a biotin moiety; an antigenic epitope and an antibody or immunologically reactive fragment thereof; an antibody and a hapten; a digoxigen moiety and an anti-digoxigenin moiety. ) antibodies; fluorescein moieties and anti-fluorescein antibodies; operators and repressors; nucleases and nucleotides; lectins and polysaccharides; steroids and steroid binding proteins; active compounds and receptors for active compounds; hormones and hormone receptors; enzymes and substrates; immunoglobulins and Protein A; oligonucleotides or polynucleotides and their corresponding complements, and the like, or combinations thereof.
[0144] In some embodiments, a library of nucleic acids is modified to include one or more polynucleotides of known composition, including, but not limited to, identifiers (e.g., tags, index tags), capture sequences, labels, adapters, restriction enzyme sites, promoters, enhancers, origins of replication, stem-loops, complementary sequences (e.g., primer binding sites, annealing sites), suitable integration sites (e.g., transposons, viral integration sites), modified nucleotides, etc., or combinations thereof. Polynucleotides of known sequence can be added to any suitable position, such as the 5' end, 3' end, or internal position of the nucleic acid sequence. Polynucleotides of known sequence can be the same or different sequences. In some embodiments, polynucleotides of known sequence are configured to hybridize to one or more oligonucleotides immobilized on a surface (e.g., a surface in a flow cell). For example, a nucleic acid molecule containing a 5' known sequence can be hybridized to a first plurality of oligonucleotides, while the 3' known sequence of the molecule can be hybridized to a second plurality of oligonucleotides. In some embodiments, the nucleic acid library can include chromosome-specific tags, capture sequences, labels, and / or adapters. In some embodiments, the nucleic acid library includes one or more detectable labels. In some embodiments, one or more detectable labels can be incorporated into the nucleic acid library at the 5' end, the 3' end, and / or at any nucleotide position within the nucleic acids in the library. In some embodiments, the nucleic acid library includes hybridized oligonucleotides. In certain embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, the nucleic acid library includes hybridized oligonucleotide probes prior to immobilization on a solid phase.
[0145] In some embodiments, the polynucleotide of known sequence comprises a universal sequence. A universal sequence is a specific nucleotide sequence that is incorporated into two or more nucleic acid molecules, or two or more subsets of nucleic acid molecules, and the universal sequence is the same for all molecules in the molecules or subsets into which it is incorporated. Universal sequences are often designed to hybridize to and / or amplify multiple different sequences using a single universal primer that is complementary to the universal sequence. In some embodiments, two (e.g., pairs) or more universal sequences and / or universal primers are used. Universal primers often comprise a universal sequence. In some embodiments, an adapter (e.g., a universal adapter) comprises a universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple species or subsets of nucleic acids.
[0146] In certain embodiments of nucleic acid library preparation (e.g., in the case of specific sequencing by synthesis procedures), nucleic acids are size-selected and / or fragmented to lengths of a few hundred base pairs or less (e.g., in preparation for library generation). In some embodiments, library preparation is performed without fragmentation (e.g., when using cell-free DNA).
[0147] In certain embodiments, ligation-based library preparation methods are used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego, CA). Ligation-based library preparation methods often utilize adapter (e.g., methylated adapter) design, which can incorporate an index sequence (e.g., a sample index sequence for identifying the sample origin for the nucleic acid sequence) in the initial ligation step and can often be used to prepare samples for single-end sequencing, double-end sequencing, and multiplex sequencing. For example, end repair of nucleic acids (e.g., fragmented nucleic acids or cell-free DNA) can be performed by a fill-in reaction, an exonuclease reaction, or a combination thereof. In some embodiments, the resulting blunt-end repaired nucleic acid can then be extended with a single nucleotide that is complementary to the single-nucleotide overhang on the 3' end of the adapter / primer. Any nucleotide can be used for the extension / overhang nucleotide.
[0148] In some embodiments, preparing a nucleic acid library involves ligating an adapter oligonucleotide (e.g., to a sample nucleic acid, a sample nucleic acid fragment, or a template nucleic acid). The adapter oligonucleotide often exhibits complementarity to a flow cell anchor and is sometimes used, for example, to immobilize a nucleic acid library to a solid support, such as the inner surface of a flow cell. In some embodiments, the adapter oligonucleotide comprises an identifier, one or more sequencing primer hybridization sites (e.g., a sequence exhibiting complementarity to a universal sequencing primer, a single-end sequencing primer, a double-end sequencing primer, a multiplex sequencing primer, etc.), or a combination thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing). In some embodiments, the adapter oligonucleotide comprises one or more of a primer annealing polynucleotide (e.g., for annealing with a flow cell-coupled oligonucleotide and / or with a free amplification primer), an index polynucleotide (e.g., a sample index sequence for tracking nucleic acids from different samples, also called a sample ID), and a barcode polynucleotide (e.g., a single molecule barcode (SMB), also called a molecular barcode, for tracking individual molecules of sample nucleic acid to be amplified prior to sequencing). In some embodiments, the primer annealing component of the adapter oligonucleotide comprises one or more universal sequences (e.g., a sequence complementary to one or more universal amplification primers). In some embodiments, the index polynucleotide (e.g., a sample index, sample ID) is a component of the adapter oligonucleotide. In some embodiments, the index polynucleotide (e.g., a sample index, sample ID) is a component of a universal amplification primer sequence.
[0149] In some embodiments, the adapter oligonucleotides are designed to generate a library construct that includes one or more of a universal sequence, a molecular barcode, a sample ID sequence, a spacer sequence, and a sample nucleic acid sequence when used in combination with an amplification primer (e.g., a universal amplification primer). In some embodiments, the adapter oligonucleotides are designed to generate a library construct that includes an ordered combination of one or more of a universal sequence, a molecular barcode, a sample ID sequence, a spacer sequence, and a sample nucleic acid sequence when used in combination with a universal amplification primer. For example, a library construct may include a first universal sequence, followed by a second universal sequence, followed by a first molecular barcode, followed by a spacer sequence, followed by a template sequence (e.g., a sample nucleic acid sequence), followed by a spacer sequence, followed by a second molecular barcode, followed by a third universal sequence, followed by a sample ID, followed by a fourth universal sequence. In some embodiments, the adapter oligonucleotides are designed to generate a library construct for each strand of a template molecule (e.g., a sample nucleic acid molecule) when used in combination with an amplification primer (e.g., a universal amplification primer). In some embodiments, the adapter oligonucleotide is a double-stranded adapter oligonucleotide.
[0150] An identifier is a suitable detectable label incorporated into or tethered to a nucleic acid (e.g., a polynucleotide), allowing the detection and / or identification of the nucleic acid containing it. In some embodiments, the identifier is incorporated into or tethered to a nucleic acid (e.g., by a polymerase) during a sequencing method. Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indexes or barcodes, radiolabels (e.g., isotopes), metal labels, fluorescent labels, chemiluminescent labels, phosphorescent labels, fluorophore quenchers, dyes, proteins (e.g., enzymes, antibodies or parts thereof, linkers, members of binding pairs), etc., or combinations thereof. In some embodiments, the identifier (e.g., nucleic acid index or barcode) is a unique, known, and / or identifiable sequence of nucleotides or nucleotide analogs. In some embodiments, the identifier is six or more adjacent nucleotides. Numerous fluorophores are available with a variety of different excitation and emission spectra. Any suitable type and / or number of fluorophores can be used as identifiers. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, or fifty or more different identifiers are utilized in the methods described herein (e.g., nucleic acid detection and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are linked to each nucleic acid in the library.Detection and / or quantification of the identifiers may be performed by any suitable method, device or machine, non-limiting examples of which include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, luminometer, fluorometer, spectrophotometer, suitable gene chip or microarray analysis, Western blot, mass spectrometry, chromatography, cytofluorimetric analysis, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, electric field suspension, suitable nucleic acid sequencing methods and / or nucleic acid sequencing devices, etc., and combinations thereof.
[0151] In some embodiments, transposon-based library preparation methods are used (e.g., EPICENTRE NEXTERA, Epicentre, Madison WI). Transposon-based methods typically use in vitro transposition to simultaneously fragment and tag DNA in a single-tube reaction (often allowing the incorporation of platform-specific tags and optional barcodes) to prepare sequencing-compatible libraries.
[0152] In some embodiments, the nucleic acid library, or portions thereof, are amplified (e.g., amplified by a PCR-based method). In some embodiments, sequencing methods involve amplification of the nucleic acid library. The nucleic acid library can be amplified before or after immobilization on a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification involves a process that amplifies or increases the number of nucleic acid templates and / or their complements present (e.g., in a nucleic acid library) by generating one or more copies of the templates and / or their complements. Amplification can be carried out by any suitable method. The nucleic acid library can be amplified by thermocycling or isothermal amplification. In some embodiments, rolling circle amplification is used. In some embodiments, amplification occurs on a solid support (e.g., inside a flow cell) to which the nucleic acid library, or portions thereof, are immobilized. In certain sequencing methods, the nucleic acid library is added to a flow cell and immobilized by hybridization to anchors under appropriate conditions. This type of nucleic acid amplification is often referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or part of the amplification products are synthesized by extension initiated from immobilized primers. Solid-phase amplification reactions are similar to standard solution-phase amplification, except that at least one of the amplification oligonucleotides (e.g., primers) is immobilized on a solid support. In some embodiments, modified nucleic acids (e.g., nucleic acids modified by the addition of adapters) are amplified.
[0153] In some embodiments, solid-phase amplification includes nucleic acid amplification reactions that include only one type of oligonucleotide primer immobilized on a surface. In certain embodiments, solid-phase amplification includes multiple different immobilized oligonucleotide primer species. In some embodiments, solid-phase amplification can include nucleic acid amplification reactions that include one type of oligonucleotide primer immobilized on a solid surface and a second, different oligonucleotide primer species in solution. Multiple different species of immobilized or solution-based primers can be used. Non-limiting examples of solid-phase nucleic acid amplification reactions include interface amplification, bridge amplification, emulsion PCR, WildFire amplification (e.g., U.S. Patent Application Publication No. 2013 / 0012399), etc., or combinations thereof.
[0154] nucleic acid capture In some embodiments, the sample nucleic acid (or sample nucleic acid library) is subjected to a target capture process. Generally, the target capture process is carried out by contacting the sample nucleic acid (or sample nucleic acid library) with a set of probe oligonucleotides under hybridization conditions. The set of probe oligonucleotides (e.g., capture oligonucleotides) generally comprises a plurality of probe oligonucleotides having sequences complementary or substantially complementary to sequences in the sample nucleic acid. The plurality of probe oligonucleotides may comprise about 10 probe oligonucleotide species, about 50 probe oligonucleotide species, about 100 probe oligonucleotide species, about 500 probe oligonucleotide species, about 1,000 probe oligonucleotide species, 2,000 probe oligonucleotide species, 3,000 probe oligonucleotide species, 4,000 probe oligonucleotide species, 5,000 probe oligonucleotide species, 10,000 probe oligonucleotide species, or more. Generally, a first probe oligonucleotide species has a different nucleotide sequence from a second probe oligonucleotide species, and different species of probe oligonucleotides in the set have different nucleotide sequences.
[0155] A probe oligonucleotide typically comprises a nucleotide sequence capable of hybridizing or annealing to a nucleic acid fragment (e.g., a target fragment) or portion thereof of interest. Probe oligonucleotides may be naturally occurring or synthetic, and may be DNA- or RNA-based. Probe oligonucleotides may, for example, enable specific separation of a target fragment from other fragments in a nucleic acid sample. As used herein, the term "specific" or "specificity" refers to the binding or hybridization of one molecule, such as an oligonucleotide for a target polynucleotide, to another molecule. "Specific" or "specificity" refers to the recognition, contact, and formation of a stable complex between two molecules, compared to substantially less recognition, contact, or complex formation between either of the two molecules with other molecules. As used herein, the terms "annealing" and "hybridizing" refer to the formation of a stable complex between two molecules. The terms "probe," "probe oligonucleotide," "capture probe," "capture oligonucleotide," "capture oligo," "oligo," or "oligonucleotide" can be used interchangeably throughout this document when referring to a probe oligonucleotide.
[0156] Probe oligonucleotides can be designed and synthesized using suitable processes and can be of any length suitable for hybridizing with a nucleotide sequence of interest and for carrying out the separation and / or analysis processes described herein. Oligonucleotides can be designed based on the nucleotide sequence of interest (e.g., a target fragment sequence, a genomic sequence, or a gene sequence). Oligonucleotides (e.g., probe oligonucleotides) can, in some embodiments, be about 10 to about 300 nucleotides, about 50 to about 200 nucleotides, about 75 to about 150 nucleotides, about 110 to about 130 nucleotides, or about 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, or 129 nucleotides in length. Oligonucleotides can be composed of naturally occurring and / or non-naturally occurring nucleotides (e.g., labeled nucleotides), or mixtures thereof. Known techniques can be used to synthesize and label oligonucleotides suitable for use with the embodiments described herein.Oligonucleotides can be chemically synthesized using an automated synthesizer according to the solid-phase phosphoramidite triester method first described by Beaucage and Caruthers (1981) Tetrahedron Letts. 22:1859-1862, and / or as described in Needham-VanDevanter et al. (1984) Nucleic Acids Res. 12:6159-6168.Oligonucleotide purification can be achieved by native acrylamide gel electrophoresis or anion-exchange high-performance liquid chromatography (HPLC), for example, as described in Pearson and Regnier (1983) J. Chrom. 255:137-149.
[0157] In some embodiments, all or a portion of the probe oligonucleotide sequence (naturally occurring or synthetic) can be substantially complementary to the target sequence or a portion thereof. As referred to herein, "substantially complementary" with respect to sequences refers to nucleotide sequences that hybridize with each other. The stringency of hybridization conditions can be varied to allow for varying amounts of sequence mismatch. 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more of each other or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more complementary.
[0158] A probe oligonucleotide that is substantially complementary to a nucleotide sequence of interest (e.g., a target sequence) or a portion thereof is also substantially similar to the complement of the target sequence or a relevant portion thereof (e.g., substantially similar to the antisense strand of a nucleic acid). One test for determining whether two nucleotide sequences are substantially similar is to determine the percent of identical nucleotide sequences that are shared. As used herein, "substantially similar" in reference to sequences means 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, "A" refers to a nucleotide sequence that is 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more identical to a nucleotide sequence.
[0159] Hybridization conditions (e.g., annealing conditions) can be determined and / or adjusted depending on the characteristics of the oligonucleotides used in the assay. The sequence and / or length of the oligonucleotide can sometimes affect hybridization with the nucleic acid sequence of interest. Annealing can be achieved using low, medium, or high stringency conditions, depending on the degree of mismatch between the oligonucleotide and the nucleic acid of interest. As used herein, the term "stringent conditions" refers to hybridization and washing conditions. Methods for optimizing the temperature conditions of hybridization reactions are known in the art and can be found in Current Protocols in Molecular Biology, John Wiley & Sons, NY, 6.3.1-6.3.6 (1989). Aqueous and non-aqueous methods are described in the references, and either method can be used. A non-limiting example of stringent hybridization conditions is hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% SDS at 50° C. Another example of stringent hybridization conditions is hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% SDS at 55° C. A further example of stringent hybridization conditions is hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% SDS at 60° C. Stringent hybridization conditions are often hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45° C., followed by one or more washes in 0.2×SSC, 0.1% SDS at 65° C. More often, stringent conditions are 0.5 M sodium phosphate, 7% SDS at 65° C., followed by one or more washes in 0.2×SSC, 1% SDS at 65° C.Stringent hybridization temperatures can also be altered (i.e., lowered) using, for example, the addition of certain organic solvents, such as formamide, which reduce the thermal stability of double-stranded polynucleotides, thereby allowing hybridizations to be performed at lower temperatures while maintaining stringent conditions, extending the useful life of potentially thermolabile nucleic acids.
[0160] In some embodiments, one or more probe oligonucleotides are associated with an affinity ligand or antigen, such as a member of a binding pair (e.g., biotin), which can bind to a capture agent, such as avidin, streptavidin, an antibody, or a receptor. For example, the probe oligonucleotides may be biotinylated so that they can be captured on streptavidin-coated beads.
[0161] In some embodiments, one or more probe oligonucleotides and / or capture agents are operatively linked to a solid support or substrate, which may be any physically separable solid to which probe oligonucleotides are directly or indirectly attached, including, but not limited to, microarrays and wells, and surfaces provided by particles, such as beads (e.g., paramagnetic beads, magnetic beads, microbeads, nanobeads), microparticles, and nanoparticles. Solid supports also include, for example, chips, columns, optical fibers, wipes, filters (e.g., flat surface filters), one or more capillaries, glass and modified or functionalized glass (e.g., controlled-pore glass (CPG)), quartz, mica, diazotized membranes (paper or nylon), polyformaldehyde, cellulose, cellulose acetate, paper, ceramics, metals, metalloids, semiconductor materials, quantum dots, coated beads or particles, other chromatographic materials, magnetic particles, plastics (acrylic, polystyrene, copolymers of styrene or other materials, polybutylene, polyurethane, TEFLON®, polyethylene, polypropylene, polyamide , polyester, polyvinylidene difluoride (PVDF), etc.), polysaccharides, nylon or nitrocellulose, resins, silica or silica-based materials including silicon, silica gel and modified silicon, Sephadex®, Sepharose®, carbon, metals (e.g., steel, gold, silver, aluminum, silicon, and copper), inorganic glass, conducting polymers (including polymers such as polypyrrole and polyindole), microstructured or nanostructured surfaces such as nucleic acid tiling arrays, nanotube, nanowire or nanoparticle decorated surfaces, or porous surfaces or gels such as methacrylates, acrylamides, sugar polymers, cellulose, silicates, or other fibrous or stranded polymers.In some embodiments, the solid support or substrate may be coated using a passive or chemically derivatized coating with any number of materials, including polymers such as dextran, acrylamide, gelatin, or agarose. The beads and / or particles may be free or associated (e.g., sintered) with one another. In some embodiments, the solid phase may be a collection of particles. In some embodiments, the particles may comprise silica, which may comprise silicon dioxide. In some embodiments, the silica may be porous, and in certain embodiments, the silica may be non-porous. In some embodiments, the particles further comprise a substance that imparts paramagnetic properties to the particles. In certain embodiments, the substance comprises a metal, and in certain embodiments, the substance is a metal oxide (e.g., iron or iron oxide, where the iron oxide contains a mixture of Fe2+ and Fe3+). The probe oligonucleotide may be linked to the solid support by covalent or non-covalent interactions, and may be linked directly to the solid support or indirectly (e.g., via an intermediate such as a spacer molecule or biotin). The probe oligonucleotide may be linked to a solid support before, during, or after nucleic acid capture.
[0162] Modified nucleic acids, such as those modified by the addition of an adapter sequence as described herein, can be captured. In some embodiments, unmodified nucleic acids are captured. In some embodiments, nucleic acids may be amplified before and / or after capture by an amplification process such as PCR. The term "captured nucleic acid" generally includes nucleic acids that have been captured, including nucleic acids that have been captured and amplified. In some embodiments, captured nucleic acids can be subjected to additional rounds of capture and amplification. Captured nucleic acids can be sequenced, such as by a sequencing process described herein.
[0163] Nucleic Acid Sequencing and Processing The methods provided herein generally involve nucleic acid sequencing and analysis. In some embodiments, nucleic acids are sequenced, and the sequencing products (e.g., a collection of sequence reads) are processed before or together with the analysis of the sequenced nucleic acids. For example, the sequence reads can be processed according to one or more of the following: aligning, mapping, filtering portions, selecting portions, counting, normalizing, weighting, creating profiles, etc., and combinations thereof. Certain processing steps can be performed in any order, and certain processing steps can be repeated. For example, partial filtering can be performed, followed by normalizing sequence read counts, or in certain embodiments, normalizing sequence read counts can be performed, followed by partial filtering. In some embodiments, the partial filtering step is followed by sequence read count normalization, followed by another partial filtering step. Certain sequencing methods and processing steps are described in more detail below.
[0164] Sequencing In some embodiments, nucleic acids (e.g., nucleic acid fragments, sample nucleic acids, cell-free nucleic acids) are sequenced. In certain instances, complete or substantially complete sequences are obtained, and sometimes partial sequences are obtained. Nucleic acid sequencing generally results in a collection of sequence reads. As used herein, a "read" (e.g., a "read," a "sequence read") is a short nucleotide sequence generated by any sequencing process described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read"), and sometimes is generated from both ends of a nucleic acid fragment (e.g., a double-end read, a two-end read).
[0165] The length of a sequence read is often associated with a particular sequencing technique. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). For example, nanopore sequencing can provide sequence reads that can vary in size from tens to hundreds or thousands of base pairs. In some embodiments, the mean, median, average, or absolute length of the sequence reads is about 15 bp to about 900 bp long. In certain embodiments, the mean, median, average, or absolute length of the sequence reads is about 1000 bp or greater. In some embodiments, the sequence reads are about 1500, 2000, 2500, 3000, 3500, 4000, 4500, or 5000 bp or more in mean, median, average, or absolute length. In some embodiments, the sequence reads are about 100 bp to about 200 bp in mean, median, average, or absolute length. In some embodiments, the sequence reads are of an average, median, mean, or absolute length of about 140 bp to about 160 bp. For example, the sequence reads can be of an average, median, mean, or absolute length of about 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, or 160 bp.
[0166] In some embodiments, the nominal, average, mean, or absolute length of a read from a single end is sometimes from about 10 contiguous nucleotides to about 250 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 200 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 150 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 125 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 100 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 75 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 60 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 50 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 40 or more contiguous nucleotides, and sometimes about 15 contiguous nucleotides, or about 36 or more contiguous nucleotides. In certain embodiments, the nominal, average, mean, or absolute length of a read from a single end is about 20 to about 30 bases long, or about 24 to about 28 bases long. In certain embodiments, the nominal, average, mean, or absolute length of a read from a single end is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, or about 29 bases long or more. In certain embodiments, the nominal, average, mean, or absolute length of a read from a single end is about 20 to about 200 bases, about 100 to about 200 bases, or about 140 to about 160 bases long. In certain embodiments, the nominal, average, mean or absolute length of a read from a single end is about 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190 or about 200 bases or more in length.In certain embodiments, the nominal, average, mean, or absolute length of the reads from both ends is, optionally, from about 10 contiguous nucleotides to about 25 contiguous nucleotides or more (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length or more), from about 15 contiguous nucleotides to about 20 contiguous nucleotides or more, and optionally about 17 contiguous nucleotides, or about 18 contiguous nucleotides. In certain embodiments, the nominal, average, mean, or absolute length of the reads read from both ends is from about 25 contiguous nucleotides to about 400 contiguous nucleotides or more (e.g., about 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, or 400 nucleotides in length or more). , about 50 contiguous nucleotides to about 350 or more contiguous nucleotides, about 100 contiguous nucleotides to about 325 contiguous nucleotides, about 150 contiguous nucleotides to about 325 contiguous nucleotides, about 200 contiguous nucleotides to about 325 contiguous nucleotides, about 275 contiguous nucleotides to about 310 contiguous nucleotides, about 100 contiguous nucleotides to about 200 contiguous nucleotides, about 100 contiguous nucleotides to about 175 contiguous nucleotides, about 125 contiguous nucleotides to about 175 contiguous nucleotides, and optionally about 140 contiguous nucleotides to about 160 contiguous nucleotides. In certain embodiments, the nominal, average, mean, or absolute length of the reads read from both ends is about 150 contiguous nucleotides, and optionally 150 contiguous nucleotides.
[0167] In some embodiments, the nucleotide sequence read obtained from a sample is a partial nucleotide sequence read. As used herein, "partial nucleotide sequence read" refers to a sequence read of any length with incomplete sequence information, also referred to as sequence ambiguity. A partial nucleotide sequence read may lack information about nucleobase identity and / or nucleobase position or order. A partial nucleotide sequence read generally contains only incomplete sequence information (or less than all of the bases are sequenced or determined), but does not include sequence reads due to accidental or unintentional sequencing errors. Such sequencing errors may not be specific to a particular sequencing process, for example, inaccurate calls about nucleobase identity and missing or extra nucleobases. Therefore, in the context of a partial nucleotide sequence read herein, certain information about the sequence is often intentionally excluded. That is, sequence information is intentionally obtained for less than all of the nucleobases, which would otherwise be characterized as or may be a sequencing error. In some embodiments, the partial nucleotide sequence reads can span a portion of the nucleic acid fragment. In some embodiments, the partial nucleotide sequence reads can span the entire length of the nucleic acid fragment. Partial nucleotide sequence reads are described, for example, in International Patent Application Publication No. WO2013 / 052907, the entire contents of which, including all text, tables, formulas, and figures, are incorporated herein by reference.
[0168] A read is generally a physical nucleic acid representation of a nucleotide sequence. For example, in a read containing a sequence described as ATGC, "A" represents an adenine nucleotide, "T" represents a thymine nucleotide, "G" represents a guanine nucleotide, and "C" represents a cytosine nucleotide as a physical nucleic acid. A sequence read obtained from a sample from a subject can be a read derived from a mixture of low-abundance and high-abundance nucleic acids. For example, a sequence read obtained from the blood of a cancer patient can be a read derived from a mixture of cancerous and non-cancerous nucleic acids. In another example, a sequence read obtained from the blood of a pregnant female can be a read derived from a mixture of fetal and maternal nucleic acids. A mixture of relatively short reads can be converted into a representation of the genomic nucleic acids present in a subject and / or the genomic nucleic acids present in a tumor or fetus by the process described herein. In certain cases, a mixture of relatively short reads can be converted into a representation of, for example, copy number alteration, genetic variation / genetic alteration, or aneuploidy. In one example, a readout of a mixture of cancerous and non-cancerous nucleic acids can be converted into a representation of a composite chromosome or portion thereof that contains features of one or both of the cancerous and non-cancerous cell chromosomes. In another example, a readout of a mixture of maternal and fetal nucleic acids can be converted into a representation of a composite chromosome or portion thereof that contains features of one or both of the maternal and fetal chromosomes.
[0169] In some cases, circulating cell-free nucleic acid fragments (CCF fragments) obtained from a cancer patient include nucleic acid fragments originating from normal cells (i.e., non-cancerous fragments) and nucleic acid fragments originating from cancer cells (i.e., cancerous fragments). Sequence reads derived from CCF fragments originating from normal cells (i.e., non-cancerous cells) are referred to herein as "non-cancerous reads." Sequence reads derived from CCF fragments originating from cancer cells are referred to herein as "cancer reads." CCF fragments from which non-cancerous reads are obtained are sometimes referred to herein as non-cancerous templates, and CCF fragments from which cancer reads are obtained are sometimes referred to herein as cancer templates.
[0170] In some cases, circulating cell-free nucleic acid fragments (CCF fragments) obtained from a pregnant female include nucleic acid fragments originating from fetal cells (i.e., fetal fragments) and nucleic acid fragments originating from maternal cells (i.e., maternal fragments). Sequence readings derived from CCF fragments originating from a fetus are referred to herein as "fetal readings." Sequence readings derived from CCF fragments originating from the genome of a pregnant female (e.g., mother) carrying a fetus are referred to herein as "maternal readings." The CCF fragments from which fetal readings are obtained are referred to herein as fetal templates, and the CCF fragments from which maternal readings are obtained are referred to herein as maternal templates.
[0171] In certain embodiments, "obtaining" a nucleic acid sequence read of a sample obtained from a subject and / or "obtaining" a nucleic acid sequence read of a biological specimen obtained from one or more reference individuals can include directly sequencing the nucleic acid to obtain the sequence information. In some embodiments, "obtaining" can include receiving sequence information obtained directly from the nucleic acid by another person.
[0172] In some embodiments, some or all of the nucleic acids in the sample are enriched and / or amplified (e.g., non-specifically, e.g., by PCR-based methods) before or during sequencing. In certain embodiments, specific nucleic acid species or subsets in the sample are enriched and / or amplified before or during sequencing. In some embodiments, species or subsets of a preselected pool of nucleic acids are randomly sequenced. In some embodiments, nucleic acids in the sample are not enriched and / or amplified before or during sequencing.
[0173] In some embodiments, a representative fraction of the genome is sequenced, sometimes referred to as "coverage" or "coverage factor." For example, 1x coverage indicates that approximately 100% of the nucleotide sequence of the genome is represented by the reads. In some cases, coverage factor is referred to as (and is directly proportional to) "sequencing depth." In some embodiments, "coverage factor" is a term that refers to and compares a previous sequencing run as a reference. For example, a second sequencing run may have half the coverage of the first sequencing run. In some embodiments, genomes are sequenced with redundancy, where a given region of the genome can be covered by two or more reads or overlapping reads (e.g., a "coverage factor" greater than 1, e.g., 2x coverage). In some embodiments, the genome (e.g., the entire genome) is sequenced at about 0.01x to about 100x coverage, about 0.1x to 20x coverage, or about 0.1x to about 1x coverage (e.g., about 0.015, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90x or more coverage). In some embodiments, specific portions of the genome (e.g., genome portions derived from targeted and / or probe-based methods) are sequenced, and the coverage fold value generally refers to the fraction of the specific genome portion sequenced (i.e., the coverage fold value does not refer to the entire genome). In some cases, the specific genome portion is sequenced at 1000x coverage or greater. For example, the specific genome portion may be sequenced at 2000x, 5,000x, 10,000x, 20,000x, 30,000x, 40,000x, or 50,000x coverage. In some embodiments, the sequencing is at about 1,000x to about 100,000x coverage. In some embodiments, the sequencing is at about 10,000x to about 70,000x coverage. In some embodiments, the sequencing is at about 20,000-fold to about 60,000-fold coverage.In some embodiments, the sequencing is at about 30,000-fold to about 50,000-fold coverage.
[0174] In some embodiments, the nucleic acid sample obtained from one individual is sequenced.In certain embodiments, the nucleic acid obtained from each of two or more samples is sequenced, where the sample is obtained from one individual or from different individuals.In certain embodiments, the nucleic acid samples obtained from two or more biological samples are pooled, where each biological sample is obtained from one individual or two or more individuals, and the pooled sample is sequenced.In the latter embodiment, the nucleic acid sample obtained from each biological sample is often identified by one or more unique identifiers.
[0175] In some embodiments, sequencing method utilizes identifiers that allow multiplexing of sequencing reaction in sequencing process.The more unique identifiers there are, for example, the more samples and / or chromosomes that can be detected in sequencing process can be multiplexed.Can use any suitable number of unique identifiers (for example, 4, 8, 12, 24, 48, 96 or more) to perform sequencing process.
[0176] The sequencing process sometimes uses a solid phase, sometimes including a flow cell, onto which nucleic acids from a library can be tethered and through which reagents can flow and contact the tethered nucleic acids. Flow cells sometimes include flow cell lanes, and the use of identifiers can facilitate the analysis of several samples in each lane. Flow cells are often solid supports that can be configured to hold bound analytes and / or allow reagent solutions to pass in an orderly fashion over the bound analytes. Flow cells are often planar, optically transparent, generally millimeter or submillimeter scale, and often contain channels or lanes within which analyte-reagent interactions occur. In some embodiments, the number of samples analyzed in a given lane of a flow cell depends on the number of unique identifiers utilized during library preparation and / or probe design. For example, multiplexing using 12 identifiers allows for the simultaneous analysis of 96 samples in an 8-lane flow cell (e.g., equivalent to the number of wells in a 96-well microwell plate). Similarly, multiplexing using, for example, 48 identifiers allows for the simultaneous analysis of 384 samples in an 8-lane flow cell (e.g., equivalent to the number of wells in a 384-well microwell plate). Non-limiting examples of commercially available multiplex sequencing kits include Illumina's Multiplexed Sample Preparation Oligonucleotide Kit, and Multiplexed Sequencing Primer and PhiX Control Kit (e.g., Illumina catalog numbers PE-400-1001 and PE-400-1002, respectively).
[0177] Any suitable method for sequencing nucleic acids can be used, including, but not limited to, Maxim & Gilbert, chain termination, sequencing by synthesis, sequencing by ligation, mass spectrometry sequencing, microscopy-based techniques, and the like, or a combination thereof. In some embodiments, the methods provided herein can use first-generation techniques, such as Sanger sequencing (including automated Sanger sequencing, including microfluidic Sanger sequencing). In some embodiments, sequencing techniques can be used that involve the use of nucleic acid imaging techniques (e.g., transmission electron microscopy (TEM) and atomic force microscopy (AFM)). In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods generally involve clonal amplification of DNA templates or single DNA molecules, and sequencing of these templates or molecules in a massively parallel manner, sometimes inside a flow cell. Next-generation (e.g., second- and third-generation) sequencing techniques capable of massively parallel DNA sequencing can be used for the methods described herein, and are collectively referred to herein as "massively parallel sequencing" (MPS). In some embodiments, MPS sequencing methods utilize a targeted approach, in which specific chromosomes, genes, or regions of interest are sequenced. In certain embodiments, a non-targeted approach is used, in which most or all nucleic acids in a sample are randomly sequenced, amplified, and / or captured.
[0178] In some embodiments, targeted approaches for enrichment, amplification, and / or sequencing are used. Targeting approaches often involve isolating, selecting, and / or enriching a subset of nucleic acids in a sample for further processing using sequence-specific oligonucleotides. In some embodiments, a library of sequence-specific oligonucleotides is utilized to target (e.g., hybridize to) one or more sets of nucleic acids in a sample. Often, the sequence-specific oligonucleotides and / or primers are selective for specific sequences (e.g., unique nucleic acid sequences) present in one or more chromosomes, genes, exons, introns, and / or regulatory regions of interest. Any suitable method or combination of methods can be used to enrich, amplify, and / or sequence one or more subsets of targeted nucleic acids. In some embodiments, the targeted sequences are isolated and / or enriched by capturing them on a solid phase (e.g., a flow cell, beads) using one or more sequence-specific anchors. In some embodiments, targeted sequences are enriched and / or amplified by polymerase-based methods (e.g., PCR-based methods with any suitable polymerase-based extension) using sequence-specific primers and / or primer sets. Sequence-specific anchors can often be used as sequence-specific primers.
[0179] MPS sequencing sometimes uses sequencing by synthesis and specific visualization process.The nucleic acid sequencing technology that can be used in the method described herein is sequencing by synthesis and reversible chain-terminating nucleotide-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ2000; HISEQ2500 (Illumina, San Diego CA)).This technology can perform parallel sequencing on millions of nucleic acid (e.g., DNA) fragments.One example of this type of sequencing technology uses a flow cell containing an optically transparent slide with eight individual lanes, and oligonucleotide anchors (e.g., adapter primers) are attached to their surfaces.
[0180] Sequencing by synthesis is generally performed by the iterative addition (e.g., by covalent addition) of nucleotides to a primer or an existing nucleic acid strand in a template-guided manner. After each iterative nucleotide addition, detection is performed, and this process is repeated multiple times until the sequence of the nucleic acid strand is obtained. The length of the resulting sequence depends, in part, on the number of addition and detection steps performed. In some sequencing-by-synthesis embodiments, one, two, three, or more nucleotides of the same type (e.g., A, G, C, or T) are added and detected in a single nucleotide addition. Nucleotides can be added by any suitable (e.g., enzymatic or chemical) method. For example, in some embodiments, a polymerase or ligase adds nucleotides to a primer or an existing nucleic acid strand in a template-guided manner. Some sequencing-by-synthesis embodiments use different types of nucleotides, nucleotide analogs, and / or identifiers. In some embodiments, reversible chain-terminating nucleotides and / or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In certain embodiments, sequencing by synthesis includes cleavage (e.g., cleavage and removal of the identifier) and / or a washing step. In some embodiments, the addition of one or more nucleotides is detected by a suitable method described herein or known in the art, non-limiting examples of which include any suitable imaging device, suitable camera, digital camera, CCD (charge coupled device)-based imaging device (e.g., CCD camera), CMOS (complementary metal oxide silicon)-based imaging device (e.g., CMOS camera), photodiode (e.g., photomultiplier tube), electron microscopy, field effect transistor (e.g., DNA field effect transistor), ISFET ion sensor (e.g., CHEMFET sensor), etc., or combinations thereof.
[0181] Any MPS method, system or technology platform suitable for the practice described herein can be used to obtain nucleic acid sequencing reads. Non-limiting examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, Helicos True Single Molecule Sequencing, Ion Torrent and ion semiconductor-based sequencing (e.g., developed by Life Technologies), WildFire, 5500, 5500xl W and / or 5500xl W Genetic Analyzer-based technology (e.g., developed and sold by Life Technologies, U.S. Patent Application Publication No. 2013 / 0012399); polony sequencing, pyrosequencing, massively parallel signature sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen systems and methods, nanopore-based platforms, chemically sensitive field effect transistor (CHEMFET) arrays, electron microscopy-based sequencing (e.g., ZS Genetics, Halcyon Other sequencing methods that may be used to practice the methods herein include digital PCR, sequencing by hybridization, nanopore sequencing, chromosome-specific sequencing (e.g., using DANSR (Digital Analysis of Selected Regions)) techniques.
[0182] In some embodiments, the sequence module generates, obtains, collects, accumulates, manipulates, converts, processes, and / or provides sequence reads. A machine including a sequence module can be any suitable machine and / or device that determines the sequence of a nucleic acid using sequencing techniques known in the art. In some embodiments, the sequence module can align, accumulate, fragment, complement, reverse complement, and / or error check (e.g., error correct the sequence reads).
[0183] Read Mapping Sequence reads can be mapped, and the number of reads that map to a particular nucleic acid region (e.g., a chromosome, or a portion thereof) is referred to as a count. Any suitable mapping method (e.g., a process, an algorithm, a program, software, a module, etc., or a combination thereof) can be used. Specific aspects of the mapping process are described below.
[0184] Mapping of nucleotide sequence reads (i.e., sequence information obtained from fragments whose physical location in the genome is unknown) can be performed in several ways, and often involves aligning the obtained sequence reads with matching sequences in a reference genome. In such alignment, the sequence reads are generally aligned to a reference sequence, and the aligned reads are referred to as "mapped" or "mapped sequence reads." In certain embodiments, mapped sequence reads are referred to as "hits" or "counts." In some embodiments, mapped sequence reads are grouped together and assigned to specific genome portions according to various parameters, which are discussed in more detail below.
[0185] The terms "aligned," "alignment," or "aligning" generally refer to two or more nucleic acid sequences that can be identified as identical (e.g., 100% identical) or partially identical. Alignment can be performed manually or by computer (e.g., software, program, module, or algorithm), non-limiting examples of which include the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysis pipeline. Alignment of sequence reads can be 100% sequence identical. In some cases, alignment is less than 100% sequence identical (i.e., incomplete match, partial match, partial alignment). In some embodiments, the alignment is about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% identical. In some embodiments, the alignment includes mismatches. In some embodiments, the alignment includes 1, 2, 3, 4, or 5 mismatches. Two or more sequences can be aligned using either strand (e.g., the sense or antisense strand). In certain embodiments, a nucleic acid sequence is aligned with the reverse complement of another nucleic acid sequence.
[0186] Various computational methods can be used to map each read of a sequence to a portion. Non-limiting examples of computer algorithms that can be used to align sequences include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE1, BOWTIE2, ELAND, MAQ, PROBEMATCH, SOAP, BWA, or SEQMAP, or variations or combinations thereof. In some embodiments, the read of a sequence can be aligned with a sequence in a reference genome. In some embodiments, the read of a sequence can be aligned with a sequence in a reference genome, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Library), or other similar databases. The sequences can be found in and / or aligned with sequences in nucleic acid databases known in the art, including the DNA Databank of Japan (DDBJ) and the DNA Databank of Japan (DDBJ). BLAST or a similar tool can be used to search the identified sequences against sequence databases. The identified sequences can then be sorted into appropriate parts, for example, using the search hits (as described below).
[0187] In some embodiments, reads can be uniquely or non-uniquely mapped to portions in a reference genome. A read is considered "uniquely mapped" if it aligns with a single sequence in the reference genome. A read is considered "non-uniquely mapped" if it aligns with two or more sequences in the reference genome. In some embodiments, non-uniquely mapped reads are excluded from further analysis (e.g., quantification). In certain embodiments, a specific, low degree of mismatch (0-1) may be accounted for as a single nucleotide polymorphism that may exist between the reference genome and the reads obtained from the individual samples being mapped. In some embodiments, no degree of mismatch is allowed for reads mapped to the reference sequence.
[0188] As used herein, the term "reference genome" can refer to any particular known sequenced or characterized genome of any organism or virus, whether a partial sequence or a complete sequence, that can be used to reference identified sequences from a subject. For example, reference genomes for human subjects and many other organisms can be found at the National Center for Biotechnology Information at the World Wide Web URL ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus, expressed as a nucleic acid sequence. As used herein, a reference sequence or reference genome is often a compiled or partially compiled genome sequence obtained from one or more individuals. In some embodiments, a reference genome is a compiled or partially compiled genome sequence obtained from one or more human individuals. In some embodiments, a reference genome includes sequences assigned to chromosomes.
[0189] In certain embodiments, mappability is evaluated for genome region (for example, part, genome part).Mappability is that nucleotide sequence reading can be unambiguously aligned to a part of reference genome, typically with only a certain number of mismatches, including for example, 0, 1, 2 or more mismatches.For a given genome region, a sliding window approach of preset read length can be used, and the obtained read-level mappability values can be averaged to estimate expected mappability.Genome region that contains a stretch of unique nucleotide sequence sometimes has high mappability value.
[0190] For sequencing reads from both ends, reads may be mapped to a reference genome by use of a suitable mapping and / or alignment program, non-limiting examples of which include BWA (Li H. and Durbin R. (2009) Bioinformatics 25, 1754-60), Novoalign [Novocraft (2010)], Bowtie (Langmead B et al. (2009) Genome Biol. 10:R25), SOAP2 (Li R et al. (2009) Bioinformatics 25, 1966-67), BFAST (Homer N et al. (2009) PLoS ONE 4, e7767), GASSST (Rizk, G. and Lavenier, D. (2010) Bioinformatics 26, 2534-2540) and MPscan (Rivals (2009) Lecture Notes in Computer Science 5724, 246-260). Reads from both ends can be mapped and / or aligned using a suitable short read alignment program.Non-limiting examples of short read alignment programs include BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, BWA, CASHX, CUDA-EC, CUSHAW, CUSHAW2, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP, and Geneious. Examples of suitable methods include Assembler, iSAAC, LAST, MAQ, mrFAST, mrsFAST, MOSAIK, MPscan, Novoalign, NovoalignCS, Novocraft, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RTG, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3, SOCS, SSAHA, SSAHA2, Stampy, SToRM, Subread, Subjunc, Taipan, UGENE, VelociMapper, TimeLogic, XpressAlign, ZOOM, etc., or combinations thereof. Reads from both ends are often mapped to opposite ends of the same polynucleotide fragment according to a reference genome. In some embodiments, read mates are independently mapped. In some embodiments, information from both sequence reads (i.e., from each end) is incorporated into the mapping process. Reference genome is often used to determine and / or infer the sequence of the nucleic acid located between the read mates from both ends.As used herein, the term " discordant read pair" refers to the read from both ends that includes a pair of read mates, and one or both read mates cannot be clearly mapped to the same region of reference genome, which is defined by a segment of consecutive nucleotides.In some embodiments, discordant read pair is the read mate from both ends that is mapped to an unexpected position in reference genome.Non-limiting examples of unexpected positions in the reference genome include (i) two different chromosomes, (ii) positions separated by more than a predetermined fragment size (e.g., more than 300 bp, more than 500 bp, more than 1000 bp, more than 5000 bp, or more than 10,000 bp), (iii) orientations that do not match the reference sequence (e.g., opposite orientations), etc., or combinations thereof. In some embodiments, discordant read mates are identified according to the length (e.g., average length, predetermined fragment size) or expected length of the template polynucleotide fragments in the sample. For example, read mates that map to positions that are separated by more than the average or expected length of the polynucleotide fragments in the sample may be identified as discordant read pairs. Read pairs that map in opposite orientations may also be determined by taking the reverse complement of one of the reads and comparing the alignment of both reads using the same strand of the reference sequence. Discordant read pairs can be identified by any suitable method and / or algorithm known in the art or described herein (e.g., SVDetect, Lumpy, BreakDancer, BreakDancerMax, CREST, DELLY, etc., or combinations thereof).
[0191] portion In some embodiments, mapped sequence reads are grouped together according to various parameters and assigned to specific genome portions (e.g., portions of a reference genome). A "portion" may also be referred to herein as a "genome section," "bin," "section," "portion of a reference genome," "portion of a chromosome," or "genome portion."
[0192] The portions are often defined by partitioning the genome according to one or more characteristics. Non-limiting examples of certain partitioning characteristics include length (e.g., fixed length, non-fixed length) and other structural characteristics. Genomic portions sometimes comprise one or more of the following characteristics: fixed length, non-fixed length, random length, non-random length, equal length, unequal length (e.g., at least two of the genome portions are of unequal length), non-overlapping (e.g., the 3' end of a genome portion sometimes abuts the 5' end of an adjacent genome portion), overlapping (e.g., at least two of the genome portions overlap), contiguous, consecutive, non-contiguous, and non-contiguous. The genomic portion is sometimes about 1 to about 1,000 kilobases in length (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900 kilobases in length), about 5 to about 500 kilobases in length, about 10 to about 100 kilobases in length, or about 40 to about 60 kilobases in length.
[0193] The partitioning is sometimes based, or at least partially based, on certain informational features, such as, for example, information content and information gain. Non-limiting examples of certain informational features include alignment speed and / or convenience, variability in sequencing coverage, GC content (e.g., stratified GC content, specific GC content, high or low GC content), GC content heterogeneity, other measures of sequence content (e.g., fraction of individual nucleotides, fraction of pyrimidines or purines, fraction of natural versus non-natural nucleic acids, fraction of methylated nucleotides and CpG content), methylation status, duplex melting temperature, amenability to sequencing or PCR, uncertainty values assigned to individual portions of the reference genome, and / or search results targeting specific features. In some embodiments, information content can be quantified using p-value profiles that measure the significance of specific genomic locations for distinguishing between confirmed normal and confirmed abnormal subjects (e.g., euploid and trisomic subjects, respectively).
[0194] In some embodiments, the genome can be partitioned to eliminate similar regions (e.g., identical or homologous regions or identical or homologous sequences) across the genome and retain only unique regions. The regions excluded during partitioning can be within a single chromosome, one or more chromosomes, or across multiple chromosomes. In some embodiments, the partitioned genome is reduced and optimized for rapid alignment, often allowing for a focus on uniquely identifiable sequences.
[0195] In some embodiments, the genome portions are derived from non-overlapping, fixed-size based partitioning, resulting in contiguous, non-overlapping portions of fixed length. Such portions are often shorter than chromosomes, and often shorter than regions of copy number variation (or copy number alteration) (e.g., duplicated or deleted regions), which may be referred to as segments. A "segment" or "genomic segment" often includes two or more fixed-length genome portions, and often includes two or more contiguous, fixed-length portions (e.g., about 2 to about 100 such portions (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 such portions)).
[0196] Sometimes, multiple parts in a group are analyzed, and sometimes, the reads mapped to the parts are quantified according to a specific group of genome parts. When parts are divided by structural features and correspond to regions in the genome, the parts are sometimes grouped into one or more segments and / or one or more regions. Non-limiting examples of regions include partial chromosomes (i.e., shorter than a chromosome), chromosomes, autosomes, sex chromosomes, and combinations thereof. One or more partial chromosome regions are sometimes genes, gene fragments, regulatory sequences, introns, exons, segments (e.g., segments spanning copy number alteration regions, segments spanning copy number variation regions), microduplications, microdeletions, etc. A region is sometimes smaller than or the same size as a target chromosome, and sometimes smaller than or the same size as a reference chromosome.
[0197] Filtering and / or Selection Parts In some embodiments, the one or more processing steps may include one or more portion filtering and / or portion selection steps. As used herein, the term "filtering" refers to removing a portion or portions of a reference genome from consideration. In certain embodiments, one or more portions are filtered (e.g., subjected to a filtering process), thereby providing filtered portions. In some embodiments, the filtering process removes certain portions, leaving portions (e.g., a subset of portions). After the filtering process, the portions retained are often referred to herein as filtered portions.
[0198] Portions of the reference genome can be selected for removal based on any suitable criteria, including, but not limited to, redundant data (e.g., duplicated or overlapping mapped reads), uninformative data (e.g., portions of the reference genome with a median count of zero), portions of the reference genome with over- or under-represented sequences, noisy data, etc., or combinations of the above. The filtering process often involves removing one or more portions of the reference genome from consideration and subtracting the counts in one or more portions of the reference genome selected for removal from the counts tallied or summed for the reference genome, one or more chromosomes, or portion of the genome under consideration. In some embodiments, portions of the reference genome can be removed sequentially (e.g., removed one by one, allowing for evaluation of the effect of removing each individual portion), and in certain embodiments, all portions of the reference genome marked for removal can be removed simultaneously. In some embodiments, portions of the reference genome characterized by variance above or below a certain level are removed, sometimes referred to herein as filtering out "noisy" portions of the reference genome. In certain embodiments, the filtering process comprises obtaining data points from the dataset that deviate from the mean profile level of the portion, chromosome, or portion of a chromosome by a predetermined multiple of the profile variance, and in certain embodiments, the filtering process comprises removing data points from the dataset that do not deviate from the mean profile level of the portion, chromosome, or portion of a chromosome by a predetermined multiple of the profile variance. In some embodiments, the filtering process is utilized to reduce the number of candidate portions of the reference genome to be analyzed for the presence or absence of genetic variations / alterations and / or copy number alterations (e.g., aneuploidy, microdeletions, microduplications).Reducing the number of candidate portions of the reference genome that are analyzed for the presence or absence of genetic variations / alterations and / or copy number alterations often reduces the complexity and / or dimensionality of the dataset, sometimes increasing the speed of searching for and / or identifying genetic variations / alterations and / or copy number alterations by two orders of magnitude or more.
[0199] The portions can be processed (e.g., filtered and / or selected) by any suitable method and according to any suitable parameters. Non-limiting examples of features and / or parameters that can be used to filter and / or select portions include redundant data (e.g., redundant or overlapping mapped reads), non-informative data (e.g., portions of the reference genome with zero mapped counts), portions of the reference genome with over- or under-represented sequences, noise data, counts, count variability, coverage, mappability, variability, reproducibility measures, read density, read density variability, level of uncertainty, guanine-cytosine (GC) content, CCF fragment length and / or read length (e.g., fragment length ratio (FLR), fetal ratio statistic (FRS)), DNase I sensitivity, methylation status, acetylation, histone distribution, chromatin structure, percent repeats, etc., or combinations thereof. The portions can be filtered and / or selected according to any suitable feature or parameter that correlates with the features or parameters listed or described herein. Portions can be filtered and / or selected according to features or parameters that are specific to the portion (e.g., as determined for a single portion according to multiple samples) and / or that are specific to the sample (e.g., as determined for multiple portions within a sample). In some embodiments, portions are filtered and / or removed according to relatively low mappability, relatively high variability, high level of uncertainty, relatively long CCF fragment length (e.g., low FRS, low FLR), relatively large fraction of repetitive sequences, high GC content, low GC content, low counts, zero counts, high counts, etc., or combinations thereof. In some embodiments, portions (e.g., subsets of portions) are selected according to a suitable level of mappability, variability, level of uncertainty, fraction of repetitive sequences, counts, GC content, etc., or combinations thereof. In some embodiments, portions (e.g., subsets of portions) are selected according to relatively short CCF fragment length (e.g., high FRS, high FLR).The counts and / or reads mapped to the portions are sometimes processed (e.g., normalized) before and / or after filtering or selecting the portions (e.g., subsets of the portions). In some embodiments, the counts and / or reads mapped to the portions are not processed before and / or after filtering or selecting the portions (e.g., subsets of the portions).
[0200] In some embodiments, portions may be filtered according to an error measure (e.g., standard deviation, standard error, calculated variance, p-value, mean absolute error (MAE), mean absolute deviation, and / or mean absolute deviation (MAD)). In particular examples, the error measure may refer to the variability of the counts. In some example embodiments, portions are filtered according to the variability of the counts. In particular embodiments, the variability of the counts is a measure of error determined for counts mapped to a portion (i.e., portion) of the reference genome for multiple samples (e.g., multiple samples obtained from multiple subjects, e.g., 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more, or 10,000 or more subjects). In some embodiments, portions with variability of counts above a predetermined upper range are filtered (e.g., eliminated from consideration). In some embodiments, moieties having count variability below a predetermined lower range are filtered (e.g., eliminated from consideration). In some embodiments, moieties having count variability outside a predetermined range are filtered (e.g., eliminated from consideration). In some embodiments, moieties having count variability within a predetermined range are selected (e.g., used to determine the presence or absence of copy number alterations). In some embodiments, the count variability of the moieties exhibits a distribution (e.g., a normal distribution). In some embodiments, moieties are selected within a quantile of the distribution. In some embodiments, moieties are selected within the 99% quantile of the distribution of count variability.
[0201] Sequence readings from any suitable number of samples can be used to identify a subset of parts that meet one or more criteria, parameters and / or characteristics described herein.Sequence readings from a group of samples from multiple subjects are sometimes used.In some embodiments, the multiple subjects include pregnant females.In some embodiments, the multiple subjects include healthy subjects.In some embodiments, the multiple subjects include cancer patients. One or more samples from each of a plurality of subjects can be handled (e.g., 1 to about 20 samples from each subject (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 samples)), and any suitable number of subjects can be handled (e.g., about 2 to about 10,000 subjects (e.g., about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000 subjects)). In some embodiments, sequence reads from the same test sample(s) from the same subject are mapped to portions in the reference genome and used to generate a subset of portions.
[0202] The portions can be selected and / or filtered by any suitable method. In some embodiments, the portions are selected according to visual inspection of the data, graphs, plots, and / or charts. In certain embodiments, the portions are selected and / or filtered (e.g., in part) by a system or machine including one or more microprocessors and a memory. In some embodiments, the portions are selected and / or filtered (e.g., in part) by a non-transitory computer-readable storage medium having an executable program stored thereon, the program directing a microprocessor to perform the selection and / or filtering.
[0203] In some embodiments, the sequence readings from sample are mapped to all or most of the reference genome, and then a preselected subset of the parts is selected.For example, under a certain length threshold, the subset of the parts to which the readings from fragments are preferentially mapped can be selected.A specific method for preselecting a subset of parts is described in US Patent Application Publication No. 2014 / 0180594, which is incorporated herein by reference.For example, in the further step of determining the presence or absence of genetic variation or genetic alteration, the readings from the selected subset of parts are often utilized.The readings from the parts are often not selected and are not utilized in the further step of determining the presence or absence of genetic variation or genetic alteration (for example, the readings in the unselected parts are removed or filtered).
[0204] In some embodiments, portions associated with the read density (e.g., when the read density is the read density for the portion) are filtered out, and the read density associated with the excluded portion is not included in determining the presence or absence of copy number alterations (e.g., chromosomal aneuploidy, microduplication, microdeletion). In some embodiments, the read density profile comprises and / or consists of the read density of the filtered portion. The portion is optionally filtered according to the distribution of counts and / or the distribution of read densities. In some embodiments, the portion is filtered according to the distribution of counts and / or the read densities when the counts and / or read densities are obtained from one or more reference samples. Herein, the one or more reference samples may be referred to as a training set. In some embodiments, the portion is filtered according to the distribution of counts and / or the read densities when the counts and / or read densities are obtained from one or more test samples. In some embodiments, the portion is filtered according to a measure of uncertainty about the read density distribution. In certain embodiments, portions that support large deviations in read density are filtered out. For example, if each read density in the distribution maps to the same portion, a read density distribution (e.g., a mean read density, an average read density, or a median read density distribution) can be determined. If each portion of the genome is associated with a measure of uncertainty, the read density distribution can be compared across multiple samples to determine a measure of uncertainty (e.g., MAD). Following the previous example, portions can be filtered according to the measure of uncertainty associated with each portion (e.g., standard deviation (SD), MAD) and a predetermined threshold. In certain cases, portions with MAD values within an acceptable range are retained, and portions with MAD values outside the acceptable range are filtered out from consideration. In some embodiments, following the previous example, portions with read density values (e.g., median, average, or mean read density) outside a predetermined measure of uncertainty are often filtered out from consideration.In some embodiments, portions including read density values outside the interquartile range of the distribution (e.g., median, mean, or average read density) are filtered out from consideration. In some embodiments, portions including read density values that fall outside the interquartile range of the distribution by more than 2, 3, 4, or 5 times are filtered out from consideration. In some embodiments, portions including read density values that fall outside the interquartile range of the distribution by more than 2 sigma, 3 sigma, 4 sigma, 5 sigma, 6 sigma, 7 sigma, or 8 sigma (e.g., where sigma is the range defined by the standard deviation) are filtered out from consideration.
[0205] Sequence read quantification In some embodiments, the sequence reads that are mapped or partitioned based on selected features or variables can be quantified to determine the amount or number of reads that map to one or more portions (e.g., portions of a reference genome). In certain embodiments, the amount of sequence reads that map to a portion or segment is referred to as the count or read density.
[0206] Count number is often associated with genome part.In some embodiments, count number is determined from some or all of the sequence readings that are mapped to (i.e., associated with) part.In certain embodiments, count number is determined from some or all of the sequence readings that are mapped to a group of part (for example, the part in segment or region (described herein)).
[0207] The count can be determined by a suitable method, operation, or mathematical process. The count sometimes is a direct sum of all sequence reads that map to a genome portion or group of genome portions corresponding to a segment, a group of portions corresponding to a subregion of the genome (e.g., a region of copy number variation, a region of copy number alteration, a region of copy number duplication, a region of copy number deletion, a region of microduplication, a region of microdeletion, a chromosomal region, an autosomal region, a sex chromosome region), and / or sometimes is a group of portions corresponding to the genome. The read quantification sometimes is a ratio, sometimes a ratio of the quantification of a portion(s) in region a to the quantification of a portion(s) in region b. Region a sometimes is a portion, a segment region, a region of copy number variation, a region of copy number alteration, a region of copy number duplication, a region of copy number deletion, a region of microduplication, a microdeletion, a chromosomal region, an autosomal region, and / or a sex chromosome region. Region b can independently be a portion, a segment, a region of copy number variation, a region of copy number alteration, a region of copy number duplication, a region of copy number deletion, a region of microduplication, a region of microdeletion, a chromosomal region, an autosomal region, a sex chromosome region, a region including all autosomes, a region including sex chromosomes, and / or a region including all chromosomes.
[0208] In some embodiments, the counts are derived from raw sequence reads and / or filtered sequence reads. In certain embodiments, the counts are the average, mean, or sum of the sequence reads mapped to a genome portion or a group of genome portions (e.g., genome portions in a region). In some embodiments, the counts are associated with an uncertainty value. The counts are sometimes adjusted. The counts can be weighted, removed, filtered, normalized, adjusted, averaged, derived as the mean, derived as the median, added, or any combination thereof, adjusted according to the sequence reads associated with the genome portion or group of portions.
[0209] Sequence read quantification is sometimes read density. Read density can be determined and / or generated for one or more segments of a genome. In certain cases, read density can be determined and / or generated for one or more chromosomes. In some embodiments, read density includes a quantitative measure of the number of counts of sequence reads mapped to a segment or portion of a reference genome. Read density can be determined by a suitable process. In some embodiments, read density is determined by a suitable distribution and / or a suitable distribution function. Non-limiting examples of distribution functions include any suitable distribution or combination thereof, such as a probability function, a probability distribution function, a probability density function (PDF), a kernel density function (kernel density estimate), a cumulative distribution function, a probability mass function, a discrete probability distribution, an absolute continuous univariate distribution, etc. Read density can be a density estimate derived from a suitable probability density function. Density estimates are constructs of observed data-based estimates of the underlying probability density function. In some embodiments, read density includes a density estimate (e.g., a probability density estimate, a kernel density estimate). Read density can be generated according to a process that includes generating density estimates for each of one or more parts of a genome, where each part comprises the count number of sequence reads.Read density can be generated for normalized and / or weighted count numbers that are mapped to parts or segments.In some cases, each read that is mapped to a part or segment can contribute to read density, a value (e.g., count number) that is equal to its weight obtained from the normalization process described herein.In some embodiments, read density is adjusted for one or more parts or segments.Read density can be adjusted by suitable methods.For example, the read density for one or more parts can be weighted and / or normalized.
[0210] The reads quantified for a given portion or segment can be from one source or different sources. In one example, the reads can be obtained from nucleic acid derived from a subject with or suspected of having cancer. In such situations, the reads mapped to one or more portions are often representative of both healthy cells (i.e., non-cancerous cells) and cancer cells (e.g., tumor cells). In certain embodiments, some of the reads mapped to a portion are derived from nucleic acid derived from cancerous cells, and some of the reads mapped to the same portion are derived from nucleic acid derived from non-cancerous cells. In another example, the reads can be obtained from a nucleic acid sample derived from a pregnant female carrying a fetus. In such situations, the reads mapped to one or more portions are often reads representing both the fetus and the fetus's mother (e.g., of a pregnant female subject). In certain embodiments, some of the reads mapped to a portion are derived from the fetus's genome, and some of the reads mapped to the same portion are derived from the maternal genome.
[0211] level In some embodiments, a value (e.g., a number, a quantitative value) is assigned to the level. The level can be determined by a suitable method, operation, or mathematical process (e.g., a processed level). The level is often or is derived from the counts (e.g., normalized counts) for a set of portions. In some embodiments, the level of a portion is substantially equal to the total number of counts (e.g., counts, normalized counts) that mapped to the portion. The level is often determined from counts that have been processed, transformed, or manipulated by a suitable method, operation, or mathematical process known in the art. In some embodiments, the level is derived from processed counts, non-limiting examples of processed counts include weighted, filtered, normalized, adjusted, averaged, derived as an average (e.g., a mean level), added, subtracted, transformed counts, or combinations thereof. In some embodiments, the level comprises normalized counts (e.g., normalized counts of a portion). The level may be for a count number normalized by appropriate processing, non-limiting examples of which are described herein. The level may include a normalized count number or a relative amount of a count number. In some embodiments, the level is a level for two or more portions of an average count number or a normalized count number, and the level is referred to as a mean level. In some embodiments, the level is a level for a set of portions having an average count number or an average of normalized count numbers, and this is referred to as a mean level. In some embodiments, the level is derived for portions including raw count numbers and / or filtered count numbers. In some embodiments, the level is based on a count number that is a raw count number. In some embodiments, the level is associated with an uncertainty value (e.g., standard deviation, MAD). In some embodiments, the level is expressed by a Z-score or p-value.
[0212] As used herein, the level of one or more parts is synonymous with the "level of genome division." As used herein, the term "level" is sometimes synonymous with the term "elevation." The meaning of the term "level" can be determined from the context in which it is used. For example, when used in the context of part, profile, reading, and / or count number, the term "level" often means elevation. When used in the context of a substance or composition, the term "level" (e.g., RNA level, plexing level) often refers to quantity. When used in the context of uncertainty (e.g., error level, confidence level, deviation level, uncertainty level), the term "level" often refers to quantity.
[0213] The normalized or non-normalized counts for two or more levels (e.g., levels in two or more profiles) can optionally be mathematically manipulated (e.g., added to, multiplied to, averaged to, normalized to, etc., or combinations thereof) according to the levels. For example, the normalized or non-normalized counts for two or more levels can be normalized according to one, some, or all of the levels in the profile. In some embodiments, the normalized or non-normalized counts for all levels in the profile are normalized according to one level in the profile. In some embodiments, the normalized or non-normalized counts for a first level in the profile are normalized according to the normalized or non-normalized counts for a second level in the profile.
[0214] Non-limiting examples of levels (e.g., first level, second level) are: a level for a set of parts that includes processed counts; a level for a set of parts that includes the mean, median, or average of counts; a level for a set of parts that includes normalized counts, or any combination thereof. In some embodiments, the first level and the second level in the profile are derived from the counts of parts that map to the same chromosome. In some embodiments, the first level and the second level in the profile are derived from the counts of parts that map to different chromosomes.
[0215] In some embodiments, the level is determined from normalized or non-normalized counts mapped to one or more portions. In some embodiments, the level is determined from normalized or non-normalized counts mapped to two or more portions, where the normalized counts in each portion are often approximately the same. There may be variation in the counts (e.g., normalized counts) within a set of portions for a level. Within a set of portions for a level, there may be one or more portions (e.g., peaks and / or dips) where the counts are significantly different from those within the other portions of the set. Any suitable number of normalized or non-normalized counts associated with any suitable number of portions may define a level.
[0216] In some embodiments, one or more levels can be determined from normalized or non-normalized counts of all or a portion of a genome portion. Often, levels can be determined from all or a portion of normalized or non-normalized counts of a chromosome or portion thereof. In some embodiments, the level is determined from two or more counts derived from two or more portions (e.g., a set of portions). In some embodiments, the level is determined from two or more counts (e.g., counts from two or more portions). In some embodiments, the level is determined from counts from 2 to about 100,000 portions. In some embodiments, the level is determined by counts derived from 2 to about 50,000, 2 to about 40,000, 2 to about 30,000, 2 to about 20,000, 2 to about 10,000, 2 to about 5000, 2 to about 2500, 2 to about 1250, 2 to about 1000, 2 to about 500, 2 to about 250, 2 to about 100, or 2 to about 60 fractions. In some embodiments, the level is determined by counts derived from about 10 to about 50 fractions. In some embodiments, the level is determined by counts derived from about 20 to about 40 or more fractions. In some embodiments, a level comprises counts from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60 or more portions. In some embodiments, a level corresponds to a set of portions (e.g., a set of portions of a reference genome, a set of portions of chromosomes, or a set of portions of portions of chromosomes).
[0217] In some embodiments, the level is determined for normalized or non-normalized counts of contiguous portions. In some embodiments, the contiguous portions (e.g., a set of portions) represent adjacent regions of a genome or adjacent regions of a chromosome or gene. For example, two or more contiguous portions may represent a sequence assembly of a DNA sequence longer than each portion when sequenced by integrating the portions end-to-end. For example, two or more contiguous portions may represent an intact genome, a chromosome, a gene, an intron, an exon, or a portion thereof. In some embodiments, the level is determined from a collection (e.g., a set) of contiguous and / or non-contiguous portions.
[0218] Data Processing and Normalization Herein, the mapped sequence reads that have been counted are referred to as raw data because they represent unmanipulated counts (e.g., raw counts). In some embodiments, the sequence read data in a dataset can be further processed (e.g., mathematically and / or statistically manipulated) and / or displayed to facilitate an outcome. In certain embodiments, datasets, including larger datasets, may benefit from preprocessing to facilitate further analysis. Preprocessing of a dataset sometimes includes removing duplicated and / or uninformative portions or portions of the reference genome (e.g., portions of the reference genome with uninformative data, duplicated mapped reads, portions with a median count of zero, over- or under-represented sequences). Without being limited by theory, data processing and / or preprocessing can (i) remove noisy data, (ii) remove uninformative data, (iii) remove redundant data, (iv) reduce the complexity of a larger data set, and / or (v) facilitate the conversion of data from one form to one or more other forms. As used herein, the terms "preprocessing" and "processing" are collectively referred to as "processing" when used in reference to data or data sets. Processing can make data more suitable for further analysis and, in some embodiments, can produce an outcome. In some embodiments, one or more or all of the processing methods (e.g., normalization methods, filtering portions, mapping, validation, etc., or a combination thereof) are performed by a processor, microprocessor, computer in conjunction with memory, and / or by a microprocessor-controlled device.
[0219] The term "noisy data," as used herein, refers to (a) data that, when analyzed or plotted, shows significant variance between data points, (b) data with significant standard deviations (e.g., greater than 3 standard deviations), (c) data with significant standard errors of the mean, and the like, as well as combinations of the above. Noisy data sometimes arises due to the quantity and / or quality of the starting material (e.g., nucleic acid sample), and sometimes arises from part of the process for preparing or replicating the DNA used to obtain the sequence reads. In certain embodiments, noise results from overrepresentation of certain sequences when prepared using PCR-based methods. The methods described herein can reduce or eliminate the contribution of noisy data, thus reducing the effect of noisy data on the resulting outcome.
[0220] The terms "non-informative data," "non-informative portion of the reference genome," and "non-informative portion," as used herein, refer to portions or data derived therefrom having a numerical value significantly different from a predetermined threshold value or a numerical value that falls outside a predefined limit range of values. The terms "threshold" and "threshold value," as used herein, refer to any number calculated using a qualified dataset and serving as a limit for diagnosing a genetic variation or genetic alteration (e.g., copy number alteration, aneuploidy, microduplication, microdeletion, chromosomal abnormality, etc.). In certain embodiments, the results obtained by the methods described herein exceed the threshold value, and the subject is diagnosed with a copy number alteration. In some embodiments, the threshold value or value range is often calculated by mathematically and / or statistically manipulating sequence read data (e.g., obtained from the reference and / or subject); in certain embodiments, the sequence read data manipulated to obtain the threshold value or value range are sequence read data (e.g., obtained from the reference and / or subject). In some embodiments, an uncertainty value is determined. The uncertainty value is generally a measure of variance or error and may be any suitable measure of variance or error. In some embodiments, the uncertainty value is a standard deviation, a standard error, a calculated variance, a p-value, or a mean absolute deviation (MAD). In some embodiments, the uncertainty value may be calculated according to the formulas described herein.
[0221] Any suitable procedure can be used to process the datasets described herein.Non-limiting examples of suitable procedures for processing datasets include filtering, normalizing, weighting, peak height monitoring, peak area monitoring, peak edge monitoring, peak level analysis, peak width analysis, peak edge position analysis, peak lateral tolerance, determining area ratio, mathematically processing data, statistically processing data, applying statistical algorithms, analyzing with a fixed variable, analyzing with an optimized variable, plotting data, identifying patterns or trends, and further processing, etc., and combinations thereof.In some embodiments, datasets are processed based on various features (e.g., GC content, overlapping, mapped reads, centromeric regions, telomeric regions, etc., and combinations thereof) and / or variables (e.g., subject sex, subject age, subject ploidy, percent contribution of cancer cell nucleic acid, fetal sex, maternal age, maternal ploidy, percent contribution of fetal nucleic acid, etc., or combinations thereof). In certain embodiments, the complexity and / or dimensionality of large and / or complex datasets can be reduced by processing the dataset as described herein.Non-limiting examples of complex datasets include sequence read data generated from one or more test subjects and multiple reference subjects of different age and ethnic backgrounds.In some embodiments, the dataset can include thousands to millions of sequence reads for each test subject and / or reference subject.
[0222] In certain embodiments, data processing can be performed in any number of steps. For example, in some embodiments, data can be processed using only a single processing procedure, and in certain embodiments, data can be processed using one or more, five or more, ten or more, or twenty or more processing steps (e.g., one or more processing steps, two or more processing steps, three or more processing steps, four or more processing steps, five or more processing steps, six or more processing steps, seven or more processing steps, eight or more processing steps, nine or more processing steps, ten or more processing steps, eleven or more processing steps, twelve or more processing steps, thirteen or more processing steps, fourteen or more processing steps, fifteen or more processing steps, sixteen or more processing steps, seventeen or more processing steps, eighteen or more processing steps, nineteen or more processing steps, or twenty or more processing steps). In some embodiments, the processing step can be the same step repeated two or more times (e.g., filtering two or more times, normalizing two or more times), while in certain embodiments, the processing step can be two or more different processing steps performed simultaneously or sequentially (e.g., filtering, normalizing; normalizing, monitoring peak height and edges; filtering, normalizing, normalizing to a reference, statistically manipulating, determining p-values, etc.). In some embodiments, any suitable number and / or combination of the same or different processing steps can be utilized to process sequence read data to facilitate obtaining an outcome. In certain embodiments, processing a dataset according to the criteria described herein can reduce the complexity and / or dimensionality of the dataset.
[0223] In some embodiments, one or more processing steps may include one or more normalization steps. Normalization may be performed by any suitable method described herein or known in the art. In certain embodiments, normalization involves adjusting values measured on different scales to a conceptually common scale. In certain embodiments, normalization involves sophisticated mathematical adjustments to bring the probability distributions of the adjusted values into alignment. In some embodiments, normalization involves fitting the distributions to a normal distribution. In certain embodiments, normalization involves mathematical adjustments that allow for comparison of corresponding normalized values for different data sets in a manner that eliminates the effects of certain global influences (e.g., errors and anomalies). In certain embodiments, normalization involves scaling. Normalization sometimes involves division of one or more data sets by a predetermined variable or formula. Normalization sometimes involves division of one or more data sets by a predetermined variable or formula. Non-limiting examples of normalization methods include sectional normalization, GC content normalization, median count (median bin count, median sectional count) normalization, linear least squares regression and non-linear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted scatterplot flattening), principal component normalization, repeat masking (RM), GC normalized repeat masking (GCRM), cQn, and / or combinations thereof. In some embodiments, determining the presence or absence of copy number alteration (for example, aneuploidy, microduplication, microdeletion) utilizes normalization methods (for example, normalization for parts, normalization by GC content, normalization by median count (median bin count, median part count), linear least squares regression and non-linear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted scatterplot flattening), principal component normalization, repeat masking (RM), GC normalized repeat masking (GCRM), cQn, normalization methods known in the art, and / or combinations thereof).For example, certain examples of available normalization processes, such as LOESS normalization, principal component normalization and hybrid normalization methods, are described in more detail below.Certain normalization process aspects are also described, for example, in International Patent Application Publication Nos. WO2013 / 052913 and WO2015 / 051163, each of which is incorporated herein by reference.
[0224] Any suitable number of normalizations can be used. In some embodiments, a dataset can be normalized one or more times, five or more times, ten or more times, or even twenty or more times. A dataset can be normalized to a value (e.g., normalization value) that represents any suitable feature or variable (e.g., sample data, reference data, or both). Non-limiting examples of the types of data normalization that can be used include: normalizing the raw count data for one or more selected test or reference portions to the total number of counts that are mapped to the chromosome or whole genome to which the selected portion or section is mapped; normalizing the raw count data for one or more selected portions to the median of the reference counts for one or more portions or chromosomes to which the selected portion is mapped; normalizing the raw count data to pre-normalized data or their derivatives; and normalizing the pre-normalized data to one or more other predetermined normalization variables. Normalizing a dataset sometimes has the effect of isolating statistical errors, depending on the feature or characteristic selected as the predetermined normalization variable. Also, normalizing a dataset sometimes allows for comparison of data features of data with different scales by giving the data a common scale (e.g., a predetermined normalization variable). In some embodiments, one or more normalizations to a statistically derived value can be used to minimize data differences and reduce the importance of outlier data. Normalizing a portion or a portion of a reference genome with respect to a normalization value is sometimes referred to as "partial normalization."
[0225] In certain embodiments, the processing step may include one or more mathematical and / or statistical operations. Any suitable mathematical and / or statistical operations, alone or in combination, may be used to analyze and / or manipulate the datasets described herein. Any suitable number of mathematical and / or statistical operations may be used. In some embodiments, the dataset may be mathematically and / or statistically manipulated one or more times, five or more times, ten or more times, or twenty or more times. Non-limiting examples of mathematical and statistical operations that may be used include addition, subtraction, multiplication, division, algebraic functions, least squares estimators, curve fitting, differential equations, rational polynomials, double polynomials, orthogonal polynomials, z-scores, p-values, chi values, phi values, analyzing peak levels, determining the location of peak edges, calculating peak area ratios, analyzing chromosome-level medians, calculating the mean absolute deviation, sum of squared residuals, mean, standard deviation, standard error, etc., or combinations thereof. Mathematical and / or statistical manipulations can be performed on all or a portion of the sequence read data or their processed products. Non-limiting examples of variables or features of a dataset that can be statistically manipulated include raw counts, filtered counts, normalized counts, peak height, peak width, peak area, peak edge, lateral tolerance, P-value, median level, mean level, distribution of counts within a genomic region, relative representation of nucleic acid species, etc., or combinations thereof.
[0226] In some embodiments, the processing step can include the use of one or more statistical algorithms. Any suitable statistical algorithm can be used alone or in combination to analyze and / or manipulate the datasets described herein. Any suitable number of statistical algorithms can be used. In some embodiments, one or more, five or more, ten or more, or twenty or more statistical algorithms can be used to analyze the dataset. Non-limiting examples of statistical algorithms suitable for use with the methods described herein include principal component analysis, decision trees, alternative hypothesis, multiple comparisons, omnibus tests, Behrens-Fisher tests, bootstrapping, Fisher's method for combining independence tests of significance, null hypothesis, type I error, type II error, exact test, one-sample Z-test, two-sample Z-test, one-sample t-test, paired t-test, two-sample pooled t-test with equal variances, two-sample unpooled t-test with unequal variances, one proportion z-test, two-proportion z-test pooled, two-proportion z-test unpooled, one-sample chi-square test, two-sample F-test for homogeneity of variances, confidence interval, credible interval, significance, meta-analysis, simple regression, robust linear regression, etc., or a combination of the above. Non-limiting examples of variables or features of a dataset that can be analyzed using statistical algorithms include raw counts, filtered counts, normalized counts, peak height, peak width, peak edge, lateral tolerance, P-value, median level, mean level, distribution of counts within a genomic region, relative representation of nucleic acid species, and the like, or combinations thereof.
[0227] In certain embodiments, a dataset can be analyzed by utilizing multiple (e.g., two or more) statistical algorithms (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, bagging, neural networks, support vector machine models, random forests, classification tree models, K-nearest neighbors, logistic regression, and / or smoothing) and / or mathematical and / or statistical operations (e.g., referred to herein as operations). In some embodiments, the use of multiple operations can generate an N-dimensional space that can be used to generate an outcome. In certain embodiments, analyzing a dataset by utilizing multiple operations can reduce the complexity and / or dimensionality of the dataset. For example, multiple operations can be used on a reference dataset to generate an N-dimensional space (e.g., a probability plot) that can be used to display the presence or absence of genetic variations / alterations and / or copy number alterations depending on the status of the reference sample (e.g., positive or negative for selected genetic variation copy number alterations). Analysis of test samples using a substantially similar set of operations can be used to generate N-dimensional points for each of the test samples. The complexity and / or dimensionality of the test subject data set is sometimes simplified to a single value or N-dimensional point that can be easily compared with the N-dimensional space generated from the reference data. The test sample data that belongs to the N-dimensional space in which the reference subject's data exists exhibits a genetic status that is substantially similar to the genetic status of the reference subject. The test sample data that exists outside the N-dimensional space in which the reference subject's data exists exhibits a genetic status that is substantially dissimilar to the genetic status of the reference subject. In some embodiments, the reference is euploid or otherwise does not have genetic variations / genetic alterations and / or copy number alterations and / or medical conditions.
[0228] In some embodiments, after the datasets have been counted, optionally filtered, normalized, and weighted as needed, these processed datasets can be further manipulated by one or more filtering, normalizing, and / or weighting procedures. In certain embodiments, a profile can be generated using datasets that have been further manipulated by one or more filtering, normalizing, and / or weighting procedures. In some embodiments, the complexity and / or dimensionality of a dataset can sometimes be reduced by one or more filtering, normalizing, and / or weighting procedures. An outcome can be provided based on the dataset with reduced complexity and / or dimensionality. In some embodiments, a profile plot of the processed data that has been further manipulated, for example, by weighting, is generated to facilitate classification and / or providing an outcome. For example, an outcome can be provided based on a plot of a profile of the weighted data.
[0229] Filtering or weighting of part can be carried out at one or more suitable points in analysis.For example, before or after mapping sequence readings to the part of reference genome, part can be filtered or weighted.In some embodiments, before or after determining the experimental bias of each genome part, part can be filtered or weighted.In certain embodiments, before or after calculating level, part can be filtered or weighted.
[0230] In some embodiments, after the datasets have been counted, optionally filtered, normalized, and optionally weighted, these processed datasets can be manipulated by one or more mathematical and / or statistical (e.g., statistical functions or algorithms) operations. In certain embodiments, the processed datasets can be further manipulated by calculating Z-scores for one or more selected portions, chromosomes, or portions of chromosomes. In some embodiments, the processed datasets can be further manipulated by calculating P-values. In certain embodiments, the mathematical and / or statistical operations include one or more assumptions regarding ploidy and / or fractions of minor species (e.g., fraction of cancer cell nucleic acids:fetal fraction). In some embodiments, a plot of a profile of the processed data further manipulated by one or more statistical and / or mathematical operations is generated to facilitate classification and / or providing an outcome. An outcome can be provided based on the plot of the profile of the statistically and / or mathematically manipulated data. Outcomes derived based on plots of statistically and / or mathematically manipulated data profiles often include one or more assumptions regarding ploidy and / or fractions of minor species (e.g., fraction of cancer cell nucleic acids:fetal fraction).
[0231] In some embodiments, data analysis and processing can include the use of one or more assumptions.Any number or type of assumptions can be used to analyze or process a data set.Non-limiting examples of assumptions that can be used for data processing and / or analysis include: subject ploidy, cancer cell contribution, maternal ploidy, fetal contribution, the prevalence of specific sequences in a reference population, ethnic background, the prevalence of selected medical conditions in related family members, the parallelism between the raw count profiles obtained from different patients and / or runs after GC normalization repeat masking (e.g., GCRM), identical matches (e.g., identical base positions) that represent PCR artifacts, assumptions specific to nucleic acid quantification assays (e.g., fetal quantification assays (FQA)), assumptions about twins (e.g., if only one of the twins is affected, the effective fetal fraction is only 50% of the total measured fetal fraction (similar for triplets, quadruplets, etc.)), cell-free DNA (e.g., cfDNA) that uniformly covers the entire genome, and combinations thereof.
[0232] In cases where the quality and / or depth of the mapped sequence reads do not allow for the prediction of the outcome of the presence or absence of genetic variation / genetic alteration and / or copy number alteration based on the normalized count profile with a desired level of confidence (e.g., 95% or more confidence level), one or more additional mathematical manipulation algorithms and / or statistical prediction algorithms can be utilized to generate additional numerical values useful for data analysis and / or providing outcomes. The term "normalized count profile" as used herein refers to a profile generated using normalized counts. Examples of methods that can be used to generate normalized counts and normalized count profiles are described herein. As mentioned above, the mapped sequence reads that result in the counts can be normalized with respect to the counts of the test sample or the counts of the reference sample. In some embodiments, the normalized count profile can be plotted and displayed.
[0233] Non-limiting examples of processing steps and normalization methods that can be used, such as normalizing over a window (static or sliding), weighting, determining bias relationships, LOESS normalization, principal component normalization, hybrid normalization, generating profiles and performing comparisons, are described in more detail herein below.
[0234] Normalization to a window (static or sliding) In certain embodiments, the processing step includes normalizing to a stationary window; in some embodiments, the processing step includes normalizing to a moving or sliding window. The term "window," as used herein, refers to one or more portions selected for analysis and is sometimes used as a reference for comparison (e.g., for normalization and / or other mathematical or statistical operations). The term "normalizing to a stationary window," as used herein, refers to a normalization process that uses one or more selected portions to compare a test dataset with a reference dataset. In some embodiments, the selected portions are used to generate a profile. A stationary window generally includes a predetermined set of portions that do not change during manipulation and / or analysis. The terms "normalizing to a moving window" and "normalizing to a sliding window," as used herein, refer to normalization performed to portions (e.g., immediate surrounding portions, adjacent portions, or sections, etc.) localized to a genomic region of the selected test portion, where one or more selected test portions are normalized to the immediate surrounding portions of the selected test portion. In certain embodiments, the selected portions are used to generate a profile. Sliding window or moving window normalization often involves iteratively moving or sliding toward adjacent test portions and normalizing the newly selected test portion to portions immediately surrounding or adjacent to the newly selected test portion, where the adjacent windows have one or more portions in common. In certain embodiments, multiple selected test portions and / or chromosomes can be analyzed using sliding window processing.
[0235] In some embodiments, one or more values can be generated by normalizing over a sliding or moving window, where each value represents the result of normalization over a different set of reference portions selected from different regions (e.g., chromosomes) of the genome. In certain embodiments, the generated one or more values are cumulative sums (e.g., a numerical estimate of the integral of the normalized count profile over the selected portion, domain (e.g., part of a chromosome), or chromosome). The values generated by the sliding or moving window process can be used to generate profiles and facilitate arriving at outcomes. In some embodiments, the cumulative sum of one or more portions can be displayed as a function of genomic location. Sometimes, moving or sliding window analysis is used to analyze a genome for the presence or absence of microdeletions and / or microduplications. In certain embodiments, displaying the cumulative sum of one or more portions is used to identify the presence or absence of regions of copy number alterations (e.g., microdeletions, microduplications).
[0236] weighted In some embodiments, the processing step includes weighting. The terms "weighted," "weighting," or "weighting function," or grammatical derivatives or equivalents thereof, as used herein, refer to a mathematical manipulation of part or all of a dataset that may be utilized to vary the influence of a particular dataset feature or variable relative to other dataset features or variables (e.g., to increase or decrease the significance and / or contribution of data contained in one or more portions or portions of a reference genome based on the quality or usefulness of the data in the selected portion or portions of the reference genome). In some embodiments, a weighting function may be used to increase the influence of data with a relatively small measurement variance and / or decrease the influence of data with a relatively large measurement variance. For example, portions of a reference genome with underrepresented or low-quality sequence data may be "weighted down" to minimize their influence on the dataset, while selected portions of a reference genome may be "weighted up" to increase their influence on the dataset. A non-limiting example of a weighting function is [1 / (standard deviation) 2 ]. Weighting portions sometimes removes portion dependencies. In some embodiments, one or more portions are weighted by an eigen function (e.g., an eigenfunction). In some embodiments, the eigen function includes replacing portions by orthogonal eigen portions. The weighting step sometimes is performed substantially similarly to the normalization step. In some embodiments, the dataset is adjusted (e.g., divided, multiplied, added, subtracted) by a predetermined variable (e.g., a weighting variable). In some embodiments, the dataset is divided by a predetermined variable (e.g., a weighting variable). Often, a predetermined variable (e.g., a minimization objective function, Phi) is selected to weight different parts of the dataset differently (e.g., to increase the influence of certain types of data while decreasing the influence of other types of data).
[0237] bias relationship In some embodiments, the processing step includes determining a bias relationship. For example, one or more relationships can be generated between local genomic bias estimates and bias frequencies. As used herein, the term "relationship" refers to a mathematical and / or graphical relationship between two or more variables or values. The relationship can be generated by appropriate mathematical and / or graphical processing. Non-limiting examples of relationships include mathematical and / or graphical representations of functions, correlations, distributions, linear or nonlinear equations, lines, regressions, fitted regressions, etc., or combinations thereof. In some embodiments, the relationship includes a fitted relationship. In some embodiments, the fitted relationship includes a fitted regression. In some embodiments, the relationship includes two or more variables or values and includes weighted variables or weighted values. In some embodiments, the relationship includes a fitted regression, where one or more variables or values of the relationship are weighted. In some embodiments, the regression is fitted in a weighted manner. In some embodiments, the regression is fitted unweighted. In certain embodiments, generating the relationship includes plotting or graphing.
[0238] In certain embodiments, a relationship is generated between GC density and GC density frequency. In some embodiments, a sample GC density relationship is displayed by generating a relationship between (i) GC density and (ii) GC density frequency for the sample. In some embodiments, a reference GC density relationship is displayed by generating a relationship between (i) GC density and (ii) GC density frequency for the reference. In some embodiments, when the estimate of local genomic bias is GC density, the sample bias relationship is a sample GC density relationship, and the reference bias relationship is a reference GC density relationship. The GC density of the reference GC density relationship and / or the sample GC density relationship is often an indication (e.g., a mathematical or quantitative indication) of the local GC content.
[0239] In some embodiments, the relationship between the local genomic bias estimate and the bias frequency comprises a distribution. In some embodiments, the relationship between the local genomic bias estimate and the bias frequency comprises a fitted relationship (e.g., a fitted regression). In some embodiments, the relationship between the local genomic bias estimate and the bias frequency comprises a linear fitted regression or a non-linear fitted regression (e.g., a polynomial regression). In certain embodiments, the relationship between the local genomic bias estimate and the bias frequency comprises a weighted relationship, wherein the local genomic bias estimate and / or the bias frequency are weighted by a suitable process. In some embodiments, the weighted fitted relationship (e.g., a weighted fit) can be obtained by a process comprising quartile regression, a parameterized probability distribution, or an empirical distribution with interpolation. In certain embodiments, the relationship between the local genomic bias estimate and the bias frequency for the test sample, the reference standard, or a portion thereof comprises a polynomial regression, and the local genomic bias estimate is weighted. In some embodiments, the weighted fit model comprises weighting the distribution values. The distribution value can be weighted by appropriate processing. In some embodiments, the value located near the tail of the distribution is weighted less than the value near the distribution median. For example, for the distribution of local genomic bias estimate (e.g., GC density) and bias frequency (e.g., GC density frequency), weight is determined according to the bias frequency for a given local genomic bias estimate, and the local genomic bias estimate that includes a bias frequency close to the mean of the distribution is weighted more than the local genomic bias estimate that includes a bias frequency far from the mean.
[0240] In some embodiments, the processing step comprises normalizing the sequence read counts by comparing the local genomic bias estimate of the sequence reads of the test sample to the local genomic bias estimate of a reference standard (e.g., a reference genome or portion thereof). In some embodiments, the sequence read counts are normalized by comparing the bias frequency of the local genomic bias estimate of the test sample to the bias frequency of the local genomic bias estimate of the reference standard. In some embodiments, the sequence read counts are normalized by comparing the sample bias relationship to the reference bias relationship, thereby generating a comparison.
[0241] The counts of sequence reads can be normalized according to the comparison of two or more relationships. In certain embodiments, two or more relationships are compared, thereby presenting a comparison that is used to reduce local bias in sequence reads (for example, normalize counts). Two or more relationships can be compared by any suitable method. In some embodiments, the comparison includes adding the second relationship to the first relationship, subtracting the second relationship from the first relationship, multiplying the first relationship by the second relationship, and / or dividing the first relationship by the second relationship. In certain embodiments, the comparison of two or more relationships includes the use of suitable linear regression and / or non-linear regression. In certain embodiments, the comparison of two or more relationships includes suitable polynomial regression (for example, third-order polynomial regression). In some embodiments, the comparison comprises adding a second regression to a first regression, subtracting a second regression from the first regression, multiplying the first regression by the second regression, and / or dividing the first regression by the second regression. In some embodiments, two or more relationships are compared using a process comprising a multiple regression inference framework. In some embodiments, two or more relationships are compared using a process comprising a suitable multivariate analysis. In some embodiments, two or more relationships are compared using a process comprising a basis function (e.g., a blending function, e.g., a polynomial basis, a Fourier basis, etc.), a spline, a radial basis function, and / or a wavelet.
[0242] In certain embodiments, the distribution of local genomic bias estimates, including bias frequencies, for test samples and reference standards are compared by a process comprising polynomial regression, where the local genomic bias estimates are weighted.In some embodiments, a polynomial regression is generated between (i) a ratio, each of which includes the bias frequency of the local genomic bias estimate of the reference standard and the bias frequency of the local genomic bias estimate of the sample, and (ii) the local genomic bias estimates.In some embodiments, a polynomial regression is generated between (i) the ratio of the bias frequency of the local genomic bias estimate of the reference standard to the bias frequency of the local genomic bias estimate of the sample, and (ii) the local genomic bias estimates.In some embodiments, comparing the distribution of local genomic bias estimates for the reads of the test sample and the reference standard comprises determining the logarithmic ratio (e.g., log2 ratio) of the bias frequencies of the local genomic bias estimates for the reference standard and the sample. In some embodiments, comparing the distribution of local genomic bias estimates comprises dividing the log ratio (e.g., log2 ratio) of the bias frequency of the local genomic bias estimates for the reference standard by the log ratio (e.g., log2 ratio) of the bias frequency of the local genomic bias estimates for the sample.
[0243] Normalizing the counts according to the comparison typically involves adjusting some counts but not others. Normalizing the counts sometimes involves adjusting all counts, and sometimes involves not adjusting the counts of any sequence reads. The counts for sequence reads are sometimes normalized by a process that includes determining a weighting factor, and sometimes the process does not involve directly generating and utilizing a weighting factor. Normalizing the counts according to the comparison sometimes involves determining a weighting factor for the counts of each sequence read. The weighting factor is often specific to the sequence read and is applied to the counts of a specific sequence read. The weighting factor is often determined according to a comparison of two or more bias relationships (e.g., a sample bias relationship compared to a reference bias relationship). The normalized counts are often determined by adjusting the count values according to the weighting factor. Adjusting the counts according to the weighting factor may include adding the weighting factor to the counts for the sequence reads, subtracting the weighting factor from the counts for the sequence reads, multiplying the counts for the sequence reads by the weighting factor, and / or dividing the counts for the sequence reads by the weighting factor. The weighting factor and / or normalized counts may be determined from a regression (e.g., a regression line). The normalized counts may be obtained directly from a regression line (e.g., a fitted regression line) obtained as a result of a comparison between the bias frequency of the local genomic bias estimate of the reference standard (e.g., a reference genome) and the bias frequency of the local genomic bias estimate of the test sample. In some embodiments, each count of the sample read is presented as a normalized count value according to a comparison of (i) the bias frequency of the read's local genomic bias estimate compared to (ii) the bias frequency of the local genomic bias estimate of the reference standard. In certain embodiments, the sequence read counts obtained for a sample are normalized to reduce bias in the sequence reads.
[0244] LOESS normalization In some embodiments, the processing step includes LOESS normalization. LOESS is a regression modeling method known in the art that combines multiple regression models in a k-nearest neighbor-based meta-model. LOESS is sometimes referred to as locally weighted polynomial regression. In some embodiments, GC LOESS applies a LOESS model to the relationship between fragment counts (e.g., sequence reads, sequence counts) and GC composition for a reference genome portion. Plotting a smooth curve through a set of data points using LOESS is sometimes referred to as a LOESS curve, particularly when each smoothed value is given by a weighted quadratic least-squares regression over the interval of values of the scatterplot reference variable on the y-axis. For each point in the dataset, the LOESS method fits a low-order polynomial to a subset of data whose explanatory variable values are in the vicinity of the point at which the response is estimated. The polynomial is fitted using a weighted least squares method, which gives greater weight to points near the point whose response is being estimated and less weight to points further away. A regression function value for a point is then obtained by evaluating the local polynomial using the explanatory variable values for that data point. The LOESS fit is sometimes considered complete after a regression function value has been calculated for each of the data points. Many of the details of this method, such as the order and weights of the polynomial model, are adaptive.
[0245] Principal component analysis In some embodiments, the processing step comprises principal component analysis (PCA). In some embodiments, the counts of sequence reads (e.g., the counts of sequence reads of test samples) are adjusted according to principal component analysis (PCA). The read density profile of one or more reference samples and / or the read density profile of the test subject can be adjusted according to PCA. In this specification, removing bias from the read density profile through PCA-related processing is sometimes referred to as adjusting the profile. PCA can be performed by a suitable PCA method or its variants. Non-limiting examples of PCA methods include canonical correlation analysis (CCA), Karhunen-Loeve (KL) transform (KLT), Hotelling transform, proper orthogonal decomposition (POD), singular value decomposition (SVD) of X, eigenvalue decomposition (EVD) of XTX, factor analysis, Eckert-Young theorem, Schmidt-Mirsky theorem, empirical orthogonal functions (EOF), empirical eigenfunction decomposition, empirical component analysis, quasi-harmonic modes, spectral decomposition, empirical mode analysis, and variations or combinations thereof. PCA often identifies and / or adjusts one or more biases in a read density profile. Herein, biases identified and / or adjusted by PCA are sometimes referred to as principal components. In some embodiments, one or more biases can be eliminated by adjusting the read density profile according to one or more principal components using an appropriate method. The read density profile can be adjusted by adding one or more principal components to the read density profile, subtracting one or more principal components from the read density profile, multiplying the read density profile by one or more principal components, and / or dividing the read density profile by one or more principal components. In some embodiments, one or more biases can be removed from the read density profile by subtracting one or more principal components from the read density profile. While biases in a read density profile are often identified and / or quantified by PCA of the profile, principal components are often subtracted from the profile at the level of read density.PCA often identifies one or more principal components. In some embodiments, PCA identifies principal components in the order of first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, and tenth, or more. In certain embodiments, one, two, three, four, five, six, seven, eight, nine, ten, or more principal components are used to adjust the profile. In certain embodiments, five principal components are used to adjust the profile. Principal components are often used to adjust the profile in the order of their appearance in PCA. For example, when three principal components are subtracted from a read density profile, the first, second, and third principal components are used. In some cases, the bias identified by the principal components includes features of the profile that are not used to adjust the profile. For example, PCA identifies copy number alterations (e.g., aneuploidy, microduplication, microdeletion, deletion, translocation, insertion) and / or sex differences as principal components. Thus, in some embodiments, one or more principal components are not used to adjust the profile. For example, in some cases, the first, second, and fourth principal components are used to adjust the profile, where the third principal component is not used to adjust the profile.
[0246] Principal components can be obtained from PCA using any suitable sample or reference standard. In some embodiments, the principal components are obtained from a test sample (e.g., a test subject). In some embodiments, the principal components are obtained from one or more reference standards (e.g., a reference sample, a reference sequence, a reference set). In certain cases, PCA is performed on a median read density profile obtained from a training set including a plurality of samples, which results in the identification of a first principal component and a second principal component. In some embodiments, the principal components are obtained from a set of subjects that lack the copy number alteration in question. In some embodiments, the principal components are obtained from a set of known euploids. Principal components are often identified according to PCA performed using one or more read density profiles of a reference standard (e.g., a training set). One or more principal components obtained from the reference standard are often subtracted from the read density profile of the test subject, thereby presenting an adjusted profile.
[0247] Hybrid Normalization In some embodiments, the processing step includes a hybrid normalization method. In certain cases, the hybrid normalization method can reduce bias (e.g., GC bias). In some embodiments, hybrid normalization involves (i) analyzing the relationship between two variables (e.g., counts and GC content) and (ii) selecting and applying a normalization method according to the analysis. In certain embodiments, hybrid normalization involves (i) regression (e.g., regression analysis) and (ii) selecting and applying a normalization method according to the regression. In some embodiments, the counts obtained for a first sample (e.g., a first set of samples) are normalized by a different method than the counts obtained from another sample (e.g., a second set of samples). In some embodiments, the counts obtained for a first sample (e.g., a first set of samples) are normalized by a first normalization method, and the counts obtained from a second sample (e.g., a second set of samples) are normalized by a second normalization method. For example, in certain embodiments, the first normalization method includes the use of linear regression and the second normalization method includes the use of non-linear regression (e.g., LOESS, GC-LOESS, LOWESS regression, LOESS smoothing).
[0248] In some embodiments, a hybrid normalization method is used to normalize sequence reads (e.g., counts, mapped counts, mapped reads) mapped to portions of the genome or chromosomes. In certain embodiments, raw counts are normalized, and in some embodiments, adjusted, weighted, filtered, or already normalized counts are normalized using the hybrid normalization method. In certain embodiments, levels or Z-scores are normalized. In some embodiments, counts mapped to selected genome portions or chromosomes are normalized using the hybrid normalization method. The counts may refer to an appropriate measure of sequence reads mapped to portions of the genome, non-limiting examples of which include raw counts (e.g., unprocessed counts), normalized counts (e.g., normalized by LOESS, principal components, or an appropriate method), portion levels (e.g., mean levels, average levels, median levels, etc.), Z-scores, etc., or combinations thereof. The counts may be raw counts or processed counts from one or more samples (e.g., test samples, samples from pregnant females). In some embodiments, the counts are obtained from one or more samples obtained from one or more subjects.
[0249] In some embodiments, the normalization method (e.g., type of normalization method) is selected according to regression (e.g., regression analysis) and / or correlation coefficient. Regression analysis refers to a statistical technique for estimating the relationship between variables (e.g., counts and GC content). In some embodiments, the regression is generated according to counts and a measure of GC content for each of the multiple portions of the reference genome. Suitable measures of GC content, non-limiting examples of which include measures of guanine content, cytosine content, adenine content, thymine content, purine (GC) content, or pyrimidine (AT or ATU) content, melting temperature (T m) (e.g., denaturation temperature, annealing temperature, hybridization temperature), free energy measures, etc., or combinations thereof, can be used. Measures of guanine (G) content, cytosine (C) content, adenine (A) content, thymine (T) content, purine (GC) content, or pyrimidine (AT or ATU) content can be expressed as a ratio or percentage. In some embodiments, any suitable ratio or percentage is used, non-limiting examples of which include GC / AT, GC / total nucleotides, GC / A, GC / T, AT / total nucleotides, AT / GC, AT / G, AT / C, G / A, C / A, G / T, G / A, G / AT, C / T, etc., or combinations thereof. In some embodiments, the measure of GC content is a ratio or percentage of GC content to total nucleotide content. In some embodiments, the measure of GC content is a ratio or percentage of GC content to total nucleotide content for sequence reads mapped to portions of a reference genome. In certain embodiments, GC content is determined according to the sequence readings that are mapped to each reference genome portion and / or from the sequence readings that are mapped to each reference genome portion, and sequence readings are obtained from sample.In some embodiments, the measure of GC content is not determined according to the sequence readings and / or from the sequence readings.In certain embodiments, the measure of GC content is determined for one or more samples that are obtained from one or more subjects.
[0250] In some embodiments, generating a regression includes generating a regression analysis or a correlation analysis. Any suitable regression can be used, non-limiting examples of which include regression analysis (e.g., linear regression analysis), goodness-of-fit analysis, Pearson correlation analysis, rank correlation, percent unexplained variance, Nash-Sutcliffe (NS) model efficiency analysis, regression model validation, proportional reduction in loss (PRL), root mean square deviation, etc., or combinations thereof. In some embodiments, a regression line is generated. In certain embodiments, generating a regression includes generating a linear regression. In certain embodiments, generating a regression includes generating a nonlinear regression (e.g., LOESS regression, LOWESS regression).
[0251] In some embodiments, the regression determines the presence or absence of a correlation (e.g., a linear correlation), for example, between counts and a measure of GC content. In some embodiments, a regression (e.g., a linear regression) is generated and a correlation coefficient is determined. In some embodiments, a non-limiting example of this is the coefficient of determination, R 2 Determine appropriate correlation coefficients, including values, Pearson correlation coefficients, etc.
[0252] In some embodiments, the goodness of fit is determined for a regression (e.g., regression analysis, linear regression). The goodness of fit is optionally determined by visual analysis or mathematical analysis. The evaluation optionally includes determining whether the goodness of fit is greater for a nonlinear regression or a linear regression. In some embodiments, a correlation coefficient is a measure of the goodness of fit. In some embodiments, the evaluation of the goodness of fit for a regression is determined according to the correlation coefficient and / or a correlation coefficient cutoff value. In some embodiments, the evaluation of the goodness of fit includes comparing the correlation coefficient to a correlation coefficient cutoff value. In some embodiments, the evaluation of the goodness of fit for a regression is indicative of a linear regression. For example, in certain embodiments, the goodness of fit is greater for a linear regression than for a nonlinear regression, and the evaluation of the goodness of fit is indicative of a linear regression. In some embodiments, the evaluation is indicative of a linear regression, and the linear regression is used to normalize the counts. In some embodiments, the evaluation of the goodness of fit for a regression is indicative of a nonlinear regression. For example, in certain embodiments, the goodness of fit is greater for a nonlinear regression than for a linear regression, and the evaluation of the goodness of fit is indicative of a nonlinear regression. In some embodiments, the evaluation refers to a non-linear regression, and the non-linear regression is used to normalize the counts.
[0253] In some embodiments, the assessment of goodness of fit indicates linear regression when the correlation coefficient is equal to or greater than the correlation coefficient cutoff. In some embodiments, the assessment of goodness of fit indicates nonlinear regression when the correlation coefficient is less than the correlation coefficient cutoff. In some embodiments, the correlation coefficient cutoff is a predetermined cutoff. In some embodiments, the correlation coefficient cutoff is about 0.5 or greater, about 0.55 or greater, about 0.6 or greater, about 0.65 or greater, about 0.7 or greater, about 0.75 or greater, about 0.8 or greater, or about 0.85 or greater.
[0254] In some embodiments, a specific type of regression (e.g., linear or nonlinear regression) is selected, and after generating the regression, the counts are normalized by subtracting the regression from the counts. In some embodiments, subtracting the regression from the counts provides normalized counts with reduced bias (e.g., GC bias). In some embodiments, a linear regression is subtracted from the counts. In some embodiments, a nonlinear regression (e.g., LOESS, GC-LOESS, LOWESS regression) is subtracted from the counts. Any suitable method can be used to subtract the regression line from the counts. For example, counts x are derived from portion i (e.g., portion i) containing a GC content of 0.5, and the regression line determines counts y at a GC content of 0.5, where xy = normalized counts for portion i. In some embodiments, counts are normalized before and / or after subtracting the regression. In some embodiments, counts normalized by a hybrid normalization method are used to generate levels, Z-scores, levels and / or profiles of a genome or portion thereof. In certain embodiments, counts normalized by the hybrid normalization method are analyzed by methods described herein to determine the presence or absence of genetic variations or genetic alterations (e.g., copy number alterations).
[0255] In some embodiments, the hybrid normalization method includes filtering or weighting one or more portions before or after normalization. Appropriate portion filtering methods can be used, including portion (e.g., reference genome portion) filtering methods described herein. In some embodiments, portions (e.g., reference genome portions) are filtered before applying the hybrid normalization method. In some embodiments, only the counts of sequencing reads mapped to selected portions (e.g., portions selected according to the variability of the counts) are normalized by hybrid normalization. In some embodiments, the counts of sequencing reads mapped to filtered reference genome portions (e.g., portions filtered according to the variability of the counts) are excluded before applying the hybrid normalization method. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., reference genome portions) according to an appropriate method (e.g., a method described herein). In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., reference genome portions) according to an uncertainty value for the counts mapped to each of the portions for multiple test samples. In some embodiments, the hybrid normalization method involves selecting or filtering portions (e.g., reference genome ...
Claims
1. 1. A method for assessing the extent of genetic mosaicism in a circulating cell-free nucleic acid genetic screening test in a pregnant female subject, comprising: identifying, by a computing system, regions of gene copy number variation for a sample obtained from the pregnant female subject based on data from the genetic screening test of the circulating cell-free nucleic acid, wherein the genetic screening test is a non-invasive prenatal test (NIPT), the regions of gene copy number variation comprise copy number variation, the sample comprises the circulating cell-free nucleic acid from the pregnant female subject, and the circulating cell-free nucleic acid comprises maternal nucleic acid and fetal nucleic acid; identifying, by the computing system, a positive screening result for the presence of one or more aneuploidies in the sample based on the data from the genetic screening test; determining, by the computing system, the fraction of nucleic acids having the copy number variation in the circulating cell-free nucleic acids; determining, by the computing system, the fraction of fetal nucleic acid in the circulating cell-free nucleic acid; generating, by the computing system, a mosaicism ratio, the mosaicism ratio being the fraction of nucleic acids having the copy number variation in the circulating cell-free nucleic acid divided by the fraction of fetal nucleic acid in the circulating cell-free nucleic acid; The computing system (i) classifies the presence of genetic mosaicism for the region of gene copy number variation if the mosaicism ratio is between 0.1 and 0.7, and (ii) provides the absence or no classification of genetic mosaicism for the region of gene copy number variation if the mosaicism ratio is greater than 0.7 or less than 0.1; providing, if the computing system classifies the presence of genetic mosaicism for the region of gene copy number variation, interpreting the positive screening result from the genetic screening test as positive with a comment regarding the likelihood of mosaicism present; A method comprising:
2. 2. The method of claim 1, further comprising the step of classifying, by the computing system, the absence of genetic mosaicism for the region of gene copy number variation if the mosaicism ratio is greater than 0.
7.
3. the fraction of nucleic acids having the copy number variation in the circulating cell-free nucleic acids is (i) sequencing-based fraction estimation; or (ii) the allelic ratio of the polymorphic sequence; or (iii) Quantification of methylation variable nucleic acids The method according to claim 1 or 2, wherein the formula is determined according to
4. 3. The method of claim 1 or 2, wherein the fraction of nucleic acids having the copy number variation in the circulating cell-free nucleic acid is a fetal fraction determined for the region of gene copy number variation.
5. the fetal fraction of nucleic acids having the copy number variations in the circulating cell-free nucleic acids, (i) sequencing-based fetal fraction estimation; or (ii) the allelic ratio of the polymorphic sequence in the fetal nucleic acid and the maternal nucleic acid; or (iii) Quantification of methylation-variable fetal and maternal nucleic acids The method of claim 4 , wherein the value is determined according to:
6. The method of any one of claims 1 to 5, wherein the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined for a genomic region that is larger than the region of gene copy number variation.
7. The method of any one of claims 1 to 5, wherein the fraction of fetal nucleic acid in the circulating cell-free nucleic acid is determined for a genomic region that is different from the region of gene copy number variation.
8. the fraction of fetal nucleic acids in the circulating cell-free nucleic acids is (i) sequencing-based fetal fraction estimation; or (ii) the allelic ratio of the polymorphic sequence in the fetal nucleic acid and the maternal nucleic acid; or (iii) Quantification of methylation-variable fetal and maternal nucleic acids The method according to any one of claims 1 to 7, wherein the method is determined according to
9. 9. The method of any of claims 1-8, further comprising providing, by the computing system, if no classification is provided and the mosaicism ratio is less than 0.1, interpreting the positive screening result from the genetic screening test as a negative result or the absence of the one or more aneuploidies.
10. The method of claim 2, or any of claims 3 to 8 when directly or indirectly dependent on claim 2, further comprising the step of: providing, by the computing system, interpreting the positive screening result from the genetic screening test as excessive or indeterminate if the absence of genetic mosaicism is classified for the region of gene copy number variation and the mosaicism ratio is greater than about 1.
3.
11. 9. The method of claim 2, or any of claims 3 to 8 when directly or indirectly dependent on claim 2, further comprising the step of: providing, by the computing system, if the absence of genetic mosaicism is classified for the region of gene copy number variation and the mosaicism ratio is less than about 1.3, interpreting the positive screening result from the genetic screening test as a true positive for the presence of the one or more aneuploidies in the sample.
12. 12. A system comprising one or more processors and a memory, wherein the memory comprises instructions executable by the one or more processors, the instructions executable by the one or more processors being configured to perform a method according to any one of claims 1 to 11.
13. 12. A machine comprising one or more processors and a memory, wherein the memory comprises instructions executable by the one or more processors, the instructions executable by the one or more processors configured to perform a method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors of a computing system, cause the computing system to implement a method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Identification of differentially presented fetal or maternal genomic regions and their use
JP2013530727A
Analysis of genomic fractions using polymorphism counting
JP2014512817A
Generating cell-free DNA libraries directly from blood
US20140274740A1
Methods for screening and diagnosing genetic conditions
US20150064695A1
Detecting and classifying copy number variation
US9260745B2