Methods and procedures for non-invasive assessment of genetic variation

By acquiring and analyzing circulating cell-free nucleic acid sequence reads from pregnant women, mapping them to a reference genome and performing normalization, and using wavelet transform and decomposition graph analysis, we address the problem of accuracy in detecting fetal chromosomal aneuploidy, microduplications, or microdeletions, and achieve prenatal diagnosis with low false negatives and false positives.

CN112575075BActive Publication Date: 2025-09-16SEQUENOM INC
View PDF 19 Cites 0 Cited by

Patent Information

Application Number
CN202011163273.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2013-05-24
Filing Date
2014-05-23
Publication Date
2025-09-16
Estimated Expiration
2035-04-12

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately detect fetal chromosomal aneuploidy, microduplication or microdeletion with low false negative and false positive results, resulting in insufficient accuracy and reliability of prenatal diagnosis.

Method used

By obtaining circulating cell-free nucleic acid sequence reads from pregnant women, mapping them to a reference genome, and performing normalization, a genomic segment-level profile is generated. Wavelet transform and decomposition graph analysis are then used to determine whether the fetus has chromosomal aneuploidy, microduplication, or microdeletion.

Benefits of technology

It achieves low false negative and low false positive detection of fetal chromosomal abnormalities, improving the accuracy and reliability of prenatal diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112575075B_ABST
    Figure CN112575075B_ABST
Patent Text Reader

Abstract

Provided herein are methods, processes, and apparatus for non-invasively assessing genetic variation using decision analysis, which sometimes includes segmentation analysis and / or odds ratio analysis.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related patent applications

[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 61 / 827,385, filed May 24, 2013, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS,” inventors Zeljko Dzakula et al., and docket No. SEQ-6068-PV.This patent application is related to U.S. Patent No. 13 / 669,136, filed on November 5, 2012, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS” (inventors Cosmin Deciu, Zeljko Dzakula, Mathias Ehrich and Sung Kim, docket number SEQ-6034-CTt), which is related to International PCT Application No. PCT / US2012 / 059123, filed on October 5, 2012, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS” (inventors Cosmin Deciu, Zeljko Dzakula, Mathias Ehrich and Sung Kim, docket number SEQ-6034-CTt). Kim, Docket No. SEQ-6034-PC); which (i) claims the benefit of U.S. Provisional Patent Application No. 61 / 709,899, filed October 4, 2012, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS,” with Cosmin Deciu, Zeljko Dzakula, Mathias Ehrich, and Sung Kim, et al., Docket No. SEQ-6034-PV3; and (ii) claims the benefit of U.S. Provisional Patent Application No. 61 / 709,899, filed October 4, 2012, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS,” with Zeljko Dzakula and Mathias Ehrich, et al., as inventors. Ehrich et al., docket number SEQ-6034-PV2; and (iii) claim the benefit of U.S. Provisional Patent Application No. 61 / 663,477, filed October 6, 2011, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS,” inventors Zeljko Dzakula and Mathias Ehrich et al., docket number SEQ-6034-PV.The entire contents of the aforementioned patent application are incorporated herein by reference, including its text, tables, and figures.

[0003] field

[0004] The techniques presented herein relate, in part, to methods, processes, and devices for the non-invasive assessment of genetic variation. background

[0005] The genetic information of living organisms (such as animals, plants, and microorganisms) and other forms of replicating genetic information (such as viruses) is encoded in deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a series of nucleotides or modified nucleotides that represent the primary structure of a chemical or hypothetical nucleic acid. The complete human genome contains approximately 30,000 genes located on twenty-four (24) chromosomes (see The Human Genome, T. Strachan, BIOS Science Publishers, 1992). Each gene encodes a specific protein, which, after expression through transcription and translation, carries out a specific biochemical function in living cells.

[0006] Many medical conditions are caused by one or more genetic variations. Certain genetic variations cause medical conditions, including, for example, hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease and cystic fibrosis (CF) (Human Genome Mutations, D.N. Cooper and M. Krawczak, BIOS Publishers, 1993). This type of genetic disease can come from the addition, substitution or deletion of a single nucleotide in a specific gene DNA. Some birth defects are caused by chromosomal abnormalities (also referred to as aneuploidy), such as trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), X monosomy (Turner syndrome) and some sex chromosome aneuploidies such as Klinefelter syndrome (XXY). Other genetic variations are fetal sex, which can usually be determined based on sex chromosomes X and Y. Some genetic variations predispose individuals to or cause any of many diseases, such as diabetes, arteriosclerosis, obesity, various autoimmune diseases and cancer (such as colorectal cancer, breast cancer, ovarian cancer, lung cancer).

[0007] Identifying one or more genetic variations or changes can form a diagnosis or tendency determination to a specific medical condition. Identifying genetic variations can help medical decision-making and / or use auxiliary medical regimens. In certain embodiments, identifying one or more genetic variations or changes involves analyzing cell-free DNA. Cell-free DNA (CF-DNA) is composed of DNA fragments from cell death and peripheral blood circulation. High concentrations of CF-DNA can indicate certain clinical conditions, such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infection and other diseases. In addition, cell-free fetal DNA (CFF-DNA) can be detected in maternal bloodstream and is used for multiple non-invasive prenatal diagnosis.

[0008] Overview

[0009] Certain aspects of the present disclosure provide methods for determining the presence or absence of a chromosomal aneuploidy, microduplication, or microdeletion in a fetus with low false negatives and low false positives, the methods comprising (a) obtaining counts of nucleic acid sequence reads mapped to portions of a reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acid from a pregnant female, (b) normalizing the counts mapped to each portion to provide calculated genomic section levels, (c) generating a profile for the genomic segments based on the calculated genomic section levels, (d) segmenting the profile to provide two or more decomposed graphs, and (e) determining the presence or absence of a chromosomal aneuploidy, microduplication, or microdeletion in the fetus based on the two or more decomposed graphs with low false negatives and low false positives.

[0010] Certain aspects of the present invention also provide methods for determining the presence or absence of a wavelet event with low false negatives and low false positives, the methods comprising: (a) obtaining counts of nucleic acid sequence reads mapped to portions of a reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acid from a pregnant female, (b) normalizing the counts mapped to each portion to provide a calculated genomic section level, (c) segmenting the group of portions into a plurality of subsets of portions, (d) determining a level for each subset based on the calculated genomic section levels, (e) determining a significance level for each of the levels, and (f) determining the presence or absence of a wavelet event with low false negatives and low false positives based on the significance level determined for each of the levels.

[0011] Certain aspects of the present invention also provide a method for determining whether a fetus has a chromosomal aneuploidy, microduplication or microdeletion with low false negative and low false positive results, the method comprising

[0012] (a) obtaining counts of nucleic acid sequence reads mapped to portions of a reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acid from a pregnant female, (b) normalizing the counts mapped to each portion to provide calculated genomic section levels, (c) selecting segments of the genome to provide groups of portions, (d) recursively partitioning the groups of portions to provide two or more subgroups of portions, (e) determining a level for each of the two or more subgroups of portions, and (f) determining the presence or absence of a chromosomal aneuploidy, microduplication, or microdeletion in the fetus based on the levels determined in (e) with low false negatives and low false positives for the sample.

[0013] Also provided herein is a system comprising one or more processors and a memory, wherein the memory comprises instructions executable by the one or more processors, and the memory comprises counts of nucleic acid sequence reads mapped to portions of a reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acid from a pregnant female, and wherein the instructions executable by the one or more processors are configured to

[0014] (a) obtaining counts of nucleic acid sequence reads mapped to portions of a reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acid from a pregnant female, (b) normalizing the counts mapped to each portion to provide a calculated genomic section level, (c) generating a profile for the genomic segment based on the calculated genomic section level, (d) segmenting the profile to provide two or more decomposition graphs, and (e) determining the presence or absence of a chromosomal aneuploidy, microduplication, or microdeletion in a fetus based on the two or more decomposition graphs with low false negatives and low false positives.

[0015] Certain technical aspects are further described in the following description, examples, claims, and figures. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings illustrate embodiments of the present technology but are not limiting. For clarity and convenience of illustration, the accompanying drawings are not made to scale, and in some cases, various aspects may be exaggerated or enlarged to assist in understanding specific embodiments.

[0017] Figure 1 Schematic diagram showing the wavelet method. Normalized portion count data (upper right) is wavelet transformed to produce a wavelet-smoothed profile (lower right). Non-uniform events are clearly observed after wavelet denoising.

[0018] Figure 2 Shows the effect of leveling without thresholding. The optimal level can be determined by events of the desired size.

[0019] Figure 3A profile with a non-uniform profile (top) and a wavelet transformed profile (middle) is shown for a sample of chromosome 13. The bottom panel shows the null edge height distribution obtained from multiple euploid reference samples of chromosome 13. In the middle panel, the two larger differences (circles) correspond to the boundaries of non-uniform events.

[0020] Figure 4 Examples showing the emergence of segments after wavelet or CBS. The three originally divided segments (left panel, right half of the chromosome) are merged into a single long stretch (right panel, right half of the chromosome), making the microduplications clearly visible.

[0021] Figure 5 A- Figure 5 E shows the chromosome profile, which is wavelet smoothed ( Figure 5 B), CBS smooth ( Figure 5 C) and segment merged ( Figure 5 D and Figure 5 E) The two best segments from the two methods are compared with each other and "cross-validated".

[0022] Figure 6A and Figure 6B Non-limiting examples of decision analysis are shown. The same elements (e.g., blocks) of the flow chart shown are optional. In some embodiments, additional elements (e.g., validation) are added.

[0023] Figure 7 A non-limiting example of a comparison extending from 650 is shown.

[0024] Figure 8 A non-limiting example of a comparison of two wavelet events represented by 631 and 632 is shown.

[0025] Figure 9 Shown are the chromosome profiles (A), which were smoothed and merged by wavelet (B) and by CBS (C). After comparison, the two best segments from the two methods "cross-reject" each other.

[0026] Figure 10Segment profiles of chromosome 22 associated with genetic variations are shown, and the genetic variations are associated with DiGeorge syndrome. Genetic microdeletions and microduplications associated with DiGeorge syndrome have been mapped to this region. The left side (Figure AG) overview is Haar wavelet and CBS segmented, smoothed, merged and compared. The composite profile is shown in the right figure (A'-G'). The differences in sample load per cell flow are shown in the figure: A-A', 0.5-plex; B-B', 1-plex; C-C', 2-plex, D-D', 3-plex; E-E', 4-plex; F-F', 5-plex and G-G', 6-plex. Even if the sample read coverage is reduced by 10 times (see, for example, Figure F'), DiGeorge microdeletions are still detected.

[0027] Figure 11A Composite wavelet events are shown, indicating that a microdeletion was detected in the profile of chromosome 1. Figure 11B Composite WT events are shown, indicating that a microduplication was detected in the profile of chromosome 2.

[0028] Figure 12 Representative examples are shown, demonstrating the detection of the location of microduplications in chromosome 12 using the maximum entropy method.

[0029] Figure 13 Magnified views of the DiGeorge region are shown for 16 samples, labeled with number pairs indicating their location on the plate. Sample pairs 3_4 (second to last) and 9_10 (fifth to last) belong to the DiGeorge pregnancy of the infant. All other samples were chromosomally typed as euploid. The highlighted box (grey area) shows the overlap between the DiGeorge region and a selection of PERUN sections (portions of the reference genome, chr22_368-chr22_451).

[0030] Figure 14 The Z scores of the DiGeorge areas are shown. Each data point is derived from the sum of two profiles obtained from two separate aliquots of each patient. Z normalization was performed based on all 16 patients, including the two affected cases.

[0031] Figure 15-16 Representative histograms for samples 3_4 (DiGeorge) and 1_2 (euploid) are shown, respectively. Each histogram shows the distribution of Z scores obtained for a 15x15 grid area contained within the DiGeorge area. The area was selected by sliding the left and right edges of the DiGeorge area by one fraction, starting from the outer edge and moving inward. Histograms for samples 3_4 and 9_10 (not shown) Figure 1 The histogram of sample 13_14 (not shown) shows the absence of Figure 1 The results indicate overrepresentation, with only a few regions receiving Z scores below 3. All other samples (e.g., 1_2) are confined to the Z score range of [-3, 3].

[0032] Figure 17 The median Z score and its ±3 MAD confidence interval are shown for each of the 16 samples. Each median Z score was determined for a 15x15 grid area (225 areas) obtained from the sliding edge. For the vast majority of DiGeorge subregions, the Z scores of the known DiGeorge samples (3_4 and 9_10) were below -3. The apparent duplication in sample 13_14 was confirmed by the fact that most of its Z scores exceeded 3. The Z scores of all other samples were confined to the [-3, 3] segment.

[0033] Figure 18-19 The representative histograms of sample 3_4 (DiGeorge) and 1_2 (euploid) are shown respectively. Each histogram shows the distribution of the Z scores obtained for the DiGeorge region. Each Z score is calculated using 16 different groups of reference samples, using the "leave one out" method. The histogram (not shown) of sample 9_10 confirms exhaustion. Depending on the reference setting, sample 3_4 is either completely exhausted or has a boundary Z score. The histogram of sample 13_14 (not shown) shows excessive presentation, which has a few boundary Z scores. All other samples (including 1_2) are confined within the Z score segment of [-3,3].

[0034] Figure 20 The median Z score for each sample is shown, which is represented by Figure 18-19 Median Z scores and their ±3 MAD confidence intervals were calculated for 16 different reference sample groups identified using the "leave one out" approach. For the vast majority of reference sample subsets, the known DiGeorge samples (3_4 and 9_10) had Z scores below -3. Significant duplication in sample 13_14 was confirmed by the fact that most of its Z scores exceeded 3. All other sample Z scores were confined to the [-3, 3] range.

[0035] Figure 21 Shown is a comparison of the median Z scores obtained using a 15x15 grid of DiGeorge subregions (x-axis) and the median Z scores generated using a "leave one out" technique (y-axis) for each of the 16 samples. The diagonal line represents ideal agreement (slope = 1, intercept = 0).

[0036] Figure 22-23The representative histograms of sample 3_4 (DiGeorge) and 1_2 (euploid) are shown respectively. Each histogram shows the distribution of the Z score obtained with respect to the subregion in DiGeorge region, using the reference samples of 16 different groups. The subregion is randomly selected from 225 subregions of 15x15 grid. "Leave one out" analysis confirms the depletion of sample 3_4 (Figure 37) and 9_10 (not shown). The histogram of sample 13_14 confirms excessive presentation (not shown). All other samples (including 1_2) are confined in the Z score section of [-3,3].

[0037] Figure 24 Shown for each of the 16 samples, the median Z score and its ± 3 MAD confidence interval for a sub-region of the DiGeorge region randomly selected using the "leave one out" method. For most reference samples, the Z scores of the known DiGeorge samples (3_4 and 9_10) were below -3. The obvious duplication in sample 13_14 was indicated by the fact that most of its Z scores exceeded 3. With the exception of sample 17_18, the Z scores of all other samples were confined to the [-3, 3] segment.

[0038] Figures 25-26 Representative histograms are shown for samples 3_4 (DiGeorge) and 1_2 (Euploid), respectively, representing the distribution of Z scores obtained for all 225 subregions of the DiGeorge region, using 16 different reference samples. For each sample, 225 subregions were generated on a 15x15 grid using the sliding edge method. Sliding edges were used in combination with a "leave one out" analysis. The results confirmed that both affected samples 3_4 and 9_10 (not shown) were depleted. The histogram for sample 13_14 confirmed overrepresentation (not shown). All other samples (including 1_2) were confined to the Z score bin of [-3,3], with the exception of sporadic exceptions in 17_18 (not shown).

[0039] Figure 27 Shown are the median Z scores obtained using a 15x15 grid of DiGeorge subregions combined with the "leave one out" method compared to the median Z scores obtained using the 15x15 grid alone. The diagonal line represents ideal agreement (slope = 1, intercept = 0).

[0040] Figure 28 Shown is the MAD of Z scores obtained using a 15x15 grid of DiGeorge subregions combined with a "leave one out" technique compared to the MAD of Z scores obtained using the 15x15 grid alone. The diagonal line represents ideal agreement (slope = 1, intercept = 0).

[0041] Figure 29Shown are median Z scores and their ±3 MAD confidence intervals, evaluated on a full 15x15 grid of subregions of the canonic DiGeorge region using a leave-one-out approach. For most reference samples, the Z scores for the known DiGeorge samples (3_4 and 9_10) were below -3. Significant duplication in sample 13_14 is indicated by the fact that most of its Z scores exceeded 3. With the exception of sample 17_18, the Z scores for all other samples were confined to the [-3, 3] segment.

[0042] Figure 30 Exemplary embodiments of systems are shown in which certain embodiments of the technology may be implemented.

[0043] Figure 31 Classification results for LDTv2 male samples are shown using the log odds ratio (LOR) method.

[0044] Figure 32 Certain aspects of Equation 23 described in Example 6 are illustrated.

[0045] Figure 33 An embodiment of the GC density provided by the Epanechnikov kernel is shown (bandwidth = 200 bp).

[0046] Figure 34 A graph showing the GC density (y-axis) of the HTRA1 gene is shown, wherein the GC density is normalized across the entire genome. The genomic position is shown on the x-axis.

[0047] Figure 35 The local genome bias estimate (e.g., GC density, x-axis) is shown for a reference genome (solid line) and sequence reads obtained for a sample (dashed line). The bias frequency (e.g., density frequency) is shown on the y-axis. The GC density estimate is normalized across the entire genome. In this example, the sample has more high GC content reads than would be expected from the reference.

[0048] Figure 36 Displays the distribution of GC density estimates for a reference genome and for sample sequence reads, using a weighted third-order polynomial fit. GC density estimates (x-axis) are normalized across the entire genome. GC density frequencies are represented on the y-axis as the log2 of the ratio of the reference density frequency to the sample density frequency.

[0049] Figure 37A The distribution of median GC density (x-axis) across all parts of the genome is shown. Figure 37BThe median absolute deviation (MAD) values ​​determined from the GC density distribution of multiple samples are displayed (x-axis). GC density frequencies are shown on the y-axis. Portions are filtered based on the median GC density distribution of multiple reference samples (e.g., a training set) and the MAD values ​​determined from the GC density distribution of multiple samples. Portions containing GC densities exceeding a predetermined threshold (e.g., four times the interquartile range of the MAD) are removed from consideration according to the filtering method.

[0050] Figure 38A A read density profile of a sample across a genome is displayed, including the median read density in the genome (y-axis, e.g., read density / portion) and the relative position of each genomic portion (x-axis, index of portion). Figure 38B Display the first principal component (PC1), Figure 38C Shown are the second principal components (PC2), obtained from principal component analysis of read density profiles obtained in a training set of 500 euploids.

[0051] Figure 39A -C shows an example of a read density profile for a sample of a genome including a trisomy of chromosome 21 (e.g., bracketed by two vertical lines). The relative position of each genomic portion is shown on the x-axis. The read density is shown on the y-axis. Figure 39A Displays raw (e.g., uncalibrated) read density profiles. Figure 39B The profile of 39A including the first adjustment (including subtracting the median profile) is shown. Figure 39C 39B shows a profile including a second adjustment. The second adjustment includes subtracting 8x the principal component profiles, weighted based on their representation found in the sample (e.g., to build a model). For example, the sample profile = A*PC1 + B*PC2 + C*PC3..., while the corrected profile (e.g., shown in 39C) = sample profile - A*PC1 + B*PC2 + C*PC3...

[0052] Figure 40 A QQ plot showing the test p-values ​​for the bootstrapped training samples of the T21 test. QQ plots are often used to compare two distributions. Figure 40 The ChAI scores of the test samples (y-axis) are compared to the uniform distribution (i.e., the expected distribution of p-values, x-axis). Each point represents the score of the log-p value of a single test sample. The samples are sorted based on the uniform distribution and assigned an "expected" value (x-axis). The lower dotted line represents the diagonal, and the upper line represents the Bonferroni threshold. Samples that follow the uniform distribution are expected to fall on the lower diagonal (lower dotted line). Due to correlations in the parts (e.g., offsets), the values ​​are away from the diagonal, indicating that the sample has a higher score than expected (low p-value). The methods described herein (e.g., ChAI, see, e.g., Example 7) can correct for this observed offset.

[0053] Figure 41A Displays a read density plot showing the difference in PC2 coefficients for males and females in the training group. Figure 41B Receiver operating characteristic (ROC) curves for sex calls with PC2 coefficients are shown. Sex calls by sequencing were used for the true reference.

[0054] Figures 42A-42B Display system implementation. Detailed Description of the Invention

[0055] Provided herein is a method for determining fetal genetic variation (such as chromosome aneuploidy, microduplication or microdeletion) in a fetus, wherein the determination is partially and / or entirely based on a nucleic acid sequence. In some embodiments, the nucleic acid sequence is obtained from a sample (such as a pregnant woman's blood) of a pregnant woman. Also provided herein is an improved data manipulation method, and in some embodiments, a system, device and module for carrying out the methods described herein. In some embodiments, the methods described herein identify that genetic variation can guide the diagnosis of a specific medical condition or determine the tendency of a specific medical condition. Identifying genetic variation can help medical decision-making and / or use beneficial medical regimens.

[0056] sample

[0057] Provided herein are methods and compositions for analyzing nucleic acids. In some embodiments, nucleic acid fragments are analyzed in a mixture of nucleic acid fragments. The mixture of nucleic acids may include two or more nucleic acid fragment species having different nucleotide sequences, different fragment lengths, different sources (e.g., genomic, fetal, maternal, cell or tissue, sample, subject, etc.), or a combination thereof.

[0058] The nucleic acid or nucleic acid mixture used in the methods and devices described herein is often separated from a sample obtained from an object. The object can be any living or non-living organism, including but not limited to humans, non-human animals, plants, bacteria, fungi or protozoa. Any human or non-human animal can be selected, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (such as cattle), equines (such as horses), goats and ovines (such as sheep, goats), porcines (such as pigs), alpacas (such as camels, llamas, alpacas), monkeys, apes (such as gorillas, chimpanzees), ursids (such as bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales and sharks. The object can be male or female (e.g., women, pregnant women). The object can be of any age (e.g., embryo, fetus, infant, child, adult).

[0059] Nucleic acid can be separated from any type of suitable biological specimen or sample (e.g., test sample). Sample or test sample can be any sample separated or obtained from an object or part thereof (e.g., human object, pregnant female, fetus). Non-limiting examples of sample include the liquid or tissue of the object, including but not limited to blood or blood products (e.g., serum, plasma, etc.), cord blood, villi, amniotic fluid, cerebrospinal fluid, spinal fluid, washings (e.g., bronchoalveolar, stomach, peritoneum, duct, ear, arthroscopy), biopsy samples (e.g., from embryos before transplantation), intermembranous fluid samples, cells (blood cells, placental cells, embryonic or fetal cells, fetal nucleated cells or fetal cell residues) or parts thereof (e.g., mitochondria, nuclei, extracts, etc.), female reproductive tract washings, urine, feces, sputum, saliva, nasal mucosa, prostate fluid, lavage fluid, semen, lymph, bile, tears, sweat, breast milk, breast fluid, etc., or a combination thereof. In some embodiments, the biological sample is a cervical swab from an object. In some embodiments, the biological sample can be blood, and sometimes plasma or serum. As used herein, the term "blood" refers to a blood sample or product from a pregnant woman or a woman being tested for possible pregnancy. The term encompasses whole blood, blood products, or any portion of blood, such as serum and plasma, buffy coat, and the like, as conventionally defined. Blood or its portions often include nucleosomes (e.g., maternal and / or fetal nucleosomes). Nucleosomes include nucleic acids and are sometimes acellular or intracellular. Blood also includes a buffy coat. The buffy coat is sometimes separated using a Ficoll gradient. The buffy coat may include white blood cells (e.g., leukocytes, T cells, B cells, platelets, etc.). In some embodiments, the buffy coat includes maternal and / or fetal nucleic acids. Blood plasma refers to the portion of whole blood obtained by centrifugation of blood treated with an anticoagulant. Blood serum refers to the liquid, aqueous layer remaining after the blood sample has coagulated. Liquid or tissue samples are typically collected according to standard methods routinely followed in hospitals or clinics. In the case of blood, an appropriate amount of peripheral blood (e.g., 3-40 ml) is typically collected and stored according to standard procedures before or after preparation. The liquid or tissue sample used for nucleic acid extraction may be acellular (e.g., cell-free). In some embodiments, the fluid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, the sample may contain fetal cells or cancer cells.

[0060] A sample is typically heterogeneous, i.e., there is more than one type of nucleic acid species in the sample. For example, heterogeneous nucleic acids can include, but are not limited to, (i) fetal-derived and maternal-derived nucleic acids, (ii) cancer and non-cancer nucleic acids, (iii) pathogen and host nucleic acids, and more commonly (iv) mutant and wild-type nucleic acids. A sample can be heterogeneous because there is more than one cell type, such as fetal cells and maternal cells, cancer cells and non-cancerous cells, or pathogens and host cells. In some embodiments, there is a minority of nucleic acid species and a majority of nucleic acid species.

[0061] In the antenatal application of technology described herein, liquid or tissue sample can be collected from the women of gestational age that is suitable for test or the women that may be pregnant through test.Suitable gestational age may be different depending on the antenatal test performed.In certain embodiments, pregnant female object is sometimes in the first three months of pregnancy, sometimes in the second trimester three months or sometimes in the last three months of pregnancy.In certain embodiments, liquid or tissue are collected from the pregnant women of fetal gestation about 1- about 45 weeks (such as fetal gestation 1-4,4-8,8-12,12-16,16-20,20-24,24-28,28-32,32-36,36-40 or 40-44 weeks) and fetal gestation about 5- about 28 weeks (such as fetal gestation 6,7,8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26 or 27 weeks). In certain embodiments, a fluid or tissue sample is collected from a pregnant female during or shortly after delivery (eg, vaginal or non-vaginal (eg, surgical)).

[0062] Obtaining blood samples and DNA extraction

[0063] The methods herein involve isolating, enriching, and analyzing fetal DNA found in maternal blood as a non-invasive means of detecting the presence of maternal and / or fetal genetic variations and / or monitoring the health of the fetus and / or pregnant woman during and sometimes after pregnancy. Thus, the first step in practicing certain methods of the present invention involves obtaining a blood sample from a pregnant woman and extracting DNA from the sample.

[0064] Obtaining a blood sample

[0065] Blood sample can be obtained from the gestational age pregnant women that are suitable for adopting the test of the method for the present invention.Suitable gestational age can be different according to the disease being measured, as described below.Collecting women's blood is usually carried out according to the standard scheme that generally follows in hospital or clinic.Gather appropriate amount of peripheral blood, for example, be generally 5-50 milliliter, and preserve according to standard procedure before further preparation.Can make the degradation of existing nucleic acid amount in sample minimum or the mode of guaranteeing its quality gather, preserve or transport described blood sample.

[0066] Preparation of blood samples

[0067] Fetal DNA found in maternal blood is analyzed using, for example, whole blood, serum, or plasma. Methods for preparing serum or plasma from maternal blood are known. For example, the blood of a pregnant woman can be placed in a tube containing EDTA or a dedicated commercially available product such as Vacutainer SST (Becton Dickinson, Franklin Lakes, New Jersey) to prevent blood coagulation, and then plasma can be obtained from the whole blood by centrifugation. Serum may or may not be obtained by centrifugation after blood coagulation. If centrifugation is used, it is typically (but not limited to) carried out at a suitable speed (e.g., 1,500-3,000 times g). Plasma or serum can be subjected to other centrifugation steps before being transferred to a new tube for DNA extraction.

[0068] In addition to the acellular fraction of whole blood, DNA can also be recovered from the cellular components and enriched in the buffy coat fraction, which can be obtained by centrifuging a whole blood sample from a woman and removing the plasma.

[0069] DNA extraction

[0070] There are many known methods for extracting DNA from biological samples including blood. Conventional methods for DNA preparation can be followed (e.g., as described in Sambrook and Russell, Molecular Cloning: A Laboratory Manual, 3rd ed., 2001); various commercially available reagents or kits, such as the QIAamp circulating nucleic acid kit, QiaAmp DNA mini kit or QiaAmp DNA blood mini kit from Qiagen, Hilden, Germany; the GenomicPrep TM Blood DNA Isolation Kit (Promega, Madison, Wis.) and GFX TM The Genomic Blood DNA Purification Kit (Amersham, Piscataway, NJ) can also be used to obtain DNA from blood samples from pregnant women. Combinations of more than one of these methods can also be used.

[0071] In some embodiments, the sample can first be enriched or relatively enriched for fetal nucleic acid using one or more methods. For example, differentiation between fetal and maternal DNA can be performed using the compositions and methods of the present invention alone or in combination with other differentiating factors. Examples of such factors include, but are not limited to, single nucleotide differences between chromosomes X and Y, chromosome Y-specific sequences, polymorphisms elsewhere in the genome, size differences between fetal and maternal DNA, and differences in methylation patterns between maternal and fetal tissues.

[0072] Other methods for enriching samples for specific nucleic acid species are described in PCT Patent Application No. PCT / US07 / 69991, filed May 30, 2007, PCT Patent Application No. PCT / US2007 / 071232, filed June 15, 2007, U.S. Provisional Application Nos. 60 / 968,876 and 60 / 968,878 (assigned to the present applicant), (PCT Patent Application No. PCT / EP05 / 012707, filed November 28, 2005), all of which are incorporated herein by reference. In certain embodiments, maternal nucleic acids are selectively removed (partially, substantially, almost completely, or completely) from the sample.

[0073] The terms "nucleic acid" and "nucleic acid molecule" are used interchangeably herein. The term refers to any composition of nucleic acid, such as DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), RNA (e.g., messenger RNA (mRNA), short inhibitory RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed by the fetus or placenta, etc.), and / or DNA or RNA analogs (e.g., containing base analogs, sugar analogs and / or non-natural backbones, etc.), RNA / DNA hybrids and polyamide nucleic acids (PNA), all of which can be in single-stranded or double-stranded form and, unless otherwise specified, encompass known analogs of natural nucleotides that can function in a manner similar to naturally occurring nucleotides. In certain embodiments, the nucleic acid can be or can be derived from a plasmid, a phage, an autonomously replicating sequence (ARS), a centromere, an artificial chromosome, a chromosome, or other nucleic acid that can replicate or be replicated in vitro or in a host cell, a cell, a cell nucleus or cytoplasm. In some embodiments, template nucleic acid can be from a single chromosome (for example, a nucleic acid sample can be from a chromosome of a sample obtained from a diploid organism). Unless clearly defined, the term encompasses known analogs containing natural nucleotides that are similar in binding properties to a reference nucleic acid and that are metabolized in a similar manner to naturally occurring nucleotides. Unless otherwise indicated, a specific nucleic acid sequence also includes conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs) and complementary sequences, as well as sequences clearly indicated. Specifically, degenerate codon substitutions can be obtained by generating a sequence in which the third position of one or more selected (or all) codons is replaced by mixed bases and / or deoxyinosine residues. The term nucleic acid is used interchangeably with locus, gene, cDNA, and gene-encoded mRNA. The term can also include equivalents, derivatives, variants, and analogs of RNA or DNA synthesized from nucleotide analogs, single-stranded ("justice" or "antisense," "plus" strand or "minus" strand, "forward" reading frame or "reverse" reading frame) and double-stranded polynucleotides. The term "gene" refers to the segment of DNA involved in producing a polypeptide chain; it includes regions preceding and following the coding region (leader and trailer regions) that are involved in the transcription / translation of the gene product and the regulation of said transcription / translation, as well as intervening sequences (introns) between individual coding segments (exons).

[0074] Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. In the case of RNA, the base cytosine is replaced by uracil. Template nucleic acids can be prepared using nucleic acids obtained from a subject as a template.

[0075] Nucleic acid isolation and processing

[0076] Nucleic acids can be obtained from one or more sample sources (e.g., cells, serum, plasma, buffy coat, lymph, skin, soil, etc.) using methods known in the art. Any suitable method can be used to isolate, extract and / or purify DNA from a biological sample (e.g., from blood or blood products), non-limiting examples of which include methods for DNA preparation (e.g., as described in Sambrook and Russell, Molecular Cloning: A Laboratory Manual, 3rd ed., 2001); various commercially available reagents or kits, such as the QIAamp Circulating Nucleic Acid Kit, the QiaAmp DNA Mini Kit or the QiaAmp DNA Blood Mini Kit from Qiagen, Hilden, Germany; the Genomic Prep TM Blood DNA Isolation Kit (Promega, Madison, Wis.) and GFX TM Genomic Blood DNA Purification Kit (Amersham, Piscataway, NJ), etc., or a combination thereof.

[0077] Cell lysis methods and reagents are known in the art and can generally be performed by chemical (e.g., detergents, hypotonic solutions, enzymatic processes, etc., or combinations thereof), physical (e.g., French press, ultrasound, etc.), or electrolytic lysis methods. Any suitable lysis process can be used. For example, chemical methods typically use a lysis agent to disrupt cells and extract nucleic acids from the cells, followed by treatment with a chaotropic salt. Physical methods such as freeze / thaw followed by grinding, using a cell press, etc. are also useful. High salt lysis methods are also commonly used. For example, alkaline lysis can be used. The latter method traditionally involves the use of a phenol-chloroform solution, but an alternative phenol-chloroform-free method comprising three solutions can be used. In the latter method, one solution can contain 15 mM Tris, pH 8.0; 10 mM EDTA and 100 ug / ml RNase A; a second solution can contain 0.2 N NaOH and 1% SDS; and a third solution can contain 3 M KOAc, pH 5.5. These methods can be found in 6.3.1-6.3.6 of Current Protocols in Molecular Biology (1989), John Wiley & Sons, Inc., New York, which is incorporated herein in its entirety.

[0078] Nucleic acid can also be separated at a time point different from another nucleic acid, wherein each sample is from the same or different source. Nucleic acid can be from a nucleic acid library, such as a cDNA or RNA library. Nucleic acid can be the product of nucleic acid purification or separation and / or amplification of nucleic acid molecules in a sample. The nucleic acid provided for methods described herein can comprise nucleic acid from a sample or from two or more samples (e.g., from 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more samples).

[0079] In certain embodiments, nucleic acid may include extracellular nucleic acid. As used herein, the term "extracellular nucleic acid" refers to nucleic acid separated from a source that is substantially free of cells, also referred to as "cell-free" nucleic acid and / or "circulating cell-free" nucleic acid. Extracellular nucleic acid may be present in blood and obtained therefrom (e.g., from the blood of a pregnant woman). Extracellular nucleic acid generally does not include detectable cells and may contain cellular elements or cell residues. Non-limiting examples of cell-free sources of extracellular nucleic acid include blood, plasma, serum, and urine. As used herein, the term "obtaining circulating cell-free sample nucleic acid" includes directly obtaining a sample (e.g., collecting a sample, e.g., a test sample) or obtaining a sample from a person who has collected a sample. Without being limited by theory, extracellular nucleic acid may be the product of apoptosis and cell rupture, which often results in a series of lengths (e.g., "ladders") across a range of extracellular nucleic acids.

[0080] In certain embodiments, extracellular nucleic acid may comprise different nucleic acid species, and is thus referred to herein as "heterogeneity". For example, the blood serum or plasma of a person suffering from cancer may comprise nucleic acids from cancer cells and nucleic acids from non-cancerous cells. In another example, the blood serum or plasma of a pregnant woman may comprise maternal nucleic acids and fetal nucleic acids. In some examples, fetal nucleic acids sometimes account for about 5% to about 50% of total nucleic acids (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48 or 49% of total nucleic acids are fetal nucleic acids). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 500 base pairs or less, about 250 base pairs or less, about 200 base pairs or less, about 150 base pairs or less, about 100 base pairs or less, about 50 base pairs or less, or about 25 base pairs or less in length.

[0081] In certain embodiments, nucleic acids can be provided for performing the methods described herein without processing a nucleic acid-containing sample. In some embodiments, nucleic acids are provided for performing the methods described herein after processing a nucleic acid-containing sample. For example, nucleic acids can be extracted, separated, purified, partially purified, or amplified from a sample. As used herein, the term "isolation" refers to removing a nucleic acid from its original environment (e.g., the natural environment of a naturally occurring nucleic acid or a host cell expressing the nucleic acid exogenously), so that the nucleic acid is altered from its original environment by human intervention (e.g., "artificial"). As used herein, the term "isolated nucleic acid" refers to a nucleic acid removed from an object (e.g., a human object). An isolated nucleic acid may contain fewer non-nucleic acid components (e.g., proteins, lipids) than the component content in the source sample. A composition comprising an isolated nucleic acid may be about 50% to more than 99% free of non-nucleic acid components. A composition comprising an isolated nucleic acid may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more than 99% free of non-nucleic acid components. As used herein, the term "purified" refers to a nucleic acid provided with less non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than the amount of non-nucleic acid components present before the nucleic acid is subjected to the purification procedure. A composition comprising a purified nucleic acid can be about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater than 99% free of other non-nucleic acid components. As used herein, the term "purified" can refer to a nucleic acid provided that contains less nucleic acid material than the sample source from which it is derived. A composition comprising a purified nucleic acid can be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater than 99% free of other nucleic acid material. For example, fetal nucleic acid can be purified from a mixture containing maternal and fetal nucleic acid. In certain examples, nucleosomes containing small fragments of fetal nucleic acid can be purified from a mixture of large nucleosome complexes containing larger fragments of maternal nucleic acid.

[0082] In some embodiments, nucleic acids are fragmented or cleaved before, during, or after the methods of the invention. The fragmented or cleaved nucleic acids can have a nominal, average, or mean length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs. Fragments can be generated by suitable methods known in the art, and the average, mean, or nominal length of nucleic acid fragments can be controlled by selecting an appropriate fragment generation method.

[0083] Nucleic acid fragments can contain overlapping nucleotide sequences, and such overlapping sequences can facilitate the construction of nucleotide sequences of unfragmented corresponding nucleic acids or segments thereof. For example, one fragment can have subsequences x and y, and other fragments can have subsequences y and z, where x, y, and z are nucleotide sequences that can be 5 nucleotides or longer in length. In certain embodiments, overlapping nucleic acid y can be used to facilitate the construction of xyz nucleotide sequence from nucleic acids in a sample. In certain embodiments, nucleic acids can be partially fragmented (e.g., from incomplete or aborted specific shearing reactions) or fully fragmented.

[0084] In some embodiments, nucleic acids can be fragmented or cleaved by suitable methods, non-limiting examples of which include physical methods (e.g., shearing, such as ultrasound, French press, heat, UV irradiation, etc.), enzymatic processing (e.g., enzymatic cleavage reagents (e.g., suitable nucleases, suitable restriction enzymes, suitable methylation-sensitive restriction enzymes)), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, base hydrolysis, heat, etc., or combinations thereof), methods described in U.S. Patent Application Publication No. 20050112590, etc., or combinations thereof.

[0085] As used herein, "fragmentation" or "cleavage" refers to a process or condition by which a nucleic acid molecule (e.g., a nucleic acid template gene molecule or an amplified product thereof) can be separated into two or more smaller nucleic acid molecules. This fragmentation or cleavage can be sequence-specific, base-specific, or non-specific, and can be accomplished by any of a variety of methods, reagents, or conditions (including, for example, chemical, enzymatic, or physical fragmentation).

[0086] As used herein, "fragment", "cleavage product", "cleaved product" or grammatical variants thereof refer to nucleic acid molecules obtained by fragmentation or shearing of a nucleic acid template gene molecule or its amplified product. Although such fragments or sheared products may refer to all nucleic acids obtained by the shearing reaction, such fragments or sheared products generally refer only to nucleic acid molecules obtained by fragmentation or shearing of a nucleic acid template gene molecule or its amplified product segment (comprising the corresponding nucleotide sequence of the nucleic acid template gene molecule). As used herein, the term "amplification" refers to a process in which a target nucleic acid in a treated sample undergoes a linear or exponential process to produce amplicon nucleic acids, the nucleotide sequence of which is identical or substantially identical to the nucleotide sequence of the target nucleic acid or its segment. In certain embodiments, the term "amplification" refers to a method comprising a polymerase chain reaction (PCR). For example, an amplified product can contain one or more nucleotides more than the amplified nucleotide region of the nucleic acid template sequence (e.g., a primer can contain "extra" nucleotides, such as a transcription initiation sequence, in addition to nucleotides complementary to the nucleic acid template gene molecule, to generate an amplified product comprising "extra" nucleotides or nucleotides that do not correspond to the amplified nucleotide region of the nucleic acid template gene molecule). Thus, fragments can comprise fragments from segments or portions of amplified nucleic acid molecules that at least in part comprise nucleotide sequence information from or based on a representative nucleic acid template molecule.

[0087] The term "complementary shearing reaction" as used herein refers to a shearing reaction performed on the same nucleic acid with different shearing reagents or by changing the shearing specificity of the same shearing reagent, thereby producing different shearing patterns of the same target or reference nucleic acid or protein. In certain embodiments, nucleic acids can be treated in one or more reaction vessels with one or more specific shearing agents (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more specific shearing agents) (e.g., nucleic acids are treated in separate containers with various specific shearing agents). As used herein, the term "specific shearing agent" refers to a reagent, sometimes a chemical or enzyme that can shear nucleic acids at one or more specific sites.

[0088] Prior to providing the nucleic acid for use in the methods described herein, the nucleic acid can also be treated to modify certain nucleotides within the nucleic acid. For example, the nucleic acid can be subjected to a treatment that selectively modifies the nucleic acid based on the methylation status of the nucleotides within the nucleic acid. Additionally, conditions such as elevated temperatures, ultraviolet radiation, and x-ray radiation can induce variations in the sequence of nucleic acid molecules. The nucleic acid can be provided in any suitable form for performing appropriate sequence analysis.

[0089] The nucleic acid may be single-stranded or double-stranded. For example, double-stranded DNA may be denatured by heating or, for example, by alkali treatment to generate single-stranded DNA. In certain embodiments, the nucleic acid is a D-loop structure formed by strand invasion of a double-stranded DNA molecule with an oligonucleotide or DNA-like molecule, such as a peptide nucleic acid (PNA). Addition of E. coli RecA protein and / or alteration of salt concentration (e.g., using methods known in the art) may facilitate formation of the D-loop.

[0090] Determine fetal nucleic acid content

[0091] In some embodiments, the amount of fetal nucleic acid in nucleic acid is determined (e.g., concentration, relative amount, absolute amount, copy number, etc.). In certain embodiments, the amount of fetal nucleic acid in a sample is referred to as "fetal fraction." In some embodiments, "fetal fraction" refers to the fraction of fetal nucleic acid in circulating cell-free nucleic acid in a sample obtained from a pregnant female (e.g., a blood sample, a serum sample, a plasma sample). In certain embodiments, the amount of fetal nucleic acid is determined based on markers specific for male fetuses (e.g., Y chromosome STR markers (e.g., DYS 19, DYS 385, DYS 392 markers); RhD markers in RhD-negative females), allele ratios of polymorphic sequences, or one or more markers specific for fetal nucleic acid but not for maternal nucleic acid (e.g., differential epigenetic biomarkers between mother and fetus (e.g., methylation; described in detail below), or fetal RNA markers in maternal plasma (see, e.g., Lo, 2005, Journal of Histochemistry and Cytochemistry 53(3):293-296)).

[0092] Determining the amount of fetal nucleic acid (e.g., fetal fraction) is sometimes performed using a fetal quantification assay (FQA), as described in U.S. Patent Application Publication 2010 / 0105049, which is incorporated herein by reference. Such assays allow for the detection and quantification of fetal nucleic acid in a maternal sample based on the methylation status of the nucleic acid in the sample. In certain embodiments, the amount of fetal nucleic acid in a maternal sample can be determined relative to the total amount of nucleic acid present, thereby providing a percentage of fetal nucleic acid in the sample. In certain embodiments, the copy number of fetal nucleic acid in a maternal sample can be determined. In certain embodiments, the amount of fetal nucleic acid can be determined in a sequence-specific (or partially-specific) manner, and sometimes the sensitivity is sufficient to perform accurate chromosome dosage analysis (e.g., to detect the presence or absence of fetal aneuploidy, microduplication, or microdeletion).

[0093] Fetal quantification assays (FQA) can be performed in conjunction with any of the methods described herein. The assay can be performed by any method known in the art and / or as described in U.S. Patent Application Publication No. 2010 / 0105049, such as by methods that can distinguish between maternal and fetal DNA based on differential methylation status, and by methods that quantify fetal DNA (i.e., determine its content). Methods for distinguishing nucleic acids based on methylation status include, but are not limited to, methylation-sensitive capture (e.g., using MBD2-Fc fragments, in which the methylation binding domain of MBD2 is fused to the Fc fragment of an antibody (MBD-FC) (Gebhard et al. (2006) Cancer Res. 66(12):6118-28)); methylation-specific antibodies, bisulfite conversion methods, such as MSP (methylation-sensitive PCR), COBRA, methylation-sensitive single nucleotide primer extension (Ms-SNuPE), or Sequenom MassCLEAVE. TM Technology; and the use of methylation-sensitive restriction enzymes (e.g., digesting maternal DNA in a maternal sample with one or more methylation-sensitive restriction enzymes to enrich for fetal DNA). Methyl-sensitive enzymes can also be used to differentiate nucleic acids based on methylation status, for example, preferentially or significantly cutting or digesting when their DNA recognition sequences are not methylated. Thus, unmethylated DNA samples will be cut into smaller fragments than methylated samples, while highly methylated DNA samples will not be cut. Unless explicitly stated, any method for distinguishing nucleic acids based on methylation status can be used in the compositions and methods of the present invention. The content of fetal DNA can be determined, for example, by introducing one or more competitors of known concentration during the amplification reaction. The content of fetal DNA can also be determined by, for example, RT-PCR, primer extension, sequencing and / or counting. In some examples, the BEAMing technology described in U.S. Patent Application Publication 2007 / 0065823 can be used to determine the content of nucleic acids. In some embodiments, the restriction efficacy can be determined and the efficiency ratio can be used to further determine the amount of fetal DNA.

[0094] In certain embodiments, a fetal quantification assay (FQA) can be performed using the concentration of fetal DNA in a maternal sample, for example, by: a) determining the total amount of DNA present in the maternal sample; b) selectively digesting the maternal DNA in the maternal sample with one or more methylation-sensitive restriction enzymes to enrich for the fetal DNA; c) determining the amount of fetal DNA from step b); and d) comparing the amount of fetal DNA obtained in step c) to the total amount of DNA obtained in step a), thereby determining the concentration of fetal DNA in the maternal sample. In certain embodiments, the absolute copy number of fetal nucleic acid in a maternal sample can be determined, for example, using mass spectrometry and / or a system utilizing competitive PCR methods for absolute copy number determination. See, for example, Ding and Cantor (2003) Proc. Natl. Acad. Sci. USA 100:3059-3064, and U.S. Patent Application Publication 2004 / 0081993, both of which are incorporated herein by reference.

[0095] In certain embodiments, fetal fraction can be determined based on the allelic ratio of a polypeptide sequence (e.g., a single nucleotide polymorphism (SNP)), for example using the method described in U.S. Patent Application Publication No. 2011 / 0224087, which is incorporated herein by reference. In this method, nucleotide sequence reads are obtained for a maternal sample and the fetal fraction is determined by comparing the total number of nucleotide sequence reads mapped to a first allele with the total number of nucleotide sequence reads mapped to a second allele of a reference polymorphic site (e.g., a SNP) located in a reference genome. In certain embodiments, fetal alleles are identified by, for example, the relatively small contribution of the fetal allele to a mixture of fetal and maternal nucleic acids in a sample relative to the larger contribution of the maternal nucleic acid to the mixture. Thus, the relative abundance of fetal nucleic acid in a maternal sample can be determined as a parameter of the total number of unique sequence reads mapped to a target nucleic acid sequence on a reference genome (for each of the two alleles of the polymorphic site).

[0096] In some embodiments, fetal fraction can be determined using methods that incorporate fragment length information (e.g., fragment length ratio (FLR) analysis, fetal ratio statistics (FRS) analysis, as described in International Application Publication No. WO2013 / 177086, which is incorporated herein by reference). Cell-free fetal nucleic acid fragments are generally shorter than nucleic acid fragments of maternal origin (see, e.g., Chan et al. (2004) Clin. Chem. 50:88-92; Lo et al. (2010) Sci. Transl. Med. 2:61ra91). Thus, in some embodiments, fetal fraction can be determined by counting fragments below a specific length threshold and comparing the count to, for example, a count of fragments above a specific length threshold and / or the amount of total nucleic acid in a sample. Methods for counting nucleic acid fragments of a specific length are described in detail in International Application Publication No. WO2013 / 177086.

[0097] In some embodiments, a fetal fraction can be determined based on a portion-specific fetal fraction estimate. Without being limited by theory, the number of reads of fetal CCF fragments (e.g., fragments of a particular length or length range) is often mapped to portions (e.g., within the same sample, e.g., within the same sequencing run) along with the frequency of the reads. Furthermore, without being limited by theory, when compared across multiple samples, certain portions may have a similar representation of reads as fetal CCF fragments (e.g., fragments of a particular length or length range), and such representation may correlate with a portion-specific fetal fraction (e.g., a relative amount, percentage, or ratio of CCF fragments that are fetal in origin).

[0098] In some embodiments, a portion-specific fetal fraction estimate is determined based in part on a portion-specific parameter and its relationship to fetal fraction. A portion-specific parameter can be any suitable parameter that reflects (e.g., correlates with) the amount or proportion of reads having CCF fragment lengths of a particular size (e.g., a size range) in a portion. A portion-specific parameter can be an average, mean, or median of a portion-specific parameter determined for multiple samples. Any suitable portion-specific parameter can be used. Non-limiting examples of portion-specific parameters include FLR (e.g., FRS), the number of reads below a selected fragment length, genomic coverage (i.e., coverage), mappability, counts (e.g., counts of sequence reads mapped to the portion, e.g., normalized counts, PERUN normalized counts, ChAI normalized counts), DNase I sensitivity, methylation status, acetylation, histidine distribution, guanine-cytosine (GC) amount, chromatin structure, the like, or a combination thereof. A portion-specific parameter can be any suitable parameter that correlates with FLR and / or FRS in a portion-specific manner. In some embodiments, some or all portion-specific parameters are direct or indirect representations of FLR for a portion. In some embodiments, a portion-specific parameter is not guanine-cytosine (GC) content.

[0099] In some embodiments, a portion-specific parameter is any suitable value representing, associated with, or proportional to the amount of CCF fragment reads, wherein the length of the reads mapped to a portion is less than a selected fragment length. In some embodiments, a portion-specific parameter represents the amount of reads derived from relatively short CCF fragments (e.g., about 200 base pairs or less) mapped to a portion. CCF fragments having a length less than a selected fragment length are typically relatively short CCF fragments, and sometimes the selected fragment length is about 200 base pairs or less (e.g., CCF fragments of about 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, or 80 bases in length). The length of a CCF fragment or the number of reads derived from a CCF fragment can be determined (e.g., inferred or deduced) by any suitable method (e.g., sequencing method, hybridization method). In some embodiments, the length of a CCF fragment is determined (e.g., inferred or deduced) by reads obtained by paired-end sequencing. In some embodiments, a CCF fragment template is determined directly from the length of a read (e.g., a single-end read) derived from the CCF fragment.

[0100] Portion-specific parameters can be weighted or adjusted by one or more weighting factors. In some embodiments, weighted or adjusted portion-specific parameters can provide a portion-specific fetal fraction estimate for a sample (e.g., a test sample). In some embodiments, weighting or adjustment generally converts portion counts (e.g., reads mapped to a portion) or other portion-specific parameter into a portion-specific fetal fraction estimate, and such conversion is sometimes referred to as a transformation.

[0101] In some embodiments a weighting factor is a coefficient or constant that partially describes and / or defines a relationship between fetal fraction (e.g., fetal fraction determined from multiple samples) and portion-specific parameters for multiple samples (e.g., a training set). In some embodiments a weighting factor is determined based on a correlation between multiple fetal fraction determinations and multiple portion-specific parameters. One or more weighting factors can define a correlation, and one or more weighting factors can be determined from the correlation. In some embodiments a weighting factor (e.g., one or more weighting factors) is determined from a fitted correlation of portions based on (i) the fraction of fetal nucleic acid determined for each of multiple samples, and (ii) the portion-specific parameters for multiple samples.

[0102] The weighting factor can be any suitable coefficient, estimated coefficient or constant derived from a suitable correlation (e.g., suitable mathematical correlation, algebraic correlation, fitting correlation, regression, regression analysis, regression model). The weighting factor can be determined according to a suitable correlation, or can be derived from a suitable correlation or from a suitable correlation assessment. In some embodiments, the weighting factor is an assessment coefficient from a fitting correlation. Fitting a variety of samples to a correlation is sometimes referred to as training a model. Any suitable model and / or method for performing relationship fitting (e.g., performing model training with respect to a training group) can be used. The non-limiting examples of available suitable models include regression models, linear regression models, simple regression models, ordinary least squares regression models, multiple regression models, general multiple regression models, polynomial regression models, general linear models, generalized linear models, discrete choice regression models, logistic regression models, multinomial logit models, mixed logit models, probit models, multinomial probit models, ordered logit models, ordered probit models, Poisson (Poisson) models, multivariate response regression models, multilevel models, fixed effects models, random effects models, mixed models, nonlinear regression models, nonparametric models, semiparametric models, robust (robust) models, quantile models, isotonic models, principal component models, minimum angle models, local models, segmented models and variable error models. In some embodiments, the fitting correlation is not a regression model. In some embodiments, the fitting correlation is selected from a decision tree model, a support vector machine model and a neural network model. The result of carrying out model training (e.g., regression model, correlation) is typically a correlation that can be mathematically described, wherein the correlation includes one or more coefficients (e.g., weighting factors). More complex multivariate models can determine 1, 2, 3 or more weighting factors. In some embodiments a model is trained based on fetal fraction and two or more portion-specific parameters (coefficients) obtained from multiple samples (e.g., by fitting a relationship by matrix fitting to multiple samples).

[0103] Weighting factors can be derived from a suitable correlation by a suitable method (e.g., a suitable mathematical correlation, algebraic correlation, fitted correlation, regression, regression analysis, regression model). In some embodiments, a fitted correlation is fitted by an evaluation, non-limiting examples of which include least squares, ordinary least squares, linear, partial, total, generalized, weighted, nonlinear, iterative weighting, ridge regression, least squares, Bayesian, Bayesian multivariate, reduced rank, LASSO, weighted rank selection criterion (WRSC), rank selection criterion (RSC), elastic net estimation (e.g., elastic net regression), and combinations thereof.

[0104] The weighting factor may be determined or associated with any suitable portion of the genome. The weighting factor may be determined or associated with any suitable portion of any suitable chromosome. In some embodiments, the weighting factor may be determined or associated with some or all portions of the genome. In some embodiments, the weighting factor may be determined or associated with portions of some or all chromosomes in the genome. Sometimes, the weighting factor may be determined or associated with portions of a selected chromosome. The weighting factor may be determined or associated with portions of one or more autosomes. The weighting factor may be determined or associated with portions of a plurality of portions that include portions in autosomes or a subset thereof. In some embodiments, the weighting factor may be determined or associated with portions of sex chromosomes (such as ChrX and / or ChrY). The weighting factor may be determined or associated with portions of one or more sex chromosomes and one or more autosomes. In certain embodiments, the weighting factor may be determined or associated with portions of chromosomes X and Y and all autosomes. The weighting factor may be determined or associated with portions of a plurality of portions that do not include portions in chromosomes X and / or Y. In certain embodiments, a weighting factor is determined for or associated with a portion of a chromosome wherein the chromosome comprises an aneuploidy (e.g., a whole chromosome aneuploidy). In certain embodiments, a weighting factor is determined for or associated with a portion of a chromosome wherein the chromosome is not aneuploid (e.g., a euploid chromosome). A weighting factor can be determined for or associated with a portion of a plurality of portions that does not include portions of chromosomes 13, 18, and / or 21.

[0105] In some embodiments, weighting factors are determined for portions based on one or more samples (e.g., a training set of samples). Weighting factors are typically specific to portions. In some embodiments, one or more weighting factors are independently assigned to portions. In some embodiments, weighting factors are determined based on relationships in fetal fraction determinations for multiple samples (e.g., sample-specific fetal fraction determinations) and portion-specific parameters determined based on multiple samples. Weighting factors are typically determined from multiple samples, e.g., from about 20 to about 100,000 or more samples, from about 100 to about 100,000 or more samples, from about 500 to about 100,000 or more samples, from about 1,000 to about 100,000 or more samples, or from about 10,000 to about 100,000 or more samples. Weighting factors can be determined from euploid samples (e.g., samples from subjects containing euploid fetuses, e.g., samples without aneuploid chromosomes). In some embodiments, weighting factors are obtained from samples containing aneuploid chromosomes (e.g., samples from subjects containing euploid fetuses). In some embodiments, a weighting factor is determined from multiple samples from subjects having a euploid fetus and a subject having a trisomic fetus. A weighting factor can be derived from multiple samples from subjects having a male fetus and / or a female fetus.

[0106] The fetal fraction is typically determined for one or more samples in a training set, and the weighting factor is derived from the fetal fraction. The fetal fraction from which the weighting factor is derived is sometimes a sample-specific fetal fraction determination. The fetal fraction from which the weighting factor is determined can be determined by any suitable method described herein or known in the art. In some embodiments, determination of fetal nucleic acid content (e.g., fetal fraction) is performed using a suitable fetal quantification assay (FQA) described herein or known in the art, non-limiting examples of which include fetal fraction determination based on markers specific for male fetuses, based on allele ratios of polymorphic sequences, based on one or more markers specific for fetal nucleic acid but not specific for maternal nucleic acid, by utilizing methylation-based DNA recognition (e.g., A. Nygren, et al., (2010) Clinical Chemistry 56(10):1627–1635), by mass spectrometry and / or systems using competitive PCR methods, by methods described in U.S. Patent Application Publication No. 2010 / 0105049 (which is incorporated herein by reference), etc., or a combination thereof. Often fetal fraction is determined in part based on the level (e.g., one or more genomic segment level, profile level) of chromosome Y. In some embodiments, fetal fraction is determined according to a suitable assay for chromosome Y (e.g., by using quantitative real-time PCR to compare the amount of a fetal-specific locus (e.g., the SRY locus on the Y chromosome in male pregnancies) with the amount of a locus on any autosome that is common in both the mother and the fetus (e.g., Lo YM, et al. (1998) Am J Hum Genet 62:768–775.)).

[0107] A portion-specific parameter (e.g., for a test sample) can be weighted or adjusted by one or more weighting factors (e.g., weighting factors derived from a training set). For example, a weighting factor can be derived for a portion based on the relationship between the portion-specific parameter and fetal fraction determined for a training set of multiple samples. The portion-specific parameter for a test sample is then adjusted and / or weighted based on the weighting factor derived from the training set. In some embodiments, the portion-specific parameter from which the weighting factor is derived is the same as the portion-specific parameter (e.g., for a test sample) that is adjusted or weighted (e.g., both are FLR). In certain embodiments, the portion-specific parameter from which the weighting factor is derived is different from the portion-specific parameter (e.g., for a test sample) that is adjusted or weighted. For example, a weighting factor can be determined based on the correlation between coverage (i.e., a portion-specific parameter) and fetal fraction for a training set of samples, while the FLR of a portion of a test sample (i.e., another portion-specific parameter) can be adjusted based on the weighting factor derived from coverage. Without being bound by any theory, portion-specific parameters (e.g., of a test sample) can sometimes be adjusted and / or weighted by weighting factors derived from different portion-specific parameters (e.g., of a training set) based on the correlation and / or association between each portion-specific parameter and the common portion-specific FLR.

[0108] A portion-specific fetal fraction estimate for a sample (e.g., a test sample) can be determined by weighting a portion-specific parameter using a weighting factor determined for that portion. Weighting can include adjusting, converting, and / or transforming the portion-specific parameter based on the weighting factor by applying any suitable mathematical operation, non-limiting examples of which include multiplication, division, addition, subtraction, integration, symbolic operations, algebraic calculations, algorithms, trigonometric or geometric functions, transformations (e.g., Fourier transforms), and the like, or combinations thereof. Weighting can include adjusting, converting, and / or transforming the portion-specific parameter based on a suitable mathematical model for the weighting factor.

[0109] In some embodiments, a fetal fraction for a sample is determined based on one or more portion-specific fetal fraction estimates. In some embodiments, a fetal fraction for a sample (e.g., a test sample) is determined (e.g., estimated) based on weighting or adjusting portion-specific parameters for one or more portions. In certain embodiments, the fraction of fetal nucleic acid for a test sample is estimated based on adjusted counts or adjusted subsets of counts. In certain embodiments, the fraction of fetal nucleic acid for a test sample is estimated based on adjusted FLR, adjusted FRS, adjusted coverage, and / or adjusted mappability of portions. In some embodiments, about 1 to about 500,000, about 100 to about 300,000, about 500 to about 200,000, about 1,000 to about 200,000, about 1,500 to about 200,000, or about 1,500 to about 50,000 portion-specific parameters are weighted or adjusted.

[0110] Determining the fetal fraction (e.g., of a test sample) can be performed by any suitable method based on multiple portion-specific fetal fraction estimates (e.g., of the same test sample). In some embodiments, a method for improving the accuracy of an estimate of the fraction of fetal nucleic acid in a test sample from a pregnant female comprises determining one or more portion-specific fetal fraction estimates, wherein the estimate of the fetal fraction of the sample is determined based on the one or more portion-specific fetal fraction estimates. In some embodiments, assessing or determining the fraction of fetal nucleic acid in a sample (e.g., a test sample) comprises summing one or more portion-specific fetal fraction estimates. Summing can comprise determining an average, mean, median, AUC, or integrated value based on multiple portion-specific fetal fraction estimates.

[0111] In some embodiments, a method for improving the accuracy of an estimate of the fraction of fetal nucleic acid in a test sample from a pregnant female comprises obtaining counts of sequence reads mapped to portions of a reference genome, the sequence reads being reads of circulating, cell-free nucleic acid from a test sample from a pregnant female, wherein at least a subset of the obtained counts are derived from regions of the genome that contribute to a greater number of fetal nucleic acid counts relative to the total counts of fetal nucleic acid relative to the total counts of fetal nucleic acid relative to other regions of the genome. In some embodiments, an estimate of the fraction of fetal nucleic acid is determined based on a subset of the portions, wherein the subset of portions is selected based on portions mapped to a greater number of fetal nucleic acid counts relative to the total counts of fetal nucleic acid relative to other portions. In some embodiments, the subset of portions is selected based on portions mapped to a greater number of fetal nucleic acid counts relative to non-fetal nucleic acid relative to other portions. Counts mapped to all portions or subsets of portions can be weighted to provide weighted counts. Weighted counts can be used to estimate the fraction of fetal nucleic acid, and the counts can be weighted according to portions that map to a number of fetal nucleic acid counts that is greater than the fetal nucleic acid counts in other portions. In some embodiments, the counts are weighted according to portions that map to a number of fetal nucleic acid counts relative to non-fetal nucleic acid that is greater than the fetal nucleic acid counts in other portions.

[0112] The fetal fraction of a sample (e.g., a test sample) can be determined based on a plurality of portion-specific fetal fraction estimates for the sample, wherein the portion-specific estimates are from portions of any suitable region or segment of the genome. The portion-specific fetal fraction estimates can be determined for one or more portions of a suitable chromosome (e.g., one or more selected chromosomes, one or more autosomes, sex chromosomes (e.g., ChrX and / or ChrY), aneuploid chromosomes, euploid chromosomes, etc., or a combination thereof).

[0113] In some embodiments, determining fetal fraction comprises

[0114] (a) obtaining counts of sequence reads that map to portions of a reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acid from a test sample from a pregnant female;

[0115] (b) using a microprocessor, weighting (i) the counts of sequence reads mapped to each portion or (ii) other portion-specific parameters to the portion-specific fraction of fetal nucleic acid by independently associating a weighting factor for each portion, thereby providing a portion-specific fetal fraction estimate according to the weighting factor, wherein each weighting factor has been determined from a fitted correlation between, for each portion, (i) the fetal nucleic acid fraction for each of the plurality of samples and (ii) the counts of sequence reads mapped to each portion (or other portion-specific parameter) for the plurality of samples; and

[0116] (c) assessing the fetal nucleic acid fraction of the test sample based on the portion-specific fetal fraction estimate.

[0117] The amount of fetal nucleic acid in the extracellular nucleic acid can be quantified and can be used in conjunction with the methods described herein. Therefore, in certain embodiments, the methods of the technology described herein include an additional step of determining the amount of fetal nucleic acid. The amount of fetal nucleic acid in the nucleic acid sample of the subject can be determined before or after processing to prepare the sample nucleic acid. In certain embodiments, after the sample nucleic acid is processed and prepared, the amount of fetal nucleic acid in the sample is determined and used for further evaluation. In some embodiments, the results include decomposing the fetal nucleic acid fraction in the sample nucleic acid into factors (such as adjusting counts, removing samples, making a determination, or not making a determination).

[0118] The determining step can be performed before, during, at any time point during the methods described herein, or after certain methods described herein (e.g., aneuploidy detection, microduplication or microdeletion detection, fetal sex determination). For example, to achieve a fetal sex or aneuploidy, microduplication or microdeletion detection method with a given sensitivity or specificity, a fetal nucleic acid quantification method can be performed before, during, or after fetal sex or aneuploidy, microduplication or microdeletion determination to identify those samples having greater than about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25% or more fetal nucleic acid. In some embodiments, samples determined to have a certain fetal nucleic acid threshold amount (e.g., about 15% or more fetal nucleic acid; e.g., about 4% or more fetal nucleic acid) are further analyzed for, for example, fetal sex or aneuploidy, microduplication or microdeletion, or the presence or absence of aneuploidy or genetic variation. In certain embodiments, only samples with a certain fetal nucleic acid threshold amount (e.g., about 15% or more fetal nucleic acid; e.g., about 4% or more fetal nucleic acid) are selected (e.g., selected and informed to the patient) to determine, for example, fetal sex or the presence or absence of aneuploidy, microduplication or microdeletion.

[0119] In some embodiments, determining the fetal fraction or determining the amount of fetal nucleic acid is not necessary to identify whether there is a chromosomal aneuploidy, microduplication or microdeletion. In some embodiments, identifying whether there is a chromosomal aneuploidy, microduplication or microdeletion does not require sequence differentiation of fetal and maternal DNA. In certain embodiments, this is due to analysis of the additive contributions of maternal and fetal sequences to specific chromosomes, chromosome portions or segments thereof. In some embodiments, identifying whether there is a chromosomal aneuploidy, microduplication or microdeletion does not rely on prior sequence information that distinguishes fetal DNA from maternal DNA.

[0120] Enrichment of nucleic acids

[0121] In some embodiments, nucleic acid (such as extracellular nucleic acid) is enriched or relatively enriched for nucleic acid subgroup or material. Nucleic acid subgroup may include, for example, fetal nucleic acid, maternal nucleic acid, nucleic acid containing a fragment of a specific length or length range, or nucleic acid from a specific genomic region (such as a single chromosome, chromosome group, and / or some chromosome region). The sample of this type of enrichment can be used in conjunction with the methods described herein. Therefore, in some embodiments, the method of this technology includes an additional step of nucleic acid subgroup such as fetal nucleic acid in the enrichment sample. In some embodiments, the above-mentioned method for determining fetal fraction can also be used for enrichment of fetal nucleic acid. In some embodiments, (partial, substantially, almost completely or completely) maternal nucleic acid is selectively removed from the sample. In some embodiments, the nucleic acid (such as fetal nucleic acid) of the specific low copy number of enrichment can improve quantitative sensitivity. Methods for enriching a specific type of nucleic acid in a sample are described, for example, in U.S. Patent No. 6,927,028, International Application Publication No. WO2007 / 140417, International Application Publication No. WO2007 / 147063, International Application Publication No. WO2009 / 032779, International Application Publication No. WO2009 / 032781, International Application Publication No. WO2010 / 033639, International Application Publication No. WO2011 / 034631, International Application Publication No. WO2006 / 056480, and International Application Publication No. WO2011 / 143659, all of which are incorporated herein by reference.

[0122] In some embodiments, nucleic acid is enriched for certain target fragment types and / or reference fragment types. In certain embodiments, nucleic acid is enriched using one or more of the following length-based separation methods with respect to specific nucleic acid fragment lengths or fragment lengths or ranges. In certain embodiments, nucleic acid is enriched using one or more sequence-based separation methods described herein and / or known in the art with respect to fragments selected from genomic regions (e.g., chromosomes). Methods for nucleic acid subpopulations (e.g., fetal nucleic acid) in certain enriched samples are described in detail below.

[0123] The method for the enrichment nucleic acid subgroup (for example fetal nucleic acid) that can be used together with the inventive method comprises the method for the apparent difference that adopts between mother and fetal nucleic acid.For example can distinguish and separate fetal nucleic acid and maternal nucleic acid based on methylation difference.Based on methylated fetal nucleic acid enrichment method referring to U.S. Patent Application Publication 2010 / 0105049, it includes this paper in by reference.This method sometimes relates to the nucleic acid and unconjugated nucleic acid of binding sample nucleic acid and methylation special binding reagent (methyl CpG binding protein (MBD), methylation specific antibody etc.) and based on different methylation states separation bonded.This type of method can also comprise and use methylation sensitive restriction enzyme (for example HhaI and HpaII as mentioned above), it comes selective digestion from the nucleic acid of maternal sample thereby at least a fetal nucleic acid zone in the enrichment sample by using selectivity and the enzyme of digesting maternal nucleic acid completely or substantially, so just can the fetal nucleic acid zone in the enrichment maternal sample.

[0124] Another method for enriching a nucleic acid subpopulation (e.g., fetal nucleic acid) that can be used with the methods of the present invention is a restriction endonuclease-enhanced polymorphic sequence method, such as the method described in U.S. Patent Application Publication No. 2009 / 0317818, which is incorporated herein by reference. The method comprises cleaving a nucleic acid containing a non-target allele with a restriction endonuclease that recognizes the non-target allele but does not recognize the target allele; and amplifying the uncut nucleic acid but not the cleaved nucleic acid, wherein the uncut amplified nucleic acid represents the target nucleic acid (e.g., fetal nucleic acid) enriched relative to the non-target nucleic acid (e.g., maternal nucleic acid). In certain embodiments, the nucleic acid can be selected so that it contains an allele having a polymorphic site that is susceptible to selective digestion by, for example, a cleavage agent.

[0125] Methods for enriching nucleic acid subpopulations (e.g., fetal nucleic acid) that can be used in conjunction with the methods of the present invention include selective enzymatic degradation methods. This method involves protecting the target sequence from digestion by exonucleases, thereby facilitating the elimination of unwanted sequences (e.g., maternal DNA) in the sample. For example, in one method, the sample nucleic acid is denatured to produce single-stranded nucleic acid, which is contacted with at least one target-specific primer pair under suitable annealing conditions, and the annealed primer is extended using nucleotide polymerization to produce a double-stranded target sequence, and the single-stranded nucleic acid is digested with a nuclease that digests single-stranded (e.g., non-target) nucleic acid. In some embodiments, the method can be repeated for at least one more cycle. In some embodiments, the same target-specific primer pair can be used to initiate the first and second cycles of extension, and in some embodiments, different target-specific primer pairs are used for the first and second cycles.

[0126] Methods for enriching nucleic acid subpopulations (e.g., fetal nucleic acid) that can be used in conjunction with the methods of the present invention include massively parallel sequencing technology (MPSS). MPSS is typically a solid phase method that uses adapters (i.e., tags) to connect, which are then decoded and the nucleic acid sequence is read in small increments. Tagged PCR products are typically amplified so that each nucleic acid produces a PCR product with a unique tag. Tags are typically used to join PCR products to microbeads. For example, after several rounds of sequence determination based on the connection, sequence signatures can be identified from each bead. Each signature sequence (MPSS tag) in the MPSS database is analyzed, all other signatures are compared, and all identical signatures are counted.

[0127] In some embodiments, some enrichment methods (such as some MPS-based and / or MPSS-based enrichment methods) may include methods based on amplification (such as PCR). In some embodiments, site-specific amplification methods (such as using site-specific amplification primers) can be used. In some embodiments, multiple SNP allele PCR methods can be used. In some embodiments, multiple SNP allele PCR methods can be used in conjunction with single sequencing. For example, the method may involve using multiple PCR (MASSARRAY system) and the capture probe sequence is included in the amplicon, and then use, for example, Illumina MPSS system sequencing. In some embodiments, multiple SNP allele PCR methods can be used in conjunction with three primer systems and index sequencing. For example, the method may involve using multiple PCR (MASSARRAY system), and the primers used include the first capture probe in some site-specific forward PCR primers, and the adapter sequence is included in the site-specific reverse PCR primers, thereby producing amplicon, and then secondary PCR includes reverse capture sequence and molecular index barcode, for using, for example, the sequencing of the Illumina MPSS system. In some embodiments, multiple SNP allele PCR method can be used in conjunction with four primer systems and index sequencing.For example, the method can relate to and use multiple PCR (MASSARRAY system), primer used is incorporated into the adapter sequence into the forward and site-specific reverse PCR primer of site, and then secondary PCR is incorporated into the forward and reverse capture sequence and molecular index bar code, for using the order-checking of for example Illumina MPSS system.In some embodiments, microfluidic method can be used.In some embodiments, microfluidic method based on array can be used.For example, the method can relate to and use microfluidic array (such as Fluidigm) for low weight amplification and be incorporated into index and capture probe, then order-checking.In some embodiments, emulsion microfluidic method can be used, for example digital droplet PCR.

[0128] In certain embodiments, a universal amplification method can be used (e.g., using universal or non-site-specific amplification primers). In some embodiments, the universal amplification method can be used in conjunction with a pull-down method. In some embodiments, the method can include pulling down a biotinylated ultramer from a universal amplification sequence library (e.g., a biotinylated pull-down assay from Agilent or IDT). For example, the method can involve preparing a standard library, enriching a selected region by a pull-down assay, and a secondary universal amplification step. In certain embodiments, the pull-down method can be used in conjunction with a connection-based method. In certain embodiments, the method can include pulling down a biotinylated ultramer connected with a sequence-specific adapter (e.g., HALOPLEX PCR, HaloGenomics). For example, the method can involve using a selector probe to capture a restriction enzyme-digested fragment, then connecting the captured product and adapter, and universal amplification followed by sequencing. In certain embodiments, the pull-down method can be used in conjunction with an extension and connection-based method. In certain embodiments, the method can include molecular inversion probe (MIP) extension and connection. For example, the method can involve the use of a molecular inversion probe in combination with a sequence adapter, followed by universal amplification and sequencing. In certain embodiments, complementary DNA can be synthesized and sequenced without amplification.

[0129] In certain embodiments, the extension and ligation method can be performed without the need for pull-down components. In certain embodiments, the method can include site-specific forward and reverse primer hybridization, extension, and ligation. The method can also include universal amplification or complementary DNA synthesis without the need for amplification followed by sequencing. In certain embodiments, the method can reduce or eliminate background sequences during analysis.

[0130] In certain embodiments, the pull-down method can be used together with or without an optional amplification component. In certain embodiments, the method can include a pull-down test and a connection of modification, which is fully incorporated into the capture probe without the need for universal amplification. For example, the method can involve capturing restriction enzyme-digested fragments using a modified selector probe, then connecting the capture product and an adapter, and optionally amplifying, and sequencing. In certain embodiments, the method can include a biotinylated pull-down test, and a combination of extending and connecting an adapter sequence and being connected with a loop single-stranded sequence. For example, the method can involve capturing a region of interest (i.e., a target sequence), extending the probe, connecting the adapter, connecting a single-stranded loop, optionally amplifying, and sequencing. In certain embodiments, the analysis of sequencing results can separate target sequence and background.

[0131] In some embodiments, nucleic acid enrichment is performed on a fragment of a selected genomic region (e.g., chromosome) using one or more sequence-based separation methods described herein. Sequence-based separation is typically based on the presence of nucleotide sequences (e.g., target fragments and / or reference fragments) in the fragment of interest in the sample that are substantially absent in other fragments or that do not contain substantial amounts (e.g., 5% or less) of other fragments. In some embodiments, sequence-based separation can generate separated target fragments and / or separated reference fragments. Separated target fragments and / or separated reference fragments are typically separated from the remaining fragments in the nucleic acid sample. In certain embodiments, separated target fragments and separated reference fragments can also be separated from each other (e.g., separated in separate test compartments). In certain embodiments, separated target fragments and separated reference fragments can be separated together (e.g., separated in the same test chamber). In some embodiments, unbound fragments can be differentially removed or degraded or digested.

[0132] In some embodiments, a selective nucleic acid capture method is used to separate target fragments and / or reference fragments from a nucleic acid sample. Commercially available nucleic acid capture systems include, for example, the Nimblegen sequence capture system (Roche NimbleGen, Madison, WI); the Illumina BEADARRAY platform (Illumina, San Diego, CA); the Affymetrix GENECHIP platform (Affymetrix, Santa Clara, CA); the Agilent SureSelect target enrichment system (Agilent Technologies, Santa Clara, CA); and related platforms. The method generally involves hybridization of a capture oligonucleotide to a segment or all of the nucleotide sequence of a target fragment or reference fragment and can include the use of a solid phase (e.g., a solid phase array) and / or a solution-based platform. The capture oligonucleotide (sometimes referred to as "bait") can be selected or designed so that it preferentially hybridizes to a nucleic acid fragment of a selected genomic region or site (e.g., one of chromosomes 21, 18, 13, X, or Y, or a reference chromosome). In certain embodiments, hybridization-based methods (e.g., using oligonucleotide arrays) can be used to enrich nucleic acid sequences from certain chromosomes (e.g., potentially aneuploid chromosomes, reference chromosomes, or other chromosomes of interest) or segments of interest thereof.

[0133] In some embodiments, nucleic acid is enriched for a specific nucleic acid fragment length, a range of lengths, a length below or above a specific threshold or cutoff value using one or more length-based separation methods. Nucleic acid fragment length generally refers to the number of nucleotides in a fragment. Nucleic acid fragment length sometimes also refers to nucleic acid fragment size. In some embodiments, length-based separation methods do not require measuring the length of individual fragments. In some embodiments, length-based separation methods are combined with methods for determining the length of individual fragments. In some embodiments, length-based separation refers to size fractionation, wherein all or part of the fractionated libraries can be separated (e.g., retained) and / or analyzed. Size fractionation is known in the art (e.g., array separation, molecular sieve separation, gel electrophoresis separation, column chromatography separation (e.g., size exclusion column) and microfluidic-based methods). In certain embodiments, length-based separation methods may include, for example, fragment cyclization, chemical treatment (e.g., formaldehyde, polyethylene glycol (PEG)), mass spectrometry, and / or size-specific nucleic acid amplification.

[0134] Certain length-based separation methods that can be used with the methods of the present invention use, for example, selective sequence tagging. The term "sequence tagging" refers to the incorporation of an identifiable unique sequence into a nucleic acid or population of nucleic acids. The term "sequence tagging" as used herein is distinct from the term "sequence tag" as used later herein. In this sequence tagging method, nucleic acids of varying size (e.g., short fragments) in a sample comprising both long and short nucleic acids undergo selective sequence tagging. The method typically involves performing a nucleic acid amplification reaction using a nested primer set comprising inner and outer primers. In certain embodiments, one or both of the inner primers can be tagged to introduce a tag onto the target amplification product. The outer primers typically do not anneal to the short fragments bearing the (inner) target sequence. The inner primers can anneal to the short fragments and produce an amplification product bearing both the tag and the target sequence. Typically, tagging of long fragments is inhibited by combinatorial mechanisms, including, for example, blocked extension of the inner primers due to prior annealing and extension of the outer primers. Enrichment of tagged fragments can be achieved by any of a variety of methods, including, for example, exonuclease digestion of single-stranded nucleic acids and amplification of the tagged fragments using amplification primers specific for at least one tag.

[0135] Other length-based separation methods that can be used with the methods of the present invention involve subjecting the nucleic acid sample to polyethylene glycol (PEG) precipitation. Examples of methods include those described in International Patent Application Publication Nos. WO2007 / 140417 and WO2010 / 115016. The method generally requires contacting the nucleic acid sample with PEG in the presence of one or more monovalent salts under conditions sufficient to substantially precipitate large nucleic acids without substantially precipitating small (e.g., less than 300 nucleotides) nucleic acids.

[0136] Other size-based enrichment methods that can be used with the methods described herein involve cyclization by ligation, for example using a cyclase. Short nucleic acid fragments can generally be cyclized more efficiently than long fragments. Non-cyclized sequences can be separated from the cyclized sequences, and the enriched short fragments can be used for further analysis.

[0137] Nucleic acid library

[0138] In some embodiments, nucleic acid library is a variety of polynucleotide molecules (e.g., nucleic acid samples) prepared, assembled, and / or modified for a specific process, the non-limiting examples of which are included in solid phase (e.g., solid supports, e.g., flow cells, beads) for fixation, enrichment, amplification, cloning, detection, and / or for nucleic acid sequencing. In certain embodiments, nucleic acid library is prepared before or during a sequencing process. Nucleic acid library (e.g., sequencing library) can be prepared using suitable methods known in the art. Nucleic acid library can be prepared by targeted or non-targeted preparation processes.

[0139] In some embodiments, the nucleic acid library is modified to include a chemical moiety (e.g., a functional group) configured for immobilizing the nucleic acid to a solid support. In some embodiments, the nucleic acid library is modified to include a biomolecule (e.g., a functional group) and / or a binding pair member configured for immobilizing the library to a solid support, non-limiting examples of which include thyroxine-binding globulin, steroid binding protein, antibody, antigen, hapten, enzyme, hemagglutinin, nucleic acid, inhibitor, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding protein, receptor, carbohydrate, oligonucleotide, polynucleotide, complementary nucleic acid sequence, etc. and combinations thereof. Some examples of specific binding pairs include, but are not limited to: an avidin portion and a biotin portion; an antigenic epitope and an antibody or immunologically active fragment thereof; an antibody and a hapten; a digoxigenin portion and an anti-digoxigenin antibody; a fluorescein portion and an anti-fluorescein antibody; an operator and an inhibitor; a nuclease and a nucleoside; a lectin and a polysaccharide; a steroid and a steroid-binding protein; an active compound and an active compound receptor; a hormone and a hormone receptor; an enzyme and a substrate; an immunoglobulin and protein A; an oligonucleotide or polynucleotide and its corresponding complement; the like or a combination thereof.

[0140] In some embodiments, the nucleic acid library is modified to include one or more polynucleotides of known composition, non-limiting examples of which include identifiers (e.g., tags, index tags), capture sequences, tag adapters, restriction enzyme sites, promoters, enhancers, replication origins, stem loops, complementary sequences (e.g., primer binding sites, annealing sites), suitable integration sites (e.g., transposons, viral integration sites), modified nucleotides, etc., or combinations thereof. Polynucleotides of known sequence can be added to appropriate positions, such as the 5' end, the 3' end, or inside the nucleic acid sequence. Polynucleotides of known sequence can be the same or different sequences. In some embodiments, the known sequence polynucleotides are configured to hybridize with one or more oligonucleotides fixed on a surface (e.g., the surface of a flow cell). For example, the 5' known sequence of a nucleic acid molecule can hybridize with a first plurality of oligonucleotides, and the 3' known sequence can hybridize with a second plurality of oligonucleotides. In some embodiments, the nucleic acid library may include chromosome-specific tags, capture sequences, tags, and / or adapters. In some embodiments, the nucleic acid library includes one or more detectable labels. In some embodiments, one or more detectable labels can be incorporated into the 5' end, the 3' end, and / or any nucleotide position of the nucleic acid in the library. In some embodiments, the nucleic acid library comprises hybridized oligonucleotides. In certain embodiments, the hybridized oligonucleotides are label probes. In some embodiments, before being fixed on a solid phase, the nucleic acid library comprises hybridized oligonucleotide probes.

[0141] In some embodiments, the polynucleotide of known sequence includes universal sequence.Universal sequence is the specific nucleotide sequence that is integrated into two or more nucleic acid molecules or two or more nucleic acid molecule subgroups, wherein the universal sequence is identical for all molecules or molecule subgroups that it is integrated into.Universal sequence is usually designed to use a single universal primer complementary to the universal sequence to hybridize and / or amplify multiple different sequences. In some embodiments, two (such as a pair) or more universal sequences and / or universal primers are used.Universal primer usually includes universal sequence. In some embodiments, adapter (such as universal adapter) includes universal sequence. In some embodiments, one or more universal sequences are used to capture, identify and / or detect multiple nucleic acid substances or its subgroup.

[0142] In certain embodiments of preparing nucleic acid libraries (e.g., in certain sequencing by synthesis procedures), the nucleic acids are size selected and / or fragmented to a length of a few hundred base pairs or less (e.g., in library generation preparation). In some embodiments, library preparation is performed without fragmentation (e.g., when using ccfDNA).

[0143] In certain embodiments, a library preparation method based on connection is used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego CA). The library preparation method based on connection is typically designed using an adapter (e.g., methylated adapter), which can be incorporated into an index sequence at the initial connection step and can typically be used to prepare samples for single read sequencing, paired end sequencing, and multiple sequencing. For example, nucleic acid (e.g., fragmented nucleic acid or ccfDNA) is sometimes end-repaired by filling in (fill-in) reaction, endonuclease reaction, or a combination thereof. In some embodiments, the resulting blunt end repair nucleic acid can then be extended with a single nucleotide that is complementary to the single nucleotide at the 3' end of the adapter / primer. Any nucleotide can be used for extension / protruding nucleotides. In some embodiments, the nucleic acid library preparation includes connecting an adapter oligonucleotide. The adapter oligonucleotide is typically complementary to a flow cell anchor and is sometimes used to fix the nucleic acid library to a solid support, such as the inner surface of a flow cell. In some embodiments, the adapter oligonucleotide includes an identifier, one or more sequencing primer hybridization sites (e.g., a sequence complementary to a universal sequencing primer, a single-end sequencing primer, a paired-end sequencing primer, a multiplex sequencing primer, etc.), or a combination thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing).

[0144] The identifier can be a suitable detectable label for incorporating or engaging nucleic acid (e.g., polynucleotide) that allows detection and / or identification of nucleic acids comprising the identifier. In some embodiments, the identifier is incorporated into or engaged with nucleic acid (e.g., by polymerase) during sequencing methods. The non-limiting examples of the identifier include nucleic acid tags, nucleic acid indexes or barcodes, radiolabels (e.g., isotopes), metal labels, chemiluminescent labels, phosphorescent labels, fluorescent quenchers, dyes, proteins (e.g., enzymes, antibodies, or portions thereof, connexons, binding pairs), etc., or a combination thereof. In some embodiments, the identifier (e.g., nucleic acid indexes or barcodes) is a unique, known, and / or identifiable sequence of nucleotides or nucleotide analogs. In some embodiments, the identifier is six or more continuous nucleotides. Many fluorophores with various excitation and emission spectra are available. Any suitable type and / or quantity of fluorophores can be used as the identifier. In some embodiments, 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more or 50 or more different identifiers are used for methods described herein (e.g., nucleic acid detection and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are connected to each nucleic acid in the library. Identifier detection and / or quantification can be performed by suitable methods or devices, non-limiting examples of which include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, illuminometers, fluorimeter, spectrophotometer, suitable gene chip or microarray analysis, Western blotting, mass spectrometry, chromatography, cell fluorescence analysis, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning flow cytometry, affinity chromatography, manual batch mode separation, electric field suspension, suitable nucleic acid sequencing methods and / or nucleic acid sequencing devices, etc. and combinations thereof.

[0145] In some embodiments, a transposon-based library preparation method is used (e.g., EPICENTRE NEXTERA, Epicentre, Madison WI). Transposon-based methods typically use in vitro transposition to similar fragments or tag DNA (often allowing for the incorporation of platform-specific tags and optional barcodes) in a single-tube reaction and prepare sequencer-ready libraries.

[0146] In some embodiments, the nucleic acid library or a portion thereof is amplified (e.g., by a PCR-based method). In some embodiments, the sequencing method includes amplifying the nucleic acid library. The nucleic acid library can be amplified before or after being fixed to a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification includes a process of amplifying or increasing the number of nucleic acid templates and / or their complements present (e.g., in a nucleic acid library) by generating one or more copies of the templates and / or their complements. Amplification can be performed by a suitable method. The nucleic acid library can be amplified by thermal cycling or by isothermal amplification. In some embodiments, a rolling circle amplification method is used. In some embodiments, amplification occurs on a solid support (e.g., within a flow cell) where the nucleic acid library or a portion thereof is fixed. In certain sequencing methods, the nucleic acid library is added to a flow cell and fixed by hybridization with an anchor under suitable conditions. Such nucleic acid amplification is generally referred to as solid phase amplification. In some embodiments of solid phase amplification, all or part of the amplified products are synthesized by extension from an immobilized primer. The solid phase amplification reaction is similar to standard solution phase amplification, except that at least one of the amplification oligonucleotides (e.g., primers) is fixed to a solid support.

[0147] In some embodiments, solid phase amplification includes nucleic acid amplification reaction, which includes only one oligonucleotide primer fixed on the surface. In certain embodiments, solid phase amplification includes a plurality of different immobilized oligonucleotide primer materials. In some embodiments, solid phase amplification may include nucleic acid amplification reaction, which includes a second different oligonucleotide primer in a solution and a kind of oligonucleotide primer fixed on a solid surface. A plurality of different immobilization or solution primers can be used. Non-limiting examples of solid phase nucleic acid amplification reactions include interface amplification, bridge amplification, emulsion PCR, WildFire amplification (e.g., U.S. patent application US20130012399) etc. or a combination thereof.

[0148] Sequencing

[0149] In some embodiments, nucleic acids (e.g., nucleic acid fragments, sample nucleic acids, cell-free nucleic acids) are sequenced. In certain embodiments, a complete sequence or a substantially complete sequence is obtained, and sometimes a partial sequence is obtained.

[0150] In some embodiments, some or all nucleic acids in a sample are enriched and / or amplified before or during sequencing (e.g., non-specifically, such as by PCR-based methods). In certain embodiments, specific nucleic acid portions or subsets in a sample are enriched and / or amplified before or during sequencing. In some embodiments, portions or subsets of a preselected nucleic acid set are randomly sequenced. In some embodiments, nucleic acids in a sample are not enriched and / or amplified before or during sequencing.

[0151] As used herein, a "read" (i.e., "a read," "sequence read") is a short nucleotide sequence generated by any sequencing method described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read"), or sometimes from both ends of a nucleic acid fragment (e.g., a paired-end read, a double-end read).

[0152] The length of a sequence read is generally related to the specific sequencing technology. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). For example, nanopore sequencing provides sequence reads that can vary in size from tens to hundreds to thousands of base pairs. In some embodiments, the sequence reads are the mean, median, average, or absolute length of about 15 bp to about 900 bp in length. In certain embodiments, the sequence reads are the mean, median, average, or absolute length of about 1000 bp or longer.

[0153] In some embodiments, the nominal, average, mean or absolute length of a single-end read is sometimes about 15 contiguous nucleotides to about 50 or more contiguous nucleotides, sometimes about 15 contiguous nucleotides to about 40 or more contiguous nucleotides, and sometimes about 15 contiguous nucleotides or about 36 or more contiguous nucleotides. In certain embodiments, the nominal, average, mean or absolute length of a single-end read is about 20 to about 30 bases, or about 24 to about 28 bases. In certain embodiments, the nominal, average, mean or absolute length of a single-end read is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 28 or about 29 bases.

[0154] In certain embodiments, the nominal, average, mean or absolute length of paired-end reads is sometimes about 10 contiguous nucleotides to about 25 contiguous nucleotides or more (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotides long or more), about 15 contiguous nucleotides to about 20 contiguous nucleotides or more, and sometimes about 17 contiguous nucleotides or about 18 contiguous nucleotides.

[0155] Readings are typically representations of nucleotide sequences in physiological nucleic acids. For example, ATGC is used in readings to describe sequences, where "A" represents adenine nucleotides, "T" represents thymine nucleotides, "G" represents guanine nucleotides, and "C" represents cytosine nucleotides. Sequence readings obtained from the blood of a pregnant woman may be readings of a mixture of fetal and maternal nucleic acids. A mixture of relatively short reads can be transformed into a representation of genomic nucleic acids in a pregnant woman and / or fetus by the methods described herein. A mixture of relatively short reads can be transformed into representations of, for example, copy number variations (e.g., maternal and / or fetal copy number variations), genetic variations or aneuploidy, microduplications, or microdeletions. Readings of a mixture of maternal and fetal nucleic acids can be transformed into a composite chromosome or segment thereof representing features of one or both of the maternal and fetal chromosomes. In certain embodiments, "obtaining" nucleic acid sequence readings from a subject sample, and / or "obtaining" nucleic acid sequence readings from biological samples of one or more reference individuals can directly involve sequencing nucleic acids to obtain sequence information. In some embodiments, "obtaining" can involve receiving sequence information directly obtained from other nucleic acids.

[0156] In some embodiments, a representative component of a genome is sequenced and is sometimes referred to as "coverage" or "fold coverage." For example, 1-fold coverage indicates that approximately 100% of the nucleotide sequence of a genome is represented by reads. In some embodiments, "fold coverage" is a related term using a previous sequencing run as a reference. For example, a second round of sequencing may have 2-fold less coverage than a first round of sequencing. In some embodiments, a genome is sequenced with redundancy, wherein a given region of the genome is covered by two or more reads or overlapping reads (e.g., greater than 1 "fold coverage," e.g., 2-fold coverage).

[0157] In some embodiments, to a kind of nucleic acid sample order-checking from an individual.In certain embodiments, each nucleic acid of two or more samples is checked, wherein sample is from an individual or from different individuals.In certain embodiments, collect nucleic acid samples from two or more biological samples (wherein each biological sample is from an individual or two or more individuals), and to this set order-checking.In the embodiment of back, often identify the nucleic acid sample from each biological sample by one or more unique identification things.

[0158] In some embodiments, sequencing methods employ identifiers that allow for multiple sequence reactions during sequencing. The greater the number of unique identifiers, the greater the number of samples and / or chromosomes detected, e.g., multiple sequencing runs can be performed. The sequencing run can be performed using any suitable number of unique identifiers (e.g., 4, 8, 12, 24, 48, 96, or more).

[0159] The sequencing process sometimes uses a solid phase, and sometimes the solid phase includes a flow cell, on which nucleic acids from the library can be joined and reagents can flow and contact the joined nucleic acids. The flow cell sometimes includes a flow cell channel, and the use of an identifier can facilitate the analysis of the number of samples in each channel. A flow cell is any solid support that can be constructed to retain and / or allow reagent solutions to pass through the binding analyte in an orderly manner. The flow cell is usually planar, optically transparent, usually at the millimeter or submillimeter level, and often has a channel or passage in which the interaction of the analyte / reagent occurs. In some embodiments, the number of samples that can be analyzed in a given flow cell channel often depends on the number of unique identifiers used in library preparation and / or probe design. Single flow cell channel. Multiple use of 12 identifiers, for example, allows 96 samples (such as the number of holes in a 96-well microplate) to be analyzed simultaneously in 8 channel flow cells. Similarly, multiple use of 48 identifiers, for example, allows 384 samples (such as the number of holes in a 384-well microplate) to be analyzed simultaneously in 8 channel flow cells. Non-limiting examples of commercially available multiplex sequencing kits include Illumina's Multiplex Sample Preparation Oligonucleotide Kit and Multiplex Sequencing Primer and PhiX Control Kit (eg, Illumina catalog numbers PE-400-1001 and PE-400-1002, respectively).

[0160] Any suitable method for sequencing nucleic acids can be used, non-limiting examples of which include Maxim & Gilbert, chain termination methods, synthesis sequencing, ligation sequencing, mass spectrometry sequencing, microscope-based techniques, etc., or combinations thereof. In some embodiments, first generation sequencing techniques, such as Sanger sequencing methods, including automatic Sanger sequencing methods (including microfluidics Sanger sequencing), can be used for the methods of the present invention. In some embodiments, other sequencing techniques, such as transmission electron microscopy (TEM) and atomic force microscopy (AFM), including nucleic acid imaging techniques, are also used herein. In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods typically involve clonal amplification of DNA templates or single DNA molecules that are sometimes sequenced in a flow cell in a massively parallel manner. Next generation (e.g., second and third generation) sequencing technologies (capable of sequencing DNA in a massively parallel manner) can be used for the methods described herein and are collectively referred to herein as "massive parallel sequencing" (MPS). In some embodiments, MPS sequencing methods employ targeted approaches, wherein specific chromosomes, genes, or regions of interest are sequences. In certain embodiments, non-targeted approaches are used, wherein most or all nucleic acids in a sample are sequenced, amplified, and / or randomly captured.

[0161] In some embodiments, targeted enrichment, amplification and / or sequencing methods are used. Targeted methods are typically used to separate, select and / or enrich nucleic acid subsets in a sample for further processing by sequence-specific oligonucleotides. In some embodiments, a library of sequence-specific oligonucleotides is used to target (e.g., hybridize) one or more nucleic acid groups in a sample. Sequence-specific oligonucleotides and / or primers are typically selective for specific sequences (e.g., unique nucleic acid sequences) present in one or more chromosomes, genes, exons, introns and / or regulatory regions of interest. Any suitable method or combination of methods can be used for enrichment, amplification and / or sequencing one or more target nucleic acid subsets. In some embodiments, one or more sequence-specific anchors are used to separate and / or enrich target sequences by being captured to a solid phase (e.g., flow cell, beads). In some embodiments, sequence-specific primers and / or primer sets are used to enrich and / or amplify target sequences based on a polymerase method (e.g., based on a PCR-method, by any suitable polymerase-based extension). Sequence-specific anchors can typically be used as sequence-specific primers.

[0162] MPS sequencing sometimes uses sequencing by synthesis and certain imaging methods. The nucleic acid sequencing technologies that can be used in the methods described herein are sequencing by synthesis and sequencing based on reversible terminators (such as Illumina's Genome Analyzer and Genome Analyzer II; HISEQ 2000; HISEQ2500 (Illumina, San Diego CA)). This technology can be used to sequence millions of nucleic acid (such as DNA) fragments in parallel. In one embodiment of this sequencing technology, a flow cell comprising an optically transparent slide with 8 separate channels is used, and the surface of the flow cell is bound to oligonucleotide anchors (such as adapter primers). The flow cell is typically a solid support that can be constructed to retain and / or provide for the orderly passage of reagent solutions through the bound analyte. The flow cell is typically planar, optically transparent, typically at the millimeter or submillimeter level, and often has channels or pathways in which the analyte / reagent interaction occurs.

[0163] In some embodiments, sequencing by synthesis includes repeatedly adding (e.g., by covalent addition) nucleotides to a primer or a pre-existing nucleic acid chain in a template-guided manner. Each repeatedly added nucleotide is detected and the process is repeated multiple times until the sequence of the nucleic acid chain is obtained. The length of the sequence obtained depends in part on the number of addition and detection steps performed. In some embodiments of sequencing by synthesis, one, two, three or more nucleotides of the same type (e.g., A, G, C or T) are added and detected in the nucleotide addition round. Nucleotides can be added by any suitable method (e.g., enzyme or chemistry). For example, in some embodiments, a polymerase or ligase adds nucleotides to a primer or a pre-existing nucleic acid chain in a template-guided manner. In some embodiments of sequencing by synthesis, different types of nucleotides, nucleotide analogs and / or identifiers are used. In some embodiments, reversible terminators and / or removable (e.g., shearable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In certain embodiments, sequencing by synthesis includes cutting (e.g., cutting and removing identifiers) and / or cleaning steps. In some embodiments, the addition of one or more nucleotides is detected by a suitable method described herein or known in the art, non-limiting examples of which include any suitable imaging device, a suitable camera, a digital camera, a CCD (charge coupled device)-based imaging device (e.g., a CCD camera), a CMOS (complementary metal oxide semiconductor)-based imaging device (e.g., a CMOS camera), a photodiode (e.g., a photomultiplier tube), an electron microscope, a field effect transistor (e.g., a DNA field effect transistor), an ISFET ion sensor (e.g., a CHEMFET sensor), the like, or a combination thereof. Other sequencing methods that can be used to perform the methods described herein include digital PCR and sequencing by hybridization.

[0164] Other sequencing methods that can be used to carry out the methods described herein include digital PCR and hybridization sequencing. Digital polymerase chain reaction (digital PCR or dPCR) can be used to directly identify and quantify nucleic acids in samples. In some embodiments, digital PCR can be performed in an emulsion. For example, individual nucleic acids are separated in, for example, a microfluidic device and each nucleic acid is amplified separately by PCR. The isolated nucleic acid is such that no more than one nucleic acid is present in each well. In some embodiments, different probes can be used to distinguish multiple alleles (e.g., fetal alleles and maternal alleles). Alleles can be counted to determine copy number.

[0165] In some embodiments, hybridization sequencing can be used. Said method relates to making multiple polynucleotide sequences contact multiple polynucleotide probes, wherein said multiple polynucleotide probes are optionally connected to substrates. In some embodiments, said substrate can be a plane with a known nucleotide sequence array. The polynucleotide sequence present in the sample can be determined using a pattern of array hybridization. In some embodiments, each probe is connected to a bead (such as a magnetic bead etc.). Hybridization with said bead can be identified and used to identify the multiple polynucleotide sequences in the sample.

[0166] In some embodiments, nanopore sequencing can be used in the methods described herein. Nanopore sequencing is a single-molecule sequencing technology whereby a single nucleic acid molecule (such as DNA) is directly sequenced as it passes through a nanopore.

[0167]

[00146] A suitable MPS method, system or technology platform for performing the methods described herein can be used to obtain nucleic acid sequencing reads. Non-limiting examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, Helicos True single molecule sequencing, Ion Torrent and Ion semiconductor-based sequencing (e.g., developed by Life Technologies), WildFire, 5500, 5500xl W and / or 5500xl W genetic analyzer-based technologies (e.g., developed and sold by Life Technologies, U.S. patent application US20130012399); Polony sequencing, Pyro sequencing, massively parallel signature sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen systems and methods, nanopore-based platforms, chemically sensitive field-effect transistor (CHEMFET) arrays, electron microscopy-based sequencing (e.g., developed by ZS Genetics, Halcyon Molecular), and nanoball sequencing.

[0168] In some embodiments, chromosome-specific sequencing is performed. In some embodiments, chromosome-specific sequencing is performed using DANSR (digital analysis of selected regions). Digital analysis of selected regions can quantify hundreds of loci simultaneously through cfDNA-dependent linkage of two position-specific oligonucleotides, using interfering 'bridge' oligonucleotides to form a PCR template. In some embodiments, chromosome-specific sequencing is performed by generating a library enriched for chromosome-specific sequences. In some embodiments, sequence reads are obtained only for a selected set of chromosomes. In some embodiments, sequence reads are obtained only for chromosomes 21, 18, and 13.

[0169] Mapping reads

[0170] Sequence reads can be mapped and the number of reads mapped to a specific nucleic acid region (e.g., a chromosome, portion, or segment thereof) is referred to as a count. Any suitable mapping method (e.g., a process, algorithm, program, software, module, etc., or a combination thereof) can be used. Certain aspects of the mapping method are described below.

[0171] Mapping nucleotide sequence reads (i.e., sequence information for fragments whose physical genomic loci are unknown) can be performed in a variety of ways, which generally include aligning the obtained sequencing reads with matching sequences in a reference genome. In such alignments, sequence reads are generally aligned with a reference sequence, and those that have been aligned are referred to as "mapped," "mapped sequence reads," or "mapped reads." In certain embodiments, mapped sequence reads are referred to as "hits" or "counts." In some embodiments, mapped sequence reads are grouped together and assigned to specific portions according to various parameters, as described in detail below.

[0172] As used herein, the terms "alignment" and "alignment" refer to two or more nucleic acid sequences that can be identified as matching (e.g., 100% identity) or partial matching. The alignment can be performed manually or by computer (e.g., software, program, module, or algorithm), non-limiting examples of which include the Efficient Local Alignment of Nucleotide Data (ELAND) computer program, which is part of the Illumina genome analysis process. The alignment of sequence reads can be a 100% sequence match. In some cases, the alignment is less than a 100% sequence match (i.e., a non-perfect match, a partial match, a partial alignment). In some embodiments, the alignment is about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76% or 75% match. In some embodiments, the alignment includes mismatches. In some embodiments, the alignment includes 1, 2, 3, 4, or 5 mismatches. Two or more sequences can be aligned using either strand. In certain embodiments, a nucleic acid sequence is aligned with the reverse complement of another nucleic acid sequence.

[0173] Various computer methods can be used to map each sequence read to a portion. Non-limiting examples of computer algorithms that can be used for aligning sequences include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE 1, BOWTIE 2, ELAND, MAQ, PROBEMATCH, SOAP or SEQMAP or variants thereof or combinations thereof. In some embodiments, sequence reads can be aligned with sequences in a reference genome. In some embodiments, sequence reads can be obtained from nucleic acid databases known in the art and / or aligned with sequences therein, including, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory) and DDBJ (DNA Database of Japan). BLAST or similar tools can be used to search for identical sequences in sequence databases. Then, for example, search hits can be used to sort identical sequences into appropriate portions (as described below).

[0174] In some embodiments, mapped sequence reads and / or information associated with mapped sequence reads are stored and / or evaluated in a suitable computer-readable format on a non-transitory computer-readable medium. "Computer-readable format" sometimes refers to a format herein. In some embodiments, mapped sequence reads are stored and / or evaluated in a suitable binary format, text format, etc., or a combination thereof. A binary format is sometimes a BAM format. A text format is sometimes a sequence alignment / map (SAM) format. Non-limiting examples of binary or text formats include BAM, SAM, SRF, FASTQ, Gzip, etc., or a combination thereof. In some embodiments, mapped sequence reads are stored and / or converted to a format that requires less storage space (e.g., fewer bytes) than a traditional format (e.g., a SAM format or a BAM format). In some embodiments, mapped sequence reads in a first format are compressed into a second format that requires less storage space than the first. As used herein, the term "compression" refers to a process of data compression, source encoding, and / or bit rate reduction in which the size of a computer-readable data file is reduced. In some embodiments, the sequence reads of the mapping are compressed into a binary format from a SAM format. After file compression, some data are sometimes lost. Sometimes the compression process does not lose data. In some file compression embodiments, some data are replaced with the index and / or reference of another data file, and the other data file comprises the information related to the sequence reads of the mapping. In some embodiments, the sequence reads of the mapping are stored in a binary format, including or consisting of the following: read count, chromosome identifier (such as the chromosome mapped by the identification reading) and chromosome position identifier (such as the part on the chromosome mapped by the identification reading). In some embodiments, the binary format includes a 20-byte array, a 16-byte array, an 8-byte array, a 4-byte array or a 2-byte array. In some embodiments, the reading information of the mapping is stored in an array with a 10-byte format, a 9-byte format, an 8-byte format, a 7-byte format, a 6-byte format, a 5-byte format, a 4-byte format, a 3-byte format, or a 2-byte format. Sometimes the data readings of the mapping are stored in a 4-byte array, including a 5-byte format. In some embodiments, the binary format includes a 5-byte format, including a 1-byte chromosome ordinal and a 4-byte chromosome portion. In some embodiments mapped reads are stored in a compressed binary format that is about 100-fold, about 90-fold, about 80-fold, about 70-fold, about 60-fold, about 55-fold, about 50-fold, about 45-fold, about 40-fold or about 30-fold smaller than a sequence alignment / map (SAM) format. In some embodiments mapped reads are stored in a compressed binary format that is about 2-fold to about 50-fold (e.g., about 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6 or about 5-fold smaller than a GZip format.

[0175] In some embodiments, a system includes a compression module (e.g., 4, Figure 42A In some embodiments, mapped sequence read information stored in a computer-readable format on a non-transitory computer-readable medium is compressed by a compression module. A compression module sometimes converts mapped sequence reads to or from a suitable format. In some embodiments, a compression module can accept mapped sequence reads in a first format (e.g., 1, Figure 42A ), converting it to a compressed format (e.g., a binary format, 5) and transferring the compressed reads to another module (e.g., an offset density module 6). The compression module typically provides sequence reads in a binary format 5 (e.g., a BReads format). Non-limiting examples of compression modules include GZIP, BGZF, and BAM, etc. or variations thereof).

[0176] The following example uses Java to convert an integer into a 4-byte array:

[0177] public static final byte[]

[0178] convertToByteArray(int value)

[0179] {

[0180] return new byte[]{

[0181] (byte)(value>>>24),

[0182] (byte)(value>>>16),

[0183] (byte)(value>>>8),

[0184] (byte)value};

[0185] }

[0186] In some embodiments, a reading can be uniquely or non-uniquely mapped to a portion in a reference genome. If a reading is compared with a single sequence in a reference genome, it is referred to as a "unique mapping." If a reading is compared with two or more sequences in a reference genome, it is referred to as a "non-unique mapping." In some embodiments, the reading of a non-unique mapping is removed from further analysis (e.g., quantitative). In certain embodiments, some small degrees of mismatch (0-1) may indicate that a single nucleic acid polymorphism may exist between the reference genome and the mapped reading from an individual sample. In some embodiments, there is no mismatch to allow a reading to be mapped to a reference sequence.

[0187] As used herein, the term "reference genome" may refer to any part or all of a specifically known, sequenced, or characterized genome of any organism or virus that can be used to reference and identify subject sequences. For example, reference genomes for human subjects and many other organisms are available from the National Center for Biotechnology Information, website ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus expressed in a nucleic acid sequence. As used herein, a reference sequence or reference genome is often an assembled or partially assembled genome sequence from an individual or multiple individuals. In some embodiments, a reference genome is an assembled or partially assembled genome sequence from one or more human individuals. In some embodiments, a reference genome includes sequences assigned to chromosomes.

[0188] In certain embodiments, when sample nucleic acid is from a pregnant female, sometimes a reference sequence is not from the fetus, the mother of the fetus, or the father of the fetus, thereby being referred to herein as an "external reference." A maternal reference can be prepared and used in some embodiments. When a reference from a pregnant female is prepared based on an external reference ("maternal reference sequence"), the readings of the DNA from the pregnant female that are substantially free of fetal DNA are typically mapped to an external reference sequence and assembled. In certain embodiments, the external reference is from the DNA of an individual who is substantially of the same race as the pregnant female. The maternal reference sequence may not completely cover maternal genomic DNA (e.g., approximately 50%, 60%, 70%, 80%, 90% or more of maternal genomic DNA may be covered), and the maternal reference may not perfectly match the maternal genomic DNA sequence (e.g., the maternal reference sequence may comprise multiple mismatches).

[0189] In certain embodiments, mappability is assessed for a genomic region (e.g., a portion, a genomic portion, a portion). Mappability is the ability of a nucleotide sequence read to clearly map to a portion of a reference genome, typically with up to a specific number of mismatches, including, for example, 0, 1, 2, or more mismatches. For a given genomic region, the expected mappability can be calculated using a sliding window method with a predetermined read length and averaged to the resulting read-level mappability value. Genomic regions that include stretches of unique nucleotide sequences sometimes have high mappability values.

[0190] Part

[0191] In some embodiments, the sequence reads (i.e., sequence tags) mapped are grouped together according to various parameters and assigned to specific portions (e.g., portions with reference to a genome). Typically, the sequence reads mapped by an individual can be used to identify portions present in a sample (e.g., the presence, absence, or content of a portion). In some embodiments, the content of a portion is an indicator of the content of a large sequence (e.g., a chromosome) in a sample. The term "portion" herein may also refer to a "genomic segment," "box," "region," "partition," "portion with reference to a genome," "portion of a chromosome," or "genomic portion." In some embodiments, a portion is a whole chromosome, a chromosome segment, a reference genome segment, a segment across multiple chromosomes, multiple chromosome segments, and / or a combination thereof. In some embodiments, a portion is predefined based on specific parameters. In some embodiments, a portion is arbitrarily defined based on the division of a genome (e.g., partitions based on size, GC content, continuous regions, continuous regions of arbitrarily defined size, etc.).

[0192] In some embodiments, portions are defined based on one or more parameters, including, for example, the length or specific features of a sequence. Portions can be selected, screened, and / or removed from consideration using any suitable criteria known in the art or described herein. In some embodiments, portions are based on a specific length of a genomic sequence. In some embodiments, a method can include analyzing sequence reads from multiple mappings of a plurality of portions. Portions can have approximately the same length or portions can have different lengths. In some embodiments, portions are approximately the same length. In some embodiments, portions of different lengths are adjusted or weighted. In some embodiments, portions are about 10 kilobases (kb) to about 100 kb, about 20 kb to about 80 kb, about 30 kb to about 70 kb, about 40 kb to about 60 kb, and sometimes about 50 kb. In some embodiments, portions are about 10 kb to about 20 kb. Portions are not limited to sequences that run continuously. Thus, portions can be composed of continuous and / or non-continuous sequences. Portions are not limited to single chromosomes. In some embodiments, a portion comprises all or part of a chromosome or all or part of two or more chromosomes. In some embodiments, a portion can span one, two or more complete chromosomes. In addition, a portion can span connected or unconnected regions of multiple chromosomes.

[0193] In some embodiments, a portion can be a specific chromosomal segment within a chromosome of interest, such as a chromosome for which a genetic variation (e.g., aneuploidy of chromosomes 13, 18, and / or 21, or sex chromosomes) is to be assessed. A portion can also be a pathogenic genome (e.g., a bacterial, fungal, or viral genome) or a fragment thereof. A portion can be a gene, a gene fragment, a regulatory sequence, an intron, an exon, etc.

[0194] In some embodiments, a genome (e.g., a human genome) is divided into parts based on the information content of a particular region. In some embodiments, dividing a genome can remove similar regions (e.g., identical or homologous regions or sequences) in the genome and retain only unique regions. The regions removed during division can be within a single chromosome or can span multiple chromosomes. In some embodiments, the divided genome is trimmed down and optimized for rapid comparison, typically allowing focus on unique identifiable sequences.

[0195] In some embodiments, the division can reduce the weight of similar regions. The process of reducing the weight of some regions will be described in detail later.

[0196] In some embodiments, genome can be divided into the region beyond chromosome based on the information generated in the context (context) of classification.For example, information content can be quantitatively analyzed using a p-value overview, measuring the significance of the specific genome position of confirmed normal and abnormal objects (such as euploid and triploid objects, respectively). In some embodiments, genome can be divided into the region beyond chromosome based on any other standard, such as, speed / convenience when comparing tags, GC content (such as high or low GC content), the uniformity of GC content, other measurements of sequence content (such as individual nucleotide fractions, pyrimidine or purine fractions, natural and non-natural nucleic acid fractions, methylated nucleotide fractions and CpG content), methylation status, dual melting temperature, the compliance of order-checking or PCR, the uncertainty value assigned to the individual part of reference genome and / or the targeted search of specific features.

[0197] A "segment" of a chromosome is typically a portion of a chromosome, and is typically a portion of a chromosome that is distinct from a portion. A chromosome segment is sometimes in a different region of the chromosome than a portion, sometimes does not share polynucleotides with a portion, and sometimes includes polynucleotides in a portion. A chromosome segment often includes more nucleotides than a portion (e.g., a segment sometimes includes a portion), and sometimes a chromosome segment includes fewer nucleotides than a portion (e.g., a segment sometimes is within a portion).

[0198] count

[0199] In some embodiments, sequence reads mapped or divided based on selected features or variables can be quantified to determine the number of reads mapped to one or more portions (e.g., reference genome portions). In certain embodiments, the quantification of sequence reads mapped to a portion is referred to as a count (e.g., a count). Typically, a count is associated with a portion. In certain embodiments, the counts of two or more portions (e.g., a group of portions) are mathematically processed (e.g., averaged, summed, normalized, etc., or a combination thereof). In some embodiments, counts are determined from some or all of the sequence reads mapped to (i.e., associated with) a portion. In certain embodiments, counts are determined from a predefined subset of mapped sequence reads. Any suitable characteristic or variable can be used to define or select a predefined subset of mapped sequence reads. In some embodiments, a predefined subset of mapped sequencing reads can include 1–n sequence reads, where n represents a number equal to the sum of all sequence reads generated from a test subject or reference subject sample.

[0200] In certain embodiments, counts are derived from sequence reads processed or processed by suitable methods, operations or mathematical processes known in the art. Counts can be determined by suitable methods, operations or mathematical processes. In certain embodiments, counts are derived from sequence reads associated with a portion, wherein some or all of the sequence reads are weighted, removed, filtered, standardized, adjusted, averaged (derived a mean), added or subtracted, or a combination thereof. In some embodiments, counts are derived from raw sequence reads and / or filtered sequence reads. In some embodiments, count values ​​are determined by a mathematical process. In certain embodiments, a count value is the average, arithmetic mean or sum of sequence reads mapped to a portion. Typically, a count is the arithmetic mean of multiple counts. In some embodiments, a count is associated with an uncertainty value.

[0201] In some embodiments, counts can be processed or transformed (e.g., normalized, combined, summed, screened, selected, averaged (derived a mean), the like, or a combination thereof). In some embodiments, counts can be transformed to produce normalized counts. Counts can be processed (e.g., normalized) by methods known in the art and / or described herein (e.g., portion-wise normalization, normalization for GC content, linear and nonlinear least squares regression, GC LOESS, LOWESS, PERUN, ChAI, RM, GCRM, cQn, and / or combinations thereof).

[0202] Count (such as original, screened and / or standardized count) can be processed and standardized to one or more levels.Level and overview are described in detail below.In certain embodiments, count is processed and / or standardized to reference level.Reference level is set forth below.Count (such as processing count) according to level processing can be associated with uncertainty value (such as calculating variance, error, standard deviation, Z-score, p-value, arithmetic mean absolute deviation etc.).In some embodiments, uncertainty value limits the scope that is higher than and lower than a certain level.Deviation value can replace uncertainty value, and the non-limiting example of deviation measurement includes standard deviation, mean absolute deviation, median absolute deviation, standard score (such as Z-score, Z-score, normal score, standardized variable) etc.

[0203] Counts are typically obtained from a nucleic acid sample from a pregnant female carrying a fetus. Counts of nucleic acid sequence reads mapped to one or more portions are typically representative of counts for both the fetus and the fetus' mother (e.g., a pregnant female subject). In certain embodiments, some counts mapped to a portion are from the fetal genome and some counts mapped to the same portion are from the maternal genome.

[0204] Data processing and standardization

[0205] Mapped sequence reads that have been counted are referred to herein as raw data because the data represent unprocessed counts (e.g., raw counts). In some embodiments, sequence read data in a data set can be further processed (e.g., mathematically and / or statistically processed) and / or displayed to help provide a result. In certain embodiments, a data set (including a larger data set) can benefit from preprocessing to help further analysis. Preprocessing of a data set sometimes involves removing redundant and / or uninformative portions or portions of a reference genome (e.g., portions with uninformative data or portions of a reference genome, redundant mapped reads, portions with a median count of 0, sequences that appear too frequently or too infrequently). Without being limited by theory, data processing and / or preprocessing can (i) remove noisy data, (ii) remove uninformative data, (iii) remove redundant data, (iv) reduce the complexity of a larger data set, and / or (v) help transform the data from one form into one or more other forms. When applied to data or data sets, the terms "preprocessing" and "processing" are collectively referred to herein as "processing." In some embodiments, processing can make data easier to be further analyzed, thereby generating results. In some embodiments, one or more or all processing methods (e.g., normalization methods, partial screening, mapping, confirmation, the like, or combinations thereof) are performed by a processor, a microprocessor, a computer, a device in conjunction with a memory and / or controlled by a microprocessor.

[0206] As used herein, the term "noisy data" refers to (a) data that have significant differences between data points when analyzed or plotted, (b) data with significant standard deviations (e.g., greater than 3 standard deviations), (c) data with significant standard errors of the mean, etc., and combinations thereof. Noisy data sometimes appear due to the quantity and / or quality of the starting material (e.g., nucleic acid samples), and sometimes appear as part of the method for preparing or replicating DNA for generating sequence reads. In certain embodiments, the noise comes from certain sequences that appear too frequently when prepared using PCR-based methods. The methods described herein can reduce or eliminate the base value of the noise data, thereby reducing the impact of the noise data on the results provided.

[0207] As used herein, the terms "uninformative data," "uninformative portion of a reference genome," and "uninformative portion" refer to a portion or data derived therefrom having a value that is significantly different from a predetermined threshold or falls outside a predetermined cutoff range. The term "threshold" herein refers to any number calculated using a qualified data set that serves as a limit for diagnosing a genetic variation (e.g., copy number variation, aneuploidy, microduplication, microdeletion, chromosomal abnormality, etc.). In certain embodiments, a threshold is exceeded by the results obtained by the methods of the present invention, and the subject is diagnosed as having a genetic variation (e.g., trisomy 21). In some embodiments, a threshold or range of values ​​is often calculated by mathematically and / or statistically processing sequence read data (e.g., from a reference and / or subject), and in certain embodiments, the sequence read data processed to generate the threshold or range of values ​​is sequence read data (e.g., from a reference and / or subject). In some embodiments, an uncertainty value is determined. The uncertainty value is typically a measure of variance or error and can be any suitable measure of variance or error. In some embodiments, an uncertainty value is a standard deviation, standard error, calculated variance, p-value, or mean absolute deviation (MAD). In some embodiments, the uncertainty value can be calculated according to the formula of Example 4.

[0208] Any suitable program can be used to process the data sets described herein. Non-limiting examples of methods suitable for processing data sets include filtering, standardization, weighting, monitoring peak height, monitoring peak area, monitoring peak edge, determining area ratio, mathematical processing of data, statistical processing of data, application of mathematical algorithms, analysis using fixed variables, analysis using optimized variables, mapping data to identify patterns or trends for other processing, etc., and combinations thereof. In some embodiments, data sets are processed according to different characteristics (such as GC content, redundant positioning readings, centromere regions, telomeric regions, etc., and combinations thereof) and / or variables (such as fetal sex, maternal age, maternal ploidy, fetal nucleic acid base value percentage, etc., and combinations thereof). In certain embodiments, processing data sets described herein can reduce the complexity and / or dimensionality of large data sets and / or complex data sets. Non-limiting examples of complex data sets include sequence reads generated by one or more test subjects and a variety of reference subjects of different ages and ethnic backgrounds. In some embodiments, a data set can comprise thousands to millions of sequence reads of each test subject and / or reference subject.

[0209] In some embodiments, data processing can be performed in any number of steps. For example, in some embodiments, data can be adjusted and / or processed using only a single processing method, while in some embodiments, data can be processed using 1 or more, 5 or more, 10 or more, or 20 or more processing steps (e.g., 1 or more processing steps, 2 or more processing steps, 3 or more processing steps, 4 or more processing steps, 5 or more processing steps, 6 or more processing steps, 7 or more processing steps, 8 or more processing steps, 9 or more processing steps, 10 or more processing steps, 11 or more processing steps, 12 or more processing steps, 13 or more processing steps, 14 or more processing steps, 15 or more processing steps, 16 or more processing steps, 17 or more processing steps, 18 or more processing steps, 19 or more processing steps, or 20 or more processing steps). In some embodiments, a processing step can be the same step repeated two or more times (e.g., filtering two or more times, normalizing two or more times), while in certain embodiments, a processing step can be two or more different processing steps performed simultaneously or sequentially (e.g., filtering, normalizing; normalizing, monitoring peak heights and edges; filtering, normalizing, normalizing to a reference, statistical processing to determine p-values, etc.). In some embodiments, sequence read data can be processed using any suitable number and / or combination of the same or different processing steps to help provide an outcome. In certain embodiments, processing a data set using the criteria described herein can reduce the complexity and / or dimensionality of the data set.

[0210] In some embodiments, one or more processing steps can include one or more filtering steps. The term "filtering" as used herein refers to removing a portion or a reference genome portion from consideration. The portion or reference genome portion to be removed can be selected according to any appropriate criteria, including but not limited to redundant data (such as redundant or overlapping mapping reads), uninformative data (such as portions or portions of a reference genome with a median count of 0), portions or portions of a reference genome containing sequences that appear too frequently or too rarely, noise data, etc., and combinations thereof. Filtering methods often involve removing one or more portions of a reference genome from consideration and subtracting the counts in the selected one or more portions of the reference genome to be removed from the count or total of the reference genome, chromosome, or genome under consideration. In some embodiments, portions of the reference genome can be removed sequentially (such as one at a time to allow evaluation of the removal effect of each individual portion), and in certain embodiments, all portions marked as needing to be removed can be removed simultaneously. In some embodiments, portions of the reference genome characterized by differences above or below a certain level are removed, which is sometimes referred to as filtering the "noise" portion of the reference genome. In certain embodiments, the filtering process comprises obtaining data points from a data set derived from an average profile level of a portion, chromosome, or chromosome segment by a predetermined multiple profile change, and in certain embodiments, the filtering process comprises removing data points from a data set derived from an average profile level of a portion, chromosome, or chromosome segment by a predetermined multiple profile difference. In some embodiments, the filtering process is used to reduce the number of candidate portions in a reference genome for analyzing the presence or absence of genetic variations. Reducing the number of candidate portions in a reference genome for analyzing the presence or absence of genetic variations (e.g., microdeletions, microduplications) generally reduces the complexity and / or dimensionality of the data set, and sometimes increases the speed of searching and / or identifying genetic variations and / or genetic anomalies by two or more orders of magnitude.

[0211] In some embodiments, one or more processing steps can include one or more standardization steps. Standardization can be carried out by suitable methods as described herein or known in the art. In certain embodiments, standardization includes adjusting the measured values ​​of different magnitudes to a theoretical common magnitude. In certain embodiments, standardization includes the mathematical adjustment of complication, to introduce the probability distribution of the adjusted numerical value in comparison. In some embodiments, standardization includes comparing distribution with normal distribution. In certain embodiments, standardization includes mathematical adjustment, which allows the comparison of the corresponding standardized values ​​for different data sets in a manner that eliminates some total effects (such as errors and anomalies). In certain embodiments, standardization includes scaling. Standardization sometimes includes dividing one or more data sets by a predetermined scalar or formula. The non-limiting example of standardization method includes portion by portion standardization, by the standardization of GC content, linear and nonlinear least squares regression, LOESS, GCLOESS, LOWESS (local weighted regression scatter point smoothing method), PERUN, ChAI, repeat masking (RM), GC-standardization and repeat masking (GCRM), cQn and / or its combination. In some embodiments, the presence or absence of a genetic variation (e.g., aneuploidy, microduplication, microdeletion) is determined using a normalization method (e.g., portion-wise normalization, normalization by GC content, linear and nonlinear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted regression scatter smoothing), PERUN, ChAI, repeat masking (RM), GC-normalization and repeat masking (GCRM), cQn, normalization methods known in the art and / or combinations thereof).

[0212] Any suitable number of normalizations can be used. In some embodiments, a data set can be normalized 1 or more times, 5 or more times, 10 or more times, or even 20 or more times. The data set can be normalized for a value (e.g., a normalized value) representing any suitable characteristic or variable (e.g., sample data, reference data, or both). Non-limiting examples of available data normalization types include normalizing the raw count data of one or more selected test or reference portions to the total number of counts on the mapping of the chromosome or whole genome mapped to the selected portion or segment; normalizing the raw count data of one or more selected portions to the median reference count of one or more portions or the chromosome mapped to the selected portion or segment; normalizing the raw count data to the aforementioned normalized data or its derivative; and normalizing the aforementioned normalized data to one or more other predetermined normalization variables. Depending on the characteristics or attributes selected as the predetermined normalization variables, normalizing the data set sometimes has the effect of separating statistical errors. By converting the data to a common scale (e.g., a predetermined normalization variable), normalizing the data set sometimes also makes the data features of data of different magnitudes comparable. In some embodiments, one or more normalizations of statistically derived values ​​can be used to minimize data variability and reduce the significance of outlying data. Normalization of portions or portions of a reference genome is sometimes referred to as "portion-wise normalization" when referring to normalized values.

[0213] In certain embodiments, the processing step includes standardization, including standardization to a static window, and in some embodiments, the processing step includes standardization, including standardization to a dynamic or sliding window. The term "window" herein refers to one or more portions selected for analysis, sometimes used as a reference for comparison (such as used as standardization and / or other mathematical or statistical operations). The term "standardization to a static window" herein refers to a standardization process using one or more portions selected for comparison of a test subject and a reference subject data set. In some embodiments, selected portions are used to generate profiles. The static window typically includes a group of predetermined portions that do not change during operation and / or analysis. The term "standardization to a dynamic window" or "standardized sliding window" herein refers to standardizing the portion of the genomic region (e.g., genetically tightly surrounded, adjacent portions or segments) located in a selected test portion, wherein one or more selected test portions are standardized to the portion that tightly surrounds the selected test portion. In certain embodiments, selected portions are used to generate profiles. Sliding or dynamic window normalization typically comprises repeatedly moving or sliding to adjacent test portions, and normalizing the newly selected test portion to a portion that closely surrounds or is adjacent to the newly selected test portion, wherein the adjacent windows have one or more common portions. In certain embodiments, a plurality of selected test portions and / or chromosomes can be analyzed by a sliding window process.

[0214] In some embodiments, being standardized to a sliding or dynamic window can produce one or more values, wherein each value represents the standardization of different groups with reference to parts selected from different regions of the genome (e.g., chromosome). In certain embodiments, the one or more values ​​gained are cumulative values ​​(e.g., digital estimates of the integration of the standardized count overview of selected parts, domains (e.g., chromosome parts) or chromosomes). Sliding or dynamic window process gained values ​​can be used to produce an overview and be convenient to obtain results. In some embodiments, the cumulative value of one or more parts can be displayed as a function of genomic position. Dynamic or sliding window analysis is sometimes used to analyze whether there is microdeletion and / or micro-insertion in a genome. In certain embodiments, showing that the cumulative value of one or more parts is used to identify whether there is a genetic variation region (e.g., microdeletion, microduplication). In some embodiments, dynamic or sliding window analysis is used to identify genomic regions containing microdeletions and in certain embodiments, dynamic or sliding window analysis is used to identify genomic regions containing microduplications.

[0215] Some examples of normalization procedures that can be used are described in detail below, such as LOESS, PERUN, ChAI, and principal component normalization methods.

[0216] In some embodiments, the processing step includes weighting. The terms "weighted", "weighting" or "weighting function" as used herein or their grammatical derivatives or equivalents refer to the mathematical processing of part or all of a data set, which is sometimes used to change the influence of certain data set features or variables on other data set features or variables (such as increasing or decreasing the importance and / or base value of the data contained in one or more parts of a reference genome based on the quality or practicality of the data in one or more parts of a selected reference genome). In some embodiments, a weighting function can be used to increase the influence of data with relatively small measurement variables, and / or reduce the influence of data with relatively large measurement differences. For example, a portion of a reference genome containing too low a frequency or low amount of sequence data can be "down weighted" to minimize the influence on the data set, whereas a selected portion of a reference genome can be "up weighted" to increase the influence on the data set. A non-limiting example of a weighting function is [1 / (standard deviation)] 2

[0015] The weighting step is sometimes performed in a manner substantially similar to the normalization step. In some embodiments, the data set is divided by a predetermined variable (e.g., a weighting variable). The predetermined variable is often selected (e.g., to minimize a target function, Phi) to weight different portions of the data set differently (e.g., to increase the influence of certain data types while decreasing the influence of other data types).

[0217] In certain embodiments, the processing step can include one or more mathematical and / or statistical processing. Any suitable mathematical and / or statistical processing can be single or combined for analyzing and / or processing data sets as herein described. Any suitable number of mathematical and / or statistical processing can be used. In some embodiments, the data set can be processed mathematically and / or statistically 1 time or multiple times, 5 times or more times, 10 times or more times or 20 times or more times. The non-limiting examples of the mathematical and statistical processing that can be used include addition, subtraction, multiplication, division, algebraic function, least squares estimation, curve fitting, differential equations, rational polynomials, double polynomials, orthogonal polynomials, z-scores, p values, chi values, phi values, peak level analysis, determining peak edge position, calculating peak area ratio, analyzing median chromosome level, calculating arithmetic mean absolute deviation, residual sum of squares, average, standard deviation, standard error etc., or its combination. Sequence read data or all or part of its processed result can be subjected to mathematical and / or statistical processing. Non-limiting examples of data set variables or features that can be statistically processed include raw counts, filtered counts, normalized counts, peak height, peak width, peak area, peak edge, lateral tolerance, P value, median level, mean level, distribution of counts within a genomic region, relative value representation of nucleic acid species, etc., or a combination thereof.

[0218] In some embodiments, the processing step can include the use of one or more statistical algorithms. Any suitable statistical algorithm can be used singly or in combination to analyze and / or process the data set described herein. Any suitable number of statistical algorithms can be used. In some embodiments, one or more, five or more, ten or more, or twenty or more statistical algorithms can be used to analyze the data set. Non-limiting examples of statistical algorithms suitable for use with the methods described herein include decision trees, counternull, multiple comparisons, omnibus tests, Behrens-Fisher problem, bootstrapping, Fisher's method with independent tests of significance, null hypothesis, type I error, type II error, exact test, one-sample Z test, two-sample Z test, one-sample t-test, paired t-test, two-sample pooled t-test with equal variances, two-sample unpooled t-test with unequal variances, one-proportion z-test, two-proportion pooled z-test, two-proportion unpooled z-test, one-sample chi-square test, two-sample F-test with equal variances, confidence interval, credible interval, significance, meta-analysis, simple linear regression, robust linear regression, the like, or combinations thereof. Non-limiting examples of data set variables or features that can be analyzed using statistical algorithms include raw counts, filtered counts, normalized counts, peak height, peak width, peak edge, lateral tolerance, P value, median level, mean level, distribution of counts within a genomic region, relative representation of nucleic acid species, the like, or a combination thereof.

[0219] In certain embodiments, a data set can be analyzed using multiple (e.g., 2 or more) statistical algorithms (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, bagging, neural networks, support vector machine models, random forests, classification tree models, K-nearest neighbors, logistic regression, and / or smooth loss Smoothing) and / or mathematical and / or statistical operations (such as operations described herein). In some embodiments, multiple operations are used to produce N-dimensional space, which can be used to provide results. In certain embodiments, the complexity and / or dimension of a data set can be reduced by analyzing the data set using multiple operations. For example, multiple operations are used to produce N-dimensional space (such as probability graph) on a reference data set, which can be used to represent whether there is genetic variation, depending on the genetic status of the reference sample (such as positive or negative to selected genetic variation). Using a substantially similar operating group to analyze test samples can be used to produce the respective N-dimensional points of the sample tested. The complexity and / or dimension of the test subject data set are sometimes reduced to N-dimensional points or single values ​​that can be easily compared with the N-dimensional space of the reference data. The test sample data falling into the N-dimensional space filled by the reference subject data represent the genetic status basically similar to the reference subject. The test sample data falling into the N-dimensional space filled by the reference subject data represent the genetic status basically dissimilar to the reference subject. In some embodiments, reference is euploid or does not have genetic variation or medical symptoms.

[0220] In some embodiments, after a data set has been calculated, optionally filtered, and standardized, the processed data set can be further manipulated using one or more filtering and / or standardization procedures. In certain embodiments, a data set that can be further manipulated using one or more filtering and / or standardization procedures can be used to generate a profile. In some embodiments, one or more filtering and / or standardization procedures can sometimes reduce the complexity and / or dimensionality of a data set. Results can be provided based on a data set that has been reduced in complexity and / or dimensionality.

[0221] In some embodiments portions can be filtered according to a measure of error (e.g., standard deviation, standard error, calculated variance, p-value, mean absolute error (MAE), mean absolute deviation and / or mean absolute deviation (MAD). In certain embodiments a measure of error refers to count variability. In some embodiments portions are filtered according to count variability. In certain embodiments count variability is a measure of error determined for counts of portions (i.e., portions) mapped to a reference genome for a plurality of samples (e.g., a plurality of samples obtained from a plurality of subjects, e.g., 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more or 10,000 or more subjects). In some embodiments portions having a count variability above a predetermined upper range can be filtered (e.g., excluded from consideration). In some embodiments a predetermined upper range is equal to or greater than about 50, about 52, about 54, about 56, about 58, about 60, about 62, about 64, about 66, about 68, about 70, about 72, about 74 or equal to or greater than about 76 In some embodiments, portions having count variability below a predetermined lower range can be filtered (e.g., excluded from consideration). In some embodiments, the predetermined lower range is a MAD value equal to or less than about 40, about 35, about 30, about 25, about 20, about 15, about 10, about 5, about 1, or equal to or less than about 0. In some embodiments, portions having count variability outside a predetermined range can be filtered (e.g., excluded from consideration). In some embodiments, the predetermined range is greater than 0 and less than about 7 In some embodiments, a MAD value of less than about 6, less than about 74, less than about 72, less than about 71, less than about 70, less than about 69, less than about 68, less than about 67, less than about 66, less than about 65, less than about 64, less than about 62, less than about 60, less than about 58, less than about 56, less than about 54, less than about 52, or less than about 50 is selected. In some embodiments, a predetermined range is a MAD value greater than 0 and less than about 67.7. In some embodiments, portions are selected whose count variability is within a predetermined range (e.g., for use in determining the presence or absence of a genetic variation).

[0222] In some embodiments, the count variability of portions represents a distribution (e.g., a normal distribution). In some embodiments, portions can be selected within a quantile of a distribution. In some embodiments, portions are selected that are within or less than about 99.9%, 99.8%, 99.7%, 99.6%, 99.5%, 99.4%, 99.3%, 99.2%, 99.1%, 99.0%, 98.9%, 98.8%, 98.7%, 98.6%, 98.5%, 98.4%, 98.3%, 98.2%, 98.1%, 98.0%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 85%, 80% or about 75% of a quantile of a distribution of count variability. In some embodiments, portions are selected that are within the 99th percentile of a distribution of count variability. In some embodiments, portions with MAD>0 and MAD<67.725, within the 99th percentile, are selected to identify a stable portion group for a reference genome.

[0223] Non-limiting examples of partial filtering involving PERUN are described herein and in International Patent Application No. PCT / US12 / 59123 (WO2013 / 052913), the entirety of which is incorporated herein by reference, including all text, tables, equations and figures. Portions can be filtered based on or in part on an error measure. The error measure includes an absolute value of the deviation, such as an R-factor, which can be used for partial removal or weighting in certain embodiments. In some embodiments, the R-factor is defined as the sum of the absolute deviation of the predicted count value from the measured value divided by the predicted count value from the measured value (e.g., Formula B described herein). Although an error measure including the absolute value of the deviation can be used, a suitable error measure can also be used. In certain embodiments, an error measure that does not include the absolute value of the deviation can be used, such as a dispersion based on squares. In some embodiments, portions are filtered or weighted according to a measure of mappability (e.g., a mappability score). Portions are sometimes filtered or weighted based on a relatively low number of sequence reads mapped to the portion (e.g., 0, 1, 2, 3, 4, 5 reads mapped to the portion). Portions can be filtered or weighted based on the type of analysis being performed. For example, for an aneuploidy analysis of chromosomes 13, 18, and / or 21, the sex chromosomes can be filtered, and only the autosomes or a subset of autosomes can be analyzed.

[0224] In a specific embodiment, the following filtering process can be used. Select the same group of parts (e.g., parts of a reference genome) within a given chromosome (e.g., chromosome 21) and compare the number of reads in affected and unaffected samples. The gap involves trisomy 21 and euploid samples, which involve part groups covering most chromosomes 21. The part groups between euploid and T21 samples are the same. The distinction between part groups and single segments is not critical, as defined by the parts. Compare the same genomic regions in different patients. This process can be used as a trisomy analysis, such as T13 or T18, in addition to or instead of T21.

[0225] In some embodiments, after a data set has been calculated, optionally filtered, and standardized, the processed data set can be manipulated by weighting. In certain embodiments, one or more portions can be optionally weighted to reduce the impact of the data contained in the selected portion (e.g., noisy data, uninformative data), and in some embodiments, one or more portions can be optionally weighted to increase or enhance the impact of the data contained in the selected portion (e.g., data with small measurement variance). In some embodiments, a data set is weighted using a single weighting function that reduces the impact of data with large variance and increases the impact of data with small variance. Weighting functions are sometimes used to reduce the impact of data with large variance and increase the impact of data with small differences (e.g., [1 / (standard deviation)] 2 In some embodiments, the weighted data are further manipulated to generate a profile of the processed data for classification and / or providing results. Results can be provided based on the weighted data profile.

[0226] Filtering and weighting of portions can be performed at one or more appropriate points in the analysis. For example, portions can be filtered or weighted before or after sequence reads are mapped to portions of a reference genome. In some embodiments, portions can be filtered or weighted before or after determining experimental bias for individual genomic portions. In some embodiments, portions can be filtered or weighted before or after calculating genomic segment levels.

[0227] In some embodiments, after the data set is calculated, optionally filtered, standardized, and optionally weighted, the processed data set can be operated on by one or more mathematical and / or statistical operations (such as statistical functions or statistical algorithms). In certain embodiments, the processed data can be further operated on by calculating the Z score of one or more selected parts, chromosomes, or parts of chromosomes. In some embodiments, the processed data set can be further operated on by calculating the p-value. For an embodiment of the equation for calculating the Z score and the p-value, see Equation 1 (Example 2). In certain embodiments, the mathematical and / or statistical operation includes one or more hypotheses related to ploidy and / or fetal fraction. In some embodiments, the data set is further operated on by one or more mathematical and / or statistical operations to generate an overview of the processed data for classification and / or providing results. Results can be provided based on the overview of the mathematical and / or statistical operation data. The results provided based on the overview of the mathematical and / or statistical operation data typically include one or more hypotheses related to ploidy and / or fetal fraction.

[0228] In certain embodiments, after a data set has been calculated, optionally filtered, and normalized, various operations are performed on the processed data set to generate an N-dimensional space and / or N-dimensional points. Results can be provided based on a profile of the data set analyzed in N dimensions.

[0229] In some embodiments, a data set is processed using one or more peak level analyses, peak width analyses, peak edge position analyses, peak lateral tolerances, etc., or derivatives thereof, or combinations thereof, as part of or subsequent to a processed and / or manipulated data set. In some embodiments, a profile of the processed data is generated using one or more peak level analyses, peak width analyses, peak edge position analyses, peak lateral tolerances, etc., or derivatives thereof, or combinations thereof, to facilitate classification and / or providing results. Results may be provided based on a profile of data that has been processed using one or more peak level analyses, peak width analyses, peak edge position analyses, peak lateral tolerances, etc., or derivatives thereof, or combinations thereof.

[0230] In some embodiments, one or more reference samples that are substantially free of the genetic variation being studied can be used to generate a reference median count profile that results in a predetermined value representing the absence of the genetic variation and typically deviates from a predetermined value in an area corresponding to the genomic location where the genetic variation is located in the test subject, if the test subject has the genetic variation. In a test subject suffering from a condition associated with the genetic variation or at risk of such a condition, the numerical value of the selected portion or segment is expected to be significantly different from the predetermined value for the unaffected genomic location. In certain embodiments, one or more reference samples known to carry the genetic variation being studied can be used to generate a reference median count profile that results in a predetermined value representing the presence of the genetic variation and typically deviates from a predetermined value in an area corresponding to the genomic location that does not have the genetic variation, wherein the test subject does not have the genetic variation in the genomic location. In a test subject not suffering from a condition associated with the genetic variation or at risk of such a condition, the numerical value of the selected portion or segment is expected to be significantly different from the predetermined value for the affected genomic location.

[0231] In some embodiments, analyzing and processing data can include using one or more hypotheses. An appropriate number or type of hypothesis can be used to analyze or process a data set. Non-limiting examples of hypotheses that can be used for data processing and / or analysis include maternal ploidy, fetal base values, the prevalence of certain sequences in a reference group, ethnic background, the prevalence of selected medical conditions in related family members, correspondence between the raw count distributions from different patients and / or GC normalization and repeat masking (such as GCRM), identical matches representing PCR artifacts (such as identical base positions), inherent assumptions in fetal quantitative tests (such as FQA), assumptions about twins (e.g., if there are 2 twins and only 1 is affected, the effective fetal fraction is only 50% of the fetal fraction of all measurements (triplets, quadruplets, etc. are also similar thereto)), uniform coverage of fetal cell-free DNA (such as cfDNA) of the whole genome, and combinations thereof.

[0232] In those examples where the quality and / or depth of the mapped sequence reads cannot predict the presence or absence of a genetic variation at a desired confidence level (e.g., a 95% or higher confidence level), one or more additional mathematical processing algorithms and / or statistical prediction algorithms can be used to generate additional numerical values ​​that can be used for data analysis and / or providing an outcome based on the standardized count distribution. The term "standardized count distribution" as used herein refers to a distribution generated using standardized counts. Examples of methods that can be used to generate standardized counts and standardized count distributions are described herein. The counted located sequence reads can be standardized relative to the test sample counts or reference sample counts. In some embodiments, a standardized count profile can be represented graphically.

[0233] LOESS standardization

[0234] LOESS is a regression model known in the art, which combines multiple regression models in the nearest meta-model based on k-. LOESS sometimes refers to local weighted polynomial regression. In some embodiments, GC LOESS applies the LOESS model to the relationship between the GC composition of the part with reference to genome and fragment count (e.g., sequence reading, counting). Using LOESS to describe a smooth curve by a data point group is sometimes referred to as a LOESS curve, particularly when given each smooth value by weighted quadratic least squares regression relative to the span of the value of the y-axis scatter plot standard variable. For each point in the data set, the LOESS method fits a low-degree polynomial to the data set, and the explanatory variable value approaches the point of the evaluated response. Fit the polynomial with weighted least squares so that the point approaching the evaluated response has more weight and the point away from has less weight. Then use the explanatory variable value of these data points to obtain the regression function value of a point by evaluating the local polynomial. After the regression function value has been calculated for each data point, sometimes the LOESS fitting is considered fully. Many details of this method, such as the degree and weight of the polynomial model are flexible.

[0235] PERUN Standardization

[0236] A normalization method for reducing errors associated with nucleic acid indicators is referred to herein as parameterized error removal and unbiased normalization (PERUN), as described herein and in International Patent Application PCT / US12 / 59123 (WO2013 / 052913), which is incorporated herein by reference in its entirety, including all text, tables, equations, and figures. The PERUN method can be applied to various nucleic acid indicators (e.g., nucleic acid sequence reads) to reduce the effects of errors that confound predictions based on the indicator.

[0237] For example, the PERUN method is used for nucleic acid sequence readings from samples and reduces the influence of errors that damage the determination of the genome segment level. The application is effectively used to determine whether an object expressed as a nucleotide sequence of various levels (e.g., portion, genome segment level) has a genetic variation using nucleic acid sequence readings. Non-limiting examples of variations in portions are chromosomal aneuploidy (e.g., trisomy 21, trisomy 18, trisomy 13) and the presence of sex chromosomes (e.g., XX in females and XY in males). Autosomal (e.g., chromosomes other than sex chromosomes) trisomy can be referred to as affected autosomes. Other non-limiting examples of variations in the genome segment level include microdeletions, microinsertions, duplications, and mosaicism.

[0238] In certain applications, the PERUN method can reduce experimental bias in normalized nucleic acid reads mapped to specific portions of a reference genome, the latter being referred to as portions and sometimes as portions of a reference genome. In this application, the PERUN method typically normalizes the counts of nucleic acid reads at specific portions of a reference genome across a large number of samples in three dimensions. A detailed description of PERUN and its applications is provided in the Examples section, as well as in International Patent Application PCT / US12 / 59123 (WO2013 / 052913) and U.S. Patent Application Publication No. US20130085681, the entire contents of which are incorporated herein by reference, including all text, tables, equations, and figures.

[0239] In certain embodiments, the PERUN method comprises calculating a genomic segment level of a portion of a reference genome from the following results: (a) sequence read counts of a test sample mapped to a portion of a reference genome, (b) an experimental bias (e.g., GC bias) of the test sample, and (c) one or more fitting parameters (e.g., fitting estimates) for the fitted relationship between (i) the experimental bias of the portion of the reference genome to which the sequence reads are mapped and (ii) the counts of sequence reads mapped to the portion. The experimental bias of each portion of the reference genome can be determined in a plurality of samples based on the fitted relationship between (i) the sequence read counts mapped to each portion of the reference genome and (ii) the mapping features of each portion of the reference genome. This fitted relationship for each sample can be aggregated in three dimensions for the plurality of samples. In certain embodiments, the aggregate can be arranged according to the experimental bias, although the PERUN method can be implemented without arranging the aggregate according to the experimental bias. The fitted relationship for each sample and the fitted relationship for each portion of the reference genome can be individually fitted to a linear function or a nonlinear function by a suitable fitting process known in the art.

[0240] In some embodiments, the relationship is a geometric and / or graphical relationship. In some embodiments, the relationship is a mathematical relationship. In some embodiments, the relationship is mapped. In some embodiments, the relationship is a linear relationship. In certain embodiments, the relationship is a nonlinear relationship. In some embodiments, the relationship is a regression (e.g., a regression line). Regression can be linear regression or nonlinear regression. The relationship can be expressed by a mathematical equation. Usually, the relationship is defined in part by one or more constants. A relationship can be generated by methods known in the art. In certain embodiments, a two-dimensional relationship can be generated for one or more samples, and a variable error check or possible error check can be selected for one or more of the dimensions. For example, a relationship can be generated using mapping software known in the art, which maps two or more variable values ​​provided by the user. A relationship can be fitted using methods known in the art (e.g., mapping software). Some relationships can be fitted by linear regression, and linear regression can generate slope and intercept. Some relationships are sometimes nonlinear and can be fitted by non-linear functions, such as parabolas, hyperbolas, or exponential functions (e.g., quadratic functions).

[0241] In the PERUN method, one or more fitted relationships can be linear. For analyzing cell-free circulating nucleic acids from pregnant women, where the experimental bias is GC bias and the mapping characteristic is GC content, the fitted relationship between (i) the sequence read counts mapped to each portion of the sample and (ii) the GC content of each portion of the reference genome can be linear. For the latter fitted relationship, when the fitted relationships are aggregated across multiple samples, a slope and a GC bias coefficient related to GC bias can be determined for each sample. In this embodiment, the fitted relationship between i) the GC bias coefficient of the portion and (ii) the counts of sequence reads mapped to the portion for multiple samples and portions can also be linear. An intercept and slope can be obtained from the latter fitted relationship. In this application, the slope represents the sample-specific bias based on GC content, and the intercept represents the portion-specific decay pattern common to all samples. When calculating genomic segment-level results to provide results (e.g., whether a genetic variation is present; determining fetal sex), the PERUN method can significantly reduce sample-specific bias and portion-specific decay.

[0242] In some embodiments, PERUN normalization uses a fit to a linear function and is described by Equation A, Equation B, or a derivative thereof.

[0243] Equation A:

[0244] M=LI+GS (A)

[0245] Equation B:

[0246] L=(M–GS) / I (B)

[0247] In some embodiments, L is a PERUN normalized level or profile. In some embodiments, L is the desired output from the PERUN normalization program. In certain embodiments, L is portion-specific. In some embodiments, L is determined based on multiple portions of a reference genome, representing a PERUN normalized level for a genome, chromosome, portion, or segment thereof. Level L is typically used for further analysis (e.g., determining Z-scores, maternal deletions / duplications, fetal microdeletions / microduplications, fetal sex, sex aneuploidy, and the like). The normalization method according to Equation B is referred to as parameterized error removal and unbiased normalization (PERUN).

[0248] In some embodiments, G is a GC bias coefficient measured using a linear model, LOESS, or any equivalent method. In some embodiments, G is a slope. In some embodiments, the GC bias coefficient G is assessed as the slope of the regression of the count M (e.g., raw count) for portion i and the GC content of portion i determined from a reference genome. In some embodiments, G represents secondary information extracted from M and determined based on a relationship. In some embodiments, G represents the relationship between a portion-specific count group and a portion-specific GC content value group for a sample (e.g., a test sample). In some embodiments, portion-specific GC content is derived from a reference genome. In some embodiments, portion-specific GC content is derived from an observed or measured GC content (e.g., measured from a sample). The GC bias coefficient is typically determined for each sample in a sample set, and is typically determined for a test sample. The GC bias coefficient is typically sample-specific. In some embodiments, the GC bias coefficient is a constant. In certain embodiments, the GC bias coefficient no longer changes once obtained from a sample.

[0249] In some embodiments, S is the slope derived from a linear relationship and I is the intercept. In some embodiments, the relationship from which I and S are derived is different from the relationship from which G is derived. In some embodiments, the relationship from which I and S are derived is fixed for a given experimental setting. In some embodiments, I and S are derived from a linear relationship based on counts (e.g., raw counts) and a GC bias coefficient based on multiple samples. In some embodiments, I and S are independently derived from the test sample. In some embodiments, I and S are derived from multiple samples. I and S are typically part-specific. In some embodiments, I and S are determined in euploid samples with reference to all parts of the genome using the assumption of L=1. In some embodiments, a linear relationship is determined for euploid samples, and I and S values ​​specific to the selected part are determined (assuming L=1). In certain embodiments, the same procedure is used for all parts of the reference genome in the human genome and a group of intercepts I and slopes S is determined for each part.

[0250] In some embodiments, cross validation is applied. Cross validation sometimes refers to rotation test (rotationestimation). In some embodiments, cross validation is used to assess the accuracy of a prediction model (e.g., PERUN) in the implementation for a test sample. In some embodiments, a round of cross validation includes dividing a data sample into complementary subgroups, performing a cross validation analysis (e.g., sometimes) on a subgroup (e.g., sometimes referred to as a training group) and using another subgroup validation analysis (e.g., sometimes referred to as a validation group or test group). In certain embodiments, multiple rounds of cross validation are performed using different partition products and / or different subgroups. Non-limiting examples of cross validation include leave-one-out, sliding edge, K-fold, 2-fold, repeated random sampling, etc., or a combination thereof. In some embodiments, cross validation randomly selects a working group containing 90% of the sample group, including known euploid fetuses and using the subgroup training model. In some embodiments, random selection is repeated 100 times, and each part produces 100 groups of slopes and 100 groups of intercepts.

[0251] In some embodiments, an M value is a measurement derived from a test sample. In some embodiments, M is a raw count of a measurement for a portion. In some embodiments, when values ​​I and S are available for a portion, the measurement M is determined from the test sample and used to determine a normalized level of PERUN, L, for a genome, chromosome, segment or portion thereof according to Equation B.

[0252] Thus, the PERUN method, when applied in parallel to sequence reads from multiple samples, can significantly reduce errors caused by (i) sample-specific experimental bias (e.g., GC bias) and (ii) portion-specific attenuation common to the samples. Other methods that address these two sources of error individually or sequentially are generally not able to reduce them as effectively as the PERUN method. Without being limited by theory, it is expected that the PERUN method is more effective in reducing errors in part because its generally additive process does not amplify as much as the generally multiplicative process used in other normalization methods (e.g., GC-LOESS).

[0253] Other normalization and statistical techniques can be used in conjunction with the PERUN method. Other processes can be applied before, after, and / or during the use of the PERUN method. Non-limiting examples of processes that can be used in conjunction with the PERUN method are described below.

[0254] In some embodiments, secondary standardization or adjustment of the genomic segment level of GC content can be used in conjunction with the PERUN method. Suitable GC content adjustment or standardization procedures (e.g., GC-LOESS, GCRM) can be used. In certain embodiments, specific samples can be identified with respect to other GC standardization processes. For example, the application of the PERUN method can determine the GC bias of each sample, and samples associated with a GC bias higher than a specific threshold can be selected for use in other GC standardization processes. In this embodiment, a predetermined threshold level can be used to select the sample for use in other GC standardizations.

[0255] In certain embodiments, a partial filtering or weighting process can be used in conjunction with the PERUN method. Suitable partial filtering or weighting processes can be used, and non-limiting examples are described herein, as well as in International Patent Application PCT / US12 / 59123 (WO2013 / 052913) and U.S. Patent Application Publication No. US20130085681, which are incorporated herein by reference in their entirety, including all text, tables, equations, and figures. In some embodiments, normalization techniques that reduce associated maternal insertions, duplications, and / or deletions (e.g., maternal and / or fetal copy number variations) are used in conjunction with the PERUN method.

[0256] The genomic segment levels calculated by the PERUN method can be used directly to provide results. In some embodiments, the genomic segment levels can be used directly to provide sample results, wherein the fetal fraction is about 2%-about 6% or higher (e.g., about 4% or higher fetal fraction). The genomic segment levels calculated by the PERUN method are sometimes further processed to provide results. In some embodiments, the calculated genomic segment levels are normalized. In certain embodiments, the sum, arithmetic mean, or median of the calculated genomic segment levels of the test portion (e.g., chromosome 21) can be divided by the sum, arithmetic mean, or median of the calculated genomic segment levels of the portion other than the test portion (e.g., chromosome 21 other than the autosomes) to generate the experimental genomic segment levels. The experimental genomic segment levels or the original genomic segment levels can be used as part of a planned analysis, such as calculating a Z-score or a Z-score. The Z-score of the sample can be generated by subtracting the expected genomic segment levels from the experimental genomic segment levels or the original genomic segment levels, and the resulting value can be divided by the standard deviation of the sample. In certain embodiments, the resulting Z-scores can be distributed and analyzed for different samples, or can be correlated with other variables, such as fetal fraction and others, and analyzed to provide an outcome.

[0257] As described herein, the PERUN method is not limited to standardizing according to GC deviation and GC content itself, and can be used to reduce the error associated with other error sources. A non-limiting example of the source of non-GC content deviation is mappability. When solving the standardization parameters beyond GC deviation and content, one or more fitting relationships can be non-linear (e.g., hyperbolic, exponential). In some embodiments, for example, when experimental deviation is determined from a non-linear relationship, an analyzable experimental deviation curvature estimate can be used.

[0258] The PERUN method can be applied to various nucleic acid indicators. Non-limiting examples of nucleic acid indicators are nucleic acid sequence readings and nucleic acid levels at specific locations of a microarray. Non-limiting examples of sequence readings include those obtained from cell-free circulating DNA, cell-free circulating RNA, cellular DNA, and cellular RNA. The PERUN method can be applied to sequence readings mapped to suitable reference sequences, such as genomic reference DNA, cellular reference RNA (e.g., transcriptome), and portions thereof (e.g., portions of the genomic complement of a DNA or RNA transcriptome, portions of chromosomes).

[0259] Thus, in certain embodiments, cellular nucleic acids (e.g., DNA or RNA) can be used as nucleic acid indicators. Reads of cellular nucleic acids mapped to portions of a reference genome can be normalized using the PERUN method. Cellular nucleic acids that bind to specific proteins are sometimes referred to as chromatin immunoprecipitation (ChIP) procedures. ChIP-enriched nucleic acids are nucleic acids, such as DNA or RNA, that are associated with cellular proteins. Reads of ChIP-enriched nucleic acids can be obtained using techniques known in the art. Reads of ChIP-enriched nucleic acids can be mapped to one or more portions of a reference genome, and the results can be normalized using the PERUN method to provide results.

[0260] In certain embodiments, cellular RNA can be used as a nucleic acid indicator. Cellular RNA readings can be mapped to a reference RNA portion and normalized using the PERUN method to provide results. A known sequence of cellular RNA (referred to as a transcriptome) or a segment thereof can be used as a reference to which RNA readings from a sample can be mapped. Readings of sample RNA can be obtained using techniques known in the art. Results mapped to the reference RNA readings can be normalized using the PERUN method to provide results.

[0261] In some embodiments, microarray nucleic acid levels can be used as nucleic acid indicators. The PERUN method can be used to analyze nucleic acid levels at specific locations of a sample or hybridized nucleic acids on an array, thereby standardizing the nucleic acid indicators provided by the microarray analysis. In this way, specific locations or hybridized nucleic acids on a microarray are similar to portions of mapped nucleic acid sequence reads, and the PERUN method can be used to normalize microarray data to provide improved results.

[0262] ChAI normalization

[0263] Other normalization methods that can be used to reduce errors in associated nucleic acid indicators are referred to herein as ChAI, typically using principal component analysis. In certain embodiments, principal component analysis includes (a) filtering portions of a reference genome according to a read density distribution, thereby providing a read density profile of a test sample, including the read densities of the filtered portions, wherein the read densities include sequence reads of circulating cell-free nucleic acids from a test sample from a pregnant female, and determining a read density distribution for the read densities of portions of a plurality of samples, (b) adjusting the read density profile of the test sample according to one or more principal components, wherein the principal components are obtained from a group of known euploid samples by principal component analysis, thereby providing a test sample profile, including the adjusted read densities, and (c) comparing the test sample profile to a reference profile, thereby providing a comparative relationship. In some embodiments, principal component analysis includes (d) determining whether a genetic variation exists in the test sample based on the comparison.

[0264] Filtration part

[0265] In some embodiments, one or more parts (such as genomic parts) are removed from consideration by a filtering process. In some embodiments, one or more parts are filtered (such as undergoing a filtering process), thereby providing a filtered part. In some embodiments, a filtering process removes certain parts and retains parts (such as partial subgroups). After the filtering process, the retained part generally refers to the filtered part herein. In some embodiments, a reference genome part is filtered. In some embodiments, the part removed by the filtering process is not included in determining whether there is a genetic variation (such as chromosome aneuploidy, microduplication, microdeletion). In some embodiments, the part of the associated reading density (such as reading density is used for part) is removed by a filtering process and the reading density of the associated removed part is not included in determining whether there is a genetic variation (such as chromosome aneuploidy, microduplication, microdeletion). In some embodiments, a reading density overview includes the reading density of the filtered part and / or is composed of it. Any suitable standard and / or method known in the art or as described herein can be used to select, filter and / or remove part. Non-limiting examples of criteria for filtering portions include redundant data (e.g., redundant or overlapping mapped reads), non-informative data (e.g., portions of a reference genome with 0 mapping counts), portions of a reference genome with sequences that occur too frequently or too infrequently, GC content, noisy data, mappability, counts, count variability, read density, read density variability, measures of uncertainty, measures of repeatability, etc., or combinations thereof. Portions are sometimes filtered based on a count distribution and / or a read density distribution. In some embodiments, portions are filtered based on a count distribution and / or a read density distribution, wherein the counts and / or read densities are obtained from one or more reference samples. Sometimes one or more reference samples are referred to herein as a training set. In some embodiments, portions are filtered based on a count distribution and / or a read density distribution, wherein the counts and / or read densities are obtained from one or more test samples. In some embodiments, portions are filtered based on a measure of uncertainty in a read density distribution. In certain embodiments, portions that exhibit large deviations in read density are removed by a filtering process. For example, a distribution of read densities (e.g., a distribution of average, mean, or median read densities, e.g., Figure 37A ), where each read density in the distribution maps to the same portion. An uncertainty measure (e.g., MAD) can be determined by comparing the distributions of read densities of multiple samples, where each genomic portion is associated with an uncertainty measure. Following the above example, portions can be filtered based on an uncertainty measure (e.g., standard deviation (SD), MAD) associated with each portion and a predetermined threshold. Figure 37B The distribution of MAD values ​​for the displayed portion is determined based on the distribution of read densities for multiple samples. The predetermined threshold is shown as a vertical dashed line, which encloses the acceptable range of MAD values. Figure 37B In some embodiments, according to the foregoing examples, portions comprising MAD values ​​within an acceptable range are retained and portions comprising MAD values ​​outside an acceptable range are removed from consideration. In some embodiments, according to the foregoing examples, portions comprising read density values ​​(e.g., median, average or mean read density) that exceed a predetermined uncertainty measure are typically removed from consideration by a filtering process. In some embodiments, portions comprising read density values ​​(e.g., median, average or mean read density) that exceed the interquartile range of a distribution are typically removed from consideration by a filtering process. In some embodiments, portions comprising read density values ​​that exceed 2, 3, 4 or 5 times the interquartile range of a distribution are removed from consideration by a filtering process. In some embodiments, portions comprising read density values ​​that exceed 2σ, 3σ, 4σ, 5σ, 6σ, 7σ or 8σ (e.g., where σ is the range defined by the standard deviation) are removed from consideration by a filtering process.

[0266] In some embodiments, the system includes a filtration module (18, Figure 42A The filtering module typically receives, retrieves and / or stores portions (e.g., portions of predetermined size and / or overlap, portions located within a reference genome) and read densities of associated portions, typically from other suitable modules (e.g., distribution modules 12, Figure 42A In some embodiments, the selected portion (e.g., 20 ( Figure 42A In some embodiments, a filtering module is required to provide filtered portions and / or remove portions from consideration. In certain embodiments, a filtering module removes read densities from consideration, where the read densities are associated with the removed portions. A filtering module typically provides selected portions (e.g., filtered portions) to other appropriate modules (e.g., distribution modules 12, Figure 42A A non-limiting example of a filtration module is shown in Example 7.

[0267] Deviation Assessment

[0268] Sequencing technology is susceptible to deviations from a variety of sources. Sometimes sequencing deviations are local deviations (e.g., local genome deviations). Local biases typically occur at the sequence read level. Local genome biases can be any suitable local bias. Non-limiting examples of local biases include sequence biases (e.g., GC biases, AT biases, etc.), biases associated with DNA enzyme I sensitivity, entropy, repeat sequence biases, chromatin structure biases, polymerase error rate biases, palindrome biases, inserted repeat biases, PCR-related biases, etc., or combinations thereof. In some embodiments, the source of the local bias is undetermined or unknown.

[0269] In some embodiments, a local genome bias assessment is determined. A local genome bias assessment is sometimes referred to herein as a local genome bias assessment. A local genome bias assessment can be determined with respect to a reference genome, a segment thereof, or a portion thereof. In some embodiments, a local genome bias assessment is determined for one or more sequence reads (e.g., some or all sequence reads of a sample). A local genome bias assessment for a sequence read is typically determined based on a local genome bias assessment of a corresponding location and / or position with respect to a reference (e.g., a reference genome). In some embodiments, a local genome bias assessment comprises a quantitative measurement of sequence bias (e.g., sequence read, reference genome sequence). A local genome bias assessment can be determined by a suitable method or mathematical process. In some embodiments, a local genome bias assessment is determined by a suitable method and / or a suitable distribution function (e.g., PDF). In some embodiments, a local genome bias assessment comprises a quantitative representation of a PDF. In some embodiments, a local genome bias assessment (e.g., probability density estimation (PDE), kernel density estimation) is determined by a probability density function (e.g., PDF, e.g., kernel density function) of a local bias content. In some embodiments, a density assessment comprises a kernel density assessment. A local genome bias assessment is sometimes expressed as the average, arithmetic mean, or median of a distribution. Sometimes local genome bias estimates are expressed as sums or integrals (e.g., the area under the curve (AUC) of a suitable distribution).

[0270] A PDF (e.g., a kernel density function, such as an Epanechnikov kernel density function) typically includes a bandwidth variable (e.g., bandwidth). The bandwidth variable typically defines the size and / or length of the window from which a probability density estimate (PDE) is derived when using the PDF. The window from which the PDE is derived typically includes a defined length of a polynucleotide. In some embodiments, the window from which the PDE is derived is a portion. The portion (e.g., portion size, portion length) is typically determined based on the bandwidth variable. The bandwidth variable determines the length or size of the window used to determine the local genome bias estimate. The local genome bias estimate is determined from the length of a polynucleotide segment (e.g., a continuous segment of nucleotide bases). The PDE (e.g., read density, local genome bias estimate (e.g., GC density)) can be determined using any suitable bandwidth, non-limiting examples of which include a bandwidth of about 5 bases to about 100,000 bases, about 5 bases to about 50,000 bases, about 5 bases to about 25,000 bases, about 5 bases to about 10,000 bases, about 5 bases to about 5,000 bases, about 5 bases to about 2,500 bases, about 5 bases to about 1000 bases, about 5 bases to about 500 bases, about 5 bases to about 250 bases, about 20 bases to about 250 bases, or the like. In some embodiments, a local genome bias estimate (e.g., GC density) is determined using a bandwidth of about 400 bases or less, about 350 bases or less, about 300 bases or less, about 250 bases or less, about 225 bases or less, about 200 bases or less, about 175 bases or less, about 150 bases or less, about 125 bases or less, about 100 bases or less, about 75 bases or less, about 50 bases or less, or about 25 bases or less. In certain embodiments, a local genome bias estimate (e.g., GC density) is determined using a bandwidth that is determined based on the average, mean, median, or maximum read length of sequence reads obtained for a given subject and / or sample. Sometimes a local genome bias estimate (e.g., GC density) is determined using a bandwidth that is approximately equal to the average, mean, median, or maximum read length of sequence reads obtained for a given subject and / or sample. In some embodiments a local genomic bias estimate (e.g., GC density) is determined using a bandwidth of about 250, 240, 230, 220, 210, 200, 190, 180, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20 or about 10 bases.

[0271] Local genome bias assessment can be determined at single base resolution, although local genome bias assessment (e.g., local GC content) can be determined at a lower resolution. In some embodiments, local genome bias assessment is determined with respect to local bias content. Local genome bias assessment is typically determined using a window (e.g., using a PDF). In some embodiments, local genome bias assessment comprises using a window comprising a preselected number of bases. Sometimes a window comprises a continuous base segment. Sometimes a window comprises one or more portions of non-continuous bases. Sometimes a window comprises one or more portions (e.g., portions of a genome). Window size or length is typically determined by bandwidth and according to PDF. In some embodiments, a window is approximately 10 or more, 8 or more, 7 or more, 6 or more, 5 or more, 4 or more, 3 or more, or approximately 2 or more times the length of the bandwidth. When using a PDF (e.g., a kernel density function) to determine density assessment, a window is sometimes twice the length of the selected bandwidth. A window may comprise any suitable number of bases. In some embodiments, a window comprises about 5 bases to about 100,000 bases, about 5 bases to about 50,000 bases, about 5 bases to about 25,000 bases, about 5 bases to about 10,000 bases, about 5 bases to about 5,000 bases, about 5 bases to about 2,500 bases, about 5 bases to about 1000 bases, about 5 bases to about 500 bases, about 5 bases to about 250 bases, or about 20 bases to about 250 bases. In some embodiments, a genome or its section is divided into a plurality of windows. The windows encompassing a genomic region may overlap or not overlap. In some embodiments, windows are located at positions equidistant from each other. In some embodiments, windows are located at positions unequally spaced from each other. In certain embodiments, a genome or its section is divided into a plurality of sliding windows, wherein windows incrementally slide across a genome or its section, wherein each window of each increment comprises a local genome preference assessment (e.g., local GC density). In some embodiments, the window can slide across the genome with any suitable increment according to any digital form or according to the sequence defined by any mathematics (athematic). In some embodiments, for local genome preference assessment, it is determined that the window slides across the genome, or its section, with following increments: about 10,000bp or more, about 5,000bp or more, about 2,500bp or more, about 1,000bp or more, about 750bp or more, about 500bp or more, about 400 bases or more, about 250bp or more, about 100bp or more, about 50bp or more, or about 25bp or more. In some embodiments, for local genome preference assessment, it is determined that the window slides across the genome, or its section, with following increments: about 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or about 1bp. For example, for a local genome bias estimate determination, the window may comprise approximately 400 bp (eg, 200 bp bandwidth) and may be slid across the genome in 1 bp increments.In some embodiments, a local genomic bias estimate is determined for each base in a genome or segment thereof using a kernel density function and a bandwidth of approximately 200 bp.

[0272] In some embodiments, a local genomic bias estimate is a local GC content and / or represents a local GC content. The term "local" herein (e.g., used to describe a local bias, a local preference estimate, a local preference content, a local genomic bias, a local GC content, etc.) refers to a polynucleotide segment of 10,000 bp or less. In some embodiments, the term "local" refers to a polynucleotide segment of 5000 bp or less, 4000 bp or less, 3000 bp or less, 2000 bp or less, 1000 bp or less, 500 bp or less, 250 bp or less, 200 bp or less, 175 bp or less, 150 bp or less, 100 bp or less, 75 bp or less, or 50 bp or less. Local GC content generally represents (e.g., mathematically, quantitatively represents) the GC content of a local segment of a genome, sequence reads, sequence read assemblies (e.g., contigs, profiles, etc.). For example, local GC content can be a local GC bias estimate or GC density.

[0273] Typically, one or more GC densities are determined for the polynucleotides of a reference or sample (e.g., a test sample). In some embodiments, GC density is an expression (e.g., a mathematical, quantitative expression) of local GC content (e.g., a polynucleotide segment of 5000 bp or less). In some embodiments, GC density is an estimate of local genomic bias. GC density can be determined using a suitable procedure described herein and / or known in the art. GC density can be determined using a suitable PDF (e.g., a kernel density function (e.g., an Epanechnikov kernel density function, e.g., see Figure 33 )) determines GC density. In some embodiments GC density is a PDE (e.g., core density evaluation). In certain embodiments, GC density is defined by the presence or absence of one or more guanine (G) and / or cytosine (C) nucleotides. In certain embodiments, GC density is defined by the presence or absence of one or more adenine (A) and / or thymine (T) nucleotides. In some embodiments, the GC density of local GC content is determined based on the GC content of an entire genome or a segment thereof (e.g., an autosome, a chromosome group, a single chromosome, a gene, e.g., see Figure 34) determined GC density standardization. One or more GC densities of the polynucleotides of a sample (such as a test sample) or with reference to a sample can be determined. Conventionally, the GC density of a reference genome is determined. In some embodiments, the GC density of a sequence read is determined according to the reference genome. Conventionally, the GC density of a read is determined according to the corresponding site and / or the GC density determined by the position of the reference genome mapped to read. In some embodiments, the GC density determined with reference to the positioning on the genome is distributed and / or provided to the reading, wherein the reading or its section is mapped to the same site of the reference genome. Any suitable method can be used to determine the GC density of the site for generating a read with reference to the site of the mapped read on the genome. In some embodiments, the median position of the mapped reading is determined with reference to the site on the genome, from which the GC density of the read is determined. For example, when the median position of the reading is mapped to the base number x at the chromosome 12 of the reference genome, the GC density of the read is conventionally provided with reference to the GC density determined by the core density evaluation of the position at the base number x at or near the chromosome 12 of the genome. In some embodiments, the GC density of some or all base positions of a read is determined according to the reference genome. Sometimes the GC density of a read comprises the average, sum, median, or integral of two or more GC densities determined for multiple base positions on a reference genome.

[0274] In some embodiments, a local genome bias evaluation (e.g., GC density) is quantitatively and / or provides a numerical value. A local genome bias evaluation (e.g., GC density) is sometimes expressed as an average, arithmetic mean, and / or median. A local genome bias evaluation (e.g., GC density) is sometimes expressed as the maximum peak height of a PDE. Sometimes a local genome bias evaluation (e.g., GC density) is expressed as the sum or integral (e.g., area under the curve (AUC)) of a suitable PDE. In some embodiments, GC density includes core weighting. In certain embodiments, the GC density of a read includes a value approximately equal to: average, arithmetic mean, sum, median, maximum peak height, or core weighted integral.

[0275] Deviation frequency

[0276] Bias frequencies are sometimes determined based on one or more local genome bias assessments (e.g., GC density). Bias frequencies are sometimes counts or sums of local genome bias assessments for a sample, a reference (e.g., a reference genome, a reference sequence), or a portion thereof. Bias frequencies are sometimes counts or sums of local genome bias assessments (e.g., each local genome bias assessment) for a sample, a reference, or a portion thereof. In some embodiments, bias frequencies are GC density frequencies. GC density frequencies are typically determined based on one or more GC densities. For example, a GC density frequency may represent the multiple represented by a GC density of value x relative to the entire genome or a segment thereof. Bias frequencies are typically distributions of local genome bias assessments, where the number of occurrences of each local genome bias assessment represents a bias frequency (e.g., see Figure 35 ). Preference frequencies are sometimes mathematically processed and / or standardized. Preference frequencies can be mathematically processed and / or standardized by suitable methods. In some embodiments, preference frequencies are standardized according to the representation (e.g., component, percentage) of each local genome preference assessment for a sample, a reference, or a portion thereof (e.g., an autosome, a chromosome subgroup, a single chromosome, or its reading). The deviation frequencies of some or all local genome preference assessments for a sample or reference can be determined. In some embodiments, the deviation frequencies of the local genome preference assessments for some or all sequence reads of a test sample can be determined.

[0277] In some embodiments a system comprises a bias density module 6. A bias density module can accept, retrieve and / or store mapped sequence reads 5 and a reference sequence 2 in any suitable format and generate local genome bias estimates, local genome bias distributions, bias frequencies, GC densities, GC density distributions and / or GC density frequencies (collectively represented by box 7). In some embodiments a bias density module transfers data and / or information (e.g., 7) to another suitable module (e.g., a relationship module 8).

[0278] relation

[0279] In some embodiments, one or more relationships are formed between a local genome bias assessment and a bias frequency. The term "relationship" herein refers to a mathematical and / or geometric relationship between two or more variables or values. A relationship can be generated by a suitable mathematical and / or geometric process. Non-limiting examples of relationships include mathematical and / or geometric process representations: functions, correlations, distributions, linear or non-linear equations, lines, regressions, fitted regressions, etc., or combinations thereof. Sometimes a relationship comprises a fitted relationship. In some embodiments, a fitted relationship comprises a fitted regression. Sometimes a relationship comprises two or more weighted variables or values. In some embodiments, a relationship comprises a fitted regression, wherein one or more variables or values ​​of the relationship are weighted. Sometimes regression is fitted in a weighted form. Sometimes regression is fitted without weighting. In certain embodiments, generating a relationship comprises drawing or charting.

[0280] In some embodiments, a suitable relationship is determined between a local genome preference assessment and a preference frequency. In some embodiments, a relationship is generated between (i) a local genome preference assessment and (ii) a preference frequency of a sample to provide a sample preference relationship. In some embodiments, a relationship is generated between (i) a local genome preference assessment and (ii) a preference frequency of a reference to provide a reference preference relationship. In certain embodiments, a relationship is generated between GC density and GC density frequency. In some embodiments, a relationship is generated between (i) a GC density and (ii) a GC density frequency of a sample to provide a sample GC density relationship. In some embodiments, a relationship is generated between (i) a GC density and (ii) a GC density frequency of a reference to provide a sample GC density relationship. In some embodiments, a relationship is generated between (i) a GC density and (ii) a GC density frequency of a reference to provide a reference GC density relationship. In some embodiments, when a local genome preference assessment is GC density, the sample preference relationship is the sample GC density relationship and is with reference to the preference relationship. The GC density with reference to the GC density relationship and / or the sample GC density relationship is typically an expression (such as mathematical or quantitative representation) of a local GC content. In some embodiments, the relationship between a local genome preference assessment and a preference frequency comprises distribution. In some embodiments, the relationship between a local genome preference assessment and a preference frequency comprises a fitting relationship (such as a fitting regression). In some embodiments, the relationship between a local genome preference assessment and a preference frequency comprises a fitted linear or non-linear regression (e.g., polynomial regression). In certain embodiments, the relationship between a local genome preference assessment and a preference frequency comprises a weighted relationship, wherein the local genome preference assessment and / or the preference frequency are weighted by a suitable process. In some embodiments, a weighted fitting relationship (e.g., weighted fitting) can be obtained by a process comprising a quantile regression, a parameterized distribution, or an empirical distribution of interpolation. In certain embodiments, a test sample, a reference, or a portion thereof, and the relationship between a local genome preference assessment and a preference frequency comprises a polynomial regression, wherein the local genome preference assessment is weighted. In some embodiments, a weighted fitting model comprises a weighted distribution value. The value of the distribution can be weighted by a suitable process. In some embodiments, the value near the end of the distribution provides less weight than the value near the distribution median. For example, for the distribution between a local genome preference assessment (e.g., GC density) and a preference frequency (e.g., GC density frequency), weight is determined according to the preference frequency of a given local genome preference assessment, wherein the local genome preference assessment comprising a preference frequency close to the distribution arithmetic mean provides more weight than the local genome preference assessment comprising a preference frequency relatively far from the arithmetic mean.

[0281] In some embodiments, system includes a preference relationship module 8. The relationship module can generate a relationship and define the function, coefficient, constant and variable of the relationship. The relationship module can accept, store and / or reclaim data and / or information (e.g., 7) and generate a relationship from a suitable module (e.g., preference density module 6). The relationship module usually generates and compares the distribution of local genome preference assessments. The relationship module can compare data groups and sometimes generate regression and / or fitting relationships. In some embodiments, the relationship module compares one or more distributions (e.g., samples and / or with reference to the distribution of local genome preference assessments) and provides weighting factors and / or weighted distribution 9 of the counting of sequence readings to other suitable modules (e.g., preference correction modules). Sometimes the relationship module directly provides the standardized count of sequence readings to distribution module 21, where counting is standardized according to relationship and / or comparison.

[0282] Generating comparisons and their applications

[0283] In some embodiments, the local preference in reducing sequence reading comprises standardization sequence reading count.Sequence reading count is usually standardized according to the comparison of test sample and reference.For example, sometimes sequence reading count is standardized by the local genome preference assessment of the sequence reading of comparative test sample and the local genome preference assessment (for example, with reference to genome or its part) with reference.In some embodiments, sequence reading count is standardized by the preference frequency of the local genome preference assessment of comparative test sample and the preference frequency of the local genome preference assessment with reference to.In some embodiments, sequence reading count is standardized by comparative sample preference relationship and with reference to preference relationship, thereby generates comparison.

[0284] Sequence read counts are typically standardized based on the comparison of two or more relationships. In certain embodiments, two or more relationships are compared, thereby providing a comparison of the local bias (e.g., standardized count) for reducing sequence reads. Two or more relationships can be compared by a suitable method. In some embodiments, the comparison includes adding, subtracting, multiplying, and / or dividing the first relationship by the second relationship. In certain embodiments, comparing two or more relationships includes using suitable linear regression and / or non-linear regression. In certain embodiments, comparing two or more relationships includes suitable polynomial regression (e.g., 3rd order polynomial regression). In some embodiments, comparing includes adding, subtracting, multiplying, and / or dividing the first regression by the second regression. In some embodiments, two or more relationships are compared by a process comprising an inference framework of multiple regressions. In some embodiments, two or more relationships are compared by a process comprising a suitable multivariate analysis. In some embodiments, two or more relationships are compared by a process comprising a basis function (e.g., a mixed function, e.g., a polynomial basis, a Fourier basis, or the like), a spline function, a radial basis function, and / or a wavelet.

[0285] In certain embodiments, the distribution of local genome bias estimates comprising bias frequencies for a test sample and a reference is compared by a process comprising polynomial regression, wherein the local genome bias estimates are weighted. In some embodiments, a polynomial regression is generated between (i) ratios, each ratio comprising the bias frequencies of a local genome bias estimate for a reference and the bias frequencies of a local genome bias estimate for a sample and (ii) the local genome bias estimates. In some embodiments, a polynomial regression is generated between (i) the ratio of the bias frequencies of a local genome bias estimate for a reference to the bias frequencies of a local genome bias estimate for a sample and (ii) the local genome bias estimates. In some embodiments, the comparison of the distribution of local genome bias estimates for a test sample and a reference comprises determining the logarithmic ratio (e.g., log2 ratio) of the bias frequencies of the local genome bias estimates for a reference and a sample. In some embodiments, the comparison of the distribution of local genome bias estimates comprises dividing the log ratio (e.g., log2 ratio) of the bias frequencies of the local genome bias estimates for a reference by the log ratio (e.g., log2 ratio) of the bias frequencies of the local genome bias estimates for a sample (e.g., see Examples 7 and 8). Figure 36 ).

[0286] According to the standardized counts for comparison, some counts are usually adjusted without adjusting others. Standardized counts sometimes adjust all counts and sometimes do not adjust any sequence read counts. Sequence read counts are sometimes standardized by a process including determining a weighting factor and sometimes the process does not include directly generating and adopting a weighting factor. According to the standardized counts for comparison, sometimes the weighting factor for determining each sequence read count is included. The weighting factor is usually specific to sequence reads and applied to specific sequence read counts. The weighting factor is usually determined based on a comparison of two or more preference relationships (such as a sample preference relationship compared with a reference preference relationship). Standardized counts are usually determined by adjusting the count value according to the weighting factor. Adjusting counts according to the weighting factor sometimes includes sequence read counts adding, subtracting, multiplying, and / or dividing the weighting factor. Weighting factors and / or standardized counts are sometimes determined from regression (such as a regression line). Standardized counts are sometimes directly obtained from a regression line (such as a fitted regression line) obtained by comparing the preference frequencies of the local genome preference assessments of the reference (such as a reference genome) and the test sample. In some embodiments, each count of a read of a sample provides a normalized count value based on a comparison of (i) the bias frequency of a local genome bias estimate of the read compared to (ii) the bias frequency of a local genome bias estimate of a reference. In certain embodiments, counts of sequence reads obtained for a sample are normalized and bias in sequence reads is reduced.

[0287] Sometimes a system includes a bias correction module 10. In some embodiments, the functions of a bias correction module are performed by a relationship modeling module 8. A bias correction module can receive, retrieve, and / or store mapped sequence reads and weighting factors (e.g., 9) from an appropriate module (e.g., relationship module 8, compression module 4). In some embodiments, a bias correction module provides counts to mapped reads. In some embodiments, a bias correction module applies weighted assignments and / or bias correction factors to sequence read counts to provide normalized and / or adjusted counts. A bias correction module typically provides normalized counts to another appropriate module (e.g., distribution module 21).

[0288] In some embodiments, standardized count comprises one or more features outside the factorized GC density, and standardized sequence read count.In some embodiments, standardized count comprises one or more different local genome preference assessments of factorized, and standardized sequence read count.In some embodiments, sequence read count is weighted according to the weighting determined by one or more features (such as one or more preferences).In some embodiments, according to one or more combined weighted standardized counts.Sometimes according to one or more combined weighted factorized one or more features and / or standardized counts by including the process using a multivariate model.Any suitable multivariate model can be used for standardized count.The non-limiting example of multivariate model comprises multivariate linear regression, multivariate quantile regression, multivariate interpolation of empirical data, non-linear multivariate model etc. or its combination.

[0289] In some embodiments, a system includes a multivariate correction module 13. The multivariate correction module can perform the functions of a bias density module 6, a relationship module 8, and / or a bias correction module 10 multiple times to adjust counts for multiple biases. In some embodiments, a multivariate correction module includes one or more bias density modules 6, relationship modules 8, and / or bias correction modules 10. The bias correction module sometimes provides normalized counts 11 to other appropriate modules (e.g., a distribution module 21).

[0290] Weighted part

[0291] In some embodiments, portions are weighted. In some embodiments, one or more portions are weighted to provide weighted portions. Weighted portions sometimes remove portion dependencies. Portions can be weighted by a suitable process. In some embodiments, one or more portions are weighted by an eigenfunction (e.g., a characteristic function). In some embodiments, an eigenfunction comprises replacing a portion with an orthogonal eigenportion. In some embodiments, a system comprises a portion weighting module 42. In some embodiments, a weighting module accepts, retrieves, and / or stores read densities, read density profiles, and / or adjusted read density profiles. In some embodiments, weighted portions are provided by a portion weighting module. In some embodiments, a weighting module is required to weight portions. A weighting module can weight portions by one or more weighting methods known in the art or described herein. A weighting module typically provides weighted portions to other suitable modules (e.g., a scoring module 46, a PCA statistics module 33, a profile generation module 26, etc.).

[0292] Principal component analysis

[0293] In some embodiments a read density profile (e.g., a read density profile of a test sample (e.g., Figure 39A) is adjusted according to principal component analysis (PCA). The reading density profiles of one or more reference samples and / or the reading density profiles of the test subjects can be adjusted according to PCA. Removing bias from a reading density profile by a PCA-related process is sometimes referred to herein as adjusting the profile. PCA can be performed by a suitable PCA method or a variant thereof. Non-limiting examples of PCA methods include classical correlation analysis (CCA), Karhunen–Loève transform (KLT), Hotelling transform, proper orthogonal decomposition (POD), singular value decomposition (SVD) of X, XTX eigenvalue decomposition (EVD), factor analysis, Eckart–Young theorem, Schmidt–Mirsky theorem, empirical orthogonal function (EOF), empirical eigenfunction decomposition, empirical component analysis, quasi-harmonic patterns, spectral analysis, empirical pattern analysis, etc., their variants or combinations. PCA typically identifies one or more biases in a reading density profile. The biases identified by PCA are sometimes referred to herein as principal components. In some embodiments, one or more biases can be removed by adjusting the reading density profile according to one or more principal components using a suitable method. A reading density profile can be adjusted by adding, subtracting, multiplying, and / or dividing the reading density profile by one or more principal components. In some embodiments, one or more preferences can be removed from a reading density profile by subtracting one or more principal components from the reading density profile. Although preferences in a reading density profile are typically identified and / or quantified by PCA of the profile, principal components are typically subtracted from the profile at the reading density level. PCA typically identifies one or more principal components. In some embodiments PCA identifies the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th and 10th or more principal components. In certain embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more principal components are used to adjust the profile. Typically, principal components are used to adjust the profile in the order in which they appear in the PCA. For example, when three principal components are subtracted from the reading density profile, the 1st, 2nd, and 3rd principal components are used. Sometimes the preferences identified by the principal components include profile features that are not used to adjust the profile. For example, PCA can identify genetic variations (e.g., aneuploidies, microduplications, microdeletions, deletions, translocations, insertions) and / or sex differences (e.g., see Figure 38C ) as principal components. Thus, in some embodiments, one or more principal components are not used to adjust a profile. For example, sometimes the 1st, 2nd, and 4th principal components are used to adjust a profile, when the 3rd principal component is not used to adjust a profile. Principal components can be obtained from a PCA using any suitable sample or reference. In some embodiments, principal components are obtained from a test sample (e.g., a test subject). In some embodiments, principal components are obtained from one or more references (e.g., a reference sample, a reference sequence, a reference group). For example, Figure 38A -C, PCA was performed on the median read density profiles obtained from a training set comprising multiple samples ( Figure 38A ), and get the first principal component ( Figure 38B ) and the second principal component ( Figure 38C ). In some embodiments, principal components are obtained from a group of subjects known to be free of the genetic variation under investigation. In some embodiments, principal components are obtained from a group of known euploids. Principal components are often identified based on PCA performed using one or more read density profiles of a reference (e.g., a training set). One or more principal components obtained from a reference are often subtracted from the read density profile of a test subject (e.g., a Figure 39B ), thus providing an overview of the adjustments (e.g. Figure 39C ).

[0294] In some embodiments, the system includes a PCA statistics module 33. The PCA statistics module can receive and / or retrieve reading density profiles from other suitable modules (e.g., profile generation module 26). PCA is typically performed by a PCA statistics module. The PCA statistics module typically receives, retrieves, and / or stores reading density profiles and processes reading density profiles from a reference group 32, a training group 30, and / or from one or more test subjects 28. The PCA statistics module can generate and / or provide principal components and / or adjust reading density profiles based on one or more principal components. Adjusted reading density profiles (e.g., 40, 38) are typically provided by the PCA statistics module. The PCA statistics module can provide and / or transfer adjusted reading density profiles (e.g., 38, 40) to other suitable modules (e.g., partial weighting module 42, scoring module 46). In some embodiments, the PCA statistics module can provide gender determination 36. Gender determination is sometimes determined based on PCA and / or based on one or more principal components to determine fetal gender. In some embodiments, the PCA statistics module includes some, all, or modifications of the R code shown below. R code to calculate principal components usually starts with cleaning the data (e.g., subtracting the median, filtering out parts, and removing extreme values):

[0295] #Clean the data outliers for PCA

[0296] dclean<-(dat-m)[mask,]

[0297] for(j in 1:ncol(dclean))

[0298] {

[0299] q<-quantile(dclean[,j],c(.25,.75))

[0300] qmin<-q[1]-4*(q[2]-q[1])

[0301] qmax<-q[2]+4*(q[2]-q[1])

[0302] dclean[dclean[,j] <qmin,j]<-qmin

[0303] dclean[dclean[,j]>qmax,j]<-qmax

[0304] }

[0305] Then calculate the principal components:

[0306] #Compute principal components

[0307] pc<-prcomp(dclean)$x

[0308] Finally, the PCA adjusted profile for each sample was calculated using:

[0309] #Compute residuals

[0310] mm<-model.matrix(~pc[,1:numpc])

[0311] for(j in 1:ncol(dclean))

[0312] dclean[,j]<-dclean[,j]-predict(lm(dclean[,j]~mm))

[0313] Comparative Overview

[0314] In some embodiments, determining a result includes comparing. In certain embodiments, a read density profile or a portion thereof is used to provide a result. In some embodiments, determining a result (e.g., determining whether a genetic variation exists) includes comparing two or more read density profiles. Comparing read density profiles typically includes comparing the read density profiles generated for a selected segment of a genome. For example, a test profile is typically compared with a reference profile, wherein the test and reference profiles are determined for substantially the same genomic segment (e.g., a reference genome). Comparing read density profiles sometimes includes comparing subgroups of two or more read density profile portions. A partial subgroup of a read density profile can represent a genomic segment (e.g., a chromosome or its segment). A read density profile can include a partial subgroup of any amount. Sometimes a read density profile includes 2 or more, 3 or more, 4 or more, or 5 or more subgroups. In certain embodiments, a read density profile includes a portion of two subgroups, wherein each portion represents an adjacent reference genomic segment. In some embodiments, a test profile can be compared to a reference profile, wherein both the test profile and the reference profile comprise a first subset of portions and a second subset of portions, wherein the first subset and the second subset represent different segments of a genome. Some subsets of portions of a read density profile can comprise genetic variation, while other subsets of portions sometimes are substantially free of genetic variation. Sometimes all subsets of portions of a profile (e.g., a test profile) are substantially free of genetic variation. Sometimes all subsets of portions of a profile (e.g., a test profile) comprise genetic variation. In some embodiments, a test profile can comprise a first subset of portions comprising a genetic variation and a second subset of portions that are substantially free of genetic variation.

[0315] In some embodiments, the methods described herein include making a comparison (e.g., comparing a test profile to a reference profile). Two or more data sets, two or more relationships, and / or two or more profiles can be compared by a suitable method. Non-limiting examples of statistical methods suitable for comparing data sets, relationships, and / or profiles include the Behrens-Fisher method, bootstrapping, Fisher's method of combining significant independent tests, Neyman-Pearson test, confirmatory data analysis, exploratory data analysis, exact test, F-test, Z-test, T-test, calculating and / or comparing uncertainty measures, null hypothesis, calculating nulls, etc., chi-square test, comprehensive test, calculating and / or comparing levels of significance (e.g., statistical significance), meta-analysis, multivariate analysis, regression, simple linear regression, enhanced linear regression, etc., or combinations thereof. In certain embodiments, comparing two or more data sets, relationships, and / or profiles includes determining and / or comparing uncertainty measures. As used herein, "uncertainty measures" refer to measures of significance (e.g., statistical significance), measures of error, measures of variance, measures of confidence, etc., or combinations thereof. The uncertainty measure can be a value (e.g., a threshold) or a range of values ​​(e.g., an interval, a confidence interval, a Bayesian confidence interval, a threshold range). Non-limiting examples of uncertainty measures include p-values, suitable variance measures (e.g., standard deviation, σ, absolute deviation, mean absolute deviation, etc.), suitable error measures (e.g., standard error, mean square error, root mean square error, etc.), suitable variance measures, suitable standard scores (e.g., standard deviation, cumulative percentage, percentage equivalence, Z-scores, T-scores, R-scores, standard nines (standard nines), percentages in standard nines, etc.), etc., or combinations thereof. In some embodiments, determining a significance level comprises determining a measure of uncertainty (e.g., a p-value). In certain embodiments, two or more data sets, relationships and / or profiles can be analyzed and / or compared using multiple (e.g., 2 or more) statistical methods (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, bagging, neural networks, support vector machine models, random forests, classification tree models, k-nearest neighbors, logistic regression and / or loss smoothing and / or any suitable mathematical and / or statistical operations (e.g., operations described herein).

[0316] In certain embodiments, comparing two or more read density profiles includes determining and / or comparing uncertainty measurements for two or more read density profiles. Read density profiles and / or associated uncertainty amounts are sometimes compared to facilitate the mathematical and / or statistical processing of a data set and / or to provide results. The read density profile generation of a test subject is sometimes compared with a read density profile generated for one or more references (e.g., a reference sample, a reference subject, etc.). In some embodiments, a result is provided by comparing the read density profile of a test subject with respect to a chromosome, a portion thereof, or a segment thereof with a read density profile of a reference, wherein the reference read density profile is obtained from a reference subject group (e.g., a reference) that is known to have no genetic variation. In some embodiments, a result is provided by comparing the read density profile of a test subject with respect to a chromosome, a portion thereof, or a segment thereof with a read density profile of a reference, wherein the reference read density profile is obtained from a reference subject group (e.g., a reference) that is known to have no genetic variation.

[0317] In certain embodiments, a read density profile for a test subject is compared to a predetermined value representation in the absence of a genetic variation, and sometimes deviates from a predetermined value at one or more genomic sites (e.g., portions) corresponding to the genomic site at which the genetic variation is located. For example, a read density profile in a test subject (e.g., a subject having or at risk for a medical condition associated with the genetic variation) is expected to be significantly different from a read density profile of a selected portion of a reference (e.g., a reference sequence, a reference subject, a reference group) of a test subject containing the genetic variation under study. The test subject read density profile is typically substantially the same as the read density profile of a selected portion of a reference (e.g., a reference sequence, a reference subject, a reference group) of a test subject that does not contain the genetic variation under study. The read density profile is typically compared to a predetermined threshold and / or threshold range (e.g., see Figure 40 ). As used herein, the term "threshold value" refers to any number calculated using a data set that meets the requirements and used as a limit for diagnosing genetic variation (e.g., copy number variation, aneuploidy, microduplication, microdeletion, chromosomal abnormality, etc.). In certain embodiments, the threshold value exceeds the result obtained by the method of the present invention, and the object is diagnosed as having a genetic variation (e.g., trisomy). In some embodiments, the threshold value or threshold value range is typically calculated by mathematically and / or statistically processing sequence reading data (e.g., from a reference and / or an object). The predetermined threshold value or threshold value range indicating whether there is a genetic variation may be different, but still provide a result that can be used to determine whether there is a genetic variation. In certain embodiments, a reading density profile including standardized reading density and / or standardized counts is generated to facilitate classification and / or provide results. The result can be provided based on a reading density profile including standardized counts (e.g., using the reading density profile graph).

[0318] In some embodiments, a system includes a scoring module 46. The scoring module can accept, retrieve and / or store read density profiles (e.g., adjusted, standardized read density profiles) from other suitable modules (e.g., profile generation module 26, PCA statistics module 33, portion weighting module 42, etc.). The scoring module can accept, retrieve, store and / or compare two or more read density profiles (e.g., test profiles, reference profiles, training groups, test subjects). The scoring module can typically provide scores (e.g., graphs, profile statistics, comparisons (e.g., differences between two or more profiles), Z-scores, uncertainty measures, call zones, sample calls 50 (e.g., determining whether a genetic variation exists), and / or results). The scoring module can provide scores to an end user and / or to other suitable modules (e.g., displays, printers, etc.). In some embodiments, the scoring module includes some, all, or modifications of the following R code, which includes an R function for calculating chi-square statistics for a specific test (e.g., high-chr21 counts).

[0319] The three parameters are:

[0320] x = sample reading data (part x sample)

[0321] m = median of the part

[0322] y = test vector (e.g. false for all parts and true for chr21)

[0323] getChisqP<-function(x,m,y)

[0324] {

[0325] ahigh<-apply(x[!y,],2,function(x)sum((x>m[!y])))

[0326] alow<-sum((!y))-ahigh

[0327] bhigh<-apply(x[y,],2,function(x)sum((x>m[y])))

[0328] blow<-sum(y)-bhigh

[0329] p<-sapply(1:length(ahigh),function(i){

[0330] p<-chisq.test(matrix(c(ahigh[i],alow[i],bhigh[i],blow[i]),2))$p.value / 2

[0331] if(ahigh[i] / alow[i]>bhigh[i] / blow[i])p<-max(p,1-p)

[0332] else p<-min(p,1-p;p})

[0333] return(p)

[0334] Hybrid regression standardization

[0335] In some embodiments, hybridization standardization is used. In some embodiments, hybridization standardization method reduces deviation (such as GC deviation). In some embodiments, hybridization standardization comprises the relationship of (i) analyzing two variables (such as counting and GC content) and (ii) selecting and applying standardization method according to said analysis. In certain embodiments, hybridization standardization comprises (i) regression (such as regression analysis) and (ii) selecting and applying standardization method according to said regression. In some embodiments, the counting obtained from the first sample (such as the first group of samples) is standardized by a method different from the counting obtained from other samples (such as the second group of samples). In some embodiments, the counting obtained from the first sample (such as the first group of samples) is standardized by the first standardization method, and the counting obtained from the second sample (such as the second group of samples) is standardized by the second standardization method. For example, in certain embodiments the first standardization method comprises using linear regression and the second standardization method comprises using non-linear regression (such as LOESS, GC-LOESS, LOWESS regression, LOESS smoothing).

[0336] In some embodiments, hybridization normalization method is used to standardize the sequence reading (such as count, mapped count, mapped reading) mapped to the portion of the genome or chromosome. In certain embodiments, the original count is standardized, and in some embodiments, the count adjusted, weighted, filtered or previously standardized is standardized by hybridization normalization method. In certain embodiments, the genome segment level or Z-score is standardized. In some embodiments, the count mapped to the portion of the selected genome or chromosome is standardized by hybridization normalization method. Count can refer to a suitable measurement of the sequence reading mapped to the portion of the genome, its non-limiting example includes original count (such as uncompressed count), standardized count (such as PERUN, ChAI or standardization of suitable method), portion level (such as average level, arithmetic average level, median level, or etc.), Z-score, etc., or its combination. Count can be the original count or processed count of one or more samples (such as test sample, sample from pregnant women). In some embodiments, count is obtained from one or more samples from one or more objects.

[0337] In some embodiments, a normalization method (e.g., a normalization method of the type described) is selected based on a regression (e.g., regression analysis) and / or a correlation coefficient. Regression analysis refers to a statistical technique for evaluating the relationship between variables (e.g., counts and GC content). In some embodiments, a regression is generated based on the counts and GC content measurements of each portion in a plurality of portions of a reference genome. Suitable GC content measurements can be used, non-limiting examples of which include measuring the content of guanine, cytosine, adenine, thymine, purine (GC), or pyrimidine (AT or ATU), melting temperature (Tm) (e.g., denaturation temperature, annealing temperature, hybridization temperature), measuring free energy, etc., or a combination thereof. Measurements of guanine (G), cytosine (C), adenine (A), thymine (T), purine (GC), or pyrimidine (AT or ATU) content can be expressed as a ratio or percentage. In some embodiments, any suitable ratio or percentage is used, non-limiting examples of which include GC / AT, GC / total nucleotides, GC / A, GC / T, AT / total nucleotides, AT / GC, AT / G, AT / C, G / A, C / A, G / T, G / A, G / AT, C / T, etc. or combinations thereof. In some embodiments, a measurement of GC content is a ratio or percentage of GC to total nucleotide content. In some embodiments, a measurement of GC content is a ratio or percentage of GC to total nucleotide content in terms of sequence reads mapped to a reference genome portion. In certain embodiments, GC content is determined based on and / or from sequence reads mapped to portions of a reference genome, and the sequence reads are obtained from a sample (e.g., a sample obtained from a pregnant female). In some embodiments, a GC content measurement is not determined based on and / or from sequence reads. In certain embodiments, a GC content measurement is determined for one or more samples obtained from one or more subjects.

[0338] In some embodiments, generating a regression comprises generating a regression analysis or a correlation analysis. Suitable regression can be used, non-limiting examples of which include regression analysis, (e.g., linear regression analysis), goodness of fit analysis, Pearson correlation analysis, hierarchical correlation, unexplained variance components, Nash–Sutcliffe model validity analysis, regression model validation, proportional reduction loss, root mean square difference, etc. or a combination thereof. In some embodiments, a regression line is generated. In certain embodiments, generating a regression comprises generating a linear regression. In certain embodiments, generating a regression comprises generating a non-linear regression (e.g., LOESS regression, LOWESS regression).

[0339] In some embodiments, a regression is used to determine whether a correlation (e.g., a linear correlation) exists, such as a correlation between counts and GC content measurements. In some embodiments, a regression (e.g., a linear regression) is generated and a correlation coefficient is determined. In some embodiments, a suitable correlation coefficient is determined, non-limiting examples of which include a coefficient of determination, an R2 value, a Pearson correlation coefficient, and the like.

[0340] In some embodiments, the goodness of fit of a regression (e.g., a regression analysis linear regression) is determined. Goodness of fit is sometimes determined by observation or mathematical analysis. Assessment sometimes includes determining whether the goodness of fit of a non-linear regression or a linear regression is greater. In some embodiments, a correlation coefficient is a measure of goodness of fit. In some embodiments, the fitness of an assessment regression is determined according to a correlation coefficient and / or a correlation coefficient cutoff value. In some embodiments, a goodness of fit assessment includes comparing a correlation coefficient with a correlation coefficient cutoff value. In some embodiments, the goodness of fit of an assessment regression indicates a linear regression. For example, in certain embodiments, the goodness of fit of a linear regression is greater than the goodness of fit of a non-linear regression, and the goodness of fit assessment indicates a linear regression. In some embodiments, an assessment indicates a linear regression and a linear regression is used for normalized counts. In some embodiments, the goodness of fit of an assessment regression indicates a nonlinear regression. For example, in certain embodiments, the goodness of fit of a non-linear regression is greater than the goodness of fit of a linear regression, and the goodness of fit assessment indicates a nonlinear regression. In some embodiments, an assessment indicates a nonlinear regression and a nonlinear regression is used for normalized counts.

[0341] In some embodiments, when the correlation coefficient is equal to or greater than a correlation coefficient cutoff value, the goodness of fit assessment indicates a linear regression. In some embodiments, when the correlation coefficient is less than a correlation coefficient cutoff value, the goodness of fit assessment indicates a nonlinear regression. In some embodiments, the correlation coefficient cutoff value is predetermined. In some embodiments, the correlation coefficient cutoff value is about 0.5 or greater, about 0.55 or greater, about 0.6 or greater, about 0.65 or greater, about 0.7 or greater, about 0.75 or greater, about 0.8 or greater, or about 0.85 or greater.

[0342] In some embodiments, when the correlation coefficient is equal to or greater than about 0.6, a standardization method comprising linear regression is used. In some embodiments, when the correlation coefficient is equal to or greater than a correlation coefficient cutoff value of 0.6, sample counts (e.g., counts of each portion with reference to genome, counts of each portion) are standardized according to linear regression, otherwise counts are standardized according to non-linear regression (e.g., when the coefficient is less than a correlation coefficient cutoff value of 0.6). In some embodiments, the standardization process comprises generating linear regression or non-linear regression using (i) counts and (ii) GC content for each portion with reference to a plurality of portions of genome. In some embodiments, when the correlation coefficient is less than a correlation coefficient cutoff value of 0.6, a standardization method comprising non-linear regression (e.g., LOWESS, LOESS) is used. In some embodiments, when the correlation coefficient (e.g., correlation coefficient) is less than about 0.7, less than about 0.65, less than about 0.6, less than about 0.55, or less than a correlation coefficient cutoff value of about 0.5, a standardization method comprising non-linear regression (e.g., LOWESS) is used. For example, in some embodiments, a normalization method comprising non-linear regression (eg, LOWESS, LOESS) is used when the correlation coefficient is less than a correlation coefficient cutoff of about 0.6.

[0343] In some embodiments, a specific type of regression (e.g., linear or non-linear regression) is selected, and after generating the regression, the counts are standardized by subtracting the regression from the counts. In some embodiments, subtracting the regression from the counts provides standardized counts with reduced deviations (e.g., GC deviations). In some embodiments, a linear regression is subtracted from the counts. In some embodiments, a non-linear regression (e.g., LOESS, GC-LOESS, LOWESS regression) is subtracted from the counts. Any suitable method can be used to subtract the regression line from the counts. For example, if the count x is derived from part i (e.g., part i) comprising a GC content of 0.5 and the regression line determines the count y at a GC content of 0.5, then xy = the standardized counts of part i. In some embodiments, the counts are standardized before and / or after the regression is subtracted. In some embodiments, counts standardized by a hybrid normalization method are used to generate genomic segment levels, Z-scores, levels and / or profiles of a genome or its segments. In certain embodiments, counts standardized by a hybrid normalization method are analyzed by methods described herein to determine whether there is a genetic variation (e.g., in a fetus).

[0344] In some embodiments, hybridization normalization methods include filtering or weighting one or more parts before or after normalization. Suitable methods as described herein can be used to filter parts, including methods for filtering parts (e.g., parts with reference to a genome). In some embodiments, parts (e.g., parts with reference to a genome) are filtered before applying hybridization normalization methods. In some embodiments, sequencing read counts that are only mapped to selected parts (e.g., parts selected according to count variability) are standardized by hybridization normalization. In some embodiments, sequencing read counts that are mapped to filtered parts (e.g., parts filtered according to count variability) of a reference genome are removed before using hybridization normalization methods. In some embodiments, hybridization normalization methods include selecting or filtering parts (e.g., parts with reference to a genome) according to suitable methods (e.g., methods described herein). In some embodiments, hybridization normalization methods include selecting or filtering parts (e.g., parts with reference to a genome) according to the uncertainty value of the counts of each part mapped to a variety of test samples. In some embodiments, hybridization normalization methods include selecting or filtering parts (e.g., parts with reference to a genome) according to count variability. In some embodiments a hybridization normalization method comprises selecting or filtering portions (e.g., portions of a reference genome) based on GC content, repetitive elements, repeat sequences, introns, exons, the like, or a combination thereof.

[0345] For example, in some embodiments, multiple samples of multiple pregnant female subjects are analyzed and a subset of portions (e.g., portions with reference to a genome) are selected based on count variability. In certain embodiments, linear regression is used to determine the correlation coefficient of (i) counts and (ii) GC content for each selected portion of a sample obtained from a pregnant female subject. In some embodiments, a correlation coefficient greater than a predetermined correlation cutoff value (e.g., about 0.6) is determined, a goodness of fit assessment indicates a linear regression, and the counts are standardized by subtracting the linear regression from the counts. In certain embodiments, a correlation coefficient less than a predetermined correlation cutoff value (e.g., about 0.6) is determined, a goodness of fit assessment indicates a nonlinear regression, a LOESS regression is generated, and the counts are standardized by subtracting the LOESS regression from the counts.

[0346] Overview

[0347] In some embodiments, a processing step may include generating one or more profiles (e.g., profile graphs) from various data sets or derivatives thereof (e.g., as a result of one or more mathematical and / or statistical data processing steps known in the art and / or described herein).

[0348] The term "overview" herein refers to the result of the mathematical and / or statistical operations of data, which can facilitate the identification of patterns and / or dependencies in a large amount of data. A "overview" generally includes the values ​​obtained from one or more operations on data or data sets based on one or more criteria. An overview generally includes a variety of data points. Any suitable number of data points may be included in the overview, depending on the nature and / or complexity of the data set. In certain embodiments, an overview may include 2 or more data points, 3 or more data points, 5 or more data points, 10 or more data points, 24 or more data points, 25 or more data points, 50 or more data points, 100 or more data points, 500 or more data points, 1000 or more data points, 5000 or more data points, 10,000 or more data points, or 100,000 or more data points.

[0349] In some embodiments, an overview is a representation of an entire data set, and in certain embodiments, an overview is a representation of a portion or subset of a data set. That is, an overview sometimes includes data points representative of data that have not been filtered to remove any data, or is generated therefrom, and sometimes an overview includes data points representative of data that have been filtered to remove unwanted data, or is generated therefrom. In some embodiments, a data point in an overview represents the results of a data operation on a portion. In certain embodiments, a data point in an overview includes the results of a data operation on groups of portions. In some embodiments, groups of portions may be adjacent to each other, and in certain embodiments, groups of portions may be from different parts of a chromosome or genome.

[0350] The data points in the profile derived from the data set can represent any suitable data classification. Non-limiting examples of categories that data can be grouped to generate profile data points include: parts based on size, parts based on sequence characteristics (e.g., GC content, AT content, position on chromosome (e.g., short arm, long arm, centromere, telomere), etc.), expression level, chromosome, etc., or a combination thereof. In some embodiments, a profile can be generated from data points obtained from other profiles (e.g., a normalized data profile that is normalized again to a different normalized value to generate a re-normalized data profile). In certain embodiments, a profile generated from data points of other profiles reduces the number of data points and / or the complexity of the data set. Reducing the number of data points and / or the complexity of the data set is generally beneficial for interpreting the data and / or for providing results.

[0351] An overview (such as a genome overview, a chromosome overview, a chromosome segment overview) is typically a set of standardized or non-standardized counts of two or more parts. An overview typically includes at least one level (such as a genome segment level), typically includes two or more levels (such as an overview typically has multiple levels). Levels are typically used for groups of parts with approximately the same count or standardized count. This paper describes levels in detail. In certain embodiments, an overview includes one or more parts, the parts can be weighted, removed, filtered, standardized, adjusted, averaged (derived mean), added, subtracted, or processed or transformed in any combination thereof. An overview typically includes a standardized count mapped to a part defining two or more levels, wherein counting is further standardized according to one of the levels by a suitable method. Typically, an overview count (such as an overview level) is associated with an uncertain value.

[0352] The profile comprising one or more levels is sometimes filled (e.g., hole filled).Filling (e.g., hole filled) refers to the process of identifying and adjusting the level of maternal microdeletion or maternal duplication (e.g., copy number variation) in the profile.In some embodiments, the level of fetal microdeletion or fetal microdeletion is filled.In some embodiments, microdeletion or microdeletion can artificially raise or reduce the overall level of the profile (e.g., chromosome profile) in the profile, resulting in false positive or false negative determined by chromosome aneuploidy (e.g., trisomy).In some embodiments, the level of microdeletion and / or deletion in the profile is identified and adjusted (e.g., filled and / or removed) by the process sometimes referred to as filling or hole filling.In certain embodiments, the profile includes one or more first levels that are significantly different from the second level in the profile, and each of the one or more first levels includes maternal copy number variation, fetal copy number variation, or maternal copy number variation and fetal copy number variation, and one or more of the first levels are adjusted.

[0353] A profile comprising one or more levels may include a first level and a second level. In some embodiments, the first level is different from (e.g., significantly different from) the second level. In some embodiments, the first level comprises a first set of portions, the second level comprises a second set of portions, and the first set of portions is not a subset of the second set of portions. In certain embodiments, the first set of portions is different from the second set of portions, from which the first and second levels are determined. In some embodiments, a profile may have multiple first levels that are different from (e.g., significantly different, e.g., having significantly different values) the second level within the profile. In some embodiments, a profile comprises one or more first levels that are significantly different from the second level within the profile, and the one or more first levels are adjusted. In some embodiments, a profile comprises one or more first levels that are significantly different from the second level within the profile, each of the one or more first levels comprising maternal copy number variation, fetal copy number variation, or maternal copy number variation and fetal copy number variation, and the one or more first levels are adjusted. In some embodiments, the first level in the profile is removed from the profile or adjusted (e.g., padded). A profile may comprise multiple levels, the multiple levels comprising one or more first levels that are significantly different from the one or more second levels, typically the predominant level in the profile being the second level, wherein the second levels are approximately equal to each other. In some embodiments greater than 50%, greater than 60%, greater than 70%, greater than 80%, greater than 90% or greater than 95% of the levels in a profile are a second level.

[0354] An overview is sometimes displayed as a graph. For example, one or more levels representing the counts of a portion (e.g., standardized counts) can be mapped and visualized. Non-limiting examples of profiles that can be generated include raw counts (e.g., raw count profiles or original profiles), standardized counts, portions-weighted, Z-scores, p-values, area ratios and fitted ploidy, median levels and the ratio between the fetal fractions of the fit and measurement, principal components, etc., or a combination thereof. In some embodiments, an overview graph allows observation of manipulated data. In certain embodiments, an overview graph can be used to provide a result (e.g., area ratios and fitted ploidy, median levels and the ratio between the fetal fractions of the fit and measurement, principal components). As used herein, the term "raw count overview graph" or "original overview graph" refers to a graph of counts in each portion of a region standardized to the total count of the region (e.g., a genome, a portion, a chromosome, a chromosome portion with reference to a genome, or a chromosome segment). In some embodiments, an overview can be generated using a static window process, and in certain embodiments, an overview can be generated using a sliding window process.

[0355] The profile generated for the test subject is sometimes compared with the profile generated for one or more reference subjects to facilitate the mathematical and / or statistical manipulation of the data set and / or provide a result. In some embodiments, a profile is generated based on one or more starting hypotheses (e.g., maternal nucleic acid contribution (e.g., maternal total integral), fetal nucleic acid contribution (e.g., fetal fraction), reference sample ploidy, etc., or a combination thereof). In certain embodiments, a test profile is typically centered around a predetermined value representing the absence of a genetic variation, and typically deviates from a predetermined value in the corresponding area of ​​the genomic position where the genetic variation (if the test subject has a genetic variation) is located in the test subject. In a test subject suffering from a disease associated with a genetic variation or having this risk, the numerical value of the selected portion is expected to be significantly different from the predetermined value of the unaffected genomic position. Based on the starting hypothesis (e.g., a fixed ploidy or optimal ploidy, a fixed fetal fraction or optimal fetal fraction, or a combination thereof), a predetermined threshold or cutoff value or threshold range indicating whether a genetic variation exists may be different, but it still provides a result that can be used to determine whether a genetic variation exists. In some embodiments, a profile indicates and / or represents a phenotype.

[0356] As a non-limiting example, normalized sample and / or reference count profiles can be obtained from raw sequence read data by:

[0357] (a) calculating a reference median count for a selected chromosome, portion or segment thereof from a reference set known to be free of genetic variation,

[0358] (b) removing the uninformative portion of the reference sample raw counts (e.g., filtering);

[0359] (c) normalizing the reference counts of all remaining reference genome portions to the total residual counts of the selected chromosome or selected genomic position of the reference sample (e.g., the summed remaining counts after removing uninformative reference genome portions), thereby generating a normalized reference subject profile;

[0360] (d) removing the corresponding portion from the test subject sample; and

[0361] (e) normalizing the remaining test subject counts at one or more selected genomic positions to the sum of the residual reference median counts for the chromosome or chromosomes containing the selected genomic positions, thereby generating a normalized test subject profile. In certain embodiments, an additional normalization step involving the entire genome (reduced by the filtered portion in (b)) can be included between (c) and (d).

[0362] A data set profile can be generated by one or more processing of the count-mapped sequence read data. Some embodiments include the following. Sequence reads are mapped and the number of sequence tags mapped to each genomic portion is determined (e.g., count). A raw count profile is generated from the counted mapped sequence reads. In certain embodiments, an outcome is provided by comparing the raw count profile of a test subject to a reference median count profile of a chromosome, portion, or segment thereof for a set of reference subjects known to be free of genetic variation.

[0363] In some embodiments, sequence read data is optionally filtered to remove noisy data or uninformative portions. After filtering, the remaining counts are typically summed to generate a filtered data set. In certain embodiments, a filtered count profile is generated from a filtered data set.

[0364] Sequence read data can be counted and optionally filtered, and the data set can be standardized to generate a level or overview. The data set can be standardized by standardizing one or more selected parts to a suitable standardized reference value. In some embodiments, the standardized reference value represents the total count of the chromosome from which a part is selected. In certain embodiments, the standardized reference value represents one or more corresponding parts of the chromosome of the reference data set prepared by the reference subject group known to not contain genetic variation. In some embodiments, the standardized reference value represents one or more corresponding parts of the chromosome of the test subject data set prepared by the test subject for analyzing whether there is a genetic variation. In certain embodiments, the standardization process is carried out using a static window method and in some embodiments, the standardization process is carried out using a mobile or sliding window method. In certain embodiments, generating an overview including standardized counts is convenient for classification and / or providing results. The result can be provided based on an overview diagram (e.g., using this overview diagram) including standardized counts.

[0365] level

[0366] In some embodiments, value (such as numerical value, quantitative value) is attributed to level.Count can be determined by suitable method, operation or mathematical process (such as processed level).Level is typically or is derived from the count (such as standardized count) of the group of part.In some embodiments the level of part is substantially equal to the count total (such as count, standardized count) mapped to part.Conventionally, the count of suitable method known in the art, operation or mathematical process processing, conversion or processing is determined to level.In some embodiments, level is derived from processed count, and the non-limiting example of the count of processing includes weighting, removal, filtering, standardization, adjustment, average, deriving arithmetic mean (such as arithmetic mean level), adding, subtracting, transforming count or its combination.In some embodiments, level includes standardized count (such as standardized count of part).Level can be used for count standardization by suitable process, and its non-limiting example includes the standardization of portion standardization, GC content, linear and non-linear least squares regression, GC LOESS, LOWESS, PERUN, ChAI, RM, GCRM, cQn etc. and / or its combination).Level can include standardized count or the relative amount of counting. In some embodiments, a level is used for the counts or standardized counts of two or more portions averaged and the level refers to the average level. In some embodiments, a level is used for the counts or the group of portions with the arithmetic mean of the standardized counts, which is referred to as the arithmetic mean level. In some embodiments, a level is derived from a portion of the counts comprising raw and / or filtered counts. In some embodiments, a level is based on raw counts. In some embodiments, a level is associated with an uncertainty value (e.g., standard deviation, MAD). In some embodiments, a level is represented by a Z-score or a p-value. The level of one or more portions herein is synonymous with "genomic segment level."

[0367] Standardized or non-standardized counts of two or more levels (e.g., two or more levels in a profile) can sometimes be subjected to mathematical operations (e.g., addition, multiplication, average, standardization, etc., or combinations thereof) based on the levels. For example, standardized or non-standardized counts of two or more levels can be standardized based on one, some, or all of the levels in a profile. In some embodiments, standardized or non-standardized counts of all levels in a profile are standardized based on one level in a profile. In some embodiments, standardized or non-standardized counts of a first level in a profile are standardized based on standardized or non-standardized counts of a second level in a profile.

[0368] Non-limiting examples of levels (e.g., first level, second level) are group levels comprising processed counts, levels comprising the mean, median or average of counts, levels comprising normalized counts, the like, or any combination thereof. In some embodiments, a first level and a second level in a profile are derived from counts of portions mapped to the same chromosome. In some embodiments, a first level and a second level in a profile are derived from counts of portions mapped to different chromosomes.

[0369] In some embodiments a level is determined from standardized or non-standardized counts mapped to one or more portions. In some embodiments a level is determined from standardized or non-standardized counts mapped to two or more portions, wherein the standardized counts for each portion are generally about the same. For a level, there can be differences in the counts (e.g., standardized counts) within a group of portions. For a level, there can be one or more portions within a group of portions that have counts that are significantly different (e.g., peaks and / or slopes) from other portions of the group. Any suitable number of standardized or non-standardized counts associated with any suitable number of portions can define a level.

[0370] In some embodiments, one or more levels can be determined from the standardized or non-standardized counts of all or some genomic portions. Typically, levels can be determined from all or some standardized or non-standardized counts of a chromosome or its segment. In some embodiments, two or more counts derived from two or more portions (e.g., groups of portions) determine a level. In some embodiments, two or more counts (e.g., from the counts of two or more portions) determine a level. In some embodiments, the counts of 2-about 100,000 portions determine a level. In some embodiments, the counts of 2-about 50,000, 2-about 40,000, 2-about 30,000, 2-about 20,000, 2-about 10,000, 2-about 5000, 2-about 2500, 2-about 1250, 2-about 1000, 2-about 500, 2-about 250, 2-about 100, or 2-about 60 portions determine a level. In some embodiments, the counts of about 10-about 50 portions determine a level. In some embodiments, counts from about 20 to about 40 or more portions determine a level. In some embodiments, a level comprises counts from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60 or more portions. In some embodiments, a level corresponds to a group of portions (e.g., a group of portions of a reference genome, a group of portions of a chromosome, or a group of portions of a chromosome segment).

[0371] In some embodiments, levels are determined based on normalized or non-normalized counts of adjacent portions. In some embodiments, adjacent portions (e.g., a group of portions) represent adjacent segments of a genome or adjacent segments of a chromosome or gene.

[0372] For example, when segments are combined tail-to-tail, two or more adjacent segments may represent a sequence set of a DNA sequence that is longer than each segment.

[0373] For example, two or more contiguous portions can represent an entire genome, chromosome, gene, intron, exon, or segment thereof. In some embodiments, a level is determined from a collection (e.g., group) of contiguous portions and / or non-contiguous portions.

[0374] Different levels

[0375] In some embodiments, a normalized count profile comprises a level (e.g., a first level) that is significantly different from other levels (e.g., a second level) within the profile. The first level may be higher or lower than the second level. In some embodiments, a first level is used for a group comprising one or more reads comprising portions of a copy number variation (e.g., maternal copy number variation, fetal copy number variation, or maternal copy number variation and fetal copy number variation) and a second level is used for a group comprising portions of reads having substantially no copy number variation. In some embodiments, significantly different refers to an observable difference. In some embodiments, significantly different refers to a statistically different or statistically significantly different. A statistically significant difference is sometimes a statistical estimate of an observable difference. Statistically significant differences can be estimated using methods suitable in the art. Any suitable threshold or range can be used to determine two levels that are significantly different. In certain embodiments, two levels (e.g., mean levels) differ by about 0.01% or more (e.g., 0.01% of one or another level value) and are significantly different. In some embodiments, two levels (e.g., mean levels) differ by about 0.1% or more and are significantly different. In some embodiments, two levels (e.g., mean levels) differ by about 0.5% or more and are significantly different. In some embodiments, two levels (e.g., mean levels) differ by about 0.5, 0.75, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or more than 10%. In some embodiments, two levels (e.g., mean levels) differ significantly and there is no overlap in the levels and / or there is no overlap within a range defined by an uncertainty value calculated for one or both levels. In certain embodiments, the uncertainty value is a standard deviation, expressed as σ. In some embodiments, two levels (e.g., mean levels) differ significantly when they differ by about 1 or more times the uncertainty value (e.g., σ). In some embodiments, two levels (e.g., mean levels) differ significantly when they differ by about 2 or more times the uncertainty value (e.g., σ), about 3 or more, about 4 or more, about 5 or more, about 6 or more, about 7 or more, about 8 or more, about 9 or more, or about 10 or more times the uncertainty value. In some embodiments, two levels (e.g., mean levels) are significantly different when their difference is about 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, or 4.0 times the uncertainty or more. In some embodiments, the confidence level increases with an increase in the difference between the two levels. In certain embodiments, the confidence level decreases with a decrease in the difference between the two levels and / or an increase in the uncertainty or more.For example, sometimes the confidence level increases in proportion to the difference between the level and the standard deviation (eg, MAD).

[0376] One or more prediction algorithms can be used to determine significance or give meaning to the test data collected under variable conditions, and their weights can be independent or interdependent. As used herein, the term "variable" refers to a factor, quantity, or function in an algorithm that has a certain value or set of values.

[0377] In some embodiments, a first set of portions generally includes portions that are different from (e.g., do not overlap with) a second set of portions. For example, sometimes a first level of normalized counts is significantly different from a second level of normalized counts in a profile, and the first level is for the first set of portions, the second level is for the second set of portions, and the portions do not overlap between the first and second sets of portions. In certain embodiments, the first set of portions is not a subset of the second set of portions from which the first level and the second level are respectively determined. In some embodiments, the first set of portions is different and / or differs from the second set of portions from which the first level and the second level are respectively determined.

[0378] In some embodiments a first set of portions is a subset of a second set of portions in a profile. For example, sometimes a second level of normalized counts for a second set of portions in a profile comprises a first level of normalized counts for a first set of portions in a profile and the first set of portions is a subset of the second set of portions in a profile. In some embodiments an average, mean or median level is derived from a second level, wherein the second level comprises the first level. In some embodiments a second level comprises a second set of portions representing an entire chromosome and the first level comprises the first set of portions, wherein the first set is a subset of the second set of portions and the first level represents maternal copy number variation, fetal copy number variation, or both maternal copy number variation and fetal copy number variation present in a chromosome.

[0379] In some embodiments, the value of the second level is closer to the arithmetic mean, average or median of the count overview of a ...

Claims

1. A system for determining whether a fetus has chromosomal aneuploidy, the system comprising modules for: (a) determining a chromosome count representation based on counts of nucleic acid sequence reads that map to portions of a reference genome, wherein the sequence reads are counts of circulating cell-free nucleic acid from a test sample from a pregnant female carrying a fetus; (b) determining the fetal fraction of the test sample; (c) calculating a log odds ratio (LOR), wherein the LOR is the logarithm of the quotient of (i) a first product between (1) the conditional probability of having the chromosomal aneuploidy and (2) the prior probability of having the chromosomal aneuploidy, and (ii) a second product between (1) the conditional probability of not having the chromosomal aneuploidy and (2) the prior probability of not having the chromosomal aneuploidy, wherein the conditional probability of having the chromosomal aneuploidy is determined based on the fetal fraction of (b) and the count representation of (a); and (d) Identifying the presence or absence of chromosomal aneuploidy based on the LOR and the chromosome count representation.

2. The system of claim 1, wherein the chromosome count representation is the count of all parts in the chromosome divided by the count of all parts in the autosomes.

3. The system of claim 1 or 2, further comprising a module for providing z-score quantification of the chromosome count representation.

4. The system of claim 3, wherein the z-score is the result of subtracting the median of (i) the test sample chromosome count representation minus (ii) the euploid count representation divided by (iii) the MAD of the euploid count representation, wherein: (i) the test sample chromosome count representation is the ratio of the count of the portion in the chromosome divided by the count of the portion in the autosome, and (ii) the median of the euploid count representation is the median of the ratio of the count of the portion in the chromosome divided by the count of the portion in the autosome for euploids.

5. The system of claim 1 , wherein the conditional probability of having the chromosomal aneuploidy is determined based on: the fetal fraction determined for the test sample in (b), the z-score of the chromosome count representation determined for the test sample in (a), and a fetal fraction-specific distribution of z-scores of the chromosome count representation.

6. The system of claim 5, wherein the conditional probability of having the chromosomal aneuploidy is determined by the relationship in Equation 23: where f is the fetal fraction, X is the summed fraction of chromosomes, X ∼ f(μX,σX), where μX and σX are the mean and standard deviation of X, respectively, and f(·) is the distribution function.

7. A system as described in claim 5 or 6, wherein the conditional probability of having the chromosomal aneuploidy is the intersection of the z-score for the test sample chromosome count representation of (a) and the fetal fraction-specific distribution of z-scores for the chromosome count representation.

8. The system of claim 1 , wherein the conditional probability of not having the chromosomal aneuploidy is determined based on the chromosome count representation of (a) and the count representation for euploidy.

9. The system of claim 8, wherein the conditional probability of not having the chromosomal aneuploidy is the intersection of the z-score of the chromosome count representation and the distribution of z-scores of chromosome count representations in euploids.

10. The system of claim 1, wherein the prior probability of having the chromosome aneuploidy and the prior probability of not having the chromosome aneuploidy are determined from a plurality of samples that do not include the test subject.

11. The system of claim 1 , comprising determining whether the LOR is greater than or less than zero.

12. The system of claim 1, wherein the counts of nucleic acid sequence reads mapped to portions of a reference genome are normalized counts.

13. The system of claim 12, wherein the counts are normalized by a normalization method comprising GC-LOESS normalization.

14. The system of claim 12 or 13, wherein the counts are normalized by a normalization method including principal component normalization.

15. The system of claim 12, wherein the counts are normalized by a normalization method comprising GC-LOESS normalization followed by principal component normalization.

16. The system of claim 12, wherein the counts are normalized by a normalization method comprising the steps of: (1) determining guanine and cytosine (GC) bias coefficients for the test sample based on a fitted correlation between (i) counts of sequence reads mapped to each portion and (ii) the GC content of each of the portions, wherein the GC bias coefficient is an estimate of the slope of a linear fit correlation or the curvature of a nonlinear fit correlation; and (2) calculating, using a microprocessor, a genomic section level for each of the portions based on the counts in (a), the GC bias coefficient in (b), and a fitted correlation for each of the portions, thereby providing a calculated genomic section level, wherein the fitted correlation is a fitted correlation between (i) the GC bias coefficient for each of the plurality of samples and (ii) the counts of sequence reads mapped to each of the portions in the plurality of samples.

17. The system of claim 1, further comprising a module for determining a z-score quantification of the chromosome count representation and determining whether it is less than, greater than, or equal to 3.

95.

18. A system as described in claim 17, further comprising a method for determining the presence of chromosomal aneuploidy based on, or at least in part based on, a determination that (i) the z-score representation of the chromosome count for the test sample is greater than or equal to 3.95 and (ii) the LOR is greater than 0.

19. A system as described in claim 17, further comprising a method for determining the absence of chromosomal aneuploidy based on, or at least in part based on, a determination that (i) the z-score of the chromosome representation of the test sample is less than 3.95 and / or (ii) the LOR is less than 0.

20. The system of claim 18 or 19, wherein the chromosomal aneuploidy is trisomy or monosomy.

21. The system of claim 1, wherein the count representation is a normalized count representation.

22. The system of claim 1, wherein one or more or all of (a), (b), (c), and (d) are performed by a microprocessor.

23. The system of claim 1, wherein one or more or all of (a), (b), (c), and (d) are performed by a computer.

24. The system of claim 1, wherein one or more or all of (a), (b), (c), and (d) are performed in conjunction with a memory.

25. The system of claim 1, further comprising a module for sequencing nucleic acids in a sample obtained from the pregnant female prior to (a), thereby providing nucleic acid sequence reads.

26. The system of claim 1, further comprising a module for mapping the nucleic acid sequence reads to portions of a reference genome prior to (a).

Citation Information

Patent Citations

  • Fragmentation-based methods and systems for sequence variation detection and discovery

    US20050112590A1

  • Method and compositions for detection and enumeration of genetic variations

    US20070065823A1

  • Restriction endonuclease enhanced polymorphic sequence detection

    US20090317818A1

  • Processes and compositions for methylation-based enrichment of fetal nucleic acid from a maternal sample useful for non invasive prenatal diagnoses

    US20100105049A1

  • Simultaneous determination of aneuploidy and fetal fraction

    US20110224087A1