Method and process for non-invasive assessment of genetic variations

By acquiring and analyzing circulating cell-free nucleic acid sequence reads from pregnant women, mapping them to a reference genome, and performing decomposition map analysis, the problem of low false negative and low false positive detection of fetal chromosomal abnormalities was solved, thus improving the accuracy of prenatal diagnosis.

CN121320518APending Publication Date: 2026-01-13SEQUENOM INC
View PDF 19 Cites 0 Cited by

Patent Information

Application Number
CN202511411937.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2013-05-24
Filing Date
2014-05-23
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Current technologies are insufficient to accurately detect fetal chromosomal aneuploidy, microduplication, or microdeletion with low false negative and low false positive rates, resulting in inadequate accuracy in prenatal diagnosis.

Method used

By acquiring circulating cell-free nucleic acid sequence reads from pregnant women, mapping them to a reference genome, standardizing the counts, generating genomic segment levels, segmenting profiles, and performing decomposition map analysis, the presence of chromosomal aneuploidy, microduplications, or microdeletions in the fetus can be determined with low false negatives and low false positives.

Benefits of technology

It enables accurate detection of fetal chromosomal aneuploidy, microduplication, or microdeletion with low false negative and low false positive rates, improving the accuracy and reliability of prenatal diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121320518A_ABST
    Figure CN121320518A_ABST
Patent Text Reader

Abstract

Methods, processes, and apparatus for non-invasive assessment of genetic variations using decision analysis are provided herein. The decision analyses sometimes include segmentation analyses and / or concession ratio analyses.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related patent applications

[0002] This patent application claims the rights of U.S. Provisional Patent Application 61 / 827,385, filed May 24, 2013, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS”, inventors Zeljko Dzakula et al., file number SEQ-6068-PV.This patent application relates to U.S. Patent 13 / 669,136, filed November 5, 2012, entitled "Methods and Procedures for Non-Invasive Assessment of Genetic Variations" (inventors: Cosmin Deciu, Zeljko Dzakula, Mathias Ehrich, and Sung Kim, file number SEQ-6034-CTt), which is in turn the subject of International PCT Application PCT / US2012 / 059123, filed October 5, 2012, entitled "Methods and Procedures for Non-Invasive Assessment of Genetic Variations" (inventors: Cosmin Deciu, Zeljko Dzakula, Mathias Ehrich, and Sung Kim). Kim, file number SEQ-6034-PC); it claims (i) the right to claim U.S. Provisional Patent Application 61 / 709,899, filed October 4, 2012, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS”, inventors Cosmin Deciu, Zeljko Dzakula, Mathias Ehrich, and Sung Kim, file number SEQ-6034-PV3; and (ii) the right to claim ... (iii) Claims the rights to U.S. Provisional Patent Application 61 / 663,477, filed October 6, 2011, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETICVARIATIONS”, inventors Zeljko Dzakula and Mathias Ehrich et al., filed SEQ-6034-PV.The entire contents of the aforementioned patent application are incorporated herein by reference, including its text, tables and figures. field

[0003] The technical aspects of this article relate to non-invasive methods, processes, and equipment for assessing genetic variations. background

[0004] The genetic information of living organisms (such as animals, plants, and microorganisms) and other forms of replicating genetic information (such as viruses) are encoded as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a series of nucleotides or modified nucleotides that represent the primary structure of a chemical or presumed nucleic acid. The complete human genome contains approximately 30,000 genes located on twenty-four (24) chromosomes (see The Human Genome, T. Strachan, BIOS Science Press, 1992). Each gene encodes a specific protein, which performs a specific biochemical function in living cells after being expressed through transcription and translation.

[0005] Many medical conditions are caused by one or more genetic variations. Some genetic variations cause medical conditions, including, for example, hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease, and cystic fibrosis (CF) (Human Genome Mutations, DN Cooper and M. Krawczak, BIOS Publishers, 1993). These genetic diseases can result from the addition, substitution, or deletion of a single nucleotide in the DNA of a specific gene. Some birth defects are caused by chromosomal abnormalities (also known as aneuploidy), such as trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), X monosomy (Turner syndrome), and certain sex chromosome aneuploidies such as Klinefelter syndrome (XXY). Other genetic variations are the sex of the fetus, which is usually determined based on the X and Y sex chromosomes. Some genetic variations predispose an individual to or cause any of many diseases, such as diabetes, arteriosclerosis, obesity, various autoimmune diseases, and cancers (such as colorectal cancer, breast cancer, ovarian cancer, and lung cancer).

[0006] Identifying one or more genetic variations or alterations can lead to a diagnosis or predisposition to a specific medical condition. Identifying genetic variations can aid in medical decision-making and / or the use of complementary medical interventions. In some implementations, identifying one or more genetic variations or alterations involves analyzing cell-free DNA. Cell-free DNA (CF-DNA) consists of DNA fragments derived from cell death and peripheral blood circulation. High concentrations of CF-DNA can indicate certain clinical conditions such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infection, and other diseases. Furthermore, cell-free fetal DNA (CFF-DNA) can be detected in maternal bloodstream and is used in a variety of non-invasive prenatal diagnostic procedures. Overview

[0007] This document provides, in certain aspects, a method for determining the presence of chromosomal aneuploidy, microduplication, or microdeletion in a fetus with low false negatives and low false positives, the method comprising (a) acquiring counts of nucleic acid sequence readings mapped to portions of a reference genome, wherein the sequence readings are readings of circulating cell-free nucleic acids from a pregnant woman; (b) normalizing the counts mapped to each portion to provide a calculated genomic segment level; (c) generating a profile of genomic regions based on the calculated genomic segment level; (d) segmenting the profile to provide two or more decomposed maps; and (e) determining the presence of chromosomal aneuploidy, microduplication, or microdeletion in the fetus with low false negatives and low false positives based on the two or more decomposed maps.

[0008] This document also provides, in certain aspects, a method for determining the presence of wavelet events with low false negatives and low false positives, the method comprising: (a) acquiring counts of nucleic acid sequence readings mapped to portions of a reference genome, wherein the sequence readings are readings of circulating cell-free nucleic acids from pregnant women; (b) normalizing the counts mapped to each portion to provide a calculated genomic segment level; (c) dividing the portions into subgroups of multiple portions; (d) determining the level of each subgroup based on the calculated genomic segment level; (e) determining the significance level of each level; and (f) determining the presence of wavelet events with low false negatives and low false positives based on the significance level determined for each level.

[0009] This article also provides, in some aspects, a method for determining the presence of chromosomal aneuploidy, microduplication, or microdeletion in a fetus with low false negatives and low false positives, the method comprising:

[0010] (a) Obtaining counts of nucleic acid sequence reads mapped to a portion of the reference genome, wherein the sequence reads are readings of circulating cell-free nucleic acids from a pregnant woman; (b) Standardizing the counts mapped to each portion to provide a calculated genomic segment level; (c) Selecting segments of the genome to provide groups of the portion; (d) Recursively dividing the groups of the portion to provide two or more subgroups of the portion; (e) Determining the level of each of the two or more subgroups of the portion; (f) Determining, based on the level determined in (e), the presence of chromosomal aneuploidy, microduplication, or microdeletion in the fetus with low false negatives and low false positives for the sample.

[0011] This document also provides a system comprising one or more processors and memory, wherein the memory contains instructions executable by the one or more processors, and the memory contains a count of nucleic acid sequence reads mapped to portions of a reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acids from pregnant women, and wherein the one or more processor-executable instructions are configured to...

[0012] (a) Obtaining counts of nucleic acid sequence reads mapped to portions of a reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acids from a pregnant woman; (b) Standardizing the counts mapped to each portion to provide a calculated genome segment level; (c) Generating a profile of genome segments based on the calculated genome segment level; (d) Segmenting the profile to provide two or more decomposed maps; and (e) Determining, based on the two or more decomposed maps, the presence of chromosomal aneuploidy, microduplication, or microdeletion in the fetus with low false negatives and low false positives.

[0013] Certain technical aspects are further described in the following description, embodiments, claims and drawings. Brief description of the attached figures

[0014] The accompanying drawings illustrate embodiments of the present technology but are not limiting. For clarity and convenience, the drawings are not made to scale and in some cases, various aspects may be exaggerated or enlarged to aid in understanding the specific embodiments.

[0015] Figure 1 This diagram illustrates the wavelet method. Standardized partial count data (upper right figure) undergoes wavelet transform to produce a wavelet-smoothed profile (lower right figure). Non-uniform events are clearly observed after wavelet denoising.

[0016] Figure 2 This displays the effect of leveling without thresholding. The optimal level can be determined by the desired event size.

[0017] Figure 3The figure shows the profiles of non-uniformity (top) and wavelet transform (middle) of the chromosome 13 sample. The figure below shows the height distribution of empty margins obtained from multiple euploid reference samples of chromosome 13. In the middle figure, the two larger differences (circles) correspond to the boundaries of non-uniform events.

[0018] Figure 4 Examples of segments appearing after wavelet or CBS are shown. The three original segments (left image, right half of the chromosome) are merged into a single long extension (right image, right half of the chromosome), making microreplication clearly visible.

[0019] Figure 5 A- Figure 5 E shows the chromosome profile, which is wavelet-smoothed. Figure 5 B) CBS smoothing ( Figure 5 C) and the merging of sections ( Figure 5 D and Figure 5 E). The two optimal segments of the two methods are compared with each other and "cross-confirmed".

[0020] Figure 6A and Figure 6B This shows a non-limiting example of decision analysis. The same elements (e.g., boxes) in the flowcharts shown are optional. Additional elements (e.g., verification) may be added in some implementations.

[0021] Figure 7 Showing a non-restrictive example of a comparison extending from 650.

[0022] Figure 8 This is a non-restrictive example showing a comparison of the two wavelet events represented by 631 and 632.

[0023] Figure 9 The chromosome profile is shown in (A), which is obtained by wavelet smoothing and merging (B) and CBS smoothing and merging (C). After comparison, the two optimal segments of the two methods are found to be "cross-restricted" from each other.

[0024] Figure 10This shows a segmental profile of chromosome 22 associated with a genetic variant linked to DiGeorge syndrome. Genetic microdeletions and microduplications associated with DiGeorge syndrome have been mapped to this region. The left-hand profile (Fig. AG) ​​is segmented, smoothed, merged, and compared using Haar wavelet and CBS methods. The complex profile is shown in the right-hand profile (A'-G'). Differences in sample load per cell stream are shown in the figures: A-A', 0.5-plex; B-B', 1-plex; C-C', 2-plex; D-D', 3-plex; E-E', 4-plex; F-F', 5-plex; and G-G', 6-plex. DiGeorge microdeletions were detected even with a 10-fold reduction in sample read coverage (see, for example, Fig. F').

[0025] Figure 11A The display of complex wavelet events indicates that microdeletions were detected in the profile of chromosome 1. Figure 11B The display of complex wavelet events indicates that micro-replication was detected in the profile of chromosome 2.

[0026] Figure 12 A representative example is shown, demonstrating the location of microreplications in chromosome 12 detected using the maximum entropy method.

[0027] Figure 13 This displays magnified views of the DiGeorge regions of 16 samples, labeled with numerical pairs indicating their position on the plate. Sample pairs 3_4 (second to last) and 9_10 (fifth to last) belong to infant DiGeorge pregnancies. All other samples are euploid. The highlighted boxes (gray areas) show the overlap between the DiGeorge regions and the PERUN partial selections (referencing portions of genomes chr22_368-chr22_451).

[0028] Figure 14 The Z-scores for the DiGeorge region are displayed. Each data point is derived from the sum of two profiles, each derived from two separate equivalences for each patient. Z-standardization was performed based on all 16 patients, including the two affected cases.

[0029] Figure 15-16 Representative histograms for samples 3_4 (DiGeorge) and 1_2 (euploid) are shown. Each histogram shows the distribution of Z-scores obtained from the 15x15 grid region contained within the DiGeorge region. The region is selected by sliding a portion along the left and right edges of the DiGeorge region, moving inwards from the outer edge. Histograms for samples 3_4 and 9_10 (not shown) are also shown. Figure 1 The histograms of samples 13 and 14 (not shown) showed a significant loss, with only a few Z-scores exceeding Z = -3 in samples 3 and 4. Figure 1 This indicates overrepresentation, with only a few regions achieving Z-scores below 3. All other samples (e.g., 1_2) are confined to the Z-score range of [-3, 3].

[0030] Figure 17 The median Z-score and its ±3 MAD confidence interval are shown for each of the 16 samples. The median Z-score was determined from a 15x15 grid area (225 regions) obtained from the sliding edge. For the vast majority of the DiGeorge subregions, the known DiGeorge samples (3_4 and 9_10) have Z-scores below -3. The obvious repetition in sample 13_14 is confirmed by the fact that most of its Z-scores exceed 3. The Z-scores of all other samples are constrained within the [-3, 3] range.

[0031] Figure 18-19 Representative histograms for samples 3_4 (DiGeorge) and 1_2 (euploid) are shown. Each histogram shows the distribution of Z-scores obtained for the DiGeorge region. Each Z-score was calculated using 16 different sets of reference samples, using the "leave one out" method. The histogram for sample 9_10 (not shown) confirms depletion. Depending on the reference settings, sample 3_4 is either completely depleted or has boundary Z-scores. The histogram for sample 13_14 (not shown) indicates over-presentation, with a few boundary Z-scores. All other samples (including 1_2) are confined to the Z-score range of [-3, 3].

[0032] Figure 20 Displays the median Z-score for each sample, which is represented by... Figure 18-19 The median Z-score and its ±3MAD confidence interval were calculated from 16 different reference samples identified using the leave-one-out method. For the vast majority of the reference sample subgroups, the known DiGeorge samples (3_4 and 9_10) had Z-scores below -3. The obvious repetition in samples 13_14 was confirmed by the fact that most of their Z-scores exceeded 3. The Z-scores of all other samples were constrained within the [-3, 3] range.

[0033] Figure 21 This shows a comparison of the median Z-score obtained using the 15x15 grid DiGeorge subregion (x-axis) with the median Z-score generated by the leave-one-out technique (y-axis) for each of the 16 samples. The diagonal represents ideal consistency (slope = 1, intercept = 0).

[0034] Figure 22-23Representative histograms for samples 3_4 (DiGeorge) and 1_2 (euploid) are shown. Each histogram shows the distribution of Z-scores obtained for subregions within the DiGeorge region, using 16 different groups of reference samples. The subregions were randomly selected from 225 subregions in a 15x15 grid. Leave-one-out analysis confirmed depletion for samples 3_4 (Fig. 37) and 9_10 (not shown). The histogram for sample 13_14 confirmed over-presentation (not shown). All other samples (including 1_2) were confined to the Z-score range of [-3, 3].

[0035] Figure 24 This shows the median Z-score and its ±3 MAD confidence interval for each of the 16 samples, randomly selected using the "leave one out" method for the subregions of the DiGeorge region. For most reference samples, the known DiGeorge samples (3_4 and 9_10) have Z-scores below -3. The obvious repetition in samples 13_14 is indicated by the fact that most of their Z-scores are above 3. Except for samples 17_18, the Z-scores of all other samples are confined to the [-3, 3] range.

[0036] Figure 25-26 Representative histograms for samples 3_4 (DiGeorge) and 1_2 (Euploid) are shown, representing the distribution of Z-scores obtained across all 225 subregions of the DiGeorge region, using 16 different sets of reference samples. For each sample, 225 subregions were generated on a 15x15 grid using the sliding edge method. The sliding edge method was used in combination with leave-one-out analysis. The results confirmed that both affected samples 3_4 and 9_10 (not shown) were depleted. The histogram for sample 13_14 confirmed over-representation (not shown). All other samples (including 1_2) were confined to the Z-score range of [-3,3], with the exception of a few in 17_18 (not shown).

[0037] Figure 27 This shows a comparison between the median Z-score obtained using a 15x15 grid with the "leave-one-out" method in the DiGeorge subregion and the median Z-score obtained using only a 15x15 grid. The diagonal represents ideal consistency (slope = 1, intercept = 0).

[0038] Figure 28 This shows a comparison of the MAD of the Z-score obtained using a combination of a 15x15 grid with the "leave-one-out" technique in the DiGeorge subregion with the MAD of the Z-score obtained using only a 15x15 grid. The diagonal represents ideal consistency (slope = 1, intercept = 0).

[0039] Figure 29The median Z-score and its ±3 MAD confidence interval are displayed, evaluated using a combined leave-one-out method on a full 15x15 grid of the canonical DiGeorge subregion. For most reference samples, the known DiGeorge samples (3_4 and 9_10) have Z-scores below -3. The apparent repetition in sample 13_14 is indicated by the fact that most of its Z-scores exceed 3. Except for sample 17_18, the Z-scores of all other samples are confined to the [-3, 3] range.

[0040] Figure 30 Exemplary implementations of the display system, wherein certain implementations of the technology may be carried out.

[0041] Figure 31 The classification results of male LDTv2 samples are shown using the logarithmic concession ratio (LOR) method.

[0042] Figure 32 Some aspects of Equation 23 as illustrated in Embodiment 6.

[0043] Figure 33 This shows an implementation of the GC density provided by the Epanechnikov kernel (bandwidth = 200bp).

[0044] Figure 34 This plot shows the GC density (y-axis) of the HTRA1 gene, where the GC density is normalized across the entire genome. Genomic locations are shown on the x-axis.

[0045] Figure 35 The local genomic offset assessment (e.g., GC density, x-axis) shows the reference genome (solid line) and the sequence reads obtained from the sample (dashed line). Offset frequencies (e.g., density frequencies) are shown on the y-axis. The GC density assessment is normalized across the entire genome. In this embodiment, the sample has more high GC content reads than expected from the reference.

[0046] Figure 36 The GC density assessment distributions of the reference genome and sample sequence reads are displayed, using a weighted third-order polynomial fit. GC density assessments (x-axis) are normalized across the entire genome. GC density frequencies are represented on the y-axis as log2 of the ratio of the reference density frequency to the sample density frequency.

[0047] Figure 37A This shows the distribution of median GC density (x-axis) across all parts of the genome. Figure 37BDisplays the median absolute deviation (MAD) value (x-axis) determined based on the GC density distributions of multiple samples. GC density frequencies are shown on the y-axis. A subset is selected based on the median GC density distribution of multiple reference samples (e.g., the training group) and the MAD value determined based on the GC density distributions of multiple samples. GC densities exceeding a predetermined threshold (e.g., four times the interquartile range of MAD) are removed from consideration according to the selection method.

[0048] Figure 38A This displays an overview of the sample's read density, including the median read density within the genome (y-axis, e.g., read density / part) and the relative position of each genomic part (x-axis, part index). Figure 38B The first principal component (PC1) is shown. Figure 38C The second principal components (PC2) are shown, which are obtained from the principal component analysis of the reading density profiles of the training group of 500 euploids.

[0049] Figure 39A -C shows an example of a sample read density profile of a genome, including trisomy 21 (e.g., enclosed by two vertical lines). The relative positions of the genomic portions are shown on the x-axis. The read density is shown on the y-axis. Figure 39A Displays a summary of the density of raw (e.g., uncalibrated) readings. Figure 39B Shows the overview of 39A, including the first adjustment (including the median deduction overview). Figure 39C The profile shown in 39B includes a second adjustment. This second adjustment includes subtracting 8x the principal component profile, weighted based on its representativeness found in the sample (e.g., modeling). For example, the sample profile = A*PC1 + B*PC2 + C*PC3…, while the corrected profile (e.g., shown in 39C) = sample profile - A*PC1 + B*PC2 + C*PC3….

[0050] Figure 40 This displays a QQ plot showing the test p-values ​​of the bootstrapped training samples for the T21 test. QQ plots are typically used to compare two distributions. Figure 40 This displays a comparison of the ChAI score (y-axis) of the test samples with a uniform distribution (i.e., the expected distribution of p-values, x-axis). Each point represents the score of the log-p value for a single test sample. Samples are sorted and assigned "expected" values ​​(x-axis) based on a uniform distribution. The lower dashed line represents the diagonal, and the upper line represents the Bonferroni threshold. Samples following a uniform distribution are expected to fall on the lower diagonal (lower dashed line). Due to correlations (e.g., offsets) in some parts, values ​​deviate from the diagonal, indicating that the sample's score is higher than expected (lower p-value). The methods described herein (e.g., ChAI, see Example 7, for example) can correct for this observed offset.

[0051] Figure 41A Display a reading density plot showing the difference in PC2 coefficients between men and women in the training group. Figure 41B The receiver operating characteristic (ROC) curves for sex call with PC2 coefficients are shown. Sex call performed via sequencing was used as a true reference.

[0052] Figures 42A-42B Implementation methods for display systems. Invention Details

[0053] This document provides methods for identifying fetal genetic variations (e.g., chromosomal aneuploidy, microduplication, or microdeletion) in a fetus, wherein the identification is partially and / or entirely based on nucleic acid sequences. In some embodiments, the nucleic acid sequences are obtained from samples taken from pregnant women (e.g., blood from pregnant women). This document also provides improved data manipulation methods, and in some embodiments, systems, apparatus, and modules for performing the methods described herein. In some embodiments, the identification of genetic variations by the methods described herein can guide the diagnosis of a specific medical condition or determine a predisposition to a specific medical condition. Identifying genetic variations can aid in medical decision-making and / or the use of beneficial medical treatments.

[0054] sample

[0055] This document provides methods and compositions for analyzing nucleic acids. In some embodiments, nucleic acid fragments in a mixture of nucleic acid fragments are analyzed. The nucleic acid mixture may include two or more types of nucleic acid fragments having different nucleotide sequences, different fragment lengths, different sources (e.g., genomic, fetal and maternal, cell or tissue, sample, object, etc.) or combinations thereof.

[0056] The nucleic acids or mixtures of nucleic acids used in the methods and apparatus described herein are often isolated from samples obtained from a subject. The subject can be any living or non-living organism, including but not limited to humans, non-human animals, plants, bacteria, fungi, or protozoa. Any human or non-human animal can be selected, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovids (e.g., cattle), equines (e.g., horses), goats and sheep (e.g., sheep, goats), suidae (e.g., pigs), alpacas (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), bears (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The subject can be male or female (e.g., women, pregnant women). The subject can be of any age (e.g., embryos, fetuses, infants, children, adults).

[0057] Nucleic acids can be isolated from any type of suitable biological sample or specimen (e.g., a test sample). The sample or test sample can be any specimen isolated from or obtained from an object or a portion thereof (e.g., a human object, a pregnant woman, a fetus). Non-limiting examples of samples include liquids or tissues of the object, including but not limited to blood or blood products (e.g., serum, plasma, etc.), cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, cerebrospinal fluid, lavage fluid (e.g., bronchoalveolar, gastric, peritoneum, catheter, ear, arthroscopy), biopsy samples (e.g., from pre-implantation embryos), intermembranous fluid samples, cells (blood cells, placental cells, embryonic or fetal cells, fetal nucleated cells, or fetal cell remnants) or portions thereof (e.g., mitochondria, nuclei, extracts, etc.), female genital tract lavage fluid, urine, feces, sputum, saliva, nasal mucosa, prostatic fluid, lavage fluid, semen, lymph, bile, tears, sweat, breast milk, mammary gland fluid, etc., or combinations thereof. In some embodiments, the biological sample is a cervical swab from the object. In some embodiments, the biological sample can be blood, and sometimes plasma or serum. As used herein, the term "blood" refers to a blood sample or product derived from a pregnant woman or a woman being tested for a possible pregnancy. The term encompasses whole blood, blood products, or any portion of blood, such as serum and plasma as conventionally defined, tannins, etc. Blood or portions thereof often include nucleosomes (e.g., maternal and / or fetal nucleosomes). Nucleosomes include nucleic acids and are sometimes cell-free or intracellular. Blood also includes a tannin. The tannin is sometimes separated using a Ficoll gradient. The tannin may include leukocytes (e.g., white blood cells, T cells, B cells, platelets, etc.). In some embodiments, the tannin includes maternal and / or fetal nucleic acids. Blood plasma refers to the portion of whole blood obtained by centrifugation of blood treated with an anticoagulant. Blood serum refers to the liquid aqueous layer retained after a blood sample has clotted. Liquid or tissue samples are typically collected according to standard methods followed in hospitals or clinical practice. In the case of blood, an appropriate amount of peripheral blood (e.g., 3-40 ml) is typically collected and preserved according to standard procedures before or after preparation. Liquid or tissue samples used for nucleic acid extraction may be cell-free (e.g., cell-free). In some embodiments, the liquid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, the sample may contain fetal cells or cancer cells.

[0058] The samples are typically heterogeneous, meaning they contain more than one type of nucleic acid material. For example, heterogeneous nucleic acids can include, but are not limited to, (i) fetal and maternal nucleic acids, (ii) cancer and non-cancer nucleic acids, (iii) pathogen and host nucleic acids, and more commonly, (iv) mutated and wild-type nucleic acids. Samples can be heterogeneous because they contain more than one cell type, such as fetal and maternal cells, cancer and non-cancer cells, or pathogen and host cells. In some embodiments, a few and a majority of nucleic acid materials are present.

[0059] For the prenatal application of the techniques described herein, liquid or tissue samples may be collected from women of gestational age suitable for testing or women who have been tested and are likely pregnant. Suitable gestational age may vary depending on the prenatal test performed. In some embodiments, the pregnant woman is sometimes in the first trimester, sometimes in the second trimester, or sometimes in the last trimester. In some embodiments, the liquid or tissue is collected from pregnant women whose fetus is approximately 1–approximately 45 weeks pregnant (e.g., fetal gestational age 1–4, 4–8, 8–12, 12–16, 16–20, 20–24, 24–28, 28–32, 32–36, 36–40, or 40–44 weeks) and sometimes whose fetus is approximately 5–approximately 28 weeks pregnant (e.g., fetal gestational age 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 weeks). In some implementations, fluid or tissue samples are collected from pregnant women during or immediately after childbirth (e.g., vaginal or non-vaginal delivery, such as cesarean delivery) or 0-72 hours after delivery.

[0060] Obtaining blood samples and DNA extraction

[0061] The methods described herein include isolating, enriching, and analyzing fetal DNA found in maternal blood as a non-invasive means of detecting the presence of maternal and / or fetal genetic variations and / or monitoring the health of the fetus and / or the pregnant woman during pregnancy and sometimes after pregnancy. Therefore, a first step in implementing certain methods of the present invention includes obtaining a blood sample from a pregnant woman and extracting DNA from the sample.

[0062] Obtaining blood samples

[0063] Blood samples can be obtained from pregnant women of gestational age suitable for testing using the methods described in this invention. The appropriate gestational age may vary depending on the disease being tested, as described below. Blood collection from women is typically performed according to standard protocols generally followed in hospitals or clinics. An appropriate amount of peripheral blood is collected, typically 5-50 ml, and preserved according to standard procedures before further preparation. The blood samples can be collected, preserved, or transported in a manner that minimizes degradation of the nucleic acids present in the sample or ensures their quality.

[0064] Preparation of blood samples

[0065] Fetal DNA found in maternal blood is analyzed using, for example, whole blood, serum, or plasma. Methods for preparing serum or plasma from maternal blood are known. For example, blood from a pregnant woman can be placed in a tube containing EDTA to prevent blood clotting, or a commercially available product such as Vacutainer SST (Becton Dickinson, Franklin Lake, New Jersey), and plasma can then be obtained from the whole blood by centrifugation. Serum can be obtained, with or without centrifugation after blood clotting. If centrifugation is used, it is typically (but not limited to) performed at a suitable speed (e.g., 1,500–3,000 g). Plasma or serum may undergo additional centrifugation steps before being transferred to new tubes for DNA extraction.

[0066] In addition to the non-cellular portion of whole blood, DNA can also be recovered from the cellular components and enriched in the tannin layer, which can be obtained by centrifuging a woman's whole blood sample and removing the plasma.

[0067] DNA extraction

[0068] There are several known methods for extracting DNA from biological samples, including blood. These can be done using standard DNA preparation methods (e.g., described in Sambrook and Russell, *Molecular Cloning: A Laboratory Manual*, 3rd edition, 2001); or using a variety of commercially available reagents or kits, such as Qiagen's QIAamp Cyclic Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiagen, Heldon, Germany), GenomicPrep... TM Blood DNA Separation Kit (Promega, Madison, Wisconsin) and GFX TM The Genomic Blood DNA Purification Kit (Amersham, Piscateway, NJ) can also be used to obtain DNA from blood samples from pregnant women. Combinations of more than one of these methods can also be used.

[0069] In some embodiments, the sample may first be enriched or relatively enriched for fetal nucleic acids using one or more methods. For example, the differentiation between fetal and maternal DNA may be performed using the compositions and methods described in this invention alone or in combination with other differentiating factors. Examples of such factors include, but are not limited to, single nucleotide differences in chromosomes X and Y, chromosome Y-specific sequences, polymorphisms elsewhere in the genome, size differences between fetal and maternal DNA, and differences in methylation forms between maternal and fetal tissues.

[0070] Other methods for enriching samples with specific nucleic acid material are described in PCT patent application No. PCT / US07 / 69991, filed May 30, 2007; PCT patent application No. PCT / US2007 / 071232, filed June 15, 2007; U.S. Provisional Applications Nos. 60 / 968,876 and 60 / 968,878 (as assigned to the applicant); and PCT patent application No. PCT / EP05 / 012707, filed November 28, 2005, all of which are incorporated herein by reference. In some embodiments, the parent nucleic acid is selectively removed (partially, substantially, almost completely, or completely) from the sample.

[0071] The terms “nucleic acid” and “nucleic acid molecule” are used interchangeably herein. The term refers to nucleic acids in any composite form, derived from, for example: DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), RNA (e.g., messenger RNA (mRNA), short repressor RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed in the fetus or placenta, etc.), and / or DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and / or non-natural backbones, etc.), RNA / DNA hybrids, and polyamide nucleic acids (PNAs), all of which may be in single-stranded or double-stranded form, and unless otherwise specified, may encompass known analogs of natural nucleotides that function in a manner similar to naturally occurring nucleotides. In some embodiments, nucleic acids may be or may be derived from: plasmids, bacteriophages, autonomously replicating sequences (ARS), centromeres, artificial chromosomes, chromosomes, or other nucleic acids capable of replicating or being replicated in vitro or in a host cell, cell, cell nucleus, or cytoplasm. In some implementations, the template nucleic acid may be derived from a single chromosome (e.g., the nucleic acid sample may be derived from a chromosome of a sample obtained from a diploid organism). Unless explicitly defined, the term covers known analogs containing a reference nucleic acid with similar binding properties and metabolized in a manner similar to that of naturally occurring nucleotides. Unless otherwise stated, a particular nucleic acid sequence also includes its conserved modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences, as well as explicitly indicated sequences. Specifically, degenerate codon substitutions can be obtained by producing a sequence in which the third position of one or more selected (or all) codons is substituted with a mixture of bases and / or deoxyinosine residues. The term nucleic acid is used interchangeably with locus, gene, cDNA, and mRNA encoded by a gene. The term may also include equivalents, derivatives, variants, and analogs of RNA or DNA synthesized from nucleotide analogs, single-stranded ("sense" or "antisense", "positive" or "negative", "positive" reading frame or "reverse" reading frame), and double-stranded polynucleotides. The term “gene” refers to a segment of DNA involved in the production of a polypeptide chain; it includes regions before and after the coding region (leader and tail regions) involved in the transcription / translation of the gene product and the regulation of said transcription / translation, as well as insertion sequences (introns) between individual coding segments (exons).

[0072] Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. For RNA, the cytosine base is replaced with uracil. Template nucleic acids can be prepared using nucleic acids obtained from the target organism as templates.

[0073] Nucleic acid isolation and processing

[0074] Nucleic acids can be obtained from one or more sample sources (such as cells, serum, plasma, ochre layer, lymph, skin, soil, etc.) using methods known in the art. DNA can be isolated, extracted, and / or purified from biological samples (e.g., from blood or blood products) using any suitable method. Non-limiting examples include methods for DNA preparation (e.g., described in Sambrook and Russell, *Molecular Cloning: A Laboratory Manual*, 3rd edition, 2001); various commercially available reagents or kits, such as Qiagen's QIAamp Cyclic Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiaagen, Heldon, Germany), GenomicPrep... TM Blood DNA Separation Kit (Promega, Madison, Wisconsin) and GFX TM Genomic blood DNA purification kit (Amersham, Piscavenge, NJ) or combinations thereof.

[0075] Cell lysis methods and reagents are known in the art and can generally be performed by chemical (e.g., detergents, hypotonic solutions, enzymatic processes, etc., or combinations thereof), physical (e.g., French pressure filtration, sonication, etc.), or electrolytic lysis methods. Any suitable lysis process can be used. For example, chemical methods typically use a lysing agent to disrupt cells and extract nucleic acids from them, followed by treatment with a dissociative salt. Physical methods, such as freezing / thawing followed by grinding, and cell pressure filtration, are also useful. High-salt lysis is also commonly used. For example, alkaline lysis can be employed. The latter method conventionally involves the use of a phenol-chloroform solution, and alternatively, a phenol-chloroform-free method comprising three solutions can be used. In the latter method, one solution may contain 15 mM Tris, pH 8.0; 10 mM EDTA and 100 μg / ml RNase A; the second solution may contain 0.2 N NaOH and 1% SDS; and the third solution may contain 3 M KOAc, pH 5.5. These methods can be found in 6.3.1–6.3.6 (1989) of the Current Protocols in Molecular Biology published by John Wiley & Sons, Inc., New York, which is included in this paper in its entirety.

[0076] Nucleic acids can also be isolated at different time points than other nucleic acids, with each sample originating from the same or different sources. Nucleic acids can be derived from nucleic acid libraries, such as cDNA or RNA libraries. Nucleic acids can be products of nucleic acid purification or isolation and / or amplification of nucleic acid molecules in a sample. Nucleic acids provided for the methods described herein may comprise nucleic acids from one sample or from two or more samples (e.g., from 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more samples).

[0077] In some embodiments, nucleic acids may include extracellular nucleic acids. As used herein, the term "extracellular nucleic acid" refers to nucleic acids isolated from a substantially cell-free source, also known as "cell-free" nucleic acids and / or "circulating cell-free" nucleic acids. Extracellular nucleic acids may be present in and obtained from blood (e.g., from the blood of a pregnant woman). Extracellular nucleic acids typically do not contain detectable cells and may contain cellular elements or cellular remnants. Non-limiting examples of cell-free sources of extracellular nucleic acids include blood, plasma, serum, and urine. As used herein, the term "obtaining circulating cell-free sample nucleic acid" includes obtaining a sample directly (e.g., collecting a sample, such as a test sample) or obtaining a sample from a person who has already collected a sample. Without being theoretically limited, extracellular nucleic acids may be products of apoptosis and cell lysis, which often results in extracellular nucleic acids having a range of lengths (e.g., "ladders").

[0078] In some implementations, extracellular nucleic acids may contain different nucleic acid substances, and are therefore referred to herein as "heterogeneity." For example, the blood serum or plasma of a person with cancer may contain nucleic acids from cancer cells and nucleic acids from non-cancer cells. In another example, the blood serum or plasma of a pregnant woman may contain maternal nucleic acids and fetal nucleic acids. In some examples, fetal nucleic acids sometimes account for about 5% to about 50% of all nucleic acids (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49% of the total nucleic acids are fetal nucleic acids). In some embodiments, the length of most fetal nucleic acids in the nucleic acid is about 500 base pairs or less, about 250 base pairs or less, about 200 base pairs or less, about 150 base pairs or less, about 100 base pairs or less, about 50 base pairs or less, or about 25 base pairs or less.

[0079] In some embodiments, nucleic acids may be provided for performing the methods described herein without processing the nucleic acid-containing sample. In some embodiments, nucleic acids are provided for performing the methods described herein after processing the nucleic acid-containing sample. For example, nucleic acids may be extracted, isolated, purified, partially purified, or amplified from the sample. As used herein, the term "isolation" means the removal of nucleic acids from their original environment (e.g., the natural environment in which nucleic acids are naturally produced or the host cell expressing exogenous nucleic acids), thus altering the nucleic acids from their original environment through human intervention (e.g., "artificial"). As used herein, the term "isolated nucleic acid" refers to nucleic acids removed from an object (e.g., a human object). Isolated nucleic acids may contain fewer non-nucleic acid components (e.g., proteins, lipids) compared to the component content present in the source sample. Compositions containing isolated nucleic acids may be about 50% to more than 99% free of non-nucleic acid components. Compositions containing isolated nucleic acids may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more than 99% free of non-nucleic acid components. As used herein, the term "purified" refers to nucleic acids containing fewer non-nucleic acid components (e.g., proteins, lipids, carbohydrates) compared to the amount of non-nucleic acid components present prior to the purification process. Compositions containing purified nucleic acids may be approximately 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other non-nucleic acid components. The term "purified" as used herein may also refer to nucleic acids containing fewer nucleic acid substances compared to the sample source from which they are derived. Compositions containing purified nucleic acids may be approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other nucleic acid substances. For example, fetal nucleic acids can be purified from a mixture containing maternal and fetal nucleic acids. In some examples, nucleosomes containing small fragments of fetal nucleic acid can be purified from a mixture of macronucleosome complexes containing larger fragments of maternal nucleic acid.

[0080] In some embodiments, nucleic acids are fragmented or cleaved before, during, or after the method of the present invention. The fragmented or cleaved nucleic acids may have a nominal, average, or mean length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs. Fragments can be generated by suitable methods known in the art, and the average, geometric mean, or nominal length of the nucleic acid fragments can be controlled by selecting an appropriate fragment generation method.

[0081] Nucleic acid fragments may contain overlapping nucleotide sequences, which can facilitate the construction of nucleotide sequences of unfragmented corresponding nucleic acids or segments thereof. For example, one fragment may have subsequences x and y, and other fragments may have subsequences y and z, where x, y, and z are nucleotide sequences of 5 nucleotides or longer. In some embodiments, overlapping nucleic acid y can be used to facilitate the construction of the xyz nucleotide sequence from the nucleic acid of a sample. In some embodiments, the nucleic acid may be partially fragmented (e.g., from an incomplete or terminated specific splicing reaction) or fully fragmented.

[0082] In some embodiments, nucleic acids may be fragmented or cleaved by suitable methods, non-limiting examples of which include physical methods (e.g., shearing, sonication, French press filtration, heat, UV irradiation, etc.), enzyme processing (e.g., enzyme cleavage reagents (e.g., suitable nucleases, suitable restriction enzymes, suitable methylation-sensitive restriction enzymes)), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, alkaline hydrolysis, heat, etc. or combinations thereof), the methods described in U.S. Patent Application Publication 20050112590, etc., or combinations thereof.

[0083] As used in this article, “fragmentation” or “splitting” refers to a method or condition that allows a nucleic acid molecule (such as a nucleic acid template gene molecule or its amplification product) to be divided into two or more smaller nucleic acid molecules. Such fragmentation or splicing can be sequence-specific, base-specific, or non-specific, and can be accomplished by any different methods, reagents, or conditions (including, for example, chemical, enzymatic, and physical fragmentation).

[0084] As used herein, the terms “fragment,” “cut product,” “cut product,” or their grammatical variations refer to nucleic acid molecules obtained by fragmentation or cutting of a nucleic acid template gene molecule or its amplification product. While such fragments or cut products can refer to all nucleic acids obtained by a cutting reaction, they generally refer only to nucleic acid molecules obtained by fragmentation or cutting of a segment of a nucleic acid template gene molecule or its amplification product (containing the corresponding nucleotide sequence of the nucleic acid template gene molecule). As used herein, the term “amplification” refers to the process of generating amplicon nucleic acids in a processed sample in a linear or exponential manner, the nucleotide sequence of which is identical or substantially identical to the nucleotide sequence of the target nucleic acid or its segment. In some embodiments, the term “amplification” refers to a method including polymerase chain reaction (PCR). For example, the amplification product can contain one or more additional nucleotides than the amplified nucleotide region of the nucleic acid template sequence (e.g., primers can contain “extra” nucleotides, such as transcription initiation sequences, in addition to nucleotides complementary to those of the nucleic acid template gene molecule, resulting in an amplification product containing “extra” nucleotides or nucleotides not corresponding to the amplified nucleotide region of the nucleic acid template gene molecule). Therefore, a fragment can contain a segment or portion of an amplified nucleic acid molecule, which at least partially contains nucleotide sequence information from or based on a representative nucleic acid template molecule.

[0085] As used herein, the term "complementary shearing reaction" refers to a shearing reaction on the same nucleic acid using different shearing agents or by altering the shearing specificity of the same shearing agent, thereby producing different shearing patterns of the same target or reference nucleic acid or protein. In some embodiments, nucleic acids may be treated in one or more reaction vessels using one or more specific shearing agents (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more specific shearing agents). As used herein, the term "specific shearing agent" refers to a reagent, sometimes a chemical or enzyme that can cleave nucleic acids at one or more specific sites.

[0086] Before being used in the methods described herein, nucleic acids may be treated to modify certain nucleotides within them. For example, the nucleic acid may be subjected to a treatment that selectively modifies the nucleic acid based on the methylation state of the nucleotides. Furthermore, conditions such as high temperature, ultraviolet radiation, and X-ray radiation can induce variations in the nucleic acid molecular sequence. Nucleic acids may be provided in any suitable form for appropriate sequence analysis.

[0087] Nucleic acids can be single-stranded or double-stranded. For example, single-stranded DNA can be generated by denaturing double-stranded DNA through heating or, for example, treatment with an alkali. In some embodiments, the nucleic acid is a D-loop structure, formed by the invasion of an oligonucleotide or DNA-like molecule, such as peptide nucleic acid (PNA), into the middle strand of a double-stranded DNA molecule. Adding E. coli RecA protein and / or altering the salt concentration (e.g., using methods known in the art) facilitates the formation of D-loops.

[0088] Determine fetal nucleic acid content

[0089] In some embodiments, the amount of fetal nucleic acid in the nucleic acid is determined (e.g., concentration, relative amount, absolute amount, copy number, etc.). In some embodiments, the amount of fetal nucleic acid in the sample is referred to as the “fetal fraction”. In some embodiments, the “fetal fraction” refers to the fetal nucleic acid fraction in circulating cell-free nucleic acid obtained from a sample (e.g., blood sample, serum sample, plasma sample) obtained from a pregnant woman. In some embodiments, the amount of fetal nucleic acid is determined based on: male fetus-specific markers (e.g., Y chromosome STR markers (e.g., DYS19, DYS 385, DYS 392 markers); RhD markers in RhD-negative women), the allele ratio of polymorphic sequences, or one or more markers specific to fetal nucleic acids but not specific to maternal nucleic acids (e.g., differential epigenetic biomarkers between mother and fetus (e.g., methylation; detailed below), or fetal RNA markers in maternal plasma (see, for example, Lo, 2005, Journal of Histochemistry and Cytochemistry 53(3):293-296)).

[0090] Determining fetal nucleic acid content (e.g., fetal fraction) is sometimes performed using fetal quantification assays (FQAs), as described in U.S. Patent Application Publication 2010 / 0105049, which is incorporated herein by reference. Such assays allow for the detection and quantification of fetal nucleic acids in maternal samples based on the methylation status of nucleic acids in the sample. In some embodiments, the content of fetal nucleic acids in the maternal sample can be determined relative to the total amount of nucleic acids present, thereby providing a percentage of fetal nucleic acids in the sample. In some embodiments, the copy number of fetal nucleic acids in the maternal sample can be determined. In some embodiments, the amount of fetal nucleic acids can be determined in a sequence-specific (or partial-specific) manner, and sometimes with sufficient sensitivity for precise chromosome dosing analysis (e.g., to detect the presence or absence of fetal aneuploidy, microduplication, or microdeletion).

[0091] Fetal Quantification Assay (FQA) can be performed in conjunction with any of the methods described herein. This assay can be performed by any method known in the art and / or as described in U.S. Patent Application Publication 2010 / 0105049, such as methods that can distinguish maternal and fetal DNA based on differential methylation status, and methods that quantify fetal DNA (i.e., determine its content). Methods for distinguishing nucleic acids based on methylation status include, but are not limited to, methylation-sensitive capture (e.g., using the MBD2-Fc fragment, wherein the methylation-binding domain of MBD2 is fused to the Fc fragment of an antibody (MBD-FC) (Gebhard et al. (2006) Cancer Res. 66(12): 6118-28)); methylation-specific antibodies; bisulfite conversion methods, such as MSP (methylation-sensitive PCR), COBRA, methylation-sensitive single nucleotide primer extension (Ms-SNuPE), or Sequenom MassCLEAVE. TM The invention relates to techniques and the application of methylation-sensitive restriction enzymes (e.g., digesting maternal DNA in a maternal sample with one or more methylation-sensitive restriction enzymes to enrich fetal DNA). Methylation-sensitive enzymes can also be used to distinguish nucleic acids based on methylation status, such as preferential or significant cleavage or digestion when their DNA recognition sequences are unmethylated. Thus, unmethylated DNA samples are cleaved into smaller fragments than methylated samples, while highly methylated DNA samples are not cleaved. Unless explicitly stated otherwise, any method for distinguishing nucleic acids based on methylation status can be used in the compositions and methods of the present invention. The amount of fetal DNA can be determined, for example, by introducing one or more competing agents at known concentrations during the amplification reaction. The amount of fetal DNA can also be determined, for example, by RT-PCR, primer extension, sequencing, and / or counting. In some examples, the BEAMing technique described in U.S. Patent Application Publication 2007 / 0065823 can be used to determine the amount of nucleic acid. In some embodiments, a restriction efficacy can be determined and the amount of fetal DNA can be further determined using this efficiency ratio.

[0092] In some embodiments, fetal quantification (FQA) can be determined using the concentration of fetal DNA in a maternal sample, for example by: a) determining the total amount of DNA present in the maternal sample; b) selectively digesting the maternal DNA in the maternal sample with one or more methylation-sensitive restriction enzymes to enrich the fetal DNA; c) determining the amount of fetal DNA from step b); and d) comparing the amount of fetal DNA obtained in step c) with the total amount of DNA obtained in step a) to determine the concentration of fetal DNA in the maternal sample. In some embodiments, the absolute copy number of fetal nucleic acid in the maternal sample can be determined, for example, using mass spectrometry and / or a system utilizing a competitive PCR method for determining the absolute copy number. See, for example, Ding and Cantor (2003) Proc. Natl. Acad. Sci. USA 100:3059-3064, and U.S. Patent Application Publication 2004 / 0081993, both of which are incorporated herein by reference.

[0093] In some embodiments, fetal fractions can be determined based on the allele ratio of polypeptide sequences (e.g., single nucleotide polymorphisms (SNPs)), for example using the method described in U.S. Patent Application Publication 2011 / 0224087, which is incorporated herein by reference. In this method, nucleotide sequence reads are obtained from a maternal sample, and the fetal fraction is determined by comparing the total number of nucleotide sequence reads mapped to a first allele with the total number of nucleotide sequence reads mapped to a second allele located at a reference polymorphic site (e.g., an SNP) in a reference genome. In some embodiments, fetal alleles are identified, for example, by the relatively small contribution of fetal alleles to a mixture of fetal and maternal nucleic acids in the sample, relative to the larger contribution of maternal nucleic acids to the mixture. Therefore, the relative abundance of fetal nucleic acids in the maternal sample can be determined as a parameter of the total number of unique sequence reads (for each of the two alleles at the polymorphic site) mapped to a target nucleic acid sequence on a reference genome.

[0094] In some embodiments, fetal fractions can be determined using methods that incorporate fragment length information (e.g., fragment length ratio (FLR) analysis, fetal ratio statistics (FRS) analysis, as described in International Application Publication WO2013 / 177086, which is incorporated herein by reference). Cell-free fetal nucleic acid fragments are typically shorter than maternally derived nucleic acid fragments (see, for example, Chan et al. (2004) Clin. Chem. 50:88-92; Lo et al. (2010) Sci. Transl. Med. 2:61ra91). Therefore, in some embodiments, fetal fractions can be determined by counting fragments below a specific length threshold and comparing said counts with, for example, counting fragments above the specific length threshold and / or the total nucleic acid content in the sample. Methods for counting nucleic acid fragments of a specific length are detailed in International Application Publication WO2013 / 177086.

[0095] In some implementations, the fetal fraction can be determined based on a part-specific fetal fraction estimate. Without being limited by any theory, the number of reads of fetal CCF fragments (e.g., fragments of a specific length or length range) is typically mapped to a fraction (e.g., within the same sample, or within the same sequencing run) along with the ranging frequency. Furthermore, without being limited by any theory, when comparing multiple samples, certain fractions may have similar reading representations to fetal CCF fragments (e.g., fragments of a specific length or length range), and these representations may be associated with a part-specific fetal fraction (e.g., the content, percentage, or proportion of fetal CCF fragments).

[0096] In some implementations, the part-specific fetal score estimate is determined in part based on part-specific parameters and their relationship to the fetal score. Part-specific parameters can be any suitable parameter reflecting the amount or proportion (e.g., related to) readings of a specific size (e.g., size range) of CCF fragment length within the fraction. Part-specific parameters can be the average, arithmetic mean, or median of part-specific parameters determined from multiple samples. Any suitable part-specific parameter can be used. Non-limiting examples of part-specific parameters include FLR (e.g., FRS), the amount of readings below the selected fragment length, genome coverage (i.e., coverage), mappability, counts (e.g., counts of sequence readings mapped to said fraction, such as normalized counts, PERUN normalized counts, ChAI normalized counts), DNase I sensitivity, methylation status, acetylation, histidine distribution, guanine-cytosine (GC) content, chromatin structure, etc., or combinations thereof. Partially specific parameters can be any suitable parameter that partially relates the FLR and / or FRS. In some embodiments, some or all of the partially specific parameters are direct or indirect representations of the FLR in relation to a portion. In some embodiments, the partially specific parameter is not the guanine-cytosine (GC) content.

[0097] In some embodiments, the part-specific parameter is any suitable value representing, associated with, or proportional to the amount of CCF fragment readings, wherein the length of said reading mapped to the part is shorter than the length of the selected fragment. In some embodiments, the part-specific parameter represents the amount of readings derived from a relatively short CCF fragment (e.g., about 200 base pairs or less) mapped to the part. CCF fragments shorter than the length of the selected fragment are typically relatively short CCF fragments, and sometimes the selected fragment length is about 200 base pairs or less (e.g., CCF fragments about 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, or 80 base pairs). The length of the CCF fragment or the readings derived from the CCF fragment can be determined (e.g., inferred or derived) by any suitable method (e.g., sequencing method, hybridization method). In some embodiments, the length of the CCF fragment is determined (e.g., inferred or derived) by readings obtained by paired-end sequencing. In some implementations, the CCF fragment template is determined directly from the length of a reading derived from the CCF fragment (such as a single-end reading).

[0098] Partial-specific parameters may be weighted or adjusted by one or more weighting factors. In some embodiments, weighted or adjusted partial-specific parameters provide a partial-specific fetal score estimate for a given sample (e.g., a test sample). In some embodiments, weighting or adjustment generally transforms a partial count (e.g., a reading mapped to a partial) or other partial-specific parameters into a partial-specific fetal score estimate; such transformations are sometimes referred to as alterations.

[0099] In some embodiments, the weighting factor is a coefficient or constant that partially describes and / or defines the relationship between fetal scores (e.g., fetal scores determined from multiple samples) and partial-specific parameters of multiple samples (e.g., a training group). In some embodiments, the weighting factor is determined based on the correlation between multiple fetal scores and multiple partial-specific parameters. One or more weighting factors may define the correlation, and one or more weighting factors may be determined from the correlation. In some embodiments, the weighting factor (such as one or more weighting factors) is determined from a partial fit correlation based on (i) the scores of each determined fetal nucleic acid in multiple samples, and (ii) the partial-specific parameters of multiple samples.

[0100] The weighting factor can be any suitable coefficient, estimated coefficient, or constant derived from a suitable correlation (e.g., suitable mathematical correlation, algebraic correlation, fitting correlation, regression, regression analysis, regression model). The weighting factor can be determined based on a suitable correlation, or it can be derived from or evaluated from a suitable correlation. In some implementations, the weighting factor is an evaluation coefficient derived from the fitting correlation. Fitting multiple samples to correlations is sometimes referred to as training the model. Any suitable model and / or method for fitting relationships (e.g., training the model on a training set) can be used. Non-limiting examples of suitable models include regression models, linear regression models, simple regression models, ordinary least squares regression models, multiple regression models, general multiple regression models, multinomial regression models, general linear models, generalized linear models, discrete choice regression models, logistic regression models, multinomial logit models, mixed logit models, probit models, multinomial probit models, ordered logit models, ordered probit models, Poisson models, multivariate response regression models, multilevel models, fixed effects models, random effects models, mixed models, nonlinear regression models, nonparametric models, semiparametric models, robust models, quantile models, isotonic models, principal component models, least angle models, local models, piecewise models, and variable error models. In some implementations, the fitted correlation is not a regression model. In some implementations, the fitted correlation is selected from decision tree models, support vector machine models, and neural network models. The result of training a model (e.g., a regression model, a correlation) is typically a mathematically describable correlation, where the correlation includes one or more coefficients (such as weighting factors). More complex multivariate models may determine 1, 2, 3, or more weighting factors. In some implementations, the model is trained based on fetal scores and two or more part-specific parameters (coefficients) obtained from multiple samples (e.g., by matrix fitting to the fitting relationship of multiple samples).

[0101] Weighting factors can be derived from suitable correlations through appropriate methods (e.g., suitable mathematical correlations, algebraic correlations, fit correlations, regression, regression analysis, regression models). In some implementations, fit correlations are fitted by evaluation, and non-limiting examples include least squares, ordinary least squares, linear, partial, total, generalized, weighted, nonlinear, iterative weighted, ridge regression, least squares, Bayesian, Bayesian multivariate, reduced-rank, LASSO, weighted rank selection criterion (WRSC), rank selection criterion (RSC), elastic network estimation (e.g., elastic network regression), and combinations thereof.

[0102] Weighting factors can be determined or associated with any suitable portion of the genome. Weighting factors can be determined or associated with any suitable portion of any suitable chromosome. In some embodiments, weighting factors can be determined or associated with some or all portions of the genome. In some embodiments, weighting factors can be determined or associated with portions of some or all chromosomes in the genome. Sometimes weighting factors can be determined or associated with selected portions of chromosomes. Weighting factors can be determined or associated with portions of one or more autosomes. Weighting factors can be determined or associated with portions of multiple portions that include portions of autosomes or subgroups thereof. In some embodiments, weighting factors can be determined or associated with portions of sex chromosomes (such as ChrX and / or ChrY). Weighting factors can be determined or associated with portions of one or more sex chromosomes and one or more autosomes. In some embodiments, weighting factors can be determined or associated with multiple portions of chromosomes X and Y and all autosomes. Weighting factors can be determined or associated with portions of multiple portions that do not include portions of chromosomes X and / or Y. In some embodiments, the weighting factor is determined or associated with a portion of a chromosome, wherein the chromosome contains aneuploidy (e.g., whole-chromosome aneuploidy). In some embodiments, the weighting factor is determined or associated with a portion of a chromosome, wherein the chromosome is not aneuploid (e.g., euploid chromosome). The weighting factor may be determined or associated with portions of a plurality of portions that do not include portions of chromosomes 13, 18, and / or 21.

[0103] In some embodiments, weighting factors are determined for portions based on one or more samples (e.g., a training set of samples). Weighting factors are typically portion-specific. In some embodiments, one or more weighting factors are assigned independently to portions. In some embodiments, weighting factors are determined based on relationships in fetal score determinations (e.g., sample-specific fetal score determinations) across multiple samples and based on portion-specific parameters determined across multiple samples. Weighting factors are typically determined from multiple samples, such as from about 20 to about 100,000 or more samples, from about 100 to about 100,000 or more samples, from about 500 to about 100,000 or more samples, from about 1,000 to about 100,000 or more samples, or from about 10,000 to about 100,000 or more samples. Weighting factors can be determined from euploid samples (e.g., samples from subjects with euploid fetuses, such as samples without aneuploid chromosomes). In some embodiments, weighting factors are obtained from samples containing aneuploid chromosomes (e.g., samples from subjects with euploid fetuses). In some implementations, a weighting factor is determined from multiple samples, said samples being from subjects with euploid fetuses and subjects with trisomy fetuses. The weighting factor may originate from a variety of samples, said samples being from subjects with male fetuses and / or female fetuses.

[0104] Fetal scores are typically determined from one or more samples in the training group, with weighting factors derived from these fetal scores. Sometimes, the fetal scores from which weighting factors are derived are sample-specific fetal scores. The fetal scores from which weighting factors are determined can be determined by any suitable method described herein or known in the art. In some embodiments, the determination of fetal nucleic acid content (e.g., fetal score) is performed using a suitable fetal quantification assay (FQA) described herein or known in the art, non-limiting examples of which include determining fetal scores based on: male fetal-specific markers, allele ratios based on polymorphic sequences, one or more markers specific to fetal nucleic acids but not specific to maternal nucleic acids, by utilizing methylation-based DNA recognition (e.g., A. Nygren, et al., (2010) Clinical Chemistry 56(10):1627–1635), by mass spectrometry methods and / or systems using competitive PCR methods, by the method described in U.S. Patent Application Publication No. 2010 / 0105049 (which is incorporated herein by reference), etc., or combinations thereof. Typically, the fetal score is determined at the Y chromosome level (e.g., at the level of one or more genomic segments, or at the profile level). In some implementations, the fetal score is determined based on an appropriate test of the Y chromosome (e.g., by comparing the amount of a fetal-specific locus (e.g., the SRY locus on the Y chromosome in male pregnancies) with the amount of any autosomal locus common in both the mother and fetus using quantitative real-time PCR (e.g., Lo YM, et al. (1998) Am J Hum Genet 62:768–775.)).

[0105] Partial-specific parameters (e.g., those of the test sample) can be weighted or adjusted by one or more weighting factors (e.g., weighting factors derived from the training group). For example, weighting factors can be derived for a part based on the relationship between partial-specific parameters and fetal scores determined for a training group of multiple samples. The partial-specific parameters of the test sample are then adjusted and / or weighted according to the weighting factors derived from the training group. In some embodiments, the partial-specific parameters for which the weighting factors are derived are the same as the adjusted or weighted (e.g., those of the test sample) partial-specific parameters (e.g., both are FLR). In some embodiments, the partial-specific parameters for which the weighting factors are derived are different from the adjusted or weighted (e.g., those of the test sample) partial-specific parameters. For example, the weighting factors can be determined by the correlation between coverage (i.e., the partial-specific parameter) and fetal scores for the training group of the sample, while the FLR of a portion of the test sample (i.e., another partial-specific parameter) can be adjusted according to a weighting factor derived from coverage. Without any theoretical constraints, (e.g., for test samples) part-specific parameters can sometimes be adjusted and / or weighted by weighting factors derived from different (e.g., training groups) part-specific parameters based on the correlation and / or association between each part-specific parameter and common part-specific FLRs.

[0106] A partial-specific fetal score estimate for a sample (e.g., a test sample) can be determined by weighting the partial-specific parameters with a weighting factor determined by that portion. Weighting may include adjusting, transforming, and / or altering the partial-specific parameters according to the weighting factor by applying any suitable mathematical operation. Non-limiting examples of such operations include multiplication, division, addition, subtraction, integration, symbolic arithmetic, algebraic computation, algorithms, trigonometric or geometric functions, transformations (such as Fourier transforms), etc., or combinations thereof. Weighting may include adjusting, transforming, and / or altering the partial-specific parameters according to a suitable mathematical model based on the weighting factor.

[0107] In some embodiments, the fetal score of a sample is determined based on one or more partial-specific fetal score estimates. In some embodiments, the fetal score of a sample (e.g., a test sample) is determined (e.g., evaluated) based on weighted or adjusted partial-specific parameters of one or more portions. In some embodiments, the fetal nucleic acid score of a test sample is evaluated based on adjusted counts or adjusted count subgroups. In some embodiments, the fetal nucleic acid score of a test sample is evaluated based on adjusted FLR, adjusted FRS, adjusted coverage, and / or adjusted mappability of portions. In some embodiments, weighted or adjusted partial-specific parameters are used, such as about 1 to about 500,000, about 100 to about 300,000, about 500 to about 200,000, about 1,000 to about 200,000, about 1,500 to about 200,000, or about 1,500 to about 50,000.

[0108] Determining the fetal score (e.g., of a test sample) can be done by any suitable method based on multiple part-specific fetal score estimates (e.g., of the same test sample). In some embodiments, methods for improving the accuracy of assessing the fetal nucleic acid score in a test sample from a pregnant woman include determining one or more part-specific fetal score estimates, wherein the assessment of the fetal score of the sample is determined based on the one or more part-specific fetal score estimates. In some embodiments, assessing or determining the fetal nucleic acid score of a sample (e.g., a test sample) includes summing one or more part-specific fetal score estimates. Summation may include determining a mean, arithmetic mean, median, AUC, or integral value based on multiple part-specific fetal score estimates.

[0109] In some embodiments, methods for improving the accuracy of assessing the fraction of fetal nucleic acids in test samples from pregnant women include obtaining counts of sequence readings mapped to a portion of a reference genome, said sequence readings being readings of circulating cell-free nucleic acids from the test sample of the pregnant woman, wherein at least a subgroup of the obtained counts originates from a region of said genome that is advantageous in obtaining a larger number of fetal nucleic acid counts relative to the total count of other regions of the genome than the total count of that region. In some embodiments, the estimate of the fraction of fetal nucleic acids is determined based on said subgroup of the portion, said subgroup of the portion being selected based on a portion mapped to a number of fetal nucleic acid counts that are larger than the fetal nucleic acid counts of other portions. In some embodiments, said subgroup of the portion is selected based on a portion mapped to a number of fetal nucleic acid counts relative to non-fetal nucleic acids that are larger than the fetal nucleic acid counts of other portions. Counts mapped to all portions or subgroups of portions may be weighted to provide weighted counts. Weighted counts can be used to assess fetal nucleic acid scores, and the counts can be weighted according to a portion of the fetal nucleic acid count mapped to a number of fetal nucleic acid counts that are larger than other portions of the fetal nucleic acid counts. In some embodiments, the counts are weighted according to a portion of the fetal nucleic acid count mapped to a number of relative non-fetal nucleic acid counts that are larger than other portions of the relative non-fetal nucleic acid counts.

[0110] The fetal fraction of a sample (such as a test sample) can be determined based on multiple part-specific fetal fraction estimates, wherein the part-specific estimates are derived from any suitable region or segment of the genome. The part-specific fetal fraction estimate can be determined for one or more portions of a suitable chromosome (e.g., one or more selected chromosomes, one or more autosomes, sex chromosomes (such as ChrX and / or ChrY), aneuploid chromosomes, euploid chromosomes, etc., or combinations thereof).

[0111] In some implementations, determining the fetal fraction includes

[0112] (a) Obtaining a count of sequence reads mapped to a portion of the reference genome, wherein the sequence reads are readings of circulating cell-free nucleic acids from test samples from pregnant women;

[0113] (b) Using a microprocessor, by independently associating weighting factors for each part, (i) the counts of sequence reads mapped to each part or (ii) other part-specific parameters weighted to the part-specific score of fetal nucleic acid, thereby providing an estimate of the part-specific fetal score based on said weighting factors, wherein each weighting factor has been determined from the fit correlation between (i) the fetal nucleic acid score of each of the multiple samples and (ii) the counts of sequence reads mapped to each part (or other part-specific parameters) of the multiple samples; and

[0114] (c) Assess the fetal nucleic acid score of the test sample based on the part-specific fetal score estimate.

[0115] The amount of fetal nucleic acid in extracellular nucleic acids can be quantified and can be used in conjunction with the methods described herein. Therefore, in some embodiments, the methods of the techniques described herein include an additional step of determining the amount of fetal nucleic acid. The amount of fetal nucleic acid in the nucleic acid sample of the subject can be determined before or after processing to prepare the sample nucleic acid. In some embodiments, the amount of fetal nucleic acid in the sample is determined after the sample nucleic acid has been processed and prepared, and used for further evaluation. In some embodiments, the results include decomposing the fetal nucleic acid fraction in the sample nucleic acid into factors (e.g., adjusting the count, removing the sample, making a decision, or not making a decision).

[0116] The determination step can be performed before, during, or at any point in time within the methods described herein, or after certain methods described herein (e.g., aneuploidy detection, microduplication or microdeletion detection, fetal sex determination). For example, to achieve a fetal sex or aneuploidy, microduplication or microdeletion detection method with a given sensitivity or specificity, fetal nucleic acid quantification methods can be performed before, during, or after fetal sex or aneuploidy, microduplication or microdeletion determination to identify samples containing more than about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25% or more fetal nucleic acid. In some embodiments, samples determined to have a certain fetal nucleic acid threshold (e.g., about 15% or more fetal nucleic acid; or about 4% or more fetal nucleic acid) are further used to analyze, for example, fetal sex or aneuploidy, microduplication or microdeletion, or the presence of aneuploidy or genetic variation. In some embodiments, only samples having a certain fetal nucleic acid threshold (e.g., about 15% or more fetal nucleic acid; or about 4% or more fetal nucleic acid) are selected (e.g., selected and informed to the patient) to determine, for example, fetal sex or the presence of aneuploidy, microduplication or microdeletion.

[0117] In some embodiments, determining the fetal fraction or the amount of fetal nucleic acid is not necessary for identifying the presence of chromosomal aneuploidy, microduplication, or microdeletion. In some embodiments, identifying the presence of chromosomal aneuploidy, microduplication, or microdeletion does not require sequence differentiation between fetal and maternal DNA. In some embodiments, this is because the additive contribution of maternal and fetal sequences to specific chromosomes, chromosomal portions, or segments is analyzed. In some embodiments, identifying the presence of chromosomal aneuploidy, microduplication, or microdeletion does not rely on prior sequence information that distinguishes fetal DNA from maternal DNA.

[0118] Enriched nucleic acids

[0119] In some embodiments, nucleic acids (e.g., extracellular nucleic acids) are enriched or relatively enriched against nucleic acid subsets or substances. Nucleic acid subsets may include, for example, fetal nucleic acids, maternal nucleic acids, nucleic acids containing fragments of a specific length or length range, or nucleic acids derived from specific genomic regions (e.g., a single chromosome, a set of chromosomes, and / or certain chromosomal regions). Such enriched samples may be used in conjunction with the methods described herein. Therefore, in some embodiments, the method of this technique includes the additional step of enriching nucleic acid subsets, such as fetal nucleic acids, in a sample. In some embodiments, the methods described above for determining fetal fractions can also be used to enrich fetal nucleic acids. In some embodiments, maternal nucleic acids are selectively removed (partially, substantially, almost completely, or completely) from the sample. In some embodiments, enriching specific low-copy-number nucleic acids (e.g., fetal nucleic acids) can improve quantitative sensitivity. Methods for enriching specific types of nucleic acids in samples, such as those described below, are incorporated herein by reference in U.S. Patent No. 6,927,028, International Application Publication No. WO2007 / 140417, International Application Publication No. WO2007 / 147063, International Application Publication No. WO2009 / 032779, International Application Publication No. WO2009 / 032781, International Application Publication No. WO2010 / 033639, International Application Publication No. WO2011 / 034631, International Application Publication No. WO2006 / 056480, and International Application Publication No. WO2011 / 143659.

[0120] In some embodiments, nucleic acids are enriched for certain target fragment types and / or reference fragment types. In some embodiments, nucleic acid enrichment is performed on specific nucleic acid fragment lengths or fragment lengths or ranges using one or more length-based separation methods described herein and / or known in the art. In some embodiments, nucleic acid enrichment is performed on fragments selected from genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein and / or known in the art. Methods for enriching certain nucleic acid subsets (e.g., fetal nucleic acids) in samples are detailed below.

[0121] Methods for enriching nucleic acid subpopulations (e.g., fetal nucleic acids) that can be used in conjunction with the methods of the present invention include methods that utilize the epigenetic differences between maternal and fetal nucleic acids. For example, fetal and maternal nucleic acids can be distinguished and separated based on methylation differences. A method for enriching fetal nucleic acids based on methylation is described in U.S. Patent Application Publication 2010 / 0105049, which is incorporated herein by reference. Such methods sometimes involve binding sample nucleic acids to methylation-specific binding agents (methyl CpG-binding proteins (MBD), methylation-specific antibodies, etc.) and separating bound and unbound nucleic acids based on different methylation states. Such methods may also include the use of methylation-sensitive restriction enzymes (e.g., HhaI and HpaII as described above), which selectively digest nucleic acids from maternal samples using enzymes that selectively and completely or substantially digest maternal nucleic acids to enrich at least one fetal nucleic acid region in the sample, thus enabling the enrichment of fetal nucleic acid regions in maternal samples.

[0122] Other methods for enriching nucleic acid subsets (e.g., fetal nucleic acids) that can be used in conjunction with the method of this invention include restriction endonuclease-enhanced polymorphic sequencing, such as the method described in U.S. Patent Application Publication 2009 / 0317818, which is incorporated herein by reference. This method includes cleaving the nucleic acid containing the non-target allele with a restriction endonuclease that recognizes the non-target allele but does not recognize the target allele; and amplifying the uncut nucleic acid but not the cleaved nucleic acid, wherein the uncut amplified nucleic acid represents a target nucleic acid (e.g., fetal nucleic acid) enriched relative to the non-target nucleic acid (e.g., maternal nucleic acid). In some embodiments, the nucleic acid may be selected such that it contains an allele having a polymorphic site that is readily digested, for example, by a cleavage agent.

[0123] Methods for enriching nucleic acid subsets (e.g., fetal nucleic acids) that can be used in conjunction with the methods of the present invention include selective enzymatic degradation. These methods involve protecting the target sequence from digestion by exonucleases, thereby facilitating the removal of unwanted sequences (e.g., maternal DNA) from the sample. For example, in one method, sample nucleic acids are denatured to produce single-stranded nucleic acids, which are then contacted with at least one target-specific primer pair under suitable annealing conditions. The annealed primers are extended using nucleotide polymerization to produce a double-stranded target sequence, and the single-stranded nucleic acids are digested with a nuclease that digests single-stranded (e.g., non-target) nucleic acids. In some embodiments, the method can be repeated at least one more cycle. In some embodiments, the same target-specific primer pair can be used to initiate the first and second cycles of extension, and in some embodiments, different target-specific primer pairs are used for the first and second cycles.

[0124] Methods for enriching nucleic acid subsets (e.g., fetal nucleic acids) that can be used in conjunction with the methods of this invention include massively parallel sequencing (MPSS). MPSS is typically a solid-phase method that uses adaptors (i.e., tags) to ligate nucleic acid sequences, which are then decoded and read in small increments. Tagged PCR products are typically amplified, resulting in PCR products with unique tags for each nucleic acid. Tags are typically used to conjugate PCR products to microbeads. For example, sequence signatures can be identified from each bead after several rounds of sequence determination based on ligation. Each signature sequence (MPSS tag) in the MPSS database is analyzed, all other signatures are compared, and all identical signatures are counted.

[0125] In some embodiments, certain enrichment methods (such as certain MPS-based and / or MPSS-based enrichment methods) may include amplification-based methods (such as PCR). In some embodiments, site-specific amplification methods (e.g., using site-specific amplification primers) may be used. In some embodiments, multiplex SNP allele PCR methods may be used. In some embodiments, multiplex SNP allele PCR methods may be used in conjunction with singlet sequencing. For example, this method may involve using multiplex PCR (MASSARRAY system) and incorporating the capture probe sequence into the amplicons, followed by sequencing using, for example, an Illumina MPSS system. In some embodiments, multiplex SNP allele PCR methods may be used in conjunction with a three-primer system and index sequencing. For example, this method may involve using multiplex PCR (MASSARRAY system) where primers incorporate a first capture probe into site-specific forward PCR primers and an adaptor sequence into site-specific reverse PCR primers to generate an amplicons, followed by secondary PCR incorporating the reverse capture sequence and molecular index barcode for sequencing using, for example, an Illumina MPSS system. In some embodiments, multiplex SNP allele PCR methods can be used in conjunction with a four-primer system and index sequencing. For example, the method may involve using multiplex PCR (MASSARRAY system) where primers incorporate adaptor sequences into site-specific forward and reverse PCR primers, followed by secondary PCR to incorporate forward and reverse capture sequences and molecular index barcodes for sequencing using, for example, an Illumina MPSS system. In some embodiments, microfluidic methods can be used. In some embodiments, array-based microfluidic methods can be used. For example, the method may involve using a microfluidic array (such as Fluidigm) for low-repetition amplification and incorporation of index and capture probes, followed by sequencing. In some embodiments, emulsion microfluidic methods, such as digital droplet PCR, can be used.

[0126] In some embodiments, universal amplification methods (e.g., using universal or non-site-specific amplification primers) may be used. In some embodiments, universal amplification methods may be combined with pull-down methods. In some embodiments, the method may include pulling down biotinylated ultramers from a universal amplification sequence library (e.g., biotinylation pull-down assays from Agilent or IDT). For example, the method may involve preparing a standard library, enriching selected regions by pull-down assays, and a secondary universal amplification step. In some embodiments, pull-down methods may be combined with ligation-based methods. In some embodiments, the method may include pulling down biotinylated ultramers ligated with sequence-specific adaptors (e.g., HALOPLEX PCR, HaloGenomics). For example, the method may involve using selector probes to capture restriction enzyme-digested fragments, then ligating the captured product and adaptor, and universal amplification followed by sequencing. In some embodiments, pull-down methods may be combined with extension and ligation-based methods. In some embodiments, the method may include molecular inverted probe (MIP) extension and ligation. For example, the method may involve using molecular inverted probes in combination with sequence adaptors, followed by universal amplification and sequencing. In some implementations, complementary DNA can be synthesized and sequenced without amplification.

[0127] In some implementations, extension and ligation methods can be performed without pulling down components. In some implementations, the method may include site-specific forward and reverse primer hybridization, extension, and ligation. The method may also include universal amplification or complementary DNA synthesis without amplification followed by sequencing. In some implementations, the method may reduce or exclude background sequences during analysis.

[0128] In some embodiments, pull-down assays may be used with or without an optional amplification component. In some embodiments, the method may include modified pull-down assays and ligations that fully incorporate the capture probe without requiring universal amplification. For example, the method may involve using a modified selector probe to capture a restriction enzyme-digested fragment, then ligating the capture product and adaptor, and optionally amplifying, and sequencing. In some embodiments, the method may include a biotinylated pull-down assay, and a combination of extension and ligation of the adaptor sequence with single-stranded circular ligation. For example, the method may involve using a selector probe to capture the region of interest (i.e., the target sequence), extending the probe, ligating the adaptor, single-stranded circular ligation, optionally amplifying, and sequencing. In some embodiments, analysis of the sequencing results may separate the target sequence from the background.

[0129] In some embodiments, nucleic acid enrichment is performed on fragments of selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein. Sequence-based separation is typically based on nucleotide sequences (e.g., target fragments and / or reference fragments) present in the fragment of interest in the sample but substantially absent (e.g., 5% or less) in other fragments or other segments. In some embodiments, sequence-based separation can generate isolated target fragments and / or isolated reference fragments. The isolated target fragments and / or isolated reference fragments are typically separated from the remaining fragments in the nucleic acid sample. In some embodiments, the isolated target fragments and isolated reference fragments may also be separated from each other (e.g., in separate laboratory compartments). In some embodiments, the isolated target fragments and isolated reference fragments may be separated together (e.g., in the same laboratory). In some embodiments, unbound fragments may be differentially removed, degraded, or digested.

[0130] In some implementations, selective nucleic acid capture methods are used to separate target fragments and / or reference fragments from a nucleic acid sample. Commercially available nucleic acid capture systems include, for example, the Nimblegen sequence capture system (Roche NimbleGen, Madison, WI); the Illumina BEADARRAY platform (Illumina, San Diego, CA); the Affymetrix GENECHIP platform (Affymetrix, Santa Clara, CA); the Agilent SureSelect target enrichment system (Agilent Technologies, Santa Clara, CA); and related platforms. The method typically involves the hybridization of a captured oligonucleotide with a segment or all of the nucleotide sequence of the target or reference fragment and may include the use of solid-phase (e.g., solid-phase arrays) and / or solution-based platforms. The captured oligonucleotide (sometimes referred to as a “bait”) may be selected or engineered so that it preferentially hybridizes to nucleic acid fragments of selected genomic regions or sites (e.g., one of chromosomes 21, 18, 13, X, or Y, or a reference chromosome). In some implementations, hybridization-based methods (e.g., using oligonucleotide arrays) can be used to enrich nucleic acid sequences from certain chromosomes (e.g., possible aneuploid chromosomes, reference chromosomes, or other chromosomes of interest) or regions of interest.

[0131] In some embodiments, one or more length-based separation methods are used to enrich nucleic acids for specific fragment lengths, length ranges, or lengths below or above a specific threshold or cutoff value. Fragment length typically refers to the number of nucleotides in a fragment. Fragment length sometimes also refers to fragment size. In some embodiments, length-based separation methods do not require measuring the length of individual fragments. In some embodiments, length-based separation methods are combined with methods for determining the length of individual fragments. In some embodiments, length-based separation refers to size fractionation, where all or part of the fractionated library can be separated (e.g., retained) and / or analyzed. Size fractionation is known in the art (e.g., array separation, molecular sieve separation, gel electrophoresis separation, column chromatography separation (e.g., size exclusion columns), and microfluidics-based methods). In some embodiments, length-based separation methods may include, for example, fragment cyclization, chemical treatment (e.g., formaldehyde, polyethylene glycol (PEG)), mass spectrometry, and / or size-specific nucleic acid amplification.

[0132] Certain length-based separation methods that can be used with the methods of this invention employ, for example, selective sequence tagging. The term "sequence tagging" refers to incorporating an identifiable, unique sequence into a nucleic acid or nucleic acid group. The term "sequence tagging" as used herein differs in meaning from the term "sequence tag" as described later herein. In this sequence tagging method, nucleic acids of various fragment sizes (e.g., short fragments) in a sample, including both long and short nucleic acids, are selectively sequence-tagged. This method typically involves a nucleic acid amplification reaction using a nested primer set, comprising internal and external primers. In some embodiments, one or both of the internal primers may be tagged to introduce a tag onto the target amplification product. External primers are typically not annealed to short fragments carrying the (internal) target sequence. Internal primers may anneal to short fragments and produce an amplification product carrying both the tag and the target sequence. Typically, tagging of long fragments is inhibited by combinatorial mechanisms, including, for example, the inhibition of internal primer extension caused by prior annealing and extension of the external primers. Enrichment of tagged fragments can be achieved by any of a variety of methods, including, for example, digestion of single-stranded nucleic acids with exonucleases and amplification of tagged fragments using amplification primers specific to at least one tag.

[0133] Other length-based separation methods that can be used with the method of this invention involve precipitation of nucleic acid samples with polyethylene glycol (PEG). Examples of methods include those described in International Patent Application Publications WO2007 / 140417 and WO2010 / 115016. These methods typically require contacting the nucleic acid sample with PEG in the presence of one or more monovalent salts under conditions sufficient to precipitate large nucleic acids in large quantities without precipitating small (e.g., less than 300 nucleotides) nucleic acids in large quantities.

[0134] Other size-based enrichment methods that can be used with the methods described herein involve cyclization via ligation, such as using cyclases. Short nucleic acid fragments are generally more efficiently cyclized than long fragments. Non-cyclized sequences can be separated from cyclized sequences, and the enriched short fragments can be used for further analysis.

[0135] Nucleic acid library

[0136] In some embodiments, a nucleic acid library is a variety of polynucleotide molecules (e.g., nucleic acid samples) prepared, assembled, and / or modified for a specific process, non-limiting examples of which include immobilization, enrichment, amplification, cloning, detection, and / or use for nucleic acid sequencing on a solid phase (e.g., a solid support, such as a flow cell, beads). In some embodiments, the nucleic acid library is prepared before or during the sequencing process. Nucleic acid libraries (e.g., sequencing libraries) can be prepared using suitable methods known in the art. Nucleic acid libraries can be prepared via targeted or non-targeted preparation processes.

[0137] In some embodiments, the nucleic acid library is modified to include chemical portions (e.g., functional groups) configured to immobilize nucleic acids to a solid support. In some embodiments, the nucleic acid library is modified to include biomolecules (e.g., functional groups) and / or binding pair members configured to immobilize the library to a solid support. Non-limiting examples include thyroxine-binding globulin, steroid-binding proteins, antibodies, antigens, haptens, enzymes, hemagglutinins, nucleic acids, inhibitors, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding proteins, receptors, carbohydrates, oligonucleotides, polynucleotides, complementary nucleic acid sequences, and combinations thereof. Examples of specific binding pairs include, but are not limited to: anti-biotin moieties and biotin moieties; antigenic epitopes and antibodies or their immunologically active fragments; antibodies and haptens; digoxigenin moieties and anti-digoxigenin antibodies; luciferin moieties and anti-luciferin antibodies; operons and inhibitors; nucleases and nucleosides; lectins and polysaccharides; steroids and steroid-binding proteins; active compounds and active compound receptors; hormones and hormone receptors; enzymes and substrates; immunoglobulins and protein A; oligonucleotides or polynucleotides and their corresponding complements; and combinations thereof.

[0138] In some embodiments, the nucleic acid library is modified to include one or more polynucleotides of known composition. Non-limiting examples include identifiers (e.g., tags, index tags), capture sequences, labeled adaptors, restriction enzyme sites, promoters, enhancers, origins of replication, stem-loops, complementary sequences (e.g., primer binding sites, annealing sites), suitable integration sites (e.g., transposons, viral integration sites), modified nucleotides, and combinations thereof. Polynucleotides of known sequences can be inserted at suitable positions, such as the 5′ end, 3′ end, or within the nucleic acid sequence. Polynucleotides of known sequences can be the same or different sequences. In some embodiments, polynucleotides of known sequences are configured to hybridize with one or more oligonucleotides immobilized on a surface (e.g., the surface of a flow cell). For example, a known 5′ sequence of a nucleic acid molecule can hybridize with a first set of oligonucleotides, while a known 3′ sequence can hybridize with a second set of oligonucleotides. In some embodiments, the nucleic acid library may include chromosome-specific tags, capture sequences, labels, and / or adaptors. In some embodiments, the nucleic acid library includes one or more detectable markers. In some embodiments, one or more detectable markers may be incorporated into the 5′ end, 3′ end, and / or any nucleotide position of the nucleic acid in the library. In some embodiments, the nucleic acid library includes hybridized oligonucleotides. In some embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, the nucleic acid library includes hybridized oligonucleotide probes before immobilization on a solid phase.

[0139] In some embodiments, the polynucleotide of a known sequence includes a universal sequence. A universal sequence is a specific nucleotide sequence integrated into two or more nucleic acid molecules or subgroups of two or more nucleic acid molecules, wherein the universal sequence is identical with respect to all the molecules or subgroups into which it is integrated. Universal sequences are typically designed to hybridize and / or amplify multiple different sequences using a single universal primer complementary to the universal sequence. In some embodiments, two (e.g., a pair) or more universal sequences and / or universal primers are used. Universal primers typically include a universal sequence. In some embodiments, an adaptor (e.g., a universal adaptor) includes a universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple nucleic acid substances or subgroups thereof.

[0140] In some embodiments of nucleic acid library preparation (e.g., in certain sequencing processes of the synthesis procedure), the size of the nucleic acids is selected and / or fragmented to a length of several hundred base pairs or less (e.g., in library generation preparation). In some embodiments, library preparation is not required (e.g., when using ccfDNA).

[0141] In some embodiments, ligation-based library preparation methods are used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego, CA). Ligation-based library preparation methods typically employ adaptors (e.g., methylated adaptors) designed to incorporate an index sequence at the initial ligation step and are generally used to prepare samples for single-read sequencing, paired-end sequencing, and multiplex sequencing. For example, sometimes nucleic acids (e.g., fragmented nucleic acids or ccfDNA) are end-repaired via fill-in reactions, endonuclease reactions, or combinations thereof. In some embodiments, the resulting blunt-end-repaired nucleic acid can subsequently be extended by a single nucleotide, which is complementary to a single nucleotide overhang at the 3' end of the adaptor / primer. Any nucleotide can be used for the extended / overhanging nucleotide. In some embodiments, nucleic acid library preparation includes linking adaptor oligonucleotides. Adaptor oligonucleotides are typically complementary to flow cell anchors and are sometimes used to immobilize the nucleic acid library to a solid support, such as the inner surface of a flow cell. In some embodiments, the adaptor oligonucleotide includes an identifier, one or more sequencing primer hybridization sites (e.g., sequences complementary to universal sequencing primers, single-end sequencing primers, paired-end sequencing primers, multiplex sequencing primers, etc.) or combinations thereof (e.g., adaptor / sequencing, adaptor / identifier, adaptor / identifier / sequencing).

[0142] The identifier may be a suitable detectable tag incorporating or conjugating a nucleic acid (e.g., a polynucleotide), which allows for the detection and / or identification of nucleic acids including the identifier. In some embodiments, the identifier is incorporating or conjugating the nucleic acid during sequencing methods (e.g., by polymerase). Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indexes or barcodes, radioactive tags (e.g., isotopes), metallic tags, chemiluminescent tags, phosphorescent tags, fluorescence quenchers, dyes, proteins (e.g., enzymes, antibodies or portions thereof, linkers, members of binding pairs), and combinations thereof. In some embodiments, the identifier (e.g., a nucleic acid index or barcode) is a unique, known, and / or identifiable sequence of a nucleotide or nucleotide analogue. In some embodiments, the identifier is six or more consecutive nucleotides. Many fluorophores with various different excitation and emission spectra are available. Any suitable type and / or number of fluorophores can be used as identifiers. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, or fifty or more different identifiers are used in the methods described herein (e.g., nucleic acid detection and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are linked to each nucleic acid in the library. Identifier detection and / or quantification can be performed by suitable methods or apparatus, non-limiting examples of which include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, illuminometer, fluorometer, spectrophotometer, suitable gene chip or microarray analysis, Western blotting, mass spectrometry, chromatography, cellular fluorescence analysis, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning flow cytometry, affinity chromatography, manual batch separation, electric field suspension, suitable nucleic acid sequencing methods and / or nucleic acid sequencing apparatus, and combinations thereof.

[0143] In some implementations, transposon-based library preparation methods are used (e.g., EPICENTRE NEXTERA, Epicentre, Madison WI). Transposon-based methods typically use in vitro translocation to similar fragments or tagged DNA (often allowing for the inclusion of platform-specific tags and optional barcodes) in a single-tube reaction to prepare a sequencer-ready library.

[0144] In some embodiments, a nucleic acid library or a portion thereof is amplified (e.g., by PCR-based methods). In some embodiments, sequencing methods include amplifying a nucleic acid library. The nucleic acid library may be amplified before or after immobilization onto a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification includes the process of amplifying or increasing (e.g., in the nucleic acid library) the amount of a present nucleic acid template and / or its complement, said process being achieved by generating one or more copies of the template and / or its complement. Amplification may be performed by suitable methods. The nucleic acid library may be amplified by thermal cycling or by isothermal amplification. In some embodiments, rolling circle amplification is used. In some embodiments, amplification occurs on a solid support (e.g., within a flow cell) where a nucleic acid library or a portion thereof is immobilized. In some sequencing methods, the nucleic acid library is added to a flow cell and immobilized by hybridization with an anchor under suitable conditions. Such nucleic acid amplification is generally referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or part of the amplification product is synthesized by extension starting from immobilized primers. Solid-phase amplification reactions are similar to standard solution-phase amplification, except that at least one of the amplified oligonucleotides (e.g., primers) is immobilized on a solid support.

[0145] In some embodiments, solid-phase amplification includes a nucleic acid amplification reaction comprising only one oligonucleotide primer immobilized on a surface. In some embodiments, solid-phase amplification includes multiple different immobilized oligonucleotide primer materials. In some embodiments, solid-phase amplification may include a nucleic acid amplification reaction comprising one oligonucleotide primer immobilized on a solid surface and a second different oligonucleotide primer in solution. Multiple different immobilized or solution primers may be used. Non-limiting examples of solid-phase nucleic acid amplification reactions include interfacial amplification, bridging amplification, emulsion PCR, WildFire amplification (e.g., U.S. Patent Application US20130012399), and combinations thereof.

[0146] sequencing

[0147] In some implementations, nucleic acids (e.g., nucleic acid fragments, sample nucleic acids, cell-free nucleic acids) are sequenced. In some implementations, a full sequence or nearly full sequence is obtained, and sometimes a partial sequence is obtained.

[0148] In some embodiments, some or all nucleic acids in the sample are enriched and / or amplified before or during sequencing (e.g., non-specific, such as by PCR-based methods). In some embodiments, specific portions or subgroups of nucleic acids in the sample are enriched and / or amplified before or during sequencing. In some embodiments, a portion or subset of a preselected set of nucleic acids is randomly sequenced. In some embodiments, nucleic acids in the sample are not enriched and / or amplified before or during sequencing.

[0149] As used herein, a “reading” (i.e., “a reading” or “sequence reading”) is a short nucleotide sequence generated by any sequencing method described herein or known in the art. Readings can be generated from one end of a nucleic acid fragment (“single-end reading”), while sometimes they are generated from both ends of a nucleic acid fragment (e.g., paired-end reading, double-end reading).

[0150] The length of sequence reads is typically related to the specific sequencing technology. For example, high-throughput methods provide sequence reads ranging in size from tens to hundreds of base pairs (bp). Nanopore sequencing, for instance, provides sequence reads ranging in size from tens to hundreds to thousands of base pairs. In some embodiments, sequence reads are arithmetic mean, median, average, or absolute lengths of approximately 15 bp to approximately 900 bp. In some embodiments, the sequence reads are arithmetic mean, median, average, or absolute lengths of approximately 1000 bp or longer.

[0151] In some embodiments, the nominal, average, arithmetic mean, or absolute length of a single-end reading is sometimes about 15 consecutive nucleotides to about 50 or more consecutive nucleotides, sometimes about 15 consecutive nucleotides to about 40 or more consecutive nucleotides, and sometimes about 15 consecutive nucleotides or about 36 or more consecutive nucleotides. In some embodiments, the nominal, average, arithmetic mean, or absolute length of a single-end reading is about 20 to about 30 bases, or about 24 to about 28 bases. In some embodiments, the nominal, average, arithmetic mean, or absolute length of a single-end reading is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 28, or about 29 bases.

[0152] In some implementations, the nominal, average, arithmetic mean, or absolute length of the paired end readings is sometimes about 10 consecutive nucleotides to about 25 consecutive nucleotides or more (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides or more), about 15 consecutive nucleotides to about 20 consecutive nucleotides or more, and sometimes about 17 consecutive nucleotides or about 18 consecutive nucleotides.

[0153] Readings are typically representations of nucleotide sequences in physiological nucleic acids. For example, sequences are described in readings using ATGC, where "A" represents adenine nucleotide, "T" represents thymine nucleotide, "G" represents guanine nucleotide, and "C" represents cytosine nucleotide. Sequence readings obtained from the blood of pregnant women can be readings of a mixture of fetal and maternal nucleic acids. Mixtures of relatively short readings can be transformed into representations of genomic nucleic acids in pregnant women and / or fetuses using the methods described herein. Mixtures of relatively short readings can be transformed into representations of, for example, copy number variations (e.g., maternal and / or fetal copy number variations), genetic variations or aneuploidy, microduplications, or microdeletions. Readings of mixtures of maternal and fetal nucleic acids can be transformed into representations of complex chromosomes or segments thereof containing features of one or both maternal and fetal chromosomes. In some embodiments, “obtaining” nucleic acid sequence readings from a subject sample and / or from biological samples of one or more reference individuals can directly involve sequencing nucleic acids to obtain sequence information. In some embodiments, “obtaining” can involve receiving sequence information directly obtained from other nucleic acids.

[0154] In some implementations, the representational components of the genome are sequenced and are sometimes referred to as “coverage” or “fold coverage.” For example, 1-fold coverage indicates that approximately 100% of the nucleotide sequence of the genome is represented by reads. In some implementations, “fold coverage” is a related term used with reference to a previous sequencing run. For example, a second round of sequencing may have 2-fold less coverage than the first round. In some implementations, the genome is sequenced with redundancy, wherein a given region of the genome is covered by two or more reads or overlapping reads (e.g., greater than 1 “fold coverage,” such as 2-fold coverage).

[0155] In some embodiments, a nucleic acid sample from a single individual is sequenced. In some embodiments, the nucleic acids of each of two or more samples are sequenced, wherein the samples are from one individual or from different individuals. In some embodiments, nucleic acid samples from two or more biological samples (where each biological sample is from one or more individuals) are collected, and the collection is sequenced. In later embodiments, the nucleic acid samples from each biological sample are often identified by one or more unique identifiers.

[0156] In some implementations, the sequencing method employs unique identifiers that allow for multiple sequence reactions during the sequencing process. The greater the number of unique identifiers, the more samples and / or chromosomes can be detected; for example, multiplexing can be performed during the sequencing process. The sequencing process can be performed using any suitable number of unique identifiers (e.g., 4, 8, 12, 24, 48, 96, or more).

[0157] Sequencing processes sometimes use a solid phase, which may include a flow cell on which nucleic acids from a library can be bound and on which reagents can flow and contact the bound nucleic acids. Flow cells sometimes include flow cell channels, and the use of identifiers facilitates the analysis of the number of samples in each channel. A flow cell is any solid support that can be constructed to retain and / or allow reagent solutions to pass through in an orderly manner with bound analytes. Flow cells are typically planar, optically transparent, usually in the millimeter or sub-millimeter range, and often have channels or pathways in which analyte / reagent interactions occur. In some implementations, the number of samples that can be analyzed in a given flow cell channel often depends on the number of unique identifiers used in library preparation and / or probe design. Single flow cell channel. Multiplexing of 12 identifiers, for example, may allow simultaneous analysis of 96 samples (e.g., the number of wells in a 96-well microplate) in an 8-channel flow cell. Similarly, multiplexing of 48 identifiers, for example, may allow simultaneous analysis of 384 samples (e.g., the number of wells in a 384-well microplate) in an 8-channel flow cell. Non-limiting examples of commercially available multiplex sequencing kits include Illumina’s multiplex sample preparation oligonucleotide kit and multiplex sequencing primer and PhiX control kit (e.g., Illumina catalog numbers PE-400-1001 and PE-400-1002, respectively).

[0158] Any suitable method for sequencing nucleic acids can be used, non-limiting examples of which include Maxim & Gilbert chain termination methods, sequencing by synthesis, ligation sequencing, mass spectrometry sequencing, microscopy-based techniques, and combinations thereof. In some embodiments, first-generation sequencing technologies such as Sanger sequencing methods, including automated Sanger sequencing methods (including microfluidic Sanger sequencing), can be used in the methods of the present invention. In some embodiments, other sequencing technologies, including nucleic acid imaging techniques (such as transmission electron microscopy (TEM) and atomic force microscopy (AFM)), are also used herein. In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods typically involve clonal amplification of a DNA template or a single DNA molecule, sometimes sequenced in a flow cell in a massively parallel manner. Next-generation (e.g., second and third generation) sequencing technologies (capable of sequencing DNA in massively parallel manner) can be used in the methods described herein and are collectively referred to herein as “massively parallel sequencing” (MPS). In some embodiments, MPS sequencing methods employ a targeted approach, wherein a specific chromosome, gene, or region of interest is the sequence. In some embodiments, a non-targeted approach is used, wherein most or all nucleic acids in a sample are sequenced, amplified, and / or randomly captured.

[0159] In some embodiments, targeted enrichment, amplification, and / or sequencing methods are used. Targeting methods typically isolate, select, and / or enrich nucleic acid subgroups in a sample for further processing using sequence-specific oligonucleotides. In some embodiments, a library of sequence-specific oligonucleotides is employed to target (e.g., hybridize) one or more nucleic acid subgroups in a sample. Sequence-specific oligonucleotides and / or primers are typically selective for specific sequences (e.g., unique nucleic acid sequences) present in one or more regions of interest, such as chromosomes, genes, exons, introns, and / or regulatory regions. Any suitable method or combination of methods can be used to enrich, amplify, and / or sequence one or more target nucleic acid subgroups. In some embodiments, target sequences are isolated and / or enriched by capture to a solid phase (e.g., flow cell, bead) using one or more sequence-specific anchors. In some embodiments, target sequences are enriched and / or amplified using sequence-specific primers and / or primer sets via polymerase-based methods (e.g., PCR-based methods, via any suitable polymerase-based extension). Sequence-specific anchors are typically used as sequence-specific primers.

[0160] MPS sequencing sometimes employs sequencing via synthesis and certain imaging methods. Nucleic acid sequencing technologies that can be used in the methods described herein are synthetic sequencing and reversible terminator-based sequencing (such as Illumina's Genome Analyzer and Genome Analyzer II; HISEQ 2000; HISEQ2500 (Illumina, San Diego, CA)). This technology allows for the parallel sequencing of millions of nucleic acid (e.g., DNA) fragments. In one embodiment of this sequencing technology, a flow cell is used comprising an optically clear slide with eight individual channels, the surface of which is bound to oligonucleotide anchors (e.g., adaptor primers). The flow cell is typically constructed to retain and / or allow reagent solutions to pass through in an orderly manner, binding the analyte. Flow cells are typically planar, optically clear, typically in the millimeter or sub-millimeter range, and often contain channels or pathways in which analyte / reagent interactions occur.

[0161] In some embodiments, synthetic sequencing involves repeatedly adding (e.g., covalently adding) nucleotides to primers or a pre-existing nucleic acid chain in a template-guided manner. Each repeatedly added nucleotide is detected, and the process is repeated multiple times until a sequence of the nucleic acid chain is obtained. The length of the obtained sequence depends in part on the number of addition and detection steps performed. In some embodiments of synthetic sequencing, one, two, three, or more nucleotides of the same type (e.g., A, G, C, or T) are added and detected in a nucleotide addition round. Nucleotides can be added by any suitable method (e.g., enzymatic or chemical). For example, in some embodiments, polymerases or ligases add nucleotides to primers or a pre-existing nucleic acid chain in a template-guided manner. In some embodiments of synthetic sequencing, different types of nucleotides, nucleotide analogs, and / or identifiers are used. In some embodiments, reversible terminators and / or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In some embodiments, synthetic sequencing includes cleavage (e.g., cleavage and removal of identifiers) and / or washing steps. In some embodiments, the addition of one or more nucleotides is detected by methods described herein or known in the art. Non-limiting examples include any suitable imaging device, a suitable camera, a digital camera, a CCD (charge-coupled device) based imaging device (e.g., a CCD camera), a CMOS (complementary metal-oxide-semiconductor) based imaging device (e.g., a CMOS camera), a photodiode (e.g., a photomultiplier tube), an electron microscope, a field-effect transistor (e.g., a DNA field-effect transistor), an ISFET ion sensor (e.g., a CHEMFET sensor), and combinations thereof. Other sequencing methods that can be used to perform the methods described herein include digital PCR and hybridization sequencing.

[0162] Other sequencing methods that can be used to perform the methods described herein include digital PCR and hybridization sequencing. Digital polymerase chain reaction (digital PCR or dPCR) can be used to directly identify and quantify nucleic acids in a sample. In some embodiments, digital PCR can be performed in an emulsion. For example, individual nucleic acids are isolated in, for example, a microfluidic device and each nucleic acid is amplified individually by PCR. Nucleic acids are isolated such that no more than one nucleic acid is contained in each well. In some embodiments, different probes can be used to distinguish multiple alleles (e.g., fetal alleles and maternal alleles). Alleles can be counted to determine copy number.

[0163] In some embodiments, hybridization sequencing can be used. The method involves contacting multiple polynucleotide sequences with multiple polynucleotide probes, each of which is optionally attached to a substrate. In some embodiments, the substrate may be a plane with an array of known nucleotide sequences. The pattern of hybridization with the array can be used to determine the polynucleotide sequences present in the sample. In some embodiments, each probe is attached to a bead (such as a magnetic bead). Hybridization with the bead can be identified and used to identify multiple polynucleotide sequences in the sample.

[0164] In some implementations, nanopore sequencing can be used in the methods described herein. Nanopore sequencing is a single-molecule sequencing technology whereby a single nucleic acid molecule (such as DNA) is directly sequenced as it passes through a nanopore.

[0165] The methods described herein can be used to obtain nucleic acid sequencing reads by employing suitable non-human MPS methods, systems, or technology platforms. Non-limiting examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, Helicos True single-molecule sequencing, particle-to-electricity (Ion Torrent) and Ion semiconductor-based sequencing (e.g., developed by Life Technologies), technologies based on WildFire, 5500, 5500xl W and / or 5500xl W genetic analyzers (e.g., US Patent Application US20130012399, developed and marketed by Life Technologies); Polony sequencing, Pyro sequencing, massively parallel signature sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen systems and methods, nanopore-based platforms, chemically sensitive field-effect transistor (CHEMFET) arrays, electron microscopy-based sequencing (e.g., developed by ZS Genetics and Halcyon Molecular), and nanosphere sequencing.

[0166] In some embodiments, chromosome-specific sequencing is performed. In some embodiments, chromosome-specific sequencing is performed using DANSR (Digital Analysis of Selected Regions). Digital analysis of selected regions can simultaneously quantify hundreds of sites by using interfering 'bridging' oligonucleotides to form PCR templates through cfDNA-dependent linkage of two site-specific oligonucleotides. In some embodiments, chromosome-specific sequencing is performed by generating a library enriched with chromosome-specific sequences. In some embodiments, sequence reads are obtained only for selected chromosome sets. In some embodiments, sequence reads are obtained only for chromosomes 21, 18, and 13.

[0167] Mapped readings

[0168] The number of sequence reads that can be mapped to a specific nucleic acid region (e.g., a chromosome, a part, or a segment thereof) is called a count. Any suitable mapping method (e.g., procedure, algorithm, program, software, module, etc., or a combination thereof) can be used. Some aspects of mapping methods are described below.

[0169] Mapped nucleotide sequence reads (i.e., sequence information of fragments at unknown physical genomic sites) can be performed in various ways, typically involving aligning the obtained sequencing reads with matching sequences in a reference genome. In this alignment, the sequence reads are usually compared to a reference sequence; those that are aligned are referred to as "mapped," "mapped sequence reads," or simply "mapped readings." In some embodiments, mapped sequence reads are referred to as "hit" or "count." In some embodiments, mapped sequence reads are aggregated and assigned to specific portions based on various parameters, as detailed below.

[0170] As used herein, the terms "alignment" and "alignment" refer to two or more nucleic acid sequences that can be identified as a match (e.g., 100% identity) or a partial match. Alignment can be performed manually or by a computer (e.g., software, program, module, or algorithm), and non-limiting examples include the Nucleotide Data Effective Local Alignment (ELAND) computer program, which is part of the Illumina genome analysis workflow. Alignment of sequence readings can be 100% sequence match. In some cases, alignment is less than 100% sequence match (i.e., imperfect match, partial match, partial alignment). In some embodiments, alignment is approximately 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% match. In some embodiments, alignment includes mismatches. In some implementations, the alignment includes 1, 2, 3, 4, or 5 mismatches. Two or more sequences can be aligned using either strand. In some implementations, the nucleic acid sequence is aligned with the reverse complementary strand of another nucleic acid sequence.

[0171] Various computer methods can be used to map sequence reads to portions. Non-limiting examples of computer algorithms that can be used for sequence alignment include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE 1, BOWTIE 2, ELAND, MAQ, PROBEMATCH, SOAP, or SEQMAP, or variations thereof, or combinations thereof. In some embodiments, sequence reads can be aligned to sequences in a reference genome. In some embodiments, sequence reads can be obtained from and / or aligned to sequences in nucleic acid databases known in the art, including, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (Japan DNA Database). BLAST or similar tools can be used to search for identical sequences against a sequence database. Search hits can then, for example, be used to sort identical sequences into appropriate portions (as described below).

[0172] In some embodiments, the mapped sequence reads and / or information associated with the mapped sequence reads are stored and / or evaluated on a non-transitory computer-readable medium in a suitable computer-readable form. "Computer-readable form" here sometimes refers to a format. In some embodiments, the mapped sequence reads are stored and / or evaluated in a suitable binary format, text format, etc., or a combination thereof. The binary format is sometimes BAM format. The text format is sometimes Sequence Alignment / Mapping (SAM) format. Non-limiting examples of binary or text formats include BAM, SAM, SRF, FASTQ, Gzip, etc., or combinations thereof. In some embodiments, the mapped sequence reads are stored in and / or converted to a format that requires less storage space (e.g., fewer bytes) than a conventional format (e.g., SAM or BAM format). In some embodiments, the mapped sequence reads in a first format are compressed to a second format that requires less storage space than the first. The term "compression" as used herein refers to the process of data compression, source encoding, and / or bitrate reduction, wherein the size of the computer-readable data file is reduced. In some implementations, mapped sequence reads are compressed from SAM format to binary format. Sometimes some data is lost during file compression. Sometimes the compression process does not result in data loss. In some file compression implementations, some data is replaced with an index and / or reference to another data file containing information relating to the mapped sequence reads. In some implementations, the mapped sequence reads are stored in binary format, including or consisting of: read counts, chromosome identifiers (e.g., identifying the chromosome mapped by the reads), and chromosome location identifiers (e.g., identifying the portion of the chromosome mapped by the reads). In some implementations, the binary format includes 20-byte arrays, 16-byte arrays, 8-byte arrays, 4-byte arrays, or 2-byte arrays. In some implementations, the mapped read information is stored in arrays in 10-byte, 9-byte, 8-byte, 7-byte, 6-byte, 5-byte, 4-byte, 3-byte, or 2-byte formats. Sometimes the mapped data reads are stored in 4-byte arrays, including 5-byte formats. In some implementations, the binary format includes a 5-byte format, including a 1-byte chromosome ordinal number and a 4-byte chromosome portion. In some implementations, the mapped readings are stored in a compressed binary format that is approximately 100 times, 90 times, 80 times, 70 times, 60 times, 55 times, 50 times, 45 times, 40 times, or 30 times smaller than the Sequence Alignment / Mapping (SAM) format. In some implementations, the mapped readings are stored in a compressed binary format that is approximately 2 to 50 times smaller than the GZip format (e.g., approximately 30, 25, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, or approximately 5 times smaller).

[0173] In some implementations, the system includes a compression module (e.g., 4, Figure 42A In some embodiments, the sequence readings of the mapped data, stored in a computer-readable format on a non-transitory computer-readable medium, are compressed by a compression module. The compression module may sometimes transform the sequence readings of the mapped data to a suitable format or from a suitable format. In some embodiments, the compression module may accept sequence readings of the mapped data in a first format (e.g., 1, Figure 42A The module converts the data into a compressed format (e.g., binary format, 5) and transfers the compressed readings to another module (e.g., offset density module 6). The compression module typically provides sequence readings in binary format 5 (e.g., BReads format). Non-limiting examples of compression modules include GZIP, BGZF, and BAM, or variations thereof.

[0174] The following provides an example of converting an integer into a 4-byte array using Java:

[0175] public static final byte[]

[0176] convertToByteArray(int value)

[0177] {

[0178] return new byte[]{

[0179] (byte)(value>>>24),

[0180] (byte)(value>>>16),

[0181] (byte)(value>>>8),

[0182] (byte)value};

[0183] }

[0184] In some implementations, reads may be uniquely or non-uniquely mapped to portions of a reference genome. A read that aligns to a single sequence in the reference genome is called a “unique mapping.” A read that aligns to two or more sequences in the reference genome is called a “non-unique mapping.” In some implementations, non-uniquely mapped reads are removed from further analysis (e.g., quantification). In some implementations, a small degree of mismatch (0-1) may indicate a possible single nucleic acid polymorphism between the reference genome and the mapped reads from an individual sample. In some implementations, no mismatches may allow reads to be mapped to a reference sequence.

[0185] As used herein, the term "reference genome" can refer to a genome of any organism or virus in which any part or all of it is specifically known, sequenced, or characterized, and can be used as a reference for identifying target sequences. For example, reference genomes used for human subjects and many other organisms are available from the National Center for Biotechnology Information, ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus expressed in a nucleic acid sequence. Reference sequences or reference genomes used herein are often assembled or partially assembled genome sequences from one or more individuals. In some embodiments, the reference genome is an assembled or partially assembled genome sequence from one or more human individuals. In some embodiments, the reference genome includes sequences assigned to chromosomes.

[0186] In some embodiments, when the sample nucleic acid is derived from a pregnant woman, the reference sequence may not originate from the fetus, the mother, or the father, and is thus referred to herein as an "external reference." In some embodiments, a maternal reference may be prepared and used. When preparing a reference from a pregnant woman based on an external reference ("maternal reference sequence"), readings of DNA from the pregnant woman that are substantially free of fetal DNA are typically mapped to and assembled from the external reference sequence. In some embodiments, the external reference originates from DNA from an individual substantially of the same ethnicity as the pregnant woman. The maternal reference sequence may not completely cover the maternal genomic DNA (e.g., it may cover approximately 50%, 60%, 70%, 80%, 90%, or more of the maternal genomic DNA), and the maternal reference may not perfectly match the maternal genomic DNA sequence (e.g., the maternal reference sequence may contain multiple mismatches).

[0187] In some implementations, mappability is evaluated for genomic regions (e.g., portions, genomic parts, or sub-regions). Mappability is the ability of nucleotide sequence reads to clearly align to a portion of a reference genome, typically with multiple to a specific number of mismatches, including, for example, 0, 1, 2, or more mismatches. For a given genomic region, the expected mappability can be calculated using a sliding window method with a predetermined read length and averaged to obtain a mappability value at the read level. Extended genomic regions that include unique nucleotide sequences sometimes have high mappability values.

[0188] Part

[0189] In some implementations, mapped sequence reads (i.e., sequence tags) are grouped together according to various parameters and assigned to specific segments (e.g., segments of a reference genome). Typically, individually mapped sequence reads can be used to identify segments present in a sample (e.g., the presence, deletion, or abundance of a segment). In some implementations, the abundance of a segment is an indicator of the abundance of large sequences (e.g., chromosomes) in the sample. The term "segment" may also refer to a "genomic segment," "box," "region," "segmentation," "segment of a reference genome," "segment of a chromosome," or "genome segment." In some implementations, a segment is an entire chromosome, a chromosomal segment, a reference genome segment, a segment spanning multiple chromosomes, multiple chromosomal segments, and / or combinations thereof. In some implementations, segments are predefined based on specific parameters. In some implementations, segments are arbitrarily defined based on genome partitioning (e.g., partitions based on size, GC content, contiguous regions, contiguous regions of arbitrarily defined sizes, etc.).

[0190] In some embodiments, a portion is defined based on one or more parameters, including, for example, sequence length or specific characteristics. Portions can be selected, screened, and / or removed from consideration using any suitable criteria known in the art or described herein. In some embodiments, a portion is based on a specific length of the genome sequence. In some embodiments, the method may include analyzing sequence reads from multiple mappings for multiple portions. Portions may have substantially the same length or portions may have different lengths. In some embodiments, portions are approximately the same length. In some embodiments, portions of different lengths are adjusted or weighted. In some embodiments, portions are about 10 kb to about 100 kb, about 20 kb to about 80 kb, about 30 kb to about 70 kb, about 40 kb to about 60 kb, and sometimes about 50 kb. In some embodiments, portions are about 10 kb to about 20 kb. Portions are not limited to continuously running sequences. Therefore, portions may consist of continuous and / or non-continuous sequences. Portions are not limited to a single chromosome. In some embodiments, a portion comprises all or part of a chromosome, or all or part of two or more chromosomes. In some embodiments, a portion may span one, two, or more complete chromosomes. Furthermore, a portion may span the connecting or disconnecting regions of multiple chromosomes.

[0191] In some embodiments, a portion may be a specific chromosomal region of the chromosome of interest, such as a chromosome used to assess genetic variation (e.g., aneuploidy of chromosomes 13, 18, and / or 21, or sex chromosomes). A portion may also be a pathogenic genome (e.g., bacteria, fungi, or viruses) or a fragment thereof. A portion may be a gene, gene fragment, regulatory sequence, intron, exon, etc.

[0192] In some implementations, the genome (e.g., the human genome) is divided into sections based on the information content of specific regions. In some implementations, genome partitioning may remove similar regions (e.g., identical or homologous regions or sequences) and retain only unique regions. Regions removed during partitioning may be within a single chromosome or may span multiple chromosomes. In some implementations, the partitioned genome is down-trimmed and optimized for rapid alignment, typically allowing focus on uniquely identifiable sequences.

[0193] In some implementations, the weights of similar regions can be reduced. The process of reducing some weights will be described in detail later.

[0194] In some implementations, the genome can be divided into extrachromosomal regions based on information generated within the classification context. For example, the information content can be quantified using p-values, measuring the significance of specific genomic locations in confirmed normal and abnormal subjects (e.g., euploid and triploid subjects, respectively). In some implementations, the genome can be divided into extrachromosomal regions based on any other criteria, such as speed / convenience of tag alignment, GC content (e.g., high or low GC content), uniformity of GC content, other measurements of sequence content (e.g., individual nucleotide fraction, pyrimidine or purine fraction, fraction of natural and non-natural nucleic acids, fraction of methylated nucleotides, and CpG content), methylation status, double melting temperature, compliance with sequencing or PCR, uncertainty in the allocation of individual portions to a reference genome, and / or targeted search for specific features.

[0195] A "segment" of a chromosome is usually a part of the chromosome, and usually a part of the chromosome that is different from the part. Chromosomal segments are sometimes located in different regions of the chromosome from the part, sometimes do not share polynucleotides with the part, and sometimes include polynucleotides in the part. Chromosomal segments usually contain more nucleotides than the part (e.g., segments sometimes include the part), and sometimes chromosomal segments contain fewer nucleotides than the part (e.g., segments sometimes are within the part).

[0196] count

[0197] In some implementations, sequence reads mapped or partitioned based on selected features or variables can be quantified to determine the number of reads mapped to one or more parts (e.g., a reference genome part). In some implementations, the quantification of sequence reads mapped to a part is called counting (e.g., a count). Typically, the count is associated with the part. In some implementations, the counts of two or more parts (e.g., groups of parts) are mathematically processed (e.g., averaging, summing, normalization, etc., or combinations thereof). In some implementations, the count is determined from some or all of the sequence reads mapped to (i.e., associated with) the part. In some implementations, the count is determined from a predefined subgroup of mapped sequence reads. The predefined subgroup of mapped sequence reads can be defined or selected using any suitable feature or variable. In some implementations, the predefined subgroup of mapped sequencing reads can contain 1–n sequence reads, where n represents the number equal to the sum of all sequence reads generated from the test or reference sample.

[0198] In some embodiments, the count is derived from sequence readings processed or manipulated by suitable methods, operations, or mathematical processes known in the art. The count can be determined by suitable methods, operations, or mathematical processes. In some embodiments, the count is derived from sequence readings associated with a portion, some or all of which have undergone weighting, removal, filtering, standardization, adjustment, averaging (to obtain a mean), addition or subtraction, or combinations thereof. In some embodiments, the count is derived from the original sequence readings and / or filtered sequence readings. In some embodiments, the count value is determined by mathematical processes. In some embodiments, the count value is the average, arithmetic mean, or sum of sequence readings mapped to a portion. Typically, the count is the arithmetic mean of multiple counts. In some embodiments, the count is associated with an uncertain value.

[0199] In some embodiments, the counts may be processed or transformed (e.g., standardization, combination, summation, screening, selection, averaging (to obtain a mean), etc., or combinations thereof). In some embodiments, the counts may be transformed to produce standardized counts. The counts may be processed (e.g., standardization) by methods known in the art and / or described herein (e.g., sample-by-sample standardization, GC content standardization, linear and nonlinear least squares regression, GC LOESS, LOWESS, PERUN, ChAI, RM, GCRM, cQn, and / or combinations thereof).

[0200] Counts (e.g., raw, screened, and / or standardized counts) can be processed and standardized to one or more levels. Levels and profiles are detailed below. In some embodiments, counts are processed and / or standardized to a reference level. The reference level is described below. Counts processed according to a level (e.g., processed counts) can be associated with an uncertainty (e.g., calculated variance, error, standard deviation, Z-score, p-value, arithmetic mean absolute deviation, etc.). In some embodiments, the uncertainty defines a range above and below a certain level. Deviation values ​​can replace uncertainty values; non-limiting examples of deviation measurements include standard deviation, mean absolute deviation, median absolute deviation, standardized scores (e.g., Z-score, Z-score, normalization score, standardized variable, etc.).

[0201] Counts are typically obtained from nucleic acid samples from pregnant women carrying a fetus. Counts mapped to one or more portions of the nucleic acid sequence are typically represented by counts from both the fetus and the mother (e.g., pregnant woman subjects). In some implementations, some counts mapped to portions are derived from the fetal genome, and some counts mapped to the same portions are derived from the maternal genome.

[0202] Data processing and standardization

[0203] The mapped sequence reads that have already been counted are referred to herein as raw data because the data represents unprocessed counts (such as raw counts). In some embodiments, the sequence read data in a dataset can be further processed (such as mathematical and / or statistical processing) and / or displayed to help provide results. In some embodiments, datasets (including large datasets) may benefit from preprocessing to aid further analysis. Preprocessing of datasets sometimes involves removing redundant and / or uninformative portions or portions of a reference genome (such as portions with uninformative data or portions of a reference genome, redundant mapped reads, portions with a median count of 0, sequences that occur too frequently or too infrequently). Without being theoretically limited, data processing and / or preprocessing can (i) remove noisy data, (ii) remove uninformative data, (iii) remove redundant data, (iv) reduce the complexity of large datasets, and / or (v) help transform the data from one form to one or more other forms. When used with data or datasets, the terms “preprocessing” and “processing” are collectively referred to herein as “processing.” In some embodiments, processing can make the data more readily available for further analysis, thereby generating results. In some implementations, one or more processing methods (e.g., standardization methods, partial screening, mapping, verification, etc. or combinations thereof) are performed by a processor, microprocessor, computer, device associated with memory and / or controlled by a microprocessor.

[0204] As used herein, the term "noisy data" refers to (a) data that shows significant differences between data points during analysis or plotting, (b) data with significant standard deviation (e.g., greater than 3 standard deviations), (c) data with significant mean standard error, and combinations thereof. Noisy data sometimes arises due to the quantity and / or quality of the starting material (e.g., nucleic acid samples), and it sometimes appears as part of the method for preparing or replicating DNA used to generate sequence reads. In some implementations, the noise originates from certain sequences that occur too frequently when prepared using PCR-based methods. The methods described herein can reduce or eliminate the baseline of noisy data, thereby reducing the impact of noisy data on the results provided.

[0205] The terms “informative data,” “informative portion of the reference genome,” and “informative portion” are used herein to refer to data whose values ​​differ significantly from a predetermined threshold or fall outside a predetermined cutoff range, or data derived therefrom. The term “threshold” refers to any number calculated using a qualifying dataset as a limitation for diagnosing genetic variations (e.g., copy number variation, aneuploidy, microduplication, microdeletion, chromosomal abnormalities, etc.). In some embodiments, a threshold exceeding the results obtained by the method of the present invention leads to a diagnosis of a genetic variation (e.g., trisomy 21) in the subject. In some embodiments, the threshold or range of values ​​is calculated using mathematical and / or statistical processing of sequence read data (e.g., from a reference and / or subject), while in other embodiments, the sequence read data processed to generate the range of thresholds or values ​​is sequence read data (e.g., from a reference and / or subject). In some embodiments, an uncertainty is determined. The uncertainty is typically a measure of variance or error and can be any suitable measure of variation or error. In some embodiments, the uncertainty is a standard deviation, standard error, calculated variance, p-value, or arithmetic mean absolute deviation (MAD). In some implementations, the uncertainty value can be calculated according to the formula in Example 4.

[0206] Any suitable procedure may be used to process the data set described herein. Non-limiting examples of methods suitable for processing the data set include filtering, standardization, weighting, monitoring peak height, monitoring peak area, monitoring peak margins, determining area ratios, mathematical processing of the data, statistical processing of the data, application of mathematical algorithms, analysis using fixed variables, analysis using optimized variables, plotting the data to identify patterns or trends for further processing, and combinations thereof. In some implementations, the data set is processed according to different characteristics (such as GC content, redundant location reads, centromere regions, telomere regions, etc., and combinations thereof) and / or variables (such as fetal sex, maternal age, maternal ploidy, fetal nucleotide baseline percentage, etc., and combinations thereof). In some implementations, processing the data set described herein can reduce the complexity and / or dimensionality of large and / or complex data sets. Non-limiting examples of complex data sets include sequence reads generated from one or more test subjects and multiple reference subjects of different ages and ethnic backgrounds. In some implementations, the data set can contain thousands to millions of sequence reads from each test subject and / or reference subject.

[0207] In some implementations, data processing can be performed in any number of steps. For example, in some implementations, data can be adjusted and / or processed using only a single processing method, while in other implementations, data can be processed using one or more, five or more, ten or more, or twenty or more processing steps (e.g., one or more processing steps, two or more processing steps, three or more processing steps, four or more processing steps, five or more processing steps, six or more processing steps, seven or more processing steps, eight or more processing steps, nine or more processing steps, ten or more processing steps, eleven or more processing steps, twelve or more processing steps, thirteen or more processing steps, fourteen or more processing steps, fifteen or more processing steps, sixteen or more processing steps, seventeen or more processing steps, eighteen or more processing steps, nineteen or more processing steps, or twenty or more processing steps). In some implementations, the processing step may be the same step repeated two or more times (e.g., filtering two or more times, standardizing two or more times), while in other implementations, the processing step may be two or more different processing steps performed simultaneously or sequentially (e.g., filtering, standardizing; standardizing, monitoring peak height and edges; filtering, standardizing, standardizing against a reference, statistical processing to determine p-values, etc.). In some implementations, any suitable number and / or combination of the same or different processing steps may be used to process sequence readout data to aid in providing results. In some implementations, processing the data set using the standards described herein can reduce the complexity and / or dimensionality of the data set.

[0208] In some implementations, one or more processing steps may include one or more filtering steps. As used herein, the term "filtering" refers to removing a portion or reference genome from consideration. The portion or reference genome to be removed can be selected according to any suitable criteria, including but not limited to redundant data (such as redundant or overlapping mapping reads), non-informative data (such as portions or reference genomes with a median count of 0), portions or reference genomes containing sequences that occur too frequently or too infrequently, noisy data, and combinations thereof. Filtering methods often involve removing one or more portions of the reference genome from consideration and subtracting the counts of the selected portions or portions of the reference genome to be removed from the counts or totals of the considered reference genome, chromosomes, or genomes. In some implementations, portions of the reference genome can be removed sequentially (e.g., one at a time to allow evaluation of the removal effect of each individual portion), while in other implementations, all portions marked for removal can be removed simultaneously. In some implementations, portions of the reference genome characterized by differences above or below a certain level are removed; these are sometimes referred to as filtering "noisy" portions of the reference genome. In some embodiments, the filtering process includes extracting data points from a dataset of average profile levels originating from a part, chromosome, or chromosomal segment through predetermined multiple profile variations, and in other embodiments, the filtering process includes removing data points from a dataset of average profile levels originating from a part, chromosome, or chromosomal segment through predetermined multiple profile differences. In some embodiments, the filtering process is used to reduce the number of candidate parts in a reference genome used to analyze the presence or absence of genetic variations. Reducing the number of candidate parts in a reference genome used to analyze the presence or absence of genetic variations (e.g., microdeletions, microduplications) typically reduces the complexity and / or dimensionality of the dataset and sometimes increases the speed of searching for and / or identifying genetic variations and / or genetic abnormalities by two or more orders of magnitude.

[0209] In some implementations, one or more processing steps may include one or more standardization steps. Standardization can be performed by suitable methods described herein or known in the art. In some implementations, standardization includes adjusting measured values ​​of different orders of magnitude to a theoretically common order of magnitude. In some implementations, standardization includes complex mathematical adjustments to introduce the probability distribution of the adjusted values ​​in the comparison. In some implementations, standardization includes comparing the distribution to a normal distribution. In some implementations, standardization includes mathematical adjustments that allow for comparison of corresponding standardized values ​​for different data sets in a manner that eliminates certain overall effects (e.g., errors and outliers). In some implementations, standardization includes scaling. Standardization sometimes includes partitioning one or more data sets by a predetermined scalar or formula. Non-limiting examples of standardization methods include component-by-component standardization, standardization by GC content, linear and nonlinear least squares regression, LOESS, GCLOESS, LOWESS (locally weighted regression scatter smoothing), PERUN, ChAI, repetition masking (RM), GC-standardization and repetition masking (GCRM), cQn and / or combinations thereof. In some implementations, the presence or absence of genetic variations (e.g., aneuploidy, microduplication, microdeletion) is determined using normalization methods (e.g., part-by-part normalization, normalization by GC content, linear and nonlinear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted regression scatter smoothing), PERUN, ChAI, duplication masking (RM), GC-normalization and duplication masking (GCRM), cQn, normalization methods known in the art, and / or combinations thereof).

[0210] Any suitable number of standardization times can be used. In some implementations, the dataset can be standardized one or more times, five or more times, ten or more times, or even twenty or more times. The dataset can be standardized to values ​​(e.g., standardized values) that represent any suitable characteristic or variable (e.g., sample data, reference data, or both). Non-limiting examples of available data standardization types include standardizing raw count data of one or more selected test or reference portions to the total counts mapped to the selected portion or segment of the chromosome or the whole genome; standardizing raw count data of one or more selected portions to the median reference count mapped to one or more portions or the chromosome of the selected portion or segment; standardizing raw count data to the aforementioned standardized data or its derivatives; and standardizing the aforementioned standardized data to one or more other predetermined standardization variables. Standardizing the dataset can sometimes separate statistical errors, depending on the characteristics or properties selected as predetermined standardization variables. Standardizing the dataset can also sometimes make data characteristics of data of different magnitudes comparable by converting the data to a common scale (e.g., predetermined standardization variables). In some implementations, one or more standardizations of statistically derived values ​​can be used to minimize data variance and reduce the importance of outlying data. When it comes to standardized values, partial standardization of a portion or a reference genome is sometimes referred to as "partial standardization."

[0211] In some embodiments, the processing steps include standardization, including standardization to a static window, and in some embodiments, the processing steps include standardization, including standardization to a dynamic or sliding window. The term "window" herein refers to the selection of one or more portions for analysis, sometimes used as a reference for comparison (e.g., for standardization and / or other mathematical or statistical operations). The term "standardization to a static window" herein refers to a standardization process using one or more portions selected for comparing test and reference datasets. In some embodiments, the selected portions are used to generate profiles. A static window typically includes a predetermined set of portions that does not change during operation and / or analysis. The terms "standardization to a dynamic window" or "standardization sliding window" herein refer to the standardization of portions located within a genomic region (e.g., genetically closely surrounding, contiguous portions or segments) of a selected test portion, wherein one or more of the selected test portions are standardized to portions closely surrounding the selected test portion. In some embodiments, the selected portions are used to generate profiles. Sliding or dynamic window normalization typically involves repeatedly moving or sliding to adjacent test portions and normalizing newly selected test portions to portions closely surrounding or adjacent to said newly selected test portions, wherein adjacent windows have one or more shared portions. In some implementations, multiple selected test portions and / or chromosomes can be analyzed via a sliding window process.

[0212] In some implementations, normalization to a sliding or dynamic window can produce one or more values, each representing normalization for different sets of reference portions selected from different regions of the genome (e.g., chromosomes). In some implementations, the resulting one or more values ​​are cumulative values ​​(e.g., numerical estimates of the integral of a normalized count profile over a selected portion, domain (e.g., a portion of a chromosome), or chromosome). Values ​​obtained from the sliding or dynamic window procedure can be used to generate profiles and facilitate obtaining results. In some implementations, the cumulative sum of one or more portions can be displayed as a function of genomic location. Dynamic or sliding window analysis is sometimes used to analyze the presence of microdeletions and / or microinsertions in the genome. In some implementations, displaying the cumulative sum of one or more portions is used to identify regions of genetic variation (e.g., microdeletions, microduplications). In some implementations, dynamic or sliding window analysis is used to identify genomic regions containing microdeletions, and in some implementations, dynamic or sliding window analysis is used to identify genomic regions containing microduplications.

[0213] The following describes in detail some examples of the standardization processes that can be used, such as LOESS, PERUN, ChAI, and principal component standardization methods.

[0214] In some implementations, the processing steps include weighting. As used herein, the terms “weighted,” “weighted,” or “weighting function,” or their syntactic derivatives or equivalents, refer to a mathematical processing of part or all of a dataset, which is sometimes used to alter the influence of certain dataset characteristics or variables on other dataset characteristics or variables (e.g., increasing or decreasing the importance and / or baseline of data contained in one or more portions of a selected reference genome based on the quality or utility of the data). In some implementations, weighting functions can be used to increase the influence of data with relatively small measurement variables and / or decrease the influence of data with relatively large measurement differences. For example, portions of the reference genome containing excessively low-frequency or low-quantity sequence data can be “downweighted” to minimize their influence on the dataset, while selected portions of the reference genome can be “upweighted” to increase their influence on the dataset. A non-limiting example of a weighting function is [1 / (standard deviation)]. 2 The weighting step is sometimes performed in a manner substantially similar to the standardization step. In some implementations, the data set is divided by a predetermined variable (such as a weighting variable). Often, the predetermined variable (such as minimizing the target function, Phi) is chosen to selectively weight different parts of the data set (e.g., increasing the influence of certain data types while decreasing the influence of others).

[0215] In some implementations, the processing steps may include one or more mathematical and / or statistical processing steps. Any suitable mathematical and / or statistical processing step may be used alone or in combination to analyze and / or process the data set described herein. Any suitable number of mathematical and / or statistical processing steps can be used. In some implementations, the data set may be subjected to mathematical and / or statistical processing one or more times, five or more times, ten or more times, or twenty or more times. Non-limiting examples of mathematical and statistical processing steps that can be used include addition, subtraction, multiplication, division, algebraic functions, least squares estimation, curve fitting, differential equations, rational polynomials, double polynomials, orthogonal polynomials, z-scores, p-values, chi-values, phi-values, peak level analysis, determining peak margin positions, calculating peak area ratios, analyzing median chromosome levels, calculating arithmetic mean absolute deviation, residual sum of squares, average, standard deviation, standard error, etc., or combinations thereof. Mathematical and / or statistical processing can be performed on all or part of the sequence read data or the processed results thereof. Non-restrictive examples of statistically manipulated data set variables or characteristics include raw counts, filtered counts, standardized counts, peak height, peak width, peak area, peak margin, lateral tolerance, p-value, median level, average level, count distribution within genomic regions, relative representation of nucleic acid content, or combinations thereof.

[0216] In some implementations, the processing steps may include using one or more statistical algorithms. Any suitable statistical algorithm may be used alone or in combination to analyze and / or process the data set described herein. Any suitable number of statistical algorithms may be used. In some implementations, one or more, five or more, ten or more, or twenty or more statistical algorithms may be used to analyze the data set. Non-limiting examples of statistical algorithms suitable for use with the methods described herein include decision trees, counternull, multiple comparisons, comprehensive tests, the Behrens-Fischer problem, bootstrapping, Fisher's method combined with a significance test for independence, null hypothesis, Type I error, Type II error, exact test, one-sample Z-test, two-sample Z-test, one-sample t-test, paired t-test, two-sample pooled t-test with equal variances, two-sample unpooled t-test with unequal variances, single proportion Z-test, pooled two-proportion Z-test, unpooled two-proportion Z-test, one-sample chi-square test, two-sample F-test with equal variances, confidence intervals, confidence intervals, significance, meta-analysis, simple linear regression, strong linear regression, or a combination thereof. Non-restricted examples of data set variables or features that can be analyzed using statistical algorithms include raw counts, filtered counts, standardized counts, peak height, peak width, peak margin, lateral tolerance, p-value, median level, average level, count distribution within genomic regions, relative representation of nucleic acid content, or combinations thereof.

[0217] In some implementations, the dataset can be analyzed using multiple (e.g., two or more) statistical algorithms, such as least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, Bagging, neural networks, support vector machine models, random forests, classification tree models, k-nearest neighbors, logistic regression, and / or smoothing loss. Smoothing) and / or mathematical and / or statistical operations (such as those described herein). In some embodiments, using multiple operations can generate an N-dimensional space that can be used to provide results. In some embodiments, analyzing a dataset using multiple operations can reduce the complexity and / or dimensionality of the dataset. For example, using multiple operations on a reference dataset can generate an N-dimensional space (e.g., a probability plot) that can be used to represent the presence or absence of genetic variation depending on the genetic status of the reference sample (e.g., positive or negative for a selected genetic variation). Analyzing test samples using substantially similar sets of operations can generate N-dimensional points for each of the tested samples. The complexity and / or dimensionality of the test dataset is sometimes reduced to N-dimensional points or single values ​​that can be easily compared to the N-dimensional space of the reference data. Test sample data falling within the N-dimensional space filled by the reference data indicates a genetic status substantially similar to that of the reference. Test sample data falling outside the N-dimensional space filled by the reference data indicates a genetic status substantially dissimilar to that of the reference. In some embodiments, the reference is euploid or does not have genetic variation or medical symptoms.

[0218] In some implementations, after computation, optional filtering, and standardization, the processed data set can be further manipulated using one or more filtering and / or standardization procedures. In some implementations, data sets that can be further manipulated using one or more filtering and / or standardization procedures can be used to generate profiles. In some implementations, one or more filtering and / or standardization procedures can sometimes reduce the complexity and / or dimensionality of the data set. Results can be provided based on the data set with reduced complexity and / or dimensionality.

[0219] In some implementations, portions may be filtered based on error measurements (e.g., standard deviation, standard error, calculated variance, p-value, arithmetic mean absolute error (MAE), mean absolute deviation, and / or arithmetic mean absolute deviation (MAD). In some implementations, the error measurement refers to count variability. In some implementations, portions are filtered based on count variability. In some implementations, count variability is an error measurement determined for the counts of portions (i.e., portions) mapped to a reference genome from multiple samples (e.g., multiple samples obtained from multiple objects, such as 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more, or 10,000 or more objects). In some implementations, portions with count variability exceeding a predetermined upper limit may be filtered (e.g., excluded from consideration). In some implementations, the predetermined upper limit is equal to or greater than about 50, about 52, about 54, about 56, about 58, about 60, about 62, about 64, about 66, about 68, about 70, about 72, about 74, or equal to or greater than about 76. The MAD value. In some embodiments, portions of count variability below a predetermined lower limit range may be filtered (e.g., excluded from consideration). In some embodiments, the predetermined lower limit range is a MAD value equal to or less than about 40, about 35, about 30, about 25, about 20, about 15, about 10, about 5, about 1, or equal to or less than about 0. In some embodiments, portions of count variability exceeding the predetermined range may be filtered (e.g., excluded from consideration). In some embodiments, the predetermined range is greater than 0 and less than about 7. 6. MAD values ​​less than approximately 74, less than approximately 72, less than approximately 71, less than approximately 70, less than approximately 69, less than approximately 68, less than approximately 67, less than approximately 66, less than approximately 65, less than approximately 64, less than approximately 62, less than approximately 60, less than approximately 58, less than approximately 56, less than approximately 54, less than approximately 52, and less than approximately 50. In some embodiments, the predetermined range is MAD values ​​greater than 0 and less than approximately 67.7. In some embodiments, the portion of the count variability within the predetermined range is selected (e.g., for determining the presence of genetic variation).

[0220] In some embodiments, the portion of the count variability represents a distribution (e.g., a normal distribution). In some embodiments, a portion may be selected within the quantiles of the distribution. In some embodiments, the quantiles of the distribution are selected to be equal to or less than about 99.9%, 99.8%, 99.7%, 99.6%, 99.5%, 99.4%, 99.3%, 99.2%, 99.1%, 99.0%, 98.9%, 98.8%, 98.7%, 98.6%, 98.5%, 98.4%, 98.3%, 98.2%, 98.1%, 98.0%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 85%, 80%, or equal to or less than about 75%. In some embodiments, the quantiles of the distribution of count variability are selected to be within 99%. In some implementations, portions with MAD>0 and MAD<67.725, within the 99th percentile, are selected to identify stable subsets of the reference genome.

[0221] Non-limiting examples of partial filtering involving PERUN (e.g.) are described herein and International Patent Application No. PCT / US12 / 59123 (WO2013 / 052913), the entire contents of which are incorporated herein by reference, including all text, tables, equations, and figures. Partial filtering may be based on or partially on error measurements. Error measurements include the absolute value of the deviation, such as an R-factor, which in some embodiments may be used for partial removal or weighting. In some embodiments, the R-factor is defined as the sum of the absolute deviations of predicted and measured values ​​divided by the predicted counts from the measured values ​​(e.g., Formula B described herein). While error measurements including the absolute value of the deviation may be used, suitable error measurements may also be used. In some embodiments, error measurements excluding the absolute value of the deviation may be used, such as square-based dispersions. In some embodiments, partial filtering or weighting is performed based on a measure of mappability (e.g., a mappability score). Sometimes, portions are filtered or weighted based on a relatively low number of sequence reads mapped to the portion (e.g., 0, 1, 2, 3, 4, 5 reads mapped to the portion). Portions can also be filtered or weighted based on the type of analysis being performed. For example, for aneuploidy analysis of chromosomes 13, 18, and / or 21, sex chromosomes can be filtered out, and only autosomes or autosomal subgroups can be analyzed.

[0222] In a specific implementation, the following filtering procedure can be used. Select portions of the same group within a given chromosome (e.g., a portion of the reference genome) and compare the number of reads in affected and unaffected samples. The gap involves trisomy 21 and euploid samples, which involves portions covering most of chromosome 21. The portions are identical between euploid and T21 samples. The difference between portions and individual segments is not critical, as defined by the portion. Compare identical genomic regions in different patients. This procedure can be used for trisomy analysis, such as T13 or T18, in addition to or instead of T21.

[0223] In some implementations, after computation, optional filtering, and standardization, the processed data set can be manipulated by weighting. In some implementations, one or more portions may be selectively weighted to reduce the influence of data contained in the selected portions (e.g., noisy data, uninformative data), and in some implementations, one or more portions may be selectively weighted to increase or strengthen the influence of data contained in the selected portions (e.g., data with small measurement variance). In some implementations, a single weighting function is used to weight the data set, which reduces the influence of data with large variance and increases the influence of data with small variance. Weighting functions are sometimes used to reduce the influence of data with large variance and increase the influence of data with small variance (e.g., [1 / (standard deviation)]). 2 In some implementations, the data is further processed by weighting to generate a profile of the processed data to facilitate classification and / or provide results. Results may be provided based on the profile of the weighted data.

[0224] Partial filtering and weighting can be performed at one or more suitable points during the analysis. For example, partial filtering or weighting can be performed before or after the sequence reads are mapped to a portion of the reference genome. In some embodiments, partial filtering or weighting can be performed before or after determining experimental bias in individual genome portions. In some embodiments, partial filtering or weighting can be performed before or after calculating at the genome segment level.

[0225] In some embodiments, after computation, optional filtering, standardization, and optional weighting, the processed data set can be manipulated by one or more mathematical and / or statistical operations (such as statistical functions or statistical algorithms). In some embodiments, the processed data can be further manipulated by calculating Z-scores for one or more selected portions, chromosomes, or portions of chromosomes. In some embodiments, the processed data set can be further manipulated by calculating p-values. See Equation 1 (Example 2) for an implementation of the equation for calculating Z-scores and p-values. In some embodiments, the mathematical and / or statistical operations include one or more assumptions related to ploidy and / or fetal fraction. In some embodiments, further manipulation by one or more mathematical and / or statistical operations produces a profile of the processed data to facilitate classification and / or provide results. Results can be provided based on a profile of the data from mathematical and / or statistical operations. Results provided based on a profile of the data from mathematical and / or statistical operations typically include one or more assumptions related to ploidy and / or fetal fraction.

[0226] In some implementations, the data set is computed, optionally filtered, and standardized, and then various operations are performed on the processed data set to produce an N-dimensional space and / or N-dimensional points. Results can be provided based on an overview plot of the data set analyzed in N dimensions.

[0227] In some implementations, the data set is processed using one or more peak level analyses, peak width analyses, peak edge position analyses, peak lateral tolerance analyses, or their derivative analyses, or combinations thereof, as part of or after the processed and / or manipulated data set. In some implementations, one or more peak level analyses, peak width analyses, peak edge position analyses, peak lateral tolerance analyses, or their derivative analyses, or combinations thereof, are used to generate a profile of the processed data to facilitate classification and / or provide results. Results may be provided based on a profile of the data, which has been processed using one or more peak level analyses, peak width analyses, peak edge position analyses, peak lateral tolerance analyses, or their derivative analyses, or combinations thereof.

[0228] In some embodiments, one or more reference samples substantially free of the genetic variant under investigation can be used to generate a reference median count profile, which yields a predetermined value representing the absence of the genetic variant and typically deviates from a predetermined value in the region corresponding to a genomic location in the test subject where the genetic variant is located, if the test subject has the genetic variant. In test subjects with a condition associated with the genetic variant or at risk of such a variant, the numerical value of the selected portion or segment is expected to differ significantly from the predetermined value for the unaffected genomic location. In some embodiments, one or more reference samples known to carry the genetic variant under investigation can be used to generate a reference median count profile, which yields a predetermined value representing the presence of the genetic variant and typically deviates from a predetermined value in the area corresponding to a genomic location where the test subject does not have the genetic variant. In test subjects without a condition associated with the genetic variant or at no risk of such a variant, the numerical value of the selected portion or segment is expected to differ significantly from the predetermined value for the affected genomic location.

[0229] In some implementations, the analysis and processing of data can include the use of one or more hypotheses. An appropriate number or type of hypotheses may be used to analyze or process the dataset. Non-limiting examples of hypotheses that can be used for data processing and / or analysis include maternal ploidy, fetal baseline, prevalence of certain sequences in a reference population, racial background, prevalence of medical conditions selected from relevant family members, correspondence between raw count distributions from different patients and / or runs after GC normalization and repetition masking (e.g., GCRM), identical matches (e.g., identical base positions) representing PCR artifacts, inherent assumptions in fetal quantification assays (e.g., FQA), assumptions about twins (e.g., if there are two twins and only one is affected, the effective fetal score is only 50% of the total fetal score measured (similar to triplets, quadruplets, etc.)), uniformly covering the entire genome of fetal cell-free DNA (e.g., cfDNA), and combinations thereof.

[0230] In those examples where the quality and / or depth of the mapped sequence reads cannot predict the presence of genetic variation at the desired confidence level (e.g., 95% or higher), one or more additional mathematical processing algorithms and / or statistical prediction algorithms can be used, based on a standardized count distribution, to generate additional numerical values ​​that can be used for data analysis and / or to provide results. The term "standardized count distribution" as used herein refers to a distribution generated using standardized counts. This document describes examples of methods that can be used to generate standardized counts and standardized count distributions. The already counted location sequence reads can be standardized relative to the test sample count or the reference sample count. In some embodiments, the standardized count profile can be represented graphically.

[0231] LOESS Standardization

[0232] LOESS is a known regression model in the art that combines multiple regression models in a k-nearest neighbor-based meta-model. LOESS sometimes refers to locally weighted polynomial regression. In some implementations, GC LOESS applies the LOESS model to the relationship between GC composition and fragment counts (e.g., sequence reads, counts) of a portion of a reference genome. Plotting a smooth curve using LOESS over a set of data points is sometimes called a LOESS curve, especially when smoothing values ​​are given by a weighted quadratic least squares regression relative to the span of the values ​​of the standard variables in a y-axis scatter plot. For each point in the dataset, the LOESS method fits a low-degree polynomial to the dataset, indicating points where the variable values ​​are close to the evaluated response. A weighted least squares polynomial is fitted such that points close to the evaluated response have more weight and points far away have less weight. The regression function value for each point is then obtained by evaluating the locally weighted polynomial using the indicator variable values ​​for that data point. Sometimes, the LOESS fit is fully considered after the regression function values ​​have been calculated for each data point. Many details of this method, such as the degree and weights of the polynomial model, are flexible.

[0233] PERUN Standardization

[0234] A standardized method for reducing errors associated with nucleic acid indicators, referred to herein as Parametric Error Removal and Unbiased Standardization (PERUN), is described herein and as set forth in International Patent Application PCT / US12 / 59123 (WO2013 / 052913), the entire text of which is incorporated herein by reference, including all text, tables, equations, and figures. The PERUN method can be used with various nucleic acid indicators (e.g., nucleic acid sequence readings) to reduce the influence of errors that confound predictions based on these indicators.

[0235] For example, the PERUN method is used for nucleic acid sequence readings from samples and reduces the error impact of impaired genomic segment-level determination. This application is effective for using nucleic acid sequence readings to determine the presence of genetic variation in objects exhibiting nucleotide sequences at various levels (e.g., partial, genomic segment level). Non-limiting examples of variation in partial sequences are chromosomal aneuploidy (e.g., trisomy 21, trisomy 18, trisomy 13) and the presence of sex chromosomes (e.g., XX in females and XY in males). Autosomal trisomy (e.g., chromosomes other than sex chromosomes) can be referred to as affected autosomes. Other non-limiting examples of variation at the genomic segment level include microdeletions, microinsertions, duplications, and mosaicism.

[0236] In some applications, the PERUN method can reduce experimental bias in normalized nucleic acid reads mapped to specific portions of a reference genome, referred to as portions and sometimes as the reference genome. In this application, the PERUN method typically normalizes the counts of nucleic acid reads at specific portions of a reference genome spanning a large number of samples in three dimensions. A detailed description of PERUN and its applications can be found in the Examples section, as well as in International Patent Application PCT / US12 / 59123 (WO2013 / 052913) and U.S. Patent Application Publication No. US20130085681, the entire contents of which are incorporated herein by reference, including all text, tables, equations, and figures.

[0237] In some implementations, the PERUN method includes calculating genomic segment levels of a reference genome portion from the following results: (a) sequence read counts of the test sample mapped to the reference genome portion, (b) experimental deviations (e.g., GC deviations) of the test sample from the test sample, and (c) one or more fitting parameters (e.g., fit estimates) for the fit relationship between (i) the experimental deviation of the reference genome portion mapped to the sequence reads and (ii) the counts of sequence reads mapped to said portion. The experimental deviation of each reference genome portion can be determined across multiple samples based on the fit relationship between (i) the sequence read counts mapped to each reference genome portion and (ii) the mapping characteristics of each reference genome portion. Such fit relationships for each sample can be aggregated across multiple samples in a three-dimensional direction. In some implementations, this aggregate can be arranged according to experimental deviations, although the PERUN method can be implemented without arranging the aggregate according to experimental deviations. The fit relationships for each sample and the fit relationships for each portion of the reference genome can be individually fitted to linear or nonlinear functions using suitable fitting procedures known in the art.

[0238] In some implementations, the relationship is a geometric and / or graphical relationship. In some implementations, the relationship is a mathematical relationship. In some implementations, the relationship is plotted. In some implementations, the relationship is a linear relationship. In some implementations, the relationship is a non-linear relationship. In some implementations, the relationship is a regression (e.g., a regression line). The regression can be linear or non-linear. The relationship can be expressed by a mathematical equation. Typically, the relationship is partially defined by one or more constants. The relationship can be generated by methods known in the art. In some implementations, a two-dimensional relationship can be generated for one or more samples, and error tests or probable error tests can be selected for one or more of the said dimensions. For example, the relationship can be generated using plotting software known in the art, which plots two or more variable values ​​provided by the user. The relationship can be fitted using methods known in the art (e.g., plotting software). Some relationships can be fitted by linear regression, and linear regression can generate slopes and intercepts. Some relationships are sometimes non-linear and can be fitted by non-linear functions, such as parabolas, hyperbolas, or exponential functions (e.g., quadratic functions).

[0239] In the PERUN method, one or more fitting relationships can be linear. To analyze cell-free circulating nucleic acids in pregnant women, where the experimental bias is GC bias and the mapping characteristic is GC content, the fitting relationship between (i) the sequence read counts mapped to each part of the sample and (ii) the GC content of each part of the reference genome can be linear. For the latter fitting relationship, when fitting a ensemble relationship among multiple samples, a slope and GC bias coefficient involving GC bias can be determined for each sample. In this embodiment, the fitting relationship between the GC bias coefficients of the i) parts of multiple samples and parts, and (ii) the sequence read counts mapped to the parts, can also be linear. The intercept and slope can be obtained from the latter fitting relationship. In this application, the slope represents sample-specific bias based on GC content, and the intercept represents a part-specific decay pattern present in all samples. When calculating at the genomic segment level to provide results (e.g., the presence of genetic variation; determining fetal sex), the PERUN method can significantly reduce sample-specific bias and part-specific decay.

[0240] In some implementations, PERUN standardization uses a fitting to a linear function and is described as in Equations A, B, or their derivatives.

[0241] Equation A:

[0242] M = LI + GS(A)

[0243] Equation B:

[0244] L=(M–GS) / I(B)

[0245] In some implementations, L is the PERUN normalization level or profile. In some implementations, L is the output required from the PERUN normalization procedure. In some implementations, L is part-specific. In some implementations, L is determined based on multiple parts of the reference genome, representing the PERUN normalization level of the genome, chromosome, its parts, or segments. The level L is often used for further analysis (e.g., determining Z-values, maternal deletions / duplications, fetal microdeletions / microduplications, fetal sex, sex aneuploidy, etc.). The normalization method according to Equation B is called Parametric Error Removal and Unbiased Normalization (PERUN).

[0246] In some implementations, G is a GC bias coefficient measured using a linear model, LOESS, or any equivalent method. In some implementations, G is a slope. In some implementations, the GC bias coefficient G is evaluated as the slope of the regression between the count M (e.g., raw count) for part i and the GC content of part i determined from a reference genome. In some implementations, G represents secondary information extracted from M and determined based on the relationship. In some implementations, G represents the relationship between a set of part-specific counts and a set of part-specific GC content values ​​for a sample (e.g., a test sample). In some implementations, the part-specific GC content is derived from a reference genome. In some implementations, the part-specific GC content is derived from observed or measured GC content (e.g., measured from a sample). The GC bias coefficient is typically determined for each sample in a sample set and typically for the test sample. The GC bias coefficient is typically sample-specific. In some implementations, the GC bias coefficient is a constant. In some implementations, the GC bias coefficient does not change once obtained from the sample.

[0247] In some implementations, S is the slope derived from a linear relationship and I is the intercept. In some implementations, the relationships from which I and S are derived differ from the relationships from which G is derived. In some implementations, the relationships from which I and S are derived are fixed for a given experimental setting. In some implementations, I and S are derived from a linear relationship based on counts (e.g., raw counts) and GC bias coefficients based on multiple samples. In some implementations, I and S are independently derived from the test sample. In some implementations, I and S are derived from multiple samples. I and S are typically part-specific. In some implementations, I and S are determined for all parts of the reference genome in euploid samples using the assumption L=1. In some implementations, a linear relationship is determined for euploid samples, and I and S values ​​specific to selected parts are determined (assuming L=1). In some implementations, the same procedure is applied to all parts of the reference genome in the human genome, and sets of intercepts I and slopes S are determined for each part.

[0248] In some implementations, cross-validation is applied. Cross-validation sometimes refers to rotation estimation. In some implementations, cross-validation is used to evaluate the accuracy of a predictive model (e.g., PERUN) implemented on test samples. In some implementations, a round of cross-validation involves dividing the data samples into complementary subgroups, performing cross-validation analysis on the subgroups (e.g., sometimes called the training group), and performing validation analysis using another subgroup (e.g., sometimes called the validation or test group). In some implementations, multiple rounds of cross-validation are performed using different partitioning products and / or different subgroups. Non-limiting examples of cross-validation methods include leave-one-out, sliding margin, K-fold, 2-fold, repeated random sampling, etc., or combinations thereof. In some implementations, cross-validation randomly selects a working group containing 90% of the sample groups, including known euploid fetuses, and uses that subgroup to train the model. In some implementations, random selection is repeated 100 times, with each portion generating 100 slopes and 100 intercepts.

[0249] In some implementations, the M value is a measurement derived from the test sample. In some implementations, M is a raw count for a portion of the measurement. In some implementations, where values ​​I and S are available for a portion, the M measurement is determined from the test sample and used to determine the PERUN normalization level L of the genome, chromosome, its segment, or portion according to Equation B.

[0250] Therefore, applying the PERUN method in parallel to sequence readings of multiple samples can significantly reduce errors caused by (i) sample-specific experimental bias (e.g., GC bias) and (ii) sample-specific attenuation common to both sources. Other methods that address these two sources of error individually or sequentially typically cannot reduce them as effectively as the PERUN method. Unbound by theory, the PERUN method is expected to reduce errors more effectively, partly because its general addition process does not expand as dramatically as the general multiplication process employed in other normalization methods (e.g., GC-LOESS).

[0251] Other standardization and statistical techniques can be used in conjunction with the PERUN method. Other procedures can be applied before, after, and / or during the use of the PERUN method. Non-limiting examples of procedures that can be used in conjunction with the PERUN method are described below.

[0252] In some implementations, secondary normalization or adjustment of GC content at the genomic segment level can be combined with the PERUN method. Appropriate GC content adjustment or normalization procedures (e.g., GC-LOESS, GCRM) can be used. In some implementations, specific samples can be identified for other GC normalization processes. For example, the application of the PERUN method can determine the GC bias of each sample, and samples with GC biases above a certain threshold can be selected for other GC normalization processes. In this implementation, a predetermined threshold level can be used to select the sample for other GC normalization processes.

[0253] In some embodiments, partial filtering or weighting processes may be used in conjunction with the PERUN method. Suitable partial filtering or weighting processes may be used, and non-limiting examples are described herein, as well as in International Patent Application PCT / US12 / 59123 (WO2013 / 052913) and U.S. Patent Application Publication No. US20130085681, the entire contents of which are incorporated herein by reference, including all text, tables, equations, and figures. In some embodiments, normalization techniques for reducing associated maternal insertions, duplications, and / or deletions (e.g., maternal and / or fetal copy number variations) are used in conjunction with the PERUN method.

[0254] Genomic segment levels calculated using the PERUN method can be used directly to provide results. In some implementations, genomic segment levels can be used directly to provide sample results where the fetal fraction is approximately 2% to approximately 6% or higher (e.g., approximately 4% or higher). Genomic segment levels calculated using the PERUN method are sometimes further processed to provide results. In some implementations, the calculated genomic segment levels are normalized. In some implementations, the sum, arithmetic mean, or median of the calculated genomic segment levels for the test portion (e.g., chromosome 21) can be divided by the sum, arithmetic mean, or median of the calculated genomic segment levels for portions other than the test portion (e.g., chromosome 21 other than autosomes) to generate experimental genomic segment levels. Experimental genomic segment levels or raw genomic segment levels can be used as part of a programmed analysis, such as calculating Z-scores. Z-scores for samples can be generated by subtracting the expected genomic segment levels from experimental or raw genomic segment levels, and the resulting value can be divided by the standard deviation of the sample. In some implementations, the resulting Z-scores can be distributed and analyzed across different samples, or correlated with other variables, such as fetal scores and others, and analyzed to provide results.

[0255] As described herein, the PERUN method is not limited to standardization based on GC bias and GC content alone, and can be used to reduce errors associated with other sources of error. A non-limiting example of a source of non-GC content bias is mappability. When addressing standardized parameters other than GC bias and content, one or more fitting relationships can be non-linear (e.g., hyperbolic, exponential). In some implementations, for example, when experimental bias is determined from a non-linear relationship, experimental bias curvature estimates can be analyzed.

[0256] The PERUN method can be applied to a variety of nucleic acid indicators. Non-limiting examples of nucleic acid indicators are nucleic acid sequence reads and nucleic acid levels at specific locations on a microarray. Non-limiting examples of sequence reads include those obtained from cell-free circulating DNA, cell-free circulating RNA, cellular DNA, and cellular RNA. The PERUN method can be applied to sequence reads mapped to suitable reference sequences, such as genomic reference DNA, cellular reference RNA (e.g., transcriptome), and portions thereof (e.g., portions of genomic complements of DNA or RNA transcriptomes, portions of chromosomes).

[0257] Therefore, in some implementations, cellular nucleic acids (e.g., DNA or RNA) can be used as nucleic acid indicators. Cellular nucleic acid readings mapped to a reference genome portion can be normalized using the PERUN method. Cellular nucleic acids bound to specific proteins sometimes refer to the chromatin immunoprecipitation (ChIP) process. ChIP-enriched nucleic acids are nucleic acids, such as DNA or RNA, associated with cellular proteins. ChIP-enriched nucleic acid readings can be obtained using techniques known in the art. ChIP-enriched nucleic acid readings can be mapped to one or more portions of a reference genome, and the results can be normalized using the PERUN method to provide the outcome.

[0258] In some implementations, cellular RNA can be used as a nucleic acid indicator. Cellular RNA readings can be mapped to a reference RNA portion and normalized using the PERUN method to provide results. A known sequence of cellular RNA (referred to as the transcriptome) or a segment thereof can be used as a reference, to which RNA readings from the sample can be mapped. Sample RNA readings can be obtained using techniques known in the art. The results of mapping RNA readings to a reference can be normalized using the PERUN method to provide results.

[0259] In some implementations, microarray nucleic acid levels can be used as nucleic acid indicators. The PERUN method can be used to analyze the nucleic acid levels or hybrid nucleic acids at specific locations on the array, thereby standardizing the nucleic acid indicators provided by microarray analysis. In this way, specific locations or hybrid nucleic acids on the microarray are similar to portions of the mapped nucleic acid sequence readings, and the PERUN method can be used to standardize microarray data to provide improved results.

[0260] ChAI Standardization

[0261] Other standardized methods that can be used to reduce errors in associated nucleic acid indicators, referred to herein as ChAI, typically employ principal component analysis. In some embodiments, principal component analysis includes (a) filtering portions of a reference genome based on a read density distribution to provide a read density profile of the test sample, including the read densities of the filtered portions, wherein the read densities include sequence reads of circulating cell-free nucleic acids from test samples from pregnant women, and determining a read density distribution for portions of multiple samples; (b) adjusting the read density profile of the test sample based on one or more principal components, wherein the principal components are obtained from a group of known euploid samples via principal component analysis, to provide a test sample profile, including the adjusted read densities; and (c) comparing the test sample profile with a reference profile to provide a comparative relationship. In some embodiments, principal component analysis includes (d) determining the presence of genetic variation in the test sample based on the comparison.

[0262] Filter section

[0263] In some embodiments, a filtering process removes one or more portions (e.g., portions of the genome) from consideration. In some embodiments, one or more portions are filtered (e.g., undergo a filtering process) to provide filtered portions. In some embodiments, the filtering process removes certain portions and retains portions (e.g., subsets of portions). After the filtering process, the retained portions generally refer to the filtered portions herein. In some embodiments, portions of the reference genome are filtered. In some embodiments, portions of the reference genome removed by the filtering process are not included in determining the presence of genetic variations (e.g., chromosomal aneuploidy, microduplication, microdeletion). In some embodiments, portions associated with read densities (e.g., read densities for portions) are removed by the filtering process, and the read densities of the removed portions are not included in determining the presence of genetic variations (e.g., chromosomal aneuploidy, microduplication, microdeletion). In some embodiments, the read density profile includes and / or comprises the read densities of the filtered portions. Any suitable criteria and / or methods known in the art or described herein can be used to select, filter, and / or remove portions from consideration. Non-limiting examples of standards for filtering portions include redundant data (e.g., readings with redundant or overlapping mappings), non-informative data (e.g., portions of a reference genome with zero mapping counts), portions of a reference genome with sequences that appear too frequently or too infrequently, GC content, noisy data, mappability, counts, count variability, reading density, reading density variability, measurements of uncertainty, measurements of repeatability, etc., or combinations thereof. Portions are sometimes filtered based on count distributions and / or reading density distributions. In some embodiments, portions are filtered based on count distributions and / or reading density distributions, wherein the counts and / or reading densities are obtained from one or more reference samples. Sometimes one or more reference samples refer to the training group herein. In some embodiments, portions are filtered based on count distributions and / or reading density distributions, wherein the counts and / or reading densities are obtained from one or more test samples. In some embodiments, portions are filtered based on measurements of uncertainty in the reading density distribution. In some embodiments, portions showing large deviations in the reading density are removed by the filtering process. For example, the distribution of reading density (e.g., the distribution of average, arithmetic mean, or median reading density, e.g.) can be determined. Figure 37A In this distribution, each reading density is mapped to the same segment. Uncertainty measures (e.g., MAD) can be determined by comparing the distributions of reading densities across multiple samples, where segments of the genome are associated with the uncertainty measure. Following the foregoing example, segments can be filtered based on uncertainty measures (e.g., standard deviation (SD), MAD) that associate each segment with a predetermined threshold. Figure 37B The distribution of MAD values ​​is shown in the display section, determined based on the reading density distribution of various samples. Predetermined thresholds are indicated by vertical dashed lines, which enclose an acceptable range of MAD values. Figure 37B In the examples, the filtering process retains portions of MAD values ​​within the acceptable range and removes portions of MAD values ​​outside the acceptable range from consideration. In some embodiments, according to the foregoing examples, portions of reading density values ​​(e.g., median, average, or arithmetic mean reading densities) exceeding a predetermined uncertainty measurement are typically removed from consideration by the filtering process. In some embodiments, portions of reading density values ​​(e.g., median, average, or arithmetic mean reading densities) exceeding the interquartile range of the distribution are typically removed from consideration by the filtering process. In some embodiments, portions of reading density values ​​exceeding 2, 3, 4, or 5 times the interquartile range of the distribution are removed from consideration by the filtering process. In some embodiments, portions of reading density values ​​exceeding 2σ, 3σ, 4σ, 5σ, 6σ, 7σ, or 8σ (e.g., where σ is a range defined by the standard deviation) are removed from consideration by the filtering process.

[0264] In some implementations, the system includes a filtering module (18, Figure 42A The filtering module typically accepts, retrieves, and / or stores the read densities of portions (e.g., portions of predetermined size and / or overlapping regions, referring to the location of portions within the genome) and associated portions, usually derived from other suitable modules (e.g., distribution module 12). Figure 42A In some implementations, the selected portion (e.g., 20( Figure 42A (e.g., the filtered portion), provided by a filtering module. In some embodiments, a filtering module is required to provide the filtered portion and / or the portion removed from the consideration. In some embodiments, the filtering module removes the reading density from the consideration, where the reading density is associated with the removed portion. The filtering module typically provides selected portions (e.g., the filtered portion) to other suitable modules (e.g., distribution module 12, ...). Figure 42A A non-limiting example of a filtering module is shown in Example 7.

[0265] Deviation Assessment

[0266] Sequencing technologies are susceptible to deviations from various sources. Sometimes these deviations are local (e.g., local genomic deviations). Local biases typically occur at the sequence read level. Local genomic biases can be any suitable local bias. Non-limiting examples of local biases include sequence biases (e.g., GC bias, AT bias, etc.), biases associated with DNase I sensitivity, entropy, repetitive sequence biases, chromatin structure biases, polymerase error rate biases, palindromic biases, insertion repeat biases, PCR-related biases, etc., or combinations thereof. In some implementations, the source of the local bias is undetermined or unknown.

[0267] In some embodiments, a local genomic bias assessment is determined. Local genomic bias assessment is sometimes referred to herein as local genomic deviation assessment. Local genomic bias assessment may be determined with respect to a reference genome, its segments, or portions. In some embodiments, a local genomic deviation assessment is determined for one or more sequence reads (e.g., some or all sequence reads of a sample). Local genomic deviation assessment of sequence reads is typically determined based on a local genomic deviation assessment of the corresponding location and / or position of a reference (e.g., a reference genome). In some embodiments, local genomic deviation assessment includes quantitatively measuring sequence deviation (e.g., sequence reads, reference genome sequences). Local genomic deviation assessments can be determined by suitable methods or mathematical processes. In some embodiments, local genomic deviation assessments are determined by suitable methods and / or suitable distribution functions (e.g., PDFs). In some embodiments, local genomic bias assessments include a quantitative representation of the PDF. In some embodiments, local genomic bias assessments (e.g., probability density assessments (PDEs), core density assessments) are determined by a probability density function of the local bias content (e.g., the PDF, e.g., a core density function). In some embodiments, density assessments include core density assessments. Local genomic bias assessments are sometimes expressed as the mean, arithmetic mean, or median of a distribution. Sometimes local genomic preference assessments are expressed as sums or integrals (e.g., the area under the curve (AUC) of a suitable distribution).

[0268] PDFs (e.g., core density functions, such as the Epanechnikov core density function) typically include bandwidth variables (e.g., bandwidth). The bandwidth variable typically defines the size and / or length of the window from which the probability density evaluation (PDE) is derived when using the PDF. The window from which the PDE is derived typically includes a defined length of polynucleotides. In some implementations, the window from which the PDE is derived is a portion. The portion is typically determined based on the bandwidth variable (e.g., portion size, portion length). The bandwidth variable determines the length or size of the window used to determine the local genome preference evaluation. The local genome preference evaluation is determined from the length of the polynucleotide segment (e.g., a continuous segment of nucleotide bases). PDEs (e.g., read density, local genomic preference assessment (e.g., GC density)) can be determined using any suitable bandwidth, non-limiting examples of which include bandwidths of about 5 bases to about 100,000 bases, about 5 bases to about 50,000 bases, about 5 bases to about 25,000 bases, about 5 bases to about 10,000 bases, about 5 bases to about 5,000 bases, about 5 bases to about 2,500 bases, about 5 bases to about 1,000 bases, about 5 bases to about 500 bases, about 5 bases to about 250 bases, about 20 bases to about 250 bases, or the like. In some implementations, bandwidths of about 400 bases or less, about 350 bases or less, about 300 bases or less, about 250 bases or less, about 225 bases or less, about 200 bases or less, about 175 bases or less, about 150 bases or less, about 125 bases or less, about 100 bases or less, about 75 bases or less, about 50 bases or less, or about 25 bases or less are used to determine local genomic discrepancy assessments (e.g., GC density). In some implementations, bandwidth is used to determine local genomic discrepancy assessments (e.g., GC density) based on the average, arithmetic mean, median, or maximum read length of sequence reads obtained for a given subject and / or sample. Sometimes, bandwidth is used to determine local genomic bias assessments (e.g., GC density) where the bandwidth is approximately equal to the average, arithmetic mean, median, or maximum read length of sequence reads obtained for a given subject and / or sample. In some implementations, bandwidths of approximately 250, 240, 230, 220, 210, 200, 190, 180, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or approximately 10 bases are used to determine local genomic deviation assessments (e.g., GC density).

[0269] Local genomic discrepancy assessments can be determined at single-base resolution, although local genomic discrepancy assessments (e.g., local GC content) can be determined at even lower resolutions. In some embodiments, local genomic discrepancy assessments are determined based on local discrepancy content. Local genomic bias assessments are typically determined using windows (e.g., using PDF determinations). In some embodiments, local genomic bias assessments include the use of windows comprising a pre-selected number of bases. Sometimes the window comprises a continuous base segment. Sometimes the window comprises one or more portions of non-continuous bases. Sometimes the window comprises one or more portions (e.g., portions of the genome). Window size or length is typically determined by bandwidth and according to the PDF. In some embodiments, the window is about 10 or more, 8 or more, 7 or more, 6 or more, 5 or more, 4 or more, 3 or more, or about 2 or more times the bandwidth length. When determining density assessments using PDFs (e.g., core density functions), the window is sometimes twice the length of the selected bandwidth. The window can comprise any suitable number of bases. In some embodiments, windows comprise about 5 bases to about 100,000 bases, about 5 bases to about 50,000 bases, about 5 bases to about 25,000 bases, about 5 bases to about 10,000 bases, about 5 bases to about 5,000 bases, about 5 bases to about 2,500 bases, about 5 bases to about 1,000 bases, about 5 bases to about 500 bases, about 5 bases to about 250 bases, or about 20 bases to about 250 bases. In some embodiments, the genome or a segment thereof is divided into multiple windows. Windows covering genomic regions may overlap or not overlap. In some embodiments, windows are located at equal distances from each other. In some embodiments, windows are located at unequal distances from each other. In some embodiments, the genome or a segment thereof is divided into multiple sliding windows, wherein windows slide incrementally across the genome or a segment thereof, wherein each incremental window includes a local genomic preference assessment (e.g., local GC density). The window can slide across the genome in any suitable increment according to any numerical form or according to any mathematically defined sequence. In some implementations, for the determination of a local genome preference assessment, the window slides across the genome, or a segment thereof, in increments of: about 10,000 bp or more, about 5,000 bp or more, about 2,500 bp or more, about 1,000 bp or more, about 750 bp or more, about 500 bp or more, about 400 bases or more, about 250 bp or more, about 100 bp or more, about 50 bp or more, or about 25 bp or more. In some implementations, for the determination of a local genome preference assessment, the window slides across the genome, or a segment thereof, in increments of: about 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or about 1 bp. For example, for determining local genome preference assessment, the window may include approximately 400 bp (e.g., 200 bp bandwidth) and may slide across the genome in 1 bp increments.In some implementations, local genomic deviation assessments of individual bases within the genome or its segments are determined using a core density function and approximately 200 bp bandwidth.

[0270] In some embodiments, local genomic deviation assessment is local GC content and / or an expression of local GC content. The term "local" (e.g., used to describe local deviation, local preference assessment, local preference content, local genomic preference, local GC content, etc.) refers to a polynucleotide segment of 10,000 bp or less. In some embodiments, the term "local" refers to a polynucleotide segment of 5000 bp or less, 4000 bp or less, 3000 bp or less, 2000 bp or less, 1000 bp or less, 500 bp or less, 250 bp or less, 200 bp or less, 175 bp or less, 150 bp or less, 100 bp or less, 75 bp or less, or 50 bp or less. Local GC content typically represents (e.g., mathematically, quantitatively) the GC content, sequence reads, or sequence read assemblies (e.g., contiguous groups, profiles, etc.) of a local segment of the genome. For example, local GC content may be a local GC deviation assessment or GC density.

[0271] Typically, the GC density of one or more polynucleotides in a reference or sample (e.g., a test sample) is determined. In some embodiments, GC density is a representation (e.g., a mathematical, quantitative representation) of the local GC content (e.g., a polynucleotide segment of 5000 bp or less). In some embodiments, GC density is a local genomic deviation assessment. GC density can be determined using suitable procedures described herein and / or known in the art. Suitable PDFs (e.g., core density functions, such as the Epanechnikov core density function, see [link to relevant documentation]) can be used. Figure 33 GC density is determined. In some embodiments, GC density is a PDE (e.g., core density assessment). In some embodiments, GC density is defined by the presence of one or more guanine (G) and / or cytosine (C) nucleotides. In some embodiments, GC density is defined by the presence of one or more adenine (A) and / or thymine (T) nucleotides. In some embodiments, the GC density of local GC content is determined based on the complete genome or its segments (e.g., autosomes, chromosome sets, single chromosomes, genes, e.g., see [link to relevant documentation]). Figure 34The GC density is normalized as determined by the reference genome. One or more GC densities of polynucleotides in a sample (e.g., a test sample) or a reference sample can be determined. Typically, the GC density of a reference genome is determined. In some embodiments, the GC density of a sequence read is determined based on a reference genome. The GC density of a read is typically determined based on the GC density determined by the corresponding site and / or position in the reference genome to which the read is mapped. In some embodiments, GC densities determined for location on the reference genome are assigned and / or provided to the read, wherein the read or a segment thereof is mapped to the same site in the reference genome. Any suitable method can be used to determine the site on the reference genome to which the read is mapped for generating the GC density of the read. In some embodiments, the median position of the mapped read determines the site on the reference genome from which the GC density of the read is determined. For example, when the median position of the read is mapped to base x on chromosome 12 of the reference genome, the GC density of the read is typically provided by evaluating the GC density determined by the core density at or near base x on chromosome 12 of the reference genome. In some embodiments, the GC density of some or all base positions of the read is determined based on the reference genome. Sometimes the GC density of a reading includes the average, sum, median, or integral of two or more GC densities determined by multiple base positions on a reference genome.

[0272] In some implementations, local genomic deviation assessments (e.g., GC density) are quantified and / or provided numerically. Local genomic bias assessments (e.g., GC density) are sometimes expressed as mean, arithmetic mean, and / or median. Local genomic bias assessments (e.g., GC density) are sometimes expressed as the maximum peak height of the PDE. Sometimes local genomic bias assessments (e.g., GC density) are expressed as the sum or integral of suitable PDEs (e.g., area under the curve (AUC)). In some implementations, GC density includes core weighting. In some implementations, the GC density of readings includes values ​​approximately equal to: mean, arithmetic mean, sum, median, maximum peak height, or core-weighted integral.

[0273] Deviation frequency

[0274] Deviation frequencies are sometimes determined based on one or more local genomic preference assessments (e.g., GC density). Deviation frequencies are sometimes the count or sum of local genomic deviation assessments of a sample, reference (e.g., reference genome, reference sequence), or a portion thereof. In some implementations, the deviation frequency is the GC density frequency. GC density frequencies are typically determined based on one or more GC densities. For example, a GC density frequency may represent the fold of a GC density of value x relative to the entire genome or its segment. Preference frequencies are typically the distribution of local genomic preference assessments, where the number of occurrences of each local genomic preference assessment represents the preference frequency (see, for example, [link to relevant documentation]). Figure 35 Preference frequencies are sometimes mathematically processed and / or standardized. Preference frequencies can be mathematically processed and / or standardized using appropriate methods. In some embodiments, preference frequencies are standardized based on a representation (e.g., component, percentage) of local genomic preference assessments for a sample, reference, or a portion thereof (e.g., autosome, chromosomal subgroup, single chromosome, or its readings). Deviation frequencies of some or all local genomic preference assessments for a sample or reference can be determined. In some embodiments, deviation frequencies of local genomic preference assessments for some or all sequence readings of a test sample can be determined.

[0275] In some implementations, the system includes a preference density module 6. The preference density module can accept, retrieve, and / or store mapped sequence reads 5 and reference sequence 2 in any suitable format and generate local genome preference assessments, local genome preference distributions, preference frequencies, GC densities, GC density distributions, and / or GC density frequencies (uniformly represented by box 7). In some implementations, the preference density module transfers data and / or information (e.g., 7) to other suitable modules (e.g., relationship module 8).

[0276] relation

[0277] In some implementations, one or more relationships are formed between local genomic preference assessments and preference frequencies. The term "relationship" herein refers to a mathematical and / or geometric relationship between two or more variables or values. Relationships can be generated through suitable mathematical and / or geometric processes. Non-limiting examples of relationships include mathematical and / or geometric representations: functions, correlations, distributions, linear or non-linear equations, lines, regressions, fitted regressions, and combinations thereof. Sometimes a relationship includes a fitted relationship. In some implementations, a fitted relationship includes a fitted regression. Sometimes a relationship includes weighted two or more variables or values. In some implementations, a relationship includes a fitted regression where one or more variables or values ​​of the relationship are weighted. Sometimes regressions are fitted in a weighted form. Sometimes regressions are fitted without weighting. In some implementations, generating a relationship includes plotting or graphing.

[0278] In some implementations, a suitable relationship is determined between local genomic preference assessment and preference frequency. In some implementations, generating a relationship between (i) local genomic preference assessment and (ii) preference frequency of a sample provides a sample preference relationship. In some implementations, generating a relationship between (i) local genomic preference assessment and (ii) preference frequency of a reference provides a reference preference relationship. In some implementations, a relationship is generated between GC density and GC density frequency. In some implementations, generating a relationship between (i) GC density and (ii) GC density frequency of a sample provides a sample GC density relationship. In some implementations, generating a relationship between (i) GC density and (ii) GC density frequency of a reference provides a reference GC density relationship. In some implementations, when the local genomic preference assessment is GC density, the sample preference relationship is the sample GC density relationship and the reference preference relationship is the reference GC density relationship. The GC density of the reference GC density relationship and / or the sample GC density relationship is typically a representation of local GC content (e.g., a mathematical or quantitative representation). In some implementations, the relationship between local genomic preference assessment and preference frequency includes a distribution. In some implementations, the relationship between local genomic preference assessment and preference frequency includes a fitting relationship (e.g., a fitted regression). In some implementations, the relationship between local genomic preference assessments and preference frequencies includes fitting a linear or nonlinear regression (e.g., multinomial regression). In some implementations, the relationship between local genomic preference assessments and preference frequencies includes a weighted relationship, where the local genomic preference assessments and / or preference frequencies are weighted through an appropriate process. In some implementations, the weighted fit relationship (e.g., a weighted fit) can be obtained through a process including quantile regression, parameterized distributions, or interpolation of an empirical distribution. In some implementations, the relationship between local genomic preference assessments and preference frequencies for a test sample, reference, or a portion thereof includes multinomial regression, where the local genomic preference assessments are weighted. In some implementations, the weighted fit model includes values ​​from a weighted distribution. The values ​​of the distribution can be weighted through an appropriate process. In some implementations, values ​​closer to the end of the distribution receive less weight than values ​​closer to the middle of the distribution. For example, for a distribution between local genomic preference assessments (e.g., GC density) and preference frequencies (e.g., GC density frequencies), weights are determined based on the preference frequencies of a given local genomic preference assessment, including local genomic preference assessments with preference frequencies close to the arithmetic mean of the distribution receiving more weight than local genomic preference assessments with preference frequencies farther from the arithmetic mean.

[0279] In some implementations, the system includes a preference relation module 8. The relation module can generate relations and define the functions, coefficients, constants, and variables of the relations. The relation module can receive, store, and / or retrieve data and / or information (e.g., 7) from appropriate modules (e.g., preference density module 6) and generate relations. The relation module typically generates and compares distributions of local genomic preference assessments. The relation module can compare datasets and sometimes generate regression and / or fitting relations. In some implementations, the relation module compares one or more distributions (e.g., sample and / or reference distributions of local genomic preference assessments) and provides weighting factors and / or weighted assignments 9 of sequence read counts to other appropriate modules (e.g., preference correction modules). Sometimes the relation module directly provides standardized counts of sequence reads to distribution module 21, where the counts are standardized based on relations and / or comparisons.

[0280] Generate comparisons and their applications

[0281] In some implementations, reducing local bias in sequence reads includes standardizing sequence read counts. Sequence read counts are typically standardized based on a comparison of the test sample with a reference. For example, sometimes sequence read counts are standardized by comparing a local genomic bias assessment of the test sample's sequence reads with a local genomic bias assessment of a reference (e.g., a reference genome or a portion thereof). In some implementations, sequence read counts are standardized by comparing the bias frequencies of the test sample's local genomic bias assessment with the bias frequencies of the reference's local genomic bias assessment. In some implementations, sequence read counts are standardized by comparing sample bias relationships and reference bias relationships to generate a comparison.

[0282] Sequence reading counts are typically standardized based on a comparison of two or more relations. In some embodiments, two or more relations are compared to provide a comparison for reducing local bias in sequence readings (e.g., standardized counts). Two or more relations can be compared using suitable methods. In some embodiments, the comparison includes adding, subtracting, multiplying, and / or dividing the second relation by the first relation. In some embodiments, comparing two or more relations includes using suitable linear and / or nonlinear regression. In some embodiments, comparing two or more relations includes suitable polynomial regression (e.g., third-order polynomial regression). In some embodiments, the comparison includes adding, subtracting, multiplying, and / or dividing the second regression by the first regression. In some embodiments, two or more relations are compared through a process that includes an inference framework comprising multiple regressions. In some embodiments, two or more relations are compared through a process that includes suitable multivariate analysis. In some embodiments, two or more relations are compared through a process that includes basis functions (e.g., mixture functions, such as polynomial bases, Fourier bases, or the like), spline functions, radial basis functions, and / or wavelets.

[0283] In some embodiments, the distributions of local genomic preference assessments, including the preference frequencies of test samples and references, are compared via a process including multinomial regression, wherein the local genomic preference assessments are weighted. In some embodiments, a multinomial regression is generated between (i) ratios, each ratio including the preference frequency of the reference local genomic preference assessment and the preference frequency of the sample local genomic preference assessment, and (ii) the local genomic preference assessment. In some embodiments, a multinomial regression is generated between (i) the ratio of the preference frequency of the reference local genomic preference assessment to the preference frequency of the sample local genomic preference assessment and (ii) the local genomic preference assessment. In some embodiments, comparing the distributions of local genomic preference assessments of test samples and references includes determining a logarithmic ratio (e.g., log2 ratio) of the preference frequencies of the reference and sample local genomic preference assessments. In some embodiments, comparing the distributions of local genomic preference assessments includes dividing the log ratio (e.g., log2 ratio) of the preference frequency of the reference local genomic preference assessment by the log ratio (e.g., log2 ratio) of the preference frequency of the sample local genomic preference assessment (see, for example, Example 7 and...). Figure 36 ).

[0284] Standardized counts based on comparisons typically adjust some counts without adjusting others. Standardized counts sometimes adjust all counts and sometimes do not adjust any sequence read counts. Sequence read counts are sometimes standardized by a process that includes determining weighting factors, and sometimes this process does not include directly generating and employing weighting factors. Standardized counts based on comparisons sometimes include determining weighting factors for each sequence read count. Weighting factors are typically sequence read-specific and applied to specific sequence read counts. Weighting factors are typically determined based on a comparison of two or more preference relationships (e.g., a sample preference relationship compared to a reference preference relationship). Standardized counts are typically determined by adjusting count values ​​according to weighting factors. Adjusting counts according to weighting factors sometimes involves adding, subtracting, multiplying, and / or dividing sequence read counts by weighting factors. Weighting factors and / or standardized counts are sometimes determined from regressions (e.g., regression lines). Standardized counts are sometimes directly derived from a regression line (e.g., a fitted regression line) obtained by comparing preference frequencies between a reference (e.g., a reference genome) and a local genomic preference assessment of the test sample. In some embodiments, the counts of sample readings are provided as standardized count values ​​based on a comparison between (i) the preference frequency assessed by local genomic preference of the readings and (ii) the preference frequency assessed by a reference local genomic preference. In some embodiments, the obtained sample sequence reading counts are standardized and the discrepancies in the sequence readings are reduced.

[0285] Sometimes the system includes a preference correction module 10. In some embodiments, the function of the preference correction module is performed via a relational simulation module 8. The preference correction module may accept, retrieve, and / or store mapped sequence readings and weighting factors (e.g., 9) from appropriate modules (e.g., relational module 8, compression module 4). In some embodiments, the preference correction module provides counts to the mapped readings. In some embodiments, the preference correction module applies weighting assignments and / or preference correction factors to the sequence reading counts to provide standardized and / or adjusted counts. The preference correction module typically provides standardized counts to other appropriate modules (e.g., distribution module 21).

[0286] In some implementations, normalized counting includes factoring one or more features other than GC density and normalizing sequence read counts. In some implementations, normalized counting includes factoring one or more different local genomic preference assessments and normalizing sequence read counts. In some implementations, sequence read counts are weighted according to weights determined by one or more features (e.g., one or more preferences). In some implementations, counts are normalized according to one or more combined weights. Sometimes factoring one or more features and / or normalizing counts according to one or more combined weights involves a process including the use of a multivariate model. Any suitable multivariate model can be used for normalized counting. Non-limiting examples of multivariate models include multivariate linear regression, multivariate quantile regression, multivariate interpolation of empirical data, nonlinear multivariate models, and combinations thereof.

[0287] In some embodiments, the system includes a multivariate correction module 13. The multivariate correction module can perform functions of the preference density module 6, relation module 8, and / or preference correction module 10 multiple times to adjust the counts of multiple preferences. In some embodiments, the multivariate correction module includes one or more preference density modules 6, relation modules 8, and / or preference correction modules 10. The preference correction module sometimes provides standardized counts 11 to other suitable modules (e.g., distribution module 21).

[0288] Weighted part

[0289] In some implementations, portions are weighted. In some implementations, one or more portions are weighted to provide a weighted portion. Weighted portions sometimes remove partial dependencies. Portions can be weighted through suitable processes. In some implementations, one or more portions are weighted by eigenfunctions (e.g., characteristic functions). In some implementations, eigenfunctions include replacing portions with orthogonal eigenfunctions. In some implementations, the system includes a partial weighting module 42. In some implementations, the weighting module receives, retrieves, and / or stores reading densities, reading density profiles, and / or adjusted reading density profiles. In some implementations, weighted portions are provided by a partial weighting module. In some implementations, a weighting module is required to weight portions. The weighting module can weight portions using one or more weighting methods known in the art or described herein. The weighting module typically provides weighted portions to other suitable modules (e.g., scoring module 46, PCA statistics module 33, profile generation module 26, etc.).

[0290] Principal component analysis

[0291] In some implementations, the reading density profile (e.g., the reading density profile of the test sample) is used. Figure 39AAdjustments are made according to Principal Component Analysis (PCA). The reading density profiles of one or more reference samples and / or the reading density profiles of the test object can be adjusted according to PCA. Removing deviations from the reading density profiles via PCA-related procedures sometimes refers to adjusting the profiles. PCA can be performed using suitable PCA methods or variations thereof. Non-limiting examples of PCA methods include classical correlation analysis (CCA), Karhunen–Loève transform (KLT), Hotelling transform, appropriate orthogonal decomposition (POD), singular value decomposition of X (SVD), eigenvalue decomposition of XTX (EVD), factor analysis, Eckart–Young theorem, Schmidt–Mirsky theorem, empirical orthogonal function (EOF), empirical eigenfunction decomposition, empirical component analysis, quasi-harmonic modes, spectral analysis, empirical mode analysis, etc., and variations or combinations thereof. PCA typically identifies one or more deviations in the reading density profiles. Deviations identified by PCA sometimes refer to principal components herein. In some implementations, one or more preferences can be removed using suitable methods based on one or more principal components by adjusting the reading density profiles. A reading density profile can be adjusted by adding, subtracting, multiplying, and / or dividing the reading density profile by one or more principal components. In some embodiments, one or more preferences can be removed from the reading density profile by subtracting one or more principal components from the reading density profile. While preferences in the reading density profile are typically identified and / or quantified by PCA of the profile, principal components are typically subtracted from the profile at the reading density level. PCA typically identifies one or more principal components. In some embodiments, PCA identifies the 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, and 10th or more principal components. In some embodiments, 1st, 2nd, 3rd, 4th, 5th, 6th, 7th, 8th, 9th, 10th or more principal components are used to adjust the profile. Typically, principal components are used to adjust the profile in the order they appear in the PCA. For example, when three principal components are subtracted from the reading density profile, the 1st, 2nd, and 3rd principal components are used. Sometimes the preferences identified by the principal components include profile features that are not used to adjust the profile. For example, PCA can identify genetic variations (e.g., aneuploidy, microduplication, microdeletion, deletion, translocation, insertion) and / or sex differences (e.g., see [link to PCA]). Figure 38C Principal components are used as principal components. Therefore, in some embodiments, one or more principal components are not used for profile adjustment. For example, sometimes principal components 1, 2, and 4 are used for profile adjustment when principal component 3 is not used. Principal components can be obtained from PCA using any suitable sample or reference. In some embodiments, principal components are obtained from test samples (e.g., test objects). In some embodiments, principal components are obtained from one or more references (e.g., reference samples, reference sequences, reference groups). For example, such as... Figure 38A As shown in -C, PCA was performed on the median reading density profile obtained from a training group that included a variety of samples. Figure 38A ), thus obtaining the first principal component ( Figure 38B ) and the second principal component ( Figure 38C The identification of principal components. In some embodiments, principal components are obtained from a group of subjects known to have no under-investigation genetic variations. In some embodiments, principal components are obtained from a group of known euploid subjects. Principal components are typically identified using PCA with one or more reading density profiles of a reference (e.g., a training group). One or more principal components obtained from a reference are typically subtracted from the reading density profile of the test subjects (e.g., ...). Figure 39B This provides an overview of the adjustments (e.g.) Figure 39C ).

[0292] In some embodiments, the system includes a PCA statistics module 33. The PCA statistics module can receive and / or retrieve reading density profiles from other suitable modules (e.g., profile generation module 26). PCA is typically performed by the PCA statistics module. The PCA statistics module typically receives, retrieves, and / or stores and processes reading density profiles from a reference group 32, a training group 30, and / or from one or more test subjects 28. The PCA statistics module can generate and / or provide principal components and / or adjust reading density profiles based on one or more principal components. Adjusted reading density profiles (e.g., 40, 38) are typically provided by the PCA statistics module. The PCA statistics module can provide and / or transfer adjusted reading density profiles (e.g., 38, 40) to other suitable modules (e.g., partially weighted module 42, scoring module 46). In some embodiments, the PCA statistics module can provide sex determination 36. Sex determination is sometimes based on PCA and / or based on one or more principal components to determine fetal sex. In some embodiments, the PCA statistics module includes some, all, or modified R codes shown below. R code for calculating principal components typically begins with data cleanup (e.g., subtracting the median, filtering parts, and cleaning up extrema):

[0293] #Clean the data outliers for PCA

[0294] dclean<-(dat-m)[mask,

[0295] for(j in 1:ncol(dclean))

[0296] {

[0297] q<-quantile(dclean[,j],c(.25,.75))

[0298] qmin<-q[1]-4*(q[2]-q[1])

[0299] qmax <- q[2] + 4*(q[2] - q[1])

[0300] dclean[dclean[,j] <qmin,j]<-qmin

[0301] dclean[dclean[,j]>qmax,j]<-qmax

[0302] }

[0303] Then calculate the principal components:

[0304] #Compute principal components

[0305] pc<-prcomp(dclean)$x

[0306] Finally, the PCA adjustment profile for each sample was calculated as follows:

[0307] #Compute residuals

[0308] mm<-model.matrix(~pc[,1:numpc])

[0309] for(j in 1:ncol(dclean))

[0310] dclean[,j]<-dclean[,j]-predict(lm(dclean[,j]~mm))

[0311] Comparative Overview

[0312] In some embodiments, determining the result includes comparison. In some embodiments, a read density profile or a portion thereof is used to provide the result. In some embodiments, determining the result (e.g., determining the presence of genetic variation) includes comparing two or more read density profiles. Comparing read density profiles typically involves comparing read density profiles generated for selected segments of the genome. For example, a test profile is typically compared to a reference profile, where the test and reference profiles determine substantially the same genomic segment (e.g., a reference genome). Comparing read density profiles sometimes includes comparing subgroups of two or more read density profile portions. Subgroups of read density profiles may represent genomic segments (e.g., chromosomes or segments thereof). Read density profiles may include any number of subgroups. Sometimes read density profiles include 2 or more, 3 or more, 4 or more, or 5 or more subgroups. In some embodiments, a read density profile includes portions of two subgroups, where each portion represents an adjacent reference genomic segment. In some implementations, a test profile may be compared to a reference profile, wherein both the test profile and the reference profile include portions of a first subgroup and a portion of a second subgroup, wherein the first and second subgroups represent different segments of the genome. Some portions of a read density profile may include genetic variation, while some other subgroups may sometimes be substantially devoid of genetic variation. Sometimes all portions of a profile (e.g., a test profile) are substantially devoid of genetic variation. Sometimes all portions of a profile (e.g., a test profile) contain genetic variation. In some implementations, a test profile may include a portion of a first subgroup containing genetic variation and a portion of a second subgroup substantially devoid of genetic variation.

[0313] In some embodiments, the methods described herein include comparisons (e.g., comparing test profiles with reference profiles). Two or more data sets, two or more relationships, and / or two or more profiles can be compared using suitable methods. Non-limiting examples of statistical methods suitable for comparing data sets, relationships, and / or profiles include the Behrens-Fisher method, bootstrap method, Fisher's method with combined significant independence tests, Neyman-Pearson test, confirmatory data analysis, probing data analysis, exact test, F-test, Z-test, T-test, calculating and / or comparing measures of uncertainty, null hypothesis, calculating counternulls, chi-square test, comprehensive test, calculating and / or comparing levels of significance (e.g., statistical significance), meta-analysis, multivariate analysis, regression, simple linear regression, reinforced linear regression, and combinations thereof. In some embodiments, comparing two or more data sets, relationships, and / or profiles includes determining and / or comparing measures of uncertainty. As used herein, "measure of uncertainty" refers to a measure of significance (e.g., statistical significance), an error measure, a measure of variance, a confidence measure, and combinations thereof. Uncertainty measures can be values ​​(e.g., thresholds) or ranges of values ​​(e.g., intervals, confidence intervals, Bayesian confidence intervals, threshold ranges). Non-limiting examples of uncertainty measures include p-values, suitable measures of difference (e.g., standard deviation, σ, absolute deviation, arithmetic mean absolute deviation, etc.), suitable measures of error (e.g., standard error, mean squared error, root mean squared error, etc.), suitable measures of variance, suitable standard scores (e.g., standard deviation, cumulative percentage, percentage equivalence, Z-score, T-score, R-score, standard nines (standard nines), percentage in a standard nine, etc.), and combinations thereof. In some implementations, determining the significance level includes determining the uncertainty measure (e.g., p-value). In some implementations, two or more sets of data, relationships, and / or profiles may be analyzed and / or compared using a variety of (e.g., two or more) statistical methods (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, Bagging, neural networks, support vector machine models, random forests, classification tree models, k-nearest neighbors, logistic regression and / or loss smoothing) and / or any suitable mathematical and / or statistical operations (e.g., those described herein).

[0314] In some embodiments, comparing two or more reading density profiles includes determining and / or comparing uncertainty measures for the two or more reading density profiles. Reading density profiles and / or associated uncertainties are sometimes compared to facilitate the explanation of mathematical and / or statistical processing of the data set and / or to provide results. The generation of a reading density profile for a test subject is sometimes compared with reading density profiles generated for one or more references (e.g., reference samples, reference subjects, etc.). In some embodiments, results are provided by comparing the reading density profile of a test subject with a reference reading density profile for a chromosome, a portion thereof, or a segment thereof, wherein the reference reading density profile is obtained from a reference subject group (e.g., a reference) known to have no genetic variation. In some embodiments, results are provided by comparing the reading density profile of a test subject with a reference reading density profile for a chromosome, a portion thereof, or a segment thereof, wherein the reference reading density profile is obtained from a reference subject group known to contain specific genetic variations (e.g., chromosomal aneuploidy, trisomy, microduplication, microdeletion).

[0315] In some implementations, the read density profile of the test subjects is compared to a predetermined value for subjects without genetic variation, and sometimes deviates from the predetermined value at one or more genomic loci (e.g., portions) corresponding to the genomic loci where the genetic variation is located. For example, the read density profile of test subjects (e.g., those with a medical condition associated with the genetic variation or at risk of such a condition) is expected to be significantly different from the read density profile of a selected portion of a reference (e.g., reference sequence, reference subject, reference group) of test subjects containing the investigational genetic variation. The read density profile of the test subjects is generally substantially the same as the read density profile of a selected portion of a reference (e.g., reference sequence, reference subject, reference group) of test subjects not containing the investigational genetic variation. The read density profile is generally compared to a predetermined threshold and / or threshold range (e.g., see...). Figure 40 As used herein, the term "threshold" refers to any number calculated using a qualifying dataset and used as a limitation for diagnosing genetic variations (e.g., copy number variation, aneuploidy, microduplication, microdeletion, chromosomal abnormalities, etc.). In some embodiments, a threshold exceeding the results obtained by the method of the present invention leads to a diagnosis of a genetic variation (e.g., trisomy). In some embodiments, the threshold or threshold range is typically calculated by mathematically and / or statistically processing sequence read data (e.g., from references and / or subjects). The predetermined threshold or threshold range indicating the presence of a genetic variation may vary, but still provides results that can be used to determine the presence of a genetic variation. In some embodiments, a read density profile including standardized read densities and / or standardized counts is generated to facilitate classification and / or provide results. Results may be provided based on a read density profile including standardized counts (e.g., using a read density profile plot).

[0316] In some implementations, the system includes a scoring module 46. The scoring module may accept, retrieve, and / or store reading density profiles (e.g., adjusted, standardized reading density profiles) from other suitable modules (e.g., profile generation module 26, PCA statistics module 33, partial weighting module 42, etc.). The scoring module may accept, retrieve, store, and / or compare two or more reading density profiles (e.g., test profiles, reference profiles, training groups, test subjects). The scoring module may typically provide scores (e.g., graphs, profile statistics, comparisons (e.g., differences between two or more profiles), Z-scores, uncertainty measurements, decision regions, sample decision 50 (e.g., determining the presence of genetic variation), and / or results). The scoring module may provide scores to the end user and / or to other suitable modules (e.g., displays, printers, etc.). In some implementations, the scoring module includes some, all, or modified R codes, including R functions (e.g., high-chr21 counts) for calculating chi-square statistics of specific tests.

[0317] These three parameters are:

[0318] x = Sample reading data (part of sample x)

[0319] m = median of a portion

[0320] y = test vector (e.g., false for all parts, but true for chr21)

[0321] getChisqP<-function(x,m,y)

[0322] {

[0323] ahigh<-apply(x[!y,],2,function(x)sum((x>m[!y])))

[0324] alow<-sum((!y))-ahigh

[0325] bhigh<-apply(x[y,],2,function(x)sum((x>m[y])))

[0326] blow<-sum(y)-bhigh

[0327] p<-sapply(1:length(ahigh),function(i){

[0328] p<-chisq.test(matrix(c(ahigh[i],alow[i],bhigh[i],blow[i]),2))$p.value / 2

[0329] if(ahigh[i] / alow[i]>bhigh[i] / blow[i])p<-max(p,1-p)

[0330] else p <- min(p, 1-p); p})

[0331] return(p)

[0332] Hybridization regression standardization

[0333] In some implementations, hybridization standardization is used. In some implementations, hybridization standardization methods reduce deviations (e.g., GC deviations). In some implementations, hybridization standardization includes (i) analyzing the relationship between bivariates (e.g., counts and GC content) and (ii) selecting and applying a standardization method based on said analysis. In some implementations, hybridization standardization includes (i) regression (e.g., regression analysis) and (ii) selecting and applying a standardization method based on said regression. In some implementations, counts obtained from a first sample (e.g., a first group of samples) are standardized using a different method than counts obtained from other samples (e.g., a second group of samples). In some implementations, counts obtained from a first sample (e.g., a first group of samples) are standardized using a first standardization method, and counts obtained from a second sample (e.g., a second group of samples) are standardized using a second standardization method. For example, in some implementations, the first standardization method includes using linear regression while the second standardization method includes using non-linear regression (e.g., LOESS, GC-LOESS, LOWESS regression, LOESS smoothing).

[0334] In some embodiments, hybridization normalization methods are used to normalize sequence readings mapped to portions of the genome or chromosomes (e.g., counts, mapped counts, mapped readings). In some embodiments, raw counts are normalized, and in some embodiments, adjusted, weighted, filtered, or previously normalized counts are normalized using hybridization normalization methods. In some embodiments, genomic segment level or Z-scores are normalized. In some embodiments, counts mapped to selected portions of the genome or chromosomes are normalized using hybridization normalization methods. A count can refer to a suitable measurement of sequence readings mapped to portions of the genome; non-limiting examples include raw counts (e.g., uncompressed counts), normalized counts (e.g., normalization using PERUN, ChAI, or suitable methods), segmental levels (e.g., average level, arithmetic average level, median level, or etc.), Z-scores, etc., or combinations thereof. A count can be a raw count or processed count of one or more samples (e.g., test samples, samples from pregnant women). In some embodiments, counts are obtained from one or more samples from one or more subjects.

[0335] In some implementations, the standardization method (e.g., the type described above) is selected based on regression (e.g., regression analysis) and / or correlation coefficients. Regression analysis refers to a statistical technique that assesses the relationship between variables (e.g., counts and GC content). In some implementations, regressions are generated based on counts and GC content measurements of various parts of a reference genome. Suitable GC content measurements can be used, and non-limiting examples include measurements of guanine, cytosine, adenine, thymine, purine (GC), or pyrimidine (AT or ATU) content, melting temperatures (Tm) (e.g., denaturation temperatures, annealing temperatures, hybridization temperatures), measurements of free energy, etc., or combinations thereof. Measurements of guanine (G), cytosine (C), adenine (A), thymine (T), purine (GC), or pyrimidine (AT or ATU) content can be expressed as proportions or percentages. In some embodiments, any suitable ratio or percentage is used, non-limiting examples of which include GC / AT, GC / total nucleotides, GC / A, GC / T, AT / total nucleotides, AT / GC, AT / G, AT / C, G / A, C / A, G / T, G / A, G / AT, C / T, etc., or combinations thereof. In some embodiments, GC content is measured as a ratio or percentage of GC to total nucleotide content. In some embodiments, GC content is measured as a ratio or percentage of GC to total nucleotide content with respect to sequence readings mapped to a reference genome portion. In some embodiments, GC content is determined based on and / or from sequence readings mapped to portions of a reference genome, and the sequence readings are obtained from a sample (e.g., a sample obtained from a pregnant woman). In some embodiments, GC content measurement is not based on and / or determined from sequence readings. In some embodiments, GC content measurement is determined from one or more samples obtained from one or more subjects.

[0336] In some implementations, generating regression includes generating regression analysis or correlation analysis. Suitable regression methods may be used, and non-limiting examples include regression analysis (e.g., linear regression analysis), goodness-of-fit analysis, Pearson correlation analysis, tiered correlation, unexplained variance components, Nash–Sutcliffe model validity analysis, regression model validation, proportional reduction loss, root mean square error, etc., or combinations thereof. In some implementations, a regression line is generated. In some implementations, generating regression includes generating linear regression. In some implementations, generating regression includes generating non-linear regression (e.g., LOESS regression, LOWESS regression).

[0337] In some implementations, regression is used to determine whether a correlation exists (e.g., a linear correlation), such as the correlation between a count and a GC content measurement. In some implementations, a regression (e.g., linear regression) is generated and a correlation coefficient is determined. In some implementations, an appropriate correlation coefficient is determined; non-limiting examples include the coefficient of determination, R² value, Pearson correlation coefficient, etc.

[0338] In some implementations, the goodness of fit of a regression (e.g., linear regression in regression analysis) is determined. Goodness of fit is sometimes determined by observation or mathematical analysis. Evaluation sometimes includes determining whether the goodness of fit of a nonlinear or linear regression is greater. In some implementations, the correlation coefficient is a measure of goodness of fit. In some implementations, the goodness of fit of a regression is evaluated based on the correlation coefficient and / or a correlation coefficient cutoff value. In some implementations, goodness of fit evaluation includes comparing the correlation coefficient to a correlation coefficient cutoff value. In some implementations, the evaluation of the goodness of fit of a regression indicates linear regression. For example, in some implementations, the goodness of fit of a linear regression is greater than that of a nonlinear regression, and the goodness of fit evaluation indicates linear regression. In some implementations, the evaluation indicates linear regression, and linear regression is used for standardized counting. In some implementations, the evaluation of the goodness of fit of a regression indicates nonlinear regression. For example, in some implementations, the goodness of fit of a nonlinear regression is greater than that of a linear regression, and the goodness of fit evaluation indicates nonlinear regression. In some implementations, the evaluation indicates nonlinear regression, and nonlinear regression is used for standardized counting.

[0339] In some implementations, when the correlation coefficient is equal to or greater than a correlation coefficient cutoff value, the goodness-of-fit assessment indicates linear regression. In some implementations, when the correlation coefficient is less than a correlation coefficient cutoff value, the goodness-of-fit assessment indicates nonlinear regression. In some implementations, the correlation coefficient cutoff value is predetermined. In some implementations, the correlation coefficient cutoff value is approximately 0.5 or greater, approximately 0.55 or greater, approximately 0.6 or greater, approximately 0.65 or greater, approximately 0.7 or greater, approximately 0.75 or greater, approximately 0.8 or greater, or approximately 0.85 or greater.

[0340] For example, in some embodiments, when the correlation coefficient is equal to or greater than about 0.6, a standardization method including linear regression is used. In some embodiments, when the correlation coefficient is equal to or greater than a correlation coefficient cutoff value of 0.6, sample counts (e.g., counts of each part of the reference genome, counts of each part) are standardized according to linear regression; otherwise, the counts are standardized according to non-linear regression (e.g., when the coefficient is less than the correlation coefficient cutoff value of 0.6). In some embodiments, the standardization process includes generating linear or non-linear regressions for (i) counts and (ii) GC content of each part of a plurality of parts of the reference genome. In some embodiments, when the correlation coefficient is less than the correlation coefficient cutoff value of 0.6, a standardization method including non-linear regression (e.g., LOWESS, LOESS) is used. In some embodiments, when the correlation coefficient (e.g., the correlation coefficient) is less than about 0.7, less than about 0.65, less than about 0.6, less than about 0.55, or less than a correlation coefficient cutoff value of about 0.5, a standardization method including non-linear regression (e.g., LOWESS) is used. For example, in some implementations, when the correlation coefficient is less than the correlation coefficient cutoff of about 0.6, a standardization method including nonlinear regression (e.g., LOWESS, LOESS) is used.

[0341] In some implementations, a specific type of regression (e.g., linear or nonlinear regression) is selected, and the counts are standardized by subtracting the regression from the counts after the regression is generated. In some implementations, subtracting the regression from the counts provides standardized counts with reduced deviations (e.g., GC deviations). In some implementations, linear regression is subtracted from the counts. In some implementations, nonlinear regression (e.g., LOESS, GC-LOESS, LOWESS regression) is subtracted from the counts. Any suitable method can be used to subtract the regression line from the counts. For example, if count x originates from a portion i (e.g., portion i) including a GC content of 0.5 and the regression line determines count y at a GC content of 0.5, then xy = the standardized count of portion i. In some implementations, standardized counts are subtracted before and / or after the regression. In some implementations, counts standardized by hybridization normalization methods are used to generate genomic segment levels, Z-scores, genomic or segmental levels, and / or profiles. In some implementations, counts standardized by hybridization normalization methods are analyzed by the methods described herein to determine the presence of genetic variation (e.g., in the fetus).

[0342] In some embodiments, the hybridization normalization method includes filtering or weighting one or more portions before or after normalization. Portions can be filtered using suitable methods described herein, including methods for filtering portions (e.g., portions of a reference genome). In some embodiments, portions (e.g., portions of a reference genome) are filtered before applying the hybridization normalization method. In some embodiments, sequencing read counts mapped only to selected portions (e.g., portions selected based on count variability) are normalized by hybridization normalization. In some embodiments, sequencing read counts mapped to filtered portions of the reference genome (e.g., portions filtered based on count variability) are removed before using the hybridization normalization method. In some embodiments, the hybridization normalization method includes selecting or filtering portions (e.g., portions of a reference genome) according to suitable methods (e.g., methods described herein). In some embodiments, the hybridization normalization method includes selecting or filtering portions (e.g., portions of a reference genome) based on uncertainties in counts mapped to various portions of multiple test samples. In some embodiments, the hybridization normalization method includes selecting or filtering portions (e.g., portions of a reference genome) based on count variability. In some implementations, hybridization normalization methods include selecting or filtering portions (e.g., referencing portions of the genome) based on GC content, repetitive elements, repetitive sequences, introns, exons, etc., or combinations thereof.

[0343] For example, in some embodiments, multiple samples from multiple pregnant women are analyzed, and partial subgroups (e.g., portions of a reference genome) are selected based on count variability. In some embodiments, linear regression is used to determine the correlation coefficient between (i) counts and (ii) GC content of each selected portion of the samples obtained from the pregnant women. In some embodiments, a correlation coefficient greater than a predetermined correlation cutoff (e.g., about 0.6) is determined, a goodness-of-fit assessment instructs linear regression, and the counts are standardized by subtracting the linear regression from the counts. In some embodiments, a correlation coefficient less than a predetermined correlation cutoff (e.g., about 0.6) is determined, a goodness-of-fit assessment instructs nonlinear regression, a LOESS regression is generated, and the counts are standardized by subtracting the LOESS regression from the counts.

[0344] Overview

[0345] In some implementations, the processing steps may include generating one or more profiles (e.g., profile plots) from various data sets or their derivatives (e.g., the results of one or more mathematical and / or statistical data processing steps known in the art and / or described herein).

[0346] The term "profile" in this document refers to the result of mathematical and / or statistical operations on data that facilitate the identification of patterns and / or correlations in large datasets. A profile typically includes values ​​obtained from one or more operations on data or a group of data based on one or more criteria. A profile typically includes multiple data points. Any suitable number of data points can be included in a profile, depending on the nature and / or complexity of the data group. In some implementations, a profile may include 2 or more data points, 3 or more data points, 5 or more data points, 10 or more data points, 24 or more data points, 25 or more data points, 50 or more data points, 100 or more data points, 500 or more data points, 1000 or more data points, 5000 or more data points, 10,000 or more data points, or 100,000 or more data points.

[0347] In some implementations, the profile is a representation of the entire data set, and in other implementations, the profile is a representation of a portion or subgroup of the data set. That is, a profile sometimes includes data points representing or generated from data that has not been filtered to remove any data, and sometimes a profile includes data points representing or generated from data that has been filtered to remove unwanted data. In some implementations, the data points in the profile represent partial data operation results. In some implementations, the data points in the profile include partial group data operation results. In some implementations, partial groups may be adjacent to each other, and in some implementations, partial groups may come from different parts of a chromosome or genome.

[0348] Data points derived from a profile of a data set can represent any suitable data classification. Non-limiting examples of data grouping to generate profile data point categories include: size-based portions, sequence feature-based portions (e.g., GC content, AT content, chromosomal location (e.g., short arm, long arm, centromere, telomere), etc.), expression levels, chromosomes, etc., or combinations thereof. In some implementations, a profile can be generated from data points derived from other profiles (e.g., normalized data profiles re-normalized to different normalization values ​​to generate renormalized data profiles). In some implementations, profiles generated from data points derived from other profiles reduce the number of data points and / or the complexity of the data set. Reducing the number of data points and / or the complexity of the data set generally facilitates data interpretation and / or results delivery.

[0349] A profile (e.g., a genome profile, chromosome profile, chromosomal segment profile) is typically a collection of normalized or non-normalized counts of two or more parts. A profile typically includes at least one level (e.g., a genome segment level) and typically includes two or more levels (e.g., a profile often has multiple levels). Levels are typically used for groups of parts having approximately the same count or normalized count. Levels are described in detail herein. In some embodiments, a profile includes one or more parts that may be processed or transformed by weighting, removal, filtering, normalization, adjustment, averaging (to obtain a mean), addition, subtraction, or any combination thereof. A profile typically includes normalized counts mapped to parts defining two or more levels, wherein the counts are further normalized according to one of the levels using a suitable method. Typically, profile counts (e.g., profile levels) are associated with uncertain values.

[0350] Profiles including one or more levels are sometimes filled (e.g., well-filled). Filling (e.g., well-filled) refers to the process of identifying and adjusting the levels in a profile that originate from maternal microdeletions or maternal duplications (e.g., copy number variations). In some embodiments, levels originating from fetal microduplications or fetal microdeletions are filled. In some embodiments, the overall level of microduplications or microdeletions in a profile may be artificially increased or decreased, resulting in false positives or false negatives in the determination of chromosomal aneuploidy (e.g., trisomy). In some embodiments, levels in a profile that originate from microduplications and / or deletions are identified and adjusted (e.g., filled and / or removed) by a process sometimes referred to as filling or well-filled. In some embodiments, a profile includes one or more first levels that are distinctly different from a second level within the profile, each of which includes maternal copy number variation, fetal copy number variation, or both maternal and fetal copy number variation, and one or more of the first levels are adjusted.

[0351] A profile including one or more levels may include a first level and a second level. In some embodiments, the first level is different from (e.g., significantly different from) the second level. In some embodiments, the first level includes a first group portion, the second level includes a second group portion, and the first group portion is not a subgroup of the second group portion. In some embodiments, the first group portion is different from the second group portion, thereby defining the first and second levels. In some embodiments, a profile may have multiple first levels that are different from (e.g., significantly different, for example, having significantly different values) the second level within the profile. In some embodiments, the profile includes one or more first levels that are significantly different from the second level within the profile, and said one or more first levels are adjusted. In some embodiments, the profile includes one or more first levels that are significantly different from the second level within the profile, each of said one or more first levels including maternal copy number variation, fetal copy number variation, or maternal copy number variation and fetal copy number variation, and said one or more first levels are adjusted. In some embodiments, the first level in the profile is removed from the profile or adjusted (e.g., padded). A profile may include multiple levels, said multiple levels including one or more first levels that are significantly different from one or more second levels, typically the dominant level in the profile is the second level, wherein the second levels are approximately equal to each other. In some implementations, levels greater than 50%, 60%, 70%, 80%, 90%, or 95% in the overview are considered second levels.

[0352] A profile is sometimes displayed as a graph. For example, one or more levels representing partial counts (e.g., standardized counts) can be plotted and visualized. Non-limiting examples of profile graphs that can be generated include raw counts (e.g., raw count profile or raw profile), standardized counts, partial-weighted, Z-scores, p-values, area ratios and fit ploidy, the ratio between the median level and the fitted and measured fetal score, principal components, etc., or combinations thereof. In some implementations, profile graphs allow for observation of manipulated data. In some implementations, profile graphs can be used to provide results (e.g., area ratios and fit ploidy, the ratio between the median level and the fitted and measured fetal score, principal components). As used herein, the term "raw count profile graph" or "raw profile graph" refers to a graph of counts in different parts of a region normalized to the total regional count (e.g., genome, part, chromosome, chromosomal part of a reference genome, or chromosomal segment). In some implementations, a static window procedure can be used to generate the profile, and in some implementations, a sliding window procedure can be used to generate the profile.

[0353] Profiles generated for test subjects are sometimes compared with profiles generated for one or more reference subjects to facilitate the explanation of mathematical and / or statistical operations on the data set and / or to provide results. In some implementations, profiles are generated based on one or more initial hypotheses (e.g., maternal nucleic acid contribution (e.g., total maternal score), fetal nucleic acid contribution (e.g., fetal score), reference sample ploidy, etc., or combinations thereof). In some implementations, test profiles are typically centered on a predetermined value representing the absence of genetic variation, and deviated from the predetermined value in the corresponding area of ​​the genomic location in the test subject where genetic variation is typically located (if the test subject has genetic variation). In test subjects with conditions associated with or at risk of such genetic variation, the numerical values ​​of the selected portion are expected to differ significantly from the predetermined values ​​for unaffected genomic locations. Based on initial hypotheses (e.g., fixed ploidy or optimal ploidy, fixed fetal score or optimal fetal score, or combinations thereof), predetermined thresholds or cutoff values ​​or threshold ranges indicating the presence of genetic variation may vary, but they still provide results that can be used to determine the presence of genetic variation. In some implementations, profiles indicate and / or represent phenotypes.

[0354] As a non-limiting example, standardized sample and / or reference count profiles can be obtained from raw sequence reading data through the following steps:

[0355] (a) Calculate the reference median count of the selected chromosome, its portion, or segment from a reference group known to contain no genetic variation.

[0356] (b) Remove uninformative portions from the original counts of the reference sample (e.g., filter);

[0357] (c) Normalize the reference counts of all remaining reference genome portions to the total residual count of the selected chromosome or selected genomic location of the reference sample (e.g., summation of residual counts after removing uninformative portions of the reference genome) to generate a normalized reference object profile.

[0358] (d) Remove the corresponding portion from the test sample; and

[0359] (e) The remaining test subject counts at one or more selected genomic locations are normalized to the sum of the residual reference median counts of chromosomes or chromosomes containing the selected genomic locations, thereby generating a normalized test subject profile. In some embodiments, additional normalization steps involving the entire genome (reduced by the filtering portion in (b)) may be included between (c) and (d).

[0360] A dataset profile can be generated through one or more processing methods on sequence read data mapped by counting. Some implementations include the following: Mapping sequence reads and determining the number of sequence tags mapped to each genomic segment (e.g., counting). Generating a raw count profile from the counted mapped sequence reads. In some implementations, results are provided by comparing the raw count profile of the test subject with a reference median count profile of chromosomes, portions, or segments of a reference subject group known to be free of genetic variation.

[0361] In some implementations, the sequence reading data may optionally be filtered to remove noisy or uninformative portions. After filtering, the remaining counts are typically summed to generate a filtered set of data. In some implementations, a filtered count profile is generated from the filtered set of data.

[0362] After counting and optional filtering, sequence read data can be standardized to generate levels or profiles. Data sets can be standardized by standardizing one or more selected portions to a suitable standardization reference value. In some embodiments, the standardization reference value represents the total count of chromosomes from a selected portion. In some embodiments, the standardization reference value represents one or more corresponding portions of chromosomes in a reference data set prepared from a reference group known to be free of genetic variation. In some embodiments, the standardization reference value represents one or more corresponding portions of chromosomes in a test subject data set prepared from test subjects analyzed for the presence or absence of genetic variation. In some embodiments, the standardization process uses a static window method, and in some embodiments, the standardization process uses a moving or sliding window method. In some embodiments, generating a profile including standardized counts facilitates classification and / or provides results. Results can be provided based on a profile plot including standardized counts (e.g., using this profile plot).

[0363] level

[0364] In some implementations, values ​​(e.g., numerical values, quantitative values) are assigned to levels. Counts can be determined by suitable methods, operations, or mathematical processes (e.g., processed levels). Levels are typically or derived from counts of a subset of a group (e.g., standardized counts). In some implementations, the level of a subset is substantially equal to the total number of counts mapped to the subset (e.g., counts, standardized counts). Levels are typically determined from counts that have been processed, transformed, or manipulated by suitable methods, operations, or mathematical processes known in the art. In some implementations, levels are derived from processed counts, non-limiting examples of which include weighted, removed, filtered, standardized, adjusted, averaged, derived arithmetic mean (e.g., arithmetic average level), added, subtracted, transformed counts, or combinations thereof. In some implementations, levels include standardized counts (e.g., partially standardized counts). Levels can be standardized by suitable processes, non-limiting examples of which include component-by-component standardization, GC content standardization, linear and nonlinear least squares regression, GC LOESS, LOWESS, PERUN, ChAI, RM, GCRM, cQn, etc., and / or combinations thereof. Levels may include standardized counts or relative quantities of counts. In some implementations, a level is used for averaged counts or standardized counts of two or more portions, and the level refers to an average level. In some implementations, a level is used for a group of counts or portions of an arithmetic mean of standardized counts, referred to as the arithmetic average level. In some implementations, the level is derived from portions of counts that include both raw and / or filtered counts. In some implementations, the level is based on the raw counts. In some implementations, the level is associated with an uncertainty (e.g., standard deviation, MAD). In some implementations, the level is represented by a Z-score or p-value. The level of one or more portions herein is synonymous with "genomic segment level".

[0365] Standardized or unstandardized counts of two or more levels (e.g., two or more levels in a profile) can sometimes be standardized based on the levels using mathematical operations (e.g., addition, multiplication, averaging, standardization, etc., or combinations thereof). For example, standardized or unstandardized counts of two or more levels can be standardized based on one, some, or all of the levels in the profile. In some embodiments, standardized or unstandardized counts of all levels in the profile are standardized based on one level in the profile. In some embodiments, standardized or unstandardized counts of a first level in the profile are standardized based on standardized or unstandardized counts of a second level in the profile.

[0366] Non-limiting examples of levels (e.g., first level, second level) are group levels that include portions of processed counts, group levels that include portions of the arithmetic mean, median, or average of counts, group levels that include portions of standardized counts, and any combination thereof. In some embodiments, the first and second levels in the overview are derived from counts mapped to portions of the same chromosome. In some embodiments, the first and second levels in the overview are derived from counts mapped to portions of different chromosomes.

[0367] In some implementations, the level is determined from normalized or unnormalized counts mapped to one or more portions. In some implementations, the level is determined from normalized or unnormalized counts mapped to two or more portions, wherein the normalized counts of each portion are generally approximately the same. For a given level, the counts (e.g., normalized counts) within a group of portions may differ. For a given level, there may be one or more portions within a group that have counts significantly different from those of other portions of the group (e.g., peaks and / or sloping). Any suitable number of normalized or unnormalized counts associated with any suitable number of portions can define the level.

[0368] In some embodiments, one or more levels can be determined from normalized or non-normalized counts of all or some portions of the genome. Typically, levels can be determined from all or some normalized or non-normalized counts of chromosomes or segments thereof. In some embodiments, levels are determined from two or more counts derived from two or more portions (e.g., groups of portions). In some embodiments, levels are determined from two or more counts (e.g., counts from two or more portions). In some embodiments, levels are determined from 2 to about 100,000 portions. In some embodiments, levels are determined from 2 to about 50,000, 2 to about 40,000, 2 to about 30,000, 2 to about 20,000, 2 to about 10,000, 2 to about 5000, 2 to about 2500, 2 to about 1250, 2 to about 1000, 2 to about 500, 2 to about 250, 2 to about 100, or 2 to about 60 portions. In some embodiments, levels are determined from about 10 to about 50 portions. In some embodiments, counts of about 20 to about 40 or more portions determine the level. In some embodiments, the level includes counts from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60 or more portions. In some embodiments, the level corresponds to a group of portions (e.g., a group of portions of a reference genome, a group of chromosomal portions, or a group of chromosomal segment portions).

[0369] In some implementations, the level is determined by the normalized or non-normalized count of adjacent portions. In some implementations, adjacent portions (e.g., groups of portions) represent adjacent segments of the genome or adjacent segments of chromosomes or genes.

[0370] For example, when merging portions tail-to-tail, two or more adjacent portions can represent a set of DNA sequences that are longer than each portion.

[0371] For example, two or more adjacent portions may represent an entire genome, chromosome, gene, intron, exon, or segment thereof. In some implementations, the level is determined from a set (e.g., a group) of adjacent and / or non-adjacent portions.

[0372] Different levels

[0373] In some embodiments, the standardized count profile includes a level (e.g., a first level) that is significantly different from other levels (e.g., a second level) within the profile. The first level may be higher or lower than the second level. In some embodiments, the first level is used for a group comprising portions of one or more readings that include copy number variation (e.g., maternal copy number variation, fetal copy number variation, or both maternal and fetal copy number variation), and the second level is used for a group comprising portions of readings that are substantially without copy number variation. In some embodiments, significant difference refers to an observable difference. In some embodiments, significant difference refers to statistical difference or statistically significant difference. Statistically significant difference is sometimes a statistical estimate of an observable difference. Statistically significant difference can be estimated using methods suitable in the art. Any suitable threshold or range can be used to determine two significantly different levels. In some embodiments, a difference of about 0.01% or more (e.g., 0.01% of one or the other level value) between two levels (e.g., the arithmetic mean) is considered significantly different. In some embodiments, a difference of about 0.1% or more between two levels (e.g., the arithmetic mean) is considered significantly different. In some embodiments, a difference of about 0.5% or more between two levels (e.g., the arithmetic mean) is considered significantly different. In some implementations, a difference of approximately 0.5, 0.75, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or greater than 10% between two levels is considered significantly different. In some implementations, two levels (e.g., the arithmetic mean) are significantly different and there is no overlap between the levels and / or no overlap within the defined range of the uncertainty calculated for one or both levels. In some implementations, the uncertainty is a standard deviation, expressed as... In some implementations, the two levels (e.g., the arithmetic mean) are significantly different, differing by about one or more times the uncertainty value (e.g. In some implementations, the two levels (e.g., the arithmetic mean) are significantly different, differing by about two or more times the uncertainty value (e.g., ...). The confidence level can be approximately 3 or more, approximately 4 or more, approximately 5 or more, approximately 6 or more, approximately 7 or more, approximately 8 or more, approximately 9 or more, or approximately 10 or more times the uncertainty. In some embodiments, the confidence level is significantly different when the difference between two levels (e.g., the arithmetic mean) is approximately 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, or 4.0 times or more of the uncertainty. In some embodiments, the confidence level increases with the increase in the difference between the two levels. In some embodiments, the confidence level decreases with the decrease in the difference between the two levels and / or the increase in the uncertainty. For example, sometimes the confidence level increases proportionally to the difference between the level and the standard deviation (e.g., MAD).

[0374] One or more prediction algorithms can be used to determine significance or to give meaning to the collected detection data under variable conditions. Their weights can be independent or interdependent. The term "variable" used in this paper refers to a factor, quantity, or function in the algorithm that has one or more values.

[0375] In some implementations, the first group typically includes portions that are different from (e.g., do not overlap with) the second group. For example, sometimes a first level of normalized counts is significantly different from a second level of normalized counts in a profile, and the first level pertains to a first group, the second level to a second group, and the portions do not overlap between the first and second groups. In some implementations, the first group is not a subgroup of the second group, thereby determining the first and second levels respectively. In some implementations, the first group differs from and / or varies with the second group, thereby determining the first and second levels respectively.

[0376] In some embodiments, the first group portion is a subgroup of the second group portion in the profile. For example, sometimes a second level of standardized counts in the second group portion of the profile includes a first level of standardized counts in the first group portion of the profile, and the first group portion is a subgroup of the second group portion of the profile. In some embodiments, the mean, arithmetic mean, or median level is derived from the second level, wherein the second level includes the first level. In some embodiments, the second level includes a second group portion representing the entire chromosome, and the first level includes the first group portion, wherein the first group is a subgroup of the second group portion, and the first level represents maternal copy number variation, fetal copy number variation, or maternal copy number variation and fetal copy number variation present in the chromosome.

[0377] In some embodiments, the value of the second level is closer to the arithmetic mean, average, or median of the count profile of a chromosome or its segment than the first level. In some embodiments, the second level is the arithmetic average level of the levels of a chromosome, a portion of a chromosome, or a segment of a chromosome. In some embodiments, the first level is significantly different from the dominant level (e.g., the second level) representing a chromosome or its segment. A profile may include multiple first levels that are significantly different from the second level, and each first level may be independently higher or lower than the second level. In some embodiments, the first and second levels originate from the same chromosome, and the first level is higher or lower than the second level, where the second level is the dominant level of the chromosome. In some embodiments, the first and second levels originate from the same chromosome, the first level indicates copy number variations (e.g., maternal and / or fetal copy number variations, deletions, insertions, duplications), and the second level is the arithmetic average level or dominant level of a portion of a chromosome or its segment.

[0378] In some implementations, the readings in the second set of portions of the second level substantially exclude genetic variations (e.g., copy number variations, maternal and / or fetal copy number variations). Typically, the second set of portions of the second level includes some variability (e.g., level variability, portion count variability). In some implementations, one or more portions of a set of portions associated with a level substantially free of copy number variations include readings of on...

Claims

1. A method for determining whether microreplications or microdeletions exist in a test sample, comprising: (a) Obtaining a count of nucleic acid sequence reads mapped to a portion of the reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acids from heterogeneous test samples; (b) Counting the portion or the subgroups of the portion to provide one or more discrete segments; (c) Identify candidate segments among the one or more discrete segments; (d) Calculate the logarithmic concession ratio (LOR) of the candidate segment, where LOR is the logarithm of the quotient of (i) and (ii) below, where (i) is the first product between (1) the conditional probability of having micro-replication or micro-deletion and (2) the prior probability of having said micro-replication or micro-deletion, and (ii) is the second product between (1) the conditional probability of not having micro-replication or micro-deletion and (2) the prior probability of not having said micro-replication or micro-deletion; as well as (d) Determine whether micro-replications or micro-deletions exist in the test sample based on the LOR determined for the candidate segment.

2. The method of claim 1, further comprising standardizing the count obtained in (a).

3. The method of claim 1 or 2, further comprising generating a quantitative representation of the candidate segments, wherein the quantitative representation is a count representation of the candidate segments.

4. The method of claim 3, wherein the quantification is a z-score quantification represented by the count of the candidate segments.

5. The method of claim 3, wherein the z-score is the result of subtracting the median of the test sample count representation from the euploid count representation with respect to the candidate segment, divided by the MAD of the euploid count representation, wherein: (i) The test sample count is the proportion of the total count of the test samples divided by the total autosome count, and (ii) the euploid median count is the median proportion of the total count of euploid samples divided by the total autosome count.

6. The method of any one of claims 1-5, wherein the method includes generating a quantitative representation of the chromosome on which the candidate segment is located.

7. The method of claim 6, wherein the quantification of chromosome representation is z-fraction quantification.

8. The method of claim 7, wherein the z-score is the result of subtracting the median of the test sample count representation from the euploid count representation with respect to the chromosome, divided by the MAD of the euploid count representation, wherein: (i) The test sample count is the proportion of the total count of the candidate segment in the chromosome in which the test sample is located to the total autosome count, and (ii) the median euploid count is the median proportion of the total count of the candidate segment in the chromosome in which the euploid sample is located to the total autosome count.

9. The method of any one of claims 3-8, wherein the conditional probability of not having micro-replication or micro-deletion is the intersection of the z-score of the count representation of the candidate segment determined with respect to the test sample and the distribution of the z-score of the count representation of the candidate segment in euploids.

10. The method of any one of claims 1-9, wherein the prior probability of having the micro-replication or micro-deletion and the prior probability of not having the micro-replication or micro-deletion are determined from a plurality of samples excluding the test subject.

11. The method of any one of claims 1-10, the method comprising determining whether the LOR is greater than 0 or less than 0.

12. The method of claim 11, wherein the method comprises, with respect to the test sample, determining the presence of microdeletions or microduplications based at least in part on: (i) a z-score quantitatively expressed as a count of the candidate segment being greater than or equal to 3.95, and (ii) an LOR greater than 0.

13. The method of claim 11, wherein the method comprises, with respect to the test sample, determining the absence of microdeletions or microduplications based at least in part on: (i) a z-score quantitatively expressed as a count of the candidate segment being less than 3.95, and / or (ii) an LOR being less than 0.

14. The method of any one of claims 1-13, wherein the microdeletion is associated with DiGeorge syndrome.

15. The method of any one of claims 3-14, wherein the count representation of the candidate segment is a normalized count representation.

16. The method of any one of claims 1-15, wherein one or more or all of (a), (b), (c) and (d) are performed by a microprocessor in the system.

17. The method of any one of claims 1-16, wherein one or more or all of (a), (b), (c) and (d) are performed by a computer.

18. The method of any one of claims 1-17, wherein one or more or all of (a), (b), (c) and (d) are performed in conjunction with a memory.

19. The method of any one of claims 1-18, wherein the method comprises, prior to (a), sequencing of nucleic acids obtained from the test sample to provide nucleic acid sequence readings.

20. The method of any one of claims 1-19, wherein the method includes, prior to (a), a portion of mapping the nucleic acid sequence reading to a reference genome.

21. The method of any one of claims 1-20, wherein the heterogeneous test sample comprises cancer nucleic acids and non-cancer nucleic acids.

22. The method of any one of claims 1-20, wherein the heterogeneous test sample comprises fetal-derived nucleic acids and maternal-derived nucleic acids.

Citation Information

Patent Citations

  • Fragmentation-based methods and systems for sequence variation detection and discovery

    US20050112590A1

  • Method and compositions for detection and enumeration of genetic variations

    US20070065823A1

  • Restriction endonuclease enhanced polymorphic sequence detection

    US20090317818A1

  • Processes and compositions for methylation-based enrichment of fetal nucleic acid from a maternal sample useful for non invasive prenatal diagnoses

    US20100105049A1

  • Simultaneous determination of aneuploidy and fetal fraction

    US20110224087A1