Methods and processes for non-invasive assessment of genetic variation

By combining sequence and length separation and analysis methods, and using microprocessors for weighted factor fitting, the efficiency and cost issues in fetal nucleic acid assessment are solved, and the accuracy of non-invasive prenatal diagnosis is improved.

CN114724627BActive Publication Date: 2026-05-08SEQUENOM INC
View PDF 20 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SEQUENOM INC
Filing Date
2014-06-20
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency, high cost, and data redundancy when assessing fetal nucleic acid scores, especially in non-invasive prenatal diagnosis where it is difficult to accurately distinguish between maternal and fetal nucleic acids.

Method used

A sequence- and length-based separation method combined with analytical techniques was employed, and a microprocessor was used for weighted factor fitting to assess fetal nucleic acid fraction, reducing reliance on nucleotide sequence determination and improving assessment accuracy.

Benefits of technology

This technology enables efficient and low-cost differentiation of fetal nucleic acids from maternal blood samples, improving the accuracy and efficiency of non-invasive prenatal diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003603576890000981
    Figure BDA0003603576890000981
  • Figure BDA0003603576890001451
    Figure BDA0003603576890001451
  • Figure BDA0003603576890001461
    Figure BDA0003603576890001461
Patent Text Reader

Abstract

Provided herein are methods and processes for non-invasive assessment of genetic variations and operations, systems, apparatuses, and devices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related patent applications

[0002] This patent application claims the benefit of U.S. Provisional Patent Application 61 / 838,048, filed June 21, 2013, entitled “METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETICVARIATIONS,” inventors Sung K. Kim et al., and case number SEQ-6071-PV. The entire contents of the aforementioned patent application are incorporated herein by reference, including its text, tables, and figures.

[0003] field

[0004] The technical aspects of this article relate to methods, procedures, and devices for non-invasive assessment of genetic variation. background

[0005] The genetic information of living organisms (such as animals, plants, and microorganisms) and other forms of replicating genetic information (such as viruses) is encoded as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a series of nucleotides or modified nucleotides that represent the primary structure of a chemical or presumed nucleic acid. The complete human genome contains approximately 30,000 genes located on twenty-four (24) chromosomes (see The Human Genome, T. Strachan, BIOS Science Press, 1992). Each gene encodes a specific protein, which, after being expressed through transcription and translation, performs a specific biochemical function in living cells.

[0006] Many medical conditions are caused by one or more genetic variations. Some genetic variations cause medical conditions, including, for example, hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease, and cystic fibrosis (CF) (Human Genome Mutations, DN Cooper and M. Krawczak, BIOS Publishing, 1993). These genetic diseases may be caused by the addition, substitution, or deletion of a single nucleotide in the DNA of a specific gene. Some birth defects are caused by chromosomal abnormalities (also known as aneuploidy), such as trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), X monosomy (Turner syndrome), and certain sex chromosome aneuploidies such as Klinefelter syndrome (XXY). Other genetic variations are fetal sex, which is usually determined based on the X and Y sex chromosomes. Some genetic variations can predispose an individual to or cause any of many diseases, such as diabetes, arteriosclerosis, obesity, various autoimmune diseases, and cancers (such as colorectal cancer, breast cancer, ovarian cancer, and lung cancer).

[0007] The identification of one or more genetic variations or changes can aid in the diagnosis of specific medical conditions or in determining the underlying causes of those conditions. Identifying genetic variations can also assist in medical decision-making and / or the use of beneficial medical treatments. In some implementations, the identification of one or more genetic variations or changes involves the analysis of cell-free DNA. Cell-free DNA (CF-DNA) consists of DNA fragments derived from cell death and peripheral blood circulation. High concentrations of CF-DNA can indicate certain clinical conditions such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infection, and other diseases. Furthermore, cell-free fetal DNA (CFF-DNA) can be detected in maternal bloodstream and is used in a variety of non-invasive prenatal diagnostic procedures.

[0008] The presence of fetal nucleic acids in maternal plasma enables non-invasive prenatal diagnosis through analysis of maternal blood samples. For example, quantitative abnormalities in fetal DNA in maternal plasma can be associated with a variety of pregnancy-related conditions, including preeclampsia, premature birth, antepartum hemorrhage, invasive placenta formation, fetal Down syndrome, and other fetal chromosomal aneuploidy. Therefore, analyzing fetal nucleic acids in maternal plasma can be a useful mechanism for monitoring maternal and infant health.

[0009] Overview

[0010] In some respects, this document provides a method for assessing the fraction of fetal nucleic acids in a test sample from a pregnant woman, the method comprising: (a) obtaining a count of sequence reads mapped to a portion of a reference genome, wherein the sequence reads are reads of circulating cell-free nucleic acids from the test sample from the pregnant woman; (b) using a microprocessor to weight (i) the count of sequence reads mapped to each portion, or (ii) other portion-specific parameters, with a portion-specific fetal nucleic acid fraction by weighting factors independently associated with each portion, thereby providing a portion-specific fetal fraction estimate based on said weighting factors, wherein each weighting factor has been determined by fit correlations between: (i) the fetal nucleic acid fraction of each sample in a plurality of samples and (ii) the count of sequence reads mapped to each portion or other portion-specific parameters of the plurality of samples; and (c) assessing the fetal nucleic acid fraction of the test sample based on said portion-specific fetal fraction estimate.

[0011] This article also provides a method for assessing the fraction of fetal nucleic acids in a test sample from a pregnant woman, the method comprising (a) obtaining a count of sequence readings mapped to a portion of a reference genome, wherein the sequence readings are readings of circulating cell-free nucleic acids from the test sample from the pregnant woman; (b)(i) using a microprocessor to adjust the count of the sequence readings mapped to each portion according to weighting factors independently assigned to each portion, thereby providing an adjusted count for said portion; or (b)(ii) using a microprocessor to select a subset of the portion, thereby providing a subset of counts, wherein the adjustment in (b)(i) or the selection in (b)(ii) is based on a portion in which the amount of fetal nucleic acid readings mapped to said portion increases; and (c) assessing the fetal nucleic acid fraction of the test sample based on the adjusted count or subset of counts.

[0012] This article also provides a method for improving the accuracy of assessing the fraction of fetal nucleic acids in test samples from pregnant women, the method comprising obtaining a count of sequence readings mapped to a portion of a reference genome, the sequence readings being readings of circulating cell-free nucleic acids from a test sample from a pregnant woman, wherein at least a subset of the obtained counts originates from a region of the genome that facilitates obtaining a larger number of fetal nucleic acid counts relative to the total count of other regions of the genome than the total fetal nucleic acid count relative to the total count of other regions of the genome.

[0013] This document also provides a system, apparatus, or device comprising one or more microprocessors and a memory, wherein the memory contains instructions executable by the one or more microprocessors, and wherein the instructions executable by the one or more microprocessors are configured to (a) access nucleotide sequence readings mapped to portions of a reference genome, wherein the sequence readings are readings of circulating cell-free nucleic acids from test samples from pregnant women; (b) weight (i) a count of sequence readings mapped to each portion, or (ii) other portion-specific parameters, based on weighting factors independently associated with each portion, thereby providing a portion-specific fetal score estimate based on the weighting factors, wherein each weighting factor has been determined by the fit correlation between each portion and (i) the fetal nucleic acid score of each sample in a plurality of samples, and (ii) the count of sequence readings mapped to each portion of a plurality of samples, or other portion-specific parameters; and (c) evaluate the fetal nucleic acid score of the test sample based on the portion-specific fetal score estimate.

[0014] This document also provides an apparatus comprising one or more microprocessors and a memory, wherein the memory contains instructions executable by the one or more microprocessors, and wherein the memory contains nucleotide sequence readings mapped to portions of a reference genome, wherein the sequence readings are readings of circulating cell-free nucleic acids from test samples from pregnant women, and wherein the instructions executable by the one or more microprocessors are used to (a) employ the microprocessors, based on weighting factors independently associated with each portion, to weight (i) a count of sequence readings mapped to each portion, or (ii) other portion-specific parameters, with a portion-specific fetal nucleic acid score, thereby providing a portion-specific fetal score estimate based on said weighting factors, wherein each of said weighting factors has been determined by the fit correlation between each portion and (i) the fetal nucleic acid scores of each of the plurality of samples, and (ii) the counts of sequence readings mapped to each portion of the plurality of samples, or other portion-specific parameters, and (b) to assess the fetal nucleic acid score of the test sample based on said portion-specific fetal score estimate.

[0015] This document also provides a non-transitory computer-readable storage medium having stored thereon an executable program, wherein the program instructs a microprocessor to perform the following operations: (a) accessing nucleotide sequence readings mapped to a portion of a reference genome, wherein the sequence readings are readings of circulating cell-free nucleic acids from test samples from pregnant women; (b) using the microprocessor, based on weighting factors independently associated with each portion, weighting (i) the counts of sequence readings mapped to each portion, or (ii) other portion-specific parameters, with a portion-specific fetal nucleic acid score, thereby providing a portion-specific fetal score estimate based on said weighting factors, wherein each of said weighting factors has been determined by the fit correlation between each portion and (i) the fetal nucleic acid scores of each of the plurality of samples, and (ii) the counts of sequence readings mapped to each portion of the plurality of samples, or other portion-specific parameters; and (c) assessing the fetal nucleic acid score of the test sample based on said portion-specific fetal score estimate.

[0016] Certain technical aspects are further described in the following description, embodiments, claims and drawings. Brief description of the attached figures

[0017] The accompanying drawings illustrate embodiments of the present technology but are not limiting. For clarity and convenience, the drawings are not made to scale and in some cases, various aspects may be exaggerated or enlarged to aid in understanding the specific embodiments.

[0018] Figure 1 Paired comparisons of the FRS (left y-axis, top histogram) and the number of exons per 50 kb bin for chromosome 13 (right y-axis, bottom histogram) are shown. Partially shown on the bottom horizontal X-axis.

[0019] Figure 2 Paired comparisons of FRS (left y-axis, top histogram) and GC content per 50 kb bin (right y-axis, bottom histogram) for chromosome 13 are shown. Partially shown on the bottom horizontal X-axis.

[0020] Figure 3 This shows a paired comparison of the number of exons per 50kb region of chromosome 13 (left y-axis, top histogram) and the GC content per 50kb region (right y-axis, bottom histogram). The regions are shown on the bottom horizontal X-axis.

[0021] Figure 4 Paired comparisons of the FRS (left vertical axis, top histogram) and the number of exons per 50kb segment of chromosome 13 (right vertical axis, bottom histogram). The segments are shown on the bottom horizontal X-axis.

[0022] Figure 5Paired comparisons of the FRS (left y-axis, top histogram) for chromosome 18 and the GC content per 50kb segment (right y-axis, bottom histogram) are shown. The segments are shown on the bottom horizontal X-axis.

[0023] Figure 6 This shows a paired comparison of the number of exons per 50kb region of chromosome 18 (left y-axis, top histogram) and the GC content per 50kb region (right y-axis, bottom histogram). The regions are shown on the bottom horizontal X-axis.

[0024] Figure 7 Paired comparisons of the FRS (left vertical axis, top histogram) and the number of exons per 50kb segment of chromosome 21 (right vertical axis, bottom histogram). The segments are shown on the bottom horizontal X-axis.

[0025] Figure 8 Paired comparisons of the FRS (left y-axis, top histogram) and GC content per 50kb segment of chromosome 21 are shown. The segments are shown on the bottom horizontal X-axis.

[0026] Figure 9 This shows a paired comparison of the number of exons per 50kb region of chromosome 21 (left y-axis, top histogram) and the GC content per 50kb region (right y-axis, bottom histogram). The regions are shown on the bottom horizontal X-axis.

[0027] Figure 10 The PERUN PAD (X-axis) showing chromosome 21 with a LOESS Z score is compared to the PERUN PAD (Y-axis) showing a LOESS Z score based on the "fetal unenriched" portion. The four quadrants represent concordance and inconsistency. A quarter line is drawn at Z=3. The upper right and lower left quadrants are separated by gray diagonal dashed lines. The dashed-dot line is the regression line used only for non-T21 samples. The dotted line is the regression line used for T21 samples based on the high FRS portion.

[0028] Figure 11 The perun pad (X-axis) showing chromosome 21 with a LOESS Z score is compared to the perun pad (Y-axis) showing a LOESS Z score based on the "fetal enrichment" portion (i.e., the portion with high FRS). The four quadrants represent concordance and inconsistency. A quarter line is drawn at Z=3. The upper right and lower left quadrants are separated by gray diagonal dashed lines. The dashed-dot line is the regression line used only for non-T21 samples. The dotted line is the regression line used for T21 samples based on the high FRS portion.

[0029] Figure 12The method for determining nucleic acid fragment length is shown, comprising the following steps: 1) hybridizing a probe (P; dotted line) with a fragment (solid line), 2) trimming the probe, and 3) measuring the probe length. The fragment size determination results for fetal-derived fragments (F) and maternal-derived fragments (M) are shown.

[0030] Figure 13 This shows the fragment length distribution for three different library preparation methods. These include an enzymatic method with automated bead removal, an enzymatic method without automated bead removal, and the TRUSEQ method with automated bead removal. Vertical lines represent the sizes of 143-base and 166-base fragments.

[0031] Figure 14 A schematic diagram of chromosome 13 without using a fragment size filter is shown.

[0032] Figure 15 A schematic diagram of chromosome 13 showing a fragment-sized filter with 150 base pairs.

[0033] Figure 16 A schematic diagram of chromosome 18 without using a fragment size filter.

[0034] Figure 17 A schematic diagram of chromosome 18 showing a fragment-sized filter with 150 base pairs.

[0035] Figure 18 A schematic diagram of chromosome 21 without using a fragment size filter is shown.

[0036] Figure 19 A schematic diagram of chromosome 21 showing a fragment-sized filter with 150 base pairs.

[0037] Figure 20 A schematic diagram of chromosome 13 with a variable fragment size filter (PERUN PAD with LOESS).

[0038] Figure 21 This diagram shows chromosome 18 with a variable fragment size filter (PERUN PAD with LOESS).

[0039] Figure 22 A schematic diagram of chromosome 21 with a variable fragment size filter (PERUN PAD with LOESS).

[0040] Figure 23 A table that displays descriptions of the data used for certain analyses.

[0041] Figure 24 Exemplary implementations of the display system, wherein certain implementations of the technology may be carried out.

[0042] Figure 25A This shows the correlation between the mean FRS (x-axis) of a subset of chromosome 21 from samples obtained from pregnant women with trisomy 21 fetuses (indicated by asterisks) or euploid fetuses (indicated by circles) and the Z-score (y-axis) of the PERUN-normalized count. (Selected for...) Figure 25A The FRS of each part of the subset is greater than the median FRS measured for all parts of chromosome 21 from which the count was obtained. Figure 25B This shows the correlation between the FQA fetal score estimate (x-axis) of all portions of chromosome 21 obtained from the count of chromosome 21 obtained from pregnant women with trisomy 21 fetuses (indicated by asterisks) or euploid fetuses (indicated by circles) and the Z-score (y-axis) of the PERUN standardized count.

[0043] Figure 26 The correlation between the GC content (x-axis) of each reading of the readings (shown in the lower right inset) indicating the range of fragment lengths of chromosome 21 and the cumulative distribution function (CDF, ​​y-axis) based on the reading length.

[0044] Figure 27 Displays the distribution of PERUN intercepts (x-axis) based on the FRS of each box, divided into quantiles (high, high-middle, low-middle, and low).

[0045] Figure 28 Displays the distribution of PERUN maximum cross-validation error (x-axis) based on the fractional (high, medium-high, medium-low, and low) of the FRS for each bin.

[0046] Figure 29 The correlation between the predicted fetal fraction percentage (x-axis) from 19,312 test samples based on the BFF model with 6,000 training samples and the fetal fraction percentage (ChrFF, y-axis) determined by the chromosome Y level is shown (R = 0.81, RMedSE = 1.5).

[0047] Figure 30 The display shows the relative prediction errors (x-axis) based on FRS for multiple bins (i.e., multiple portions) with high fetal fraction content (distribution shown on the left) and low fetal fraction content (distribution shown on the right). The bins with high fetal fraction content show better performance and lower errors. The predicted scores are based on elastic-net regression, using bootstrapping to obtain the density profile.

[0048] Figure 31This displays four distribution plots of model coefficients (x-axis) determined using elastic net regression for subsets of multiple bins separated by fetal fraction content (e.g., low, low-medium, high-medium, high). Bins with higher fetal fraction content tend to produce larger coefficients (positive or negative).

[0049] Figure 32 Two distribution plots are shown for fetal score estimates (x-axis) of female and male test samples determined using the BFF method. The two distribution plots largely overlap. Male and female fetuses show no difference in fetal score distribution (KS-test, P = 0.49). Invention Details

[0050] This document provides methods for analyzing polynucleotides in a mixture of nucleic acids, including, for example, methods for determining the presence of genetic variation. The assessment of genetic variation (e.g., fetal aneuploidy) in a maternal sample typically involves: sequencing the nucleic acids present in the sample, mapping the sequence reads to certain regions in the genome, quantifying the sequence reads of the sample, and analyzing the quantification results. These methods typically involve directly analyzing the nucleic acids in the sample and obtaining nucleotide sequence reads of all or substantially all of the nucleic acids in the sample, which can be expensive and produce redundant and / or irrelevant data. However, combining certain sequence-based and / or length-based separation methods with certain sequence-based and / or length-based analyses can produce specific information about a target genomic region (e.g., a specific chromosome) and, in some examples, can distinguish the origin of nucleic acid fragments, such as maternal origin versus fetal origin. Some methods may include employing sequencing methods, enrichment techniques, and length-based analyses. Some methods described herein, in some embodiments, can be applied without determining the nucleotide sequence of the nucleic acid fragment. This article provides a method for analyzing polynucleotides in a mixture of nucleic acids (e.g., determining the presence of fetal aneuploidy) by combining sequence-based and / or length-based separation and analysis methods.

[0051] This document also provides methods, processing methods, and apparatus for identifying genetic variations. Identification of genetic variations sometimes involves detecting copy number variations and / or sometimes involves adjusting the level of copy number variations. In some embodiments, a level is adjusted to provide identification of one or more genetic variations or changes in a manner that reduces the likelihood of false positive or false negative diagnoses. In some embodiments, the methods described herein for identifying genetic variations can guide the diagnosis of a specific medical condition or determine the predisposition to a specific medical condition. Identifying genetic variations can aid in medical decision-making and / or the use of beneficial medical treatments.

[0052] This document also provides systems, apparatus, and modules that are used in the methods described herein in some embodiments.

[0053] sample

[0054] This document provides methods and compositions for analyzing nucleic acids. In some embodiments, nucleic acid fragments in a mixture of nucleic acid fragments are analyzed. The nucleic acid mixture may include two or more types of nucleic acid fragments having different nucleotide sequences, different fragment lengths, different sources (e.g., genomic sources, fetal and maternal sources, cell or tissue sources, sample sources, object sources, etc.) or combinations thereof.

[0055] The nucleic acids or mixtures of nucleic acids used in the methods and apparatus described herein are often isolated from samples obtained from a subject. The subject can be any living or non-living organism, including but not limited to humans, non-human animals, plants, bacteria, fungi, or protozoa. Any human or non-human animal can be selected, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovids (e.g., cattle), equines (e.g., horses), goats and sheep (e.g., sheep, goats), suidae (e.g., pigs), alpacas (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), bears (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The subject can be male or female (e.g., women, pregnant women). The subject can be of any age (e.g., embryos, fetuses, infants, children, adults).

[0056] Nucleic acids can be isolated from any type of suitable biological sample or specimen (e.g., a test sample). The sample or test sample can be any specimen isolated from or obtained from a subject or a portion thereof (e.g., a human subject, a pregnant woman, a fetus). Non-limiting examples of samples include liquids or tissues of the subject, including but not limited to blood or blood products (e.g., serum, plasma, etc.), cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, cerebrospinal fluid, lavage fluid (e.g., bronchoalveolar, gastric, peritoneum, catheter, ear, arthroscopy), biopsy samples (e.g., from pre-implantation embryos), intermembranous fluid samples, cells (blood cells, placental cells, embryonic or fetal cells, fetal nucleated cells, or fetal cell remnants) or portions thereof (e.g., mitochondria, nuclei, extracts, etc.), female genital tract lavage fluid, urine, feces, sputum, saliva, nasal mucosa, prostatic fluid, lavage fluid, semen, lymph, bile, tears, sweat, breast milk, mammary gland fluid, etc., or combinations thereof. In some embodiments, the biological sample is a cervical swab from the subject. In some embodiments, the biological sample can be blood, and sometimes plasma or serum. As used herein, the term "blood" refers to a blood sample or product derived from a pregnant woman or a woman being tested for a possible pregnancy. The term encompasses whole blood, blood products, or any portion of blood, such as serum and plasma as conventionally defined, tannins, etc. Blood or portions thereof often include nucleosomes (e.g., maternal and / or fetal nucleosomes). Nucleosomes contain nucleic acids and are sometimes cell-free or intracellular. Blood also includes a tannin. The tannin is sometimes separated using a Ficoll gradient. The tannin may include leukocytes (e.g., white blood cells, T cells, B cells, platelets, etc.). In some embodiments, the tannin includes maternal and / or fetal nucleic acids. Blood plasma refers to the portion of whole blood obtained by centrifugation of blood treated with an anticoagulant. Blood serum refers to the liquid aqueous layer retained after a blood sample has clotted. Liquid or tissue samples are typically collected according to standard methods followed in hospitals or clinical practice. In the case of blood, an appropriate amount of peripheral blood (e.g., 3-40 ml) is typically collected and preserved according to standard procedures before or after preparation. Liquid or tissue samples used for nucleic acid extraction may be cell-free (e.g., cell-free). In some embodiments, the liquid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, the sample may contain fetal cells or cancer cells.

[0057] The samples are typically heterogeneous, meaning they contain more than one type of nucleic acid material. For example, heterogeneous nucleic acids can include, but are not limited to, (i) fetal and maternal nucleic acids, (ii) cancer and non-cancer nucleic acids, (iii) pathogen and host nucleic acids, and more commonly, (iv) mutated and wild-type nucleic acids. Samples can be heterogeneous because they contain more than one cell type, such as fetal and maternal cells, cancer and non-cancer cells, or pathogen and host cells. In some embodiments, a few and a majority of nucleic acid materials are present.

[0058] For the prenatal application of the techniques described herein, liquid or tissue samples may be collected from women of gestational age suitable for testing or women who have been tested and are likely pregnant. Suitable gestational age may vary depending on the prenatal test performed. In some embodiments, the pregnant woman is sometimes in the first trimester, sometimes in the second trimester, or sometimes in the last trimester. In some embodiments, the liquid or tissue is collected from pregnant women whose fetus is approximately 1–approximately 45 weeks pregnant (e.g., fetal gestational age 1–4, 4–8, 8–12, 12–16, 16–20, 20–24, 24–28, 28–32, 32–36, 36–40, or 40–44 weeks) and sometimes whose fetus is approximately 5–approximately 28 weeks pregnant (e.g., fetal gestational age 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 weeks). In some implementations, fluid or tissue samples are collected from pregnant women during or immediately after childbirth (e.g., vaginal or non-vaginal delivery, such as cesarean delivery) or 0-72 hours after delivery.

[0059] Obtaining blood samples and DNA extraction

[0060] The methods described herein include isolating, enriching, and analyzing fetal DNA found in maternal blood as a non-invasive means of detecting the presence of maternal and / or fetal genetic variations and / or monitoring the health of the fetus and / or the pregnant woman during pregnancy and sometimes after pregnancy. Therefore, a first step in implementing certain methods of the present invention includes obtaining a blood sample from the pregnant woman and extracting DNA from the sample.

[0061] Obtaining blood samples

[0062] Blood samples can be obtained from pregnant women of gestational age suitable for testing using the methods described in this invention. The appropriate gestational age may vary depending on the disease being tested, as described below. Blood collection from women is typically performed according to standard protocols generally followed in hospitals or clinics. An appropriate amount of peripheral blood is collected, typically 5-50 ml, and preserved according to standard procedures before further preparation. The blood samples can be collected, preserved, or transported in a manner that minimizes degradation of the nucleic acids present in the sample or ensures their quality.

[0063] Preparation of blood samples

[0064] Fetal DNA found in maternal blood is analyzed using, for example, whole blood, serum, or plasma. Methods for preparing serum or plasma from maternal blood are known. For example, blood from a pregnant woman can be placed in a tube containing EDTA to prevent blood clotting, or a commercially available product such as Vacutainer SST (Becton Dickinson, Franklin Lake, New Jersey), and plasma can then be obtained from the whole blood by centrifugation. Serum can be obtained, with or without centrifugation after blood clotting. If centrifugation is used, it is typically (but not limited to) performed at a suitable speed (e.g., 1,500–3,000 g). Plasma or serum may undergo additional centrifugation steps before being transferred to new tubes for DNA extraction.

[0065] In addition to the non-cellular portion of whole blood, DNA can also be recovered from the cellular components and enriched in the tannin layer, which can be obtained by centrifuging a woman's whole blood sample and removing the plasma.

[0066] DNA extraction

[0067] There are several known methods for extracting DNA from biological samples, including blood. These can be done using standard DNA preparation methods (e.g., described in Sambrook and Russell, *Molecular Cloning: A Laboratory Manual*, 3rd edition, 2001); or using a variety of commercially available reagents or kits, such as Qiagen's QIAamp Cyclic Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiaagen, Heldon, Germany), and Genome Prep. TM Blood DNA Separation Kit (Promega, Madison, Wisconsin) and GFX TM The Genomic Blood DNA Purification Kit (Amersham, Piscateway, NJ) can also be used to obtain DNA from blood samples from pregnant women. Combinations of more than one of these methods can also be used.

[0068] In some embodiments, the sample may first be enriched or relatively enriched for fetal nucleic acids using one or more methods. For example, the differentiation between fetal and maternal DNA may be performed using the compositions and methods described in this invention alone or in combination with other differentiating factors. Examples of such factors include, but are not limited to, single nucleotide differences in chromosomes X and Y, chromosome Y-specific sequences, polymorphisms elsewhere in the genome, size differences between fetal and maternal DNA, and differences in methylation forms between maternal and fetal tissues.

[0069] Other methods for enriching samples with specific nucleic acid material are described in PCT patent application No. PCT / US07 / 69991, filed May 30, 2007; PCT patent application No. PCT / US2007 / 071232, filed June 15, 2007; U.S. Provisional Applications Nos. 60 / 968,876 and 60 / 968,878 (as designated to the applicant); and PCT patent application No. PCT / EP05 / 012707, filed November 28, 2005, all of which are incorporated herein by reference. In some embodiments, the parent nucleic acid is selectively removed (partially, substantially, almost completely, or completely) from the sample.

[0070] The terms “nucleic acid” and “nucleic acid molecule” are used interchangeably herein. The term refers to nucleic acids in any composite form, derived from, for example: DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), RNA (e.g., messenger RNA (mRNA), short repressor RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed in the fetus or placenta, etc.), and / or DNA or RNA analogs (e.g., containing base analogs, sugar analogs, and / or non-natural backbones, etc.), RNA / DNA hybrids, and polyamide nucleic acids (PNAs), all of which may be in single-stranded or double-stranded form, and unless otherwise specified, may encompass known analogs of natural nucleotides that function in a manner similar to naturally occurring nucleotides. In some embodiments, nucleic acids may be or may be derived from: plasmids, bacteriophages, autonomously replicating sequences (ARS), centromeres, artificial chromosomes, chromosomes, or other nucleic acids capable of replicating or being replicated in vitro or in a host cell, cell, cell nucleus, or cytoplasm. In some implementations, the template nucleic acid may be derived from a single chromosome (e.g., the nucleic acid sample may be derived from a chromosome of a sample obtained from a diploid organism). Unless explicitly defined, the term covers known analogs containing a reference nucleic acid with similar binding properties and metabolized in a manner similar to that of naturally occurring nucleotides. Unless otherwise stated, a particular nucleic acid sequence also includes its conserved modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences, as well as explicitly indicated sequences. Specifically, degenerate codon substitutions can be obtained by producing a sequence in which the third position of one or more selected (or all) codons is substituted with a mixture of bases and / or deoxyinosine residues. The term nucleic acid is used interchangeably with locus, gene, cDNA, and mRNA encoded by a gene. The term may also include equivalents, derivatives, variants, and analogs of RNA or DNA synthesized from nucleotide analogs, single-stranded ("sense" or "antisense", "positive" or "negative", "positive" reading frame or "reverse" reading frame), and double-stranded polynucleotides. The term “gene” refers to a segment of DNA involved in the production of a polypeptide chain; it includes regions before and after the coding region (leader and tail regions) involved in the transcription / translation of the gene product and the regulation of said transcription / translation, as well as insertion sequences (introns) between individual coding segments (exons).

[0071] Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. For RNA, the cytosine base is replaced with uracil. Template nucleic acids can be prepared using nucleic acids obtained from the target organism as templates.

[0072] Nucleic acid isolation and processing

[0073] Nucleic acids can be obtained from one or more sample sources (such as cells, serum, plasma, ochre layer, lymph, skin, soil, etc.) using methods known in the art. DNA can be isolated, extracted, and / or purified from biological samples (e.g., from blood or blood products) using any suitable method. Non-limiting examples include methods for DNA preparation (e.g., described in Sambrook and Russell, *Molecular Cloning: A Laboratory Manual*, 3rd edition, 2001); various commercially available reagents or kits, such as Qiagen's QIAamp Cyclic Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qagen, Heldon, Germany), and Genome Prep. TM Blood DNA Separation Kit (Promega, Madison, Wisconsin) and GFX TM Genomic blood DNA purification kit (Amersham, Piscavenge, NJ) or combinations thereof.

[0074] Cell lysis methods and reagents are known in the art and can generally be performed by chemical (e.g., detergents, hypotonic solutions, enzymatic processes, etc., or combinations thereof), physical (e.g., French pressure filtration, sonication, etc.), or electrolytic lysis methods. Any suitable lysis process can be used. For example, chemical methods typically use a lysing agent to disrupt cells and extract nucleic acids from them, followed by treatment with a dissociative salt. Physical methods, such as freezing / thawing followed by grinding, and cell pressure filtration, are also useful. High-salt lysis is also commonly used. For example, alkaline lysis can be employed. The latter method conventionally involves the use of a phenol-chloroform solution, and alternatively, a phenol-chloroform-free method comprising three solutions can be used. In the latter method, one solution may contain 15 mM Tris, pH 8.0; 10 mM EDTA and 100 μg / ml RNase A; the second solution may contain 0.2 N NaOH and 1% SDS; and the third solution may contain 3 M KOAc, pH 5.5. These methods can be found in sections 6.3.1–6.3.6 (1989) of the *Current Protocols in Molecular Biology*, published by John Wiley & Sons, Inc., New York, which are included in full in this paper.

[0075] Nucleic acids can also be isolated at different time points than other nucleic acids, with each sample originating from the same or different sources. Nucleic acids can be derived from nucleic acid libraries, such as cDNA or RNA libraries. Nucleic acids can be products of nucleic acid purification or isolation and / or amplification of nucleic acid molecules in a sample. Nucleic acids provided for the methods described herein may comprise nucleic acids from one sample or from two or more samples (e.g., from 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 11 or more, 12 or more, 13 or more, 14 or more, 15 or more, 16 or more, 17 or more, 18 or more, 19 or more, 20 or more samples).

[0076] In some embodiments, nucleic acids may include extracellular nucleic acids. As used herein, the term "extracellular nucleic acid" refers to nucleic acids isolated from a substantially cell-free source, and is also referred to as "cell-free" nucleic acids, "circulating cell-free nucleic acids" (e.g., CCF fragments), and / or "cell-free circulating nucleic acids." Extracellular nucleic acids may be present in and obtained from blood (e.g., from the blood of a pregnant woman). Extracellular nucleic acids typically do not contain detectable cells and may contain cellular elements or cellular remnants. Non-limiting examples of cell-free sources of extracellular nucleic acids include blood, plasma, serum, and urine. As used herein, the term "obtaining circulating cell-free sample nucleic acid" includes obtaining a sample directly (e.g., collecting a sample, such as a test sample) or obtaining a sample from a person who has already collected a sample. Without being theoretically limited, extracellular nucleic acids may be products of apoptosis and cell lysis, which often results in extracellular nucleic acids having a range of lengths (e.g., "ladders").

[0077] In some implementations, extracellular nucleic acids may contain different nucleic acid substances, and are therefore referred to herein as "heterogeneity." For example, the blood serum or plasma of a person with cancer may contain nucleic acids from cancer cells and nucleic acids from non-cancer cells. In another example, the blood serum or plasma of a pregnant woman may contain maternal nucleic acids and fetal nucleic acids. In some examples, fetal nucleic acids sometimes account for about 5% to about 50% of all nucleic acids (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49% of the total nucleic acids are fetal nucleic acids). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 500 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 500 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 250 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 250 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 200 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 200 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 150 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 150 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 100 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 100 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 50 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 50 base pairs or less). In some embodiments, the majority of the fetal nucleic acid in the nucleic acid is about 25 base pairs or less in length (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the fetal nucleic acid length is about 25 base pairs or less).

[0078] In some embodiments, nucleic acids may be provided for performing the methods described herein without processing the nucleic acid-containing sample. In some embodiments, nucleic acids are provided for performing the methods described herein after processing the nucleic acid-containing sample. For example, nucleic acids may be extracted, isolated, purified, partially purified, or amplified from the sample. As used herein, the term "isolation" means the removal of nucleic acids from their original environment (e.g., the natural environment in which nucleic acids are naturally produced or the host cell expressing exogenous nucleic acids), thus altering the nucleic acids from their original environment through human intervention (e.g., "artificial"). As used herein, the term "isolated nucleic acid" refers to nucleic acids removed from an object (e.g., a human object). Isolated nucleic acids may contain fewer non-nucleic acid components (e.g., proteins, lipids) compared to the component content present in the source sample. Compositions containing isolated nucleic acids may be about 50% to more than 99% free of non-nucleic acid components. Compositions containing isolated nucleic acids may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more than 99% free of non-nucleic acid components. As used herein, the term "purified" refers to nucleic acids containing fewer non-nucleic acid components (e.g., proteins, lipids, carbohydrates) compared to the amount of non-nucleic acid components present prior to the purification process. Compositions containing purified nucleic acids may be about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other non-nucleic acid components. The term "purified" as used herein may also refer to nucleic acids containing fewer nucleic acid substances compared to the sample source from which they are derived. Compositions containing purified nucleic acids may be about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater than 99% free of other nucleic acid substances. For example, fetal nucleic acids can be purified from a mixture containing maternal and fetal nucleic acids. In some examples, nucleosomes containing small fragments of fetal nucleic acid can be purified from a mixture of macronucleosome complexes containing larger fragments of maternal nucleic acid.

[0079] In some embodiments, nucleic acids are fragmented or cleaved before, during, or after the method of the present invention. The fragmented or cleaved nucleic acids may have a nominal, average, or arithmetic mean length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs. Fragments can be generated by suitable methods known in the art, and the average, geometric mean, or nominal length of the nucleic acid fragments can be controlled by selecting an appropriate fragment generation method.

[0080] Nucleic acid fragments may contain overlapping nucleotide sequences, which can facilitate the construction of nucleotide sequences of unfragmented corresponding nucleic acids or segments thereof. For example, one fragment may have subsequences x and y, and other fragments may have subsequences y and z, where x, y, and z are nucleotide sequences of 5 nucleotides or longer. In some embodiments, overlapping nucleic acid y can be used to facilitate the construction of the xyz nucleotide sequence from the nucleic acid of a sample. In some embodiments, the nucleic acid may be partially fragmented (e.g., from an incomplete or terminated specific splicing reaction) or fully fragmented.

[0081] In some embodiments, nucleic acids may be fragmented or cleaved by suitable methods, non-limiting examples of which include physical methods (e.g., shearing, sonication, French press filtration, heat, ultraviolet irradiation, etc.), enzyme processing (e.g., enzyme cleavage reagents (e.g., suitable nucleases, suitable restriction enzymes, suitable methylation-sensitive restriction enzymes)), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, alkaline hydrolysis, heat, etc. or combinations thereof), the methods described in U.S. Patent Application Publication 20050112590, etc., or combinations thereof.

[0082] As used in this article, “fragmentation” or “splitting” refers to a method or condition that allows a nucleic acid molecule (such as a nucleic acid template gene molecule or its amplification product) to be divided into two or more smaller nucleic acid molecules. Such fragmentation or splicing can be sequence-specific, base-specific, or non-specific, and can be accomplished by any different methods, reagents, or conditions (including, for example, chemical, enzymatic, and physical fragmentation).

[0083] As used herein, the terms “fragment,” “cut product,” “cut product,” or their grammatical variations refer to nucleic acid molecules obtained by fragmentation or cutting of a nucleic acid template gene molecule or its amplification product. While such fragments or cut products can refer to all nucleic acids obtained by a cutting reaction, they generally refer only to nucleic acid molecules obtained by fragmentation or cutting of a segment of a nucleic acid template gene molecule or its amplification product (containing the corresponding nucleotide sequence of the nucleic acid template gene molecule). As used herein, the term “amplification” refers to the process of generating amplicon nucleic acids in a processed sample in a linear or exponential manner, the nucleotide sequence of which is identical or substantially identical to the nucleotide sequence of the target nucleic acid or its segment. In some embodiments, the term “amplification” refers to a method including polymerase chain reaction (PCR). For example, the amplification product can contain one or more additional nucleotides than the amplified nucleotide region of the nucleic acid template sequence (e.g., primers can contain “extra” nucleotides, such as transcription initiation sequences, in addition to nucleotides complementary to those of the nucleic acid template gene molecule, resulting in an amplification product containing “extra” nucleotides or nucleotides not corresponding to the amplified nucleotide region of the nucleic acid template gene molecule). Therefore, a fragment can contain a segment or portion of an amplified nucleic acid molecule, which at least partially contains nucleotide sequence information from or based on a representative nucleic acid template molecule.

[0084] As used herein, the term "complementary shearing reaction" refers to a shearing reaction on the same nucleic acid using different shearing agents or by altering the shearing specificity of the same shearing agent, thereby producing different shearing patterns of the same target or reference nucleic acid or protein. In some embodiments, nucleic acids may be treated in one or more reaction vessels using one or more specific shearing agents (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more specific shearing agents). As used herein, the term "specific shearing agent" refers to a reagent, sometimes a chemical or enzyme that can cleave nucleic acids at one or more specific sites.

[0085] Before being used in the methods described herein, nucleic acids may be treated to modify certain nucleotides within them. For example, the nucleic acid may be subjected to a treatment that selectively modifies the nucleic acid based on the methylation state of the nucleotides. Furthermore, conditions such as high temperature, ultraviolet radiation, and X-ray radiation can induce variations in the nucleic acid molecular sequence. Nucleic acids may be provided in any suitable form for appropriate sequence analysis.

[0086] Nucleic acids can be single-stranded or double-stranded. For example, single-stranded DNA can be generated by denaturing double-stranded DNA through heating or, for example, treatment with an alkali. In some embodiments, the nucleic acid is a D-loop structure, formed by the invasion of an oligonucleotide or DNA-like molecule, such as peptide nucleic acid (PNA), into the middle strand of a double-stranded DNA molecule. Adding E. coli RecA protein and / or altering the salt concentration (e.g., using methods known in the art) facilitates the formation of D-loops.

[0087] Genome targets

[0088] In some embodiments, the target nucleic acid, also referred to herein as the target fragment, includes a multinucleotide fragment from multiple genomic regions (e.g., a single chromosome, a set of chromosomes, and / or certain chromosomal regions) of a specific genomic region. In some embodiments, the genomic region may be associated with fetal genetic abnormalities (e.g., aneuploidy) and other genetic variations, including but not limited to: mutations (e.g., point mutations), insertions, additions, deletions, translocations, trinucleotide repeat disorders, and / or single nucleotide polymorphisms (SNPs). In some embodiments, the reference nucleic acid, also referred herein as the reference fragment, includes a multinucleotide fragment from one or more genomic regions not associated with fetal genetic abnormalities. In some embodiments, the target nucleic acid and / or reference nucleic acid (i.e., the target fragment and / or reference fragment) contain nucleotide sequences substantially proprietary to the chromosome of interest or the reference chromosome (e.g., nucleotide sequences that are not found elsewhere in the genome or are substantially similar).

[0089] In some embodiments, fragments from multiple genomic regions are analyzed. In some embodiments, target fragments and reference fragments from multiple genomic regions are analyzed. In some embodiments, fragments from multiple genomic regions are analyzed to determine, for example, the presence, amount (e.g., relative amount), or proportion of a chromosome of interest. In some embodiments, the chromosome of interest is a chromosome suspected of euploidy and may be referred to herein as a "test chromosome." In some embodiments, fragments from multiple genomic regions of a presumed euploid chromosome are analyzed. This chromosome may be referred to herein as a "reference chromosome." In some embodiments, multiple test chromosomes are analyzed. In some embodiments, the test chromosomes are selected from chromosome 13 (chromosome 13), chromosome 18 (chromosome 18), and chromosome 21 (chromosome 21). In some embodiments, the reference chromosomes are selected from chromosomes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, X, and Y, and sometimes, the reference chromosomes are selected from autosomes (i.e., not X and Y). In some embodiments, chromosome 20 (Chr20) is selected as the reference chromosome. In some embodiments, chromosome 14 is selected as the reference chromosome. In some embodiments, chromosome 9 is selected as the reference chromosome. In some embodiments, the test chromosome and the reference chromosome are from the same individual. In some embodiments, the test chromosome and the reference chromosome are from different individuals.

[0090] In some embodiments, the test and / or reference chromosome is analyzed for fragments from at least one genomic region. In some embodiments, the test and / or reference chromosome is analyzed for fragments from at least 10 genomic regions (e.g., about 20, 30, 40, 50, 60, 70, 80, or 90 genomic regions). In some embodiments, the test and / or reference chromosome is analyzed for fragments from at least 100 genomic regions (e.g., about 200, 300, 400, 500, 600, 700, 800, or 900 genomic regions). In some embodiments, the test and / or reference chromosome is analyzed for fragments from at least 1,000 genomic regions (e.g., about 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, or 9,000 genomic regions). In some embodiments, the test chromosome and / or reference chromosome are analyzed for fragments from at least 10,000 genomic regions (e.g., approximately 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, or 90,000 genomic regions). In some embodiments, the test chromosome and / or reference chromosome are analyzed for fragments from at least 100,000 genomic regions (e.g., approximately 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, or 900,000 genomic regions).

[0091] Enrichment and isolation of nucleic acid subgroups

[0092] In some embodiments, nucleic acids (e.g., extracellular nucleic acids) are enriched or relatively enriched against nucleic acid subsets or substances. Nucleic acid subsets may include, for example, fetal nucleic acids, maternal nucleic acids, nucleic acids containing fragments of a specific length or length range, or nucleic acids derived from specific genomic regions (e.g., a single chromosome, a set of chromosomes, and / or certain chromosomal regions). Such enriched samples may be used in conjunction with the methods described herein. Therefore, in some embodiments, the technique includes the additional step of enriching nucleic acid subsets, such as fetal nucleic acids, in a sample. In some embodiments, the methods described herein for determining fetal fractions may also be used to enrich fetal nucleic acids. In some embodiments, maternal nucleic acids are selectively removed (partially, substantially, almost completely, or completely) from the sample. In some embodiments, enriching specific low-copy-number nucleic acids (e.g., fetal nucleic acids) may improve quantitative sensitivity. Methods for enriching specific types of nucleic acids in samples, such as those described below, are incorporated herein by reference in U.S. Patent Nos. 6,927,028, WO2007 / 140417, WO2007 / 147063, WO2009 / 032779, WO2009 / 032781, WO2010 / 033639, WO2011 / 034631, WO2006 / 056480, and WO2011 / 143659.

[0093] In some embodiments, nucleic acids are enriched for certain target fragment types and / or reference fragment types. In some embodiments, nucleic acid enrichment is performed on specific nucleic acid fragment lengths or fragment lengths or ranges using one or more length-based separation methods described herein and / or known in the art. In some embodiments, nucleic acid enrichment is performed on fragments selected from genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein and / or known in the art. Methods for enriching certain nucleic acid subsets (e.g., fetal nucleic acids) in samples are detailed below.

[0094] Methods for enriching nucleic acid subpopulations (e.g., fetal nucleic acids) that can be used in conjunction with the methods of the present invention include methods that utilize the epigenetic differences between maternal and fetal nucleic acids. For example, fetal and maternal nucleic acids can be distinguished and separated based on methylation differences. A method for enriching fetal nucleic acids based on methylation is described in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference. Such methods sometimes involve binding sample nucleic acids to methylation-specific binding agents (methyl CpG-binding proteins (MBD), methylation-specific antibodies, etc.) and separating bound and unbound nucleic acids based on different methylation states. Such methods may also include the use of methylation-sensitive restriction enzymes (e.g., HhaI and HpaII as described above), which selectively digest nucleic acids from a maternal sample by using enzymes that selectively and completely or substantially digest maternal nucleic acids to enrich at least one fetal nucleic acid region in the sample, thus enabling the enrichment of fetal nucleic acid regions in the maternal sample.

[0095] Other methods for enriching nucleic acid subsets (e.g., fetal nucleic acids) that can be used in conjunction with the method of this invention include restriction endonuclease-enhanced polymorphic sequencing, such as the method described in U.S. Patent Application Publication No. 2009 / 0317818, which is incorporated herein by reference. This method includes cleaving the nucleic acid containing the non-target allele with a restriction endonuclease that recognizes the non-target allele but does not recognize the target allele; and amplifying the uncut nucleic acid but not the cleaved nucleic acid, wherein the uncut amplified nucleic acid represents a target nucleic acid (e.g., fetal nucleic acid) enriched relative to the non-target nucleic acid (e.g., maternal nucleic acid). In some embodiments, the nucleic acid may be selected such that it contains an allele having a polymorphic site that is readily digested, for example, by a cleavage agent.

[0096] Methods for enriching nucleic acid subsets (e.g., fetal nucleic acids) that can be used in conjunction with the methods of the present invention include selective enzymatic degradation. These methods involve protecting the target sequence from digestion by exonucleases, thereby facilitating the removal of unwanted sequences (e.g., maternal DNA) from the sample. For example, in one method, sample nucleic acids are denatured to produce single-stranded nucleic acids, which are then contacted with at least one target-specific primer pair under suitable annealing conditions. The annealed primers are extended using nucleotide polymerization to produce a double-stranded target sequence, and the single-stranded nucleic acids are digested with a nuclease that digests single-stranded (e.g., non-target) nucleic acids. In some embodiments, the method can be repeated at least one more cycle. In some embodiments, the same target-specific primer pair can be used to initiate the first and second cycles of extension, and in some embodiments, different target-specific primer pairs are used for the first and second cycles.

[0097] In some embodiments, nucleic acid enrichment is performed on fragments of selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein. In some embodiments, nucleic acids are enriched for specific polynucleotide fragment lengths or fragment length ranges and for fragments from selected genomic regions (e.g., chromosomes) by combining length-based and sequence-based separation methods. The length-based and sequence-based separation methods are described in further detail below.

[0098] Methods for enriching nucleic acid subsets (e.g., fetal nucleic acids) that can be used in conjunction with the methods of this invention include massively parallel sequencing (MPSS). MPSS is typically a solid-phase method that uses adaptors (i.e., tags) to ligate nucleic acid sequences, which are then decoded and read in small increments. Tagged PCR products are typically amplified, resulting in PCR products with unique tags for each nucleic acid. Tags are typically used to conjugate PCR products to microbeads. For example, sequence signatures can be identified from each bead after several rounds of sequence determination based on ligation. Each signature sequence (MPSS tag) in the MPSS database is analyzed, all other signatures are compared, and all identical signatures are counted.

[0099] In some embodiments, certain enrichment methods (such as certain MPS-based and / or MPSS-based enrichment methods) may include amplification-based methods (such as PCR). In some embodiments, site-specific amplification methods (e.g., using site-specific amplification primers) may be used. In some embodiments, multiplex SNP allele PCR methods may be used. In some embodiments, multiplex SNP allele PCR methods may be used in conjunction with singlet sequencing. For example, this method may involve using multiplex PCR (MASSARRAY system) and incorporating the capture probe sequence into the amplicons, followed by sequencing using, for example, an Illumina MPSS system. In some embodiments, multiplex SNP allele PCR methods may be used in conjunction with a three-primer system and index sequencing. For example, this method may involve using multiplex PCR (MASSARRAY system) where primers incorporate a first capture probe into site-specific forward PCR primers and an adaptor sequence into site-specific reverse PCR primers to generate an amplicons, followed by secondary PCR incorporating the reverse capture sequence and molecular index barcode for sequencing using, for example, an Illumina MPSS system. In some implementations, multiplex SNP allele PCR methods can be used in conjunction with a four-primer system and index sequencing. For example, this method may involve using multiplex PCR (MASSARRAY system) with primers that incorporate adaptor sequences into site-specific forward and reverse PCR primers, followed by secondary PCR to incorporate forward and reverse capture sequences and molecular index barcodes for sequencing using, for example, the Eminenta MPSS system. In some implementations, microfluidic methods can be used. In some implementations, array-based microfluidic methods can be used. For example, this method may involve using a microfluidic array (such as Fluidigm) for low-repetition amplification and incorporation of index and capture probes, followed by sequencing. In some implementations, emulsion microfluidic methods, such as digital droplet PCR, can be used.

[0100] In some embodiments, universal amplification methods (e.g., using universal or non-site-specific amplification primers) may be used. In some embodiments, universal amplification methods may be combined with pull-down methods. In some embodiments, the method may include pulling down biotinylated ultramers from a universal amplification sequence library (e.g., biotinylated pull-down assays from Agilent or IDT). For example, the method may involve preparing a standard library, enriching selected regions by pull-down assays, and a secondary universal amplification step. In some embodiments, pull-down methods may be used in conjunction with ligation-based methods. In some embodiments, the method may include pulling down biotinylated ultramers ligated with sequence-specific adaptors (e.g., HALOPLEX PCR, HaloGenomics). For example, the method may involve using selector probes to capture restriction enzyme-digested fragments, then ligating the captured product and adaptor, and universal amplification followed by sequencing. In some embodiments, pull-down methods may be used in conjunction with extension and ligation-based methods. In some embodiments, the method may include molecular inverted probe (MIP) extension and ligation. For example, this method may involve using a molecular inverted probe in combination with a sequence adaptor, followed by universal amplification and sequencing. In some implementations, complementary DNA can be synthesized and sequenced without amplification.

[0101] In some implementations, extension and ligation methods can be performed without pulling down components. In some implementations, the method may include site-specific forward and reverse primer hybridization, extension, and ligation. The method may also include universal amplification or complementary DNA synthesis without amplification followed by sequencing. In some implementations, the method may reduce or exclude background sequences during analysis.

[0102] In some embodiments, pull-down assays may be used with or without optional amplification components. In some embodiments, the method may include modified pull-down assays and ligations that fully incorporate the capture probe without universal amplification. For example, the method may involve using a modified selector probe to capture a restriction enzyme-digested fragment, then ligating the capture product and adaptor, and optionally amplifying, and sequencing. In some embodiments, the method may include a biotinylated pull-down assay, and a combination of extension and ligation of the adaptor sequence with circular single-stranded ligation. For example, the method may involve using a selector probe to capture the region of interest (i.e., the target sequence), extending the probe, ligating the adaptor, single-stranded circular ligation, optional amplification, and sequencing. In some embodiments, analysis of the sequencing results may separate the target sequence from the background.

[0103] In some embodiments, nucleic acid enrichment is performed on fragments of selected genomic regions (e.g., chromosomes) using one or more sequence-based separation methods described herein. Sequence-based separation is typically based on nucleotide sequences (e.g., target fragments and / or reference fragments) present in the fragment of interest in the sample but substantially absent (e.g., 5% or less) in other fragments or other segments. In some embodiments, sequence-based separation can generate isolated target fragments and / or isolated reference fragments. The isolated target fragments and / or isolated reference fragments are typically separated from the remaining fragments in the nucleic acid sample. In some embodiments, the isolated target fragments and isolated reference fragments may also be separated from each other (e.g., in separate laboratory compartments). In some embodiments, the isolated target fragments and isolated reference fragments may be separated together (e.g., in the same laboratory). In some embodiments, unbound fragments may be differentially removed, degraded, or digested.

[0104] In some implementations, selective nucleic acid capture methods are used to separate target fragments and / or reference fragments from a nucleic acid sample. Commercially available nucleic acid capture systems include, for example, the Nimblegen sequence capture system (NimbleGen of Roche, Madison, Wisconsin); the BEADARRAY platform (Illumina, San Diego, California); the GENECHIP platform (Affymetrix, Santa Clara, California); the Agilent SureSelect target enrichment system (Agilent Technologies, Santa Clara, California); and related platforms. The method typically involves the hybridization of a captured oligonucleotide with a segment or all of the nucleotide sequence of the target or reference fragment and may include the use of solid-phase (e.g., solid-phase arrays) and / or solution-based platforms. The captured oligonucleotide (sometimes referred to as a “bait”) may be selected or engineered so that it preferentially hybridizes to nucleic acid fragments of selected genomic regions or sites (e.g., one of chromosomes 21, 18, 13, X, or Y, or a reference chromosome). In some implementations, hybridization-based methods (e.g., using oligonucleotide arrays) can be used to enrich nucleic acid sequences from certain chromosomes (e.g., possible aneuploid chromosomes, reference chromosomes, or other chromosomes of interest) or regions of interest.

[0105] Capture oligonucleotides typically contain nucleotide sequences capable of hybridizing or annealing to or from a portion of a nucleic acid fragment of interest (e.g., a target fragment, a reference fragment). Capture oligonucleotides can be naturally occurring or synthetic and can be based on DNA or RNA. Capture oligonucleotides allow for specific separations, such as separating target fragments and / or reference fragments from other fragments in a nucleic acid sample. The term “specific” or “specific” as used herein refers to the binding or hybridization of one molecule with another molecule, such as an oligonucleotide targeting a target polynucleotide. “Specific” or “specific” refers to the recognition, contact, and formation of a stable complex between two molecules, compared to a significantly lower recognition, contact, or complex formation between either molecule and other molecules. The term “annealing” as used herein refers to the formation of a stable complex between two molecules. When referring to capture oligonucleotides, the terms “capture oligonucleotide,” “capture oligo,” “oligonucleotide,” or “oligonucleotide” are used interchangeably throughout the text. The following characteristics of oligonucleotides can be applied to primers and other oligonucleotides, such as the probes provided herein.

[0106] Capture oligonucleotides can be designed and synthesized using suitable methods, and these capture oligonucleotides can have any length suitable for hybridization to the nucleotide sequence of interest and for the isolation and / or analytical processing described herein. The oligonucleotides can be designed based on the nucleotide sequence of interest (e.g., a target fragment sequence, a reference fragment sequence). In some embodiments, the length of the oligonucleotide can be about 10 to about 300 nucleotides, about 10 to about 100 nucleotides, about 10 to about 70 nucleotides, about 10 to about 50 nucleotides, about 15 to about 30 nucleotides, or about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides. The oligonucleotides can consist of naturally occurring and / or non-naturally occurring nucleotides (e.g., labeled nucleotides) or mixtures thereof. The oligonucleotides suitable for the embodiments described herein can be synthesized and labeled using known techniques. The oligonucleotides can be chemically synthesized via a solid-phase phosphoramidite-triester method, first described by Beaucage and Caruthers in Tetrahedron Letts. 22:1859-1862 (1981), for example, using an automated synthesizer as described by Needham-VanDevanter et al. in Nucleic Acids Res. 12:6159-6168, 1984. Purification of the oligonucleotides can be achieved by natural acrylamide gel electrophoresis or by anion-exchange high-performance liquid chromatography (HPLC), for example, as described by Pearson and Reanier in J. Chrom. 255:137-149 (1983).

[0107] In some implementations, all or part of the (naturally generated or synthetic) oligonucleotide sequence may be substantially complementary to the target fragment and / or reference fragment sequence or a portion thereof. "Substantially complementary" as used herein refers to nucleotide sequences capable of hybridizing with each other. The stringency of the hybridization conditions can be varied to allow for different amounts of sequence mismatches. The complementarity between the included target / reference sequence and oligonucleotide sequence is 55% or higher, 56% or higher, 57% or higher, 58% or higher, 59% or higher, 60% or higher, 61% or higher, 62% or higher, 63% or higher, 64% or higher, 65% or higher, 66% or higher, 67% or higher, 68% or higher, 69% or higher, 70% or higher, 71% or higher, 72% or higher, 73% or higher, 74% or higher, 75% or higher. High, 76% or higher, 77% or higher, 78% or higher, 79% or higher, 80% or higher, 81% or higher, 82% or higher, 83% or higher, 84% or higher, 85% or higher, 86% or higher, 87% or higher, 88% or higher, 89% or higher, 90% or higher, 91% or higher, 92% or higher, 93% or higher, 94% or higher, 95% or higher, 96% or higher, 97% or higher, 98% or higher, or 99% or higher.

[0108] Oligonucleotides substantially complementary to a nucleic acid sequence of interest (e.g., a target fragment sequence, a reference fragment sequence) or a portion thereof are also substantially similar to the complementary sequence of the target nucleic acid sequence or a related portion thereof (e.g., substantially similar to the antisense strand of the nucleic acid). One test for determining whether two nucleotide sequences are substantially similar is to determine the percentage of shared identical nucleotide sequences. "Substantially similar" as used herein means that the nucleotide sequences are 55% or higher, 56% or higher, 57% or higher, 58% or higher, 59% or higher, 60% or higher, 61% or higher, 62% or higher, 63% or higher, 64% or higher, 65% or higher, 66% or higher, 67% or higher, 68% or higher, 69% or higher, 70% or higher, 71% or higher, 72% or higher, 73% or higher, 74% or higher, or 75% or higher. or higher, 76% or higher, 77% or higher, 78% or higher, 79% or higher, 80% or higher, 81% or higher, 82% or higher, 83% or higher, 84% or higher, 85% or higher, 86% or higher, 87% or higher, 88% or higher, 89% or higher, 90% or higher, 91% or higher, 92% or higher, 93% or higher, 94% or higher, 95% or higher, 96% or higher, 97% or higher, 98% or higher, or 99% or higher.

[0109] Annealing conditions (e.g., hybridization conditions) can be determined and / or adjusted based on the characteristics of the oligonucleotides used in the analysis. The oligonucleotide sequence and / or length can sometimes affect the hybridization of the nucleic acid sequence of interest. Depending on the degree of mismatch between the oligonucleotide and the nucleic acid of interest, low, medium, or high stringency conditions can be used to achieve annealing. As used herein, the term "stringency conditions" refers to the conditions for hybridization and washing. Methods for optimizing hybridization reaction temperature conditions are known in the art and can be found in sections 6.3.1–6.3.6 of *Current Protocols in Molecular Biology* (1989), published by John Wiley & Sons, NY. Aqueous and non-aqueous methods described in that reference can be used. A non-limiting example of stringency hybridization conditions is hybridization at approximately 45°C in 6X sodium chloride / sodium citrate (SSC), followed by one or more washes at 50°C in 0.2X SSC, 0.1% SDS. Another example of stringent hybridization conditions is hybridization in 6X sodium chloride / sodium citrate (SSC) at approximately 45°C, followed by one or more washes in 0.2X SSC and 0.1% SDS at 55°C. Another example of stringent hybridization conditions is hybridization in 6X sodium chloride / sodium citrate (SSC) at approximately 45°C, followed by one or more washes in 0.2X SSC and 0.1% SDS at 60°C. Typically, stringent hybridization conditions are hybridization in 6X sodium chloride / sodium citrate (SSC) at approximately 45°C, followed by one or more washes in 0.2X SSC and 0.1% SDS at 65°C. More commonly, stringent conditions involve treatment with 0.5M sodium phosphate and 7% SDS at 65°C, followed by one or more washes in 0.2X SSC and 1% SDS at 65°C. The stringent hybridization temperature can also be altered (i.e., lowered) by adding certain organic solvents such as formamide. Organic solvents (such as formamide) reduce the thermal stability of double-stranded polynucleotides, thereby allowing hybridization to be performed at lower temperatures while still maintaining stringent conditions and extending the lifespan of potentially heat-sensitive nucleic acids.

[0110] As used herein, “hybridization” or its grammatical variations refer to the annealing of a first nucleic acid molecule with a second nucleic acid molecule under strict low, medium, or high conditions, or under nucleic acid synthesis conditions. Hybridization can include situations where a first nucleic acid molecule anneals to a second nucleic acid molecule, wherein the first and second nucleic acid molecules are complementary. As used herein, “specific hybridization” refers to the preferential hybridization of an oligonucleotide with a sequence complementary to that oligonucleotide under nucleic acid synthesis conditions, compared to hybridization with a nucleic acid molecule that does not have a complementary sequence. For example, specific hybridization includes hybridization that captures an oligonucleotide and a target fragment sequence complementary to that oligonucleotide.

[0111] In some embodiments, one or more capturing oligonucleotides are associated with an affinity ligand (e.g., a member of a binding pair (e.g., biotin)) or an antigen capable of binding to a capturing agent (e.g., avidin, streptavidin, antibody, or receptor). For example, the capturing oligonucleotide may be biotinylated so that it can be captured onto a streptavidin-coated bead.

[0112] In some embodiments, one or more capturing oligonucleotides and / or capturing agents are efficiently linked to a solid support or matrix. The solid support or matrix can be any practically separable solid to which the capturing oligonucleotides can be directly or indirectly attached, including but not limited to surfaces provided by microarrays and pores, and particles such as beads (e.g., paramagnetic beads, magnetic beads, microbeads, nanobeads), microparticles, and nanoparticles. Solid supports may also include, for example, chips, columns, optical fibers, swabs, filters (e.g., flat surface filters), one or more capillaries, glass and modified or functionalized glass (e.g., controlled-pore glass (CPG)), quartz, mica, diazotized components (paper or nylon), polyoxymethylene, cellulose, cellulose acetate, paper, ceramics, metals, metalloids, semiconductor materials, quantum dots, coated beads or particles, other chromatographic materials, magnetic particles; plastics (including acrylic resins, polystyrene, copolymers of styrene and other materials, polybutene, polyurethane, TEFLON). TM Polyethylene, polypropylene, polyamide, polyester, polyvinylidene fluoride (PVDF), polysaccharides, nylon or nitrocellulose, resins, silica or silica-based materials, including silicon, silicone, and modified silicon. Carbon, metals (e.g., steel, gold, silver, aluminum, silicon, and copper), inorganic glass, conductive polymers (including polymers such as polypyrrole and polyindole); surfaces of micro or nanostructures, such as nucleic acid tile arrays, nanotubes, nanowires, or nanoparticle-modified surfaces; or porous surfaces or gels such as methacrylates, acrylamide, sugar polymers, cellulose, silicates, or other fibrous or stranded polymers. In some embodiments, the solid support or matrix may be coated with a passively or chemically derivatized coating having any number of materials, including polymers such as dextran, acrylamide, gelatin, or agarose. Beads and / or particles may be free or interconnected (e.g., sintered). In some embodiments, the solid support may be an aggregate of beads. In some embodiments, the particles may comprise silica, and the silica may comprise silicon dioxide. In some embodiments, the silica may be porous, and in some embodiments, the silica may be non-porous. In some embodiments, the particles also contain an agent that imparts paramagnetic properties to the particles. In some embodiments, the reagent comprises a metal, and in other embodiments, the reagent is a metal oxide (e.g., iron or an oxide of iron, wherein the oxide of iron comprises a mixture of Fe2+ and Fe3+). The oligonucleotide can be attached to the solid support via covalent or non-covalent interactions, and can be attached directly or indirectly (e.g., via a medium such as a spacer molecule or biotin) to the solid support. The probe can be attached to the solid support before, during, or after nucleic acid capture.

[0113] In some embodiments, one or more length-based separation methods are used to enrich nucleic acids for specific fragment lengths, length ranges, or lengths below or above a specific threshold or cutoff value. Fragment length typically refers to the number of nucleotides in a fragment. Fragment length sometimes also refers to fragment size. In some embodiments, length-based separation methods do not require measuring the length of individual fragments. In some embodiments, length-based separation methods are combined with methods for determining the length of individual fragments. In some embodiments, length-based separation refers to size fractionation, where all or part of the fractionated library can be separated (e.g., retained) and / or analyzed. Size fractionation is known in the art (e.g., array separation, molecular sieve separation, gel electrophoresis separation, column chromatography separation (e.g., size exclusion columns), and microfluidics-based methods). In some embodiments, length-based separation methods may include, for example, fragment cyclization, chemical treatment (e.g., formaldehyde, polyethylene glycol (PEG)), mass spectrometry, and / or size-specific nucleic acid amplification.

[0114] In some embodiments, nucleic acid fragments having a certain length, a length range, or a length below or above a specific threshold or cutoff value are isolated from the sample. In some embodiments, fragments having a length below a specific threshold or cutoff value (e.g., 500 bp, 400 bp, 300 bp, 200 bp, 150 bp, 100 bp) are referred to as “short” fragments, while fragments having a length above a specific threshold or cutoff value (e.g., 500 bp, 400 bp, 300 bp, 200 bp, 150 bp, 100 bp) are referred to as “long” fragments. In some embodiments, fragments having a certain length, a length range, or a length below or above a specific threshold or cutoff value are retained for analysis, while fragments having different lengths, length ranges, or lengths above or below the threshold or cutoff value are not retained for analysis. In some embodiments, fragments less than about 500 bp are retained. In some embodiments, fragments less than about 400 bp are retained. In some embodiments, fragments less than about 300 bp are retained. In some embodiments, fragments less than about 200 bp are retained. In some implementations, segments smaller than approximately 150 bp are retained. For example, segments smaller than approximately 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, 110 bp, or 100 bp are retained. In some implementations, segments between approximately 100 bp and approximately 200 bp are retained. For example, segments between approximately 190 bp, 180 bp, 170 bp, 160 bp, 150 bp, 140 bp, 130 bp, 120 bp, or 110 bp are retained. In some implementations, segments within the range of approximately 100 bp to approximately 200 bp are retained. For example, segments within the range of approximately 110 bp to approximately 190 bp, 130 bp to approximately 180 bp, 140 bp to approximately 170 bp, 140 bp to approximately 150 bp, 150 bp to approximately 160 bp, or 145 bp to approximately 155 bp are retained. In some embodiments, segments that are approximately 10 bp to approximately 30 bp shorter than other segments of a certain length or length range are retained. In some embodiments, segments that are approximately 10 bp to approximately 20 bp shorter than other segments of a certain length or length range are retained. In some embodiments, segments that are approximately 10 bp to approximately 15 bp shorter than other segments of a certain length or length range are retained.

[0115] In some implementations, nucleic acids are enriched using one or more bioinformatics-based (e.g., insilico) methods targeting specific nucleic acid fragment lengths, length ranges, or lengths below or above specific thresholds or cutoff values. For example, suitable nucleotide sequencing methods can be used to obtain nucleotide sequence reads of nucleic acid fragments. In some examples, such as when using paired-end sequencing, the length of a specific fragment can be determined based on the position of the mapped sequence reads obtained from the ends of the fragment. Sequence reads used for specific analyses (e.g., determining the presence or absence of genetic variation) can be enriched or filtered against one or more selected fragment lengths or corresponding fragment length thresholds, as described in further detail herein.

[0116] Some length-based separation methods that can be used with the methods of this invention sometimes employ, for example, selective sequence tagging. The term "sequence tagging" refers to incorporating an identifiable, unique sequence into a nucleic acid or nucleic acid group. The term "sequence tagging" as used herein differs in meaning from the term "sequence tag" as described later herein. In this sequence tagging method, nucleic acids of various fragment sizes (e.g., short fragments) in a sample, including both long and short nucleic acids, are selectively sequence-tagged. This method typically involves a nucleic acid amplification reaction using a nested primer set, comprising internal and external primers. In some embodiments, one or both of the internal primers may be tagged to introduce a tag onto the target amplification product. External primers are typically not annealed to short fragments carrying the (internal) target sequence. Internal primers may anneal to the short fragments and produce an amplification product carrying both the tag and the target sequence. Typically, tagging of long fragments is inhibited by combinatorial mechanisms, including, for example, the inhibition of internal primer extension caused by prior annealing and extension of the external primers. Enrichment of tagged fragments can be achieved by any of a variety of methods, including, for example, digestion of single-stranded nucleic acids with exonucleases and amplification of tagged fragments using amplification primers specific to at least one tag.

[0117] Other length-based separation methods that can be used with the method of this invention involve precipitation of nucleic acid samples with polyethylene glycol (PEG). Examples of methods include those described in International Patent Application Publications WO2007 / 140417 and WO2010 / 115016. These methods typically require contacting the nucleic acid sample with PEG in the presence of one or more monovalent salts under conditions sufficient to precipitate large nucleic acids in large quantities without precipitating small (e.g., less than 300 nucleotides) nucleic acids in large quantities.

[0118] Other size-based enrichment methods that can be used with the methods described herein involve cyclization via ligation, such as using cyclases. Short nucleic acid fragments are generally more efficiently cyclized than long fragments. Non-cyclized sequences can be separated from cyclized sequences, and the enriched short fragments can be used for further analysis.

[0119] Determine the segment length

[0120] In some embodiments, the length of one or more nucleic acid fragments is determined. In some embodiments, the length of one or more target fragments is determined, thereby identifying one or more target fragment size classes. In some embodiments, the lengths of one or more target fragments and one or more reference fragments are determined, thereby identifying one or more target fragment length classes and one or more reference fragment length classes. In some embodiments, fragment length is determined by measuring the length of a probe hybridized with the fragment, which will be discussed in further detail below. The length of a nucleic acid fragment or probe can be determined using any method suitable for determining the length of a nucleic acid fragment in the art, such as mass-sensitive methods (e.g., mass spectrometry (e.g., matrix-assisted laser desorption / ionization (MALDI) mass spectrometry and electrospray (ES) mass spectrometry), electrophoresis (e.g., capillary electrophoresis), microscopy techniques (scanning tunneling microscopy, atomic force microscopy), length measurement using nanopores, and sequence-based length determination methods (e.g., paired-end sequencing). In some embodiments, fragment or probe length can be determined without using fragment charge-based separation methods. In some embodiments, fragment or probe length can be determined without using electrophoresis. In some embodiments, fragment or probe length can be determined without using nucleotide sequencing.

[0121] mass spectrometry

[0122] In some embodiments, mass spectrometry is used to determine the length of nucleic acid fragments. Mass spectrometry methods are commonly used to determine the mass of molecules (e.g., nucleic acid fragments). In some embodiments, the length of a nucleic acid fragment can be inferred from the mass of the fragment. In some embodiments, the range of nucleic acid fragment lengths can be predicted from the mass of the fragment. In some embodiments, the length of a nucleic acid fragment can be inferred from the mass of the probe hybridizing with the fragment, which is described in more detail below. In some embodiments, the presence of a target nucleic acid and / or reference nucleic acid of a given length can be verified by comparing the signal quality of the detection of the target fragment and / or reference fragment with the expected quality. Relative signal intensity, for example, a mass peak on the spectrum of a specific nucleic acid fragment and / or fragment length, can sometimes indicate the relative population of the fragment species among other nucleic acids in the sample (see, for example, Jurinke et al. (2004) Molmicrobiol. 15:1165-1167).

[0123] Mass spectrometry typically works by ionizing chemical compounds to produce charged molecules or molecular fragments, and then detecting their mass-to-charge ratio. A typical mass spectrometry process involves several steps, including (1) loading a sample onto a mass spectrometer and then vaporizing it, (2) ionizing the sample components by any of a variety of methods (e.g., bombarding with an electron beam) to generate charged particles (ions), (3) separating the ions in the analyzer according to their mass-to-charge ratio by an electromagnetic field, (4) detecting the ions (e.g., by quantitative methods), and (5) processing the ion signal into a mass spectrum.

[0124] Mass spectrometry methods are well known in the art (see, for example, Burlingame et al., Anal. Chem. 70:647R-716R (1998)), and include, for example, quadrupole mass spectrometry, ion trap mass spectrometry, time-of-flight mass spectrometry, gas chromatography-mass spectrometry, and tandem mass spectrometry, which can be used in the methods described herein. The fundamental process associated with mass spectrometry methods is the generation of gaseous ions from a sample and the detection of their mass. The movement of gaseous ions can be precisely controlled using electromagnetic fields generated in the mass spectrometer. The movement of ions in these electromagnetic fields is proportional to the ion's m / z (mass-to-charge ratio), which forms the basis for detecting m / z, and thus the mass of the sample. The movement of ions in these electromagnetic fields allows for the retention and focusing of ions, which contributes to the high sensitivity of mass spectrometry. During m / z measurements, ions are efficiently transported to a particle detector, which records the arrival of these ions. The amount of ions at each m / z is observed as peaks on a graph, where the x-axis is m / z and the y-axis is relative abundance. Different mass spectrometers have different resolution levels, that is, their ability to distinguish peaks between ions that are closely related in mass. Resolution is defined as R = m / δm, where m is the ion mass and δm is the mass difference between two peaks in the mass spectrum. For example, a mass spectrometer with a resolution of 1000 can distinguish an ion with m / z of 100.0 from an ion with m / z of 100.1.

[0125] Some mass spectrometry methods can employ various combinations of ion sources and mass analyzers, allowing for flexibility in the design of customized detection protocols. In some embodiments, the mass spectrometer can be programmed to sequentially or simultaneously transfer all ions from the ion source to the mass spectrometer. In some embodiments, the mass spectrometer can be programmed to select ions of a specific mass and transfer them into the mass spectrometer while blocking other ions.

[0126] Several types of mass spectrometers are available, or can be produced by employing multiple configurations. Generally, a mass spectrometer has the following main components: sample inlet, ion source, mass analyzer, detector, vacuum system, instrument control system, and data system. Differences in the sample inlet, ion source, and mass analyzer typically determine the type and capabilities of the instrument. For example, the inlet can be a capillary column liquid chromatography source, or it can be a direct probe or stage, such as those used in matrix-assisted laser desorption / sorption spectroscopy (MALS). Commonly used ion sources include, for example, electrospraying, including nanospraying and microspraying, or MALS. Mass analyzers include, for example, quadrupole mass filters, ion trap mass analyzers, and time-of-flight mass analyzers.

[0127] Ion formation is the starting point for mass spectrometry analysis. Several ionization methods are available, and the choice of ionization method depends on the sample being analyzed. For example, for the analysis of peptides, a relatively mild ionization method (such as electrospray ionization (ESI)) may be desired. In ESI, a solution containing the sample is passed through a fine needle at a high potential, which generates a strong electric field, producing a fine spray of highly charged droplets that are introduced into the mass spectrometer. Other ionization methods include, for example, fast atom bombardment (FAB), which uses a beam of high-energy neutral atoms to bombard a solid sample, causing desorption and ionization. Matrix-assisted laser desorption / ionization (MALDI) is a method in which a laser pulse is used to bombard a sample that has been crystallized in a matrix of ultraviolet-absorbing compounds (e.g., 2,5-dihydroxybenzoic acid, α-cyano-4-hydroxycitric acid, 3-hydroxypicolinic acid (3-HPA), ammonium dicitrate (DAC), and combinations thereof). Other ionization methods known in the art include, for example, plasma and glow discharge, plasma desorption ionization, resonant ionization, and secondary ionization.

[0128] Different mass analyzers can be paired with different ion sources. Different mass analyzers have different advantages, which is known in the art and as described herein. The choice of mass spectrometer and method for detection depends on the specific experiment; for example, a more sensitive mass analyzer may be used when detecting small amounts of ions produced. Several types of mass analyzers and mass spectrometry methods are described below.

[0129] Ion mobility mass (IM) spectroscopy is a gas-phase separation method. IM separates gas-phase ions based on their collision cross-sections and can be coupled with time-of-flight (TOF) mass spectrometry. IM-MS is discussed in more detail by Verbeck et al. in the Journal of Biomolecular Techniques (Vol. 13, No. 2, pp. 56-61).

[0130] Quadrupole mass spectrometry employs a quadrupole mass filter or analyzer. This type of mass analyzer consists of four rods arranged in two sets of two electrically connected rods. A combination of RF and DC voltages is applied to each pair of rods, creating a field that causes the ions to vibrate and move from the start to the end of the mass filter. These fields result in a high-pass mass filter in one pair of rods and a low-pass filter in the other pair. The overlap between the high-pass and low-pass filters leaves a defined m / z that passes through both filters and spans the length of the quadrupole. This m / z is selected and remains stable in the quadrupole mass filter, while all other m / z have unstable trajectories and do not remain in the mass filter. The mass spectrum is generated by tilting the applied field, thereby selecting the increased m / z to pass through the mass filter and reach the detector. Furthermore, the quadrupole can be configured to include and transport ions of all m / z by applying a field with only RF. This allows the quadrupole to function as a lens or focusing system in the mass spectrometer region where ion transport is required without mass filtration.

[0131] The quadrupole mass spectrometer described herein, as well as other mass spectrometers, can be programmed to analyze within a defined m / z or mass range. Given that the required mass range for nucleic acid fragments is known, in some examples, the mass spectrometer can be programmed to transmit ions projected within the correct mass range while excluding ions in higher or lower mass ranges. The ability to select a mass range reduces background noise in the analysis and thus increases the signal-to-noise ratio. Therefore, in some examples, the mass spectrometer is able to perform the separation step and detect and identify certain mass-distinguished nucleic acid fragments.

[0132] Ion trap mass spectrometry employs an ion trap mass analyzer. Typically, a field is applied so that all ions with a specific m / z are initially trapped and oscillate within the mass analyzer. Ions enter the ion trap from the ion source through a focusing device (e.g., an octagonal lens system). The ion trap appears in the trapping region before excitation and ejection through electrodes to the detector. Mass analysis is performed by sequentially applying a voltage that increases the oscillation in such a way that ions with the increased m / z are ejected from the trap and enter the detector. Unlike quadrupole mass spectrometry, all ions are retained in the field of the mass analyzer, except for those with a selected m / z. Control of the ion quantity can be achieved by varying the ion implantation time into the trap.

[0133] Time-of-flight mass spectrometry (TOF-MS) employs a time-of-flight mass analyzer. Typically, ions first gain a fixed amount of kinetic energy by being accelerated in an electric field (generated by a high voltage). After acceleration, the ions enter a field-free or "drift" region, where they move at a speed inversely proportional to their m / z. Therefore, ions with a low m / z move faster than ions with a high m / z. The time required for an ion to traverse the length of the field-free region is measured and used to calculate the ion's m / z.

[0134] Gas chromatography-mass spectrometry (GC-MS) typically enables real-time detection of targets. The gas chromatography (GC) section of the system separates the chemical mixture into multiple analytes, which are then identified and quantified by mass spectrometry (MS).

[0135] Tandem mass spectrometry can employ a combination of the aforementioned mass analyzers. A tandem mass spectrometer can use a first mass analyzer to separate ions based on their m / z, thus separating the ions of interest for further analysis. The separated ions of interest are then broken into fragments (called collision-activated dissociation or collision-induced dissociation), and these fragments are analyzed by a second mass analyzer. These types of tandem mass spectrometry systems are called tandem in a spatial system because the two mass analyzers are spatially separated, typically by a collision chamber. Tandem mass spectrometry systems also include tandem in a temporal system, where a single mass analyzer is used sequentially to separate ions, induce fragmentation, and then perform mass analysis.

[0136] Tandem mass spectrometers in the spatial domain have more than one mass analyzer. For example, a tandem quadrupole mass spectrometer system may have a first quadrupole mass filter, followed by a collision chamber, then a second quadrupole mass filter, and finally a detector. Another arrangement employs a quadrupole mass filter for the first mass analyzer and a time-of-flight mass analyzer for the second mass analyzer, wherein the collision chamber separates the two mass analyzers. Other tandem systems are known in the art, including reflective time-of-flight, tandem sector, and sector-quadrupole mass spectrometers.

[0137] A time-series tandem mass spectrometer has a mass analyzer that performs different functions at different times. For example, an ion trap mass spectrometer can be used to capture ions of all m / z values. A series of RF scans are applied, which eject ions of all m / z values ​​except for the ion of interest from the trap. After the m / z of interest has been separated, RF pulses are applied to induce collisions with gas molecules in the trap, thereby inducing ion fragmentation. The m / z values ​​of the fragmented ions are then measured by the mass analyzer. Ion cyclotron resonance instruments, also known as Fourier transform mass spectrometers, are examples of time-series tandem systems.

[0138] Several types of tandem mass spectrometry experiments can be performed by controlling the selected ions at each stage of the experiment. Different types of experiments employ different operating modes, sometimes referred to as "scans" of multiple mass analyzers. In a first example, called a mass spectrometric scan, the first mass analyzer and the collision chamber transmit all ions for mass analysis to a second mass analyzer. In a second example, called a product ion scan, the ions of interest are selected based on mass in the first mass analyzer and then fragmented in the collision chamber. The resulting ions are then mass-analyzed by scanning the second mass analyzer. In a third example, called a precursor ion scan, the first mass analyzer is scanned to sequentially transmit ions of analyzed mass into the collision chamber for fragmentation. The second mass analyzer selects the resulting ions of interest based on mass for transmission to the detector. Thus, the detector signal is the result of all precursor ions that can be fragmented into common product ions. Other experimental formats include neutral loss scans, where constant mass differences are accounted for in the mass scan.

[0139] For quantification, control measures can be employed that provide a signal relating to the amount of nucleic acid fragments (e.g., present or introduced nucleic acid fragments). Control measures that allow the conversion of relative quality signals into absolute quantities can be achieved by adding known amounts of quality tags or quality markers to each sample prior to the detection of the nucleic acid fragment. See, for example, Ding and Cantor (2003) PNAS US A. Mar 18; 100(6):3059-64. The quality signal can be normalized using any quality tag that does not interfere with the detection of the fragment. Typically, the standard is dissociative, differs from any of the molecular tags in the sample, and may have the same or different quality signatures.

[0140] Sometimes, separation steps can be used to remove salts, enzymes, or other buffer components from nucleic acid samples. Several methods well known in the art, such as chromatography, gel electrophoresis, or precipitation, can be used to remove samples. For example, size exclusion chromatography or affinity chromatography can be used to remove salts from samples. The choice of separation method can be based on the amount of sample. For example, microaffinity chromatography can be used when small amounts of sample are available or when miniaturized equipment is used. Furthermore, the necessity of a separation step, and the choice of separation method, can depend on the detection method used. Sometimes, salts can absorb energy from the laser in matrix-assisted laser desorption / ionization (MALAD) and result in low ionization efficiency. Therefore, sometimes the efficiency of MALAD and electrospray ionization can be improved by removing salts from the sample.

[0141] Electrophoresis

[0142] In some embodiments, electrophoresis is used to determine the length of nucleic acid fragments. In some embodiments, electrophoresis is not used to determine the length of nucleic acid fragments. In some embodiments, electrophoresis is used to determine the length of the corresponding probe (e.g., the corresponding trimmed probe described herein). In some embodiments, electrophoresis may also be used as a length-based separation method, as described herein. Any electrophoresis method known in the art (thereby separating nucleic acids by length) can be coupled with the methods provided herein, including but not limited to standard electrophoresis techniques and specialized electrophoresis techniques, such as capillary electrophoresis. Examples in the art exist of methods for separating nucleic acids and measuring the length of nucleic acid fragments using standard electrophoresis techniques. Non-limiting examples are provided herein. After running a nucleic acid sample in an agarose or polyacrylamide gel, the gel can be labeled (e.g., stained) with ethidium bromide (see Sambrook and Russell, *Molecular Cloning: A Laboratory Manual*, 3rd edition, 2001). The presence of a band of the same size as a standard control indicates the presence of a specific nucleic acid sequence length, and the length of the nucleic acid sequence of interest can then be detected and quantified based on a comparison of the band intensity with the control.

[0143] In some implementations, capillary electrophoresis is used to separate, identify, and sometimes quantify nucleic acid fragments. Capillary electrophoresis (CE) encompasses a family of related separation techniques that use narrow-pore fused silica capillaries to separate arrays of complexes of large and small molecules (e.g., nucleic acids of different lengths). Nucleic acid molecules can be separated based on their different charges, sizes, and hydrophobicities using high electric field strengths. Sample introduction is accomplished by immersing the end of the capillary into a sample vial and applying pressure, vacuum, or voltage. Depending on the type of capillary and electrolyte used, the CE techniques can be categorized into several separation techniques, any of which can be adapted to the methods provided herein. Non-limiting examples of these include capillary zone electrophoresis (CZE), also known as non-liquid phase CE (FSCE), capillary isoelectric focusing (CIEF), isovelocity electrophoresis (ITP), electrokinetic chromatography (EKC), micellar electrokinetic capillary chromatography (MECC or MEKC), microemulsion electrokinetic chromatography (MEEKC), non-aqueous capillary electrophoresis (NACE), and capillary electrochromatography (CEC).

[0144] Any apparatus, instrument, or machine capable of performing capillary electrophoresis can be coupled with the method described herein. The main components of a capillary electrophoresis system typically include sample vials, source vials and destination vials, a capillary, electrodes, a high-voltage power supply, a detector, data output, and manipulation devices. The source vials, destination vials, and capillary are filled with an electrolyte (e.g., an aqueous buffer solution). To introduce the sample, the capillary inlet is placed inside the vial containing the sample and then returned to the source vial (the sample is introduced into the capillary by capillary action, pressure, or siphon). The migration of the analyte (i.e., nucleic acids) is then initiated by an electric field applied between the source and destination vials and supplied to the electrodes via a high-voltage power supply. Cations or anions are introduced into the capillary in the same direction via electroosmosis. The analyte (i.e., nucleic acids) separates due to its electrophoretic mobility and is detected near the capillary outlet. The output of the detector is sent to a data output device and manipulation device, such as an integrator or computer. The data is then displayed as an electrophoresis plot, which reports the detector response over time. The isolated nucleic acids are shown as peaks appearing at different migration times in the electrophoresis plot.

[0145] Separation by capillary electrophoresis can be detected by several detection devices. Most commercially available systems use ultraviolet (UV) or UV-Vis absorbance as their primary detection mode. In these systems, a section of the capillary itself serves as the detection cell. On-tube detection allows for the detection of separated analytes without loss of resolution. Typically, the capillaries used in capillary electrophoresis are coated with polymers to increase stability. The portion of the capillary used for UV detection is usually optically transparent. The path length of the detection cell in capillary electrophoresis (approximately 50 micrometers) is much smaller than that of a conventional UV cell (approximately 1 cm). According to the Beer-Lambert law, the sensitivity of the detector is proportional to the path length of the cell. To increase sensitivity, the path length can be increased, although this results in a loss of resolution. The capillary itself can be expanded at the detection point to form a "bubble cell" with a longer path length, or additional tubing can be added at the detection point. However, both of these methods may reduce the resolution of the separation.

[0146] Fluorescence can also be used in capillary electrophoresis to detect naturally fluorescent or chemically modified samples containing fluorescent tags (e.g., labeled nucleic acid fragments or probes as described herein). This detection modality offers high sensitivity and improved selectivity for these samples. The method requires focusing a light beam onto the capillary. Laser-induced fluorescence can be used in CE systems, with detection limits as low as 10⁻¹⁸ to 10⁻²¹ mol. The sensitivity of this technique is attributed to the high incident light intensity and the ability to accurately focus the light onto the capillary.

[0147] Several capillary electrophoresis instruments are known in the art and can be coupled with the methods described herein. These include, but are not limited to, the CALIPER LAB CHIP GX (Caliper LifeSciences, Mountain View, California), the P / ACE 2000 series (Beckman Coulter, Bully, California), the HP G1600A CE (Hewlett-Packard, Palo Alto, California), the AGILENT 7100CE (Agilent Technologies, Santa Clara, California), and the ABI PRISM genetic analyzer (Applied Biosystems, Carlsbad, California).

[0148] Microscopic observation

[0149] In some embodiments, the length of the nucleic acid fragment is determined using an imaging-based method, such as microscopic observation. In some embodiments, the length of the corresponding probe (e.g., the corresponding trimmed probe described herein) is determined using an imaging-based method. In some embodiments, the fragment length can be determined by microscopic observation of a single nucleic acid fragment (see, for example, U.S. Patent Nos. 5,720,928). In some embodiments, the nucleic acid fragment is fixed in an elongated state to a surface (e.g., a modified glass surface), stained, and observed under a microscope. Images of the fragment can be collected and processed (e.g., length measurement). In some embodiments, the imaging and image analysis steps can be automated. Methods for directly observing nucleic acid fragments using a microscope are known in the art (see, for example, Lai et al. (1999) Nat Genet. 23(3):309-13; Aston et al. (1999) Trends Biotechnol. 17(7):297-302; Aston et al. (1999) Methods Enzymol. 303:55-73; Jing et al. (1998) Proc Natl Acad Sci USA. 95(14):8046-51; and U.S. Patent Nos. 5,720,928). Other microscopic observation methods that can be used with the methods described herein include, but are not limited to, scanning tunneling microscopy (STM), atomic force microscopy (ATM), scanning force microscopy (SFM), photon scanning microscopy (PSTM), scanning tunneling potential measurement (STP), magnetic force microscopy (MFM), scanning probe microscopy, scanning voltage microscopy, photoconductive atomic force microscopy, electrochemical scanning tunneling microscopy, electron microscopy, spin-polarized scanning tunneling microscopy (SPSTM), scanning thermal microscopy, scanning Joule expansion microscopy, photothermal microspectroscopy, etc.

[0150] In some implementations, scanning tunneling microscopy (STM) can be used to determine the length of nucleic acid fragments. STM methods typically produce atomic-level images of molecules, such as nucleic acid fragments. STM can be performed, for example, in air, water, ultra-high vacuum, various other liquid or gaseous environments, and at temperatures ranging from, for example, close to 0 Kelvin to several hundred degrees Celsius. Typically, the components of an STM system include a scanning tip, piezoelectrically controlled height and x, y scanners, coarse-tuning sample-tip control, a vibration isolation system, and a computer. STM methods are generally based on the concept of quantum tunneling. For example, when a conductive tip is brought close to the surface of a molecule (e.g., a nucleic acid fragment), applying a deviation (i.e., a potential difference) between the two allows electrons to open channels through the vacuum in between. The resulting tunneling current is a function of the tip position, the applied voltage, and the local density of states (LDOS) of the sample. Information is acquired by monitoring the position of the tip across the surface and can be displayed in image form. If the tip passes through the sample in the xy plane, changes in the surface height and density of states result in changes in the current. These changes can be mapped in the image. Sometimes, the change in current at a position can be detected on its own, or the corresponding tip height z corresponding to a constant current can be detected. These two modes are often referred to as constant height mode and constant current mode, respectively.

[0151] In some implementations, atomic force microscopy (AFM) can be used to determine the length of nucleic acid fragments. Generally, AFM is a high-resolution form of nanoscale microscopy. Typically, information about an object (e.g., a nucleic acid fragment) is collected by using a mechanical probe to "sense" the surface. The ability to move tiny but precise piezoelectric elements with electronic commands facilitates very accurate scanning. In some variations, an operating cantilever can be used to scan the potential. Typically, the components of an AFM system include:

[0152] A cantilever with a pointed tip (i.e., probe) at its end is used to scan the surface of a sample (e.g., a nucleic acid fragment). The cantilever is typically silicon or silicon nitride, with curved tooth tip radii on the nanometer scale. When the tip is brought close to the sample surface, the force between the tip and the sample causes the cantilever to deflect based on Hooke's law. Depending on the specific conditions, the forces detected in AFM include, for example, mechanical contact forces, van der Waals forces, capillary forces, chemical bonds, electrostatic forces, magnetic forces, Casimir forces, and solvent forces. Typically, deflection is detected using a laser spot reflected from the top surface of the cantilever into an array of photodiodes. Other available methods include optical interferometry, capacitive sensing, or piezoelectric AFM cantilever.

[0153] Nanopores

[0154] In some embodiments, the length of the nucleic acid fragment is determined using a nanopore. In some embodiments, the length of the corresponding probe (e.g., the corresponding modified probe described herein) is determined using a nanopore. A nanopore is typically a small pore or channel with a diameter on the order of 1 nanometer. Certain transmembrane cell proteins (such as α-hemolysin) can function as nanopores. In some embodiments, the nanopore may be synthetic (e.g., using a silicon platform). Immersing the nanopore in a conductive fluid and applying a voltage through the fluid results in a slight current, attributed to ion conduction through the nanopore. The amount of current flowing is sensitive to the size of the nanopore. When a nucleic acid fragment passes through the nanopore, the nucleic acid molecule blocks the nanopore to a certain extent, resulting in a change in current. The duration of this current change as the nucleic acid fragment passes through the nanopore can be measured. In some embodiments, the length of the nucleic acid fragment can be determined based on this detection result.

[0155] In some embodiments, the length of the nucleic acid fragment can be determined as a function of time. Sometimes, longer nucleic acid fragments may require a relatively long time to pass through a nanopore, while other times, shorter nucleic acid fragments may require a relatively short time. Therefore, in some embodiments, the relative length of the fragment can be determined based on the nanopore passage time. In some embodiments, an approximate or absolute fragment length can be determined by comparing the nanopore passage time of the target fragment and / or reference fragment with the passage time of a set of standards (i.e., having known lengths).

[0156] probe

[0157] In some embodiments, the fragment length is determined using one or more probes. In some embodiments, probes are designed such that each hybridizes to a nucleic acid of interest in the sample. For example, a probe may contain a polynucleotide sequence complementary to the nucleic acid of interest, or may contain a series of monomers that bind to the nucleic acid of interest. Probes may have any length suitable for hybridization (e.g., complete hybridization) to one or more nucleic acid fragments of interest. For example, probes may have any length that extends or prolongs the length of the nucleic acid fragment to which they hybridize. The length of a probe may be about 100 bp or longer. For example, the length of a probe may be at least about 200, 300, 400, 500, 600, 700, 800, 900, or 1000 bp.

[0158] In some embodiments, the probe may comprise a polynucleotide sequence complementary to the nucleic acid of interest, and one or more polynucleotide sequences not complementary to the nucleic acid of interest (i.e., non-complementary sequences). The non-complementary sequences may be located, for example, at the 5' and / or 3' ends of the probe. In some embodiments, the non-complementary sequences may comprise nucleotide sequences not present in the organism of interest and / or sequences that cannot hybridize with any sequence in the human genome. For example, the non-complementary sequences may be derived from any non-human genome known in the art, such as a non-mammalian genome, plant genome, fungal genome, bacterial genome, or viral genome. In some embodiments, the non-complementary sequence is derived from the PhiX 174 genome. In some embodiments, the non-complementary sequence may comprise modified or synthetic nucleotides that cannot hybridize with the complementary nucleotides.

[0159] The probes may be designed and synthesized according to methods known in the art and are described herein as oligonucleotides (e.g., capture oligonucleotides). The probes may also include any properties known in the art and described herein for use with oligonucleotides. The probes described herein may be designed to contain nucleotides (e.g., adenine (A), thymine (T), cytosine (C), guanine (G), and uracil (U)), modified nucleotides (e.g., pseudouridine, dihydrouridine, inosine (I), and 7-methylguanosine), synthetic nucleotides, degenerate bases (e.g., 6H,8H-3,4-dihydropyrimidino[4,5-c][1,2]oxazine-7-one (P), 2-amino-6-methoxyaminopurine (K), N6-methoxyadenine (Z), and hypoxanthine (I)), generic bases other than nucleotides, modified nucleotides, or synthetic nucleotides, or combinations thereof, and are generally designed to have an initial length longer than the fragment to which they hybridize.

[0160] In some embodiments, the probe comprises a plurality of monomers capable of hybridizing with any one of naturally occurring or modified forms of nucleotides such as adenine (A), thymine (T), cytosine (C), guanine (G), and uracil (U). In some embodiments, the probe comprises a plurality of monomers capable of hybridizing with at least three of adenine, thymine, cytosine, and guanine. For example, the probe may include monomer species capable of hybridizing to A, T, and C; A, T, and G; G, C, and T; or G, C, and A. In some embodiments, the probe comprises a plurality of monomers capable of hybridizing with all of adenine, thymine, cytosine, and guanine. For example, the probe may include monomer species capable of hybridizing with all of A, T, C, and G. In some embodiments, hybridization conditions (e.g., stringency) may be adjusted according to the methods described herein, for example, to facilitate hybridization of certain monomer species with different nucleotide species. In some embodiments, the monomers comprise nucleotides. In some embodiments, the monomers comprise naturally occurring nucleotides. In some embodiments, the monomers comprise modified nucleotides.

[0161] In some embodiments, the probe monomer comprises inosine. Inosine is a common nucleotide in tRNA and, in some examples, is capable of hybridizing with A, T, and C. Example 9 herein describes a method for determining nucleic acid fragment size using a polyinosine probe. In some embodiments, the polyinosine probe is hybridized to the nucleic acid fragment under low-strict or non-strict hybridization conditions (e.g., low temperature and / or high salt compared to the strict hybridization conditions described herein). In some embodiments, the nucleic acid fragment is treated with sodium bisulfite, which results in the deamination of unmethylated cytosine residues in the fragment to form uracil residues. In some embodiments, the nucleic acid fragment treated with sodium bisulfite is amplified (e.g., PCR amplification) and then treated with sodium bisulfite. In some embodiments, the nucleic acid fragment is ligated to a sequence containing a universal amplification primer site that does not contain cytosine residues. A complementary second strand can then be generated, for example, using a universal amplification primer and an extension reaction. Typically, uracil residues in the first strand generate complementary adenine residues in the second strand. Thus, a second strand without guanine residues can be generated. In some cases, the guanine-free complementary second strand can hybridize to the polyinosine probe under stringent hybridization conditions.

[0162] In some embodiments, the probe monomer comprises a universal base monomer. Generally, a universal base monomer is a nucleobase analog or a synthetic monomer capable of non-selectively hybridizing to each natural base (e.g., A, G, C, T). Therefore, probes containing universal base monomers can sometimes hybridize to nucleic acid fragments regardless of the nucleotide sequence. The common base may include, but is not limited to, 3-nitropyrrole, 4-nitroindole, 5-nitroindole, 6-nitroindole, 3-methyl-7-propynyl isoquinolone (PIM), 3-methyl isoquinolone (MICS), and 5-methyl isoquinolone (5MICS) (see, for example, Nichols et al. (1994) Nature 369, 492-493; Bergstrom et al. (1995) J. Am. Chem. Soc. 117, 1201-1209; Loakes and Brown (1994) Nucleic Acids Res. 22, 4039-4043; Lin and Brown (1992) Nucleic Acids Res. 20, 5149-5152; Lin and Brown (1989) Nucleic Acids Res. 17, 10383; Brown and Lin (1991) Carbohydrate Research 216, 129-139; Berger et al. (2000) Nucleic Acids Res. 28(15):2911–2914).

[0163] In some embodiments, the monomer of the probe comprises a non-nucleotide monomer. In some embodiments, the monomer comprises a subunit of a synthetic polymer. In some embodiments, the monomer comprises pyrrolidone. Pyrrolidone is a monomer of the synthetic polymer polypyrrolidone and in some examples is capable of hybridizing to all of A, T, G, and C.

[0164] In some embodiments, the method for determining fragment length includes the steps of contacting a nucleic acid fragment (e.g., a target fragment and / or a reference fragment) with a variety of probes capable of annealing to the fragment under annealing conditions, thereby generating multiple fragment-probe types, such as target-probe types and reference-probe types. Probe and / or hybridization conditions (e.g., stringency) may be optimized to facilitate complete or substantially complete fragment binding (e.g., high stringency). Complete or substantially complete fragment-probe hybridization generally comprises a double strand, wherein the fragment does not contain unhybridized portions, and the probe may contain unhybridized portions, as described in further detail below.

[0165] In some implementations, for example, when the probe length is longer than the fragment length, the target-probe type and / or reference-probe type may each include an unhybridized probe portion (i.e., a single-stranded probe portion; see, for example, Figure 12 The unhybridized probe portion may be located at one end of the probe (e.g., the 3' or 5' end of the probe) or at both ends of the probe (i.e., the 3' and 5' ends of the probe), and may contain any number of monomers. In some embodiments, the unhybridized probe portion may contain about 1 to about 500 monomers. For example, the unhybridized probe portion may contain about 5, 10, 20, 30, 40, 50, 100, 200, 300, or 400 monomers.

[0166] In some embodiments, unhybridized probe portions can be removed from the target-probe type and / or reference-probe type, thereby producing a trimmed probe. Removal of the unhybridized probe portions can be achieved by any method known in the art for cleaving and / or digesting polymers, for example, methods for cleaving or digesting single-stranded nucleic acids. The unhybridized probe portions can be removed from the 5' end and / or the 3' end of the probe. The methods may include chemical and / or enzymatic cleavage or digestion. In some embodiments, enzymes capable of cleaving phosphodiester bonds between nucleotide subunits of nucleic acids are used to remove the unhybridized probe portions. The enzymes described may include, but are not limited to, ribozymes (e.g., DNase I, RNase I), endonucleases (e.g., mung bean sprout nuclease, S1 ribozyme, etc.), restriction nucleases, exonucleases (e.g., exonucleases I, III, T, T7, λ), phosphodiesterases (e.g., phosphodiesterase II, calf spleen phosphodiesterase, snake venom phosphodiesterase, etc.), deoxyribonucleases (DNases), ribonucleases (RNases), flanking endonucleases, 5' nucleases, 3' nucleases, 3'-5' exonucleases, 5'-3' exonucleases, etc., or combinations thereof. The trimmed probe typically has the same or substantially the same length as the fragment it hybridizes to. Therefore, determining the length of the trimmed probe described herein provides a measurement of the length of the corresponding nucleic acid fragment. The length of the trimmed probe can be measured using any method known in the art or described herein for determining the length of a nucleic acid fragment. In some embodiments, the probe may contain detectable molecules or entities to facilitate detection and / or determination of length (e.g., fluorophores, radioisotopes, colorimetric agents, particles, enzymes, etc.). The length of the trimmed probe may be assessed using, or not using, the isolated product after removal of the unhybridized portion.

[0167] In some embodiments, the trimmed probe is dissociated (i.e., separated) from its corresponding nucleic acid fragment. The probe can be separated from its corresponding nucleic acid fragment using any method known in the art, including but not limited to thermal denaturation. The trimmed probe can be distinguished from the corresponding nucleic acid fragment by methods known in the art or described herein for labeling and / or separating molecular species in a mixture. For example, the probe and / or nucleic acid fragment may have detectable properties that allow the probe to be distinguished from the nucleic acid it hybridizes to. Non-limiting examples of detectable properties include optical properties, electrical properties, magnetic properties, chemical properties, and time and / or velocity through an opening of known size. In some embodiments, the probe and the sample nucleic acid fragment are physically separated from each other. Separation can be accomplished, for example, using a capture ligand, such as biotin or other affinity ligands, and a capture reagent, such as avidin, streptavidin, an antibody, or a receptor. The probe or nucleic acid fragment may contain a capture ligand that has specific binding activity to the capture reagent. For example, fragments from nucleic acid samples may be biotinylated or liganded to affinity ligands using methods well known in the art, and separated from the probe using a pull-down assay with streptavidin-coated beads. In some embodiments, capture ligands and capture reagents or any other components (e.g., mass tags) may be used to increase the mass of the nucleic acid fragment, thereby enabling it to be excluded from the mass range of the probe detected by the mass spectrometer. In some embodiments, the mass of the probe is increased by the addition of the monomer itself and / or the mass tag to shift the mass range away from the mass range of the nucleic acid fragment.

[0168] Nucleic acid library

[0169] In some embodiments, a nucleic acid library is a variety of polynucleotide molecules (e.g., nucleic acid samples) prepared, assembled, and / or modified for a specific process, non-limiting examples of which include immobilization, enrichment, amplification, cloning, detection, and / or use for nucleic acid sequencing on a solid phase (e.g., a solid support, such as a flow cell, beads). In some embodiments, the nucleic acid library is prepared before or during the sequencing process. Nucleic acid libraries (e.g., sequencing libraries) can be prepared using suitable methods known in the art. Nucleic acid libraries can be prepared via targeted or non-targeted preparation processes.

[0170] In some embodiments, the nucleic acid library is modified to include chemical portions (e.g., functional groups) configured to immobilize nucleic acids to a solid support. In some embodiments, the nucleic acid library is modified to include biomolecules (e.g., functional groups) and / or binding pair members configured to immobilize the library to a solid support. Non-limiting examples include thyroxine-binding globulin, steroid-binding proteins, antibodies, antigens, haptens, enzymes, hemagglutinins, nucleic acids, inhibitors, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding proteins, receptors, carbohydrates, oligonucleotides, polynucleotides, complementary nucleic acid sequences, and combinations thereof. Examples of specific binding pairs include, but are not limited to: anti-biotin moieties and biotin moieties; antigenic epitopes and antibodies or their immunologically active fragments; antibodies and haptens; digoxigenin moieties and anti-digoxigenin antibodies; luciferin moieties and anti-luciferin antibodies; operons and inhibitors; nucleases and nucleosides; lectins and polysaccharides; steroids and steroid-binding proteins; active compounds and active compound receptors; hormones and hormone receptors; enzymes and substrates; immunoglobulins and protein A; oligonucleotides or polynucleotides and their corresponding complements; and combinations thereof.

[0171] In some embodiments, the nucleic acid library is modified to include one or more polynucleotides of known composition. Non-limiting examples include identifiers (e.g., tags, index tags), capture sequences, labeled adaptors, restriction enzyme sites, promoters, enhancers, origins of replication, stem-loops, complementary sequences (e.g., primer binding sites, annealing sites), suitable integration sites (e.g., transposons, viral integration sites), modified nucleotides, and combinations thereof. Polynucleotides of known sequences can be inserted at suitable positions, such as the 5′ end, 3′ end, or within the nucleic acid sequence. Polynucleotides of known sequences can be the same or different sequences. In some embodiments, polynucleotides of known sequences are configured to hybridize with one or more oligonucleotides immobilized on a surface (e.g., the surface of a flow cell). For example, a known 5′ sequence of a nucleic acid molecule can hybridize with a first set of oligonucleotides, while a known 3′ sequence can hybridize with a second set of oligonucleotides. In some embodiments, the nucleic acid library may include chromosome-specific tags, capture sequences, labels, and / or adaptors. In some embodiments, the nucleic acid library includes one or more detectable markers. In some embodiments, one or more detectable markers may be incorporated into the 5′ end, 3′ end, and / or any nucleotide position of the nucleic acid in the library. In some embodiments, the nucleic acid library includes hybridized oligonucleotides. In some embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, the nucleic acid library includes hybridized oligonucleotide probes before immobilization on a solid phase.

[0172] In some embodiments, the polynucleotide of a known sequence includes a universal sequence. A universal sequence is a specific nucleotide sequence that integrates into two or more nucleic acid molecules or subsets of two or more nucleic acid molecules, wherein the universal sequence is identical with respect to all the molecules or subsets into which it is integrated. Universal sequences are typically designed to hybridize and / or amplify multiple different sequences using a single universal primer complementary to the universal sequence. In some embodiments, two (e.g., a pair) or more universal sequences and / or universal primers are used. Universal primers typically include a universal sequence. In some embodiments, an adaptor (e.g., a universal adaptor) includes a universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple nucleic acid substances or subsets thereof.

[0173] In some embodiments of nucleic acid library preparation (e.g., in certain sequencing processes of a synthesis procedure), the size of the nucleic acids is selected and / or fragmented to a length of several hundred base pairs or less (e.g., in library generation preparation). In some embodiments, library preparation is not required (e.g., when using ccfDNA).

[0174] In some embodiments, ligation-based library preparation methods are used (e.g., ILLUMINA TRUSEQ, Eminenthal, San Diego, CA). Ligation-based library preparation methods typically employ adaptors (e.g., methylated adaptors) designed to incorporate an index sequence at the initial ligation step and are generally used to prepare samples for single-read sequencing, paired-end sequencing, and multiplex sequencing. For example, sometimes nucleic acids (e.g., fragmented nucleic acids or ccfDNA) undergo end repair via fill-in reactions, endonuclease reactions, or combinations thereof. In some embodiments, the resulting blunt-end repaired nucleic acid can subsequently be extended by a single nucleotide, which is complementary to a single nucleotide overhang at the 3' end of the adaptor / primer. Any nucleotide can be used for the extended / overhanging nucleotide. In some embodiments, nucleic acid library preparation includes linking adaptor oligonucleotides. These adaptor oligonucleotides are typically complementary to flow cell anchors and are sometimes used to immobilize the nucleic acid library to a solid support, such as the inner surface of a flow cell. In some embodiments, the adaptor oligonucleotide includes an identifier, one or more sequencing primer hybridization sites (e.g., sequences complementary to universal sequencing primers, single-end sequencing primers, paired-end sequencing primers, multiplex sequencing primers, etc.) or combinations thereof (e.g., adaptor / sequencing, adaptor / identifier, adaptor / identifier / sequencing).

[0175] The identifier may be a suitable detectable tag incorporating or conjugating a nucleic acid (e.g., a polynucleotide), which allows for the detection and / or identification of nucleic acids including the identifier. In some embodiments, the identifier is incorporating or conjugating the nucleic acid during sequencing methods (e.g., by polymerase). Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indexes or barcodes, radioactive tags (e.g., isotopes), metallic tags, chemiluminescent tags, phosphorescent tags, fluorescence quenchers, dyes, proteins (e.g., enzymes, antibodies or portions thereof, linkers, members of binding pairs), and combinations thereof. In some embodiments, the identifier (e.g., a nucleic acid index or barcode) is a unique, known, and / or identifiable sequence of a nucleotide or nucleotide analogue. In some embodiments, the identifier is six or more consecutive nucleotides. Many fluorophores with various different excitation and emission spectra are available. Any suitable type and / or number of fluorophores can be used as identifiers. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, or fifty or more different identifiers are used in the methods described herein (e.g., nucleic acid detection and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are linked to each nucleic acid in the library. Identifier detection and / or quantification can be performed by suitable methods or apparatus, non-limiting examples of which include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, illuminometer, fluorometer, spectrophotometer, suitable gene chip or microarray analysis, Western blotting, mass spectrometry, chromatography, cellular fluorescence analysis, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning flow cytometry, affinity chromatography, manual batch separation, electric field suspension, suitable nucleic acid sequencing methods and / or nucleic acid sequencing apparatus, and combinations thereof.

[0176] In some implementations, transposon-based library preparation methods are used (e.g., EPICENTRE NEXTERA, Epicentre, Madison, Wisconsin). Transposon-based methods typically use in vitro translocation to similar fragments or tagged DNA (often allowing for the inclusion of platform-specific tags and optional barcodes) in a single-tube reaction to prepare a sequencer-ready library.

[0177] In some embodiments, a nucleic acid library or a portion thereof is amplified (e.g., by PCR-based methods). In some embodiments, sequencing methods include amplifying a nucleic acid library. The nucleic acid library may be amplified before or after immobilization onto a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification includes the process of amplifying or increasing (e.g., in the nucleic acid library) the amount of a present nucleic acid template and / or its complement, said process being achieved by generating one or more copies of the template and / or its complement. Amplification may be performed by suitable methods. The nucleic acid library may be amplified by thermal cycling or by isothermal amplification. In some embodiments, rolling circle amplification is used. In some embodiments, amplification occurs on a solid support (e.g., within a flow cell) where a nucleic acid library or a portion thereof is immobilized. In some sequencing methods, the nucleic acid library is added to a flow cell and immobilized by hybridization with an anchor under suitable conditions. Such nucleic acid amplification is generally referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or part of the amplification product is synthesized by extension starting from immobilized primers. Solid-phase amplification reactions are similar to standard solution-phase amplification, except that at least one of the amplified oligonucleotides (e.g., primers) is immobilized on a solid support.

[0178] In some embodiments, solid-phase amplification includes a nucleic acid amplification reaction comprising only one oligonucleotide primer immobilized on a surface. In some embodiments, solid-phase amplification includes multiple different immobilized oligonucleotide primer materials. In some embodiments, solid-phase amplification may include a nucleic acid amplification reaction comprising one oligonucleotide primer immobilized on a solid surface and a second different oligonucleotide primer in solution. Multiple different immobilized or solution primers may be used. Non-limiting examples of solid-phase nucleic acid amplification reactions include interfacial amplification, bridging amplification, emulsion PCR, WildFire amplification (e.g., U.S. Patent Application US20130012399), and combinations thereof.

[0179] sequencing

[0180] In some embodiments, nucleic acids (e.g., nucleic acid fragments, sample nucleic acids, cell-free nucleic acids) may be sequenced. In some embodiments, a full or nearly full sequence is obtained, and sometimes a partial sequence is obtained. In some embodiments, nucleic acids are not sequenced when the methods described herein are performed, and the sequence of nucleic acids is not determined by sequencing methods. In some embodiments, fragment length is determined by sequencing methods. In some embodiments, fragment length is not determined by sequencing methods. Sequencing, localization, and correlation analysis methods are as described herein or are known in the art (e.g., U.S. Patent Application Publication US2009 / 0029377, incorporated herein by reference). Certain aspects of such methods are described below.

[0181] In some implementations, fragment length is determined using sequencing methods. In some implementations, fragment length is determined using a paired-end sequencing platform. This platform involves sequencing the paired ends of a nucleic acid fragment. Generally, sequences corresponding to the paired ends of the fragment can be mapped to a reference genome (e.g., a reference human genome). In some implementations, both ends are sequenced with read lengths sufficient, individually, for each fragment end to be mapped to the reference genome. Examples of paired-end sequence read lengths are shown below. In some implementations, all or part of the sequence reads can be mapped to the reference genome without mismatch. In some implementations, each read is mapped independently. In some implementations, information from both sequence reads (i.e., from each end) is included in the mapping process. For example, fragment length can be determined by calculating the differences between genomic equivalents of the paired-end reads assigned to each mapping.

[0182] In some implementations, fragment length can be determined using sequencing methods to obtain a complete or substantially complete nucleotide sequence of the fragment. These sequencing methods include platforms that produce relatively long read lengths (e.g., Roche 454, Ion Torrent, single-molecule sequencing (Pacific Biosciences), real-time SMRT technology, etc.).

[0183] In some embodiments, some or all nucleic acids in the sample are enriched and / or amplified before or during sequencing (e.g., non-specific, such as by PCR-based methods). In some embodiments, a specific portion or subset of nucleic acids in the sample is enriched and / or amplified before or during sequencing. In some embodiments, a portion or subset of a preselected set of nucleic acids is randomly sequenced. In some embodiments, nucleic acids in the sample are not enriched and / or amplified before or during sequencing.

[0184] As used herein, a “reading” (i.e., “a reading”, “sequence reading”) is a short nucleotide sequence generated by any sequencing method described herein or known in the art. Readings can be generated from one end of a nucleic acid fragment (“single-end reading”), while sometimes they are generated from both ends of a nucleic acid fragment (e.g., paired-end reading, double-end read).

[0185] The length of sequence reads is typically related to the specific sequencing technology. For example, high-throughput methods provide sequence reads ranging in size from tens to hundreds of base pairs (bp). Nanopore sequencing, for instance, provides sequence reads ranging in size from tens to hundreds to thousands of base pairs. In some embodiments, sequence reads are arithmetic mean, median, average, or absolute lengths of approximately 15 bp to approximately 900 bp. In some embodiments, the sequence reads are arithmetic mean, median, average, or absolute lengths of approximately 1000 bp or longer.

[0186] In some embodiments, the nominal, average, arithmetic mean, or absolute length of a single-end reading is sometimes about 1 nucleotide to about 500 consecutive nucleotides, about 15 consecutive nucleotides to about 50 consecutive nucleotides, about 30 consecutive nucleotides to about 40 consecutive nucleotides, and sometimes about 35 consecutive nucleotides or about 36 consecutive nucleotides. In some embodiments, the nominal, average, arithmetic mean, or absolute length of a single-end reading is about 20 to about 30 bases, or about 24 to about 28 bases. In some embodiments, the nominal, average, arithmetic mean, or absolute length of the single-end reading is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, or 49 bases.

[0187] In some implementations, the nominal, average, arithmetic mean, or absolute length of the paired-end readings is sometimes about 10 consecutive nucleotides to about 25 consecutive nucleotides (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long), about 15 consecutive nucleotides to about 20 consecutive nucleotides, and sometimes about 17 consecutive nucleotides, about 18 consecutive nucleotides, about 20 consecutive nucleotides, about 25 consecutive nucleotides, about 36 consecutive nucleotides, or about 45 consecutive nucleotides.

[0188] Readings are typically representations of nucleotide sequences in physiological nucleic acids. For example, sequences are described in readings using ATGC, where "A" represents adenine nucleotide, "T" represents thymine nucleotide, "G" represents guanine nucleotide, and "C" represents cytosine nucleotide. Sequence readings obtained from the blood of pregnant women can be readings of a mixture of fetal and maternal nucleic acids. Mixtures of relatively short readings can be transformed into representations of genomic nucleic acids in pregnant women and / or fetuses using the methods described herein. Mixtures of relatively short readings can be transformed into representations of, for example, copy number variations (e.g., maternal and / or fetal copy number variations), genetic variations, or aneuploidy. Readings of mixtures of maternal and fetal nucleic acids can be transformed into representations of complex chromosomes or segments thereof containing features of one or both maternal and fetal chromosomes. In some embodiments, “obtaining” nucleic acid sequence readings from a subject sample and / or “obtaining” nucleic acid sequence readings from biological samples of one or more reference individuals can directly involve sequencing nucleic acids to obtain sequence information. In some embodiments, “obtaining” can involve receiving sequence information directly obtained from other nucleic acids.

[0189] In some implementations, partial sequencing of the genome is sometimes expressed as the amount of genome covered by the determined nucleotide sequence (e.g., less than 1 "fold" coverage). When sequencing the genome with approximately 1 fold coverage, the reading represents approximately 100% of the genome's nucleotide sequence. Genome sequencing can also be performed using redundancy, where a given region of the genome can be covered by two or more readings or overlapping readings (e.g., greater than 1 "fold" coverage). In some implementations, the genome is sequenced with a coverage of about 0.1-100 times, about 0.2-20 times, or about 0.2-1 times (e.g., about 0.02-, 0.03-, 0.04-, 0.05-, 0.06-, 0.07-, 0.08-, 0.09-, 0.1-, 0.2-, 0.3-, 0.4-, 0.5-, 0.6-, 0.7-, 0.8-, 0.9-, 1-, 2-, 3-, 4-, 5-, 6-, 7-, 8-, 9-, 10-, 15-, 20-, 30-, 40-, 50-, 60-, 70-, 80-, 90- times).

[0190] In some embodiments, genome coverage or sequence coverage is proportional to the total sequence read count. For example, experiments that generate and / or analyze a larger number of sequence read counts are typically associated with a higher level of sequence coverage. Experiments that generate and / or analyze fewer sequence read counts are typically associated with a lower level of sequence coverage. In some embodiments, sequence coverage and / or sequence read counts may be reduced without significantly reducing the accuracy (e.g., sensitivity and / or specificity) of the methods described herein. A significant reduction in accuracy may be a reduction of about 1% to about 20% compared to a method that does not use reduced sequence read counts. For example, a significant reduction in accuracy may be a reduction of about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, or more. In some embodiments, sequence coverage and / or sequence read counts are reduced by about 50% or more. For example, sequence coverage and / or sequence read counts may be reduced by about 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more. In some embodiments, sequence coverage and / or sequence read counts are reduced by about 60% to about 85%. For example, sequence coverage and / or sequence read counts may be reduced by about 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, or 84%. In some embodiments, sequence coverage and / or sequence read counts may be reduced by removing certain sequence reads. In some examples, sequence reads from fragments longer than a specific length (e.g., fragments longer than about 160 bases) are removed.

[0191] In some implementations, a subset of readings is selected for analysis, while sometimes, portions of the readings are removed from the analysis. In some cases, selecting a subset of readings can enrich the types of nucleic acids (e.g., fetal nucleic acids). Enrichment of readings from fetal nucleic acids, for example, typically improves the accuracy of the methods described herein (e.g., fetal aneuploidy detection). However, selecting and removing readings from the analysis typically reduces the accuracy of the methods described herein (e.g., due to increased variability). Therefore, without theoretical limitations, generally speaking, in methods that involve selecting and / or removing readings (e.g., fragments from a specific size range), a trade-off is required between the increased accuracy associated with enrichment of fetal readings and the decreased accuracy associated with a reduced number of reads. In some implementations, the method includes selecting a subset of readings that enrich readings from fetal nucleic acids without significantly reducing the accuracy of the method. Regardless of this apparent trade-off, it has been determined that, as described herein, employing a subset of nucleotide sequence readings (e.g., readings from relatively short fragments) can improve or maintain the accuracy of fetal genetic analysis. For example, in some implementations, about 80% or more of the nucleotide sequence readings may be discarded, and the sensitivity and specificity values ​​may be maintained at values ​​similar to those of a method that does not discard the nucleotide sequence readings.

[0192] In some embodiments, a subset of nucleic acid fragments is selected prior to sequencing. In some embodiments, hybridization-based techniques (e.g., using oligonucleotide arrays) may be used for initial selection of nucleic acid sequences from certain chromosomes (e.g., sex chromosomes and / or chromosomes not involved in aneuploidy testing, potentially aneuploid chromosomes, and other chromosomes). In some embodiments, nucleic acids may be separated by size (e.g., by gel electrophoresis, size exclusion chromatography, or by microfluidic-based methods), while in some examples, fetal nucleic acids may be enriched by selecting those with lower molecular weights (e.g., less than 300 base pairs, less than 200 base pairs, less than 150 base pairs, less than 100 base pairs). In some embodiments, fetal nucleic acids may be enriched by suppressing maternal background nucleic acids (e.g., by adding formaldehyde). In some embodiments, a portion or subset of the preselected nucleic acid fragment set is randomly sequenced. In some embodiments, the nucleic acids are amplified prior to sequencing. In some embodiments, a portion or subset of the nucleic acids is amplified prior to sequencing.

[0193] In some embodiments, a nucleic acid sample from a single individual is sequenced. In some embodiments, the nucleic acids of each of two or more samples are sequenced, wherein the samples are from one individual or from different individuals. In some embodiments, nucleic acid samples from two or more biological samples (where each biological sample is from one or more individuals) are collected, and the collection is sequenced. In later embodiments, the nucleic acid samples from each biological sample are often identified by one or more unique identifiers or identifier tags.

[0194] In some implementations, the sequencing method employs unique identifiers that allow for multiple sequence reactions during the sequencing process. The greater the number of unique identifiers, the more samples and / or chromosomes can be detected; for example, multiplexing can be performed during the sequencing process. The sequencing process can be performed using any suitable number of unique identifiers (e.g., 4, 8, 12, 24, 48, 96, or more).

[0195] Sequencing processes sometimes use a solid phase, which may include a flow cell on which nucleic acids from a library can be bound and on which reagents can flow and contact the bound nucleic acids. Flow cells sometimes include flow cell channels, and the use of identifiers facilitates the analysis of the number of samples in each channel. Flow cells are typically constructed to retain and / or allow reagent solutions to pass through in an orderly manner with bound analytes. Flow cells are typically planar, optically transparent, usually in the millimeter or sub-millimeter range, and often have channels or pathways in which analyte / reagent interactions occur. In some implementations, the number of samples that can be analyzed in a given flow cell channel often depends on the number of unique identifiers used in library preparation and / or probe design. Single flow cell channel. Multiplexing of 12 identifiers, for example, may allow simultaneous analysis of 96 samples (e.g., the number of wells in a 96-well microplate) in an 8-channel flow cell. Similarly, multiplexing of 48 identifiers, for example, may allow simultaneous analysis of 384 samples (e.g., the number of wells in a 384-well microplate) in an 8-channel flow cell. Non-limiting examples of commercially available multiplex sequencing kits include Emindek's multiplex sample preparation oligonucleotide kit and multiplex sequencing primer and PhiX control kit (e.g., Emindek catalog numbers PE-400-1001 and PE-400-1002, respectively).

[0196] Any suitable method for sequencing nucleic acids can be used, non-limiting examples of which include Maxim and Gilbert, chain termination methods, sequencing by synthesis, ligation sequencing, mass spectrometry sequencing, microscopy-based techniques, and combinations thereof. In some embodiments, first-generation sequencing technologies such as Sanger sequencing methods, including automated Sanger sequencing methods (including microfluidic Sanger sequencing), can be used in the methods of the present invention. In some embodiments, other sequencing technologies, including nucleic acid imaging techniques (such as transmission electron microscopy (TEM) and atomic force microscopy (AFM)), are also used herein. In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods typically involve clonal amplification of a DNA template or a single DNA molecule, sometimes sequenced in a flow cell in a massively parallel manner. Next-generation (e.g., second and third generation) sequencing technologies (capable of sequencing DNA in massively parallel manner) can be used in the methods described herein and are collectively referred to herein as “massively parallel sequencing” (MPS). In some embodiments, MPS sequencing methods employ a targeted approach, wherein a specific chromosome, gene, or region of interest is the sequence. In some embodiments, a non-targeted approach is used, wherein most or all nucleic acids in a sample are sequenced, amplified, and / or randomly captured.

[0197] In some embodiments, targeted enrichment, amplification, and / or sequencing methods are used. Targeting methods typically isolate, select, and / or enrich subsets of nucleic acids in a sample for further processing using sequence-specific oligonucleotides. In some embodiments, a library of sequence-specific oligonucleotides is employed to target (e.g., hybridize) one or more sets of nucleic acids in a sample. Sequence-specific oligonucleotides and / or primers are typically selective for specific sequences (e.g., unique nucleic acid sequences) present in one or more regions of interest within chromosomes, genes, exons, introns, and / or regulatory regions. Any suitable method or combination of methods can be used to enrich, amplify, and / or sequence one or more target nucleic acid subsets. In some embodiments, target sequences are isolated and / or enriched by capturing to a solid phase (e.g., flow cell, beads) using one or more sequence-specific anchors. In some embodiments, target sequences are enriched and / or amplified using sequence-specific primers and / or primer sets via polymerase-based methods (e.g., PCR-based methods, via any suitable polymerase-based extension). Sequence-specific anchors are typically used as sequence-specific primers.

[0198] MPS sequencing sometimes uses sequencing via synthesis and certain imaging methods. Nucleic acid sequencing technologies that can be used in the methods described herein are synthetic sequencing and sequencing based on reversible terminators (such as Illumina's Genome Analyzer and Genome Analyzer II; HISEQ 2000; HISEQ 2500 (Illumina, San Diego, California)). This technology allows for the parallel sequencing of millions of nucleic acid (such as DNA) fragments. In one embodiment of this sequencing technology, a flow cell containing an optically clear slide with eight individual channels is used, the surface of which is bound to oligonucleotide anchors (such as adaptor primers). The flow cell is typically constructed to retain and / or allow reagent solutions to pass through in an orderly manner, binding the analyte. The flow cell is typically planar, optically clear, usually in the millimeter or sub-millimeter range, and often has channels or pathways in which analyte / reagent interactions occur.

[0199] In some embodiments, synthetic sequencing involves repeatedly adding (e.g., covalently adding) nucleotides to primers or a pre-existing nucleic acid chain in a template-guided manner. Each repeatedly added nucleotide is detected, and the process is repeated multiple times until a sequence of the nucleic acid chain is obtained. The length of the obtained sequence depends in part on the number of addition and detection steps performed. In some embodiments of synthetic sequencing, one, two, three, or more nucleotides of the same type (e.g., A, G, C, or T) are added and detected in a nucleotide addition round. Nucleotides can be added by any suitable method (e.g., enzymatic or chemical). For example, in some embodiments, polymerases or ligases add nucleotides to primers or a pre-existing nucleic acid chain in a template-guided manner. In some embodiments of synthetic sequencing, different types of nucleotides, nucleotide analogs, and / or identifiers are used. In some embodiments, reversible terminators and / or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In some embodiments, synthetic sequencing includes cutting (e.g., cutting and removing identifiers) and / or washing steps. In some embodiments, the addition of one or more nucleotides is detected by methods described herein or known in the art. Non-limiting examples include any suitable imaging device, a suitable camera, a digital camera, a CCD (charge-coupled device) based imaging device (e.g., a CCD camera), a CMOS (complementary metal-oxide-semiconductor) based imaging device (e.g., a CMOS camera), a photodiode (e.g., a photomultiplier tube), an electron microscope, a field-effect transistor (e.g., a DNA field-effect transistor), an ISFET ion sensor (e.g., a CHEMFET sensor), and combinations thereof. Other sequencing methods that can be used to perform the methods described herein include digital PCR and hybridization sequencing.

[0200] Other sequencing methods that can be used to perform the methods described herein include digital PCR and hybridization sequencing. Digital polymerase chain reaction (digital PCR or dPCR) can be used to directly identify and quantify nucleic acids in a sample. In some embodiments, digital PCR can be performed in an emulsion. For example, individual nucleic acids are isolated in, for example, a microfluidic device and each nucleic acid is amplified individually by PCR. Nucleic acids are isolated such that no more than one nucleic acid is contained in each well. In some embodiments, different probes can be used to distinguish multiple alleles (e.g., fetal alleles and maternal alleles). Alleles can be counted to determine copy number.

[0201] In some embodiments, hybridization sequencing can be used. The method involves contacting multiple polynucleotide sequences with multiple polynucleotide probes, each of which is optionally attached to a substrate. In some embodiments, the substrate may be a plane with an array of known nucleotide sequences. The pattern of hybridization with the array can be used to determine the polynucleotide sequences present in the sample. In some embodiments, each probe is attached to a bead (such as a magnetic bead). Hybridization with the bead can be identified and used to identify multiple polynucleotide sequences in the sample.

[0202] In some implementations, nanopore sequencing can be used in the methods described herein. Nanopore sequencing is a single-molecule sequencing technology whereby a single nucleic acid molecule (such as DNA) is directly sequenced as it passes through a nanopore.

[0203] The methods described herein can be used to obtain nucleic acid sequencing reads by employing suitable non-human MPS methods, systems, or technology platforms. Non-limiting examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, Helicos True single-molecule sequencing, particle-to-torrent (PTT) and ion-semiconductor (ITS) sequencing (e.g., developed by Life Technologies), technologies based on WildFire, 5500, 5500xl W and / or 5500xl W genetic analyzers (e.g., developed and marketed by Life Technologies, US patent application US20130012399); polymerase cloning sequencing, pyrosequencing, massively parallel signature sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen systems and methods, nanopore-based platforms, chemically sensitive field-effect transistor (CHEMFET) arrays, electron microscopy-based sequencing (e.g., developed by ZS Genetics, Halcyon Molecular), and nanosphere sequencing.

[0204] In some embodiments, chromosome-specific sequencing is performed. In some embodiments, chromosome-specific sequencing is performed using DANSR (Digital Analysis of Selected Regions). Digital analysis of selected regions can simultaneously quantify hundreds of sites by using interfering 'bridging' oligonucleotides to form PCR templates through cfDNA-dependent linkage of two site-specific oligonucleotides. In some embodiments, chromosome-specific sequencing is performed by generating a library enriched with chromosome-specific sequences. In some embodiments, sequence reads are obtained only for selected chromosome sets. In some embodiments, sequence reads are obtained only for chromosomes 21, 18, and 13.

[0205] Mapped readings

[0206] The number of sequence reads that can be mapped to a specific nucleic acid region (e.g., a chromosome, a part, or a segment thereof) is called a count. Any suitable mapping method (e.g., procedure, algorithm, program, software, module, etc., or a combination thereof) can be used. Some aspects of mapping methods are described below.

[0207] Mapped nucleotide sequence reads (i.e., sequence information of fragments at unknown physical genomic sites) can be performed in various ways, typically involving aligning the obtained sequencing reads with matching sequences in a reference genome. In this alignment, the sequence reads are usually compared to a reference sequence; those that are aligned are referred to as "mapped," "mapped sequence reads," or simply "mapped readings." In some embodiments, mapped sequence reads are referred to as "hit" or "count." In some embodiments, mapped sequence reads are aggregated and assigned to specific portions based on various parameters, as detailed below.

[0208] As used herein, the terms "alignment" and "alignment" refer to two or more nucleic acid sequences that can be identified as a match (e.g., 100% identical) or a partial match. Alignment can be performed manually or by a computer (e.g., software, program, module, or algorithm), and non-limiting examples include the Nucleotide Data Effective Local Alignment (ELAND) computer program, which is part of the Illumina genome analysis workflow. Alignment of sequence readings can be 100% sequence match. In some cases, alignment is less than 100% sequence match (i.e., imperfect match, partial match, partial alignment). In some embodiments, alignment is approximately 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% match. In some embodiments, alignment includes mismatches. In some implementations, the alignment includes 1, 2, 3, 4, or 5 mismatches. Two or more sequences can be aligned using either strand. In some implementations, the nucleic acid sequence is aligned with the reverse complementary strand of another nucleic acid sequence.

[0209] Various computer methods can be used to map sequence reads to portions. Non-limiting examples of computer algorithms that can be used for sequence alignment include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE 1, BOWTIE 2, ELAND, MAQ, probe MATCH, SOAP, or SEQMAP, or variations thereof, or combinations thereof. In some embodiments, sequence reads can be aligned to sequences in a reference genome. In some embodiments, sequence reads can be obtained from and / or aligned to sequences in nucleic acid databases known in the art, including, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (Japan DNA Database). BLAST or similar tools can be used to search for identical sequences against a sequence database. Search hits can then, for example, be used to sort identical sequences into appropriate portions (as described below).

[0210] In some implementations, reads may be uniquely or non-uniquely mapped to portions of a reference genome. A read that aligns to a single sequence in the reference genome is called a “unique mapping.” A read that aligns to two or more sequences in the reference genome is called a “non-unique mapping.” In some implementations, non-uniquely mapped reads are removed from further analysis (e.g., quantification). In some implementations, a small degree of mismatch (0-1) may indicate a possible single nucleic acid polymorphism between the reference genome and the mapped reads from an individual sample. In some implementations, no mismatches may allow reads to be mapped to a reference sequence.

[0211] As used herein, the term "reference genome" can refer to a genome of any organism or virus in which any part or all of it is specifically known, sequenced, or characterized, and can be used as a reference for identifying target sequences. For example, reference genomes used for human subjects and many other organisms are available from the National Center for Biotechnology Information, www.ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus expressed in a nucleic acid sequence. Reference sequences or reference genomes used herein are often assembled or partially assembled genome sequences from one or more individuals. In some embodiments, the reference genome is an assembled or partially assembled genome sequence from one or more human individuals. In some embodiments, the reference genome includes sequences assigned to chromosomes.

[0212] In some embodiments, when the sample nucleic acid is derived from a pregnant woman, the reference sequence may not be derived from the fetus, the mother, or the father, and is thus referred to herein as an "external reference." In some embodiments, a maternal reference may be prepared and used. When preparing a reference from a pregnant woman based on an external reference ("maternal reference sequence"), readings of DNA from the pregnant woman that are substantially free of fetal DNA are typically mapped to and assembled from the external reference sequence. In some embodiments, the external reference is derived from DNA from an individual substantially of the same ethnicity as the pregnant woman. The maternal reference sequence may not completely cover the maternal genomic DNA (e.g., it may cover approximately 50%, 60%, 70%, 80%, 90%, or more of the maternal genomic DNA), and the maternal reference may not be a perfect match to the maternal genomic DNA sequence (e.g., the maternal reference sequence may contain multiple mismatches).

[0213] In some implementations, mappability is evaluated for genomic regions (e.g., portions, genomic parts, or sub-regions). Mappability is the ability of nucleotide sequence reads to clearly align to a portion of a reference genome, typically with multiple to a specific number of mismatches, including, for example, 0, 1, 2, or more mismatches. For a given genomic region, the expected mappability can be calculated using a sliding window method with a predetermined read length and averaged to obtain a mappability value at the read level. Extended genomic regions that include unique nucleotide sequences sometimes have high mappability values.

[0214] Part

[0215] In some implementations, mapped sequence reads (i.e., sequence tags) are grouped together according to various parameters and assigned to specific portions (e.g., portions of a reference genome). Typically, individually mapped sequence reads can be used to identify portions present in a sample (e.g., the presence, deletion, or abundance of a portion). In some implementations, the abundance of a portion is an indicator of the abundance of large sequences (e.g., chromosomes) in the sample. The term "portion" may also refer to "genomic segment," "box," "region," "segmentation," "portion of a reference genome," "portion of a chromosome," or "genomic portion." In some implementations, a portion is an entire chromosome, a chromosomal segment, a reference genome segment, a segment spanning multiple chromosomes, a segment spanning multiple chromosomes, and / or combinations thereof. In some implementations, portions are predefined based on specific parameters (e.g., indicators). In some implementations, portions are arbitrarily or non-arbitrarily defined based on genome partitioning (e.g., partitioning based on size, GC content, contiguous regions, contiguous regions of arbitrarily defined size, etc.). In some implementations, portions are selected from discontinuous genome bins, genome bins with contiguous sequences of predetermined lengths, variable-size bins, point-based views of smooth coverage maps, and / or combinations thereof.

[0216] In some embodiments, a portion is defined based on one or more parameters, including, for example, sequence length or specific characteristics. Portions can be selected, screened, and / or removed from consideration using any suitable criteria known in the art or described herein. In some embodiments, a portion is based on a specific length of the genome sequence. In some embodiments, the method may include analyzing sequence reads from multiple mappings for multiple portions. Portions may have substantially the same length or portions may have different lengths. In some embodiments, portions are approximately the same length. In some embodiments, portions of different lengths are adjusted or weighted. In some embodiments, a portion is about 10 kilobases (kb) – about 20 kb, about 10 kb – about 100 kb, about 20 kb – about 80 kb, about 30 kb – about 70 kb, about 40 kb – about 60 kb. In some embodiments, the length of a portion is about 10 kb, 20 kb, 30 kb, 40 kb, 50 kb, or about 60 kb. Portions are not limited to continuously running sequences. Therefore, a portion may consist of continuous and / or non-continuous sequences. A portion is not limited to a single chromosome. In some embodiments, a portion includes all or part of a chromosome or all or part of two or more chromosomes. In some embodiments, a portion may span one, two, or more complete chromosomes. Furthermore, a portion may span connected or unconnected regions of multiple chromosomes.

[0217] In some embodiments, a portion may be a specific chromosomal region of the chromosome of interest, such as a chromosome used to assess genetic variation (e.g., aneuploidy of chromosomes 13, 18, and / or 21, or sex chromosomes). A portion may also be a pathogenic genome (e.g., bacteria, fungi, or viruses) or a fragment thereof. A portion may be a gene, gene fragment, regulatory sequence, intron, exon, etc.

[0218] In some implementations, the genome (e.g., the human genome) is divided into sections based on the information content of specific regions. In some implementations, genome partitioning may remove similar regions (e.g., identical or homologous regions or sequences) and retain only unique regions. Regions removed during partitioning may be within a single chromosome or may span multiple chromosomes. In some implementations, the partitioned genome is down-trimmed and optimized for rapid alignment, typically allowing focus on uniquely identifiable sequences.

[0219] In some implementations, the weights of similar regions can be reduced. The process of reducing some weights will be described in detail later.

[0220] In some implementations, the genome can be divided into extrachromosomal regions based on information generated within the classification context. For example, the information content can be quantified using p-values, measuring the significance of specific genomic locations in confirmed normal and abnormal subjects (e.g., euploid and triploid subjects, respectively). In some implementations, the genome can be divided into extrachromosomal regions based on any other criteria, such as speed / convenience of tag alignment, GC content (e.g., high or low GC content), uniformity of GC content, other measurements of sequence content (e.g., individual nucleotide fraction, pyrimidine or purine fraction, fraction of natural and non-natural nucleic acids, fraction of methylated nucleotides, and CpG content), methylation status, double melting temperature, compliance with sequencing or PCR, uncertainty in the allocation of individual portions to a reference genome, and / or targeted search for specific features.

[0221] A "segment" of a chromosome is usually a part of the chromosome, and usually a part of the chromosome that is different from the part. Chromosomal segments are sometimes located in different regions of the chromosome from the part, sometimes do not share polynucleotides with the part, and sometimes include polynucleotides from the part. Chromosomal segments usually contain more nucleotides than the part (e.g., segments sometimes include parts), and sometimes chromosomal segments contain fewer nucleotides than the part (e.g., segments sometimes are within parts).

[0222] Partial filtering and / or selection

[0223] Sometimes, segments are processed (e.g., normalization, filtering, selection, etc., or combinations thereof) based on one or more features, parameters, criteria, and / or methods described herein or known in the art. A segment can be processed by any suitable method and based on any suitable parameter. Non-limiting examples of features and / or parameters that can be used to filter and / or select segments include counts, coverage, mappability, variability, uncertainty level, guanine-cytosine (GC) content, CCF fragment length and / or read length (e.g., fragment length ratio (FLR), fetal ratio statistics (FRS)), DNase I sensitivity, methylation status, acetylation, histone distribution, chromatin structure, etc., or combinations thereof. A segment can be filtered and / or selected based on any suitable feature or parameter associated with the listed or described features or parameters herein. A segment can be filtered and / or selected based on features or parameters specific to a particular segment (e.g., identification of a single segment from multiple samples) and / or features or parameters specific to a particular sample (e.g., identification of multiple segments within a sample). In some implementations, portions are filtered and / or removed based on relatively low mappability, relatively high variability, high levels of uncertainty, relatively long CCF fragment lengths (e.g., low FRS, low FLR), relatively large repeat sequence fractions, high GC content, low GC content, low counts, zero counts, high counts, etc., or combinations thereof. In some implementations, portions (e.g., subsets of multiple portions) are selected based on appropriate levels of mappability, variability, uncertainty, repeat sequence fractions, counts, GC content, etc., or combinations thereof. In some implementations, portions (e.g., subsets of multiple portions) are selected based on relatively short CCF fragment lengths (e.g., high FRS, high FLR). Counts and / or readings mapped to portions are sometimes processed (e.g., normalized) before and / or after filtering or selecting portions (e.g., subsets of multiple portions). In some implementations, counts and / or readings mapped to portions are not processed before and / or after filtering or selecting portions (e.g., subsets of multiple portions).

[0224] Sequence readings from any suitable number of samples can be used to identify subsets of multiple parts that meet one or more criteria, parameters, and / or characteristics described herein. Sometimes, sequence readings from a set of samples derived from multiple pregnant women are used. It can process one or more samples from each of multiple pregnant women (e.g., 1 to about 20 samples from each pregnant woman (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18 or 19 samples)), and can select an appropriate number of pregnant women (e.g., about 2 to about 10,000 pregnant women (e.g., about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000)). In some implementations, sequence reads from the same test sample derived from the same pregnant woman are mapped to portions of a reference genome and used to generate subsets of multiple portions.

[0225] It has been observed that circulating cell-free nucleic acid fragments (CCF fragments) obtained from pregnant women generally contain nucleic acid fragments from fetal cells (i.e., fetal fragments) and nucleic acid fragments from maternal cells (i.e., maternal fragments). Sequence reads derived from fetal CCF fragments are referred to herein as “fetal reads.” Sequence reads derived from the genome of a pregnant woman carrying a fetus (e.g., the mother) are referred to herein as “maternal reads.” The CCF fragment from which fetal reads are derived is referred to herein as the fetal template, and the CCF fragment from which maternal reads are derived is referred to herein as the maternal template.

[0226] It has also been observed that fetal fragments in CCF fragments are typically relatively short (e.g., approximately 200 base pairs or less in length), and maternal fragments include both such relatively short and relatively long fragments. A subset of multiple portions mapped with a significant amount of readings from the relatively short fragments can be selected and / or identified. Without being theoretically limited, it is anticipated that the readings mapped to said portions are enriched with fetal readings, which can improve the accuracy of fetal genetic analysis (e.g., detecting the presence of fetal genetic variations (e.g., fetal chromosomal aneuploidy (e.g., T21, T18, and / or T13))).

[0227] However, when fetal genetic analysis is based on subsets of readings, a significant number of readings are typically not considered. For fetal genetic analysis, the selection of subsets of readings mapped to selected subsets of multiple parts, and the removal of readings from non-selected parts, can reduce the accuracy of the genetic analysis due to, for example, increased variation. In some embodiments, for fetal genetic analysis, when considering the selection of subsets of multiple parts, approximately 30% to approximately 70% (e.g., approximately 35%, 40%, 45%, 50%, 55%, 60%, or 65%) of sequencing readings obtained from the subject or sample are removed. In some embodiments, for fetal genetic analysis, approximately 30% to approximately 70% (e.g., approximately 35%, 40%, 45%, 50%, 55%, 60%, or 65%) of sequencing readings obtained from the subject or sample are used as subsets mapped to multiple parts.

[0228] Therefore, without theoretical limitations, for fetal genetic analysis, a trade-off is typically required between the increased accuracy associated with enriched fetal readings and the decreased accuracy associated with reduced reading data (e.g., partial and / or removal of readings). In some implementations, the method includes selecting a subset of multiple portions of readings rich in fetal nucleic acids (e.g., fetal readings) that can improve, or not significantly reduce, the accuracy of fetal genetic analysis. Despite this apparent trade-off, as described herein, it has been determined that employing a subset of multiple portions mapped to a significant amount of readings from relatively short segments can improve the accuracy of fetal genetic analysis.

[0229] In some implementations, a subset of multiple portions is selected based on readings from fragments of the CCF, wherein the portions to which the readings are mapped are shorter than the length of the selected fragment. Sometimes, a subset of multiple portions is selected by filtering out portions that do not meet this criterion. In some implementations, a subset of multiple portions is selected based on the amount of readings originating from relatively short CCF fragments (e.g., about 200 base pairs or less) mapped to the portions. Any suitable method may be used to identify and / or select portions that map to a significant amount of readings from CCF fragments shorter than the length of the selected fragment (e.g., a first selected fragment length). CCF fragments shorter than the selected fragment length are typically relatively short CCF fragments, and sometimes the selected fragment length is about 200 base pairs or less (e.g., CCF fragments about 190, 180, 170, 160, 150, 140, 130, 120, 110, 100, 90, or 80 base pairs). The length of a CCF fragment can be determined (e.g., inferred or derived) by mapping two or more reads (e.g., paired-end reads) derived from the fragment to a reference genome. For paired-end reads derived from the CCF fragment, for example, mapping the reads to a reference genome allows determining the length of the genomic sequence between the mapped reads, and the sum of the lengths of the two reads and the length of the genomic sequence between the reads equals the length of the CCF fragment. Sometimes, the length of the CCF fragment template is directly determined by the length of the reads derived from the fragment (e.g., single-end reads).

[0230] In some embodiments, a subset of multiple portions of readings mapped to significant amounts from CCF segments shorter than selected segment lengths is selected and / or identified based on whether the amount of readings from mappings of CCF segments shorter than a first selected segment length is greater than the amount of readings from mappings of CCF segments shorter than a second selected segment length. In some embodiments, a subset of multiple portions of readings mapped to significant amounts from CCF segments shorter than selected segment lengths is selected and / or identified based on whether the amount of readings from mappings of CCF segments shorter than a certain portion of a first selected segment length is greater than the average, arithmetic mean, or median amount of readings from mappings of CCF segments shorter than the analyzed portion of a second selected segment length. In some embodiments, a subset of portions of readings mapped to significant amounts from CCF segments shorter than selected segment lengths is selected and / or identified based on a segment length ratio (FLR) determined for each portion. The “segment length ratio” is also referred to herein as the fetal ratio statistic (FRS).

[0231] In some implementations, the FLR is determined in part based on the amount of readings mapped to portions of the CCF fragment shorter than a selected fragment length. In some implementations, the FLR value is typically a ratio of X to Y, where X is the amount of readings derived from CCF fragments shorter than a first selected fragment length, and Y is the amount of readings derived from CCF fragments shorter than a second selected fragment length. The selection of the first selected fragment length is typically independent of the second selected fragment length, and vice versa, while the second selected fragment length is typically greater than the first selected fragment length. The first selected fragment length can be about 200 bases or less to about 30 bases or less. In some embodiments, the length of the first selected fragment is about 200, 190, 180, 170, 160, 155, 150, 145, 140, 135, 130, 125, 120, 115, 110, 105, 100, 95, 90, 85, 80, 75, 70, 65, 60, 55, or 50 bases. In some embodiments, the length of the first selected fragment is about 170 to about 130 bases, and sometimes, about 160 to about 140 bases. In some embodiments, the length of the second selected fragment is about 2000 to about 200 bases. In some embodiments, the length of the second selected fragment is about 1000, 950, 800, 850, 800, 750, 700, 650, 600, 550, 500, 450, 400, 350, 300, or 250 bases. In some embodiments, the length of the first selected fragment is about 140 to about 160 bases (e.g., about 150 bases), and the length of the second selected fragment is about 500 to about 700 bases (e.g., about 600 bases). In some embodiments, the length of the first selected fragment is about 150 bases, and the length of the second selected fragment is about 600 bases.

[0232] In some implementations, the FLR is the average, arithmetic mean, or median of multiple FLR values. For example, sometimes the FLR for a given portion is: the average, arithmetic mean, or median FLR value for (i) two or more test samples, (ii) two or more subjects, or (iii) two or more test samples and two or more subjects. In some implementations, the average, arithmetic mean, or median FLR is derived from the FLR values ​​of two or more portions of a genome, chromosome, or segment thereof. In some implementations, the average, arithmetic mean, or median FLR is associated with uncertainty (e.g., standard deviation, median absolute deviation).

[0233] In some implementations, subsets of multiple parts are selected and / or identified based on one or more FLR values ​​(e.g., a comparison of one or more FLR values). In some implementations, subsets of multiple parts are selected and / or identified based on FLR and a threshold (e.g., a comparison of FLR to a threshold). In some implementations, the average, arithmetic mean, or median FLR from a given part is compared with the average, arithmetic mean, or median FLR from two or more parts derived from a genome, chromosome, or segment thereof. For example, sometimes the average FLR of a given part is compared with the median FLR of the given part. In some implementations, parts are selected and / or identified based on the average, arithmetic mean, or median FLR determined for a part and the average, arithmetic mean, or median FLR determined for a set of multiple parts (e.g., from a genome, chromosome, or segment thereof). In some implementations, if the average FLR of a part is below a threshold determined based on the median FLR, that part is removed and not considered (e.g., in fetal genetic analysis). In some implementations, if the mean, arithmetic mean, or median FLR of a portion is higher than a threshold determined based on the mean, arithmetic mean, or median FLR of a genome, chromosome, or segment thereof, then that portion is selected and / or added to a subset of multiple portions for consideration (e.g., when determining the presence of genetic variation). In some implementations, if the FLR of a portion is equal to or greater than about 0.15 to about 0.30 (e.g., about 0.16, 0.17, 0.18, 0.19, 0.20, 0.21, 0.21, 0.22, 0.23, 0.24, 0.25, 0.26, 0.27, 0.28, 0.29), then that portion is selected for consideration (e.g., added to or included in a subset of multiple portions for fetal genetic analysis). In some implementations, if the FLR of a certain portion is equal to or lower than about 0.20 to about 0.10 (e.g., about 0.19, 0.18, 0.17, 0.16, 0.15, 0.14, 0.13, 0.12, 0.11), then that portion is removed and not taken into consideration (e.g., filtered out).

[0234] Sometimes, portions within a subset are partially selected and / or identified based on whether significant readings from CCF fragments shorter than the selected fragment length map to portions (e.g., according to FLR). In some embodiments, portions within a subset may be selected and / or identified based on one or more characteristics or criteria, and the amount of sequence readings mapped from fragments shorter than the selected fragment length. In some embodiments, subsets of multiple portions are selected and / or identified based on whether significant readings from CCF fragments shorter than the selected fragment length map to portions (e.g., according to FLR) and one or more other features. Non-limiting examples of other features include the number of exons and / or GC content of one or more of the genome, chromosome, or segments thereof, and / or said portions. Therefore, sometimes, portions selected and / or identified based on whether significant readings from CCF fragments shorter than the selected fragment length map to portions of a subset (e.g., according to FLR) are further selected or removed based on the GC content and / or the number of exons in said portions. In some implementations, if the GC content and / or exon number in a certain portion are not related to the FLR of that portion, then that portion is not selected or is not considered.

[0235] In some embodiments, a subset of multiple portions consists of, substantially consists of, or includes portions that conform to one or more specific criteria described herein (e.g., portions characterized by an FLR equal to or greater than a certain value). In some embodiments, portions that do not conform to the criteria are included in a subset of multiple portions that conform to the criteria, for example, to improve the accuracy of fetal genetic analysis. In some embodiments, in a subset of multiple portions that are “substantially composed” of portions selected according to a certain criterion (e.g., an FLR equal to or greater than a certain value), about 90% or more (e.g., about 91%, 91%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) of said portions conform to the criterion, and about 10% or less (e.g., about 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, about 1% or less) of said portions do not conform to the criterion.

[0236] Parts can be selected and / or filtered out by any suitable method. In some embodiments, parts are selected based on the examination of data, charts, curves, and / or graphs. In some embodiments, parts are selected and / or filtered out (e.g., partially) by a system or apparatus including one or more microprocessors and memory. In some embodiments, parts are selected and / or filtered out (e.g., partially) by a non-transitory computer-readable storage medium having an executable program stored thereon, wherein the program provides instructions to the microprocessor to perform the selection and / or filtering.

[0237] A subset of multiple parts selected by the methods described herein can be used for fetal genetic analysis in different ways. In some embodiments, readings derived from the sample are used in a mapping method employing a pre-selected subset of the multiple parts described herein, without employing all or most portions of the reference genome. Those readings mapped to the pre-selected subset of multiple parts are typically used in further steps of fetal genetic analysis, and readings not mapped to the pre-selected subset of multiple parts are typically not used in further steps of fetal genetic analysis (e.g., unmapped readings are removed or filtered out).

[0238] In some implementations, sequence reads derived from the sample are mapped to all or most portions of a reference genome, and then a subset of a pre-selected plurality of portions as described herein is selected. Reads from the selected subset of portions are typically used in further steps of fetal genetic analysis. In later implementations, reads from unselected portions are typically not used in further steps of fetal genetic analysis (e.g., reads from unselected portions are removed or filtered out).

[0239] count

[0240] In some implementations, sequence reads mapped or partitioned based on selected features or variables can be quantified to determine the number of reads mapped to one or more parts (e.g., a reference genome part). In some implementations, the quantification of sequence reads mapped to a part is called counting (e.g., a count). Typically, the count is associated with the part. In some implementations, the counts of two or more parts (e.g., groups of parts) are mathematically processed (e.g., averaging, summing, normalization, etc., or combinations thereof). In some implementations, the count is determined from some or all of the sequence reads mapped to (i.e., associated with) the part. In some implementations, the count is determined from a predefined subset of the mapped sequence reads. The predefined subset of the mapped sequence reads can be defined or selected using any suitable feature or variable. In some implementations, the predefined subset of the mapped sequencing reads can contain 1–n sequence reads, where n represents the number equal to the sum of all sequence reads generated from the test or reference sample.

[0241] In some embodiments, the count is derived from sequence readings processed or manipulated by suitable methods, operations, or mathematical processes known in the art. The count can be determined by suitable methods, operations, or mathematical processes. In some embodiments, the count is derived from sequence readings associated with a portion, some or all of which have undergone weighting, removal, filtering, standardization, adjustment, averaging (to obtain a mean), addition or subtraction, or combinations thereof. In some embodiments, the count is derived from the original sequence readings and / or filtered sequence readings. In some embodiments, the count value is determined by mathematical processes. In some embodiments, the count value is the average, arithmetic mean, or sum of sequence readings mapped to a portion. Typically, the count is the arithmetic mean of multiple counts. In some embodiments, the count is associated with an uncertain value.

[0242] In some embodiments, the counts may be processed or transformed (e.g., standardization, combination, summation, screening, selection, averaging (to obtain a mean), etc., or combinations thereof). In some embodiments, the counts may be transformed to produce standardized counts. The counts may be processed (e.g., standardization) by methods known in the art and / or described herein (e.g., sample-by-sample standardization, GC content standardization, linear and nonlinear least squares regression, GC LOESS, LOWESS, PERUN, RM, GCRM, cQn, and / or combinations thereof).

[0243] Counts (e.g., raw, screened, and / or standardized counts) can be processed and standardized to one or more levels. Levels and profiles are detailed below. In some embodiments, counts are processed and / or standardized to a reference level. The reference level is described below. Counts processed according to a level (e.g., processed counts) can be associated with an uncertainty (e.g., calculated variance, error, standard deviation, Z-score, p-value, arithmetic mean absolute deviation, etc.). In some embodiments, the uncertainty defines a range above and below a certain level. Deviation values ​​can replace uncertainty values; non-limiting examples of deviation measurements include standard deviation, mean absolute deviation, median absolute deviation, standardized scores (e.g., Z-score, Z-score, normalization score, standardized variable, etc.).

[0244] Counts are typically obtained from nucleic acid samples from pregnant women carrying a fetus. Counts mapped to one or more portions of the nucleic acid sequence are typically represented by counts from both the fetus and the mother (e.g., pregnant woman subjects). In some implementations, some counts mapped to portions are derived from the fetal genome, and some counts mapped to the same portions are derived from the maternal genome.

[0245] Data processing and standardization

[0246] The mapped sequence reads that have already been counted are referred to herein as raw data, because the data represents unprocessed counts (such as raw counts). In some embodiments, the sequence read data in a dataset can be further processed (such as mathematical and / or statistical processing) and / or displayed to help provide results. In some embodiments, datasets (including large datasets) may benefit from preprocessing to aid further analysis. Preprocessing of datasets sometimes involves removing redundant and / or uninformative portions or portions of a reference genome (such as portions with uninformative data or portions of a reference genome, redundant mapped reads, portions with a median count of 0, sequences that occur too frequently or too infrequently). Without being theoretically limited, data processing and / or preprocessing can (i) remove noisy data, (ii) remove uninformative data, (iii) remove redundant data, (iv) reduce the complexity of large datasets, and / or (v) help transform the data from one form to one or more other forms. When used with data or datasets, the terms “preprocessing” and “processing” are collectively referred to herein as “processing.” In some embodiments, processing can make the data more readily available for further analysis, thereby generating results. In some implementations, one or more processing methods (e.g., standardization methods, partial screening, mapping, verification, etc. or combinations thereof) are performed by a processor, microprocessor, computer, device associated with memory and / or controlled by a microprocessor.

[0247] As used herein, the term "noisy data" refers to (a) data that shows significant differences between data points during analysis or plotting, (b) data with significant standard deviation (e.g., greater than 3 standard deviations), (c) data with significant mean standard error, and combinations thereof. Noisy data sometimes arises due to the quantity and / or quality of the starting material (e.g., nucleic acid samples), and sometimes it appears as part of the method for preparing or replicating DNA used to generate sequence reads. In some embodiments, the noise originates from certain sequences that occur too frequently when prepared using PCR-based methods. The methods described herein can reduce or eliminate the baseline of noisy data, thereby reducing the impact of noisy data on the results provided.

[0248] The terms “informative data,” “informative portion of the reference genome,” and “informative portion” are used herein to refer to data whose values ​​differ significantly from a predetermined threshold or fall outside a predetermined cutoff range, or data derived therefrom. The term “threshold” refers to any number calculated using a qualifying dataset as a limitation for diagnosing genetic variations (e.g., copy number variation, aneuploidy, chromosomal abnormalities, etc.). In some embodiments, a threshold exceeding the results obtained by the method of the present invention leads to a diagnosis of a genetic variation (e.g., trisomy 21) in the subject. In some embodiments, the threshold or range of values ​​is calculated using mathematical and / or statistical processing of sequence read data (e.g., from a reference and / or subject), while in other embodiments, the sequence read data processed to generate the range of thresholds or values ​​is sequence read data (e.g., from a reference and / or subject). In some embodiments, an uncertainty is determined. The uncertainty is typically a measure of variance or error and can be any suitable measure of variation or error. In some embodiments, the uncertainty is a standard deviation, standard error, calculated variance, p-value, or arithmetic mean absolute deviation (MAD). In some embodiments, the uncertainty can be calculated according to the formula of Example 4.

[0249] Any suitable procedure may be used to process the data set described herein. Non-limiting examples of methods suitable for processing the data set include filtering, standardization, weighting, monitoring peak height, monitoring peak area, monitoring peak margins, determining area ratios, mathematical processing of the data, statistical processing of the data, application of mathematical algorithms, analysis using fixed variables, analysis using optimized variables, plotting the data to identify patterns or trends for further processing, and combinations thereof. In some implementations, the data set is processed according to different characteristics (such as GC content, redundant location reads, centromere regions, telomere regions, etc., and combinations thereof) and / or variables (such as fetal sex, maternal age, maternal ploidy, fetal nucleotide baseline percentage, etc., and combinations thereof). In some implementations, processing the data set described herein can reduce the complexity and / or dimensionality of large and / or complex data sets. Non-limiting examples of complex data sets include sequence reads generated from one or more test subjects and multiple reference subjects of different ages and ethnic backgrounds. In some implementations, the data set can contain thousands to millions of sequence reads from each test subject and / or reference subject.

[0250] In some implementations, data processing can be performed in any number of steps. For example, in some implementations, data can be adjusted and / or processed using only a single processing method, while in other implementations, data can be processed using one or more, five or more, ten or more, or twenty or more processing steps (e.g., one or more processing steps, two or more processing steps, three or more processing steps, four or more processing steps, five or more processing steps, six or more processing steps, seven or more processing steps, eight or more processing steps, nine or more processing steps, ten or more processing steps, eleven or more processing steps, twelve or more processing steps, thirteen or more processing steps, fourteen or more processing steps, fifteen or more processing steps, sixteen or more processing steps, seventeen or more processing steps, eighteen or more processing steps, nineteen or more processing steps, or twenty or more processing steps). In some implementations, the processing step may be the same step repeated two or more times (e.g., filtering two or more times, standardizing two or more times), while in other implementations, the processing step may be two or more different processing steps performed simultaneously or sequentially (e.g., filtering, standardizing; standardizing, monitoring peak height and edges; filtering, standardizing, standardizing against a reference, statistical processing to determine p-values, etc.). In some implementations, any suitable number and / or combination of the same or different processing steps may be used to process sequence readout data to aid in providing results. In some implementations, processing the data set using the standards described herein can reduce the complexity and / or dimensionality of the data set.

[0251] In some implementations, one or more processing steps may include one or more filtering steps. As used herein, the term "filtering" refers to removing a portion or reference genome from consideration. The portion or reference genome to be removed can be selected according to any suitable criteria, including but not limited to redundant data (such as redundant or overlapping mapping reads), non-informative data (such as portions or portions of the reference genome with a median count of 0), portions or portions of the reference genome containing sequences that occur too frequently or too infrequently, noisy data, and combinations thereof. Filtering methods often involve removing one or more portions of the reference genome from consideration and subtracting the counts of the selected portions or portions of the reference genome to be removed from the counts or totals of the considered reference genome, chromosome, or genome. In some implementations, portions of the reference genome can be removed sequentially (e.g., one at a time to allow evaluation of the impact of removal of each individual portion), while in other implementations, all portions marked for removal can be removed simultaneously. In some implementations, portions of the reference genome characterized by differences above or below a certain level are removed; these are sometimes referred to as filtering "noisy" portions of the reference genome. In some embodiments, the filtering process includes extracting data points from a dataset of average profile levels originating from a part, chromosome, or chromosomal segment through predetermined multiple profile variations, and in other embodiments, the filtering process includes removing data points from a dataset of average profile levels originating from a part, chromosome, or chromosomal segment through predetermined multiple profile differences. In some embodiments, the filtering process is used to reduce the number of candidate parts in a reference genome used to analyze the presence or absence of genetic variations. Reducing the number of candidate parts in a reference genome used to analyze the presence or absence of genetic variations (e.g., microdeletions, microduplications) typically reduces the complexity and / or dimensionality of the dataset and sometimes increases the speed of searching for and / or identifying genetic variations and / or genetic abnormalities by two or more orders of magnitude.

[0252] In some implementations, one or more processing steps may include one or more standardization steps. Standardization can be performed by suitable methods described herein or known in the art. In some implementations, standardization includes adjusting measured values ​​of different orders of magnitude to a theoretically common order of magnitude. In some implementations, standardization includes complex mathematical adjustments to introduce a adjusted probability distribution of values ​​in the comparison. In some implementations, standardization includes comparing the distribution to a normal distribution. In some implementations, standardization includes mathematical adjustments that allow for comparison of corresponding standardized values ​​for different data sets in a manner that eliminates certain overall effects (e.g., errors and outliers). In some implementations, standardization includes scaling. Standardization sometimes includes dividing one or more data sets by a predetermined scalar or formula. Non-limiting examples of standardization methods include component-by-component standardization, standardization by GC content, linear and nonlinear least squares regression, LOESS, GCLOESS, LOWESS (locally weighted regression scatter smoothing), PERUN, repetition masking (RM), GC-standardization and repetition masking (GCRM), conditional quantile standardization (cQn), and / or combinations thereof. In some embodiments, the determination of the presence or absence of genetic variations (e.g., aneuploidy) employs standardization methods (e.g., part-by-part standardization, standardization by GC content, linear and nonlinear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted regression scatter smoothing), PERUN, repeat masking (RM), GC-standardization and repeat masking (GCRM), cQn, standardization methods known in the art, and / or combinations thereof). In some embodiments, the counts are standardized.

[0253] For example, LOESS is a known regression model in the art that combines multiple regression models in a k-nearest neighbor-based meta-model. LOESS sometimes refers to locally weighted polynomial regression. In some implementations, GC LOESS applies the LOESS model to the relationship between GC composition and fragment counts (e.g., sequence reads, counts) of a portion of a reference genome. Plotting a smooth curve using LOESS over a set of data points is sometimes called a LOESS curve, especially when smoothing values ​​are given by a weighted quadratic least squares regression relative to the span of the values ​​of the standard variables in a y-axis scatter plot. For each point in the data set, the LOESS method fits a low-degree polynomial to the data set, indicating points where the variable values ​​are close to the evaluated response. A weighted least squares polynomial is fitted such that points close to the evaluated response have more weight and points far away have less weight. The regression function value for each point is then obtained by evaluating the locally weighted polynomial using the indicator variable values ​​for that data point. Sometimes, the LOESS fit is fully considered after the regression function values ​​have been calculated for each data point. Many details of this method, such as the degree and weights of the polynomial model, are flexible.

[0254] Any suitable number of standardization times can be used. In some implementations, the dataset can be standardized one or more times, five or more times, ten or more times, or even twenty or more times. The dataset can be standardized to values ​​(e.g., standardized values) that represent any suitable characteristic or variable (e.g., sample data, reference data, or both). Non-limiting examples of available data standardization types include standardizing raw count data from one or more selected test or reference portions to the total counts mapped to the selected portion or segment of the chromosome or the whole genome; standardizing raw count data from one or more selected portions to the median reference count mapped to one or more portions or the chromosome of the selected portion or segment; standardizing raw count data to the aforementioned standardized data or its derivatives; and standardizing the aforementioned standardized data to one or more other predetermined standardization variables. Standardizing the dataset can sometimes separate statistical errors, depending on the characteristics or properties selected as predetermined standardization variables. Standardizing the dataset can also sometimes make data characteristics of data of different magnitudes comparable by converting the data to a common scale (e.g., predetermined standardization variables). In some implementations, one or more standardizations of statistically derived values ​​can be used to minimize data variance and reduce the importance of outlying data. When it comes to standardized values, partial standardization of a portion or a reference genome is sometimes referred to as "partial standardization."

[0255] In some embodiments, the processing steps include standardization, including standardization to a static window, and in some embodiments, the processing steps include standardization, including standardization to a dynamic or sliding window. The term "window" herein refers to the selection of one or more portions for analysis, sometimes used as a reference for comparison (e.g., for standardization and / or other mathematical or statistical operations). The term "standardization to a static window" herein refers to a standardization process using one or more portions selected for comparing test and reference datasets. In some embodiments, the selected portions are used to generate profiles. A static window typically includes a predetermined set of portions that does not change during operation and / or analysis. The terms "standardization to a dynamic window" or "standardization sliding window" herein refer to the standardization of portions located within a genomic region (e.g., genetically closely surrounding, contiguous portions or segments) of a selected test portion, wherein one or more of the selected test portions are standardized to portions closely surrounding the selected test portion. In some embodiments, the selected portions are used to generate profiles. Sliding or dynamic window normalization typically involves repeatedly moving or sliding to adjacent test portions and normalizing newly selected test portions to portions closely surrounding or adjacent to said newly selected test portions, wherein adjacent windows have one or more shared portions. In some implementations, multiple selected test portions and / or chromosomes can be analyzed via a sliding window process.

[0256] In some implementations, normalization to a sliding or dynamic window can produce one or more values, each representing normalization for different sets of reference portions selected from different regions of the genome (e.g., chromosomes). In some implementations, the resulting one or more values ​​are cumulative values ​​(e.g., numerical estimates of the integral of a normalized count profile over a selected portion, domain (e.g., a portion of a chromosome), or chromosome). Values ​​obtained from the sliding or dynamic window procedure can be used to generate profiles and facilitate obtaining results. In some implementations, the cumulative sum of one or more portions can be displayed as a function of genomic location. Dynamic or sliding window analysis is sometimes used to analyze the presence of microdeletions and / or microinsertions in the genome. In some implementations, displaying the cumulative sum of one or more portions is used to identify regions of genetic variation (e.g., microdeletions, microduplications). In some implementations, dynamic or sliding window analysis is used to identify genomic regions containing microdeletions, and in some implementations, dynamic or sliding window analysis is used to identify genomic regions containing microduplications.

[0257] A useful standardization method for reducing errors associated with nucleic acid indicators is referred to herein as Parametric Error Removal and Unbiased Standardization (PERUN), and is incorporated herein in its entirety by reference, including all text, tables, equations, and figures, as described herein and, for example, by U.S. Patent Application No. 13 / 669,136 and International Patent Application PCT / US12 / 59123 (WO2013 / 052913). The PERUN method can be used with a variety of nucleic acid indicators (e.g., nucleic acid sequence readings) to reduce the influence of errors that confound predictions based on such indicators.

[0258] For example, the PERUN method is used for nucleic acid sequence readings from samples and reduces the impact of errors in determining genomic segment levels. This application is effective for using nucleic acid sequence readings to determine the presence of genetic variations in objects exhibiting various levels of nucleotide sequences (e.g., partial, genomic segment level). Non-limiting examples of variation in partial sequences are chromosomal aneuploidy (e.g., trisomy 21, trisomy 18, trisomy 13) and the presence of sex chromosomes (e.g., XX in females and XY in males). Autosomal trisomy (e.g., chromosomes other than sex chromosomes) can be referred to as affected autosomes. Other non-limiting examples of variation at the genomic segment level include microdeletions, microinsertions, duplications, and mosaicism.

[0259] In some applications, the PERUN method can reduce experimental bias by standardizing nucleic acid metrics of specific genomic units (i.e., portions). A portion includes a suitable set of nucleic acid metrics; non-limiting examples include the length of consecutive nucleotides referred to herein as a genomic segment or a portion of the reference genome. A bin may include other nucleic acid metrics, as described herein. In such applications, the PERUN method generally standardizes nucleic acid metrics of multiple samples in a three-dimensional model across specific bins.

[0260] In some applications, the PERUN method can reduce experimental and / or systematic bias by standardizing nucleic acid indicators (e.g., counts, readings) mapped to specific segments (e.g., portions) of a reference genome. In this application, the PERUN method typically standardizes nucleic acid readings at specific portions of a reference genome across a large number of samples in three dimensions. A detailed description of PERUN and its applications can be found in the Examples section, as well as in International Patent Application PCT / US12 / 59123 (WO2013 / 052913) and U.S. Patent Application Publication No. US20130085681, the entire contents of which are incorporated herein by reference, including all text, tables, equations, and figures.

[0261] In some implementations, the PERUN method includes calculating genomic segment-level data for a portion of the reference genome from the following results: (a) the number of sequence reads mapped from the test sample to the portion of the reference genome, (b) the experimental bias (e.g., GC bias) of the test sample, and (c) one or more fitting parameters (e.g., fit estimates) for the fit between (i) the experimental bias of the reference genome portion mapped to the sequence reads and (ii) the number of sequence reads mapped to said portion. The experimental bias of each portion of the reference genome can be determined across a variety of samples based on the fit between (i) the number of sequence reads mapped to each portion of the reference genome and (ii) the mapping characteristics of each portion of the reference genome. Such fits for each sample can be aggregated across a variety of samples in a three-dimensional direction. In some implementations, this aggregate can be arranged according to the experimental bias, although the PERUN method can be implemented without arranging the aggregate according to the experimental bias. The fits for each sample and the fits for each portion of the reference genome can be individually fitted to linear or nonlinear functions using suitable fitting methods known in the art (e.g., fitting models). Non-restrictive examples of suitable models that can be used for relationship fitting include linear regression models, simple regression models, ordinary least squares regression models, multiple regression models, general multiple regression models, polynomial regression models, general linear models, generalized linear models, discrete choice regression models, logistic regression models, polynomial logarithmic models, mixed logarithmic models, probabilistic unit models, polynomial probabilistic unit models, ordered logarithmic models, ordered probabilistic unit models, Poisson (Poisson) models, multivariate response regression models, multilevel models, fixed effects models, random effects models, mixed models, nonlinear regression models, nonparametric models, semiparametric models, robust models, quantile models, isotonic models, principal component models, least angle models, local models, piecewise models, and variable error models.

[0262] In some implementations, the relationship is a geometric and / or graphical relationship. The terms "relationship" and "correlation" are used synonymously herein. In some implementations, the relationship is a mathematical relationship. In some implementations, the relationship is graphical. In some implementations, the relationship is a linear relationship. In some implementations, the relationship is a non-linear relationship. In some implementations, the relationship is a regression (e.g., a regression line). The regression can be linear or non-linear. The relationship can be expressed by a mathematical equation. Relationships are typically defined in part by one or more constants and / or one or more variables. Relationships can be generated by methods known in the art. In some implementations, a two-dimensional relationship can be generated for one or more samples, and error tests or probable error tests can be selected for one or more of the said dimensions. For example, relationships can be generated using graphing software known in the art, which plots two or more variable values ​​provided by the user. Relationships can be fitted using methods known in the art (e.g., by performing regression, regression analysis, for example, by a suitable regression procedure, for example, software). Some relationships can be fitted by linear regression, and linear regression can generate slopes and intercepts. Some relationships are sometimes nonlinear and can be fitted by nonlinear functions, such as parabolas, hyperbolas, or exponential functions (e.g., quadratic functions).

[0263] In the PERUN method, one or more fitting relationships can be linear. To analyze cell-free circulating nucleic acids in pregnant women, where the experimental bias is GC bias and the mapping characteristic is GC content, the fitting relationship between (i) the sequence read counts mapped to each part of the sample and (ii) the GC content of each part of the reference genome can be linear. For the latter fitting relationship, when ensemble fitting relationships are made across multiple samples, a slope and GC bias coefficient involving GC bias can be determined for each sample. In this embodiment, the fitting relationship between the GC bias coefficients of the parts described in (i) and (ii) the sequence read counts mapped to the parts can also be linear. The intercept and slope can be obtained from the latter fitting relationship. In this application, the slope represents sample-specific bias based on GC content, and the intercept represents a part-specific decay pattern present in all samples. When calculating at the genomic segment level to provide results (e.g., the presence of genetic variation; determining fetal sex), the PERUN method can significantly reduce sample-specific bias and part-specific decay.

[0264] In some implementations, PERUN normalization uses a fitting to a linear function and is described in Equations A, B, or their derivatives.

[0265] Equation A:

[0266] M = LI + GS(A)

[0267] Equation B:

[0268] L=(M–GS) / I(B)

[0269] In some implementations, L is the PERUN normalization level or profile. In some implementations, L is the output required from the PERUN normalization procedure. In some implementations, L is part-specific. In some implementations, L is determined based on multiple parts of the reference genome, representing the PERUN normalization level of the genome, chromosome, its parts, or segments. The level L is typically used for further analysis (e.g., determining Z-values, maternal deletions / duplications, fetal microdeletions / microduplications, fetal sex, sex aneuploidy, etc.). The normalization method according to Equation B is called Parametric Error Removal and Unbiased Normalization (PERUN).

[0270] In some implementations, G is a GC bias coefficient measured using a linear model, LOESS, or any equivalent method. In some implementations, G is a slope. In some implementations, the GC bias coefficient G is evaluated as the slope of the regression between the count M (e.g., raw count) for part i and the GC content of part i determined from a reference genome. In some implementations, G represents secondary information extracted from M and determined based on the relationship. In some implementations, G represents the relationship between a set of part-specific counts and a set of part-specific GC content values ​​for a sample (e.g., a test sample). In some implementations, the part-specific GC content is derived from a reference genome. In some implementations, the part-specific GC content is derived from observed or measured GC content (e.g., measured from a sample). The GC bias coefficient is typically determined for each sample in a sample set and typically for the test sample. The GC bias coefficient is typically sample-specific. In some implementations, the GC bias coefficient is a constant. In some implementations, the GC bias coefficient does not change once obtained from the sample.

[0271] In some implementations, S is the slope derived from a linear relationship and I is the intercept. In some implementations, the relationships from which I and S are derived differ from the relationships from which G is derived. In some implementations, the relationships from which I and S are derived are fixed for a given experimental setting. In some implementations, I and S are derived from a linear relationship based on counts (e.g., raw counts) and GC bias coefficients based on multiple samples. In some implementations, I and S are independently derived from the test sample. In some implementations, I and S are derived from multiple samples. I and S are typically part-specific. In some implementations, I and S are determined for all parts of the reference genome in euploid samples using the assumption L=1. In some implementations, a linear relationship is determined for euploid samples, and I and S values ​​specific to selected parts are determined (assuming L=1). In some implementations, the same procedure is applied to all parts of the reference genome in the human genome, and sets of intercepts I and slopes S are determined for each part.

[0272] In some implementations, cross-validation is applied. Cross-validation is sometimes referred to as rotation estimation. In some implementations, cross-validation is used to evaluate the accuracy of a predictive model (e.g., PERUN) implemented on a test sample. In some implementations, a round of cross-validation involves partitioning the data sample into complementary subsets, performing cross-validation analysis on the subsets (e.g., sometimes called the training set), and performing validation analysis using another subset (e.g., sometimes called the validation set or test set). In some implementations, multiple rounds of cross-validation are performed using different partition products and / or different subsets. Non-limiting examples of cross-validation methods include leave-one-out, sliding margin, K-fold, 2-fold, repeated random sampling, etc., or combinations thereof. In some implementations, cross-validation randomly selects a working group containing 90% of the sample groups, including known euploid fetuses, and uses this subset to train the model. In some implementations, random selection is repeated 100 times, with each portion generating 100 slopes and 100 intercepts.

[0273] In some implementations, the M value is a measurement derived from the test sample. In some implementations, M is a raw count for a portion of the measurement. In some implementations, where values ​​I and S are available for a portion, the M measurement is determined from the test sample and used to determine the PERUN normalization level L of the genome, chromosome, its segment, or portion according to Equation B.

[0274] Therefore, applying the PERUN method in parallel to sequence readings of multiple samples can significantly reduce errors caused by (i) sample-specific experimental bias (e.g., GC bias) and (ii) sample-specific attenuation common to both sources. Other methods that address these two sources of error individually or sequentially typically cannot reduce them as effectively as the PERUN method. Unrestricted by theory, the PERUN method is expected to reduce error more effectively because its general addition process does not amplify as dramatically as the general multiplication process used in other normalization methods (e.g., GC-LOESS).

[0275] Other standardization and statistical techniques can be used in conjunction with the PERUN method. Other procedures can be applied before, after, and / or during the use of the PERUN method. Non-limiting examples of procedures that can be used in conjunction with the PERUN method are described below.

[0276] In some implementations, secondary normalization or adjustment of GC content at the genomic segment level can be combined with the PERUN method. Appropriate GC content adjustment or normalization procedures (e.g., GC-LOESS, GCRM) can be used. In some implementations, additional GC normalization processes can be applied to select and / or identify specific samples. For example, the application of the PERUN method can determine the GC bias of each sample, and samples with GC biases above a certain threshold can be selected for further GC normalization processes. In this implementation, a predetermined threshold level can be used to select the sample for further GC normalization.

[0277] In some embodiments, partial filtering or weighting processes may be used in conjunction with the PERUN method. Suitable partial filtering or weighting processes may be used, and non-limiting examples are described herein, as well as in International Patent Application PCT / US12 / 59123 (WO2013 / 052913) and U.S. Patent Application Publication No. US20130085681, the entire contents of which are incorporated herein by reference, including all text, tables, equations, and figures. In some embodiments, normalization techniques for reducing associated maternal insertions, duplications, and / or deletions (e.g., maternal and / or fetal copy number variations) are used in conjunction with the PERUN method.

[0278] Genomic segment levels calculated using the PERUN method can be used directly to provide results. In some implementations, genomic segment levels can be used directly to provide sample results where the fetal fraction is approximately 2% to approximately 6% or higher (e.g., approximately 4% or higher). Genomic segment levels calculated using the PERUN method are sometimes further processed to provide results. In some implementations, the calculated genomic segment levels are normalized. In some implementations, the sum, arithmetic mean, or median of the calculated genomic segment levels for the test portion (e.g., chromosome 21) can be divided by the sum, arithmetic mean, or median of the calculated genomic segment levels for portions other than the test portion (e.g., chromosome 21 other than autosomes) to generate the experimental genomic segment levels. The experimental or raw genomic segment levels can be used as part of a programmed analysis, such as calculating Z-scores. The Z-score for a sample can be generated by subtracting the expected genomic segment level from the experimental or raw genomic segment levels, and the resulting value can be divided by the standard deviation of the sample. In some implementations, the resulting Z-scores can be distributed and analyzed across different samples, or correlated with other variables, such as fetal scores and others, and analyzed to provide results.

[0279] As described herein, the PERUN method is not limited to standardization based on GC bias and GC content alone, and can be used to reduce errors associated with other sources of error. A non-limiting example of a source of non-GC content bias is mappability. When addressing standardized parameters other than GC bias and content, one or more fitting relationships can be non-linear (e.g., hyperbolic, exponential). In some implementations, for example, when experimental bias is determined from a non-linear relationship, experimental bias curvature estimates can be analyzed.

[0280] The PERUN method can be applied to a variety of nucleic acid indicators. Non-limiting examples of nucleic acid indicators are nucleic acid sequence reads and nucleic acid levels at specific locations on a microarray. Non-limiting examples of sequence reads include those obtained from cell-free circulating DNA, cell-free circulating RNA, cellular DNA, and cellular RNA. The PERUN method can be applied to sequence reads mapped to suitable reference sequences, such as genomic reference DNA, cellular reference RNA (e.g., transcriptome), and portions thereof (e.g., portions of genomic complements of DNA or RNA transcriptomes, portions of chromosomes).

[0281] Therefore, in some implementations, cellular nucleic acids (e.g., DNA or RNA) can be used as nucleic acid indicators. Cellular nucleic acid readings mapped to a reference genome portion can be normalized using the PERUN method. Cellular nucleic acids bound to specific proteins sometimes refer to the chromatin immunoprecipitation (ChIP) process. ChIP-enriched nucleic acids are nucleic acids, such as DNA or RNA, associated with cellular proteins. ChIP-enriched nucleic acid readings can be obtained using techniques known in the art. ChIP-enriched nucleic acid readings can be mapped to one or more portions of a reference genome, and the results can be normalized using the PERUN method to provide the outcome.

[0282] In some implementations, cellular RNA can be used as a nucleic acid indicator. Cellular RNA readings can be mapped to a reference RNA portion and normalized using the PERUN method to provide results. A known sequence of cellular RNA (referred to as the transcriptome) or a segment thereof can be used as a reference, to which RNA readings from the sample can be mapped. Sample RNA readings can be obtained using techniques known in the art. The results of mapping RNA readings to a reference can be normalized using the PERUN method to provide results.

[0283] In some implementations, microarray nucleic acid levels can be used as nucleic acid indicators. The PERUN method can be used to analyze the nucleic acid levels or hybrid nucleic acids at specific locations on the array, thereby standardizing the nucleic acid indicators provided by microarray analysis. In this way, specific locations or hybrid nucleic acids on the microarray are similar to portions of the mapped nucleic acid sequence readings, and the PERUN method can be used to standardize microarray data to provide improved results.

[0284] In some implementations, the processing steps include weighting. As used herein, the terms “weighted,” “weighted,” or “weighting function,” or their syntactic derivatives or equivalents, refer to a mathematical processing of part or all of a dataset, which is sometimes used to alter the influence of certain dataset characteristics or variables on other dataset characteristics or variables (e.g., increasing or decreasing the importance and / or baseline of data contained in one or more portions of a selected reference genome based on the quality or utility of the data). In some implementations, weighting functions can be used to increase the influence of data with relatively small measurement variables and / or decrease the influence of data with relatively large measurement differences. For example, portions of the reference genome containing excessively low-frequency or low-quantity sequence data can be “downweighted” to minimize their influence on the dataset, while selected portions of the reference genome can be “upweighted” to increase their influence on the dataset. A non-limiting example of a weighting function is [1 / (standard deviation)]. 2 The weighting step is sometimes performed in a manner substantially similar to the standardization step. In some implementations, the data set is divided by a predetermined variable (such as a weighting variable). Often, a predetermined variable (such as minimizing a target function, Φ) is chosen to selectively weight different parts of the data set (e.g., increasing the influence of certain data types while decreasing the influence of others).

[0285] In some implementations, the processing steps may include one or more mathematical and / or statistical processes. Any suitable mathematical and / or statistical process may be used alone or in combination to analyze and / or process the data set described herein. Any suitable number of mathematical and / or statistical processes can be used. In some implementations, the data set may be subjected to mathematical and / or statistical processes one or more times, five or more times, ten or more times, or twenty or more times. Non-limiting examples of mathematical and statistical processes that can be used include addition, subtraction, multiplication, division, algebraic functions, least squares estimation, curve fitting, differential equations, rational polynomials, double polynomials, orthogonal polynomials, z-scores, p-values, χ-values, etc. This includes the ability to perform mathematical and / or statistical processing on sequence read data or its processed results, including peak level analysis, determining peak edge positions, calculating peak area ratios, analyzing median chromosome levels, calculating arithmetic mean absolute deviation, residual sum of squares, mean, standard deviation, standard error, etc. Non-restrictive examples of statistically processable data set variables or characteristics include raw counts, filtered counts, standardized counts, peak height, peak width, peak area, peak edge, lateral tolerance, p-value, median level, average level, count distribution within genomic regions, relative representation of nucleic acid content, etc., or combinations thereof.

[0286] In some implementations, the processing steps may include using one or more statistical algorithms. Any suitable statistical algorithm may be used alone or in combination to analyze and / or process the data set described herein. Any suitable number of statistical algorithms may be used. In some implementations, one or more, five or more, ten or more, or twenty or more statistical algorithms may be used to analyze the data set. Non-limiting examples of statistical algorithms suitable for use with the methods described herein include decision trees, count null values, multiple comparisons, comprehensive tests, the Behrens-Fischer problem, bootstrapping, Fisher's method combined with a significance test for independence, null hypothesis, Type I error, Type II error, exact test, one-sample Z-test, two-sample Z-test, one-sample t-test, paired t-test, two-sample pooled t-test with equal variances, two-sample unpooled t-test with unequal variances, single proportion Z-test, pooled two-proportion Z-test, unpooled two-proportion Z-test, one-sample chi-square test, two-sample F-test with equal variances, confidence intervals, confidence intervals, significance, meta-analysis, simple linear regression, strong linear regression, or a combination thereof. Non-restricted examples of data set variables or features that can be analyzed using statistical algorithms include raw counts, filtered counts, standardized counts, peak height, peak width, peak margin, lateral tolerance, p-value, median level, average level, count distribution within genomic regions, relative representation of nucleic acid content, or combinations thereof.

[0287] In some implementations, the dataset can be analyzed using multiple (e.g., two or more) statistical algorithms, such as least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, Bagging, neural networks, support vector machine models, random forests, classification tree models, k-nearest neighbors, logistic regression, and / or smoothing loss. Smoothing) and / or mathematical and / or statistical operations (such as those described herein). In some embodiments, using multiple operations can generate an N-dimensional space that can be used to provide results. In some embodiments, analyzing a dataset using multiple operations can reduce the complexity and / or dimensionality of the dataset. For example, using multiple operations on a reference dataset can generate an N-dimensional space (e.g., a probability plot) that can be used to represent the presence or absence of genetic variation depending on the genetic status of the reference sample (e.g., positive or negative for a selected genetic variation). Analyzing test samples using substantially similar sets of operations can generate N-dimensional points for each of the tested samples. The complexity and / or dimensionality of the test dataset is sometimes reduced to N-dimensional points or single values ​​that can be easily compared to the N-dimensional space of the reference data. Test sample data falling within the N-dimensional space filled by the reference data indicates a genetic status substantially similar to that of the reference. Test sample data falling outside the N-dimensional space filled by the reference data indicates a genetic status substantially dissimilar to that of the reference. In some embodiments, the reference is euploid or does not have genetic variation or medical symptoms.

[0288] In some implementations, after computation, optional filtering, and standardization, the processed data set can be further manipulated using one or more filtering and / or standardization procedures. In some implementations, data sets that can be further manipulated using one or more filtering and / or standardization procedures can be used to generate profiles. In some implementations, one or more filtering and / or standardization procedures can sometimes reduce the complexity and / or dimensionality of the data set. Results can be provided based on the data set with reduced complexity and / or dimensionality.

[0289] In some implementations, portions may be filtered based on error measurements (e.g., standard deviation, standard error, calculated variance, p-value, arithmetic mean absolute error (MAE), mean absolute deviation, and / or arithmetic mean absolute deviation (MAD). In some implementations, the error measurement refers to count variability. In some implementations, portions are filtered based on count variability. In some implementations, count variability is an error measurement determined for the counts of portions (i.e., portions) mapped to a reference genome from multiple samples (e.g., multiple samples obtained from multiple objects, such as 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more, or 10,000 or more objects). In some implementations, portions with count variability exceeding a predetermined upper limit may be filtered (e.g., excluded from consideration). In some implementations, the predetermined upper limit is equal to or greater than about 50, about 52, about 54, about 56, about 58, about 60, about 62, about 64, about 66, about 68, about 70, about 72, about 74, or equal to or greater than about 7. The MAD value is 6. In some embodiments, portions of count variability below a predetermined lower limit range can be filtered (e.g., excluded from consideration). In some embodiments, the predetermined lower limit range is a MAD value equal to or less than about 40, about 35, about 30, about 25, about 20, about 15, about 10, about 5, or about 1 equal to or less than about 0. In some embodiments, portions of count variability exceeding the predetermined range can be filtered (e.g., excluded from consideration). In some embodiments, the predetermined range is greater than 0 and less than about 76. The MAD values ​​are less than approximately 74, less than approximately 72, less than approximately 71, less than approximately 70, less than approximately 69, less than approximately 68, less than approximately 67, less than approximately 66, less than approximately 65, less than approximately 64, less than approximately 62, less than approximately 60, less than approximately 58, less than approximately 56, less than approximately 54, less than approximately 52, and less than approximately 50. In some embodiments, the predetermined range is MAD values ​​greater than 0 and less than approximately 67.7. In some embodiments, the portion of the count variability within the predetermined range is selected (e.g., for determining the presence of genetic variation).

[0290] In some embodiments, the portion of the count variability represents a distribution (e.g., a normal distribution). In some embodiments, a portion may be selected within the quantiles of the distribution. In some embodiments, the quantiles of the distribution are selected to be equal to or less than about 99.9%, 99.8%, 99.7%, 99.6%, 99.5%, 99.4%, 99.3%, 99.2%, 99.1%, 99.0%, 98.9%, 98.8%, 98.7%, 98.6%, 98.5%, 98.4%, 98.3%, 98.2%, 98.1%, 98.0%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 85%, 80%, or equal to or less than about 75%. In some embodiments, the quantiles of the distribution of count variability are selected to be within 99%. In some implementations, portions with MAD>0 and MAD<67.725, within the 99th percentile, are selected to identify stable subsets of the reference genome.

[0291] Non-limiting examples of partial filtering involving PERUN are described herein and in full by reference in International Patent Application No. PCT / US12 / 59123 (WO2013 / 052913), all of which are incorporated herein by reference, including all text, tables, equations, and figures. Partial filtering may be based on or in part on error measurements. Error measurements include the absolute value of the deviation, such as an R-factor, which in some embodiments may be used for partial removal or weighting. In some embodiments, the R-factor is defined as the sum of the absolute deviations of predicted and measured values ​​divided by the predicted counts from the measured values ​​(e.g., Formula B described herein). While error measurements including the absolute value of the deviation may be used, suitable error measurements may also be used. In some embodiments, error measurements excluding the absolute value of the deviation may be used, such as a square-based dispersion. In some embodiments, partial filtering or weighting is based on a measure of mappability (e.g., a mappability score). Sometimes partial filtering or weighting is based on a relatively low number of sequence readings mapped to said partial (e.g., 0, 1, 2, 3, 4, 5 readings mapped to said partial). Depending on the type of analysis being performed, some parts can be filtered or weighted. For example, for aneuploidy analysis of chromosomes 13, 18, and / or 21, sex chromosomes can be filtered out, and only autosomes or subsets of autosomes can be analyzed. For determining fetal sex, autosomes can be filtered out, and only sex chromosomes (X and Y), or one of the sex chromosomes (X or Y), can be analyzed.

[0292] In a specific implementation, the following filtering procedure can be used. Select portions of the same group within a given chromosome (e.g., a portion of the reference genome) and compare the number of reads in affected and unaffected samples. The gap involves trisomy 21 and euploid samples, which involves portions covering most of chromosome 21. The portions are identical between euploid and T21 samples. The difference between portions and individual segments is not critical, as defined by the portion. Compare identical genomic regions in different patients. This procedure can be used for trisomy analysis, such as T13 or T18, in addition to or instead of T21.

[0293] In some implementations, after computation, optional filtering, and standardization, the processed data set can be manipulated by weighting. In some implementations, one or more portions may be selectively weighted to reduce the influence of data contained in the selected portions (e.g., noisy data, uninformative data), and in some implementations, one or more portions may be selectively weighted to increase or strengthen the influence of data contained in the selected portions (e.g., data with small measurement variance). In some implementations, a single weighting function is used to weight the data set, which reduces the influence of data with large variance and increases the influence of data with small variance. Weighting functions are sometimes used to reduce the influence of data with large variance and increase the influence of data with small variance (e.g., [1 / (standard deviation)]). 2 In some implementations, the data is further processed by weighting to generate a profile of the processed data to facilitate classification and / or provide results. Results may be provided based on the profile of the weighted data.

[0294] Partial filtering and weighting can be performed at one or more suitable points during the analysis. For example, partial filtering or weighting can be performed before or after the sequence reads are mapped to a portion of the reference genome. In some embodiments, partial filtering or weighting can be performed before or after determining experimental bias in individual genome portions. In some embodiments, partial filtering or weighting can be performed before or after calculating at the genomic segment level.

[0295] In some embodiments, after computation, optional filtering, standardization, and optional weighting, the processed data set can be manipulated by one or more mathematical and / or statistical operations (such as statistical functions or statistical algorithms). In some embodiments, the processed data can be further manipulated by calculating Z-scores for one or more selected portions, chromosomes, or portions of chromosomes. In some embodiments, the processed data set can be further manipulated by calculating p-values. See Equation 1 (Example 2) for an implementation of the equation for calculating Z-scores and p-values. In some embodiments, the mathematical and / or statistical operations include one or more assumptions related to ploidy and / or fetal fraction. In some embodiments, further manipulation by one or more mathematical and / or statistical operations produces a profile of the processed data to facilitate classification and / or provide results. Results can be provided based on a profile of the data from mathematical and / or statistical operations. Results provided based on a profile of the data from mathematical and / or statistical operations typically include one or more assumptions related to ploidy and / or fetal fraction.

[0296] In some implementations, the data set is computed, optionally filtered, and standardized, and then various operations are performed on the processed data set to produce an N-dimensional space and / or N-dimensional points. Results can be provided based on an overview plot of the data set analyzed in N dimensions.

[0297] In some implementations, the data set is processed using one or more peak level analyses, peak width analyses, peak edge position analyses, peak lateral tolerance analyses, or their derivative analyses, or combinations thereof, as part of or after the processed and / or manipulated data set. In some implementations, one or more peak level analyses, peak width analyses, peak edge position analyses, peak lateral tolerance analyses, or their derivative analyses, or combinations thereof, are used to generate a profile of the processed data to facilitate classification and / or provide results. Results may be provided based on a profile of the data, which has been processed using one or more peak level analyses, peak width analyses, peak edge position analyses, peak lateral tolerance analyses, or their derivative analyses, or combinations thereof.

[0298] In some embodiments, one or more reference samples substantially free of the genetic variant under study can be used to generate a reference median count profile, which yields a predetermined value representing the absence of the genetic variant and typically deviates from a predetermined value in the area corresponding to a genomic location in the test subject where the genetic variant is located, if the test subject has the genetic variant. In test subjects with a condition associated with the genetic variant or at risk of such a variant, the numerical value of the selected portion or segment is expected to differ significantly from the predetermined value for the unaffected genomic location. In some embodiments, one or more reference samples known to carry the genetic variant under study can be used to generate a reference median count profile, which yields a predetermined value representing the presence of the genetic variant and typically deviates from a predetermined value in the area corresponding to a genomic location where the test subject does not have the genetic variant. In test subjects without a condition associated with the genetic variant or at no risk of such a variant, the numerical value of the selected portion or segment is expected to differ significantly from the predetermined value for the affected genomic location.

[0299] In some implementations, the analysis and processing of data can include the use of one or more hypotheses. An appropriate number or type of hypotheses may be used to analyze or process the dataset. Non-limiting examples of hypotheses that can be used for data processing and / or analysis include maternal ploidy, fetal baseline, prevalence of certain sequences in a reference population, racial background, prevalence of medical conditions selected from relevant family members, correspondence between raw count distributions from different patients and / or runs after GC normalization and repetition masking (e.g., GCRM), identical matches (e.g., identical base positions) representing PCR artifacts, inherent assumptions in fetal quantification assays (e.g., FQA), assumptions about twins (e.g., if there are two twins and only one is affected, the effective fetal score is only 50% of the total fetal score measured (similar to triplets, quadruplets, etc.)), uniformly covering the entire genome of fetal cell-free DNA (e.g., cfDNA), and combinations thereof.

[0300] In those examples where the quality and / or depth of the mapped sequence reads cannot predict the presence of genetic variation at the desired confidence level (e.g., 95% or higher), one or more additional mathematical processing algorithms and / or statistical prediction algorithms can be used, based on a standardized count distribution, to generate additional numerical values ​​that can be used for data analysis and / or to provide results. The term "standardized count distribution" as used herein refers to a distribution generated using standardized counts. This document describes examples of methods that can be used to generate standardized counts and standardized count distributions. The already counted location sequence reads can be standardized relative to the test sample count or the reference sample count. In some embodiments, the standardized count profile can be represented graphically.

[0301] Overview

[0302] In some implementations, the processing steps may include generating one or more profiles (e.g., profile plots) from various data sets or their derivatives (e.g., the results of one or more mathematical and / or statistical data processing steps known in the art and / or described herein).

[0303] The term "profile" in this document refers to the result of mathematical and / or statistical operations on data that facilitate the identification of patterns and / or correlations in large datasets. A profile typically includes values ​​obtained from one or more operations on data or a group of data based on one or more criteria. A profile typically includes multiple data points. Any suitable number of data points can be included in a profile, depending on the nature and / or complexity of the data group. In some implementations, a profile may include 2 or more data points, 3 or more data points, 5 or more data points, 10 or more data points, 24 or more data points, 25 or more data points, 50 or more data points, 100 or more data points, 500 or more data points, 1000 or more data points, 5000 or more data points, 10,000 or more data points, or 100,000 or more data points.

[0304] In some implementations, the profile is a representation of the entire data set, and in other implementations, the profile is a representation of a portion or subset of the data set. That is, a profile sometimes includes data points representing or generated from data that has not been filtered to remove any data, and sometimes includes data points representing or generated from data that has been filtered to remove unwanted data. In some implementations, the data points in the profile represent the results of data operations on a portion of the data. In some implementations, the data points in the profile include the results of data operations on a portion of the data set. In some implementations, the portions of the data set may be adjacent to each other, and in some implementations, the portions of the data set may originate from different parts of a chromosome or genome.

[0305] Data points derived from a profile of a data set can represent any suitable data classification. Non-limiting examples of data grouping to generate profile data point categories include: size-based portions, sequence feature-based portions (e.g., GC content, AT content, chromosome position (e.g., short arm, long arm, centromere, telomere), etc.), expression levels, chromosomes, etc., or combinations thereof. In some implementations, a profile can be generated from data points derived from other profiles (e.g., normalized data profiles re-normalized to different normalization values ​​to generate renormalized data profiles). In some implementations, profiles generated from data points derived from other profiles reduce the number of data points and / or the complexity of the data set. Reducing the number of data points and / or the complexity of the data set generally facilitates data interpretation and / or results delivery.

[0306] A profile (e.g., a genome profile, chromosome profile, chromosomal segment profile) is typically a collection of normalized or non-normalized counts of two or more parts. A profile typically includes at least one level (e.g., a genome segment level) and typically includes two or more levels (e.g., a profile often has multiple levels). Levels are typically used for groups of parts having approximately the same count or normalized count. Levels are described in detail herein. In some embodiments, a profile includes one or more parts that may be processed or transformed by weighting, removal, filtering, normalization, adjustment, averaging (to obtain a mean), addition, subtraction, or any combination thereof. A profile typically includes normalized counts mapped to parts defining two or more levels, wherein the counts are further normalized according to one of the levels using a suitable method. Typically, profile counts (e.g., profile levels) are associated with uncertain values.

[0307] Profiles including one or more levels are sometimes filled (e.g., well-filled). Filling (e.g., well-filled) refers to the process of identifying and adjusting the levels in a profile that originate from maternal microdeletions or maternal duplications (e.g., copy number variations). In some embodiments, levels originating from fetal microduplications or fetal microdeletions are filled. In some embodiments, the overall level of microduplications or microdeletions in a profile may be artificially increased or decreased, resulting in false positives or false negatives in the determination of chromosomal aneuploidy (e.g., trisomy). In some embodiments, levels in a profile that originate from microduplications and / or deletions are identified and adjusted (e.g., filled and / or removed) by a process sometimes referred to as filling or well-filled. In some embodiments, a profile includes one or more first levels that are distinctly different from second levels within the profile, each of the one or more first levels including maternal copy number variation, fetal copy number variation, or both maternal copy number variation and fetal copy number variation, and one or more of the first levels are adjusted.

[0308] A profile including one or more levels may include a first level and a second level. In some embodiments, the first level is different from (e.g., significantly different from) the second level. In some embodiments, the first level includes a first set of portions, the second level includes a second set of portions, and the first set of portions is not a subset of the second set of portions. In some embodiments, the first set of portions is different from the second set of portions, thereby determining the first and second levels. In some embodiments, a profile may have multiple first levels that are different from (e.g., significantly different, for example, having significantly different values) the second level within the profile. In some embodiments, the profile includes one or more first levels that are significantly different from the second level within the profile, and said one or more first levels are adjusted. In some embodiments, the profile includes one or more first levels that are significantly different from the second level within the profile, each of said one or more first levels including maternal copy number variation, fetal copy number variation, or maternal copy number variation and fetal copy number variation, and said one or more first levels are adjusted. In some embodiments, the first levels in the profile are removed from the profile or adjusted (e.g., padded). A profile may include multiple levels, said multiple levels including one or more first levels that are significantly different from one or more second levels, typically the dominant level in the profile is the second level, wherein the second levels are approximately equal to each other. In some implementations, levels greater than 50%, 60%, 70%, 80%, 90%, or 95% in the overview are considered second levels.

[0309] A profile is sometimes displayed as a graph. For example, one or more levels representing partial counts (e.g., standardized counts) can be plotted and visualized. Non-limiting examples of profile graphs that can be generated include raw counts (e.g., raw count profile or raw profile), standardized counts, partial-weighted, Z-scores, p-values, area ratios and fit ploidy, the ratio between the median level and the fitted and measured fetal score, principal components, etc., or combinations thereof. In some implementations, profile graphs allow for observation of manipulated data. In some implementations, profile graphs can be used to provide results (e.g., area ratios and fit ploidy, the ratio between the median level and the fitted and measured fetal score, principal components). As used herein, the term "raw count profile graph" or "raw profile graph" refers to a graph of counts in different parts of a region normalized to the total regional count (e.g., genome, part, chromosome, chromosomal part of a reference genome, or chromosomal segment). In some implementations, a static window procedure can be used to generate the profile, and in some implementations, a sliding window procedure can be used to generate the profile.

[0310] Profiles generated for test subjects are sometimes compared with profiles generated for one or more reference subjects to facilitate the explanation of mathematical and / or statistical operations on the data set and / or to provide results. In some implementations, profiles are generated based on one or more initial hypotheses (e.g., maternal nucleic acid contribution (e.g., total maternal score), fetal nucleic acid contribution (e.g., fetal score), reference sample ploidy, etc., or combinations thereof). In some implementations, test profiles are typically centered on a predetermined value representing the absence of genetic variation, and deviated from the predetermined value in the corresponding area of ​​the genomic location in the test subject where genetic variation is typically located (if the test subject has genetic variation). In test subjects with a condition associated with or at risk of such genetic variation, the numerical values ​​of the selected portion are expected to differ significantly from the predetermined values ​​for unaffected genomic locations. Based on initial hypotheses (e.g., fixed ploidy or optimal ploidy, fixed fetal score or optimal fetal score, or combinations thereof), predetermined thresholds or cutoff values ​​or threshold ranges indicating the presence of genetic variation may vary, but they still provide results that can be used to determine the presence of genetic variation. In some implementations, profiles indicate and / or represent phenotypes.

[0311] As a non-limiting example, a standardized sample and / or reference count profile can be obtained from raw sequence reading data by: (a) calculating the median reference count of selected chromosomes, portions, or segments from a set of references known to not carry genetic variation; (b) removing information-insensitive portions from the raw counts of the reference samples (e.g., filtering); (c) standardizing the remaining total counts of the entire remaining portion of the reference genome for the selected chromosomes or selected genomic locations of the reference samples (e.g., the sum of the remaining counts after removing information-insensitive portions of the reference genome), thereby producing a standardized reference profile; (d) removing the corresponding portions from the test sample; and (e) standardizing the remaining test sample counts for one or more selected genomic locations for the sum of the remaining median reference counts of the chromosomes containing the selected genomic locations, thereby producing a standardized test sample profile. In some embodiments, additional standardization steps involving the entire genome (reduced by the filtered portions in (b)) may be included between (c) and (d).

[0312] A dataset profile can be generated through one or more processing methods on sequence read data mapped by counting. Some implementations include the following: Mapping sequence reads and determining the number of counts (i.e., sequence tags) mapped to each genomic segment (e.g., counts). Generating a raw count profile from the counted mapped sequence reads. In some implementations, results are provided by comparing the raw count profile of the test subject with a reference median count profile of chromosomes, portions, or segments of a reference subject group known to be free of genetic variation.

[0313] In some implementations, the sequence reading data may optionally be filtered to remove noisy or uninformative portions. After filtering, the remaining counts are typically summed to generate a filtered set of data. In some implementations, a filtered count profile is generated from the filtered set of data.

[0314] After counting and optional filtering, sequence read data can be standardized to generate levels or profiles. Data sets can be standardized by standardizing one or more selected portions to a suitable standardization reference value. In some embodiments, the standardization reference value represents the total count of chromosomes from a selected portion. In some embodiments, the standardization reference value represents one or more corresponding portions of chromosomes in a reference data set prepared from a reference group known to be free of genetic variation. In some embodiments, the standardization reference value represents one or more corresponding portions of chromosomes in a test subject data set prepared from test subjects analyzed for the presence or absence of genetic variation. In some embodiments, the standardization process uses a static window method, and in some embodiments, the standardization process uses a moving or sliding window method. In some embodiments, generating a profile including standardized counts facilitates classification and / or provides results. Results can be provided based on a profile plot including standardized counts (e.g., using this profile plot).

[0315] level

[0316] In some implementations, values ​​(e.g., numerical values, quantitative values) are assigned to levels. Counts can be determined by suitable methods, operations, or mathematical processes (e.g., processed levels). Levels are typically or derived from a subset of counts (e.g., standardized counts). In some implementations, the level of a subset is substantially equal to the total number of counts mapped to the subset (e.g., counts, standardized counts). Levels are typically determined from counts that have been processed, transformed, or manipulated by suitable methods, operations, or mathematical processes known in the art. In some implementations, levels are derived from processed counts, non-limiting examples of which include weighted, removed, filtered, standardized, adjusted, averaged, derived arithmetic mean (e.g., arithmetic average level), added, subtracted, transformed counts, or combinations thereof. In some implementations, levels include standardized counts (e.g., partially standardized counts). Levels can be standardized by suitable processes, non-limiting examples of which include component-by-component standardization, GC content standardization, linear and nonlinear least squares regression, GC LOESS, LOWESS, PERUN, RM, GCRM, cQn, etc., and / or combinations thereof. Levels may include standardized counts or relative quantities of counts. In some embodiments, the level is used for averaged counts or standardized counts of two or more portions, and the level refers to the average level. In some embodiments, the level is used for a group of counts or portions of an arithmetic mean of standardized counts, referred to as the arithmetic average level. In some embodiments, the level is derived from portions of counts that include both raw and / or filtered counts. In some embodiments, the level is based on the raw counts. In some embodiments, the level is associated with an uncertainty (e.g., standard deviation, MAD). In some embodiments, the level is represented by a Z-score or p-value.

[0317] In this article, the term "level" in one or more sections is synonymous with "genomic segment level." Sometimes, the term "level" is used synonymously with the term "elevation." The meaning of the term "level" can be determined by its context. For example, the term "level" in context referring to genomic segments, profiles, readings, and / or counts typically indicates elevation. The term "level" in context referring to substances or components (e.g., RNA level, cluster level) typically indicates quantity. The term "level" in context referring to uncertainty (e.g., error level, confidence level, bias level, uncertainty level) typically indicates quantity.

[0318] Standardized or unstandardized counts of two or more levels (e.g., two or more levels in a profile) can sometimes be standardized based on the levels using mathematical operations (e.g., addition, multiplication, averaging, standardization, etc., or combinations thereof). For example, standardized or unstandardized counts of two or more levels can be standardized based on one, some, or all of the levels in the profile. In some embodiments, standardized or unstandardized counts of all levels in the profile are standardized based on one level in the profile. In some embodiments, standardized or unstandardized counts of a first level in the profile are standardized based on standardized or unstandardized counts of a second level in the profile.

[0319] Non-limiting examples of levels (e.g., first level, second level) are group levels that include portions of processed counts, group levels that include portions of the arithmetic mean, median, or average of counts, group levels that include portions of standardized counts, and any combination thereof. In some embodiments, the first and second levels in the overview are derived from counts mapped to portions of the same chromosome. In some embodiments, the first and second levels in the overview are derived from counts mapped to portions of different chromosomes.

[0320] In some implementations, the level is determined from normalized or unnormalized counts mapped to one or more portions. In some implementations, the level is determined from normalized or unnormalized counts mapped to two or more portions, wherein the normalized counts of each portion are generally approximately the same. For a given level, the counts (e.g., normalized counts) within a group of portions may differ. For a given level, there may be one or more portions within a group that have counts significantly different from those of other portions of the group (e.g., peaks and / or sloping). Any suitable number of normalized or unnormalized counts associated with any suitable number of portions can define the level.

[0321] In some embodiments, one or more levels can be determined from normalized or non-normalized counts of all or some portions of the genome. Typically, levels can be determined from all or some normalized or non-normalized counts of chromosomes or segments thereof. In some embodiments, levels are determined from two or more counts derived from two or more portions (e.g., groups of portions). In some embodiments, levels are determined from two or more counts (e.g., counts from two or more portions). In some embodiments, levels are determined from 2 to about 100,000 portions. In some embodiments, levels are determined from 2 to about 50,000, 2 to about 40,000, 2 to about 30,000, 2 to about 20,000, 2 to about 10,000, 2 to about 5000, 2 to about 2500, 2 to about 1250, 2 to about 1000, 2 to about 500, 2 to about 250, 2 to about 100, or 2 to about 60 portions. In some embodiments, levels are determined from about 10 to about 50 portions. In some embodiments, counts of about 20 to about 40 or more portions determine the level. In some embodiments, the level includes counts from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60 or more portions. In some embodiments, the level corresponds to a group of portions (e.g., a group of portions of a reference genome, a group of chromosomal portions, or a group of chromosomal segment portions).

[0322] In some embodiments, the level is determined based on the normalized or non-normalized counts of adjacent portions. In some embodiments, adjacent portions (e.g., groups of portions) represent adjacent segments of the genome or adjacent segments of chromosomes or genes. For example, when merging portions tail-to-tail, two or more adjacent portions may represent a set of sequences of DNA sequences longer than each portion. For example, two or more adjacent portions may represent an entire genome, chromosome, gene, intron, exon, or segment thereof. In some embodiments, the level is determined from a set (e.g., groups) of adjacent portions and / or non-adjacent portions.

[0323] Different levels

[0324] In some embodiments, the standardized count profile includes a level (e.g., a first level) that is significantly different from other levels (e.g., a second level) within the profile. The first level may be higher or lower than the second level. In some embodiments, the first level is used for a group comprising portions of one or more readings that include copy number variation (e.g., maternal copy number variation, fetal copy number variation, or both maternal and fetal copy number variation), and the second level is used for a group comprising portions of readings that are substantially without copy number variation. In some embodiments, significant difference refers to an observable difference. In some embodiments, significant difference refers to statistical difference or statistically significant difference. Statistically significant difference is sometimes a statistical estimate of an observable difference. Statistically significant difference can be estimated using methods suitable in the art. Any suitable threshold or range can be used to determine two significantly different levels. In some embodiments, a difference of about 0.01% or more (e.g., 0.01% of one or the other level value) between two levels (e.g., the arithmetic mean) is considered significantly different. In some embodiments, a difference of about 0.1% or more between two levels (e.g., the arithmetic mean) is considered significantly different. In some embodiments, a difference of about 0.5% or more between two levels (e.g., the arithmetic mean) is considered significantly different. In some implementations, a difference of approximately 0.5, 0.75, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or greater than 10% between two levels is considered significantly different. In some implementations, two levels (e.g., the arithmetic mean) are significantly different and there is no overlap between the levels and / or no overlap within the defined range of the uncertainty calculated for one or both levels. In some implementations, the uncertainty is a standard deviation, expressed as... In some implementations, the two levels (e.g., the arithmetic mean) are significantly different, differing by about one or more times the uncertainty value (e.g. In some implementations, the two levels (e.g., the arithmetic mean) are significantly different, differing by about two or more times the uncertainty value (e.g., ...). The confidence level can be approximately 3 or more, approximately 4 or more, approximately 5 or more, approximately 6 or more, approximately 7 or more, approximately 8 or more, approximately 9 or more, or approximately 10 or more times the uncertainty. In some embodiments, the confidence level is significantly different when the difference between two levels (e.g., the arithmetic mean) is approximately 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, or 4.0 times or more of the uncertainty. In some embodiments, the confidence level increases with the increase in the difference between the two levels. In some embodiments, the confidence level decreases with the decrease in the difference between the two levels and / or the increase in the uncertainty. For example, sometimes the confidence level increases proportionally to the difference between the level and the standard deviation (e.g., MAD).

[0325] One or more prediction algorithms can be used to determine significance or to give meaning to the collected detection data under variable conditions. Their weights can be independent or interdependent. The term "variable" used in this paper refers to a factor, quantity, or function in the algorithm that has one or more values.

[0326] In some implementations, the first group typically includes portions that are different from (e.g., do not overlap with) the second group. For example, sometimes a first level of normalized counts is significantly different from a second level of normalized counts in a profile, and the first level applies to the first group, the second level applies to the second group, and the portions do not overlap between the first and second groups. In some implementations, the first group is not a subset of the second group, thereby determining the first and second levels separately. In some implementations, the first group differs from and / or varies with the second group, thereby determining the first and second levels separately.

[0327] In some implementations, the first set of portions is a subset of the second set of portions in the profile. For example, sometimes a second level of standardized counts for the second set of portions in the profile includes a first level of standardized counts for the first set of portions in the profile, and the first set of portions is a subset of the second set of portions in the profile. In some implementations, the mean, arithmetic mean, or median level is derived from the second level, wherein the second level includes the first level. In some implementations, the second level includes a second set of portions representing the entire chromosome, and the first level includes the first set of portions, wherein the first set is a subset of the second set of portions, and the first level represents maternal copy number variation, fetal copy number variation, or maternal copy number variation and fetal copy number variation present in the chromosome.

[0328] In some embodiments, the value of the second level is closer to the arithmetic mean, average, or median of the count profile of a chromosome or its segment than the first level. In some embodiments, the second level is the arithmetic average of the levels of a chromosome, a portion of a chromosome, or a segment of a chromosome. In some embodiments, the first level is significantly different from the dominant level (e.g., the second level) representing a chromosome or its segment. A profile may include multiple first levels that are significantly different from the second level, and each first level may be independently higher or lower than the second level. In some embodiments, the first and second levels originate from the same chromosome, and the first level is higher or lower than the second level, where the second level is the dominant level of the chromosome. In some embodiments, the first and second levels originate from the same chromosome, the first level indicates copy number variation (e.g., maternal and / or fetal copy number variation, deletion, insertion, duplication), and the second level is the arithmetic average or dominant level of a portion of a chromosome or its segment.

[0329] In some implementations, the readings in the second set of portions of the second level substantially exclude genetic variations (e.g., copy number variations, maternal and / or fetal copy number variations). Typically, the second set of portions of the second level includes some variability (e.g., level variability, portion count variability). In some implementations, one or more portions of a set of portions associated with a level substantially free of copy number variations include readings of one or more copy number variations present in the maternal and / or fetal genome. For example, sometimes a set of portions includes copy number variations present in small chromosomal segments (e.g., fewer than 10 portions) and the set of portions is used for a level associated with substantially free copy number variations. Therefore, a set of portions substantially excluding copy number variations may still include copy number variations present in fewer than about 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 portions of the level.

[0330] In some embodiments, a first level is used for a first group portion and a second level is used for a second group portion, and the first group portion and the second group portion are adjacent (e.g., nucleic acid sequences relating to a chromosome or a segment thereof are adjacent). In some embodiments, the first group portion and the second group portion are not adjacent.

[0331] Relatively short sequence reads of a mixture of fetal and maternal nucleic acids can be used to provide counts, which can be transformed into levels and / or profiles. Counts, levels, and profiles can be described in electronic or tangible form and can be visualized. Counts mapped to portions (e.g., representing levels and / or profiles) can provide a visual representation of the fetal and / or maternal genome, chromosomes, or chromosomal portions or segments present in the fetus and / or pregnant woman.

[0332] Reference levels and standardized reference values

[0333] In some embodiments, the profile includes a reference level (e.g., a level used as a reference). Typically, the profile of normalized counts provides a reference level from which expected levels and expected ranges are determined (see Expected Levels and Ranges below). The reference level is typically used for a portion of the normalized counts that includes mapping readings from the mother and fetus. The reference level is typically the sum of normalized counts of mapping readings from the fetus and mother (e.g., a pregnant woman). In some embodiments, the reference level includes a portion of the mapping readings from an euploid mother and / or an euploid fetus. In some embodiments, the reference level is used for a portion of the mapping readings that include fetal and / or maternal genetic variations (e.g., aneuploidy (e.g., trisomy), copy number variation, microduplication, microdeletion, insertion). In some embodiments, the reference level is used for a portion that substantially excludes fetal and / or maternal genetic variations (e.g., aneuploidy (e.g., trisomy), copy number variation, microduplication, microdeletion, insertion). In some embodiments, a second level is used as the reference level. In some embodiments, the profile includes a first level of normalized counts and a second level of normalized counts, the first level being significantly different from the second level and the second level being the reference level. In some implementations, the profile includes a first level of standardized counts for a first group portion, a second level of standardized counts for a second group portion, the first group portion including mapped readings with maternal and / or fetal copy number variations, the second group portion including mapped readings substantially without maternal and / or fetal copy number variations, and the second level being a reference level.

[0334] In some implementations, counts mapping one or more levels of a profile to portions are standardized based on counts of a reference level. In some implementations, standardizing level counts based on a reference level count involves dividing the level count by the reference level count or a multiple or fraction thereof. Counts standardized based on a reference level count are typically standardized according to other processes (e.g., PERUN), and the reference level count is also typically standardized (e.g., via PERUN). In some implementations, level counts are standardized based on a reference level count, and the reference level count may be expanded to a suitable value before or after standardization. The process of expanding the reference level count may include any suitable constant (i.e., a number), and any suitable mathematical operation may be used for the reference level count.

[0335] A standardized reference value (NRV) is typically determined based on the count of a standardized reference level. Determining the NRV may include any suitable standardization process (e.g., mathematical operation) used for the count of the reference level, where the same standardization process is used to standardize the counts of other levels within the same profile. Determining the NRV typically involves dividing the reference level by itself. Determining the NRV typically involves dividing the reference level by a multiple of itself. Determining the NRV typically involves dividing the reference level by the sum or difference of the reference level and a constant (such as any number).

[0336] NRV sometimes refers to a null value. NRV can be any suitable value. In some implementations, NRV is any value other than 0. In some implementations, NRV is an integer. In some implementations, NRV is a positive integer. In some implementations, NRV is 1, 10, 100, or 1000. Typically, NRV equals 1. In some implementations, NRV equals 0. The reference level count can be normalized to any suitable NRV. In some implementations, the reference level count is normalized to an NRV of 0. Typically, the reference level count is normalized to an NRV of 1.

[0337] Expected level

[0338] Sometimes, the expected level is a predefined level (e.g., theoretical level, predicted level). Sometimes, the “expected level” is referred to herein as a “predetermined level value.” In some embodiments, the expected level is a predicted value of the level of a standardized count of a set that includes portions of copy number variation. In some embodiments, the expected level is determined for a set that substantially does not contain portions of copy number variation. Expected levels for chromosome ploidy (e.g., 0, 1, 2 (i.e., diploid), 3, or 4 chromosomes) or microploidy (homozygous or heterozygous deletions, duplications, insertions, or their absence) can be determined. Typically, the expected level for maternal microploidy (e.g., maternal and / or fetal copy number variation) is determined.

[0339] The expected level of genetic variation or copy number variation can be determined by any suitable method. Typically, the expected level is determined by applying appropriate mathematical processing to the level (e.g., mapping a count of a subset of the set to a given level). In some implementations, the expected level is sometimes determined using a constant (called the expected level constant). Sometimes, the expected level of copy number variation is calculated by multiplying, adding, subtracting, dividing by, or combining a reference level, NRV, or a standardized count of the reference level by the expected level constant. Typically, the expected level for the same subject, sample, or test group (e.g., the expected level of maternal and / or fetal copy number variation) is determined based on the same reference level or NRV.

[0340] Typically, the expected level is determined by multiplying a reference level, its normalized count, or NRV by an expected level constant, where the reference level, its normalized count, or NRV is not equal to zero. In some embodiments, the expected level is determined by adding the expected level constant to an NRV, the reference level, or its normalized count that is equal to zero. In some embodiments, the expected level, the normalized count of the reference level, the NRV, and the expected level constant are scalable. The scaling process can include any suitable constant (i.e., a number) and any suitable mathematical operations, wherein the same scaling process is applied to all values ​​considered.

[0341] Expected level constant

[0342] The expected level constant can be determined by a suitable method. In some embodiments, the expected level constant is determined arbitrarily. Typically, the expected level constant is determined based on experience. In some embodiments, the expected level constant is determined according to mathematical operations. In some embodiments, the expected level constant is determined based on a reference (e.g., a reference genome, a reference sample, reference test data). In some embodiments, the expected level constant is predetermined for the level representing the presence or absence of genetic variation or copy number variation (e.g., duplication, insertion, or deletion). In some embodiments, the expected level constant is predetermined for the level representing the presence or absence of maternal copy number variation, fetal copy number variation, or both maternal and fetal copy number variation. The expected level constant for copy number variation can be any suitable constant or set of constants.

[0343] In some embodiments, the expected level constant for homozygous repeats (e.g., homozygous repeats) may be about 1.6 to about 2.4, about 1.7 to about 2.3, about 1.8 to about 2.2, or about 1.9 to about 2.1. In some embodiments, the expected level constant for homozygous repeats is about 1.6, 1.7, 1.8, 1.9, 2.0, 2.1, 2.2, 2.3, or about 2.4. Typically, the expected level constant for homozygous repeats is about 1.90, 1.92, 1.94, 1.96, 1.98, 2.0, 2.02, 2.04, 2.06, 2.08, or about 2.10. Typically, the expected level constant for homozygous repeats is about 2.

[0344] In some embodiments, the expected level constant for heterozygous repeats (e.g., homozygous repeats) is about 1.2 to about 1.8, about 1.3 to about 1.7, or about 1.4 to about 1.6. In some embodiments, the expected level constant for heterozygous repeats is about 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, or about 1.8. Typically, the expected level constant for heterozygous repeats is about 1.40, 1.42, 1.44, 1.46, 1.48, 1.5, 1.52, 1.54, 1.56, 1.58, or about 1.60. In some embodiments, the expected level constant for heterozygous repeats is about 1.5.

[0345] In some embodiments, the expected level constant for the absence of copy number variation (e.g., absence of maternal copy number variation and / or fetal copy number variation) is about 1.3 to about 0.7, about 1.2 to about 0.8, or about 1.1 to about 0.9. In some embodiments, the expected level constant for the absence of copy number variation is about 1.3, 1.2, 1.1, 1.0, 0.9, 0.8, or about 0.7. Typically, the expected level constant for the absence of copy number variation is about 1.09, 1.08, 1.06, 1.04, 1.02, 1.0, 0.98, 0.96, 0.94, or about 0.92. In some embodiments, the expected level constant for the absence of copy number variation is about 1.

[0346] In some embodiments, the expected level constant for heterozygous deletion (e.g., maternal, fetal, or maternal and fetal heterozygous deletion) is about 0.2 to about 0.8, about 0.3 to about 0.7, or about 0.4 to about 0.6. In some embodiments, the expected level constant for heterozygous deletion is about 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, or about 0.8. Typically, the expected level constant for heterozygous deletion is about 0.40, 0.42, 0.44, 0.46, 0.48, 0.5, 0.52, 0.54, 0.56, 0.58, or about 0.60. In some embodiments, the expected level constant for heterozygous deletion is about 0.5.

[0347] In some embodiments, the expected level constant for homozygous deletions (e.g., homozygous deletions) may be about -0.4 to about 0.4, about -0.3 to about 0.3, about -0.2 to about 0.2, or about -0.1 to about 0.1. In some embodiments, the expected level constant for homozygous deletions is about -0.4, -0.3, -0.2, -0.1, 0.0, 0.1, 0.2, 0.3, or about 0.4. Typically, the expected level constant for homozygous deletions is about -0.1, -0.08, -0.06, -0.04, -0.02, 0.0, 0.02, 0.04, 0.06, 0.08, or about 0.10. Typically, the expected level constant for homozygous deletions is about 0.

[0348] Expected range

[0349] In some implementations, the presence or absence of genetic variation or copy number variation (e.g., maternal copy number variation, fetal copy number variation, or both maternal and fetal copy number variation) is determined by levels falling within or outside a expected range. The expected range is typically determined based on the expected level. In some implementations, the expected range is determined by levels that include essentially no genetic variation or essentially no copy number variation. Suitable methods may be used to determine the expected range.

[0350] In some implementations, the expected range of levels is determined based on an appropriate uncertainty value calculated for each level. Non-limiting examples of uncertainty values ​​include standard deviation, standard error, calculated variance, p-value, or arithmetic mean absolute deviation (MAD). In some implementations, the expected range of genetic variation or copy number variation is partially determined by calculating the uncertainty values ​​of levels (e.g., a first level, a second level, a first level, and a second level). In some implementations, the expected range of levels is determined based on uncertainty values ​​calculated for a profile (e.g., a profile of standardized counts of chromosomes or their segments). In some implementations, uncertainty values ​​are calculated for levels that contain substantially no genetic variation or substantially no copy number variation. In some implementations, uncertainty values ​​are calculated for a first level, a second level, or a first level and a second level. In some implementations, uncertainty values ​​are determined for a first level, a second level, or a second level that includes the first level.

[0351] Sometimes, the expected level range is calculated by multiplying, adding, subtracting, or dividing the uncertainty value by a constant (e.g., a predetermined constant) n. Suitable combinations of mathematical operations or processes may be used. Sometimes, the constant n (...

Claims

1. A system comprising one or more microprocessors and a memory, wherein the memory contains instructions executable by the one or more microprocessors, and wherein the instructions executable by the one or more microprocessors are used for: (a) Obtain the count of sequence reads mapped to various parts of the reference genome, where the sequence reads are readings of circulating cell-free (CCF) nucleic acids from test samples from pregnant women. (b) Select a subset of the subset, thereby providing a count subset, in which (i) The selection is based on the portion of the fetal nucleic acid readings that have been mapped to an increased number, and (ii) The portion of the mapping with an increased number of fetal nucleic acid readings is determined based on the ratio of X to Y, wherein, X is the number of readings derived from CCF segments shorter than the first selected segment length, and Y is the number of readings derived from CCF segments shorter than the second selected segment length; and (c) Evaluate the fetal nucleic acid score of the test sample based on the count subset.

2. The system as claimed in claim 1, wherein, The ratio is the average ratio of multiple samples.

3. The system as described in claim 2, wherein, The selection is based on the fact that the average ratio of a certain part is greater than the average ratio of the average of all parts.

4. The system as claimed in claim 1, wherein, The first selected fragment is about 140 to about 160 bases in length and the second selected fragment is about 500 to about 700 bases in length.

5. The system as described in claim 4, wherein, The first selected fragment is approximately 150 bases long and the second selected fragment is approximately 600 bases long.

6. The system as described in any one of claims 1 to 5, wherein, The counts are standardized counts.

7. The system of claim 6, wherein, The counts are standardized based on guanine-cytosine (GC) content.

8. The system as claimed in any one of claims 1 to 7, wherein, A subset of the portion is a portion of one or more autosomes.

9. The system of claim 8, wherein, A subset of the portion is a portion of one or more euploid chromosomes.

10. A non-transitory computer-readable storage medium comprising instructions, wherein when the instructions are executed, a microprocessor performs the following operations: (a) Obtain the count of sequence reads mapped to various parts of the reference genome, where the sequence reads are readings of circulating cell-free (CCF) nucleic acids from test samples from pregnant women. (b) Select a subset of the subset, thereby providing a count subset, in which (i) The selection is based on the portion of the fetal nucleic acid readings that have been mapped to an increased number, and (ii) The portion of the mapping with an increased number of fetal nucleic acid readings is determined based on the ratio of X to Y, wherein, X is the number of readings derived from CCF segments shorter than the first selected segment length, and Y is the number of readings derived from CCF segments shorter than the second selected segment length; and (c) Evaluate the fetal nucleic acid score of the test sample based on the count subset.

11. The non-transitory computer-readable storage medium of claim 10, wherein, The ratio is the average ratio of multiple samples.

12. The non-transitory computer-readable storage medium of claim 11, wherein, The selection is based on the fact that the average ratio of a certain part is greater than the average ratio of the average of all parts.

13. The non-transitory computer-readable storage medium of claim 10, wherein, The first selected fragment is about 140 to about 160 bases in length and the second selected fragment is about 500 to about 700 bases in length.

14. The non-transitory computer-readable storage medium of claim 13, wherein, The first selected fragment is approximately 150 bases long and the second selected fragment is approximately 600 bases long.

15. The non-transitory computer-readable storage medium as claimed in any one of claims 10 to 14, wherein, The counts are standardized counts.

16. The non-transitory computer-readable storage medium of claim 15, wherein, The counts are standardized based on guanine-cytosine (GC) content.

17. The non-transitory computer-readable storage medium as claimed in any one of claims 10 to 16, wherein, A subset of the portion is a portion of one or more autosomes.

18. The non-transitory computer-readable storage medium of claim 17, wherein, A subset of the portion is a portion of one or more euploid chromosomes.

Citation Information

Patent Citations

  • Fragmentation-based methods and systems for sequence variation detection and discovery

    US20050112590A1

  • Method and compositions for detection and enumeration of genetic variations

    US20070065823A1

  • Diagnosing fetal chromosomal aneuploidy using massively parallel genomic sequencing

    US20090029377A1

  • Restriction endonuclease enhanced polymorphic sequence detection

    US20090317818A1

  • Processes and compositions for methylation-based enrichment of fetal nucleic acid from a maternal sample useful for non invasive prenatal diagnoses

    US20100105049A1