Methods and processes for non-invasive assessment of genetic variations
The system addresses invasive and biased sequencing issues by normalizing genomic bias and adjusting read density, enabling non-invasive and cost-effective genetic mutation detection.
Patent Information
- Application Number
- JP2025066232
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2013-10-04
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-23
AI Technical Summary
Current methods for identifying genetic mutations and variations are often invasive, costly, and prone to sequencing biases, which can lead to inaccurate results in diagnosing medical conditions or determining predispositions.
A system utilizing microprocessors and memory to reduce sequencing bias by generating and normalizing relationships between genomic bias and GC density, and adjusting read density profiles using principal component analysis to accurately detect gene mutations and aneuploidy.
This approach enables non-invasive, cost-effective, and accurate identification of genetic mutations and aneuploidy, improving diagnostic accuracy and reducing sequencing costs.
Smart Images

Figure 2025108574000001_ABST
Abstract
Description
Technical Field
[0001] Related Patent Application This patent application was filed on October 4, 2013, with the title "METHODS AND PROCESSES FOR NON-INVASIVE ASSESSMENT OF GENETIC VARIATIONS", named Gregory Hannum as the inventor, and claims the benefit of U.S. Provisional Patent Application No. 61 / 887,081, designated by docket number SEQ-6073-PV. The entire content of the above application, including all text, tables and drawings, is incorporated herein by reference.
[0002] The technology provided herein relates, in part, to methods, processes and machines for non-invasive assessment of genetic variations.
Background Art
[0003] The genetic information of living organisms (e.g., animals, plants and microorganisms) as well as other forms that replicate genetic information (e.g., viruses) is encoded in deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a continuous nucleotide or modified nucleotide, which represents the primary structure of a chemical or hypothetical nucleic acid. In the case of humans, the complete genome contains approximately 30,000 genes located on 24 chromosomes (see Non-Patent Document 1). Each gene encodes a specific protein, which, after being expressed through transcription and translation in living cells, performs specific biochemical functions.
[0004] Many medical conditions are caused by mutations in one or more genes. Mutations in certain specific genes cause medical conditions, such as, for example, hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease, and cystic fibrosis (CF) (Non-Patent Document 2). Such hereditary diseases can result from the addition, substitution, or deletion of a single nucleotide in the DNA of a specific gene. For example, certain congenital defects are caused by chromosomal abnormalities, also called aneuploidies, such as, by way of example, trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), trisomy 16, and trisomy 22, monosomy X (Turner syndrome), and certain sex chromosome aneuploidies, such as, by way of example, Klinefelter syndrome (XXY). Mutations in another gene are the sex of the fetus, which can often be determined based on the X and Y sex chromosomes. Mutations in some genes can make an individual more likely to develop, or be at risk of developing, one of several diseases, such as, for example, diabetes, arteriosclerosis, obesity, various autoimmune diseases, and cancer (e.g., colorectal cancer, breast cancer, ovarian cancer, lung cancer).
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0006] The identification of mutations or variations in one or more genes can lead to the diagnosis of a particular medical condition or the determination of a predisposition to such a condition. The identification of gene variations can result in the facilitation of medical decisions and / or the utilization of useful medical procedures. In certain embodiments, the identification of mutations or variations in one or more genes includes the analysis of cell-free DNA. Cell-free DNA (CF-DNA) results from cell death and consists of DNA fragments that circulate in peripheral blood. High concentrations of CF-DNA can be indicative of certain clinical conditions, such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infection, and other diseases. Additionally, cell-free fetal DNA (CFF-DNA) can be detected in the maternal bloodstream and used for various non-invasive prenatal diagnostic methods.
Means for Solving the Problem
[0007] In certain aspects herein, a system comprising a memory and one or more microprocessors, wherein the one or more microprocessors are configured to perform processing to reduce bias in reads about the sequence of a sample according to instructions in the memory, the processing comprising: (a) generating a relationship between (i) an estimated value of local genomic bias and (ii) a bias frequency for reads of the sequence of a test sample, thereby generating a sample bias relationship, wherein the reads of the sequence are those of cell-free circulating nucleic acid derived from the test sample and are mapped against a reference genome; (b) comparing the sample bias relationship with a reference bias relationship, thereby generating a comparison, wherein the reference bias relationship is between (i) an estimated value of local genomic bias and (ii) a bias frequency for the reference; and (c) normalizing the count number of reads of the sequence of the sample according to the comparison determined in (b), thereby reducing the bias in the reads of the sequence with respect to the sample. A system is provided that includes: (a) generating a relationship between (i) an estimated value of local genomic bias and (ii) a bias frequency for reads of the sequence of a test sample, thereby generating a sample bias relationship, wherein the reads of the sequence are those of cell-free circulating nucleic acid derived from the test sample and are mapped against a reference genome; (b) comparing the sample bias relationship with a reference bias relationship, thereby generating a comparison, wherein the reference bias relationship is between (i) an estimated value of local genomic bias and (ii) a bias frequency for the reference; and (c) normalizing the count number of reads of the sequence of the sample according to the comparison determined in (b), thereby reducing the bias in the reads of the sequence with respect to the sample.
[0008] In this specification, in certain embodiments, a system includes a memory and one or more microprocessors, the one or more microprocessors being configured to perform processing to reduce bias in reading sequences for a sample according to instructions in the memory, the processing including: (a) generating a relationship between (i) the guanine and cytosine (GC) density and (ii) the GC density frequency for reads of the sequence of a test sample, thereby generating a sample GC density relationship, wherein the reads of the sequence are of cell-free circulating nucleic acids derived from the test sample and the reads of the sequence are mapped to a reference genome; (b) comparing the sample GC density relationship with a reference GC density relationship, thereby generating a comparison, wherein the reference GC density relationship is between (i) the GC density and (ii) the GC density frequency for the reference; and (c) normalizing the count number of reads of the sequence for the sample according to the comparison determined in (b), thereby reducing bias in the reads of the sequence for the sample. A system is provided that includes these steps.
[0009] In certain embodiments herein, a system includes a memory and one or more microprocessors, where the one or more microprocessors are configured to perform a process for determining the presence or absence of aneuploidy for a sample according to instructions in the memory, the process comprising: (a) filtering a portion of a reference genome according to a read density distribution, thereby providing a read density profile for a test sample that includes the read density of the filtered portion, where the read density includes reads of sequences of cell-free circulating nucleic acids derived from a test sample from a pregnant female, and the read density distribution is determined for portions of read density for a plurality of samples; (b) adjusting the read density profile for the test sample according to one or more principal components obtained from a series of known euploid samples by principal component analysis, thereby providing a test sample profile that includes the adjusted read density; (c) comparing the test sample profile to a reference profile, thereby providing a comparison; and (d) determining the presence or absence of chromosomal aneuploidy for the test sample according to the comparison. A system is also provided that includes these steps.
[0010] Certain embodiments of this technology are further described in the following description, examples, claims, and drawings.
[0011] The drawings illustrate, but do not limit, embodiments of the technology. The drawings are not drawn to scale for clarity and ease of understanding, and in some instances, various aspects may be shown exaggerated or enlarged to facilitate understanding of particular embodiments.
Brief Description of the Drawings
[0012]
Figure 1
[0013]
Figure 2
[0014]
Figure 3
[0015]
Figure 4
[0016]
Figure 5
[0017]
Figure 6
[0018]
Figure 7
[0019]
Figure 8
[0020]
Figure 9
[0021]
Figure 10A
Figure 10B
[0022]
Figure 11
[0023]
Figure 12
Best Mode for Carrying Out the Invention
[0024] Next-generation sequencing enables the sequencing of nucleic acids on a genome-wide scale by methods that are faster and less expensive than traditional methods of sequencing. The methods, systems, and products provided herein can utilize advanced sequencing technologies to locate and identify gene mutations as well as / or associated diseases and disorders. The methods, systems, and products provided herein can often provide a non-invasive assessment of a subject genome (e.g., a fetal genome) using a blood sample or a portion thereof, and are often safer, faster, and / or less expensive than more invasive techniques (e.g., amniocentesis, biopsy). In some embodiments, provided herein is a method that in part includes obtaining a read of the sequence of nucleic acids present in a sample, where the read of the sequence is often mapped to a reference sequence, processing the count number of the read of the sequence, and determining the presence or absence of a gene mutation. The systems, methods, and products provided herein are useful for locating and / or identifying gene mutations and for diagnosing and treating diseases, disorders, and disabilities associated with mutations in certain genes.
[0025] Also, in some embodiments herein, a data manipulation method is provided for reducing and / or removing sequencing bias introduced by various aspects of sequencing technology. Sequencing bias often contributes to non-uniform distribution of reads across a genome or segment thereof, and / or variation in read quality. Sequencing bias can corrupt genomic sequencing data, impair valid data analysis, contaminate results, and prevent accurate data interpretation. Sometimes, sequencing bias can be reduced by increasing sequencing coverage, but this approach often inflates the sequencing cost and has very limited effectiveness. The data manipulation methods described herein can reduce and / or remove sequencing bias, thereby improving the quality of sequence read data without increasing the sequencing cost. Also, in some embodiments herein, a system, machine, apparatus, product, and module for implementing the methods described herein are provided.
[0026] sample Methods and compositions for analyzing nucleic acids are provided herein. In some embodiments, nucleic acid fragments in a mixture of nucleic acid fragments are analyzed. The mixture of nucleic acids can include two or more nucleic acid fragment species having different nucleotide sequences, different fragment lengths, different origins (e.g., genomic origin, fetal origin versus maternal origin, cell origin or tissue origin, sample origin, subject origin, etc.), or combinations thereof.
[0027] The nucleic acids or nucleic acid mixtures utilized in the methods, systems, machines, and / or devices described herein are often isolated from samples obtained from a subject (e.g., a test subject). The subject from which a specimen or sample is obtained is sometimes referred to herein as the test subject. The subject can be any living or non-living organism, including but not limited to humans, non-human animals, plants, bacteria, fungi, viruses, or protists. Any human or non-human animal can be selected, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, Bovidae (e.g., cows), Equidae (e.g., horses), Caprine and Ovine (e.g., sheep, goats), Swine (e.g., pigs), Camelidae (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), Ursidae (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The subject can be male or female (e.g., female, pregnant female, pregnant female animal). The subject can be of any age (e.g., embryo, fetus, neonate, infant, adult).
[0028] Nucleic acids can be isolated from any suitable biological specimen or sample of any type (e.g., a test sample). A sample or test sample can be any specimen isolated or obtained from a subject or a part thereof (e.g., a human subject, a pregnant female, a fetus). A test sample is often obtained from a test subject. A test sample is often obtained from a pregnant female (e.g., a pregnant human female). Non-limiting examples of specimens include body fluids or tissues obtained from a subject, including, without limitation, blood or blood products (e.g., serum, plasma, etc.), cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, spinal fluid, lavage fluids (e.g., those derived from bronchoalveolar, gastric, peritoneal, ductal, ear, arthroscopic procedures), biopsy samples (e.g., samples obtained from pre-implantation embryos), abdominal puncture samples, cells (blood cells, placental cells, embryonic or fetal cells, fetal nucleated cells or fetal cell remnants) or parts thereof (e.g., mitochondria, nuclei, extracts, etc.), washings of the female genital tract, urine, feces, sputum, saliva, nasal mucus, prostatic fluid, wash fluids, semen, lymph fluid, bile, tears, sweat, milk, breast fluid, etc., or combinations thereof. A test sample can include blood or blood products (e.g., plasma, serum, lymphocytes, platelets, buffy coat). A test sample sometimes includes serum obtained from a pregnant female. A test sample sometimes includes plasma obtained from a pregnant female. In some embodiments, the biological sample is a cervical swab obtained from a subject. In some embodiments, the biological sample can be blood and sometimes can be plasma or serum. The term "blood" as used herein refers to a blood sample or preparation from a subject (e.g., a test subject, e.g., a pregnant woman or a woman being tested for pregnancy). This term includes whole blood, blood products or any fraction of blood, including, by way of example, serum, plasma, buffy coat, etc. according to conventional definitions. Blood or its fractions often contain nucleosomes (e.g., maternal and / or fetal nucleosomes). Nucleosomes contain nucleic acids and are sometimes cell-free or intracellular nucleosomes. Blood also includes buffy coat. Buffy coat is sometimes isolated by using a Ficoll gradient. Buffy coat can contain white blood cells (e.g., leukocytes, T cells, B cells, platelets, etc.).In certain embodiments, the buffy coat contains maternal nucleic acids and / or fetal nucleic acids. Plasma refers to a fraction of whole blood obtained as a result of centrifugation of blood treated with an anticoagulant. Serum refers to the aqueous liquid portion remaining after the blood sample has coagulated. Body fluids or tissue samples are often collected according to standard protocols commonly followed in hospitals or outpatient clinics. In the case of blood, an appropriate amount of venipuncture blood (e.g., 3 to 40 milliliters) is often collected and can be stored according to standard procedures either before or after preparation. Body fluids or tissue samples from which nucleic acids are to be extracted may be cell-free (e.g., acellular). In some embodiments, body fluids or tissue samples may contain cellular elements or cell remnants. In some embodiments, fetal cells or cancerous cells may be included in the sample.
[0029] Often, the sample is heterogeneous, which means that more than one type of nucleic acid species is present in the sample. For example, heterogeneous nucleic acids include, but are not limited to, (i) nucleic acids derived from the fetus and nucleic acids derived from the mother, (ii) cancerous nucleic acids and non-cancerous nucleic acids, (iii) pathogen nucleic acids and host nucleic acids, and more generally, (iv) mutated nucleic acids and wild-type nucleic acids. The sample can be heterogeneous because more than one cell type, by way of example, fetal cells and maternal cells, cancerous cells and non-cancerous cells, or pathogen cells and host cells, are present. In some embodiments, both a small amount of nucleic acid species and a large amount of nucleic acid species are present.
[0030] When applying the techniques described herein prenatally, a body fluid or tissue sample can be collected from a female at a gestational age appropriate for testing, or from a female being tested for pregnancy. The appropriate gestational age can vary depending on the prenatal test being performed. In certain embodiments, the female subject during pregnancy is sometimes in the first trimester, sometimes in the second trimester, or sometimes in the third trimester. In certain embodiments, the body fluid or tissue is collected from a pregnant female at about 1 to about 45 weeks of gestation (e.g., 1 - 4, 4 - 8, 8 - 12, 12 - 16, 16 - 20, 20 - 24, 24 - 28, 28 - 32, 32 - 36, 36 - 40, or 40 - 44 weeks of gestation), and sometimes at about 5 to about 28 weeks of gestation (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, or 27 weeks of gestation). In certain embodiments, the body fluid or tissue sample is collected from a pregnant female during or immediately after delivery (e.g., vaginal delivery or non - vaginal delivery (e.g., surgical delivery)) (e.g., 0 - 72 hours later).
[0031] Obtaining a blood sample and extracting DNA The methods herein often include the isolation, enrichment, and analysis of fetal DNA found in maternal blood as a non - invasive means for detecting the presence or absence of genetic mutations in the mother and / or fetus during and sometimes after pregnancy, and / or for monitoring the health status of the fetus and / or the pregnant female. Thus, the first step in performing certain methods herein often includes obtaining a blood sample from a pregnant woman and extracting DNA from the sample.
[0032] Obtaining a blood sample A blood sample can be obtained from a pregnant woman at a gestational age appropriate for testing using the method according to the present technique. The appropriate gestational age can be varied according to the obstacles to be tested, as discussed below. Collection of blood from a woman is often carried out according to the standard protocol commonly followed by hospitals or clinics. An appropriate amount of venous blood, for example, typically 5 - 50 ml, can often be collected and stored according to standard procedures before further preparation. The blood sample can be collected, stored, or transported in a manner that minimizes degradation of the quality of the nucleic acids present in the sample.
[0033] Preparation of blood sample Analysis of fetal DNA found in maternal blood can be performed using, for example, whole blood, serum, or plasma. Methods for preparing serum or plasma from maternal blood are known. For example, the blood of a pregnant woman can be placed in a tube containing EDTA or a special commercial product such as Vacutainer SST (Becton Dickinson, Franklin Lakes, N.J.) to prevent blood clotting, and then plasma can be obtained from the whole blood by centrifugation. Serum can be obtained with or without centrifugation after blood clotting. When using centrifugation, it is typically carried out at an appropriate speed, for example, 1,500 - 3,000 g, but not necessarily so. The plasma or serum may be subjected to an additional centrifugation step before transferring to a new tube for DNA extraction.
[0034] In addition to the cell-free portion of whole blood, DNA can also be recovered from the cell fraction and concentrated in the buffy coat portion, which can be obtained by centrifuging a whole blood sample obtained from a woman and removing the plasma.
[0035] DNA extraction There are numerous known methods for extracting DNA from biological samples, including blood. One can follow the general methods of DNA preparation (e.g., as described by Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3d ed. 2001), and can also obtain DNA from blood samples obtained from pregnant women using various commercially available reagents or kits, such as Qiagen's QIAamp Circulating Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp DNA Blood Mini Kit (Qiagen, Hilden, Germany), GenomicPrep™ Blood DNA Isolation Kit (Promega, Madison, Wis.), and GFX™ Genomic Blood DNA Purification Kit (Amersham, Piscataway, N.J.). Additionally, combinations of more than one of these methods can be used.
[0036] In some embodiments, the sample can first be enriched, or enriched to some extent, for fetal nucleic acids by one or more methods. For example, the compositions and processes of the present technology can be used alone or in combination with other discriminative factors to discriminate between fetal DNA and maternal DNA. Examples of these factors include, but are not limited to, single nucleotide differences between the X and Y chromosomes, sequences specific to the Y chromosome, polymorphisms located elsewhere in the genome, size differences between fetal DNA and maternal DNA, and differences in methylation patterns between maternal and fetal tissues.
[0037] Other methods for concentrating a sample for a particular species of nucleic acid are described in PCT Patent Application No. PCT / US07 / 69991, filed May 30, 2007, PCT Patent Application No. PCT / US2007 / 071232, filed Jun. 15, 2007, U.S. Provisional Application Nos. 60 / 968,876 and 60 / 968,878 (assigned to the present applicant) (PCT Patent Application No. PCT / EP05 / 012707, filed Nov. 28, 2005), all of which are hereby incorporated by reference herein. In certain embodiments, the parental nucleic acid is selectively (partially, substantially, almost completely, or completely) removed from the sample.
[0038] The terms "nucleic acid" and "nucleic acid molecule" can be used interchangeably throughout the present disclosure. These terms refer to DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), RNA (e.g., messenger RNA (mRNA), small interfering RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed in fetus or placenta, etc.), and / or analogs of DNA or RNA (e.g., those containing analogs of bases, sugars and / or externally added backbones, etc.), RNA / DNA hybrids, polyamide nucleic acids (PNA), etc., all of which can be in single-stranded or double-stranded form and can include known analogs of natural nucleotides that function in a manner similar to naturally occurring nucleotides unless otherwise limited. In certain embodiments, the nucleic acid may be a plasmid, phage, autonomously replicating sequence (ARS), centromere, artificial chromosome, chromosome, or other nucleic acid that can replicate or be replicated in vitro or in a host cell, cell, cell nucleus or cytoplasm of a cell, or may be derived therefrom. The template nucleic acid may, in some embodiments, be derived from a single chromosome (e.g., a nucleic acid sample may be derived from one chromosome of a sample obtained from a diploid organism). Unless otherwise limited, this term includes nucleic acids containing known analogs of natural nucleotides that have binding properties similar to the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise stated, a particular nucleic acid sequence includes not only the explicitly shown sequence, but also its conservatively modified variants (e.g., degenerate codon substituents), alleles, orthologs, single nucleotide polymorphisms (SNPs) and complementary sequences implicitly. Specifically, degenerate codon substituents can be obtained by generating a sequence in which the third position of one or more selected (or all) codons is substituted with a residue of a wobble base and / or a deoxyinosine residue. The term nucleic acid is used interchangeably with locus, gene, cDNA, and mRNA encoded by the gene.This term can also include, as equivalents, derivatives, variants and analogs of RNA or DNA synthesized from nucleotide analogs, single-stranded (the "sense" strand or "antisense" strand, "plus" strand or "minus" strand, "forward" reading frame or "reverse" reading frame), and double-stranded polynucleotides. The term "gene" means a segment of DNA involved in the production of a polypeptide chain, which includes regions preceding and following the coding region (leader and trailer) involved in the transcription / translation of the gene product and the regulation of transcription / translation, as well as intervening sequences (introns) between individual coding segments (exons).
[0039] Deoxyribonucleotides include deoxyadenosine, deoxycytidine, deoxyguanosine, and deoxythymidine. In the case of RNA, the base cytosine is replaced by uracil. Nucleic acids can be prepared using nucleic acids obtained from a subject as a template.
[0040] Isolation and Manipulation of Nucleic Acids Nucleic acids can be obtained from one or more sources (e.g., cells, serum, plasma, buffy coat, lymph, skin, soil, etc.) by methods known in the art. Nucleic acids are often isolated from test samples. DNA can be isolated, extracted, and / or purified from biological samples (e.g., blood or blood products) using any suitable method, non-limiting examples of which include methods for DNA preparation (e.g., as described by Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3d ed. 2001), various commercially available reagents or kits, such as Qiagen's QIAamp Circulating Nucleic Acid Kit, QiaAmp DNA Mini Kit, or QiaAmp Examples include the DNA Blood Mini Kit (Qiagen, Hilden, Germany), GenomicPrep™ Blood DNA Isolation Kit (Promega, Madison, Wis.), and GFX™ Genomic Blood DNA Purification Kit (Amersham, Piscataway, N.J.), or combinations thereof.
[0041] Cell lysis procedures and reagents are known in the art and generally can be performed by chemical methods (e.g., detergents, hypotonic solutions, enzymatic procedures, etc., or combinations thereof), physical methods (e.g., French press, sonication, etc.), or lysis by electrolytes. Any suitable lysis procedure can be utilized. For example, chemical methods generally utilize a lysing agent to break open the cells, extract the nucleic acids from the cells, and subsequently treat with chaotropic salts. Physical methods, such as freeze / thaw followed by grinding; use of a cell press, etc. are also useful. Lysis procedures with high salt concentrations are also commonly used. For example, lysis procedures using alkali can be utilized. The latter procedures have conventionally incorporated the use of a phenol-chloroform solution, and alternative phenol-chloroform-free procedures involving three solutions can also be utilized. In the case of the latter procedures, one solution can contain 15 mM Tris, pH 8.0; 10 mM EDTA, and 100 μg / ml ribonuclease A; a second solution can contain 0.2 N NaOH and 1% SDS; and a third solution can contain 3 M KOAc, pH 5.5. These procedures can be found in Current Protocols in Molecular Biology, John Wiley & Sons, N.Y., 6.3.1 - 6.3.6 (1989), which is incorporated herein by reference in its entirety.
[0042] When comparing nucleic acids to another nucleic acid, they can be isolated at different times, and each of the samples can be from the same source or different sources. For example, the nucleic acids can be from a nucleic acid library, such as a cDNA library or an RNA library. The nucleic acids can be the result of nucleic acid purification or isolation, and / or amplification of nucleic acid molecules obtained from a sample. The nucleic acids provided in the processes described herein can be nucleic acids from one sample, or can contain nucleic acids from two or more samples (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more, nineteen or more, or twenty or more samples).
[0043] In certain embodiments, the nucleic acids can include extracellular nucleic acids. As used herein, the term "extracellular nucleic acid" can refer to nucleic acids isolated from a source that is substantially cell-free, and is also referred to as "cell-free" nucleic acid and / or "cell-free circulating" nucleic acid. Extracellular nucleic acids are present in blood (e.g., blood of a pregnant female) and can be obtained therefrom. Extracellular nucleic acids often do not contain detectable cells and may contain cellular elements or cell remnants. Non-limiting examples of cell-free sources for obtaining extracellular nucleic acids are blood, plasma, serum, and urine. As used herein, the term "obtaining cell-free circulating sample nucleic acids" includes directly obtaining a sample (e.g., collecting a sample, such as a test sample), or obtaining a sample from another person who has collected the sample. Without being limited by theory, extracellular nucleic acids can be products of cell apoptosis and cell breakdown, which underlie extracellular nucleic acids that often have a range of lengths across a spectrum (e.g., a "ladder").
[0044] In certain embodiments, extracellular nucleic acids can contain different nucleic acid species and are thus referred to herein as "heterogeneous." For example, serum or plasma obtained from a person with cancer may contain nucleic acids derived from cancerous cells and nucleic acids derived from non-cancerous cells. In another example, serum or plasma obtained from a pregnant female may contain maternal nucleic acids and fetal nucleic acids. In some cases, fetal nucleic acids are sometimes about 5% to about 50% of the total nucleic acids (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48 or 49% of all nucleic acids are fetal nucleic acids). In some embodiments, the majority of the length of fetal nucleic acids in the nucleic acids is about 500 base pairs or less, about 250 base pairs or less, about 200 base pairs or less, about 150 base pairs or less, about 100 base pairs or less, about 50 base pairs or less, or about 25 base pairs or less.
[0045] In certain embodiments, the methods described herein can be performed by providing a nucleic acid without processing a sample containing the nucleic acid. In some embodiments, a sample containing the nucleic acid is processed and then the nucleic acid is provided to perform the methods described herein. For example, the nucleic acid can be extracted, isolated, purified, partially purified, or amplified from the sample. As used herein, the term "isolated" refers to removing a nucleic acid from its original environment (e.g., its natural environment if it occurs naturally, or the host cell if it is expressed exogenously), such that the nucleic acid has been altered in that it has been removed from its original environment by human intervention (e.g., "by the hand of man"). The term "isolated nucleic acid" as used herein can refer to a nucleic acid that has been removed from a subject (e.g., a human subject). An isolated nucleic acid can be provided with fewer non-nucleic acid components (e.g., proteins, lipids) than are present in the source sample. A composition containing an isolated nucleic acid may contain less than about 50% to greater than 99% non-nucleic acid components. A composition containing an isolated nucleic acid may contain about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater than 99% non-nucleic acid components. As used herein, the term "purified" can refer to providing a nucleic acid that contains fewer non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than were present before the nucleic acid was subjected to a purification procedure. A composition containing a purified nucleic acid may contain about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater than 99% other non-nucleic acid components. As used herein, the term "purified" can refer to providing a nucleic acid that contains fewer nucleic acid species than are present in the sample source from which the nucleic acid is derived. A composition containing a purified nucleic acid may contain about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater than 99% other nucleic acid species. For example, fetal nucleic acids can be purified from a mixture containing maternal and fetal nucleic acids.In certain instances, nucleosomes containing small fragments of fetal nucleic acids can be purified from a mixture of larger nucleosome complexes containing larger fragments of maternal nucleic acids.
[0046] In some embodiments, the nucleic acid is fragmented or cleaved before, during, or after the methods described herein. The fragmented or cleaved nucleic acid can have a nominal, average, or mean length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to about 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs. The fragments can be generated by suitable methods known in the art, and the average, mean, or nominal length of the nucleic acid fragments can be controlled by selecting appropriate fragment generation procedures.
[0047] The nucleic acid fragments can contain overlapping nucleotide sequences, and such overlapping sequences can facilitate the construction of the nucleotide sequences of the corresponding nucleic acid, or segments thereof, that are not fragmented. For example, one fragment may have sub-sequences x and y, and another fragment may have sub-sequences y and z, where x, y, and z are nucleotide sequences that can be 5 nucleotides or longer. In certain embodiments, the overlapping sequence y can be utilized to facilitate the construction of the x-y-z nucleotide sequence in the nucleic acid derived from the sample. In certain embodiments, the nucleic acid may be fragmented partially (e.g., from an incomplete or aborted particular cleavage reaction), or may be fragmented completely.
[0048] In some embodiments, the nucleic acids are fragmented or cleaved by suitable methods, non-limiting examples of which include physical methods (e.g., shearing, e.g., sonication, French press, heating, UV irradiation, etc.), enzymatic treatments (e.g., enzymatic cleaving agents (e.g., suitable nucleases, suitable restriction enzymes, suitable methylation-sensitive restriction enzymes)), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, base hydrolysis, heating, etc., or combinations thereof), the treatments described in U.S. Patent Application Publication No. 20050112590, etc., or combinations thereof.
[0049] As used herein, "fragmentation" or "cleavage" refers to a procedure or condition that can break a nucleic acid molecule, such as a nucleic acid template gene molecule or an amplification product thereof, into two or more smaller nucleic acid molecules. Such fragmentation or cleavage can be sequence-specific, base-specific, or non-specific and can be achieved by any of a variety of methods, reagents, or conditions, including chemical, enzymatic, and physical fragmentation.
[0050] As used herein, the terms "fragment", "cleavage product", "cleaved product", or grammatical variants thereof refer to nucleic acid molecules obtained as a result of fragmentation or cleavage of a nucleic acid template gene molecule, or amplification products thereof. Such fragments or cleaved products may refer to all nucleic acid molecules obtained as a result of a cleavage reaction, but typically, such fragments or cleaved products refer only to nucleic acid molecules or amplification product segments thereof that contain the corresponding nucleotide sequence of the nucleic acid template gene molecule, resulting from fragmentation or cleavage of the nucleic acid template gene molecule. The term "amplification", as used herein, refers to subjecting a target nucleic acid in a sample to a process that linearly or exponentially generates an amplicon nucleic acid having the same or substantially the same nucleotide sequence as the target nucleic acid or a segment thereof. In certain embodiments, the term "amplification" refers to methods including polymerase chain reaction (PCR). For example, an amplification product can contain one or more additional nucleotides than the amplified nucleotide region of the nucleic acid template sequence (e.g., a primer can contain "extra" nucleotides, such as a transcription initiation sequence, in addition to nucleotides complementary to the nucleic acid template gene molecule, resulting in an amplification product that contains "extra" nucleotides or nucleotides that do not correspond to the amplified nucleotide region of the nucleic acid template gene molecule). Thus, a fragment can include a segment or portion of an amplified nucleic acid molecule that contains, at least in part, nucleotide sequence information obtained from or based on the indicated nucleic acid template molecule.
[0051] As used herein, the term "complementary cleavage reaction" refers to cleavage reactions performed on the same nucleic acid using different cleavage reagents or by varying the cleavage specificity of the same cleavage reagent, and thus generates alternative cleavage patterns of the same target or reference nucleic acid or protein. In certain embodiments, a nucleic acid can be treated in one or more reaction vessels with one or more specific cleavage agents (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more specific cleavage agents) (e.g., the nucleic acid can be treated with each specific cleavage agent in a separate vessel). The term "specific cleavage agent" as used herein refers to an agent, sometimes a chemical or an enzyme, that can cleave a nucleic acid at one or more specific sites.
[0052] Also, prior to providing a nucleic acid to the methods described herein, the nucleic acid can be exposed to a process that modifies certain nucleotides in the nucleic acid. For example, a process of selectively modifying a nucleic acid based on the methylation status of the nucleotides therein can be applied to the nucleic acid. Additionally, conditions such as high temperature, ultraviolet radiation, X-rays, etc. can cause changes in the sequence of a nucleic acid molecule. The nucleic acid can be provided in any suitable form useful for performing appropriate sequence analysis.
[0053] The nucleic acid can be single-stranded or double-stranded. For example, single-stranded DNA can be generated by denaturing double-stranded DNA, for example, by treatment with heat or alkali. In certain embodiments, the nucleic acid adopts a D-loop structure formed by invading an oligonucleotide into the strand of a double-stranded DNA molecule, or is a DNA-like molecule, such as a peptide nucleic acid (PNA). The formation of the D-loop can be promoted by adding the E. coli RecA protein and / or by changing the salt concentration, for example, using methods known in the art.
[0054] Determination of the content of fetal nucleic acid In some embodiments, the amount of fetal nucleic acid in a nucleic acid (e.g., concentration, relative amount, absolute amount, copy number, etc.) is determined. In certain embodiments, the amount of fetal nucleic acid in a sample is referred to as the "fetal fraction". In some embodiments, the "fetal fraction" refers to the fraction of fetal nucleic acid in cell-free circulating nucleic acid in a sample (e.g., a blood sample, a serum sample, a plasma sample) obtained from a pregnant female. In certain embodiments, the amount of fetal nucleic acid is determined according to a marker specific to male fetuses (e.g., Y-chromosome STR markers (e.g., DYS19, DYS385, DYS392 markers); RhD markers in RhD-negative females), according to the ratio of alleles of polymorphic sequences, or according to one or more markers that are specific to fetal nucleic acid and not to maternal nucleic acid (e.g., epigenetic biomarker differences between mother and fetus (e.g., methylation; described in more detail below), or fetal RNA markers in maternal plasma (see, e.g., Lo, 2005, Journal of Histochemistry and Cytochemistry, 53(3):293-296)).
[0055] The determination of the content of fetal nucleic acid (e.g., fetal fraction) is sometimes performed using a fetal quantifier assay (FQA), for example, according to the description in US Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference. This type of assay enables the detection and quantification of fetal nucleic acid in a maternal sample based on the methylation status of the nucleic acid in the sample. In certain embodiments, the amount of fetal nucleic acid derived from a maternal sample can be determined relative to the total amount of nucleic acid present, whereby the percentage of fetal nucleic acid in the sample is obtained. In certain embodiments, the copy number of fetal nucleic acid in a maternal sample can be determined. In certain embodiments, the amount of fetal nucleic acid can be determined in a sequence-specific (or partially specific) manner, sometimes with sufficient sensitivity to enable accurate chromosomal dosage analysis (e.g., detecting the presence or absence of fetal aneuploidy).
[0056] A fetal quantitative assay (FQA) can be performed in conjunction with any of the methods described herein. By any method known in the art and / or as described in U.S. Patent Application Publication No. 2010 / 0105049, for example, based on differences in methylation status, maternal DNA and fetal DNA can be distinguished, and fetal DNA can be quantified (e.g., its amount can be determined). Such assays can be performed by methods such as those that can distinguish nucleic acids based on methylation status, including but not limited to, capture using methylation sensitivity, e.g., using an MBD2-Fc fragment (the methyl-binding domain of MBD2 is fused to the Fc fragment of an antibody (MBD-FC)) (Gebhard et al. (2006) Cancer Res. 66(12):6118-28); methylation-specific antibodies; methods of conversion with bisulfite, e.g., MSP (methylation-sensitive PCR), COBRA, extension of primers with methylation-sensitive single nucleotides (Ms-SNuPE), or Sequenom MassCLEAVE™ technology; and the use of methylation-sensitive restriction enzymes (e.g., digesting maternal DNA in a maternal sample with one or more methylation-sensitive restriction enzymes, thereby enriching fetal DNA). Also, methyl-sensitive enzymes can be used to distinguish nucleic acids based on methylation status, and these enzymes can, for example, perform preferential or substantial cleavage or digestion at their DNA recognition sequences when the latter are not methylated. Thus, unmethylated DNA samples are cut into smaller fragments than methylated DNA samples, and highly methylated DNA samples are not cut. In the absence of a clear description, any method for differentiating nucleic acids based on methylation status can be used in combination with the techniques and methods herein. The amount of fetal DNA can be determined during an amplification reaction, for example, by introducing one or more competitor substances at known concentrations. The determination of the amount of fetal DNA can also be performed, for example, by RT-PCR, primer extension, sequencing, and / or counting. In certain cases, the amount of nucleic acid can be determined using BEAMing technology according to the description in U.S. Patent Application Publication No. 2007 / 0065823.In certain embodiments, the restriction efficiency can be determined and the ratio of efficiencies can be used to further determine the amount of fetal DNA.
[0057] In certain embodiments, a fetal quantification assay (FQA) can be used to determine the concentration of fetal DNA in a maternal sample, for example, by the following method: a) determining the total amount of DNA present in the maternal sample; b) selectively digesting the maternal DNA in the maternal sample using one or more methylation-sensitive restriction enzymes, thereby enriching the fetal DNA; c) determining the amount of fetal DNA obtained from step b); d) comparing the amount of fetal DNA obtained from step c) with the total amount of DNA obtained from step a), thereby determining the concentration of fetal DNA in the maternal sample. In certain embodiments, the absolute copy number of fetal nucleic acids in the maternal sample can be determined using, for example, mass spectrometry and / or a system that uses a competitive PCR approach to measure the absolute copy number. See, for example, Ding and Cantor (2003) PNAS, USA, Vol. 100: pp. 3059-3064, and U.S. Patent Application Publication No. 2004 / 0081993, both of which are incorporated herein by reference.
[0058] In certain embodiments, the fetal fraction can be determined based on the ratio of alleles of a polymorphic array (e.g., a single nucleotide polymorphism (SNP)), using, for example, the methods described in U.S. Patent Application Publication No. 2011 / 0224087, which is incorporated herein by reference. In such methods, nucleotide sequence reads are obtained for a maternal sample and compared at polymorphic sites (e.g., SNPs) that provide information in the reference genome to determine the total number of nucleotide sequence reads that map to a first allele and the total number of nucleotide sequence reads that map to a second allele, thereby determining the fetal fraction. In certain embodiments, for example, for a mixture of fetal and maternal nucleic acids in a sample, the maternal nucleic acids contribute significantly to such a mixture, and the fetal allele contribution is relatively small compared thereto, such that the fetal allele is identified. Thus, the relative abundance of fetal nucleic acids in the maternal sample can be determined as a parameter of the total number of unique sequence reads mapped to the target nucleic acid sequences on the reference genome for each of those two alleles of the polymorphic site.
[0059] In certain embodiments, the fetal fraction can be determined based on one or more levels. The determination results of the fetal fraction according to the levels are described, for example, in International Application Publication No. WO2014 / 055774, the entire content of which is incorporated herein by reference including all documents, tables, formulas, and drawings. In some embodiments, the fetal fraction is determined according to levels classified as those indicating variations in maternal and / or fetal copy number. For example, the determination of the fetal fraction can include an assessment of the expected level of the variation in maternal and / or fetal copy number utilized in the determination of the fetal fraction. In some embodiments, the fetal fraction is determined for a level (e.g., a first level) classified as indicating a variation in copy number according to a range of expected levels determined for the same type of copy number variation. The fetal fraction can be determined according to the observed level within the range of expected levels, whereby it is classified as a variation in maternal and / or fetal copy number. In some embodiments, the fetal fraction is determined when the observed level (e.g., a first level) classified as a variation in maternal and / or fetal copy number is different from the expected level determined for the same maternal and / or fetal copy number variation. The fetal fraction can be provided as a percentage. For example, the fetal fraction can be divided by 100, thereby obtaining a percentage value. For example, for a first level that indicates a homozygous duplication of the mother and is at a level of 155, and an expected level for the homozygous duplication of the mother, which is at a level of 150, the fetal fraction can be determined as 10% (e.g., (fetal fraction = 2×(155 - 150)).
[0060] In combination with the methods provided herein, the amount of fetal nucleic acid in extracellular nucleic acid can be quantified and used. Thus, in certain embodiments, the methods of the techniques described herein include an additional step of determining the amount of fetal nucleic acid. The amount of fetal nucleic acid in a nucleic acid sample obtained from a subject can be determined before or after the processing for preparing the sample nucleic acid. In certain embodiments, after the sample nucleic acid is processed and prepared, the amount of fetal nucleic acid in the sample is determined and this amount is utilized for further evaluation. In some embodiments, the outcome includes factoring the fraction of fetal nucleic acid in the sample nucleic acid (e.g., adjusting the count number, removing the sample, making a determination, or not making a determination). In certain embodiments, the methods provided herein can be used in combination with a method for determining the fetal fraction. For example, a method for determining the fetal fraction, including a normalization process, can include one or more normalization methods provided herein (e.g., principal component normalization).
[0061] The step of determination can be performed before, during, at any point within, or after a method described herein, or after a method for a particular determination (e.g., detection of aneuploidy, determination of fetal gender) described herein. For example, to perform a method for determining fetal gender or aneuploidy with a given sensitivity or specificity, a method for quantifying fetal nucleic acid is performed before, during, or after the determination of fetal gender or aneuploidy to identify samples having greater than about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25% or more fetal nucleic acid. In some embodiments, for example, samples determined to have a particular threshold amount of fetal nucleic acid (e.g., about 15% or more fetal nucleic acid; about 4% or more fetal nucleic acid) are further analyzed for the determination of fetal gender or aneuploidy, or for the presence or absence of aneuploidy or genetic mutations. In certain embodiments, the determination result of the presence or absence of fetal gender or aneuploidy is selected (e.g., selected and communicated to the patient) only if the sample has a particular threshold amount of fetal nucleic acid (e.g., about 15% or more fetal nucleic acid; about 4% or more fetal nucleic acid).
[0062] In some embodiments, determination of fetal fraction or determination of the amount of fetal nucleic acid is not required, nor necessary, to identify the presence or absence of chromosomal aneuploidy. In some embodiments, identification of the presence or absence of chromosomal aneuploidy does not require discrimination between the sequences of fetal DNA and maternal DNA. In certain embodiments, this is because the combined contributions of both maternal and fetal sequences in a particular chromosome, chromosomal portion, or segment thereof are analyzed. In some embodiments, identification of the presence or absence of chromosomal aneuploidy does not rely on prior sequence information that would distinguish fetal DNA from maternal DNA.
[0063] Concentration of Nucleic Acid In some embodiments, a nucleic acid (e.g., an extracellular nucleic acid) is concentrated or relatively concentrated to obtain a subpopulation or species of nucleic acids. The subpopulation of nucleic acids can include, for example, fetal nucleic acids, maternal nucleic acids, nucleic acids containing fragments of a particular length or range of lengths, or nucleic acids derived from a particular genomic region (e.g., a single chromosome, a series of chromosomes and / or a particular chromosomal region). Such concentrated samples can be used in conjunction with the methods provided herein. Thus, in certain embodiments, the methods of the technology include an additional step of concentrating a subpopulation of nucleic acids in a sample, such as fetal nucleic acids and the like. In certain embodiments, the methods described above for determining the fetal fraction can also be used to concentrate and obtain fetal nucleic acids. In certain embodiments, maternal nucleic acids are selectively (partially, substantially, almost completely or completely) removed from the sample. In certain embodiments, quantitative sensitivity can be improved by concentrating to obtain nucleic acids of a particular low copy number species (e.g., fetal nucleic acids). Methods for concentrating a sample for a particular species of nucleic acid are described, for example, in U.S. Patent No. 6,927,028, International Patent Application Publication No. WO2007 / 140417, International Patent Application Publication No. WO2007 / 147063, International Patent Application Publication No. WO2009 / 032779, International Patent Application Publication No. WO2009 / 032781, International Patent Application Publication No. WO2010 / 033639, International Patent Application Publication No. WO2011 / 034631, International Patent Application Publication No. WO2006 / 056480, and International Patent Application Publication No. WO2011 / 143659, the entire contents of each of which are incorporated herein by reference, including all documents, tables, formulas, and drawings.
[0064] In some embodiments, nucleic acids are concentrated to obtain certain target fragment species and / or reference fragment species. In certain embodiments, one or more length-based separation methods described below are used to concentrate nucleic acids to obtain a fragment length or range of fragment lengths of a particular nucleic acid. In certain embodiments, one or more sequence-based separation methods described herein and / or known in the art are used to concentrate nucleic acids to obtain fragments derived from a selected genomic region (e.g., a chromosome). Certain methods for concentrating a subpopulation of nucleic acids in a sample (e.g., fetal nucleic acids) are described in detail below.
[0065] Some methods for concentrating a subpopulation of nucleic acids (e.g., fetal nucleic acids) that can be used with the methods described herein include methods that exploit epigenetic differences between maternal and fetal nucleic acids. For example, based on differences in methylation, fetal nucleic acids can be differentiated from maternal nucleic acids and then separated. A method for concentrating fetal nucleic acids based on methylation is described in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference. Such methods sometimes include binding the sample nucleic acids to a methylation-specific binding agent (methyl-CpG binding protein (MBD), methylation-specific antibody, etc.) and separating the bound nucleic acids from the unbound nucleic acids based on differences in methylation status. Such methods can also include the use of methylation-sensitive restriction enzymes (described above; e.g., HhaI and HpaII), which can be used to selectively digest the maternal nucleic acids and enrich the sample for at least one region of fetal nucleic acids by selectively digesting the nucleic acids derived from the maternal sample with an enzyme that digests the maternal nucleic acids to enrich the region of fetal nucleic acids in the maternal sample.
[0066] Another method for enriching a subpopulation of nucleic acids (e.g., fetal nucleic acids) that can be used in conjunction with the methods described herein is an approach that enhances polymorphic sequences with a restriction endonuclease such as the method described in U.S. Patent Application Publication No. 2009 / 0317818, which is incorporated herein by reference. Such methods include cleaving a nucleic acid containing a non-target allele with a restriction endonuclease that recognizes a nucleic acid containing a non-target allele but not a target allele, and amplifying the uncleaved nucleic acid without amplifying the cleaved nucleic acid, wherein the uncleaved, amplified nucleic acid is a target nucleic acid (e.g., fetal nucleic acid) enriched relative to a non-target nucleic acid (e.g., maternal nucleic acid). In certain embodiments, for example, the nucleic acid can be selected to include an allele having a polymorphic site that is susceptible to selective digestion by a cleavage agent.
[0067] Some methods for enriching a subpopulation of nucleic acids (e.g., fetal nucleic acids) that can be used in conjunction with the methods described herein include an approach of selective enzymatic degradation. Such methods include protecting the target sequence from exonuclease digestion, thereby facilitating the elimination of unwanted sequences (e.g., maternal DNA) in the sample. For example, in one approach, the sample nucleic acid is denatured to produce single-stranded nucleic acids, the single-stranded nucleic acids are contacted with at least one pair of target-specific primers under appropriate annealing conditions and annealed, the annealed primers are extended by nucleotide polymerization to produce a double-stranded target sequence, and the single-stranded (e.g., non-target) nucleic acids are digested using a nuclease that digests single-stranded nucleic acids. In certain embodiments, this method can be repeated in at least one additional cycle. In certain embodiments, the same pair of target-specific primers is used to extend the primers in each of the first and second cycles, and in certain embodiments, different pairs of target-specific primers are used for the first and second cycles.
[0068] Some methods for enriching subpopulations of nucleic acids (e.g., fetal nucleic acids) that can be used with the methods described herein include the approach of massively parallel signature sequencing (MPSS). MPSS is typically a solid-phase method that uses ligation of adapters (e.g., tags), followed by decoding of the adapters to fragment and read nucleic acid sequences. Typically, tagged PCR products are amplified, resulting in the generation of PCR products with unique tags from each nucleic acid. Often, tags are used to attach the PCR products to microbeads. After performing ligation-based sequencing several times, for example, the signature of the sequence can be identified from each bead. Each signature sequence (MPSS tag) in the MPSS dataset is analyzed, compared to all other signatures, and all identical signatures are counted.
[0069] In certain embodiments, certain enrichment methods (e.g., certain enrichment methods based on MPS and / or MPSS) can include an approach based on amplification (e.g., PCR). In certain embodiments, locus-specific amplification methods can be used (e.g., using locus-specific amplification primers). In certain embodiments, a multiplex SNP allele PCR approach can be used. In certain embodiments, the multiplex SNP allele PCR approach can be used in combination with uniplex sequencing. For example, such an approach can include the use of multiplex PCR (e.g., MASSARRAY system), and incorporation of capture probe sequences into the amplicons, followed by sequencing using, for example, the Illumina MPSS system. In certain embodiments, the multiplex SNP allele PCR approach can be used in combination with a three-primer system and index sequencing. For example, such an approach can include using multiplex PCR (e.g., MASSARRAY system) with primers having a first capture probe incorporated into a locus-specific forward PCR primer and an adapter sequence incorporated into a locus-specific reverse PCR primer for sequencing using, for example, the Illumina MPSS system, thereby generating amplicons, followed by performing a second PCR for incorporating reverse capture sequences and molecular index barcodes. In certain embodiments, the multiplex SNP allele PCR approach can be used in combination with a four-primer system and index sequencing.For example, such an approach can include using multiplex PCR (e.g., the MASSARRAY system) that uses primers having adapter sequences incorporated into both a locus-specific forward PCR primer and a locus-specific reverse PCR primer for sequencing, for example, using the Illumina MPSS system, followed by performing a second PCR to incorporate both a forward capture sequence and a reverse capture sequence as well as a molecular index barcode. In certain embodiments, an approach using microfluidic technology can be used. In certain embodiments, an approach using array-based microfluidic technology can be used. For example, such an approach can include using an array by microfluidic technology (e.g., Fluidigm) to perform low-plex amplification as well as incorporation of indices and capture probes, followed by performing sequencing. In certain embodiments, an approach using emulsion microfluidic technology, such as digital droplet PCR, for example, can be used.
[0070] In certain embodiments, a universal amplification method can be used (e.g., using universal primers or amplification primers that are not locus-specific). In certain embodiments, the universal amplification method can be used in combination with a pull-down approach. In certain embodiments, the method can include a pull-down with biotinylated ultramers from a universally amplified sequencing library (e.g., a biotinylated pull-down assay from Agilent or IDT). For example, such an approach can include the preparation of a standard library, enrichment of selected regions by a pull-down assay, and a second universal amplification step. In certain embodiments, the pull-down approach can be used in combination with a ligation-based method. In certain embodiments, the method can include a pull-down with biotinylated ultramers using ligation of sequence-specific adapters (e.g., HALOPLEX PCR, Halo Genomics). For example, such an approach can include the use of selector probes to capture restriction enzyme digestion fragments, followed by ligation of the captured products to adapters, and universal amplification, followed by sequencing. In certain embodiments, the pull-down approach can be used in combination with methods based on extension and ligation. In certain embodiments, the method can include extension and ligation by molecular inversion probes (MIPs). For example, such an approach can include the use of molecular inversion probes in combination with sequence adapters, followed by universal amplification and sequencing. In certain embodiments, complementary DNA can be synthesized and sequenced without amplification.
[0071] In certain embodiments, the elongation and ligation approach can be performed without a pull-down component. In certain embodiments, the method can include hybridization, elongation, and ligation with locus-specific forward and reverse primers. Such methods can further include universal amplification, or complementary DNA synthesis without amplification, followed by sequencing. In certain embodiments, such methods can reduce or eliminate background sequences during analysis.
[0072] In certain embodiments, the pull-down approach can be used with optional amplification components or without amplification components. In certain embodiments, the method can include a modified pull-down assay and ligation, fully incorporating the capture probe and not performing universal amplification. For example, such an approach can include the use of a modified selector probe to capture restriction enzyme digestion fragments, followed by ligation of the captured product to an adapter, optional amplification, and sequencing. In certain embodiments, the method can include a biotinylated pull-down assay involving elongation and ligation of adapter sequences in combination with circular single-strand ligation. For example, such an approach can include the use of a selector probe for a capture region of interest (e.g., a target sequence), elongation of the probe, ligation of the adapter, single-strand circular ligation, optional amplification, and sequencing. In certain embodiments, the target sequence can be separated from the background by analysis of the sequencing results.
[0073] In some embodiments, one or more sequence-based separation methods described herein are used to enrich nucleic acids to obtain fragments derived from a selected genomic region (e.g., a chromosome). Sequence-based separation generally relies on the fact that the nucleotide sequence is present in the fragment of interest (e.g., the target and / or reference fragment) and substantially absent, or present in only trace amounts (e.g., 5% or less), in other fragments of the sample. In some embodiments, sequence-based separation can be used to separate target fragments and / or reference fragments. The separated target fragments and / or separated reference fragments are often isolated and removed from the remaining fragments in the nucleic acid sample. In certain embodiments, the separated target fragments and the separated reference fragments are also isolated and removed from each other (e.g., isolated as compartments of a separation assay). In certain embodiments, the separated target fragments and the separated reference fragments are isolated together (e.g., isolated as the same assay compartment). In some embodiments, unbound fragments can be differentially removed or degraded or digested.
[0074] In some embodiments, a process of selectively capturing nucleic acids is used to isolate and retrieve target fragments and / or reference fragments from a nucleic acid sample. Commercially available systems for capturing nucleic acids include, for example, the Nimblegen Sequence Capture System (Roche NimbleGen, Madison, WI); the Illumina BEADARRAY platform (Illumina, San Diego, CA); the Affymetrix GENECHIP platform (Affymetrix, Santa Clara, CA); the Agilent SureSelect Target Enrichment System (Agilent Technologies, Santa Clara, CA); and related platforms. Such methods typically involve hybridization of capture oligonucleotides to segments or all of the nucleotide sequences of the target or reference fragments, and can include the use of solid-phase (e.g., solid-phase arrays) and / or solution-based platforms. Capture oligonucleotides (sometimes referred to as "baits") are selected or designed to preferentially hybridize to nucleic acid fragments derived from selected genomic regions or loci (e.g., one of chromosomes 21, 18, 13, X or Y, or a reference chromosome). In certain embodiments, hybridization-based methods (e.g., using oligonucleotide arrays) are used to enrich for nucleic acid sequences derived from a particular chromosome (e.g., a chromosome that may be aneuploid, a reference chromosome, or other chromosome of interest), or segments thereof.
[0075] In some embodiments, one or more methods of length-based separation are used to enrich nucleic acids for lengths of a particular nucleic acid fragment, a range of lengths, or lengths below or above a particular threshold or cutoff. The length of a nucleic acid fragment typically refers to the number of nucleotides in the fragment. Also, the length of a nucleic acid fragment is sometimes also referred to as the size of the nucleic acid fragment. In some embodiments, the length-based separation method is performed without measuring the length of individual fragments. In some embodiments, the length-based separation method is performed in conjunction with a method for determining the length of individual fragments. In some embodiments, length-based separation refers to a size fractionation procedure, and all or a portion of the fractionated pool can be isolated (e.g., retained) and / or analyzed. Size fractionation procedures are known in the art (e.g., separation on an array, separation by a molecular sieve, separation by gel electrophoresis, separation by column chromatography (e.g., a molecular sieve column), and approaches based on microfluidic technology). In certain embodiments, examples of length-based separation approaches can include, for example, circularization of fragments, treatment with chemicals (e.g., formaldehyde, polyethylene glycol (PEG)), mass spectrometry, and / or size-specific nucleic acid amplification.
[0076] Separation methods based on certain lengths that can be used in conjunction with the methods described herein utilize, for example, a tagging approach with selective arrays. The term "tagging with an array" refers to incorporating recognizable and distinctively different arrays into a nucleic acid or population of nucleic acids. The term "tagging with an array" as used herein has a different meaning than the term "array tag" described later herein. In such tagging with an array methods, nucleic acids of a certain fragment size species (e.g., short fragments) are subjected to tagging with a selective array in a sample containing long and short nucleic acids. Such methods typically include the step of performing a nucleic acid amplification reaction using a set of nested primers including inner and outer primers. In certain embodiments, one or both of the inner primers are tagged such that the tag can be introduced onto the amplification product of the target. The outer primers generally do not anneal to short fragments carrying the (inner) target sequence. The inner primers can anneal to short fragments and generate amplification products carrying the tag and the target sequence. Typically, tagging of long fragments is inhibited through a combination of mechanisms including, for example, blocking of the extension of the inner primers by prior annealing and extension of the outer primers. Enrichment for the tagged fragments can be performed by any of a variety of methods including, for example, exonuclease digestion of single-stranded nucleic acids and amplification of the tagged fragments using amplification primers specific for at least one tag.
[0077] Another method of separation based on length that can be used in conjunction with the methods described herein involves subjecting a nucleic acid sample to polyethylene glycol (PEG) precipitation. Examples of methods are those described in International Patent Application Publications WO2007 / 140417 and WO2010 / 115016, the entire contents of each of which are incorporated herein by reference, including all documents, tables, formulas, and drawings. This method generally requires contacting a nucleic acid sample with PEG in the presence of one or more monovalent salts under conditions sufficient to substantially precipitate large nucleic acids without substantially precipitating small (e.g., less than 300 nucleotides) nucleic acids.
[0078] Another method of enrichment based on size that can be used in conjunction with the methods described herein involves ligation, e.g., circularization by ligation using circligase. Short nucleic acid fragments can typically be circularized with higher efficiency than long fragments. Sequences that did not circularize can be separated from the circularized sequences, and the enriched short fragments can be used for further analysis.
[0079] Nucleic acid library In some embodiments, a nucleic acid library is a plurality of polynucleotide molecules (e.g., a sample of nucleic acids) that are prepared, collected, and / or modified for a particular process (non-limiting examples of which include immobilization on a solid phase (e.g., a solid support, e.g., a flow cell, beads), enrichment, amplification, cloning, detection) and / or for nucleic acid sequencing. In certain embodiments, the nucleic acid library is prepared before or during the process of nucleic acid sequencing. The nucleic acid library (e.g., a sequencing library) can be prepared by any suitable method known in the art. The nucleic acid library can be prepared by a preparation process that is targeted or not targeted.
[0080] In some embodiments, a library of nucleic acids is modified to include chemical moieties (e.g., functional groups) configured for immobilization of the nucleic acids to a solid support. In some embodiments, a library of nucleic acids is modified to include biological molecules (e.g., functional groups) and / or members of binding pairs configured for immobilization of the library to a solid support, non-limiting examples of which include thyroxine-binding globulin, steroid-binding protein, antibody, antigen, hapten, enzyme, lectin, nucleic acid, repressor, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding protein, receptor, carbohydrate, oligonucleotide, polynucleotide, complementary nucleic acid sequences, etc., and combinations thereof. Some examples of specific binding pairs include, without limitation, an avidin moiety and a biotin moiety; an antigenic epitope and an antibody or an immunologically reactive fragment thereof; an antibody and a hapten; a digoxigen moiety and an anti-digoxigen ) antibody; a fluorescein moiety and an anti-fluorescein antibody; an operator and a repressor; a nuclease and a nucleotide; a lectin and a polysaccharide; a steroid and a steroid-binding protein; an active compound and a receptor for the active compound; a hormone and a hormone receptor; an enzyme and a substrate; an immunoglobulin and protein A; an oligonucleotide or polynucleotide and its corresponding complement, etc., or combinations thereof.
[0081] In some embodiments, a library of nucleic acids is modified to include one or more polynucleotides of known composition, non-limiting examples of which include identifiers (e.g., tags, index tags), capture sequences, labels, adapters, restriction enzyme sites, promoters, enhancers, origins of replication, stem loops, complementary sequences (e.g., primer binding sites, annealing sites), appropriate integration sites (e.g., transposons, viral integration sites), modified nucleotides, etc., or combinations thereof. Polynucleotides of known sequences can be added at appropriate positions, e.g., the 5' end, 3' end, or internally of a nucleic acid sequence. Polynucleotides of known sequences can be of the same sequence or different sequences. In some embodiments, polynucleotides of known sequences are configured to hybridize to one or more oligonucleotides immobilized on a surface (e.g., a surface in a flow cell). For example, a nucleic acid molecule containing a 5' known sequence can be hybridized to a first plurality of oligonucleotides, while the 3' known sequence of that molecule can be hybridized to a second plurality of oligonucleotides. In some embodiments, a library of nucleic acids can include chromosome-specific tags, capture sequences, labels, and / or adapters. In some embodiments, a library of nucleic acids includes one or more detectable labels. In some embodiments, one or more detectable labels can be incorporated into the nucleic acid library at the 5' end, at the 3' end, and / or at the position of any nucleotide internal to the nucleic acids in the library. In some embodiments, a library of nucleic acids includes hybridized oligonucleotides. In certain embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, a library of nucleic acids includes oligonucleotide probes hybridized prior to immobilization on a solid phase.
[0082] In some embodiments, the polynucleotide of a known sequence comprises a universal sequence. The universal sequence is a specific nucleotide sequence incorporated into two or more nucleic acid molecules, or two or more subsets of nucleic acid molecules, and the universal sequence is the same for all of the molecules or subsets of molecules into which it is incorporated. Universal sequences are often designed to hybridize to and / or amplify multiple different sequences using a single universal primer that is complementary to the universal sequence. In some embodiments, two (e.g., pairs) or more universal sequences and / or universal primers are used. Universal primers often contain the universal sequence. In some embodiments, an adapter (e.g., a universal adapter) contains the universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple species or subsets of nucleic acids.
[0083] In certain embodiments of the preparation of a nucleic acid library (e.g., in certain sequencing by synthesis procedures), the nucleic acid is selected and / or fragmented by size to obtain a length of several hundred base pairs or less (e.g., in the case of preparation for library generation). In some embodiments, the library preparation is performed without fragmentation (e.g., when using ccfDNA).
[0084] In certain embodiments, methods for preparing libraries based on ligation are used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego CA). Methods for preparing libraries based on ligation often utilize the design of adapters (e.g., methylated adapters), which can incorporate index sequences in the first ligation step and can often be used to prepare samples for single-end read sequencing, paired-end sequencing, and multiplex sequencing. For example, end repair of nucleic acids (e.g., fragmented nucleic acids or ccfDNA) is sometimes performed by a fill-in reaction, an exonuclease reaction, or a combination thereof. In some embodiments, the resulting blunt-ended repaired nucleic acids can then be extended with a single nucleotide that is complementary to the single nucleotide overhang on the 3' end of the adapter / primer. Any nucleotide can be used for the extension / overhang nucleotide. In some embodiments, the preparation of the nucleic acid library includes the ligation of adapter oligonucleotides. Adapter oligonucleotides often exhibit complementarity to a flow cell anchor and are sometimes utilized, for example, to immobilize a nucleic acid library on a solid support, such as the inner surface of a flow cell. In some embodiments, the adapter oligonucleotide includes an identifier, one or more sequencing primer hybridization sites (e.g., a sequence complementary to a universal sequencing primer, a single-end sequencing primer, a paired-end sequencing primer, a multiplex sequencing primer, etc.), or a combination thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing).
[0085] An identifier is a suitable detectable label that is incorporated into or attached to a nucleic acid (e.g., a polynucleotide), and the identifier enables the detection and / or identification of the nucleic acid containing it. In some embodiments, the identifier is incorporated into or attached to the nucleic acid (e.g., by a polymerase) during a sequencing method. Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indices or barcodes, radiolabels (e.g., isotopes), metal labels, fluorescent labels, chemiluminescent labels, phosphorescent labels, fluorophore quenchers, dyes, proteins (e.g., enzymes, antibodies or portions thereof, linkers, members of binding pairs), etc., or combinations thereof. In some embodiments, the identifier (e.g., a nucleic acid index or barcode) is a nucleotide or nucleotide analog of a unique, known, and / or identifiable sequence. In some embodiments, the identifier is six or more contiguous nucleotides. A number of fluorophores with diverse different excitation and emission spectra are available. Any suitable type and / or number of fluorophores can be used as an identifier. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more, or fifty or more different identifiers are utilized in the methods described herein (e.g., nucleic acid detection and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are ligated to each nucleic acid in a library.The detection and / or quantification of identifiers can be performed by suitable methods, machines or devices, and non-limiting examples thereof include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, luminometers, fluorometers, spectrophotometers, analysis by suitable gene chips or microarrays, Western blot, mass spectrometry, chromatography, analysis by cell fluorometry, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning cell counting, affinity chromatography, separation by manual batch mode, field suspension, suitable nucleic acid sequencing methods and / or nucleic acid sequencing devices, etc., as well as combinations thereof.
[0086] In some embodiments, a method for preparing a transposon-based library is used (e.g., EPICENTRE NEXTERA, Epicentre, Madison WI). Transposon-based methods typically use in vitro transposition to simultaneously fragment and tag DNA in a reaction in a single tube (often allowing incorporation of platform-specific tags and optional barcodes) to prepare a library that can be used with a sequencing device.
[0087] In some embodiments, a nucleic acid library or a portion thereof is amplified (e.g., by a PCR-based method). In some embodiments, the sequencing method includes amplification of the nucleic acid library. The nucleic acid library can be amplified before or after immobilization onto a solid support (e.g., a solid support in a flow cell). Nucleic acid amplification includes a process of amplifying or increasing the number of nucleic acid templates and / or their complements (e.g., present in the nucleic acid library) by generating one or more copies of the template and / or its complement. Amplification can be performed by an appropriate method. The nucleic acid library can be amplified by a thermocycling method or an isothermal amplification method. In some embodiments, a rolling circle amplification method is used. In some embodiments, amplification occurs on a solid support (e.g., inside a flow cell) to which the nucleic acid library or a portion thereof is immobilized. In certain sequencing methods, the nucleic acid library is added to a flow cell and immobilized by hybridization to an anchor under appropriate conditions. This type of nucleic acid amplification is often referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or part of the amplification products are synthesized by extension starting from an immobilized primer. The solid-phase amplification reaction is similar to standard solution-phase amplification, except that at least one of the amplification oligonucleotides (e.g., primers) is immobilized on a solid support.
[0088] In some embodiments, solid-phase amplification involves a nucleic acid amplification reaction that includes only one species of oligonucleotide primer immobilized on a surface. In certain embodiments, solid-phase amplification includes multiple different immobilized oligonucleotide primer species. In some embodiments, solid-phase amplification can include a nucleic acid amplification reaction that includes one species of oligonucleotide primer immobilized on a solid surface and a second different oligonucleotide primer species in solution. Multiple different species of immobilized or solution-based primers can be used. Non-limiting examples of solid-phase nucleic acid amplification reactions include surface amplification, bridge amplification, emulsion PCR, WildFire amplification (e.g., U.S. Patent Publication No. US20130012399), etc., or combinations thereof.
[0089] Sequencing In some embodiments, nucleic acids (e.g., nucleic acid fragments, sample nucleic acids, cell-free nucleic acids) are sequenced. In certain embodiments, complete or substantially complete sequences are obtained, and sometimes partial sequences are obtained.
[0090] In some embodiments, some or all of the nucleic acids in a sample are concentrated and / or amplified before or during sequencing (e.g., non-specifically, e.g., by a PCR-based method). In certain embodiments, specific portions or subsets of nucleic acids in a sample are concentrated and / or amplified before or during sequencing. In some embodiments, sequencing of a randomly selected portion or subset of a pre-selected pool of nucleic acids is performed. In some embodiments, no concentration and / or amplification of nucleic acids in the sample is performed before or during sequencing.
[0091] As used herein, "reads" (e.g., "a read", "a sequence read") are short nucleotide sequences generated by any sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment ("reads from a single end"), and sometimes from both ends of the nucleic acid (e.g., paired-end reads, reads from two ends).
[0092] The length of sequence reads is often associated with a particular sequencing technique. For example, high-throughput methods provide sequence reads with sizes in base pairs (bp) that can vary from dozens to hundreds. For example, nanopore sequencing can provide sequence reads with sizes in base pairs that can vary from dozens to hundreds or thousands. In some embodiments, the average, median, mean length, or absolute length of sequence reads is from about 15 bp to about 900 bp in length. In certain embodiments, the average, median, mean length, or absolute length of sequence reads is about 1000 bp or more.
[0093] In some embodiments, the nominal, average, mean length, or absolute length of reads from a single end is sometimes from about 15 contiguous nucleotides to about 50 or more contiguous nucleotides, from about 15 contiguous nucleotides to about 40 or more contiguous nucleotides, and sometimes about 15 contiguous nucleotides, or about 36 or more contiguous nucleotides. In certain embodiments, the nominal, average, mean length, or absolute length of reads from a single end is from about 20 to about 30 bases in length, or from about 24 to about 28 bases in length. In certain embodiments, the nominal, average, mean length, or absolute length of reads from a single end is about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, or about 29 bases in length or more.
[0094] In certain embodiments, the nominal, average, mean length or absolute length of the reads from both ends can sometimes be from about 10 adjacent nucleotides to about 25 adjacent nucleotides or more (e.g., about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24 or about 25 nucleotides long or more), from about 15 adjacent nucleotides to about 20 adjacent nucleotides or more, and sometimes about 17 adjacent nucleotides, or about 18 adjacent nucleotides.
[0095] Reads are generally representations of nucleotide sequences as physical nucleic acids. For example, in a read containing a sequence depicted as ATGC, as a physical nucleic acid, "A" represents an adenine nucleotide, "T" represents a thymine nucleotide, "G" represents a guanine nucleotide, and "C" represents a cytosine nucleotide. Reads of sequences obtained from the blood of a pregnant female can be reads derived from a mixture of fetal and maternal nucleic acids. A mixture of relatively short reads can be converted by the processes described herein into a representation of the genomic nucleic acids present in the pregnant female and / or the fetus. A mixture of relatively short reads can be converted, for example, into a copy number variation (e.g., a maternal and / or fetal copy number variation), a gene mutation, or an aneuploidy representation. Reads of a mixture of maternal and fetal nucleic acids can be converted into a representation of a composite chromosome or a segment thereof that includes features of one or both of the maternal and fetal chromosomes. In certain embodiments, "obtaining" a read of the nucleic acid sequence of a sample obtained from a subject and / or "obtaining" a read of the nucleic acid sequence of a biological specimen obtained from one or more reference individuals can include directly performing nucleic acid sequencing to obtain sequence information. In some embodiments, "obtaining" can include receiving sequence information directly obtained from nucleic acids by another.
[0096] In some embodiments, the percentage of the genome that is represented is sequenced and is sometimes referred to as "coverage" or "coverage depth". For example, 1X coverage indicates that approximately 100% of the nucleotide sequence of the genome is represented by reads. In some embodiments, "coverage depth" is a term used to compare relative to a previous sequencing run as a reference. For example, a second sequencing run may have 1 / 2 the coverage of a first sequencing run. In some embodiments, redundancy is introduced such that a given region of the genome can be covered by two or more reads, or overlapping reads (e.g., a coverage depth greater than 1, e.g., 2X coverage).
[0097] In some embodiments, nucleic acids from a single nucleic acid sample obtained from a single individual are sequenced. In certain embodiments, nucleic acids from each of two or more samples are sequenced, where the samples are obtained from a single individual or from different individuals. In certain embodiments, nucleic acid samples obtained from two or more biological samples are pooled, where each biological sample is obtained from a single individual or from two or more individuals, and the pooled sample is sequenced. In the latter embodiments, nucleic acid samples obtained from each biological sample are often identified by one or more unique identifiers.
[0098] In some embodiments, the sequencing method utilizes identifiers that enable multiplexing of the sequencing reactions in the sequencing process. The greater the number of unique identifiers, the greater the number of samples and / or chromosomes that can be multiplexed, for example, in the sequencing process. Any suitable number (e.g., 4, 8, 12, 24, 48, 96 or more) of unique identifiers can be used to perform the sequencing process.
[0099] Array determination processes sometimes use a solid phase, which sometimes includes a flow cell onto which nucleic acids from a library can be attached and reagents can be flowed and brought into contact with the attached nucleic acids. A flow cell sometimes includes lanes of the flow cell, and the use of identifiers can facilitate the analysis of several samples in each lane. A flow cell is often a solid support configured to hold the bound analyte and / or allow a reagent solution to pass orderly over the bound analyte. A flow cell often has a planar shape, is optically transparent, is generally on the millimeter or sub-millimeter scale, and often has channels or lanes in which interactions between the analyte and the reagent occur. In some embodiments, the number of samples analyzed in a given lane of a flow cell depends on the number of unique identifiers utilized during library preparation and / or probe design. A single lane of a flow cell. For example, multiplexing using 12 identifiers enables the simultaneous analysis of 96 samples (e.g., equal to the number of wells in a 96-well microtiter plate) in an 8-lane flow cell. Similarly, for example, multiplexing using 48 identifiers enables the simultaneous analysis of 384 samples (e.g., equal to the number of wells in a 384-well microtiter plate) in an 8-lane flow cell. Non-limiting examples of commercially available multiplex sequencing kits include Illumina's Multiplexing Sample Preparation Oligonucleotide Kit, as well as the Multiplexing Sequencing Primer and PhiX Control Kit (e.g., Illumina catalog numbers PE-400~1001 and PE-400~1002, respectively).
[0100] Any suitable method for nucleic acid sequencing can be used, and non-limiting examples thereof include Maxim & Gilbert, chain termination methods, sequencing by synthesis, sequencing by ligation, sequencing by mass spectrometry, microscopy-based techniques, etc., or combinations thereof. In some embodiments, in the methods provided herein, first-generation techniques, such as Sanger sequencing methods (including automated Sanger sequencing methods, including microfluidic Sanger sequencing) can be used. In some embodiments, sequencing techniques can be used that include the use of nucleic acid imaging techniques (e.g., transmission electron microscopy (TEM) and atomic force microscopy (AFM)). In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods generally involve clonally amplifying a DNA template or a single DNA molecule and sequencing these templates or molecules in parallel on a large scale, sometimes inside a flow cell. Next-generation (e.g., second-generation and third-generation) sequencing techniques capable of sequencing DNA in parallel on a large scale can be used for the methods described herein, and these are collectively referred to herein as "massively parallel sequencing" (MPS). In some embodiments, MPS sequencing utilizes a targeted approach, in which case a specific chromosome, gene, or region of interest is sequenced. In certain embodiments, a non-targeted approach is used, in which case, randomly, most or all of the nucleic acids in a sample are sequenced, amplified, and / or captured.
[0101] In some embodiments, targeted approaches for enrichment, amplification, and / or sequencing are used. Targeted approaches often isolate, select, and / or enrich subsets of nucleic acids in a sample for further processing using sequence-specific oligonucleotides. In some embodiments, a library of sequence-specific oligonucleotides is utilized to target (e.g., hybridize to) one or more sets of nucleic acids in a sample. Often, the sequence-specific oligonucleotides and / or primers are selected to have specific sequences (e.g., unique nucleic acid sequences) present in one or more of the chromosomes, genes, exons, introns, and / or regulatory regions of interest. Any suitable method or combination of methods can be used to enrich, amplify, and / or sequence one or more subsets of the targeted nucleic acids. In some embodiments, the targeted sequences are isolated and / or enriched by capturing them on a solid phase (e.g., a flow cell, beads) using one or more sequence-specific anchors. In some embodiments, polymerase-based methods (e.g., methods based on any suitable extension by polymerase such as PCR-based methods) using sequence-specific primers and / or primer sets are used to enrich and / or amplify the targeted sequences. Sequence-specific anchors can often be used as sequence-specific primers.
[0102] MPS sequencing sometimes uses sequencing by synthesis and certain visualization processes. Nucleic acid sequencing techniques that can be used in the methods described herein include sequencing by synthesis and reversible chain-termination nucleotide-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ2000; HISEQ2500 (Illumina, San Diego It is (CA). Using this technique, sequencing can be performed in parallel on millions of nucleic acid (e.g., DNA) fragments. In one example of this type of sequencing technology, a flow cell containing an optically transparent slide with eight individual lanes is used, and oligonucleotide anchors (e.g., adapter primers) are bound on their surfaces. A flow cell is often a solid support configured to hold the bound analyte and / or allow a reagent solution to pass orderly over the bound analyte. Flow cells often have a planar shape, are optically transparent, generally on the millimeter or sub-millimeter scale, and often have channels or lanes in which the interaction between the analyte and the reagent occurs.
[0103] In some embodiments, synthesis-based sequencing involves guiding the synthesis to add nucleotides iteratively (e.g., by covalent addition) to a primer or an existing nucleic acid strand. Each time a nucleotide is added iteratively, detection is performed, and this process is repeated multiple times until the sequence of the nucleic acid strand is obtained. The length of the resulting sequence depends, inter alia, on the number of addition and detection steps performed. In some embodiments of synthesis-based sequencing, one, two, three or more nucleotides of the same type (e.g., A, G, C or T) are added and detected in a single nucleotide addition. Nucleotides can be added by any suitable (e.g., enzymatic or chemical) method. For example, in some embodiments, a polymerase or ligase is guided to add nucleotides to a primer or an existing nucleic acid strand. In some embodiments of synthesis-based sequencing, different types of nucleotides, nucleotide analogs and / or identifiers are used. In some embodiments, reversible chain-terminating nucleotides and / or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In certain embodiments, synthesis-based sequencing involves cleavage (e.g., cleavage and removal of an identifier) and / or washing steps. In some embodiments, the addition of one or more nucleotides is detected by any suitable method described herein or known in the art, non-limiting examples of which include any suitable imaging device, suitable camera, digital camera, CCD (charge-coupled device)-based imaging device (e.g., a CCD camera), CMOS (complementary metal oxide semiconductor)-based imaging device (e.g., a CMOS camera), photodiode (e.g., a photomultiplier tube), electron microscopy, field effect transistor (e.g., a DNA field effect transistor), ISFET ion sensor (e.g., a CHEMFET sensor), etc., or combinations thereof. Other sequencing methods that can be used to practice the methods herein include digital PCR and hybridization-based sequencing.
[0104] Other sequencing methods that can be used to practice the methods of this specification include digital PCR and sequencing by hybridization. Digital polymerase chain reaction (digital PCR or dPCR) can be used to directly identify and quantify nucleic acids in a sample. In some embodiments, digital PCR can be performed in an emulsion. For example, individual nucleic acids can be separated, e.g., in a microfluidic chamber device, and each nucleic acid can be individually amplified by PCR. The nucleic acids can be separated such that only one nucleic acid per well is present. In some embodiments, different probes can be used to distinguish various alleles (e.g., fetal alleles and maternal alleles). The alleles can be counted to determine the copy number.
[0105] In certain embodiments, sequencing by hybridization can be used. This method includes contacting a plurality of polynucleotide sequences with a plurality of polynucleotide probes, and each of the plurality of polynucleotide probes can optionally be ligated to a substrate. In some embodiments, the substrate can be a flat surface having an array of known nucleotide sequences. The hybridization pattern to the array can be used to determine the polynucleotide sequences present in the sample. In some embodiments, each probe is ligated to a bead, e.g., an electromagnetic bead, etc. The hybridization to the beads can be identified and used to identify the plurality of polynucleotide sequences within the sample.
[0106] In some embodiments, nanopore sequencing can be used in the methods described herein. Nanopore sequencing is a single molecule sequencing technology by which the sequence is directly determined each time a single nucleic acid molecule (e.g., DNA) passes through a nanopore.
[0107] Using an MPS method, system, or technology platform suitable for the implementation methods described in this specification, reads of nucleic acid sequences can be obtained. Non-limiting examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, Helicos True Single Molecule Sequencing, Ion Torrent and ion semiconductor-based sequencing (e.g., developed by Life Technologies), technologies based on WildFire, 5500, 5500xl W and / or 5500xl W Genetic Analyzer (e.g., developed and sold by Life Technologies, U.S. Patent Publication No. US20130012399); polony sequencing, pyrosequencing, massively parallel signature sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen's systems and methods, nanopore-based platforms, chemoresistive field effect transistor (CHEMFET) arrays, electron microscopy-based sequencing (e.g., developed by ZS Genetics, Halcyon Molecular), nanoparticle sequencing, etc., or combinations thereof.
[0108] In some embodiments, sequence determination specific to a chromosome is performed. In some embodiments, DANSR (Digital Analysis of Selected Regions) is utilized to perform sequence determination specific to a chromosome. By performing digital analysis of selected regions by cfDNA-dependent concatenation of two locus-specific oligonucleotides via intervening "bridge" oligonucleotides to form a PCR template, it becomes possible to simultaneously quantify hundreds of loci. In some embodiments, sequence determination specific to a chromosome is performed by generating a library enriched with sequences specific to the chromosome. In some embodiments, sequence reads are obtained for only a selected set of chromosomes. In some embodiments, sequence reads are obtained for only chromosomes 21, 18, and 13. In some embodiments, sequence reads are obtained for and / or mapped against all or segments of a reference genome.
[0109] In some embodiments, sequence reads are generated, obtained, collected, integrated, manipulated, transformed, processed, and / or provided by a sequence module. A machine including the sequence module can be a suitable machine and / or device for determining the sequence of nucleic acids by leveraging sequencing techniques known in the art. In some embodiments, the sequencing module can align, integrate, fragment, complement, reverse complement, and / or perform error checking (correcting errors in sequence reads).
[0110] In some embodiments, the reads of the nucleotide sequences obtained from a sample are reads of partial nucleotide sequences. As used herein, a "read of a partial nucleotide sequence" refers to a read of an array of any length with incomplete sequence information, also referred to as sequence ambiguity. A read of a partial nucleotide sequence may lack information regarding the identity of the nucleobases and / or the position or order of the nucleobases. A read of a partial nucleotide sequence generally does not include reads of sequences that are due to accidental or unintended sequencing errors where only incomplete sequence information (or less than all of the bases are sequenced or determined) is present. Such sequencing errors may be specific to a particular sequencing process and may include, for example, inaccurate determination of the identity of the nucleobases and of missing or extra nucleobases. Thus, for the reads of partial nucleotide sequences herein, certain information about the sequence is often carefully excluded. That is, sequence information regarding less than all of the nucleobases, or that can otherwise be characterized or is a sequencing error, is carefully obtained. In some embodiments, the reads of partial nucleotide sequences can extend over a portion of the nucleic acid fragment. In some embodiments, the reads of partial nucleotide sequences can extend over the entire length of the nucleic acid fragment. Reads of partial nucleotide sequences are described, for example, in International Patent Application Publication No. WO2013 / 052907, the entire contents of which are incorporated herein by reference in their entirety, including all documents, tables, formulas, and drawings.
[0111] Mapping of Reads The reads of the sequences can be mapped. Any suitable mapping method (e.g., process, algorithm, program, software, module, etc., or combinations thereof) can be used, and certain aspects of the mapping process are described below.
[0112] Mapping of nucleotide sequence reads (e.g., sequence information obtained from a fragment whose physical location in the genome is unknown) can be performed in several ways, which often involves alignment of the obtained sequence reads with the matching sequences in the reference genome. In such an alignment, the sequence reads are generally aligned against the reference sequence, and the aligned reads are referred to as "mapped", "mapped sequence reads" or "mapped reads".
[0113] As used herein, the terms "aligned", "alignment" or "aligning" refer to two or more nucleic acid sequences that can be identified as a match (e.g., 100% identical) or a partial match. Alignment can be performed manually or by a computer (e.g., software, program, module or algorithm), and non-limiting examples thereof include the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysis pipeline. Alignment of sequence reads can be 100% sequence identity. In some cases, the alignment is lower than 100% sequence identity (e.g., imperfect match, partial match, partial alignment). In some embodiments, the alignment is about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76% or 75% identity. In some embodiments, the alignment includes mismatches. In some embodiments, the alignment includes 1, 2, 3, 4 or 5 mismatches. Two or more sequences can be aligned using either strand. In certain embodiments, a nucleic acid sequence is aligned with the reverse complement of another nucleic acid sequence.
[0114] Using various computational methods, the reads of the array can be mapped and / or aligned against a reference genome. Non-limiting examples of computer algorithms that can be used to align the array include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE1, BOWTIE2, ELAND, MAQ, PROBEMATCH, SOAP or SEQMAP, or modifications or combinations thereof. In some embodiments, the reads of the array can be aligned with a reference sequence and / or a sequence in a reference genome. In some embodiments, the reads of the array can be found in and / or aligned with sequences in nucleic acid databases known in the art, including, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (DNA Databank of Japan). The identified sequences can be searched against a sequence database using BLAST or a similar tool.
[0115] In some embodiments, the reads of the mapped array and / or information associated with the reads of the mapped array are stored on and / or accessed from a non-transitory computer-readable storage medium in a suitable computer-readable format. As used herein, the term "computer-readable format" is sometimes referred to broadly as a format. In some embodiments, the reads of the mapped array are stored on and / or accessed in a suitable binary format, text format, etc., or a combination thereof. The binary format is sometimes the BAM format. The text format is sometimes the Sequence Alignment / Map (SAM) format. Non-limiting examples of binary formats and / or text formats include BAM, SAM, SRF, FASTQ, Gzip, etc., or combinations thereof. In some embodiments, the reads of the mapped array are stored on and / or converted to a format that requires less storage space (e.g., fewer bytes) than conventional formats (e.g., SAM format or BAM format). In some embodiments, the reads of the mapped array in a first format are compressed into a second format that requires less storage space than the first format. As used herein, the term "compressed" refers to the process of data compression, source coding, and / or bitrate reduction that reduces the size of a computer-readable data file. In some embodiments, the reads of the mapped array are compressed from the SAM format of the binary format. When a file is compressed, some data is sometimes lost. Sometimes, no data is lost in the compression process. In some embodiments of file compression, some data is replaced with an index and / or reference to another data file containing information about the reads of the mapped array.In some embodiments, the reads of the mapped array are stored in a binary format that includes or consists of the read count, an identifier of the chromosome (e.g., identifying the chromosome to which the read is mapped), and an identifier of the chromosomal position (e.g., identifying the position on the chromosome to which the read is mapped). In some embodiments, the binary format includes a 20-byte array, a 16-byte array, an 8-byte array, a 4-byte array, or a 2-byte array. In some embodiments, the mapped read information is stored in an array in a 10-byte format, a 9-byte format, an 8-byte format, a 7-byte format, a 6-byte format, a 5-byte format, a 4-byte format, a 3-byte format, or a 2-byte format. Sometimes, the mapped read data is stored in a 4-byte array that includes a 5-byte format. In some embodiments, the binary format includes a 5-byte format that includes a 1-byte ordinal number of the chromosome and a 4-byte position of the chromosome. In some embodiments, the mapped reads are stored in a compressed binary format that is about 1 / 100, about 1 / 90, about 1 / 80, about 1 / 70, about 1 / 60, about 1 / 55, about 1 / 50, about 1 / 45, about 1 / 40, or about 1 / 30 of the Sequence Alignment / Map (SAM) format. In some embodiments, the mapped reads are stored in a compressed binary format that is about 1 / 2 to about 1 / 50 (e.g., about 1 / 30, 1 / 25, 1 / 20, 1 / 19, 1 / 18, 1 / 17, 1 / 16, 1 / 15, 1 / 14, 1 / 13, 1 / 12, 1 / 11, 1 / 10, 1 / 9, 1 / 8, 1 / 7, 1 / 6, or about 1 / 5) of the GZip format.
[0116] In some embodiments, the system includes a compression module (e.g., 4 in FIG. 10A). In some embodiments, the read information of the mapped array stored on a non-transitory computer-readable storage medium in a computer-readable format is compressed by the compression module. The compression module sometimes converts the read of the mapped array to or from an appropriate format. In some embodiments, the compression module receives reads of the mapped array in a first format (e.g., 1), converts them to a compressed format (e.g., binary format, 5), and can transfer the compressed reads to another module (e.g., the bias density module, 6). The compression module often provides the read of the array in binary format, 5 (e.g., BReads format). Non-limiting examples of the compression module include GZIP, BGZF, and BAM, etc., or modified forms thereof. The following shows an example of conversion to a 4-byte array of integers using java.
Number
[0117] In some embodiments, reads can be mapped uniquely or non-uniquely to a reference genome. In the case of alignment with a single sequence in the reference genome, the read is considered to be "uniquely mapped". In the case of alignment with two or more sequences in the reference genome, the read is considered to be "non-uniquely mapped". In some embodiments, non-uniquely mapped reads are excluded from further analysis (e.g., quantification). In certain embodiments, a particular, low degree of mismatch (0 to 1) may be explained as a single nucleotide polymorphism that may exist between the reference genome and the reads obtained from the individual samples being mapped. In some embodiments, no degree of mismatch is allowed for reads mapped to the reference sequence.
[0118] As used herein, the term "reference genome" can refer to any particular known sequenced or characterized genome of any organism or virus, whether a partial sequence or a complete sequence, that can be used to reference an identified sequence from a subject. A reference genome can sometimes refer to a segment of a reference genome (e.g., a chromosome or a part thereof, e.g., one or more portions of a reference genome). The human genome, a human genome assembly, and / or a genome from any other organism can be used as a reference genome. One or more human genomes, human genome assemblies, and genomes of other organisms can be found at the National Center for Biotechnology Information at www.ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus, represented as a nucleic acid sequence. As used herein, a reference sequence or reference genome is often a genomic sequence obtained from, collected from, or partially collected from one or more individuals. In some embodiments, the reference genome is a genomic sequence obtained from, collected from, or partially collected from one or more human individuals. In some embodiments, the reference genome includes sequences assigned to chromosomes. The term "reference sequence" as used herein refers to one or more polynucleotide sequences of one or more reference samples. In some embodiments, the reference sequence includes reads of sequences obtained from a reference sample. In some embodiments, the reference sequence includes reads of sequences obtained from one or more reference samples, assemblies of the reads, consensus DNA sequences (e.g., sequence contigs), read density, and / or read density profiles. A read density profile obtained from a reference sample is sometimes referred to herein as a reference profile. A read density profile obtained from a test sample and / or a test subject is sometimes referred to herein as a test profile. In some embodiments, the reference sample is obtained from a reference subject that is substantially free of genetic mutations (e.g., mutations in the gene in question). In some embodiments, the reference sample is obtained from a reference subject that includes known genetic mutations.As used herein, the term "reference" can refer to a reference genome, reference sequence, reference sample, and / or reference target.
[0119] In certain embodiments, when the sample nucleic acid is from a pregnant female, the reference sequence sometimes does not derive from the fetus, nor from the fetus's mother, nor from the fetus's father, and this is referred to herein as an "external reference". In some embodiments, a maternal reference can be prepared and used. When preparing a reference from a pregnant female (a "maternal reference sequence") based on an external reference, reads obtained from the DNA of the pregnant female that do not substantially contain fetal DNA are often mapped to and collected against the external reference sequence. In certain embodiments, the external reference is derived from the DNA of an individual having substantially the same ethnicity as the pregnant female. The maternal reference sequence may not completely cover the maternal genomic DNA (e.g., it may cover about 50%, 60%, 70%, 80%, 90% or more of the maternal genomic DNA), and the maternal reference may not exactly match the maternal genomic DNA sequence (e.g., the maternal reference sequence may contain multiple mismatches).
[0120] In certain embodiments, mapability is evaluated for a genomic region (e.g., a portion, a genomic portion). Mapability is the ability to unambiguously align reads of a nucleotide sequence to a portion of the reference genome, typically with a certain number of mismatches, including, for example, 0, 1, 2 or more mismatches. In some embodiments, mapability is provided as a score or value generated by an appropriate mapping algorithm or computer mapping software. For a given genomic region, the expected mapability can be estimated by using a sliding window approach of a pre-set read length and averaging the resulting read-level mapability values. Genomic regions containing stretches of unique nucleotide sequences sometimes have high mapability values.
[0121] The reads of the array can be mapped by a mapping module or a machine including the mapping module, and the mapping module generally maps the reads to a reference genome or a segment thereof. The mapping module can map the reads of the array by suitable methods known in the art. In some embodiments, the mapping module or the machine including the mapping module is required to provide the mapped reads of the array.
[0122] Count number The reads of the mapped array can be quantified to determine the number of reads mapped to a region or portion of the reference genome. In certain embodiments, the reads mapped to the reference genome, or a region, portion, or segment thereof, are referred to as count numbers. In some embodiments, the count number includes a value. In certain embodiments, the value of the count number is determined by mathematical processing. The count number can be determined by suitable methods, operations, or mathematical processing. In certain embodiments, the count number is processed by weighting, removal, filtering, normalization, adjustment, averaging, addition or subtraction, or a combination thereof. In certain embodiments, the count number is derived from the reads of the array that are processed or operated by suitable methods, operations, or mathematical processing described herein or known in the art. For example, the count number is often normalized and / or weighted by one or more biases associated with the reads of the array. In some embodiments, the count number is normalized and / or weighted according to the GC bias associated with the reads of the array. In some embodiments, the count number is derived from the unprocessed reads of the array and / or the filtered reads of the array. In some embodiments, one or more count numbers are not mathematically manipulated. The terms "raw count" and "raw counts" as used herein refer to one or more count numbers that have not been mathematically manipulated.
[0123] In some embodiments, the count is determined for some or all of the reads of the array mapped to a reference genome, or a region, portion, or segment thereof. In certain embodiments, the count is determined from a predefined subset of the reads of the mapped array. The predefined subset of the reads of the mapped array (e.g., the selected subset) can be defined or selected by leveraging any suitable feature or variable. In some embodiments, the predefined subset of the reads of the array to be mapped can include from 1 to n reads of the array, where n means a number equal to the total of all the reads of the array generated from the sample of the subject under test or the reference subject.
[0124] The count is often derived from the reads of the array obtained from a subject (e.g., the subject under test). The count is sometimes derived from the reads of the array obtained from a nucleic acid sample derived from a pregnant female who gives birth to a fetus. The count of the reads of the nucleic acid sequence is often the count that represents both the fetus and the mother of the fetus (e.g., of the pregnant female subject). In certain embodiments, when the subject is a pregnant female, some of the counts are derived from the genome of the fetus and some of the counts are derived from the genome of the mother.
[0125] Read density The count number of reads of an array (e.g., weighted count number) is displayed as read density. Read density is often determined and / or generated for one or more portions of a genome. In certain embodiments, read density is determined and / or generated for one or more chromosomes. In some embodiments, read density includes a quantitative measure of the count number of reads of sequences mapped against portions of a reference genome. Read density can be determined by appropriate processing. In some embodiments, read density is determined by an appropriate distribution and / or an appropriate distribution function. Non-limiting examples of distribution functions include any appropriate distribution, such as a probability function, a probability distribution function, a probability density function (PDF), a kernel density function (kernel density estimation), a cumulative distribution function, a probability mass function, a discrete probability distribution, an absolutely continuous univariate distribution, or a combination thereof. In certain embodiments, the PDF includes a kernel density function (kernel density estimation). Non-limiting examples of kernel density functions that can be used to generate an estimate of local genomic bias include a uniform kernel density function (uniform kernel), a Gaussian kernel density function (Gaussian kernel), a triangular kernel density function (triangular kernel), a biweight kernel density function (biweight kernel), a tricube kernel density function (tricube kernel), a triweight kernel density function (triweight kernel), a cosine kernel function (cosine kernel), an Epanechnikov kernel density function (Epanechnikov kernel), a normal kernel density function (normal kernel), or a combination thereof. Read density is often a density estimate derived from an appropriate probability density function. Density estimation is the construction of an estimate of a underlying probability density function based on observed data. In some embodiments, read density includes a density estimate (e.g., probability density estimation, kernel density estimation). Density estimation often includes kernel density estimation. In some embodiments, read density is a kernel density estimate determined according to a kernel density function. Read density is often a step of generating a density estimate for each of one or more portions of a genome, each portion being generated by a process that includes a step of including a count number of reads of a sequence.Read density is often generated for normalized and / or weighted counts mapped to portions. In some embodiments, each read mapped to a portion often contributes a value (e.g., a count) equal to the read density, its weight obtained from the normalization process described herein. In some embodiments, the read density for one or more portions is adjusted. The read density can be adjusted by suitable methods. For example, the read density for one or more portions can be weighted and / or normalized.
[0126] In some embodiments, the system includes a distribution module 12. The distribution module often generates and / or provides read densities (e.g., 22, 24) for portions of the genome (e.g., filtered portions). The distribution module can provide read densities, read density distributions 14, and / or associated measures of uncertainty (e.g., MAD, quantiles) for one or more reference samples, training sets (e.g., 3), and / or test samples. The distribution module can receive, retrieve, and / or store reads of an array (e.g., 1, 3, 5) and / or count numbers (e.g., normalized count numbers 11, weighted count numbers). The distribution module often receives (e.g., user input and user parameters for a portion), retrieves, generates, and / or stores portions (e.g., unfiltered or filtered portions). Sometimes, the distribution module receives and / or retrieves portions (e.g., filtered portions and / or selected portions 20) from a filtering module 18. In some embodiments, the distribution module includes code and / or source code (e.g., a collection of standard or custom scripts) that implements the functions of the distribution module, and / or instructions (e.g., algorithms, scripts) for a microprocessor in the form of one or more software packages (e.g., a statistical software package). In some embodiments, the distribution module includes code (e.g., a script) written in java, S, or R that utilizes an appropriate package (e.g., an S package, an R package). A non-limiting example of the distribution module is provided in Example 2.
[0127] In some embodiments, a read density profile is determined. In some embodiments, the read density profile includes at least one read density, and often includes two or more read densities (e.g., the read density profile often includes a plurality of read densities). In some embodiments, the read density profile includes an appropriate quantitative value (e.g., mean, median, Z-score, etc.). The read density profile often includes values obtained as a result of one or more read densities. The read density profile includes values obtained as a result of one or more operations on the read density based on one or more adjustments (e.g., normalization). In some embodiments, the read density profile includes unoperated read densities. In some embodiments, one or more read density profiles are generated from various aspects of a dataset containing read densities or a derivative thereof (e.g., the results of one or more mathematical and / or statistical data processing steps known in the art and / or described herein). In certain embodiments, the read density profile includes normalized read densities. In some embodiments, the read density profile includes adjusted read densities. In certain embodiments, the read density profile includes unprocessed read densities (e.g., unoperated, unadjusted or unnormalized), normalized read densities, weighted read densities, filtered partial read densities, Z-scores of read densities, p-values of read densities, integer values of read densities (e.g., area under the curve), mean values of read densities, mean or median, principal components, etc., or combinations thereof. The read density of the read density profile and / or the read density profile is often associated with a measure of uncertainty (e.g., MAD). In certain embodiments, the read density profile includes the distribution of read density medians. In some embodiments, the read density profile includes relationships (e.g., fitted relationships, regression, etc.) of a plurality of read densities. For example, sometimes the read density profile includes the relationship between a read density (e.g., a read density value) and a genomic location (e.g., a portion, a location of the portion).In some embodiments, the read density profile is generated using a stationary window process, and in certain embodiments, the read density profile is generated using a sliding window process. As used herein, the term "density read profile" refers to the result of mathematical and / or statistical manipulation of read density that can facilitate identification of patterns and / or correlations in read data of a large number of arrays.
[0128] In some embodiments, the read density profile is printed and / or displayed (e.g., visually displayed, e.g., displayed as a plot or graph).
[0129] A read density profile often includes a plurality of data points, where each data point represents a quantitative value of one or more read densities. Any suitable number of data points can be incorporated into the read density profile depending on the nature and / or complexity of the data set. In certain embodiments, the read density profile can include two or more data points, three or more data points, five or more data points, ten or more data points, twenty-four or more data points, twenty-five or more data points, fifty or more data points, one hundred or more data points, five hundred or more data points, one thousand or more data points, five thousand or more data points, ten thousand or more data points, one hundred thousand or more data points, or one million or more data points. In some embodiments, the data points are quantitative values and / or estimated values of the count of reads of an array mapped to and / or associated with one or more portions. In some embodiments, the data points in the read density profile include the results of data manipulation of the counts mapped to one or more portions. In certain embodiments, the data points are often quantitative values and / or estimated values of one or more read densities (e.g., average read density). A read density profile often includes a plurality of read densities associated with and / or mapped to multiple portions of a reference genome. In some embodiments, the read density profile includes read densities derived from 2 to about 1,000,000 portions. In some embodiments, read densities derived from 2 to about 500,000, 2 to about 100,000, 2 to about 50,000, 2 to about 40,000, 2 to about 30,000, 2 to about 20,000, 2 to about 10,000, 2 to about 5000, 2 to about 2500, 2 to about 1250, 2 to about 1000, 2 to about 500, 2 to about 250, 2 to about 100, or 2 to about 60 portions determine the read density profile. In some embodiments, read densities derived from about 10 to about 50 portions determine the read density profile.
[0130] In some embodiments, a read density profile corresponds to a series of portions (e.g., a series of portions of a reference genome, a series of portions of a chromosome, or a subset of portions of a segment of a chromosome). In some embodiments, a read density profile includes read densities and / or count numbers associated with a collection of portions (e.g., a set, a subset). In some embodiments, a read density profile is determined for read densities of contiguous portions. In some embodiments, contiguous portions include gaps that include segments of reference sequences and / or reads of sequences not included in the density profile (e.g., portions removed by filtering). Sometimes, contiguous portions (e.g., a series of portions) represent adjacent segments of a genome or adjacent segments of a chromosome or gene. For example, two or more contiguous portions, when aligned by integrating the portions end-to-end, may represent an assembly of a DNA sequence longer than each portion. For example, two or more contiguous portions may represent an intact genome, chromosome, gene, intron, exon, or segment thereof. Sometimes, a read density profile is determined from a collection (e.g., a set, a subset) of contiguous and / or non-contiguous portions. In some cases, a read density profile includes one or more portions that can be transformed by weighting, excluding, filtering, normalizing, adjusting, averaging, deriving as an average, adding, subtracting, processing, or any combination thereof.
[0131] In some embodiments, a read density profile includes read density for a portion of a genome that includes a mutation of a gene. In some embodiments, a read density profile includes read density for a portion of a genome that does not include a mutation of a gene (e.g., a portion of a genome that substantially does not include a mutation of a gene). In certain embodiments, a read density profile includes read density for a portion of a genome that includes a mutation of a gene and read density for a portion of a genome that substantially does not include a mutation of a gene.
[0132] Read density profiles are often determined for a sample and / or a reference (e.g., a reference sample). Read density profiles are sometimes generated for an entire genome, for one or more chromosomes, or for a part or segment of a genome or chromosome. In some embodiments, one or more read density profiles are determined for a genome or segment thereof. In some embodiments, a read density profile is an overall representation of the series of read densities of a sample, and in certain embodiments, a read density profile is a representation of a part or subset of the read densities of a sample. That is, a read density profile may in some cases include or be generated from read densities that display data that has not been filtered to exclude any data, and a read density profile may in some cases include or be generated from data points that display data that has been filtered to exclude unwanted data.
[0133] In some embodiments, a read density profile is determined for a reference (e.g., a reference sample, a training set). A read density profile for a reference is sometimes referred to herein as a reference profile. In some embodiments, a reference profile includes read densities obtained from one or more references (e.g., a reference sequence, a reference sample). In some embodiments, a reference profile includes read densities determined for one or more (e.g., a series of) known euploid samples. In some embodiments, a reference profile includes read densities of a filtered portion. In some embodiments, a reference profile includes read densities adjusted by one or more principal components.
[0134] In some embodiments, the system includes a profile generation module (e.g., 26). The profile generation module often receives, collects, and / or stores read densities (e.g., 22, 24). The profile generation module can receive and / or collect read densities (e.g., adjusted, weighted, normalized, averaged, integrated read densities) from another suitable module (e.g., the distribution module). The profile generation module can receive and / or collect read densities from a suitable source (e.g., one or more references, training sets, one or more test subjects, etc.). The profile generation module often generates and / or provides read density profiles (e.g., 32, 30, 28) to another suitable module (e.g., the PCA statistics module 33, the partial weighting module 42, the scoring module 46), and / or to the user (e.g., by plotting, graphing, and / or printing). An example of, or a portion of, the profile generation module is provided in Example 2.
[0135] portion In some embodiments, the reads and / or count numbers of the mapped array are grouped together according to various parameters and assigned to specific segments and / or regions of a reference genome referred to herein as "portions" or "a portion". In some embodiments, a portion is the entire chromosome, a segment of a chromosome, a segment of the reference genome, a segment spanning multiple chromosomes, multiple segments of chromosomes, and / or combinations thereof. In some embodiments, a portion is predefined based on certain parameters (e.g., a predetermined length, a predetermined interval, a predetermined GC content, or any other suitable parameter). In some embodiments, a portion is arbitrarily defined based on genome partitioning (e.g., partitioning by size, GC content, neighboring regions, neighboring regions of arbitrarily defined size, etc.). In some embodiments, a portion is described based on one or more parameters, including, for example, the length of the sequence or one or more specific features. In some embodiments, a portion is based on a specific length of the genomic sequence. The portions may be approximately the same length or the portions may be of different lengths. In some embodiments, the portions are of approximately equal length. In some embodiments, portions of different lengths are adjusted or weighted. The portions can be of any suitable length. In some embodiments, the portions are about 10 kilobases (kb) to about 100 kb, about 20 kb to about 80 kb, about 30 kb to about 70 kb, about 40 kb to about 60 kb, and sometimes about 50 kb. In some embodiments, the portions are about 10 kb to about 20 kb. The portions are not limited to contiguous runs of the sequence. Thus, a portion can be composed of contiguous and / or non-contiguous sequences.
[0136] In some embodiments, a portion includes a window that contains a preselected number of bases. The window can contain any suitable number of bases determined by the length of the portion. In some embodiments, the genome or a segment thereof is partitioned into a plurality of windows. Windows encompassing regions of the genome may or may not overlap. In some embodiments, windows are positioned equidistant from each other. In some embodiments, windows are positioned at different distances from each other. In certain embodiments, the genome or a segment thereof is partitioned into a plurality of sliding windows that slide the window gradually over the genome or segment thereof. The window can also be slid over the genome at any suitable increment, or according to any numerical pattern or any arbitrary defined sequence. In some embodiments, the window is slid over the genome or a segment thereof in an increment of about 100,000 bp or less, about 50,000 bp or less, about 25,000 bp or less, about 10,000 bp or less, about 5,000 bp or less, about 1,000 bp or less, about 500 bp or less, or about 100 bp or less. For example, the window may contain about 100,000 bp and may be slid over the genome in an increment of 50,000 bp.
[0137] In some embodiments, the portion can be a specific chromosomal segment in a chromosome of interest, such as a chromosome evaluating a genetic mutation (e.g., trisomy of chromosomes 13, 18, and / or 21, or sex chromosome aneuploidy). The portion is not limited to a single chromosome. In some embodiments, one or more portions include all or part of one chromosome, or all or part of two or more chromosomes. In some embodiments, one or more portions can span an entire chromosome, one, two, or more chromosomes. Further, the portion can also span contiguous or scattered regions of multiple chromosomes. The portion can be a gene, a fragment of a gene, a regulatory sequence, an intron, an exon, etc.
[0138] In some embodiments, certain regions of the genome are filtered before partitioning the genome or segments thereof into parts. Regions of the genome can be selected for exclusion from the partitioning process using any suitable method. Regions containing similar regions (e.g., identical or homologous regions or sequences, e.g., repetitive regions) are often removed and / or filtered. Sometimes, unmappable regions are excluded. In some embodiments, only unique regions are retained. Regions removed during partitioning may belong to a single chromosome or span multiple chromosomes. In some embodiments, the partitioned genome can be trimmed, optimized, and often focused on sequences that can be uniquely identified for faster alignment. In some embodiments, partitioning of the genome into regions (e.g., regions beyond the limits of a chromosome) can be based on the gain of information generated in the context of classification. For example, the information content can be quantified using a p-value profile that measures the significance of specific genomic locations for distinguishing between a group of subjects confirmed as normal and a group of subjects confirmed as abnormal (e.g., a group of euploid subjects and a group of trisomic subjects, respectively). In some embodiments, partitioning of the genome into regions (e.g., regions beyond the limits of a chromosome) can be based on any other criterion, such as speed / convenience in aligning reads, GC content (e.g., high or low GC content), uniformity of GC content, other measures of sequence content (e.g., the proportion of individual nucleotides, the proportion of pyrimidines or purines, the proportion of natural nucleic acids to non-natural nucleic acids, the proportion of methylated nucleotides, and CpG content), methylation status, melting temperature of the double strand, compliance with sequencing or PCR, a measure of uncertainty assigned to individual parts of the reference genome, and / or a search targeting specific features.
[0139] A "segment" of a genome is a region that sometimes includes one or more chromosomes, or a part of a chromosome. A "segment" is typically a part of the genome that is different from a portion. A "segment" of a genome and / or chromosome is sometimes in a region of the genome or chromosome that is different from a portion, sometimes does not share polynucleotides with a portion, and sometimes includes polynucleotides that are in a portion. A segment of a genome or chromosome often contains a larger number of nucleotides than a portion (e.g., a segment sometimes includes one or more portions), and a segment of a chromosome sometimes contains a smaller number of nucleotides than a portion (e.g., a segment is sometimes inside a portion).
[0140] Filtering of portions In certain embodiments, one or more portions (e.g., genomic portions) are excluded from consideration by a filtering process. In certain embodiments, one or more portions are filtered (e.g., subjected to a filtering process), thereby presenting the filtered portions. In some embodiments, a particular portion is excluded by a filtering process and portions (e.g., a subset of portions) are retained. Herein, the portions retained after the filtering process are often referred to as the filtered portions. In some embodiments, reference genomic portions are filtered. In some embodiments, reference genomic portions excluded by the filtering process are not included in the determination of the presence or absence of a genetic mutation (e.g., aneuploidy). In some embodiments, portions of chromosomes in the reference genome are filtered. In some embodiments, portions associated with read density (e.g., where the read density is the read density for the portion) are excluded by the filtering process, and the read density associated with the excluded portion is not included in the determination of the presence or absence of a genetic mutation (e.g., aneuploidy). In some embodiments, the read density profile comprises and / or consists of the read density of the filtered portions. Portions can be selected, filtered, and / or excluded from consideration using any suitable criteria and / or methods known in the art or described herein. Non-limiting examples of criteria used for filtering portions include redundant data (e.g., redundancy or duplication of mapped reads), uninformative data (e.g., reference genomic portions with a mapped count number of zero), reference genomic portions having overrepresented or underrepresented sequences, GC content, noise data, mappability, count number, variability of count number, read density, variability of read density, measure of uncertainty, measure of reproducibility, etc., or combinations of the foregoing. Portions are optionally filtered according to the distribution of count number and / or the distribution of read density. In some embodiments, portions are filtered according to the distribution of count number and / or read density when the count number and / or read density are obtained from one or more reference samples.In this specification, in some cases, one or more reference samples are referred to as a training set. In some embodiments, portions are filtered according to the distribution of the count number and / or the read density as obtained from one or more test samples and / or according to the read density. In some embodiments, portions are filtered according to a measure of uncertainty about the read density distribution. In certain embodiments, portions that back up a large deviation in the read density are excluded by a filtering process. For example, when each read density in the distribution is mapped to the same portion, the distribution of the read density (e.g., the distribution of the average value of the read density, the mean of the read density, or the median of the read density; e.g., the distribution in FIG. 5A) can be determined. When each portion of the genome is associated with a measure of uncertainty, the measure of uncertainty (e.g., MAD) can be determined by comparing the distribution of the read density for a plurality of samples. According to the previous example, portions can be filtered according to the measure of uncertainty (e.g., standard deviation (SD), MAD) associated with each portion and a predetermined threshold value. FIG. 5B shows the distribution of the MAD values for portions, which is a distribution determined according to the read density distribution for a plurality of samples. The predetermined threshold value is indicated by a vertical dashed line surrounding the range of acceptable MAD values. In the example of FIG. 5B, portions containing MAD values within the acceptable range are retained, and portions containing MAD values outside the acceptable range are excluded from consideration by a filtering process. In some embodiments, according to the previous example, portions containing read density values outside a predetermined measure of uncertainty (e.g., the median, mean, or average of the read density) are often excluded from consideration by a filtering process. In some embodiments, portions containing read density values outside the interquartile range of the distribution (e.g., the median, mean, or average of the read density) are excluded from consideration by a filtering process. In some embodiments, portions containing read density values that deviate from the interquartile range of the distribution by more than 2-fold, 3-fold, 4-fold, or 5-fold are excluded from consideration by a filtering process.In some embodiments, portions including read density values that deviate beyond 2 sigma, 3 sigma, 4 sigma, 5 sigma, 6 sigma, 7 sigma, or 8 sigma (e.g., where sigma is a range defined by the standard deviation) are excluded from consideration by filtering processing.
[0141] In some embodiments, the system includes a filtering module 18. The filtering module often receives, retrieves, and / or stores a portion (e.g., a portion of a predetermined size and / or overlapping, the position of the portion in the reference genome) and a read density associated with the portion, which often originates from another suitable module (e.g., the distribution module 12). In some embodiments, a selected portion (e.g., 20, e.g., the filtered portion) is presented by the filtering module. In some embodiments, the filtering module is requested to present the filtered portion and / or to exclude the portion from consideration. In certain embodiments, when the read density is associated with an excluded portion, the filtering module excludes the read density from consideration. The filtering module often presents the selected portion (e.g., the filtered portion) to another suitable module (e.g., the distribution module 21). Non-limiting examples of the filtering module are presented in Example 3.
[0142] Estimated value of bias Sequencing techniques can be vulnerable to multiple sources of bias. In some cases, the bias in sequencing is a local bias (e.g., local genomic bias). Local bias often manifests at the level of reads of the sequence. Local genomic bias can be any suitable local bias. Non-limiting examples of local bias include biases correlated with sequence bias (e.g., GC bias, AT bias, etc.), DNase I sensitivity, entropy, repetitive sequence bias, chromatin structure bias, polymerase error rate bias, palindromic sequence bias, inverted repeat bias, PCR-related bias, etc., or combinations thereof. In some embodiments, the source of local bias is undetermined or unknown.
[0143] In some embodiments, an estimate of the local genomic bias is determined. As used herein, the estimate of the local genomic bias is sometimes referred to as the estimation of the local genomic bias. The estimate of the local genomic bias can be determined for a reference genome, a segment or portion thereof. In certain embodiments, the estimate of the local genomic bias is determined for one or more chromosomes in the reference genome. In some embodiments, the estimate of the local genomic bias is determined for one or more sequence reads (e.g., reads of a part or all of a sample). The estimate of the local genomic bias is often determined for the sequence reads according to the estimation of the local genomic bias at the corresponding positions and / or loci of the reference (e.g., reference genome, chromosome in the reference genome). In some embodiments, the estimate of the local genomic bias includes a quantitative measure of the bias of the sequence (e.g., read of the reference genome sequence, sequence). The estimation of the local genomic bias can be determined by an appropriate method or mathematical process. In some embodiments, the estimate of the local genomic bias is determined by an appropriate distribution and / or an appropriate distribution function (e.g., PDF). In some embodiments, the estimate of the local genomic bias includes a quantitative representation of the PDF. In some embodiments, the estimate of the local genomic bias (e.g., probability density estimation (PDE), kernel density estimation) is determined by a probability density function of the local bias content (e.g., PDF: probability density function, e.g., kernel density function). In some embodiments, the density estimation includes kernel density estimation. The estimate of the local genomic bias is sometimes represented as the mean value, average, or median of the distribution. Sometimes, the estimate of the local genomic bias is represented as the sum or integral of an appropriate distribution (e.g., area under a curve (AUC)).
[0144] A PDF (e.g., a kernel density function, e.g., an Epanechnikov kernel density function) often includes a bandwidth variable (e.g., bandwidth). The bandwidth variable often defines the size and / or length of a window that derives a probability density estimate (PDE) when using the PDF. The window for deriving the PDE often includes a polynucleotide of a defined length. In some embodiments, the window for deriving the PDE is a portion. The portion (e.g., the size of the portion, the length of the portion) often depends on the bandwidth variable. The bandwidth variable determines the length or size of a window used to determine an estimate of local genomic bias, from which an estimate of local genomic bias is determined, which is the length of a polynucleotide segment (e.g., a continuous segment of nucleotide bases). The length or size of the window is determined. Non-limiting examples thereof include any suitable bandwidth, such as a bandwidth of about 5 bases to about 100,000 bases, about 5 bases to about 50,000 bases, about 5 bases to about 25,000 bases, about 5 bases to about 10,000 bases, about 5 bases to about 5,000 bases, about 5 bases to about 2,500 bases, about 5 bases to about 1000 bases, about 5 bases to about 500 bases, about 5 bases to about 250 bases, about 20 bases to about 250 bases, etc., to determine a PDE (e.g., read density, an estimate of local genomic bias (e.g., GC density)). In some embodiments, an estimate of local genomic bias (e.g., GC density) is determined using a bandwidth of about 400 bases or less, about 350 bases or less, about 300 bases or less, about 250 bases or less, about 225 bases or less, about 200 bases or less, about 175 bases or less, about 150 bases or less, about 125 bases or less, about 100 bases or less, about 75 bases or less, about 50 bases or less, or about 25 bases or less. In certain embodiments, an estimate of local genomic bias (e.g., GC density) is determined using a bandwidth determined according to the average read length, average read length, median read length, or maximum read length of reads of the sequence obtained for a given subject and / or sample.In some cases, an estimate of local genomic bias (e.g., GC density) is determined using a bin width that is approximately equal to the average read length, mean read length, median read length, or maximum read length of the reads of the sequences obtained for a given subject and / or sample. In some embodiments, an estimate of local genomic bias (e.g., GC density) is determined using a bin width of about 250, 240, 230, 220, 210, 200, 190, 180, 160, 150, 140, 130, 120, 110, 100, 90, 80, 70, 60, 50, 40, 30, 20, or about 10 bases.
[0145] Estimates of local genomic bias can be determined at single-base resolution, but estimates of local genomic bias (e.g., local GC content) can also be determined at low resolution. In some embodiments, an estimate of local genomic bias is determined for the content of local bias. Estimates of local genomic bias (e.g., determined using a PDF) are often determined using windows. In some embodiments, an estimate of local genomic bias involves the use of a window that includes a preselected number of bases. Optionally, the window includes a segment of contiguous bases. Optionally, the window includes one or more non-contiguous portions of bases. Optionally, the window includes one or more portions (e.g., genomic portions). The size or length of the window is often determined by the bandwidth and according to a PDF. In some embodiments, the window is about 10 times or more, 8 times or more, 7 times or more, 6 times or more, 5 times or more, 4 times or more, 3 times or more, or about 2 times or more the length of the bandwidth. When using a PDF (e.g., a kernel density function) to determine a density estimate, the window is optionally twice the length of the selected bandwidth. The window can include any suitable number of bases. In some embodiments, the window includes from about 5 bases to about 100,000 bases, from about 5 bases to about 50,000 bases, from about 5 bases to about 25,000 bases, from about 5 bases to about 10,000 bases, from about 5 bases to about 5,000 bases, from about 5 bases to about 2,500 bases, from about 5 bases to about 1000 bases, from about 5 bases to about 500 bases, from about 5 bases to about 250 bases, or from about 20 bases to about 250 bases. In some embodiments, the genome or a segment thereof is partitioned into a plurality of windows. Windows that encompass regions of the genome may or may not overlap. In some embodiments, the windows are spaced equidistant from each other. In some embodiments, the windows are spaced at different distances from each other. In certain embodiments, the genome or a segment thereof is partitioned into a plurality of sliding windows where the windows are gradually slid across the genome or segment thereof.Each window of each increment contains an estimate of local genomic bias (e.g., local GC density). The window can also be slid across the genome with any suitable increment, can be slid according to any numerical pattern, or can be slid according to any non-subject defined sequence. In some embodiments, to determine the estimate of local genomic bias, the window is slid across the genome or a segment thereof with an increment of about 10,000 bp or more, about 5,000 bp or more, about 2,500 bp or more, about 1,000 bp or more, about 750 bp or more, about 500 bp or more, about 400 bases or more, about 250 bp or more, about 100 bp or more, about 50 bp or more, or about 25 bp or more. In some embodiments, to determine the estimate of local genomic bias, the window is slid across the genome or a segment thereof with an increment of about 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, 2, or about 1 bp. For example, to determine the estimate of local genomic bias, the window can include about 400 bp (e.g., a 200 bp bandwidth) and can be slid across the genome with a 1 bp increment. In some embodiments, a kernel density function and a bandwidth of about 200 bp are used to determine the estimate of local genomic bias for each base within the genome or a segment thereof.
[0146] In some embodiments, the estimate of local genomic bias is a measure of local GC content and / or an indication of local GC content. As used herein, the term "local" (e.g., as used to describe local bias, estimate of local bias, local bias content, local genomic bias, local GC content, etc.) refers to a polynucleotide segment of 10,000 bp or less. In some embodiments, the term "local" refers to a polynucleotide segment of 5000 bp or less, 4000 bp or less, 3000 bp or less, 2000 bp or less, 1000 bp or less, 500 bp or less, 250 bp or less, 200 bp or less, 175 bp or less, 150 bp or less, 100 bp or less, 75 bp or less, or 50 bp or less. Local GC content is often an indication (e.g., a mathematical indication, a quantitative indication) of GC content for a local segment of a genome, sequence read, sequence read assembly (e.g., contig, profile, etc.). For example, local GC content may be an estimate of local GC bias, or may be local GC density.
[0147] One or more GC densities are often determined for the polynucleotides of a reference or a sample (e.g., a test sample). In some embodiments, the GC density is an indication (e.g., a mathematical indication, a quantitative indication) of the local GC content (e.g., for a polynucleotide segment of 5000 bp or less). In some embodiments, the GC density is an estimate of the local genomic bias. The GC density can be determined using appropriate processing described herein and / or appropriate processing known in the art. The GC density can be determined using an appropriate PDF (e.g., a kernel density function (e.g., an Epanechnikov kernel density function, see e.g., FIG. 1)). In some embodiments, the GC density is a PDE (e.g., kernel density estimation). In certain embodiments, the GC density is defined by the presence or absence of one or more guanine (G) nucleotides and / or cytosine (C) nucleotides. Conversely, in some embodiments, the GC density can also be defined by the presence or absence of one or more adenine (A) nucleotides and / or thymidine (T) nucleotides. In some embodiments, the GC density for the local GC content is normalized according to the GC density determined for the whole genome or a segment thereof (e.g., an autosome, a set of chromosomes, a single chromosome, a gene; see e.g., FIG. 2). One or more GC densities can be determined for the polynucleotides of a sample (e.g., a test sample) or a reference sample. The GC density is often determined for a reference genome. In some embodiments, the GC density is determined for sequence reads according to a reference genome. The GC density of a read is often determined according to the GC density determined for the corresponding position and / or locus of the reference genome to which the read maps. In some embodiments, the GC density determined for a position on the reference genome is assigned to and / or presented for a read, where the read or a segment thereof maps to a position on the same reference genome. The position of a read mapped on the reference genome can be determined for the purpose of generating the GC density for the read using any suitable method.In some embodiments, the median point of the mapped reads determines a position on the reference genome at which the GC density for the reads derived therefrom is determined. For example, if the median point of a read maps to chromosome 12 at base number x of the reference genome, the GC density of the read is often presented as the GC density determined by kernel density estimation for a location on chromosome 12 at or near base number x of the reference genome. In some embodiments, the GC density is determined for some or all of the base locations of the read according to the reference genome. In some cases, the GC density of a read includes the mean, sum, median, or integral of two or more GC densities determined for multiple base locations on the reference genome.
[0148] In some embodiments, the estimation of local genomic bias (e.g., GC density) is quantified and / or presented as a value. The estimation of local genomic bias (e.g., GC density) is sometimes represented as a mean value, average, and / or median. The estimation of local genomic bias (e.g., GC density) is sometimes represented as the maximum peak height of a PDE. In some cases, the estimation of local genomic bias (e.g., GC density) is represented as the sum or integral of an appropriate PDE (e.g., area under the curve (AUC)). In some embodiments, the GC density includes kernel weights. In certain embodiments, the GC density of a read includes a value approximately equal to the mean value, average, sum, median, maximum peak height, or integral of the kernel weights.
[0149] Bias frequency The bias frequency is, in some cases, determined according to an estimate of the bias of one or more local genomes (e.g., GC density). The bias frequency is, in some cases, a count or sum of the occurrences of an estimate of the bias of a local genome for a sample, a reference (e.g., a reference genome, a reference sequence, a chromosome in a reference genome), or a portion thereof. The bias frequency is, in some cases, a count or sum of the occurrences of an estimate of the bias of a local genome (e.g., an estimate of the bias of each local genome) for a sample, a reference, or a portion thereof. In some embodiments, the bias frequency is a GC density frequency. The GC density frequency is often determined according to one or more GC densities. For example, the GC density frequency can indicate the number of times a GC density of value x is represented across the entire genome or a segment thereof. The bias frequency is often a distribution of estimates of the bias of local genomes, where the number of occurrences of an estimate of the bias of each local genome is represented as the bias frequency (see, e.g., FIG. 3). The bias frequency is, in some cases, mathematically manipulated and / or normalized. The bias frequency can be mathematically manipulated and / or normalized by an appropriate method. In some embodiments, the bias frequency is normalized according to the representation (e.g., fraction, percentage) of an estimate of the bias of each local genome (e.g., an autosome, a subset of chromosomes, a single chromosome, or a read thereof) for a sample, a reference, or a portion thereof. The bias frequency can be determined for estimates of the bias of some or all of the local genomes of a sample or a reference. In some embodiments, the bias frequency can be determined for estimates of the bias of local genomes for some or all of the sequence reads of a test sample.
[0150] In some embodiments, the system includes a bias density module 6. The bias density module receives, retrieves, and / or stores the mapped array of reads 5 and the reference array 2 in any suitable format, and is capable of generating an estimate of the local genomic bias, the local genomic bias distribution, the bias frequency, the GC density, the GC density distribution, and / or the GC density frequency (collectively, shown by box 7). In some embodiments, the bias density module transfers data and / or information (e.g., 7) to another suitable module (e.g., the relationship module 8).
[0151] Relationship In some embodiments, one or more relationships are generated between an estimate of the local genomic bias and the bias frequency. As used herein, the term "relationship" refers to a mathematical and / or graphical relationship between two or more variables or values. A relationship can be generated by appropriate mathematical and / or graphical processing. Non-limiting examples of relationships include mathematical representations and / or graphical representations of functions, correlations, distributions, linear or non-linear, straight lines, regressions, fitted regressions, etc., or combinations thereof. In some cases, the relationship includes a fitted relationship. In some embodiments, the fitted relationship includes a fitted regression. In some cases, the relationship is between two or more variables or values that include weighted variables or weighted values. In some embodiments, the relationship includes a fitted regression where one or more variables or values of the relationship are weighted. In some cases, the regression is fitted in a weighted manner. In some cases, the regression is fitted without weighting. In certain embodiments, the generation of the relationship includes plotting or graphing.
[0152] In some embodiments, an appropriate relationship is determined between an estimate of local genomic bias and a bias frequency. In some embodiments, a sample bias relationship is presented by generating a relationship between (i) an estimate of local genomic bias and (ii) a bias frequency for a sample. In some embodiments, a reference bias relationship is presented by generating a relationship between (i) an estimate of local genomic bias and (ii) a bias frequency for a reference. In certain embodiments, the relationship is generated between a GC density and a GC density frequency. In some embodiments, a sample GC density relationship is presented by generating a relationship between (i) a GC density and (ii) a GC density frequency for a sample. In some embodiments, a reference GC density relationship is presented by generating a relationship between (i) a GC density and (ii) a GC density frequency for a reference. In some embodiments, when the estimate of local genomic bias is a GC density, the sample bias relationship is the sample GC density relationship and the reference bias relationship is the reference GC density relationship. The GC density of the reference GC density relationship and / or the sample GC density relationship is often a representation (e.g., a mathematical or quantitative representation) of the local GC content. In some embodiments, the relationship between the estimate of local genomic bias and the bias frequency includes a distribution. In some embodiments, the relationship between the estimate of local genomic bias and the bias frequency includes a fitted relationship (e.g., a fitted regression). In some embodiments, the relationship between the estimate of local genomic bias and the bias frequency includes a linear or non-linear fitted regression (e.g., a polynomial regression). In certain embodiments, the relationship between the estimate of local genomic bias and the bias frequency includes a weighted relationship, where the estimate of local genomic bias and / or the bias frequency is weighted by an appropriate process. In some embodiments, a weighted fitted relationship (e.g., a weighted fit) can be obtained by a process that includes a quartile regression, a parameterized probability distribution, or an empirical distribution with interpolation. In certain embodiments, the relationship between the estimate of local genomic bias and the bias frequency for a test sample, a reference, or a portion thereof includes a polynomial regression, and the estimate of local genomic bias is weighted. In some embodiments, the weighted fit model includes weighting distribution values.The distribution values can be weighted by appropriate processing. In some embodiments, values located near the tails of the distribution are given a smaller weight than values closer to the median of the distribution. For example, for the distribution of an estimate of local genomic bias (e.g., GC density) and the bias frequency (e.g., GC density frequency), the weight is determined according to the bias frequency for a given estimate of local genomic bias, where an estimate of local genomic bias that includes a bias frequency closer to the mean of the distribution is given a larger weight than an estimate of local genomic bias that includes a bias frequency farther from the mean.
[0153] In some embodiments, the system includes a relationship module 8. The relationship module can generate, among other things, functions, coefficients, constants, and variables that define the relationship. The relationship module can receive, store, and / or retrieve data and / or information (e.g., 7) from an appropriate module (e.g., the bias density module 6) and generate a relationship. The relationship module often generates and compares distributions of estimates of local genomic bias. The relationship module can compare data sets and, in some cases, generate regression and / or fitted relationships. In some embodiments, the relationship module compares one or more distributions (e.g., the distributions of estimates of local genomic bias of a sample and / or a reference) and presents a weighting factor and / or weight assignment 9 for the count number of array reads to another appropriate module (e.g., the bias correction module). In some cases, the relationship module directly presents the count number of reads of a normalized array to the distribution module 21, where the count number is normalized according to relationships and / or comparisons.
[0154] Generation and Use of Comparisons In some embodiments, the process for reducing local bias during array reads includes normalizing the count of array reads. The count of array reads is often normalized according to a comparison with a reference of the test sample. For example, in some cases, the count of array reads is normalized by comparing an estimated value of the local genomic bias of the array reads of the test sample with an estimated value of the local genomic bias of a reference (e.g., a reference genome or a part thereof). In some embodiments, the count of array reads is normalized by comparing the bias frequency of the estimated value of the local genomic bias of the test sample with the bias frequency of the estimated value of the local genomic bias of the reference. In some embodiments, the count of array reads is normalized by comparing a sample bias relationship with a reference bias relationship, thereby generating a comparison.
[0155] The count of array reads is often normalized according to the comparison of two or more relationships. In certain embodiments, two or more relationships are compared, thereby presenting the comparisons used to reduce local bias (e.g., normalize the count) during array reads. By appropriate methods, two or more relationships can be compared. In some embodiments, the comparison includes adding the second relationship to the first relationship, subtracting the second relationship from the first relationship, multiplying the first relationship by the second relationship, and / or dividing the first relationship by the second relationship. In certain embodiments, the comparison of two or more relationships includes the use of appropriate linear regression and / or non-linear regression. In certain embodiments, the comparison of two or more relationships includes appropriate polynomial regression (e.g., cubic polynomial regression). In some embodiments, the comparison includes adding the second regression to the first regression, subtracting the second regression from the first regression, multiplying the first regression by the second regression, and / or dividing the first regression by the second regression. In some embodiments, two or more relationships are compared by a process including an inference framework of multiple regression. In some embodiments, two or more relationships are compared by a process including appropriate multivariate analysis. In some embodiments, two or more relationships are compared by a process including basis functions (e.g., blending functions, e.g., polynomial basis, Fourier basis, etc.), splines, radial basis functions, and / or wavelets.
[0156] In certain embodiments, the distribution of estimates of local genomic bias, including the bias frequencies for the test sample and the reference, is compared by a process that includes polynomial regression, where the estimates of local genomic bias are weighted. In some embodiments, the polynomial regression is generated between (i) a ratio each of whose terms includes the bias frequency of the estimate of local genomic bias of the reference and the bias frequency of the estimate of local genomic bias of the sample, and (ii) the estimate of local genomic bias. In some embodiments, the polynomial regression is generated between (i) the ratio of the bias frequency of the estimate of local genomic bias of the reference to the bias frequency of the estimate of local genomic bias of the sample, and (ii) the estimate of local genomic bias. In some embodiments, the comparison of the distributions of estimates of local genomic bias for the test sample and the reference reads includes determining the log ratio (e.g., log2 ratio) of the bias frequencies of the estimates of local genomic bias for the reference and the sample. In some embodiments, the comparison of the distributions of estimates of local genomic bias includes dividing the log ratio (e.g., log2 ratio) of the bias frequency of the estimate of local genomic bias for the reference by the log ratio (e.g., log2 ratio) of the bias frequency of the estimate of local genomic bias for the sample (see, e.g., Example 1 and FIG. 4).
[0157] Typically, when normalizing the counts according to a comparison, some counts are adjusted while others are not. When normalizing the counts, in some cases all counts are adjusted, and in some cases the counts of the reads of any array are not adjusted. The counts for the reads of an array are normalized in some cases by a process that includes determining a weighting factor, and in some cases the process does not include the direct generation and utilization of a weighting factor. Normalizing the counts according to a comparison sometimes includes determining a weighting factor for the counts of the reads of each array. The weighting factor is specific to a read of an array and is often applied to the count of the read of the specific array. The weighting factor is often determined according to a comparison of two or more bias relationships (e.g., a sample bias relationship compared to a reference bias relationship). The normalized counts are often determined by adjusting the count values according to the weighting factor. Adjusting the counts according to the weighting factor sometimes includes adding the weighting factor to the count of the read of an array, subtracting the weighting factor from the count of the read of an array, multiplying the count of the read of an array by the weighting factor, and / or dividing the count of the read of an array by the weighting factor. The weighting factor and / or the normalized counts are sometimes determined from a regression (e.g., a regression line). The normalized counts are sometimes obtained directly from a regression line (e.g., a fitted regression line) that results from a comparison of the bias frequency of an estimate of the local genomic bias of a reference (e.g., a reference genome, a chromosome in the reference genome) with the bias frequency of an estimate of the local genomic bias of a test sample. In some embodiments, each count of the reads of a sample is presented as a normalized count value according to a comparison of (i) the bias frequency of an estimate of the local genomic bias of the read with (ii) the bias frequency of an estimate of the local genomic bias of a reference. In certain embodiments, the counts of the reads of an array obtained for a sample are normalized to reduce the bias in the reads of the array.
[0158] In some cases, the system includes a bias correction module 10. In some embodiments, the functionality of the bias correction module is performed by the relationship modeling module 8. The bias correction module can receive, retrieve, and / or store the mapped array reads and weighting factors (e.g., 9) from an appropriate module (e.g., relationship module 8, compression module 4). In some embodiments, the bias correction module presents a count number to the mapped reads. In some embodiments, the bias correction module applies a weighting and / or bias correction factor to the count number of the array reads, thereby presenting a normalized and / or adjusted count number. The bias correction module often presents the normalized count number to another appropriate module (e.g., distribution module 21).
[0159] In certain embodiments, normalizing the count number includes factoring one or more features in addition to the GC density and normalizing the count number of the array reads. In certain embodiments, normalizing the count number includes factoring an estimate of one or more different local genomic biases and normalizing the count number of the array reads. In certain embodiments, the count number of the array reads is weighted according to a weighting determined according to one or more features (e.g., one or more biases). In some embodiments, the count number is normalized according to one or more combined weights. In some cases, factoring one or more features and / or normalizing the count number according to one or more combined weights involves a process that includes the use of a multivariate model. Any suitable multivariate model can be used to normalize the count number. Non-limiting examples of multivariate models include multivariate linear regression, multivariate quantile regression, multivariate interpolation of empirical data, non-linear multivariate models, etc., or combinations thereof.
[0160] In some embodiments, the system includes a multivariate correction module 13. The multivariate correction module performs the functions of the bias density module 6, the relationship module 8, and / or the bias correction module 10 multiple times, thereby enabling adjustment of the counts for multiple biases. In some embodiments, the multivariate correction module includes one or more of the bias density modules 6, the relationship module 8, and / or the bias correction module 10. Optionally, the multivariate correction module presents the normalized counts 11 to another suitable module (e.g., the distribution module 21).
[0161] Weighted portion In some embodiments, a portion is weighted. In some embodiments, one or more portions are weighted, thereby presenting a weighted portion. The weighted portion optionally removes portion dependencies. The portion can be weighted by appropriate processing. In some embodiments, one or more portions are weighted by an eigen function (or eigenfunction). In some embodiments, the eigen function includes replacing the portion with orthogonal eigen portions. In some embodiments, the system includes a portion weighting module 42. In some embodiments, the weighting module receives, retrieves, and / or stores a lead density, a lead density profile, and / or an adjusted lead density profile. In some embodiments, the weighted portion is presented by the portion weighting module. In some embodiments, the weighting module is requested to weight the portion. The weighting module can weight the portion by one or more weighting methods known in the art or described herein. The weighting module often presents the weighted portion to another suitable module (e.g., the scoring module 46, the PCA statistics module 33, the profile generation module 26, etc.).
[0162] Principal component analysis In some embodiments, a read density profile (e.g., the read density profile of a test sample (e.g., FIG. 7A)) is adjusted according to principal component analysis (PCA). The read density profiles of one or more reference samples and / or the read density profile of the object under test can be adjusted according to PCA. The read density profile for a genome, a part of the genome, a chromosome, or a segment of a chromosome can be adjusted by PCA. In this specification, in some cases, the removal of bias from the read density profile through PCA-related processing is referred to as adjustment of the profile. PCA can be performed by an appropriate PCA method or a variation thereof. Non-limiting examples of PCA methods include canonical correlation analysis (CCA), Karhunen-Loeve (KL) transform (KLT), Hotelling transform, proper orthogonal decomposition (POD), singular value decomposition (SVD) of X, eigenvalue decomposition (EVD) of XTX, factor analysis, Eckart-Young theorem, Schmidt-Mirsky theorem, empirical orthogonal function (EOF), empirical eigenfunction decomposition, empirical component analysis, quasi-harmonic mode, spectral decomposition, empirical mode analysis, etc., including variations or combinations of these. PCA often identifies one or more biases in the read density profile. In this specification, in some cases, the bias identified by PCA is referred to as a principal component. In some embodiments, one or more biases can be excluded by adjusting the read density profile according to one or more principal components using an appropriate method. The read density profile can be adjusted by adding one or more principal components to the read density profile, subtracting one or more principal components from the read density profile, multiplying the read density profile by one or more principal components, and / or dividing the read density profile by one or more principal components. In some embodiments, one or more biases can be excluded from the read density profile by subtracting one or more principal components from the read density profile.Biases in read density profiles are often identified and / or quantified by PCA of the profiles, but the principal components are often subtracted from the profiles at the level of read density. Biases or features of read density profiles that are identified and / or quantified by PCA of the profiles include, but are not limited to, fetal gender, sequence bias (e.g., guanine and cytosine (GC) bias), fetal fraction, bias correlated with DNase I sensitivity, entropy, repetitive sequence bias, chromatin structure bias, polymerase error rate bias, palindromic sequence bias, inverted repeat bias, PCR amplification bias, and hidden copy number variations.
[0163] PCA is often used to identify one or more principal components. In some embodiments, PCA is used to identify principal components ranked first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, and tenth, or more. In certain embodiments, one, two, three, four, five, six, seven, eight, nine, ten or more principal components are used to adjust the profile. In certain embodiments, five principal components are used to adjust the profile. The principal components are often used to adjust the profile in the order of their appearance in the PCA. For example, when subtracting three principal components from the read density profile, the first, second, and third principal components are used. In some cases, the biases identified by the principal components are features of the profile and include features that are not used to adjust the profile. For example, PCA can identify gene mutations (e.g., aneuploidy, deletion, translocation, insertion) and / or sex differences (e.g., as seen in FIG. 6C) as principal components. Thus, in some embodiments, one or more principal components are not used to adjust the profile. For example, in some cases, the first, second, and fourth principal components are used to adjust the profile, where the third principal component is not used to adjust the profile. The principal components can be obtained from the PCA using any suitable sample or reference. In some embodiments, the principal components are obtained from a test sample (e.g., a test subject). In some embodiments, the principal components are obtained from one or more references (e.g., reference sample, reference sequence, reference set). For example, as shown in FIG. 6, PCA is performed on the median read density profile obtained from a training set (FIG. 6A) that includes a plurality of samples that result in the identification of a first principal component (FIG. 6B) and a second principal component (FIG. 6C). In some embodiments, the principal components are obtained from a set of subjects known to lack the mutations of the gene in question. In some embodiments, the principal components are obtained from a set of known euploidies. The principal components are often identified according to a PCA performed using one or more read density profiles of the reference (e.g., the training set).One or more principal components obtained from a reference are subtracted from the lead density profile of the subject under test (e.g., FIG. 7B), thereby often presenting an adjusted profile (e.g., FIG. 7C).
[0164] In some embodiments, the system includes a PCA statistics module 33. The PCA statistics module can receive and / or retrieve a lead density profile from another suitable module (e.g., the profile generation module 26). PCA is often performed by the PCA statistics module. The PCA statistics module often receives, retrieves, and / or stores a lead density profile and processes the lead density profile from the reference set 32, the training set 30, and / or one or more subjects under test 28. The PCA statistics module can generate and / or present principal components and / or adjust the lead density profile according to one or more principal components. The adjusted lead density profile (e.g., 40, 38) is often provided by the PCA statistics module. The PCA statistics module can present and / or transfer the adjusted lead density profile (e.g., 38, 40) to another suitable module (e.g., the partial weighting module 42, the scoring module 46). In some embodiments, the PCA statistics module can present a gender determination 36. The gender determination is, in some cases, a determination of the gender of the fetus determined according to PCA and / or according to one or more principal components. In some embodiments, the PCA statistics module includes some, all, or one modification of the R code shown below. The R code for calculating the principal components generally begins with data cleaning (e.g., subtracting the median, filtering parts, and trimming extreme values).
Number
Number
Number
[0165] Comparison of profiles In some embodiments, the determination of the outcome involves a comparison. In certain embodiments, a read density profile or a portion thereof is utilized to present the outcome. In certain embodiments, a read density profile for a genome, a portion of a genome, a chromosome, or a segment of a chromosome is utilized for providing the outcome. In some embodiments, the determination of the outcome (e.g., determination of the presence or absence of a gene mutation) involves a comparison of two or more read density profiles. The comparison of read density profiles often involves a comparison of read density profiles made for a selected segment of the genome. For example, a test profile is often compared to a reference profile, and the test profile and the reference profile are determined for segments of the genome (e.g., a reference genome) that are substantially the same segment. The comparison of read density profiles may in some cases involve a comparison of two or more subsets of portions of the read density profile. A subset of a portion of a read density profile may represent a segment of the genome (e.g., a chromosome or a segment thereof). A read density profile may include any amount of subsets of portions. In some cases, a read density profile includes two or more, three or more, four or more, or five or more subsets. In certain embodiments, a read density profile includes two subsets of portions, where each portion represents an adjacent segment of the reference genome. In some embodiments, a test profile can be compared to a reference profile, where the test profile and the reference profile both include a first subset of portions and a second subset of portions, where the first subset and the second subset represent different segments of the genome. A certain subset of a portion of a read density profile may be capable of including a gene mutation, and other subsets of portions may in some cases substantially not include a gene mutation. In some cases, all subsets of portions of a profile (e.g., a test profile) substantially do not include a gene mutation. In some cases, all subsets of portions of a profile (e.g., a test profile) include a gene mutation.In some embodiments, the test profile may include a first subset of portions that include mutations of genes, and a second subset of portions that substantially do not include mutations of genes.
[0166] In some embodiments, the methods described herein include pre-forming a comparison (e.g., comparing a test profile to a reference profile). By appropriate methods, comparisons can be made for two or more data sets, two or more relationships, and / or two or more profiles. Non-limiting examples of statistical methods suitable for comparing data sets, relationships, and / or profiles include the Behrens-Fisher method, the bootstrap method, Fisher's method for combining independent significance tests, the Neyman-Pearson test, confirmatory data analysis, exploratory data analysis, exact tests, F-tests, Z-tests, T-tests, measures of uncertainty, null hypotheses, counter-null hypotheses, etc., calculations and / or comparisons such as chi-square tests, omnibus tests, calculations and / or comparisons of levels of significance (e.g., statistical significance), meta-analysis, multivariate analysis, regression, simple linear regression, robust linear regression, etc., or combinations of the foregoing. In certain embodiments, the comparison of two or more data sets, relationships, and / or profiles includes determining and / or comparing a measure of uncertainty. As used herein, "measure of uncertainty" refers to a measure of significance (e.g., statistical significance), measure of error, measure of variance, measure of reliability, etc., or combinations thereof. A measure of uncertainty may be a value (e.g., a threshold) or a range of values (e.g., an interval, a confidence interval, a Bayesian confidence interval, a threshold range). Non-limiting examples of measures of uncertainty include p-values, appropriate measures of deviation (e.g., standard deviation, sigma, absolute deviation, mean absolute deviation, etc.), appropriate measures of error (e.g., standard error, root mean square error, root mean square error of the second order, etc.), appropriate measures of variance, appropriate standard scores (e.g., standard deviation, cumulative percentage, percentile equivalent, Z-score, T-score, R-score, standard nine-point method (stanine), stanine percentile, etc.), etc., or combinations thereof. In some embodiments, determining a level of significance includes determining a measure of uncertainty (e.g., a p-value).In certain embodiments, two or more datasets, relationships, and / or profiles can be analyzed and / or compared by leveraging multiple (e.g., two or more) statistical methods (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, bagging, neural networks, support vector machine models, random forests, classification tree models, k-nearest neighbors, logistic regression, and / or LOESS smoothing), and / or any suitable mathematical and / or statistical operations (e.g., referred to herein as operations).
[0167] In certain embodiments, the comparison of two or more lead density profiles includes determining and / or comparing measures of uncertainty for the two or more lead density profiles. Optionally, the lead density profiles and / or associated measures of uncertainty are compared to facilitate the interpretation of mathematical and / or statistical operations on the dataset and / or to present an outcome. Optionally, the lead density profile generated for a test subject is compared to the lead density profile generated for one or more references (e.g., reference sample, reference subject, etc.). In some embodiments, the outcome is presented by comparing the lead density profile from the test subject for a chromosome, portion, or segments thereof to the lead density profile from a reference, where the reference lead density profile is obtained from a set of reference subjects (e.g., a reference) known to not carry a mutation in the gene. In some embodiments, the outcome is presented by comparing the lead density profile from the test subject for a chromosome, portion, or segments thereof to the lead density profile from a reference, where the reference lead density profile is obtained from a set of reference subjects known to carry a mutation in a specific gene (e.g., aneuploidy, trisomy of a chromosome).
[0168] In certain embodiments, the read density profile of a subject is compared to a predetermined value indicative of the absence of a gene mutation, and in some cases, deviates from the predetermined value at one or more genomic positions (e.g., portions) corresponding to the genomic position where the gene mutation is located. For example, in a subject (e.g., a subject at risk of or suffering from a medical condition associated with a gene mutation), the read density profile is expected to be significantly different from a reference read density profile (e.g., a reference sequence, reference subject, reference set) for a selected portion if the subject contains the gene mutation in question. The read density profile of the subject is often substantially the same as the reference read density profile (e.g., a reference sequence, reference subject, reference set) for a selected portion if the subject does not contain the gene mutation in question. The read density profile is often compared to a predetermined threshold and / or threshold range (e.g., see FIG. 8). As used herein, the term "threshold" refers to any number calculated using a qualitative data set and used as a limit of diagnosis for a gene mutation (e.g., copy number variation, aneuploidy, chromosomal abnormality, etc.). In certain embodiments, the threshold is exceeded by the results obtained by the methods described herein, and the subject is diagnosed as having a gene mutation (e.g., trisomy). In some embodiments, the value or range of values of the threshold is often calculated by mathematically and / or statistically manipulating sequence read data (e.g., derived from a reference and / or subject). A predetermined threshold or range of thresholds indicating the presence or absence of a gene mutation can also vary while still presenting a useful outcome for determining the presence or absence of a gene mutation. In certain embodiments, a read density profile including normalized read density and / or normalized count number is generated to facilitate classification and / or presentation of the outcome. The outcome can be presented based on a plot of the read density profile including the normalized count number (e.g., using such a plot of the read density profile).
[0169] In some embodiments, the system includes a scoring module 46. The scoring module can receive, retrieve, and / or store a lead density profile (e.g., an adjusted, normalized lead density profile) from another suitable module (e.g., the profile generation module 26, the PCA statistics module 33, the partial weighting module 42, etc.). The scoring module can receive, retrieve, store, and / or compare two or more lead density profiles (e.g., a test profile, a reference profile, a training set, a test subject). The scoring module can often present a score (e.g., a plot, profile statistics, a comparison (e.g., the difference between two or more profiles), a Z-score, a measure of uncertainty, a decision region, a sample determination 50 (e.g., determination of the presence or absence of a gene mutation), and / or an outcome). The scoring module can present the score to an end user and / or another suitable module (e.g., a display, a printer, etc.). In some embodiments, the scoring module is R code as shown below and includes part, all, or one modification of the R code that includes an R function for calculating a chi-square statistic for a specific test (e.g., a large count of chromosome 21). The three parameters are x = the lead data of the sample (sample of part x) m = the median for the part y = the test vector (e.g., false for all parts except true for chromosome 21) are.
Number
[0170] Experimental conditions In certain embodiments, the principal component normalization process can be adjusted for biases related to experimental conditions. Data processing that takes experimental conditions into account is described, for example, in International Patent Application Publication No. WO2013 / 109981, the entire content of which is incorporated herein by reference in its entirety, including all documents, tables, formulas, and drawings.
[0171] In certain cases, samples can be affected by common experimental conditions. Samples processed at substantially the same time, or using substantially the same conditions and / or reagents, sometimes exhibit data variability (e.g., bias) induced by similar experimental conditions (e.g., common experimental conditions) when compared to other samples processed at different times and / or using different conditions and / or reagents and / or at the same time. There are often practical considerations that limit the number of samples that can be prepared, processed, and / or analyzed at any given time during an experimental procedure. In certain embodiments, the time frame for processing a sample from a starting material to generate an outcome can sometimes be days, weeks, or even months. High-throughput experiments that analyze a large number of samples can generate batch effects or experimental condition-induced data variability due to the time between isolation and final analysis. Experimental condition-induced data variability often includes any data variability that results from sample isolation, storage, preparation, and / or analysis. Non-limiting examples of experimental condition-induced variability include overrepresentation or underrepresentation of sequences; noisy data; false data points or outlier data points, reagent effects, personnel effects, flow cell-based variability and / or plate-based variability including laboratory condition effects, etc. Experimental condition-induced variability sometimes occurs in subpopulations of samples in a dataset (e.g., batch effects). A batch is often samples processed using substantially the same reagents, samples processed on the same sample preparation plate (e.g., sample preparation; e.g., a microwell plate used for nucleic acid isolation), samples developed for analysis on the same development plate (e.g., a microwell plate used to organize samples prior to loading on a flow cell), samples processed at substantially the same time, samples processed by the same personnel, and / or samples processed under substantially the same experimental conditions (e.g., temperature, CO2 level, ozone level, etc., or combinations thereof). Experimental condition batch effects sometimes affect samples that are analyzed on the same flow cell, prepared on the same reagent plate or microwell plate, and / or developed for analysis on the same reagent plate or microwell plate (e.g., preparing a nucleic acid library for sequencing).Sources of additional variability can include the quality of the isolated nucleic acid, the amount of isolated nucleic acid, the time from isolation of the nucleic acid until storage, the time during storage, the storage temperature, etc., and combinations thereof. The variability of data points within a batch (e.g., a subpopulation of samples in a dataset processed at the same time and / or using the same reagents and / or experimental conditions) can sometimes be greater than the variability of data points seen between batches. This data variability can sometimes include spurious data or outlier data that can, on a scale, render the interpretation of some or all of the other data in the dataset impossible. Portions or all of the dataset can be adjusted for experimental conditions using data processing steps described herein and known in the art; for example, normalization to the median absolute deviation calculated for all samples analyzed in a flow cell or processed in a microwell plate. Data processing that takes experimental conditions into account is described, for example, in International Patent Application Publication No. WO2013 / 109981, the entire contents of which are incorporated herein by reference, including all documents, tables, formulas, and drawings.
[0172] Detection of heterogeneity using comparisons In some embodiments, principal component normalization processing is used in conjunction with a method for determining the presence or absence of heterogeneity according to a comparison. Detection of heterogeneity using comparisons is described, for example, in International Patent Application Publication No. WO2014 / 116598, the entire contents of which are incorporated herein by reference, including all documents, tables, formulas, and drawings.
[0173] In this section, the comparison of ratios, or ratios, or ratio values, ploidy assessment, and ploidy assessment values are collectively referred to as "comparisons". In some embodiments, the presence or absence of aneuploidy in a subject is determined according to one or more comparisons. In some embodiments, the presence or absence of aneuploidy in a subject is determined according to one or more comparisons for three selected autosomes (e.g., when one or more of the three selected autosomes are test chromosomes). In some embodiments, the presence or absence of aneuploidy is determined according to one or more comparisons generated for a series of different chromosomes, euploid regions, aneuploid regions, or euploid regions and aneuploid regions. In some embodiments, the presence or absence of aneuploidy (e.g., aneuploidy in a fetus) in a subject is determined according to the comparisons obtained for the subject and euploid regions and / or aneuploid regions (e.g., euploid regions and aneuploid regions determined for a reference set). In certain embodiments, the presence or absence of aneuploidy is determined according to the relationship between the comparisons obtained for the subject and euploid regions and / or aneuploid regions. For example, in some embodiments, the presence or absence of aneuploidy is determined according to whether the comparison is within a euploid region or an aneuploid region, or how far the ploidy assessment value is from a euploid region or an aneuploid region. In some embodiments, the relationship is proximity or distance (e.g., mathematical difference and / or graph distance, e.g., distance between a point and a region). The relationship can be determined by methods known in the art or appropriate methods described herein, non-limiting examples of which include probability distribution, probability density function, cumulative distribution function, likelihood function, Bayesian model comparison, Bayesian factor, information criterion of deviation, chi-square test, Euclidean distance, spatial analysis, Mahalanobis distance, Manhattan distance, Chebyshev distance, Minkowski distance, Bregman divergence, Bhattacharyya distance, Hellinger distance, metric space, Canberra distance, convex hull (e.g., even-odd bending rule), etc., or combinations thereof.
[0174] In some embodiments, the absence of aneuploidy is determined according to comparisons and regions of euploidy. In some embodiments, the absence of aneuploidy is determined according to the relationship between the comparison and the region of euploidy. In some embodiments, a comparison within, in, or near the region of euploidy is a determination of euploid chromosomes (e.g., the absence of aneuploid chromosomes). In some embodiments, a comparison in or near the region of euploidy indicates that each chromosome for which the comparison was determined is euploid. For example, sometimes, a comparison generated according to the count numbers mapped to ChrA, ChrB, and ChrC is within the region of euploidy (e.g., the region of euploidy determined according to the count numbers mapped to ChrA, ChrB, and ChrC), and the absence of aneuploidy is determined. In some embodiments, when the absence of aneuploidy is determined according to a comparison, it indicates that each chromosome (e.g., each chromosome for which a ploidy evaluation value was derived) is euploid (e.g., euploid in the mother and / or the fetus).
[0175] In some embodiments, a comparison outside the region of aneuploidy is a determination of one or more euploid chromosomes. In some embodiments, a comparison outside the region of euploidy indicates that one or more chromosomes for which the comparison was determined are euploid. For example, sometimes, a comparison generated according to the count numbers mapped to ChrA, ChrB, and ChrC is outside the region of euploidy (e.g., the region of euploidy determined according to the count numbers mapped to ChrA, ChrB, and ChrC), and the absence of aneuploidy is determined. In some embodiments, a comparison outside the region of euploidy is used for comparison or evaluation and indicates that two of the three chromosomes for which the comparison was determined are euploid.
[0176] In some embodiments, the comparison is within the range of the aneuploid region and the one or more chromosomes for which the comparison is determined are euploid. For example, sometimes the comparison generated according to the counts mapped to ChrA, ChrB, and ChrC is within the range of the aneuploid region (e.g., the aneuploid region determined according to the counts mapped to ChrA, ChrB, and ChrC), and the absence of chromosomal aneuploidy is determined for two of the three chromosomes.
[0177] In some embodiments, the presence of chromosomal aneuploidy is determined according to the comparison and the euploid region. In certain embodiments, the presence of chromosomal aneuploidy is determined according to the relationship between the comparison and the euploid region. In some embodiments, a comparison outside the euploid region is a determination of an aneuploid chromosome (e.g., the presence of an aneuploid somatic chromosome). In some embodiments, a comparison outside the euploid region indicates that the one or more chromosomes for which the comparison is determined are aneuploid. For example, sometimes the comparison generated according to the counts mapped to ChrA, ChrB, and ChrC is outside the euploid region (e.g., the euploid region determined according to the counts mapped to ChrA, ChrB, and ChrC), and the presence of chromosomal aneuploidy is determined.
[0178] In some embodiments, a comparison within, in, or near an aneuploid region is a determination of an aneuploid chromosome (e.g., the presence of an aneuploid chromosome). In some embodiments, a comparison in or near an aneuploid region indicates that one or more chromosomes for which a ploidy assessment value has been determined are aneuploid. In some embodiments, a comparison in or near an aneuploid region indicates that 1, 2, 3, 4, and / or 5 chromosomes for which the comparison has been determined are aneuploid. In some embodiments, a comparison in or near an aneuploid region indicates that one of three chromosomes for which the comparison has been determined is aneuploid. For example, sometimes a comparison generated according to the counts mapped to ChrA, ChrB, and ChrC is within an aneuploid region (e.g., an aneuploid region determined according to the counts mapped to ChrA, ChrB, and ChrC), and one of the chromosomes is an aneuploid chromosome.
[0179] In some embodiments, a comparison near an aneuploid region is a determination of an aneuploid chromosome (e.g., the presence of an aneuploid chromosome). In some embodiments, a comparison near an aneuploid region indicates that one or more chromosomes for which the comparison has been determined are aneuploid. In some embodiments, a reference plot includes a defined euploid region and three defined aneuploid regions (e.g., aneuploid for Chr13, Chr18, or Chr21), and a determination of the presence of aneuploidy is made according to a comparison closest to one of the aneuploid regions. For example, a comparison closer to an aneuploid region for Chr21 than to another region (e.g., an aneuploid region for Chr13 or Chr18, or the euploid region) may indicate the presence of aneuploidy for Chr21.
[0180] In some embodiments, the comparison generated according to the counts mapped to Chr13, Chr18, and Chr21 is within the range of an aneuploid region (e.g., an aneuploid region determined according to the counts mapped to Chr13, Chr18, and Chr21), and one of the chromosomes is an aneuploid chromosome. In some embodiments, the comparison generated according to the counts mapped to Chr13, Chr18, and Chr21 is within the range of an aneuploid region (e.g., an aneuploid region determined according to the counts mapped to Chr13, Chr18, and Chr21), Chr18 and Chr21 are determined to be euploid, and Chr13 is determined to be aneuploid. In some embodiments, the comparison generated according to the counts mapped to Chr13, Chr18, and Chr21 is within the range of an aneuploid region (e.g., an aneuploid region determined according to the counts mapped to Chr13, Chr18, and Chr21), Chr13 and Chr21 are determined to be euploid, and Chr18 is determined to be aneuploid. In some embodiments, the comparison generated according to the counts mapped to Chr13, Chr18, and Chr21 is within the range of an aneuploid region (e.g., an aneuploid region determined according to the counts mapped to Chr13, Chr18, and Chr21), Chr18 and Chr13 are determined to be euploid, and Chr21 is determined to be aneuploid.
[0181] In some embodiments, the presence or absence of aneuploidy is determined according to a first comparison and a second comparison, where both comparisons are generated from reads of sequences mapped to the same set of two or more chromosomes. In some embodiments, the presence or absence of aneuploidy in a subject is determined according to the relationship (e.g., distance) between a first comparison generated for the subject and a second comparison generated for a second subject. In some embodiments, the second comparison is a series of comparisons (e.g., regions) generated for one or more subjects. In some embodiments, the presence or absence of aneuploidy in a subject is determined according to the relationship (e.g., distance) between a first comparison generated for the subject and a reference set of comparisons generated for one or more subjects. In some embodiments, the first comparison is a comparison for the subject and the second comparison is a comparison or series of comparisons that display one or more euploid fetuses. In some embodiments, the second comparison is a value or series of values (e.g., regions) expected for a euploid fetus. In some embodiments, the second comparison is a value or series of values generated for a subject (e.g., a pregnant female subject) for which it is known that the fetus is euploid for one or more of the chromosomes for which the comparison was generated. In some embodiments, the distance is determined according to an uncertainty value (e.g., standard deviation or MAD). In some embodiments, the distance between the first comparison and the second comparison (e.g., the second comparison that displays one or more euploid subjects) is 1, 2, 3, 4, 5, 6 times, or more than the associated uncertainty, and the first comparison is determined to be aneuploid. In some embodiments, the distance between the first comparison and the second comparison (e.g., the second comparison that displays one or more euploid subjects) is 3 times, or more than the associated uncertainty, and the first comparison is determined to display an aneuploid chromosome.
[0182] In some embodiments, the presence or absence of aneuploidy is determined according to a comparison generated according to the count numbers mapped to one or more specific chromosomes, and according to the regions of euploidy, aneuploidy, or the regions of euploidy and aneuploidy. In some embodiments, the presence or absence of aneuploidy is determined according to a comparison generated according to the sequence reads mapped to one or more specific chromosomes, and the sequence reads mapped to other chromosomes are not required for the determination. In some embodiments, the presence or absence of aneuploidy is determined according to a comparison generated according to the sequence reads mapped to 2, 3, 4, 5, or 6 different chromosomes, and the count numbers mapped to other chromosomes are not obtained or required for the determination. In some embodiments, the presence or absence of aneuploidy is determined according to a comparison generated according to 3 different chromosomes or segments thereof, and the determination is not based on chromosomes other than one of the 3 different chromosomes. For example, when ChrA, ChrB, and ChrC represent 3 different chromosomes or segments thereof, the presence or absence of aneuploidy is sometimes determined according to a comparison generated according to ChrA, ChrB, and ChrC, and the determination is not based on chromosomes other than ChrA, ChrB, or ChrC. In some embodiments, ChrA, ChrB, and ChrC represent Chr13, Chr21, and Chr18, respectively.
[0183] Sex chromosome karyotype In some embodiments, principal component normalization processing is used in combination with a method for determining a sex chromosome karyotype. The method for determining a sex chromosome karyotype is described, for example, in International Patent Application Publication No. WO2013 / 192562, the entire content of which is incorporated herein by reference, including all documents, tables, formulas, and drawings.
[0184] In some embodiments, the counts of reads of sequences mapped to one or more sex chromosomes (i.e., chromosome X, chromosome Y) are normalized. In some embodiments, the normalization includes principal component normalization. In some embodiments, the normalization determines an experimental bias for portions of the reference genome. In some embodiments, the experimental bias can be determined for multiple samples from a first fitted relationship (e.g., a fitted linear relationship, a fitted non-linear relationship) for each sample between the count of reads of sequences mapped to each portion of the reference genome and a mapping characteristic (e.g., GC content) for each portion. The slope of the fitted relationship (e.g., linear relationship) is generally determined by linear regression. In some embodiments, each experimental bias is represented by an experimental bias coefficient. The experimental bias coefficient is, for example, the slope of a linear relationship between (i) the count of reads of sequences mapped to each portion of the reference genome and (ii) the mapping characteristic for each portion. In some embodiments, the experimental bias can include an estimation of the curvature of the experimental bias.
[0185] In some embodiments, the method further includes calculating, for each portion of the genome, at the level of genomic bins (e.g., elevation, level) from a second fitted relationship (e.g., a fitted linear relationship, a fitted non-linear relationship) between the experimental bias and the count of reads of sequences mapped to each portion, where the slope of the relationship can be determined by linear regression. For example, if the first fitted relationship is linear and the second fitted relationship is linear, the level L i of each portion of the reference genome can be determined according to Equation α: L i =(m i -G i S)I -1 Equation α
[0186] where G i is the experimental bias, I is the intercept of the second fitted relationship, S is the slope of the second relationship, and m iis the measured count mapped to each part of the reference genome, and i is the sample.
[0187] In some embodiments, a secondary normalization process is applied at the level of one or more calculated genomic segments. In some embodiments, secondary normalization includes GC normalization and sometimes includes the use of the PERUN method. In some embodiments, secondary normalization includes principal component normalization.
[0188] Determination of fetal ploidy In some embodiments, the principal component normalization process is used in conjunction with a method for determining fetal ploidy. A method for determining fetal ploidy is described, for example, in U.S. Patent Application Publication No. 2013 / 0288244, the entire contents of which are hereby incorporated by reference in their entirety, including all documents, tables, formulas, and drawings.
[0189] Fetal ploidy can be determined in part from a measure of fetal fraction, and determination of fetal ploidy is used to determine the presence or absence of genetic mutations (e.g., aneuploidy, trisomy). Fetal ploidy can be determined in part from a measure of fetal fraction determined by any suitable method of determining fetal fraction, including the methods described herein. In some embodiments, this method requests the calculated reference count F i (sometimes also denoted as f i ) determined for a portion of the genome (i.e., bin i) for a plurality of samples, where the fetal ploidy for portion i of the genome is known to be euploid. In some embodiments, an uncertainty value (e.g., standard deviation, σ) is determined for the reference count f i for the reference count f iThe uncertainty value, the test sample count number, and / or the measured fetal fraction (F) are used to determine the ploidy of the fetus. In some embodiments, a reference count number (e.g., a reference count number based on an average value, mean, or median) is normalized by principal component normalization and / or other normalization, such as binwise normalization, normalization by GC content, linear least squares regression and non-linear least squares regression, LOESS, GC LOESS, LOWESS, PERUN, RM, GCRM, and / or combinations thereof. In some embodiments, when the reference count number is normalized by principal component normalization, the reference count number of a segment of the genome known to be aneuploid is equal to 1. In some embodiments, both the reference count number for a portion or segment of the genome (e.g., for a fetus known to be aneuploid) and the count number of the test sample are normalized by principal component normalization, and the reference count number is equal to 1. In some embodiments, when the reference count number is normalized by PERUN, the reference count number of a segment of the genome known to be aneuploid is equal to 1. In some embodiments, both the reference count number for a portion or segment of the genome (e.g., for a fetus known to be aneuploid) and the count number of the test sample are normalized by PERUN, and the reference count number is equal to 1. Similarly, in some embodiments, when the count number is normalized by the median of the reference count number (i.e., divided by the median of the reference count number), the reference count number of a portion or segment of the genome known to be aneuploid is equal to 1. For example, in some embodiments, both the reference count number for a portion or segment of the genome (e.g., for a fetus known to be aneuploid) and the count number of the test sample are normalized by the median of the reference count number, the normalized reference count number is equal to 1, and the test sample count number is normalized by the median of the reference count number (e.g., divided by the median of the reference count number). In some embodiments, both the reference count number for a portion or segment of the genome (e.g., for a fetus known to be aneuploid) and the count number of the test sample are normalized by principal component normalization, GCRM, GC, RM, or an appropriate method.In some embodiments, the reference count is the reference count by an average value, an average, or a median. The reference count is often the normalized count for a bin (e.g., at the level of a normalized genomic segment). In some embodiments, the reference count and the count for the test sample are the raw counts. In some embodiments, the reference count is determined from a count profile by an average value, an average, or a median. In some embodiments, the reference count is at the level of a calculated genomic segment. In some embodiments, the reference count of the reference sample and the count of the test sample (e.g., a patient sample, e.g., y. i ) are normalized by the same method or process.
[0190] Additional data processing and normalization In this specification, the reads of the mapped arrays that have been counted are referred to as raw data, because these data represent unprocessed counts (e.g., raw counts). In some embodiments, the data of the reads of the arrays in a dataset can be further processed (e.g., manipulated mathematically and / or statistically) and / or presented to facilitate obtaining an outcome. In certain embodiments, including larger datasets, preprocessing of the dataset may be useful to facilitate further analysis. Preprocessing of a dataset sometimes involves removing portions of the reference genome that are redundant and / or uninformative (e.g., portions of the reference genome with uninformative data, overlapping, mapped reads, portions where the median count is zero, arrays that are overrepresented or underrepresented). Without being limited by theory, processing and / or preprocessing of data can (i) remove noisy data, (ii) remove uninformative data, (iii) remove redundant data, (iv) reduce the complexity of a larger dataset, and / or (v) facilitate conversion of data from one form to one or more other forms. As used in this specification, the terms “preprocessing” and “processing,” when used with respect to data or a dataset, are collectively referred to as “processing.” Processing can make the data in a more suitable state for further analysis and, in some embodiments, can result in an outcome. In some embodiments, one or more or all of the processing methods (e.g., methods of normalization, partial filtering, mapping, validation, etc., or combinations thereof) are performed by a processor in conjunction with memory, a microprocessor, a computer, and / or a device controlled by a microprocessor.
[0191] As used herein, the term "noisy data" refers to (a) data that exhibits significant dispersion between data points when analyzed or plotted, (b) data having a significant standard deviation (e.g., greater than 3 standard deviations), (c) data having a significant standard error of the mean, etc., and combinations of the above. Noisy data sometimes arises due to the quantity and / or quality of the starting material (e.g., nucleic acid sample), and sometimes arises from part of the process for preparing or replicating the DNA used to obtain sequence reads. In certain embodiments, the noise results from certain sequences that are overrepresented when prepared using PCR-based methods. The methods described herein can reduce or eliminate the contribution of noisy data, and thus reduce the effect of noisy data on the resulting outcome.
[0192] As used herein, the terms "data that does not provide information", "portion of the reference genome that does not provide information", and "portion that does not provide information" refer to portions having numerical values that are significantly different from the value of a predetermined threshold, or numerical values that exist outside a predetermined limit range of values, or data derived therefrom. The terms "threshold" and "threshold value" refer, herein, to any number calculated using a qualified data set and serve as the limit for the diagnosis of gene mutations (e.g., copy number mutations, aneuploidy, microduplications, microdeletions, chromosomal abnormalities, etc.). In certain embodiments, the results obtained by the methods described herein exceed the threshold and the subject is diagnosed as having a gene mutation (e.g., trisomy 21). In some embodiments, the threshold value or range of values is often calculated by mathematically and / or statistically manipulating the data of the sequence reads (e.g., obtained from a reference and / or a subject), and in certain embodiments, the data of the sequence reads that are manipulated to obtain the threshold value or range of values are the data of the sequence reads (e.g., obtained from a reference and / or a subject). In some embodiments, a value of uncertainty is determined. The value of uncertainty is generally a measure of variance or error and may be any suitable measure of variance or error. In some embodiments, the value of uncertainty is the standard deviation, standard error, calculated variance, p-value, or mean absolute deviation (MAD). In some embodiments, the value of uncertainty can be calculated according to the methods described herein.
[0193] Any suitable procedure can be utilized to process the datasets described herein. Non-limiting examples of procedures suitable for use in processing datasets include filtering, normalizing, weighting, monitoring peak height, monitoring peak area, monitoring peak edges, determining area ratios, mathematically processing data, statistically processing data, applying statistical algorithms, analyzing using certain variables, analyzing using optimized variables, plotting data, identifying patterns or trends for further processing, and combinations of the above. In some embodiments, the datasets are processed based on various features (e.g., GC content, repeats, mapped reads, centromere regions, telomere regions, etc., and combinations thereof), and / or variables (e.g., fetal gender, maternal age, maternal ploidy, percent contribution of fetal nucleic acids, etc., or combinations thereof). In certain embodiments, processing the datasets according to the description herein can reduce the complexity and / or dimensionality of large and / or complex datasets. Non-limiting examples of complex datasets include data of sequence reads generated from one or more test subjects and multiple reference subjects with different age and ethnicity backgrounds. In some embodiments, the datasets can include thousands to millions of sequence reads for each test subject and / or reference subject.
[0194] In certain embodiments, data processing can be performed in any number of steps. For example, in some embodiments, data can be processed using only a single processing procedure, and in certain embodiments, one or more, five or more, ten or more, or twenty or more processing steps (e.g., one or more processing steps, two or more processing steps, three or more processing steps, four or more processing steps, five or more processing steps, six or more processing steps, seven or more processing steps, eight or more processing steps, nine or more processing steps, ten or more processing steps, eleven or more processing steps, twelve or more processing steps, thirteen or more processing steps, fourteen or more processing steps, fifteen or more processing steps, sixteen or more processing steps, seventeen or more processing steps, eighteen or more processing steps, nineteen or more processing steps, or twenty or more processing steps) can be used to process the data. In some embodiments, the processing steps can be the same steps repeated two or more times (e.g., filtering two or more times, normalizing two or more times), and in certain embodiments, the processing steps can be two or more different processing steps performed simultaneously or sequentially (e.g., filtering and normalizing; normalizing and monitoring peak height and edges; filtering, normalizing, normalizing against a reference, statistically manipulating to determine a p-value, etc.). In some embodiments, any suitable number and / or combination of the same or different processing steps can be utilized to facilitate processing the data of the array reads to obtain an outcome. In certain embodiments, the complexity and / or dimensionality of a dataset can be reduced by processing the dataset according to the criteria described herein.
[0195] In some embodiments, one or more processing steps can include one or more filtering steps. As used herein, the term "filtering" refers to removing portions of a partial or reference genome from consideration. This can include, but is not limited to, duplicate data (e.g., mapped reads that are duplicated or overlapping), data with no information (e.g., portions of the reference genome where the median count is zero), portions of the reference genome with overrepresented or underrepresented arrays, noisy data, etc., or any suitable criteria, including combinations thereof, to select and remove portions of the reference genome. Filtering processes often involve removing one or more portions of the reference genome from consideration and subtracting the counts in the one or more portions of the reference genome selected for removal from the counts counted or summed for the reference genome, one or more chromosomes, or portions of the genome under consideration. In some embodiments, portions of the reference genome can be removed sequentially (e.g., one at a time to allow evaluation of the effect of removing each individual portion), and in certain embodiments, all portions of the reference genome marked for removal can be removed simultaneously. In some embodiments, portions of the reference genome characterized by a variance above or below a certain level are removed, which is sometimes referred to herein as filtering the "noisy" portions of the reference genome. In certain embodiments, the filtering process includes obtaining from the data set data points that deviate from the average profile level of a portion, chromosome, or chromosomal segment by a predetermined multiple of the profile variance, and in certain embodiments, the filtering process includes removing from the data set data points that do not deviate from the average profile level of a portion, chromosome, or chromosomal segment by a predetermined multiple of the profile variance. In some embodiments, the filtering process is utilized to reduce the number of candidate portions of the reference genome for analyzing the presence or absence of gene mutations.By reducing the number of candidate portions of a reference genome to analyze for the presence or absence of gene mutations (e.g., microdeletions, microduplications), often the complexity and / or dimensionality of a dataset is reduced, and sometimes the speed of searching for and / or identifying gene mutations and / or gene abnormalities is increased by two digits or more.
[0196] In some embodiments, one or more processing steps can include one or more normalization steps. Normalization can be performed by suitable methods described herein or known in the art. In certain embodiments, normalization includes adjusting values measured on different scales to a conceptually common scale. In certain embodiments, normalization includes advanced mathematical adjustments to bring the probability distributions of the adjusted values into alignment. In some embodiments, normalization includes fitting the distribution to a normal distribution. In certain embodiments, normalization includes mathematical adjustments that enable comparison of corresponding values normalized for different datasets in a way that eliminates the effects of certain overall influences (e.g., errors and anomalies). In certain embodiments, normalization includes scaling. Normalization sometimes includes division of one or more datasets by a given variable or expression. Normalization sometimes includes subtraction of one or more datasets by a given variable or expression. Non-limiting examples of normalization methods include normalization by part, normalization by GC content, normalization of count number medians (bin count number median, partial count number median), linear and non-linear least squares regression, LOESS, GC LOESS, LOWESS (local weighted scatterplot smoothing), PERUN, ChAI, principal component normalization, repeat masking (RM), GC-normalized repeat masking (GCRM), cQn, and / or combinations thereof. In some embodiments, determination of the presence or absence of a gene mutation (e.g., aneuploidy, microduplication, microdeletion) utilizes a normalization method (e.g., normalization by part, normalization by GC content, normalization of count number medians (bin count number median, partial count number median), linear and non-linear least squares regression, LOESS, GC LOESS, LOWESS (local weighted scatterplot smoothing), PERUN, ChAI, principal component normalization, repeat masking (RM), GC-normalized repeat masking (GCRM), cQn, normalization methods known in the art, and / or combinations thereof).In some embodiments, determination of the presence or absence of a gene mutation (e.g., aneuploidy, microduplication, microdeletion) utilizes one or more of LOESS, normalization of the median count (bin count median, partial count median), and principal component normalization. In some embodiments, determination of the presence or absence of a gene mutation utilizes LOESS and then utilizes normalization of the median count (bin count median, partial count median). In some embodiments, determination of the presence or absence of a gene mutation utilizes LOESS, then utilizes normalization of the median count (bin count median, partial count median), and then utilizes principal component normalization.
[0197] Any suitable number of normalizations can be used. In some embodiments, the dataset can be normalized once or multiple times, 5 or more times, 10 or more times, or even 20 or more times. The dataset can be normalized against values (e.g., normalized values) that display any suitable feature or variable (e.g., sample data, reference data, or both). Non-limiting examples of types of data normalization that can be used include normalizing the raw count data for one or more selected test portions or reference portions against the total number of counts mapped against the entire chromosome or genome to which the selected portion or segment is mapped; normalizing the raw count data for one or more selected portions against the median of the reference counts for one or more portions or chromosomes to which the selected portion or segment is mapped; normalizing the raw count data against pre-normalized data or their derived values; and normalizing the pre-normalized data against one or more other predetermined normalization variables. Normalization of the dataset sometimes has the effect of isolating statistical errors, depending on the feature or characteristic selected as the predetermined normalization variable. Also, normalization of the dataset sometimes enables comparison of the features of data having different scales by giving the data a common scale (e.g., a predetermined normalization variable). In some embodiments, one or more normalizations against statistically derived values can be utilized to minimize differences in the data and reduce the significance of outlier data. Normalizing a portion or a portion of the reference genome with respect to the normalized value is sometimes referred to as "normalization with respect to the portion".
[0198] In certain embodiments, the processing step including normalization includes normalizing with respect to a stationary window, and in some embodiments, the processing step including normalization includes normalizing with respect to a moving window or a sliding window. As used herein, the term "window" refers to one or more portions selected for analysis and is sometimes used as a reference for comparison (e.g., used for normalization and / or other mathematical or statistical operations). The term "normalizing with respect to a stationary window" as used herein refers to the process of normalization using one or more portions selected to compare a test dataset and a reference dataset. In some embodiments, a profile is generated using the selected portions. A stationary window generally includes a predetermined series of portions that do not change during an operation and / or analysis. The terms "normalizing with respect to a moving window" and "normalizing with respect to a sliding window" as used herein refer to normalization performed on portions limited to the genomic region of a selected test portion (e.g., adjacent portions or segments around the immediate vicinity of a gene, etc.), where one or more selected test portions are normalized with respect to portions around the immediate vicinity of the selected test portion. In certain embodiments, a profile is generated using the selected portions. Normalization of a sliding window or a moving window often includes repeatedly moving or sliding towards adjacent test portions and normalizing a newly selected test portion with respect to portions around the immediate vicinity of the newly selected test portion or adjacent to the newly selected test portion, where adjacent windows have one or more common portions. In certain embodiments, a plurality of selected test portions and / or chromosomes can be analyzed by a sliding window process.
[0199] In some embodiments, one or more values can be generated by normalizing against a sliding window or a moving window, where each value represents the result of normalization against a different set of reference portions selected from different regions of the genome (e.g., chromosomes). In certain embodiments, the one or more generated values are cumulative sums (e.g., numerical estimates of the integral of a count profile normalized across selected portions, domains (e.g., part of a chromosome) or chromosomes). The values generated by processing the sliding window or the moving window can be used to generate a profile and facilitate reaching an outcome. In some embodiments, the cumulative sum of one or more portions can be shown as a function of the genomic position. Sometimes, the analysis of the moving window or the sliding window is used to analyze the genome for the presence or absence of microdeletions and / or microinsertions. In certain embodiments, the presence or absence of regions of gene mutations (e.g., microdeletions, microduplications) is identified using the indication of the cumulative sum of one or more portions. In some embodiments, the analysis of the moving window or the sliding window is used to identify genomic regions containing microdeletions, and in certain embodiments, the analysis of the moving window or the sliding window is used to identify genomic regions containing microduplications.
[0200] Specific examples of normalization processes that can be utilized, such as LOESS, PERUN, ChAI, and principal component normalization methods, etc., are described in more detail below.
[0201] In some embodiments, the processing step includes weighting. As used herein, the terms "weighted," "weighting," or "weighting function," or their grammatical derivatives or equivalents, refer to mathematical operations on some or all of a dataset that may be used to vary the influence of a particular dataset feature or variable relative to other dataset features or variables (e.g., increasing or decreasing the significance and / or contribution of data contained in one or more portions or parts of a reference genome relative to the quality or utility of data in selected one or more portions of the reference genome). In some embodiments, a weighting function can be used to increase the influence of data having a relatively small measurement variance and / or decrease the influence of data having a relatively large measurement variance. For example, the "weight is reduced" for portions of a reference genome having underrepresented or low-quality array data to minimize the influence on the dataset, while the "weight is increased" for selected portions of the reference genome to increase the influence on the dataset. A non-limiting example of a weighting function is [1 / (standard deviation) 2 . The weighting step is sometimes performed in a manner substantially similar to the normalization step. In some embodiments, the dataset is divided by a predetermined variable (e.g., a weighting variable). Often, a predetermined variable (e.g., a minimization objective function, Phi) is selected to apply different weightings to different parts of the dataset (e.g., increasing the influence of a particular data type while decreasing the influence of other data types).
[0202] In certain embodiments, the processing steps can include one or more mathematical and / or statistical operations. Any suitable mathematical and / or statistical operations can be used, either alone or in combination, to analyze and / or manipulate the datasets described herein. Any suitable number of mathematical and / or statistical operations can be used. In some embodiments, the dataset can be mathematically and / or statistically manipulated one or more times, five or more times, ten or more times, or twenty or more times. Non-limiting examples of mathematical and statistical operations that can be used include addition, subtraction, multiplication, division, algebraic functions, least squares estimators, curve fitting, differential equations, rational polynomials, double polynomials, orthogonal polynomials, z-scores, p-values, chi values, phi values, peak level analysis, determination of peak edge locations, calculation of peak area ratios, analysis of chromosomal level medians, calculation of mean absolute deviation, sum of squared residuals, mean, standard deviation, standard error, etc., or combinations thereof. The mathematical and / or statistical operations can be performed on all or part of the array read data or their processed products. Non-limiting examples of variables or features of the dataset that can be statistically manipulated include unprocessed count numbers, filtered count numbers, normalized count numbers, peak height, peak width, peak area, peak edge, lateral tolerance, P-value, level median, average level, distribution of count numbers within genomic regions, relative representation of nucleic acid species, etc., or combinations thereof.
[0203] In some embodiments, the processing steps can include the use of one or more statistical algorithms. Any suitable statistical algorithm can be used, either alone or in combination, to analyze and / or manipulate the datasets described herein. Any suitable number of statistical algorithms can be used. In some embodiments, one or more, five or more, ten or more, or twenty or more statistical algorithms can be used to analyze the dataset. Non-limiting examples of statistical algorithms suitable for use with the methods described herein include decision trees, hypothesis testing, multiple comparisons, omnibus tests, Bonferroni-Fisher tests, bootstrap methods, Fisher's method for combining independent significance tests, null hypotheses, type I errors, type II errors, exact tests, one-sample Z-tests, two-sample Z-tests, one-sample t-tests, paired t-tests, two-sample pooled t-tests with equal variances, two-sample unpooled t-tests with unequal variances, one-proportion z-tests, two-proportion z-tests pooled, two-proportion z-tests unpooled, one-sample chi-square tests, two-sample F-tests for equality of variances, confidence intervals, credible intervals, significance, meta-analysis, simple linear regression, robust linear regression, etc., or combinations thereof. Non-limiting examples of variables or features of the dataset that can be analyzed using statistical algorithms include raw count numbers, filtered count numbers, normalized count numbers, peak height, peak width, peak edges, lateral tolerance, P-values, median levels, average levels, distribution of count numbers within genomic regions, relative representation of nucleic acid species, etc., or combinations thereof.
[0204] In certain embodiments, a dataset can be analyzed by utilizing multiple (e.g., two or more) statistical algorithms (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, bagging, neural networks, support vector machine models, random forests, classification tree models, k-nearest neighbor method, logistic regression, and / or loss smoothing), and / or mathematical and / or statistical operations (e.g., referred to herein as manipulations). In some embodiments, the use of multiple manipulations can generate an N-dimensional space that can be used to yield an outcome. In certain embodiments, the complexity and / or dimensionality of a dataset can be reduced by analyzing the dataset by utilizing multiple manipulations. For example, by using multiple manipulations on a reference dataset, an N-dimensional space (e.g., a probability plot) can be generated that can be used to indicate the presence or absence of a gene mutation depending on the genetic status of a reference sample (e.g., positive or negative for a mutation in a selected gene). Using the analysis of test samples that use a substantially similar set of manipulations, an N-dimensional point can be generated for each of the test samples. The complexity and / or dimensionality of the dataset being tested can sometimes be simplified to a single value or an N-dimensional point that can be readily compared to the N-dimensional space generated from the reference data. Test sample data that belongs to the N-dimensional space where the reference data exists indicates a genetic status that is substantially similar to the genetic status of the reference gene. Test sample data that exists outside of the N-dimensional space where the reference data exists indicates a genetic status that is not substantially similar to the genetic status of the reference gene. In some embodiments, the reference is a euploid and otherwise has no gene mutations or medical conditions.
[0205] In some embodiments, after the dataset is counted, optionally filtered and normalized, the processed datasets can be further manipulated by one or more procedures that filter and / or normalize. In certain embodiments, profiles can be generated using datasets that are further manipulated by one or more procedures that filter and / or normalize. In some embodiments, sometimes, the complexity and / or dimensionality of the dataset can be reduced by one or more procedures that filter and / or normalize. Outcomes can be provided based on the dataset with reduced complexity and / or dimensionality.
[0206] In some embodiments, portions can be filtered according to a measure of error (e.g., standard deviation, standard error, calculated variance, p-value, mean absolute error (MAE), mean absolute deviation, and / or mean absolute deviation (MAD)). In certain embodiments, the measure of error refers to the variability of the count number. In some embodiments, portions are filtered according to the variability of the count number. In certain embodiments, the variability of the count number is a measure of error determined for the count number mapped to a portion (i.e., a portion) of a reference genome for a plurality of samples (e.g., a plurality of subjects, e.g., more than 50 subjects / animals, more than 100 subjects / animals, more than 500 subjects / animals, more than 1000 subjects / animals, more than 5000 subjects / animals, or more than 10,000 subjects / animals). In some embodiments, portions having a variability of the count number that exceeds a predetermined upper range are filtered out (e.g., excluded from consideration). In some embodiments, the predetermined upper range is a MAD value equal to or greater than about 50, equal to or greater than about 52, equal to or greater than about 54, equal to or greater than about 56, equal to or greater than about 58, equal to or greater than about 60, equal to or greater than about 62, equal to or greater than about 64, equal to or greater than about 66, equal to or greater than about 68, equal to or greater than about 70, equal to or greater than about 72, equal to or greater than about 74, or equal to or greater than about 76. In some embodiments, portions having a variability of the count number that is below a predetermined lower range are filtered out (e.g., excluded from consideration). In some embodiments, the predetermined lower range is a MAD value equal to or less than about 40, equal to or less than about 35, equal to or less than about 30, equal to or less than about 25, equal to or less than about 20, equal to or less than about 15, equal to or less than about 10, equal to or less than about 5, equal to or less than about 1, or equal to or less than about 0.In some embodiments, a portion having variability in the count number outside a predetermined range is filtered out (e.g., excluded from consideration). In some embodiments, the predetermined range is a MAD value from greater than zero to less than about 76, less than about 74, less than about 73, less than about 72, less than about 71, less than about 70, less than about 69, less than about 68, less than about 67, less than about 66, less than about 65, less than about 64, less than about 62, less than about 60, less than about 58, less than about 56, less than about 54, less than about 52, or less than about 50. In some embodiments, the predetermined range is a MAD value from greater than zero to less than about 67.7. In some embodiments, a portion having variability in the count number within the predetermined range is selected (e.g., used to determine the presence or absence of a gene mutation).
[0207] In some embodiments, the variability in the count of the portions exhibits a distribution (e.g., a normal distribution). In some embodiments, the portions are selected within the quartiles of the distribution. In some embodiments, portions equal to or less than about 99.9%, about 99.8%, about 99.7%, about 99.6%, about 99.5%, about 99.4%, about 99.3%, about 99.2%, about 99.1%, about 99.0%, about 98.9%, about 98.8%, about 98.7%, about 98.6%, about 98.5%, about 98.4%, about 98.3%, about 98.2%, about 98.1%, about 98.0%, about 97%, about 96%, about 95%, about 94%, about 93%, about 92%, about 91%, about 90%, about 85%, about 80%, or about 75% of the distribution of the variability in the count are selected. In some embodiments, portions within the 99% quartile of the distribution of the variability in the count are selected. In some embodiments, within the 99% quartile, portions with MAD>0 and portions with MAD<67.725 are selected, and as a result, a series of stable portions of the reference genome are identified.
[0208] Non-limiting examples of filtering parts with respect to PERUN are shown, for example, in this specification and International Patent Application No. PCT / US12 / 59123 (WO2013 / 052913), the entire content of which, including all documents, tables, formulas and drawings, is incorporated herein by reference. Parts can be filtered based on a measure of error or based on a part of the measure of error. In certain embodiments, a measure of error that includes the absolute value of a deviation, such as an R factor, can be used to remove parts or weight parts. The R factor is defined, in some embodiments, as the result of dividing the sum of the absolute deviations of the values of the counts predicted from the actual measurements by the values of the counts predicted from the actual measurements. A measure of error that includes the absolute value of the deviation can be used, but an appropriate measure of error can also be used instead. In certain embodiments, a measure of error that does not include the absolute value of the deviation, for example, a variance based on the square, can be used. In some embodiments, parts are filtered or weighted according to a measure of mapability (e.g., a mapability score). Sometimes, the part is filtered or weighted according to a relatively low number of array reads mapped to the part (e.g., 0, 1, 2, 3, 4, 5 reads mapped to the part). Parts can be filtered or weighted according to the type of analysis being performed. For example, in the case of aneuploidy analysis of chromosomes 13, 18 and / or 21, the sex chromosomes can be filtered and only the autosomes or a subset of the autosomes can be analyzed.
[0209] In certain embodiments, the following filtering process can be utilized. Select the same series of portions (e.g., portions of the reference genome) within a given chromosome (e.g., chromosome 21), and compare the number of reads between the affected sample and the non - affected sample. The gaps relate the trisomy 21 sample to the euploid sample, and include a series of portions that cover most of chromosome 21. These series of portions are the same between the euploid sample and the T21 sample. Since the portions can be defined, the distinction between a series of portions and a single segment is not very important. Compare the same genomic regions in different patients. This process can be utilized for the analysis of trisomy, for example, for T13 or T18 in addition to or instead of T21.
[0210] In some embodiments, after the dataset is counted, optionally filtered and normalized, these processed datasets can be manipulated by weighting. In certain embodiments, one or more portions can be selected and weighted to reduce the impact of data (e.g., noisy data, data that gives no information) contained within the selected portions. In some embodiments, one or more portions can be selected and weighted to enhance or increase the impact of data (e.g., data with a small variance measured) contained within the selected portions. In some embodiments, a single weighting function that decreases the impact of data with a large variance and increases the impact of data with a small variance is utilized to weight the dataset. Sometimes, a weighting function is used to decrease the impact of data with a large variance and increase the impact of data with a small variance (e.g., [1 / (standard deviation) 2 ). In some embodiments, further manipulation by weighting is performed to generate a plot of the profile of the processed data to facilitate classification and / or the provision of an outcome. An outcome can be brought about based on the plot of the profile of the weighted data.
[0211] Filtering or weighting of portions can be done at one or more appropriate points in the analysis. For example, the portions can be filtered or weighted before or after mapping the reads of the array to portions of the reference genome. In some embodiments, the portions can be filtered or weighted before or after determining the bias of the experiments for individual genomic portions. In certain embodiments, the portions can be filtered or weighted before or after calculating the levels of genomic bins.
[0212] In some embodiments, after the dataset has been counted, optionally filtered, normalized, and optionally weighted, these processed datasets can be manipulated by one or more mathematical and / or statistical (e.g., by statistical functions or statistical algorithms) operations. In certain embodiments, the processed datasets can be further manipulated by calculating Z - scores for one or more selected portions, chromosomes, or portions of chromosomes. In some embodiments, the processed datasets can be further manipulated by calculating P - values. In certain embodiments, the mathematical and / or statistical operations include one or more assumptions regarding ploidy and / or fetal fraction. In some embodiments, one or more statistical and / or mathematical operations are performed to further manipulate the processed data to generate a plot of the profile of the data to facilitate classification and / or providing an outcome. An outcome can be provided based on the plot of the profile of the statistically and / or mathematically manipulated data. The outcome provided based on the plot of the profile of the statistically and / or mathematically manipulated data often includes one or more assumptions regarding ploidy and / or fetal fraction.
[0213] In certain embodiments, after the data set is counted, optionally filtered and normalized, a plurality of operations are performed on the processed data set to generate an N-dimensional space and / or N-dimensional points. An outcome can be provided based on a plot of the profile of the data set analyzed in N-dimensions.
[0214] In some embodiments, as part of and / or subsequent to the processing and / or manipulation of the data set, one or more of peak level analysis, peak width analysis, location of peak edges analysis, peak lateral tolerance, etc., their derivatives, or combinations of the above are utilized to process the data set. In some embodiments, a plot of the profile of the data processed using one or more of peak level analysis, peak width analysis, location of peak edges analysis, peak lateral tolerance, etc., their derivatives, or combinations of the above is generated to facilitate classification and / or provision of an outcome. An outcome can be provided based on a plot of the profile of the data processed using one or more of peak level analysis, peak width analysis, location of peak edges analysis, peak lateral tolerance, etc., their derivatives, or combinations of the above.
[0215] In some embodiments, one or more reference samples that substantially do not contain mutations of the gene in question can be used to obtain a reference median count profile, which can be a predetermined value indicating the absence of gene mutations. Often, if the test subject carries a gene mutation, the profile will deviate from the predetermined value in the region corresponding to the genomic location where the gene mutation is located in the test subject. In a test subject at risk of or suffering from a medical condition associated with a gene mutation, the numerical values for the selected portion or segment are expected to be significantly different from the predetermined values for the genomic location when not suffering from the condition. In certain embodiments, one or more reference samples known to carry mutations of the gene in question can be used to obtain a reference median count profile, which can be a predetermined value indicating the presence of gene mutations. Often, the profile will deviate from the predetermined value in the region corresponding to the genomic location where the test subject does not carry the gene mutation. In a test subject without risk of or not suffering from a medical condition associated with a gene mutation, the numerical values for the selected portion or segment are expected to be significantly different from the predetermined values for the genomic location when suffering from the condition.
[0216] In some embodiments, the analysis and processing of data can include the use of one or more assumptions. An appropriate number or type of assumptions can be utilized to analyze or process a dataset. Non-limiting examples of assumptions that can be used for data processing and / or analysis include ploidy of a population, fetal contribution, prevalence of a particular sequence in a reference population, ethnic background, prevalence of a selected medical condition in related families by blood, parallelism between profiles of raw count numbers obtained from different patients and / or runs after GC normalization repeat masking (e.g., GCRM), identical matches (e.g., at the position of the same base) that imply unnatural results of PCR, assumptions specific to a fetal quantification assay (e.g., FQA), assumptions regarding twins (e.g., if only one of the twins is affected, the effective fetal fraction is only 50% of the measured total fetal fraction (similarly for triplets, quadruplets, etc.)), cell-free fetal DNA (e.g., cfDNA) that uniformly covers the entire genome, and combinations thereof.
[0217] Based on a normalized count number profile, in cases where it is not possible at the desired level of confidence (e.g., a confidence level of 95% or more) to predict the presence or absence of a genetic mutation outcome by the quality and / or depth of reads of the mapped sequences, one or more additional mathematical operation algorithms and / or statistical prediction algorithms can be utilized to generate additional numerical values useful for data analysis and / or outcome provision. The term "normalized count number profile" as used herein refers to a profile generated using normalized count numbers. Examples of normalized count numbers and methods that can be used to generate normalized count number profiles are described herein. As described above, the reads of the mapped sequences that have been counted can be normalized with respect to the count numbers of the test sample or the reference sample. In some embodiments, the normalized count number profile can be plotted and shown.
[0218] LOESS normalization LOESS is a regression modeling method known in the art, which is a regression modeling method that combines a multiple regression model within a k-nearest neighbor-based metamodel. LOESS is sometimes referred to as locally weighted polynomial regression. In some embodiments, in GC LOESS, the LOESS model is applied to the relationship between the fragment count (e.g., array reads, array counts) and the GC composition for a reference genomic portion. Plotting a smooth curve through a set of data points, plotting using LOESS is sometimes called a LOESS curve, especially when each smoothed value is given by weighted quadratic least squares regression over an interval of the values of the scatter plot reference variable on the y-axis. For each point in the dataset, the LOESS method fits a low-order polynomial to a subset of the data where the explanatory variable values are near the point whose response is to be estimated. The polynomial is fit using a weighted least squares method that gives large weights to points near the point whose response is to be estimated and small weights to distant points. The regression function value for the point is then obtained by evaluating the value of the local polynomial using the explanatory variable value for that data point. The LOESS fit is considered complete, in some cases, after calculating the regression function value for each of the data points. Many of the details of this method, such as the degree and weights of the polynomial model, are adaptable.
[0219] PERUN normalization As used herein, a normalization method for reducing errors associated with nucleic acid metrics is referred to as PERUN (parameterized error removal and unbiased normalization), which is described in International Patent Application Publication No. WO2013 / 052913, the entire contents of which, including this specification and all of its text, tables, formulas, and drawings, are incorporated herein by reference. The PERUN method can be applied to various nucleic acid metrics (e.g., nucleic acid sequence reads) for the purpose of reducing the effect of errors that confound predictions based on such metrics.
[0220] In certain embodiments, the PERUN method calculates the level of genomic segments for a reference genomic segment from (a) the count of sequence reads mapped to the reference genomic segment for a test sample, (b) the experimental bias (e.g., GC bias) for the test sample, and (c) one or more fitted parameters (e.g., an estimate of fit) for a fitted relationship between (i) the experimental bias for the reference genomic segment to which the sequence reads are mapped and (ii) the count of sequence reads mapped to the segment. The experimental bias for each reference genomic segment is a fitted relationship for each sample over a plurality of samples, which can be determined according to the relationship between (i) the count of sequence reads mapped to each of the reference genomic segments and (ii) the mapping characteristics for each of the reference genomic segments. The fitted relationship for each sample can be assembled in three dimensions for a plurality of samples. In certain embodiments, the assembly can be sorted according to the experimental bias, but the PERUN method can also be performed without sorting the assembly according to the experimental bias. The fitted relationship for each sample and the fitted relationship for each part of the reference genome can be independently fitted to a linear or non-linear function by appropriate fitting processes known in the art.
[0221] Hybrid normalization of regression In some embodiments, a hybrid normalization method is used. In some embodiments, the hybrid normalization method reduces bias (e.g., GC bias). In some embodiments, hybrid normalization includes (i) an analysis of the relationship between two variables (e.g., count number and GC content), and (ii) the selection and application of a normalization method according to the analysis. In certain embodiments, hybrid normalization includes (i) regression (e.g., regression analysis), and (ii) the selection and application of a normalization method according to the regression. In some embodiments, the count numbers obtained for a first sample (e.g., a first sample set) are normalized in a different way than the count numbers obtained from another sample (e.g., a second sample set). In some embodiments, the count numbers obtained for a first sample (e.g., a first sample set) are normalized by a first normalization method, and the count numbers obtained from a second sample (e.g., a second sample set) are normalized by a second normalization method. For example, in certain embodiments, the first normalization method includes the use of linear regression, and the second normalization method includes the use of non-linear regression (e.g., LOESS, GC-LOESS, LOWESS regression, LOESS smoothing).
[0222] In some embodiments, a hybrid normalization method is used to normalize reads of sequences mapped to a portion of a genome or chromosome (e.g., count numbers, mapped count numbers, mapped reads). In certain embodiments, raw count numbers are normalized, and in some embodiments, adjusted, weighted, filtered, or already-normalized count numbers are normalized by the hybrid normalization method. In certain embodiments, the genomic bin level or Z-score is normalized. In some embodiments, count numbers mapped to a selected genomic portion or chromosome are normalized by the hybrid normalization method. Count numbers are a suitable measure of reads of sequences mapped to a portion of a genome, non-limiting examples of which include raw count numbers (e.g., unprocessed count numbers), normalized count numbers (e.g., normalized by PERUN, ChAI, principal component normalization, or a suitable method), sub-levels (e.g., mean level, average level, median level, etc.), Z-scores, etc., or measures including combinations thereof. Count numbers may be raw count numbers derived from one or more samples (e.g., test samples, samples from pregnant females), or processed count numbers. In some embodiments, count numbers are obtained from one or more samples obtained from one or more subjects.
[0223] In some embodiments, a normalization method (e.g., type of normalization method) is selected according to regression (e.g., regression analysis) and / or a correlation coefficient. Regression analysis refers to a statistical technique for estimating the relationship between variables (e.g., count numbers and GC content). In some embodiments, regression is generated according to count numbers and measures of GC content for each of multiple portions of a reference genome. Suitable measures of GC content, non-limiting examples of which include measures of guanine content, cytosine content, adenine content, thymine content, purine (GC) content, or pyrimidine (AT or ATU) content, melting temperature (T m)(For example, denaturation temperature, annealing temperature, hybridization temperature), a measure of free energy, etc., or a measure including a combination thereof can be used. Measures of guanine (G) content, cytosine (C) content, adenine (A) content, thymine (T) content, purine (GC) content, or pyrimidine (AT or ATU) content can be expressed as ratios or percentages. In some embodiments, any suitable ratio or percentage, non-limiting examples of which are GC / AT, GC / total nucleotides, GC / A, GC / T, AT / total nucleotides, AT / GC, AT / G, AT / C, G / A, C / A, G / T, G / A, G / AT, C / T, etc., or ratios or percentages including combinations thereof are used. In some embodiments, the measure of GC content is a ratio or percentage of the GC content to the total nucleotide content. In some embodiments, the measure of GC content is a ratio or percentage of the GC content to the total nucleotide content for reads of the sequence mapped to a portion of the reference genome. In certain embodiments, the GC content is determined according to and / or from reads of the sequence mapped to each portion of the reference genome, and the reads of the sequence are obtained from a sample (e.g., a sample obtained from a pregnant female). In some embodiments, the measure of GC content is not determined according to and / or from reads of the sequence. In certain embodiments, the measure of GC content is determined for one or more samples obtained from one or more subjects.
[0224] In some embodiments, generating a regression includes generating a regression analysis or a correlation analysis. Non-limiting examples thereof include regression analysis (e.g., linear regression analysis), analysis of goodness of fit, Pearson correlation analysis, rank correlation, proportion of unexplained variance, efficiency analysis by the NS (Nash-Sutcliffe) model, confirmation of the validity of a regression model, PRL (proportional reduction in loss), root mean square deviation, etc., or a combination thereof, and appropriate regression can be used. In some embodiments, a regression line is generated. In certain embodiments, generating a regression includes generating a linear regression. In certain embodiments, generating a regression includes generating a non-linear regression (e.g., LOESS regression, LOWESS regression).
[0225] In some embodiments, regression is used to determine, for example, the presence or absence of a correlation (e.g., linear correlation) between the count number of GC content and a scale. In some embodiments, a regression (e.g., linear regression) is generated and a correlation coefficient is determined. In some embodiments, non-limiting examples thereof include determining an appropriate correlation coefficient including a coefficient of determination, R 2 value, Pearson correlation coefficient, etc.
[0226] In some embodiments, goodness of fit is determined for regression (e.g., regression analysis, linear regression). Goodness of fit is determined, in some cases, by visual analysis or mathematical analysis. The evaluation includes, in some cases, determining whether the goodness of fit is greater for non-linear regression or for linear regression. In some embodiments, the correlation coefficient is a measure of goodness of fit. In some embodiments, the evaluation of goodness of fit for regression is determined according to the correlation coefficient and / or a cut-off value of the correlation coefficient. In some embodiments, the evaluation of goodness of fit includes a comparison of the correlation coefficient with the cut-off value of the correlation coefficient. In some embodiments, the evaluation of goodness of fit for regression refers to linear regression. For example, in a particular embodiment, the goodness of fit is greater for linear regression than for non-linear regression, and the evaluation of goodness of fit refers to linear regression. In some embodiments, the evaluation refers to linear regression and linear regression is used to normalize the count number. In some embodiments, the evaluation of goodness of fit for regression refers to non-linear regression. For example, in a particular embodiment, the goodness of fit is greater for non-linear regression than for linear regression, and the evaluation of goodness of fit refers to non-linear regression. In some embodiments, the evaluation refers to non-linear regression and non-linear regression is used to normalize the count number.
[0227] In some embodiments, the evaluation of goodness of fit indicates linear regression if the correlation coefficient is equal to or greater than the correlation coefficient cut-off. In some embodiments, the evaluation of goodness of fit indicates non-linear regression if the correlation coefficient is less than the correlation coefficient cut-off. In some embodiments, the correlation coefficient cut-off is a predetermined cut-off. In some embodiments, the correlation coefficient cut-off is about 0.5 or greater, about 0.55 or greater, about 0.6 or greater, about 0.65 or greater, about 0.7 or greater, about 0.75 or greater, about 0.8 or greater, or about 0.85 or greater.
[0228] For example, in certain embodiments, a normalization method including linear regression is used when the correlation coefficient is equal to or greater than about 0.6. In certain embodiments, when the correlation coefficient is equal to or greater than a correlation coefficient cutoff of 0.6, the count number of a sample (e.g., the count number per reference genome portion, the count number per portion) is normalized according to linear regression, and otherwise, the count number is normalized according to non-linear regression (e.g., when the coefficient is less than the correlation coefficient cutoff of 0.6). In some embodiments, the normalization process includes generating (i) a count number and (ii) a linear regression or non-linear regression for each of a plurality of portions of the reference genome with respect to GC content. In certain embodiments, when the correlation coefficient is less than a correlation coefficient cutoff of 0.6, a normalization method including non-linear regression (e.g., LOWESS, LOESS) is used. In some embodiments, when the correlation coefficient (e.g., the correlation coefficient) is less than a correlation coefficient cutoff of about 0.7, less than about 0.65, less than about 0.6, less than about 0.55, or less than about 0.5, a normalization method including non-linear regression (e.g., LOWESS) is used. For example, in some embodiments, when the correlation coefficient is less than a correlation coefficient cutoff of about 0.6, a normalization method including non-linear regression (e.g., LOWESS, LOESS) is used.
[0229] In some embodiments, after selecting a specific type of regression (e.g., linear or non-linear regression) and generating the regression, the count number is normalized by subtracting the regression from the count number. In some embodiments, by subtracting the regression from the count number, a normalized count number with reduced bias (e.g., GC bias) is presented. In some embodiments, a linear regression is subtracted from the count number. In some embodiments, a non-linear regression (e.g., LOESS, GC-LOESS, LOWESS regression) is subtracted from the count number. Any suitable method can be used to subtract the regression line from the count number. For example, the count number x is derived from part i containing 0.5 GC content, and the regression line is used to determine the count number y when the GC content is 0.5, so x - y = the normalized count number for part i. In some embodiments, the count number is normalized before and / or after subtracting the regression. In some embodiments, the count number normalized by the hybrid normalization method is used to generate a genome segment level, Z-score, genome or its segment level and / or profile. In certain embodiments, the count number normalized by the hybrid normalization method is analyzed by the methods described herein to determine the presence or absence of gene mutations (e.g., in a fetus).
[0230] In some embodiments, the hybrid normalization method includes filtering or weighting one or more portions before or after normalization. Appropriate portion filtering methods can be used, including the filtering methods of the portions described herein (e.g., reference genome portions). In some embodiments, the portions (e.g., reference genome portions) are filtered before applying the hybrid normalization method. In some embodiments, only the counts of the sequencing reads mapped to selected portions (e.g., portions selected according to the variability of the counts) are normalized by hybrid normalization. In some embodiments, the counts of the sequencing reads mapped to the filtered reference genome portions (e.g., portions filtered according to the variability of the counts) are excluded before leveraging the hybrid normalization method. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., reference genome portions) according to an appropriate method (e.g., the methods described herein). In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., reference genome portions) according to the uncertainty values for the counts mapped to each of the portions for a plurality of test samples. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., reference genome portions) according to the variability of the counts. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., reference genome portions) according to the GC content, repetitive elements, repetitive sequences, introns, exons, etc., or combinations thereof.
[0231] For example, in some embodiments, a plurality of samples from a plurality of pregnant female subjects are analyzed, and a subset of a portion (e.g., a reference genomic portion) is selected according to the variability of the count number. In certain embodiments, linear regression is used to determine, for each of the selected portions of the samples obtained from pregnant female subjects, the correlation coefficient for (i) the count number and (ii) the GC content. In some embodiments, a correlation coefficient exceeding a predetermined correlation cutoff value (e.g., a correlation cutoff value of about 0.6) is determined, and by evaluating the goodness of fit, the linear regression is indicated, and the count number is normalized by subtracting the linear regression from the count number. In certain embodiments, a correlation coefficient less than a predetermined correlation cutoff value (e.g., a correlation cutoff value of about 0.6) is determined, and by evaluating the goodness of fit, a non-linear regression is indicated, a LOESS regression is generated, and the count number is normalized by subtracting the LOESS regression from the count number.
[0232] Profile In some embodiments, the step of processing may include generating (e.g., plotting a profile) one or more profiles from various aspects of a dataset or a derivative thereof (e.g., the results of one or more mathematical data processing steps and / or statistical data processing steps known in the art and / or described herein).
[0233] As used herein, the term "profile" refers to the result of mathematical and / or statistical operations on data that can facilitate the identification of patterns and / or correlations in a large amount of data. A "profile" often includes values obtained as a result of one or more operations on data or a dataset, based on one or more reference criteria. A profile often includes a plurality of data points. Depending on the nature and / or complexity of the dataset, any suitable number of data points can be incorporated into the profile. In certain embodiments, the profile can incorporate two or more data points, three or more data points, five or more data points, ten or more data points, twenty-four or more data points, twenty-five or more data points, fifty or more data points, one hundred or more data points, five hundred or more data points, one thousand or more data points, five thousand or more data points, ten thousand or more data points, or one hundred thousand or more data points.
[0234] In some embodiments, the profile displays the entire dataset, and in certain embodiments, the profile displays a part or subset of the dataset. That is, the profile may in some cases include or be generated from data points that display data that has not been filtered to exclude any data, and the profile may in some cases include or be generated from data points that display data that has been filtered to exclude unwanted data. In some embodiments, the data points in the profile display the results of data operations on parts. In certain embodiments, the data points in the profile include the results of data operations on groups of parts. In some embodiments, the groups of parts can be adjacent to each other, and in certain embodiments, the groups of parts can be from different parts of a chromosome or genome.
[0235] Data points in a profile derived from a dataset can represent any suitable categorization of data. Non-limiting examples of categories into which data can be grouped to generate profile data points include portions based on size, portions based on array features (e.g., GC content, AT content, location on a chromosome (e.g., short arm, long arm, centromere, telomere), etc.), levels of expression, chromosomes, etc., or combinations thereof. In some embodiments, a profile can be generated from data points obtained from another profile (e.g., a normalized data profile that has been renormalized according to different normalization values to generate a renormalized data profile). In certain embodiments, a profile generated from data points obtained from another profile reduces the number of data points and / or the complexity of the dataset. Reducing the number of data points and / or the complexity of the dataset often facilitates the interpretation of the data and / or the presentation of the outcome.
[0236] A profile (e.g., a genomic profile, a chromosomal profile, a profile of segments of a chromosome) is often a collection of normalized or non-normalized count numbers of two or more parts. A profile often includes at least one level (e.g., at the level of genomic segments) and often includes two or more levels (e.g., a profile often has multiple levels). A level is generally a level for a set of parts having approximately the same count number or normalized count number. Levels are described in more detail herein. In certain embodiments, a profile includes one or more parts that can be transformed by weighting, excluding, filtering, normalizing, adjusting, averaging, deriving as an average, adding, subtracting, processing, or any combination thereof. A profile often includes normalized count numbers mapped to parts that define two or more levels, where the count numbers are further normalized according to one of the levels by an appropriate method. The count numbers of a profile (e.g., profile level) are often associated with uncertain values.
[0237] Profiles that include one or more levels are sometimes filled in (e.g., hole-filled). Filling in (e.g., hole-filling) refers to the process of identifying and adjusting levels in a profile due to microdeletions in the parent or duplication in the parent (e.g., copy number variation). In some embodiments, levels due to fetal microduplication or fetal microdeletion are filled in. In some embodiments, microduplication or microdeletion in a profile can artificially increase or decrease the overall level of the profile (e.g., a chromosomal profile), resulting in false positive or false negative determinations regarding chromosomal aneuploidy (e.g., trisomy). In some embodiments, levels in a profile due to microduplication and / or deletion are identified and optionally adjusted (e.g., filled in and / or excluded) by a process referred to as filling in or hole-filling. In certain embodiments, the profile includes one or more first levels that are significantly different from a second level in the profile, and each of the one or more first levels includes a copy number variation in the parent, a copy number variation in the fetus, or a copy number variation in the parent and a copy number variation in the fetus, and one or more of the first levels are adjusted.
[0238] A profile that includes one or more levels may include a first level and a second level. In some embodiments, the first level is different (e.g., significantly different) from the second level. In some embodiments, the first level includes a first set of portions, the second level includes a second set of portions, and the first set of portions is not a subset of the second set of portions. In certain embodiments, the first set of portions is different from the second set of portions, and the first level and the second level are determined therefrom. In some embodiments, the profile may have a plurality of first levels that are different (e.g., significantly different, e.g., have significantly different values) from the second level in the profile. In some embodiments, the profile includes one or more first levels that are significantly different from the second level in the profile and adjusts one or more of the first levels. In some embodiments, the profile includes one or more first levels that are significantly different from the second level in the profile, and each of the one or more first levels includes a maternal copy number variation, a fetal copy number variation, or a maternal copy number variation and a fetal copy number variation, and adjusts one or more of the first levels. In some embodiments, the first level in the profile is excluded from or adjusted (e.g., filled in) the profile. The profile can include a plurality of levels including one or more first levels that are significantly different from one or more second levels, and most of the levels in the profile are often second levels that are approximately equal to each other. In some embodiments, more than 50%, more than 60%, more than 70%, more than 80%, more than 90% or more than 95% of the levels in the profile are second levels.
[0239] Profiles may, in some cases, be presented as plots. For example, one or more levels displaying the count number of parts (e.g., normalized count number) can be plotted and visualized. Non-limiting examples of plots of profiles that can be generated include raw count numbers (e.g., raw count number profiles or raw profiles), normalized count numbers, part weights, z-scores, p-values, area ratios compared to fitted ploidy, median levels compared to the ratio of the fitted fetal fraction to the measured fetal fraction, principal components, etc., or combinations thereof. In some embodiments, visualization of the manipulation data is enabled by plotting the profiles. In certain embodiments, plots of profiles can be utilized to present outcomes (e.g., area ratios compared to fitted ploidy, median levels compared to the ratio of the fitted fetal fraction to the measured fetal fraction, principal components). As used herein, the terms "plot of raw count number profile" or "plot of raw profile" refer to a plot of the count number in each part (e.g., genome, part, chromosome, chromosomal part of the reference genome, or segment of a chromosome) in a region, normalized according to the total count number in the region. In some embodiments, profiles can be generated using static window processing, and in certain embodiments, profiles can be generated using sliding window processing.
[0240] The profiles generated for the test subjects, in some cases, facilitate the interpretation of mathematical and / or statistical operations on the dataset and / or present an outcome, when compared to the profiles generated for one or more reference subjects. In some embodiments, the profiles are generated based on one or more starting assumptions (e.g., the nucleic acid contribution of the mother (e.g., the maternal fraction), the nucleic acid contribution of the fetus (e.g., the fetal fraction), the ploidy of the reference sample, etc., or combinations thereof). In certain embodiments, the test profile often centers around a predetermined value indicating the absence of a gene mutation, and when the test subject is assumed to carry a gene mutation, it often deviates from the predetermined value in an area corresponding to the genomic location where the gene mutation is located in the test subject. For a test subject at risk of or suffering from a medical condition associated with a gene mutation, it is expected that the numerical values for the selected portion will vary significantly from the predetermined values for genomic locations not affected. Depending on the starting assumptions (e.g., a certain ploidy or an optimized ploidy, a certain fetal fraction or an optimized fetal fraction, or combinations thereof), the predetermined threshold or cutoff value or range of thresholds indicating the presence or absence of a gene mutation can also vary while still presenting an outcome useful for determining the presence or absence of a gene mutation. In some embodiments, the profiles indicate and / or display a phenotype.
[0241] By way of non-limiting example, a normalized sample and / or reference count profile can be obtained from raw array read data by: (a) calculating the median reference count for chromosomes, portions, or segments thereof selected from a set of references known to not carry mutations in the gene; (b) excluding (e.g., filtering) portions of the reference sample that do not provide information from the raw counts; (c) normalizing the reference counts for all remaining reference genome portions according to the total remaining counts (e.g., the sum of the remaining counts after excluding reference genome portions that do not provide information) for the reference sample, selected chromosome, or selected genomic position, thereby generating a normalized reference profile; (d) excluding corresponding portions from the test sample; and (e) normalizing the remaining test counts for one or more selected genomic positions according to the sum of the median remaining reference counts for one or more chromosomes containing the selected genomic position, thereby generating a normalized test profile. In certain embodiments, a further normalization step for the entire genome reduced by filtering of portions in (b) can be incorporated between (c) and (d).
[0242] A dataset profile can be generated by one or more operations on the counted mapped array read data. Some embodiments include: mapping array reads and determining (e.g., counting) the number of array tags mapped to each genomic portion. Generating a raw count profile from the counted mapped array reads. In certain embodiments, presenting an outcome by comparing a raw count profile from a test subject to a median reference count profile for chromosomes, portions, or segments thereof from a set of reference subjects known to not carry mutations in the gene.
[0243] In some embodiments, the read data of the array is optionally filtered to exclude noise data or portions that do not provide information. After filtering, it is typical to sum the remaining counts to generate a filtered data set. In certain embodiments, a filtered count profile is generated from the filtered data set.
[0244] After counting and optionally filtering the read data of the array, the data set can be normalized to generate a level or profile. The data set can be normalized by normalizing one or more selected portions according to appropriate normalized reference values. In some embodiments, the normalized reference value represents the total count for one or more chromosomes for which a portion is selected. In certain embodiments, the normalized reference value represents one or more corresponding portions that are portions of one or more chromosomes derived from a reference data set prepared from a set of reference subjects known to not carry gene mutations. In some embodiments, the normalized reference value represents one or more corresponding portions that are portions of one or more chromosomes derived from a test subject data set prepared from a test subject being analyzed for the presence or absence of gene mutations. In certain embodiments, the normalization process is performed using the static window method, and in some embodiments, the normalization process is performed using the moving window method or the sliding window method. In certain embodiments, a profile including the normalized counts is generated to facilitate the classification and / or presentation of the outcome. The outcome can be presented based on a plot of the profile including the normalized counts (e.g., using such a plot of the profile).
[0245] Level In some embodiments, a value (e.g., a number, a quantitative value) is attributed to a level. The level can be determined by an appropriate method, operation, or mathematical process (e.g., a processed level). The level is often a count number (e.g., a normalized count number) for a subset of parts or is derived therefrom. In some embodiments, the level of a part is substantially equal to the total number of count numbers (e.g., count numbers, normalized count numbers) mapped to the part. The level is often determined from count numbers that have been processed, transformed, or manipulated by an appropriate method, operation, or mathematical process known in the art. In some embodiments, the level is derived from the processed count numbers, and non-limiting examples of processed count numbers include being weighted, excluded, filtered, normalized, adjusted, averaged, derived as an average (e.g., an average level), added, subtracted, transformed count numbers, or combinations thereof. In some embodiments, the level includes a normalized count number (e.g., a normalized count number of a part). The level can be a level for a normalized count number that has been normalized by an appropriate process, the non-limiting examples of which include normalization with respect to a part, normalization by GC content, count number median normalization, linear least squares regression and non-linear least squares regression, LOESS (e.g., GC LOESS), LOWESS, PERUN, ChAI, principal component normalization, RM, GCRM, cQn, etc., and / or combi...
Claims
【Claim 1】 The apparatus described in the specification.
Citation Information
Patent Citations
Detecting and classifying copy number variation
US20130096011A1
Diagnostic processes that factor experimental conditions
US20130150253A1
Methods and processes for non-invasive assessment of genetic variations
US20130261983A1
Normalizing chromosomes for the determination and verification of common and rare chromosomal aneuploidies
WO2012141712A1
Noninvasive detection of fetal genetic abnormality
WO2013000100A1