Methods and processes for assessment of genetic variations
The method for classifying gene copy number variations in test samples through segmentation and quantitative sequence read analysis addresses the limitations of current techniques, providing improved accuracy and diagnostic capabilities for non-invasive prenatal and oncological testing.
Patent Information
- Application Number
- JP2025022284
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2017-01-24
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2038-01-24
AI Technical Summary
Current methods for non-invasive classification of gene copy number variations (CNVs) in test samples are limited in accuracy and efficiency, particularly for prenatal and oncological testing.
A method involving the identification of copy number variant segments using a segmentation process and providing quantitative values of sequence reads for sub-chromosomal regions, allowing for classification of the presence or absence of CNVs based on changes relative to a reference sample set.
This approach enables accurate and non-invasive classification of CNVs, improving diagnostic capabilities for prenatal and oncological testing by enhancing sensitivity and specificity.
Smart Images

Figure 2025087718000001_ABST
Abstract
Description
Technical Field
[0001] Related Patent Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 449,766, filed on January 24, 2017. The entire contents of the provisional patent application are hereby incorporated by reference in their entirety for all purposes.
[0002] Field The technology provided herein relates in part to methods, systems, devices, and computer program products for non-invasive classification of gene copy number variations (CNVs) for test samples. The technology provided herein is useful for classifying gene CNVs for samples, for example, as part of non-invasive prenatal (NIPT) testing and oncological testing.
Background Art
[0003] Background The genetic information of living organisms (e.g., animals, plants, and microorganisms) and other forms that replicate genetic information (e.g., viruses) is encoded in deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). Genetic information is a sequence of nucleotides or modified nucleotides corresponding to the primary structure of a chemical nucleic acid or a hypothetical nucleic acid. The entire human genome contains approximately 30,000 genes located on 24 chromosomes (i.e., 22 autosomes, the X chromosome, and the Y chromosome; see The Human Genome, T. Strachan, BIOS Scientific Publishers, 1992). Each gene encodes a specific protein, which performs a specific biochemical function in a living cell after expression through transcription and translation.
[0004] Many medical conditions are caused by one or more genetic mutations and / or genetic changes. Certain genetic mutations and / or genetic changes cause medical conditions such as, for example, hemophilia, thalassemia, Duchenne muscular dystrophy (DMD), Huntington's disease (HD), Alzheimer's disease, and cystic fibrosis (CF) (Human Genome Mutations, D.N. Cooper and M. Krawczak, BIOS Publishers, 1993). Such genetic disorders can result from the addition, substitution, or deletion of a single nucleotide in the DNA of a particular gene. Certain birth defects are caused by chromosomal abnormalities, also known as aneuploidies, such as trisomy 21 (Down syndrome), trisomy 13 (Patau syndrome), trisomy 18 (Edwards syndrome), monosomy X (Turner syndrome), and certain sex chromosome aneuploidies, such as Klinefelter syndrome (XXY). Another genetic mutation is the sex of the fetus, which can often be determined based on the sex chromosomes X and Y. Some genetic mutations can make an individual susceptible to, or cause, any of several diseases, such as diabetes, atherosclerosis, obesity, various autoimmune diseases, and cancer (e.g., colorectal cancer, breast cancer, ovarian cancer, lung cancer, bladder cancer, gastric cancer, cervical cancer, kidney cancer, prostate cancer, brain cancer, and esophageal cancer).
[0005] When one or more gene mutations and / or gene changes (e.g., copy number changes, copy number variations, single nucleotide changes, single nucleotide variations, chromosomal changes, translocations, deletions, insertions, etc.) or genetic dispersions are identified, it is possible to diagnose certain medical symptoms or determine a predisposition to certain medical symptoms. When a genetic dispersion is identified, it is possible to prompt a medical decision and / or use beneficial medical procedures. In certain embodiments, the identification of one or more gene mutations and / or gene changes requires the analysis of circulating cell-free nucleic acids. Circulating cell-free nucleic acids (CCF-NA), e.g., cell-free DNA (CCF-DNA), are composed of DNA fragments that are derived from cell death and circulate in peripheral blood. High concentrations of CF-DNA can suggest certain clinical symptoms, such as cancer, trauma, burns, myocardial infarction, stroke, sepsis, infections, and other diseases. Additionally, cell-free fetal DNA (CFF-DNA) can be detected in the maternal bloodstream and can be used for various non-invasive prenatal diagnoses.
Summary of the Invention
Means for Solving the Problems
[0006] Abstract A method is provided herein for classifying the presence or absence of copy number variations in a subchromosomal region for a test sample, the method comprising: a) identifying the presence or absence of copy number variant segments in a region comprising a first subset of genomic portions using a method comprising a segmentation process; b) providing a quantitative value of sequence reads for a sub-region within a subchromosomal region comprising a second subset of genomic portions, wherein the second set is a predetermined subset of genomic portions, and the genomic portions in (a) and (b) comprise portions of a reference genome to which sequence reads obtained for nucleic acids in the test sample are mapped, the method comprising classifying the presence or absence of copy number variations in the subchromosomal region for the test sample according to (a), or (b), or (a) and (b).
[0007] In certain embodiments, provided is a method for classifying the presence or absence of a copy number variation in a subchromosomal region for a test sample, the method comprising: a) identifying the presence or absence of a copy number variation segment in a region comprising a first subset of genomic portions using a method comprising a segmentation process; b) providing a quantitative value of sequence reads for a sub-region within a subchromosomal region comprising a second subset of genomic portions, wherein the second set is a predetermined subset of genomic portions, and the genomic portions in (a) and (b) comprise portions of a reference genome to which sequence reads obtained for nucleic acids in the test sample are mapped; and c) providing a classification of the presence or absence of a copy number variation in the subchromosomal region for the test sample based on changes within the region of (a), within the sub-region of (b), or both, relative to a reference sample set. The region in (a) may encompass the subchromosomal region or may overlap with the subchromosomal region.
[0008] In some embodiments, the first subset of genomic portions is a portion within a region on a chromosome where a copy number variation associated with a phenotype of interest is expected to be present. Such genomic portions are often obtainable by mining public disease databases such as the International Standards of Cytogenomic Arrays database (ISCA). In one embodiment, the phenotype is a microdeletion syndrome. In one embodiment, the first subset of genomic portions is one or more genomic portions selected from 1p36, 22q11.2, 15q11-13, 8q23.2-24.1, 11q24.1, 4p13.3, 17p13.3, and 7q11.23.
[0009] In certain embodiments, a method is provided for classifying the presence or absence of a copy number variation in a sub-chromosomal region for a test sample, the method comprising: a) providing a quantitative value of sequence reads for a sub-region within a sub-chromosomal region that includes a subset of genomic portions, where: i) the genomic portion includes a portion of a reference genome to which sequence reads obtained for nucleic acids in the test sample are mapped; ii) the subset is a predetermined subset of genomic portions; iii) the predetermined subset of genomic portions is identified by a process that includes: 1) providing a plurality of candidate sub-regions within the sub-chromosomal region; 2) providing one or more accuracy measures for each of the plurality of candidate sub-regions for a plurality of samples in a training set, where each of the plurality of samples is classified as having a copy number variation in the sub-chromosomal region; and 3) identifying the sub-region in (a) as a sub-region that provides an accuracy measure equal to or exceeding a predetermined threshold; and b) providing a classification of the presence or absence of a copy number variation in the sub-chromosomal region for the test sample according to the quantitative value of the sequence reads in (a) relative to the quantitative value of the sequence reads for a reference sample set.
[0010] Also provided herein is a system comprising one or more processors and a memory, the memory including instructions executable by the one or more processors, the instructions executable by the one or more processors comprising: a) configured to identify the presence or absence of a copy number variation segment in a region including a first subset of genomic portions using a method including a segmentation process; and / or
[0011] b) configured to provide a quantitative value of sequence reads for a sub-region within a sub-chromosomal region containing a second genomic subset, where the second set is a predetermined genomic subset, and the genomic subsets in (a) and (b) include the portions of the reference genome to which the sequence reads obtained for the nucleic acids in the test sample are mapped; c) configured to provide a classification of the presence or absence of copy number variations in the sub-chromosomal region for the test sample based on changes within the region of (a), within the sub-region of (b), or both, relative to a set of reference samples.
[0012] A computer program product as a computer-readable storage medium is also provided herein, the product including instructions programmed for a computer to: a) identify the presence or absence of copy number variant segments in a region containing a first genomic subset using a method including a segmentation process; and / or b) provide a quantitative value of sequence reads for a sub-region within a sub-chromosomal region containing a second genomic subset, where the second set is a predetermined genomic subset, and the genomic subsets in (a) and (b) include the portions of the reference genome to which the sequence reads obtained for the nucleic acids in the test sample are mapped; and c) provide a classification of the presence or absence of copy number variations in the sub-chromosomal region for the test sample based on changes within the region of (a), within the sub-region of (b), or both, relative to a set of reference samples.
[0013] Certain embodiments are further described in the following description, examples, claims, and drawings.
[0014] The drawings illustrate certain embodiments of the technology and are not limiting. For clarity and simplicity of illustration, the drawings are not drawn to scale and in some cases, various aspects may be exaggerated or enlarged to facilitate understanding of certain embodiments.
Brief Description of the Drawings
[0015]
Figure 1
[0016]
Figure 2
[0017]
Figure 3
[0018]
Figure 4
[0019]
Figure 5
[0020]
Figure 6
Best Mode for Carrying Out the Invention
[0021] Detailed Description Methods are provided herein that are useful for classifying the presence or absence of copy number variations in sub-chromosomal regions for a test sample. In some embodiments, the sample nucleic acid subjected to a sequencing process and the resulting sequence reads are further analyzed to determine the presence or absence of copy number variations. In some embodiments, the presence or absence of copy number variations is classified according to a genome-wide sequencing analysis. In some embodiments, the presence or absence of copy number variations is classified according to a focused sequencing analysis (e.g., analysis of sequence reads for a predetermined genomic sub-region). A focused sequencing analysis can improve the accuracy (e.g., sensitivity) for detecting copy number variations in a particular type of sample. In some embodiments, the presence or absence of copy number variations is classified according to a genome-wide sequencing analysis and a focused sequencing analysis.
[0022] In some embodiments, systems, apparatuses, and computer program products are also provided for performing the methods or portions of the methods described herein. Classification of Copy Number Variations Using Genome-Wide Sequence Analysis and / or Focused Sequence Analysis
[0023] Methods and processes are provided herein for classifying the presence or absence of copy number variations (e.g., microdeletions, microduplications) in sub-chromosomal regions. As used herein, microdeletions and microduplications collectively refer to deletions or duplications that are smaller than 5 million base pairs. Microdeletions and microduplications are typically too small to be detected by conventional cytogenetic methods or high-resolution karyotyping. By using the methods and systems of the present disclosure, both microdeletions and microduplications can be accurately detected.
[0024] In some embodiments, the presence or absence of a copy number variation is classified according to an array read set. In some embodiments, the array reads are obtained for nucleic acids in a test sample. In some embodiments, the array reads are mapped to genomic portions in a reference genome. In some embodiments, classifying the presence or absence of a copy number variation in a subchromosomal region includes identifying the presence or absence of a copy number variation segment. As used herein, a copy number variation segment is a segment in a chromosome that includes a copy number variation. In some embodiments, a copy number variation segment is identified using a method that includes a segmentation process. A method that includes a segmentation process can include a determination analysis, such as the determination analysis described herein. A method that includes a segmentation process can be part of a genome-wide sequence analysis method. A method that includes a segmentation process can be part of a sequence analysis of nucleic acids captured by probe oligonucleotides. In some embodiments, classifying the presence or absence of a copy number variation in a subchromosomal region includes providing a quantitative value of array reads for subregions within the subchromosomal region. As an illustrative example, the subregion is the region defined by the gray dashed line in FIG. 4.
[0025] In some embodiments, the sub-region comprises a predetermined subset of genomic portions. Providing a quantitative value of array reads for the sub-region can be part of intensive array analysis. Providing a quantitative value of array reads for the sub-region can be part of intensive array analysis of nucleic acids captured by probe oligonucleotides. In some embodiments, classification of the presence or absence of a copy number variation in a sub-chromosomal region is provided according to the presence or absence of a copy number variation segment. In some embodiments, classification of the presence or absence of a copy number variation in a sub-chromosomal region is provided according to a quantitative value of array reads for a sub-region within the sub-chromosomal region. In some embodiments, classification of the presence or absence of a copy number variation in a sub-chromosomal region is provided according to the presence or absence of a copy number variation segment and according to a quantitative value of array reads for a sub-region within the sub-chromosomal region.
[0026] In some embodiments, classifying the presence or absence of a copy number variation in a sub-chromosomal region involves providing a quantitative value of sequence reads for sub-regions within the sub-chromosomal region, where the sub-regions include a predetermined subset of the genome. The predetermined subset of the genome can be identified according to one or more accuracy metrics for a plurality of samples (e.g., a plurality of samples in a training set). Generally, each sample in the set of the plurality of samples (e.g., the training set) is classified as having a copy number variation in the target sub-chromosomal region. The samples in the set of the plurality of samples can be obtained from one or more subjects known to have a copy number variation and / or can be generated by adding genomic DNA having a copy number variation to a reference sample and / or can be generated according to in silico modeling. Having a copy number variation in the target sub-chromosomal region can include a copy number variation identified at genomic coordinates within the target sub-chromosomal region, a copy number variation identified at genomic coordinates overlapping the target sub-chromosomal region, a copy number variation identified at genomic coordinates adjacent to the target sub-chromosomal region (e.g., within about 1 megabase of the target sub-chromosomal region), etc. The copy number variations in the set of the plurality of samples can include duplications, microduplications, deletions, and microdeletions. Duplications and deletions can be of any size, while microduplications and microdeletions generally refer to duplications and deletions smaller than 5 million bases that are usually too small to be detected by conventional cytogenetic methods or high-resolution karyotyping.
[0027] The accuracy measure for a plurality of samples may include any suitable accuracy measure for determining the presence or absence of copy number variation for the plurality of samples. The accuracy measure may include sensitivity, specificity, standard deviation, median absolute deviation (MAD), measure of certainty, measure of confidence, measure of certainty or confidence that a value obtained for a test sample is inside or outside a particular value range, measure of uncertainty, measure of uncertainty that a value obtained for a test sample is inside or outside a particular value range, coefficient of variation (CV), confidence level, confidence interval (e.g., about 95% confidence interval), standard score (e.g., z-score), chi value, phi value, result of a t-test, p-value, ploidy value, fitted minor allele ratio, area ratio, median level, etc. or combinations thereof. In some embodiments, the accuracy measure includes sensitivity.
[0028] Typically, each of the plurality of samples (e.g., the plurality of samples in the training set) has a known copy number variation, so the accuracy of detecting copy number variations can be evaluated. In some embodiments, the accuracy of detecting copy number variations for the plurality of samples can be optimized. In some embodiments, the accuracy of detecting copy number variations for the plurality of samples can be optimized by identifying a genomic subset that provides an optimal accuracy measure for classifying the presence of copy number variations for the plurality of samples. As disclosed herein, the term "optimal accuracy" refers to an accuracy measure that is equal to or higher than a predetermined threshold. That predetermined threshold is considered the minimum requirement for detecting the presence or absence of copy number variations with reasonable accuracy. One of ordinary skill in the art can readily determine what the predetermined threshold is for any particular accuracy measure required for a particular assay. In some embodiments, the accuracy of detecting copy number variations for the plurality of samples can be optimized by identifying a genomic subset that provides an optimal sensitivity for classifying the presence of copy number variations for the plurality of samples. In some embodiments, the genomic subset that provides an optimal accuracy measure (e.g., optimal sensitivity) is referred to as a predetermined genomic subset or a predetermined subregion. In some embodiments, the genomic subset that provides an optimal accuracy measure (e.g., optimal sensitivity) is identified by a process that includes: 1) providing a plurality of candidate subregions within a target subchromosomal region (e.g., a subchromosomal region having a possible copy number variation); 2) providing one or more accuracy measures (e.g., sensitivity values) for each of the plurality of candidate subregions for the plurality of samples (e.g., in the training set); and 3) identifying the genomic subset in the subregion that provides the optimal accuracy (e.g., optimal sensitivity) according to that one or more accuracy measures. The plurality of candidate subregions provided for identifying the genomic subset that provides the optimal accuracy measure typically includes subregions having one or more genomic coordinates that are different from each other.For example, a candidate sub-region may have unique genomic coordinates at its 5' end, may have unique genomic coordinates at its 3' end, or may have unique genomic coordinates at both its 5' and 3' ends. The candidate sub-regions may be of the same length as each other, may be of different lengths, or may be a combination of both.
[0029] In some embodiments, one or more accuracy metrics include a sensitivity metric. Sensitivity can be determined as the number or percentage of samples identified as having a copy number variation, where the samples are from a plurality of sample sets having copy number variations. In some embodiments, the sensitivity for classifying each of a plurality of samples (e.g., within a training set) as having a copy number variation in a target sub-chromosomal region is at least about 70%. For example, the sensitivity for classifying each of a plurality of samples (e.g., within a training set) as having a copy number variation in a target sub-chromosomal region can be at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. In some embodiments, the sensitivity for classifying each of a plurality of samples (e.g., within a training set) as having a copy number variation in a target sub-chromosomal region is at least about 75%. In some embodiments, the sensitivity for classifying each of a plurality of samples (e.g., within a training set) as having a copy number variation in a target sub-chromosomal region is at least about 80%. In some embodiments, the sensitivity for classifying each of a plurality of samples (e.g., within a training set) as having a copy number variation in a target sub-chromosomal region is at least about 85%. In some embodiments, the sensitivity for classifying each of a plurality of samples (e.g., within a training set) as having a copy number variation in a target sub-chromosomal region is at least about 90%. In some embodiments, the sensitivity for classifying each of a plurality of samples (e.g., within a training set) as having a copy number variation in a target sub-chromosomal region is at least about 95%. In some embodiments, the sensitivity for classifying each of a plurality of samples (e.g., within a training set) as having a copy number variation in a target sub-chromosomal region is at least about 97%.
[0030] In some embodiments, classifying the presence or absence of a copy number variation in a sub-chromosomal region involves providing a quantitative value of sequence reads for the sub-region (e.g., the sub-regions described above). The quantitative value of sequence reads for the sub-region can be a sequence read count (e.g., the direct sum of read counts, raw read count, normalized read count, filtered read count, read density, weighted read count, read count ratio, average read count, mean read count value, adjusted read count, etc. and combinations thereof). In some embodiments, the quantitative value of sequence reads for the sub-region is a quantitative value of normalized sequence reads generated by a normalization process. The normalization process can include any suitable normalization that normalizes GC bias and / or other biases. Examples of certain normalization processes are described herein. In some embodiments, the normalization process includes LOESS normalization. In some embodiments, the normalization process includes principal component normalization. Classifying the presence or absence of a copy number variation in a sub-chromosomal region can be based on a change in the quantitative value of sequence reads relative to a reference sample set. For the purposes of the present disclosure, the reference sample set can be any sample identified as not having the copy number variation to be detected in the test sample. The reference sample can be from a subject without a copy number variation and of a similar tissue type and / or a similar population type.
[0031] In some embodiments, the quantitative value of sequence reads for the sub-region is a standard score. In some embodiments, the quantitative value of sequence reads for the sub-region is a z-score. The z-score can be for the sub-region or can be assigned to each genomic portion included in the sub-region. The z-score is as follows: Z SUB =(SUB scq -SUB mcq ) / MAD and can be generated for the sub-region according to (Z SUB ).
[0032] where SUBscq is the test sample count quantification value of the sub-region (e.g., SUB scq can be the result of dividing the normalized total count in the sub-region for the test sample by the normalized total count of the autosome); SUB mcq is the median of the count quantification values for the sub-region generated for the reference sample set; MAD is the median absolute deviation determined for the count quantification values of the sub-region for the reference sample set. In a particular case, SUB mcq is the mean of the count quantification values for the sub-region generated for the reference sample set; the denominator of the above equation is the standard deviation determined for the count quantification values of the sub-region for the reference sample set. In a particular case, SUB scq can be the result of dividing the total count in the sub-region for the test sample by the total count of the autosome. The total count of the autosome can be normalized (e.g., can be GC-normalized), filtered (e.g., repeat regions can be filtered out, low mapping regions can be filtered out, and / or other regions can be filtered out as described herein), or can be normalized and filtered. In a particular case, SUB scqIt can be the result of dividing the total count (e.g., normalized total count) in the sub-region by the total count (e.g., normalized total count) for the genomic subset of the test sample. Examples of genomic subsets can include, for example, all autosomes, a part of all autosomes, a particular autosome, a part of a particular autosome, etc. and combinations thereof. The reference sample set can include samples classified as not having copy number variations. In some embodiments, the reference sample consists of samples classified as not having copy number variations. Thus, in some embodiments, the reference sample includes or consists of samples in which each chromosome and each chromosomal region being tested is euploid. The reference sample can be from a human subject. In some embodiments, the reference sample is from a female subject. In some embodiments, the reference sample is from a male subject. In some embodiments, the reference sample is from male and female subjects. The reference sample can include samples from one subject or samples from multiple subjects. The reference sample can include one reference sample, but often includes multiple samples. For example, the reference sample can include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more samples. Instead of the z-score, other quantitative values can be utilized, and non-limiting examples thereof include normal scores, z-values, standardized variables, and t-statistics.
[0033] In some embodiments, the presence or absence of a copy number variation for a sub-region is classified according to a z-score cut-off. The z-score cut-off can be determined according to a preferred level of sensitivity and / or specificity for determining the presence or absence of a copy number variation for a test sample. In some embodiments, the z-score cut-off value is set to an absolute value of about 2 to about 4. For example, the z-score cut-off value can be set to an absolute value of about 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9 or 4.0. In some embodiments, the z-score cut-off value is set to an absolute value of about 3 to about 5. For example, the z-score cut-off value can be set to an absolute value of about 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9 or 5.0. In some embodiments, the z-score cut-off value is set to an absolute value of about 3.9 to about 4.0. For example, the z-score cut-off value can be set to an absolute value of about 3.90, 3.91, 3.92, 3.93, 3.94, 3.95, 3.96, 3.97, 3.98, 3.99 or 4.0. In some embodiments, the z-score cut-off value is set to an absolute value of about 3.95. If the absolute value of one or more z-scores for a sub-region is greater than the selected cut-off value, the presence or absence of a copy number variation for the test sample can be determined. In some embodiments, the classification of the presence or absence of a copy number variation in a sub-chromosomal region for a test sample is provided according to a quantitative value of sequence reads for the sub-region (e.g., z-score). In some embodiments, if the z-score generated using the methods described herein is less than -3, less than -3.2 or less than -3.5, e.g., less than -3.95, the classification of the presence of a deletion in the sub-chromosomal region is determined. In some embodiments, if the z-score is greater than 3, greater than 3.2, greater than 3.5, e.g., greater than 3.95, the classification of the presence of a duplication in the sub-chromosomal region is made.
[0034] In some embodiments, classifying the presence or absence of a copy number variation in a subchromosomal region involves identifying the presence or absence of a copy number variation segment. In some embodiments, the copy number variation segment is identified using a method that includes a segmentation process. A method that includes a segmentation process can include a determination analysis such as the determination analysis described herein. For example, a determination analysis can involve applying one or more results, evaluations, and one or more methods that result in a series of determinations, based on the possible consequences of those results, evaluations, and / or those determinations, and can end in a significant aspect of the process where the final determination is made. In some embodiments, the determination analysis is a decision tree. In some embodiments, the presence or absence of a copy number variation segment is identified according to a determination analysis that includes a segmentation process or a segmenting process.
[0035] In some embodiments, a segmentation process is applied to identify segments (e.g., segments spanning copy number variations; copy number variation segments). Any suitable segmentation process may be utilized, including but not limited to the circular binary segmentation (CBS) process. CBS generally functions by repeatedly dividing a single chromosome into regions of equal copy number using likelihood ratio statistics. CBS is described, for example, in Olshen et al. (2004) Biostatistics 5:557-72; Venkatraman et al. (2007) Bioinformatics 23:657-63; Lai et al. (2005) Bioinformatics 21:3763-70; Willenbrock et al. (2005) Bioinformatics 21:4084-91. Instead of or in addition to CBS, other processes may be utilized, non-limiting examples of which include wavelet segmentation (e.g., Haar wavelet segmentation), Fourier transform, sliding window z-score, and Markov chain models.
[0036] In some embodiments, the method of classifying the presence or absence of copy number variations uses genome-wide analysis, i.e., analysis based on the Circular Binary Segmentation (CBS) method to find events, such as the edges of microdeletions or microduplications, within a genomic window that includes a target region, e.g., 22q11.2. CBS is useful for detecting small deletions. In some embodiments, the method of classifying the presence or absence of copy number variations uses focused analysis, i.e., analysis using a predefined region within the target region. Generally, when the test sample includes a low fetal fraction, e.g., less than 10% fetal fraction, focused sequencing analysis is more reliable and / or more sensitive, while when the test sample includes a high fetal fraction, e.g., more than 10% fetal fraction, genome-wide sequencing analysis can be more sensitive and is thus preferred. Illustrative embodiments are shown in FIG. 2. In certain embodiments, the method uses both genome-wide analysis and focused sequencing analysis, and maximizes sensitivity by using the edge detection ability of CBS, which allows for the identification of small deletions and improvement of sensitivity at low fetal fractions by focused sequencing analysis.
[0037] In some embodiments, a quantitative value is generated for a copy number variant segment identified by a segmentation process. In some embodiments, the segmentation process generates a quantitative value for the copy number variant segment. The quantitative value for the copy number variant segment may include the quantitative value of the array reads. The quantitative value of the array reads for the copy number variant segment may be an array read count (e.g., the direct sum of read counts, raw read count, normalized read count, filtered read count, read density, weighted read count, read count ratio, average read count, read count average value, adjusted read count, etc. and combinations thereof). In some embodiments, the quantitative value of the array reads for the copy number variant segment is the quantitative value of the normalized array reads generated by a normalization process. The normalization process may include any suitable normalization that normalizes GC bias and / or other biases. Examples of certain normalization processes are described herein. In some embodiments, the normalization process includes LOESS normalization. In some embodiments, the normalization process includes principal component normalization.
[0038] In some embodiments, the quantitative value for the copy number variant segment is a standard score. In some embodiments, the quantitative value for the copy number variant segment is a z-score. The z-score may sometimes be for the segment and sometimes be assigned to each genomic portion included in the segment. The z-score is as follows: Z SEG =(SEG scq -SEG mcq ) / MAD and can be generated for the copy number variant segment according to (Z SEG ).
[0039] Wherein, SEG scq is the test sample count quantitative value of the segment (e.g., SEG scqcan be the result of dividing the total normalized count in the segment for the test sample by the total normalized count of the autosomes); SEG mcq is the median of the count quantification values for the segments generated for the reference sample set; MAD is the median absolute deviation determined for the count quantification values of the segments for the reference sample set. In certain cases, SEG mcq is the mean of the count quantification values for the segments generated for the reference sample set; the denominator of the above equation is the standard deviation determined for the count quantification values of the segments for the reference sample set. In certain cases, SEG scq can be the result of dividing the total count in the sub-region for the test sample by the total count of the autosomes. The total count of the autosomes can be normalized (e.g., can be GC-normalized), filtered (e.g., repeat regions can be filtered out, low mapping regions can be filtered out, and / or other regions can be filtered out as described herein), or normalized and filtered. In certain cases, SEG scq can be the result of dividing the total count (e.g., the total normalized count) in the sub-region for the test sample by the total count (e.g., the total normalized count) for the genomic subset. Examples of genomic subsets can include, for example, all autosomes, a portion of all autosomes, a particular autosome, a portion of a particular autosome, etc. and combinations thereof. The reference sample set can be any suitable reference set and can include the reference sample sets described herein.
[0040] Non-limiting examples of methodologies useful for generating z-score copy number quantification values based on segmentation (e.g., CBS) are described in Zhao et al., Clin.Chem. 61:4:608-616 (2015); Lefkowitz et al., American Journal of Obstetrics & Gynecology 1.e1 (2016); and International Patent Application No. PCT / US2014 / 039389 (filed May 23, 2014, and published Nov. 27, 2014 as WO2014 / 190286). Instead of the z-score, other normalized CNV quantification values may be utilized, non-limiting examples of which include the normal score, z-value, standardized variable, and t-statistic.
[0041] In some embodiments, the presence or absence of a copy number variation for a segment is classified according to a z-score cut-off. The z-score cut-off can be determined according to a preferred level of sensitivity and / or specificity for determining the presence or absence of a copy number variation for a test sample. In some embodiments, the z-score cut-off value is set to an absolute value of from about 2 to about 4. For example, the z-score cut-off value can be set to an absolute value of about 2.0, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 2.7, 2.8, 2.9, 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9 or 4.0. In some embodiments, the z-score cut-off value is set to an absolute value of from about 3 to about 5. For example, the z-score cut-off value can be set to an absolute value of about 3.0, 3.1, 3.2, 3.3, 3.4, 3.5, 3.6, 3.7, 3.8, 3.9, 4.0, 4.1, 4.2, 4.3, 4.4, 4.5, 4.6, 4.7, 4.8, 4.9 or 5.0. In some embodiments, the z-score cut-off value is set to an absolute value of from about 3.9 to about 4.0. For example, the z-score cut-off value can be set to an absolute value of about 3.90, 3.91, 3.92, 3.93, 3.94, 3.95, 3.96, 3.97, 3.98, 3.99 or 4.0. In some embodiments, the z-score cut-off value is set to an absolute value of about 3.95. If the absolute value of one or more z-scores for a segment is greater than the selected cut-off value, the presence or absence of a copy number variation for the test sample can be determined. In some embodiments, the classification of the presence or absence of a copy number variation in a subchromosomal region for a test sample is provided according to a quantitative value (e.g., z-score) for the copy number variation segment.
[0042] In some embodiments, the classification of the presence or absence of a copy number variation in a subchromosomal region for a test sample is provided according to a quantitative value (e.g., z-score) for the copy number variation segment and a quantitative value (e.g., z-score) for the sequence reads for the subregion. In some embodiments, the classification of the presence or absence of a copy number variation in a subchromosomal region for a test sample is provided according to a quantitative value (e.g., z-score) for the copy number variation segment or a quantitative value (e.g., z-score) for the sequence reads for the subregion. Thus, in certain cases, the classification is provided according to quantitative values (e.g., z-scores) for both the segment and the subregion, and in certain cases, the classification is provided according to either the quantitative value (e.g., z-score) for the segment or the quantitative value (e.g., z-score) for the subregion.
[0043] In some embodiments, a segment includes a first subset of genomic portions and a sub-region includes a second subset of genomic portions. In some embodiments, the first subset of genomic portions and the second subset of genomic portions include the same genomic portion. In some embodiments, the first subset of genomic portions and the second subset of genomic portions consist of the same genomic portion. In some embodiments, the first subset of genomic portions and the second subset of genomic portions include different genomic portions. In some embodiments, the first subset of genomic portions and the second subset of genomic portions include some genomic portions that are the same and some genomic portions that are different. In some embodiments, the second subset of genomic portions is a subset of the first subset of genomic portions. In some embodiments, the first subset of genomic portions is a subset of the second subset of genomic portions. In some embodiments, the second subset of genomic portions overlaps with the first subset of genomic portions. In some embodiments, the second subset of genomic portions partially overlaps with the first subset of genomic portions. In some embodiments, the second subset of genomic portions includes fewer genomic portions than the first subset of genomic portions. In some embodiments, the second subset of genomic portions includes more genomic portions than the first subset of genomic portions.
[0044] In some embodiments, the methods herein include classifying the presence or absence of microduplications in subchromosomal regions. The microduplications can be duplications in chromosomes selected from chromosome 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, the X chromosome, and the Y chromosome. In some embodiments, the methods herein include classifying the presence or absence of microdeletions in subchromosomal regions. The microdeletions can be deletions in chromosomes selected from chromosome 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, the X chromosome, and the Y chromosome. In some embodiments, the microdeletions are deletions in a genomic region or a part of a genomic region selected from 1p36, 22q11.2, 15q11-13, 8q23.2-24.1, 11q24.1, 4p13.3, 17p13.3, and 7q11.23. In some embodiments, the microdeletions or microduplications are associated with a disease or syndrome. Examples of syndromes that can be associated with certain microdeletions and / or microduplications include 1p36 syndrome, DiGeorge syndrome, Prader-Willi syndrome, Angelman syndrome, Langer-Giedion syndrome, Jacobsen syndrome, Wolf-Hirschhorn syndrome, Miller-Dieker syndrome, and Williams-Beuren syndrome. A non-limiting list of known and / or possible associations between copy number variations in certain genomic regions and syndromes is provided in Table 1 below. [Table 1]
[0045] In some embodiments, copy number variations in subchromosomal regions are characterized by their size (i.e., length). The length of a copy number variation in a subchromosomal region refers to the number of consecutive nucleotide bases that are deleted (e.g., in the case of a microdeletion) or duplicated (e.g., in the case of a microduplication). In some embodiments, the length of a copy number variation in a subchromosomal region is about 1 megabase or less. For example, the length of a copy number variation in a subchromosomal region can be about 900 kilobases (kb), 800 kb, 700 kb, 600 kb, 500 kb, 400 kb, 300 kb, 200 kb, or 100 kb. In some embodiments, the length of a copy number variation in a subchromosomal region is from about 1 megabase to about 40 megabases. For example, the length of a copy number variation in a subchromosomal region can be from about 1 megabase to about 2 megabases, 1 megabase to about 3 megabases, 1 megabase to about 4 megabases, 1 megabase to about 5 megabases, 1 megabase to about 6 megabases, 1 megabase to about 7 megabases, 1 megabase to about 8 megabases, 1 megabase to about 9 megabases, 1 megabase to about 10 megabases, 1 megabase to about 11 megabases, 1 megabase to about 12 megabases, 1 megabase to about 13 megabases, 1 megabase to about 14 megabases, 1 megabase to about 15 megabases, 1 megabase to about 16 megabases, 1 megabase to about 17 megabases, 1 megabase to about 18 megabases, 1 megabase to about 19 megabases, 1 megabase to about 20 megabases, 1 megabase to about 25 megabases, 1 megabase to about 30 megabases, 1 megabase to about 35 megabases, or 1 megabase to about 40 megabases. In some embodiments, the length of a copy number variation in a subchromosomal region is from about 1 megabase to about 20 megabases. In some embodiments, the length of a copy number variation in a subchromosomal region is from about 1 megabase to about 10 megabases. In some embodiments, the length of a copy number variation in a subchromosomal region is from about 1 megabase to about 7 megabases.
[0046] In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with a sensitivity of at least about 70%. For example, the presence or absence of a copy number variation in a subchromosomal region for a test sample can be classified with a sensitivity of at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with a sensitivity of at least about 75%. In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with a sensitivity of at least about 80%. In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with a sensitivity of at least about 85%. In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with a sensitivity of at least about 90%. In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with a sensitivity of at least about 95%. In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with a sensitivity of at least about 97%.
[0047] In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with at least about 90% specificity. For example, the presence or absence of a copy number variation in a subchromosomal region for a test sample can be classified with at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9% or 100% specificity. In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with at least about 99% specificity. In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with at least about 99.9% specificity. In some embodiments, the presence or absence of a copy number variation in a subchromosomal region for a test sample is classified with about 100% specificity.
[0048] In some embodiments, the nucleic acid in the test sample is derived from a test subject. In some embodiments, the nucleic acid in the test sample comprises cell-free circulating nucleic acid. In some embodiments, the cell-free circulating nucleic acid is derived from the plasma or serum of the test subject. In some embodiments, the test subject is a male. In some embodiments, the test subject is a human male. In some embodiments, the test subject is a female. In some embodiments, the test subject is a human female. In some embodiments, the test subject is a pregnant woman. In some embodiments, the nucleic acid in the test sample comprises maternal nucleic acid and fetal nucleic acid. In some embodiments, the ratio of fetal nucleic acid in the test sample is less than about 25%. For example, the ratio of fetal nucleic acid in the test sample can be about 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or 1%. In some embodiments, the ratio of fetal nucleic acid in the test sample is less than about 10%. In some embodiments, the ratio of fetal nucleic acid in the test sample is less than about 5%. In some embodiments, the test subject is a cancer patient or a subject being tested or screened for cancer. In some embodiments, the nucleic acid in the test sample comprises patient / host nucleic acid and nucleic acid from a tumor or cancer cells. In some embodiments, the ratio of tumor / cancer nucleic acid in the test sample is less than about 25%. For example, the ratio of tumor / cancer nucleic acid in the test sample can be about 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2% or 1%. In some embodiments, the ratio of tumor / cancer nucleic acid in the test sample is less than about 10%. In some embodiments, the ratio of tumor / cancer nucleic acid in the test sample is less than about 5%.
[0049] Sample Systems, methods, and products for analyzing nucleic acids are provided herein. In some embodiments, nucleic acid fragments in a mixture of nucleic acid fragments are analyzed. The nucleic acid fragments may be referred to as nucleic acid templates, and these terms may be used interchangeably herein. The mixture of nucleic acids can include two or more nucleic acid fragment species having the same or different nucleotide sequences, different fragment lengths, different origins (e.g., genomic origin, fetal origin vs. maternal origin, cell or tissue origin, cancer origin vs. non-cancer origin, tumor origin vs. non-tumor origin, sample origin, subject origin, etc.) or combinations thereof.
[0050] The nucleic acids or nucleic acid mixtures used in the systems, methods, and products described herein are often isolated from a sample obtained from a subject (e.g., a test subject). The subject can be any living or non-living entity, including but not limited to humans, non-human animals, plants, bacteria, fungi, protests, or pathogens. Any human or non-human animal can be selected, such as, for example, mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, cattle (e.g., cows), horses (e.g., horses), goats and sheep (e.g., sheep, goats), pigs (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), bears, poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. The subject can be a male or female entity (e.g., female, pregnant woman). The subject can be of any age (e.g., embryo, fetus, infant, child, adult). The subject can be a cancer patient, a patient suspected of having cancer, a patient in remission, a patient with a family history of cancer, and / or a subject undergoing cancer screening. In some embodiments, the test subject is a female entity. In some embodiments, the test subject is a human female. In some embodiments, the test subject is a male entity. In some embodiments, the test subject is a human male.
[0051] Nucleic acids can be isolated from any suitable biological specimen or sample (e.g., a test sample). A sample or test sample can be any specimen isolated or obtained from a subject or a part thereof (e.g., a human subject, a pregnant woman, a cancer patient, a fetus, a tumor). Non-limiting examples of specimens include blood or blood products (e.g., serum, plasma, etc.), cord blood, chorionic villi, amniotic fluid, cerebrospinal fluid, bone marrow fluid, lavage fluids (e.g., bronchoalveolar lavage fluid, gastric lavage fluid, peritoneal lavage fluid, tube lavage fluid, ear lavage fluid, arthroscopic lavage fluid), biopsy samples (e.g., pre-implantation embryos; cancer biopsy materials), celocentesis samples, cells (blood cells, placental cells, embryonic cells, or fetal cells, nucleated fetal cells or fetal cellular remnants, normal cells, abnormal cells (e.g., cancer cells)) or portions thereof (e.g., mitochondria, nuclei, extracts, etc.), washings of the female reproductive tract, urine, feces, sputum, saliva, nasal mucosa, prostatic fluid, lavage fluid, semen, lymph fluid, bile, tears, sweat, breast milk, milk, etc. or combinations thereof, and fluids or tissues derived from a subject. In some embodiments, the biological sample is a cervical swab from a subject. The fluid or tissue sample from which the nucleic acid is extracted may be cell-free (e.g., acellular). In some embodiments, the fluid or tissue sample may contain cellular elements or cellular remnants. In some embodiments, fetal cells or cancer cells may be included in the sample.
[0052] The sample can be a liquid sample. The liquid sample can contain extracellular nucleic acids (e.g., cell-free circulating DNA). Non-limiting examples of liquid samples include blood or blood products (e.g., serum, plasma, etc.), urine, biopsy samples (e.g., liquid biopsy materials for detecting cancer), the liquid samples described above, etc. or combinations thereof. In certain embodiments, the sample is a liquid biopsy material, which broadly refers to the evaluation of a liquid sample from a subject for the presence or absence, progression or remission of a disease (e.g., cancer). Liquid biopsy materials can be used together with solid biopsy materials (e.g., tumor biopsy materials) or as an alternative to solid biopsy materials. In certain cases, extracellular nucleic acids are analyzed in the liquid biopsy material.
[0053] In some embodiments, the biological sample can be blood, plasma, or serum. The term "blood" encompasses whole blood, blood products, or any fraction of blood, such as serum, plasma, buffy coat, etc., as conventionally defined. Blood or its fractions often contain nucleosomes. Nucleosomes contain nucleic acids and can be cell-free or intracellular. Blood also includes the buffy coat. The buffy coat can be isolated by using a ficoll gradient. The buffy coat can contain white blood cells (e.g., leukocytes, T cells, B cells, platelets, etc.). Plasma refers to the fraction of whole blood resulting from centrifugation of anticoagulated blood. Serum refers to the watery portion of the fluid remaining after a blood sample has coagulated. Fluid or tissue samples are often collected according to standard protocols commonly followed by hospitals or clinics. In the case of blood, an appropriate amount of peripheral blood (e.g., 3 - 40 milliliters, 5 - 50 milliliters) is often collected, which can be stored according to standard procedures before or after preparation.
[0054] Analysis of nucleic acids found in the blood of a subject can be performed, for example, using whole blood, serum, or plasma. Analysis of fetal DNA found in the blood of a mother can be performed, for example, using whole blood, serum, or plasma. Analysis of tumor DNA found in the blood of a patient can be performed, for example, using whole blood, serum, or plasma. Methods for preparing serum or plasma from blood obtained from a subject (e.g., a maternal subject; a cancer patient) are known. For example, the blood of a subject (e.g., the blood of a pregnant woman; the blood of a cancer patient) can be placed in a tube containing EDTA or a dedicated commercial product such as Vacutainer SST (Becton Dickinson, Franklin Lakes, N.J.) to prevent blood clotting, and then plasma can be obtained from the whole blood by centrifugation. Serum can be obtained with or without blood clotting after centrifugation. When using centrifugation, the centrifugation is usually performed at an appropriate speed, for example, 1,500 to 3,000 × g, but is not limited thereto. Plasma or serum can be subjected to a further centrifugation step and then transferred to a new tube for nucleic acid extraction. In addition to the cell-free portion of whole blood, nucleic acids can also be recovered from the cell fraction concentrated in the buffy coat portion that can be obtained after centrifugation of a whole blood sample from a subject and removal of plasma.
[0055] Samples can be heterogeneous. For example, a sample can contain more than one cell type and / or one or more nucleic acid species. In some cases, a sample can contain (i) fetal and maternal cells, (ii) cancer and non-cancer cells, and / or (iii) pathogenic and host cells. In some cases, a sample can contain (i) nucleic acids from cancer and non-cancer, (ii) nucleic acids from pathogens and hosts, (iii) nucleic acids from the fetus and the mother, and / or more generally, (iv) mutant and wild-type nucleic acids. In some cases, a sample can contain minority and majority nucleic acid species as described in more detail below. In some cases, a sample can contain cells and / or nucleic acids from a single subject or can contain cells and / or nucleic acids from multiple subjects.
[0056] Cell type As used herein, "cell type" refers to a type of cell that can be distinguished from another type of cell. Extracellular nucleic acids can contain nucleic acids from several different cell types. Non-limiting examples of cell types that can lead to nucleic acids in circulating cell-free nucleic acids include liver cells (e.g., hepatocytes), lung cells, spleen cells, pancreatic cells, colon cells, skin cells, bladder cells, eye cells, brain cells, esophageal cells, head cells, neck cells, ovarian cells, testicular cells, prostate cells, placental cells, epithelial cells, endothelial cells, fat cells, kidney / renal cells, heart cells, muscle cells, blood cells (e.g., leukocytes), central nervous system (CNS) cells, etc. and combinations of the foregoing cells. In some embodiments, cell types that lead to nucleic acids in the circulating cell-free nucleic acids to be analyzed include leukocytes, endothelial cells, and hepatocyte liver cells. As described in more detail herein, various cell types can be screened as part of identifying and selecting loci of nucleic acids for which the marker status is the same or substantially the same for cell types in subjects with medical conditions and cell types in subjects without medical conditions.
[0057] Certain cell types may remain the same or substantially the same in subjects with a medical condition and in subjects without a medical condition. In non-limiting examples, the number of live or viable cells of a particular cell type may be decreased in a cell degeneration condition, and the living viable cells are not modified or are not significantly modified in the subject having the medical condition.
[0058] Certain cell types may be modified as part of a medical condition and may have one or more characteristics different from their original state. In non-limiting examples, a particular cell type may, as part of a cancer condition, proliferate at a rate faster than normal, may become cancerous into cells with different morphologies, may become cancerous into cells expressing one or more different cell surface markers, and / or may become part of a tumor. In embodiments where a particular cell type (i.e., a progenitor cell) is modified as part of a medical condition, the state of the marker for each of one or more markers being assayed is often the same or substantially the same for that particular cell type in a subject having the medical condition and for that particular cell type in a subject without the medical condition. Thus, the term "cell type" may relate to the type of cell in a subject without a particular medical condition and to the modified version of that cell in a subject having the medical condition. In some embodiments, "cell type" is only the progenitor cell and not the modified version arising from the progenitor cell. "Cell type" may relate to the progenitor cell and to the modified cells arising from the progenitor cell. In such embodiments, the state of the marker for the marker being assayed is often the same or substantially the same for the cell type in a subject having a medical condition and for the cell type in a subject without the medical condition.
[0059] In certain embodiments, the cell type is a cancer cell. Certain types of cancer cells include, for example, leukemia cells (e.g., acute myeloid leukemia, acute lymphoblastic leukemia, chronic myeloid leukemia, chronic lymphoblastic leukemia); cancerous kidney / renal cells (e.g., renal cell carcinoma (clear cell, papillary type 1, papillary type 2, chromophobe, oncocytoma, collecting duct), renal adenocarcinoma, adrenal tumor, Wilms tumor, transitional cell carcinoma); brain tumor cells (e.g., acoustic neuroma, astrocytoma (grade I: pilocytic astrocytoma, grade II: low-grade astrocytoma, grade III: anaplastic astrocytoma, grade IV: glioblastoma multiforme (GBM)), chordoma, CNS lymphoma, craniopharyngioma, glioma (brainstem glioma, ependymoma, mixed glioma, optic nerve glioma, subependymoma), medulloblastoma, meningioma, metastatic brain tumor, oligodendroglioma, pituitary tumor, primitive neuroectodermal tumor (PNET), schwannoma, juvenile pilocytic astrocytoma (JPA), pineal tumor, rhabdoid tumor).
[0060] Different cell types can be distinguished by any suitable characteristic, such as one or more different cell surface markers, one or more different morphological features, one or more different functions, one or more different protein (e.g., histone) modifications, and one or more different nucleic acid markers, but are not limited thereto. Non-limiting examples of nucleic acid markers include single nucleotide polymorphisms (SNPs), methylation status of nucleic acid loci, short tandem repeats, insertions (e.g., microinsertions), deletions (microdeletions), etc. and combinations thereof. Non-limiting examples of protein (e.g., histone) modifications include acetylation, methylation, ubiquitination, phosphorylation, SUMOylation, etc. and combinations thereof.
[0061] As used herein, the term "related cell type" refers to a cell type having a plurality of characteristics in common with another cell type. In related cell types, 75% or more of the cell surface markers may be common to that cell type (e.g., about 80%, 85%, 90% or 95% or more of the cell surface markers are common to the related cell type).
[0062] nucleic acid Methods for analyzing nucleic acids are provided herein. The terms "nucleic acid", "nucleic acid molecule", "nucleic acid fragment", and "nucleic acid template" may be used interchangeably throughout the present disclosure. These terms refer to, for example, DNA (e.g., complementary DNA (cDNA), genomic DNA (gDNA), etc.), RNA (e.g., messenger RNA (mRNA), small interfering RNA (siRNA), ribosomal RNA (rRNA), tRNA, microRNA, RNA highly expressed by the fetus or placenta, etc.), and / or DNA analogs or RNA analogs (e.g., base analogs, sugar analogs, and / or those containing non-natural backbones, etc.), RNA / DNA hybrids, and nucleic acids of any composition from peptide nucleic acids (PNA), all of which can be in single-stranded or double-stranded form and, unless otherwise limited, can function in a manner similar to naturally occurring nucleotides and can include known analogs of naturally occurring nucleotides. Nucleic acids can be, in certain embodiments, plasmids, phages, viruses, bacteria, autonomously replicating sequences (ARS), mitochondria, centromeres, artificial chromosomes, chromosomes, or other nucleic acids that can replicate or can be replicated in vitro or in a host cell, cell, cell nucleus, or cytoplasm of a cell, or can be derived therefrom. In some embodiments, the template nucleic acid can be derived from a single chromosome (e.g., a nucleic acid sample can be derived from one chromosome of a sample obtained from a diploid organism). Unless specifically limited, the term encompasses nucleic acids having binding properties similar to the reference nucleic acid and being metabolized in a manner similar to naturally occurring nucleotides, including known analogs of naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence implicitly encompasses its conservatively modified variants (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNP), and complementary sequences, as well as the explicitly shown sequences. Specifically, degenerate codon substitutions can be achieved by creating a sequence in which the third position of one or more selected (or all) codons is substituted with a mixed base and / or a deoxyinosine residue.The term "nucleic acid" is used interchangeably with locus, gene, cDNA, and mRNA encoded by a gene. This term includes single-stranded polynucleotides ("sense" or "antisense", "plus" strand or "minus" strand, "forward" reading frame or "reverse" reading frame) and double-stranded polynucleotides as equivalents, derivatives, variants and analogs of RNA or DNA synthesized from nucleotide analogs, and may also include single-stranded polynucleotides. The term "gene" refers to a region of DNA involved in the production of a polypeptide chain; this term generally includes regions before and after the coding region (leader and trailer) involved in the transcription / translation of the gene product and the control of transcription / translation, as well as intervening sequences (introns) between individual coding regions (exons). Nucleotides or bases generally refer to the purine and pyrimidine molecular units of nucleic acids (e.g., adenine (A), thymine (T), guanine (G) and cytosine (C)). In the case of RNA, the base thymine is replaced by uracil. The length or size of a nucleic acid can be expressed as the number of bases.
[0063] A nucleic acid can be single-stranded or double-stranded. For example, single-stranded DNA can be prepared, for example, by denaturing double-stranded DNA by treatment with heat or alkali. In certain embodiments, the nucleic acid is a D-loop structure formed by strand invasion of a double-stranded DNA molecule by an oligonucleotide or DNA-like molecule, such as a peptide nucleic acid (PNA). The formation of the D-loop can be facilitated using methods known in the art, for example, by the addition of E. coli RecA protein and / or by changing the salt concentration.
[0064] The nucleic acids provided for the processes described herein can include nucleic acids from one sample or two or more samples (e.g., one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, eleven or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, sixteen or more, seventeen or more, eighteen or more, nineteen or more or twenty or more samples).
[0065] The nucleic acids can be obtained from one or more sources (e.g., biological samples, blood, cells, serum, plasma, buffy coat, urine, lymph, skin, soil, etc.) by methods known in the art. Any suitable method can be used to isolate, extract and / or purify DNA from biological samples (e.g., blood or blood products), non-limiting examples of which include methods of DNA preparation (e.g., those described in Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3d ed., 2001), various commercially available reagents or kits, such as Qiagen's QIAamp Circulating Nucleic Acid Kit, QiaAmp DNA Mini Kit or QiaAmp DNA Blood Mini Kit (Qiagen, Hilden, Germany), GenomicPrep TM Blood DNA Isolation Kit (Promega, Madison, Wis.) and GFX TM Genomic Blood DNA Purification Kit (Amersham, Piscataway, N.J.), etc. or combinations thereof can be mentioned.
[0066] In some embodiments, the nucleic acid is extracted from cells using a cell lysis procedure. Cell lysis procedures and reagents are known in the art and generally can be performed by chemical lysis methods (e.g., detergents, hypotonic solutions, enzymatic procedures, etc. or combinations thereof), physical lysis methods (e.g., French press, sonication, etc.) or lysis methods by electrolysis. Any suitable lysis procedure can be used. For example, chemical methods generally use a lysing agent to disrupt cells, extract nucleic acids from the cells, and then treat with chaotropic salts. Physical methods such as grinding after freeze / thaw, use of a cell press, etc. are also useful. In some cases, high salt lysis procedures and / or alkaline lysis procedures can be used.
[0067] In certain embodiments, the nucleic acid can include extracellular nucleic acid. As used herein, the term "extracellular nucleic acid" can refer to nucleic acids isolated from a source substantially free of cells and is also referred to as "cell-free" nucleic acid, "circulating cell-free nucleic acid" (e.g., CCF fragments, ccfDNA) and / or "cell-free circulating nucleic acid". Extracellular nucleic acid can be present in blood (e.g., the blood of a human subject) and can be obtained therefrom. Extracellular nucleic acid often does not contain detectable cells and may contain cellular elements or cell remnants. Non-limiting examples of cell-free sources for extracellular nucleic acid are blood, plasma, serum and urine. As used herein, the term "obtaining cell-free circulating sample nucleic acid" includes obtaining the sample directly (e.g., collecting a sample, e.g., a test sample) or obtaining the sample from another person who has collected the sample. Without being limited by theory, extracellular nucleic acid can be the product of apoptosis and cell destruction of cells, which often results in extracellular nucleic acid having a range of lengths (e.g., a "ladder"). In some embodiments, the sample nucleic acid from a test subject is circulating cell-free nucleic acid. In some embodiments, the circulating cell-free nucleic acid is derived from the plasma or serum of the test subject.
[0068] Extracellular nucleic acids can contain various nucleic acid species and are thus referred to herein as "heterogeneous" in certain embodiments. For example, the serum or plasma of a person with cancer can contain nucleic acids from cancer cells (e.g., tumors, neoplasms) and nucleic acids from non-cancer cells. In another example, the serum or plasma from a pregnant woman can contain maternal nucleic acids and fetal nucleic acids. In some cases, the cancer nucleic acids or fetal nucleic acids can be about 5% to about 50% of the total nucleic acids (e.g., about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48 or 49% of the total nucleic acids are cancer nucleic acids or fetal nucleic acids).
[0069] At least two different nucleic acid species can be present in different amounts as extracellular nucleic acids, and they are sometimes referred to as minor species and major species. In certain cases, the minor species of nucleic acids are derived from diseased cell types (e.g., cancer cells, wasting cells, cells attacked by the immune system). In certain embodiments, genetic mutations or genetic changes (e.g., copy number changes, copy number variations, single nucleotide changes, single nucleotide variations, chromosomal changes and / or translocations) are determined for the minor species of nucleic acids. In certain embodiments, genetic mutations or genetic changes are determined for the major species of nucleic acids. Generally, the terms "minor" or "major" are not intended to be strictly defined at any point. In one aspect, the nucleic acids considered "minor" can have an abundance of, for example, at least about 0.1% to less than 50% of the total nucleic acids in the sample. In some embodiments, the minor nucleic acids can have an abundance of at least about 1% to about 40% of the total nucleic acids in the sample. In some embodiments, the minor nucleic acids can have an abundance of at least about 2% to about 30% of the total nucleic acids in the sample. In some embodiments, the minor nucleic acids can have an abundance of at least about 3% to about 25% of the total nucleic acids in the sample. For example, the minor nucleic acids can have an abundance of about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29% or 30% of the total nucleic acids in the sample. In some cases, the minor species of extracellular nucleic acids can be about 1% to about 40% of the total nucleic acids (e.g., about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39% or 40% of that nucleic acid is the minor species nucleic acid). In some embodiments, the minor nucleic acids are extracellular DNA.In some embodiments, the minority nucleic acids are extracellular DNA derived from apoptotic tissues. In some embodiments, the minority nucleic acids are extracellular DNA derived from tissues affected by cell proliferative disorders. In some embodiments, the minority nucleic acids are extracellular DNA derived from tumor cells. In some embodiments, the minority nucleic acids are extracellular fetal DNA.
[0070] In another aspect, the nucleic acids considered to be "majority" can have an abundance of, for example, more than 50% to about 99.9% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acids can have an abundance of at least about 60% to about 99% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acids can have an abundance of at least about 70% to about 98% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acids can have an abundance of at least about 75% to about 97% of the total nucleic acids in the sample. For example, the majority nucleic acids can have an abundance of at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of the total nucleic acids in the sample. In some embodiments, the majority nucleic acids are extracellular DNA. In some embodiments, the majority nucleic acids are extracellular maternal DNA. In some embodiments, the majority nucleic acids are DNA derived from healthy tissues. In some embodiments, the majority nucleic acids are DNA derived from non-tumor cells.
[0071] In some embodiments, the minority species of extracellular nucleic acids are of a length of about 500 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species of nucleic acids are of a length of about 500 base pairs or less). In some embodiments, the minority species of extracellular nucleic acids are of a length of about 300 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species of nucleic acids are of a length of about 300 base pairs or less). In some embodiments, the minority species of extracellular nucleic acids are of a length of about 250 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species of nucleic acids are of a length of about 250 base pairs or less). In some embodiments, the minority species of extracellular nucleic acids are of a length of about 200 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species of nucleic acids are of a length of about 200 base pairs or less). In some embodiments, the minority species of extracellular nucleic acids are of a length of about 150 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species of nucleic acids are of a length of about 150 base pairs or less). In some embodiments, the minority species of extracellular nucleic acids are of a length of about 100 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species of nucleic acids are of a length of about 100 base pairs or less). In some embodiments, the minority species of extracellular nucleic acids are of a length of about 50 base pairs or less (e.g., about 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100% of the minority species of nucleic acids are of a length of about 50 base pairs or less).
[0072] Nucleic acids can be provided for performing the methods described herein, with or without processing of the sample containing the nucleic acid. In some embodiments, the nucleic acids are provided for performing the methods described herein after processing of the sample containing the nucleic acid. For example, the nucleic acids can be extracted from, isolated from, purified from, partially purified from, or amplified from a sample. The term "isolated," as used herein, refers to a nucleic acid that has been removed from its original environment (e.g., its natural environment if it occurs naturally, or the host cell if it is exogenously expressed), and thus has been altered from its original environment by human intervention (e.g., "by the hand of man"). The term "isolated nucleic acid" can refer to a nucleic acid that has been removed from a subject (e.g., a human subject). An isolated nucleic acid can be provided with fewer non-nucleic acid components (e.g., proteins, lipids) than were present in the source sample. A composition containing an isolated nucleic acid can be free of non-nucleic acid components in an amount of about 50% to greater than 99%. A composition containing an isolated nucleic acid can be free of non-nucleic acid components in an amount of about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater than 99%. The term "purified," as used herein, can refer to a provided nucleic acid that contains fewer non-nucleic acid components (e.g., proteins, lipids, carbohydrates) than were present before the nucleic acid was subjected to a purification procedure. A composition containing a purified nucleic acid can be free of other non-nucleic acid components in an amount of about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater than 99%. The term "purified," as used herein, can refer to a provided nucleic acid that contains fewer nucleic acid species than the sample source from which the nucleic acid is derived. A composition containing a purified nucleic acid can be free of other nucleic acid species in an amount of about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater than 99%. For example, fetal nucleic acids can be purified from a mixture containing maternal and fetal nucleic acids.In certain instances, small fragments of fetal nucleic acids (e.g., 30-500 bp fragments) can be purified or partially purified from a mixture containing both fetal and maternal nucleic acid fragments. In certain instances, nucleosomes containing smaller fragments of fetal nucleic acids can be purified from a mixture of larger nucleosome complexes containing larger fragments of maternal nucleic acids. In certain instances, nucleic acids of cancer cells can be purified from a mixture containing nucleic acids of cancer cells and non-cancer cells. In certain instances, nucleosomes containing small fragments of nucleic acids of cancer cells can be purified from a mixture of larger nucleosome complexes containing larger fragments of non-cancer nucleic acids. In some embodiments, the nucleic acids are provided for performing the methods described herein without prior processing of the sample containing the nucleic acids. For example, the nucleic acids can be analyzed directly from the sample without prior extraction, purification, partial purification, and / or amplification.
[0073] In some embodiments, a nucleic acid, e.g., a nucleic acid of a cell, is sheared or cleaved before, during, or after the methods described herein. The terms "shearing" or "cleaving" generally refer to a procedure or condition by which a nucleic acid molecule (e.g., a nucleic acid template gene molecule or an amplification product thereof) can be separated into two (or more) smaller nucleic acid molecules. Such shearing or cleavage can be sequence-specific, base-specific, or non-specific and can be accomplished by any of a variety of methods, reagents, or conditions, including, for example, chemical, enzymatic, physical shearing (e.g., physical fragmentation). The sheared or cleaved nucleic acids can have a nominal length, average length, or average value of the length of about 5 to about 10,000 base pairs, about 100 to about 1,000 base pairs, about 100 to about 500 base pairs, or about 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, or 9000 base pairs.
[0074] The sheared or cleaved nucleic acid can be produced by suitable methods, non-limiting examples of which include physical methods (e.g., shearing, such as sonication, French press, heating, UV irradiation, etc.), enzymatic processes (e.g., enzymatic cleavage agents, such as suitable nucleases, suitable restriction enzymes, suitable methylation-sensitive restriction enzymes), chemical methods (e.g., alkylation, DMS, piperidine, acid hydrolysis, base hydrolysis, heating, etc. or combinations thereof), the processes described in U.S. Patent Application Publication No. 2005 / 0112590, etc. or combinations thereof. The average length, mean length or nominal length of the resulting nucleic acid fragments can be controlled by selecting an appropriate method for producing the fragments.
[0075] The term "amplified", as used herein, refers to subjecting a target nucleic acid in a sample to a process that linearly or exponentially generates an amplicon nucleic acid or a part thereof having the same or substantially the same nucleotide sequence as the target nucleic acid. In certain embodiments, the term "amplified" refers to methods including polymerase chain reaction (PCR). In certain cases, the amplification product may contain one or more nucleotides more than the amplified nucleotide region of the nucleic acid template sequence (e.g., the primer may contain "extra" nucleotides, such as a transcription initiation sequence, in addition to nucleotides complementary to the nucleic acid template gene molecule, resulting in an amplification product that contains "extra" nucleotides or nucleotides not corresponding to the amplified nucleotide region of the nucleic acid template gene molecule).
[0076] The nucleic acid can also be exposed to a process that modifies certain nucleotides in the nucleic acid before providing the nucleic acid for the methods described herein. For example, a process that selectively modifies the nucleic acid based on the methylation state of the nucleotides in the nucleic acid can be applied to the nucleic acid. Further, conditions such as high temperature, ultraviolet light, x-rays, etc. can induce changes in the sequence of the nucleic acid molecule. The nucleic acid can be provided in any suitable form useful for performing sequence analysis.
[0077] Nucleic Acid Enrichment In some embodiments, nucleic acids (e.g., extracellular nucleic acids) are enriched for or relatively enriched for a subpopulation or species of nucleic acids. Subpopulations of nucleic acids can include, for example, fetal nucleic acids, maternal nucleic acids, cancer nucleic acids, patient nucleic acids, nucleic acids comprising fragments of a particular length or range of lengths, or nucleic acids derived from a particular genomic region (e.g., a single chromosome, a set of chromosomes, and / or a particular chromosomal region). Such enriched samples can be used in conjunction with the methods provided herein. Thus, in certain embodiments, the methods of the technology include an additional step of enriching for a subpopulation of nucleic acids in a sample, such as cancer nucleic acids or fetal nucleic acids. In certain embodiments, methods for measuring the cancer cell nucleic acid ratio or fetal ratio can also be used to enrich for cancer nucleic acids or fetal nucleic acids. In certain embodiments, nucleic acids from normal tissue (e.g., non-cancer cells) are selectively removed (partially, substantially, almost completely, or completely) from the sample. In certain embodiments, maternal nucleic acids are selectively removed (partially, substantially, almost completely, or completely) from the sample. In certain embodiments, quantitative sensitivity can be improved by enriching for certain low-copy-number species of nucleic acids (e.g., cancer nucleic acids or fetal nucleic acids). Methods for enriching a sample for a particular nucleic acid species are described, for example, in U.S. Patent No. 6,927,028, International Patent Application Publication No. WO2007 / 140417, International Patent Application Publication No. WO2007 / 147063, International Patent Application Publication No. WO2009 / 032779, International Patent Application Publication No. WO2009 / 032781, International Patent Application Publication No. WO2010 / 033639, International Patent Application Publication No. WO2011 / 034631, International Patent Application Publication No. WO2006 / 056480, and International Patent Application Publication No. WO2011 / 143659, the entire contents of each of which, including all text, tables, formulas, and drawings, are incorporated herein by reference.
[0078] In some embodiments, the nucleic acids are enriched for certain target fragment species and / or reference fragment species. In certain embodiments, the nucleic acids are enriched for a particular nucleic acid fragment length or range of fragment lengths using one or more separation methods based on length as described below. In certain embodiments, the nucleic acids are enriched for fragments from selected genomic regions (e.g., chromosomes) using one or more separation methods based on sequence as described herein and / or known in the art.
[0079] Non-limiting examples of methods for enriching a subpopulation of nucleic acids in a sample include methods that utilize epigenetic differences between nucleic acid species (e.g., a method for enriching methylated fetal nucleic acids as described in U.S. Patent Application Publication No. 2010 / 0105049, incorporated herein by reference); polymorphic sequence approaches enhanced by restriction endonucleases (e.g., the method described in U.S. Patent Application Publication No. 2009 / 0317818, incorporated herein by reference); selective enzymatic digestion approaches; massively parallel signature sequencing (MPSS) approaches; amplification (e.g., PCR)-based approaches (e.g., locus-specific amplification methods, multiplex SNP allele PCR approaches; universal amplification methods); pull-down approaches (e.g., biotinylated ultramer pull-down methods); methods based on extension and ligation (e.g., extension and ligation of molecular inversion probes (MIPs)); and combinations thereof.
[0080] In some embodiments, the nucleic acid is enriched for fragments from a selected genomic region (e.g., a chromosome) using one or more sequence-based isolation methods described herein. Sequence-based isolation is generally based on nucleotide sequences that are present in the fragment of interest (e.g., the target fragment and / or the reference fragment) and substantially absent or present in only trace amounts (e.g., 5% or less) in other fragments of the sample. In some embodiments, sequence-based isolation can generate isolated target fragments and / or isolated reference fragments. The isolated target fragments and / or isolated reference fragments are often isolated from the remaining fragments in the nucleic acid sample. In certain embodiments, the isolated target fragments and the isolated reference fragments are also isolated from each other (e.g., isolated into separate assay compartments). In certain embodiments, the isolated target fragments and the isolated reference fragments are isolated together (e.g., isolated into the same assay compartment). In some embodiments, unbound fragments can be differentially removed, or degraded, or digested.
[0081] In some embodiments, a selective nucleic acid capture process is used to isolate target fragments and / or reference fragments from a nucleic acid sample. Commercially available nucleic acid capture systems include, for example, the Nimblegen Sequence Capture System (Roche NimbleGen, Madison, WI); the Illumina BEADARRAY platform (Illumina, San Diego, CA); the Affymetrix GENECHIP platform (Affymetrix, Santa Clara, CA); Agilent SureSelect Target Enrichment System (Agilent Technologies, Santa Clara, CA); and related platforms. Such methods typically involve hybridization of capture oligonucleotides with some or all of the nucleotide sequences of target or reference fragments, and may involve the use of solid-phase (e.g., solid-phase arrays) and / or solution-based platforms. Capture oligonucleotides (sometimes referred to as "baits") can be selected or designed to preferentially hybridize to nucleic acid fragments derived from selected genomic regions or loci (e.g., one of chromosomes 21, 18, 13, X or Y or a reference chromosome). In certain embodiments, hybridization-based methods (e.g., methods using oligonucleotide arrays) can be used to enrich nucleic acid sequences, the genes or regions of interest, from a particular chromosome (e.g., a potentially aneuploid chromosome, a reference chromosome or other chromosome of interest). Thus, in some embodiments, a nucleic acid sample is optionally enriched by capturing a subset of fragments, e.g., using capture oligonucleotides complementary to selected genes in the sample nucleic acid. In certain cases, the captured fragments are amplified. For example, captured fragments containing adapters can be amplified using primers complementary to the adapter oligonucleotides to form a set of amplified fragments indexed according to the adapter sequences. In some embodiments, nucleic acids are selected from a selected genomic region (e.g., a chromosome, a gene) by amplifying one or more regions of interest using oligonucleotides (e.g., PCR primers) complementary to sequences in fragments containing the region of interest or a portion thereof, thereby enriching for fragments.
[0082] In some embodiments, the nucleic acids are enriched for the length of a particular nucleic acid fragment, a range of lengths, or lengths below or above a particular threshold or cutoff, using one or more separation methods based on length. The length of a nucleic acid fragment generally refers to the number of nucleotides in that fragment. The length of a nucleic acid fragment is sometimes referred to as the size of the nucleic acid fragment. In some embodiments, the length-based separation method is performed without measuring the length of individual fragments. In some embodiments, the length-based separation method is performed with a method for measuring the length of individual fragments. In some embodiments, length-based separation refers to a size fractionation procedure by which all or a portion of a fractionated pool can be isolated (e.g., retained) and / or analyzed. Size fractionation procedures are known in the art (e.g., separation on an array, separation by a molecular sieve, separation by gel electrophoresis, separation by column chromatography (e.g., size exclusion column), and microfluidics-based approaches). In certain cases, length-based separation approaches can include, for example, selective sequence tagging approaches, circularization of fragments, chemical treatments (e.g., formaldehyde, polyethylene glycol (PEG) precipitation), mass spectrometry, and / or size-specific nucleic acid amplification).
[0083] Quantification of Nucleic Acids The amount of nucleic acid in a sample (e.g., concentration, relative amount, absolute amount, copy number, etc.) can be measured. In some embodiments, the amount of a minority nucleic acid in a nucleic acid (e.g., concentration, relative amount, absolute amount, copy number, etc.) is measured. In certain embodiments, the amount of a minority nucleic acid species in a sample is referred to as the "minority species ratio". In some embodiments, the "minority species ratio" refers to the ratio of a minority nucleic acid species in cell-free circulating nucleic acids in a sample obtained from a subject (e.g., a blood sample, a serum sample, a plasma sample, a urine sample).
[0084] The amount of minority nucleic acids in extracellular nucleic acids can be quantified and used with the methods provided herein. Thus, in certain embodiments, the methods described herein include an additional step of measuring the amount of minority nucleic acids. The amount of minority nucleic acids in a sample from a subject can be measured before or after the processing to prepare the sample nucleic acids. In certain embodiments, the amount of minority nucleic acids in the sample after the sample nucleic acids are processed and prepared is measured and that amount is used for further evaluation. In some embodiments, the outcome includes considering the minority species ratio in the sample nucleic acids (e.g., adjusting the count, removing the sample, generating a call or not generating a call).
[0085] Measurement of the minority species ratio can be performed before, during, or at any point in time prior to, or after a particular method described herein (e.g., detection of a genetic mutation or genetic change). For example, for performing a method of measuring a genetic mutation / genetic change with a particular sensitivity or specificity, a minority nucleic acid quantification method is performed before, during, or after the measurement of the genetic mutation / genetic change to identify those samples that contain minority nucleic acids that are about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25% or more. In some embodiments, samples that are measured to have a particular threshold amount of minority nucleic acids (e.g., minority nucleic acids that are about 15% or more; minority nucleic acids that are about 4% or more) are further analyzed, for example, for a genetic mutation / genetic change, or for the presence or absence of a genetic mutation / genetic change. In certain embodiments, for example, the measurement of a genetic mutation or genetic change is selected (e.g., selected and the patient is contacted) only for samples that have a particular threshold amount of minority nucleic acids (e.g., minority nucleic acids that are about 15% or more; minority nucleic acids that are about 4% or more).
[0086] In some embodiments, the amount of cancer cell nucleic acid in a nucleic acid (e.g., concentration, relative amount, absolute amount, copy number, etc.) is measured. In certain cases, the amount of cancer cell nucleic acid in a sample is referred to as the "ratio of cancer cell nucleic acid" and may be referred to as the "cancer ratio" or "tumor ratio". In some embodiments, the "ratio of cancer cell nucleic acid" refers to the ratio of cancer cell nucleic acid in cell-free circulating nucleic acid in a sample obtained from a subject (e.g., a blood sample, a serum sample, a plasma sample, a urine sample).
[0087] In some embodiments, the amount of fetal nucleic acid in a nucleic acid (e.g., concentration, relative amount, absolute amount, copy number, etc.) is measured. In certain embodiments, the amount of fetal nucleic acid in a sample is referred to as the "fetal ratio". In some embodiments, the "fetal ratio" refers to the ratio of fetal nucleic acid in cell-free circulating nucleic acid in a sample obtained from a pregnant woman (e.g., a blood sample, a serum sample, a plasma sample, a urine sample). Certain methods described herein or known in the art for measuring the fetal ratio can be used to measure the ratio of cancer cell nucleic acid and / or the minority species ratio.
[0088] In certain cases, the fetal fraction can be measured according to a marker specific to male fetuses (e.g., Y chromosome STR markers (e.g., DYS19, DYS385, DYS392 markers); RhD markers in RhD-negative females), according to the ratio of alleles of a polymorphic sequence, or according to one or more markers specific to fetal nucleic acids and not specific to maternal nucleic acids (e.g., differential epigenetic biomarkers (e.g., methylation) between the mother and the fetus or fetal RNA markers in maternal plasma (see, e.g., Lo, 2005, Journal of Histochemistry and Cytochemistry 53(3):293-296)). The measurement of the fetal fraction can be performed, for example, using a fetal quantity assay (FQA) as described in U.S. Patent Application Publication No. 2010 / 0105049, which is incorporated herein by reference. This type of assay makes it possible to detect and quantify fetal nucleic acids in a maternal sample based on the methylation status of the nucleic acids in the sample.
[0089] In certain embodiments, the minority fraction can be measured based on the ratio of alleles of a polymorphic sequence (e.g., single nucleotide polymorphism (SNP)), for example, using the method described in U.S. Patent Application Publication No. 2011 / 0224087, which is incorporated herein by reference. In such a method for measuring the fetal fraction, for example, nucleotide sequence reads for a maternal sample are obtained and the fetal fraction is measured by comparing the total number of nucleotide sequence reads that map to a first allele and the total number of nucleotide sequence reads that map to a second allele at an informative polymorphic site (e.g., SNP) in a reference genome.
[0090] The minority species ratio can be measured in some embodiments using methods that incorporate information obtained from chromosomal abnormalities, such as those described in International Patent Application Publication No. WO2014 / 055774, which is incorporated herein by reference. The minority species ratio can be measured in some embodiments using methods that incorporate information obtained from sex chromosomes, such as those described in U.S. Patent Application Publication Nos. 2013 / 0288244 and 2013 / 0338933, each of which is incorporated herein by reference.
[0091] The minority species ratio can be measured in some embodiments using methods that incorporate information on fragment length (e.g., analysis of fragment length ratio (FLR), analysis of fetal ratio statistic (FRS), such as those described in International Patent Application Publication No. 2013 / 177086, which is incorporated herein by reference). Cell-free fetal nucleic acid fragments are typically shorter than maternal-derived nucleic acid fragments (see, e.g., Chan et al. (2004) Clin. Chem. 50:88-92; Lo et al. (2010) Sci. Transl. Med. 2:61ra91). Thus, the fetal ratio can be measured in some embodiments by counting fragments below a threshold length and comparing that number to, for example, the number of fragments above a threshold length and / or the amount of total nucleic acid in the sample. Methods for counting nucleic acid fragments of a particular length are described in more detail in International Patent Application Publication No. WO2013 / 177086.
[0092] The minority species ratio can, in some embodiments, be measured according to a partially specific ratio estimation (e.g., as described in International Patent Application Publication No. WO2014 / 205401, incorporated herein by reference). Without being bound by theory, the amount of reads from fetal CCF fragments (e.g., fragments of a particular length or length range) often maps to portions with varying frequencies (e.g., within the same sample, e.g., within the same sequencing run). Also without being bound by theory, certain portions tend to have a similar presentation of reads from fetal CCF fragments (e.g., fragments of a particular length or length range) when compared across multiple samples, and that presentation correlates with a partially specific fetal ratio (e.g., the relative amount, percentage, or ratio of CCF fragments of fetal origin). Partially specific fetal ratio estimates are typically measured according to partially specific parameters and their relationship to those fetal ratios.
[0093] In some embodiments, the measurement of the minority species ratio (e.g., the ratio of cancer cell nucleic acids; the fetal ratio) is not required or not necessary for the identification of the presence or absence of a genetic mutation or genetic change. In some embodiments, the identification of the presence or absence of a genetic mutation or genetic change does not require the discrimination of the sequences of the majority nucleic acids and the minority nucleic acids. In certain embodiments, this is because the total contribution of both the minority and majority sequences in a particular chromosome, chromosomal portion, or part thereof is analyzed. In some embodiments, the identification of the presence or absence of a genetic mutation or genetic change does not rely on putative sequence information that can distinguish the minority nucleic acids from the majority nucleic acids.
[0094] Nucleic acid library In some embodiments, a nucleic acid library is a plurality of polynucleotide molecules (e.g., a sample of nucleic acids) that are prepared, assembled, and / or modified for a particular process, non-limiting examples of which include immobilization on a solid phase (e.g., a solid support, a flow cell, beads), enrichment, amplification, cloning, detection, and / or nucleic acid sequencing. In certain embodiments, the nucleic acid library is prepared before or during a sequencing process. The nucleic acid library (e.g., a sequencing library) can be prepared by suitable methods as known in the art. The nucleic acid library can be prepared by a targeted or non-targeted preparation process.
[0095] In some embodiments, a library of nucleic acids is modified to include a chemical moiety (e.g., a functional group) configured to immobilize the nucleic acid to a solid support. In some embodiments, a library of nucleic acids is modified to include a biomolecule (e.g., a functional group) and / or a member of a binding pair configured to immobilize the library to a solid support, non-limiting examples of which include thyroxine-binding globulin, steroid-binding protein, antibody, antigen, hapten, enzyme, lectin, nucleic acid, repressor, protein A, protein G, avidin, streptavidin, biotin, complement component C1q, nucleic acid-binding protein, receptor, carbohydrate, oligonucleotide, polynucleotide, complementary nucleic acid sequence, etc. and combinations thereof. Some examples of specific binding pairs include an avidin moiety and a biotin moiety; an antigenic epitope and an antibody or an immunologically reactive fragment thereof; an antibody and a hapten; a digoxigenin moiety and an anti-digoxigenin antibody; a fluorescein moiety and an anti-fluorescein antibody; an operator and a repressor; a nuclease and a nucleotide; a lectin and a polysaccharide; a steroid and a steroid-binding protein; an active compound and a receptor for the active compound; a hormone and a hormone receptor; an enzyme and a substrate; an immunoglobulin and protein A; an oligonucleotide or polynucleotide and its corresponding complementary strand; etc. or combinations thereof, but are not limited thereto.
[0096] In some embodiments, the nucleic acid library is modified to include one or more polynucleotides of known composition, non-limiting examples of which include identifiers (e.g., tags, index tags), capture sequences, labels, adapters, restriction enzyme sites, promoters, enhancers, origins of replication, stem loops, complimentary sequences (e.g., primer binding sites, annealing sites), suitable integration sites (e.g., transposons, viral integration sites), modified nucleotides, etc. or combinations thereof. The polynucleotides of known sequence can be added at suitable positions, e.g., at the 5' end, 3' end or within the nucleic acid sequence. The polynucleotides of known sequence can be of the same or different sequences. In some embodiments, the polynucleotides of known sequence are configured to hybridize to one or more oligonucleotides immobilized on a surface (e.g., a surface within a flow cell). For example, a nucleic acid molecule containing a known 5' sequence can hybridize to a first plurality of oligonucleotides, whereas a known 3' sequence can hybridize to a second plurality of oligonucleotides. In some embodiments, the nucleic acid library can include chromosome-specific tags, capture sequences, labels and / or adapters. In some embodiments, the nucleic acid library includes one or more detectable labels. In some embodiments, one or more detectable labels can be incorporated into the nucleic acid library at the 5' end, 3' end and / or any nucleotide position within the nucleic acids in the library. In some embodiments, the nucleic acid library includes hybridized oligonucleotides. In certain embodiments, the hybridized oligonucleotides are labeled probes. In some embodiments, the nucleic acid library includes hybridized oligonucleotide probes prior to immobilization on a solid phase.
[0097] In some embodiments, the polynucleotide of the known sequence comprises a universal sequence. The universal sequence is a specific nucleotide acid sequence integrated into two or more nucleic acid molecules or two or more subsets of nucleic acid molecules, where the universal sequence is the same for all molecules or subsets of molecules into which it is integrated. The universal sequence is often designed to hybridize to multiple different sequences and / or to amplify multiple different sequences using a single universal primer complementary to the universal sequence. In some embodiments, two (e.g., a pair) or more universal sequences and / or universal primers are used. The universal primer often contains the universal sequence. In some embodiments, an adapter (e.g., a universal adapter) contains the universal sequence. In some embodiments, one or more universal sequences are used to capture, identify, and / or detect multiple nucleic acid species or subsets of nucleic acids.
[0098] In certain embodiments of preparing a nucleic acid library (e.g., in certain sequencing by synthesis procedures), the nucleic acids are size selected and / or fragmented to lengths of several hundred base pairs or less (e.g., in preparation for library construction). In some embodiments, the library preparation is performed without fragmentation (e.g., when using cell-free DNA).
[0099] In certain embodiments, a ligation-based library preparation method is used (e.g., ILLUMINA TRUSEQ, Illumina, San Diego CA). Ligation-based library preparation methods often utilize adapter (e.g., methylated adapter) designs that can incorporate index sequences (e.g., sample index sequences that identify the origin of a sample relative to a nucleic acid sequence) in an initial ligation step and can often be used to prepare samples for single-read sequencing, paired-end sequencing, and multiplexed sequencing. For example, nucleic acids (e.g., fragmented nucleic acids or cell-free DNA) can have their ends repaired by a fill-in reaction, an exonuclease reaction, or a combination thereof. In some embodiments, the resulting blunt-ended repaired nucleic acids can then be extended by only a single nucleotide complementary to the single nucleotide overhang at the 3' end of the adapter / primer. Any nucleotide can be used for the extension / overhang nucleotide.
[0100] In some embodiments, the preparation of the nucleic acid library involves ligating adapter oligonucleotides (e.g., to sample nucleic acids, sample nucleic acid fragments, template nucleic acids). Adapter oligonucleotides are often complementary to flow cell anchors and may be used to immobilize the nucleic acid library on a solid support (e.g., the inner surface of a flow cell). In some embodiments, the adapter oligonucleotide includes an identifier, one or more sequencing primer hybridization sites (e.g., a sequence complementary to a universal sequencing primer, a single-end sequencing primer, a paired-end sequencing primer, a multiplexed sequencing primer, etc.) or a combination thereof (e.g., adapter / sequencing, adapter / identifier, adapter / identifier / sequencing). In some embodiments, the adapter oligonucleotide includes one or more of a primer annealing polynucleotide (e.g., for annealing to an oligonucleotide attached to a flow cell and / or a free amplification primer), an index polynucleotide (e.g., a sample index sequence for tracking nucleic acids from various samples; also referred to as a sample ID), and a barcode polynucleotide (e.g., a single molecule barcode (SMB) for tracking individual sample nucleic acid molecules amplified prior to sequencing; also referred to as a molecular barcode). In some embodiments, the primer annealing component of the adapter oligonucleotide includes one or more universal sequences (e.g., a sequence complementary to one or more universal amplification primers). In some embodiments, the index polynucleotide (e.g., a sample index; a sample ID) is a component of the adapter oligonucleotide. In some embodiments, the index polynucleotide (e.g., a sample index; a sample ID) is a component of the universal amplification primer sequence.
[0101] In some embodiments, an adapter oligonucleotide is designed to generate a library construct that includes one or more of a universal sequence, a molecular barcode, a sample ID sequence, a spacer sequence, and a sample nucleic acid sequence when used in combination with an amplification primer (e.g., a universal amplification primer). In some embodiments, an adapter oligonucleotide is designed to generate a library construct that includes an ordered combination of one or more of a universal sequence, a molecular barcode, a sample ID sequence, a spacer sequence, and a sample nucleic acid sequence when used in combination with a universal amplification primer. For example, the library construct can include a first universal sequence, followed by a second universal sequence, followed by a first molecular barcode, followed by a spacer sequence, followed by a template sequence (e.g., a sample nucleic acid sequence), followed by a spacer sequence, followed by a second molecular barcode, followed by a third universal sequence, followed by a sample ID, followed by a fourth universal sequence. In some embodiments, an adapter oligonucleotide is designed to generate a library construct for each strand of a template molecule (e.g., a sample nucleic acid molecule) when used in combination with an amplification primer (e.g., a universal amplification primer). In some embodiments, the adapter oligonucleotide is a double-stranded adapter oligonucleotide.
[0102] An identifier can be a suitable detectable label incorporated into or attached to a nucleic acid (e.g., a polynucleotide) that enables the detection and / or identification of the nucleic acid containing the identifier. In some embodiments, the identifier is incorporated into or attached to the nucleic acid during a sequencing method (e.g., by polymerase). Non-limiting examples of identifiers include nucleic acid tags, nucleic acid indexes or barcodes, radiolabels (e.g., isotopes), metal labels, fluorescent labels, chemiluminescent labels, phosphorescent labels, fluorophore quenchers, dyes, proteins (e.g., enzymes, antibodies or portions thereof, linkers, members of binding pairs), etc. or combinations thereof. In some embodiments, an identifier (e.g., a nucleic acid index or barcode) is a unique sequence, a known sequence and / or a distinguishable sequence of nucleotides or nucleotide analogs. In some embodiments, the identifier is 6 or more consecutive nucleotides. A number of fluorophores with various different excitation and emission spectra are available. Any suitable type and / or number of fluorophores can be used as an identifier. In some embodiments, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twenty or more, thirty or more or fifty or more different identifiers are used in the methods described herein (e.g., nucleic acid detection methods and / or sequencing methods). In some embodiments, one or two types of identifiers (e.g., fluorescent labels) are ligated to each nucleic acid in a library.The detection and / or quantification of the identifier can be performed by suitable methods, apparatuses or instruments, and non-limiting examples thereof include flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, luminometers, fluorometers, spectrophotometers, suitable gene chips or microarray analysis, Western blot, mass spectrometry, chromatography, cell fluorescence analysis, fluorescence microscopy, suitable fluorescence or digital imaging methods, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, dielectrophoresis, suitable nucleic acid sequencing methods and / or nucleic acid sequencing apparatuses, etc. and combinations thereof.
[0103] In some embodiments, a transposon-based library preparation method is used (e.g., EPICENTRE NEXTERA, Epicentre, Madison WI). Transposon-based methods typically use in vitro transposition to simultaneously fragment and tag DNA in a single tube reaction (often enabling the incorporation of platform-specific tags and optional barcodes) to prepare a sequencer-compatible library.
[0104] In some embodiments, the nucleic acid library or a portion thereof is amplified (e.g., amplified by a PCR-based method). In some embodiments, the sequencing method includes amplification of the nucleic acid library. The nucleic acid library can be amplified before or after immobilization onto a solid support (e.g., a solid support within a flow cell). Nucleic acid amplification includes a process of amplifying or increasing the number of nucleic acid templates present (e.g., present in the nucleic acid library) and / or their complementary strands by generating one copy or more copies of the template and / or its complementary strand. Amplification can be performed by a suitable method. The nucleic acid library can be amplified by a thermocycling method or an isothermal amplification method. In some embodiments, a rolling circle amplification method is used. In some embodiments, amplification is performed on a solid support (e.g., within a flow cell) on which the nucleic acid library or a portion thereof is immobilized. In certain sequencing methods, the nucleic acid library is added to a flow cell and immobilized by hybridization to an anchor under suitable conditions. This type of nucleic acid amplification is often referred to as solid-phase amplification. In some embodiments of solid-phase amplification, all or a portion of the amplification product is synthesized by extension starting from an immobilized primer. The solid-phase amplification reaction is similar to standard solution-phase amplification except that at least one of the amplification oligonucleotides (e.g., primers) is immobilized on the solid support. In some embodiments, modified nucleic acids (e.g., nucleic acids modified by the addition of adapters) are amplified.
[0105] In some embodiments, solid-phase amplification includes a nucleic acid amplification reaction that includes only one species of oligonucleotide primer immobilized on a surface. In certain embodiments, solid-phase amplification includes a plurality of different immobilized oligonucleotide primer species. In some embodiments, solid-phase amplification can include a nucleic acid amplification reaction that includes one species of oligonucleotide primer immobilized on a solid surface and a second, different oligonucleotide primer species in solution. A plurality of different species of immobilized primers or solution-based primers can be used. Non-limiting examples of solid-phase nucleic acid amplification reactions include surface amplification, bridge amplification, emulsion PCR, WildFire amplification (e.g., U.S. Patent Application Publication No. 2013 / 0012399), etc. or combinations thereof.
[0106] Capture of Nucleic Acids In some embodiments, a sample nucleic acid (or sample nucleic acid library) is subjected to a target capture process. Generally, the target capture process is performed by contacting the sample nucleic acid (or sample nucleic acid library) with a set of probe oligonucleotides under hybridization conditions. The set of probe oligonucleotides (e.g., capture oligonucleotides) generally includes a plurality of probe oligonucleotides having sequences complementary or substantially complementary to sequences in the sample nucleic acid. The plurality of probe oligonucleotides can include about 10 probe oligonucleotide species, about 50 probe oligonucleotide species, about 100 probe oligonucleotide species, about 500 probe oligonucleotide species, about 1,000 probe oligonucleotide species, 2,000 probe oligonucleotide species, 3,000 probe oligonucleotide species, 4,000 probe oligonucleotide species, 5,000 probe oligonucleotide species, 10,000 probe oligonucleotide species or more probe oligonucleotide species. Usually, the first probe oligonucleotide species has a nucleotide sequence different from that of the second probe oligonucleotide species, and different species of probe oligonucleotides in a set each have a different nucleotide sequence.
[0107] Probe oligonucleotides typically contain a nucleotide sequence that can hybridize or anneal to a nucleic acid fragment of interest (e.g., a target fragment) or a portion thereof. Probe oligonucleotides can be naturally occurring or synthetic and can be DNA- or RNA-based. Probe oligonucleotides can, for example, be capable of specifically separating a target fragment from other fragments in a nucleic acid sample. As used herein, the terms "specific" or "specificity" refer to the binding or hybridization of one molecule to another molecule (e.g., an oligonucleotide to a target polynucleotide). "Specific" or "specificity" refers to the recognition, contact, and formation of a stable complex between two molecules, as compared to substantially low recognition, contact, or complex formation by either of the two molecules with other molecules. As used herein, the terms "anneal" and "hybridize" refer to the formation of a stable complex between two molecules. The terms "probe", "probe oligonucleotide", "capture probe", "capture oligonucleotide", "capture oligo", "oligo", or "oligonucleotide" can be used interchangeably throughout this document when referring to probe oligonucleotides.
[0108] Probe oligonucleotides can be designed and synthesized using suitable processes and can be of any length suitable for hybridizing to a nucleotide sequence of interest and for performing the isolation and / or analysis processes described herein. The oligonucleotides can be designed based on a nucleotide sequence of interest (e.g., a target fragment sequence, a genomic sequence, a gene sequence). In some embodiments, the oligonucleotides (e.g., probe oligonucleotides) can be about 10 to about 300 nucleotides, about 50 to about 200 nucleotides, about 75 to about 150 nucleotides, about 110 to about 130 nucleotides, or about 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128 or 129 nucleotides in length. The oligonucleotides can be composed of naturally occurring and / or non-naturally occurring nucleotides (e.g., labeled nucleotides) or mixtures thereof. Oligonucleotides suitable for use in the embodiments described herein can be synthesized and labeled using known techniques. Oligonucleotides can be chemically synthesized according to the solid-phase phosphoramidite triester method first reported by Beaucage and Caruthers (1981) Tetrahedron Letts. 22:1859-1862, using an automated synthesizer and / or as described by Needham-VanDevanter et al. (1984) Nucleic Acids Res. 12:6159-6168. Purification of the oligonucleotides can be carried out by denaturing acrylamide gel electrophoresis or by anion-exchange high performance liquid chromatography (HPLC) as described, for example, by Pearson and Regnier (1983) J. Chrom. 255:137-149.
[0109] All or part of a probe oligonucleotide sequence (either naturally occurring or synthetic) may, in some embodiments, be substantially complementary to a target sequence or a part thereof. "Substantially complementary", when referred to in the context of sequences herein, refers to nucleotide sequences that hybridize to each other. The stringency of the hybridization conditions can be varied to allow for various amounts of sequence mismatches. Target sequences and oligonucleotide sequences that are 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more complementary are included.
[0110] A probe oligonucleotide that is substantially complementary to a nucleotide sequence of interest (e.g., a target sequence) or a part thereof is also substantially similar to the complementary strand of the target sequence or a related part thereof (e.g., substantially similar to the antisense strand of the nucleic acid). One test for determining whether two nucleotide sequences are substantially similar is to measure the percentage of the same nucleotide sequence that is shared. "Substantially similar" when referred to in this specification with respect to sequences means 55% or more, 56% or more, 57% or more, 58% or more, 59% or more, 60% or more, 61% or more, 62% or more, 63% or more, 64% or more, 65% or more, 66% or more, 67% or more, 68% or more, 69% or more, 70% or more, 71% or more, 72% or more, 73% or more, 74% or more, 75% or more, 76% or more, 77% or more, 78% or more, 79% or more, 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more identical nucleotide sequences.
[0111] Hybridization conditions (e.g., annealing conditions) can be determined and / or adjusted according to the characteristics of the oligonucleotides used in the assay. The sequence and / or length of the oligonucleotide can sometimes affect hybridization to the nucleic acid sequence of interest. Depending on the degree of mismatch between the oligonucleotide and the nucleic acid of interest, low, medium or high stringency conditions can be used to achieve annealing. As used herein, the term "stringent conditions" refers to the conditions for hybridization and washing. Methods for optimizing the temperature conditions of the hybridization reaction are known in the art and can be found in Current Protocols in Molecular Biology, John Wiley & Sons, N.Y., 6.3.1-6.3.6 (1989). Aqueous and non-aqueous methods are described in that reference and either can be used. Non-limiting examples of stringent hybridization conditions are one or more washes in 0.2× SSC, 0.1% SDS at 50°C following hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45°C. Another example of stringent hybridization conditions is one or more washes in 0.2× SSC, 0.1% SDS at 55°C following hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45°C. Further examples of stringent hybridization conditions are one or more washes in 0.2× SSC, 0.1% SDS at 60°C following hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45°C. Stringent hybridization conditions are often one or more washes in 0.2× SSC, 0.1% SDS at 65°C following hybridization in 6× sodium chloride / sodium citrate (SSC) at about 45°C. Stringency conditions are more often one or more washes in 0.2× SSC, 1% SDS at 65°C following 0.5 M sodium phosphate, 7% SDS at 65°C.The stringent hybridization temperature can also be altered (i.e., decreased) by adding a particular organic solvent, such as formamide. Organic solvents such as formamide lower the thermal stability of double-stranded polynucleotides, allowing hybridization to occur at lower temperatures while maintaining stringent conditions and extending the lifespan of useful nucleic acids that may be thermally unstable.
[0112] In some embodiments, one or more probe oligonucleotides associate with an affinity ligand (e.g., a member of a binding pair (e.g., biotin) or an antigen that can bind to a capture agent such as avidin, streptavidin, an antibody, or a receptor). For example, a probe oligonucleotide can be biotinylated so that it can be captured by beads coated with streptavidin.
[0113] In some embodiments, one or more probe oligonucleotides and / or capture agents are effectively linked to a solid support or substrate. The solid support or substrate can be any physically separable solid to which the probe oligonucleotides can attach directly or indirectly, including surfaces provided by microarrays and wells, and particles such as beads (e.g., paramagnetic beads, magnetic beads, microbeads, nanobeads), microparticles, and nanoparticles, but not limited thereto. Solid supports include, for example, chips, columns, optical fibers, wipes (swabbing papers), filters (e.g., flat surface filters), one or more capillaries, glass and processed or functionalized glass (e.g., controlled-pore glass (CPG)), quartz, mica, diazotized membranes (papers or nylons), polyformaldehyde, cellulose, cellulose acetate, paper, ceramics, metals, metalloids, semiconductor materials, quantum dots, coated beads or particles, other chromatographic materials, magnetic particles; plastics (including acrylic resins, polystyrene, copolymers of styrene or other materials, polybutylene, polyurethane, TEFLON®, polyethylene, polypropylene, polyamide, polyester, polyvinylidene difluoride (PVDF), etc.), polysaccharides, nylon or nitrocellulose, resins, silica or silica-based materials (including silicon, silica gel, and modified silicon), Sephadex®, Sepharose®, carbon, metals (e.g., steel, gold, silver, aluminum, silicon, and copper), inorganic glass, conductive polymers (including polymers such as polypyrrole and polyindole); microstructured or nanostructured surfaces (e.g., surfaces decorated with nucleic acid tiling arrays, nanotubes, nanowires, or nanoparticles); or porous surfaces or gels (e.g., methacrylate, acrylamide, sugar polymers, cellulose, silicates, or other fibrous or chain polymers) may also be included.In some embodiments, the solid support or substrate can be coated with a passively or chemically derivatized coating by any number of materials including polymers such as dextran, acrylamide, gelatin or agarose. The beads and / or particles can be free or connected to each other (e.g., sintered). In some embodiments, the solid phase can be an aggregate of particles. In some embodiments, the particles can include silica, which can include silicon dioxide. In some embodiments, the silica can be porous, and in certain embodiments, the silica can be non-porous. In some embodiments, the particles further include an agent that imparts paramagnetism to the particle. In certain embodiments, the agent includes a metal, and in certain embodiments, the agent is a metal oxide (e.g., iron or iron oxide, where the iron oxide includes a mixture of Fe2+ and Fe3+). The probe oligonucleotide can be linked to the solid support by covalent or non-covalent interactions and can be linked directly or indirectly (e.g., via an intervening substance such as a spacer molecule or biotin) to the solid support. The probe oligonucleotide can be linked to the solid support before, during or after nucleic acid capture.
[0114] Nucleic acids modified, such as by the addition of adapter sequences described herein, can be captured. In some embodiments, unmodified nucleic acids are captured. The nucleic acids can be amplified, in some embodiments, by an amplification process such as PCR, before and / or after capture. The term "captured nucleic acid" typically includes the captured nucleic acid and the captured and amplified nucleic acid. In some embodiments, the captured nucleic acid can be subjected to additional rounds of capture and amplification. The captured nucleic acid can be sequenced by a sequencing process described herein.
[0115] Detection of copy number variations in captured nucleic acids Methods and processes are provided herein for classifying the presence or absence of copy number variations (e.g., microduplications, microdeletions). In some embodiments, determination of the presence or absence of a copy number variation is determined according to a set of sequence reads. In some embodiments, determination of the presence or absence of a copy number variation is determined according to quantitative values of sequence reads for segments and / or sub-regions described herein. In some embodiments, the sequence reads are obtained from circulating cell-free sample nucleic acids from a test subject captured by probe oligonucleotides under hybridization conditions. In some embodiments, the presence or absence of a copy number variation is determined according to a set of consensus sequences generated from the sequence reads. In some embodiments, the presence or absence of a copy number variation is determined according to quantitative values of probe coverage. In some embodiments, determination of the presence or absence of a copy number variation is determined according to quantitative probe coverage values for segments and / or sub-regions described herein. The quantitative probe coverage value can be a quantitative value of sequence reads for each probe oligonucleotide. The quantitative probe coverage value can be a quantitative value of a consensus sequence for each probe oligonucleotide. In some embodiments, the presence or absence of a copy number variation is determined according to normalized quantitative probe coverage values (e.g., normalized quantitative probe coverage values of sequence reads for each probe oligonucleotide; normalized quantitative probe coverage values of consensus sequences for each probe oligonucleotide). In some embodiments, determination of the presence or absence of a copy number variation includes a segmentation process. In some embodiments, determination of the presence or absence of a copy number variation includes a filtering process.
[0116] In some embodiments, determination of the presence or absence of a copy number variation is based on a probe coverage quantification value or a normalized probe coverage quantification value. In some embodiments, "based on" may include other factors (e.g., segments, filtered segments, copy number measurements or estimations, copy number increase or decrease measurements or estimations, filtered copy number measurements or estimations, filtered copy number increase or decrease measurements or estimations). The presence or absence of a copy number variation may, in some embodiments, be determined according to a probe coverage quantification value or a normalized probe coverage quantification value for a single probe oligonucleotide. The presence or absence of a copy number variation may, in some embodiments, be determined according to a probe coverage quantification value or a normalized probe coverage quantification value for a plurality of probe oligonucleotides.
[0117] In some embodiments, sample nucleic acid is captured by a probe oligonucleotide. Typically, in such embodiments, the sample nucleic acid is contacted with the probe oligonucleotide under hybridization conditions. The sample nucleic acid may (or may consist of) a sample polynucleotide, and the probe oligonucleotide may include a probe polynucleotide complementary to the sample polynucleotide in the sample nucleic acid. In some embodiments, the probe polynucleotide is complementary to a sequence in a subchromosomal region, segment, and / or subregion of interest described herein. In some embodiments, the stringency of the hybridization conditions allows only probe polynucleotides having 100% complementarity (i.e., no mismatches) to hybridize to the sample nucleic acid. In some embodiments, the stringency of the hybridization conditions allows probe polynucleotides having one or two mismatches to hybridize to the sample nucleic acid.
[0118] In some embodiments, the sequence reads are mapped to a reference genome portion. Certain methods for mapping sequence reads to a reference genome portion are described herein. In some embodiments, the genome portion is of a predetermined length. In some embodiments, the genome portions are of equal length. In some embodiments, the genome portion is about 50 kilobases in length. In some embodiments, at least two genome portions are of unequal length. In some embodiments, the genome portions do not overlap. In some embodiments, the 3' end of a genome portion is adjacent to the 5' end of each adjacent downstream genome portion. In some embodiments, at least two genome portions overlap.
[0119] In some embodiments, the sequence reads that map to the reference genome match the probe sequence and are identified as on-target reads. In some embodiments, the methods herein include the step of identifying on-target reads. In some embodiments, when a read aligns to a genomic region corresponding to the probe oligonucleotide sequence, that read is identified as on-target. As described in more detail herein, the probe oligonucleotide sequence typically aligns (i.e., corresponds) to a particular region of the genome (e.g., the reference genome) and often includes a nucleotide sequence corresponding to a particular genomic sequence of interest (e.g., a sequence in a subchromosomal region, segment, and / or sub-region of interest described herein). A read that aligns to the genomic region to which the probe oligonucleotide aligns is considered an on-target read. In some embodiments, when the entire read length aligns to the genomic region that aligns to the probe oligonucleotide, the sequence read may be considered on-target. In some embodiments, when a portion of the read aligns to the genomic region corresponding to the probe oligonucleotide sequence and a portion of the read aligns within a genomic region adjacent to the genomic region corresponding to the probe oligonucleotide sequence, that read is identified as on-target. Generally, in such cases, the read aligns to a contiguous genomic sequence that includes 1) a portion of the genomic region corresponding to the probe oligonucleotide sequence, and 2) a genomic region adjacent to the genomic region corresponding to the probe oligonucleotide sequence. The latter genomic region may be located upstream or downstream of the genomic region corresponding to the probe oligonucleotide sequence.For example, when a portion of a read (e.g., at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% of the read) having a genomic region corresponding to a probe oligonucleotide sequence and the remainder of the read align to a genomic sequence immediately upstream or downstream of the genomic region corresponding to the probe oligonucleotide sequence, the sequence read may be considered on-target. In some embodiments, when a portion of the read does not align to the probe sequence and the entire read length aligns to a genomic sequence immediately upstream or downstream of the genomic region corresponding to the probe oligonucleotide sequence, the sequence read may be considered on-target.
[0120] A sequence that includes the probe sequence (i.e., the genomic sequence corresponding to the probe sequence) and additional genomic sequences upstream and / or downstream of the probe sequence may be referred to as a padded probe sequence. A set of padded probe sequences may be referred to as a padded panel. In some embodiments, the padded probe sequence includes at least 1 nucleotide of the genomic sequence immediately upstream and / or downstream of the genomic sequence corresponding to the probe sequence. For example, the padded probe sequence may include at least about 5, 10, 20, 30, 40, 50, 100, 150, 200, 250, 300, 400, 500, or 1000 nucleotides of the genomic sequence immediately upstream and / or downstream of the genomic sequence corresponding to the probe sequence. In some embodiments, the padded probe sequence includes a 250 nucleotide genomic sequence immediately upstream and a 250 nucleotide genomic sequence immediately downstream of the genomic sequence corresponding to the probe sequence.
[0121] Probe oligonucleotide arrays can be stored in a database as an array panel. In some embodiments, reads are directly aligned with a probe oligonucleotide array (e.g., a probe oligonucleotide array stored in a table or database with or without adjacent genomic region sequences as described above), and such reads are identified as on-target reads. For example, sequence reads can be aligned to an array panel in a database without first mapping to a reference genome. In some embodiments, when the entire read length aligns to the probe array, the sequence read can be considered on-target. In some embodiments, sequence reads are directly aligned to the padded probe array as described above. For example, in some embodiments, when a portion of the read (e.g., at least about 5% of the read, 10% of the read, 20% of the read, 30% of the read, 40% of the read, 50% of the read, 60% of the read, 70% of the read, 80% of the read, 90% of the read) aligns to the probe array and the remainder of the read aligns to the genomic sequence immediately upstream or downstream of the probe array, that sequence read can be considered on-target. In some embodiments, when a portion of the read does not align to the probe array and the entire read length aligns to the genomic sequence immediately upstream or downstream of the probe array, the sequence read can be considered on-target.
[0122] In some embodiments, the consensus sequence is generated from sequence reads. In some embodiments, the consensus sequence is generated from sequence reads identified as "on-target" reads. Generally, a consensus is generated by collapsing a set of sequence reads (e.g., reads within a read group) to generate a single nucleotide sequence corresponding to the unique nucleic acid molecules in the sample from which the sequence reads were generated. The consensus sequence can be generated from the read group by any suitable method, such as linear or non-linear methods for consensus generation derived from, for example, digital communication theory, information theory, or bioinformatics (e.g., averaging, voting, statistical, dynamic programming, maximum a posteriori or maximum likelihood detection methods, Bayesian methods, hidden Markov methods, or support vector machine methods, etc.).
[0123] In some embodiments, the determination of the presence or absence of a copy number variation is determined according to a probe coverage quantification value (e.g., a probe coverage quantification value for a segment and / or sub-region described herein; a probe coverage quantification value for an array in a segment and / or an array in a sub-region described herein). Probe coverage generally refers to the quantification value of an array read or consensus sequence mapped to each nucleotide position in a probe oligonucleotide. In some embodiments, measuring the probe coverage quantification value includes measuring the number of array reads mapped to each nucleotide position in the probe oligonucleotide. The array read can be shorter in length than the probe oligonucleotide and / or can partially overlap the probe oligonucleotide sequence. Thus, the quantification value of the array reads mapped to each nucleotide in the probe can vary depending on the length of the probe oligonucleotide. Thus, in some embodiments, measuring the probe coverage quantification value includes measuring the quantile estimate value of the population of array reads mapped to each nucleotide position in the probe. Examples of quantile estimate values can include, for example, the median, mean, mode, range, etc. In some embodiments, measuring the probe coverage quantification value includes measuring the median of the number of array reads mapped to each nucleotide position in the probe. In some embodiments, the median of the number of array reads mapped to each nucleotide position for each probe oligonucleotide is the probe coverage quantification value for each probe oligonucleotide. In some embodiments, measuring the probe coverage quantification value includes measuring the number of consensus sequences mapped to each nucleotide position in the probe oligonucleotide. The consensus sequence can be shorter in length than the probe oligonucleotide and / or can partially overlap the probe oligonucleotide sequence. Thus, the quantification value of the consensus sequences mapped to each nucleotide in the probe can vary depending on the length of the probe oligonucleotide.Thus, in some embodiments, the measurement of the probe coverage quantification value includes the measurement of the median number of consensus sequences mapped to each nucleotide position in the probe.
[0124] In some embodiments, the determination of the presence or absence of a copy number variation is determined according to the normalized probe coverage quantification value. The probe coverage quantification value can be normalized using a suitable normalization process such as the normalization process described herein. In some embodiments, normalization includes scaling the probe coverage quantification value for each probe oligonucleotide for the test sample. Scaling the probe coverage quantification value for each probe oligonucleotide generates a scaled probe coverage quantification value for each probe oligonucleotide. In some embodiments, the probe coverage quantification value for each probe is scaled according to the median of the probe coverage quantification values for all probe oligonucleotides for the test sample. For example, the probe coverage quantification value for each probe oligonucleotide can be divided by the median of the probe coverage quantification values.
[0125] In some embodiments, normalization includes normalizing the probe coverage quantification value according to the guanine-cytosine (GC) content for each probe oligonucleotide for the test sample. In some embodiments, normalization includes normalizing the scaled probe coverage quantification value according to the guanine-cytosine (GC) content for each probe oligonucleotide for the test sample. Normalizing the probe coverage quantification value according to the GC content for each probe oligonucleotide generates a GC-normalized probe coverage quantification value for each probe oligonucleotide. In some embodiments, the probe coverage quantification value is normalized by LOESS normalization. LOESS normalization (e.g., GC LOESS) is described in further detail herein.
[0126] In some embodiments, normalization includes normalizing the probe coverage quantification value for a test sample according to the probe coverage quantification value obtained from a reference sample. The reference sample may include samples classified as having no copy number variations. In some embodiments, the reference sample consists of samples classified as having no copy number variations. Thus, in some embodiments, the reference sample includes or consists of samples that are euploid for each chromosome and each chromosomal region being tested. The reference sample can be from a human subject. In some embodiments, the reference sample is from a female subject. In some embodiments, the reference sample is from a male subject. In some embodiments, the reference sample is from male and female subjects. The reference sample may include samples from one subject or may include samples from multiple subjects. The reference sample may include one reference sample, and more often includes multiple samples. For example, the reference sample may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100 or more samples.
[0127] In some embodiments, the probe coverage quantification value for a test sample is normalized according to the probe coverage quantification value obtained from a reference sample. In some embodiments, the scaled probe coverage quantification value for a test sample is normalized according to the probe coverage quantification value obtained from a reference sample. In some embodiments, the GC-normalized probe coverage quantification value for a test sample is normalized according to the probe coverage quantification value obtained from a reference sample. In some embodiments, the probe coverage quantification value for a test sample is normalized according to the median probe coverage for each probe oligonucleotide obtained from a reference sample. In some embodiments, the scaled probe coverage quantification value for a test sample is normalized according to the median probe coverage for each probe oligonucleotide obtained from a reference sample. In some embodiments, the GC-normalized probe coverage quantification value for a test sample is normalized according to the median probe coverage for each probe oligonucleotide obtained from a reference sample. The median probe coverage is often measured according to the probe coverage quantification value for the same probe across multiple reference samples. In some embodiments, the median probe coverage is measured according to the normalized (e.g., GC-normalized) probe coverage quantification value for the same probe across multiple reference samples. The normalization of the probe coverage quantification value (e.g., median probe coverage) for each probe oligonucleotide according to the probe coverage quantification value obtained from a reference sample generates a reference-sample-normalized probe coverage quantification value for each probe oligonucleotide for the test sample.
[0128] In some embodiments, normalizing according to the probe coverage median (e.g., the probe coverage median for each probe oligonucleotide obtained from a reference sample) includes dividing each probe coverage quantification value for each probe oligonucleotide (i.e., for a test sample) by the probe coverage median for each probe oligonucleotide obtained from a reference sample. In some embodiments, normalizing according to the probe coverage median (e.g., the probe coverage median for each probe oligonucleotide obtained from a reference sample) includes dividing each scaled probe coverage quantification value for each probe oligonucleotide (i.e., for a test sample) by the probe coverage median for each probe oligonucleotide obtained from a reference sample. In some embodiments, normalizing according to the probe coverage median (e.g., the probe coverage median for each probe oligonucleotide obtained from a reference sample) includes dividing each GC-normalized probe coverage quantification value for each probe oligonucleotide (i.e., for a test sample) by the probe coverage median for each probe oligonucleotide obtained from a reference sample. In such embodiments, normalizing according to the probe coverage median generates a ratio for each probe oligonucleotide.
[0129] In some embodiments, the probe coverage quantification value is logarithmically transformed. For example, the probe coverage quantification value normalized by a reference sample for each probe oligonucleotide can be logarithmically transformed. By logarithmically transforming the probe coverage quantification value normalized by a reference sample for each probe oligonucleotide, a logarithmically transformed, reference sample-normalized probe coverage quantification value for each probe oligonucleotide is generated. In certain embodiments, the ratio for each probe oligonucleotide is logarithmically transformed. By logarithmically transforming the ratio for each probe oligonucleotide, a logarithmically transformed ratio for each probe oligonucleotide is generated. In some embodiments, the logarithmic transformation is a log2 transformation. Thus, in some embodiments, a log2-transformed, reference sample-normalized probe coverage quantification value for each probe oligonucleotide is generated. In some embodiments, a log2 ratio for each probe oligonucleotide is generated. In certain cases, the log2 ratio of the probe coverage quantification value is proportional to the log2 ratio for an increase or decrease in copy number (CN) as shown, for example, in Equation A:
Number
[0130] where "test coverage" refers to the probe coverage quantification value (e.g., scaled probe coverage quantification value, normalized probe coverage quantification value) for the probe oligonucleotide for the test sample; "normal coverage" refers to the probe coverage quantification value (e.g., probe coverage median) for the probe oligonucleotide obtained from the reference sample; and CN is an increase or decrease in copy number for the segment represented by the probe oligonucleotide (i.e., the segment containing the same or substantially the same sequence as the probe oligonucleotide sequence).
[0131] The normalized probe coverage quantification value for a probe oligonucleotide with respect to a test sample can refer to any normalized probe coverage quantification value described herein or any suitable variation thereof. For example, the normalized probe coverage quantification value for a probe oligonucleotide with respect to a test sample can refer to the scaled probe coverage quantification value for the probe oligonucleotide with respect to the test sample. In certain cases, the normalized probe coverage quantification value for a probe oligonucleotide with respect to a test sample can refer to the GC-normalized probe coverage quantification value for the probe oligonucleotide with respect to the test sample. In certain cases, the normalized probe coverage quantification value for a probe oligonucleotide with respect to a test sample can refer to the probe coverage quantification value for the probe oligonucleotide with respect to the test sample normalized by a reference sample. In certain cases, the normalized probe coverage quantification value for a probe oligonucleotide with respect to a test sample can refer to the ratio for the probe oligonucleotide with respect to the test sample. In certain cases, the normalized probe coverage quantification value for a probe oligonucleotide with respect to a test sample can refer to the logarithmically transformed, reference sample-normalized probe coverage quantification value for the probe oligonucleotide with respect to the test sample. In certain cases, the normalized probe coverage quantification value for a probe oligonucleotide with respect to a test sample can refer to the logarithmically transformed ratio for the probe oligonucleotide with respect to the test sample. In certain cases, the normalized probe coverage quantification value for a probe oligonucleotide with respect to a test sample can refer to the log2-transformed, reference sample-normalized probe coverage quantification value for the probe oligonucleotide with respect to the test sample.In certain cases, the normalized probe coverage quantification value for a probe oligonucleotide with respect to a test sample may refer to the log2 ratio of the probe oligonucleotide with respect to the test sample.
[0132] In some embodiments, the segmentation process is applied to identify segments (e.g., segments spanning copy number variations). Any suitable segmentation process may be utilized, including but not limited to the circular binary segmentation (CBS) process. Instead of or in addition to CBS, other processes may be utilized, non-limiting examples of which include wavelet segmentation (e.g., Haar wavelet segmentation), Fourier transform, sliding window z-score, and Markov chain models.
[0133] In some embodiments, the segmentation process is applied to identify segments according to the probe coverage quantification value for each probe oligonucleotide. In some embodiments, the segmentation process is applied to identify segments according to the normalized probe coverage quantification value for each probe oligonucleotide. A segment may include a plurality of probe oligonucleotides (i.e., a plurality of probe oligonucleotides having a probe coverage quantification value suggesting an increase or decrease in copy number variation). The segmentation process may provide a start position and an end position for each segment (e.g., a start position and an end position according to genomic coordinates; a start position and an end position according to probe index), a quantification value of copy number variation for the segment, and optionally a measure of confidence for the segment. In some embodiments, the position (e.g., the position according to probe index) and the probe coverage quantification value for each end of each segment are provided for each segment. In some embodiments, the position (e.g., the position according to probe index) and the normalized probe coverage quantification value for each end of each segment are provided for each segment. In some embodiments, one or more genes overlapping with each segment are identified.
[0134] In some embodiments, the copy number of each segment is determined or estimated according to the probe coverage quantification value associated with each segment. In some embodiments, the copy number of each segment is determined or estimated according to the normalized probe coverage quantification value associated with each segment. The determination or estimation of the copy number of each segment provides a copy number (CN) increase or a copy number (CN) decrease for each segment. In some embodiments, the copy number (CN) increase or the copy number (CN) decrease for each segment is determined or estimated according to the conversion of the segment median coverage for each segment. Thus, in a particular case, the segment median coverage is determined according to the probe coverage quantification value for the probe oligonucleotide in the segment. In some embodiments, the copy number (CN) increase or the copy number (CN) decrease for each segment is determined or estimated according to the conversion of the segment median coverage log2 ratio for each segment. Thus, in a particular case, the median of the log2 ratio is determined according to the probe coverage quantification value for the probe oligonucleotide in the segment. In other words, the median of the log2 ratio for the probe oligonucleotide in the segment is used for the determination or estimation of the copy number (CN) increase or the copy number (CN) decrease for the segment. For example, the copy number increase or the copy number decrease for the segment is given by Equation B: CN = 2 * (2 (セグメント.中央値.log2比) - 1) Equation B (where CN is the copy number increase or the copy number decrease for each segment) can be determined or estimated according to.
[0135] In some embodiments, segments are filtered (e.g., removed from those to be considered). Segments can be filtered according to a probe coverage quantification value associated with the segment, a normalized probe coverage quantification value associated with the segment, and one or more of a copy number increase or a copy number decrease for the segment. Typically, filtering of segments provides a set of segments that are filtered and retained. Segments are often paired with corresponding copy number quantification values, and segments with an absolute value of the corresponding copy number quantification value of 0 to about 1 (for duplication candidates) or 0 to about 0.9 (for deletion candidates) are often filtered out as part of a noise reduction filtering process. In some embodiments, segments with a copy number increase of 1 or more (for duplication candidates) can be retained as filtered segments. For example, segments with a copy number increase of 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more can be retained as filtered segments. In some embodiments, segments with a copy number decrease of 0.9 or more (for deletion candidates) can be retained as filtered segments. Since the copy number quantification value for deletion candidates typically drops below zero, "a copy number decrease of 0.9 or more" corresponds to the absolute value of the copy number quantification value for deletion candidates. Thus, for example, segments with a copy number decrease of 0.9, 1.0, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 or 2 can be retained as filtered segments. In other words, segments having a copy number quantification value of -0.9, -1.0, -1.2, -1.3, -1.4, -1.5, -1.6, -1.7, -1.8, -1.9 or -2 can be retained as filtered segments.
[0136] Nucleic Acid Sequencing and Processing The methods provided herein typically include nucleic acid sequencing and analysis. In some embodiments, the nucleic acid is sequenced and the products of the sequencing (e.g., a set of sequence reads) are processed either before or simultaneously with the analysis of the sequenced nucleic acid. For example, the sequence reads can be processed according to one or more of the following steps: aligning, mapping, filtering portions, selecting portions, counting, normalizing, weighting, generating profiles, etc., and combinations thereof. Certain processing steps can be performed in any order and certain processing steps can be repeated. For example, after portions are filtered, the sequence read count can be normalized, and in certain embodiments, the sequence read count can be normalized and then portions can be filtered. In some embodiments, following the step of filtering portions, normalization of the sequence read count is followed by a further step of filtering portions. Certain sequencing methods and processing steps are described in further detail below.
[0137] Sequencing In some embodiments, a nucleic acid (e.g., a nucleic acid fragment, a sample nucleic acid, a cell-free nucleic acid) is sequenced. In certain cases, a complete or substantially complete sequence is obtained, and in some cases, a partial sequence is obtained. Nucleic acid sequencing typically generates a set of sequence reads. As used herein, "reads" (e.g., "a read", "sequence reads") are short nucleotide sequences generated by any sequencing process described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment ("single-end reads") or sometimes from both ends of a nucleic acid fragment (e.g., paired-end reads, double-end reads).
[0138] The length of an array read is often related to a specific array determination technique. For example, high-throughput methods provide array reads that can vary in size from dozens to hundreds of base pairs (bp). For example, nanopore sequencing can provide array reads that can vary in size from dozens, hundreds to thousands of base pairs. In some embodiments, the array read is the average value, median, average length or absolute value of the length from about 15 bp to about 900 bp. In certain embodiments, the array read is the average value, median, average length or absolute value of a length of about 1000 bp or more. In some embodiments, the array read is the average value, median, average length or absolute value of a length of about 1500, 2000, 2500, 3000, 3500, 4000, 4500 or 5000 bp or more. In some embodiments, the array read is the average value, median, average length or absolute value of a length from about 100 bp to about 200 bp. In some embodiments, the array read is the average value, median, average length or absolute value of a length from about 140 bp to about 160 bp. For example, the array read can be the average value, median, average length or absolute value of a length of about 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159 or 160 bp.
[0139] In some embodiments, the nominal length, average length, mean value of the length, or absolute value of the length of a single-end read is about 10 consecutive nucleotides to about 250 or more consecutive nucleotides, about 15 consecutive nucleotides to about 200 or more consecutive nucleotides, about 15 consecutive nucleotides to about 150 or more consecutive nucleotides, about 15 consecutive nucleotides to about 125 or more consecutive nucleotides, about 15 consecutive nucleotides to about 100 or more consecutive nucleotides, about 15 consecutive nucleotides to about 75 or more consecutive nucleotides, about 15 consecutive nucleotides to about 60 or more consecutive nucleotides, 15 consecutive nucleotides to about 50 or more consecutive nucleotides, about 15 consecutive nucleotides to about 40 or more consecutive nucleotides, and may be about 15 consecutive nucleotides or about 36 or more consecutive nucleotides. In certain embodiments, the nominal length, average length, mean value of the length, or absolute value of the length of a single-end read is about 20 to about 30 bases in length or about 24 to about 28 bases in length. In certain embodiments, the nominal length, average length, mean value of the length, or absolute value of the length of a single-end read is about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 21, 22, 23, 24, 25, 26, 27, 28, or about 29 bases in length or longer. In certain embodiments, the nominal length, average length, mean value of the length, or absolute value of the length of a single-end read is about 20 to about 200 bases in length, about 100 to about 200 bases in length, or about 140 to about 160 bases in length. In certain embodiments, the nominal length, average length, mean value of the length, or absolute value of the length of a single-end read is about 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or about 200 bases in length or longer.In certain embodiments, the nominal length, average length, mean length or absolute value of the length of the paired-end reads is from about 10 consecutive nucleotides to about 25 consecutive nucleotides or more (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 nucleotide lengths or more), sometimes from about 15 consecutive nucleotides to about 20 consecutive nucleotides or more, and sometimes about 17 consecutive nucleotides or about 18 consecutive nucleotides. In certain embodiments, the nominal length, average length, mean length or absolute value of the length of the paired-end reads is from about 25 consecutive nucleotides to about 400 consecutive nucleotides or more (e.g., about 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390 or 400 nucleotide lengths or more), sometimes from about 50 consecutive nucleotides to about 350 consecutive nucleotides or more, sometimes from about 100 consecutive nucleotides to about 325 consecutive nucleotides, sometimes from about 150 consecutive nucleotides to about 325 consecutive nucleotides, sometimes from about 200 consecutive nucleotides to about 325 consecutive nucleotides, sometimes from about 275 consecutive nucleotides to about 310 consecutive nucleotides, sometimes from about 100 consecutive nucleotides to about 200 consecutive nucleotides, sometimes from about 100 consecutive nucleotides to about 175 consecutive nucleotides, sometimes from about 125 consecutive nucleotides to about 175 consecutive nucleotides, and sometimes from about 140 consecutive nucleotides to about 160 consecutive nucleotides. In certain embodiments, the nominal length, average length, mean length or absolute value of the length of the paired-end reads is about 150 consecutive nucleotides, and sometimes 150 consecutive nucleotides.
[0140] In some embodiments, the nucleotide sequence reads obtained from a sample are partial nucleotide sequence reads. As used herein, a "partial nucleotide sequence read" refers to an array read of any length having incomplete sequence information, also referred to as sequence ambiguity. A partial nucleotide sequence read may lack information regarding the identity of the nucleobases and / or the position or order of the nucleobases. A partial nucleotide sequence read generally does not include array reads in which the incomplete sequence information (or fewer bases than all of those bases have been sequenced or determined) is due to inadvertent or non-intentional sequencing errors. Such sequencing errors may be specific to a particular sequencing process and include, for example, inaccurate calls for the identity of nucleobases and missing or extra nucleobases. Thus, for partial nucleotide sequence reads herein, certain information regarding the sequence is often deliberately excluded. That is, sequence information regarding fewer nucleobases than all nucleobases, or sequence information that may be characterized separately as a sequencing error or may be a sequencing error, is deliberately obtained. In some embodiments, a partial nucleotide sequence read may extend to a portion of a nucleic acid fragment. In some embodiments, a partial nucleotide sequence read may extend to the entire length of a nucleic acid fragment. Partial nucleotide sequence reads are described, for example, in International Patent Application Publication No. WO2013 / 052907, the entire contents of which, including the text, tables, formulas, and drawings, are incorporated herein by reference.
[0141] A read is generally the presentation of a nucleotide sequence in a physical nucleic acid. For example, in a read containing a sequence depicted as ATGC, in the physical nucleic acid, "A" represents an adenine nucleotide, "T" represents a thymine nucleotide, "G" represents a guanine nucleotide, and "C" represents a cytosine nucleotide. Sequence reads obtained from a sample derived from a subject can be reads from a mixture of minority and majority nucleic acids. For example, a sequence read obtained from the blood of a cancer patient can be a read from a mixture of cancer nucleic acids and non-cancer nucleic acids. In another example, a sequence read obtained from the blood of a pregnant woman can be a read from a mixture of fetal nucleic acids and maternal nucleic acids. A mixture of relatively short reads can be converted by the processes described herein into a presentation of the genomic nucleic acids present in the subject and / or a presentation of the genomic nucleic acids present in a tumor or a fetus. In certain cases, a mixture of relatively short reads can be converted, for example, into a presentation of copy number variations, gene mutations / gene changes, or aneuploidy. In one example, a read of a mixture of cancer nucleic acids and non-cancer nucleic acids can be converted into a presentation of a composite chromosome or a portion thereof that includes features of the chromosomes of one or both of cancer cells and non-cancer cells. In another example, a read of a mixture of maternal nucleic acids and fetal nucleic acids can be converted into a presentation of a composite chromosome or a portion thereof that includes features of the chromosomes of one or both of the mother and the fetus.
[0142] In some cases, circulating cell-free nucleic acid fragments (CCF fragments) obtained from a cancer patient contain nucleic acid fragments of normal cell origin (i.e., non-cancer fragments) and nucleic acid fragments of cancer cell origin (i.e., cancer fragments). Sequence reads derived from CCF fragments of normal cell (i.e., non-cancerous cell) origin are referred to herein as "non-cancer reads". Sequence reads derived from CCF fragments of cancer cell origin are referred to herein as "cancer reads". CCF fragments from which non-cancer reads are obtained can be referred to herein as non-cancer templates, and CCF fragments from which cancer reads are obtained can be referred to herein as cancer templates.
[0143] In some cases, circulating cell-free nucleic acid fragments (CCF fragments) obtained from a pregnant woman include nucleic acid fragments of fetal origin (i.e., fetal fragments) and nucleic acid fragments of maternal origin (i.e., maternal fragments). Sequence reads derived from CCF fragments of fetal origin are referred to herein as "fetal reads." Sequence reads derived from CCF fragments of genomic origin from a pregnant woman with a fetus (e.g., the mother) are referred to herein as "maternal reads." CCF fragments from which fetal reads are obtained are referred to herein as fetal templates, and CCF fragments from which maternal reads are obtained are referred to herein as maternal templates.
[0144] In certain embodiments, "obtaining" nucleic acid sequence reads from a subject sample and / or "obtaining" nucleic acid sequence reads from a biological sample from one or more reference individuals can include directly sequencing the nucleic acid to obtain sequence information. In some embodiments, "obtaining" can include receiving sequence information directly obtained from the nucleic acid by another.
[0145] In some embodiments, some or all of the nucleic acids in a sample are concentrated and / or amplified (e.g., non-specifically, e.g., by a PCR-based method) before or during sequencing. In certain embodiments, a particular nucleic acid species or subset in a sample is concentrated and / or amplified before or during sequencing. In some embodiments, a preselected species or subset of a nucleic acid pool is randomly sequenced. In some embodiments, the nucleic acids in a sample are not concentrated and / or amplified before or during sequencing.
[0146] In some embodiments, a representative portion of the genome is sequenced, which may be referred to as "coverage" or "fold coverage". For example, 1-fold coverage implies that approximately 100% of the nucleotide sequence of the genome is represented by the reads. In some cases, fold coverage is referred to as "sequencing depth" (and is proportional to the "sequencing depth"). In some embodiments, "fold coverage" is a relative term that refers to a prior sequencing run. For example, a second sequencing run may have 2-fold less coverage than a first sequencing run. In some embodiments, the genome is sequenced repeatedly, where a given genomic region can be covered by two or more reads or overlapping reads (e.g., a "fold coverage" greater than 1, e.g., 2-fold coverage). In some embodiments, the genome (e.g., the entire genome) is sequenced at about 0.01-fold to about 100-fold coverage, about 0.1-fold to 20-fold coverage, or about 0.1-fold to about 1-fold coverage (e.g., about 0.015, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.08, 0.09, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 40, 50, 60, 70, 80, 90-fold or more coverage). In some embodiments, a particular portion of the genome (e.g., a genomic portion by a targeting method and / or a probe-based method) is sequenced, and the fold coverage value typically refers to a part of that particular genomic portion that has been sequenced (i.e., the fold coverage value does not refer to the entire genome). In some cases, a particular genomic portion is sequenced at a fold coverage of 1000-fold or more. For example, a particular genomic portion can be sequenced at 2000-fold, 5,000-fold, 10,000-fold, 20,000-fold, 30,000-fold, 40,000-fold, or 50,000-fold coverage. In some embodiments, the sequencing is performed at about 1,000-fold to about 100,000-fold coverage.In some embodiments, sequencing is performed at a coverage of about 10,000-fold to about 70,000-fold. In some embodiments, sequencing is performed at a coverage of about 20,000-fold to about 60,000-fold. In some embodiments, sequencing is performed at a coverage of about 30,000-fold to about 50,000-fold.
[0147] In some embodiments, one nucleic acid sample from one individual is sequenced. In certain embodiments, nucleic acids from each of two or more samples are sequenced, where the samples are from one individual or from different individuals. In certain embodiments, nucleic acid samples from two or more biological samples are pooled, where each biological sample is from one individual or from two or more individuals, and the pool is sequenced. In the latter embodiment, nucleic acid samples from each biological sample are often identified by one or more unique identifiers.
[0148] In some embodiments, the sequencing method uses identifiers that enable multiplexing of the sequencing reactions in the sequencing process. The greater the number of unique identifiers, the greater the number of samples and / or chromosomes that can be multiplexed in the sequencing process, e.g., for detection. The sequencing process can be performed using any suitable number of (e.g., 4, 8, 12, 24, 48, 96 or more) unique identifiers.
[0149] The array determination process may utilize a solid phase, which may include a flow cell to which nucleic acids from a library can be attached, through which reagents can be flowed, and which can come into contact with the attached nucleic acids. The flow cell may have flow cell lanes, and the use of identifiers may facilitate the analysis of several samples in each lane. The flow cell is often a solid support configured to hold a reagent solution over the bound analyte and / or to pass a reagent solution over the bound analyte in an orderly manner. The flow cell is often planar in shape, optically transparent, generally on a scale of millimeters or less, and often has channels or lanes where analyte / reagent interactions occur. In some embodiments, the number of samples analyzed in a given flow cell lane depends on the number of unique identifiers used during library preparation and / or probe design. Multiplexing using 12 identifiers, for example, enables the simultaneous analysis of 96 samples (equal to the number of wells in a 96-well microtiter plate), e.g., in an 8-lane flow cell. Similarly, multiplexing using 48 identifiers enables the simultaneous analysis of 384 samples (equal to the number of wells in a 384-well microtiter plate), e.g., in an 8-lane flow cell. Non-limiting examples of commercially available multiplex array determination kits include Illumina's Multiplex Sample Preparation Oligonucleotide Kit and Multiplex Array Determination Primer and PhiX Control Kit (e.g., Illumina catalog numbers PE-400-1001 and PE-400-1002, respectively).
[0150] Any suitable method for sequencing nucleic acids can be used, non-limiting examples of which include Maxim & Gilbert, chain-termination methods, sequencing by synthesis, sequencing by ligation, sequencing by mass spectrometry, microscopy-based techniques, etc. or combinations thereof. In some embodiments, Sanger sequencing methods, including first-generation techniques, such as automated Sanger sequencing methods including microfluidic Sanger sequencing, can be used in the methods provided herein. In some embodiments, sequencing techniques including the use of nucleic acid imaging techniques (e.g., transmission electron microscopy (TEM) and atomic force microscopy (AFM)) can be used. In some embodiments, high-throughput sequencing methods are used. High-throughput sequencing methods generally require clonally amplified DNA templates or single DNA molecules to be sequenced, sometimes in a flow cell, in a massively parallel processing format. Next-generation (e.g., second and third generation) sequencing methods that can sequence DNA in a massively parallel processing format can be used for the methods described herein and are collectively referred to herein as "massively parallel processing sequencing" (MPS). In some embodiments, MPS sequencing methods use a targeted approach where a specific chromosome, gene, or region of interest is sequenced. In certain embodiments, a non-targeted approach is used where most or all of the nucleic acids in a sample are randomly sequenced, amplified, and / or captured.
[0151] In some embodiments, targeted enrichment, amplification and / or sequencing approaches are used. Targeted approaches often isolate, select and / or enrich a subset of nucleic acids in a sample for further processing by using sequence-specific oligonucleotides. In some embodiments, a library of sequence-specific oligonucleotides is used to target (e.g., hybridize to) one or more nucleic acid sets in a sample. The sequence-specific oligonucleotides and / or primers are often selective for specific sequences (e.g., unique nucleic acid sequences) present in one or more target chromosomes, genes, exons, introns and / or regulatory regions of interest. Any suitable method or combination of methods can be used for the enrichment, amplification and / or sequencing of one or more subsets of targeted nucleic acids. In some embodiments, the targeted sequences are isolated and / or enriched by capture on a solid phase (e.g., flow cell, bead) using one or more sequence-specific anchors. In some embodiments, the targeted sequences are enriched and / or amplified by polymerase-based methods (e.g., PCR-based methods, extension based on any suitable polymerase) using sequence-specific primers and / or primer sets. The sequence-specific anchors often can be used as sequence-specific primers.
[0152] MPS array determination may sometimes utilize array determination by synthesis and certain imaging processes. Nucleic acid sequencing techniques that can be used in the methods described herein are sequencing by synthesis and reversible terminator-based sequencing (e.g., Illumina’s Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 2500 (Illumina, San Diego CA)). Using this technology, millions of nucleic acid (e.g., DNA) fragments can be sequenced in parallel. In one example of this type of sequencing technology, a flow cell is used that comprises an optically transparent slide with eight individual lanes on a surface to which oligonucleotide anchors (e.g., adapter primers) are attached.
[0153] Sequence determination by synthesis is typically performed by iteratively adding nucleotides (e.g., covalently) to a primer or an existing nucleic acid strand in a template-specific manner. Each iterative addition of a nucleotide is detected, and the process is repeated multiple times until the sequence of the nucleic acid strand is obtained. The length of the resulting sequence depends in part on the number of addition and detection steps performed. In some embodiments of sequence determination by synthesis, one, two, three or more nucleotides of the same type (e.g., A, G, C or T) are added and detected in a single nucleotide addition. Nucleotides can be added by any suitable method (e.g., enzymatically or chemically). For example, in some embodiments, a polymerase or ligase adds nucleotides to a primer or an existing nucleic acid strand in a template-specific manner. In some embodiments of sequence determination by synthesis, different types of nucleotides, nucleotide analogs and / or identifiers are used. In some embodiments, reversible terminators and / or removable (e.g., cleavable) identifiers are used. In some embodiments, fluorescently labeled nucleotides and / or nucleotide analogs are used. In certain embodiments, sequence determination by synthesis includes cleavage (e.g., cleavage and removal of an identifier) and / or washing steps. In some embodiments, the addition of one or more nucleotides is detected by any suitable method described herein or known in the art, non-limiting examples of which include any suitable imaging device, suitable camera, digital camera, CCD (charge-coupled device)-based imaging device (e.g., a CCD camera), CMOS (complementary metal oxide semiconductor)-based imaging device (e.g., a CMOS camera), photodiode (e.g., a photomultiplier tube), electron microscopy, field effect transistor (e.g., a DNA field effect transistor), ISFET ion sensor (e.g., a CHEMFET sensor), etc. or combinations thereof.
[0154] Any suitable MPS method, system, or technical platform for performing the methods described herein can be used to obtain nucleic acid sequence reads. Non-limiting examples of MPS platforms include Illumina / Solex / HiSeq (e.g., Illumina’s Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ), SOLiD, Roche / 454, PACBIO and / or SMRT, Helicos True Single Molecule Sequencing, Ion Torrent and Ion semiconductor-based sequencing (e.g., developed by Life Technologies), WildFire, 5500, 5500xl W and / or 5500xl W Genetic Analyzer-based technologies (e.g., developed and sold by Life Technologies, U.S. Patent Application Publication No. 2013 / 0012399); polony sequencing, pyrosequencing, massively parallel signature sequencing (MPSS), RNA polymerase (RNAP) sequencing, LaserGen systems and methods, nanopore-based platforms, chemoresistive field effect transistor (CHEMFET) arrays, electron microscopy-based sequencing (e.g., developed by ZS Genetics, Halcyon Molecular), nanoparticle sequencing, etc. or combinations thereof. Other sequencing methods that can be used to perform the methods herein include digital PCR, sequencing by hybridization, nanopore sequencing, chromosome-specific sequencing (e.g., using DANSR (digital analysis of selected regions) technology).
[0155] In some embodiments, the sequence reads are made, obtained, collected, assembled, manipulated, transformed, processed, and / or provided by an array module. A device comprising the array module can be a suitable device and / or apparatus for determining the sequence of a nucleic acid using sequencing techniques known in the art. In some embodiments, the array module can align, assemble, fragment, generate a complement, generate a reverse complement, and / or perform error checking (e.g., correct errors in the sequence reads). complement) and / or perform error checking (e.g., correct errors in the sequence reads).
[0156] Mapping of Reads The sequence reads can be mapped, and the number of reads that map to a particular nucleic acid region (e.g., a chromosome or a part thereof) is referred to as a count. Any suitable mapping method (e.g., a process, algorithm, program, software, module, etc. or a combination thereof) can be used. Certain aspects of the mapping process are described later in this specification.
[0157] The mapping of nucleotide sequence reads (i.e., sequence information from fragments of unknown physical genomic location) can be done in several ways and often involves aligning the resulting sequence reads with a matching sequence in a reference genome. In such an alignment, the sequence reads are typically aligned to the reference sequence, and the aligned sequence reads are referred to as "mapped", "mapped sequence reads" or "mapped reads". In certain embodiments, the mapped sequence reads are referred to as "hits" or "counts". In some embodiments, the mapped sequence reads are grouped together according to various parameters and assigned to specific genomic portions that are discussed in more detail below.
[0158] The terms "aligned," "alignment," or "aligning" generally refer to two or more nucleic acid sequences that can be identified as a match (e.g., 100% identity) or a partial match. Alignment can be performed manually or by a computer (e.g., software, program, module, or algorithm), non-limiting examples of which include the Efficient Local Alignment of Nucleotide Data (ELAND) computer program distributed as part of the Illumina Genomics Analysis pipeline. Alignment of sequence reads can be a 100% sequence match. In some cases, the alignment is a sequence match of less than 100% (i.e., an incomplete match, a partial match, a partial alignment). In some embodiments, the alignment is about 99%, 98%, 97%, 96%, 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, 79%, 78%, 77%, 76%, or 75% match. In some embodiments, the alignment includes mismatches. In some embodiments, the alignment includes 1, 2, 3, 4, or 5 mismatches. Two or more sequences can be aligned using either strand (e.g., the sense strand or the antisense strand). In certain embodiments, a nucleic acid sequence is aligned with the reverse complement strand of another nucleic acid sequence.
[0159] Various computer-implemented methods can be used to map each array read to a portion. Non-limiting examples of computer algorithms that can be used to align arrays include, but are not limited to, BLAST, BLITZ, FASTA, BOWTIE 1, BOWTIE 2, ELAND, MAQ, PROBEMATCH, SOAP, BWA, or SEQMAP, or variants or combinations thereof. In some embodiments, the array reads can be aligned with sequences within a reference genome. In some embodiments, the array reads can be found within nucleic acid databases known in the art and / or aligned with sequences within nucleic acid databases known in the art, such as, for example, GenBank, dbEST, dbSTS, EMBL (European Molecular Biology Laboratory), and DDBJ (DNA Databank of Japan). BLAST or similar tools can be used to search a specified sequence against an array database. The search hits can then be used, for example, to select the specified sequence to an appropriate portion (described later in this specification).
[0160] In some embodiments, the reads can map uniquely or non-uniquely to portions within the reference genome. When a read aligns with a single sequence within the reference genome, that read is considered to be "uniquely mapped." When a read aligns with two or more sequences within the reference genome, that read is considered to be "non-uniquely mapped." In some embodiments, non-uniquely mapped reads are excluded from further analysis (e.g., quantification). In certain embodiments, a certain small number of mismatches (0 to 1) can be tolerated to account for single nucleotide polymorphisms that may exist between the reference genome and the reads from the individual samples being mapped. In some embodiments, no degree of mismatch is tolerated for reads that map to the reference sequence.
[0161] As used herein, the term "reference genome" can refer to any particular known, sequenced, or characterized genome of any organism or virus that can be used to reference a specified sequence, whether in whole or in part, derived from a subject. For example, reference genomes used for human subjects as well as many other organisms can be found at the National Center for Biotechnology Information at the World Wide Web URL ncbi.nlm.nih.gov. "Genome" refers to the complete genetic information of an organism or virus, expressed as a nucleic acid sequence. As used herein, a reference sequence or reference genome is often an assembled genomic sequence or a partially assembled genomic sequence from one or more individuals. In some embodiments, the reference genome is an assembled or partially assembled genomic sequence from one or more human individuals. In some embodiments, the reference genome includes sequences assigned to chromosomes.
[0162] In certain embodiments, mappability is evaluated for a genomic region (e.g., a portion, genomic segment). Mappability is the ability to unambiguously align a nucleotide sequence read to a portion of the reference genome, typically allowing for a specified number of mismatches (e.g., including 0, 1, 2, or more mismatches). For a given genomic region, the expected mappability can be estimated by averaging the resulting read-level mappability values using a sliding-window approach with a pre-set read length. Genomic regions containing contiguous unique nucleotide sequences can sometimes have high mappability values.
[0163] In the case of paired-end sequencing, reads can be mapped to a reference genome by using a suitable mapping program and / or alignment program. Non-limiting examples of such programs include BWA (Li H. and Durbin R. (2009) Bioinformatics 25, 1754-60), Novoalign [Novocraft (2010)], Bowtie (Langmead B et al. (2009) Genome Biol. 10:R25), SOAP2 (Li R et al. (2009) Bioinformatics 25, 1966-67), BFAST (Homer N et al. (2009) PLoS ONE 4, e7767), GASSST (Rizk, G. and Lavenier, D. (2010) Bioinformatics 26, 2534-2540), and MPscan (Rivals E. et al. (2009) Lecture Notes in Computer Science 5724, 246-260), etc. Paired-end reads can be mapped and / or aligned using a suitable short-read alignment program. Non-limiting examples of short-read alignment programs include BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, BWA, CASHX, CUDA-EC, CUSHAW, CUSHAW2, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP, Geneious Assembler, iSAAC, LAST, MAQ, mrFAST, mrsFAST, MOSAIK, MPscan, Novoalign, NovoalignCS, Novocraft, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RTG, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3, SOCS, SSAHA, SSAHA2, Stampy, STORM, Subread, Subjunc, Taipan, UGENE, VelociMapper, TimeLogic, XpressAlign, ZOOM, etc. or combinations thereof. Paired-end reads are often mapped to opposite ends of the same polynucleotide fragment according to the reference genome. In some embodiments, read mates are mapped independently. In some embodiments, information from both sequence reads (i.e., from each end) is considered in the mapping process.A reference genome is often used to determine and / or infer the sequence of nucleic acids located between paired-end read mates. The term "discordant read pair", as used herein, refers to a pair of end reads that contains a pair of read mates that do not clearly map to the same region of the reference genome that is partially defined by a segment of contiguous nucleotides. In some embodiments, a discordant read pair is a paired-end read mate that maps to an unexpected location in the reference genome. Non-limiting examples of unexpected locations in the reference genome include (i) two different chromosomes, (ii) locations that are separated by more than a predetermined fragment size (e.g., more than 300 bp, more than 500 bp, more than 1000 bp, more than 5000 bp, or more than 10,000 bp), (iii) an orientation that does not match the reference sequence (e.g., the reverse orientation), or combinations thereof. In some embodiments, discordant read mates are identified according to the length (e.g., average length, predetermined fragment size) or expected length of the template polynucleotide fragment in the sample. For example, a read mate that maps to a location that is separated by more than the average or expected length of the polynucleotide fragment in the sample may sometimes be identified as a discordant read pair. A read pair that maps in the reverse orientation may sometimes be determined by obtaining the reverse complement of one of those reads and comparing the alignment of both reads using the same strand of the reference sequence. Discordant read pairs can be identified by any suitable method and / or algorithm known in the art or described herein (e.g., SVDetect, Lumpy, BreakDancer, BreakDancerMax, CREST, DELLLY, etc., or combinations thereof).
[0164] portion In some embodiments, the mapped array reads are grouped together according to various parameters and assigned to specific genomic portions (e.g., portions of a reference genome). A "portion" may also be referred to herein as a "genomic section", "bin", "partition", "portion of a reference genome", "portion of a chromosome", or "genomic portion".
[0165] Portions are often defined by dividing the genome according to one or more characteristics. Non-limiting examples of certain characteristics of the division include length (e.g., a defined length, an undefined length) and other structural characteristics. Genomic portions can include one or more of the following characteristics: a defined length, an undefined length, a random length, a non-random length, an equal length, an unequal length (e.g., at least two of the genomic portions have unequal lengths), non-overlapping (e.g., when the 3' end of a genomic portion may be adjacent to the 5' end of an adjacent genomic portion), overlapping (e.g., at least two of the genomic portions overlap), contiguous, continuous, non-contiguous, and non-continuous. Genomic portions can be about 1 to about 1,000 kilobases in length (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900 kilobases in length), about 5 to about 500 kilobases in length, about 10 to about 100 kilobases in length, or about 40 to about 60 kilobases in length.
[0166] Partitioning may be based, or partially based, on features (e.g., information content and information content increase) regarding certain information. Non-limiting examples of features regarding certain information include alignment speed and / or convenience, variation in sequencing coverage, GC content (e.g., stratified GC content, specific GC content, high or low GC content), uniformity of GC content, other measures of sequence content (e.g., ratio of individual nucleotides, ratio of pyrimidines or purines, ratio of natural nucleic acids to unnatural nucleic acids, ratio of methylated nucleotides and CpG content), methylation state, melting temperature of the double strand, amenability to sequencing or PCR, uncertainty values assigned to individual parts of the reference genome, and / or targeted searches for specific features. In some embodiments, the information content can be quantified using a p-value profile that measures the significance of specific genomic positions to distinguish between groups of verified normal and abnormal subjects (e.g., subjects with normal ploidy and trisomy subjects, respectively).
[0167] In some embodiments, by partitioning the genome, similar regions (e.g., identical or homologous regions or sequences) across the genome can be excluded and only unique regions can be maintained. Regions removed in partitioning can be present within a single chromosome, can be one or more chromosomes, or can span multiple chromosomes. In some embodiments, the partitioned genome is reduced and optimized for faster alignment, often focusing on uniquely distinguishable sequences.
[0168] In some embodiments, the genomic portions are generated by dividing based on a predetermined size that does not overlap, thereby resulting in non-overlapping portions of a predetermined length that are contiguous. Such portions are often shorter than a chromosome and often shorter than regions of copy number variation (or copy number change) (e.g., regions that are duplicated or deleted), which may be referred to as segments. A "segment" or "genomic segment" often includes genomic portions of two or more predetermined lengths and often includes two or more contiguous portions of a predetermined length (e.g., about 2 to about 100 such portions (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90 such portions)).
[0169] The plurality of portions may sometimes be analyzed in groups, and the reads mapped to the portions may sometimes be quantified according to a particular group of genomic portions. When the portions are divided by structural features and correspond to regions in the genome, the portions may sometimes be grouped into one or more segments and / or one or more regions. Non-limiting examples of regions include sub-chromosomes (i.e., shorter than a chromosome), chromosomes, autosomes, sex chromosomes, and combinations thereof. One or more sub-chromosomal regions may be genes, gene fragments, regulatory sequences, introns, exons, segments (e.g., segments spanning a copy number change region; segments spanning a copy number variation region), microduplications, microdeletions, etc. The region may sometimes be smaller than or the same size as the chromosome of interest and may sometimes be smaller than or the same size as the reference chromosome.
[0170] Filtering and / or selection of portions In some embodiments, one or more processing steps may include one or more sub-filtering steps and / or sub-selection steps. The term "filtering," as used herein, refers to removing portions of a partial or reference genome from consideration. In certain embodiments, one or more portions are filtered (e.g., subjected to a filtering process) to provide a filtered portion. In some embodiments, the filtering process removes certain portions and retains portions (e.g., a subset of the portions). After the filtering process, the retained portions are often referred to herein as the filtered portions.
[0171] The portions of the reference genome can be selected for removal based on any suitable criteria, such as redundant data (e.g., redundant or overlapping mapped reads), uninformative data (e.g., portions of the reference genome where the median count is zero), portions of the reference genome that may be over-represented or under-represented, noisy data, etc. or combinations of the foregoing, but are not limited thereto. The filtering process often involves removing one or more portions of the reference genome from those to be considered, and subtracting the counts in the one or more portions of the reference genome selected for removal from the counted or summed counts for the reference genome, chromosome, or portion of the genome under consideration. In some embodiments, the portions of the reference genome can be removed sequentially (e.g., one at a time to enable assessment of the impact of removal of each individual portion), and in certain embodiments, all portions of the reference genome marked for removal can be removed simultaneously. In some embodiments, portions of the reference genome characterized by a variance above or below a certain level are removed, which is sometimes referred to herein as filtering of the "noisy" portions of the reference genome. In certain embodiments, the filtering process includes obtaining from the dataset data points that deviate from the average value of the profile level of a portion, chromosome, or part of a chromosome for each variance of a predetermined plurality of profiles, and in certain embodiments, the filtering process includes removing from the dataset data points that do not deviate from the average value of the profile level of a portion, chromosome, or part of a chromosome for each variance of a predetermined plurality of profiles. In some embodiments, the filtering process is used to reduce the number of candidate portions of the reference genome that are analyzed for the presence or absence of genetic mutations / gene changes and / or copy number changes (e.g., aneuploidy, microdeletions, microduplications).A decrease in the number of candidate portions of a reference genome analyzed for the presence or absence of gene mutations / gene changes and / or copy number changes often reduces the complexity and / or dimensionality of a dataset and can increase by two or more orders of magnitude the speed of searching for and / or identifying gene mutations / gene changes and / or copy number changes.
[0172] The portion can be processed (e.g., filtered and / or selected) by any suitable method and according to any suitable parameters. Non-limiting examples of features and / or parameters that can be used to filter and / or select the portion include redundant data (e.g., redundant or overlapping mapped reads), data with no information (e.g., portions of the reference genome with a mapped count of 0), portions of the reference genome containing overrepresented or underrepresented arrays, noisy data, counts, count variations, coverage, mapping quality, variations, measures of repeatability, read density, read density variations, levels of uncertainty, guanine-cytosine (GC) content, CCF fragment length and / or read length (e.g., fragment length ratio (FLR), fetal ratio statistic (FRS)), DNaseI sensitivity, methylation status, acetylation, histone distribution, chromatin structure, percentage of repeats, etc. or combinations thereof. The portion can be filtered and / or selected according to any suitable features or parameters that correlate with the features or parameters listed or described herein. The portion can be filtered and / or selected according to features or parameters specific to the portion (e.g., when measured for a single portion across multiple samples) and / or features or parameters specific to the sample (e.g., when measured for multiple portions within one sample). In some embodiments, the portion is filtered and / or removed according to relatively low mapping quality, relatively large variations, high levels of uncertainty, relatively long CCF fragment length (e.g., low FRS, low FLR), relatively high ratio of repetitive sequences, high GC content, low GC content, low count, zero count, high count, etc. or combinations thereof. In some embodiments, the portion (e.g., a subset of the portions) is selected according to a suitable level of mapping quality, variations, levels of uncertainty, ratio of repetitive sequences, count, GC content, etc. or combinations thereof.In some embodiments, the portions (e.g., a subset of the portions) are selected according to a relatively short CCF fragment length (e.g., high FRS, high FLR). The counts and / or reads mapped to the portions may be processed (e.g., normalized) before and / or after filtering or selecting the portions (e.g., a subset of the portions). In some embodiments, the counts and / or reads mapped to the portions are not processed before and / or after filtering or selecting the portions (e.g., a subset of the portions).
[0173] In some embodiments, the portions may be filtered according to a measure of error (e.g., standard deviation, standard error, calculated variance, p-value, mean absolute error (MAE), mean absolute deviation and / or mean value of the absolute deviation (MAD)). In certain cases, the measure of error may refer to the variation in counts. In some embodiments, the portions are filtered according to the variation in counts. In a particular embodiment, the variation in counts is a measure of error determined for the counts mapped to a portion of the reference genome (i.e., the portion) for a plurality of samples (e.g., a plurality of subjects, e.g., 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more or 10,000 or more subjects). In some embodiments, portions having a variation in counts above a predetermined upper range are filtered (e.g., excluded from consideration). In some embodiments, portions having a variation in counts below a predetermined lower range are filtered (e.g., excluded from consideration). In some embodiments, portions having a variation in counts outside a predetermined range are filtered (e.g., excluded from consideration). In some embodiments, portions having a variation in counts within a predetermined range are selected (e.g., used to determine the presence or absence of a copy number change). In some embodiments, the variation in counts of the portions exhibits a distribution (e.g., a normal distribution). In some embodiments, portions within a certain quantile of that distribution are selected. In some embodiments, portions within the 99th quantile of the distribution of the variation in counts are selected.
[0174] In some embodiments, a method of classifying the presence or absence of a copy number variation in a sub-chromosomal region for a test sample includes a step of identifying using a segmentation process. In some embodiments, the presence or absence of a copy number variation segment may be in a region including a first subset of genomic portions, the region including at least a part of the sub-chromosomal region of interest. As an illustrative example, the region including the first subset of genomic portions is the region surrounded by the black dashed line in FIG. 4. In some embodiments, the first subset of genomic portions is a portion within a region on a chromosome where a copy number variation associated with the phenotype of interest is expected to exist. In some embodiments, such genomic portions can often be obtained by mining public disease databases such as the International Standards of Cytogenomic Arrays database (ISCA). In some embodiments, the genomic portions used herein can be identified by the circular binary segmentation (CBS) algorithm within the sub-chromosomal region of interest. In one embodiment, the phenotype is a microdeletion syndrome. In one embodiment, the first subset of genomic portions is one or more genomic portions selected from 1p36, 22q11.2, 15q11-13, 8q23.2-24.1, 11q24.1, 4p13.3, 17p13.3, and 7q11.23.
[0175] In some embodiments, a method of classifying the presence or absence of a copy number variation in a sub-chromosomal region for a test sample includes a step of providing a quantitative value of sequence reads for a sub-region within a sub-chromosomal region including a subset of genomic portions. The genomic portion includes a portion of a reference genome to which sequence reads obtained for nucleic acids in the test sample are mapped. In some embodiments, the set is a predetermined subset of genomic portions. As an illustrative example, the sub-chromosomal region is the region surrounded by the black dashed line in FIG. 4.
[0176] In some embodiments, a predetermined genomic subset is identified according to one or more accuracy metrics for a plurality of samples in a training set, where each of the plurality of samples in the training set is classified as having a copy number variation in a target subchromosomal region. As described in detail herein, accuracy metrics can include sensitivity, specificity, standard deviation, median absolute deviation (MAD), measures of certainty, measures of confidence, measures of certainty or confidence that a value obtained for a test sample is inside or outside a particular value range, measures of uncertainty, measures of uncertainty that a value obtained for a test sample is inside or outside a particular value range, coefficient of variation (CV), confidence level, confidence interval (e.g., approximately 95% confidence interval), standard score (e.g., z-score), chi value, phi value, results of a t-test, p-value, ploidy value, fitted minor allele ratio, area ratio, median level, etc., or combinations thereof, but are not limited thereto. In some embodiments, the accuracy metric includes sensitivity. The genomic portion is selected based on the genomic portion providing an accuracy metric that is considered optimal, i.e., an accuracy metric equal to or higher than a predetermined threshold (which is considered a minimum requirement for detecting the presence or absence of a copy number variation with reasonable accuracy). For example, when sensitivity is used as the accuracy metric, the threshold can be any value from 70% to 100%, such as 75% to 99%, 80% to 98%, or 85% to 95%.
[0177] In one embodiment, a predetermined genomic subset is identified by a process that includes: 1) providing a plurality of candidate subregions within a subchromosomal region; 2) providing one or more accuracy metrics for each of the plurality of candidate subregions for a plurality of samples in a training set, where each of the plurality of samples is classified as having a copy number variation in the subchromosomal region; and 3) identifying, according to one or more accuracy metrics, the subregion in (a) as the subregion that provides optimal accuracy.
[0178] Array reads from any suitable number of samples can be used to identify a subset of portions that meet one or more criteria, parameters, and / or features described herein. Array reads from a group of samples from multiple subjects may sometimes be used. In some embodiments, the multiple subjects include pregnant women. In some embodiments, the multiple subjects include healthy subjects. In some embodiments, the multiple subjects include cancer patients. One or more samples from each of the multiple subjects (e.g., 1 to about 20 samples from each subject (e.g., about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 samples)) can be addressed, and a suitable number of subjects (e.g., about 2 to about 10,000 subjects (e.g., about 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000 subjects)) can be addressed. In some embodiments, array reads from the same test sample from the same subject are mapped to portions in the reference genome and used to generate a subset of portions.
[0179] The portions can be selected and / or filtered by any suitable method. In some embodiments, the portions are selected according to a visual inspection of data, graphs, plots, and / or charts. In certain embodiments, the portions are selected and / or filtered (e.g., in part) by a system or device comprising one or more microprocessors and memory. In some embodiments, the portions are selected and / or filtered (e.g., in part) by a non-transitory computer-readable storage medium storing an executable program, where the program instructs the microprocessor to perform the selection and / or filtering.
[0180] In some embodiments, sequence reads derived from a sample are mapped to all or most parts of a reference genome, and then a subset of a preselected part is selected. For example, a subset of the parts to which reads from fragments at a threshold of a specific length preferentially map can be selected. A certain method for preselecting a subset of parts is described in U.S. Patent Application Publication No. 2014 / 0180594, which is incorporated herein by reference. Reads from the selected subset of parts are often used, for example, in a further step of determining the presence or absence of a genetic mutation or genetic change. Often, reads from unselected parts are not used in a further step of determining the presence or absence of a genetic mutation or genetic change (e.g., reads in unselected parts are removed or filtered).
[0181] In some embodiments, portions related to read density (e.g., where the read density is for a particular portion) are removed by a filtering process, and the read density related to the removed portions is not included in the determination of the presence or absence of copy number variations (e.g., aneuploidy, microduplications, microdeletions). In some embodiments, the read density profile includes and / or consists of the read density of the filtered portions. The portions may be filtered according to the distribution of counts and / or the distribution of read density. In some embodiments, the portions are filtered according to the distribution of counts and / or read density, where those counts and / or read density are obtained from one or more reference samples. One or more reference samples may be referred to herein as a training set. In some embodiments, the portions are filtered according to the distribution of counts and / or read density, where those counts and / or read density are obtained from one or more test samples. In some embodiments, the portions are filtered according to a measure of uncertainty with respect to the read density distribution. In a particular embodiment, portions showing large deviations in read density are removed by the filtering process. For example, the distribution of read density (e.g., the distribution of the mean of the means or the median of the read density) can be determined, where each read density in the distribution maps to the same portion. The measure of uncertainty (e.g., MAD) can be determined by comparing the distribution of read density for multiple samples, where each portion of the genome is associated with a measure of uncertainty. According to the foregoing example, the portions can be filtered according to the measure of uncertainty (e.g., standard deviation (SD), MAD) associated with each portion and a predetermined threshold. In a particular case, portions including MAD values within an acceptable range are retained, and portions including MAD values outside the acceptable range are removed from consideration by the filtering process.In some embodiments, according to the foregoing examples, portions including lead density values outside a given measure of uncertainty (e.g., median, mean, or average of lead density) are often removed from what is to be considered by a filtering process. In some embodiments, portions including lead density values outside the interquartile range of a distribution (e.g., median, mean, or average of lead density) are removed from what is to be considered by a filtering process. In some embodiments, portions including lead density values more than two, three, four, or five times outside the interquartile range of a distribution are removed from what is to be considered by a filtering process. In some embodiments, portions including lead density values more than two, three, four, five, six, seven, or eight sigmas outside (e.g., sigma is a range defined by standard deviation) are removed from what is to be considered by a filtering process.
[0182] Quantitative value of array reads Array reads mapped or partitioned based on selected features or variables can, in some embodiments, be quantified to measure the amount or number of reads mapped to one or more portions (e.g., portions of a reference genome). In certain embodiments, the amount of array reads mapped to a particular portion or segment is referred to as a count or lead density.
[0183] Counts are often associated with genomic portions. In some embodiments, counts are measured from some or all of the array reads mapped to (i.e., associated with) a portion. In certain embodiments, counts are measured from some or all of the array reads mapped to a group of portions (e.g., portions within a segment or region (described herein)).
[0184] The count can be measured by a suitable method, operation, or mathematical process. The count can be the direct sum of all sequence reads mapped to a genomic portion or group of genomic portions corresponding to a segment, or a group of portions corresponding to a sub-region of the genome (e.g., a copy number variant region, a copy number change region, a copy number duplication region, a copy number deletion region, a micro-duplication region, a micro-deletion region, a chromosomal region, an autosomal region, a sex chromosomal region), and / or can be a group of portions corresponding to the genome. The quantitative value of a read can be a ratio, and can be a ratio of the quantitative value for a portion in region a to the quantitative value for a portion in region b. Region a can be one portion, a segment region, a copy number variant region, a copy number change region, a copy number duplication region, a copy number deletion region, a micro-duplication region, a micro-deletion region, a chromosomal region, an autosomal region, and / or a sex chromosomal region. Region b can independently be one portion, a segment region, a copy number variant region, a copy number change region, a copy number duplication region, a copy number deletion region, a micro-duplication region, a micro-deletion region, a chromosomal region, an autosomal region, a sex chromosomal region, a region including all autosomes, a region including a sex chromosome, and / or a region including all chromosomes.
[0185] In some embodiments, the count is obtained from raw sequence reads and / or filtered sequence reads. In certain embodiments, the count is the average, mean value, or sum of sequence reads mapped to a genomic portion or group of genomic portions (e.g., genomic portions within a region). In some embodiments, the count is associated with an uncertainty value. The count can be adjusted. The count can be adjusted according to sequence reads associated with a genomic portion or group of portions for which weighting, removal, filtering, normalization, adjustment, averaging, derivation as a mean value, derivation as a median value, addition, or combinations thereof have been performed.
[0186] The quantitative value of an array read can sometimes be the read density. The read density can be measured and / or generated for one or more segments of the genome. In certain cases, the read density can be measured and / or generated for one or more chromosomes. In some embodiments, the read density includes a quantitative measure of the count of array reads mapped to a segment or portion of a reference genome. The read density can be measured by a suitable process. In some embodiments, the read density is measured by a suitable distribution and / or a suitable distribution function. Non-limiting examples of distribution functions include any suitable distribution or combinations thereof such as a probability function, a probability distribution function, a probability density function (PDF), a kernel density function (kernel density estimation), a cumulative distribution function, a probability mass function, a discrete probability distribution, an absolutely continuous univariate distribution, etc. The read density can be a density estimate derived from a suitable probability density function. A density estimate is the construction of an estimate based on the observed data of a potential probability density function. In some embodiments, the read density includes a density estimate (e.g., a probability density estimate, a kernel density estimate). The read density can be generated according to a process that includes generating a density estimate for each of one or more portions of the genome, where each portion includes a count of array reads. The read density can be generated for a normalized and / or weighted count mapped to a portion or segment. In some cases, each read mapped to a portion or segment can contribute to the read density with a value (e.g., a count) equal to its weight obtained from the normalization process described herein. In some embodiments, the read density for one or more portions or segments is adjusted. The read density can be adjusted by a suitable method. For example, the read density for one or more portions can be weighted and / or normalized.
[0187] The reads quantified for a given portion or segment can be from one origin or different origins. In one example, the reads may be obtained from nucleic acids from a subject having or suspected of having cancer. In such a situation, reads mapped to one or more portions are often reads that represent both healthy cells (i.e., non-cancer cells) and cancer cells (e.g., tumor cells). In certain embodiments, some of the reads mapped to a portion are from cancer cell nucleic acids and some of the reads mapped to the same portion are from non-cancer cell nucleic acids. In another example, the reads may be obtained from a nucleic acid sample from a pregnant woman having a fetus. In such a situation, reads mapped to one or more portions are often reads that represent both the fetus and the mother of the fetus (e.g., the pregnant subject). In certain embodiments, some of the reads mapped to a portion are from the genome of the fetus and some of the reads mapped to the same portion are from the genome of the mother.
[0188] Level In some embodiments, a value (e.g., a number, a quantitative value) is binned into levels. The levels can be determined by a suitable method, operation, or mathematical process (e.g., a processed level). The level is often a count for a subset (e.g., a normalized count) or is derived from such a count. In some embodiments, the level of a portion is substantially equal to the total count (e.g., a count, a normalized count) mapped to that portion. The level is often determined from a count that has been processed, transformed, or manipulated by a suitable method, operation, or mathematical process known in the art. In some embodiments, a level is derived from a processed count, and non-limiting examples of processed counts include weighted counts, removed counts, filtered counts, normalized counts, adjusted counts, averaged counts, counts derived as an average value (e.g., an average value level), added counts, subtracted counts, transformed counts, or combinations thereof. In some embodiments, a level includes a normalized count (e.g., a normalized count of a portion). A level can be with respect to a count normalized by a suitable process, non-limiting examples of which are described herein. A level can include a normalized count or a relative amount of a count. In some embodiments, a level is with respect to the average of two or more portion counts or normalized counts, and that level is referred to as an average level. In some embodiments, a level is with respect to a subset having an average value of a count or an average value of a normalized count, which is referred to as an average value level. In some embodiments, a level is derived for a portion that includes raw counts and / or filtered counts. In some embodiments, a level is based on raw counts. In some embodiments, a level is related to an uncertainty value (e.g., a standard deviation, MAD). In some embodiments, a level is represented by a Z-score or a p-value.
[0189] The level for one or more parts is synonymous with "genomic segment level" herein. The term "level" may be synonymous with the term "height" when used herein. The determination of the meaning of the term "level" can be made from the context in which it is used. For example, the term "level" often means height when used in the context of a part, profile, read, and / or count. The term "level" often refers to an amount when used in the context of a substance or composition (e.g., the level of RNA, the plexing level, the amount). The term "level" often refers to an amount when used in the context of uncertainty (e.g., the level of error, the level of confidence, the level of deviation, the level of uncertainty).
[0190] The normalized or non-normalized counts for two or more levels (e.g., two or more levels in a profile) may sometimes be mathematically manipulated according to the levels (e.g., added, multiplied, averaged, normalized, etc. or combinations thereof). For example, the normalized or non-normalized counts for two or more levels may be normalized according to one, some, or all of the levels in a profile. In some embodiments, the normalized or non-normalized counts for all levels in a profile are normalized according to one level in that profile. In some embodiments, the normalized or non-normalized counts for a first level in a profile are normalized according to the normalized or non-normalized counts for a second level in that profile.
[0191] Non-limiting examples of levels (e.g., a first level, a second level) include levels for subsets that include processed counts, levels for subsets that include the average value, median, or mean of counts, levels for subsets that include normalized counts, etc. or any combination thereof. In some embodiments, the first level and the second level in a profile are derived from counts of portions mapped to the same chromosome. In some embodiments, the first level and the second level in a profile are derived from counts of portions mapped to different chromosomes.
[0192] In some embodiments, a level is determined from normalized or non-normalized counts mapped to one or more portions. In some embodiments, a level is determined from normalized or non-normalized counts mapped to two or more portions, where the normalized counts for each portion are often approximately the same. Variations in counts (e.g., normalized counts) can exist in a subset for a level. In a subset for a level, there can be one or more portions having counts that are significantly different from other portions of the set (e.g., peaks and / or dips). Any suitable number of normalized or non-normalized counts associated with any suitable number of portions can define a level.
[0193] In some embodiments, one or more levels can be determined from all or some of the normalized or non-normalized counts of a portion of a genome. Often, a level can be determined from all or some of the normalized or non-normalized counts of a chromosome or a portion thereof. In some embodiments, two or more counts derived from two or more portions (e.g., subsets) are used to determine a level. In some embodiments, two or more counts (e.g., counts from two or more portions) are used to determine a level. In some embodiments, counts from 2 to about 100,000 portions are used to determine a level. In some embodiments, counts from 2 to about 50,000, 2 to about 40,000, 2 to about 30,000, 2 to about 20,000, 2 to about 10,000, 2 to about 5000, 2 to about 2500, 2 to about 1250, 2 to about 1000, 2 to about 500, 2 to about 250, 2 to about 100, or 2 to about 60 portions are used to determine a level. In some embodiments, counts from about 10 to about 50 portions are used to determine a level. In some embodiments, counts from about 20 to about 40 or more portions are used to determine a level. In some embodiments, a level includes counts from about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60 or more portions. In some embodiments, a level corresponds to a subset (e.g., a subset of a reference genome, a subset of a chromosome, or a subset of a portion of a chromosome).
[0194] In some embodiments, a level is determined for a normalized or non-normalized count of consecutive portions. In some embodiments, consecutive portions (e.g., subsets) correspond to adjacent regions of a genome or adjacent regions of a chromosome or gene. For example, two or more consecutive portions can be a sequence assembly of a DNA sequence longer than each portion when they are aligned by attaching and merging those portions end to end. For example, two or more consecutive portions can be an intact genome, chromosome, gene, intron, exon, or a part thereof. In some embodiments, a level is determined from a set of consecutive portions and / or non-consecutive portions (e.g., a set).
[0195] Data processing and normalization The counted and mapped array reads are referred to herein as raw data. This is because the data corresponds to an unmanipulated count (e.g., a raw count). In some embodiments, the data of the array reads within a dataset can be further processed (e.g., mathematically and / or statistically manipulated) and / or displayed to facilitate the provision of an outcome. In certain embodiments, a dataset containing a larger dataset can benefit from preprocessing to facilitate further analysis. Preprocessing of a dataset can include the removal of redundant and / or uninformative portions or portions of the reference genome (e.g., portions of the reference genome with uninformative data, redundant mapped reads, portions where the median count is zero, overrepresented or underrepresented arrays). Without being limited to theory, data processing and / or preprocessing can (i) remove noisy data, (ii) remove uninformative data, (iii) remove redundant data, (iv) reduce the complexity of a larger dataset, and / or (v) facilitate the conversion of data from one form to one or more other forms. The terms "preprocessing" and "processing" when used with respect to data or a dataset are collectively referred to herein as "processing". Processing can make the data applicable to further analysis and, in some embodiments, can generate an outcome. In some embodiments, one or more processing methods or all processing methods (e.g., normalization methods, filtering of portions, mapping, validation, etc. or combinations thereof) are performed by a processor connected to memory, a microprocessor, a computer, and / or a device controlled by a microprocessor.
[0196] As used herein, the term "noisy data" refers to (a) data that has significant variance between data points when analyzed or plotted, (b) data that has a significant standard deviation (e.g., a standard deviation greater than 3), (c) data that has a significant standard error of the mean, etc., and combinations of the foregoing. Noisy data can sometimes result from the amount and / or quality of the starting material (e.g., nucleic acid sample) and can sometimes occur as part of the process for preparing or replicating the DNA used to generate sequence reads. In certain embodiments, the noise is due to certain sequences that are overrepresented when prepared using PCR-based methods. The methods described herein can reduce or eliminate the involvement of noisy data and, thus, reduce the impact of noisy data on the provided outcome.
[0197] The terms "data without information value", "reference genome portion without information value", and "portion without information value", when used herein, refer to a portion having a numerical value that is significantly different from a predetermined threshold or does not fall within a cut-off range of a predetermined value, or data derived therefrom. The terms "threshold value" and "threshold" in this specification refer to any number calculated using a qualified dataset and serving as a limit for the diagnosis of gene mutations or gene changes (e.g., copy number changes, aneuploidy, microduplications, microdeletions, chromosomal abnormalities, etc.). In certain embodiments, the threshold value is exceeded by the result obtained by the method described herein, and the subject is diagnosed with a copy number change. The threshold value or range of values is often calculated in some embodiments by mathematically and / or statistically manipulating sequence read data (e.g., sequence read data from a reference and / or a subject), and in certain embodiments, the sequence read data manipulated to generate the threshold value or range of values is sequence read data (e.g., sequence read data from a reference and / or a subject). In some embodiments, an uncertainty value is determined. The uncertainty value is generally a measure of variance or error and can be any suitable measure of variance or error. In some embodiments, the uncertainty value is the standard deviation, standard error, calculated variance, p-value, or mean absolute deviation (MAD). In some embodiments, the uncertainty value can be calculated according to the formulas described herein.
[0198] Any suitable procedure may be used to process the datasets described herein. Non-limiting examples of procedures suitable for use in processing the datasets include filtering, normalization, weighting, monitoring peak height, monitoring peak area, monitoring peak ends, peak level analysis, peak width analysis, peak end position analysis, peak lateral tolerances, measurement of area ratios, mathematical processing of data, statistical processing of data, application of statistical algorithms, analysis using fixed variables, analysis using optimized variables, plotting of data to identify patterns or trends for further processing, and combinations of the foregoing. In some embodiments, the datasets are processed based on various features (e.g., GC content, mapped redundant reads, centromere regions, telomere regions, etc. and combinations thereof) and / or variables (e.g., subject gender, subject age, subject ploidy, percent contribution of cancer cell nucleic acid, fetal gender, maternal age, maternal ploidy, percent contribution of fetal nucleic acid, etc. or combinations thereof). In certain embodiments, the processing of the datasets as described herein can reduce the complexity and / or dimensionality of large and / or complex datasets. Non-limiting examples of complex datasets include sequence read data generated from one or more test subjects and multiple reference subjects of different ages and ethnic backgrounds. In some embodiments, the datasets can include thousands to millions of sequence reads for each test subject and / or each reference subject.
[0199] Data processing can be performed in any number of steps in certain embodiments. For example, in some embodiments, data can be processed using only one processing procedure, and in certain embodiments, data can be processed using one or more, five or more, ten or more, or twenty or more processing steps (e.g., one or more processing steps, two or more processing steps, three or more processing steps, four or more processing steps, five or more processing steps, six or more processing steps, seven or more processing steps, eight or more processing steps, nine or more processing steps, ten or more processing steps, eleven or more processing steps, twelve or more processing steps, thirteen or more processing steps, fourteen or more processing steps, fifteen or more processing steps, sixteen or more processing steps, seventeen or more processing steps, eighteen or more processing steps, nineteen or more processing steps, or twenty or more processing steps). In some embodiments, the processing steps can be the same steps repeated two or more times (e.g., filtering two or more times, normalizing two or more times), and in certain embodiments, the processing steps can be two or more different processing steps performed simultaneously or sequentially (e.g., filtering, normalizing; normalizing peak height and peak end, monitoring; filtering, normalizing, normalizing against a reference, statistical operations to determine p-values, etc.). In some embodiments, any suitable number and / or combination of the same or different processing steps can be used to process the array read data to facilitate the provision of an outcome. In certain embodiments, the processing of a data set according to the criteria described herein can reduce the complexity and / or dimensionality of the data set.
[0200] In some embodiments, one or more processing steps may include one or more normalization steps. Normalization may be performed by any suitable method described herein or known in the art. In certain embodiments, normalization includes adjusting values measured on different scales to a conceptually common scale. In certain embodiments, normalization includes sophisticated mathematical adjustment to align the probability distribution of the adjusted values. In some embodiments, normalization includes aligning the distribution to a normal distribution. In certain embodiments, normalization includes mathematical adjustment that enables comparison of corresponding normalized values for different datasets so as to eliminate the effect of certain overall effects (e.g., errors and exceptions). In certain embodiments, normalization includes scaling. Normalization may sometimes include division of one or more datasets by a given variable or formula. Normalization may sometimes include subtraction of one or more datasets by a given variable or formula. Non-limiting examples of normalization methods include per-segment normalization, normalization by GC content, normalization of the median count (median bin count, median sub-count), linear and non-linear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted scatterplot smoothing), principal component normalization, repeat mask (RM), GC normalization and repeat mask (GCRM), cQn and / or combinations thereof. In some embodiments, determination of the presence or absence of copy number variations (e.g., aneuploidy, microduplications, microdeletions) is made using a normalization method (e.g., per-segment normalization, normalization by GC content, normalization of the median count (median bin count, median sub-count), linear and non-linear least squares regression, LOESS, GC LOESS, LOWESS (locally weighted scatterplot smoothing), principal component normalization, repeat mask (RM), GC normalization and repeat mask (GCRM), cQn, normalization methods known in the art, and / or combinations thereof). Certain examples of normalization processes that may be used, e.g., LOESS normalization, principal component normalization, and hybrid normalization methods, are described in more detail later in this specification.Aspects of certain normalization processes are also described, for example, in International Patent Application Publication No. WO2013 / 052913 and International Patent Application Publication No. WO2015 / 051163, each of which is incorporated herein by reference.
[0201] Any suitable number of normalizations can be used. In some embodiments, the dataset can be normalized one or more times, five or more times, ten or more times, or even twenty or more times. The dataset can be normalized with respect to values (e.g., normalized values) representing any suitable feature or variable (e.g., sample data, reference data, or both). Non-limiting examples of types of data normalization that can be used include normalizing raw count data for one or more selected test portions or reference portions to the total count mapped to the chromosome or entire genome to which the selected portion or bin is mapped; normalizing raw count data for one or more selected portions to the median of the reference counts for one or more portions or chromosomes to which the selected portion is mapped; normalizing raw count data with respect to pre-normalized data or its derivative; and normalizing pre-normalized data with respect to one or more other predetermined normalization variables. Normalization of a dataset can sometimes have the effect of decoupling statistical errors, depending on the feature or characteristic selected as the predetermined normalization variable. Normalization of a dataset can sometimes also enable comparison of data characteristics of data having different scales by bringing the data to a common scale (e.g., a predetermined normalization variable). In some embodiments, one or more normalizations with respect to statistically derived values can be used to minimize data differences and reduce the significance of out-of-range data. Normalizing a portion or a portion of a reference genome with respect to a normalized value is sometimes referred to as "per-portion normalization."
[0202] In certain embodiments, the processing steps may include one or more mathematical and / or statistical operations. Any suitable mathematical and / or statistical operations may be used alone or in combination to analyze and / or manipulate the data sets described herein. Any suitable number of mathematical and / or statistical operations may be used. In some embodiments, the data set may be mathematically and / or statistically manipulated one or more times, five or more times, ten or more times, or twenty or more times. Non-limiting examples of mathematical and statistical operations that may be used include addition, subtraction, multiplication, division, algebraic functions, least squares estimators, curve fitting, differential equations, rational polynomials, double polynomials, orthogonal polynomials, z-scores, p-values, chi values, phi values, peak level analysis, determination of peak end positions, calculation of peak area ratios, analysis of chromosomal level medians, calculation of mean absolute deviations, sum of squared residuals, mean values, standard deviations, standard errors, etc. or combinations thereof. The mathematical and / or statistical operations may be performed on all or part of the array read data or its processed version. Non-limiting examples of variables or features of the data set that may be statistically manipulated include raw counts, filtered counts, normalized counts, peak heights, peak widths, peak areas, peak ends, lateral tolerance, P-values, median levels, mean levels, distribution of counts within genomic regions, relative representation of nucleic acid species, etc. or combinations thereof.
[0203] In some embodiments, the processing step may include the use of one or more statistical algorithms. Any suitable statistical algorithm may be used alone or in combination to analyze and / or manipulate the datasets described herein. Any suitable number of statistical algorithms can be used. In some embodiments, the dataset may be analyzed using one or more, five or more, ten or more, or twenty or more statistical algorithms. Non-limiting examples of statistical algorithms suitable for use with the methods described herein include principal component analysis, decision trees, null hypothesis, multiple comparisons, omnibus tests, Behrens-Fisher problem, bootstrapping, Fisher's method for combining independent significance tests, null hypothesis, type I error, type II error, exact tests, one-sample Z test, two-sample Z test, one-sample t test, paired t test, pooled two-sample t test with equal variances, unpooled two-sample t test with unequal variances, one-proportion z test, pooled two-proportion z test, unpooled two-proportion z test, one-sample chi-square test, two-sample F test for equalizing variances, confidence intervals, credibility intervals, significance, meta-analysis, linear simple regression, robust linear regression, etc. or combinations of the foregoing. Non-limiting examples of variables or features of the dataset that may be analyzed using statistical algorithms include raw counts, filtered counts, normalized counts, peak height, peak width, peak ends, lateral tolerance, P values, median levels, mean levels, distribution of counts within genomic regions, relative representation of nucleic acid species, etc. or combinations thereof.
[0204] In certain embodiments, a dataset may be analyzed by using a plurality of (e.g., two or more) statistical algorithms (e.g., least squares regression, principal component analysis, linear discriminant analysis, quadratic discriminant analysis, bagging, neural networks, support vector machine models, random forests, classification tree models, k-nearest neighbor, logistic regression, and / or smoothing methods) and / or mathematical and / or statistical operations (e.g., those referred to as operations herein). In some embodiments, the use of a plurality of operations may generate an N-dimensional space that can be used to provide an outcome. In certain embodiments, the analysis of a dataset by using a plurality of operations may reduce the complexity and / or dimensionality of that dataset. For example, by using a plurality of operations on a reference dataset, an N-dimensional space (e.g., a probability plot) can be generated that can be used to represent the presence or absence of gene mutations / gene changes and / or copy number changes depending on the state of the reference sample (e.g., positive or negative for a selected copy number change). Analysis of a test sample using a substantially similar set of operations can be used to generate an N-dimensional point for each test sample. The complexity and / or dimensionality of a dataset of a test subject may sometimes be reduced to a single value or an N-dimensional point that can be easily compared to the N-dimensional space generated from the reference data. Data of a test sample that falls within the N-dimensional space occupied by the data of a reference subject suggests a genetic state that is substantially similar to the genetic state of the reference subject. Data of a test sample that does not fall within the N-dimensional space occupied by the data of a reference subject suggests a genetic state that is substantially different from the genetic state of the reference subject. In some embodiments, the reference is euploid or otherwise has no gene mutations / gene changes and / or copy number changes and / or medical symptoms.
[0205] After the dataset is counted, filtered if necessary, normalized, and weighted if necessary, the processed dataset can, in some embodiments, be further manipulated by one or more filtering procedures and / or normalization procedures and / or weighting procedures. The dataset further manipulated by one or more filtering procedures and / or normalization procedures and / or weighting procedures can, in certain embodiments, be used to generate a profile. One or more filtering procedures and / or normalization procedures and / or weighting procedures can, in some embodiments, sometimes reduce the complexity and / or dimensionality of the dataset. The outcome can be provided based on the dataset with reduced complexity and / or dimensionality. In some embodiments, for example, a plot of the profile of the processed data further manipulated by weighting is generated to facilitate classification and / or outcome provision. The outcome can, for example, be provided based on a plot of the profile of the weighted data.
[0206] Partial filtering or weighting can be performed at one or more suitable points in the analysis. For example, the portion can be filtered or weighted before or after the array reads are mapped to a portion of the reference genome. The portion can, in some embodiments, be filtered or weighted before or after the experimental bias for an individual genomic portion is determined. In certain embodiments, the portion can be filtered or weighted before or after the level is calculated.
[0207] After the dataset is counted, filtered if necessary, normalized, and weighted if necessary, the processed dataset can, in some embodiments, be manipulated by one or more mathematical and / or statistical operations (e.g., statistical functions or statistical algorithms). In certain embodiments, the processed dataset can be further manipulated by calculating Z-scores for one or more selected portions, chromosomes, or portions of chromosomes. In some embodiments, the processed dataset can be further manipulated by calculating P-values. In certain embodiments, the mathematical and / or statistical operations include one or more hypotheses regarding ploidy and / or the ratio of minority species (e.g., the ratio of cancer cell nucleic acids; the fetal ratio). In some embodiments, a plot of the profile of the processed data further manipulated by one or more statistical and / or mathematical operations is generated to facilitate classification and / or the provision of an outcome. The outcome can be provided based on a plot of the profile of the statistically and / or mathematically manipulated data. The outcome provided based on a plot of the profile of the statistically and / or mathematically manipulated data often includes one or more hypotheses regarding ploidy and / or the ratio of minority species (e.g., the ratio of cancer cell nucleic acids; the fetal ratio).
[0208] In some embodiments, the analysis and processing of data can include the use of one or more assumptions. A suitable number or type of assumptions can be used to analyze or process a data set. Non-limiting examples of assumptions that can be used for data processing and / or analysis include ploidy of a subject, contribution of cancer cells, maternal ploidy, fetal contribution, prevalence of a particular sequence in a reference population, ethnic background, prevalence of selected medical conditions in related families, similarity between raw count profiles from different patients and / or between runs after GC normalization and repeat masking (e.g., GCRM), perfect matches representing PCR artifacts (e.g., at the same base position), assumptions specific to nucleic acid quantification assays (e.g., fetal quantification assay (FQA)), assumptions regarding twins (e.g., when both and only one of the twins are affected, the effective fetal ratio is only 50% of the sum of the measured fetal ratios (similarly for triplets, quadruplets, etc.)), cell-free DNA (e.g., cfDNA) covering the entire genome uniformly, and combinations thereof.
[0209] If the quality and / or depth of the mapped array reads do not enable prediction of the presence or absence of gene mutations / gene changes and / or copy number changes at a desired confidence level (e.g., 95% or higher confidence level) based on the normalized count profile, one or more additional mathematical operation algorithms and / or statistical prediction algorithms may be used to generate additional numerical values useful for data analysis and / or providing an outcome. The term "normalized count profile", as used herein, refers to a profile generated using normalized counts. Examples of methods that may be used to generate normalized counts and normalized count profiles are described herein. As stated, the mapped and counted array reads may be normalized with respect to the counts of the test sample or the counts of the reference sample. In some embodiments, the normalized count profile may be presented as a plot.
[0210] Non-limiting examples of processing steps and normalization methods that may be used, such as normalization for a window (static or sliding), weighting, determination of bias relationships, LOESS normalization, principal component normalization, hybrid normalization, generation and comparison of profiles, are described in more detail later in this specification.
[0211] Normalization for a window (static or sliding) In certain embodiments, the processing step includes normalization for a static window, and in some embodiments, the processing step includes normalization for a moving or sliding window. As used herein, the term "window" refers to one or more portions that are selected for analysis and used as a reference for comparison (e.g., for normalization and / or other mathematical or statistical operations). The term "normalization for a static window" as used herein refers to a normalization process that uses one or more portions selected for comparison of a dataset of a test subject and a dataset of a reference subject. In some embodiments, the selected portions are used to generate a profile. A static window generally includes a predetermined subset that does not change during operation and / or analysis. The terms "normalization for a moving window" and "normalization for a sliding window" as used herein refer to normalization performed on portions located in the genomic region of a selected test portion (e.g., portions immediately surrounding, adjacent to, or segments thereof), where one or more selected test portions are normalized with respect to portions immediately surrounding the selected test portion. In certain embodiments, the selected portions are used to generate a profile. Sliding window normalization or moving window normalization often involves repeatedly moving or sliding to adjacent test portions and normalizing the newly selected test portion with respect to portions immediately surrounding or adjacent to the newly selected test portion, where adjacent windows share one or more portions. In certain embodiments, multiple selected test portions and / or chromosomes can be analyzed by a sliding window process.
[0212] In some embodiments, normalization for a sliding window or moving window may generate one or more values, where each value corresponds to normalization against a different subset of reference portions selected from different genomic regions (e.g., chromosomes). In certain embodiments, the one or more generated values are estimated numerical values of the integral of a normalized count profile for a cumulative sum (e.g., a selected portion, domain (e.g., a part of a chromosome) or chromosome). The values generated by the sliding window or moving window process can be used to generate a profile and facilitate reaching an outcome. In some embodiments, the cumulative sum of one or more portions can be displayed as a function of genomic position. Moving window analysis or sliding window analysis may sometimes be used to analyze the genome for the presence or absence of microdeletions and / or microduplications. In certain embodiments, the display of the cumulative sum of one or more portions is used to identify the presence or absence of regions of copy number variation (e.g., microdeletions, microduplications). Weighting
[0213] In some embodiments, the processing step includes weighting. The terms "weighted," "weighting," or "weight function" or their grammatical derivatives or equivalents, when used herein, change the influence of the features or variables of a particular dataset on the features or variables of other datasets (e.g., increase or decrease the significance and / or contribution of the data contained in one or more parts or portions of the reference genome based on the quality or usefulness of the data in the selected part or portion of the reference genome). It refers to some or all of the mathematical operations of a dataset that may be used to). The weighting function can be used in some embodiments to increase the influence of data with a relatively small variance in measurements and / or to decrease the influence of data with a relatively large variance in measurements. For example, a portion of the reference genome with underrepresented or low-quality array data can be "weighted down" to minimize its impact on the dataset, while a selected portion of the reference genome can be "weighted up" to increase its impact on the dataset. A non-limiting example of a weighting function is [1 / (standard deviation) 2 . Partial weighting can sometimes eliminate partial dependencies. In some embodiments, one or more parts are weighted by a specific function (e.g., an eigenfunction). In some embodiments, a specific eigenfunction includes replacing a part with an orthogonal eigenpart. The weighting step may be performed in a manner substantially similar to the normalization step. In some embodiments, the dataset is adjusted (e.g., divided, multiplied, added, subtracted) with a predetermined variable (e.g., a weighting variable). In some embodiments, the dataset is divided by a predetermined variable (e.g., a weighting variable). A predetermined variable (e.g., a minimized objective function, Phi) is often selected to differentially weight different parts of the dataset (e.g., increase the influence of a particular data type while decreasing the influence of other data types).
[0214] Relationship of bias In some embodiments, the processing step includes determining the relationship of the bias. For example, one or more relationships are generated between the local genomic bias estimate and the bias frequency. The term "relationship", as used herein, refers to a mathematical and / or graphical relationship between two or more variables or values. A certain relationship can be generated by a suitable mathematical process and / or graphical process. Non-limiting examples of relationships include mathematical representations and / or graphical representations such as functions, correlations, distributions, linear equations or non-linear equations, lines, regressions, fitted regressions, etc. or combinations thereof. A relationship may sometimes include a fitting relationship. In some embodiments, the fitting relationship includes a fitted regression. A relationship may sometimes include two or more weighted variables or values. In some embodiments, a certain relationship includes a fitted regression in which one or more variables or values of the relationship are weighted. The regression may sometimes be applied in a weighted form. The regression may sometimes be applied without weighting. In a particular embodiment, generating a relationship includes plotting or graphically representing.
[0215] In a particular embodiment, a relationship is generated between the GC density and the GC density frequency. In some embodiments, a sample GC density relationship is provided by generating a relationship between (i) the GC density and (ii) the GC density frequency for a sample. In some embodiments, a reference GC density relationship is provided by generating a relationship between (i) the GC density and (ii) the GC density frequency for a reference. In some embodiments, when the local genomic bias estimate is the GC density, the sample bias relationship is the sample GC density relationship and the reference bias relationship is the reference GC density relationship. The GC density of the reference GC density relationship and / or the sample GC density relationship is often a presentation (e.g., a mathematical presentation or a quantitative presentation) of the local GC content.
[0216] In some embodiments, the relationship between the local genomic bias estimate and the bias frequency includes a distribution. In some embodiments, the relationship between the local genomic bias estimate and the bias frequency includes a fitting relationship (e.g., fitting regression). In some embodiments, the relationship between the local genomic bias estimate and the bias frequency includes a fitted linear or non-linear regression (e.g., polynomial regression). In certain embodiments, when the local genomic bias estimate and / or the bias frequency are weighted by a suitable process, the relationship between the local genomic bias estimate and the bias frequency includes a weighted relationship. In some embodiments, a weighted fitting relationship (e.g., weighted fitting) can be obtained by a process that includes quantile regression, a parameterized distribution, or an empirical distribution using interpolation. In certain embodiments, when the local genomic bias estimate is weighted, the relationship between the local genomic bias estimate and the bias frequency for a test sample, a reference, or a portion thereof includes polynomial regression. In some embodiments, the weighted fitting model includes weighting of the values of the distribution. The values of the distribution can be weighted by a suitable process. In some embodiments, values located near the tails of the distribution are provided with a smaller weight than values closer to the median of the distribution. For example, in the case of a distribution between a local genomic bias estimate (e.g., GC density) and a bias frequency (e.g., GC density frequency), the weight is determined according to the bias frequency for a given local genomic bias estimate, where a local genomic bias estimate that includes a bias frequency closer to the mean value of the distribution is provided with a larger weight than a local genomic bias estimate that includes a bias frequency farther from the mean value.
[0217] In some embodiments, the processing step includes normalizing the array read count by comparing a local genomic bias estimate of the array reads of the test sample to a local genomic bias estimate of a reference (e.g., a reference genome or a portion thereof). In some embodiments, the count of the array reads is normalized by comparing the bias frequency of the local genomic bias estimate of the test sample to the bias frequency of the local genomic bias estimate of the reference. In some embodiments, the count of the array reads is normalized by comparing the sample bias relationship to the reference bias relationship, thereby generating a comparison result.
[0218] The count of the array reads can be normalized according to the comparison results of two or more relationships. In certain embodiments, two or more relationships are compared, thereby providing a comparison result that is used to reduce local bias in the array reads (e.g., normalize the count). The two or more relationships can be compared by a suitable method. In some embodiments, the comparison result includes addition, subtraction, multiplication, and / or division of a first relationship and a second relationship. In certain embodiments, the comparison of two or more relationships includes the use of a suitable linear regression and / or non-linear regression. In certain embodiments, the comparison of two or more relationships includes a suitable polynomial regression (e.g., a cubic polynomial regression). In some embodiments, the comparison result includes addition, subtraction, multiplication, and / or division of a first regression and a second regression. In some embodiments, two or more relationships are compared by a process that includes an inference framework of multiple regressions. In some embodiments, two or more relationships are compared by a process that includes a suitable multivariate analysis. In some embodiments, two or more relationships are compared by a process that includes basis functions (e.g., blending functions, e.g., polynomial basis, Fourier basis, etc.), splines, radial basis functions, and / or wavelets.
[0219] In certain embodiments, the distribution of local genomic bias estimates, including the bias frequencies for the test sample and the reference, is compared by a process that includes polynomial regression in which the local genomic bias estimates are weighted. In some embodiments, the polynomial regression is generated between (i) ratios, each of which includes the bias frequency of the local genomic bias estimate of the reference and the bias frequency of the local genomic bias estimate of the sample, and (ii) the local genomic bias estimates. In some embodiments, the polynomial regression is generated between (i) the ratio of the bias frequency of the local genomic bias estimate of the reference to the bias frequency of the local genomic bias estimate of the sample, and (ii) the local genomic bias estimates. In some embodiments, the comparison of the distributions of local genomic bias estimates for the test sample and the reference reads includes measuring the log ratio (e.g., log2 ratio) of the bias frequencies of the local genomic bias estimates for the reference and the sample. In some embodiments, the comparison of the distributions of local genomic bias estimates includes dividing the log ratio (e.g., log2 ratio) of the bias frequency of the local genomic bias estimate for the reference by the log ratio (e.g., log2 ratio) of the bias frequency of the local genomic bias estimate for the sample.
[0220] Normalizing the counts according to the comparison results typically involves adjusting some counts and not others. Count normalization may sometimes adjust all counts and sometimes not adjust any of the array read counts. The counts for the array reads may sometimes be normalized by a process that includes a step of determining a weighting factor, and that process may sometimes not include the step of directly generating and using the weighting factor. Normalizing the counts according to the comparison results may sometimes involve determining a weighting factor for each count of the array reads. The weighting factor is often specific to the array read and is applied to the count of the specific array read. The weighting factor is often determined according to two or more comparison results of bias relationships (e.g., a sample bias relationship compared to a reference bias relationship). The normalized count is often determined by adjusting the count value according to the weighting factor. Adjusting the count according to the weighting factor may sometimes include adding the weighting factor to the count for the array read, subtracting the weighting factor from the count for the array read, multiplying the count for the array read by the weighting factor, and / or dividing the count for the array read by the weighting factor. The weighting factor and / or the normalized count may sometimes be determined from a regression (e.g., a regression line). The normalized count may sometimes be directly obtained from a regression line (e.g., a fitted regression line) resulting from the comparison of the bias frequency of the local genomic bias estimate of the reference (e.g., the reference genome) and the bias frequency of the local genomic bias estimate of the test sample. In some embodiments, each count of the reads of the sample is provided with a normalized count value according to the comparison result of (i) the bias frequency of the local genomic bias estimate of the read compared to (ii) the bias frequency of the local genomic bias estimate of the reference. In a particular embodiment, the counts of the array reads obtained for the sample are normalized and the bias in those array reads is reduced.
[0221] LOESS normalization In some embodiments, the processing step includes LOESS normalization. LOESS is a regression modeling method known in the art for combining multiple regression models in a meta-model based on the k-nearest neighbor method. LOESS is sometimes referred to as local weighted polynomial regression. GC LOESS, in some embodiments, applies the LOESS model to the relationship between fragment counts (e.g., array reads, counts) for portions of the reference genome and the GC composition. Plotting a smooth curve through a data point set using LOESS is sometimes called the LOESS curve, especially when each smoothed value is given by weighted quadratic least squares regression over the range of values of the response variable in the scatter plot on the y-axis. For each point in a dataset, the LOESS method fits a low-order polynomial to a subset of that data, where the explanatory variable values are close to the point for which the response is being estimated. The polynomial is fit using weighted least squares, with points closer to the point for which the response is being estimated given a larger weight and points further away given a smaller weight. The value of the regression function for a point is then obtained by evaluating the local polynomial using the explanatory variable values for that data point. The fitting of LOESS is sometimes considered complete after the regression function values have been calculated for each data point. Many details of this method (e.g., polynomial model and degree of weighting) are flexible.
[0222] Principal component analysis In some embodiments, the processing step includes principal component analysis (PCA). In some embodiments, the array read count (e.g., the array read count of a test sample) is adjusted according to principal component analysis (PCA). In some embodiments, the read density profile (e.g., the read density profile of a test sample) is adjusted according to principal component analysis (PCA). The read density profiles of one or more reference samples and / or the read density profile of a test subject may be adjusted according to PCA. Removing bias from the read density profile by a PCA-related process is sometimes referred to herein as profile adjustment. PCA can be performed by a suitable PCA method or a variant thereof. Non-limiting examples of PCA methods include canonical correlation analysis (CCA), Karhunen-Loeve transform (KLT), Hotelling transform, proper orthogonal decomposition (POD), singular value decomposition of X (SVD), eigenvalue decomposition of XTX (EVD), factor analysis, Eckart-Young theorem, Schmidt-Mirsky theorem, empirical orthogonal function (EOF), empirical eigenfunction decomposition, empirical component analysis, quasi-harmonic mode, spectral decomposition, empirical modal analysis, etc., their variants or combinations. PCA often identifies and / or adjusts one or more biases in the read density profile. The bias identified and / or adjusted by PCA is sometimes referred to herein as the principal component. In some embodiments, one or more biases can be removed by adjusting the read density profile according to one or more principal components using a suitable method. The read density profile can be adjusted by addition, subtraction, multiplication, and / or division of one or more principal components and the read density profile. In some embodiments, one or more biases can be removed from the read density profile by subtracting one or more principal components from the read density profile. The bias in the read density profile is often identified and / or quantified by PCA of the profile, but the principal component is often subtracted from the profile at the level of the read density.PCA often identifies one or more principal components. In some embodiments, PCA identifies the first, second, third, fourth, fifth, sixth, seventh, eighth, ninth, and tenth or more principal components. In certain embodiments, one, two, three, four, five, six, seven, eight, nine, ten or more principal components are used to adjust the profile. In certain embodiments, five principal components are used to adjust the profile. The principal components are often used to adjust the profile in the order of appearance in the PCA. For example, if three principal components are subtracted from the read density profile, the first, second, and third principal components are used. The bias identified by the principal components may sometimes include features of the profile that are not used to adjust the profile. For example, PCA may identify copy number variations (e.g., aneuploidy, microduplications, microdeletions, deletions, translocations, insertions) and / or sexual differences as principal components. Thus, in some embodiments, one or more principal components are not used to adjust the profile. For example, if the third principal component is not used to adjust the profile, the first, second, and fourth principal components may be used to adjust the profile.
[0223] The principal components can be obtained from PCA using any suitable sample or reference. In some embodiments, the principal components are obtained from a test sample (e.g., a test subject). In some embodiments, the principal components are obtained from one or more references (e.g., a reference sample, a reference array, a reference set). In certain cases, PCA is performed on the median of the read density profiles obtained from a training set comprising a plurality of samples, and the first principal component and the second principal component are identified. In some embodiments, the principal components are obtained from a set of subjects lacking copy number variations of interest. In some embodiments, the principal components are obtained from a set of known euploidies. The principal components are often identified according to PCA performed using one or more read density profiles of a reference (e.g., a training set). One or more principal components obtained from the reference are often subtracted from the read density profile of the test subject, thereby providing an adjusted profile.
[0224] Hybrid normalization In some embodiments, the processing step includes a hybrid normalization method. The hybrid normalization method can, in certain cases, reduce a bias (e.g., a GC bias). Hybrid normalization includes, in some embodiments, (i) the analysis of the relationship between two variables (e.g., count and GC content), and (ii) the selection and application of a normalization method according to that analysis. Hybrid normalization includes, in certain embodiments, (i) regression (e.g., regression analysis) and (ii) the selection and application of a normalization method according to that regression. In some embodiments, the counts obtained for a first sample (e.g., a first sample set) are normalized in a different way than the counts obtained from another sample (e.g., a second sample set). In some embodiments, the counts obtained for a first sample (e.g., a first sample set) are normalized by a first normalization method, and the counts obtained from a second sample (e.g., a second sample set) are normalized by a second normalization method. For example, in certain embodiments, the first normalization method includes the use of linear regression, and the second normalization method includes the use of non-linear regression (e.g., LOESS, GC-LOESS, LOWESS regression, LOESS smoothing).
[0225] In some embodiments, the hybrid normalization method is used to normalize sequence reads (e.g., counts, mapped counts, mapped reads) mapped to portions of a genome or chromosome. In certain embodiments, raw counts are normalized, and in some embodiments, adjusted, weighted, filtered, or pre-normalized counts are normalized by the hybrid normalization method. In certain embodiments, levels or Z-scores are normalized. In some embodiments, counts mapped to selected portions of a genome or chromosome are normalized by a hybrid normalization approach. Counts can refer to a suitable measure of sequence reads mapped to portions of a genome, non-limiting examples of which include raw counts (e.g., unprocessed counts), normalized counts (e.g., normalized by LOESS, principal components, or a suitable method), sub-levels (e.g., average level, mean value level, median level, etc.), Z-scores, etc. or combinations thereof. Those counts can be raw or processed counts from one or more samples (e.g., test samples, samples from pregnant women). In some embodiments, the counts are obtained from one or more samples obtained from one or more subjects.
[0226] In some embodiments, the normalization method (e.g., type of normalization method) is selected according to regression (e.g., regression analysis) and / or correlation coefficients. Regression analysis refers to a statistical technique for estimating the relationship between variables (e.g., counts and GC content). In some embodiments, the regression is generated according to measures of counts and GC content for each of multiple portions of a reference genome. Suitable measures of GC content can be used, non-limiting examples of which include measures of guanine, cytosine, adenine, thymine, purine (GC), or pyrimidine (AT or ATU) content, melting temperature (T m)(e.g., denaturation temperature, annealing temperature, hybridization temperature), a measure of free energy, etc. or a combination thereof. Measures of guanine (G), cytosine (C), adenine (A), thymine (T), purine (GC) or pyrimidine (AT or ATU) content can be expressed as ratios or percentages. In some embodiments, any suitable ratio or percentage is used, non-limiting examples of which include GC / AT, GC / total nucleotides, GC / A, GC / T, AT / total nucleotides, AT / GC, AT / G, AT / C, G / A, C / A, G / T, G / A, G / AT, C / T, etc. or combinations thereof. In some embodiments, the measure of GC content is the ratio or percentage of GC to total nucleotide content. In some embodiments, the measure of GC content is the ratio or percentage of GC to total nucleotide content for sequence reads mapped to a portion of the reference genome. In certain embodiments, the GC content is measured according to and / or from sequence reads mapped to each portion of the reference genome, and those sequence reads are obtained from a sample. In some embodiments, the measure of GC content is not determined according to and / or from sequence reads. In certain embodiments, the measure of GC content is determined for one or more samples obtained from one or more subjects.
[0227] In some embodiments, the generation of the regression includes the generation of a regression analysis or a correlation analysis. Suitable regressions can be used, and non-limiting examples thereof include regression analysis (e.g., linear regression analysis), goodness-of-fit analysis, Pearson correlation analysis, rank correlation, fraction of variance unexplained, Nash-Sutcliffe model efficiency analysis, regression model verification, proportional reduction in loss, root mean square deviation, etc. or combinations thereof. In some embodiments, a regression line is generated. In certain embodiments, the generation of the regression includes the generation of a linear regression. In certain embodiments, the generation of the regression includes the generation of a non-linear regression (e.g., LOESS regression, LOWESS regression).
[0228] In some embodiments, the regression determines the presence or absence of a correlation (e.g., a linear correlation) between, for example, a count and a measure of GC content. In some embodiments, a regression (e.g., a linear regression) is generated and a correlation coefficient is determined. In some embodiments, a suitable correlation coefficient is determined, and non-limiting examples thereof include the coefficient of determination, R 2 value, Pearson correlation coefficient, etc.
[0229] In some embodiments, the goodness of fit is measured with respect to a regression (e.g., regression analysis, linear regression). The goodness of fit may sometimes be measured by visual analysis or mathematical analysis. The evaluation may sometimes include determining whether the goodness of fit is higher for a non - linear regression or higher for a linear regression. In some embodiments, the correlation coefficient is a measure of the goodness of fit. In some embodiments, the evaluation of the goodness of fit for a regression is revealed according to a correlation coefficient and / or a cut - off value of the correlation coefficient. In some embodiments, the evaluation of the goodness of fit includes comparing the correlation coefficient with the cut - off value of the correlation coefficient. In some embodiments, the evaluation of the goodness of fit for a regression suggests a linear regression. For example, in a particular embodiment, the goodness of fit is higher for a linear regression than for a non - linear regression, and the evaluation of that goodness of fit suggests a linear regression. In some embodiments, the evaluation suggests a linear regression and linear regression is used to normalize the count. In some embodiments, the evaluation of the goodness of fit for a regression suggests a non - linear regression. For example, in a particular embodiment, the goodness of fit is higher for a non - linear regression than for a linear regression, and the evaluation of that goodness of fit suggests a non - linear regression. In some embodiments, the evaluation suggests a non - linear regression and non - linear regression is used to normalize the count.
[0230] In some embodiments, when the correlation coefficient is equal to or exceeds the cut - off of the correlation coefficient, the evaluation of the goodness of fit suggests a linear regression. In some embodiments, when the correlation coefficient is less than the cut - off of the correlation coefficient, the evaluation of the goodness of fit suggests a non - linear regression. In some embodiments, the cut - off of the correlation coefficient is pre - determined. In some embodiments, the cut - off of the correlation coefficient is about 0.5 or more, about 0.55 or more, about 0.6 or more, about 0.65 or more, about 0.7 or more, about 0.75 or more, about 0.8 or more, or about 0.85 or more.
[0231] In some embodiments, a particular type of regression is selected (e.g., linear or non-linear regression), and after that regression is generated, the count is normalized by subtracting that regression from the count. In some embodiments, subtracting the regression from the count provides a normalized count with reduced bias (e.g., GC bias). In some embodiments, a linear regression is subtracted from the count. In some embodiments, a non-linear regression (e.g., LOESS, GC-LOESS, LOWESS regression) is subtracted from the count. Any suitable method may be used to subtract the regression line from the count. For example, if count x is derived from part i (e.g., part i) with a GC content of 0.5 and the regression line determines count y at a GC content of 0.5, then for part i, x - y = the normalized count. In some embodiments, the count is normalized before and / or after subtraction of the regression. In some embodiments, the count normalized by a hybrid normalization approach is used to generate a genome or a level, Z-score, level and / or profile of a part thereof. In certain embodiments, the count normalized by a hybrid normalization approach is analyzed by the methods described herein to determine the presence or absence of a genetic mutation or genetic change (e.g., copy number change).
[0232] In some embodiments, the hybrid normalization method includes filtering or weighting one or more portions before or after normalization. Suitable methods for filtering portions (e.g., portions of a reference genome) described herein can be used, including methods for filtering the portions described herein. In some embodiments, the portions (e.g., portions of the reference genome) are filtered before applying the hybrid normalization method. In some embodiments, only the counts of sequencing reads mapped to selected portions (e.g., portions selected according to the variance of the counts) are normalized by hybrid normalization. In some embodiments, the counts of sequencing reads mapped to the filtered reference genome portions (e.g., portions filtered according to the variance of the counts) are removed before using the hybrid normalization method. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., portions of the reference genome) according to a suitable method (e.g., a method described herein). In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., portions of the reference genome) according to uncertainty values for the counts mapped to each portion for a plurality of test samples. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., portions of the reference genome) according to the variance of the counts. In some embodiments, the hybrid normalization method includes selecting or filtering portions (e.g., portions of the reference genome) according to GC content, repetitive elements, repetitive sequences, introns, exons, etc. or combinations thereof. Profile
[0233] In some embodiments, the processing step includes generating one or more profiles (e.g., profile plots) from various aspects of a dataset or its derivatives (e.g., the result of one or more mathematical and / or statistical data processing steps known in the art and / or described herein).
[0234] As used herein, the term "profile" refers to the result of a mathematical and / or statistical manipulation of data that may facilitate the identification of patterns and / or correlations in a large amount of data. A "profile" often includes values resulting from one or more operations on data or a dataset based on one or more criteria. A profile often includes a plurality of data points. Depending on the nature and / or complexity of the dataset, any suitable number of data points may be included in the profile. In certain embodiments, the profile may include two or more data points, three or more data points, five or more data points, ten or more data points, twenty-four or more data points, twenty-five or more data points, fifty or more data points, one hundred or more data points, five hundred or more data points, one thousand or more data points, five thousand or more data points, ten thousand or more data points, or one hundred thousand or more data points.
[0235] In some embodiments, the profile represents the entire dataset, and in certain embodiments, the profile represents a part or subset of the dataset. That is, the profile may include or be generated from data points that represent unfiltered data for removing any data, and the profile may include or be generated from data points that represent filtered data for removing unwanted data. In some embodiments, the data points in a profile correspond to the result of data manipulation on a part. In certain embodiments, the data points in a profile include the result of data manipulation on a group of parts. In some embodiments, the groups of parts may be adjacent to each other, and in certain embodiments, the groups of parts may be from different parts of a chromosome or genome.
[0236] The data points in a profile derived from a dataset can represent any suitable categorization of data. Non-limiting examples of categories into which data can be grouped to generate profile data points include parts based on size, parts based on sequence features (e.g., GC content, AT content, location on a chromosome (e.g., short arm, long arm, centromere, telomere), etc.), expression levels, chromosomes, etc., or combinations thereof. In some embodiments, a profile can be generated from data points obtained from another profile (e.g., a normalized data profile that has been renormalized for different normalization values to generate a renormalized data profile). In certain embodiments, a profile generated from data points obtained from another profile reduces the number of data points and / or the complexity of the dataset. The reduction in the number of data points and / or the complexity of the dataset often facilitates the interpretation of the data and / or the provision of an outcome.
[0237] A profile (e.g., a genomic profile, a chromosomal profile, a profile of a portion of a chromosome) is often a set of normalized or non-normalized counts for two or more portions. A profile often includes at least one level and often includes two or more levels (e.g., a profile often has multiple levels). A level is generally for a set of portions having approximately the same count or normalized count. Levels are described in more detail herein. In certain embodiments, a profile includes one or more portions, and those portions can be weighted, removed, filtered, normalized, adjusted, averaged, derived as an average value, added, subtracted, processed, or transformed by any combination thereof. A profile often includes normalized counts mapped to portions that define two or more levels, where those counts are further normalized according to one of those levels by a suitable method. Counts of a profile (e.g., a profile level) are often associated with uncertain values.
[0238] Profiles that include one or more levels may sometimes be padded (e.g., hole padding). Padding (e.g., hole padding) refers to the process of identifying and adjusting levels in a profile due to copy number variations (e.g., microduplications or microdeletions in a patient's genome, maternal microduplications or microdeletions). In some embodiments, levels resulting from microduplications or microdeletions in a tumor or fetus are padded. Microduplications or microdeletions in a profile may, in some embodiments, artificially increase or decrease the overall level of a profile (e.g., a chromosomal profile) that results in a false positive or false negative determination of aneuploidy (e.g., trisomy). In some embodiments, levels in a profile due to microduplications and / or deletions are identified and adjusted (e.g., padded and / or removed) by a process that may sometimes be referred to as padding or hole padding.
[0239] A profile including one or more levels may include a first level and a second level. In some embodiments, the first level is different (e.g., significantly different) from the second level. In some embodiments, the first level includes a first subset, the second level includes a second subset, and the first subset is not a subset of the second subset. In certain embodiments, the first subset is different from the second subset in which the first and second levels are measured. In some embodiments, a profile may have a plurality of first levels that are different (e.g., significantly different, e.g., have significantly different values) from a second level within that profile. In some embodiments, a profile includes one or more first levels that are significantly different from a second level within that profile, and the one or more first levels are adjusted. In some embodiments, a first level within a profile is removed from or adjusted (e.g., padded) in that profile. A profile may include a plurality of levels including one or more first levels that are significantly different from one or more second levels, and most of the levels in a profile are often the second levels, and the second levels are approximately equal to each other. In some embodiments, more than 50%, more than 60%, more than 70%, more than 80%, more than 90% or more than 95% of the levels in a profile are the second levels.
[0240] Profiles may sometimes be displayed as plots. For example, one or more levels representing a count of portions (e.g., normalized count) may be plotted and visualized. Non-limiting examples of plots of profiles that may be generated include raw counts (e.g., raw count profiles or raw profiles), normalized counts, z-scores weighted by portions, p-values, area ratios for fitted ploidy, median levels for the ratio of the fitted minority ratio to the measured minority ratio, principal components, etc. or combinations thereof. Plots of profiles, in some embodiments, enable visualization of the manipulated data. In certain embodiments, plots of profiles may be used to provide outcomes (e.g., area ratios for fitted ploidy, median levels for the ratio of the fitted minority ratio to the measured minority ratio, principal components). The terms “raw count profile plot” or “raw profile plot,” as used herein, refer to a plot of the counts in each portion in a region, normalized to the total counts in a region (e.g., a genome, portion, chromosome, chromosomal portion of a reference genome or part of a chromosome). In some embodiments, profiles may be generated using a static window process, and in certain embodiments, profiles may be generated using a sliding window process.
[0241] The profiles generated for a test subject may be compared to profiles generated for one or more reference subjects to facilitate the mathematical and / or statistical manipulation and interpretation of the dataset and / or to provide an outcome. In some embodiments, the profiles are generated based on one or more starting assumptions, such as the assumptions described herein. In certain embodiments, test profiles often center around a predetermined value representative of the absence of copy number variation, and when the test subject has a copy number variation, the copy number variation often deviates from the predetermined value in the region corresponding to the genomic location within the test subject. In test subjects at risk for or suffering from medical conditions associated with copy number variation, the numerical values for the selected portions are expected to vary significantly from the predetermined values for unaffected genomic locations. Depending on the starting assumptions (e.g., a given ploidy or optimized ploidy, a given ratio of cancer cell nucleic acid or optimized ratio of cancer cell nucleic acid, a given fetal ratio or optimized fetal ratio, or combinations thereof), the predetermined threshold or cutoff value or range of threshold values suggesting the presence or absence of copy number variation may vary, but still provides an outcome useful for determining the presence or absence of copy number variation. In some embodiments, the profiles suggest and / or represent a phenotype.
[0242] In some embodiments, the use of one or more reference samples that do not substantially include a change in the copy number of the subject can be used to generate a reference count profile (e.g., a median profile of the reference counts), which can result in a predetermined value representative of the absence of a copy number change. When a test subject has a copy number change, the copy number change often deviates from the predetermined value in the region corresponding to the genomic position located within that test subject. In a test subject at risk of or suffering from a medical condition associated with a copy number change, the numerical value for the selected portion or segment is expected to vary significantly from the predetermined value for a genomic position not affected. In certain embodiments, the use of one or more reference samples that have been determined to have a copy number change in the subject can be used to generate a reference count profile (a median profile of the reference counts), which can result in a predetermined value representative of the presence of a copy number change. When a test subject does not have a copy number change, the copy number change often deviates from the predetermined value in the region corresponding to the genomic position. In a test subject not at risk of or suffering from a medical condition associated with a copy number change, the numerical value for the selected portion or segment is expected to vary significantly from the predetermined value for an affected genomic position.
[0243] As a non-limiting example, a normalized sample count profile and / or a normalized reference count profile can be obtained from raw sequence read data by: (a) calculating the median of the reference counts for a selected chromosome, a portion or part thereof, from a set of references that have been found to have no copy number variations; (b) removing uninformative portions from the raw counts of the reference samples (e.g., filtering); (c) normalizing the reference counts for all remaining portions of the reference genome relative to the total remaining counts for the selected chromosome or selected genomic position of the reference sample (e.g., the sum of the counts remaining after removing uninformative portions of the reference genome), thereby generating a normalized reference subject profile; (d) removing the corresponding portions from the samples of the test subject; (e) normalizing the remaining test subject counts for one or more selected genomic positions relative to the sum of the medians of the remaining reference counts for the chromosome containing the selected genomic position, thereby generating a normalized test subject profile. In certain embodiments, in (b), an additional normalization step for the entire genome, reduced by the filtered portion, can be included between (c) and (d).
[0244] In some embodiments, a lead density profile is measured. In some embodiments, the lead density profile includes at least one lead density and often includes two or more lead densities (e.g., the lead density profile often includes a plurality of lead densities). In some embodiments, the lead density profile includes a suitable quantitative value (e.g., an average value, a median value, a Z-score, etc.). The lead density profile often includes a value resulting from one or more lead densities. The lead density profile may sometimes include a value resulting from one or more operations on the lead density based on one or more adjustments (e.g., normalization). In some embodiments, the lead density profile includes unoperated lead density. In some embodiments, one or more lead density profiles are generated from various aspects of a dataset that includes a lead density or its derivative (e.g., the result of one or more mathematical and / or statistical data processing steps known in the art and / or described herein). In a particular embodiment, the lead density profile includes a normalized lead density. In some embodiments, the lead density profile includes an adjusted lead density. In a particular embodiment, the lead density profile includes raw lead density (e.g., unoperated, unadjusted, or unnormalized lead density), normalized lead density, weighted lead density, filtered portion of lead density, Z-score of lead density, p-value of lead density, integral value of lead density (e.g., area under the curve), mean, average value or median value of lead density, principal component, etc. or combinations thereof. The lead density of the lead density profile and / or the lead density profile are often related to a measure of uncertainty (e.g., MAD). In a particular embodiment, the lead density profile includes the distribution of the median value of the lead density. In some embodiments, the lead density profile includes the relationship of a plurality of lead densities (e.g., a fitting relationship, regression, etc.).For example, a read density profile may include a relationship between a read density (e.g., a value of the read density) and a genomic position (e.g., a portion, a position of the portion). In some embodiments, the read density profile is generated using a static window process, and in certain embodiments, the read density profile is generated using a sliding window process. In some embodiments, the read density profile may be printed and / or displayed (e.g., visually displayed, e.g., displayed as a plot or graph).
[0245] In some embodiments, the read density profile corresponds to a subset (e.g., a subset of a reference genome, a subset of a chromosome, or a sub - subset of a part of a chromosome). In some embodiments, the read density profile includes read density and / or read count assoc...
Claims
[Claim 1] The invention as depicted in the drawings.
Citation Information
Patent Citations
Methods and processes for non-invasive assessment of gene variations
JP2016526879A
Methods and procedures for non-invasive assessment of genetic variation
JP2016533173A
Methods and processes for non-invasive assessment of genetic variations
US20160034640A1
Chromosome representation determinations
WO2015183872A1